Friday, April 24, 2015

Python for Data Scientists - IPython


Introduction


Having learned some basic packages of Python, you probably started to wonder that working through the python console is not very productive. In R we have RStudio, of which we've already talked in first articles. Good folks of Python community have developed an IPython - an interactive Python console and web environment.

Installation


As usual we are using python pip package manager to install the package:
pip install ipython

Usage


Once the package is installed, you can launch the console version by simply typing it's name in the console:
ipython
Once the application is started, one can simply type python commands and observe the results. Though it doesn't look that far different from the ordinary python console, it provides auto-quoting, code completion, search of previously executed commands, output caching and many more.
As I said previously, there are two modes of running the IPython - console and web. To launch the web interface, one must add the "notebook" parameter:
ipython notebook
This will open http://localhost:8888/tree URL using a default browser. You can then create new IPython files, with ipynb extension, or upload one from the local file system.

The graphical interface is no doubt much more comfortable to use and adds numerous editing and flow control features on top of those supported by the console mode.



If you have chosen to work with Python language for your data project, IPython is a must to have tool. Make sure to have it in your toolkit.
Monday, April 13, 2015

Python for Data Scientists - Pandas


Introduction

Having learnt NumPy and SciPy in previous articles, let's discuss our next package, called pandas.
Pandas provides rich data structures and functions designed to make working with structured data fast, easy, and expressive. It is, as you will see, one of the critical ingredients enabling Python to be a powerful and productive data analysis environment.

Pandas combines the high performance array-computing features of NumPy with the flexible data manipulation capabilities of spreadsheets and relational databases (such as SQL), mingling DataFrame - a two-dimensional tabular, column-oriented data structure with both row and column labels. It provides sophisticated indexing functionality to make it easy to reshape, slice and dice, perform aggregations, and select subsets of data.

For users of the R language for statistical computing, the DataFrame name will be familiar, as the object was named after the similar R data.frame object. However the functionality provided by R is merely a subset of that provided by the pandas DataFrame.

Installation

Installation of pandas is as everything in Python ecosystem, a piece of cake. For those working with Python distribution, it's been pre-packed for you. To install it manually using python package manager

pip install pandas

Data Structures

To get started with pandas, you will need to get comfortable with its two data structures used throughout the library: Series and DataFrame. I won't get into details about Panel and Panel4D, somewhat less-used containers, of which you can read on pandas site.

A Series is a one-dimensional array-like object containing an array of data (of any NumPy data type) and an associated array of data labels, called its index.

import pandas as pd
import numpy as np
pd.Series([1,3,5,np.nan,6,8])

# 0     1
# 1     3
# 2     5
# 3   NaN
# 4     6
# 5     8

A DataFrame represents a tabular, spreadsheet-like data structure containing an or- dered collection of columns, each of which can be a different value type (numeric, string, boolean, etc.).

dates = pd.date_range('20130101',periods=6)
pd.DataFrame(np.random.randn(6,4),index=dates,columns=list('ABCD'))
#                    A         B         C         D
# 2013-01-01  0.469112 -0.282863 -1.509059 -1.135632
# 2013-01-02  1.212112 -0.173215  0.119209 -1.044236
# 2013-01-03 -0.861849 -2.104569 -0.494929  1.071804
# 2013-01-04  0.721555 -0.706771 -1.039575  0.271860
# 2013-01-05 -0.424972  0.567020  0.276232 -1.087401
# 2013-01-06 -0.673690  0.113648 -1.478427  0.524988

As you can see DataFrame has both a row and column index; it can be thought of as a dict of Series.

Examples

Index Objects

As we've seen in the DataFrame example, we can provide a custom index to both DataFrame and Series objects and then reference items through by using it.
s = Series(range(3), index=['a', 'b', 'c'])
s['a'] # 0
s['b'] = 16.5

Reindexing

Consider a case, where you'd like to change the indices of your data resulting alternation, addition or removal of entities.
index = ['a', 'c', 'd']
columns = ['GBP', 'USD', 'EUR']
data = np.arange(9).reshape((3, 3))      # reshape 9x1 to 3x3
frame = DataFrame(data, index=index, columns=columns)
#   GBP USD EUR
# a 0   1   2 
# c 3   4   5
# d 6   7   8

frame.reindex(columns=['GBP', 'JPY', 'EUR'])
#   GBP JPY EUR
# a 0   NaN   2 
# c 3   NaN   5
# d 6   NaN   8

frame.drop('GBP')
#   JPY EUR
# a NaN   2 
# c NaN   5
# d NaN   8

As you can see, both adding and removing indices is as easy as breathing. Both methods support array parameters, so bulk data alternation is also possible and even advisable from optimization purposes.

Arithmetic and data alignment

Another important pandas feature is arithmetic behavior between objects with different indexes. When adding together objects, if any index pairs are not the same, the respective index in the result will be the union of the index pairs.

s1 = Series([7.3, -2.5, 3.4, 1.5], index=['a', 'c', 'd', 'e'])
s2 = Series([-2.1, 3.6, -1.5, 4, 3.1], index=['a', 'c', 'e', 'f', 'g'])
s1 + s2
# a 5.2
# c 1.1
# d NaN
# e 0.0
# f NaN
# g NaN

Here column d, f and g were converted to NaN as they didn't have a match in both series. The same applies to DataFrames of course. One thing you might find useful is filling the NaN values with some defaults. This can be achieved using filling function, supported by all corresponding methods: add, sub, div and mul.

Merge

What about morphing 2 data objects. No worries - pandas comes to rescue. It provides various facilities for easily combining together objects with various kinds of set logic for the indexes and relational algebra functionality in the case of join / merge-type operations.

key = ['foo', 'foo']
left = pd.DataFrame({'key': key, 'lval': [1, 2]})
#    key  lval
# 0  foo     1
# 1  foo     2

right = pd.DataFrame({'key': key, 'rval': [4, 5]})
#    key  rval
# 0  foo     4
# 1  foo     5

pd.concat([left,right])
#    key  lval  rval
# 0  foo     1   NaN
# 1  foo     2   NaN
# 0  foo   NaN     4
# 1  foo   NaN     5

merged = pd.merge(left, right, on='key')
#    key  lval  rval
# 0  foo     1     4
# 1  foo     1     5
# 2  foo     2     4
# 3  foo     2     5

merged.groupby('key').sum()
#      lval  rval
# key            
# foo     6    18

Handling Missing Data

Very often, if not always, we deal with incomplete data, either by it's nature like sensor data or as a result of human error like spreadsheets. Pandas provides various functionality to deal with such situations.

from numpy import nan as NA
data = Series([1, NA, 3.5, NA, 7])
data.dropna() # similar to data[data.notnull()]
# 0 1.0
# 2 3.5
# 4 7.0

data.fillna(0)
# 0 1.0
# 1 0.0
# 2 3.5
# 3 0.0
# 4 7.0

With every evolving API, pandas provides numerous functionality, which will ease any data scientist life. Make sure you keep yourself updated with the features of every release.

Monday, March 30, 2015

Python for Data Scientists - SciPy


Introduction


This article continues the Python for Data Scientists series by talking about SciPy. It is built on top of NumPy, of which we've already talked in the previous article. SciPy provides many user-friendly and efficient numerical routines addressing a number of different standard problem domains in scientific computing such as integration, differential and sparse linear system solvers, optimizers and root finding algorithms, Fourier Transforms, various standard continuous and discrete probability distributions and many more. Together NumPy and SciPy form a reasonably complete computational replacement for much of MATLAB along with some of its add-on toolboxes.

Installation


Installation of SciPy is trivial. In many cases, it will be already supplied to you with python distribution, or as usual may be installed manually using python package manager
pip install scipy
Depending on the running OS, you might be needing to install gfortran, prior to SciPy installation.

Examples

Optimization

Very often we want to find a maxima or minima of the function, that is find a solution for optimization problem. Let's see how to do this with SciPy, by finding maxima of Bessel function. Since optimization is a process of finding a minima, we are negating the function:
from scipy import special, optimize
f = lambda x: -special.jv(3, x) # define a function
sol = optimize.minimize(f, 1.0) # optimize the function

Statistics

Today we cannot imagine ourselves without statistics. From generating random variables to emitting some events at known probability, statistics has deeply ingrained itself into any developer's toolset.
from scipy.stats import percentileofscore
list = [1, 2, 3, 4]
percentileofscore(list, 3) # what percentage lies beneath 3 => 75

Singular Value Decomposition

SVD or Singular Value Decomposition has many useful applications in signal processing and statistics. As a data scientist you will be meeting it a lot! Dimension reduction, collaborative filtering, you name it, it is always there. Let's see how to calculate one:
from scipy import linalg
a = np.random.randn(9, 6) + 1.j*np.random.randn(9, 6)
U, s, Vh = linalg.svd(a)

Interpolation

We'll finish our overview with an example of interpolation. Very often we want to approximate a continues function by evaluating point at constant rate. SciPy provides a handful of functions to do so in multiple dimensions.
from scipy.interpolate import interp1d
import numpy as np
x = np.linspace(0, 10, 10)
y = np.cos(-x**2/8.0)
f = interp1d(x, y)
SciPy contains numerous functions from various domain of science. Be sure to overview them all in the documentation, as most probably your next task is already fully implemented, tested and optimized by one of the provided functions of this wonderful package.
Tuesday, March 10, 2015

Python for Data Scientists - NumPy



Introduction


We'll start our Python for Data Scientists series with NumPy, short for Numerical Python, which is the foundational package for scientific computing in Python. One of its primary purposes with regards to data analysis is as the primary container for data to be passed between algorithms. For numerical data, NumPy arrays are a much more efficient way of storing and manipulating data than the other built-in Python data structures. Also, libraries written in a lower-level language, such as C or Fortran, can operate on the data stored in a NumPy array without copying any data. Here are some of the things it provides:
  • A fast and efficient multidimensional array object ndarray
  • Functions for performing element-wise computations arrays
  • Tools for reading and writing array-based data sets to disk
  • Linear algebra operations, Fourier transform, and random number generation
  • Tools for integrating connecting C, C++, and Fortran code to Python 
Knowing Numpy is fundamental and while by itself it does not provide very much high-level data analytical functionality, having an understanding of NumPy arrays and array-oriented computing will help you use tools like pandas much more effectively.

Installation


Since everyone uses Python for different applications, there is no single solution for setting up Python and required add-on packages. Personally I recommend using one of the following base Python distributions:
  • Enthought Python Distribution: a scientific-oriented Python distribution from Enthought. This includes Canopy Express, a free base scientific distribution (with NumPy, SciPy, matplotlib, Chaco, and IPython) and Canopy Full, a comprehensive suite of more than 300 scientific packages across many domains.
  • Python(x,y): A free scientific-oriented Python distribution for Windows.
If you'd rather install your packages by yourself, then the following code will do the trick:
pip install numpy

Features

ndarray: A Multidimensional Array Object

One of the key features of NumPy is its N-dimensional array object, or ndarray, which is a fast, flexible container for large data sets in Python. Arrays enable you to perform mathematical operations on whole blocks of data using similar syntax to the equivalent operations between scalar elements. This is important because they enable you to express batch operations on data without writing any for loops. This is usually called vectorization. Consider the next snippet:
import numpy as np
arr = np.arange(15) # returns numbers from 0 to 15, but as an array
arr[5:8] = 12       # assign 12 to items indexed from 5 to 8
arr.sort()          # sorts the array
arr = 1 / arr       # self assignment of 1 divided by each array item
arr.reshape((3, 5)) # reshapes array into 3x5 matrix
arr[arr < 5] = 0    # zeroes elements greater than 5
The code is self explanatory and gives you a little taste of what you can do with NumPy. Let us take a step further.

Universal functions

A universal function, or ufunc, is a function that performs elementwise operations on data in ndarrays. You can think of them as fast vectorized wrappers for simple functions that take one or more scalar values and produce one or more scalar results. Look at the next examples of some them. For more details, have a look at it's page.
x = np.sqrt(arr)    # element-wise square root
y = np.random.randn(8) * 100
y = np.floor(y)     # floors each element of the array
np.maximum(x, y)    # element-wise maximum

Storing Arrays on Disk in Binary Format

np.save and np.load are the two workhorse functions for efficiently saving and loading array data on disk. Arrays are saved by default in an uncompressed raw binary format with file extension .npy.
arr1 = np.arange(10)
np.save('some_array', arr2)
arr2 = np.load('some_array.npy')
np.array_equal(arr1, arr2)
Loading text from files is a fairly standard task. It will at times be useful to load data into vanilla NumPy arrays using np.loadtxt or the more specialized np.genfromtxt. These functions have many options allowing you to specify different delimiters, converter functions for certain columns, skipping rows, and other things.

Linear Algebra

Linear algebra, like matrix multiplication, decompositions, determinants, and others are the building block of nearly every data algorithm. numpy.linalg has a standard set of matrix decompositions and things like inverse and determinant. These are implemented under the hood using the same industry-standard Fortran libraries used in other languages like MATLAB and R, such as like BLAS, LAPACK, or the Intel MKL.
import dot from np, allclose
import randn from np.random
import svd from np.linalg

a = randn(9, 6)
b = randn(9, 6)
c = a + 1j*b                         # initiate complex matrix
U, s, V = svd(a, full_matrices=True) # perform svd decomposition
S = np.zeros((9, 6), dtype=complex)  # 9x6 complex zero matrix
S[:6, :6] = np.diag(s)               # swap diagonals
allclose(a, dot(U, dot(S, V)))       # equal within a tolerance
This will conclude the tutorial about NumPy and feel free to check it's documentation in depth. Next time we'll be taking a deeper look into Python Data Science tool kit with an overview about SciPy.
Friday, February 20, 2015

R dynamic report generation with Knitr


Let us change our traditional attitude to the construction of programs: Instead of imagining that our main task is to instruct a computer what to do, let us concentrate rather on explaining to humans what we want the computer to do.

              Donald E. Knuth, Literate Programming, 1984


Overview and Motivation


So what is dynamic documentation and why do we need it. As opposed to usual programming, R programs were intended to used as report for not development oriented folks, whether they are data scientists, statisticians or managers. Moreover by nature, R programs don't tend to be huge spanning across hundreds of thousands of code lines. All these led to a huge demand good documentation framework. But how do you document a report? One may of course is to write a passage and then paste a copied graph into it, however once something changes one must re-copy all the graph, which is of course very tedious and non-rewarding procedure.

Knitr


Knitr is an R package that allows straightforward integration of R code for writing reports and is developed by Yihui Xie. It is very powerful and easy to get started with and has potentially a lot of uses.
All you need is to create a rmd file and then use a regular markdown syntax with some additional features. For example:
## Loading and preprocessing the data
```{r echo = TRUE}
data = read.csv("activity.csv")
summary(data)
```
Here we create a header and then insert a code snippet wrapped in ```{r} ```. echo = TRUE tells knitr you want to render the code and the results. That's it - it's that easy. The variables you create in one section are visible in others, so you can write your program as usual using any packages or functions you want. In the end, the knitting process will parse your document, run all the R snippents and append the code and the results, where needed, to the generated HTML or PDF file.

Knitr if fully integrated with RStudio, however if you're using another IDE or just a fan of a console, you can always knit your program by running the following command:
Rscript -e "library(knitr); knit('./file-here.rmd')"
For more advanced options and more detailed examples please read the documentation and demos sections in knitr site.

Publish


When it comes to sharing your report, it has always been an obstacle. An endless email thread with pdf attachments - sounds familiar? One of the best features of RStudio is an ability to publish the Knitr reports at rpubs.com. You can share the link then with everyone you need for them to view the report. Here is an example of my whether events analysis - Simple whether events analysis
Bare in mind and reports published on rPubs are publicly available, so you probably shouldn't publish something classified there.

Hope you enjoyed the article and stay tuned for the next one, of course :)
Friday, February 6, 2015

Data Scientist Toolkit



The history of technology is the history of the invention of tools and techniques, and is similar in many ways to the history of humanity. And since data scientists are mere mortals, they also need tools to make their work more productive and even enjoyable, but that's just me. In this article we'll be talking about main languages and tools used by data scientists. For ones who have recently entered this field of science, it will be a great overview about mostly used tools.

R


A great advantage of R is that scientists adopted it as their de facto standard. As a consequence, the latest cutting-edge techniques are first available in R. It also seems to be the preference of most Kaggle competition winners. Most of R practitioners use R Studio, and while they do offer a free community version, enterprise edition is a bit expensive. I use the community version only when I compete in Kaggle, haven't won anything yet :( Commercially I find Sublime Text a good alternative, since I use R only for proof of concepts, I'll return to this point later in my Python discussion. It's a very good and much cheaper IDE fully supporting R's syntax and being bundled with lots of top notch features.

Matlab/Octave


Matlab has been a top preference of algorithmists for years. It's a full blown solution with packages for every possible scientific field. The pricing however is what drove many away from it, as even a home usage version costs around 200$. As a result Octave was created to provide an open source alternative to run Matlab code. It's completely free, but you get what you pay for. There is no IDE and it is normally used through its interactive command line interface. Paired with Sublime Text it however provides a decent alternative to Matlab, if you don't need anything too fancy.

One thing to be aware of, is the fact that Octave's developers try to make Octave syntax "superior" to Matlab's. If it tries to be "better", it thus tries to be different, which is not in line with the reasons most people use it for. In my experience, running stuff developed in Matlab doesn't ever work in one go, except for the really simple, really short stuff. For any sizeable function, I always have to translate a lot of stuff before it works in Octave, if not re-write it from scratch.

Scala


The data boom has been sparked by the appearance of Hadoop ecosystem. Since Hadoop and all it's supportive tools are developed in Java, it sort of makes sense to analyse the data using the same language. Java however is not a very intuitive and easy to learn language, especially for non programmers. It was created specifically with the goal of being a better language, shedding those aspects of Java which it considered restrictive, overly tedious, or frustrating for the developer. Despite some appearance in data science community, it still remains mostly for data engineers usage and with a brisk pace of Python in Big Data domain, it has lost even more of it's relevance.

Python


According to Gartner:
the need for data scientists growing at about 3x those for statisticians and BI analysts, and an anticipated 100,000+ person analytic talent shortage through 2020
And as enterprises struggle to put data to work, they're also struggling to find qualified data scientists. More often than not, however, such data scientists may already work for them and likely have some familiarity with Python. It also much easier from the development perspective to implement everything in one language, since we really want those algorithms will need to get their way into working product some day.

As Tal Yakoni pointed out:
Nothing is more annoying than parsing some text data in Python, finally getting it into the format you want internally, and then realizing you have to write it out to disk in a different format so that you can hand it off to R or MATLAB for some other set of analyses
While R and Matlab production servers do exist, they are immensely expensive. That's why for decades, Matlab programs were reimplemented in C++ or Java to cut the costs. Python, however, can be used on any machine and with services like Amazon EC2, you can pay per hour of usage, making it affordable for any budget suffocated start-up. Moreover because Python is an object-oriented programming language, it’s easier to write large-scale, maintainable, and robust code with it than with R or Matlab. Using Python, the prototype code that you write on your own computer can be used as production code if needed, thus cutting enormously time to market time.

Python still lacks some of R's richness for data analysis, but it is closing the gap really fast. There are plenty of packages for any flavour, implemented in C, making them extremely fast.

Hope you enjoyed the article and feel free to share and comment.


Thursday, January 29, 2015

Big Data Buzz Words Overview


I wanted to start this blog by a quick overview of current state of Big Data playground. There was a lot of noise during this year from everywhere making it nearly impossible for a newcomer to learn this world without being overwhelmed. So let's start.

How does Big Data differ from NoSQL?


NoSql is a type of database, which provides a mechanism for storage and retrieval of data modelled in means other than the tabular relations used in relational databases. The most prominent representatives are Cassandra, MongoDB, Neo4j and Redis each taking a different approach in data representation. It is important to notice, that NoSql databases have nothing to do with the amount of stored data, but merely it's representation. On contrary Big Data is commonly referred to technologies used to store and operate on huge amounts of data. Usually it is referred to a Apache Hadoop ecosystem. In fact Hadoop it's file system based databased, so it's also NoSql database. However if Redis can be used for storing any amount of data, Hadoop was specifically designed to store data of large amounts.

Who are Cloudera, Hortonworks and MapR?


Since Hadoop is an open source software, several companies have sprung over the years providing support and additional useful tools to the platform. These companies are Cloudera, Hortonworks and MapR, each distributes it's own distribution of Hadoop.

Cloudera has been here for the longest time since the creation of Hadoop. Hortonworks came later. While Cloudera and Hortonworks are 100 percent open source, most versions of MapR come with proprietary modules, like using proprietary file system MapR-FS instead of HDFS. 

Each vendor/distribution has its unique strength and weaknesses, each have certain overlapping features as well. If you are looking to make the most of Hadoop’s immense data processing power, you should make a comparative study in them. To help you start, have a look at comparison table to the right by various features.

For more detailed comparison, you can request a 65 page free comparison booklet from Altoros.

Hadoop Eco System


Despite being an ingenuous piece of software, Hadoop is very difficult to operate on and makes it's developers' life a living hell. Even the easiest aggregate operation require one to implement a Hadoop job using MapReduce paradigm.

To make life simpler, good people from Facebook invented and open-sourced Hive project, which converts SQL to a series of MapReduce jobs. It tries to look like MySQL by storing table schemas in it's local database.

The problem with Hive, is that it was never developed for real-time and in memory processing. It was built for offline batch processing kinda stuff. Best suited when you need long running jobs performing data heavy operations like joins on very huge datasets. For folks wishing interactivity, other tools were developed by different vendors, each serving the same purpose more or less. These include Impala from Cloudera, Presto from Facebook and Apache Drill, heavily pushed by MapR. Apache Drill has similar goals to Impala and Presto – fast interactive queries for large datasets, and like these technologies it also requires installation of worker nodes. However, unlike Impala and Presto, Drill aims to support multiple backing stores (HDFS, HBase, MongoDB).

YARN and Spark


Appetite comes with eating. And after seeing what Hadoop can do, people wanted to do even more, but much quicker. It turns out, MapReduce wasn't the best architecture for real-time processing so in 2012 a sub-project called YARN was started promising the solution. Sometimes called MapReduce 2.0, YARN is a software rewrite that decouples resource management and scheduling capabilities from the data processing component, enabling Hadoop to support more varied processing approaches and a broader array of applications, the most prominent are Spark and Storm.

Apache Spark is an in-memory distributed data analysis platform-- primarily targeted at speeding up batch analysis jobs, iterative machine learning jobs, interactive query and graph processing. One of Spark's primary distinctions is its use of RDDs or Resilient Distributed Datasets. RDDs are great for pipelining parallel operators for computation, allowing it to run programs up to 100x faster than Hadoop MapReduce in memory, or 10x faster on disk.

Apache Storm is focused on stream processing or what some call complex event processing. Storm implements a fault tolerant method for performing a computation or pipelining multiple computations on an event as it flows into a system.

Machine Learning with MLlib and Mahout


As data scientists we're interested in data insights, rather than the way it's stored. However since large amounts of data were stored in Hadoop, people needed a way to access it and more importantly to be able to learn from it. The answer didn't made itself to wait in a form of Apache Mahout for Hadoop MapReduce and Apache MLlib for Spark environment.

Hope you've found this article useful and are welcome to comment and share below.