Wednesday, April 25, 2012

Secondary Sort using Python and MRJob

MRjob is an excellent module developed as open source project by Yelp.. I chosed mrjob because of the following features


  1. It provides a seamless JSON reader and writer (i.e the mapper can read json lines and convert them into   lists)
  2. We can test hadoop job locally (in windows or unix) on a small dataset without actually using huge hdfs files (quick !!)
  3. Can orchestrate many mappers and reducers in the same code 

My task is to parse json formatted web log files and parse them, say the columns are sessionid,stepno and data. so the psuedo-code
  1. Read the json files using mrjob protocol
        DEFAULT_INPUT_PROTOCOL = 'json_value'
        DEFAULT_OUTPUT_PROTOCOL = 'repr_value'  #  output is delimited

  2.  yield sessionid, (sessionid,stepno,data) from mapper
                      the Mapreduce will make sure that all sessionids(key) goes to same mapper.. with the remaining values sent as a dictionary (value) to make it easier for us to srt in reducer

 3.  Use sorted from itertools of python module to sort by stepnumber in reducer

 def reducer(self, sessionId, details):
                sdetail = sorted(details, key=lambda x: x[1])  # sorting by stepno for each session
                for d in sdetail:
                        line_data='\t'.join(str(n) for n in d)

We are doing the secondary sort to scan through each events as the sequence is very important to do the funnel analysis of logs..

Complete code :

import sys,time
#sys.path.append('/usr/lib/python2.4/site-packages/')
from mrjob.job import MRJob
from mrjob.protocol import JSONValueProtocol
from itertools import groupby
from operator import itemgetter, attrgetter

class uet(MRJob):
        DEFAULT_INPUT_PROTOCOL = 'json_value'
        DEFAULT_OUTPUT_PROTOCOL = 'repr_value'

        def mapper(self, _, line):
                        sessionId = line['sessionId']
                        data = line['data']
                        if len(sessionId) < 13:
                                for i in range(len(data)):
                                        no = data[i]['no']
                                        yield sessionId,(sessionId, no,data)

        def reducer(self, sessionId, details):
                sdetail = sorted(details, key=lambda x: x[1])  # sorting by stepno for each session
                for d in sdetail:
                        line_data='\t'.join(str(n) for n in d)
                        print str(line_data)


if __name__ == '__main__':
    uet.run()



Tuesday, April 17, 2012

Top NBA Players - by twitter followers

I really dont have any idea of how advertising companies decide on the price for certain celebrities. Because its very hard to measure the direct relationship to the sales and to derive an ROI. Also Im not sure if any celebrity with enormous following can make a big impact. I did data collection for fun to see who is actually more popular in Twitter and have more following. I used infochimps rest api calls to get the aggregated information and formatted files to make it readable by Tableau. (used Python for ETL).. seems like SHAQ even after out of the league has a greater fan following than active player.. I havent included Kobe and Rose.. and I tried my best to get the official ids of each player..


Tuesday, April 10, 2012

DW in Hive - handling big dimensions

This is a an issue bothering our reporting data quality. How to handle big dimension tables in Hive data warehouse. How to balance performance and data-quality..

Problem statement:
Currently dimension hive table (dim_customer) is partitioned by date. The daily incremental creates the new partition so that we can improve the performance at the reporting side. This poses 2 critical issues
1. slowly changing dimension is lost
2. Compromise in data-quality
3. Have to filter by the dimension table in the reporting

Solution:
The only way to solve this problem is to have the dimension table as one big Hive table instead of partitions. But this creates issues with the refresh strategy and overhead on reporting Hive query..

The following is a recipe to solve this block.. step.1 is certainly the priority

1. Increase processing power
Hadoop is not only about mega storage it is also about mega processing ..so if we process big files then we got to have good number of nodes. say we have 30 nodes to process partitioned dimension table..we have to move to 120 nodes for single dimension file strategy. Processing power is tripled - it is directly proportional !.

2. Use SQOOP merge
We cannot extract the whole table from transactional system every time..source transactional systems might not allow.. we can only capture the change data. SQOOP merge comes handy for this purpose. We can overwrite only the incremental records in the hive table (type 1 SCD). Again this is a Map reduce program.. we need processing power..

3. Use Bucketed Hive tables
Hive performance block comes while joining. We can create a table with bucketing.. like hashing index on the customer_id.


If the tables being joined are bucketized, and the buckets are a multiple of each other, the buckets can be joined with each other. If table A has 8 buckets are table B has 4 buckets, the following join

SELECT /*+ MAPJOIN(b) */ a.KEY, a.value
FROM a JOIN b ON a.KEY = b.KEY
can be done on the mapper only. Instead of fetching B completely for each mapper of A, only the required buckets are fetched. For the query above, the mapper processing bucket 1 for A will only fetch bucket 1 of B. It is not the default behavior, and is governed by the following parameter

Bucketing and Sqoop merge requires good planning and metadata management..
Managing Hadoop from the scratch is challenging as we bump into the limits sooner, we need to adapt quickly else data might grow beyond limits. One advantage( and complexity) is that the internal processing (mapreduce) is open and its upto the developer to improve. And the biggest advantage of all is scaling out.. you can add nodes easily to really make a difference..






Thursday, March 29, 2012

Data science -- the cool scientist without white gowns

"Scientist" is a cool word during my school days, wanna become one but donno on what. All I see as scientists wore white gown with colorful liquids around. But later during college and working days scientists seemed to be boring people with no personal life, accumulated in educational institutes with college kids helping around.

Recently there is a profile called "Data scientist" all over the BI market and started appearing in every article where Hadoop / Big data is there. It must be some part what a BI/DW person is doing with some specialization. Yes it is the formula,as for me..

BI + big-data + statistics + scripting + visualization = Data scientist

Ok can be scientist and work for a corporate or invent something new for you name? possibly...
But seems like a lot to cover , learn and experience at work. Maybe not really if we are in right job. Im just listing the very higher level outline..(and it is not limited to..) and my intention is not to oversimplify, but certainly to simplify the puzzle..

BI
- DW work like ETL , databases, SQL with exposure to enterprise setup. Collecting data from heterogenous data sources. Log analysis. Dimensional modelling, DW architecture. DB performance.

Big-data
Hadoop is the first thing comes to mind for Big-data.. but good to know about noSQL dbs.
Mapreduce - shared nothing architecture - need for MR - use cases - tools available - pros and cons

Statistics
Basics - application of statistics in real-world - R programming

Scripting
Perl, Python, Java

Visualization
Reporting (I like Tableau), Complex SQLs , Ability to tell a story with data - by whatever way you effectively deliver..

I would like to list some of the coolest learning materials available for above topics..

I believe the thirst for discovery, admiring the hidden secret in the boring pile of data would make a Data Scientist..









Tuesday, February 21, 2012

Dual table in Hive

Since there is no Dual table in Hive..

I have created a dummy dual with dummy value ‘X’.

hive> CREATE TABLE dual (dummy STRING);

hive> load data local inpath '/local/user/dw/hive/dual.txt' overwrite into table dual;

Now we can use this table to select string values. Like

hive> select 'name','place','age' from dual;

to get current timestamp

hive> select unix_timestamp() from dual;

Friday, February 17, 2012

Talent is overrated - Geoff Colvin

Simply to say this book is good and can be life changing for you or may be to your kid.

I will get some new idea everyday, while taking bath or watching some movie or during half-sleep @ office or even at workout. Some of them are great, turned out to be great things which are implemented by someone, some are already existing and I'm ignorant of it. Many self analyzing thought made what I'm now. But the point is I'm not great, maybe by the scale of this society. Then how to define greatness - a person achieved enviable status in a field. The word field is the key here.

Being a Software Engineer I cannot become maddeningly wealthy like Buffet just by reading 'the intelligent investor'. Im already a hard-wired employee in a field who's time is controlled by someone else. And if you are married with kids personal time is a joke. There are success stories everywhere in Tennis, basketball, investing and especially entrepreneurship. Just by looking at it or researching will not make us the same. You might have known this but yet human greed takeover sometimes and we will start day-dreaming. Mostly the end result is disappointment.

If I have to become a cricketer I should have that thought by at least age 10 and played relentlessly until my skin is totally tanned. At this point this book makes a clear the 10000 hours point which Malcolm Gladwell's 'Outliers' is also talking about. Reading "Talent is overrated" book I can recall Arnold Schwarzenegger's quote ..

"Number one, come to America. Number two, work your butt off. And number three, marry a Kennedy."

The book says these very clearly

1. Choose your field, if you are already in a field and earning from it, its hard to leave and try new. But children has the advantage to excel enormously.
2. Deliberate practice .. it must be conscious, measurable and improving
3. Dont stick in the OK plateau if you want to become world-class
4. Never look around to excel in other area which you cannot afford to spend time on.
5. Get family support for what you are doing

Its a revealing read and can be read again if you fail to succeed in some area. The world is so competitive that only the extraordinary can win. (even in a job interview ) Can be a great book if you are a parent which can change your kids life all together.


Friday, February 10, 2012

Shell script - snippets

Loop through dates

startdate=`/bin/date --date="2007-07-01" +%Y-%m-%d`
enddate=`/bin/date --date="2011-07-01" +%Y-%m-%d`

foldate="$startdate"
until [ "$foldate" == "$enddate" ]
do
echo $foldate
foldate=`/bin/date --date="$foldate 1 month" +%Y-%m-%d`
done