{"cells":[{"metadata":{"_cell_guid":"0b086443-710e-4f84-8c79-b177a7ef04ef","_uuid":"283867da9acaa6de0d309f4d56b3104ec17d58f9"},"cell_type":"markdown","source":"# First look at the features of the 10MM-row training set\n\nHere I use a random sample of 10,000,000 rows from the training set.   \n\n1. I did the random sampling of 10,000,000 rows,\n             train_sample = train.sample(n=10000000, random_state=4321)\n2. Saved the sampled data frame into a .csv file,\n             train_sample.to_csv(\"train_10mln.csv\")\n3. Uploaded the file into Kaggle. "},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"# load libraries and check what we have in the \"../input\" directory\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nprint(os.listdir(\"../input\"))","execution_count":4,"outputs":[]},{"metadata":{"_cell_guid":"e24326b8-f63f-4453-8e97-430b0121397f","_uuid":"e610ff10e403f975c115fcfc5864cad3acd55199","trusted":true,"collapsed":true},"cell_type":"code","source":"test = pd.read_csv('../input/talkingdata-adtracking-fraud-detection/test.csv', encoding=\"utf-8\")\ntrain = pd.read_csv('../input/talkingdata-adtracking-subsample-2/train_10mln.csv', encoding=\"utf-8\")","execution_count":5,"outputs":[]},{"metadata":{"_cell_guid":"fd367f75-0fec-429e-ab8a-3c3e3f1ba759","_uuid":"130bb4bcc0f5ec26a3d4aa629afa079f18daae7f"},"cell_type":"markdown","source":"### Helper function just for printing"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"trusted":true},"cell_type":"code","source":"from IPython.display import Markdown, display_html, display, HTML\n\ndef print_style(string, color=None):\n    colorstr = \"<span style='color:{}'>{}</span>\".format(color, string)\n    display(Markdown(colorstr))\n    \n## Displays the first 3 rows \n## as well as the type, the number of na values and the number of unique values in the columns.   \ndef display_df(df, name, nrows=3):\n    df_shape = df.shape\n    display(HTML('<h4 style=color:{}><h4>{}</h4></h4>'.format('purple',name)))\n    display(HTML('<h4 style=color:{}>{}</h4>'.format('blue', df_shape)))\n    display(df.head(nrows))\n    display(pd.DataFrame([df.dtypes, np.sum(df.isnull()), df.nunique()],\n                         columns=df.dtypes.index,\n                         index=['type', 'na', 'nunique']))","execution_count":8,"outputs":[]},{"metadata":{"_cell_guid":"31dddce3-242e-4cde-badc-f1cfb354c120","_uuid":"e44a6f46e68f2901ac4fe6fa4f652056b116d1e7"},"cell_type":"markdown","source":"## First look at the data"},{"metadata":{"_cell_guid":"8b3653e6-09f2-4b8a-bce1-13abcaee7cab","_uuid":"ff6c3207b034f39438321f18a298c9145908e3c6","trusted":true},"cell_type":"code","source":"display_df(test, name='test')\ndisplay_df(train, name='train')","execution_count":9,"outputs":[]},{"metadata":{"_cell_guid":"4aa3c726-935b-4f7b-ae14-34d1981e32fb","_uuid":"6bb3abe9bf61a767ee0b2c49ee5e320e2672432a"},"cell_type":"markdown","source":"**Note,** the column *Unnamed: 0* tells what rows are selected from the original data set. Thus, \" *Unnamed: 0* is not present in the original training set.\n"},{"metadata":{"_cell_guid":"21ea8013-c432-458a-92a2-adbe1cbc52f9","_uuid":"b0bcdeac0e9acc8b5a85380eb3be15bd17a47306"},"cell_type":"markdown","source":"## Visualize is_attributed\n\nOf course, we are dealing with a highly imbalanced dataset."},{"metadata":{"_cell_guid":"cf046253-d782-4e30-a8e0-24a3801a382d","_uuid":"387bfa0885e13b1dad1cb58965424b42571aa71a","trusted":true},"cell_type":"code","source":"import seaborn as sns\n\n# Look at the classes of is_attributed variable in percentages\nfraud, normal = train['is_attributed'].value_counts(normalize=True)\nprint('fraud (class 0): {}, normal (class 1): {}'.format(fraud, normal))\nax = sns.countplot(\"is_attributed\", data = train)","execution_count":10,"outputs":[]},{"metadata":{"_cell_guid":"ed6bde66-0140-4a57-a671-8ab80ec78980","_uuid":"9fbd6953c44a8d4ec44de4e7fcd49ac58b6bcc80","collapsed":true},"cell_type":"markdown","source":"## Pairwise Relationships"},{"metadata":{"_cell_guid":"12da4b6e-b904-4ade-a038-3f74b97c5159","_uuid":"f10ab23696b8b0007a75f2441b73dc1ec104b53d","trusted":true},"cell_type":"code","source":"corr_train = train.corr()\nprint_style('<h4 style=\"font-family:courier; font-size:200%; color:green \">train</h4>')\nsns.heatmap(corr_train, \n            cmap=\"Greens\",\n            annot=True) ","execution_count":15,"outputs":[]},{"metadata":{"_cell_guid":"bffa38aa-8b56-4fe0-8f6a-97c36236bbd9","_uuid":"a9d7099e96536574685880d269fbe7c079e5820a","trusted":true},"cell_type":"code","source":"corr_test = test.corr()\nprint_style('<h4 style=\"font-family:courier; font-size:150%; color:green \">test</h4>')\nsns.heatmap(corr_test, \n            cmap=\"Greens\", \n            annot=True)","execution_count":16,"outputs":[]},{"metadata":{"_cell_guid":"6ab534cb-0c9f-4c10-980f-390c74bf0878","_uuid":"80074e1fe9c134134a09a4ba42c192307a920466"},"cell_type":"markdown","source":"### Discussion\n\n* **device** and **os** are the most correlated features in the training set, as one would imagine.\n* **app** is correlated with **device** and **os** in the training set.\n* There are almost no correlations between the features in the test set.\n    * slightly negative correlation between **app** and **channel**\n\nThere are correlated features in training set that are not correlated in the test set.  This is an indicator that there are at least couple of features that behave differently in train set compared to the test set. "},{"metadata":{"_cell_guid":"1a06fe71-5145-46ad-9584-478d252a49fe","_uuid":"d2f91d21b0a3e447e87c7e83962529b5f721c96b"},"cell_type":"markdown","source":"## Feature Distributions (unshared x and y axes)"},{"metadata":{"_cell_guid":"585cbb04-7d46-4a4b-87f0-a1af19aa7211","_uuid":"4db9bab04cd583bf68f10a76b9a44e6d68695374","trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nfeatures = [x for x in ['ip', 'app', 'device', 'os', 'channel']]\n\n# separate fraud from normal\n# fraud here are the rows that didn't result in the purchase, thus the value 0\nnormal = train[train.is_attributed == 1][features]\nfraud = train[train.is_attributed == 0][features]\n\nprint_style('<h1 style=\"font-family:courier; font-size:200%; color:green \">ip, app, device, os, channel distributions</h1>')\n\n#### Uncomment the line below and comment out the second line below \n#### to see the histograms with the same x and y axes\n#f, axes = plt.subplots(5, 2, figsize = (12,20), sharex='all', sharey = 'all')\nf, axes = plt.subplots(5, 2, figsize = (15,25))\n\n## there are 5 rows(i) and 2 columns (1 column for train, the other for test set).\nfor i in [0,1,2,3,4]:\n        #train\n        axes[i, 0].hist(fraud[features[i]], label = 'fraud', alpha = 0.5, color = 'red')\n        axes[i, 0].hist(normal[features[i]], label = 'normal', alpha = 0.5)\n        axes[i, 0].legend(loc='upper right', prop={'size': 15})\n        axes[i, 0].set_title('{} (train)\\nnunique = {}\\nnunique(fraud) = {}\\nnunique(normal) = {}'.format(features[i],\n                                                                                                          train[features[i]].nunique(),\n                                                                                                          fraud[features[i]].nunique(), \n                                                                                                          normal[features[i]].nunique()), \n                             size=17)\n        # test\n        axes[i, 1].hist(test[features[i]].values, alpha = 0.5)\n        axes[i, 1].set_title('{} (test)\\nnunique = {}'.format(features[i], \n                                                              test[features[i]].nunique()), \n                             size=17)\n\nf.subplots_adjust(hspace=0.8)","execution_count":19,"outputs":[]},{"metadata":{"_cell_guid":"6d92ee0b-6a82-4b3d-9a6b-498e03d3ffdd","_uuid":"747a1fccb45de15a7c19e93910ef1a01817a07e0"},"cell_type":"markdown","source":"## The distribution of *click_time*"},{"metadata":{"_cell_guid":"9466e89d-51f6-448b-a7c8-13dc0ce41ff3","_uuid":"411c258d26c5cb202ce687212c777bce5ef39f22","trusted":true},"cell_type":"code","source":"train[\"click_time\"] = train[\"click_time\"].astype(\"datetime64\")","execution_count":20,"outputs":[]},{"metadata":{"_cell_guid":"c6dfdd7c-b3ce-447e-b1f4-b9f643973f4f","_uuid":"544fbac0f184d1f63cbddc13aaff4631472207b6","trusted":true},"cell_type":"code","source":"print_style('<h1 style=\"font-family:courier; font-size:200%; color:green \">click_time distribution (train)</h1>')\ntrain.groupby([train['click_time'].dt.hour]).count().plot(kind=\"bar\", figsize=(20,10))","execution_count":21,"outputs":[]},{"metadata":{"_cell_guid":"d8559429-e1d4-4dfd-9c44-f85b286a0a4d","_uuid":"472d70121b58e9e759d499fcebd453e21b814775","trusted":true},"cell_type":"code","source":"test[\"click_time\"] = test[\"click_time\"].astype(\"datetime64\")","execution_count":22,"outputs":[]},{"metadata":{"_cell_guid":"292eaeaa-951e-4115-a6e1-24d8a57de91e","_uuid":"c41bf9ebb7f68719f08c286ede36d87549347957","trusted":true},"cell_type":"code","source":"print_style('<h1 style=\"font-family:courier; font-size:200%; color:green \">click_time distribution (test)</h1>')\ntest.groupby([test['click_time'].dt.hour]).count().plot(kind=\"bar\", figsize=(20,10))","execution_count":23,"outputs":[]},{"metadata":{"_cell_guid":"2306b08c-18e4-471a-88df-6f27814fa513","_uuid":"9ca29a62726f4caefb5e5d17ecd2effddd046b89","collapsed":true,"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}