{"cells":[{"metadata":{"collapsed":true,"_uuid":"a9d02f97de7d869d99d59d11001e7da1ed853d5a","_cell_guid":"075d1735-08ef-459c-9ca5-d96379c84219","trusted":true},"cell_type":"code","source":"# This is a simple data exploration notebook. \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport os\n\nimport seaborn as sns\n%matplotlib inline","execution_count":46,"outputs":[]},{"metadata":{"_uuid":"2468c37acdb7e77347fd41e8be8beada61c1bbe7","_cell_guid":"a1125b4a-6e39-4080-893f-bd0b6e8d1043"},"cell_type":"markdown","source":"It is an attempt to extend the thoughtful work of fellow kagglers - yulia (https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns) and anokas (https://www.kaggle.com/anokas/talkingdata-adtracking-eda). Some of the data structures I had to recreate to fit my code requirements. "},{"metadata":{"_uuid":"ac1af3175a49676bc4a07fdea9204e5661eacf6a","_cell_guid":"a8ffd956-bb2a-4919-99ce-58f5d7b9e0f7"},"cell_type":"markdown","source":"\n"},{"metadata":{"_uuid":"d1122005a7fd91360bf1c2a5d149366ba62af319","_cell_guid":"30519237-6f22-43c1-83cf-da139efdcf78","trusted":true},"cell_type":"code","source":"print(os.listdir('../input'))","execution_count":5,"outputs":[]},{"metadata":{"_uuid":"fbe70f610d2fd7acdf1d8a8c3989a2b8fb0935d1","_cell_guid":"95680a3e-f3a2-438a-9824-5dbc3dad19ef"},"cell_type":"markdown","source":"I'm using the random training sample provided by the organizers (train_sample) which has 100000 click records. I'm choosing the random data since it'll give a relatively unbiased view of the data. Let's first have a peek into the dataset."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train= pd.read_csv('../input/train_sample.csv')\ntrain.head()","execution_count":47,"outputs":[]},{"metadata":{"_uuid":"48ddd2aa24cc6d819c6cf0d68a77afbaa1d1dfa0","_cell_guid":"764cf2f7-a500-424d-9731-34d2908d616b","trusted":true},"cell_type":"code","source":"train.info()","execution_count":48,"outputs":[]},{"metadata":{"_uuid":"b532f7743041ce4ed0b98d0b95854b6e6dec55b0","_cell_guid":"07d51503-9ac9-4f8a-9d46-104095e0c76c"},"cell_type":"markdown","source":"Converting the datatypes to categorical to remove the notion of order from the variables with the exception of check_time and attributed_time which are converted to datetime. "},{"metadata":{"collapsed":true,"_kg_hide-input":false,"_uuid":"d9f8c21395f000b62987b350a375191166a026f1","_cell_guid":"bab9e164-6ec0-4eee-b9f7-54cb97e362a2","trusted":true},"cell_type":"code","source":"cols=train.columns\ncols_time=['click_time', 'attributed_time']\ncols_categorical=[col for col in cols if col not in cols_time]\n\nfor col in cols_categorical:\n    train[col]=train[col].astype('category')\nfor col in cols_time:\n    train[col]=pd.to_datetime(train[col])","execution_count":49,"outputs":[]},{"metadata":{"_uuid":"90eecf0c3ade295b0cd7ad5a48bec34cc9cf3020","_cell_guid":"b140caf6-c5c1-48d0-8a27-11f663b63712","trusted":true},"cell_type":"code","source":"train.describe()","execution_count":50,"outputs":[]},{"metadata":{"_uuid":"edbfdb03ba4d57600b74dadb5737839020e243bd","_cell_guid":"7a7eaa99-6666-4180-8e26-576d2e581dfb"},"cell_type":"markdown","source":"One way to look at the data - these are the entitites we are dealing with - \n\n* **user** : defined by ip, device, os\n* **mode** :  defined by app, channel\n* **user journey** : Start - click_time (user came in contact with the mode) and End - attributed_time (user was acquired)"},{"metadata":{"_kg_hide-output":false,"_uuid":"0a5022af59238e39f0faedcc5648f071553060ee","_cell_guid":"e777a675-6e19-489b-97c7-06e1fdea01b9","trusted":true},"cell_type":"code","source":"train['conversion_time']=pd.to_timedelta(train['attributed_time']-train['click_time']).astype('timedelta64[s]')\nprint(train['conversion_time'].quantile(0.9)/3600)\ntrain.describe()","execution_count":51,"outputs":[]},{"metadata":{"_uuid":"5d8cec01bece81e577ef5a8f694efbaed6457f36","_cell_guid":"78887b31-ad17-4957-ae71-82b8666a8912"},"cell_type":"markdown","source":"Amongst the users which were acquired:\n*     Min time between user coming in contact with a mode and the user getting acquired is 4 seconds.\n*     50% of the users got converted within approx 5 mins.\n*     75% of the users took 1 hour to decide and download.\n*     90% of the users took 4 hours to decide and download.\n*     Max time taken by a user after coming in contact with an ad and going for download is approx 20 hours.\nHence within a day of contact with the ad, the users who could have been acquired were acquired. This is the engagement duration. Let's look at the full distribution."},{"metadata":{"collapsed":true,"_uuid":"0a77d17af4e8e9cd4fc69698f4a1206c61f05b66","_cell_guid":"84ee66fa-7001-4dba-8f35-f9e147feb442","trusted":true},"cell_type":"code","source":"ctimes=train['conversion_time'].dropna()","execution_count":11,"outputs":[]},{"metadata":{"_uuid":"249a1e4a8201b9dd92c22b7f427448c5cfa796e3","_cell_guid":"8ba2bbed-0911-42b1-8ce7-5e9209bdf015","trusted":true},"cell_type":"code","source":"sns.distplot(ctimes)","execution_count":12,"outputs":[]},{"metadata":{"_uuid":"26878a993d07d4e73260195911435888e429f6bc","_cell_guid":"05c8b5f6-d3ac-4c19-baf9-03896ba86a39","trusted":true},"cell_type":"code","source":"sns.distplot(np.log10(ctimes))","execution_count":13,"outputs":[]},{"metadata":{"collapsed":true,"_kg_hide-input":true,"_uuid":"c43b63c5a8aa1f5aafc12c36a48d20c33684831f","_cell_guid":"482af301-165e-44e6-9363-6b5568b6f754","_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"# min_time=np.nanmin(train['click_time'])\n# max_time=np.nanmax(train['attributed_time'])\n# print(min_time)\n# print(max_time)","execution_count":14,"outputs":[]},{"metadata":{"collapsed":true,"_kg_hide-input":true,"_uuid":"abe8e5aada9af227187e8c6162dc531b38745fa2","_cell_guid":"892d25d4-2b75-44ee-9699-2fe6ee1bf857","_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"# max_engagement_window_size=str(int(np.ceil(np.nanmax(ctimes)/3600)))+'H'\n# max_engagement_window_size","execution_count":15,"outputs":[]},{"metadata":{"_uuid":"680372b5a9a82f79a636912917f4a3550e515db2","_cell_guid":"41f7f454-6501-4189-8777-0b5e2f45a25b"},"cell_type":"markdown","source":"Now let's look at how the different variables are correlated in the 2 different categories - Successful Ad campgains versus Unsuccessful ones. "},{"metadata":{"_uuid":"4d6065ac8dbd7a7869e6a0b08996bcdefc334dac","_cell_guid":"c78a60e1-fa74-458e-b5b2-020de1a107ab","trusted":true},"cell_type":"code","source":"sns.pairplot(train[train['is_attributed']==1])","execution_count":16,"outputs":[]},{"metadata":{"_uuid":"59296e9647330953f0c565db3d4709dea1dd7188","_cell_guid":"470ff1a9-b3b9-4a5d-acfb-48898e1cc57a"},"cell_type":"markdown","source":"Here are a few observations for the successful Ad campaigns:\n* The conversion times are failry distributed over the different ip addresses and the channels.\n* There are few devices which have succesful conversions, however there are many devices which don't have any conversions at all. Maybe the ads on those devices are not that engaging(say particular tablets or desktop versions). "},{"metadata":{"_kg_hide-output":false,"_kg_hide-input":false,"_uuid":"f5be862f958d4bb7b5509898a2263cf8de15f6ea","_cell_guid":"d9e7f8d5-36a3-4b95-8bb3-023cf32c87aa","trusted":true},"cell_type":"code","source":"col_list=[col for col in train.columns if col!='conversion_time']\n# print(col_list)\n\nsns.pairplot(train.loc[train['is_attributed']!=1,col_list])","execution_count":17,"outputs":[]},{"metadata":{"_uuid":"50583cba4a40d3d7f92c818dafa087e6c69f5d33","_cell_guid":"d936c900-6c69-4738-a735-e67e6c113ff1"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"2b17746c38438a1313819ebdf5b709b98cb3231e"},"cell_type":"markdown","source":"Let's build a simple classifier"},{"metadata":{"collapsed":true,"_uuid":"d173e2db942995a861dd8d2f67e5ad4263f18ddd","_cell_guid":"96865d14-a898-41bf-962a-d68bc3a152c7","trusted":true},"cell_type":"code","source":"from sklearn.model_selection import StratifiedKFold, train_test_split, GridSearchCV\nfrom sklearn.pipeline import make_pipeline\nfrom sklearn.ensemble import RandomForestClassifier\n","execution_count":24,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c116fd8556fc98053d7a59a6ea573616575095df"},"cell_type":"code","source":"train['is_attributed'].value_counts()/train.shape[0]","execution_count":53,"outputs":[]},{"metadata":{"_uuid":"1b8afabf0fad0abcde24dd6d28781834dd54809d"},"cell_type":"markdown","source":"It's a very imabalanced dataset with 99.8% of the clicks not yielding to any downloads and only 0.2% of the data has downloads!!! We'll need evaluation metrics like roc_auc_score and precision_recall_curve"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"0f704635c775604f51a143ed6aa89ea56adc3656"},"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, classification_report\nfrom sklearn.metrics import roc_auc_score, f1_score, precision_score, recall_score, precision_recall_curve, roc_curve","execution_count":54,"outputs":[]},{"metadata":{"_uuid":"a766e6e6413af02653cbdda35f53572ed691550d"},"cell_type":"markdown","source":"First let's create some features for the click time. Attributed time is NaT for all those data points which have is_attributed=0, which is 99.8% of the data!"},{"metadata":{"trusted":true,"_uuid":"9ce11be16106846e5f47101e371ce25eede3ad4c"},"cell_type":"code","source":"train['chour']=train['click_time'].dt.hour\ntrain['cminute']=train['click_time'].dt.minute\ntrain['cday']=train['click_time'].dt.day\n","execution_count":55,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"b894b6ec2e8beb1c2154fdbbfe2b251b4f076da8"},"cell_type":"code","source":"del train['click_time']\ndel train['attributed_time']","execution_count":56,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"d3ad708788c6408dc5167451648532c3932d54bc"},"cell_type":"code","source":"Check how much of the data still has missing values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cad9db98a70fe068ffa5f71494ef590fbdefe88c"},"cell_type":"code","source":"train.isnull().any()","execution_count":58,"outputs":[]},{"metadata":{"_uuid":"32293cd1f82173792a48d0ebfbb0d907108cf3f3"},"cell_type":"markdown","source":"Conversion time still has NaN values corresponding to all the cases where there was no attributed_time (which is 99.8%) of the cases). Hence it's better to drop this column as well."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"26cfef8b987b674512528be111e5733f29650e23"},"cell_type":"code","source":"del train['conversion_time']","execution_count":59,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b5163612aa179c70b2f6b2f01bc47b6b1affb143"},"cell_type":"code","source":"X=train\nY=train['is_attributed']\ndel X['is_attributed']\nX.head()","execution_count":60,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"748cd1c131e731a048465df0943b9a1a0e20b2d6"},"cell_type":"code","source":"Xtrain, Xtest, Ytrain, Ytest=train_test_split(X,Y,test_size=0.2, random_state=32)","execution_count":61,"outputs":[]},{"metadata":{"_uuid":"86aeae0942f6a59fd1ee7ac700f12d4b81f8353e"},"cell_type":"markdown","source":"**Random Forest**"},{"metadata":{"trusted":true,"_uuid":"b1caabbcfff85b06091f6bfb3bdef44f446050bf"},"cell_type":"code","source":"param_grid={'n_estimators':np.arange(10,100,20), \n           'max_depth':[None, 5, 10],\n           'max_features':['auto','sqrt']}\nrfgrid=GridSearchCV(RandomForestClassifier(),param_grid,cv=StratifiedKFold(5))","execution_count":62,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"46cb60da021aea613d258994599f7fa7f6f15f9b"},"cell_type":"code","source":"rfgrid.fit(Xtrain, Ytrain)","execution_count":63,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0088b6938ebdfaf52c4a310bd567ebf75388c7ed"},"cell_type":"code","source":"rfgrid.score(Xtest,Ytest)","execution_count":68,"outputs":[]},{"metadata":{"_uuid":"5b8ec0b10ed99b75bbb623f74088de9b732e54ed"},"cell_type":"markdown","source":"Though the score is 99.76%, the class is highly imbalanced, hence other score measures are needed.."}],"metadata":{"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}