{"cells":[{"metadata":{},"cell_type":"markdown","source":"**Handling Imbalanced datasets**    \nThere are 4 commonly used methods for handling imbalanced datasets   \n**Sampling methods for handling imbalance datasets**\n1. Downsampling (Under sampling)\n2. Upsampling (Over sampling)\n3. Upweighting\n4. Combination of over- and under-sampling\n\n1.**Downsampling(Under sampling):**  \n In this the classes are balanced by performing resampling on majority class. It reduces the size of majority class in order to match with the size of the minority class.\n There are different methods to perform this sampling. one of those method is Random Under sampler, which selectes a subset of data randomly to balance the data in target classes.\n\n2.**Upsampling (Over sampling):**\n  This method is used when there is an insufficient data. In this the resampling performed on minority class. It increases the size of the minority class in order to match with the size   of the majority class.One best example of this sampling is SMOTE (Synthetic minority over sampling Technique).It works by creating synthetic samples from the minor class to balance the   data in target classes.\n  \n3.**Upweighting:**\n  Here the classes are balanced by scaling the weight of minority class. This weight is scaled by taking the ratio of number of samples in majority class to the number of samples in       minority class. A minority class weight of 30 (say) means the model treats the minority class as 30 times as important as it would majority class of weight 1. \n  \n Generally these 3 methods are most widely used to handle imbalanced data in building ML models. For DL models there is a separate batch generator to handle highly imbalanced data.\n \n For more details about sampling methods, please check below link,  \n https://imbalanced-learn.readthedocs.io/en/stable/api.html#module-imblearn.over_sampling\n "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"#Import required libraries\nimport numpy as np \nimport pandas as pd \nimport warnings\nwarnings.filterwarnings('ignore')\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"df=pd.read_csv('../input/talkingdata-adtracking-fraud-detection/train_sample.csv')\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#convert timestamp to datatime\ndef todatetime(df):\n    df['click_time']=pd.to_datetime(df['click_time'])\n    df['click_hour']=df['click_time'].dt.hour\n    df['click_day']=df['click_time'].dt.day\n    df['click_weekday']=df['click_time'].dt.weekday\n    df['click_month']=df['click_time'].dt.month\n    df['click_year']=df['click_time'].dt.year\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df=todatetime(df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df=df.drop(['click_time'],axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df=df.drop('attributed_time',axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Shuffling observations\ndf=df.sample(frac=1)\ndf","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Distribution of classes in train dataset before sampling**"},{"metadata":{"trusted":true},"cell_type":"code","source":"import seaborn as sn\nimport matplotlib.pyplot as plt\nplt.figure(figsize=(10,8))\nsn.countplot(x='is_attributed',data=df)\nplt.ylabel('Number of clicks')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Classes:**  \n0- User will not download an app after clicking a mobile app advertisement    \n1- User will download an app after clicking a mobile app advertisement"},{"metadata":{"trusted":true},"cell_type":"code","source":"df['is_attributed'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The target label data is highly imbalanced (99.75:0.25)%. This needs to be balanced to make the model to be generalize well."},{"metadata":{"trusted":true},"cell_type":"code","source":"target_label=df['is_attributed']\ntarget_label.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ones=df[df['is_attributed']==1]\nzeros=df[df['is_attributed']==0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df=df.drop(['is_attributed','ip'],axis=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Split the dataset into train & test sets**"},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nx_train,x_test,y_train,y_test=train_test_split(df,target_label,test_size=0.2,random_state=42)\nprint(x_train.shape,y_train.shape)\nprint(x_test.shape,y_test.shape)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Create validation dataset from train dataset**"},{"metadata":{"trusted":true},"cell_type":"code","source":"x_train,x_val,y_train,y_val=train_test_split(x_train,y_train,test_size=0.1,random_state=42)\nprint(x_train.shape,y_train.shape)\nprint(x_val.shape,y_val.shape)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Building LGBM model**"},{"metadata":{"trusted":true},"cell_type":"code","source":"import lightgbm as lgb\n#load datasets in lgb formate\ntrain_data=lgb.Dataset(x_train,label=y_train,free_raw_data=False)\nvalidation_data=lgb.Dataset(x_val,label=y_val,free_raw_data=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**LGBM basemodel**"},{"metadata":{"trusted":true},"cell_type":"code","source":"#set parameters for training\nparams={ 'num_leaves':160,\n        'object':'binary',\n        'metric':['auc','binary_logloss']\n       }","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\n#Original LGB model before sampling\nnum_round=100\ndef lgb_basemodel(x_train,y_train):\n    lgb_model=lgb.train(params,train_data,num_round,valid_sets=validation_data,early_stopping_rounds=20)\n    return lgb_model","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**LGBM model with Downsampling**"},{"metadata":{"trusted":true},"cell_type":"code","source":"#LGBM model after resampling the data using Under sampling techniques\nfrom imblearn.under_sampling import RandomUnderSampler \ndef lgb_downsampling(x_train,y_train):\n    lgb_enn=RandomUnderSampler(random_state=42)\n    x_resample,y_resample=lgb_enn.fit_resample(x_train,y_train)\n    train_data=lgb.Dataset(x_resample,label=y_resample,free_raw_data=False)\n    lgb_model=lgb.train(params,train_data,num_round,valid_sets=validation_data,early_stopping_rounds=20)\n    return lgb_model,x_resample,y_resample;","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**LGBM model with Upsampling**"},{"metadata":{"trusted":true},"cell_type":"code","source":"#LGBM model after resampling the data using Up sampling techniques\nfrom imblearn.over_sampling import SMOTE  #Balances the classes by performing upsampling on minority class\ndef lgb_upsampling(x_train,y_train):\n    lgb_smote= SMOTE(random_state=42)\n    x_resample,y_resample=lgb_smote.fit_resample(x_train,y_train)\n    train_data=lgb.Dataset(x_resample,label=y_resample)\n    lgb_model=lgb.train(params,train_data,num_round,valid_sets=validation_data,early_stopping_rounds=20)\n    return lgb_model,x_resample,y_resample;","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**LGBM model with Upweighting**"},{"metadata":{"trusted":true},"cell_type":"code","source":"weight_factor=zeros.shape[0]/ones.shape[0]  # Ratio of number of samples in majority class to number of samples in minority class\nprint('Weight factor is %0.2f'%(weight_factor))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#set parameters for training\nparams1={ 'num_leaves':160,\n        'object':'binary',\n        'metric':['auc','binary_logloss'],\n        'scale_pos_weight':397.41                 #Weight of minority class\n       }","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#LGBM model using Upweighting technique for handling the imbalanced data\ndef lgb_Upweighting(x_train,y_train):\n    lgb_model=lgb.train(params1,train_data,num_round,valid_sets=validation_data,early_stopping_rounds=20)\n    return lgb_model;","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Train the LGBM models**"},{"metadata":{"trusted":true},"cell_type":"code","source":"#Basemodel\nlgb_basemodel=lgb_basemodel(x_train,y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Downsampling model\nlgb_downsampling,x_down,y_down=lgb_downsampling(x_train,y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"y_down_df=pd.DataFrame(y_down)\ny_down.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.hist(y_down);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Upsampling model\nlgb_upsampling,x_up,y_up=lgb_upsampling(x_train,y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"y_up_df=pd.DataFrame(y_up)\ny_up.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.hist(y_up);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Upweighting model \nlgb_upweighting=lgb_Upweighting(x_train,y_train);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Testing models on unseen dataset**"},{"metadata":{"trusted":true},"cell_type":"code","source":"#Basemodel\ny_base=lgb_basemodel.predict(x_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Upsampling\ny_upsampling=lgb_upsampling.predict(x_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Downsampling\ny_downsampling=lgb_downsampling.predict(x_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Upweighting\ny_upweighting=lgb_upweighting.predict(x_test)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Plot confusion matrix for all Logistic models**"},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, classification_report, roc_curve, roc_auc_score\nimport scikitplot as skplt","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Basemodel\nskplt.metrics.plot_confusion_matrix(y_test,y_base>0.5,normalize=False,figsize=(12,8),title='Confusion matrix for base model')  \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Upsampling\nskplt.metrics.plot_confusion_matrix(y_test,y_upsampling>0.5,normalize=False,figsize=(12,8),title='Confusion matrix for upsampling model')  #0.5 is threshold value\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#downsampling\nskplt.metrics.plot_confusion_matrix(y_test,y_downsampling>0.5,normalize=False,figsize=(12,8),title='Confusion matrix for downsampling model')  #0.5 is threshold value\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Upweighting\nskplt.metrics.plot_confusion_matrix(y_test,y_upweighting>0.5,normalize=False,figsize=(12,8),title='Confusion matrix for upweighting model')  #0.5 is threshold value\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Classification report for all Logistic models**"},{"metadata":{"trusted":true},"cell_type":"code","source":"#Base model\ncm_base=classification_report(y_test,y_base>0.5)\nprint(cm_base)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Upsampling model\ncm_up=classification_report(y_test,y_upsampling>0.5)\nprint(cm_up)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Downsampling model\ncm_up=classification_report(y_test,y_downsampling>0.5)\nprint(cm_up)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Upweighting model\ncm_upweight=classification_report(y_test,y_upweighting>0.5)  # 0.5 is threshold value\nprint(cm_upweight)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Conclusions:-**\n\n1. Both LGB base & upweighting models are performing same. These models are not at all learning positive labels, so always predicting negative class.\n2. LGB Upsampling (SMOTE) permormed well among all models. This model performance can be improved further by hypermeter tuning\n3. LGB downsampling model performing reasonally but not good.\n\nAll above models are data dependent, so they may perform well on some datasets but not all. It is better to build all models & choose best among for prediction.\nFor large datasets, Ensemble methods can be employed, but these are computationally expensive. Also it is good to check all other sampling methods along with SMOTE & RandomUnderSampler for better understanding of models to handle unbalanced datasets."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}