{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"## Importing the libraries:\n\n\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nimport random\n\npd.options.display.precision = 2\npd.options.display.max_rows = 100\n#pd.set_option('display.max_rows', 20)\n\nimport numpy as np\n\n# For Feature Selection\nfrom sklearn.feature_selection import SelectFromModel\n\n# For Feature Importances\n#from yellowbrick.model_selection import FeatureImportances\n\n# For metrics evaluation\nfrom sklearn.metrics import precision_recall_curve, classification_report, plot_confusion_matrix\n\n# For Data Modeling\n#from sklearn.model_selection import train_test_split ,cross_val_score,GridSearchCV ,  # To split the data in training and testing part   \nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC\n#from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\n\n# To Disable Warnings\nimport warnings\nwarnings.filterwarnings(action = \"ignore\")\n\nfrom scipy.stats import loguniform\nfrom sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold, GridSearchCV, RandomizedSearchCV\nfrom sklearn.metrics import confusion_matrix, roc_auc_score, roc_curve, classification_report, precision_recall_curve\nfrom sklearn.model_selection import RepeatedStratifiedKFold\nfrom sklearn.metrics import accuracy_score                          # For calculating the accuracy for the model\nfrom sklearn.metrics import precision_score                         # For calculating the Precision of the model\nfrom sklearn.metrics import recall_score                            # For calculating the recall of the model\n#from sklearn.metrics import precision_recall_curve                  # For precision and recall metric estimation\n#from sklearn.metrics import confusion_matrix                        # For verifying model performance using confusion matrix\nfrom sklearn.metrics import f1_score                                # For Checking the F1-Score of our model  \n#from sklearn.metrics import roc_curve  \n# For Roc-Auc metric estimation\n\n# Importing missingno for missing value plot\nimport missingno as msno\n\n# Importing datetime for using datetime\nfrom datetime import datetime\nimport warnings                                                     # Importing warning to disable runtime warnings\nwarnings.filterwarnings(\"ignore\")                                   # Warnings will appear only once","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:54:20.983578Z","iopub.execute_input":"2022-08-09T07:54:20.984059Z","iopub.status.idle":"2022-08-09T07:54:22.190969Z","shell.execute_reply.started":"2022-08-09T07:54:20.98402Z","shell.execute_reply":"2022-08-09T07:54:22.189546Z"},"trusted":true},"execution_count":3,"outputs":[]},{"cell_type":"markdown","source":"#### 1.0 Description:\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nIn this competition, you’ll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.","metadata":{}},{"cell_type":"markdown","source":"### Evaluation Metric:\nThe evaluation metric, , for this competition is the mean of two measures of rank ordering: Normalized Gini Coefficient, , and default rate captured at 4%, .\n\nThe default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n\nFor both of the sub-metrics  and , the negative labels are given a weight of 20 to adjust for downsampling.\n\nThis metric has a maximum value of 1.0.\n\nPython code for calculating this metric can be found in this Notebook.\n\nSubmission File\nFor each customer_ID in the test set, you must predict a probability for the target variable. The file should contain a header and have the following format:","metadata":{}},{"cell_type":"markdown","source":"### 2.0  Reading the train & test dataset .","metadata":{}},{"cell_type":"markdown","source":"### 2.1 Dowloading & reading the train labels dataset.","metadata":{}},{"cell_type":"code","source":"## Since the train data is huge we are considering 100000 rows.\ntrain_df = pd.read_csv('../input/amex-default-prediction/train_data.csv',nrows= 1000000)    ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:54:22.193962Z","iopub.execute_input":"2022-08-09T07:54:22.19518Z","iopub.status.idle":"2022-08-09T07:55:31.474714Z","shell.execute_reply.started":"2022-08-09T07:54:22.195118Z","shell.execute_reply":"2022-08-09T07:55:31.473272Z"},"trusted":true},"execution_count":4,"outputs":[]},{"cell_type":"markdown","source":"### Observations: \n- We have read 100000 rows & 190 cols. \n- The data set has unique customer ID's & S_2 in date format.","metadata":{}},{"cell_type":"markdown","source":"#### Getting Info of train data: ","metadata":{}},{"cell_type":"code","source":"train_df.customer_ID.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:55:31.476437Z","iopub.execute_input":"2022-08-09T07:55:31.476798Z","iopub.status.idle":"2022-08-09T07:55:31.65023Z","shell.execute_reply.started":"2022-08-09T07:55:31.476765Z","shell.execute_reply":"2022-08-09T07:55:31.648878Z"},"trusted":true},"execution_count":5,"outputs":[{"execution_count":5,"output_type":"execute_result","data":{"text/plain":"82975"},"metadata":{}}]},{"cell_type":"markdown","source":"### Observation:\n- The dataset uses a memory of 145MB.\n- Has 185 cols float type \n- Has 4 object type cols: Customer ID, S_2,D_63, D_64.\n","metadata":{}},{"cell_type":"markdown","source":"#### Describing the train data.","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv('../input/amex-default-prediction/test_data.csv', nrows= 1000000)   ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:55:31.652688Z","iopub.execute_input":"2022-08-09T07:55:31.653006Z","iopub.status.idle":"2022-08-09T07:56:41.808021Z","shell.execute_reply.started":"2022-08-09T07:55:31.652977Z","shell.execute_reply":"2022-08-09T07:56:41.806522Z"},"trusted":true},"execution_count":6,"outputs":[]},{"cell_type":"code","source":"train_label_df= pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:41.810565Z","iopub.execute_input":"2022-08-09T07:56:41.811854Z","iopub.status.idle":"2022-08-09T07:56:42.75405Z","shell.execute_reply.started":"2022-08-09T07:56:41.811792Z","shell.execute_reply":"2022-08-09T07:56:42.752582Z"},"trusted":true},"execution_count":7,"outputs":[]},{"cell_type":"code","source":"train_label_df.customer_ID.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:42.755897Z","iopub.execute_input":"2022-08-09T07:56:42.756394Z","iopub.status.idle":"2022-08-09T07:56:42.96004Z","shell.execute_reply.started":"2022-08-09T07:56:42.756345Z","shell.execute_reply":"2022-08-09T07:56:42.958543Z"},"trusted":true},"execution_count":8,"outputs":[{"execution_count":8,"output_type":"execute_result","data":{"text/plain":"458913"},"metadata":{}}]},{"cell_type":"code","source":"#train_df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:42.961751Z","iopub.execute_input":"2022-08-09T07:56:42.962722Z","iopub.status.idle":"2022-08-09T07:56:42.969Z","shell.execute_reply.started":"2022-08-09T07:56:42.962686Z","shell.execute_reply":"2022-08-09T07:56:42.967098Z"},"trusted":true},"execution_count":9,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The dataset \" count\" column shows many missing values. \n- There is skewness in the dataset.\n- Needs to perform standardization/ Normalization of the dataset columns.","metadata":{}},{"cell_type":"markdown","source":"#### Checking Info of Test data.","metadata":{}},{"cell_type":"code","source":"#test_df.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:42.970617Z","iopub.execute_input":"2022-08-09T07:56:42.971183Z","iopub.status.idle":"2022-08-09T07:56:42.981085Z","shell.execute_reply.started":"2022-08-09T07:56:42.971151Z","shell.execute_reply":"2022-08-09T07:56:42.979721Z"},"trusted":true},"execution_count":10,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The test dataset uses a memory of 145MB.\n- Has 185 cols float type \n- Has 4 object type cols: Customer ID, S_2,D_63, D_64.","metadata":{}},{"cell_type":"markdown","source":"#### Checking the describe function on test data.","metadata":{}},{"cell_type":"code","source":"#test_df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:42.983092Z","iopub.execute_input":"2022-08-09T07:56:42.983468Z","iopub.status.idle":"2022-08-09T07:56:42.993074Z","shell.execute_reply.started":"2022-08-09T07:56:42.983434Z","shell.execute_reply":"2022-08-09T07:56:42.992122Z"},"trusted":true},"execution_count":11,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The test dataset \" count\" column shows many missing values. \n- There is skewness in the dataset.\n- Needs to perform standardization/ Normalization of the dataset columns.","metadata":{}},{"cell_type":"markdown","source":"#### Info on train label data.","metadata":{}},{"cell_type":"code","source":"#train_label_df.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:42.998428Z","iopub.execute_input":"2022-08-09T07:56:42.998888Z","iopub.status.idle":"2022-08-09T07:56:43.007873Z","shell.execute_reply.started":"2022-08-09T07:56:42.998852Z","shell.execute_reply":"2022-08-09T07:56:43.006821Z"},"trusted":true},"execution_count":12,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The train_label_df has 100000 enteries & two cols.","metadata":{}},{"cell_type":"markdown","source":"### Describing the train_label_df.","metadata":{}},{"cell_type":"code","source":"#train_label_df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:43.010881Z","iopub.execute_input":"2022-08-09T07:56:43.011608Z","iopub.status.idle":"2022-08-09T07:56:43.022974Z","shell.execute_reply.started":"2022-08-09T07:56:43.01154Z","shell.execute_reply":"2022-08-09T07:56:43.021955Z"},"trusted":true},"execution_count":13,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- No missing values in train_label_df.","metadata":{}},{"cell_type":"markdown","source":"### 2.0  Merging the train datset with train label dataset on \" Customer_ID","metadata":{}},{"cell_type":"code","source":"joined = train_df.merge(train_label_df, how=\"left\", on=[\"customer_ID\"])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:43.024217Z","iopub.execute_input":"2022-08-09T07:56:43.02461Z","iopub.status.idle":"2022-08-09T07:56:46.788733Z","shell.execute_reply.started":"2022-08-09T07:56:43.024567Z","shell.execute_reply":"2022-08-09T07:56:46.787366Z"},"trusted":true},"execution_count":14,"outputs":[]},{"cell_type":"code","source":"joined.customer_ID.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:46.791073Z","iopub.execute_input":"2022-08-09T07:56:46.791593Z","iopub.status.idle":"2022-08-09T07:56:46.959568Z","shell.execute_reply.started":"2022-08-09T07:56:46.791545Z","shell.execute_reply":"2022-08-09T07:56:46.957992Z"},"trusted":true},"execution_count":15,"outputs":[{"execution_count":15,"output_type":"execute_result","data":{"text/plain":"82975"},"metadata":{}}]},{"cell_type":"markdown","source":"### 2.1  Fetching the Data.info() of train dataset.","metadata":{}},{"cell_type":"code","source":"# Finding information about DataFrame\njoined.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:46.961882Z","iopub.execute_input":"2022-08-09T07:56:46.962362Z","iopub.status.idle":"2022-08-09T07:56:46.998643Z","shell.execute_reply.started":"2022-08-09T07:56:46.962316Z","shell.execute_reply":"2022-08-09T07:56:46.997294Z"},"trusted":true},"execution_count":16,"outputs":[{"name":"stdout","text":"<class 'pandas.core.frame.DataFrame'>\nInt64Index: 1000000 entries, 0 to 999999\nData columns (total 191 columns):\n #    Column       Dtype  \n---   ------       -----  \n 0    customer_ID  object \n 1    S_2          object \n 2    P_2          float64\n 3    D_39         float64\n 4    B_1          float64\n 5    B_2          float64\n 6    R_1          float64\n 7    S_3          float64\n 8    D_41         float64\n 9    B_3          float64\n 10   D_42         float64\n 11   D_43         float64\n 12   D_44         float64\n 13   B_4          float64\n 14   D_45         float64\n 15   B_5          float64\n 16   R_2          float64\n 17   D_46         float64\n 18   D_47         float64\n 19   D_48         float64\n 20   D_49         float64\n 21   B_6          float64\n 22   B_7          float64\n 23   B_8          float64\n 24   D_50         float64\n 25   D_51         float64\n 26   B_9          float64\n 27   R_3          float64\n 28   D_52         float64\n 29   P_3          float64\n 30   B_10         float64\n 31   D_53         float64\n 32   S_5          float64\n 33   B_11         float64\n 34   S_6          float64\n 35   D_54         float64\n 36   R_4          float64\n 37   S_7          float64\n 38   B_12         float64\n 39   S_8          float64\n 40   D_55         float64\n 41   D_56         float64\n 42   B_13         float64\n 43   R_5          float64\n 44   D_58         float64\n 45   S_9          float64\n 46   B_14         float64\n 47   D_59         float64\n 48   D_60         float64\n 49   D_61         float64\n 50   B_15         float64\n 51   S_11         float64\n 52   D_62         float64\n 53   D_63         object \n 54   D_64         object \n 55   D_65         float64\n 56   B_16         float64\n 57   B_17         float64\n 58   B_18         float64\n 59   B_19         float64\n 60   D_66         float64\n 61   B_20         float64\n 62   D_68         float64\n 63   S_12         float64\n 64   R_6          float64\n 65   S_13         float64\n 66   B_21         float64\n 67   D_69         float64\n 68   B_22         float64\n 69   D_70         float64\n 70   D_71         float64\n 71   D_72         float64\n 72   S_15         float64\n 73   B_23         float64\n 74   D_73         float64\n 75   P_4          float64\n 76   D_74         float64\n 77   D_75         float64\n 78   D_76         float64\n 79   B_24         float64\n 80   R_7          float64\n 81   D_77         float64\n 82   B_25         float64\n 83   B_26         float64\n 84   D_78         float64\n 85   D_79         float64\n 86   R_8          float64\n 87   R_9          float64\n 88   S_16         float64\n 89   D_80         float64\n 90   R_10         float64\n 91   R_11         float64\n 92   B_27         float64\n 93   D_81         float64\n 94   D_82         float64\n 95   S_17         float64\n 96   R_12         float64\n 97   B_28         float64\n 98   R_13         float64\n 99   D_83         float64\n 100  R_14         float64\n 101  R_15         float64\n 102  D_84         float64\n 103  R_16         float64\n 104  B_29         float64\n 105  B_30         float64\n 106  S_18         float64\n 107  D_86         float64\n 108  D_87         float64\n 109  R_17         float64\n 110  R_18         float64\n 111  D_88         float64\n 112  B_31         int64  \n 113  S_19         float64\n 114  R_19         float64\n 115  B_32         float64\n 116  S_20         float64\n 117  R_20         float64\n 118  R_21         float64\n 119  B_33         float64\n 120  D_89         float64\n 121  R_22         float64\n 122  R_23         float64\n 123  D_91         float64\n 124  D_92         float64\n 125  D_93         float64\n 126  D_94         float64\n 127  R_24         float64\n 128  R_25         float64\n 129  D_96         float64\n 130  S_22         float64\n 131  S_23         float64\n 132  S_24         float64\n 133  S_25         float64\n 134  S_26         float64\n 135  D_102        float64\n 136  D_103        float64\n 137  D_104        float64\n 138  D_105        float64\n 139  D_106        float64\n 140  D_107        float64\n 141  B_36         float64\n 142  B_37         float64\n 143  R_26         float64\n 144  R_27         float64\n 145  B_38         float64\n 146  D_108        float64\n 147  D_109        float64\n 148  D_110        float64\n 149  D_111        float64\n 150  B_39         float64\n 151  D_112        float64\n 152  B_40         float64\n 153  S_27         float64\n 154  D_113        float64\n 155  D_114        float64\n 156  D_115        float64\n 157  D_116        float64\n 158  D_117        float64\n 159  D_118        float64\n 160  D_119        float64\n 161  D_120        float64\n 162  D_121        float64\n 163  D_122        float64\n 164  D_123        float64\n 165  D_124        float64\n 166  D_125        float64\n 167  D_126        float64\n 168  D_127        float64\n 169  D_128        float64\n 170  D_129        float64\n 171  B_41         float64\n 172  B_42         float64\n 173  D_130        float64\n 174  D_131        float64\n 175  D_132        float64\n 176  D_133        float64\n 177  R_28         float64\n 178  D_134        float64\n 179  D_135        float64\n 180  D_136        float64\n 181  D_137        float64\n 182  D_138        float64\n 183  D_139        float64\n 184  D_140        float64\n 185  D_141        float64\n 186  D_142        float64\n 187  D_143        float64\n 188  D_144        float64\n 189  D_145        float64\n 190  target       int64  \ndtypes: float64(185), int64(2), object(4)\nmemory usage: 1.4+ GB\n","output_type":"stream"}]},{"cell_type":"markdown","source":"### Observation:\n- The Merged dataset has 190 cols.\n- 185 Cols are float type, 2 cols int64 type, 4 cols object type.","metadata":{}},{"cell_type":"code","source":"#joined.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:47.000603Z","iopub.execute_input":"2022-08-09T07:56:47.001201Z","iopub.status.idle":"2022-08-09T07:56:47.010364Z","shell.execute_reply.started":"2022-08-09T07:56:47.001141Z","shell.execute_reply":"2022-08-09T07:56:47.007871Z"},"trusted":true},"execution_count":17,"outputs":[]},{"cell_type":"markdown","source":"### Observations: \n- The merged dataset has missing values & the data set is skewed.","metadata":{}},{"cell_type":"code","source":"#test_df.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:47.012468Z","iopub.execute_input":"2022-08-09T07:56:47.014217Z","iopub.status.idle":"2022-08-09T07:56:47.02471Z","shell.execute_reply.started":"2022-08-09T07:56:47.014157Z","shell.execute_reply":"2022-08-09T07:56:47.023388Z"},"trusted":true},"execution_count":18,"outputs":[]},{"cell_type":"markdown","source":"### Observations:\n- The test dataset has 189 cols out of which 185 cols float type, 1col int64 type, 4 object type cols.","metadata":{}},{"cell_type":"code","source":"#test_df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:47.026942Z","iopub.execute_input":"2022-08-09T07:56:47.027456Z","iopub.status.idle":"2022-08-09T07:56:47.036733Z","shell.execute_reply.started":"2022-08-09T07:56:47.027414Z","shell.execute_reply":"2022-08-09T07:56:47.035545Z"},"trusted":true},"execution_count":19,"outputs":[]},{"cell_type":"markdown","source":"### 3.0 Numerical Data Distribution:","metadata":{}},{"cell_type":"code","source":"num_feature = []\n\nfor i in joined.columns.values:\n    if ((joined[i].dtype == int) | (joined[i].dtype == float)):\n        num_feature.append(i)\n    \nprint('Total Numerical Features:', len(num_feature))\nprint('Features:', num_feature)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:47.038178Z","iopub.execute_input":"2022-08-09T07:56:47.039102Z","iopub.status.idle":"2022-08-09T07:56:47.050191Z","shell.execute_reply.started":"2022-08-09T07:56:47.039064Z","shell.execute_reply":"2022-08-09T07:56:47.049252Z"},"trusted":true},"execution_count":20,"outputs":[{"name":"stdout","text":"Total Numerical Features: 187\nFeatures: ['P_2', 'D_39', 'B_1', 'B_2', 'R_1', 'S_3', 'D_41', 'B_3', 'D_42', 'D_43', 'D_44', 'B_4', 'D_45', 'B_5', 'R_2', 'D_46', 'D_47', 'D_48', 'D_49', 'B_6', 'B_7', 'B_8', 'D_50', 'D_51', 'B_9', 'R_3', 'D_52', 'P_3', 'B_10', 'D_53', 'S_5', 'B_11', 'S_6', 'D_54', 'R_4', 'S_7', 'B_12', 'S_8', 'D_55', 'D_56', 'B_13', 'R_5', 'D_58', 'S_9', 'B_14', 'D_59', 'D_60', 'D_61', 'B_15', 'S_11', 'D_62', 'D_65', 'B_16', 'B_17', 'B_18', 'B_19', 'D_66', 'B_20', 'D_68', 'S_12', 'R_6', 'S_13', 'B_21', 'D_69', 'B_22', 'D_70', 'D_71', 'D_72', 'S_15', 'B_23', 'D_73', 'P_4', 'D_74', 'D_75', 'D_76', 'B_24', 'R_7', 'D_77', 'B_25', 'B_26', 'D_78', 'D_79', 'R_8', 'R_9', 'S_16', 'D_80', 'R_10', 'R_11', 'B_27', 'D_81', 'D_82', 'S_17', 'R_12', 'B_28', 'R_13', 'D_83', 'R_14', 'R_15', 'D_84', 'R_16', 'B_29', 'B_30', 'S_18', 'D_86', 'D_87', 'R_17', 'R_18', 'D_88', 'B_31', 'S_19', 'R_19', 'B_32', 'S_20', 'R_20', 'R_21', 'B_33', 'D_89', 'R_22', 'R_23', 'D_91', 'D_92', 'D_93', 'D_94', 'R_24', 'R_25', 'D_96', 'S_22', 'S_23', 'S_24', 'S_25', 'S_26', 'D_102', 'D_103', 'D_104', 'D_105', 'D_106', 'D_107', 'B_36', 'B_37', 'R_26', 'R_27', 'B_38', 'D_108', 'D_109', 'D_110', 'D_111', 'B_39', 'D_112', 'B_40', 'S_27', 'D_113', 'D_114', 'D_115', 'D_116', 'D_117', 'D_118', 'D_119', 'D_120', 'D_121', 'D_122', 'D_123', 'D_124', 'D_125', 'D_126', 'D_127', 'D_128', 'D_129', 'B_41', 'B_42', 'D_130', 'D_131', 'D_132', 'D_133', 'R_28', 'D_134', 'D_135', 'D_136', 'D_137', 'D_138', 'D_139', 'D_140', 'D_141', 'D_142', 'D_143', 'D_144', 'D_145', 'target']\n","output_type":"stream"}]},{"cell_type":"markdown","source":"### Observations: \n- There are total 187 cols numerical .","metadata":{}},{"cell_type":"markdown","source":"## Deleting the coloumns from train data that has correlation >= 0.7","metadata":{}},{"cell_type":"code","source":"## Deleting the coloumns from train data that has correlation >= 0.7\n# Create correlation matrix\ncorr_matrix = joined.corr().abs()\n\n# Select upper triangle of correlation matrix\nupper = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k=1).astype(np.bool_))\n\n# Find features with correlation greater or equal to 0.7\nto_drop = [column for column in upper.columns if any(upper[column] >= 0.7)]\nprint('columns to drop in the train data set',to_drop)\n\n# Drop features \njoined.drop(to_drop, axis=1, inplace=True)\nlen(joined)\njoined.columns\njoined.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:56:47.051612Z","iopub.execute_input":"2022-08-09T07:56:47.052147Z","iopub.status.idle":"2022-08-09T07:58:11.38388Z","shell.execute_reply.started":"2022-08-09T07:56:47.052104Z","shell.execute_reply":"2022-08-09T07:58:11.382791Z"},"trusted":true},"execution_count":21,"outputs":[{"name":"stdout","text":"columns to drop in the train data set ['B_3', 'D_48', 'B_8', 'R_3', 'B_11', 'R_4', 'S_7', 'D_55', 'B_13', 'D_58', 'D_61', 'B_15', 'B_16', 'B_18', 'B_19', 'B_20', 'B_22', 'S_15', 'B_23', 'D_74', 'D_75', 'D_77', 'R_8', 'D_81', 'B_28', 'D_84', 'B_30', 'R_21', 'B_33', 'S_24', 'D_103', 'D_104', 'D_105', 'D_107', 'B_37', 'B_38', 'D_110', 'D_111', 'B_39', 'D_113', 'D_118', 'D_119', 'D_121', 'D_129', 'D_131', 'D_132', 'D_133', 'D_141', 'D_142', 'D_143']\n","output_type":"stream"},{"execution_count":21,"output_type":"execute_result","data":{"text/plain":"(1000000, 141)"},"metadata":{}}]},{"cell_type":"markdown","source":"### 5.0 Finding the missing values of Train data.","metadata":{}},{"cell_type":"code","source":"## Train data missing values\nmissing_values = joined.isna().sum()\npercent_missing = ((missing_values / joined.index.size) * 100)          \n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:11.385422Z","iopub.execute_input":"2022-08-09T07:58:11.385981Z","iopub.status.idle":"2022-08-09T07:58:11.854904Z","shell.execute_reply.started":"2022-08-09T07:58:11.385946Z","shell.execute_reply":"2022-08-09T07:58:11.853587Z"},"trusted":true},"execution_count":22,"outputs":[]},{"cell_type":"code","source":"## Dropping columns from train data that have more than 50% missing values\ncolumns_to_drop = list(percent_missing[percent_missing >= 50].index)\nJoined_1 = joined.drop(columns_to_drop, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:11.856356Z","iopub.execute_input":"2022-08-09T07:58:11.856705Z","iopub.status.idle":"2022-08-09T07:58:12.17112Z","shell.execute_reply.started":"2022-08-09T07:58:11.856673Z","shell.execute_reply":"2022-08-09T07:58:12.17012Z"},"trusted":true},"execution_count":23,"outputs":[]},{"cell_type":"markdown","source":"#### Observation: \n- 117 Columns & 1000000 rows are left after dropping missing values that are greater than or equal to 50.","metadata":{}},{"cell_type":"markdown","source":"## Fetching the Unique values from Train dataset.","metadata":{}},{"cell_type":"code","source":"## Getting Uique values in train dataset\n#for col in Joined_1:\n    #print(col)\n    #print(Joined_1[col].unique())\n    #print('\\n')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:12.172777Z","iopub.execute_input":"2022-08-09T07:58:12.173131Z","iopub.status.idle":"2022-08-09T07:58:12.177833Z","shell.execute_reply.started":"2022-08-09T07:58:12.173098Z","shell.execute_reply":"2022-08-09T07:58:12.176724Z"},"trusted":true},"execution_count":24,"outputs":[]},{"cell_type":"markdown","source":"#### Filling the missing values with mode & median.","metadata":{}},{"cell_type":"code","source":"## To fill the Mode in object type col\n#Joined_1['D_64'] = Joined_1['D_64'].mode()[0]\nJoined_1['D_64'].fillna(Joined_1['D_64'].mode()[0], inplace=True)\nJoined_1['D_64'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:12.180005Z","iopub.execute_input":"2022-08-09T07:58:12.180527Z","iopub.status.idle":"2022-08-09T07:58:12.327967Z","shell.execute_reply.started":"2022-08-09T07:58:12.180483Z","shell.execute_reply":"2022-08-09T07:58:12.326396Z"},"trusted":true},"execution_count":25,"outputs":[{"execution_count":25,"output_type":"execute_result","data":{"text/plain":"0"},"metadata":{}}]},{"cell_type":"code","source":"##Fill NaN values with 0 & then using mean value to impute  in col D_43 of train \nJoined_1['D_43'] = Joined_1['D_43'].fillna(0)\nJoined_1['D_43'] = Joined_1['D_43'].fillna(Joined_1['D_43'].mean())\n\nJoined_1['D_68'] = Joined_1['D_68'].fillna(0)\nJoined_1['D_68'] = Joined_1['D_68'].fillna(Joined_1['D_68'].mean())\n\nJoined_1['D_114'] = Joined_1['D_114'].fillna(0)\nJoined_1['D_114'] = Joined_1['D_114'].fillna(Joined_1['D_114'].mean())\n\nJoined_1['D_120'] = Joined_1['D_120'].fillna(0)\nJoined_1['D_120'] = Joined_1['D_120'].fillna(Joined_1['D_120'].mean())\n\nJoined_1['D_126'] = Joined_1['D_126'].fillna(0)\nJoined_1['D_126'] = Joined_1['D_126'].fillna(Joined_1['D_126'].mean())\n\nJoined_1['D_116'] = Joined_1['D_116'].fillna(0)\nJoined_1['D_116'] = Joined_1['D_116'].fillna(Joined_1['D_116'].mean())\n\nJoined_1['D_117'] = Joined_1['D_117'].fillna(0)\nJoined_1['D_117'] = Joined_1['D_117'].fillna(Joined_1['D_117'].mean())","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:12.33002Z","iopub.execute_input":"2022-08-09T07:58:12.330537Z","iopub.status.idle":"2022-08-09T07:58:12.471289Z","shell.execute_reply.started":"2022-08-09T07:58:12.330495Z","shell.execute_reply":"2022-08-09T07:58:12.470027Z"},"trusted":true},"execution_count":26,"outputs":[]},{"cell_type":"markdown","source":"#### Checking the Uique values in after filling the Nan values in train dataset.","metadata":{}},{"cell_type":"code","source":"#for col in Joined_1:\n#    print(col)\n#    print(Joined_1[col].unique())\n#    print('\\n')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:12.473118Z","iopub.execute_input":"2022-08-09T07:58:12.473607Z","iopub.status.idle":"2022-08-09T07:58:12.480874Z","shell.execute_reply.started":"2022-08-09T07:58:12.473566Z","shell.execute_reply":"2022-08-09T07:58:12.479393Z"},"trusted":true},"execution_count":27,"outputs":[]},{"cell_type":"markdown","source":"### 6.0  Value Counts of target variable & its 'count plot'.","metadata":{}},{"cell_type":"code","source":"Joined_1['target'].value_counts() /len(Joined_1['target']) \nsns.countplot(x='target', data=Joined_1, palette='hls')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:12.482459Z","iopub.execute_input":"2022-08-09T07:58:12.483502Z","iopub.status.idle":"2022-08-09T07:58:12.858677Z","shell.execute_reply.started":"2022-08-09T07:58:12.483455Z","shell.execute_reply":"2022-08-09T07:58:12.856912Z"},"trusted":true},"execution_count":28,"outputs":[{"output_type":"display_data","data":{"text/plain":"<Figure size 432x288 with 1 Axes>","image/png":"iVBORw0KGgoAAAANSUhEUgAAAZcAAAEGCAYAAACpXNjrAAAAOXRFWHRTb2Z0d2FyZQBNYXRwbG90bGliIHZlcnNpb24zLjUuMiwgaHR0cHM6Ly9tYXRwbG90bGliLm9yZy8qNh9FAAAACXBIWXMAAAsTAAALEwEAmpwYAAAVxUlEQVR4nO3df6xf9X3f8ecrOATyg9iA51Gb1Ki1UlFUCNyBm0xVG1RjWFujpkGgZvaYhRtB0kadtpJpkjdYJip1y+IscWcVB7vqQh3aFDeCeJaTLlpXE18Sys8g35Ag2wLs2gYnQSQje++P78fly+V7ry/0fL/XP54P6eh7zvt8zvl8rmTrpXPO53u+qSokSerSm2Z7AJKkk4/hIknqnOEiSeqc4SJJ6pzhIknq3JzZHsDx4txzz63FixfP9jAk6YTy4IMP/l1VzZ9cN1yaxYsXMz4+PtvDkKQTSpKnB9W9LSZJ6pzhIknqnOEiSeqc4SJJ6pzhIknqnOEiSeqc4SJJ6pzhIknqnOEiSeqc39Dv0Phvf3i2h6DjzNi6P5ztIUizwisXSVLnDBdJUucMF0lS5wwXSVLnDBdJUucMF0lS5wwXSVLnDBdJUucMF0lS54YWLkneneShvuVIko8lOTvJ9iS72+e81j5J1iWZSPJwkkv7zrWqtd+dZFVf/bIkj7Rj1iVJqw/sQ5I0GkMLl6p6sqouqapLgMuAF4EvArcCO6pqCbCjbQNcDSxpyxpgPfSCAlgLXAFcDqztC4v1wE19xy1v9an6kCSNwKhui10JfLuqngZWAJtafRNwbVtfAWyunp3A3CTnAVcB26vqUFUdBrYDy9u+s6pqZ1UVsHnSuQb1IUkagVGFy/XA59v6gqp6pq0/Cyxo6wuBPX3H7G216ep7B9Sn6+NVkqxJMp5k/MCBA6/7j5IkDTb0cElyOvBrwBcm72tXHDXM/qfro6o2VNVYVY3Nnz9/mMOQpFPKKK5crga+UVXPte3n2i0t2uf+Vt8HnN933KJWm66+aEB9uj4kSSMwinC5gVduiQFsBY7O+FoF3NtXX9lmjS0FXmi3trYBy5LMaw/ylwHb2r4jSZa2WWIrJ51rUB+SpBEY6o+FJXkb8MvAb/WV7wC2JFkNPA1c1+r3AdcAE/Rmlt0IUFWHktwO7GrtbquqQ239ZuAu4Ezg/rZM14ckaQSGGi5V9QPgnEm1g/Rmj01uW8AtU5xnI7BxQH0cuGhAfWAfkqTR8Bv6kqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTODTVcksxNck+SbyV5IsnPJzk7yfYku9vnvNY2SdYlmUjycJJL+86zqrXfnWRVX/2yJI+0Y9YlSasP7EOSNBrDvnL5FPDlqvoZ4GLgCeBWYEdVLQF2tG2Aq4ElbVkDrIdeUABrgSuAy4G1fWGxHrip77jlrT5VH5KkERhauCR5J/ALwJ0AVfWjqnoeWAFsas02Ade29RXA5urZCcxNch5wFbC9qg5V1WFgO7C87TurqnZWVQGbJ51rUB+SpBEY5pXLBcAB4HNJvpnkj5K8DVhQVc+0Ns8CC9r6QmBP3/F7W226+t4Bdabp41WSrEkynmT8wIEDb+RvlCQNMMxwmQNcCqyvqvcAP2DS7al2xVFDHMO0fVTVhqoaq6qx+fPnD3MYknRKGWa47AX2VtUDbfseemHzXLulRfvc3/bvA87vO35Rq01XXzSgzjR9SJJGYGjhUlXPAnuSvLuVrgQeB7YCR2d8rQLubetbgZVt1thS4IV2a2sbsCzJvPYgfxmwre07kmRpmyW2ctK5BvUhSRqBOUM+/0eBP0lyOvAUcCO9QNuSZDXwNHBda3sfcA0wAbzY2lJVh5LcDuxq7W6rqkNt/WbgLuBM4P62ANwxRR+SpBEYarhU1UPA2IBdVw5oW8AtU5xnI7BxQH0cuGhA/eCgPiRJo+E39CVJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdM1wkSZ0zXCRJnRtquCT5bpJHkjyUZLzVzk6yPcnu9jmv1ZNkXZKJJA8nubTvPKta+91JVvXVL2vnn2jHZro+JEmjMYorl1+qqkuqaqxt3wrsqKolwI62DXA1sKQta4D10AsKYC1wBXA5sLYvLNYDN/Udt/wYfUiSRmA2boutADa19U3AtX31zdWzE5ib5DzgKmB7VR2qqsPAdmB523dWVe2sqgI2TzrXoD4kSSMw7HAp4H8meTDJmlZbUFXPtPVngQVtfSGwp+/Yva02XX3vgPp0fbxKkjVJxpOMHzhw4HX/cZKkweYM+fz/tKr2JflHwPYk3+rfWVWVpIY5gOn6qKoNwAaAsbGxoY5Dkk4lQ71yqap97XM/8EV6z0yea7e0aJ/7W/N9wPl9hy9qtenqiwbUmaYPSdIIDC1ckrwtyTuOrgPLgEeBrcDRGV+rgHvb+lZgZZs1thR4od3a2gYsSzKvPchfBmxr+44kWdpmia2cdK5BfUiSRmCYt8UWAF9ss4PnAP+jqr6cZBewJclq4Gngutb+PuAaYAJ4EbgRoKoOJbkd2NXa3VZVh9r6zcBdwJnA/W0BuGOKPiRJIzC0cKmqp4CLB9QPAlcOqBdwyxTn2ghsHFAfBy6aaR+SpNHwG/qSpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOzShckuyYSU2SJDjGW5GTnAG8FTi3/ZZK2q6zeOUnhSVJepVjvXL/t4CPAT8BPMgr4XIE+G/DG5Yk6UQ2bbhU1aeATyX5aFV9ekRjkiSd4Gb0Y2FV9ekk7wUW9x9TVZuHNC5J0glsRuGS5I+BnwIeAn7cygUYLpKk15jpVOQx4H1VdXNVfbQtvz2TA5OcluSbSb7Uti9I8kCSiSR/muT0Vn9L255o+xf3nePjrf5kkqv66stbbSLJrX31gX1IkkZjpuHyKPCP32AfvwM80bf9+8Anq+qngcPA6lZfDRxu9U+2diS5ELge+FlgOfDZFlinAZ8BrgYuBG5obafrQ5I0AjMNl3OBx5NsS7L16HKsg5IsAv4Z8EdtO8D7gXtak03AtW19Rdum7b+ytV8B3F1VP6yq7wATwOVtmaiqp6rqR8DdwIpj9CFJGoEZPXMB/v0bPP9/Bf4N8I62fQ7wfFW93Lb38sr3ZRYCewCq6uUkL7T2C4GdfefsP2bPpPoVx+jjVZKsAdYAvOtd73r9f50kaaCZzhb7X6/3xEl+BdhfVQ8m+cXXe/woVNUGYAPA2NhYzfJwJOmkMdPZYt+jNzsM4HTgzcAPquqsaQ57H/BrSa4BzqD3rf5PAXOTzGlXFouAfa39PuB8YG+SOcA7gYN99aP6jxlUPzhNH5KkEZjRM5eqekdVndXC5EzgA8Bnj3HMx6tqUVUtpvdA/itV9ZvAV4HfaM1WAfe29a1tm7b/K1VVrX59m012AbAE+DqwC1jSZoad3vrY2o6Zqg9J0gi87rciV89fAFcdq+0Ufg/43SQT9J6P3NnqdwLntPrvAre2/h4DtgCPA18GbqmqH7erko8A2+jNRtvS2k7XhyRpBGZ6W+zX+zbfRO97Ly/NtJOq+ivgr9r6U/Rmek1u8xLwwSmO/wTwiQH1+4D7BtQH9iFJGo2Zzhb71b71l4Hv0psiLEnSa8x0ttiNwx6IJOnkMdMfC1uU5ItJ9rflz9oXJCVJeo2ZPtD/HL1ZWz/Rlr9sNUmSXmOm4TK/qj5XVS+35S5g/hDHJUk6gc00XA4m+dDRF0Ym+RC9LytKkvQaMw2XfwlcBzwLPEPvC4r/YkhjkiSd4GY6Ffk2YFVVHQZIcjbwB/RCR5KkV5nplcvPHQ0WgKo6BLxnOEOSJJ3oZhoub0oy7+hGu3KZ6VWPJOkUM9OA+M/A3yT5Qtv+IANexyJJEsz8G/qbk4zT+4VHgF+vqseHNyxJ0olsxre2WpgYKJKkY3rdr9yXJOlYDBdJUucMF0lS5wwXSVLnDBdJUucMF0lS54YWLknOSPL1JH+b5LEk/6HVL0jyQJKJJH+a5PRWf0vbnmj7F/ed6+Ot/mSSq/rqy1ttIsmtffWBfUiSRmOYVy4/BN5fVRcDlwDLkywFfh/4ZFX9NHAYWN3arwYOt/onWzuSXAhcD/wssBz47NFX/wOfAa4GLgRuaG2Zpg9J0ggMLVyq5/tt881tKXrf8r+n1TcB17b1FW2btv/KJGn1u6vqh1X1HWACuLwtE1X1VFX9CLgbWNGOmaoPSdIIDPWZS7vCeAjYD2wHvg08X1UvtyZ7gYVtfSGwB6DtfwE4p78+6Zip6udM08fk8a1JMp5k/MCBA/+Av1SS1G+o4VJVP66qS4BF9K40fmaY/b1eVbWhqsaqamz+fH+1WZK6MpLZYlX1PPBV4OeBuUmOvtNsEbCvre8Dzgdo+99J76eU/74+6Zip6gen6UOSNALDnC02P8nctn4m8MvAE/RC5jdas1XAvW19a9um7f9KVVWrX99mk10ALAG+DuwClrSZYafTe+i/tR0zVR+SpBEY5g9+nQdsarO63gRsqaovJXkcuDvJfwS+CdzZ2t8J/HGSCeAQvbCgqh5LsoXeG5lfBm6pqh8DJPkIsA04DdhYVY+1c/3eFH1IkkZgaOFSVQ8z4KeQq+opes9fJtdfovcjZIPO9QkG/DhZVd0H3DfTPiRJo+E39CVJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdG+Y39CUdJz78f8Znewg6Dv3he8eGdm6vXCRJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdM1wkSZ0zXCRJnTNcJEmdM1wkSZ0bWrgkOT/JV5M8nuSxJL/T6mcn2Z5kd/uc1+pJsi7JRJKHk1zad65Vrf3uJKv66pcleaQdsy5JputDkjQaw7xyeRn4V1V1IbAUuCXJhcCtwI6qWgLsaNsAVwNL2rIGWA+9oADWAlcAlwNr+8JiPXBT33HLW32qPiRJIzC0cKmqZ6rqG239e8ATwEJgBbCpNdsEXNvWVwCbq2cnMDfJecBVwPaqOlRVh4HtwPK276yq2llVBWyedK5BfUiSRmAkz1ySLAbeAzwALKiqZ9quZ4EFbX0hsKfvsL2tNl1974A60/QhSRqBoYdLkrcDfwZ8rKqO9O9rVxw1zP6n6yPJmiTjScYPHDgwzGFI0illqOGS5M30guVPqurPW/m5dkuL9rm/1fcB5/cdvqjVpqsvGlCfro9XqaoNVTVWVWPz589/Y3+kJOk1hjlbLMCdwBNV9V/6dm0Fjs74WgXc21df2WaNLQVeaLe2tgHLksxrD/KXAdvaviNJlra+Vk4616A+JEkjMMwfC3sf8M+BR5I81Gr/FrgD2JJkNfA0cF3bdx9wDTABvAjcCFBVh5LcDuxq7W6rqkNt/WbgLuBM4P62ME0fkqQRGFq4VNX/BjLF7isHtC/glinOtRHYOKA+Dlw0oH5wUB+SpNHwG/qSpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4ZLpKkzhkukqTOGS6SpM4NLVySbEyyP8mjfbWzk2xPsrt9zmv1JFmXZCLJw0ku7TtmVWu/O8mqvvplSR5px6xLkun6kCSNzjCvXO4Clk+q3QrsqKolwI62DXA1sKQta4D10AsKYC1wBXA5sLYvLNYDN/Udt/wYfUiSRmRo4VJVXwMOTSqvADa19U3AtX31zdWzE5ib5DzgKmB7VR2qqsPAdmB523dWVe2sqgI2TzrXoD4kSSMy6mcuC6rqmbb+LLCgrS8E9vS129tq09X3DqhP18drJFmTZDzJ+IEDB97AnyNJGmTWHui3K46azT6qakNVjVXV2Pz584c5FEk6pYw6XJ5rt7Ron/tbfR9wfl+7Ra02XX3RgPp0fUiSRmTU4bIVODrjaxVwb199ZZs1thR4od3a2gYsSzKvPchfBmxr+44kWdpmia2cdK5BfUiSRmTOsE6c5PPALwLnJtlLb9bXHcCWJKuBp4HrWvP7gGuACeBF4EaAqjqU5HZgV2t3W1UdnSRwM70ZaWcC97eFafqQJI3I0MKlqm6YYteVA9oWcMsU59kIbBxQHwcuGlA/OKgPSdLo+A19SVLnDBdJUucMF0lS5wwXSVLnDBdJUucMF0lS5wwXSVLnDBdJUucMF0lS5wwXSVLnDBdJUucMF0lS5wwXSVLnDBdJUucMF0lS5wwXSVLnDBdJUucMF0lS5wwXSVLnTtpwSbI8yZNJJpLcOtvjkaRTyUkZLklOAz4DXA1cCNyQ5MLZHZUknTpOynABLgcmquqpqvoRcDewYpbHJEmnjDmzPYAhWQjs6dveC1wxuVGSNcCatvn9JE+OYGyninOBv5vtQcy6T//32R6BXst/m01H/zp/clDxZA2XGamqDcCG2R7HySjJeFWNzfY4pMn8tzkaJ+ttsX3A+X3bi1pNkjQCJ2u47AKWJLkgyenA9cDWWR6TJJ0yTsrbYlX1cpKPANuA04CNVfXYLA/rVOPtRh2v/Lc5Aqmq2R6DJOkkc7LeFpMkzSLDRZLUOcNFnfK1OzpeJdmYZH+SR2d7LKcCw0Wd8bU7Os7dBSyf7UGcKgwXdcnX7ui4VVVfAw7N9jhOFYaLujTotTsLZ2kskmaR4SJJ6pzhoi752h1JgOGibvnaHUmA4aIOVdXLwNHX7jwBbPG1OzpeJPk88DfAu5PsTbJ6tsd0MvP1L5KkznnlIknqnOEiSeqc4SJJ6pzhIknqnOEiSeqc4SKNQJK5SW4eQT/X+rJQHQ8MF2k05gIzDpf0vJH/n9fSeyO1NKv8nos0AkmOviH6SeCrwM8B84A3A/+uqu5NspjeF1AfAC4DrgFWAh8CDtB7KeiDVfUHSX6K3s8bzAdeBG4Czga+BLzQlg9U1bdH9TdK/ebM9gCkU8StwEVVdUmSOcBbq+pIknOBnUmOviZnCbCqqnYm+SfAB4CL6YXQN4AHW7sNwIeraneSK4DPVtX723m+VFX3jPKPkyYzXKTRC/CfkvwC8P/o/SzBgrbv6ara2dbfB9xbVS8BLyX5S4AkbwfeC3whydFzvmVUg5dmwnCRRu836d3Ouqyq/m+S7wJntH0/mMHxbwKer6pLhjM86R/OB/rSaHwPeEdbfyewvwXLLwE/OcUxfw38apIz2tXKrwBU1RHgO0k+CH//8P/iAf1Is8ZwkUagqg4Cf53kUeASYCzJI/Qe2H9rimN20fvJgoeB+4FH6D2oh97Vz+okfws8xis/J3038K+TfLM99JdmhbPFpONYkrdX1feTvBX4GrCmqr4x2+OSjsVnLtLxbUP7UuQZwCaDRScKr1wkSZ3zmYskqXOGiySpc4aLJKlzhoskqXOGiySpc/8fxG0hsd4NIBgAAAAASUVORK5CYII=\n"},"metadata":{"needs_background":"light"}}]},{"cell_type":"markdown","source":"### Observations: Target variable is unbalanced. Hence need to perform  sampling techniques.","metadata":{}},{"cell_type":"markdown","source":"#### 7.0 Dropping the customer id & S2 Column from the train & test data set.","metadata":{}},{"cell_type":"code","source":"\nJoined_1.drop(['customer_ID','S_2'],axis=1, inplace= True)\nprint(Joined_1.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:12.860624Z","iopub.execute_input":"2022-08-09T07:58:12.861041Z","iopub.status.idle":"2022-08-09T07:58:13.387722Z","shell.execute_reply.started":"2022-08-09T07:58:12.861007Z","shell.execute_reply":"2022-08-09T07:58:13.386136Z"},"trusted":true},"execution_count":29,"outputs":[{"name":"stdout","text":"(1000000, 115)\n","output_type":"stream"}]},{"cell_type":"code","source":"Joined_1 = Joined_1.replace([np.inf, -np.inf], np.nan)\nJoined_1 = Joined_1.dropna()\nJoined_1 = Joined_1.reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:13.396305Z","iopub.execute_input":"2022-08-09T07:58:13.396775Z","iopub.status.idle":"2022-08-09T07:58:15.959329Z","shell.execute_reply.started":"2022-08-09T07:58:13.396735Z","shell.execute_reply":"2022-08-09T07:58:15.957998Z"},"trusted":true},"execution_count":30,"outputs":[]},{"cell_type":"markdown","source":"#### 8.0 Performing One Hot encoding on Train data.","metadata":{}},{"cell_type":"code","source":"#from sklearn.preprocessing import OneHotEncoder\n\none_hot_encoded_train = pd.get_dummies(Joined_1, columns = ['D_63', 'D_64'], drop_first='True')\n\n### For Train data.\nX = one_hot_encoded_train.drop('target',axis = 1)\ny = one_hot_encoded_train['target']\n#print(X)\n\nX_copy= X.copy()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:15.960855Z","iopub.execute_input":"2022-08-09T07:58:15.961778Z","iopub.status.idle":"2022-08-09T07:58:17.101805Z","shell.execute_reply.started":"2022-08-09T07:58:15.961741Z","shell.execute_reply":"2022-08-09T07:58:17.100384Z"},"trusted":true},"execution_count":31,"outputs":[]},{"cell_type":"markdown","source":"#### 9.0 Feature Selection Using Random Forest.","metadata":{}},{"cell_type":"code","source":"# create the classifier with n_estimators = 100\nfrom sklearn.ensemble import RandomForestClassifier\n\nclf = RandomForestClassifier(n_estimators=10, random_state=0,verbose=True)\n\n# fit the model to the training set\n\nclf.fit(X, y)\n# view the feature scores\n\nfeature_scores = pd.Series(clf.feature_importances_, index=X.columns).sort_values(ascending=False)\n\nfeature_scores.head(15)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T07:58:17.103907Z","iopub.execute_input":"2022-08-09T07:58:17.104437Z","iopub.status.idle":"2022-08-09T08:00:53.555606Z","shell.execute_reply.started":"2022-08-09T07:58:17.104389Z","shell.execute_reply":"2022-08-09T08:00:53.554348Z"},"trusted":true},"execution_count":32,"outputs":[{"name":"stderr","text":"[Parallel(n_jobs=1)]: Using backend SequentialBackend with 1 concurrent workers.\n[Parallel(n_jobs=1)]: Done  10 out of  10 | elapsed:  2.6min finished\n","output_type":"stream"},{"execution_count":32,"output_type":"execute_result","data":{"text/plain":"P_2     0.13\nD_44    0.08\nB_10    0.06\nB_9     0.04\nB_7     0.03\nB_2     0.03\nR_27    0.02\nD_45    0.02\nB_1     0.02\nD_43    0.02\nD_52    0.02\nS_3     0.01\nB_4     0.01\nR_1     0.01\nD_62    0.01\ndtype: float64"},"metadata":{}}]},{"cell_type":"markdown","source":"#### 10.0 Filtering out the required column  from train based upon feature importance.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = X.filter(['P_2','B_9','D_44','B_2','B_1','B_7','B_6','D_45','B_10','D_52','S_3','D_62','B_4','R_27','D_43'])\nX.columns\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:00:53.55704Z","iopub.execute_input":"2022-08-09T08:00:53.557443Z","iopub.status.idle":"2022-08-09T08:00:53.584161Z","shell.execute_reply.started":"2022-08-09T08:00:53.55741Z","shell.execute_reply":"2022-08-09T08:00:53.583059Z"},"trusted":true},"execution_count":33,"outputs":[{"execution_count":33,"output_type":"execute_result","data":{"text/plain":"(540944, 15)"},"metadata":{}}]},{"cell_type":"markdown","source":"#### 11.0 Split the train data into 70:30 percentage.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split \nX_train, X_test_1, y_train, y_test_1 = train_test_split(X, y, test_size = 0.3, random_state = 0,stratify= y)\n\n# describes info about train and test set\nprint(\"Number transactions X_train dataset: \", X_train.shape)\nprint(\"Number transactions y_train dataset: \", y_train.shape)\nprint(\"Number transactions X_test_1 dataset: \", X_test_1.shape)\nprint(\"Number transactions y_test_1 dataset: \", y_test_1.shape)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:00:53.585818Z","iopub.execute_input":"2022-08-09T08:00:53.586141Z","iopub.status.idle":"2022-08-09T08:00:53.977864Z","shell.execute_reply.started":"2022-08-09T08:00:53.586111Z","shell.execute_reply":"2022-08-09T08:00:53.976099Z"},"trusted":true},"execution_count":34,"outputs":[{"name":"stdout","text":"Number transactions X_train dataset:  (378660, 15)\nNumber transactions y_train dataset:  (378660,)\nNumber transactions X_test_1 dataset:  (162284, 15)\nNumber transactions y_test_1 dataset:  (162284,)\n","output_type":"stream"}]},{"cell_type":"markdown","source":"#### 12.0  Checking for the value_counts of y_train dataset.","metadata":{}},{"cell_type":"code","source":"#y_train.value_counts()/len(y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:00:53.979905Z","iopub.execute_input":"2022-08-09T08:00:53.980437Z","iopub.status.idle":"2022-08-09T08:00:53.986098Z","shell.execute_reply.started":"2022-08-09T08:00:53.980386Z","shell.execute_reply":"2022-08-09T08:00:53.985011Z"},"trusted":true},"execution_count":35,"outputs":[]},{"cell_type":"markdown","source":"### 13.0 Scaling the data with Robust scaler as data has outliers.\n### perform a robust scaler transform of the dataset.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import RobustScaler\nfrom pandas import DataFrame\n\ntrans = RobustScaler(with_centering=False, with_scaling=True)\n\n# convert the array back to a dataframe\nX_train = trans.fit_transform(X_train)\nX_test_1 = trans.transform(X_test_1)\n\nX_train_df= DataFrame(X_train)\nX_train_df.columns = X.columns\n#print(X_train_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:00:53.988012Z","iopub.execute_input":"2022-08-09T08:00:53.988761Z","iopub.status.idle":"2022-08-09T08:00:54.217744Z","shell.execute_reply.started":"2022-08-09T08:00:53.98872Z","shell.execute_reply":"2022-08-09T08:00:54.216113Z"},"trusted":true},"execution_count":36,"outputs":[]},{"cell_type":"markdown","source":"#### 14.0 Using Smote() Techniques.","metadata":{}},{"cell_type":"markdown","source":"from imblearn.over_sampling import SMOTE\nfrom collections import Counter\ncounter = Counter(y_train)\n#print('Before',counter)\n\n# oversampling the train dataset using SMOTE\nsmt = SMOTE()\nX_train_sm, y_train_sm = smt.fit_resample(X_train, y_train)\n\ncounter = Counter(y_train_sm)\n#print('After',counter)","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:34:03.704494Z","iopub.status.busy":"2022-08-08T14:34:03.704076Z","iopub.status.idle":"2022-08-08T14:35:02.945958Z","shell.execute_reply":"2022-08-08T14:35:02.944923Z","shell.execute_reply.started":"2022-08-08T14:34:03.704454Z"}}},{"cell_type":"markdown","source":"#### 15.0 Model Building - Imbalanced data","metadata":{}},{"cell_type":"code","source":"model = list()\nresample = list()\nprecision = list()\nrecall = list()\nF1score = list()\nAUCROC = list()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:00:54.219297Z","iopub.execute_input":"2022-08-09T08:00:54.219733Z","iopub.status.idle":"2022-08-09T08:00:54.227542Z","shell.execute_reply.started":"2022-08-09T08:00:54.219697Z","shell.execute_reply":"2022-08-09T08:00:54.225908Z"},"trusted":true},"execution_count":37,"outputs":[]},{"cell_type":"code","source":"def test_eval(clf_model, X_test_1, y_test_1, algo=None, sampling=None):\n    \n    # Test set prediction\n    y_prob=clf_model.predict_proba(X_test_1)\n    y_pred=clf_model.predict(X_test_1)\n    \n    print('Confusion Matrix')\n    print('='*60)\n    print(confusion_matrix(y_test_1,y_pred),\"\\n\")\n    print('Classification Report')\n    print('='*60)\n    print(classification_report(y_test_1,y_pred),\"\\n\")\n    print('AUC-ROC')\n    print('='*60)\n    print(roc_auc_score(y_test_1, y_prob[:,1]))\n          \n    model.append(algo)\n    precision.append(precision_score(y_test_1,y_pred))\n    recall.append(recall_score(y_test_1,y_pred))\n    F1score.append(f1_score(y_test_1,y_pred))\n    AUCROC.append(roc_auc_score(y_test_1, y_prob[:,1]))\n    resample.append(sampling)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:00:54.229646Z","iopub.execute_input":"2022-08-09T08:00:54.230765Z","iopub.status.idle":"2022-08-09T08:00:54.240869Z","shell.execute_reply.started":"2022-08-09T08:00:54.230727Z","shell.execute_reply":"2022-08-09T08:00:54.239337Z"},"trusted":true},"execution_count":38,"outputs":[]},{"cell_type":"markdown","source":"### 15.1 Applying the logistic regression model with cross validation & using random search.","metadata":{}},{"cell_type":"markdown","source":"# define model\nlog_model = LogisticRegression()\n\n# define evaluation\ncv = RepeatedStratifiedKFold(n_splits=4, n_repeats=2, random_state=1)\n\n# define search space\nspace = dict()\nspace['solver'] = ['newton-cg']\nspace['penalty'] = [ 'l2']\nspace['C'] = loguniform(1e-5, 100)\n\n# define search\nclf_LR = RandomizedSearchCV(log_model, space, n_iter=2, scoring='roc_auc', n_jobs=-1, cv=cv, random_state=1)\n\n# execute search\nresult = clf_LR.fit(X_train, y_train)\n\n# summarize result\nprint('Best Score: %s' % result.best_score_)\nprint('Best Hyperparameters: %s' % result.best_params_)","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:35:02.97009Z","iopub.status.busy":"2022-08-08T14:35:02.969459Z","iopub.status.idle":"2022-08-08T14:36:27.261949Z","shell.execute_reply":"2022-08-08T14:36:27.258484Z","shell.execute_reply.started":"2022-08-08T14:35:02.970054Z"}}},{"cell_type":"markdown","source":"test_eval(clf_LR, X_test_1, y_test_1, 'Logistic Regression', 'actual')","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:36:27.264607Z","iopub.status.busy":"2022-08-08T14:36:27.264202Z","iopub.status.idle":"2022-08-08T14:36:27.941823Z","shell.execute_reply":"2022-08-08T14:36:27.940571Z","shell.execute_reply.started":"2022-08-08T14:36:27.264564Z"}}},{"cell_type":"markdown","source":"#### Fitting the logistic Regression model on smote applied on train data.","metadata":{}},{"cell_type":"markdown","source":"## SMOTE Resampling\nclf_LR.fit(X_train_sm, y_train_sm)\nclf_LR.best_estimator_","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:36:27.943937Z","iopub.status.busy":"2022-08-08T14:36:27.943452Z","iopub.status.idle":"2022-08-08T14:38:37.872053Z","shell.execute_reply":"2022-08-08T14:38:37.870678Z","shell.execute_reply.started":"2022-08-08T14:36:27.943876Z"}}},{"cell_type":"markdown","source":"test_eval(clf_LR, X_test_1, y_test_1, 'Logistic Regression', 'smote')","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:38:37.873833Z","iopub.status.busy":"2022-08-08T14:38:37.873319Z","iopub.status.idle":"2022-08-08T14:38:38.767019Z","shell.execute_reply":"2022-08-08T14:38:38.7658Z","shell.execute_reply.started":"2022-08-08T14:38:37.873763Z"}}},{"cell_type":"markdown","source":"#predicting on test data\ny_pred_test_LR = clf_LR.predict(X_test_1)\ny_pred_test_LR","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:38:38.769557Z","iopub.status.busy":"2022-08-08T14:38:38.768782Z","iopub.status.idle":"2022-08-08T14:38:38.792039Z","shell.execute_reply":"2022-08-08T14:38:38.790097Z","shell.execute_reply.started":"2022-08-08T14:38:38.769495Z"}}},{"cell_type":"markdown","source":"#### 15.2  Applying Random Forest algorith with cross validation & using random search.","metadata":{}},{"cell_type":"code","source":"estimators = [30]\n# Maximum number of depth in each tree:\nmax_depth = [i for i in range(5,16,2)]\n# Minimum number of samples to consider to split a node:\nmin_samples_split = [10]        \n# Minimum number of samples to consider at each leaf node:\nmin_samples_leaf = [1, 2, 5]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:00:54.242493Z","iopub.execute_input":"2022-08-09T08:00:54.242858Z","iopub.status.idle":"2022-08-09T08:00:54.253163Z","shell.execute_reply.started":"2022-08-09T08:00:54.242824Z","shell.execute_reply":"2022-08-09T08:00:54.251747Z"},"trusted":true},"execution_count":39,"outputs":[]},{"cell_type":"code","source":"rf_model = RandomForestClassifier() \nfrom sklearn.model_selection import RandomizedSearchCV\nrf_params={'n_estimators':estimators,\n           'max_depth':max_depth,\n           'min_samples_split':min_samples_split}\n\nclf_RF = RandomizedSearchCV(rf_model, rf_params, cv=5, scoring='roc_auc', n_jobs=-1, n_iter=3, verbose=2)\nclf_RF.fit(X_train, y_train)\nclf_RF.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:00:54.254984Z","iopub.execute_input":"2022-08-09T08:00:54.256332Z","iopub.status.idle":"2022-08-09T08:05:53.599112Z","shell.execute_reply.started":"2022-08-09T08:00:54.256255Z","shell.execute_reply":"2022-08-09T08:05:53.597396Z"},"trusted":true},"execution_count":40,"outputs":[{"name":"stdout","text":"Fitting 5 folds for each of 3 candidates, totalling 15 fits\n","output_type":"stream"},{"execution_count":40,"output_type":"execute_result","data":{"text/plain":"RandomForestClassifier(max_depth=13, min_samples_split=10, n_estimators=30)"},"metadata":{}}]},{"cell_type":"code","source":"test_eval(clf_RF, X_test_1, y_test_1, 'Random Forest', 'actual')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:05:53.601252Z","iopub.execute_input":"2022-08-09T08:05:53.602653Z","iopub.status.idle":"2022-08-09T08:05:56.155736Z","shell.execute_reply.started":"2022-08-09T08:05:53.602591Z","shell.execute_reply":"2022-08-09T08:05:56.154409Z"},"trusted":true},"execution_count":41,"outputs":[{"name":"stdout","text":"Confusion Matrix\n============================================================\n[[100830  12015]\n [ 11266  38173]] \n\nClassification Report\n============================================================\n              precision    recall  f1-score   support\n\n           0       0.90      0.89      0.90    112845\n           1       0.76      0.77      0.77     49439\n\n    accuracy                           0.86    162284\n   macro avg       0.83      0.83      0.83    162284\nweighted avg       0.86      0.86      0.86    162284\n \n\nAUC-ROC\n============================================================\n0.9281534890414579\n","output_type":"stream"}]},{"cell_type":"markdown","source":"#### Fitting the Random Forest model on smote applied on train data.","metadata":{}},{"cell_type":"markdown","source":"## 2.SMOTE Resampling\nclf_RF.fit(X_train_sm, y_train_sm)\nclf_RF.best_estimator_\n#y_pred_train= clf_model.predict(X_train)\n","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:47:11.825638Z","iopub.status.busy":"2022-08-08T14:47:11.824452Z","iopub.status.idle":"2022-08-08T14:58:28.184902Z","shell.execute_reply":"2022-08-08T14:58:28.183916Z","shell.execute_reply.started":"2022-08-08T14:47:11.825595Z"}}},{"cell_type":"markdown","source":"test_eval(clf_RF, X_test_1, y_test_1, 'Random Forest', 'smote')","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:58:28.186551Z","iopub.status.busy":"2022-08-08T14:58:28.186192Z","iopub.status.idle":"2022-08-08T14:58:30.027871Z","shell.execute_reply":"2022-08-08T14:58:30.02693Z","shell.execute_reply.started":"2022-08-08T14:58:28.186515Z"}}},{"cell_type":"code","source":"#predicting on test data\ny_pred_test_RF = clf_RF.predict(X_test_1,)\ny_pred_test_RF","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:05:56.157516Z","iopub.execute_input":"2022-08-09T08:05:56.159265Z","iopub.status.idle":"2022-08-09T08:05:57.116501Z","shell.execute_reply.started":"2022-08-09T08:05:56.159221Z","shell.execute_reply":"2022-08-09T08:05:57.114794Z"},"trusted":true},"execution_count":42,"outputs":[{"execution_count":42,"output_type":"execute_result","data":{"text/plain":"array([0, 1, 0, ..., 0, 0, 1])"},"metadata":{}}]},{"cell_type":"markdown","source":"### Model Comparision.","metadata":{}},{"cell_type":"markdown","source":"clf_eval_df = pd.DataFrame({'model':model,\n                            'resample':resample,\n                            'precision':precision,\n                            'recall':recall,\n                            'f1-score':F1score,\n                            'AUC-ROC':AUCROC})","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:58:30.664447Z","iopub.status.busy":"2022-08-08T14:58:30.664063Z","iopub.status.idle":"2022-08-08T14:58:30.671978Z","shell.execute_reply":"2022-08-08T14:58:30.670439Z","shell.execute_reply.started":"2022-08-08T14:58:30.66441Z"}}},{"cell_type":"markdown","source":"clf_eval_df","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:58:30.674226Z","iopub.status.busy":"2022-08-08T14:58:30.673824Z","iopub.status.idle":"2022-08-08T14:58:30.694412Z","shell.execute_reply":"2022-08-08T14:58:30.693501Z","shell.execute_reply.started":"2022-08-08T14:58:30.674191Z"}}},{"cell_type":"markdown","source":"sns.set(font_scale=1.2)\n#sns.palplot(sns.color_palette())\ng = sns.FacetGrid(clf_eval_df, col=\"model\", height=5)\ng.map(sns.barplot, \"resample\", \"recall\", palette='twilight', order=[\"actual\", \"smote\", \"adasyn\", \"smote+tomek\", \"smote+enn\"])\ng.set_xticklabels(rotation=30)\ng.set_xlabels(' ', fontsize=14)\ng.set_ylabels('Recall', fontsize=14)","metadata":{"execution":{"iopub.execute_input":"2022-08-08T14:58:30.696758Z","iopub.status.busy":"2022-08-08T14:58:30.695784Z","iopub.status.idle":"2022-08-08T14:58:31.30661Z","shell.execute_reply":"2022-08-08T14:58:31.305701Z","shell.execute_reply.started":"2022-08-08T14:58:30.696724Z"}}},{"cell_type":"code","source":"##predictions for actual Test data.\n\ndef amex_metric_mod(y_true, y_pred):\n\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:05:57.118785Z","iopub.execute_input":"2022-08-09T08:05:57.119171Z","iopub.status.idle":"2022-08-09T08:05:57.133061Z","shell.execute_reply.started":"2022-08-09T08:05:57.119138Z","shell.execute_reply":"2022-08-09T08:05:57.131372Z"},"trusted":true},"execution_count":43,"outputs":[]},{"cell_type":"code","source":"   print(amex_metric_mod(y_test_1, y_pred_test_RF )) ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:05:57.135818Z","iopub.execute_input":"2022-08-09T08:05:57.136222Z","iopub.status.idle":"2022-08-09T08:05:57.185511Z","shell.execute_reply.started":"2022-08-09T08:05:57.136187Z","shell.execute_reply":"2022-08-09T08:05:57.183802Z"},"trusted":true},"execution_count":44,"outputs":[{"name":"stdout","text":"0.4654957284808312\n","output_type":"stream"}]},{"cell_type":"markdown","source":"### 5.3 Processing of the test data.","metadata":{}},{"cell_type":"code","source":"def create_model(test_df,submission):\n    \n    ## Missing values\n    test_df.drop('customer_ID',axis=1,inplace=True)\n    test_df.isna().sum()/len(test_df)*100\n         \n    ## Filling the missing values\n    test_df .fillna(method='ffill',inplace=True)\n    test_df .fillna(method='bfill',inplace=True)\n    \n    ## Use of Robust scaler to scale the data\n    test_df_final = trans.transform(test_df)\n\n    #Test_df= DataFrame(test_df_final)\n    #Test_df.columns = test_df.columns\n        \n    ## Predicting the target variable\n    submission['predicted']=clf_RF.predict(test_df_final)\n            \n    # Merge the prediction and customer_ID into submission dataframe\n    #submission = pd.DataFrame({\"customer_ID\":test_df.customer_ID,\"prediction\":Test_df['predicted']})\n\n    submission.to_csv('submission.csv',mode='a', index=False)\n            \n        ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:05:57.187797Z","iopub.execute_input":"2022-08-09T08:05:57.188185Z","iopub.status.idle":"2022-08-09T08:05:57.198635Z","shell.execute_reply.started":"2022-08-09T08:05:57.18815Z","shell.execute_reply":"2022-08-09T08:05:57.197262Z"},"trusted":true},"execution_count":45,"outputs":[]},{"cell_type":"code","source":"final_col=['customer_ID','P_2','B_9','D_44','B_2','B_1','B_7','B_6','D_45','B_10','D_52','S_3','D_62','B_4','R_27','D_43']","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:05:57.200791Z","iopub.execute_input":"2022-08-09T08:05:57.201226Z","iopub.status.idle":"2022-08-09T08:05:57.211321Z","shell.execute_reply.started":"2022-08-09T08:05:57.20117Z","shell.execute_reply":"2022-08-09T08:05:57.210029Z"},"trusted":true},"execution_count":46,"outputs":[]},{"cell_type":"code","source":"test_df_data = pd.read_csv('../input/amex-default-prediction/test_data.csv', chunksize=500000, iterator=True, usecols =final_col)\ntest_df_data","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:05:57.212924Z","iopub.execute_input":"2022-08-09T08:05:57.213359Z","iopub.status.idle":"2022-08-09T08:05:57.239023Z","shell.execute_reply.started":"2022-08-09T08:05:57.213273Z","shell.execute_reply":"2022-08-09T08:05:57.237785Z"},"trusted":true},"execution_count":47,"outputs":[{"execution_count":47,"output_type":"execute_result","data":{"text/plain":"<pandas.io.parsers.readers.TextFileReader at 0x7fe24ef54c50>"},"metadata":{}}]},{"cell_type":"code","source":"for iter_num, chunk in enumerate(test_df_data, 1):\n    print(iter_num)\n    subm_t=pd.DataFrame(columns=['customer_ID','predicted'])\n    subm_t[\"customer_ID\"]=  chunk[\"customer_ID\"]\n   \n    #print(chunk.info())\n    #print(subm_t.info())\n    create_model(chunk,subm_t)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:05:57.240943Z","iopub.execute_input":"2022-08-09T08:05:57.242852Z","iopub.status.idle":"2022-08-09T08:15:37.852845Z","shell.execute_reply.started":"2022-08-09T08:05:57.242805Z","shell.execute_reply":"2022-08-09T08:15:37.850801Z"},"trusted":true},"execution_count":48,"outputs":[{"name":"stdout","text":"1\n2\n3\n4\n5\n6\n7\n[CV] END .max_depth=5, min_samples_split=10, n_estimators=30; total time=  37.9s\n[CV] END .max_depth=7, min_samples_split=10, n_estimators=30; total time=  49.5s\n[CV] END max_depth=13, min_samples_split=10, n_estimators=30; total time= 1.2min\n8\n[CV] END .max_depth=5, min_samples_split=10, n_estimators=30; total time=  38.3s\n[CV] END .max_depth=7, min_samples_split=10, n_estimators=30; total time=  45.9s\n[CV] END .max_depth=7, min_samples_split=10, n_estimators=30; total time=  45.0s\n[CV] END max_depth=13, min_samples_split=10, n_estimators=30; total time= 1.0min\n[CV] END .max_depth=5, min_samples_split=10, n_estimators=30; total time=  37.5s\n[CV] END .max_depth=5, min_samples_split=10, n_estimators=30; total time=  37.6s\n[CV] END .max_depth=7, min_samples_split=10, n_estimators=30; total time=  51.5s\n[CV] END max_depth=13, min_samples_split=10, n_estimators=30; total time= 1.2min\n9\n[CV] END .max_depth=5, min_samples_split=10, n_estimators=30; total time=  39.5s\n[CV] END .max_depth=7, min_samples_split=10, n_estimators=30; total time=  46.2s\n[CV] END max_depth=13, min_samples_split=10, n_estimators=30; total time= 1.2min\n[CV] END max_depth=13, min_samples_split=10, n_estimators=30; total time= 1.0min\n10\n11\n12\n13\n14\n15\n16\n17\n18\n19\n20\n21\n22\n23\n","output_type":"stream"}]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:37.85774Z","iopub.execute_input":"2022-08-09T08:15:37.859422Z","iopub.status.idle":"2022-08-09T08:15:38.183967Z","shell.execute_reply.started":"2022-08-09T08:15:37.859332Z","shell.execute_reply":"2022-08-09T08:15:38.182951Z"},"trusted":true},"execution_count":49,"outputs":[{"execution_count":49,"output_type":"execute_result","data":{"text/plain":"2824"},"metadata":{}}]},{"cell_type":"code","source":" df_sub=pd.read_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:38.185612Z","iopub.execute_input":"2022-08-09T08:15:38.187377Z","iopub.status.idle":"2022-08-09T08:15:47.521791Z","shell.execute_reply.started":"2022-08-09T08:15:38.187295Z","shell.execute_reply":"2022-08-09T08:15:47.519939Z"},"trusted":true},"execution_count":50,"outputs":[]},{"cell_type":"code","source":"df_sub.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:47.525493Z","iopub.execute_input":"2022-08-09T08:15:47.52717Z","iopub.status.idle":"2022-08-09T08:15:47.544939Z","shell.execute_reply.started":"2022-08-09T08:15:47.52711Z","shell.execute_reply":"2022-08-09T08:15:47.543705Z"},"trusted":true},"execution_count":51,"outputs":[{"name":"stdout","text":"<class 'pandas.core.frame.DataFrame'>\nRangeIndex: 11363784 entries, 0 to 11363783\nData columns (total 2 columns):\n #   Column       Dtype \n---  ------       ----- \n 0   customer_ID  object\n 1   predicted    object\ndtypes: object(2)\nmemory usage: 173.4+ MB\n","output_type":"stream"}]},{"cell_type":"code","source":"df_sub.customer_ID.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:47.546798Z","iopub.execute_input":"2022-08-09T08:15:47.547995Z","iopub.status.idle":"2022-08-09T08:15:49.658409Z","shell.execute_reply.started":"2022-08-09T08:15:47.547953Z","shell.execute_reply":"2022-08-09T08:15:49.657033Z"},"trusted":true},"execution_count":52,"outputs":[{"execution_count":52,"output_type":"execute_result","data":{"text/plain":"924622"},"metadata":{}}]},{"cell_type":"code","source":"\ndf_sub.drop(df_sub.loc[df_sub['customer_ID']=='customer_ID'].index,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:49.661064Z","iopub.execute_input":"2022-08-09T08:15:49.662405Z","iopub.status.idle":"2022-08-09T08:15:51.992186Z","shell.execute_reply.started":"2022-08-09T08:15:49.66235Z","shell.execute_reply":"2022-08-09T08:15:51.99031Z"},"trusted":true},"execution_count":53,"outputs":[]},{"cell_type":"code","source":"\ndf_sub.customer_ID.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:51.994253Z","iopub.execute_input":"2022-08-09T08:15:51.994734Z","iopub.status.idle":"2022-08-09T08:15:54.102921Z","shell.execute_reply.started":"2022-08-09T08:15:51.994695Z","shell.execute_reply":"2022-08-09T08:15:54.101207Z"},"trusted":true},"execution_count":54,"outputs":[{"execution_count":54,"output_type":"execute_result","data":{"text/plain":"924621"},"metadata":{}}]},{"cell_type":"code","source":"df_final_sub=df_sub.groupby('customer_ID').tail(1)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:54.104611Z","iopub.execute_input":"2022-08-09T08:15:54.105013Z","iopub.status.idle":"2022-08-09T08:15:58.127895Z","shell.execute_reply.started":"2022-08-09T08:15:54.10498Z","shell.execute_reply":"2022-08-09T08:15:58.126591Z"},"trusted":true},"execution_count":55,"outputs":[]},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"code","source":"df_final_sub.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:58.129553Z","iopub.execute_input":"2022-08-09T08:15:58.129984Z","iopub.status.idle":"2022-08-09T08:15:58.245078Z","shell.execute_reply.started":"2022-08-09T08:15:58.129934Z","shell.execute_reply":"2022-08-09T08:15:58.24345Z"},"trusted":true},"execution_count":56,"outputs":[{"name":"stdout","text":"<class 'pandas.core.frame.DataFrame'>\nInt64Index: 924621 entries, 8 to 11363783\nData columns (total 2 columns):\n #   Column       Non-Null Count   Dtype \n---  ------       --------------   ----- \n 0   customer_ID  924621 non-null  object\n 1   predicted    924621 non-null  object\ndtypes: object(2)\nmemory usage: 21.2+ MB\n","output_type":"stream"}]},{"cell_type":"code","source":"df_final_sub.predicted.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:58.246684Z","iopub.execute_input":"2022-08-09T08:15:58.247064Z","iopub.status.idle":"2022-08-09T08:15:58.277452Z","shell.execute_reply.started":"2022-08-09T08:15:58.247028Z","shell.execute_reply":"2022-08-09T08:15:58.275775Z"},"trusted":true},"execution_count":57,"outputs":[{"execution_count":57,"output_type":"execute_result","data":{"text/plain":"4"},"metadata":{}}]},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"code","source":"df_final_sub.predicted= df_final_sub.predicted.astype (float)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:58.279928Z","iopub.execute_input":"2022-08-09T08:15:58.280812Z","iopub.status.idle":"2022-08-09T08:15:58.395445Z","shell.execute_reply.started":"2022-08-09T08:15:58.280753Z","shell.execute_reply":"2022-08-09T08:15:58.393844Z"},"trusted":true},"execution_count":58,"outputs":[]},{"cell_type":"code","source":"df_final_sub.predicted= df_final_sub.predicted.astype ('Int64')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:58.39827Z","iopub.execute_input":"2022-08-09T08:15:58.398753Z","iopub.status.idle":"2022-08-09T08:15:58.413133Z","shell.execute_reply.started":"2022-08-09T08:15:58.398715Z","shell.execute_reply":"2022-08-09T08:15:58.411738Z"},"trusted":true},"execution_count":59,"outputs":[]},{"cell_type":"code","source":"df_final_sub.predicted.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:58.415497Z","iopub.execute_input":"2022-08-09T08:15:58.415881Z","iopub.status.idle":"2022-08-09T08:15:58.438205Z","shell.execute_reply.started":"2022-08-09T08:15:58.415846Z","shell.execute_reply":"2022-08-09T08:15:58.437003Z"},"trusted":true},"execution_count":60,"outputs":[{"execution_count":60,"output_type":"execute_result","data":{"text/plain":"0    640347\n1    284274\nName: predicted, dtype: Int64"},"metadata":{}}]},{"cell_type":"code","source":"df_final_sub.to_csv(\"Amex_default_pallavi.csv\",header=True,index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:15:58.440762Z","iopub.execute_input":"2022-08-09T08:15:58.441875Z","iopub.status.idle":"2022-08-09T08:16:00.527019Z","shell.execute_reply.started":"2022-08-09T08:15:58.441834Z","shell.execute_reply":"2022-08-09T08:16:00.525856Z"},"trusted":true},"execution_count":61,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}