{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"## Importing the libraries:\n\n\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nimport random\n\npd.options.display.precision = 2\npd.options.display.max_rows = 100\n#pd.set_option('display.max_rows', 20)\n\nimport numpy as np\n\n# For Feature Selection\nfrom sklearn.feature_selection import SelectFromModel\n\n# For Feature Importances\nfrom yellowbrick.model_selection import FeatureImportances\n\n# For metrics evaluation\nfrom sklearn.metrics import precision_recall_curve, classification_report, plot_confusion_matrix\n\n# For Data Modeling\n#from sklearn.model_selection import train_test_split ,cross_val_score,GridSearchCV ,  # To split the data in training and testing part   \nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC\n#from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\n\n# To Disable Warnings\nimport warnings\nwarnings.filterwarnings(action = \"ignore\")\n\nfrom scipy.stats import loguniform\nfrom sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold, GridSearchCV, RandomizedSearchCV\nfrom sklearn.metrics import confusion_matrix, roc_auc_score, roc_curve, classification_report, precision_recall_curve\nfrom sklearn.model_selection import RepeatedStratifiedKFold\nfrom sklearn.metrics import accuracy_score                          # For calculating the accuracy for the model\nfrom sklearn.metrics import precision_score                         # For calculating the Precision of the model\nfrom sklearn.metrics import recall_score                            # For calculating the recall of the model\n#from sklearn.metrics import precision_recall_curve                  # For precision and recall metric estimation\n#from sklearn.metrics import confusion_matrix                        # For verifying model performance using confusion matrix\nfrom sklearn.metrics import f1_score                                # For Checking the F1-Score of our model  \n#from sklearn.metrics import roc_curve  \n# For Roc-Auc metric estimation\n\n# Importing missingno for missing value plot\nimport missingno as msno\n\n# Importing datetime for using datetime\nfrom datetime import datetime\nimport warnings                                                     # Importing warning to disable runtime warnings\nwarnings.filterwarnings(\"ignore\")                                   # Warnings will appear only once","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:48:07.876087Z","iopub.execute_input":"2022-08-09T08:48:07.876398Z","iopub.status.idle":"2022-08-09T08:48:09.200484Z","shell.execute_reply.started":"2022-08-09T08:48:07.876320Z","shell.execute_reply":"2022-08-09T08:48:09.199537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.0 Description:\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nIn this competition, you’ll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.","metadata":{}},{"cell_type":"markdown","source":"### Evaluation Metric:\nThe evaluation metric, , for this competition is the mean of two measures of rank ordering: Normalized Gini Coefficient, , and default rate captured at 4%, .\n\nThe default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n\nFor both of the sub-metrics  and , the negative labels are given a weight of 20 to adjust for downsampling.\n\nThis metric has a maximum value of 1.0.\n\nPython code for calculating this metric can be found in this Notebook.\n\nSubmission File\nFor each customer_ID in the test set, you must predict a probability for the target variable. The file should contain a header and have the following format:","metadata":{}},{"cell_type":"markdown","source":"### 2.0  Reading the train & test dataset .","metadata":{}},{"cell_type":"markdown","source":"### 2.1 Dowloading & reading the train labels dataset.","metadata":{}},{"cell_type":"code","source":"## Since the train data is huge we are considering 100000 rows.\ntrain_df = pd.read_csv('../input/amex-default-prediction/train_data.csv',nrows= 1000000)    \ntest_df = pd.read_csv('../input/amex-default-prediction/test_data.csv', nrows= 1000000)      \ntrain_label_df= pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\",nrows= 1000000)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:48:09.202457Z","iopub.execute_input":"2022-08-09T08:48:09.202809Z","iopub.status.idle":"2022-08-09T08:50:07.905172Z","shell.execute_reply.started":"2022-08-09T08:48:09.202773Z","shell.execute_reply":"2022-08-09T08:50:07.904144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observations: \n- We have read 100000 rows & 190 cols. \n- The data set has unique customer ID's & S_2 in date format.","metadata":{}},{"cell_type":"markdown","source":"#### Getting Info of train data: ","metadata":{}},{"cell_type":"code","source":"train_df.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:07.906603Z","iopub.execute_input":"2022-08-09T08:50:07.907205Z","iopub.status.idle":"2022-08-09T08:50:07.949203Z","shell.execute_reply.started":"2022-08-09T08:50:07.907167Z","shell.execute_reply":"2022-08-09T08:50:07.948343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The dataset uses a memory of 145MB.\n- Has 185 cols float type \n- Has 4 object type cols: Customer ID, S_2,D_63, D_64.\n","metadata":{}},{"cell_type":"markdown","source":"#### Describing the train data.","metadata":{}},{"cell_type":"code","source":"#train_df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:07.953335Z","iopub.execute_input":"2022-08-09T08:50:07.953708Z","iopub.status.idle":"2022-08-09T08:50:07.960620Z","shell.execute_reply.started":"2022-08-09T08:50:07.953659Z","shell.execute_reply":"2022-08-09T08:50:07.959645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The dataset \" count\" column shows many missing values. \n- There is skewness in the dataset.\n- Needs to perform standardization/ Normalization of the dataset columns.","metadata":{}},{"cell_type":"markdown","source":"#### Checking Info of Test data.","metadata":{}},{"cell_type":"code","source":"#test_df.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:07.962933Z","iopub.execute_input":"2022-08-09T08:50:07.963332Z","iopub.status.idle":"2022-08-09T08:50:07.969534Z","shell.execute_reply.started":"2022-08-09T08:50:07.963299Z","shell.execute_reply":"2022-08-09T08:50:07.968450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The test dataset uses a memory of 145MB.\n- Has 185 cols float type \n- Has 4 object type cols: Customer ID, S_2,D_63, D_64.","metadata":{}},{"cell_type":"markdown","source":"#### Checking the describe function on test data.","metadata":{}},{"cell_type":"code","source":"#test_df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:07.971132Z","iopub.execute_input":"2022-08-09T08:50:07.971506Z","iopub.status.idle":"2022-08-09T08:50:07.979929Z","shell.execute_reply.started":"2022-08-09T08:50:07.971473Z","shell.execute_reply":"2022-08-09T08:50:07.979014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The test dataset \" count\" column shows many missing values. \n- There is skewness in the dataset.\n- Needs to perform standardization/ Normalization of the dataset columns.","metadata":{}},{"cell_type":"markdown","source":"#### Info on train label data.","metadata":{}},{"cell_type":"code","source":"#train_label_df.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:07.981033Z","iopub.execute_input":"2022-08-09T08:50:07.981278Z","iopub.status.idle":"2022-08-09T08:50:07.990654Z","shell.execute_reply.started":"2022-08-09T08:50:07.981255Z","shell.execute_reply":"2022-08-09T08:50:07.989797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The train_label_df has 100000 enteries & two cols.","metadata":{}},{"cell_type":"markdown","source":"### Describing the train_label_df.","metadata":{}},{"cell_type":"code","source":"#train_label_df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:07.992455Z","iopub.execute_input":"2022-08-09T08:50:07.992953Z","iopub.status.idle":"2022-08-09T08:50:08.000703Z","shell.execute_reply.started":"2022-08-09T08:50:07.992920Z","shell.execute_reply":"2022-08-09T08:50:07.999707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- No missing values in train_label_df.","metadata":{}},{"cell_type":"markdown","source":"### 2.0  Merging the train datset with train label dataset on \" Customer_ID","metadata":{}},{"cell_type":"code","source":"joined = train_df.merge(train_label_df, how=\"left\", on=[\"customer_ID\"])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:08.003373Z","iopub.execute_input":"2022-08-09T08:50:08.003782Z","iopub.status.idle":"2022-08-09T08:50:11.455568Z","shell.execute_reply.started":"2022-08-09T08:50:08.003747Z","shell.execute_reply":"2022-08-09T08:50:11.454266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.1  Fetching the Data.info() of train dataset.","metadata":{}},{"cell_type":"code","source":"# Finding information about DataFrame\n#joined.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:11.462703Z","iopub.execute_input":"2022-08-09T08:50:11.463369Z","iopub.status.idle":"2022-08-09T08:50:11.471507Z","shell.execute_reply.started":"2022-08-09T08:50:11.463332Z","shell.execute_reply":"2022-08-09T08:50:11.470664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n- The Merged dataset has 190 cols.\n- 185 Cols are float type, 2 cols int64 type, 4 cols object type.","metadata":{}},{"cell_type":"code","source":"#joined.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:11.473682Z","iopub.execute_input":"2022-08-09T08:50:11.476604Z","iopub.status.idle":"2022-08-09T08:50:11.480630Z","shell.execute_reply.started":"2022-08-09T08:50:11.476560Z","shell.execute_reply":"2022-08-09T08:50:11.479802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observations: \n- The merged dataset has missing values & the data set is skewed.","metadata":{}},{"cell_type":"code","source":"#test_df.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:11.482503Z","iopub.execute_input":"2022-08-09T08:50:11.484899Z","iopub.status.idle":"2022-08-09T08:50:11.493067Z","shell.execute_reply.started":"2022-08-09T08:50:11.484852Z","shell.execute_reply":"2022-08-09T08:50:11.491917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observations:\n- The test dataset has 189 cols out of which 185 cols float type, 1col int64 type, 4 object type cols.","metadata":{}},{"cell_type":"code","source":"#test_df.describe().T","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:11.494673Z","iopub.execute_input":"2022-08-09T08:50:11.495650Z","iopub.status.idle":"2022-08-09T08:50:11.503314Z","shell.execute_reply.started":"2022-08-09T08:50:11.495615Z","shell.execute_reply":"2022-08-09T08:50:11.501919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3.0 Numerical Data Distribution:","metadata":{}},{"cell_type":"code","source":"num_feature = []\n\nfor i in joined.columns.values:\n    if ((joined[i].dtype == int) | (joined[i].dtype == float)):\n        num_feature.append(i)\n    \nprint('Total Numerical Features:', len(num_feature))\nprint('Features:', num_feature)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:11.504627Z","iopub.execute_input":"2022-08-09T08:50:11.505728Z","iopub.status.idle":"2022-08-09T08:50:11.530941Z","shell.execute_reply.started":"2022-08-09T08:50:11.505663Z","shell.execute_reply":"2022-08-09T08:50:11.529707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observations: \n- There are total 187 cols numerical .","metadata":{}},{"cell_type":"markdown","source":"## Deleting the coloumns from train data that has correlation >= 0.7","metadata":{}},{"cell_type":"code","source":"## Deleting the coloumns from train data that has correlation >= 0.7\n# Create correlation matrix\ncorr_matrix = joined.corr().abs()\n\n# Select upper triangle of correlation matrix\nupper = corr_matrix.where(np.triu(np.ones(corr_matrix.shape), k=1).astype(np.bool_))\n\n# Find features with correlation greater or equal to 0.7\nto_drop = [column for column in upper.columns if any(upper[column] >= 0.7)]\nprint('columns to drop in the train data set',to_drop)\n\n# Drop features \njoined.drop(to_drop, axis=1, inplace=True)\nlen(joined)\njoined.columns\njoined.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:50:11.535351Z","iopub.execute_input":"2022-08-09T08:50:11.538463Z","iopub.status.idle":"2022-08-09T08:51:23.334147Z","shell.execute_reply.started":"2022-08-09T08:50:11.536475Z","shell.execute_reply":"2022-08-09T08:51:23.333035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5.0 Finding the missing values of Train data.","metadata":{}},{"cell_type":"code","source":"## Train data missing values\nmissing_values = joined.isna().sum()\npercent_missing = ((missing_values / joined.index.size) * 100)          \n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:23.335596Z","iopub.execute_input":"2022-08-09T08:51:23.336046Z","iopub.status.idle":"2022-08-09T08:51:23.732033Z","shell.execute_reply.started":"2022-08-09T08:51:23.336012Z","shell.execute_reply":"2022-08-09T08:51:23.731098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Dropping columns from train data that have more than 50% missing values\ncolumns_to_drop = list(percent_missing[percent_missing >= 50].index)\nJoined_1 = joined.drop(columns_to_drop, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:23.733586Z","iopub.execute_input":"2022-08-09T08:51:23.733965Z","iopub.status.idle":"2022-08-09T08:51:24.026006Z","shell.execute_reply.started":"2022-08-09T08:51:23.733928Z","shell.execute_reply":"2022-08-09T08:51:24.024923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Observation: \n- 117 Columns & 1000000 rows are left after dropping missing values that are greater than or equal to 50.","metadata":{}},{"cell_type":"markdown","source":"## Fetching the Unique values from Train dataset.","metadata":{}},{"cell_type":"code","source":"## Getting Uique values in train dataset\n#for col in Joined_1:\n    #print(col)\n    #print(Joined_1[col].unique())\n    #print('\\n')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:24.027674Z","iopub.execute_input":"2022-08-09T08:51:24.028070Z","iopub.status.idle":"2022-08-09T08:51:24.032650Z","shell.execute_reply.started":"2022-08-09T08:51:24.028033Z","shell.execute_reply":"2022-08-09T08:51:24.031439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Filling the missing values with mode & median.","metadata":{}},{"cell_type":"code","source":"## To fill the Mode in object type col\n#Joined_1['D_64'] = Joined_1['D_64'].mode()[0]\nJoined_1['D_64'].fillna(Joined_1['D_64'].mode()[0], inplace=True)\nJoined_1['D_64'].isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:24.034299Z","iopub.execute_input":"2022-08-09T08:51:24.034711Z","iopub.status.idle":"2022-08-09T08:51:24.158539Z","shell.execute_reply.started":"2022-08-09T08:51:24.034676Z","shell.execute_reply":"2022-08-09T08:51:24.157456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"##Fill NaN values with 0 & then using mean value to impute  in col D_43 of train \nJoined_1['D_43'] = Joined_1['D_43'].fillna(0)\nJoined_1['D_43'] = Joined_1['D_43'].fillna(Joined_1['D_43'].mean())\n\nJoined_1['D_68'] = Joined_1['D_68'].fillna(0)\nJoined_1['D_68'] = Joined_1['D_68'].fillna(Joined_1['D_68'].mean())\n\nJoined_1['D_114'] = Joined_1['D_114'].fillna(0)\nJoined_1['D_114'] = Joined_1['D_114'].fillna(Joined_1['D_114'].mean())\n\nJoined_1['D_120'] = Joined_1['D_120'].fillna(0)\nJoined_1['D_120'] = Joined_1['D_120'].fillna(Joined_1['D_120'].mean())\n\nJoined_1['D_126'] = Joined_1['D_126'].fillna(0)\nJoined_1['D_126'] = Joined_1['D_126'].fillna(Joined_1['D_126'].mean())\n\nJoined_1['D_116'] = Joined_1['D_116'].fillna(0)\nJoined_1['D_116'] = Joined_1['D_116'].fillna(Joined_1['D_116'].mean())\n\nJoined_1['D_117'] = Joined_1['D_117'].fillna(0)\nJoined_1['D_117'] = Joined_1['D_117'].fillna(Joined_1['D_117'].mean())","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:24.160196Z","iopub.execute_input":"2022-08-09T08:51:24.160562Z","iopub.status.idle":"2022-08-09T08:51:24.242233Z","shell.execute_reply.started":"2022-08-09T08:51:24.160527Z","shell.execute_reply":"2022-08-09T08:51:24.241260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Checking the Uique values in after filling the Nan values in train dataset.","metadata":{}},{"cell_type":"code","source":"#for col in Joined_1:\n#    print(col)\n#    print(Joined_1[col].unique())\n#    print('\\n')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:24.243741Z","iopub.execute_input":"2022-08-09T08:51:24.244114Z","iopub.status.idle":"2022-08-09T08:51:24.249902Z","shell.execute_reply.started":"2022-08-09T08:51:24.244078Z","shell.execute_reply":"2022-08-09T08:51:24.248931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6.0  Value Counts of target variable & its 'count plot'.","metadata":{}},{"cell_type":"code","source":"Joined_1['target'].value_counts() /len(Joined_1['target']) \nsns.countplot(x='target', data=Joined_1, palette='hls')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:24.251294Z","iopub.execute_input":"2022-08-09T08:51:24.252549Z","iopub.status.idle":"2022-08-09T08:51:24.550334Z","shell.execute_reply.started":"2022-08-09T08:51:24.252511Z","shell.execute_reply":"2022-08-09T08:51:24.549395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observations: Target variable is unbalanced. Hence need to perform  sampling techniques.","metadata":{}},{"cell_type":"markdown","source":"#### 7.0 Dropping the customer id & S2 Column from the train & test data set.","metadata":{}},{"cell_type":"code","source":"\nJoined_1.drop(['customer_ID','S_2'],axis=1, inplace= True)\nprint(Joined_1.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:24.551703Z","iopub.execute_input":"2022-08-09T08:51:24.552618Z","iopub.status.idle":"2022-08-09T08:51:24.840006Z","shell.execute_reply.started":"2022-08-09T08:51:24.552571Z","shell.execute_reply":"2022-08-09T08:51:24.838845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Joined_1 = Joined_1.replace([np.inf, -np.inf], np.nan)\nJoined_1 = Joined_1.dropna()\nJoined_1 = Joined_1.reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:24.841399Z","iopub.execute_input":"2022-08-09T08:51:24.841932Z","iopub.status.idle":"2022-08-09T08:51:26.785488Z","shell.execute_reply.started":"2022-08-09T08:51:24.841893Z","shell.execute_reply":"2022-08-09T08:51:26.784482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 8.0 Performing One Hot encoding on Train data.","metadata":{}},{"cell_type":"code","source":"#from sklearn.preprocessing import OneHotEncoder\n\none_hot_encoded_train = pd.get_dummies(Joined_1, columns = ['D_63', 'D_64'], drop_first='True')\n\n### For Train data.\nX = one_hot_encoded_train.drop('target',axis = 1)\ny = one_hot_encoded_train['target']\n#print(X)\n\nX_copy= X.copy()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:26.786993Z","iopub.execute_input":"2022-08-09T08:51:26.787594Z","iopub.status.idle":"2022-08-09T08:51:27.821674Z","shell.execute_reply.started":"2022-08-09T08:51:26.787540Z","shell.execute_reply":"2022-08-09T08:51:27.820675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 9.0 Feature Selection Using Random Forest.","metadata":{}},{"cell_type":"code","source":"# create the classifier with n_estimators = 100\n\nclf = RandomForestClassifier(n_estimators=50, random_state=0)\n\n# fit the model to the training set\n\nclf.fit(X, y)\n# view the feature scores\n\nfeature_scores = pd.Series(clf.feature_importances_, index=X.columns).sort_values(ascending=False)\n\nfeature_scores.head(15)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T08:51:27.823323Z","iopub.execute_input":"2022-08-09T08:51:27.823702Z","iopub.status.idle":"2022-08-09T09:01:09.555928Z","shell.execute_reply.started":"2022-08-09T08:51:27.823662Z","shell.execute_reply":"2022-08-09T09:01:09.554823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 10.0 Filtering out the required column  from train based upon feature importance.","metadata":{}},{"cell_type":"code","source":"X = X.filter(['P_2','B_9','D_44','B_2','B_1','B_7','B_6','D_45','B_10','D_52','S_3','D_62','B_4','R_27','D_43'])\nX.columns\nX.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:01:09.557526Z","iopub.execute_input":"2022-08-09T09:01:09.557924Z","iopub.status.idle":"2022-08-09T09:01:09.579300Z","shell.execute_reply.started":"2022-08-09T09:01:09.557888Z","shell.execute_reply":"2022-08-09T09:01:09.578282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 11.0 Split the train data into 70:30 percentage.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split \nX_train, X_test_1, y_train, y_test_1 = train_test_split(X, y, test_size = 0.3, random_state = 0,stratify= y)\n\n# describes info about train and test set\nprint(\"Number transactions X_train dataset: \", X_train.shape)\nprint(\"Number transactions y_train dataset: \", y_train.shape)\nprint(\"Number transactions X_test_1 dataset: \", X_test_1.shape)\nprint(\"Number transactions y_test_1 dataset: \", y_test_1.shape)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:01:09.587512Z","iopub.execute_input":"2022-08-09T09:01:09.588366Z","iopub.status.idle":"2022-08-09T09:01:09.818536Z","shell.execute_reply.started":"2022-08-09T09:01:09.588332Z","shell.execute_reply":"2022-08-09T09:01:09.817379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 12.0  Checking for the value_counts of y_train dataset.","metadata":{}},{"cell_type":"code","source":"#y_train.value_counts()/len(y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:01:09.819992Z","iopub.execute_input":"2022-08-09T09:01:09.820439Z","iopub.status.idle":"2022-08-09T09:01:09.827911Z","shell.execute_reply.started":"2022-08-09T09:01:09.820402Z","shell.execute_reply":"2022-08-09T09:01:09.824251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 13.0 Scaling the data with Robust scaler as data has outliers.\n### perform a robust scaler transform of the dataset.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import RobustScaler\nfrom pandas import DataFrame\n\ntrans = RobustScaler(with_centering=False, with_scaling=True)\n\n# convert the array back to a dataframe\nX_train = trans.fit_transform(X_train)\nX_test_1 = trans.transform(X_test_1)\n\nX_train_df= DataFrame(X_train)\nX_train_df.columns = X.columns\n#print(X_train_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:01:09.829131Z","iopub.execute_input":"2022-08-09T09:01:09.829566Z","iopub.status.idle":"2022-08-09T09:01:09.983354Z","shell.execute_reply.started":"2022-08-09T09:01:09.829532Z","shell.execute_reply":"2022-08-09T09:01:09.982364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 14.0 Using Smote() Techniques.","metadata":{}},{"cell_type":"code","source":"from imblearn.over_sampling import SMOTE\nfrom collections import Counter\ncounter = Counter(y_train)\n#print('Before',counter)\n\n# oversampling the train dataset using SMOTE\nsmt = SMOTE()\nX_train_sm, y_train_sm = smt.fit_resample(X_train, y_train)\n\ncounter = Counter(y_train_sm)\n#print('After',counter)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:01:09.984623Z","iopub.execute_input":"2022-08-09T09:01:09.984988Z","iopub.status.idle":"2022-08-09T09:02:04.876902Z","shell.execute_reply.started":"2022-08-09T09:01:09.984953Z","shell.execute_reply":"2022-08-09T09:02:04.875752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 15.0 Model Building - Imbalanced data","metadata":{}},{"cell_type":"code","source":"model = list()\nresample = list()\nprecision = list()\nrecall = list()\nF1score = list()\nAUCROC = list()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:02:04.878505Z","iopub.execute_input":"2022-08-09T09:02:04.878981Z","iopub.status.idle":"2022-08-09T09:02:04.886277Z","shell.execute_reply.started":"2022-08-09T09:02:04.878945Z","shell.execute_reply":"2022-08-09T09:02:04.884114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def test_eval(clf_model, X_test_1, y_test_1, algo=None, sampling=None):\n    \n    # Test set prediction\n    y_prob=clf_model.predict_proba(X_test_1)\n    y_pred=clf_model.predict(X_test_1)\n    \n    print('Confusion Matrix')\n    print('='*60)\n    print(confusion_matrix(y_test_1,y_pred),\"\\n\")\n    print('Classification Report')\n    print('='*60)\n    print(classification_report(y_test_1,y_pred),\"\\n\")\n    print('AUC-ROC')\n    print('='*60)\n    print(roc_auc_score(y_test_1, y_prob[:,1]))\n          \n    model.append(algo)\n    precision.append(precision_score(y_test_1,y_pred))\n    recall.append(recall_score(y_test_1,y_pred))\n    F1score.append(f1_score(y_test_1,y_pred))\n    AUCROC.append(roc_auc_score(y_test_1, y_prob[:,1]))\n    resample.append(sampling)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:02:04.887783Z","iopub.execute_input":"2022-08-09T09:02:04.888892Z","iopub.status.idle":"2022-08-09T09:02:04.897837Z","shell.execute_reply.started":"2022-08-09T09:02:04.888839Z","shell.execute_reply":"2022-08-09T09:02:04.896933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 15.1 Applying the logistic regression model with cross validation & using random search.","metadata":{}},{"cell_type":"code","source":"# define model\nlog_model = LogisticRegression()\n\n# define evaluation\ncv = RepeatedStratifiedKFold(n_splits=4, n_repeats=2, random_state=1)\n\n# define search space\nspace = dict()\nspace['solver'] = ['newton-cg']\nspace['penalty'] = [ 'l2']\nspace['C'] = loguniform(1e-5, 100)\n\n# define search\nclf_LR = RandomizedSearchCV(log_model, space, n_iter=2, scoring='roc_auc', n_jobs=-1, cv=cv, random_state=1)\n\n# execute search\nresult = clf_LR.fit(X_train, y_train)\n\n# summarize result\nprint('Best Score: %s' % result.best_score_)\nprint('Best Hyperparameters: %s' % result.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:02:04.898994Z","iopub.execute_input":"2022-08-09T09:02:04.901537Z","iopub.status.idle":"2022-08-09T09:03:25.887573Z","shell.execute_reply.started":"2022-08-09T09:02:04.901509Z","shell.execute_reply":"2022-08-09T09:03:25.885995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_eval(clf_LR, X_test_1, y_test_1, 'Logistic Regression', 'actual')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:03:25.894212Z","iopub.execute_input":"2022-08-09T09:03:25.898192Z","iopub.status.idle":"2022-08-09T09:03:26.561502Z","shell.execute_reply.started":"2022-08-09T09:03:25.898126Z","shell.execute_reply":"2022-08-09T09:03:26.560502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Fitting the logistic Regression model on smote applied on train data.","metadata":{}},{"cell_type":"code","source":"## SMOTE Resampling\nclf_LR.fit(X_train_sm, y_train_sm)\nclf_LR.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:03:26.563074Z","iopub.execute_input":"2022-08-09T09:03:26.563458Z","iopub.status.idle":"2022-08-09T09:05:19.285701Z","shell.execute_reply.started":"2022-08-09T09:03:26.563421Z","shell.execute_reply":"2022-08-09T09:05:19.284410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_eval(clf_LR, X_test_1, y_test_1, 'Logistic Regression', 'smote')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:05:19.291121Z","iopub.execute_input":"2022-08-09T09:05:19.292151Z","iopub.status.idle":"2022-08-09T09:05:19.956283Z","shell.execute_reply.started":"2022-08-09T09:05:19.292115Z","shell.execute_reply":"2022-08-09T09:05:19.955223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#predicting on test data\ny_pred_test_LR = clf_LR.predict(X_test_1)\ny_pred_test_LR","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:05:19.957609Z","iopub.execute_input":"2022-08-09T09:05:19.958237Z","iopub.status.idle":"2022-08-09T09:05:19.979530Z","shell.execute_reply.started":"2022-08-09T09:05:19.958200Z","shell.execute_reply":"2022-08-09T09:05:19.978278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 15.2  Applying Random Forest algorith with cross validation & using random search.","metadata":{}},{"cell_type":"code","source":"estimators = [30]\n# Maximum number of depth in each tree:\nmax_depth = [i for i in range(5,16,2)]\n# Minimum number of samples to consider to split a node:\nmin_samples_split = [10]        \n# Minimum number of samples to consider at each leaf node:\nmin_samples_leaf = [1, 2, 5]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:05:19.984011Z","iopub.execute_input":"2022-08-09T09:05:19.985323Z","iopub.status.idle":"2022-08-09T09:05:19.999535Z","shell.execute_reply.started":"2022-08-09T09:05:19.985270Z","shell.execute_reply":"2022-08-09T09:05:19.998246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_model = RandomForestClassifier() \n\nrf_params={'n_estimators':estimators,\n           'max_depth':max_depth,\n           'min_samples_split':min_samples_split}\n\nclf_RF = RandomizedSearchCV(rf_model, rf_params, cv=cv, scoring='roc_auc', n_jobs=-1, n_iter=3, verbose=2)\nclf_RF.fit(X_train, y_train)\nclf_RF.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:05:20.001349Z","iopub.execute_input":"2022-08-09T09:05:20.004261Z","iopub.status.idle":"2022-08-09T09:13:31.229766Z","shell.execute_reply.started":"2022-08-09T09:05:20.004197Z","shell.execute_reply":"2022-08-09T09:13:31.228576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_eval(clf_RF, X_test_1, y_test_1, 'Random Forest', 'actual')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:13:31.231157Z","iopub.execute_input":"2022-08-09T09:13:31.232178Z","iopub.status.idle":"2022-08-09T09:13:33.582935Z","shell.execute_reply.started":"2022-08-09T09:13:31.232139Z","shell.execute_reply":"2022-08-09T09:13:33.581914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Fitting the Random Forest model on smote applied on train data.","metadata":{}},{"cell_type":"code","source":"## 2.SMOTE Resampling\nclf_RF.fit(X_train_sm, y_train_sm)\nclf_RF.best_estimator_\n#y_pred_train= clf_model.predict(X_train)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:13:33.584365Z","iopub.execute_input":"2022-08-09T09:13:33.585111Z","iopub.status.idle":"2022-08-09T09:31:18.070799Z","shell.execute_reply.started":"2022-08-09T09:13:33.585068Z","shell.execute_reply":"2022-08-09T09:31:18.069530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_eval(clf_RF, X_test_1, y_test_1, 'Random Forest', 'smote')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:18.072430Z","iopub.execute_input":"2022-08-09T09:31:18.073101Z","iopub.status.idle":"2022-08-09T09:31:21.158045Z","shell.execute_reply.started":"2022-08-09T09:31:18.073057Z","shell.execute_reply":"2022-08-09T09:31:21.156885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#predicting on test data\ny_pred_test_RF = clf_RF.predict(X_test_1,)\ny_pred_test_RF","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:21.160016Z","iopub.execute_input":"2022-08-09T09:31:21.161173Z","iopub.status.idle":"2022-08-09T09:31:22.253379Z","shell.execute_reply.started":"2022-08-09T09:31:21.161116Z","shell.execute_reply":"2022-08-09T09:31:22.252200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Comparision.","metadata":{}},{"cell_type":"code","source":"clf_eval_df = pd.DataFrame({'model':model,\n                            'resample':resample,\n                            'precision':precision,\n                            'recall':recall,\n                            'f1-score':F1score,\n                            'AUC-ROC':AUCROC})","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.255100Z","iopub.execute_input":"2022-08-09T09:31:22.255632Z","iopub.status.idle":"2022-08-09T09:31:22.264288Z","shell.execute_reply.started":"2022-08-09T09:31:22.255585Z","shell.execute_reply":"2022-08-09T09:31:22.262934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clf_eval_df","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.266155Z","iopub.execute_input":"2022-08-09T09:31:22.267646Z","iopub.status.idle":"2022-08-09T09:31:22.290642Z","shell.execute_reply.started":"2022-08-09T09:31:22.267586Z","shell.execute_reply":"2022-08-09T09:31:22.289356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#sns.set(font_scale=1.2)\n#sns.palplot(sns.color_palette())\n#g = sns.FacetGrid(clf_eval_df, col=\"model\", height=5)\n#g.map(sns.barplot, \"resample\", \"recall\", palette='twilight', order=[\"actual\", \"smote\", \"adasyn\", \"smote+tomek\", \"smote+enn\"])\n#g.set_xticklabels(rotation=30)\n#g.set_xlabels(' ', fontsize=14)\n#g.set_ylabels('Recall', fontsize=14)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.291993Z","iopub.execute_input":"2022-08-09T09:31:22.292471Z","iopub.status.idle":"2022-08-09T09:31:22.298700Z","shell.execute_reply.started":"2022-08-09T09:31:22.292426Z","shell.execute_reply":"2022-08-09T09:31:22.297250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"##predictions for actual Test data.\n\ndef amex_metric_mod(y_true, y_pred):\n\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.300957Z","iopub.execute_input":"2022-08-09T09:31:22.302009Z","iopub.status.idle":"2022-08-09T09:31:22.314502Z","shell.execute_reply.started":"2022-08-09T09:31:22.301840Z","shell.execute_reply":"2022-08-09T09:31:22.313341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"   print(amex_metric_mod(y_test_1, y_pred_test_RF )) ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.318425Z","iopub.execute_input":"2022-08-09T09:31:22.318772Z","iopub.status.idle":"2022-08-09T09:31:22.366918Z","shell.execute_reply.started":"2022-08-09T09:31:22.318742Z","shell.execute_reply":"2022-08-09T09:31:22.365609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5.3 Processing of the test data.","metadata":{}},{"cell_type":"code","source":"def create_model(test_df,submission):\n    \n    ## Missing values\n    test_df.drop('customer_ID',axis=1,inplace=True)\n    test_df.isna().sum()/len(test_df)*100\n         \n    ## Filling the missing values\n    test_df .fillna(method='ffill',inplace=True)\n    test_df .fillna(method='bfill',inplace=True)\n    \n    ## Use of Robust scaler to scale the data\n    test_df_final = trans.transform(test_df)\n\n    #Test_df= DataFrame(test_df_final)\n    #Test_df.columns = test_df.columns\n        \n    ## Predicting the target variable\n    submission['predicted']=clf_RF.predict(test_df_final)\n            \n    # Merge the prediction and customer_ID into submission dataframe\n    #submission = pd.DataFrame({\"customer_ID\":test_df.customer_ID,\"prediction\":Test_df['predicted']})\n\n    submission.to_csv('submission.csv',mode='a', index=False)\n            \n        ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.368924Z","iopub.execute_input":"2022-08-09T09:31:22.369418Z","iopub.status.idle":"2022-08-09T09:31:22.377972Z","shell.execute_reply.started":"2022-08-09T09:31:22.369370Z","shell.execute_reply":"2022-08-09T09:31:22.376358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_col=['customer_ID','P_2','B_9','D_44','B_2','B_1','B_7','B_6','D_45','B_10','D_52','S_3','D_62','B_4','R_27','D_43']","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.380081Z","iopub.execute_input":"2022-08-09T09:31:22.380510Z","iopub.status.idle":"2022-08-09T09:31:22.390466Z","shell.execute_reply.started":"2022-08-09T09:31:22.380469Z","shell.execute_reply":"2022-08-09T09:31:22.389183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df_data = pd.read_csv('../input/amex-default-prediction/test_data.csv', chunksize=500000, iterator=True, usecols =final_col)\ntest_df_data","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.392243Z","iopub.execute_input":"2022-08-09T09:31:22.393005Z","iopub.status.idle":"2022-08-09T09:31:22.415098Z","shell.execute_reply.started":"2022-08-09T09:31:22.392962Z","shell.execute_reply":"2022-08-09T09:31:22.413734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for iter_num, chunk in enumerate(test_df_data, 1):\n    print(iter_num)\n    subm_t=pd.DataFrame(columns=['customer_ID','predicted'])\n    subm_t[\"customer_ID\"]=  chunk[\"customer_ID\"]\n   \n    #print(chunk.info())\n    #print(subm_t.info())\n    create_model(chunk,subm_t)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:31:22.417039Z","iopub.execute_input":"2022-08-09T09:31:22.417768Z","iopub.status.idle":"2022-08-09T09:40:40.562404Z","shell.execute_reply.started":"2022-08-09T09:31:22.417725Z","shell.execute_reply":"2022-08-09T09:40:40.561341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:40.563935Z","iopub.execute_input":"2022-08-09T09:40:40.565288Z","iopub.status.idle":"2022-08-09T09:40:40.753780Z","shell.execute_reply.started":"2022-08-09T09:40:40.565244Z","shell.execute_reply":"2022-08-09T09:40:40.752786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub=pd.read_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:40.755116Z","iopub.execute_input":"2022-08-09T09:40:40.757495Z","iopub.status.idle":"2022-08-09T09:40:47.758769Z","shell.execute_reply.started":"2022-08-09T09:40:40.757457Z","shell.execute_reply":"2022-08-09T09:40:47.757702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:47.760266Z","iopub.execute_input":"2022-08-09T09:40:47.760706Z","iopub.status.idle":"2022-08-09T09:40:47.775748Z","shell.execute_reply.started":"2022-08-09T09:40:47.760658Z","shell.execute_reply":"2022-08-09T09:40:47.774551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.customer_ID.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:47.777412Z","iopub.execute_input":"2022-08-09T09:40:47.778056Z","iopub.status.idle":"2022-08-09T09:40:49.449580Z","shell.execute_reply.started":"2022-08-09T09:40:47.778018Z","shell.execute_reply":"2022-08-09T09:40:49.448574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.drop(df_sub.loc[df_sub['customer_ID']=='customer_ID'].index,inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:49.451366Z","iopub.execute_input":"2022-08-09T09:40:49.452072Z","iopub.status.idle":"2022-08-09T09:40:51.186909Z","shell.execute_reply.started":"2022-08-09T09:40:49.452028Z","shell.execute_reply":"2022-08-09T09:40:51.185925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub.customer_ID.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:51.188551Z","iopub.execute_input":"2022-08-09T09:40:51.188980Z","iopub.status.idle":"2022-08-09T09:40:53.124311Z","shell.execute_reply.started":"2022-08-09T09:40:51.188943Z","shell.execute_reply":"2022-08-09T09:40:53.123326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final_sub=df_sub.groupby('customer_ID').tail(1)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:53.125902Z","iopub.execute_input":"2022-08-09T09:40:53.126283Z","iopub.status.idle":"2022-08-09T09:40:56.390321Z","shell.execute_reply.started":"2022-08-09T09:40:53.126245Z","shell.execute_reply":"2022-08-09T09:40:56.389276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final_sub.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:56.391936Z","iopub.execute_input":"2022-08-09T09:40:56.392286Z","iopub.status.idle":"2022-08-09T09:40:56.482319Z","shell.execute_reply.started":"2022-08-09T09:40:56.392249Z","shell.execute_reply":"2022-08-09T09:40:56.481378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final_sub.predicted.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:40:56.485322Z","iopub.execute_input":"2022-08-09T09:40:56.485722Z","iopub.status.idle":"2022-08-09T09:40:56.509579Z","shell.execute_reply.started":"2022-08-09T09:40:56.485689Z","shell.execute_reply":"2022-08-09T09:40:56.508576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final_sub.predicted= df_final_sub.predicted.astype (float)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:42:07.274293Z","iopub.execute_input":"2022-08-09T09:42:07.274670Z","iopub.status.idle":"2022-08-09T09:42:07.363021Z","shell.execute_reply.started":"2022-08-09T09:42:07.274635Z","shell.execute_reply":"2022-08-09T09:42:07.362083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final_sub.predicted= df_final_sub.predicted.astype ('Int64')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:42:12.772740Z","iopub.execute_input":"2022-08-09T09:42:12.773457Z","iopub.status.idle":"2022-08-09T09:42:12.787307Z","shell.execute_reply.started":"2022-08-09T09:42:12.773417Z","shell.execute_reply":"2022-08-09T09:42:12.786022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final_sub.predicted.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:42:18.682853Z","iopub.execute_input":"2022-08-09T09:42:18.683603Z","iopub.status.idle":"2022-08-09T09:42:18.706808Z","shell.execute_reply.started":"2022-08-09T09:42:18.683564Z","shell.execute_reply":"2022-08-09T09:42:18.705545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final_sub.to_csv(\"Amex_default_with smote.csv\",header=True,index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T09:44:24.037108Z","iopub.execute_input":"2022-08-09T09:44:24.037834Z","iopub.status.idle":"2022-08-09T09:44:25.857930Z","shell.execute_reply.started":"2022-08-09T09:44:24.037800Z","shell.execute_reply":"2022-08-09T09:44:25.856457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"\n","metadata":{}}]}