{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**<center><h1> AMERICAN EXPRESS (DEFAULT PREDICATION) - OPTIMIZED RF ML MODEL</h1></center>**","metadata":{}},{"cell_type":"markdown","source":"---\n# **Table of Contents**\n---\n\n**1.** [**Problem Statement**](#Section1)<br>\n**2.** [**Objective**](#Section2)<br>\n**3.** [**Installing & Importing Libraries**](#Section3)<br>\n**4.** [**Data Acquisition & Description**](#Section4)<br>\n**5.** [**Data Pre-processing**](#Section5)<br>\n**6.** [**Model Development & Evaluation**](#Section6)<br>\n**7.** [**Conclusion**](#Section7)<br>\n**8.** [**Submission**](#Section8)<br>\n","metadata":{}},{"cell_type":"markdown","source":"---\n<a name = Section1></a>\n# **1. Problem Statement**\n---\n\n* Whether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time.\n* How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n* Credit default prediction is central to managing risk in a consumer lending business. \n * Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. \n* Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\n* The objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. \n * The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"---\n<a name = Section2></a>\n# **2. Objective**\n---\n\n-  The objective of this assignment is to **predict** credit defaulters","metadata":{}},{"cell_type":"markdown","source":"---\n<a name = Section3></a>\n# **3. Installing & Importing Libraries**\n---","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-21T04:50:37.119153Z","iopub.execute_input":"2022-08-21T04:50:37.119616Z","iopub.status.idle":"2022-08-21T04:50:37.128332Z","shell.execute_reply.started":"2022-08-21T04:50:37.119581Z","shell.execute_reply":"2022-08-21T04:50:37.127356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# calculate file size in KB, MB, GB\ndef convert_bytes(size):\n    \"\"\" Convert bytes to KB, or MB or GB\"\"\"\n    for x in ['bytes', 'KB', 'MB', 'GB', 'TB']:\n        if size < 1024.0:\n            return \"%3.1f %s\" % (size, x)\n        size /= 1024.0\n\n# display CSV file with size\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        csvfile=os.path.join(dirname, filename)\n        csvfilesize = os.path.getsize(csvfile)\n        filesize = convert_bytes(csvfilesize)\n        print(f'{csvfile} size is', filesize, 'bytes')","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:50:42.528794Z","iopub.execute_input":"2022-08-21T04:50:42.529308Z","iopub.status.idle":"2022-08-21T04:50:42.548575Z","shell.execute_reply.started":"2022-08-21T04:50:42.529272Z","shell.execute_reply":"2022-08-21T04:50:42.547220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n<a name = Section4></a>\n# **4. Data Acquisition & Description**\n---\n* **train_data.csv** - training data with multiple statement dates per customer_ID\n* **train_labels.csv** - target label for each customer_ID\n* **test_data.csv** - corresponding test data; objective is to predict the target label for each customer_ID\n* **sample_submission.csv** - a sample submission file in the correct format\n","metadata":{}},{"cell_type":"code","source":"# Importing the dataset\nfrom pathlib import Path\n\ninput_path = Path('/kaggle/input/amex-default-prediction/')","metadata":{"execution":{"iopub.status.busy":"2022-08-20T17:44:42.353369Z","iopub.execute_input":"2022-08-20T17:44:42.353730Z","iopub.status.idle":"2022-08-20T17:44:42.357371Z","shell.execute_reply.started":"2022-08-20T17:44:42.353702Z","shell.execute_reply":"2022-08-20T17:44:42.356771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import dask.dataframe as dd","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:50:48.119722Z","iopub.execute_input":"2022-08-21T04:50:48.120400Z","iopub.status.idle":"2022-08-21T04:50:48.815257Z","shell.execute_reply.started":"2022-08-21T04:50:48.120334Z","shell.execute_reply":"2022-08-21T04:50:48.814125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading dataset train_data.csv\ntrain_df_sample = pd.read_csv('../input/amex-default-prediction/train_data.csv', nrows=100000)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:50:50.385525Z","iopub.execute_input":"2022-08-21T04:50:50.385914Z","iopub.status.idle":"2022-08-21T04:50:57.109428Z","shell.execute_reply.started":"2022-08-21T04:50:50.385882Z","shell.execute_reply":"2022-08-21T04:50:57.108193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get shape of dataframe\nprint('Shape of dataset is:', train_df_sample.shape)\n\n# print summary of dataframe\ntrain_df_sample.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-20T17:44:54.564929Z","iopub.execute_input":"2022-08-20T17:44:54.565893Z","iopub.status.idle":"2022-08-20T17:44:54.595262Z","shell.execute_reply.started":"2022-08-20T17:44:54.565850Z","shell.execute_reply":"2022-08-20T17:44:54.594318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading dataset train_labels.csv\ntrain_label_df = pd.read_csv('../input/amex-default-prediction/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:51:40.942719Z","iopub.execute_input":"2022-08-21T04:51:40.943507Z","iopub.status.idle":"2022-08-21T04:51:41.842936Z","shell.execute_reply.started":"2022-08-21T04:51:40.943468Z","shell.execute_reply":"2022-08-21T04:51:41.841799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get shape of dataframe\nprint('Shape of dataset is:', train_label_df.shape)\n\n# print summary of dataframe\ntrain_label_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:51:45.580694Z","iopub.execute_input":"2022-08-21T04:51:45.581452Z","iopub.status.idle":"2022-08-21T04:51:45.633256Z","shell.execute_reply.started":"2022-08-21T04:51:45.581412Z","shell.execute_reply":"2022-08-21T04:51:45.632082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading dataset test_data.csv\ntest_df = pd.read_csv('../input/amex-default-prediction/test_data.csv', nrows=100000, index_col='customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:51:51.232783Z","iopub.execute_input":"2022-08-21T04:51:51.233213Z","iopub.status.idle":"2022-08-21T04:51:57.594119Z","shell.execute_reply.started":"2022-08-21T04:51:51.233175Z","shell.execute_reply":"2022-08-21T04:51:57.593009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# get shape of dataframe\nprint('Shape of dataset is:', test_df.shape)\n\n# print summary of dataframe\n#test_df.info(verbose=True)\ntest_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:52:01.147891Z","iopub.execute_input":"2022-08-21T04:52:01.148295Z","iopub.status.idle":"2022-08-21T04:52:01.172511Z","shell.execute_reply.started":"2022-08-21T04:52:01.148263Z","shell.execute_reply":"2022-08-21T04:52:01.171000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge of train_df_sample and train_label_df dataframe using key as customer_ID\ntrain_df = dd.merge(train_df_sample,train_label_df,on='customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:52:07.232833Z","iopub.execute_input":"2022-08-21T04:52:07.233244Z","iopub.status.idle":"2022-08-21T04:52:07.878196Z","shell.execute_reply.started":"2022-08-21T04:52:07.233211Z","shell.execute_reply":"2022-08-21T04:52:07.877236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print summary of merged dataframe\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:52:10.709860Z","iopub.execute_input":"2022-08-21T04:52:10.710671Z","iopub.status.idle":"2022-08-21T04:52:10.731240Z","shell.execute_reply.started":"2022-08-21T04:52:10.710627Z","shell.execute_reply":"2022-08-21T04:52:10.729956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# shape of train dataframe\ntrain_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:52:16.809043Z","iopub.execute_input":"2022-08-21T04:52:16.809496Z","iopub.status.idle":"2022-08-21T04:52:16.818724Z","shell.execute_reply.started":"2022-08-21T04:52:16.809459Z","shell.execute_reply":"2022-08-21T04:52:16.817737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Top 5 records of the data frame --> observe NaN values in the data frame\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:52:19.745444Z","iopub.execute_input":"2022-08-21T04:52:19.745861Z","iopub.status.idle":"2022-08-21T04:52:19.780162Z","shell.execute_reply.started":"2022-08-21T04:52:19.745826Z","shell.execute_reply":"2022-08-21T04:52:19.778793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Shape of test dataframe\ntest_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:52:52.940584Z","iopub.execute_input":"2022-08-21T04:52:52.941456Z","iopub.status.idle":"2022-08-21T04:52:52.948648Z","shell.execute_reply.started":"2022-08-21T04:52:52.941415Z","shell.execute_reply":"2022-08-21T04:52:52.947593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Top 5 records of test data frame\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:52:55.642569Z","iopub.execute_input":"2022-08-21T04:52:55.643351Z","iopub.status.idle":"2022-08-21T04:52:55.670050Z","shell.execute_reply.started":"2022-08-21T04:52:55.643299Z","shell.execute_reply":"2022-08-21T04:52:55.669115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note there are NaN values in the data frames","metadata":{}},{"cell_type":"markdown","source":"<a name = Section5></a>\n\n---\n# **5. Data Pre-Processing**\n---","metadata":{}},{"cell_type":"markdown","source":"As per sweetviz report, there are no duplicate rows. Lets check missing values.","metadata":{}},{"cell_type":"code","source":"#Check if there are null/missing values\ntrain_df.isnull()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:53:01.886421Z","iopub.execute_input":"2022-08-21T04:53:01.886844Z","iopub.status.idle":"2022-08-21T04:53:01.957371Z","shell.execute_reply.started":"2022-08-21T04:53:01.886809Z","shell.execute_reply":"2022-08-21T04:53:01.956155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are missing values or Nan values","metadata":{}},{"cell_type":"code","source":"#drop variables with missing values >=70% in the train dataframe\ni=0\nfor col in train_df.columns:\n    if (train_df[col].isnull().sum()/len(train_df[col])*100) >=70:\n        print(\"Column Dropped\", col)\n        train_df.drop(labels=col,axis=1,inplace=True)\n        i=i+1\n        \nprint(\"Total dropped columns are\", i)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:53:16.833007Z","iopub.execute_input":"2022-08-21T04:53:16.834067Z","iopub.status.idle":"2022-08-21T04:53:18.286094Z","shell.execute_reply.started":"2022-08-21T04:53:16.833989Z","shell.execute_reply":"2022-08-21T04:53:18.284910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#drop variables with missing values >=70% in the test dataframe\ni=0\nfor col in test_df.columns:\n    if (test_df[col].isnull().sum()/len(test_df[col])*100) >=70:\n        print(\"Column Dropped\", col)\n        test_df.drop(labels=col,axis=1,inplace=True)\n        i=i+1\n        \nprint(\"Total dropped columns are\", i)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:53:23.360159Z","iopub.execute_input":"2022-08-21T04:53:23.360561Z","iopub.status.idle":"2022-08-21T04:53:24.656815Z","shell.execute_reply.started":"2022-08-21T04:53:23.360526Z","shell.execute_reply":"2022-08-21T04:53:24.655494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Dropping Customer ID and S_2 column in training data\n\ndef drop_features():\n  train_df.drop(columns=['customer_id', 'S_2'], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:53:28.980073Z","iopub.execute_input":"2022-08-21T04:53:28.980522Z","iopub.status.idle":"2022-08-21T04:53:28.985837Z","shell.execute_reply.started":"2022-08-21T04:53:28.980487Z","shell.execute_reply":"2022-08-21T04:53:28.984621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Dropping Customer ID and S_2 column in training data\n\ndef drop_features():\n  test_df.drop(columns=['customer_id', 'S_2'], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:53:33.959433Z","iopub.execute_input":"2022-08-21T04:53:33.960054Z","iopub.status.idle":"2022-08-21T04:53:33.965686Z","shell.execute_reply.started":"2022-08-21T04:53:33.960003Z","shell.execute_reply":"2022-08-21T04:53:33.964709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a name = Section6></a>\n\n---\n# **6. Model Development & Evaluation**\n---","metadata":{}},{"cell_type":"code","source":"#installation\n!pip install dtale","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:53:49.004036Z","iopub.execute_input":"2022-08-21T04:53:49.004479Z","iopub.status.idle":"2022-08-21T04:54:35.100555Z","shell.execute_reply.started":"2022-08-21T04:53:49.004441Z","shell.execute_reply":"2022-08-21T04:54:35.099124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Importing Libraries\nimport dtale\nd = dtale.show(train_df)\nd = dtale.show(test_df)\nd.open_browser()","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:54:52.452936Z","iopub.execute_input":"2022-08-21T04:54:52.453976Z","iopub.status.idle":"2022-08-21T04:55:01.321214Z","shell.execute_reply.started":"2022-08-21T04:54:52.453914Z","shell.execute_reply":"2022-08-21T04:55:01.318918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Installing Library\n!pip install xverse","metadata":{"execution":{"iopub.status.busy":"2022-08-21T04:55:10.757091Z","iopub.execute_input":"2022-08-21T04:55:10.757555Z","iopub.status.idle":"2022-08-21T04:55:24.449043Z","shell.execute_reply.started":"2022-08-21T04:55:10.757517Z","shell.execute_reply":"2022-08-21T04:55:24.447594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"WOE is kind of a feature engineering technique which will take care of imputing categorical vales and null values without the need of apply of feature engineering techniques\nLink for further information: https://www.analyticsvidhya.com/blog/2021/06/understand-weight-of-evidence-and-information-value/","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"#Importing Library and applying WOE\n\nimport xverse  #run if required\n\n#Splitting the data as X and y and applying clf.fit\n\nX = train_df.drop('target',axis=1)\ny = train_df['target']\nfrom xverse.transformer import WOE\nclf = WOE()\nclf.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T05:03:11.896716Z","iopub.execute_input":"2022-08-21T05:03:11.897150Z","iopub.status.idle":"2022-08-21T05:04:45.466116Z","shell.execute_reply.started":"2022-08-21T05:03:11.897116Z","shell.execute_reply":"2022-08-21T05:04:45.464885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Applying Tranformation\nX = clf.transform(X)\nX","metadata":{"execution":{"iopub.status.busy":"2022-08-21T05:05:23.419359Z","iopub.execute_input":"2022-08-21T05:05:23.419783Z","iopub.status.idle":"2022-08-21T05:06:12.228910Z","shell.execute_reply.started":"2022-08-21T05:05:23.419750Z","shell.execute_reply":"2022-08-21T05:06:12.227692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In above data frame, check the null values and categorical features imputed automatically using WOE.","metadata":{}},{"cell_type":"markdown","source":"Cross-validation is a technique for evaluating ML models by training several ML models on subsets of the available input data and evaluating them on the complementary subset of the data. Use cross-validation to detect overfitting, ie, failing to generalize a pattern.\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"Applying the related libraries, WOE and cross-validation to check the Evalutaion metrics for the MODEL.","metadata":{}},{"cell_type":"code","source":"#RANDOM FOREST CLASSIFIER\nimport warnings\nwarnings.filterwarnings('ignore')\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.ensemble import GradientBoostingClassifier,RandomForestClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import confusion_matrix,classification_report\nXtrain, Xtest, ytrain, ytest = train_test_split(X, y, test_size= .20,random_state=20)\nclf = WOE()\nclf.fit(Xtrain, ytrain)\nXtrain = clf.transform(Xtrain)\nXtest = clf.transform(Xtest)\nrf = RandomForestClassifier(n_estimators=100,class_weight='balanced')\nrf.fit(Xtrain,ytrain)\nprint(\"Training Accuracy\")\nprint(rf.score(Xtrain,ytrain))\nprint(\"Testing Accuracy\")\nprint(rf.score(Xtest,ytest))\npredicted = rf.predict(Xtest)\nprint(confusion_matrix(ytest,predicted))\nprint(classification_report(ytest,predicted))\n\nscoresdt = cross_val_score(rf,Xtrain,ytrain,cv=10,scoring='f1')\nprint(scoresdt)\nprint(\"Average f1\")\nprint(np.mean(scoresdt))","metadata":{"execution":{"iopub.status.busy":"2022-08-21T05:07:48.561142Z","iopub.execute_input":"2022-08-21T05:07:48.561625Z","iopub.status.idle":"2022-08-21T05:13:09.366316Z","shell.execute_reply.started":"2022-08-21T05:07:48.561585Z","shell.execute_reply":"2022-08-21T05:13:09.364827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#liner regression prediction\nrf.predict(Xtest)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T05:13:09.370580Z","iopub.execute_input":"2022-08-21T05:13:09.370972Z","iopub.status.idle":"2022-08-21T05:13:10.133793Z","shell.execute_reply.started":"2022-08-21T05:13:09.370936Z","shell.execute_reply":"2022-08-21T05:13:10.132396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#liner regression predict proba\nrf.predict_proba(Xtest)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T05:13:17.736408Z","iopub.execute_input":"2022-08-21T05:13:17.737492Z","iopub.status.idle":"2022-08-21T05:13:18.486256Z","shell.execute_reply.started":"2022-08-21T05:13:17.737450Z","shell.execute_reply":"2022-08-21T05:13:18.485369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Prediction is working!","metadata":{}},{"cell_type":"markdown","source":"<a name = Section7></a>\n\n---\n# **7. Conclusion**\n---","metadata":{}},{"cell_type":"markdown","source":"Average Accuracy of the model is as follows:\n* Random Forest Classifier: 0.894688 (0.004678)\n","metadata":{}},{"cell_type":"markdown","source":"Corresponding Recall value is as follows:\n\n* Random Forest Classifier: 0.777262 (0.009059)","metadata":{}},{"cell_type":"markdown","source":"# **8. Submission****","metadata":{}},{"cell_type":"code","source":"prediction = rf.predict_proba(Xtest)\nfinal_predictions = prediction[:,1]","metadata":{"execution":{"iopub.status.busy":"2022-08-21T05:13:58.641048Z","iopub.execute_input":"2022-08-21T05:13:58.641465Z","iopub.status.idle":"2022-08-21T05:13:59.387534Z","shell.execute_reply.started":"2022-08-21T05:13:58.641430Z","shell.execute_reply":"2022-08-21T05:13:59.386559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output = pd.DataFrame({'customer_ID': Xtest.index,'prediction': ytest})\noutput.to_csv('submission.csv', index=False, header=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-21T05:14:03.557207Z","iopub.execute_input":"2022-08-21T05:14:03.557666Z","iopub.status.idle":"2022-08-21T05:14:03.599032Z","shell.execute_reply.started":"2022-08-21T05:14:03.557620Z","shell.execute_reply":"2022-08-21T05:14:03.597789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Generating submission file!","metadata":{}},{"cell_type":"markdown","source":"**<center><h1> HAPPY LEARNING !</h1></center>**","metadata":{}}]}