{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Client Default Prediction Using Machine Learning\nIn this notebook, the target is to predict if a customer has a chance to default after getting credit card loan from a bank. The target is to make a classification machine learning model for predicting credit card default chance for a given customer.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Importing Packages & Libraries for Computation\nThe first step is to load all libraries that we need for computation. ","metadata":{}},{"cell_type":"code","source":"# importing machine learning libraries\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os, time, re, tqdm, math # utility libraries for computation\nfrom sklearn.model_selection import train_test_split\nimport seaborn as sns\nfrom matplotlib import pyplot as plt\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.svm import SVR\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.metrics import confusion_matrix, classification_report\nfrom sklearn.metrics import plot_confusion_matrix, precision_score, recall_score, f1_score\nfrom sklearn.metrics import precision_recall_curve\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler, MinMaxScaler, LabelEncoder\nfrom sklearn.impute import KNNImputer\nfrom sklearn.svm import SVC","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:47:58.991920Z","iopub.execute_input":"2022-05-29T21:47:58.992373Z","iopub.status.idle":"2022-05-29T21:48:00.619231Z","shell.execute_reply.started":"2022-05-29T21:47:58.992283Z","shell.execute_reply":"2022-05-29T21:48:00.618137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# installing d tale library for performing EDA \n!pip install dtale","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:48:05.654511Z","iopub.execute_input":"2022-05-29T21:48:05.655305Z","iopub.status.idle":"2022-05-29T21:48:45.769815Z","shell.execute_reply.started":"2022-05-29T21:48:05.655261Z","shell.execute_reply":"2022-05-29T21:48:45.768853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# listing the datasets we have\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:03.374159Z","iopub.execute_input":"2022-05-29T21:49:03.377645Z","iopub.status.idle":"2022-05-29T21:49:03.391675Z","shell.execute_reply.started":"2022-05-29T21:49:03.377526Z","shell.execute_reply":"2022-05-29T21:49:03.390710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are given **4** files. One is train file, one is the testing file, one file contains labels for features and the last one is the sample submission file.","metadata":{}},{"cell_type":"markdown","source":"## Loading the Dataset\nFirst of all, we need to load dataset in memory to perform computation. The dataset is pretty much higher, i.e., 18+ GBs in size. For working in a limited computational environment, we'll load a chunk of data for now, and perform computations on that dataset. For now, we'll load 20,000 examples from data. ","metadata":{}},{"cell_type":"code","source":"start = time.time()\n\n# define path variables\ntrain_path = '../input/amex-default-prediction/train_data.csv'\ntrain_labels_path = '../input/amex-default-prediction/train_labels.csv'\n\n# define chunk size\nchunk_size = 20000\n# load dataset with given chunksize\ntrain_data = pd.read_csv(train_path, low_memory=False, chunksize=chunk_size)\ntrain_labels = pd.read_csv(train_labels_path, chunksize=chunk_size)\n\nend = time.time()\nprint('Time Taken: %.3f seconds' % (end-start))","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:08.758033Z","iopub.execute_input":"2022-05-29T21:49:08.759132Z","iopub.status.idle":"2022-05-29T21:49:08.797528Z","shell.execute_reply.started":"2022-05-29T21:49:08.759088Z","shell.execute_reply":"2022-05-29T21:49:08.796263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## EDA (Exploratory Data Analysis)\nThe first step after loading the dataset is the EDA or exploratory data analysis step. In this step, we critically analyse the dataset, draw useful insights, make decisions for modelling and check the dataset for any missing, null, NaN values, and so on. ","metadata":{}},{"cell_type":"code","source":"# convert the IO object into a pandas dataframe\ntrain_data = train_data.__next__()\ntrain_labels = train_labels.__next__()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:12.236572Z","iopub.execute_input":"2022-05-29T21:49:12.237209Z","iopub.status.idle":"2022-05-29T21:49:14.596460Z","shell.execute_reply.started":"2022-05-29T21:49:12.237161Z","shell.execute_reply":"2022-05-29T21:49:14.595524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.head(4)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:49:41.606586Z","iopub.execute_input":"2022-05-29T16:49:41.607326Z","iopub.status.idle":"2022-05-29T16:49:41.617654Z","shell.execute_reply.started":"2022-05-29T16:49:41.607281Z","shell.execute_reply":"2022-05-29T16:49:41.616674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.drop('customer_ID', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:24.133448Z","iopub.execute_input":"2022-05-29T21:49:24.134102Z","iopub.status.idle":"2022-05-29T21:49:24.155588Z","shell.execute_reply.started":"2022-05-29T21:49:24.134063Z","shell.execute_reply":"2022-05-29T21:49:24.154586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:27.357103Z","iopub.execute_input":"2022-05-29T21:49:27.357520Z","iopub.status.idle":"2022-05-29T21:49:27.367602Z","shell.execute_reply.started":"2022-05-29T21:49:27.357486Z","shell.execute_reply":"2022-05-29T21:49:27.366891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels[:4]","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:30.794100Z","iopub.execute_input":"2022-05-29T21:49:30.795311Z","iopub.status.idle":"2022-05-29T21:49:30.817149Z","shell.execute_reply.started":"2022-05-29T21:49:30.795254Z","shell.execute_reply":"2022-05-29T21:49:30.816210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# merging labels with dataset features\ntrain_data = pd.concat([train_data, train_labels], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:53:56.715178Z","iopub.execute_input":"2022-05-29T21:53:56.715623Z","iopub.status.idle":"2022-05-29T21:53:56.738014Z","shell.execute_reply.started":"2022-05-29T21:53:56.715584Z","shell.execute_reply":"2022-05-29T21:53:56.737051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:49:53.987841Z","iopub.execute_input":"2022-05-29T16:49:53.988266Z","iopub.status.idle":"2022-05-29T16:49:53.993040Z","shell.execute_reply.started":"2022-05-29T16:49:53.988230Z","shell.execute_reply":"2022-05-29T16:49:53.992024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check the shape of data\ntrain_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:54:00.876026Z","iopub.execute_input":"2022-05-29T21:54:00.876803Z","iopub.status.idle":"2022-05-29T21:54:00.882065Z","shell.execute_reply.started":"2022-05-29T21:54:00.876768Z","shell.execute_reply":"2022-05-29T21:54:00.881388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After loading **20,000** exmples, the shape of dataset is now having **20,000** training examples & **191** features.","metadata":{}},{"cell_type":"code","source":"train_data.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:54:05.195422Z","iopub.execute_input":"2022-05-29T21:54:05.195992Z","iopub.status.idle":"2022-05-29T21:54:05.221850Z","shell.execute_reply.started":"2022-05-29T21:54:05.195956Z","shell.execute_reply":"2022-05-29T21:54:05.221055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# export the dataframe\ntrain_data.to_csv('train_data_20k.csv')","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:50:50.381646Z","iopub.execute_input":"2022-05-29T16:50:50.382045Z","iopub.status.idle":"2022-05-29T16:50:54.996821Z","shell.execute_reply.started":"2022-05-29T16:50:50.382013Z","shell.execute_reply":"2022-05-29T16:50:54.995562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check the dtype of S_2 feature\ntrain_data['S_2'].dtype","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:40.309776Z","iopub.execute_input":"2022-05-29T21:49:40.310225Z","iopub.status.idle":"2022-05-29T21:49:40.319851Z","shell.execute_reply.started":"2022-05-29T21:49:40.310190Z","shell.execute_reply":"2022-05-29T21:49:40.318784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical = [i for i in train_data.columns if train_data[i].dtype == object]\nnumeric = [i for i in train_data.columns if train_data[i].dtype != object]","metadata":{"execution":{"iopub.status.busy":"2022-05-29T17:11:27.429818Z","iopub.execute_input":"2022-05-29T17:11:27.430823Z","iopub.status.idle":"2022-05-29T17:11:27.444224Z","shell.execute_reply.started":"2022-05-29T17:11:27.430782Z","shell.execute_reply":"2022-05-29T17:11:27.443223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.columns","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:51:25.671541Z","iopub.execute_input":"2022-05-29T16:51:25.671993Z","iopub.status.idle":"2022-05-29T16:51:25.678971Z","shell.execute_reply.started":"2022-05-29T16:51:25.671953Z","shell.execute_reply":"2022-05-29T16:51:25.678267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(categoricl), len(numeric)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:51:28.123236Z","iopub.execute_input":"2022-05-29T16:51:28.123688Z","iopub.status.idle":"2022-05-29T16:51:28.129370Z","shell.execute_reply.started":"2022-05-29T16:51:28.123643Z","shell.execute_reply":"2022-05-29T16:51:28.128717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We mostly have numerical attributes in our dataset, since the data is a real-world dataset. ","metadata":{}},{"cell_type":"markdown","source":"### Preprocessing & Featuring Datatime Attribute ","metadata":{}},{"cell_type":"code","source":"train_data.rename(columns={'S_2': 'Date'}, \n                  inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:50.529955Z","iopub.execute_input":"2022-05-29T21:49:50.531191Z","iopub.status.idle":"2022-05-29T21:49:50.537400Z","shell.execute_reply.started":"2022-05-29T21:49:50.531134Z","shell.execute_reply":"2022-05-29T21:49:50.536582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check the data again\ntrain_data.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:49:53.425760Z","iopub.execute_input":"2022-05-29T21:49:53.426819Z","iopub.status.idle":"2022-05-29T21:49:53.455292Z","shell.execute_reply.started":"2022-05-29T21:49:53.426770Z","shell.execute_reply":"2022-05-29T21:49:53.454287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# convert the \ntrain_data['Date'] = pd.to_datetime(train_data['Date'], \n                                         infer_datetime_format=True, format='%Y/%m/%d %H:%M:%S')","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:50:21.609616Z","iopub.execute_input":"2022-05-29T21:50:21.610057Z","iopub.status.idle":"2022-05-29T21:50:21.624585Z","shell.execute_reply.started":"2022-05-29T21:50:21.610021Z","shell.execute_reply":"2022-05-29T21:50:21.623639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head(4)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:50:28.634984Z","iopub.execute_input":"2022-05-29T21:50:28.636003Z","iopub.status.idle":"2022-05-29T21:50:28.662800Z","shell.execute_reply.started":"2022-05-29T21:50:28.635965Z","shell.execute_reply":"2022-05-29T21:50:28.661965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['Date'].dtype","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:50:32.186381Z","iopub.execute_input":"2022-05-29T21:50:32.186850Z","iopub.status.idle":"2022-05-29T21:50:32.194234Z","shell.execute_reply.started":"2022-05-29T21:50:32.186811Z","shell.execute_reply":"2022-05-29T21:50:32.193101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Analyzing a Particular Customer's ID\nFor now, we'll analyze the transaction history of a particular customer. We'll fetch the data from dataset of a partocular customer & then apply computations on that data frame","metadata":{}},{"cell_type":"code","source":"train_data['customer_ID'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:51:40.530361Z","iopub.execute_input":"2022-05-29T16:51:40.531201Z","iopub.status.idle":"2022-05-29T16:51:40.544046Z","shell.execute_reply.started":"2022-05-29T16:51:40.531150Z","shell.execute_reply":"2022-05-29T16:51:40.543084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_customer_id = '0089bf123391cdddcdc34a8ea239de5188c8fc7e4a5974e0cd3c2461c8d3dc0b'","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:51:43.813018Z","iopub.execute_input":"2022-05-29T16:51:43.813875Z","iopub.status.idle":"2022-05-29T16:51:43.818478Z","shell.execute_reply.started":"2022-05-29T16:51:43.813823Z","shell.execute_reply":"2022-05-29T16:51:43.817821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# fetch the data from dataset\nrandom_customer_data = train_data[train_data['customer_ID'] == random_customer_id]","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:51:45.039366Z","iopub.execute_input":"2022-05-29T16:51:45.039786Z","iopub.status.idle":"2022-05-29T16:51:45.070081Z","shell.execute_reply.started":"2022-05-29T16:51:45.039747Z","shell.execute_reply":"2022-05-29T16:51:45.068933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_customer_data.head(4)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:51:47.181746Z","iopub.execute_input":"2022-05-29T16:51:47.182148Z","iopub.status.idle":"2022-05-29T16:51:47.309797Z","shell.execute_reply.started":"2022-05-29T16:51:47.182115Z","shell.execute_reply":"2022-05-29T16:51:47.308650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_customer_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:51:49.894130Z","iopub.execute_input":"2022-05-29T16:51:49.894526Z","iopub.status.idle":"2022-05-29T16:51:49.904316Z","shell.execute_reply.started":"2022-05-29T16:51:49.894494Z","shell.execute_reply":"2022-05-29T16:51:49.903322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set(rc={'figure.figsize':(11.7,8.27)})\n\nsns.lineplot(x='Date', y='P_2', \n             data=train_data)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:55:38.539222Z","iopub.execute_input":"2022-05-29T16:55:38.540299Z","iopub.status.idle":"2022-05-29T16:55:46.776535Z","shell.execute_reply.started":"2022-05-29T16:55:38.540252Z","shell.execute_reply":"2022-05-29T16:55:46.775646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_customer_data['Date'].min(), random_customer_data['Date'].max()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:56:27.712808Z","iopub.execute_input":"2022-05-29T16:56:27.713994Z","iopub.status.idle":"2022-05-29T16:56:27.721901Z","shell.execute_reply.started":"2022-05-29T16:56:27.713946Z","shell.execute_reply":"2022-05-29T16:56:27.720691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.lineplot(x='Date', y='P_2', \n             data=random_customer_data)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T16:59:19.408693Z","iopub.execute_input":"2022-05-29T16:59:19.409159Z","iopub.status.idle":"2022-05-29T16:59:19.711324Z","shell.execute_reply.started":"2022-05-29T16:59:19.409123Z","shell.execute_reply":"2022-05-29T16:59:19.710140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# viewing the plot of risk factor of a random customer\nsns.lineplot(x='Date', y='R_2', \n             data=random_customer_data)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T17:03:21.513911Z","iopub.execute_input":"2022-05-29T17:03:21.515111Z","iopub.status.idle":"2022-05-29T17:03:21.850718Z","shell.execute_reply.started":"2022-05-29T17:03:21.515066Z","shell.execute_reply":"2022-05-29T17:03:21.849746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.scatter(x='Date', y='P_2', \n             data=random_customer_data)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T17:08:07.464113Z","iopub.execute_input":"2022-05-29T17:08:07.464528Z","iopub.status.idle":"2022-05-29T17:08:07.750896Z","shell.execute_reply.started":"2022-05-29T17:08:07.464495Z","shell.execute_reply":"2022-05-29T17:08:07.749895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(x=train_data['target']).set_title('Class distribution of Taregt Feature')","metadata":{"execution":{"iopub.status.busy":"2022-05-29T17:10:48.122524Z","iopub.execute_input":"2022-05-29T17:10:48.123648Z","iopub.status.idle":"2022-05-29T17:10:48.318567Z","shell.execute_reply.started":"2022-05-29T17:10:48.123596Z","shell.execute_reply":"2022-05-29T17:10:48.317425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical","metadata":{"execution":{"iopub.status.busy":"2022-05-29T17:11:32.011929Z","iopub.execute_input":"2022-05-29T17:11:32.012335Z","iopub.status.idle":"2022-05-29T17:11:32.018374Z","shell.execute_reply.started":"2022-05-29T17:11:32.012301Z","shell.execute_reply":"2022-05-29T17:11:32.017634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# checking the distribution of target column\nplt.figure(figsize=(10, 8))\ncircle = plt.Circle((0, 0), 0.7, color='white')\nplt.pie(train_data['target'].value_counts(), labels=['No Default', 'Default'], colors=['green', 'red' \n                                                                                   ])\np = plt.gcf()\np.gca().add_artist(circle)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:54:14.570590Z","iopub.execute_input":"2022-05-29T21:54:14.571384Z","iopub.status.idle":"2022-05-29T21:54:14.755967Z","shell.execute_reply.started":"2022-05-29T21:54:14.571347Z","shell.execute_reply":"2022-05-29T21:54:14.754926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"index = 0\nfor column in random_customer_data.columns:\n    if column in [\"S_2\", \"customer_ID\", \"target\"] + categorical:\n        continue\n    \n    if index % 4 == 0:\n        plt.figure(figsize=(16, 4))\n    plt.subplot(1, 4, index % 4 + 1)\n    \n    sns.histplot(data=random_customer_data, x=column, hue=\"target\", bins=20)\n    plt.ylabel(\"\")\n    \n    if index % 4 == 3:\n        plt.show()\n    \n    index += 1","metadata":{"execution":{"iopub.status.busy":"2022-05-29T17:14:47.358169Z","iopub.execute_input":"2022-05-29T17:14:47.358587Z","iopub.status.idle":"2022-05-29T17:15:36.927844Z","shell.execute_reply.started":"2022-05-29T17:14:47.358539Z","shell.execute_reply":"2022-05-29T17:15:36.926745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_col = [\n    \"B_2\", \"B_7\", \"B_18\", \"B_23\", \"B_32\", \"D_48\",\n    \"D_55\", \"D_61\", \"D_121\", \"P_2\", \"S_11\",\n    \n]","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:03:50.555165Z","iopub.execute_input":"2022-05-29T22:03:50.555622Z","iopub.status.idle":"2022-05-29T22:03:50.561238Z","shell.execute_reply.started":"2022-05-29T22:03:50.555580Z","shell.execute_reply":"2022-05-29T22:03:50.560224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# define chunk size again for macking machine learning model\nchunk_size = 2000000\n\n# load dataset with given chunksize\ndata = pd.read_csv(train_path, low_memory=False, chunksize=chunk_size, usecols=['customer_ID'] + X_col)\nlabels = pd.read_csv(train_labels_path, chunksize=200000)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:11:09.312494Z","iopub.execute_input":"2022-05-29T22:11:09.313601Z","iopub.status.idle":"2022-05-29T22:11:09.327663Z","shell.execute_reply.started":"2022-05-29T22:11:09.313560Z","shell.execute_reply":"2022-05-29T22:11:09.326930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# convert to dataframe\ndata = data.__next__()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:11:12.104627Z","iopub.execute_input":"2022-05-29T22:11:12.105596Z","iopub.status.idle":"2022-05-29T22:12:43.613180Z","shell.execute_reply.started":"2022-05-29T22:11:12.105555Z","shell.execute_reply":"2022-05-29T22:12:43.612188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlabels = labels.__next__()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:13:00.790721Z","iopub.execute_input":"2022-05-29T22:13:00.791104Z","iopub.status.idle":"2022-05-29T22:13:01.150498Z","shell.execute_reply.started":"2022-05-29T22:13:00.791075Z","shell.execute_reply":"2022-05-29T22:13:01.149248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_mean = data.groupby(\"customer_ID\")[X_col].mean().reset_index()\ndata_last = data.groupby(\"customer_ID\")[X_col].last().reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:13:07.944114Z","iopub.execute_input":"2022-05-29T22:13:07.944504Z","iopub.status.idle":"2022-05-29T22:13:09.334900Z","shell.execute_reply.started":"2022-05-29T22:13:07.944473Z","shell.execute_reply":"2022-05-29T22:13:09.333918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# labels.drop('customer_ID', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:10:12.517487Z","iopub.execute_input":"2022-05-29T22:10:12.518070Z","iopub.status.idle":"2022-05-29T22:10:12.533834Z","shell.execute_reply.started":"2022-05-29T22:10:12.518037Z","shell.execute_reply":"2022-05-29T22:10:12.532776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# merging the dataset by customer ID\nnew_data = pd.merge(\n    left=data_mean, \n    right=data_last, \n    how=\"inner\",\n    on=\"customer_ID\",\n    suffixes=(\"_mean\", \"_last\"),\n)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:13:15.245892Z","iopub.execute_input":"2022-05-29T22:13:15.246279Z","iopub.status.idle":"2022-05-29T22:13:15.422167Z","shell.execute_reply.started":"2022-05-29T22:13:15.246249Z","shell.execute_reply":"2022-05-29T22:13:15.421201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:13:17.830961Z","iopub.execute_input":"2022-05-29T22:13:17.831602Z","iopub.status.idle":"2022-05-29T22:13:17.838154Z","shell.execute_reply.started":"2022-05-29T22:13:17.831572Z","shell.execute_reply":"2022-05-29T22:13:17.837339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:13:22.815905Z","iopub.execute_input":"2022-05-29T22:13:22.816305Z","iopub.status.idle":"2022-05-29T22:13:22.822405Z","shell.execute_reply.started":"2022-05-29T22:13:22.816276Z","shell.execute_reply":"2022-05-29T22:13:22.821396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data = pd.merge(new_data, \n                    labels, \n                    on=\"customer_ID\", \n                    how=\"left\")","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:13:28.930895Z","iopub.execute_input":"2022-05-29T22:13:28.931672Z","iopub.status.idle":"2022-05-29T22:13:29.131651Z","shell.execute_reply.started":"2022-05-29T22:13:28.931637Z","shell.execute_reply":"2022-05-29T22:13:29.130629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:13:38.811918Z","iopub.execute_input":"2022-05-29T22:13:38.812374Z","iopub.status.idle":"2022-05-29T22:13:38.818646Z","shell.execute_reply.started":"2022-05-29T22:13:38.812340Z","shell.execute_reply":"2022-05-29T22:13:38.817896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.head(4)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:13:42.194637Z","iopub.execute_input":"2022-05-29T22:13:42.195266Z","iopub.status.idle":"2022-05-29T22:13:42.228462Z","shell.execute_reply.started":"2022-05-29T22:13:42.195215Z","shell.execute_reply":"2022-05-29T22:13:42.227532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.columns","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:14:26.290323Z","iopub.execute_input":"2022-05-29T22:14:26.290994Z","iopub.status.idle":"2022-05-29T22:14:26.298195Z","shell.execute_reply.started":"2022-05-29T22:14:26.290947Z","shell.execute_reply":"2022-05-29T22:14:26.297283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# converting datetime column\n'''\nnew_data['S_2'] = pd.to_datetime(new_data['S_2'], \n                                         infer_datetime_format=True, format='%Y/%m/%d %H:%M:%S')\n                                         '''","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:14:41.996924Z","iopub.execute_input":"2022-05-29T22:14:41.997331Z","iopub.status.idle":"2022-05-29T22:14:42.006740Z","shell.execute_reply.started":"2022-05-29T22:14:41.997297Z","shell.execute_reply":"2022-05-29T22:14:42.006081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# new_data['S_2'].dtype","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:14:50.016041Z","iopub.execute_input":"2022-05-29T22:14:50.016443Z","iopub.status.idle":"2022-05-29T22:14:50.020898Z","shell.execute_reply.started":"2022-05-29T22:14:50.016413Z","shell.execute_reply":"2022-05-29T22:14:50.019995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# creating additional features\n'''\nnew_data['Year of Transaction'] = new_data['S_2'].dt.year\nnew_data['Month of Transaction'] = new_data['S_2'].dt.month\nnew_data['Day of Transaction'] = new_data['S_2'].dt.day\n'''","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:14:58.289241Z","iopub.execute_input":"2022-05-29T22:14:58.290458Z","iopub.status.idle":"2022-05-29T22:14:58.296008Z","shell.execute_reply.started":"2022-05-29T22:14:58.290416Z","shell.execute_reply":"2022-05-29T22:14:58.294935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# new_data.drop('S_2', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:57:02.186981Z","iopub.execute_input":"2022-05-29T21:57:02.188127Z","iopub.status.idle":"2022-05-29T21:57:02.522330Z","shell.execute_reply.started":"2022-05-29T21:57:02.188080Z","shell.execute_reply":"2022-05-29T21:57:02.521214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# new_data.head(4)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:15:05.750178Z","iopub.execute_input":"2022-05-29T22:15:05.750570Z","iopub.status.idle":"2022-05-29T22:15:05.755311Z","shell.execute_reply.started":"2022-05-29T22:15:05.750541Z","shell.execute_reply":"2022-05-29T22:15:05.754233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols = [i for i in new_data.columns if new_data[i].dtype==object]\nnum_cols = [i for i in new_data.columns if new_data[i].dtype!=object]","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:15:10.338419Z","iopub.execute_input":"2022-05-29T22:15:10.339230Z","iopub.status.idle":"2022-05-29T22:15:10.346803Z","shell.execute_reply.started":"2022-05-29T22:15:10.339187Z","shell.execute_reply":"2022-05-29T22:15:10.345816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols\nprint(len(num_cols))","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:15:15.950651Z","iopub.execute_input":"2022-05-29T22:15:15.951416Z","iopub.status.idle":"2022-05-29T22:15:15.956814Z","shell.execute_reply.started":"2022-05-29T22:15:15.951376Z","shell.execute_reply":"2022-05-29T22:15:15.955881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:15:19.057586Z","iopub.execute_input":"2022-05-29T22:15:19.058555Z","iopub.status.idle":"2022-05-29T22:15:19.065479Z","shell.execute_reply.started":"2022-05-29T22:15:19.058513Z","shell.execute_reply":"2022-05-29T22:15:19.064346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data['D_63'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:57:16.890494Z","iopub.execute_input":"2022-05-29T21:57:16.891330Z","iopub.status.idle":"2022-05-29T21:57:16.922875Z","shell.execute_reply.started":"2022-05-29T21:57:16.891269Z","shell.execute_reply":"2022-05-29T21:57:16.922186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.isnull().sum().to_numpy()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:15:28.469963Z","iopub.execute_input":"2022-05-29T22:15:28.470548Z","iopub.status.idle":"2022-05-29T22:15:28.509275Z","shell.execute_reply.started":"2022-05-29T22:15:28.470498Z","shell.execute_reply":"2022-05-29T22:15:28.508317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Defining Imputation Functions For Dealing with Missing, NaN & Null Values","metadata":{}},{"cell_type":"code","source":"new_data.sample()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# making an imputation function\ndef random_imputation(x):\n    random_sample = new_data[x].dropna().sample(new_data[x].isna().sum(), replace=True)\n    random_sample.index = new_data[new_data[x].isnull()].index\n    new_data.loc[new_data[x].isnull(), x] = random_sample\n\n# define imputation mode    \ndef imputation_mode(x):\n    mode = new_data[x].mode()[0]\n    new_data[x] = new_data[x].fillna(mode)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:15:43.420813Z","iopub.execute_input":"2022-05-29T22:15:43.421236Z","iopub.status.idle":"2022-05-29T22:15:43.428253Z","shell.execute_reply.started":"2022-05-29T22:15:43.421201Z","shell.execute_reply":"2022-05-29T22:15:43.427173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# apply function to columns\nfor c in num_cols:\n    random_imputation(c)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:15:47.647040Z","iopub.execute_input":"2022-05-29T22:15:47.647972Z","iopub.status.idle":"2022-05-29T22:15:47.780916Z","shell.execute_reply.started":"2022-05-29T22:15:47.647929Z","shell.execute_reply":"2022-05-29T22:15:47.780228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.isnull().sum().to_numpy()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:15:59.269584Z","iopub.execute_input":"2022-05-29T22:15:59.270381Z","iopub.status.idle":"2022-05-29T22:15:59.305901Z","shell.execute_reply.started":"2022-05-29T22:15:59.270345Z","shell.execute_reply":"2022-05-29T22:15:59.304922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T17:36:33.372512Z","iopub.execute_input":"2022-05-29T17:36:33.373360Z","iopub.status.idle":"2022-05-29T17:36:33.378948Z","shell.execute_reply.started":"2022-05-29T17:36:33.373311Z","shell.execute_reply":"2022-05-29T17:36:33.378159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check for null or missing values\nnew_data[cat_cols].isna().sum().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:01.439170Z","iopub.execute_input":"2022-05-29T21:58:01.439588Z","iopub.status.idle":"2022-05-29T21:58:01.518950Z","shell.execute_reply.started":"2022-05-29T21:58:01.439553Z","shell.execute_reply":"2022-05-29T21:58:01.517755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_imputation('D_64')\n\nfor c in cat_cols:\n    imputation_mode(c)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:04.890928Z","iopub.execute_input":"2022-05-29T21:58:04.891610Z","iopub.status.idle":"2022-05-29T21:58:05.159841Z","shell.execute_reply.started":"2022-05-29T21:58:04.891571Z","shell.execute_reply":"2022-05-29T21:58:05.158827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data[cat_cols].isna().sum().sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:10.854900Z","iopub.execute_input":"2022-05-29T21:58:10.855347Z","iopub.status.idle":"2022-05-29T21:58:10.934204Z","shell.execute_reply.started":"2022-05-29T21:58:10.855310Z","shell.execute_reply":"2022-05-29T21:58:10.933202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.isnull().sum().to_numpy()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T17:38:20.885447Z","iopub.execute_input":"2022-05-29T17:38:20.885888Z","iopub.status.idle":"2022-05-29T17:38:21.037593Z","shell.execute_reply.started":"2022-05-29T17:38:20.885850Z","shell.execute_reply":"2022-05-29T17:38:21.036184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:15.173518Z","iopub.execute_input":"2022-05-29T21:58:15.173935Z","iopub.status.idle":"2022-05-29T21:58:15.179885Z","shell.execute_reply.started":"2022-05-29T21:58:15.173903Z","shell.execute_reply":"2022-05-29T21:58:15.178900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Converting Categorical Columns to Integers","metadata":{}},{"cell_type":"code","source":"D_63 = new_data[['D_63']]\nD_63 = pd.get_dummies(D_63)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:18.086076Z","iopub.execute_input":"2022-05-29T21:58:18.086450Z","iopub.status.idle":"2022-05-29T21:58:18.122614Z","shell.execute_reply.started":"2022-05-29T21:58:18.086422Z","shell.execute_reply":"2022-05-29T21:58:18.121885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"D_64 = new_data[['D_64']]\nD_64 = pd.get_dummies(D_64)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:20.445845Z","iopub.execute_input":"2022-05-29T21:58:20.446475Z","iopub.status.idle":"2022-05-29T21:58:20.476454Z","shell.execute_reply.started":"2022-05-29T21:58:20.446437Z","shell.execute_reply":"2022-05-29T21:58:20.475375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.drop(['D_63', 'D_64', 'customer_ID'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:22.728875Z","iopub.execute_input":"2022-05-29T21:58:22.729284Z","iopub.status.idle":"2022-05-29T21:58:22.828332Z","shell.execute_reply.started":"2022-05-29T21:58:22.729250Z","shell.execute_reply":"2022-05-29T21:58:22.827342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_data = pd.concat([new_data, D_63, D_64], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:24.561970Z","iopub.execute_input":"2022-05-29T21:58:24.563089Z","iopub.status.idle":"2022-05-29T21:58:24.689297Z","shell.execute_reply.started":"2022-05-29T21:58:24.563049Z","shell.execute_reply":"2022-05-29T21:58:24.688222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_data.head(4)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:27.557734Z","iopub.execute_input":"2022-05-29T21:58:27.558143Z","iopub.status.idle":"2022-05-29T21:58:27.583015Z","shell.execute_reply.started":"2022-05-29T21:58:27.558111Z","shell.execute_reply":"2022-05-29T21:58:27.581930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:58:51.430382Z","iopub.execute_input":"2022-05-29T21:58:51.431213Z","iopub.status.idle":"2022-05-29T21:58:51.437418Z","shell.execute_reply.started":"2022-05-29T21:58:51.431166Z","shell.execute_reply":"2022-05-29T21:58:51.436500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_data.target","metadata":{"execution":{"iopub.status.busy":"2022-05-29T21:59:13.517496Z","iopub.execute_input":"2022-05-29T21:59:13.517982Z","iopub.status.idle":"2022-05-29T21:59:13.526832Z","shell.execute_reply.started":"2022-05-29T21:59:13.517942Z","shell.execute_reply":"2022-05-29T21:59:13.525733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# X = final_data.drop('target', axis=1)\n\nX = new_data.drop(['customer_ID', 'target'], axis=1)\nY = new_data['target']","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:18:02.780724Z","iopub.execute_input":"2022-05-29T22:18:02.781644Z","iopub.status.idle":"2022-05-29T22:18:02.798357Z","shell.execute_reply.started":"2022-05-29T22:18:02.781603Z","shell.execute_reply":"2022-05-29T22:18:02.797571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.shape, Y.shape","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:18:05.161912Z","iopub.execute_input":"2022-05-29T22:18:05.162478Z","iopub.status.idle":"2022-05-29T22:18:05.169293Z","shell.execute_reply.started":"2022-05-29T22:18:05.162447Z","shell.execute_reply":"2022-05-29T22:18:05.168387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_data.columns[1:-1]","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:16:59.652100Z","iopub.execute_input":"2022-05-29T22:16:59.652641Z","iopub.status.idle":"2022-05-29T22:16:59.658709Z","shell.execute_reply.started":"2022-05-29T22:16:59.652603Z","shell.execute_reply":"2022-05-29T22:16:59.657754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Splicing (Splitting into Train & Test Sets)","metadata":{}},{"cell_type":"code","source":"X_train, X_val, Y_train, Y_val = train_test_split(X, \n                                                  Y, \n                                                  test_size=0.2, \n                                                  random_state=123)\n\nprint(X_train.shape)\nprint(X_val.shape)\nprint(Y_train.shape)\nprint(Y_val.shape)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:18:16.087844Z","iopub.execute_input":"2022-05-29T22:18:16.088751Z","iopub.status.idle":"2022-05-29T22:18:16.137799Z","shell.execute_reply.started":"2022-05-29T22:18:16.088710Z","shell.execute_reply":"2022-05-29T22:18:16.136837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Applying Machine Learning Models","metadata":{}},{"cell_type":"code","source":"%%time\nrf = RandomForestClassifier()\nrf.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:18:18.837953Z","iopub.execute_input":"2022-05-29T22:18:18.838453Z","iopub.status.idle":"2022-05-29T22:19:56.683144Z","shell.execute_reply.started":"2022-05-29T22:18:18.838382Z","shell.execute_reply":"2022-05-29T22:19:56.682025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf.get_params()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:20:35.715761Z","iopub.execute_input":"2022-05-29T22:20:35.716203Z","iopub.status.idle":"2022-05-29T22:20:35.723837Z","shell.execute_reply.started":"2022-05-29T22:20:35.716170Z","shell.execute_reply":"2022-05-29T22:20:35.722922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Accuracy of Random Forest on training data: %.3f' % rf.score(X_train, Y_train))\nprint('Accuracy of Random Forest on validation data: %.3f' % rf.score(X_val, Y_val))","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:22:41.083322Z","iopub.execute_input":"2022-05-29T22:22:41.084316Z","iopub.status.idle":"2022-05-29T22:22:45.770959Z","shell.execute_reply.started":"2022-05-29T22:22:41.084263Z","shell.execute_reply":"2022-05-29T22:22:45.770013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 5))\nplt.title('Feature Importance of Random Forest')\nimp = pd.Series(rf.feature_importances_, index=X.columns)\nimp.nlargest(20).plot(kind='barh')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:23:20.991813Z","iopub.execute_input":"2022-05-29T22:23:20.992232Z","iopub.status.idle":"2022-05-29T22:23:21.356066Z","shell.execute_reply.started":"2022-05-29T22:23:20.992203Z","shell.execute_reply":"2022-05-29T22:23:21.354943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have obtained an accuracy of 99% on training data but got 74% on testing data. We need to boost the accuracy. The work is in progress!","metadata":{}},{"cell_type":"code","source":"def get_scores(clf):\n    model = clf.fit(X_train, Y_train)\n    y_pred = clf.predict(X_test)\n    \n    print('============================================================================')\n    print('Classification Results of Classification Model After Training')\n    print('============================================================================')\n    print('')\n    \n    print('Accuracy of Classifier on training dataset: %.3f' % rf.score(X_train, Y_train))\n    print('Accuracy of Classifier on test dataset: %.2f' % rf.score(X_test, Y_test))\n    print('Precision of classifier: %.3f' % precision_score(Y_test, y_pred, average='weighted'))\n    print('Recall of classifer: %.3f' % recall_score(Y_test, y_pred, average='weighted'))\n    print('F1 score of classifer: %.3f' % f1_score(Y_test, y_pred, average='weighted'))","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:24:25.633777Z","iopub.execute_input":"2022-05-29T22:24:25.634276Z","iopub.status.idle":"2022-05-29T22:24:25.643073Z","shell.execute_reply.started":"2022-05-29T22:24:25.634240Z","shell.execute_reply":"2022-05-29T22:24:25.641710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# getting the scores of random forest\nget_scores(rf)","metadata":{"execution":{"iopub.status.busy":"2022-05-29T22:24:43.954750Z","iopub.execute_input":"2022-05-29T22:24:43.955222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}