{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# American Express Default Prediction","metadata":{}},{"cell_type":"markdown","source":"### 1. Competition Objective","metadata":{}},{"cell_type":"markdown","source":"American Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nThe objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. In this competition, you'll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.","metadata":{}},{"cell_type":"markdown","source":"### 2. Data Overview","metadata":{"execution":{"iopub.status.busy":"2022-07-26T05:04:12.743866Z","iopub.execute_input":"2022-07-26T05:04:12.744340Z","iopub.status.idle":"2022-07-26T05:04:12.750284Z","shell.execute_reply.started":"2022-07-26T05:04:12.744306Z","shell.execute_reply":"2022-07-26T05:04:12.748919Z"}}},{"cell_type":"markdown","source":"The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThere are a total of 190 variables in the dataset with approximately 450,000 customers in the training set and 925,000 in the test set. Due to the dataset size, I will use the compressed version of the train and test sets provided by @munumbutt's AMEX-Feather-Dataset and take the last statement for each customer.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\nD_*: Delinquency variables\nS_*: Spend variables\nP_*: Payment variables\nB_*: Balance variables\nR_*: Risk variables\nWith the following features being categorical: B_30, B_38, D_63, D_64, D_66, D_68, D_114, D_116, D_117, D_120, D_126.\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-26T05:02:05.960934Z","iopub.execute_input":"2022-07-26T05:02:05.961473Z","iopub.status.idle":"2022-07-26T05:02:06.001854Z","shell.execute_reply.started":"2022-07-26T05:02:05.961374Z","shell.execute_reply":"2022-07-26T05:02:06.000222Z"}}},{"cell_type":"markdown","source":"In the training set, the last statement of all customers was in March 2018, while in the test set the date of customers' last statements range from April through October 2019.","metadata":{}},{"cell_type":"markdown","source":"### 3. Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import train_test_split \n#\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This library is used to read feather file and process large data\n#pip install pyarrow","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Location of data files\n\ntrain = pd.read_feather('/kaggle/input/amexfeather/train_data.ftr')\ntest = pd.read_feather('/kaggle/input/amexfeather/test_data.ftr')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Train Dataset Shape: \", train.shape)\nprint(\"Test Dataset Shape: \", test.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n#Grouping Data based on Last values\n\ntemp_test = pd.DataFrame(test.groupby(['customer_ID']).tail(1))\ntemp_test.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n#Grouping Data based on Last values\n\ntemp_train_last = pd.DataFrame(train.groupby(['customer_ID']).tail(1))\ntemp_train_last.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First thing first, reduce datasize to get good inference out of the dataset on available compute resources.\n\nDroping Missing value columns (Missing values > 75%) as those columns will not add any value to EDA.\n\nFor other missing value columns(Missing values < 75%) we will impute values in later steps 4.1 and 4.2.","metadata":{}},{"cell_type":"code","source":"#Drop columns with missing values morethan 75%\ntemp_train_last.dropna(axis=1, thresh=int(0.75*len(temp_train_last)),inplace=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_train_last.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Checking the distribution of Target throught the train data ","metadata":{}},{"cell_type":"code","source":"temp_train_last['target'].value_counts().plot(kind='pie', title=\"Title\", legend=True, \\\n                   autopct='%1.1f%%', explode=(0, 0), \\\n                   shadow=True, startangle=0)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Checking categorical column's data distribution","metadata":{}},{"cell_type":"code","source":"cols = []\ncols = temp_train_last.columns\ncols = cols.sort_values()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in list(temp_train_last.select_dtypes(['category']).columns):\n    fig, ax = plt.subplots(figsize=(5,3))\n    bar = temp_train_last[col].value_counts().plot(kind='bar', title=\"Distribution of \" + col + '\\n')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As the data lables are masked we cannot come to a conclusion about it. ","metadata":{}},{"cell_type":"markdown","source":"There is more hidden information in the data, @girishkumarsahu has done extraordinary work on this.\n\nCheck out his great work on EDA and observations here https://www.kaggle.com/code/girishkumarsahu/american-express-default-prediction-eda","metadata":{}},{"cell_type":"markdown","source":"### 4. Data Preprocessing","metadata":{}},{"cell_type":"markdown","source":"Handling missing data using different approaches for categorical and continuous variables. ","metadata":{}},{"cell_type":"markdown","source":"**4.1 Handling Missing Values - Categorical Values**","metadata":{}},{"cell_type":"markdown","source":"There are categories with String data in it, Lets use label encoder to process those in numerical data","metadata":{}},{"cell_type":"code","source":"from sklearn import preprocessing\n  \n# label_encoder object knows how to understand word labels.\nlabel_encoder = preprocessing.LabelEncoder()\n  \n# Encode labels in columns 'D_63' and 'D_64'.\ntemp_train_last['D_63']= label_encoder.fit_transform(temp_train_last['D_63'])\ntemp_train_last['D_64']= label_encoder.fit_transform(temp_train_last['D_64'])  \n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in list(temp_train_last.select_dtypes(['category']).columns):\n    #print(col,temp_train_last[col].mode()[0])\n    temp_train_last[col].fillna(temp_train_last[col].mode()[0], inplace=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As Feather format reduced the data size, there is an issue of handling them with NumPy, it could not recognize the values of float16, so converting everything to float64.\n\nPlease note that this will increase the size of the dataset.","metadata":{}},{"cell_type":"code","source":"for col in list(temp_train_last.select_dtypes(['category']).columns):\n    temp_train_last[col] = temp_train_last[col].astype('Int64')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**4.2 Handling Missing Values - Missing Continuous Values**","metadata":{}},{"cell_type":"code","source":"for col in list(temp_train_last.select_dtypes(['float16','int32','int64']).columns):\n\n    temp_train_last[col].fillna(np.mean(~temp_train_last[col].isnull()), inplace=True)\n    temp_train_last[col] = temp_train_last[col].astype(np.float64)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_train_last.reset_index(inplace=True, drop=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_train_last.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finalizing the dataset after droping unnecessary columns for Machine Learning.","metadata":{}},{"cell_type":"code","source":"afterdroptrain = temp_train_last.drop(labels=['customer_ID','S_2'],axis=1)\nafterdroptrain.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"afterdroptrain.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At this stage we only have numerical values in the dataset and we do not have any missing values now.\n\nThe dataset is now ready for further processing and model building. ","metadata":{}},{"cell_type":"markdown","source":"Some Memory management tasks...","metadata":{}},{"cell_type":"code","source":"del train\ndel temp_train_last","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5. Model Building","metadata":{}},{"cell_type":"markdown","source":"We have below Amex's metric to calculate the score for final submission of the predictions","metadata":{}},{"cell_type":"code","source":"def amex_metric_mod(y_true, y_pred):\n\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Splitting the train data to process for feature importance","metadata":{}},{"cell_type":"code","source":"X = afterdroptrain.drop(labels=['target'], axis=1)\ny = afterdroptrain['target']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"X.shape: \" , X.shape, \"\\ny.shape: \" , y.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As Data is imbalanced have to use SMOTE to make data Balanced.","metadata":{}},{"cell_type":"markdown","source":"**5.1 SMOTE**","metadata":{}},{"cell_type":"code","source":"# check version number\nimport imblearn\nfrom imblearn.over_sampling import SMOTE\nprint(imblearn.__version__)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n##approx Runtime 6mins on Kaggle\n##approx Runtime 30 seconds on Local\n\n# transform the dataset\noversample = SMOTE()\nX, y = oversample.fit_resample(X, y)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# summarize the new class distribution\n\n\n# Generate and plot a synthetic imbalanced classification dataset\nfrom collections import Counter\ncounter = Counter(y)\nprint(counter)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now data is balanced, but increased in size.","metadata":{}},{"cell_type":"code","source":"print(\"X.shape:\",X.shape,\"\\ny.shape:\",y.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y.value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Splitting data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)\n\n# Display the shape of training and testing data\nprint('X_train shape: ', X_train.shape)\nprint('y_train shape: ', y_train.shape)\nprint('X_test shape: ', X_test.shape)\nprint('y_test shape: ', y_test.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**5.2 Feature Importance**","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import classification_report\nfrom sklearn.metrics import confusion_matrix","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will use 2 models for machine learning, LightGradientBoosting and XGBoost.","metadata":{}},{"cell_type":"code","source":"%%time\n\n## aprox runtime 15-30 seconds\n\nimport lightgbm as lgb\nlgb = lgb.LGBMClassifier()\nlgb.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Important Features as per LGBM\n\nlgb_imp_feat = pd.DataFrame(lgb.feature_name_)\nlgb_imp_featVal = pd.DataFrame(lgb.feature_importances_)\n\nlgb_imp_feat.columns=['Name']\nlgb_imp_featVal.columns=['Value']\n\nlgbm_features = pd.concat([lgb_imp_feat, lgb_imp_featVal], axis=1)\ndel lgb_imp_feat, lgb_imp_featVal\n\n#lgbm_features.sort_values(by = 'Value',ascending = False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlgb_y_pred_train = lgb.predict(X_train)\nlgb_y_pred_test = lgb.predict(X_test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nlgb_report = classification_report(y_test, lgb_y_pred_test, output_dict=True)\nlgb_report = pd.DataFrame(lgb_report).transpose()\n\ncm = confusion_matrix(y_test, lgb_y_pred_test)\nlgb_report['True Positives(TP)'] = cm[0,0]\nlgb_report['True Negatives(TN)'] = cm[1,1]\nlgb_report['False Positives(FP)'] = cm[0,1]\nlgb_report['False Negatives(FN)'] = cm[1,0]\nlgb_report['Train Accuracy'] = format(lgb.score(X_train, y_train),'.3f')\nlgb_report['Test Accuracy'] = format(lgb.score(X_test, y_test),'.3f')\nlgb_report['AMEX Metric Train Accuracy'] = format(amex_metric_mod(y_train, lgb_y_pred_train),'.3f')\nlgb_report['AMEX Metric Test Accuracy'] = format(amex_metric_mod(y_test, lgb_y_pred_test),'.3f')\n\nlgb_report","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n## approx runtime 51 seconds - 1 min 40 seconds\nfrom xgboost import XGBClassifier\nxgb = XGBClassifier(max_depth = 5, n_estimators=20)\nxgb.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nxgb_y_pred_train = xgb.predict(X_train)\nxgb_y_pred_test = xgb.predict(X_test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Important Features as per XGB\n\nxgb_imp_feat = pd.DataFrame(xgb.get_booster().feature_names)\nxgb_imp_featVal = pd.DataFrame(xgb.feature_importances_)\n\nxgb_imp_feat.columns=['Name']\nxgb_imp_featVal.columns=['Value']\n\nxgb_features = pd.concat([xgb_imp_feat, xgb_imp_featVal], axis=1)\ndel xgb_imp_feat, xgb_imp_featVal\nxgb_features['Name'] = xgb_features['Name'].str.replace(' ','_')\n#xgb_features.sort_values(by = 'Value',ascending = False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nxgb_report = classification_report(y_test, xgb_y_pred_test, output_dict=True)\nxgb_report = pd.DataFrame(xgb_report).transpose()\n\ncm = confusion_matrix(y_test, xgb_y_pred_test)\nxgb_report['True Positives(TP)'] = cm[0,0]\nxgb_report['True Negatives(TN)'] = cm[1,1]\nxgb_report['False Positives(FP)'] = cm[0,1]\nxgb_report['False Negatives(FN)'] = cm[1,0]\nxgb_report['Train Accuracy'] = format(xgb.score(X_train, y_train),'.3f')\nxgb_report['Test Accuracy'] = format(xgb.score(X_test, y_test),'.3f')\nxgb_report['AMEX Metric Train Accuracy'] = format(amex_metric_mod(y_train, xgb_y_pred_train),'.3f')\nxgb_report['AMEX Metric Test Accuracy'] = format(amex_metric_mod(y_test, xgb_y_pred_test),'.3f')\n\nxgb_report","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Selecting top 20 features from the result of each model's important features after LGB and XGB's important feature result below ","metadata":{}},{"cell_type":"code","source":"#Taking Top 20 Features we got from from XGB and LGBM\nxgb_features_top20 = xgb_features.sort_values(by='Value',ascending=False).head(20)\nlgbm_features_top20 = lgbm_features.sort_values(by='Value',ascending=False).head(20)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import seaborn library\nimport seaborn as sns\n \n# Declaring the cm variable by the\n# color palette from seaborn\ncm = sns.light_palette(\"green\", as_cmap=True)\n \n# Visualizing the DataFrame with set precision\nprint(\"\\nModified Stlying DataFrame:\")\n\nFeaturejoin = [xgb_features_top20,lgbm_features_top20]\n  \nresult = pd.concat(Featurejoin, axis=1,ignore_index=True)\nresult.columns=['XGB Feat','XGB Imp','LGB Feat','LGB Imp']\n\nresult.style.background_gradient(cmap=cm).set_precision(2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Using n Important features**","metadata":{}},{"cell_type":"markdown","source":"We can select any n number of features subject to compute accessibility and time taken to process the model.","metadata":{}},{"cell_type":"code","source":"imp_feature_cols = set(list(xgb_features_top20['Name']) + list(lgbm_features_top20['Name']))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"imp_feature_cols = list(imp_feature_cols)\nprint(\"Total Important Feature Columns:\",len(imp_feature_cols))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will move further with Top ~30+ features combining LGB and XGB's important feature suggestions.\n\nSelecting data based on the important features for training.","metadata":{}},{"cell_type":"code","source":"X_train_imp = X_train[imp_feature_cols]\nX_test_imp = X_test[imp_feature_cols]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport lightgbm as lgb\nlgb_imp = lgb.LGBMClassifier()\nlgb_imp.fit(X_train_imp, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n#from xgboost import XGBClassifier\nxgb_imp = XGBClassifier(max_depth = 5,n_estimators=20)\nxgb_imp.fit(X_train_imp, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# predict the results\nlgb_y_pred_train_imp=lgb_imp.predict(X_train_imp)\nlgb_y_pred_test_imp=lgb_imp.predict(X_test_imp)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# predict the results\nxgb_y_pred_train_imp=xgb_imp.predict(X_train_imp)\nxgb_y_pred_test_imp=xgb_imp.predict(X_test_imp)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nprint(\"\\n LGB Result for selected important Features\\n\")\n\nlgb_report_imp = classification_report(y_test, lgb_y_pred_test_imp, output_dict=True)\nlgb_report_imp = pd.DataFrame(lgb_report_imp).transpose()\n\ncm_lgb_imp = confusion_matrix(y_test, lgb_y_pred_test_imp)\nlgb_report_imp['True Positives(TP)'] = cm_lgb_imp[0,0]\nlgb_report_imp['True Negatives(TN)'] = cm_lgb_imp[1,1]\nlgb_report_imp['False Positives(FP)'] = cm_lgb_imp[0,1]\nlgb_report_imp['False Negatives(FN)'] = cm_lgb_imp[1,0]\nlgb_report_imp['Train Accuracy'] = format(lgb_imp.score(X_train_imp, y_train),'.3f')\nlgb_report_imp['Test Accuracy'] = format(lgb_imp.score(X_test_imp, y_test),'.3f')\nlgb_report_imp['AMEX Metric Train Accuracy'] = format(amex_metric_mod(y_train, lgb_y_pred_train_imp),'.3f')\nlgb_report_imp['AMEX Metric Test Accuracy'] = format(amex_metric_mod(y_test, lgb_y_pred_test_imp),'.3f')\n\nlgb_report_imp","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nprint(\"\\n xgb Result for selected important Features\\n\")\n\nxgb_report_imp = classification_report(y_test, xgb_y_pred_test_imp, output_dict=True)\nxgb_report_imp = pd.DataFrame(xgb_report_imp).transpose()\n\ncm_xgb_imp = confusion_matrix(y_test, xgb_y_pred_test_imp)\nxgb_report_imp['True Positives(TP)'] = cm_xgb_imp[0,0]\nxgb_report_imp['True Negatives(TN)'] = cm_xgb_imp[1,1]\nxgb_report_imp['False Positives(FP)'] = cm_xgb_imp[0,1]\nxgb_report_imp['False Negatives(FN)'] = cm_xgb_imp[1,0]\n\nxgb_report_imp['Train Accuracy'] = format(xgb_imp.score(X_train_imp, y_train),'.3f')\nxgb_report_imp['Test Accuracy'] = format(xgb_imp.score(X_test_imp, y_test),'.3f')\n\nxgb_report_imp['AMEX Metric Train Accuracy'] = format(amex_metric_mod(y_train, xgb_y_pred_train_imp),'.3f')\nxgb_report_imp['AMEX Metric Test Accuracy'] = format(amex_metric_mod(y_test, xgb_y_pred_test_imp),'.3f')\n\n\nxgb_report_imp","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6. Prediction","metadata":{}},{"cell_type":"markdown","source":"Reading test data with specific columns which we have used for training the model.","metadata":{}},{"cell_type":"code","source":"test_imp = temp_test[imp_feature_cols]\ntest_imp.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final = pd.DataFrame()\ntest_final['customer_ID'] = temp_test['customer_ID']\ntest_final.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_imp.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Taking full test data for prediction of final results. change the fraction value as per requirements.","metadata":{}},{"cell_type":"code","source":"test_imp_frac = test_imp.sample(frac=1.0, random_state=42)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_imp_frac.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_imp_frac.columns","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Handling missing values as we have handled for Train data.","metadata":{}},{"cell_type":"code","source":"#Handling Missing Value - Categorical\nif 'D_63' in test_imp_frac.columns:\n    test_imp_frac['D_63']= label_encoder.fit_transform(test_imp_frac['D_63']) \n\nfor col in list(test_imp_frac.select_dtypes(['category']).columns):\n    #print(col,temp_train_last[col].mode()[0])\n    test_imp_frac[col].fillna(test_imp_frac[col].mode()[0], inplace=True)\n    test_imp_frac[col] = test_imp_frac[col].astype('Int64')\n    ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Handling Missing Value - Continuous\nfor col in list(test_imp_frac.select_dtypes(['float16','int32','int64']).columns):\n\n    test_imp_frac[col].fillna(np.mean(~test_imp_frac[col].isnull()), inplace=True)\n    test_imp_frac[col] = test_imp_frac[col].astype(np.float64)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"LGB's Predictions","metadata":{}},{"cell_type":"code","source":"lgb_test_prediction=lgb_imp.predict(test_imp_frac)\nlgb_test_prediction\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgb_test_prediction_proba=lgb_imp.predict_proba(test_imp_frac)\nlgb_test_prediction_proba","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"XGB's Predictions","metadata":{}},{"cell_type":"code","source":"xgb_test_prediction=xgb_imp.predict(test_imp_frac) \nxgb_test_prediction","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Moving ahead with LGB's predictions as lesser time taken for processing","metadata":{}},{"cell_type":"code","source":"test_final['prediction'] = lgb_test_prediction","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final['prediction'] = test_final['prediction'].astype('Int64')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final['prediction'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_final.to_csv(\"Submission.csv\",header=True,index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Predictions distribution for Target.\n\nAs we can see from below chart that data is distributed same as the train dataset's distribution when we started with the data.  \n\n* ~70+ % - Non Default \n* ~20+ % - Default","metadata":{}},{"cell_type":"code","source":"test_final['prediction'].value_counts().plot(kind='pie', title=\"Title\", legend=True, \\\n                   autopct='%1.1f%%', explode=(0, 0), \\\n                   shadow=True, startangle=0)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"****Thanks for exploring! Give your suggesstions in comments. Any suggestions are welcomed!****","metadata":{}}]}