{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Whether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we'll pay back what we charge? That's a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nThe objective of [this competition](https://www.kaggle.com/competitions/amex-default-prediction) is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. In this competition, you'll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:22:46.839143Z","iopub.execute_input":"2022-08-09T15:22:46.841348Z","iopub.status.idle":"2022-08-09T15:22:46.892012Z","shell.execute_reply.started":"2022-08-09T15:22:46.841112Z","shell.execute_reply":"2022-08-09T15:22:46.890654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Libraries for data visulaization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nwarnings.filterwarnings('ignore') #To supress warnings\nimport scipy as sp\nimport itertools\nfrom tqdm import tqdm","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:22:49.029806Z","iopub.execute_input":"2022-08-09T15:22:49.031318Z","iopub.status.idle":"2022-08-09T15:22:50.196815Z","shell.execute_reply.started":"2022-08-09T15:22:49.031248Z","shell.execute_reply":"2022-08-09T15:22:50.195732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Loading train_labels.csv file**","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\") #Loading dataset\ntrain_labels.head() #To see first five rows of the dataset","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:22:52.748332Z","iopub.execute_input":"2022-08-09T15:22:52.748806Z","iopub.status.idle":"2022-08-09T15:22:53.910580Z","shell.execute_reply.started":"2022-08-09T15:22:52.748770Z","shell.execute_reply":"2022-08-09T15:22:53.909196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels['target'].value_counts() #Counts unique values","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Loading train_data.ftr, test_data.ftr files**\n\nThere are a total of 190 variables in the dataset with approximately 450,000 customers in the training set and 925,000 in the test set. This dataset is prodigious and reading it directly consumes the entire memory. There are two options to overcome this problem:\n\n* Read data chunk by chunk\n* Use a compressed dataset\n\nThe disadvantage of reading data chunk by chunk is we cannot explore the entire data distribution.\n\n Second option is using the compressed version of the train and test sets provided by @munumbutt's [AMEX-Feather-Dataset](https://www.kaggle.com/datasets/munumbutt/amexfeather).","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_feather(\"../input/amexfeather/train_data.ftr\")","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:22:58.280039Z","iopub.execute_input":"2022-08-09T15:22:58.280856Z","iopub.status.idle":"2022-08-09T15:23:20.771033Z","shell.execute_reply.started":"2022-08-09T15:22:58.280809Z","shell.execute_reply":"2022-08-09T15:23:20.769685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.tail() #Displays last 5 rows of a dataset","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_data_32 = pd.read_feather(\"../input/amexfeather/train_data_f32.ftr\")\n# train_data_32.tail()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = pd.read_feather(\"../input/amexfeather/test_data.ftr\")\nsample_submission = pd.read_csv(\"../input/amex-default-prediction/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:23:20.774160Z","iopub.execute_input":"2022-08-09T15:23:20.774589Z","iopub.status.idle":"2022-08-09T15:24:05.226981Z","shell.execute_reply.started":"2022-08-09T15:23:20.774550Z","shell.execute_reply":"2022-08-09T15:24:05.225554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Shape of train data:\",train_data.shape)\nprint(\"Shape of test data:\",test_data.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To get a summary of a dataframe we use info() function. This method prints information about a dataframe including the index dtypes, column dtypes, non-null values and memory usage.","metadata":{}},{"cell_type":"code","source":"train_data.info(max_cols=191 ,show_counts = True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can verify the presence of null values using isnull() function.","metadata":{}},{"cell_type":"code","source":"train_data.isnull().sum()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" I will take the latest statement for each customer.","metadata":{}},{"cell_type":"code","source":"train_data1 = train_data.groupby('customer_ID').tail(1).set_index('customer_ID')\ntest_data1 = test_data.groupby('customer_ID').tail(1).set_index('customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:05.228709Z","iopub.execute_input":"2022-08-09T15:24:05.229686Z","iopub.status.idle":"2022-08-09T15:24:13.040595Z","shell.execute_reply.started":"2022-08-09T15:24:05.229627Z","shell.execute_reply":"2022-08-09T15:24:13.039443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Exploratory Data Analysis**\n\nThe target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:  \n**`D_*`:** Delinquency variables  \n**`S_*`:** Spend variables  \n**`P_*`:** Payment variables  \n**`B_*`:** Balance variables  \n**`R_*`:** Risk variables  \nWith the following features being categorical: `B_30`, `B_38`, `D_63`, `D_64`, `D_66`, `D_68`, `D_114`, `D_116`, `D_117`, `D_120`, `D_126`. ","metadata":{}},{"cell_type":"code","source":"feat_Delinquency = [c for c in train_data.columns if c.startswith('D_')]\nfeat_Spend = [c for c in train_data.columns if c.startswith('S_')]\nfeat_Payment = [c for c in train_data.columns if c.startswith('P_')]\nfeat_Balance = [c for c in train_data.columns if c.startswith('B_')]\nfeat_Risk = [c for c in train_data.columns if c.startswith('R_')]\nprint(f'Total number of Delinquency variables: {len(feat_Delinquency)}')\nprint(f'Total number of Spend variables: {len(feat_Spend)}')\nprint(f'Total number of Payment variables: {len(feat_Payment)}')\nprint(f'Total number of Balance variables: {len(feat_Balance)}')\nprint(f'Total number of Risk variables: {len(feat_Risk)}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.express as px\nlabels=['Delinquency', 'Spend','Payment','Balance','Risk']\nvalues= [len(feat_Delinquency), len(feat_Spend),len(feat_Payment), len(feat_Balance),len(feat_Risk)]\nfig = px.pie(train_data, values=values, names=labels)\nfig.update_traces(textposition='inside', textinfo='percent+label')\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **EDA on train data**","metadata":{}},{"cell_type":"code","source":"missing_train_data = train_data.isna().sum().div(len(train_data)).mul(100).sort_values(ascending=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(2,1, figsize=(25,10))\nsns.barplot(x=missing_train_data[:100].index, y=missing_train_data[:100].values, ax=ax[0])\nsns.barplot(x=missing_train_data[100:].index, y=missing_train_data[100:].values, ax=ax[1])\nax[0].set_ylabel(\"Percentage [%]\"), ax[1].set_ylabel(\"Percentage [%]\")\nax[0].tick_params(axis='x', rotation=90); ax[1].tick_params(axis='x', rotation=90)\nplt.suptitle(\"Amount of missing data (in train data)\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target = train_data[\"target\"].value_counts()\ntarget_0 = round((target[0]/train_data['target'].count()*100),2)\ntarget_1 = round((target[1]/train_data['target'].count()*100),2)\ntarget_percentage = {'Target':['0', '1'], 'Percentage':[target_0, target_1]} \ndf_target_percentage = pd.DataFrame(target_percentage)\ngroupedvalues = df_target_percentage.groupby('Percentage').sum().reset_index()\n\nax = sns.barplot(x='Target',y='Percentage', data=df_target_percentage, errwidth=0)\nplt.title('Percentage of target_0 vs target_1 on train data')\nax.bar_label(ax.containers[0])\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above plot, it can be inferred visually that target = 0 rows are more than target = 1 rows. It means the dataset is imbalanced with majority class 'target = 0' and minority 'target = 1'.\n\n24.91% of customers had a default - it is worth talking to these two different groups and find dome key differences.","metadata":{}},{"cell_type":"code","source":"print(f'Number of unique customers: {train_data[\"customer_ID\"].nunique()}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 458913 unique customers.","metadata":{}},{"cell_type":"code","source":"customers = train_data.groupby(['customer_ID','target']).size().reset_index()\ncustomers = customers.rename(columns={0:'Count'})\ncustomers.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,1, figsize=(15,5))\nsns.histplot(x='Count', data=customers, hue='target', stat='percent', multiple=\"dodge\", bins=np.arange(0,14), ax=ax)\nax.bar_label(ax.containers[0], fmt='%.f%%')\nax.bar_label(ax.containers[1], fmt='%.f%%')\nplt.title(\"Count Distribution on train data\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del customers","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"statement = train_data.groupby('customer_ID')['S_2'].max().reset_index()\nstatement.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(20,5))\nax = plt.axes()\n\na=sns.countplot(data=statement,x=statement['S_2'])\nx_dates = statement['S_2'].dt.strftime('%Y-%m-%d').sort_values().unique()\nax.set_xticklabels(labels=x_dates, rotation=45, ha='right')\nplt.title(\"Customer's Last Date Statement's Count Distribution\")\nplt.show()\n\ndel statement","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:13.043144Z","iopub.execute_input":"2022-08-09T15:24:13.043525Z","iopub.status.idle":"2022-08-09T15:24:13.186682Z","shell.execute_reply.started":"2022-08-09T15:24:13.043492Z","shell.execute_reply":"2022-08-09T15:24:13.184851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **EDA on test data**","metadata":{}},{"cell_type":"code","source":"missing_test_data = test_data.isna().sum().div(len(test_data)).mul(100).sort_values(ascending=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(2,1, figsize=(25,10))\nsns.barplot(x=missing_test_data[:100].index, y=missing_test_data[:100].values, ax=ax[0])\nsns.barplot(x=missing_test_data[100:].index, y=missing_test_data[100:].values, ax=ax[1])\nax[0].set_ylabel(\"Percentage [%]\"), ax[1].set_ylabel(\"Percentage [%]\")\nax[0].tick_params(axis='x', rotation=90); ax[1].tick_params(axis='x', rotation=90)\nplt.suptitle(\"Amount of missing data (in test data)\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customers = test_data.groupby(['customer_ID']).size().reset_index()\ncustomers = customers.rename(columns={0:'Count'})\ncustomers.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,1, figsize=(15,5))\nsns.countplot(x = 'Count',data=customers)\nax.grid(linestyle=\"--\",axis='y',color='gray')\nplt.title(\"Count Distribution on test data\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del customers","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"statement = test_data.groupby('customer_ID')['S_2'].max().reset_index()\nfig = plt.figure(figsize=(20,5))\nax = plt.axes()\n\na=sns.countplot(data=statement,x=statement['S_2'])\nx_dates = statement['S_2'].dt.strftime('%Y-%m-%d').sort_values().unique()\nax.set_xticklabels(labels=x_dates, rotation=45)\nplt.title(\"Customer's Last Date Statement's Count Distribution on test data\")\nplt.show()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del statement","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Distribution of Continuos Deliquency Variables**","metadata":{}},{"cell_type":"code","source":"cat_var= ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\nobj_col=['customer_ID', 'S_2']","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:13.188668Z","iopub.execute_input":"2022-08-09T15:24:13.189082Z","iopub.status.idle":"2022-08-09T15:24:13.198399Z","shell.execute_reply.started":"2022-08-09T15:24:13.189035Z","shell.execute_reply":"2022-08-09T15:24:13.197397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in cat_var:\n    sns.countplot(data=train_data,x=i, hue='target')\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del_cols = [c for c in train_data1.columns if (c.startswith(('D','t'))) & (c not in cat_var)] #Delinquency\ndf_del = train_data1[del_cols]\nspd_cols = [c for c in train_data1.columns if (c.startswith(('S','t'))) & (c not in cat_var)] #Spend\ndf_spd = train_data1[spd_cols]\npay_cols = [c for c in train_data1.columns if (c.startswith(('P','t'))) & (c not in cat_var)] #Payment\ndf_pay = train_data1[pay_cols]\nbal_cols = [c for c in train_data1.columns if (c.startswith(('B','t'))) & (c not in cat_var)] #Balance\ndf_bal = train_data1[bal_cols]\nris_cols = [c for c in train_data1.columns if (c.startswith(('R','t'))) & (c not in cat_var)] #Risk\ndf_ris = train_data1[ris_cols]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import math\ndef kdeplot(cols,df,title,figsize):\n    plt_cols = 5\n    plt_rows = math.ceil(len(cols)/plt_cols)\n    \n    fig, axes = plt.subplots(plt_rows, plt_cols, figsize = figsize)\n    for i, ax in enumerate(axes.reshape(-1)):\n        if i < len(cols) - 1:\n            sns.kdeplot(x = cols[i], hue='target', data = df, fill = True, ax = ax)\n            ax.tick_params()\n            ax.xaxis.get_label()\n            ax.set_ylabel('')\n    fig.suptitle(title, fontsize = 35, x = 0.5, y = 1)\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def corr(df,title,figsize):\n    plt.figure(figsize =figsize)\n    corr = df.corr()\n    mask = np.triu(np.ones_like(corr, dtype = bool))\n    sns.heatmap(corr, mask = mask, robust = True, center = 0,square = True, linewidths =.6)\n    plt.title(title)\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kdeplot(del_cols, df_del, 'Distribution of Delinquency Variables',(35,150))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr(df_del,'Correlation of Delinquency Variables',(20,20))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kdeplot(spd_cols,df_spd,'Distribution of Spend Variables',((16,16)))\ncorr(df_spd,'Correlation of Spend Variables',(11,11))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kdeplot(pay_cols,df_pay,'Distribution of Payment Variables',(12,4))\ncorr(df_pay,'Correlation of Payment Variables',(6,6))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kdeplot(bal_cols,df_bal,'Distribution of Balance Variables',(15,24))\ncorr(df_bal,'Correlation of Balance Variables',(11,11))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kdeplot(ris_cols,df_ris,'Distribution of Risk Variables',(18,23))\ncorr(df_ris,'Correlation of Risk Variables',(11,11))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\ntarget = train_data1.corrwith(train_data1['target'], axis=0)\nval = [str(round(v ,1) *100) + '%' for v in target.values]\nfig = go.Figure()\nfig.add_trace(go.Bar(y=target.index, x= target.values, orientation='h',text = val))\nfig.update_layout(title = \"Correlation of variables with Target\",width = 700, height = 3000)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Model Training**","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\nle = LabelEncoder()\nfor cat_feat in cat_var:\n    train_data1[cat_feat] = le.fit_transform(train_data1[cat_feat])\n    test_data1[cat_feat] = le.transform(test_data1[cat_feat])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:13.200365Z","iopub.execute_input":"2022-08-09T15:24:13.200930Z","iopub.status.idle":"2022-08-09T15:24:14.815584Z","shell.execute_reply.started":"2022-08-09T15:24:13.200880Z","shell.execute_reply":"2022-08-09T15:24:14.814474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data1.drop(['S_2'],axis=1,inplace=True)\ntest_data1.drop(['S_2'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:14.817377Z","iopub.execute_input":"2022-08-09T15:24:14.818049Z","iopub.status.idle":"2022-08-09T15:24:16.416329Z","shell.execute_reply.started":"2022-08-09T15:24:14.817999Z","shell.execute_reply":"2022-08-09T15:24:16.415133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_data1.drop('target', axis=1) # Putting feature variables into X\ny = train_data1['target'] # Putting target variable to y","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:16.417729Z","iopub.execute_input":"2022-08-09T15:24:16.418075Z","iopub.status.idle":"2022-08-09T15:24:16.753685Z","shell.execute_reply.started":"2022-08-09T15:24:16.418045Z","shell.execute_reply":"2022-08-09T15:24:16.752491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split # Importing the train test split library\nX = train_data1.drop('target', axis=1) # Putting feature variables into X\ny = train_data1['target'] # Putting target variable to y\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify = y) # Splitting data into train and test set 75:25","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:16.755301Z","iopub.execute_input":"2022-08-09T15:24:16.755841Z","iopub.status.idle":"2022-08-09T15:24:18.711093Z","shell.execute_reply.started":"2022-08-09T15:24:16.755786Z","shell.execute_reply":"2022-08-09T15:24:18.708440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:32:11.735267Z","iopub.execute_input":"2022-08-09T14:32:11.736320Z","iopub.status.idle":"2022-08-09T14:32:11.857828Z","shell.execute_reply.started":"2022-08-09T14:32:11.736267Z","shell.execute_reply":"2022-08-09T14:32:11.856458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import metrics\nfrom sklearn import metrics\nfrom sklearn.metrics import confusion_matrix # Prints the correct and also incorrect values in number count .\nfrom sklearn.metrics import f1_score #Combines precision, recall into a single metric by taking the harmonic mean\nfrom sklearn.metrics import classification_report #Used to show the precision, recall, F1 Score, and support of our trained classification model.\nfrom sklearn.metrics import roc_curve, auc, roc_auc_score #Shows the trade-off between sensitivity (or TPR) and specificity (1 – FPR).","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:18.714625Z","iopub.execute_input":"2022-08-09T15:24:18.715132Z","iopub.status.idle":"2022-08-09T15:24:18.722050Z","shell.execute_reply.started":"2022-08-09T15:24:18.715088Z","shell.execute_reply":"2022-08-09T15:24:18.720999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**ROC curve**  : ROC curve (Receiver Operating Characteristics Curve) is a metric used to measure the performance of a classifier model. The ROC curve depicts the rate of true positives (The model correctly predicts the positive class) with respect to the rate of false positives (the model predicts as positive class but in actual case it is a negative class), highlighting the sensitivity (Sensitivity is a measure of how well a machine learning model can detect positive instances. It is also known as the true positive rate (TPR) or recall of the classifier model.","metadata":{}},{"cell_type":"code","source":"# ROC Curve function\n\ndef draw_roc( actual, probs ):\n    fpr, tpr, thresholds = metrics.roc_curve( actual, probs,\n                                              drop_intermediate = False )\n    auc_score = metrics.roc_auc_score( actual, probs )\n    plt.figure(figsize=(5, 5))\n    plt.plot( fpr, tpr, label='ROC curve (area = %0.2f)' % auc_score )\n    plt.plot([0, 1], [0, 1], 'k--')\n    plt.xlim([0.0, 1.0]) # X axis limit is from 0 to 1\n    plt.ylim([0.0, 1.05]) # Y axis limit is from 0 to 1.05\n    plt.xlabel('False Positive Rate or [1 - True Negative Rate]') #The actual one is negative but predicted as positive\n    plt.ylabel('True Positive Rate') #Both actual and predicted value came out negative\n    plt.title('Receiver operating characteristic example')\n    plt.legend(loc=\"lower right\")\n    plt.show()\n\n    return None","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:18.723629Z","iopub.execute_input":"2022-08-09T15:24:18.724354Z","iopub.status.idle":"2022-08-09T15:24:18.736067Z","shell.execute_reply.started":"2022-08-09T15:24:18.724309Z","shell.execute_reply":"2022-08-09T15:24:18.734723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Confusion matrix**   : A Confusion matrix is an n x n matrix used for evaluating the performance of a classification model. Here n is the number of target variables or target classes. The matrix compares the actual target values with those predicted by the machine learning model.","metadata":{}},{"cell_type":"code","source":"# Created a common function to plot confusion matrix\ndef Plot_confusion_matrix(y_test, pred_test):\n    cm = confusion_matrix(y_test, pred_test)\n    plt.clf()\n    plt.imshow(cm, interpolation='nearest', cmap=plt.cm.Accent)\n    categoryNames = ['Paid','Default']\n    plt.title('Confusion Matrix')\n    plt.ylabel('Actual label')\n    plt.xlabel('Predicted label')\n    ticks = np.arange(len(categoryNames))\n    plt.xticks(ticks, categoryNames, rotation=45)\n    plt.yticks(ticks, categoryNames)\n    s = [['TP','FN'], ['FP', 'TN']]\n    for i in range(2):\n        for j in range(2):\n            plt.text(j,i, str(s[i][j])+\" = \"+str(cm[i][j]),fontsize=12)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:18.738548Z","iopub.execute_input":"2022-08-09T15:24:18.739247Z","iopub.status.idle":"2022-08-09T15:24:18.752548Z","shell.execute_reply.started":"2022-08-09T15:24:18.739202Z","shell.execute_reply":"2022-08-09T15:24:18.750981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Precision Recall curve** : Precision-Recall is a useful measure of success of prediction when the classes are very imbalanced. In information retrieval, precision is a measure of result relevancy, while recall is a measure of how many truly relevant results are returned.\n\nThe precision-recall curve shows the tradeoff between precision and recall for different threshold. A high area under the curve represents both high recall and high precision, where high precision relates to a low false positive rate, and high recall relates to a low false negative rate. High scores for both show that the classifier is returning accurate results (high precision), as well as returning a majority of all positive results (high recall).\n\nA system with high recall but low precision returns many results, but most of its predicted labels are incorrect when compared to the training labels. A system with high precision but low recall is just the opposite, returning very few results, but most of its predicted labels are correct when compared to the training labels. An ideal system with high precision and high recall will return many results, with all results labeled correctly.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import precision_recall_curve \ndef precision_recall_plot(y_true, y_probs, label):\n   \n    p, r, thresh = precision_recall_curve(y_true, y_probs)\n    p, r, thresh = list(p), list(r), list(thresh)\n    p.pop()\n    r.pop()\n\n    fig, axis = plt.subplots(nrows=1, ncols=1, figsize=(10, 5))\n    sns.lineplot(thresh, p, estimator=None,\n                     label='Precision', ax=axis)\n    axis.set_xlabel('Threshold')\n    axis.set_ylabel('Precision')\n    axis.legend(loc='lower left')\n    axis_twin = axis.twinx()\n    sns.lineplot(thresh, r, estimator=None,color='limegreen', label='Recall', ax=axis_twin)\n    axis_twin.set_ylabel('Recall')\n    axis_twin.set_ylim(0, 1)\n    axis_twin.legend(bbox_to_anchor=(0.24, 0.18))\n    axis.set_xlim(0, 1)\n    axis.set_ylim(0, 1)\n    axis.set_title('Precision Vs Recall')\n    \n    plt.close()\n    \n    fig.subplots_adjust(wspace=5)\n    fig.tight_layout()\n    display(fig)\n    \n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:18.754273Z","iopub.execute_input":"2022-08-09T15:24:18.755301Z","iopub.status.idle":"2022-08-09T15:24:18.768500Z","shell.execute_reply.started":"2022-08-09T15:24:18.755239Z","shell.execute_reply":"2022-08-09T15:24:18.767137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def metric_All(clf_name,y,y_pred):\n    accuracy_metric = metrics.accuracy_score(y, y_pred) #To check accuracy\n    f1score_metric = f1_score(y, y_pred) #To see F1score\n    confusion = metrics.confusion_matrix(y, y_pred)\n    TP = confusion[0,0] # true positive \n    TN = confusion[1,1] # true negatives\n    FP = confusion[1,0] # false positives\n    FN = confusion[0,1] # false negatives\n    sensitivity = TP / float(TP+FN) #Measure of how well a machine learning model can detect positive instances.\n    specificity =  TN / float(TN+FP) # Metric that evaluates a model's ability to predict true negatives of each available category.\n    print(\"Classification report\")\n    print(classification_report(y, y_pred))\n    print(\"ROC curve\")\n    draw_roc(y, y_pred)\n    roc_metric = metrics.roc_auc_score(y, y_pred)\n    fpr, tpr, thresholds = metrics.roc_curve(y, y_pred)\n    Plot_confusion_matrix(y, y_pred)\n    precision_recall_plot(y, y_pred,clf_name)\n    return (accuracy_metric,f1score_metric,sensitivity,specificity,roc_metric)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:18.770393Z","iopub.execute_input":"2022-08-09T15:24:18.771793Z","iopub.status.idle":"2022-08-09T15:24:18.788485Z","shell.execute_reply.started":"2022-08-09T15:24:18.771732Z","shell.execute_reply":"2022-08-09T15:24:18.787088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from lightgbm import LGBMClassifier , early_stopping , log_evaluation\nfrom catboost import CatBoostClassifier\n# from sklearn.model_selection import StratifiedKFold","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:24:18.789946Z","iopub.execute_input":"2022-08-09T15:24:18.790994Z","iopub.status.idle":"2022-08-09T15:24:20.475916Z","shell.execute_reply.started":"2022-08-09T15:24:18.790935Z","shell.execute_reply":"2022-08-09T15:24:20.474685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# kfold = StratifiedKFold(n_splits=10, shuffle=True, random_state=21)\n# for fold, (train_idx, val_idx) in enumerate(kfold.split(X,y)):\n#     print(\"\\nFold {}\".format(fold+1))\n#     X_train, y_train = X.iloc[train_idx,:], y[train_idx]\n#     X_test, y_test = X.iloc[val_idx,:], y[val_idx]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Reference of these params took from: https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline\nparams_lgb = {'boosting_type': 'dart',\n          'n_estimators': 1000,\n          'num_leaves': 100,\n          'learning_rate': 0.01,\n          'colsample_bytree': 0.9,\n          'min_child_samples': 2000,\n          'max_bins': 500,\n          'reg_alpha': 2,\n          'objective': 'binary',\n          'random_state': 21,\n          'bagging_freq': 10,\n          'bagging_fraction': 0.50,\n          'n_jobs': -1,\n          'lambda_l2': 2}\n\nparams_cat = {\n    'max_depth': 7,\n    'od_type': 'Iter',\n    'l2_leaf_reg': 70,\n    'random_seed': 42,\n    'iterations': 2500, #20500,\n    'learning_rate': 0.03,\n    'loss_function': 'Logloss',\n    'early_stopping_rounds': 1500}","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:25:20.244801Z","iopub.execute_input":"2022-08-09T15:25:20.245350Z","iopub.status.idle":"2022-08-09T15:25:20.253898Z","shell.execute_reply.started":"2022-08-09T15:25:20.245309Z","shell.execute_reply":"2022-08-09T15:25:20.252622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"classifiers = {\n    \"Cat Boost Classifier\": CatBoostClassifier(**params_cat,nan_mode ='Min'),\n    \"Light GBM\": LGBMClassifier(**params_lgb)\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:25:20.776789Z","iopub.execute_input":"2022-08-09T15:25:20.777814Z","iopub.status.idle":"2022-08-09T15:25:20.789767Z","shell.execute_reply.started":"2022-08-09T15:25:20.777759Z","shell.execute_reply":"2022-08-09T15:25:20.787814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_Results = pd.DataFrame(columns=['Model','Accuracy','F1 score','Sensitivity','Specificity','ROC value'])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:25:24.381531Z","iopub.execute_input":"2022-08-09T15:25:24.381966Z","iopub.status.idle":"2022-08-09T15:25:24.390858Z","shell.execute_reply.started":"2022-08-09T15:25:24.381933Z","shell.execute_reply":"2022-08-09T15:25:24.389514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:25:28.595997Z","iopub.execute_input":"2022-08-09T15:25:28.596478Z","iopub.status.idle":"2022-08-09T15:25:28.757944Z","shell.execute_reply.started":"2022-08-09T15:25:28.596411Z","shell.execute_reply":"2022-08-09T15:25:28.756578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgbm_probs = []\nlgbm_preds = []\ncat_probs = []\ncat_preds = []","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:25:31.534076Z","iopub.execute_input":"2022-08-09T15:25:31.534531Z","iopub.status.idle":"2022-08-09T15:25:31.539825Z","shell.execute_reply.started":"2022-08-09T15:25:31.534485Z","shell.execute_reply":"2022-08-09T15:25:31.538662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def model(clf, df_Results,X_train,y_train, X_test, y_test):\n    for i, (clf_name,clf) in enumerate(classifiers.items()):\n#         kfold = StratifiedKFold(n_splits=5, shuffle=True, random_state=21)\n#         for fold, (train_idx, val_idx) in enumerate(kfold.split(X,y)):\n#             X_train, y_train = X.iloc[train_idx,:], y[train_idx]\n#             X_test, y_test = X.iloc[val_idx,:], y[val_idx]\n        if clf_name == \"Cat Boost Classifier\":\n            clf.fit(X_train, y_train,eval_set = [(X_test, y_test)], cat_features=cat_var,  verbose = 100)\n            cat_prob = clf.predict_proba(X_test)[:,1] #On split test set (Split from train dataset)\n            cat_probs.append(cat_prob) \n\n            #On the original test dataset\n            cat_preds.append(clf.predict_proba(test_data1)[:, 1])\n        if clf_name == \"Light GBM\":\n            clf.fit(X_train, y_train, eval_set=[(X_train, y_train), (X_test, y_test)],\n                                           callbacks=[early_stopping(100), log_evaluation(200)],\n                                           eval_metric=['auc','binary_logloss'])\n            lgbm_prob = clf.predict_proba(X_test)[:,1] #On split test set (Split from train dataset)\n            lgbm_probs.append(lgbm_prob)\n\n            #On the original test dataset\n            lgbm_preds.append(clf.predict_proba(test_data1)[:, 1])\n\n        y_pred = clf.predict(X_test) #Predictions on the split set\n        print(clf_name)\n        s = metric_All(clf_name,y_test,y_pred)\n        y_test_pred = clf.predict(test_data1)   # Predictions on the original test dataset \n        df_Results = df_Results.append(pd.DataFrame({'Model': clf_name+ ' on test data','Accuracy': s[0] ,'F1 score':s[1],'Sensitivity':s[2],'Specificity':s[3],'ROC value': s[4]}, index=[0]),ignore_index= True)\n        print('-'*60 )\n        gc.collect()\n    return df_Results","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:27:35.298642Z","iopub.execute_input":"2022-08-09T15:27:35.299074Z","iopub.status.idle":"2022-08-09T15:27:35.312371Z","shell.execute_reply.started":"2022-08-09T15:27:35.299040Z","shell.execute_reply":"2022-08-09T15:27:35.310885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_Results = model(classifiers,df_Results,X_train,y_train, X_test, y_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:27:36.128328Z","iopub.execute_input":"2022-08-09T15:27:36.128819Z","iopub.status.idle":"2022-08-09T16:17:04.449953Z","shell.execute_reply.started":"2022-08-09T15:27:36.128779Z","shell.execute_reply":"2022-08-09T16:17:04.448317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T16:17:26.618047Z","iopub.execute_input":"2022-08-09T16:17:26.618666Z","iopub.status.idle":"2022-08-09T16:17:26.865542Z","shell.execute_reply.started":"2022-08-09T16:17:26.618615Z","shell.execute_reply":"2022-08-09T16:17:26.862891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(np.mean(gbm_test_preds, axis=0).tolist())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(np.mean(cat_preds, axis=0).tolist())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_Results","metadata":{"execution":{"iopub.status.busy":"2022-08-09T16:17:29.477856Z","iopub.execute_input":"2022-08-09T16:17:29.478453Z","iopub.status.idle":"2022-08-09T16:17:29.500608Z","shell.execute_reply.started":"2022-08-09T16:17:29.478389Z","shell.execute_reply":"2022-08-09T16:17:29.499294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above, it is clear that both the classifiers performed well.","metadata":{}},{"cell_type":"markdown","source":"# Tuning","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import RandomizedSearchCV\nfrom sklearn.model_selection import KFold","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:52:54.177213Z","iopub.execute_input":"2022-08-09T14:52:54.177621Z","iopub.status.idle":"2022-08-09T14:52:54.182683Z","shell.execute_reply.started":"2022-08-09T14:52:54.177589Z","shell.execute_reply":"2022-08-09T14:52:54.181503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params_cat = {\n    'max_depth': 7,\n    'od_type': 'Iter',\n    'l2_leaf_reg': 70,\n    'random_seed': 42,\n    'iterations': 2500, #20500,\n    'learning_rate': 0.03,\n    'loss_function': 'Logloss',\n    'early_stopping_rounds': 1500}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def TuneLGBM(X,y):\n# #     params_cat = {\n# #         'max_depth': range(1,10,2),\n# #         'od_type': 'Iter',\n# #         'l2_leaf_reg': range(10,100,10),\n# #         'iterations': range(1000,5000,1000) #20500,\n# #     }\n#     params_lgb = {\n#                   'n_estimators': 1000,\n#                   'num_leaves': range(100,500,50),\n#                   'n_estimators':range(1000,5000,1000),\n#                   'min_child_samples': range(500,3000,500),\n#                   'max_bins': range(100,1000,200),\n#                   'reg_alpha': range(1,5,1)}\n\n# #                   'objective': 'binary',\n# #                   'random_state': 21,\n# #                   'bagging_freq': 10,\n# #                   'bagging_fraction': 0.50,\n# #                   'n_jobs': -1,\n# #                   'lambda_l2': 2}\n#     lgbm = LGBMClassifier(boosting_type='dart',learning_rate=0.01,colsample_bytree=0.9,objective = 'binary',random_state = 21)\n# #    cat = CatBoostClassifier(learning_rate = 0.03, loss_function = 'Logloss',early_stopping_rounds = 1500, random_state = 21)\n#     random_search_lgbm = RandomizedSearchCV(estimator = lgbm, \n#                             param_distributions = params_lgb, \n#                             scoring= 'roc_auc',\n#                             cv = KFold(n_splits=5, shuffle=True, random_state=4),\n#                             n_jobs = -1,\n#                             verbose = 1, \n#                             return_train_score=True)\n\n#       # Fit the model\n#     random_search_lgbm.fit(X, y)\n#     best_estimator = random_search_lgbm.best_estimator_\n#     return best_estimator","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:10:10.614582Z","iopub.execute_input":"2022-08-09T15:10:10.616060Z","iopub.status.idle":"2022-08-09T15:10:10.630105Z","shell.execute_reply.started":"2022-08-09T15:10:10.615993Z","shell.execute_reply":"2022-08-09T15:10:10.629075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TuneLGBM(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T15:10:16.037110Z","iopub.execute_input":"2022-08-09T15:10:16.037755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# NEW CODE 1","metadata":{}},{"cell_type":"code","source":"# from sklearn.model_selection import StratifiedKFold \n# y_valid, gbm_val_probs, gbm_test_preds= [],[],[]\n# sk_fold = StratifiedKFold(n_splits=10, shuffle=True, random_state=21)\n\n# for fold, (train_idx, val_idx) in enumerate(sk_fold.split(X, y)):\n#     print(\"\\nFold {}\".format(fold+1))\n#     X_train, y_train = X.iloc[train_idx,:], y[train_idx]\n#     X_val, y_val = X.iloc[val_idx,:], y[val_idx]\n#     print(\"Train shape: {}, {}, Valid shape: {}, {}\\n\".format(\n#         X_train.shape, y_train.shape, X_val.shape, y_val.shape))\n#     gbm = LGBMClassifier(**params).fit(X_train, y_train, \n#                            eval_set=[(X_train, y_train), (X_val, y_val)],\n#                            callbacks=[early_stopping(200), log_evaluation(500)],\n#                            eval_metric=['auc','binary_logloss'])\n#     gbm_prob = gbm.predict_proba(X_val)[:,1] #On split test set (Split from train dataset)\n#     gbm_val_probs.append(gbm_prob)\n#     y_valid.append(y_val)\n\n#     #On the original test dataset\n#     gbm_test_preds.append(gbm.predict_proba(test_data1)[:,1])\n\n#     y_pred = gbm.predict(X_val) #Predictions on the split set\n#     s = metric_All(\"Light GBM\",y_val,y_pred)\n#     y_test_pred = gbm.predict(test_data1)   # Predictions on the original test dataset \n#     print('-'*60 )\n#     gc.collect()\n#     del X_train, y_train, X_val, y_val","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# NEW CODE 2","metadata":{}},{"cell_type":"code","source":"# import numpy as np\n# import pandas as pd\n# pd.set_option('display.max_rows', 500)\n# pd.set_option('display.max_columns', 500)\n# pd.set_option('display.width', 1000)\n\n# import gc\n# import warnings\n# warnings.filterwarnings('ignore')\n# import scipy as sp\n# import joblib\n# from sklearn.preprocessing import LabelEncoder\n# import lightgbm as lgb\n# from itertools import combinations","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def amex_metric(y_true, y_pred):\n#     labels = np.transpose(np.array([y_true, y_pred]))\n#     labels = labels[labels[:, 1].argsort()[::-1]]\n#     weights = np.where(labels[:,0]==0, 20, 1)\n#     cut_vals = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n#     top_four = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n#     gini = [0,0]\n#     for i in [1,0]:\n#         labels = np.transpose(np.array([y_true, y_pred]))\n#         labels = labels[labels[:, i].argsort()[::-1]]\n#         weight = np.where(labels[:,0]==0, 20, 1)\n#         weight_random = np.cumsum(weight / np.sum(weight))\n#         total_pos = np.sum(labels[:, 0] *  weight)\n#         cum_pos_found = np.cumsum(labels[:, 0] * weight)\n#         lorentz = cum_pos_found / total_pos\n#         gini[i] = np.sum((lorentz - weight_random) * weight)\n#     return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def lgb_amex_metric(y_pred, y_true):\n#     y_true = y_true.get_label()\n#     return 'amex_metric', amex_metric(y_true, y_pred), True","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.train.html\n# https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963/notebook\n# params_lgb = {'boosting_type': 'dart',\n#           'n_estimators': 1000,\n#           'num_leaves': 100,\n#           'learning_rate': 0.01,\n#           'colsample_bytree': 0.9,\n#           'min_child_samples': 2000,\n#           'max_bins': 500,\n#           'reg_alpha': 2,\n#           'objective': 'binary',\n#           'random_state': 21,\n#           'bagging_freq': 10,\n#           'bagging_fraction': 0.50,\n#           'n_jobs': -1,\n#           'lambda_l2': 2}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import lightgbm as lgbm\n# from sklearn.model_selection import StratifiedKFold\n\n# test_predictions = np.zeros(len(test_data1))\n# kfold = StratifiedKFold(n_splits = 5, shuffle = True, random_state = 42)\n# for fold, (train_idx, val_idx) in enumerate(kfold.split(X,y)):\n#     print(\"\\nFold {}\".format(fold+1))\n#     X_train, y_train = X.iloc[train_idx,:], y[train_idx]\n#     X_val, y_val = X.iloc[val_idx,:], y[val_idx]    \n    \n#     gbm = lgbm.train(params = params_lgb,\n#                     train_set = lgbm.Dataset(X_train, y_train, categorical_feature = cat_var),\n#                      num_boost_round = 1000,\n#                      valid_sets = lgbm.Dataset(X_val, y_val, categorical_feature = cat_var),\n#                      early_stopping_rounds = 100,\n#                      feval = lgb_amex_metric\n#                     )\n    \n#     y_pred = gbm.predict(X_val)\n#     test_predictions += y_pred / 5 # 5 is number of folds here.\n#     test_pred = gbm.predict(test_data1)\n#     score = amex_metric(y_val, y_pred)\n#     del X_train, X_val, y_train, y_val","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test_df = pd.DataFrame({'customer_ID': test_data1['customer_ID'], 'prediction': test_predictions})\n# test_df.to_csv(f'/content/drive/MyDrive/Submission_3.csv', index = False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **SUBMISSION**","metadata":{}},{"cell_type":"code","source":"import plotly.graph_objects as go\nsample_submission['prediction']=np.mean(cat_preds, axis=0)\ndf=pd.DataFrame(data={'Target':sample_submission['prediction'].apply(lambda x:1 if x > 0.5 else 0)})\ndf=df.Target.value_counts(normalize=True)\ndf.rename(index={0:'Paid', 1:'Default'},inplace=True)\nfig=go.Figure()\nfig.add_trace(go.Pie(labels=df.index, values=df*100, \n                     showlegend=True, hovertemplate = \"%{label} Accounts: %{value:.2f}%<extra></extra>\"))\nfig.update_layout(title='Predicted Target Distribution', \n                  legend=dict(traceorder='reversed',y=1,x=1),\n                  uniformtext_minsize=15, uniformtext_mode='hide',width=700)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T16:18:12.426127Z","iopub.execute_input":"2022-08-09T16:18:12.426745Z","iopub.status.idle":"2022-08-09T16:18:13.122947Z","shell.execute_reply.started":"2022-08-09T16:18:12.426694Z","shell.execute_reply":"2022-08-09T16:18:13.120663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission.to_csv('submission.csv', index=False)\nsample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T16:18:19.706710Z","iopub.execute_input":"2022-08-09T16:18:19.707183Z","iopub.status.idle":"2022-08-09T16:18:23.467015Z","shell.execute_reply.started":"2022-08-09T16:18:19.707139Z","shell.execute_reply":"2022-08-09T16:18:23.465538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from subprocess import check_output\nprint(check_output([\"ls\", \"../input\"]).decode(\"utf8\"))\nprint(check_output([\"ls\", \"../working\"]).decode(\"utf8\"))\nfrom IPython.display import FileLink\nFileLink(r'submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T16:18:27.341485Z","iopub.execute_input":"2022-08-09T16:18:27.341954Z","iopub.status.idle":"2022-08-09T16:18:28.088094Z","shell.execute_reply.started":"2022-08-09T16:18:27.341914Z","shell.execute_reply":"2022-08-09T16:18:28.086488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"padding:20px;color:white;margin:0;font-size:175%;text-align:center;display:fill;border-radius:5px;background-color:#016CC9;overflow:hidden;font-weight:500\">Thank you for reading. \n    Please let me know your feedback 🙂.\n</div>","metadata":{}}]}