{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Problem Statement\n\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nIn this competition, you’ll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.","metadata":{"id":"0a1ff936"}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:47:19.51638Z","iopub.execute_input":"2022-08-31T09:47:19.517335Z","iopub.status.idle":"2022-08-31T09:47:19.538761Z","shell.execute_reply.started":"2022-08-31T09:47:19.517299Z","shell.execute_reply":"2022-08-31T09:47:19.537698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Importing Libraries","metadata":{"id":"025798f6"}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport missingno as msno\nimport catboost as cb\nfrom xgboost import XGBClassifier\nimport optuna\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\nimport gc\nfrom tabulate import tabulate\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.impute import KNNImputer\nfrom sklearn.model_selection import KFold,StratifiedKFold\nimport lightgbm as lgb","metadata":{"id":"a01d68b8","scrolled":true,"execution":{"iopub.status.busy":"2022-08-31T09:45:43.292077Z","iopub.execute_input":"2022-08-31T09:45:43.292886Z","iopub.status.idle":"2022-08-31T09:45:47.572345Z","shell.execute_reply.started":"2022-08-31T09:45:43.2928Z","shell.execute_reply":"2022-08-31T09:45:47.571083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns',None)\npd.set_option('display.max_rows',None)","metadata":{"id":"a00915ba","execution":{"iopub.status.busy":"2022-08-31T09:45:47.574585Z","iopub.execute_input":"2022-08-31T09:45:47.574924Z","iopub.status.idle":"2022-08-31T09:45:47.581914Z","shell.execute_reply.started":"2022-08-31T09:45:47.574888Z","shell.execute_reply":"2022-08-31T09:45:47.580659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\n#import os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n    #for filename in filenames:\n        #print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:47.583962Z","iopub.execute_input":"2022-08-31T09:45:47.584724Z","iopub.status.idle":"2022-08-31T09:45:47.594801Z","shell.execute_reply.started":"2022-08-31T09:45:47.584689Z","shell.execute_reply":"2022-08-31T09:45:47.593612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Path to Feather Files\n#ftrpath = \"/kaggle/input/parquet-files-amexdefault-prediction\"\n\n#Path to CSV Files\n#csvpath = \"/kaggle/input/amex-default-prediction\"","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:47.599691Z","iopub.execute_input":"2022-08-31T09:45:47.599973Z","iopub.status.idle":"2022-08-31T09:45:47.604839Z","shell.execute_reply.started":"2022-08-31T09:45:47.599948Z","shell.execute_reply":"2022-08-31T09:45:47.603769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Path to Feather Files\nftrpath = \"/Users/akansha/Documents/Data Science/Kaggle Competitions/American Express/amex-default-prediction\"\n\n#Path to CSV Files\ncsvpath = \"/Users/akansha/Documents/Data Science/Kaggle Competitions/American Express/amex-default-prediction\"","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:47.606709Z","iopub.execute_input":"2022-08-31T09:45:47.60749Z","iopub.status.idle":"2022-08-31T09:45:47.616217Z","shell.execute_reply.started":"2022-08-31T09:45:47.607449Z","shell.execute_reply":"2022-08-31T09:45:47.614827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndef import_train_data():\n        \n    \"\"\"\n    Imports Training dataset and prints the shape, Information and Unique Customers\n    :return train_data : Training dataset\n    \"\"\"\n   \n    train_data = pd.read_feather(ftrpath + '/train_data.ftr')\n    print(\"**************Shape of Dataset***********************\")\n    print(train_data.shape)\n    print(\"*****************************************************\")\n    print(\"**************Information about Dataset**************\")\n    print(train_data.info(verbose=True))\n    print(\"**************Unique Customers***********************\")\n    print(f'Number of unique customers: {train_data[\"customer_ID\"].nunique()}')\n    return train_data\n\ndef import_test_data():\n    \n    \"\"\"\n    Imports Test dataset and prints the shape, Information and Unique Customers\n    :return test_data : Test dataset\n    \"\"\"\n   \n    test_data = pd.read_feather(ftrpath + '/test_data.ftr')\n    print(\"**************Shape of Dataset***********************\")\n    print(test_data.shape)\n    print(\"*****************************************************\")\n    print(\"**************Information about Dataset**************\")\n    print(test_data.info(verbose=True))\n    print(\"**************Unique Customers***********************\")\n    print(f'Number of unique customers: {test_data[\"customer_ID\"].nunique()}')\n    return test_data\n\ndef describe_data(data: pd.DataFrame):\n    \n    \"\"\"\n    Returns the descriptive statistics of the dataset\n    :param data: Dataset for which summary statistics needs to be calculated\n    :return : Summary statistics of the Dataframe provided.\n    \"\"\"\n   \n    return(data.describe(include='all').T)\n\ndef missing_values(data: pd.DataFrame) -> pd.DataFrame:\n\n     \n    \"\"\"\n    Creates a DataFrame with following details for each feature:\n    Total Missing Values\n    Percentage of Missing values\n    No of Unique values\n    :param data: Dataset for which missing values needs to be calculated\n    :return missing_data : Top 40 features with highest percentage of missing values.\n    \"\"\"\n    \n    total_missing_values = data.isnull().sum().sort_values(ascending=False)\n    percent_missing_values = (data.isnull().sum()/data.isnull().count()).sort_values(ascending=False)\n    missing_data = pd.concat([total_missing_values, percent_missing_values], axis=1, keys=['Total','Percent'])\n    missing_data['Uniques'] = data.nunique().values\n    return(missing_data.head(40))\n\ndef correlation_values(data: pd.DataFrame,count:int=30) -> pd.DataFrame:\n    \"\"\"\n    Creates a DataFrame which contains the correlation between features in the dataset\n    It contains following details\n    Feature 1\n    Feature 2\n    CORRELALTION - Correlation between the two features\n    CORR_ABS -  Absolute value of correlation\n    :param data: Dataset for which correlation needs to be calculated\n    :param count: No of rows to be returned. Default value is set as 30\n    :return corr_df : Top \"count\" correlation in the dataset.\n    \"\"\"\n    train_data_corr =  data.corr(method='pearson')\n    corr=train_data_corr.where(np.triu(np.ones(train_data_corr.shape),k=1).astype(np.bool))\n    corr_df=corr.unstack().reset_index()\n\n    corr_df.columns = ['Variable1','Variable2','CORRELATION']\n    corr_df['CORR_ABS'] = abs(corr_df['CORRELATION'])\n    return(corr_df.sort_values('CORR_ABS', ascending=False).head(count))\n","metadata":{"id":"JIIm7M7Vq_MM","execution":{"iopub.status.busy":"2022-08-31T09:45:47.619575Z","iopub.execute_input":"2022-08-31T09:45:47.621468Z","iopub.status.idle":"2022-08-31T09:45:47.63776Z","shell.execute_reply.started":"2022-08-31T09:45:47.621432Z","shell.execute_reply":"2022-08-31T09:45:47.636643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def categorical_plot(data:pd.DataFrame, feature: str,xticks:int=0):\n    \n    \"\"\"\n    It is used to plot the distribution of feature with respect to target variable\n    - Total percentage for each category in feature\n    - Percentage of Defaulters for each category in feature\n    - Percentage of Non Defaulters for each category in feature\n    \"\"\"\n    \n    fig, axs = plt.subplots(1,3,figsize=(20,6))\n    feature_target_1_percentage = ((data[data.target == 1][feature].value_counts().sort_values(ascending=False) / data[feature].value_counts().sort_values(ascending=False))*100).round(2)\n    feature_target_0_percentage = ((data[data.target == 0][feature].value_counts().sort_values(ascending=False) / data[feature].value_counts().sort_values(ascending=False))*100).round(2)\n    total_percentage = ((data[feature].value_counts().sort_values(ascending=False) / data.shape[0])*100).round(2)\n    \n    axs[0].set(xlabel=feature, ylabel='Total Percentage')\n    sns.barplot(x= total_percentage.index,y=total_percentage.values,ax=axs[0]).set_title(\"Distribution based on \"+feature)\n    \n    sns.barplot(x= feature_target_0_percentage.index,y=feature_target_0_percentage.values,ax=axs[1]).set_title(feature+\" by Non Defaulter\")\n    axs[1].set(xlabel=feature, ylabel='Percentage of Non Defaulters')\n    sns.barplot(x= feature_target_1_percentage.index,y=feature_target_1_percentage.values,ax=axs[2]).set_title(feature+\" by Defaulter\")\n    plot3 =axs[2].set(xlabel=feature, ylabel='Percentage of Defaulters')\n    def autolabel(rects):\n      for rect in rects:\n          height = rect.get_height()\n          if height>0:\n            ax.text(rect.get_x()+0.2, rect.get_height()/2,height,\n                    ha='center', va='bottom', rotation=90, color='black',size=12,family='serif',style=\"normal\",weight=\"light\")\n    for ax in axs.flatten():\n        for label in ax.get_xticklabels():\n            label.set_rotation(xticks)\n            autolabel(ax.patches)\n\n    plt.show()\n    \ndef target_plot(data: pd.DataFrame):\n    \n    \"\"\"\n    This method is used to plot a pie chart which shows distribution of target variable\n    :param data: Dataset which contains the target variable\n    \"\"\"\n    \n    column = data['target'].value_counts(normalize=True)\n    pie, ax = plt.subplots(figsize=[10,6])\n    labels = column.keys()\n    plt.pie(x=column, autopct=\"%.1f%%\", labels=labels, pctdistance=0.5,explode=[0.05]*2)\n    plt.title(\" Distribution of Target variable\", fontsize=14);\n    \ndef statements_per_customer_plot(data: pd.DataFrame):\n    \n    \"\"\"\n    This method is used to plot a pie chart which shows number of statements for customers in the dataset\n    :param data: Dataset for which the statement counts need to be plotted\n    \"\"\"\n    \n    column = data.customer_ID.value_counts().value_counts().sort_index(ascending=False)\n    pie, ax = plt.subplots(figsize=[10,8])\n    labels = column.keys()\n    plt.pie(x=column, labels=labels, pctdistance=0.5)\n    plt.title(\" Distribution of Statements per Customer\", fontsize=14);","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:47.640195Z","iopub.execute_input":"2022-08-31T09:45:47.640714Z","iopub.status.idle":"2022-08-31T09:45:47.662464Z","shell.execute_reply.started":"2022-08-31T09:45:47.640678Z","shell.execute_reply":"2022-08-31T09:45:47.661064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def process_and_feature_engineer(data: pd.DataFrame) -> pd.DataFrame:\n    \n    \"\"\"\n    This method is used to feature engineer the dataset\n    For categorical features we aggregate the values for each customer ID and calculate count, last and nunique values\n    For numerical features we aggregate the values for each customer ID and calculate mean, std, min, max, last\n    :param data: Dataset to feature engineer\n    :returns df: Dataset after Feature Engineering\n    \"\"\"\n    \n    features = [f for f in list(data.columns) if f not in ['customer_ID','S_2','target']]\n    cat_features = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n    num_features = [feature for feature in features if feature not in cat_features]\n\n    df_num_agg = data.groupby(\"customer_ID\")[num_features].agg(['mean', 'std', 'min', 'max', 'last','first'])\n    df_num_agg.columns = ['_'.join(x) for x in df_num_agg.columns]\n    \n    \n      # Transform float64 columns to float32\n    cols = list(df_num_agg.dtypes[df_num_agg.dtypes == 'float64'].index)\n    for col in cols:\n        df_num_agg[col] = df_num_agg[col].astype(np.float32)\n    # Transform int64 columns to int32\n    cols = list(df_num_agg.dtypes[df_num_agg.dtypes == 'int64'].index)\n    for col in cols:\n        df_num_agg[col] = df_num_agg[col].astype(np.int32)\n    df_num_agg = df_num_agg.fillna(0)\n    \n    df_cat_agg = data.groupby(\"customer_ID\")[cat_features].agg(['count', 'last','first','nunique'])\n    df_cat_agg.columns = ['_'.join(x) for x in df_cat_agg.columns]\n    #df_cat_agg.fillna(\"NONE\",inplace=True)\n    df_cat_agg = df_cat_agg.fillna(\"None\")\n    df = pd.concat([df_num_agg, df_cat_agg], axis=1)\n    del df_num_agg, df_cat_agg\n    num_cols = list(df.dtypes[(df.dtypes == 'float32') | (df.dtypes == 'float64')].index)\n    for col in num_cols:\n         df[col] = df[col].round(2)\n    return df\n    print('shape after engineering', df.shape )\n","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:47.66432Z","iopub.execute_input":"2022-08-31T09:45:47.664982Z","iopub.status.idle":"2022-08-31T09:45:47.679856Z","shell.execute_reply.started":"2022-08-31T09:45:47.664942Z","shell.execute_reply":"2022-08-31T09:45:47.678289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def amex_metric(y_true: pd.DataFrame, y_pred:pd.DataFrame) -> float:\n    \n    \"\"\"\n    This method calculated the evaluation metric for this competition. It is the mean of two measures of rank ordering: \n    Normalized Gini Coefficient 𝐺\n    and default rate captured at 4%, 𝐷\n\n    𝑀=0.5⋅(𝐺+𝐷)\n    \n    The default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n    :param y_true : target values provided for training\n    :param y_pred : Predicted values\n    :return amex_metric: metric used in this competition\n    \n    \"\"\"\n    labels = np.transpose(np.array([y_true, y_pred]))\n    labels = labels[labels[:, 1].argsort()[::-1]]\n    weights = np.where(labels[:,0]==0, 20, 1)\n    cut_vals = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n    gini = [0,0]\n    for i in [1,0]:\n        labels = np.transpose(np.array([y_true, y_pred]))\n        labels = labels[labels[:, i].argsort()[::-1]]\n        weight = np.where(labels[:,0]==0, 20, 1)\n        weight_random = np.cumsum(weight / np.sum(weight))\n        total_pos = np.sum(labels[:, 0] *  weight)\n        cum_pos_found = np.cumsum(labels[:, 0] * weight)\n        lorentz = cum_pos_found / total_pos\n        gini[i] = np.sum((lorentz - weight_random) * weight)\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:47.681497Z","iopub.execute_input":"2022-08-31T09:45:47.681841Z","iopub.status.idle":"2022-08-31T09:45:47.695134Z","shell.execute_reply.started":"2022-08-31T09:45:47.681807Z","shell.execute_reply":"2022-08-31T09:45:47.694084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Reading and understanding data","metadata":{"id":"0fc4682b"}},{"cell_type":"code","source":"#Importing the training dataset\ntrain_data = import_train_data('/kaggle/input/amex-default-prediction/train_data.csv')","metadata":{"id":"1c275bed","scrolled":true,"execution":{"iopub.status.busy":"2022-08-31T09:45:47.700795Z","iopub.execute_input":"2022-08-31T09:45:47.701496Z","iopub.status.idle":"2022-08-31T09:45:48.006252Z","shell.execute_reply.started":"2022-08-31T09:45:47.701461Z","shell.execute_reply":"2022-08-31T09:45:47.992408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# First 5 items in the training dataset\ntrain_data.head()","metadata":{"id":"34f272a9","outputId":"8a4af8f5-b8ab-43db-949c-ae52aa259a3c","execution":{"iopub.status.busy":"2022-08-31T09:45:48.00981Z","iopub.status.idle":"2022-08-31T09:45:48.011137Z","shell.execute_reply.started":"2022-08-31T09:45:48.010805Z","shell.execute_reply":"2022-08-31T09:45:48.010832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Summary statistics for training dataset\ndescribe_data(train_data)","metadata":{"id":"05fa2453","outputId":"187144fd-3674-4917-9695-744ca643695d","execution":{"iopub.status.busy":"2022-08-31T09:45:48.012428Z","iopub.status.idle":"2022-08-31T09:45:48.02113Z","shell.execute_reply.started":"2022-08-31T09:45:48.018841Z","shell.execute_reply":"2022-08-31T09:45:48.018869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Missing Values","metadata":{"id":"a592c780"}},{"cell_type":"code","source":"missing_values(train_data)","metadata":{"id":"fdacaa9f","outputId":"71e5a7b5-17eb-4dc7-edbd-1810de341d53","execution":{"iopub.status.busy":"2022-08-31T09:45:48.025156Z","iopub.status.idle":"2022-08-31T09:45:48.032033Z","shell.execute_reply.started":"2022-08-31T09:45:48.031668Z","shell.execute_reply":"2022-08-31T09:45:48.031726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = import_test_data()","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.034007Z","iopub.status.idle":"2022-08-31T09:45:48.036449Z","shell.execute_reply.started":"2022-08-31T09:45:48.035312Z","shell.execute_reply":"2022-08-31T09:45:48.035342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Exploratory Data Analysis","metadata":{"id":"f9f430eb"}},{"cell_type":"markdown","source":"The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\nwith the following features being categorical:\n\n['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']","metadata":{"id":"57fcd880"}},{"cell_type":"markdown","source":"#### Feature Distribution - No of Deliquency, Spend, Payment, Risk and Balance Features in the dataset","metadata":{"id":"2976bb33"}},{"cell_type":"code","source":"features_Delinquency = [f for f in train_data.columns if f.startswith('D_')]\nfeatures_Spend = [f for f in train_data.columns if f.startswith('S_')]\nfeatures_Payment = [f for f in train_data.columns if f.startswith('P_')]\nfeatures_Balance = [f for f in train_data.columns if f.startswith('B_')]\nfeatures_Risk = [f for f in train_data.columns if f.startswith('R_')]\nprint(f'Total number of Delinquency variables: {len(features_Delinquency)}')\nprint(f'Total number of Spend variables: {len(features_Spend)}')\nprint(f'Total number of Payment variables: {len(features_Payment)}')\nprint(f'Total number of Balance variables: {len(features_Balance)}')\nprint(f'Total number of Risk variables: {len(features_Risk)}')","metadata":{"id":"f3bf4979","outputId":"0f518cc5-56fb-4c02-e115-a389c5063756","execution":{"iopub.status.busy":"2022-08-31T09:45:48.040189Z","iopub.status.idle":"2022-08-31T09:45:48.041526Z","shell.execute_reply.started":"2022-08-31T09:45:48.041254Z","shell.execute_reply":"2022-08-31T09:45:48.041279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"column = [len(features_Delinquency),len(features_Spend),len(features_Payment),len(features_Balance),len(features_Risk)]\npie, ax = plt.subplots(figsize=[10,6])\nlabels = ['Delinquency','Spend','Payment','Balance','Risk']\nplt.pie(x=column, autopct=\"%.1f%%\", labels=labels, pctdistance=0.5)\nplt.title(\"Distribution of Features\", fontsize=14);","metadata":{"id":"458ef585","outputId":"7e0afaee-9bb4-476a-d744-a42dee41ca59","execution":{"iopub.status.busy":"2022-08-31T09:45:48.046226Z","iopub.status.idle":"2022-08-31T09:45:48.049909Z","shell.execute_reply.started":"2022-08-31T09:45:48.049269Z","shell.execute_reply":"2022-08-31T09:45:48.049305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Target Variable","metadata":{"id":"d21ba97a"}},{"cell_type":"code","source":"train_labels = pd.read_csv(csvpath + '/train_labels.csv')","metadata":{"id":"b1dd627a","execution":{"iopub.status.busy":"2022-08-31T09:45:48.051737Z","iopub.status.idle":"2022-08-31T09:45:48.0541Z","shell.execute_reply.started":"2022-08-31T09:45:48.053733Z","shell.execute_reply":"2022-08-31T09:45:48.053765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.head()","metadata":{"id":"fa9c7991","outputId":"bca706d5-c6f1-4d7f-f6cf-09a6470b6db7","execution":{"iopub.status.busy":"2022-08-31T09:45:48.055681Z","iopub.status.idle":"2022-08-31T09:45:48.058463Z","shell.execute_reply.started":"2022-08-31T09:45:48.058163Z","shell.execute_reply":"2022-08-31T09:45:48.05819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.merge(train_data, train_labels, how='inner', on = 'customer_ID')","metadata":{"id":"79aadd9c","execution":{"iopub.status.busy":"2022-08-31T09:45:48.060161Z","iopub.status.idle":"2022-08-31T09:45:48.062415Z","shell.execute_reply.started":"2022-08-31T09:45:48.062147Z","shell.execute_reply":"2022-08-31T09:45:48.062172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_plot(train_data)","metadata":{"id":"5389283f","outputId":"cb473396-c998-407f-ced5-9f5767433219","execution":{"iopub.status.busy":"2022-08-31T09:45:48.063875Z","iopub.status.idle":"2022-08-31T09:45:48.067502Z","shell.execute_reply.started":"2022-08-31T09:45:48.067166Z","shell.execute_reply":"2022-08-31T09:45:48.067198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Converting S_2 into datetime column","metadata":{"id":"b860e7bb"}},{"cell_type":"code","source":"train_data['S_2']= pd.to_datetime(train_data['S_2'])","metadata":{"id":"52656186","execution":{"iopub.status.busy":"2022-08-31T09:45:48.069037Z","iopub.status.idle":"2022-08-31T09:45:48.069804Z","shell.execute_reply.started":"2022-08-31T09:45:48.069551Z","shell.execute_reply":"2022-08-31T09:45:48.069575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Analyzing Categorical columns","metadata":{"id":"601aafd0"}},{"cell_type":"code","source":"categorical_columns = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\nfor col in categorical_columns:\n    train_data[col] = train_data[col].astype(str)","metadata":{"id":"f62ee545","execution":{"iopub.status.busy":"2022-08-31T09:45:48.075453Z","iopub.status.idle":"2022-08-31T09:45:48.076308Z","shell.execute_reply.started":"2022-08-31T09:45:48.076041Z","shell.execute_reply":"2022-08-31T09:45:48.07607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in categorical_columns:\n    categorical_plot(train_data,col)","metadata":{"id":"672938f2","outputId":"ab930c28-f794-46ca-b99b-29bdf8fd9cf8","execution":{"iopub.status.busy":"2022-08-31T09:45:48.077741Z","iopub.status.idle":"2022-08-31T09:45:48.078532Z","shell.execute_reply.started":"2022-08-31T09:45:48.078263Z","shell.execute_reply":"2022-08-31T09:45:48.078287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### No of statements per customer","metadata":{"id":"81788236"}},{"cell_type":"code","source":"statements_per_customer_plot(train_data)","metadata":{"id":"7bd3d5d5","outputId":"e3d99af3-c728-4758-fc70-5c955932ea8b","execution":{"iopub.status.busy":"2022-08-31T09:45:48.079972Z","iopub.status.idle":"2022-08-31T09:45:48.082443Z","shell.execute_reply.started":"2022-08-31T09:45:48.082157Z","shell.execute_reply":"2022-08-31T09:45:48.082189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Around 84% customers have 13 statements","metadata":{"id":"6ffc6f4a"}},{"cell_type":"code","source":"tempcols = []\nfor i in train_data.columns:\n    if train_data[i].nunique() <= 2:\n        tempcols.append(i)\nprint(tempcols)","metadata":{"id":"031c8723","outputId":"6edde016-0b10-4757-a87f-08707067b03d","execution":{"iopub.status.busy":"2022-08-31T09:45:48.083989Z","iopub.status.idle":"2022-08-31T09:45:48.08685Z","shell.execute_reply.started":"2022-08-31T09:45:48.086584Z","shell.execute_reply":"2022-08-31T09:45:48.086609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data[tempcols].info()","metadata":{"id":"4c80d103","outputId":"2b7acc18-d1fa-4b52-ef98-e943d6937546","execution":{"iopub.status.busy":"2022-08-31T09:45:48.088309Z","iopub.status.idle":"2022-08-31T09:45:48.089097Z","shell.execute_reply.started":"2022-08-31T09:45:48.088824Z","shell.execute_reply":"2022-08-31T09:45:48.088848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"B_31 and D_87 have binary values.B_31 is always 0 or 1.\nD_87 is 1 or missing values","metadata":{"id":"86223eda"}},{"cell_type":"code","source":"for col in ['B_31','D_87']:\n    sns.countplot(data=train_data, x=col)\n    plt.show()","metadata":{"id":"9565003c","outputId":"5919ceeb-8a17-472a-d9f6-b168933606c4","execution":{"iopub.status.busy":"2022-08-31T09:45:48.093047Z","iopub.status.idle":"2022-08-31T09:45:48.096144Z","shell.execute_reply.started":"2022-08-31T09:45:48.09581Z","shell.execute_reply":"2022-08-31T09:45:48.095835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Delinquency Features","metadata":{"id":"BoZ93RJVNZZ8"}},{"cell_type":"code","source":"describe_data(train_data[features_Delinquency])","metadata":{"id":"4fSUH0E_M2Nn","outputId":"fdfaf1d5-e433-41a5-d3e9-e7b40da862d4","execution":{"iopub.status.busy":"2022-08-31T09:45:48.104607Z","iopub.status.idle":"2022-08-31T09:45:48.107432Z","shell.execute_reply.started":"2022-08-31T09:45:48.107142Z","shell.execute_reply":"2022-08-31T09:45:48.107168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_values(train_data[features_Delinquency])\n","metadata":{"id":"ru7Tjr_oMUqZ","execution":{"iopub.status.busy":"2022-08-31T09:45:48.108937Z","iopub.status.idle":"2022-08-31T09:45:48.110373Z","shell.execute_reply.started":"2022-08-31T09:45:48.110115Z","shell.execute_reply":"2022-08-31T09:45:48.110141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(train_data[features_Delinquency[:50]],figsize=(20,8),sort=\"ascending\", color=\"blue\")","metadata":{"id":"qNJXbh5zNEjx","execution":{"iopub.status.busy":"2022-08-31T09:45:48.117398Z","iopub.status.idle":"2022-08-31T09:45:48.118989Z","shell.execute_reply.started":"2022-08-31T09:45:48.11872Z","shell.execute_reply":"2022-08-31T09:45:48.118747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(train_data[features_Delinquency[50:]],figsize=(15,8),sort=\"ascending\", color=\"blue\")","metadata":{"id":"iUFw8hCkNQr0","execution":{"iopub.status.busy":"2022-08-31T09:45:48.121649Z","iopub.status.idle":"2022-08-31T09:45:48.123014Z","shell.execute_reply.started":"2022-08-31T09:45:48.122755Z","shell.execute_reply":"2022-08-31T09:45:48.122778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.drop(columns=[\"D_77\",\"D_104\",\"D_119\",\"D_75\"],inplace=True,axis=1)\ntest_data.drop(columns=[\"D_77\",\"D_104\",\"D_119\",\"D_75\"],inplace=True,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.125425Z","iopub.status.idle":"2022-08-31T09:45:48.129617Z","shell.execute_reply.started":"2022-08-31T09:45:48.129357Z","shell.execute_reply":"2022-08-31T09:45:48.129383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Spend Features","metadata":{"id":"D14uwJgn5gcs"}},{"cell_type":"code","source":"describe_data(train_data[features_Spend])","metadata":{"id":"AzD4Z-on5jIU","execution":{"iopub.status.busy":"2022-08-31T09:45:48.132182Z","iopub.status.idle":"2022-08-31T09:45:48.133662Z","shell.execute_reply.started":"2022-08-31T09:45:48.133403Z","shell.execute_reply":"2022-08-31T09:45:48.133427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_values(train_data[features_Spend])\n","metadata":{"id":"zLqjw1In5uHL","execution":{"iopub.status.busy":"2022-08-31T09:45:48.136186Z","iopub.status.idle":"2022-08-31T09:45:48.137909Z","shell.execute_reply.started":"2022-08-31T09:45:48.137636Z","shell.execute_reply":"2022-08-31T09:45:48.13766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(train_data[features_Spend],figsize=(15,8),sort=\"ascending\", color=\"blue\")","metadata":{"id":"WDIRGzAqM-3s","execution":{"iopub.status.busy":"2022-08-31T09:45:48.140538Z","iopub.status.idle":"2022-08-31T09:45:48.142115Z","shell.execute_reply.started":"2022-08-31T09:45:48.141803Z","shell.execute_reply":"2022-08-31T09:45:48.141831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.heatmap(train_data[features_Spend])","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.14438Z","iopub.status.idle":"2022-08-31T09:45:48.145736Z","shell.execute_reply.started":"2022-08-31T09:45:48.145483Z","shell.execute_reply":"2022-08-31T09:45:48.145507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.drop(columns=[\"S_24\",\"S_7\"],inplace=True,axis=1)\ntest_data.drop(columns=[\"S_24\",\"S_7\"],inplace=True,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.148272Z","iopub.status.idle":"2022-08-31T09:45:48.149762Z","shell.execute_reply.started":"2022-08-31T09:45:48.149493Z","shell.execute_reply":"2022-08-31T09:45:48.149519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Payment Features","metadata":{"id":"182urnAc9zaE"}},{"cell_type":"code","source":"describe_data(train_data[features_Payment])","metadata":{"id":"-fuJC7O295aB","execution":{"iopub.status.busy":"2022-08-31T09:45:48.152506Z","iopub.status.idle":"2022-08-31T09:45:48.153834Z","shell.execute_reply.started":"2022-08-31T09:45:48.15357Z","shell.execute_reply":"2022-08-31T09:45:48.153595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_values(train_data[features_Payment])","metadata":{"id":"DQXccQ3q9_cd","execution":{"iopub.status.busy":"2022-08-31T09:45:48.156329Z","iopub.status.idle":"2022-08-31T09:45:48.158705Z","shell.execute_reply.started":"2022-08-31T09:45:48.157844Z","shell.execute_reply":"2022-08-31T09:45:48.157879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(train_data[features_Payment],figsize=(8,8),sort=\"ascending\", color=\"blue\")","metadata":{"id":"aNog0q30IliD","execution":{"iopub.status.busy":"2022-08-31T09:45:48.160907Z","iopub.status.idle":"2022-08-31T09:45:48.162432Z","shell.execute_reply.started":"2022-08-31T09:45:48.162172Z","shell.execute_reply":"2022-08-31T09:45:48.162196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.heatmap(train_data[features_Payment])","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.165417Z","iopub.status.idle":"2022-08-31T09:45:48.166788Z","shell.execute_reply.started":"2022-08-31T09:45:48.166524Z","shell.execute_reply":"2022-08-31T09:45:48.166547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Risk Features","metadata":{"id":"vqfDVPsX-Nnw"}},{"cell_type":"code","source":"describe_data(train_data[features_Risk])","metadata":{"id":"QTRRkMNK-MsY","execution":{"iopub.status.busy":"2022-08-31T09:45:48.169406Z","iopub.status.idle":"2022-08-31T09:45:48.170812Z","shell.execute_reply.started":"2022-08-31T09:45:48.170505Z","shell.execute_reply":"2022-08-31T09:45:48.170529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_values(train_data[features_Risk])","metadata":{"id":"0PXgHX1J-Vat","execution":{"iopub.status.busy":"2022-08-31T09:45:48.173637Z","iopub.status.idle":"2022-08-31T09:45:48.175081Z","shell.execute_reply.started":"2022-08-31T09:45:48.174766Z","shell.execute_reply":"2022-08-31T09:45:48.174794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(train_data[features_Risk],figsize=(12,8),sort=\"ascending\", color=\"blue\")","metadata":{"id":"V_6cmUQGLgtY","execution":{"iopub.status.busy":"2022-08-31T09:45:48.17792Z","iopub.status.idle":"2022-08-31T09:45:48.180162Z","shell.execute_reply.started":"2022-08-31T09:45:48.179107Z","shell.execute_reply":"2022-08-31T09:45:48.179132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.heatmap(train_data[features_Risk])","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.18239Z","iopub.status.idle":"2022-08-31T09:45:48.183788Z","shell.execute_reply.started":"2022-08-31T09:45:48.183528Z","shell.execute_reply":"2022-08-31T09:45:48.183553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Balance Features","metadata":{"id":"0n5bbkwZ-9Qd"}},{"cell_type":"code","source":"describe_data(train_data[features_Balance])","metadata":{"id":"hhMUtspY-_eo","execution":{"iopub.status.busy":"2022-08-31T09:45:48.186498Z","iopub.status.idle":"2022-08-31T09:45:48.187883Z","shell.execute_reply.started":"2022-08-31T09:45:48.187611Z","shell.execute_reply":"2022-08-31T09:45:48.187635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_values(train_data[features_Balance])","metadata":{"id":"Fno63aCX_FH2","execution":{"iopub.status.busy":"2022-08-31T09:45:48.190166Z","iopub.status.idle":"2022-08-31T09:45:48.19208Z","shell.execute_reply.started":"2022-08-31T09:45:48.191783Z","shell.execute_reply":"2022-08-31T09:45:48.191808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(train_data[features_Balance],figsize=(15,8),sort=\"ascending\", color=\"blue\")","metadata":{"id":"pYmYsxHCMk6K","execution":{"iopub.status.busy":"2022-08-31T09:45:48.194166Z","iopub.status.idle":"2022-08-31T09:45:48.195637Z","shell.execute_reply.started":"2022-08-31T09:45:48.195384Z","shell.execute_reply":"2022-08-31T09:45:48.195409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.heatmap(train_data[features_Balance])","metadata":{"id":"58f3171b","execution":{"iopub.status.busy":"2022-08-31T09:45:48.19864Z","iopub.status.idle":"2022-08-31T09:45:48.20029Z","shell.execute_reply.started":"2022-08-31T09:45:48.199968Z","shell.execute_reply":"2022-08-31T09:45:48.199997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.drop(columns=[\"B_11\",\"B_23\"],inplace=True,axis=1)\ntest_data.drop(columns=[\"B_11\",\"B_23\"],inplace=True,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.202937Z","iopub.status.idle":"2022-08-31T09:45:48.205219Z","shell.execute_reply.started":"2022-08-31T09:45:48.204883Z","shell.execute_reply":"2022-08-31T09:45:48.204913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Engineering","metadata":{}},{"cell_type":"code","source":"train_data_final = process_and_feature_engineer(train_data)\ntest_data_final = process_and_feature_engineer(test_data)","metadata":{"id":"fa985c77","execution":{"iopub.status.busy":"2022-08-31T09:45:48.207479Z","iopub.status.idle":"2022-08-31T09:45:48.208908Z","shell.execute_reply.started":"2022-08-31T09:45:48.208647Z","shell.execute_reply":"2022-08-31T09:45:48.208672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data_final.head()","metadata":{"id":"LP3S6NFMyXOa","execution":{"iopub.status.busy":"2022-08-31T09:45:48.210976Z","iopub.status.idle":"2022-08-31T09:45:48.212462Z","shell.execute_reply.started":"2022-08-31T09:45:48.212192Z","shell.execute_reply":"2022-08-31T09:45:48.212216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data_final.shape","metadata":{"id":"CeKBEg8cy9iI","execution":{"iopub.status.busy":"2022-08-31T09:45:48.214681Z","iopub.status.idle":"2022-08-31T09:45:48.216196Z","shell.execute_reply.started":"2022-08-31T09:45:48.215928Z","shell.execute_reply":"2022-08-31T09:45:48.215952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.merge(train_data_final, train_labels, how='inner', on = 'customer_ID')","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.218362Z","iopub.status.idle":"2022-08-31T09:45:48.219965Z","shell.execute_reply.started":"2022-08-31T09:45:48.219656Z","shell.execute_reply":"2022-08-31T09:45:48.219681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.222142Z","iopub.status.idle":"2022-08-31T09:45:48.223629Z","shell.execute_reply.started":"2022-08-31T09:45:48.223372Z","shell.execute_reply":"2022-08-31T09:45:48.223403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_target = train_data['target']","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.225707Z","iopub.status.idle":"2022-08-31T09:45:48.227277Z","shell.execute_reply.started":"2022-08-31T09:45:48.227004Z","shell.execute_reply":"2022-08-31T09:45:48.227046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.drop(columns=[\"target\",\"customer_ID\"],inplace=True,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.2294Z","iopub.status.idle":"2022-08-31T09:45:48.230939Z","shell.execute_reply.started":"2022-08-31T09:45:48.230683Z","shell.execute_reply":"2022-08-31T09:45:48.230707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.233127Z","iopub.status.idle":"2022-08-31T09:45:48.234655Z","shell.execute_reply.started":"2022-08-31T09:45:48.234396Z","shell.execute_reply":"2022-08-31T09:45:48.23442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.info(verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.236981Z","iopub.status.idle":"2022-08-31T09:45:48.238748Z","shell.execute_reply.started":"2022-08-31T09:45:48.23847Z","shell.execute_reply":"2022-08-31T09:45:48.238496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data_final=test_data_final.reset_index()\n\n\noutput= pd.DataFrame()\noutput['customer_ID'] = test_data_final[\"customer_ID\"]\n\ntest_data_final.drop(columns=[\"customer_ID\"],inplace=True,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.240995Z","iopub.status.idle":"2022-08-31T09:45:48.241946Z","shell.execute_reply.started":"2022-08-31T09:45:48.241676Z","shell.execute_reply":"2022-08-31T09:45:48.2417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_columns = categorical_columns = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\ncat_features = [f\"{cf}_last\" for cf in categorical_columns]\ncat_features.extend([f\"{cf}_first\" for cf in categorical_columns])\nfor categorical_feature in cat_features:\n    \n    le_encoder = LabelEncoder()\n    train_data[categorical_feature] = train_data[categorical_feature].astype('str')\n    train_data[categorical_feature] = train_data[categorical_feature].astype('str')\n    \n    le_encoder.fit(train_data[categorical_feature])\n    test_data_final[categorical_feature]  = test_data_final[categorical_feature].map(lambda s: '<unknown>' if s not in le_encoder.classes_ else s)\n    le_encoder.classes_ = np.append(le_encoder.classes_, '<unknown>')\n    train_data[categorical_feature]  = le_encoder.transform(train_data[categorical_feature] )\n    test_data_final[categorical_feature]  = le_encoder.transform(test_data_final[categorical_feature])\n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.245197Z","iopub.status.idle":"2022-08-31T09:45:48.247094Z","shell.execute_reply.started":"2022-08-31T09:45:48.2468Z","shell.execute_reply":"2022-08-31T09:45:48.246827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def lgb_amex_metric(y_pred, y_true):\n    y_true = y_true.get_label()\n    return 'amex_metric', amex_metric(y_true, y_pred), True","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.248413Z","iopub.status.idle":"2022-08-31T09:45:48.250079Z","shell.execute_reply.started":"2022-08-31T09:45:48.249795Z","shell.execute_reply":"2022-08-31T09:45:48.249819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" train_x, valid_x, train_y, valid_y = train_test_split(train_data,y_target, test_size=0.3,random_state=42)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.252316Z","iopub.status.idle":"2022-08-31T09:45:48.254005Z","shell.execute_reply.started":"2022-08-31T09:45:48.253645Z","shell.execute_reply":"2022-08-31T09:45:48.25367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model Training","metadata":{}},{"cell_type":"code","source":"def objective(trial):\n    \n    features = [col for col in train_data.columns if col not in ['customer_ID','S_2','target']]\n   \n    param = {\n        \"objective\": 'binary',\n        \"metric\": \"binary_logloss\",\n        \"boosting\":\"dart\",\n        \"verbosity\": -1,\n        'boosting_type': 'gbdt',\n        \"learning_rate\":trial.suggest_discrete_uniform(\"learning_rate\", 0.001, 0.02, 0.001),\n        'lambda_l1': trial.suggest_loguniform('lambda_l1', 1e-8, 10.0),\n        'lambda_l2': trial.suggest_loguniform('lambda_l2', 1e-8, 10.0),\n        'num_leaves': trial.suggest_int('num_leaves', 2, 512),\n        'feature_fraction': trial.suggest_uniform('feature_fraction', 0.1, 1.0),\n        'bagging_fraction': trial.suggest_uniform('bagging_fraction', 0.1, 1.0),\n        'bagging_freq': trial.suggest_int('bagging_freq', 0, 15),\n        'min_child_samples': trial.suggest_int('min_child_samples', 1, 100),\n       \n    }\n\n  \n    lgb_train = lgb.Dataset(train_x, train_y, categorical_feature = cat_features)\n    lgb_valid = lgb.Dataset(valid_x, valid_y, categorical_feature = cat_features)\n    model = lgb.train(\n            params = param,\n            train_set = lgb_train,\n            num_boost_round = 10000,\n            valid_sets = [lgb_train, lgb_valid],\n            verbose_eval = 100,\n            feval = lgb_amex_metric\n            )\n\n    # Predict validation\n    val_pred = model.predict(valid_x)\n    # Compute out of folds metric\n    score = amex_metric(valid_y, val_pred)\n    return score","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.256382Z","iopub.status.idle":"2022-08-31T09:45:48.258003Z","shell.execute_reply.started":"2022-08-31T09:45:48.257744Z","shell.execute_reply":"2022-08-31T09:45:48.257768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"study = optuna.create_study(direction=\"maximize\")\nstudy.optimize(objective, n_trials=50, timeout=600)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.259413Z","iopub.status.idle":"2022-08-31T09:45:48.260141Z","shell.execute_reply.started":"2022-08-31T09:45:48.259885Z","shell.execute_reply":"2022-08-31T09:45:48.259908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_data.to_feather('output/train_1.feather')","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.261627Z","iopub.status.idle":"2022-08-31T09:45:48.262581Z","shell.execute_reply.started":"2022-08-31T09:45:48.262257Z","shell.execute_reply":"2022-08-31T09:45:48.262314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Best trial:\")\ntrial = study.best_trial\n\nprint(\"  Value: {}\".format(trial.value))\n\nprint(\"  Params: \")\nparam_dist = { }\nfor key, value in trial.params.items():\n    param_dist[key]=value\nparam_dist","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.263925Z","iopub.status.idle":"2022-08-31T09:45:48.264799Z","shell.execute_reply.started":"2022-08-31T09:45:48.264426Z","shell.execute_reply":"2022-08-31T09:45:48.264451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lgb_train = lgb.Dataset(train_x, train_y, categorical_feature = cat_features)\nlgb_model = lgb.train(params = param_dist,train_set = lgb_train, num_boost_round = 5000)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.266226Z","iopub.status.idle":"2022-08-31T09:45:48.267078Z","shell.execute_reply.started":"2022-08-31T09:45:48.266803Z","shell.execute_reply":"2022-08-31T09:45:48.26683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Evaluation","metadata":{}},{"cell_type":"code","source":"output[\"prediction\"]=lgb_model.predict(test_data_final)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.268421Z","iopub.status.idle":"2022-08-31T09:45:48.269158Z","shell.execute_reply.started":"2022-08-31T09:45:48.268888Z","shell.execute_reply":"2022-08-31T09:45:48.268911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.270519Z","iopub.status.idle":"2022-08-31T09:45:48.271324Z","shell.execute_reply.started":"2022-08-31T09:45:48.271067Z","shell.execute_reply":"2022-08-31T09:45:48.271092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.to_csv(\"Submission_v1.csv\",index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.272719Z","iopub.status.idle":"2022-08-31T09:45:48.27346Z","shell.execute_reply.started":"2022-08-31T09:45:48.273208Z","shell.execute_reply":"2022-08-31T09:45:48.273232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!kaggle competitions list -s amex-default-prediction","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.27512Z","iopub.status.idle":"2022-08-31T09:45:48.27607Z","shell.execute_reply.started":"2022-08-31T09:45:48.275739Z","shell.execute_reply":"2022-08-31T09:45:48.275797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!kaggle competitions submit amex-default-prediction -f Submission_v1.csv -m \"V2\"","metadata":{"execution":{"iopub.status.busy":"2022-08-31T09:45:48.277519Z","iopub.status.idle":"2022-08-31T09:45:48.278262Z","shell.execute_reply.started":"2022-08-31T09:45:48.277992Z","shell.execute_reply":"2022-08-31T09:45:48.278015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}