{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n# <p style=\"padding:15px;font-family:newtimeroman;color:#9900cc;text-align:center;font-size:110%;border-radius:40px 5px;\"> (|| American Express Default Prediction - EDA ||)</p>\n<img src=\"https://blog.bankbazaar.com/wp-content/uploads/2016/03/Surviving-a-Credit-Card-Default.png\" style=\"border-radius:5px;width:100%;height:500px\">\n\n# <p style=\"background-color:#9900cc;padding:15px;font-family:newtimeroman;color:#ffff80;font-size:110%;border-radius:40px 5px;\">1 | Competition overview</p>\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nIn this competition, you’ll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.\n\n## Data overview\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\nwith the following features being categorical:\n\n['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\nYour task is to predict, for each customer_ID, the probability of a future payment default (target = 1).\n\nNote that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.\n\n## objective\nThe objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\n## Evaluation\nThe evaluation metric, M , for this competition is the mean of two measures of rank ordering: Normalized Gini Coefficient,G, and default rate captured at 4%,D.\n\nM = 0.5(G + D)\n\nThe default rate captured at 4% is the percentage of the positive labels (defaults) captured within the highest-ranked 4% of the predictions, and represents a Sensitivity/Recall statistic.\n\nFor both of the sub-metrics  and , the negative labels are given a weight of 20 to adjust for downsampling.\n\nThis metric has a maximum value of 1.0.\n\nPython code for calculating this metric can be found in this Notebook.","metadata":{"id":"nombqTiD8FgX"}},{"cell_type":"markdown","source":"# <p style=\"background-color:#9900cc;padding:15px;font-family:newtimeroman;color:#ffff80;font-size:110%;border-radius:40px 5px;\">2 | Basic Stuffs -- import,settings,reading</p>","metadata":{"id":"6nhUmSW68ORW"}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\n# visualization tools\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\n# kaggle utils\nimport kaggle_utils_py as kaggle_utils\n\n# garbage collector\nimport gc\n\n# modeling\nimport optuna\nfrom sklearn.model_selection import StratifiedKFold \nfrom sklearn.metrics import roc_auc_score, roc_curve, auc\nfrom lightgbm import LGBMClassifier, early_stopping","metadata":{"id":"W3Gelqz48DCr","execution":{"iopub.status.busy":"2022-07-02T13:59:12.582736Z","iopub.execute_input":"2022-07-02T13:59:12.583352Z","iopub.status.idle":"2022-07-02T13:59:16.407388Z","shell.execute_reply.started":"2022-07-02T13:59:12.583256Z","shell.execute_reply":"2022-07-02T13:59:16.406033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# some basic settings for me\npd.set_option('display.max_columns', None)\n\n# set the warning off\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"id":"fRsGVjqG8QZm","execution":{"iopub.status.busy":"2022-07-02T13:59:16.409541Z","iopub.execute_input":"2022-07-02T13:59:16.410199Z","iopub.status.idle":"2022-07-02T13:59:16.419686Z","shell.execute_reply.started":"2022-07-02T13:59:16.410148Z","shell.execute_reply":"2022-07-02T13:59:16.417630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrain = pd.read_feather('../input/amex-default-prediction-feather/train.feather')\ntest = pd.read_feather('../input/amex-default-prediction-feather/test.feather')\ntrain_labels = pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\")\nsub = pd.read_csv('../input/amex-default-prediction/sample_submission.csv')","metadata":{"id":"hwYi__UQ8Rzv","outputId":"e2b5e010-8fd5-4748-ffe8-89b300baba43","execution":{"iopub.status.busy":"2022-07-02T13:59:21.584307Z","iopub.execute_input":"2022-07-02T13:59:21.584710Z","iopub.status.idle":"2022-07-02T14:00:20.430435Z","shell.execute_reply.started":"2022-07-02T13:59:21.584677Z","shell.execute_reply":"2022-07-02T14:00:20.428856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"shape of the data --->\", train.shape)\nprint(\"shape of the test data --->\", test.shape)","metadata":{"id":"J0wMslpp8h4t","outputId":"3886079c-3876-4f54-9400-5f7764400605","execution":{"iopub.status.busy":"2022-07-02T14:00:20.434056Z","iopub.execute_input":"2022-07-02T14:00:20.434933Z","iopub.status.idle":"2022-07-02T14:00:20.441818Z","shell.execute_reply.started":"2022-07-02T14:00:20.434882Z","shell.execute_reply":"2022-07-02T14:00:20.440707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"id":"ekKw-DUi_Wmg","outputId":"83f5d918-6660-46fd-dcfa-6cbf57fca700","execution":{"iopub.status.busy":"2022-07-02T14:00:20.443130Z","iopub.execute_input":"2022-07-02T14:00:20.443544Z","iopub.status.idle":"2022-07-02T14:00:20.606204Z","shell.execute_reply.started":"2022-07-02T14:00:20.443509Z","shell.execute_reply":"2022-07-02T14:00:20.605080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d_feats = [c for c in train.columns if c.startswith('D_')]\ns_feats = [c for c in train.columns if c.startswith('S_')]\np_feats = [c for c in train.columns if c.startswith('P_')]\nb_feats = [c for c in train.columns if c.startswith('B_')]\nr_feats = [c for c in train.columns if c.startswith('R_')]\nprint(f'Number of Delinquency variables: {len(d_feats)}')\nprint(f'Number of Spend variables: {len(s_feats)}')\nprint(f'Number of Payment variables: {len(p_feats)}')\nprint(f'Number of Balance variables: {len(b_feats)}')\nprint(f'Number of Risk variables: {len(r_feats)}')","metadata":{"id":"IBe1iDIY_cyc","outputId":"30526f31-3525-4601-acaf-1e1efab13742","execution":{"iopub.status.busy":"2022-07-02T14:00:20.608280Z","iopub.execute_input":"2022-07-02T14:00:20.608676Z","iopub.status.idle":"2022-07-02T14:00:20.617774Z","shell.execute_reply.started":"2022-07-02T14:00:20.608641Z","shell.execute_reply":"2022-07-02T14:00:20.616960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n# <p style=\"background-color:#9900cc;padding:15px;font-family:newtimeroman;color:#ffff80;font-size:110%;border-radius:40px 5px;\"> 3 | Understanding Customer data</p>","metadata":{"id":"z0bdBNO0OF0p"}},{"cell_type":"code","source":"unique_customer_count = len(train.groupby(\"customer_ID\")['customer_ID'].count())\nprint(\"unique customer data in training data -->\", unique_customer_count)\nunique_customer_count_test = len(test.groupby(\"customer_ID\")['customer_ID'].count())\nprint(\"unique customer data in test data -->\", unique_customer_count_test)","metadata":{"id":"z0F5SqLMYXjV","outputId":"70bc9673-a3a9-4d96-9d77-5ed67d02ad76","execution":{"iopub.status.busy":"2022-07-02T14:00:20.618719Z","iopub.execute_input":"2022-07-02T14:00:20.619124Z","iopub.status.idle":"2022-07-02T14:00:27.211369Z","shell.execute_reply.started":"2022-07-02T14:00:20.619084Z","shell.execute_reply":"2022-07-02T14:00:27.210342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# checking single customer data\ntrain.groupby(\"customer_ID\").size()","metadata":{"id":"CR8LpGRrLjKA","outputId":"d9fa093b-e204-43ff-d788-3a43ecf3ef41","execution":{"iopub.status.busy":"2022-07-02T14:00:27.212789Z","iopub.execute_input":"2022-07-02T14:00:27.213845Z","iopub.status.idle":"2022-07-02T14:00:28.837488Z","shell.execute_reply.started":"2022-07-02T14:00:27.213806Z","shell.execute_reply":"2022-07-02T14:00:28.836266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# checking one customer data\ntrain[train[\"customer_ID\"] == \"0000099d6bd597052cdcda90ffabf56573fe9d7c79be5fbac11a8ed792feb62a\"]","metadata":{"id":"2zTh34fhMxRd","outputId":"bcd034be-e014-49e3-cad8-6c5011033b9c","execution":{"iopub.status.busy":"2022-07-02T14:00:28.839144Z","iopub.execute_input":"2022-07-02T14:00:28.839647Z","iopub.status.idle":"2022-07-02T14:00:29.942589Z","shell.execute_reply.started":"2022-07-02T14:00:28.839590Z","shell.execute_reply":"2022-07-02T14:00:29.941584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = train.groupby(\"customer_ID\")['customer_ID'].count().values\ny_test = test.groupby(\"customer_ID\")['customer_ID'].count().values","metadata":{"id":"lreexodXQWHW","execution":{"iopub.status.busy":"2022-07-02T14:00:29.943661Z","iopub.execute_input":"2022-07-02T14:00:29.943983Z","iopub.status.idle":"2022-07-02T14:00:36.738785Z","shell.execute_reply.started":"2022-07-02T14:00:29.943955Z","shell.execute_reply":"2022-07-02T14:00:36.737897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Histogram(\n    y = y,\n    ybins = dict(size = 0.5),\n    marker_color= '#9900cc'))\nfig.update_layout(\n    template = \"plotly_dark\",\n    title = \"Customer profile count -- training data\",\n    yaxis_title = \"Number of months\",\n    bargap = 0.2\n)\nfig.show()\n\nfig = go.Figure()\nfig.add_trace(go.Histogram(\n    y = y_test,\n    ybins = dict(size = 0.5),\n    marker_color= '#9900cc'))\nfig.update_layout(\n    template = \"plotly_dark\",\n    title = \"Customer profile count -- test data\",\n    yaxis_title = \"Number of months\"\n)\nfig.show()","metadata":{"id":"zKUawJWfOkLD","outputId":"da303bbf-44e0-4744-f1ee-a0b0774e9449","execution":{"iopub.status.busy":"2022-07-02T14:00:36.739933Z","iopub.execute_input":"2022-07-02T14:00:36.740257Z","iopub.status.idle":"2022-07-02T14:00:37.887650Z","shell.execute_reply.started":"2022-07-02T14:00:36.740230Z","shell.execute_reply":"2022-07-02T14:00:37.886302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- From here we can see the dsitribution of profile length is common between train and test data.\n","metadata":{"id":"DVeij0N-XCdO"}},{"cell_type":"code","source":"del y\ndel y_test\ngc.collect()","metadata":{"id":"DKW0QIoftN0h","outputId":"bc0484d2-a6e4-4dab-ef4f-31507fd7e48f","execution":{"iopub.status.busy":"2022-07-02T14:00:37.890488Z","iopub.execute_input":"2022-07-02T14:00:37.890870Z","iopub.status.idle":"2022-07-02T14:00:38.140741Z","shell.execute_reply.started":"2022-07-02T14:00:37.890836Z","shell.execute_reply":"2022-07-02T14:00:38.139746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# connection between the profile length and target output\ncount = train.groupby(\"customer_ID\")['customer_ID'].count()\ncon_check_df = pd.DataFrame({\"customer_ID\":count.index, \"count\": count.values})\n# merge the data with the label data frame\ncon_check_df = con_check_df.merge(train_labels, on='customer_ID', how='left')","metadata":{"id":"nWuizSsCQnW0","execution":{"iopub.status.busy":"2022-07-02T14:00:38.141948Z","iopub.execute_input":"2022-07-02T14:00:38.142314Z","iopub.status.idle":"2022-07-02T14:00:40.718285Z","shell.execute_reply.started":"2022-07-02T14:00:38.142283Z","shell.execute_reply":"2022-07-02T14:00:40.717318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"con_check_df.head(3)","metadata":{"id":"oWS4iUXqX7Sp","outputId":"008d9d99-cca8-4e7e-d6fc-bc48f661d3f4","execution":{"iopub.status.busy":"2022-07-02T14:00:40.719482Z","iopub.execute_input":"2022-07-02T14:00:40.719793Z","iopub.status.idle":"2022-07-02T14:00:40.730349Z","shell.execute_reply.started":"2022-07-02T14:00:40.719765Z","shell.execute_reply":"2022-07-02T14:00:40.729212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nsns.countplot(data = con_check_df,y='count',hue='target', orient='h')\n","metadata":{"id":"S9rHChR5YArg","outputId":"bdd4d75c-e9ff-45a7-b0e6-f371253035a4","execution":{"iopub.status.busy":"2022-07-02T14:00:40.731561Z","iopub.execute_input":"2022-07-02T14:00:40.732222Z","iopub.status.idle":"2022-07-02T14:00:41.176320Z","shell.execute_reply.started":"2022-07-02T14:00:40.732190Z","shell.execute_reply":"2022-07-02T14:00:41.175268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- nearly 30 - 50 % of all  profile length has target 1 (default) \n- Can't get a great correlation between profile length and target -- but thinking like keeping this information may help.","metadata":{"id":"LqQkkk3qdCgN"}},{"cell_type":"code","source":"del con_check_df\ngc.collect()","metadata":{"id":"kJb2Ueo_zToo","outputId":"bac4309e-64ce-4023-f93d-38a7fff692df","execution":{"iopub.status.busy":"2022-07-02T14:00:41.177517Z","iopub.execute_input":"2022-07-02T14:00:41.177849Z","iopub.status.idle":"2022-07-02T14:00:41.342501Z","shell.execute_reply.started":"2022-07-02T14:00:41.177819Z","shell.execute_reply":"2022-07-02T14:00:41.341198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n# <p style=\"background-color:#9900cc;padding:15px;font-family:newtimeroman;color:#ffff80;font-size:110%;border-radius:40px 5px;\"> 4 | Common data Analysis</p>","metadata":{"id":"UDDl1NVdtRaj"}},{"cell_type":"code","source":"# merge the two dataset\ntrain = train.groupby('customer_ID').tail(1).set_index('customer_ID')\ndata = train.merge(train_labels, on='customer_ID', how='left')","metadata":{"id":"6cU3S67GzQNa","execution":{"iopub.status.busy":"2022-07-02T14:00:41.344130Z","iopub.execute_input":"2022-07-02T14:00:41.344599Z","iopub.status.idle":"2022-07-02T14:00:45.403658Z","shell.execute_reply.started":"2022-07-02T14:00:41.344561Z","shell.execute_reply":"2022-07-02T14:00:45.402461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns, categorical_col, numerical_col,missing_value_df = kaggle_utils.Common_data_analysis(data, missing_value_highlight_threshold=5.0, display_df = False,\n                                                                                              only_show_missing=False)\n","metadata":{"id":"fnO6VsoDacew","outputId":"be51e004-cb91-4898-fb93-ffe4c48da009","execution":{"iopub.status.busy":"2022-07-02T14:00:53.509386Z","iopub.execute_input":"2022-07-02T14:00:53.509834Z","iopub.status.idle":"2022-07-02T14:01:05.067447Z","shell.execute_reply.started":"2022-07-02T14:00:53.509792Z","shell.execute_reply":"2022-07-02T14:01:05.066451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# by dataset defenistion descrete columns are \ndescrete_cols=['B_30', 'B_38', 'D_63', 'D_64', 'D_66', 'D_68',\n          'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'target']\n\n# so numerical columns we need to check\nnumerical_col = [c for c in numerical_col if c not in descrete_cols]\n\ntarget_col = 'target'","metadata":{"id":"yXrdJIh9vsRK","execution":{"iopub.status.busy":"2022-07-02T14:01:05.069614Z","iopub.execute_input":"2022-07-02T14:01:05.070112Z","iopub.status.idle":"2022-07-02T14:01:05.075676Z","shell.execute_reply.started":"2022-07-02T14:01:05.070064Z","shell.execute_reply":"2022-07-02T14:01:05.074761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# null value analysis\nprint(\"shape of missing value df\", missing_value_df.shape)\nmissing_value_df.head()","metadata":{"id":"WUDph3qfQkJS","outputId":"1d4c8a34-04be-4ea4-80a1-d950a999fd50","execution":{"iopub.status.busy":"2022-07-02T14:01:05.077178Z","iopub.execute_input":"2022-07-02T14:01:05.077499Z","iopub.status.idle":"2022-07-02T14:01:05.107688Z","shell.execute_reply.started":"2022-07-02T14:01:05.077469Z","shell.execute_reply":"2022-07-02T14:01:05.106812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Features having missing values -->\",missing_value_df[missing_value_df['% of Missing value(NA)'] > 0.00].shape[0])","metadata":{"id":"ATdv7qXBRGmf","outputId":"833d3126-7c77-4fb4-daee-143002f81955","execution":{"iopub.status.busy":"2022-07-02T14:01:05.109820Z","iopub.execute_input":"2022-07-02T14:01:05.110407Z","iopub.status.idle":"2022-07-02T14:01:05.117321Z","shell.execute_reply.started":"2022-07-02T14:01:05.110369Z","shell.execute_reply":"2022-07-02T14:01:05.115836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nmissing_value_df = missing_value_df[missing_value_df[\"% of Missing value(NA)\"] > 0.00]\nmissing_value_df = missing_value_df.sort_values(ascending=True, by = '% of Missing value(NA)')\nfig = go.Figure()\n\n#fig.add_trace()\n\n# add line in the chart\nfor i in range(100): # gonna have 100 missing values lines\n    #print(missing_value_df.iloc[i,3])\n    fig.add_shape(dict(type = 'line', \n                  x0= 0 , y0 = i,\n                  x1 = missing_value_df.iloc[i,3], y1 = i,\n                  line = {'color': '#9900cc', 'width' : 3}))\nfig.add_trace(go.Scatter(x = missing_value_df[\"% of Missing value(NA)\"], y = missing_value_df.index, \n                         mode='markers', \n                         marker_color='#ffff80', marker_size=8))\nfig.update_layout(template='plotly_dark',\n                  title = \"Feature with missing values :(\",\n                  xaxis = dict(title = \"Missing value percentage\", zeroline=False),\n                  yaxis_showgrid=False,\n                  width = 1000,\n                  height = 1500,)","metadata":{"id":"T8JqxfVwSS-R","outputId":"484026ef-3425-4944-860e-c3eab66586bc","execution":{"iopub.status.busy":"2022-07-02T14:01:05.119132Z","iopub.execute_input":"2022-07-02T14:01:05.119530Z","iopub.status.idle":"2022-07-02T14:01:07.085712Z","shell.execute_reply.started":"2022-07-02T14:01:05.119496Z","shell.execute_reply":"2022-07-02T14:01:07.084829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del missing_value_df\ngc.collect()","metadata":{"id":"5YitLOxUeiL4","outputId":"37a8828b-e13a-494b-b14c-2adc72d7c372","execution":{"iopub.status.busy":"2022-07-02T14:01:07.086916Z","iopub.execute_input":"2022-07-02T14:01:07.087350Z","iopub.status.idle":"2022-07-02T14:01:07.253414Z","shell.execute_reply.started":"2022-07-02T14:01:07.087317Z","shell.execute_reply":"2022-07-02T14:01:07.252561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n# <p style=\"background-color:#9900cc;padding:15px;font-family:newtimeroman;color:#ffff80;font-size:110%;border-radius:40px 5px;\"> 5 | Distribution Analysis</p>","metadata":{"id":"aZF41zzxx5-d"}},{"cell_type":"code","source":"def plot_hist(data, columns, nrow, ncol, figsize, hue_value=None):\n    # find the distubution of the data. ( visualization would be so good)\n    fig, ax = plt.subplots(nrow,ncol, figsize=figsize)\n    col, row = ncol,nrow\n    col_count = 0\n    sns.set_style('dark')\n    for r in range(row):\n        for c in range(col):\n            if col_count >= len(columns):\n                ax[r,c].text(0.5, 0.5, \"no data\")\n            else:\n                sns.kdeplot(data=data, x=columns[col_count], hue=hue_value, ax=ax[r, c], palette=['#9900cc','#99ff99'],\n                                fill = True, hue_order=[1,0], legend = True)\n                ax[r,c].set(xlabel = columns[col_count], ylabel=(\"Density\" if c==0 else ''))\n                col_count +=1\n        # print(\"col count \", col_count)\n            ","metadata":{"id":"9EcxPEnGxeoy","execution":{"iopub.status.busy":"2022-07-02T14:01:07.254593Z","iopub.execute_input":"2022-07-02T14:01:07.254974Z","iopub.status.idle":"2022-07-02T14:01:07.267336Z","shell.execute_reply.started":"2022-07-02T14:01:07.254941Z","shell.execute_reply":"2022-07-02T14:01:07.266250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Delinquency variable","metadata":{"id":"1LgjiZKUAdcK"}},{"cell_type":"code","source":"# Find the distribution of Delinquency variables\nd_feats = [c for c in d_feats if c not in descrete_cols]\nplot_hist(data, d_feats, 15, 6, (50,100),hue_value=target_col)","metadata":{"id":"fD2k1ohxyjAK","execution":{"iopub.status.busy":"2022-07-02T14:01:07.268717Z","iopub.execute_input":"2022-07-02T14:01:07.269112Z","iopub.status.idle":"2022-07-02T14:03:29.586495Z","shell.execute_reply.started":"2022-07-02T14:01:07.269075Z","shell.execute_reply":"2022-07-02T14:03:29.585134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Spend variables","metadata":{"id":"ve3IZwDqC5ld"}},{"cell_type":"code","source":"# Find the distribution of Spend variables\ns_feats = [c for c in s_feats if c not in descrete_cols]\ns_feats.remove('S_2')\nplot_hist(data, s_feats, 7, 3, (50,50),hue_value=target_col)","metadata":{"id":"qL1PIpwLy-sE","execution":{"iopub.status.busy":"2022-07-02T14:03:29.587938Z","iopub.execute_input":"2022-07-02T14:03:29.588423Z","iopub.status.idle":"2022-07-02T14:04:11.824501Z","shell.execute_reply.started":"2022-07-02T14:03:29.588381Z","shell.execute_reply":"2022-07-02T14:04:11.823531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Payment and Balance variables","metadata":{"id":"g813EEinD2De"}},{"cell_type":"code","source":"# Find the distribution of Payment variables\np_feats = [c for c in p_feats if c not in descrete_cols]\nb_feats = [c for c in b_feats if c not in descrete_cols]\nplot_hist(data, p_feats + b_feats, 11, 4, (50,50),hue_value=target_col)","metadata":{"id":"qggx8_8CB8M8","execution":{"iopub.status.busy":"2022-07-02T14:04:11.826589Z","iopub.execute_input":"2022-07-02T14:04:11.827437Z","iopub.status.idle":"2022-07-02T14:05:30.486661Z","shell.execute_reply.started":"2022-07-02T14:04:11.827401Z","shell.execute_reply":"2022-07-02T14:05:30.485784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Risk variables\n","metadata":{"id":"aNsc5J6uQR_K"}},{"cell_type":"code","source":"# Find the distribution of Payment variables\nr_feats = [c for c in r_feats if c not in descrete_cols]\nplot_hist(data, r_feats, 8, 4, (50,50),hue_value=target_col)","metadata":{"id":"Hmtb_jfcEGP0","execution":{"iopub.status.busy":"2022-07-02T14:05:30.487816Z","iopub.execute_input":"2022-07-02T14:05:30.488403Z","iopub.status.idle":"2022-07-02T14:06:26.084158Z","shell.execute_reply.started":"2022-07-02T14:05:30.488352Z","shell.execute_reply":"2022-07-02T14:06:26.083059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Can't see any feature following normal distribution \n- And all features having different distributions\n- We can't use parameterised models -- best go for some non-parameterised models","metadata":{"id":"EjwiBthgeO-N"}},{"cell_type":"markdown","source":"# <p style=\"background-color:#9900cc;padding:15px;font-family:newtimeroman;color:#ffff80;font-size:110%;border-radius:40px 5px;\"> 6 | Correlation Analysis\n</p>","metadata":{"id":"5FAFME0Lest9"}},{"cell_type":"code","source":"# correlation with target\n#col = [c for c in data.columns if data[c].dtypes != 'object']\n\ncorr = data.corrwith(data[target_col], axis=0)\nval = [str(round(v ,2) *100) + '%' for v in corr.values]\n\nfig = go.Figure()\nfig.add_trace(go.Bar(y=corr.index, x= corr.values,\n                     orientation='h',\n                     marker_color = '#9900cc',\n                     text = val,\n                     textposition = 'outside',\n                     textfont_color = '#ffff80'))\nfig.update_layout(template = 'plotly_dark',\n                  title = \"Correlation with Target\",\n                  width = 800,\n                  height = 3000)\nfig.update_xaxes(range=[-2,2])","metadata":{"id":"On2jYqu8QhB5","outputId":"036e83a7-c493-41d1-8059-3eb007e7f33b","execution":{"iopub.status.busy":"2022-07-02T14:06:35.865215Z","iopub.execute_input":"2022-07-02T14:06:35.865637Z","iopub.status.idle":"2022-07-02T14:06:37.741767Z","shell.execute_reply.started":"2022-07-02T14:06:35.865605Z","shell.execute_reply":"2022-07-02T14:06:37.740792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del val,corr\ngc.collect()","metadata":{"id":"ucV_7WrKzxki","outputId":"88569767-adf9-4636-a1ba-3d7e1e366613","execution":{"iopub.status.busy":"2022-07-02T14:06:37.743225Z","iopub.execute_input":"2022-07-02T14:06:37.743572Z","iopub.status.idle":"2022-07-02T14:06:38.004795Z","shell.execute_reply.started":"2022-07-02T14:06:37.743541Z","shell.execute_reply":"2022-07-02T14:06:38.003671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"there are lots of values correlated with target, B_2 is the most negatively correlated feature with -0.56 correlation. At the same time most positively correlated feature is D_68 with 0.61 correlation\n","metadata":{"id":"K16CS6V5wfE3"}},{"cell_type":"markdown","source":"# <p style=\"background-color:#9900cc;padding:15px;font-family:newtimeroman;color:#ffff80;font-size:110%;border-radius:40px 5px;\"> 7 | Target value distibution</p>","metadata":{"id":"-4jPzznhfr5G"}},{"cell_type":"code","source":"# plot the target\ncount = data[target_col].value_counts()\nprint(count)\nprint(\"percentage of first class --- >\",count[0]/data.shape[0])\nprint(\"percentage of second class --->\", count[1]/data.shape[0])","metadata":{"id":"IwABpD9YfySa","execution":{"iopub.status.busy":"2022-07-02T14:06:38.628519Z","iopub.execute_input":"2022-07-02T14:06:38.629345Z","iopub.status.idle":"2022-07-02T14:06:38.648585Z","shell.execute_reply.started":"2022-07-02T14:06:38.629298Z","shell.execute_reply":"2022-07-02T14:06:38.647524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure()\nfig.add_trace(go.Bar(x= ['Paid', \"Default\"],y=count.values,\n                     marker_color = ['#9900cc','#ffff80'],\n                     text = [str(round(count[0]/data.shape[0],2) * 100) + '%' , str(round(count[1]/data.shape[0], 2) * 100) + '%']))\nfig.update_layout(template = 'plotly_dark',\n                  title = \"target value distribution\",\n                  width = 500,\n                  height = 500)","metadata":{"id":"trRB0E_ZjW5M","execution":{"iopub.status.busy":"2022-07-02T14:06:39.723849Z","iopub.execute_input":"2022-07-02T14:06:39.725151Z","iopub.status.idle":"2022-07-02T14:06:39.768565Z","shell.execute_reply.started":"2022-07-02T14:06:39.725110Z","shell.execute_reply":"2022-07-02T14:06:39.767633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# delete some unwanted variables\ndel d_feats,s_feats,p_feats,b_feats,r_feats\ngc.collect()","metadata":{"id":"qPV3Q18vz2A3","execution":{"iopub.status.busy":"2022-07-02T14:06:40.080743Z","iopub.execute_input":"2022-07-02T14:06:40.081184Z","iopub.status.idle":"2022-07-02T14:06:40.245224Z","shell.execute_reply.started":"2022-07-02T14:06:40.081148Z","shell.execute_reply":"2022-07-02T14:06:40.244280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"background-color:#9900cc;padding:15px;font-family:newtimeroman;color:#ffff80;font-size:110%;border-radius:40px 5px;\"> 8 | Prediction</p>","metadata":{"id":"mp9zWJtMzeL5"}},{"cell_type":"code","source":"# del categorical data\nneeded_col = [c for c in data.columns if c not in ['customer_ID','S_2']]\ndata = data[needed_col]\ntest.drop('S_2', inplace = True, axis = 1)\ntest.drop('customer_ID', inplace = True, axis = 1)","metadata":{"id":"CWiEF--KzhcV","execution":{"iopub.status.busy":"2022-07-02T14:06:41.716639Z","iopub.execute_input":"2022-07-02T14:06:41.717083Z","iopub.status.idle":"2022-07-02T14:07:01.600233Z","shell.execute_reply.started":"2022-07-02T14:06:41.717022Z","shell.execute_reply":"2022-07-02T14:07:01.599167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X=data.drop(['target'],axis=1)\ny=data['target']\n\ndel data\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-02T14:07:01.602022Z","iopub.execute_input":"2022-07-02T14:07:01.602413Z","iopub.status.idle":"2022-07-02T14:07:02.078075Z","shell.execute_reply.started":"2022-07-02T14:07:01.602378Z","shell.execute_reply":"2022-07-02T14:07:02.077244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)","metadata":{"id":"Qdnp6e8zzj-R","execution":{"iopub.status.busy":"2022-07-02T14:07:02.079279Z","iopub.execute_input":"2022-07-02T14:07:02.079621Z","iopub.status.idle":"2022-07-02T14:07:02.094583Z","shell.execute_reply.started":"2022-07-02T14:07:02.079590Z","shell.execute_reply":"2022-07-02T14:07:02.093438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # parameter tuning\n# def objective(trial, X, y):\n#     param = {\n#         \"n_estimators\": trial.suggest_int(\"n_estimators\", 5000,20000,step=10),\n#         'learning_rate' : trial.suggest_uniform('learning_rate',0.01, 0.1),\n#         \"lambda_l1\": trial.suggest_loguniform(\"reg_lambda\", 1.0, 50.0), # L1 regularization parameter\n#         \"lambda_l2\": trial.suggest_loguniform(\"lambda_l2\", 1.0, 50.0),\n#         \"max_depth\": trial.suggest_int(\"max_depth\", 5, 20), # max depth of the tree\n#         \"num_leaves\": trial.suggest_int(\"max_depth\", 30, 2000, step=10),\n#         \"subsample\": trial.suggest_loguniform(\"subsample\", 0.1, 1.0), #Denotes the fraction of observations to be randomly samples for each tree.\n#         \"bagging_fraction\": trial.suggest_loguniform(\"subsample\", 0.1, 1.0),\n#         \"bagging_freq\": trial.suggest_int(\"bagging_freq\", 0,10),\n#         \"feature_fraction\": trial.suggest_float(\"feature_fraction\", 0.2, 0.95, step=0.1),\n#         \"min_data_in_leaf\": trial.suggest_int(\"bagging_freq\", 20,200),\n#         \"boosting_type\":'gbdt',\n        \n#         \"colsample_bytree\": trial.suggest_float(\"feature_fraction\", 0.2, 0.95, step=0.1),\n#         \"max_bins\": trial.suggest_int(\"max_depth\", 30, 2000, step=10),\n#         \"objective\": \"binary\",\n#         \"random_state\": 23,\n#         \"max_bin\": 500\n#     }\n\n#     # cross - validation\n#     cv = StratifiedKFold(n_splits=2, shuffle=True, random_state=32)\n#     cross_val_score = []\n#     for fold_index, (train_id, val_id) in enumerate(cv.split(X,y)):\n#         # get the train and val set for this cross validation\n#         print(\"=\"*20, end=\" \")\n#         print(\"Fold \", fold_index, end = \" \")\n#         print(\"=\"*20, )\n#         X_train, X_val = X.iloc[train_id], X.iloc[val_id]\n#         y_train, y_val = y[train_id], y[val_id]\n\n#         # define the model\n#         model = LGBMClassifier(**param)\n#         # fit the model\n#         model.fit(X_train, y_train, eval_set=[(X_val, y_val)], early_stopping_rounds=100, verbose=200, eval_metric=[\"auc\"])\n\n#         # predict \n#         y_pred = model.predict_proba(X_val)[:,1]\n        \n#         y_pred=pd.DataFrame(data={'prediction':y_pred})\n#         y_true=pd.DataFrame(data={'target':y_val.reset_index(drop=True)})\n#         gini_score=amex_metric(y_true = y_true, y_pred = y_pred)\n    \n#         cross_val_score.append(roc_auc_score(y_val, y_pred))\n        \n#         del X_train, X_val,y_train, y_val\n#         gc.collect()\n        \n#     return np.mean(np.array(cross_val_score))\n\n\n# # strat the study\n# study = optuna.create_study(study_name=\"LGBM classifier\", direction=\"maximize\")\n# fun = lambda trial: objective(trial, X, y)\n# study.optimize(fun, n_trials=50)","metadata":{"id":"8Xmqf4P20QiF","execution":{"iopub.status.busy":"2022-07-02T14:07:02.096173Z","iopub.execute_input":"2022-07-02T14:07:02.096476Z","iopub.status.idle":"2022-07-02T14:07:02.112331Z","shell.execute_reply.started":"2022-07-02T14:07:02.096449Z","shell.execute_reply":"2022-07-02T14:07:02.111360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# best param\n# Trial 26 finished with value: 0.9596085950724972 and parameters: {'n_estimators': 13470, 'learning_rate': 0.03334140860049416, 'reg_lambda': 3.7629073371138517, 'lambda_l2': 24.60526923347014, 'max_depth': 16, 'subsample': 0.3297124328760659, 'bagging_freq': 3, 'feature_fraction': 0.2}. Best is trial 17 with value: 0.9601344000765331.\nbest_param= {\"n_estimators\":1500,\n            \"learning_rate\":0.04,\n            #\"lambda_l2\":24.60526923347014,\n            \"max_depth\":16,\n            \"subsample\":0.32,\n             \"bagging_freq\": 3,\n             #\"feature_fraction\":0.2,\n             \"random_state\": 37,\n             \"boosting_type\":'gbdt',\n             \"min_child_samples\": 2000,\n             'objective': 'binary'\n            }","metadata":{"id":"6jWZStl70VWF","execution":{"iopub.status.busy":"2022-07-02T14:07:02.113611Z","iopub.execute_input":"2022-07-02T14:07:02.113948Z","iopub.status.idle":"2022-07-02T14:07:02.127321Z","shell.execute_reply.started":"2022-07-02T14:07:02.113918Z","shell.execute_reply":"2022-07-02T14:07:02.126533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# prediction\ngbm_test_preds, gini=[],[]\nft_importance=pd.DataFrame(index=X.columns)\n# cross - validation\ncv = StratifiedKFold(n_splits=10, shuffle=True, random_state=32)\ncross_val_score = []\nfor fold_index, (train_id, val_id) in enumerate(cv.split(X,y)):\n    # get the train and val set for this cross validation\n    print(\"=\"*20, end=\" \")\n    print(\"Fold \", fold_index, end = \" \")\n    print(\"=\"*20, )\n    X_train, X_val = X.iloc[train_id], X.iloc[val_id]\n    y_train, y_val = y[train_id], y[val_id]\n\n    # define the model\n    model = LGBMClassifier(**best_param)\n    # fit the model\n    model.fit(X_train, y_train, eval_set=[(X_val, y_val)], early_stopping_rounds=100, verbose=200, eval_metric=[\"auc\"])\n\n    # predict \n    y_pred = model.predict_proba(X_val)[:,1]\n\n    y_pred=pd.DataFrame(data={'prediction':y_pred})\n    y_true=pd.DataFrame(data={'target':y_val.reset_index(drop=True)})\n    gini_score=amex_metric(y_true = y_true, y_pred = y_pred)\n\n    cross_val_score.append(roc_auc_score(y_val, y_pred))\n    print(\"Gini score {} --- cross validation score {}\".format(gini_score,cross_val_score))\n    \n    # gbm_test_preds.append(model.predict_proba(test)[:,1])\n    \n\n    del X_train, X_val,y_train, y_val\n    gc.collect()\ndel X, y\ngc.collect()","metadata":{"id":"JIKe3CQD0XX2","execution":{"iopub.status.busy":"2022-07-02T14:07:02.128488Z","iopub.execute_input":"2022-07-02T14:07:02.128915Z","iopub.status.idle":"2022-07-02T14:24:05.437391Z","shell.execute_reply.started":"2022-07-02T14:07:02.128884Z","shell.execute_reply":"2022-07-02T14:24:05.436389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}