{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7602123,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1><center><font size=\"6\">Home Credit</font></center></h1>\n<h2><center><font size=\"6\">Credit Risk Model Stability</font></center></h2>\n\n<div align=\"center\">\n    <img src=\"https://logowik.com/content/uploads/images/home-credit706.logowik.com.webp\" width=\"600\">\n  </div>\n<br>\n\n> 📌 **Competition Scope**: The goal of this competition is to predict which clients are more likely to default on their loans. The evaluation will favor solutions that are stable over time.\n\n# <a id='0'>Content</a>\n- <a href='#1'>Introduction</a>  \n- <a href='#2'>Import libraries and load data</a>  \n- <a href='#3'>Data exploration</a>\n - <a href='#31'>Check the data</a>\n - <a href='#32'>static_0</a>\n     - <a href='#33'>Numerical Features</a>  \n     - <a href='#34'>Categorical Features</a>  \n - <a href='#35'>static_cb_0</a>\n     - <a href='#36'>Numerical Features</a>  \n     - <a href='#37'>Categorical Features</a>  \n- <a href='#38'>Next steps</a>","metadata":{}},{"cell_type":"markdown","source":"# <a id='1'>Introduction</a>  \n\n### About the data analized in this notebook:\n\nThere is a **lot** of data,so in this notebook we will only focus on depth = 0 data sources, that is:\n\n* `train_base.csv`\n\nThen we will merge the data coming from internal sources:\n* `train_static_0_0.csv`\n* `train_static_0_1.csv`\n\nAnd finally we will look at the data coming from external sources:\n* `train_static_cb_0.csv` ","metadata":{}},{"cell_type":"markdown","source":"# <a id='2'>Import libraries and load data</a>  \n\n## Import libraries and packages","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nfrom tqdm.auto import tqdm\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T17:56:44.870636Z","iopub.execute_input":"2024-02-20T17:56:44.871066Z","iopub.status.idle":"2024-02-20T17:56:47.838507Z","shell.execute_reply.started":"2024-02-20T17:56:44.871026Z","shell.execute_reply":"2024-02-20T17:56:47.835355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_base = pd.read_csv('/kaggle/input/home-credit-credit-risk-model-stability/csv_files/train/train_base.csv')\npd.set_option('display.max_columns', 1000)","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:51:00.531365Z","iopub.execute_input":"2024-02-20T19:51:00.531817Z","iopub.status.idle":"2024-02-20T19:51:02.141034Z","shell.execute_reply.started":"2024-02-20T19:51:00.531784Z","shell.execute_reply":"2024-02-20T19:51:02.139893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a id='3'>Data exploration</a>  \n\n## <a id='31'>Check the data</a>  \n\nLet's check the train base file first:","metadata":{}},{"cell_type":"code","source":"(train_base.info())","metadata":{"execution":{"iopub.status.busy":"2024-02-20T17:56:49.198294Z","iopub.execute_input":"2024-02-20T17:56:49.199247Z","iopub.status.idle":"2024-02-20T17:56:49.316516Z","shell.execute_reply.started":"2024-02-20T17:56:49.199199Z","shell.execute_reply":"2024-02-20T17:56:49.315112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's keep in mind how many case_id we have for training:","metadata":{}},{"cell_type":"code","source":"(train_base\n .case_id\n .nunique())","metadata":{"execution":{"iopub.status.busy":"2024-02-20T17:56:49.319012Z","iopub.execute_input":"2024-02-20T17:56:49.319381Z","iopub.status.idle":"2024-02-20T17:56:49.405666Z","shell.execute_reply.started":"2024-02-20T17:56:49.319344Z","shell.execute_reply":"2024-02-20T17:56:49.404653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's take a look at the distribution of the target","metadata":{}},{"cell_type":"code","source":"# Plot the distribution of the binary column using seaborn\nax = sns.countplot(x='target', data=train_base)\n\n# Add count labels on top of each bar\nfor p in ax.patches:\n    ax.annotate(format(p.get_height(), '.0f'), \n                   (p.get_x() + p.get_width() / 2., p.get_height()), \n                   ha = 'center', va = 'center', \n                   xytext = (0, 5), \n                   textcoords = 'offset points')\n\n# Add labels and title\nplt.xlabel('Target')\nplt.ylabel('Count')\nplt.title('Distribution of Target')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-20T17:56:49.406887Z","iopub.execute_input":"2024-02-20T17:56:49.407834Z","iopub.status.idle":"2024-02-20T17:56:49.921604Z","shell.execute_reply.started":"2024-02-20T17:56:49.407801Z","shell.execute_reply":"2024-02-20T17:56:49.920539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So it is pretty unbalanced, wich is kind of expected for the type of problem that we are facing. We just need to remember this for when we start modeling, specially for validation.","metadata":{}},{"cell_type":"markdown","source":"## <a id='32'>static_0</a>  \n\nLet's see the first static data: \n\n* train_static_0_0.csv  \n* train_static_0_1.csv \n\nboth files contains the same data, it just happens that they separated the data to make it more manageable. So let's load them and put them together:","metadata":{}},{"cell_type":"code","source":"train_static_0_0 = pd.read_csv('/kaggle/input/home-credit-credit-risk-model-stability/csv_files/train/train_static_0_0.csv')\ntrain_static_0_1 = pd.read_csv('/kaggle/input/home-credit-credit-risk-model-stability/csv_files/train/train_static_0_1.csv')\ntrain_static = pd.concat([train_static_0_0, train_static_0_1], axis=0)\ntrain_static.reset_index(drop=True, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:45:50.011231Z","iopub.execute_input":"2024-02-20T19:45:50.011708Z","iopub.status.idle":"2024-02-20T19:46:25.866172Z","shell.execute_reply.started":"2024-02-20T19:45:50.011673Z","shell.execute_reply":"2024-02-20T19:46:25.864724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a quick look at the info of the data:","metadata":{}},{"cell_type":"code","source":"train_static.info()","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:48:45.618773Z","iopub.execute_input":"2024-02-20T19:48:45.619235Z","iopub.status.idle":"2024-02-20T19:48:45.633679Z","shell.execute_reply.started":"2024-02-20T19:48:45.619201Z","shell.execute_reply":"2024-02-20T19:48:45.632491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It occuppies too much memory, so let's optimize the memory usage.","metadata":{}},{"cell_type":"code","source":"# taked from: https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage?scriptVersionId=3684066&cellId=3\n\ndef reduce_mem_usage(df):\n    \"\"\" iterate through all the columns of a dataframe and modify the data type\n        to reduce memory usage.        \n    \"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n    \n    for col in df.columns:\n        col_type = df[col].dtype\n        \n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            else:\n                # this was giving me some problem (np.float16) to plot some of the variables\n                #if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                #    df[col] = df[col].astype(np.float16)\n                if c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n                    \n        elif col == \"date_decision\":\n            df[col] = pd.to_datetime(df[col])\n            \n        else:\n            df[col] = df[col].astype('category')\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2024-02-20T21:08:41.241307Z","iopub.execute_input":"2024-02-20T21:08:41.241812Z","iopub.status.idle":"2024-02-20T21:08:41.256302Z","shell.execute_reply.started":"2024-02-20T21:08:41.241775Z","shell.execute_reply":"2024-02-20T21:08:41.254977Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We optimize the file `train_base.csv` first:","metadata":{}},{"cell_type":"code","source":"train_base = reduce_mem_usage(train_base)","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:51:20.642279Z","iopub.execute_input":"2024-02-20T19:51:20.643297Z","iopub.status.idle":"2024-02-20T19:51:20.869249Z","shell.execute_reply.started":"2024-02-20T19:51:20.643255Z","shell.execute_reply":"2024-02-20T19:51:20.867785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then, we optimize `train_static`:","metadata":{}},{"cell_type":"code","source":"train_static = reduce_mem_usage(train_static)","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:50:12.148759Z","iopub.execute_input":"2024-02-20T19:50:12.149671Z","iopub.status.idle":"2024-02-20T19:50:19.122579Z","shell.execute_reply.started":"2024-02-20T19:50:12.149619Z","shell.execute_reply":"2024-02-20T19:50:19.120890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def missing_data(data):\n    total = data.isnull().sum()\n    percent = (data.isnull().sum()/data.isnull().count()*100)\n    tt = pd.concat([total, percent], axis=1, keys=['Total', 'Percent'])\n    types = []\n    for col in data.columns:\n        dtype = str(data[col].dtype)\n        types.append(dtype)\n    tt['Types'] = types\n    return(np.transpose(tt))","metadata":{"execution":{"iopub.status.busy":"2024-02-20T17:57:26.133719Z","iopub.execute_input":"2024-02-20T17:57:26.134063Z","iopub.status.idle":"2024-02-20T17:57:26.140203Z","shell.execute_reply.started":"2024-02-20T17:57:26.134036Z","shell.execute_reply":"2024-02-20T17:57:26.139389Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a look at the missing data:","metadata":{}},{"cell_type":"code","source":"missing_data(train_static)","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:51:34.607401Z","iopub.execute_input":"2024-02-20T19:51:34.608232Z","iopub.status.idle":"2024-02-20T19:51:35.863507Z","shell.execute_reply.started":"2024-02-20T19:51:34.608191Z","shell.execute_reply":"2024-02-20T19:51:35.862041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Despite that this comes from an internal source, there is a lot of features with big percentages of missing values. \n\nNow, let's put the together `train_base.csv` with `train_static` together.","metadata":{}},{"cell_type":"code","source":"train = train_base.merge(train_static,how='left', on='case_id')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:52:27.649602Z","iopub.execute_input":"2024-02-20T19:52:27.650103Z","iopub.status.idle":"2024-02-20T19:52:31.735483Z","shell.execute_reply.started":"2024-02-20T19:52:27.650064Z","shell.execute_reply":"2024-02-20T19:52:31.734177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check that everything looks ok in terms of numbers of `case_id` (1526659):","metadata":{}},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:52:31.737793Z","iopub.execute_input":"2024-02-20T19:52:31.738168Z","iopub.status.idle":"2024-02-20T19:52:31.803910Z","shell.execute_reply.started":"2024-02-20T19:52:31.738137Z","shell.execute_reply":"2024-02-20T19:52:31.802786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_to_exclude = ['case_id', 'MONTH', 'WEEK_NUM', 'target']\nnumerical_cols = [col for col in train.columns if (train[col].dtype not in ['datetime64[ns]', 'category']) and (col not in cols_to_exclude)]\nprint(len(numerical_cols))","metadata":{"execution":{"iopub.status.busy":"2024-02-20T21:18:47.783467Z","iopub.execute_input":"2024-02-20T21:18:47.784670Z","iopub.status.idle":"2024-02-20T21:18:47.794772Z","shell.execute_reply.started":"2024-02-20T21:18:47.784612Z","shell.execute_reply":"2024-02-20T21:18:47.793520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just 129 numerical columns...\n\n<center><img src=\"https://uploads.dailydot.com/2023/12/ralph-wiggum-im-in-danger-meme.jpg?q=65&auto=format&w=1600&ar=2:1&fit=crop\" width=500></center>","metadata":{}},{"cell_type":"markdown","source":"### Let's do some plotting\n\nI have only see EDA notebooks that plot distributions of the features and some boxplots, but didn't take in account the time aspect. So i made this plots in order to better see how some features change in time.\n\nInspirations: \n\n* The extensive EDA in this [notebook](https://www.kaggle.com/code/taichiuemura/home-credit-eda?kernelSessionId=162821340) from @taichiuemura.\n* This great [notebook](https://www.kaggle.com/code/thomasmeiner/home-credit-eda-feature-engineering-modelling?kernelSessionId=162011344) from @thomasmeiner.","metadata":{}},{"cell_type":"code","source":"def aggregated_plot(aggregated_column, plotting_column, df_, title=None):\n    df_viz = df_.groupby(aggregated_column)[plotting_column].mean()\n    ax = sns.lineplot(data=df_viz)\n    ax.set_xticklabels(ax.get_xticklabels(), rotation=45, ha='right')  # Rotate x-axis labels by 45 degrees\n    plt.tight_layout()  # Adjust layout to prevent overlap of labels\n    if title:\n        ax.set_title(title)  # Set the title if provided\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:53:22.747791Z","iopub.execute_input":"2024-02-20T19:53:22.748263Z","iopub.status.idle":"2024-02-20T19:53:22.756146Z","shell.execute_reply.started":"2024-02-20T19:53:22.748222Z","shell.execute_reply":"2024-02-20T19:53:22.754861Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is the file with the description of every feature, we will use it to put a title to every plot.","metadata":{}},{"cell_type":"code","source":"features_definitions = pd.read_csv('/kaggle/input/home-credit-credit-risk-model-stability/feature_definitions.csv')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:53:23.465494Z","iopub.execute_input":"2024-02-20T19:53:23.466013Z","iopub.status.idle":"2024-02-20T19:53:23.475073Z","shell.execute_reply.started":"2024-02-20T19:53:23.465955Z","shell.execute_reply":"2024-02-20T19:53:23.474079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a id='33'>Numerical Features</a>  ","metadata":{}},{"cell_type":"markdown","source":"**Beware**: here it comes 129 plots, i didn't went one from one making observations or annotations, this is just for having it there when modeling and coming back and maybe see or discover some quick pattern.","metadata":{}},{"cell_type":"code","source":"for col in numerical_cols:\n    description = features_definitions.loc[features_definitions['Variable'] == col, 'Description'].values[0]\n    print(f'Variable: {description} \\n')\n    aggregated_plot('date_decision', col, train)\n    print('\\n\\n\\n')","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-02-20T19:53:59.189999Z","iopub.execute_input":"2024-02-20T19:53:59.190429Z","iopub.status.idle":"2024-02-20T19:54:44.997826Z","shell.execute_reply.started":"2024-02-20T19:53:59.190396Z","shell.execute_reply":"2024-02-20T19:54:44.996673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is probably some of this features that i didnt have to plot against `date_decision`, but instead with some other date feature, but this is a work in progress, i just started and didnt have time to go look at the 129 feature descriptions.","metadata":{}},{"cell_type":"markdown","source":"## <a id='34'>Categorical features</a>  ","metadata":{}},{"cell_type":"markdown","source":"I will exclude the date features and some of the largest ones.","metadata":{}},{"cell_type":"code","source":"date_cols_to_exclude = ['datefirstoffer_1144D', 'datelastinstal40dpd_247D', 'datelastunpaid_3546854D', 'dtlastpmtallstes_4499206D',\n                  'firstclxcampaign_1125D', 'firstdatedue_489D', 'lastactivateddate_801D', 'lastapplicationdate_877D',\n                  'lastapprdate_640D', 'lastdelinqdate_224D', 'lastrejectdate_50D', 'lastrepayingdate_696D', 'maxdpdinstldate_3546855D',\n                  'payvacationpostpone_4187118D', 'validfrom_1069D']\nlarge_cols_to_exclude = ['lastapprcommoditycat_1041M', 'lastcancelreason_561M', 'lastrejectcommoditycat_161M', \n                         'lastrejectreason_759M', 'lastrejectreasonclient_4145040M', 'previouscontdistrict_112M',\n                        'lastapprcommoditytypec_5251766M', 'lastrejectcommodtypec_5251769M']\ncategorical_cols_to_plot = [col for col in train.columns if (train[col].dtype == 'category') and (col not in large_cols_to_exclude) and (col not in date_cols_to_exclude)]","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:03:49.881862Z","iopub.execute_input":"2024-02-20T20:03:49.882323Z","iopub.status.idle":"2024-02-20T20:03:49.891684Z","shell.execute_reply.started":"2024-02-20T20:03:49.882288Z","shell.execute_reply":"2024-02-20T20:03:49.890591Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_count_pairs(data_df, feature, title, hue='set'):\n    f, ax = plt.subplots(1,1,figsize = (8,4))\n    sns.countplot(x = feature, data = data_df, hue=hue)\n    plt.grid(color = \"black\", linestyle = \"-.\", linewidth = 0.5, axis=\"y\", which = \"major\")\n    ax.set_title(f\"{title}\")\n    # Adjust the position of the legend to the right and a bit higher\n    plt.legend(loc='upper right', bbox_to_anchor=(1.31, 1))\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-20T19:55:02.653442Z","iopub.execute_input":"2024-02-20T19:55:02.654556Z","iopub.status.idle":"2024-02-20T19:55:02.660954Z","shell.execute_reply.started":"2024-02-20T19:55:02.654515Z","shell.execute_reply":"2024-02-20T19:55:02.659917Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in categorical_cols_to_plot:\n    description = features_definitions.loc[features_definitions['Variable'] == col, 'Description'].values[0]\n    plot_count_pairs(train, col, description, hue='target')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:03:52.866684Z","iopub.execute_input":"2024-02-20T20:03:52.867161Z","iopub.status.idle":"2024-02-20T20:03:57.222386Z","shell.execute_reply.started":"2024-02-20T20:03:52.867124Z","shell.execute_reply":"2024-02-20T20:03:57.221145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Categorical Features that are too big","metadata":{}},{"cell_type":"code","source":"def big_plot_count_pairs(data_df, feature, title, hue='set'):\n    f, ax = plt.subplots(1,1,figsize = (8,4))\n    sns.countplot(x = feature, data = data_df, hue=hue)\n    plt.grid(color = \"black\", linestyle = \"-.\", linewidth = 0.5, axis=\"y\", which = \"major\")\n    ax.set_xticklabels(ax.get_xticklabels(), rotation=90, ha='right')  # Rotate x-axis labels by 45 degrees\n    ax.set_title(f\"{title}\")\n    # Adjust the position of the legend to the right and a bit higher\n    plt.legend(loc='upper right', bbox_to_anchor=(1.31, 1))\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-20T18:30:19.258394Z","iopub.execute_input":"2024-02-20T18:30:19.259236Z","iopub.status.idle":"2024-02-20T18:30:19.266480Z","shell.execute_reply.started":"2024-02-20T18:30:19.259196Z","shell.execute_reply":"2024-02-20T18:30:19.265295Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in large_cols_to_exclude:\n    description = features_definitions.loc[features_definitions['Variable'] == col, 'Description'].values[0]\n    big_plot_count_pairs(train, col, description, hue='target')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:04:23.663379Z","iopub.execute_input":"2024-02-20T20:04:23.664413Z","iopub.status.idle":"2024-02-20T20:04:35.777258Z","shell.execute_reply.started":"2024-02-20T20:04:23.664365Z","shell.execute_reply":"2024-02-20T20:04:35.776301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a id='35'>static_cb_0</a>  \n\n\nLet's look now at the other static data `static_cb_0`, this one comes from an external data source.","metadata":{}},{"cell_type":"code","source":"static_cb_0 = pd.read_csv('/kaggle/input/home-credit-credit-risk-model-stability/csv_files/train/train_static_cb_0.csv')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:07:44.182677Z","iopub.execute_input":"2024-02-20T20:07:44.183298Z","iopub.status.idle":"2024-02-20T20:07:52.855851Z","shell.execute_reply.started":"2024-02-20T20:07:44.183258Z","shell.execute_reply":"2024-02-20T20:07:52.854460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Reducing his memory use:","metadata":{}},{"cell_type":"code","source":"static_cb_0 = reduce_mem_usage(static_cb_0)","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:08:28.080425Z","iopub.execute_input":"2024-02-20T20:08:28.080887Z","iopub.status.idle":"2024-02-20T20:08:30.467418Z","shell.execute_reply.started":"2024-02-20T20:08:28.080850Z","shell.execute_reply":"2024-02-20T20:08:30.466258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How many missing data each feature has?","metadata":{}},{"cell_type":"code","source":"missing_data(static_cb_0)","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:08:41.778669Z","iopub.execute_input":"2024-02-20T20:08:41.779139Z","iopub.status.idle":"2024-02-20T20:08:42.124664Z","shell.execute_reply.started":"2024-02-20T20:08:41.779103Z","shell.execute_reply":"2024-02-20T20:08:42.123565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are gonna put together the `train_base.csv` with the `static_cb_0.csv` in order to plot. ","metadata":{}},{"cell_type":"code","source":"train_static_external = train_base.merge(static_cb_0,how='left', on='case_id')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:09:35.252692Z","iopub.execute_input":"2024-02-20T20:09:35.253251Z","iopub.status.idle":"2024-02-20T20:09:35.946847Z","shell.execute_reply.started":"2024-02-20T20:09:35.253189Z","shell.execute_reply":"2024-02-20T20:09:35.945880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a id='36'>Numerical Features</a>  ","metadata":{}},{"cell_type":"code","source":"cols_to_exclude = ['case_id', 'MONTH', 'WEEK_NUM', 'target']\nnumerical_cols = [col for col in train_static_external.columns if (train_static_external[col].dtype not in ['datetime64[ns]', 'category']) and (col not in cols_to_exclude)]","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:10:52.054145Z","iopub.execute_input":"2024-02-20T20:10:52.054577Z","iopub.status.idle":"2024-02-20T20:10:52.061335Z","shell.execute_reply.started":"2024-02-20T20:10:52.054545Z","shell.execute_reply":"2024-02-20T20:10:52.060219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in numerical_cols:\n    description = features_definitions.loc[features_definitions['Variable'] == col, 'Description'].values[0]\n    print(f'Variable: {description} \\n')\n    aggregated_plot('date_decision', col, train_static_external)\n    print('\\n\\n\\n')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:11:22.897572Z","iopub.execute_input":"2024-02-20T20:11:22.898251Z","iopub.status.idle":"2024-02-20T20:11:35.741790Z","shell.execute_reply.started":"2024-02-20T20:11:22.898210Z","shell.execute_reply":"2024-02-20T20:11:35.740663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a id='37'>Categorical features</a>  ","metadata":{}},{"cell_type":"code","source":"date_cols_to_exclude = ['assignmentdate_238D', 'assignmentdate_4527235D', 'assignmentdate_4955616D', 'birthdate_574D', 'dateofbirth_337D',\n                       'dateofbirth_342D', 'responsedate_1012D', 'responsedate_4527233D', 'responsedate_4917613D']\nlarge_cols_to_exclude = ['riskassesment_302T']\ncategorical_cols_to_plot = [col for col in train_static_external.columns if (train_static_external[col].dtype == 'category') and (col not in large_cols_to_exclude) and (col not in date_cols_to_exclude)]","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:25:45.116912Z","iopub.execute_input":"2024-02-20T20:25:45.117940Z","iopub.status.idle":"2024-02-20T20:25:45.123916Z","shell.execute_reply.started":"2024-02-20T20:25:45.117901Z","shell.execute_reply":"2024-02-20T20:25:45.123173Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in categorical_cols_to_plot:\n    description = features_definitions.loc[features_definitions['Variable'] == col, 'Description'].values[0]\n    plot_count_pairs(train_static_external, col, description, hue='target')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:25:45.242096Z","iopub.execute_input":"2024-02-20T20:25:45.242775Z","iopub.status.idle":"2024-02-20T20:25:49.301312Z","shell.execute_reply.started":"2024-02-20T20:25:45.242736Z","shell.execute_reply":"2024-02-20T20:25:49.300083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in large_cols_to_exclude:\n    description = features_definitions.loc[features_definitions['Variable'] == col, 'Description'].values[0]\n    big_plot_count_pairs(train_static_external, col, description, hue='target')","metadata":{"execution":{"iopub.status.busy":"2024-02-20T20:26:10.576533Z","iopub.execute_input":"2024-02-20T20:26:10.577022Z","iopub.status.idle":"2024-02-20T20:26:10.996149Z","shell.execute_reply.started":"2024-02-20T20:26:10.576969Z","shell.execute_reply":"2024-02-20T20:26:10.994964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a id='38'>Next steps</a>  \n\n* Train a baseline model.\n* Expand this data exploration based on findings with the baseline and some feature importance.\n* Do some depth = 1 EDA.\n\n\n<img src=\"https://media.makeameme.org/created/something-is-coming-b69bb58577.jpg\" width=500>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}