{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7602123,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TABLE STATIC ANALYSIS","metadata":{}},{"cell_type":"markdown","source":"This table is composed by 3 tables:\n\n- <code>train_static_0_0</code>, <code>train_static_0_1</code> that are internal data frames of home credit.\n- <code>train_tatic_cb_0</code> that is an external dataset.\n\nWe are mainly interested in the internal datasets since we will have to do a stable inference in a future and a not sure table cannot be a good predictor with this goal.\n\nWe will analyze this points:\n\n- the columns of all dataframes\n- how to merge them\n- their NA meanings and how to fill them\n- some plots","metadata":{}},{"cell_type":"markdown","source":"# 1. SETTINGS","metadata":{}},{"cell_type":"code","source":"import polars as pl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:14.923245Z","iopub.execute_input":"2024-02-14T17:26:14.924959Z","iopub.status.idle":"2024-02-14T17:26:16.252731Z","shell.execute_reply.started":"2024-02-14T17:26:14.924893Z","shell.execute_reply":"2024-02-14T17:26:16.251621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataPath = \"/kaggle/input/home-credit-credit-risk-model-stability/\"","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:16.255589Z","iopub.execute_input":"2024-02-14T17:26:16.256284Z","iopub.status.idle":"2024-02-14T17:26:16.263310Z","shell.execute_reply.started":"2024-02-14T17:26:16.256234Z","shell.execute_reply":"2024-02-14T17:26:16.261797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will import the target dataframe with the features definition in order to improve the graphics later. ","metadata":{}},{"cell_type":"code","source":"df_target = pd.read_parquet(dataPath + 'parquet_files/train/train_base.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:16.265365Z","iopub.execute_input":"2024-02-14T17:26:16.266166Z","iopub.status.idle":"2024-02-14T17:26:16.478494Z","shell.execute_reply.started":"2024-02-14T17:26:16.266117Z","shell.execute_reply":"2024-02-14T17:26:16.477277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_feature = pd.read_csv(dataPath + 'feature_definitions.csv')\ndf_feature = df_feature.set_index('Variable')\ndf_feature.index.name = None\n\ndict_feature = df_feature.to_dict()\ndict_feature = dict_feature['Description']","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:16.479757Z","iopub.execute_input":"2024-02-14T17:26:16.480358Z","iopub.status.idle":"2024-02-14T17:26:16.500110Z","shell.execute_reply.started":"2024-02-14T17:26:16.480313Z","shell.execute_reply":"2024-02-14T17:26:16.499028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_0_0 = pl.read_parquet(dataPath + \"parquet_files/train/train_static_0_0.parquet\")\ntrain_static_0_1 = pl.read_parquet(dataPath + \"parquet_files/train/train_static_0_1.parquet\")\ntrain_static_cb_1 = pl.read_parquet(dataPath + \"parquet_files/train/train_static_cb_0.parquet\")\n","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:16.504452Z","iopub.execute_input":"2024-02-14T17:26:16.505536Z","iopub.status.idle":"2024-02-14T17:26:18.914366Z","shell.execute_reply.started":"2024-02-14T17:26:16.505462Z","shell.execute_reply":"2024-02-14T17:26:18.913522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. STRUCTURE OF THE DATAFRAMES","metadata":{}},{"cell_type":"markdown","source":"Let's first see how the dataframe are made. ","metadata":{}},{"cell_type":"code","source":"train_static_0_0.shape","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:18.915831Z","iopub.execute_input":"2024-02-14T17:26:18.916269Z","iopub.status.idle":"2024-02-14T17:26:18.926447Z","shell.execute_reply.started":"2024-02-14T17:26:18.916215Z","shell.execute_reply":"2024-02-14T17:26:18.925123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_0_1.shape","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:18.927788Z","iopub.execute_input":"2024-02-14T17:26:18.928181Z","iopub.status.idle":"2024-02-14T17:26:18.942198Z","shell.execute_reply.started":"2024-02-14T17:26:18.928148Z","shell.execute_reply":"2024-02-14T17:26:18.940575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_cb_1.shape","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:18.943998Z","iopub.execute_input":"2024-02-14T17:26:18.944375Z","iopub.status.idle":"2024-02-14T17:26:18.952994Z","shell.execute_reply.started":"2024-02-14T17:26:18.944341Z","shell.execute_reply":"2024-02-14T17:26:18.951704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's quite clear that the esternal dataframe has a different structure. \n\nIn this first analys as we have said we will only analyze the internal dataset. ","metadata":{}},{"cell_type":"markdown","source":"# 3. INTERNAL DATA SOURCE ANALYSIS","metadata":{}},{"cell_type":"markdown","source":"Let's first see if the internal datasources have the same columns.","metadata":{}},{"cell_type":"code","source":"columns_0_0 = list(train_static_0_0.columns)\ncolumns_0_1 = list(train_static_0_1.columns)\n\ncolumns_0_0.sort()\ncolumns_0_1.sort()\n\ncolumns_0_0 == columns_0_1","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:18.954320Z","iopub.execute_input":"2024-02-14T17:26:18.954927Z","iopub.status.idle":"2024-02-14T17:26:18.967901Z","shell.execute_reply.started":"2024-02-14T17:26:18.954879Z","shell.execute_reply":"2024-02-14T17:26:18.966566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The two dataframe have the same columns but different rows.","metadata":{}},{"cell_type":"markdown","source":"Let's go in more details.","metadata":{}},{"cell_type":"code","source":"print(\"Number of case id in first dataframe: \", train_static_0_0[\"case_id\"].n_unique())\nprint(\"The case id are unique in the first dataframe: \", train_static_0_0[\"case_id\"].n_unique() == train_static_0_0.shape[0])\n","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:18.970279Z","iopub.execute_input":"2024-02-14T17:26:18.970799Z","iopub.status.idle":"2024-02-14T17:26:18.999041Z","shell.execute_reply.started":"2024-02-14T17:26:18.970754Z","shell.execute_reply":"2024-02-14T17:26:18.998137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Number of case id in second dataframe: \", train_static_0_1[\"case_id\"].n_unique())\nprint(\"The case id are unique in the second dataframe: \", train_static_0_1[\"case_id\"].n_unique() == train_static_0_1.shape[0])\n","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:19.000987Z","iopub.execute_input":"2024-02-14T17:26:19.001322Z","iopub.status.idle":"2024-02-14T17:26:19.014481Z","shell.execute_reply.started":"2024-02-14T17:26:19.001294Z","shell.execute_reply":"2024-02-14T17:26:19.013577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So each case id in the dataframes is unique.\n\nLet's see if the two dataframe have some case id in common.","metadata":{}},{"cell_type":"code","source":"set(train_static_0_1[\"case_id\"].unique()).intersection(set(train_static_0_0[\"case_id\"].unique()))","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:19.016178Z","iopub.execute_input":"2024-02-14T17:26:19.016827Z","iopub.status.idle":"2024-02-14T17:26:19.345336Z","shell.execute_reply.started":"2024-02-14T17:26:19.016790Z","shell.execute_reply":"2024-02-14T17:26:19.344024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We can conclude that we have two perfectly separated dataframes with each one its case ids and with the same columns. We can concated them.**","metadata":{}},{"cell_type":"code","source":"train_static_internal = pl.concat(\n    [\n        train_static_0_0, \n        train_static_0_1,\n    ],\n    how=\"vertical_relaxed\",\n)","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:19.347189Z","iopub.execute_input":"2024-02-14T17:26:19.347660Z","iopub.status.idle":"2024-02-14T17:26:20.108009Z","shell.execute_reply.started":"2024-02-14T17:26:19.347625Z","shell.execute_reply":"2024-02-14T17:26:20.106964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. DEEPER ANALYSIS ON THE COMPLETE INTERNAL TABLE","metadata":{}},{"cell_type":"markdown","source":"Let's move to a deeper analysis on the entire dataframe.","metadata":{}},{"cell_type":"markdown","source":"Create the pandas representation in order to plot it.","metadata":{}},{"cell_type":"code","source":"train_static_internal_pd = train_static_internal.to_pandas()","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:20.113121Z","iopub.execute_input":"2024-02-14T17:26:20.113749Z","iopub.status.idle":"2024-02-14T17:26:22.025309Z","shell.execute_reply.started":"2024-02-14T17:26:20.113715Z","shell.execute_reply":"2024-02-14T17:26:22.023996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.1 NULL ANALYSIS","metadata":{}},{"cell_type":"code","source":"df_nulls = (train_static_internal.null_count() / train_static_internal.shape[0]).transpose(include_header=True).sort(by=\"column_0\", descending=True).to_pandas()\ndf_nulls[\"perc_of_nulls\"] = df_nulls.iloc[:, 1] \ndf_nulls = df_nulls.drop(\"column_0\", axis = 1)\ndf_nulls","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:22.026951Z","iopub.execute_input":"2024-02-14T17:26:22.027594Z","iopub.status.idle":"2024-02-14T17:26:22.058128Z","shell.execute_reply.started":"2024-02-14T17:26:22.027550Z","shell.execute_reply":"2024-02-14T17:26:22.056729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see in case of a lot of null it seems this is a information. We have to understand how to deal with it. ","metadata":{}},{"cell_type":"code","source":"df_nulls[\"perc_of_nulls\"].hist(bins=30)","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:22.059968Z","iopub.execute_input":"2024-02-14T17:26:22.060473Z","iopub.status.idle":"2024-02-14T17:26:22.355235Z","shell.execute_reply.started":"2024-02-14T17:26:22.060425Z","shell.execute_reply":"2024-02-14T17:26:22.353724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_nulls.loc[(df_nulls[\"perc_of_nulls\"] < 0.8) & (df_nulls[\"perc_of_nulls\"]>0.6)]","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:22.356734Z","iopub.execute_input":"2024-02-14T17:26:22.357139Z","iopub.status.idle":"2024-02-14T17:26:22.375341Z","shell.execute_reply.started":"2024-02-14T17:26:22.357105Z","shell.execute_reply":"2024-02-14T17:26:22.373323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's seems in this case as well that the absence of the value is an information. ","metadata":{}},{"cell_type":"markdown","source":"## 4.2 ANALYSIS OF CATEGORICAL VS NUMERICAL ","metadata":{}},{"cell_type":"markdown","source":"We first want to split the numerical variable from the cathegorical ones.","metadata":{}},{"cell_type":"code","source":"features_num = list(train_static_internal_pd.select_dtypes('number'))\nfeatures_total = train_static_internal_pd.columns.tolist()\nfeatures_date = [el for el in features_total if el.endswith(\"D\")]\nfeatures_cat = [el for el in features_total if el not in (features_num + features_date)]\nfeatures_num.remove('case_id')","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:22.377724Z","iopub.execute_input":"2024-02-14T17:26:22.378268Z","iopub.status.idle":"2024-02-14T17:26:23.007213Z","shell.execute_reply.started":"2024-02-14T17:26:22.378223Z","shell.execute_reply":"2024-02-14T17:26:23.005924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features_date","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:23.008987Z","iopub.execute_input":"2024-02-14T17:26:23.009480Z","iopub.status.idle":"2024-02-14T17:26:23.017464Z","shell.execute_reply.started":"2024-02-14T17:26:23.009447Z","shell.execute_reply":"2024-02-14T17:26:23.016153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in features_cat:\n    print(f\"  {col}\")","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:23.019379Z","iopub.execute_input":"2024-02-14T17:26:23.019897Z","iopub.status.idle":"2024-02-14T17:26:23.028605Z","shell.execute_reply.started":"2024-02-14T17:26:23.019852Z","shell.execute_reply":"2024-02-14T17:26:23.026800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the date column we can see that we will need to compute a difference between a reference time and the considered date. ","metadata":{}},{"cell_type":"code","source":"# aesthetics\ndefault_color_1 = 'darkblue'\ndefault_color_2 = 'darkgreen'\ndefault_color_3 = 'darkred'","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:23.030232Z","iopub.execute_input":"2024-02-14T17:26:23.030680Z","iopub.status.idle":"2024-02-14T17:26:23.041397Z","shell.execute_reply.started":"2024-02-14T17:26:23.030645Z","shell.execute_reply":"2024-02-14T17:26:23.040387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_internal_pd = train_static_internal_pd.merge(df_target, on='case_id')","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:23.042498Z","iopub.execute_input":"2024-02-14T17:26:23.042963Z","iopub.status.idle":"2024-02-14T17:26:24.912528Z","shell.execute_reply.started":"2024-02-14T17:26:23.042929Z","shell.execute_reply":"2024-02-14T17:26:24.911249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see how much values are unique in the numerical and categorical variables.","metadata":{}},{"cell_type":"code","source":"for col in features_num:\n    print(col, \": \", len(train_static_internal_pd[col].unique()))","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:24.914198Z","iopub.execute_input":"2024-02-14T17:26:24.915282Z","iopub.status.idle":"2024-02-14T17:26:27.614128Z","shell.execute_reply.started":"2024-02-14T17:26:24.915242Z","shell.execute_reply":"2024-02-14T17:26:27.612837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"date_columns = []\nlong_features_cat = []\nshort_features_cat = []\nfor col in features_cat:\n    if 'date' in col or col.endswith(\"D\"):\n        date_columns.append(col)\n        features_cat.remove(col)\n    elif len(train_static_internal_pd[col].unique()) > 10:\n        long_features_cat.append(col)\n    elif len(train_static_internal_pd[col].unique()) <= 10:\n        short_features_cat.append(col)\n    else: \n        raise ValueError(\"Strange column: \", col)\n        ","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:27.615691Z","iopub.execute_input":"2024-02-14T17:26:27.617292Z","iopub.status.idle":"2024-02-14T17:26:30.637005Z","shell.execute_reply.started":"2024-02-14T17:26:27.617233Z","shell.execute_reply":"2024-02-14T17:26:30.635818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in long_features_cat:\n    print(col, \": \", len(train_static_internal_pd[col].unique()))","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:30.638255Z","iopub.execute_input":"2024-02-14T17:26:30.638627Z","iopub.status.idle":"2024-02-14T17:26:31.845565Z","shell.execute_reply.started":"2024-02-14T17:26:30.638576Z","shell.execute_reply":"2024-02-14T17:26:31.844256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in short_features_cat:\n    print(col, \": \", len(train_static_internal_pd[col].unique()))","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:31.847602Z","iopub.execute_input":"2024-02-14T17:26:31.848952Z","iopub.status.idle":"2024-02-14T17:26:32.862207Z","shell.execute_reply.started":"2024-02-14T17:26:31.848914Z","shell.execute_reply":"2024-02-14T17:26:32.860733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_internal_pd[\"isdebitcard_729L\"].unique()","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:32.863864Z","iopub.execute_input":"2024-02-14T17:26:32.864467Z","iopub.status.idle":"2024-02-14T17:26:32.945987Z","shell.execute_reply.started":"2024-02-14T17:26:32.864427Z","shell.execute_reply":"2024-02-14T17:26:32.944593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_internal_pd[\"isbidproductrequest_292L\"].unique()","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:32.947976Z","iopub.execute_input":"2024-02-14T17:26:32.948370Z","iopub.status.idle":"2024-02-14T17:26:33.025299Z","shell.execute_reply.started":"2024-02-14T17:26:32.948338Z","shell.execute_reply":"2024-02-14T17:26:33.023597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_internal_pd[\"paytype_783L\"].unique()","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:33.027099Z","iopub.execute_input":"2024-02-14T17:26:33.027557Z","iopub.status.idle":"2024-02-14T17:26:33.103210Z","shell.execute_reply.started":"2024-02-14T17:26:33.027521Z","shell.execute_reply":"2024-02-14T17:26:33.102015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_internal_pd[\"typesuite_864L\"].unique()","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:33.104948Z","iopub.execute_input":"2024-02-14T17:26:33.105312Z","iopub.status.idle":"2024-02-14T17:26:33.185900Z","shell.execute_reply.started":"2024-02-14T17:26:33.105282Z","shell.execute_reply":"2024-02-14T17:26:33.184640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_internal_pd[\"bankacctype_710L\"].unique()","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:33.187483Z","iopub.execute_input":"2024-02-14T17:26:33.188666Z","iopub.status.idle":"2024-02-14T17:26:33.263818Z","shell.execute_reply.started":"2024-02-14T17:26:33.188626Z","shell.execute_reply":"2024-02-14T17:26:33.262373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_static_internal_pd[\"isbidproduct_1095L\"].unique()","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:33.265356Z","iopub.execute_input":"2024-02-14T17:26:33.265866Z","iopub.status.idle":"2024-02-14T17:26:33.286300Z","shell.execute_reply.started":"2024-02-14T17:26:33.265825Z","shell.execute_reply":"2024-02-14T17:26:33.284964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see a very big difference among the variable cardinality.","metadata":{}},{"cell_type":"code","source":"def plot_continuous(df, feature, txt):\n    '''Plot a histogram and boxplot for the churned and retained distributions for the specified feature.'''\n    df_func = df.copy()\n    df_paid = df.loc[df[\"target\"] == 1]\n    df_default = df.loc[df[\"target\"] == 0]\n    \n    df_func['target'] = df_func['target'].astype('category')\n    fig, ax1 = plt.subplots()\n\n    for df, label in zip([df_paid,df_default], [0, 1]): \n        sns.boxplot(data=df,\n                     x=feature,\n                     bins=30,\n                     alpha=0.66,\n                     edgecolor='firebrick',\n                     label=label,\n                     kde=False,\n                     ax=ax1)\n    ax1.legend()\n    fig.text(.5, .005, txt, ha='center')\n    plt.tight_layout();","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:33.288060Z","iopub.execute_input":"2024-02-14T17:26:33.288658Z","iopub.status.idle":"2024-02-14T17:26:33.298547Z","shell.execute_reply.started":"2024-02-14T17:26:33.288614Z","shell.execute_reply":"2024-02-14T17:26:33.297043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_categorical(df, feature, txt):\n    '''For a categorical feature, plot a seaborn.countplot for the total counts of each category next to a barplot for the churn rate.'''\n    fig, ax1 = plt.subplots()\n\n    sns.countplot(x=feature,\n                  hue='target',\n                  data=df,\n                  ax=ax1)\n    ax1.set_ylabel('Count')\n    ax1.legend(labels=['paid', 'default'])\n    ax1.tick_params(axis='x', rotation=90)\n    \n    fig.text(.5, .005, txt, ha='center')\n    plt.tight_layout();\n","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:33.300619Z","iopub.execute_input":"2024-02-14T17:26:33.301051Z","iopub.status.idle":"2024-02-14T17:26:33.315834Z","shell.execute_reply.started":"2024-02-14T17:26:33.301017Z","shell.execute_reply":"2024-02-14T17:26:33.314620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in short_features_cat:\n    print(i)\n    plot_categorical(train_static_internal_pd, i, f'({dict_feature[i]})')","metadata":{"execution":{"iopub.status.busy":"2024-02-14T17:26:33.318140Z","iopub.execute_input":"2024-02-14T17:26:33.318897Z","iopub.status.idle":"2024-02-14T17:27:02.665810Z","shell.execute_reply.started":"2024-02-14T17:26:33.318860Z","shell.execute_reply":"2024-02-14T17:27:02.663944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## CONCLUSIONS","metadata":{}},{"cell_type":"markdown","source":"For this dataframe we can conclude that on a first glance, the internal dataset has a good quality.\n\nWe only have to take into account that:\n\n- The NAs seem informative and so we don't have to drop them. \n- We don't know nothing about the outlier values at the moment. Maybe they are informative so we will not drop them.\n- the date can be used to compute a time difference. ","metadata":{}}]}