{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7602123,"sourceType":"competition"},{"sourceId":7575239,"sourceType":"datasetVersion","datasetId":4405300}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!wget http://bit.ly/3ZLyF82 -O CSS.css -q\n    \nfrom IPython.core.display import HTML\nwith open('./CSS.css', 'r') as file:\n    custom_css = file.read()\n\nHTML(custom_css)\n\n!cp /kaggle/input/2024-home-credit-public-repo/HomeCredit2024Banner.png .","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-02-07T01:35:24.897571Z","iopub.execute_input":"2024-02-07T01:35:24.898033Z","iopub.status.idle":"2024-02-07T01:35:27.608360Z","shell.execute_reply.started":"2024-02-07T01:35:24.897985Z","shell.execute_reply":"2024-02-07T01:35:27.606414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Libraries</p>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import sys\nsys.path.append('/kaggle/input/2024-home-credit-public-repo')\n\nimport subprocess\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\nfrom tqdm.auto import tqdm\nfrom utils import RC, PALETTE, cS\nfrom utils import plot_count\n\nsns.set(rc=RC)\n\nimport warnings\nwarnings.filterwarnings('ignore')\npd.set_option('display.max_colwidth', None)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-02-07T01:35:27.614228Z","iopub.execute_input":"2024-02-07T01:35:27.614681Z","iopub.status.idle":"2024-02-07T01:35:29.366758Z","shell.execute_reply.started":"2024-02-07T01:35:27.614630Z","shell.execute_reply":"2024-02-07T01:35:29.365106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p align=\"right\">\n  <img src=\"https://api.monosnap.com/file/download?id=KBbSBFnbD0Vdu1aI2I2dI5m7mgTSq8\"/>\n</p>","metadata":{"_kg_hide-input":false}},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Overview</p>\n\nThe goal of this competition is to predict which clients are more likely to default on their loans. The evaluation will favor solutions that are stable over time.\n\nYour participation may offer consumer finance providers a more reliable and longer-lasting way to assess a potential client’s default risk.","metadata":{}},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Data Schema</p>","metadata":{}},{"cell_type":"code","source":"!tree '/kaggle/input/home-credit-credit-risk-model-stability' -d \n\nROOT = '/kaggle/input/home-credit-credit-risk-model-stability'\nfolders = ['csv_files/train', 'parquet_files/train', 'csv_files/test', 'parquet_files/test']\nextensions = ['.csv',  '.parquet'] * 2\n\nfor dir_, ext in zip(folders, extensions):\n    folder_path = Path(ROOT) / dir_\n    num_files = len(list(folder_path.glob(f'*{ext}')))\n    print(f\"{cS.blk}Number of .csv files in: {'./' + dir_:>21}: {cS.blu}{num_files}{cS.res}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-02-07T01:35:29.368534Z","iopub.execute_input":"2024-02-07T01:35:29.369176Z","iopub.status.idle":"2024-02-07T01:35:30.512269Z","shell.execute_reply.started":"2024-02-07T01:35:29.369138Z","shell.execute_reply":"2024-02-07T01:35:30.510281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p align=\"right\">\n  <img src=\"https://api.monosnap.com/file/download?id=QBYoGNfr4osgQfmrNPKGU4k05hUDkZ\"/>\n</p>","metadata":{"execution":{"iopub.status.busy":"2024-02-06T15:39:34.368459Z","iopub.execute_input":"2024-02-06T15:39:34.369137Z","iopub.status.idle":"2024-02-06T15:39:34.406577Z","shell.execute_reply.started":"2024-02-06T15:39:34.369090Z","shell.execute_reply":"2024-02-06T15:39:34.405212Z"}}},{"cell_type":"markdown","source":"- There are **465** features and **436** respective descriptions in `feature_definitions.csv`. There are no missing descriptions so it means some feature might have same descriptions (for example description `Number of tax deductions` for features: `pmtcount_4527229L`, `pmtcount_4955617L`, `pmtcount_693L`).\n\n- Various predictors were transformed, so to have the following notation for similar groups of transformations: **`P M A D T L`**\n\n- Above you can see that `.csv` files duplicate `.parquet` ones. Ideally, we gotta check if they are really the same as mentioned in the competition data sections.","metadata":{}},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Evaluation</p>","metadata":{}},{"cell_type":"markdown","source":"<p>Submissions are evaluated using a gini stability metric. A gini score is calculated for predictions corresponding to each <code>WEEK_NUM</code>.</p>\n\n$$\n\\text{Gini} = 2 \\times \\text{Area under ROC curve} - 1 \n$$\n\n\nA `linear regression model`, $$y = ax + b $$, is fit through the weekly Gini scores, and a <code>falling rate</code> is calculated as: $$min(0, a)$$\nThis is used to penalize models that drop off in predictive ability.\n\nFinally, the variability of the predictions are calculated by taking the standard deviation of the residuals from the above linear regression, applying a penalty to model variablity.\n\nThe final metric is calculated as:\n\n$$\n\\text{stability metric} = mean(gini) +88.0 \\cdot min(0, a) -0.5 \\cdot std(\\text{residuals})\n$$\n","metadata":{}},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Tables</p>\n\nAlright, let's pick up some files. First, we definitely want to open the above mentioned `base` file. But, let's check the sizes.","metadata":{}},{"cell_type":"code","source":"def get_disk_usage(directory):\n    cmd = f'du {directory}/* -h | sort -rh'\n    result = subprocess.run(cmd, shell=True, stdout=subprocess.PIPE, text=True)\n    output_lines = result.stdout.split('\\n')\n\n    # Extract file/directory names and sizes\n    data = [line.split('\\t') for line in output_lines if line]\n    df = pd.DataFrame(data, columns=['size', 'path'])\n    df['file_name'] = df.path.str.replace('train_|test_', '', regex=True).\\\n    apply(lambda x: Path(x).stem)\n    return df\n\ntrain_disk_usage = get_disk_usage(f'{ROOT}/csv_files/train').reset_index()\ntest_disk_usage = get_disk_usage(f'{ROOT}/csv_files/test')\n\ntrain_disk_usage.reset_index().merge(test_disk_usage, on=['file_name'],\n                                     how='outer', suffixes=['_train', '_test'])\\\n                                     .sort_values(by='index').drop(columns=['index'])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-02-07T01:35:30.516378Z","iopub.execute_input":"2024-02-07T01:35:30.516920Z","iopub.status.idle":"2024-02-07T01:35:30.625526Z","shell.execute_reply.started":"2024-02-07T01:35:30.516873Z","shell.execute_reply":"2024-02-07T01:35:30.623767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations**:\n\n- There are large train files. It looks we really need something powerful to handle them. Polars is a good choice, though I can still keep pandas for the small files.\n- `static_0_2`, `credit_bureau_a_1_4`, `credit_bureau_a_2_11`, `applprev_1_2` are missing in the train folder.","metadata":{}},{"cell_type":"code","source":"PATH_BASE_TRAIN = f'{ROOT}/csv_files/train/train_base.csv'\nPATH_BASE_TEST = f'{ROOT}/csv_files/test/test_base.csv'\n\ntrain =  pd.read_csv(PATH_BASE_TRAIN)\ntest =   pd.read_csv(PATH_BASE_TEST)\n\ndisplay(train.head(3))\ndisplay(test.head(3))\ndisplay(train.dtypes)","metadata":{"execution":{"iopub.status.busy":"2024-02-07T01:35:30.627754Z","iopub.execute_input":"2024-02-07T01:35:30.628202Z","iopub.status.idle":"2024-02-07T01:35:32.257312Z","shell.execute_reply.started":"2024-02-07T01:35:30.628158Z","shell.execute_reply":"2024-02-07T01:35:32.255760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n\nTotal of `case_id` is **1526659** with no dups. There few things I want to verify for the base file before I am done with it:\n- Check if any dates are missing in `date_decision`;\n- Check the target counts;\n- Check the target distibutions against year/month of `date_decision`.","metadata":{}},{"cell_type":"code","source":"def get_date_interval_info(df, cS):\n    df['date_decision'] = pd.to_datetime(df['date_decision'])\n    date_delta = df['date_decision'].drop_duplicates().sort_values().diff()\n    len_uniq_dates = len(df.date_decision.unique())\n    print(\n        f'\\n{cS.blk}[INFO] Actual date range:  {cS.red}{date_delta.sum().days + 1} day(s).',\n        f'\\n{cS.blk}[INFO] Total unique dates: {cS.red}{len_uniq_dates} day(s).'\n    )\n\n    print(f'\\n{cS.blk}[INFO] Min date: {cS.red}{df.date_decision.dt.date.min()}',\n          f'\\n{cS.blk}[INFO] Max date: {cS.red}{df.date_decision.dt.date.max()}')\n    \nget_date_interval_info(train, cS)\n!printf  \"\\n----------Test data----------\\n\"\nget_date_interval_info(test, cS)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-02-07T01:35:32.259938Z","iopub.execute_input":"2024-02-07T01:35:32.260364Z","iopub.status.idle":"2024-02-07T01:35:35.545982Z","shell.execute_reply.started":"2024-02-07T01:35:32.260332Z","shell.execute_reply":"2024-02-07T01:35:35.543990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_count(train, ['target'])","metadata":{"execution":{"iopub.status.busy":"2024-02-07T01:35:35.548393Z","iopub.execute_input":"2024-02-07T01:35:35.548957Z","iopub.status.idle":"2024-02-07T01:35:36.113573Z","shell.execute_reply.started":"2024-02-07T01:35:35.548902Z","shell.execute_reply":"2024-02-07T01:35:36.112292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_date_components_histogram_seaborn(df, component):\n    fig = plt.figure(figsize=(15, 5))\n    sns.histplot(data=df, x=component, bins=30, multiple=\"stack\", kde=False, hue='target',\n                 palette=[PALETTE[0], '#EC0010'])\n    plt.title(f'{component.capitalize()} Distribution')\n    plt.xlabel(component.capitalize())\n    plt.ylabel('Frequency')\n    plt.show()\n\n# train['year'] = train.date_decision.dt.year-2019\ntrain['month'] = train.date_decision.dt.month\ntrain['day_of_year'] = train.date_decision.dt.day_of_year\ntrain['day_of_week'] = train.date_decision.dt.day_of_week\n\n\nplot_date_components_histogram_seaborn(train, 'date_decision')\nplot_date_components_histogram_seaborn(train, 'month')\nplot_date_components_histogram_seaborn(train, 'day_of_year')\nplot_date_components_histogram_seaborn(train, 'day_of_week')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-02-07T01:35:36.115458Z","iopub.execute_input":"2024-02-07T01:35:36.117045Z","iopub.status.idle":"2024-02-07T01:35:43.593393Z","shell.execute_reply.started":"2024-02-07T01:35:36.116988Z","shell.execute_reply":"2024-02-07T01:35:43.591612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Notes:**\n- The dataset is extremely imbalanced, which is expectable for such domain.\n- It seems covid-19 impacted the number of observations greatly.\n- Interesting fact, the bank makes decision even on the weekends. Clearly, this process is automated and runs on a predefined schedule.","metadata":{}},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Null Values%</p>","metadata":{}},{"cell_type":"code","source":"# Code snippet for null values calculation can be found at \n# https://www.kaggle.com/datasets/sergiosaharovskiy/2024-home-credit-public-repo\n# I don't want to waste 1 min 30 second of your life staring at the tqdm bar.\ntrain_disk_usage = pd.read_csv('/kaggle/input/2024-home-credit-public-repo/files/train_disk_usage.csv')\ntrain_disk_usage.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-02-07T01:35:43.595578Z","iopub.execute_input":"2024-02-07T01:35:43.596052Z","iopub.status.idle":"2024-02-07T01:35:43.627743Z","shell.execute_reply.started":"2024-02-07T01:35:43.596015Z","shell.execute_reply":"2024-02-07T01:35:43.625834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"values = train_disk_usage['isna_%'].values.tolist()\ntotal_area = train_disk_usage['height'] * train_disk_usage['width']\ntotal_area_scaled = total_area / total_area.max()\n\nrows = 4\ncols = 8\nfig, axs = plt.subplots(rows, cols, figsize=(18, 11))\n\nfor i, ax in enumerate(axs.flat):\n    \n    outer_square_side = np.sqrt(total_area_scaled[i])\n    inner_square_side = np.sqrt(total_area_scaled[i]*values[i])\n\n    # Add the small square inside the 1x1 image\n    ax.add_patch(plt.Rectangle((0.5 - outer_square_side / 2, 0.5 - outer_square_side / 2),\n                               outer_square_side, outer_square_side,\n                               color='#F03F47', label='Total Records'))\n    \n    ax.add_patch(plt.Rectangle((0.4 - inner_square_side / 2, 0.6 - inner_square_side / 2),\n                               inner_square_side, inner_square_side,\n                               color='#645F64', label='Null Values'))\n\n    ax.set_xticks([])\n    ax.set_yticks([])\n    ax.set_aspect('equal')\n    ax.set_title(f'{train_disk_usage.file_name.iloc[i]}\\n'\n                 f'{train_disk_usage[\"size\"].iloc[i]:}\\nNull_%: {values[i]*100:.2f}')\n\nplt.legend(bbox_to_anchor=(-4, -.4), loc='lower center', ncol=2)\nplt.suptitle('\\nNull values% in Train files scaled and shaped as Squares')\n\nplt.tight_layout()\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-02-07T01:57:58.307321Z","iopub.execute_input":"2024-02-07T01:57:58.308321Z","iopub.status.idle":"2024-02-07T01:58:01.092853Z","shell.execute_reply.started":"2024-02-07T01:57:58.308279Z","shell.execute_reply":"2024-02-07T01:58:01.088062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**How to read it**:\n- The sizes of squares are scaled across all `train` files, where the biggest area is the biggest file;\n- Red squares are normalized total counts of all records across all columns in the respective `train` file;\n- Black squares are normalized total counts of **null** records across all columns in the respective `train` file.\n\nFor example, It means the `train` files are  **>30%** empty on average. The `credit_bureau` files are the biggest ones.","metadata":{}},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Distributions</p>","metadata":{}},{"cell_type":"code","source":"#TODO","metadata":{"execution":{"iopub.status.busy":"2024-02-07T01:35:46.146346Z","iopub.execute_input":"2024-02-07T01:35:46.146827Z","iopub.status.idle":"2024-02-07T01:35:46.153243Z","shell.execute_reply.started":"2024-02-07T01:35:46.146786Z","shell.execute_reply":"2024-02-07T01:35:46.151507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Correlations</p>","metadata":{}},{"cell_type":"code","source":"#TODO","metadata":{"execution":{"iopub.status.busy":"2024-02-07T01:35:46.155258Z","iopub.execute_input":"2024-02-07T01:35:46.155758Z","iopub.status.idle":"2024-02-07T01:35:46.167081Z","shell.execute_reply.started":"2024-02-07T01:35:46.155711Z","shell.execute_reply":"2024-02-07T01:35:46.165048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Hierarchical clustering</p>","metadata":{}},{"cell_type":"code","source":"#TODO","metadata":{"execution":{"iopub.status.busy":"2024-02-07T01:35:46.169412Z","iopub.execute_input":"2024-02-07T01:35:46.169944Z","iopub.status.idle":"2024-02-07T01:35:46.179600Z","shell.execute_reply.started":"2024-02-07T01:35:46.169899Z","shell.execute_reply":"2024-02-07T01:35:46.178255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#EC0010; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #EC0010\">Base XGB Model</p>","metadata":{}},{"cell_type":"code","source":"#TODO","metadata":{"execution":{"iopub.status.busy":"2024-02-07T01:35:46.182069Z","iopub.execute_input":"2024-02-07T01:35:46.182576Z","iopub.status.idle":"2024-02-07T01:35:46.194404Z","shell.execute_reply.started":"2024-02-07T01:35:46.182536Z","shell.execute_reply":"2024-02-07T01:35:46.192753Z"},"trusted":true},"execution_count":null,"outputs":[]}]}