{
  "id": 508113,
  "title": "13th place solution - pmts_year_1139T postprocess",
  "url": "/competitions/home-credit-credit-risk-model-stability/writeups/team-dic-13th-place-solution-pmts-year-1139t-postp",
  "author_name": "",
  "post_date": "2024-05-30T22:31:55.237Z",
  "votes": 38,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle and the host of the competition. To be honest, we don't fully understand what was effective, let me briefly describe our solution.</p>\n<p>Thanks <a href=\"https://www.kaggle.com/kentookumura\" target=\"_blank\">@kentookumura</a> <a href=\"https://www.kaggle.com/pegasus27\" target=\"_blank\">@pegasus27</a> for the collaboration.</p>\n<h2>Context</h2>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview</a></p>\n<p>Data context: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data</a></p>\n<h2><strong>Overview of the approach</strong></h2>\n<p>In feature engineering, we created 772 handcrafted features using tables other than credit_bureau_b1, other_1, deposit_1, debitcard_1. We performed manual and correlation-based feature selection and reduced the number of features to 411. </p>\n<p>In modeling, we built 10 models using LightGBM, XGBoost, CatBoost, and HistGradientBoostingClassifier. I stacked these outputs with a RidgeClassifier and applied probability calibration to create predicted values. Then, we created the final predicted values using random seed averaging.</p>\n<p>In postprocess, we added a simple process to take the maximum value of pmts_year_1139T in the credit_bureau_a_2 table and apply a negative correction to the score for each year.</p>\n<h2><strong>Details of the submission</strong></h2>\n<h3>Feature Engineering</h3>\n<p>The feature engineering was mainly done by <a href=\"https://www.kaggle.com/kentookumura\" target=\"_blank\">@kentookumura</a>. Initially, we decided not to use the credit_bureau_b_2, other_1, deposit_1, debitcard_1 tables which had a high rate of missing case_id. We created handcrafted features based on the results of EDA and the solutions from past competitions. We reduced the number of features from 772 to 411 through manual and correlation-based feature selection. Below, we will describe the points that we believe were effective.</p>\n<p><strong>Era</strong></p>\n<pre><code>df_base = df_base.with_columns(\n        ((pl.col() / ).floor() * ).alias().cast(pl.Int32),\n)\n</code></pre>\n<p><strong>Age at Start of Employment</strong></p>\n<pre><code>df = df.with_columns(\n        ((pl.col() - pl.col()).dt.total_days() // ).cast(pl.Int32).alias(), \n)\n</code></pre>\n<p><strong>Employment Period</strong></p>\n<pre><code>df_base = df_base.with_columns(\n        (pl.col() - pl.col()).alias(),\n)\n</code></pre>\n<p><strong>Date processing other than suffix D</strong></p>\n<pre><code> ():\n     colin df.columns:\n         col[-] (,):\n            df = df.with_columns(pl.col(col) - pl.col())\n            df = df.with_columns(pl.col(col).dt.total_days())\n            df = df.with_columns(pl.col(col).cast(pl.Float32))\n\n                   col:\n            df = df.with_columns(pl.col(col) - pl.col().dt.year())\n            df = df.with_columns(pl.col(col).cast(pl.Int32))\n</code></pre>\n<p><strong>Merge tax_registry tables</strong></p>\n<p>From some case_id with multiple provider information, we inferred the correspondence of each table column and made it into one table.</p>\n<p><strong>Aggregation of String type (mode and n_unique)</strong><br>\nWe used the process <code>pl.col(col).drop_nans().drop_nulls().mode().sort().first()</code> for reproducibility in polars.</p>\n<p><strong>Removal of features that fluctuate greatly during the training data period</strong></p>\n<p>We manually checked and removed features that fluctuate greatly with each WEEK_NUM.</p>\n<h3>Modeling</h3>\n<p><a href=\"https://www.kaggle.com/uplus26e7\" target=\"_blank\">@uplus26e7</a> mainly handled the modeling. We used StratifiedGroupKFold(k=5) based on WEEK_NUM for CV. We tried multiple GBDT models that do not require scaling of features or missing value completion, as these did not go well. In order to create diverse models, we created 10 models with multiple parameters and stacked them with RidgeClassifier. Finally, we corrected the predicted values using Scikit-Learn's CalibratedClassifierCV.</p>\n<p>The above model performed random seed averaging (5 seeds) and used it for the final inference.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Local CV AUC (average 5 seeds)</th>\n<th>Main Parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGBoost</td>\n<td>0.8569388045</td>\n<td></td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>0.8543810988</td>\n<td></td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8576141787</td>\n<td>boosting=”gbdt”, extra_tree=True</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8569003931</td>\n<td>boosting=”gbdt”</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8068982316</td>\n<td>boosting=”rf”</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.7993733596</td>\n<td>boosting=”rf”, extra_tree=True</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8546620954</td>\n<td>boosting=”dart”</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8517172791</td>\n<td>boosting=”dart”, extra_tree=True</td>\n</tr>\n<tr>\n<td>HistGradientBoostingClassifier</td>\n<td>0.8496922895</td>\n<td></td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8574351195</td>\n<td>boosting=”gbdt”, extra_tree=True, data_sample_strategy=”goss”</td>\n</tr>\n<tr>\n<td>CalibratedClassifierCV (RidgeClassifier)</td>\n<td>0.859322</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h3>Postprocess</h3>\n<p><a href=\"https://www.kaggle.com/kentookumura\" target=\"_blank\">@kentookumura</a>'s thorough EDA and experiments revealed that the 'pmts_year_1139T' in the 'credit_bureau_a_2' table is likely the most recent 'date_decision' year. It was also observed that no date column transformations were added in data changes. Based on these findings, we implemented post-processing to decrease the predicted value based on the maximum 'pmts_year_1139T' value.</p>\n<pre><code>submission = pd.read_csv()\npmts_year = ... \n    submission.loc[pmts_year == , ] = (submission.loc[pmts_year == , ] - ).clip()\n    submission.loc[pmts_year == , ] = (submission.loc[pmts_year == , ] - ).clip()\n    submission.loc[pmts_year == , ] = (submission.loc[pmts_year==, ] - ).clip()\n    submission.to_csv(, index=)\n</code></pre>",
  "messages": [
    {
      "id": "2840860",
      "postDate": "05/28/2024 09:27:59",
      "content": "<p>First of all, thanks to Kaggle and the host of the competition. To be honest, we don't fully understand what was effective, let me briefly describe our solution.</p>\n<p>Thanks <a href=\"https://www.kaggle.com/kentookumura\" target=\"_blank\">@kentookumura</a> <a href=\"https://www.kaggle.com/pegasus27\" target=\"_blank\">@pegasus27</a> for the collaboration.</p>\n<h2>Context</h2>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview</a></p>\n<p>Data context: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data</a></p>\n<h2><strong>Overview of the approach</strong></h2>\n<p>In feature engineering, we created 772 handcrafted features using tables other than credit_bureau_b1, other_1, deposit_1, debitcard_1. We performed manual and correlation-based feature selection and reduced the number of features to 411. </p>\n<p>In modeling, we built 10 models using LightGBM, XGBoost, CatBoost, and HistGradientBoostingClassifier. I stacked these outputs with a RidgeClassifier and applied probability calibration to create predicted values. Then, we created the final predicted values using random seed averaging.</p>\n<p>In postprocess, we added a simple process to take the maximum value of pmts_year_1139T in the credit_bureau_a_2 table and apply a negative correction to the score for each year.</p>\n<h2><strong>Details of the submission</strong></h2>\n<h3>Feature Engineering</h3>\n<p>The feature engineering was mainly done by <a href=\"https://www.kaggle.com/kentookumura\" target=\"_blank\">@kentookumura</a>. Initially, we decided not to use the credit_bureau_b_2, other_1, deposit_1, debitcard_1 tables which had a high rate of missing case_id. We created handcrafted features based on the results of EDA and the solutions from past competitions. We reduced the number of features from 772 to 411 through manual and correlation-based feature selection. Below, we will describe the points that we believe were effective.</p>\n<p><strong>Era</strong></p>\n<pre><code>df_base = df_base.with_columns(\n        ((pl.col() / ).floor() * ).alias().cast(pl.Int32),\n)\n</code></pre>\n<p><strong>Age at Start of Employment</strong></p>\n<pre><code>df = df.with_columns(\n        ((pl.col() - pl.col()).dt.total_days() // ).cast(pl.Int32).alias(), \n)\n</code></pre>\n<p><strong>Employment Period</strong></p>\n<pre><code>df_base = df_base.with_columns(\n        (pl.col() - pl.col()).alias(),\n)\n</code></pre>\n<p><strong>Date processing other than suffix D</strong></p>\n<pre><code> ():\n     colin df.columns:\n         col[-] (,):\n            df = df.with_columns(pl.col(col) - pl.col())\n            df = df.with_columns(pl.col(col).dt.total_days())\n            df = df.with_columns(pl.col(col).cast(pl.Float32))\n\n                   col:\n            df = df.with_columns(pl.col(col) - pl.col().dt.year())\n            df = df.with_columns(pl.col(col).cast(pl.Int32))\n</code></pre>\n<p><strong>Merge tax_registry tables</strong></p>\n<p>From some case_id with multiple provider information, we inferred the correspondence of each table column and made it into one table.</p>\n<p><strong>Aggregation of String type (mode and n_unique)</strong><br>\nWe used the process <code>pl.col(col).drop_nans().drop_nulls().mode().sort().first()</code> for reproducibility in polars.</p>\n<p><strong>Removal of features that fluctuate greatly during the training data period</strong></p>\n<p>We manually checked and removed features that fluctuate greatly with each WEEK_NUM.</p>\n<h3>Modeling</h3>\n<p><a href=\"https://www.kaggle.com/uplus26e7\" target=\"_blank\">@uplus26e7</a> mainly handled the modeling. We used StratifiedGroupKFold(k=5) based on WEEK_NUM for CV. We tried multiple GBDT models that do not require scaling of features or missing value completion, as these did not go well. In order to create diverse models, we created 10 models with multiple parameters and stacked them with RidgeClassifier. Finally, we corrected the predicted values using Scikit-Learn's CalibratedClassifierCV.</p>\n<p>The above model performed random seed averaging (5 seeds) and used it for the final inference.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Local CV AUC (average 5 seeds)</th>\n<th>Main Parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGBoost</td>\n<td>0.8569388045</td>\n<td></td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>0.8543810988</td>\n<td></td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8576141787</td>\n<td>boosting=”gbdt”, extra_tree=True</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8569003931</td>\n<td>boosting=”gbdt”</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8068982316</td>\n<td>boosting=”rf”</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.7993733596</td>\n<td>boosting=”rf”, extra_tree=True</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8546620954</td>\n<td>boosting=”dart”</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8517172791</td>\n<td>boosting=”dart”, extra_tree=True</td>\n</tr>\n<tr>\n<td>HistGradientBoostingClassifier</td>\n<td>0.8496922895</td>\n<td></td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>0.8574351195</td>\n<td>boosting=”gbdt”, extra_tree=True, data_sample_strategy=”goss”</td>\n</tr>\n<tr>\n<td>CalibratedClassifierCV (RidgeClassifier)</td>\n<td>0.859322</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h3>Postprocess</h3>\n<p><a href=\"https://www.kaggle.com/kentookumura\" target=\"_blank\">@kentookumura</a>'s thorough EDA and experiments revealed that the 'pmts_year_1139T' in the 'credit_bureau_a_2' table is likely the most recent 'date_decision' year. It was also observed that no date column transformations were added in data changes. Based on these findings, we implemented post-processing to decrease the predicted value based on the maximum 'pmts_year_1139T' value.</p>\n<pre><code>submission = pd.read_csv()\npmts_year = ... \n    submission.loc[pmts_year == , ] = (submission.loc[pmts_year == , ] - ).clip()\n    submission.loc[pmts_year == , ] = (submission.loc[pmts_year == , ] - ).clip()\n    submission.loc[pmts_year == , ] = (submission.loc[pmts_year==, ] - ).clip()\n    submission.to_csv(, index=)\n</code></pre>",
      "rawMarkdown": "First of all, thanks to Kaggle and the host of the competition. To be honest, we don't fully understand what was effective, let me briefly describe our solution.\n\nThanks @kentookumura @pegasus27 for the collaboration.\n\n## Context\n\nBusiness context: [https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview)\n\nData context: [https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data)\n\n## **Overview of the approach**\n\nIn feature engineering, we created 772 handcrafted features using tables other than credit_bureau_b1, other_1, deposit_1, debitcard_1. We performed manual and correlation-based feature selection and reduced the number of features to 411. \n\nIn modeling, we built 10 models using LightGBM, XGBoost, CatBoost, and HistGradientBoostingClassifier. I stacked these outputs with a RidgeClassifier and applied probability calibration to create predicted values. Then, we created the final predicted values using random seed averaging.\n\nIn postprocess, we added a simple process to take the maximum value of pmts_year_1139T in the credit_bureau_a_2 table and apply a negative correction to the score for each year.\n\n## **Details of the submission**\n\n### Feature Engineering\n\nThe feature engineering was mainly done by @kentookumura. Initially, we decided not to use the credit_bureau_b_2, other_1, deposit_1, debitcard_1 tables which had a high rate of missing case_id. We created handcrafted features based on the results of EDA and the solutions from past competitions. We reduced the number of features from 772 to 411 through manual and correlation-based feature selection. Below, we will describe the points that we believe were effective.\n\n**Era**\n\n```python\ndf_base = df_base.with_columns(\n\t\t((pl.col(\"first_birth_259D\") / 10).floor() * 10).alias(\"era\").cast(pl.Int32),\n)\n```\n\n**Age at Start of Employment**\n\n```python\ndf = df.with_columns(\n        ((pl.col(\"empl_employedfrom_271D\") - pl.col(\"birth_259D\")).dt.total_days() // 365).cast(pl.Int32).alias(\"agestartofemploymentA\"), \n)\n```\n\n**Employment Period**\n\n```python\ndf_base = df_base.with_columns(\n        (pl.col(\"first_birth_259D\") - pl.col(\"first_agestartofemploymentA\")).alias(\"durationofemploymentA\"),\n)\n```\n\n**Date processing other than suffix D**\n\n```python\ndef handle_dates(df):\n    for colin df.columns:\n        if col[-1]in (\"D\",):\n            df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n            df = df.with_columns(pl.col(col).dt.total_days())\n            df = df.with_columns(pl.col(col).cast(pl.Float32))\n\n\t\t\t\telif \"year\" in col:\n            df = df.with_columns(pl.col(col) - pl.col(\"date_decision\").dt.year())\n            df = df.with_columns(pl.col(col).cast(pl.Int32))\n```\n\n**Merge tax_registry tables**\n\nFrom some case_id with multiple provider information, we inferred the correspondence of each table column and made it into one table.\n\n**Aggregation of String type (mode and n_unique)**\nWe used the process `pl.col(col).drop_nans().drop_nulls().mode().sort().first()` for reproducibility in polars.\n\n**Removal of features that fluctuate greatly during the training data period**\n\nWe manually checked and removed features that fluctuate greatly with each WEEK_NUM.\n\n### Modeling\n\n@uplus26e7 mainly handled the modeling. We used StratifiedGroupKFold(k=5) based on WEEK_NUM for CV. We tried multiple GBDT models that do not require scaling of features or missing value completion, as these did not go well. In order to create diverse models, we created 10 models with multiple parameters and stacked them with RidgeClassifier. Finally, we corrected the predicted values using Scikit-Learn's CalibratedClassifierCV.\n\nThe above model performed random seed averaging (5 seeds) and used it for the final inference.\n\n| Model | Local CV AUC (average 5 seeds) | Main Parameters |\n| --- | --- | --- |\n| XGBoost | 0.8569388045 |  |\n| CatBoost | 0.8543810988 |  |\n| LightGBM | 0.8576141787 | boosting=”gbdt”, extra_tree=True |\n| LightGBM | 0.8569003931 | boosting=”gbdt” |\n| LightGBM | 0.8068982316 | boosting=”rf” |\n| LightGBM | 0.7993733596 | boosting=”rf”, extra_tree=True |\n| LightGBM | 0.8546620954 | boosting=”dart” |\n| LightGBM | 0.8517172791 | boosting=”dart”, extra_tree=True |\n| HistGradientBoostingClassifier | 0.8496922895 |  |\n| LightGBM | 0.8574351195 | boosting=”gbdt”, extra_tree=True, data_sample_strategy=”goss” |\n| CalibratedClassifierCV (RidgeClassifier) | 0.859322 |  |\n\n### Postprocess\n\n@kentookumura's thorough EDA and experiments revealed that the 'pmts_year_1139T' in the 'credit_bureau_a_2' table is likely the most recent 'date_decision' year. It was also observed that no date column transformations were added in data changes. Based on these findings, we implemented post-processing to decrease the predicted value based on the maximum 'pmts_year_1139T' value.\n\n```python\nsubmission = pd.read_csv(\"submission.csv\")\npmts_year = ... # max pmts_year_1139T group by case_id\n    submission.loc[pmts_year == 2020, \"score\"] = (submission.loc[pmts_year == 2020, \"score\"] - 0.07).clip(0)\n    submission.loc[pmts_year == 2021, \"score\"] = (submission.loc[pmts_year == 2021, \"score\"] - 0.06).clip(0)\n    submission.loc[pmts_year == 2022, \"score\"] = (submission.loc[pmts_year==2022, \"score\"] - 0.02).clip(0)\n    submission.to_csv(\"submission.csv\", index=False)\n```",
      "votes": null
    },
    {
      "id": "2840903",
      "postDate": "05/28/2024 09:49:18",
      "content": "<p>Nice writeup! However, I have some questions:</p>\n<ol>\n<li>Did you use pseudo-labeling/meta-features?</li>\n<li>For the LightGBM models, was 'dart' better than any other boosting method (gbdt, rf)? For how long did training with dart last?</li>\n</ol>",
      "rawMarkdown": "Nice writeup! However, I have some questions:\n\n1. Did you use pseudo-labeling/meta-features?\n2. For the LightGBM models, was 'dart' better than any other boosting method (gbdt, rf)? For how long did training with dart last?",
      "votes": null
    },
    {
      "id": "2841052",
      "postDate": "05/28/2024 11:28:28",
      "content": "<p>Thank you for your comment.</p>\n<ol>\n<li>Did you use pseudo-labeling/meta-features?<br>\nWe tried both, but since they didn't work on the public lb, we didn't include them in the final submission.</li>\n<li>For the LightGBM models, was 'dart' better than any other boosting method (gbdt, rf)? For how long did training with dart last?<br>\nThe best LightGBM model in local cv, public lb was gbdt with extra_tree set to true.<br>\nIn dart training, we fixed n_estimators to 1000.</li>\n</ol>",
      "rawMarkdown": "Thank you for your comment.\n\n1. Did you use pseudo-labeling/meta-features?\n    We tried both, but since they didn't work on the public lb, we didn't include them in the final submission.\n2. For the LightGBM models, was 'dart' better than any other boosting method (gbdt, rf)? For how long did training with dart last?\n    The best LightGBM model in local cv, public lb was gbdt with extra_tree set to true.\n    In dart training, we fixed n_estimators to 1000.",
      "votes": null
    },
    {
      "id": "2841113",
      "postDate": "05/28/2024 12:12:42",
      "content": "<p><a href=\"https://www.kaggle.com/uplus26e7\" target=\"_blank\">@uplus26e7</a> <br>\nHow did you implement pseudo-labels? As OOF predictions on the validation or train set?</p>",
      "rawMarkdown": "uplus26e7 \nHow did you implement pseudo-labels? As OOF predictions on the validation or train set?",
      "votes": null
    },
    {
      "id": "2841192",
      "postDate": "05/28/2024 12:44:18",
      "content": "<p>In each fold, we used the predictions of the base model to the validation data as targets.</p>",
      "rawMarkdown": "In each fold, we used the predictions of the base model to the validation data as targets.",
      "votes": null
    },
    {
      "id": "2841344",
      "postDate": "05/28/2024 14:37:14",
      "content": "<p>Glad to see you have used Calibrated results. My mistake was that I did not select the calibrated response for the final submission, else it would have boosted me up by at least 50. I learned a good lesson in this first competition - to trust my instincts. I manually implemented the isotonic calibration on the final results of my LGB and CatBoost ensemble and it gave me a 0.09 point boost on the private LB but did not select it as final submission because it did not give me any boost in public LB. </p>",
      "rawMarkdown": "Glad to see you have used Calibrated results. My mistake was that I did not select the calibrated response for the final submission, else it would have boosted me up by at least 50. I learned a good lesson in this first competition - to trust my instincts. I manually implemented the isotonic calibration on the final results of my LGB and CatBoost ensemble and it gave me a 0.09 point boost on the private LB but did not select it as final submission because it did not give me any boost in public LB.",
      "votes": null
    },
    {
      "id": "2841741",
      "postDate": "05/28/2024 17:15:21",
      "content": "<p>Thank you for your sharing.  I learned a lot from your post.<br>\nHowever, apart from using pseudo-labeling, are there any techniques I can use? </p>",
      "rawMarkdown": "Thank you for your sharing.  I learned a lot from your post.\nHowever, apart from using pseudo-labeling, are there any techniques I can use?",
      "votes": null
    },
    {
      "id": "2842147",
      "postDate": "05/28/2024 22:14:28",
      "content": "<p>In my opinion, this competition is not suitable for learning materials because the postprocess effects are too big.</p>",
      "rawMarkdown": "In my opinion, this competition is not suitable for learning materials because the postprocess effects are too big.",
      "votes": null
    },
    {
      "id": "2844419",
      "postDate": "05/30/2024 04:25:25",
      "content": "<p>Congratulations on securing 14th place in this competition. Thanks for sharing details of your solution. </p>",
      "rawMarkdown": "Congratulations on securing 14th place in this competition. Thanks for sharing details of your solution.",
      "votes": null
    },
    {
      "id": "2885338",
      "postDate": "06/23/2024 03:55:04",
      "content": "<p>Thank you for your sharing，<br>\nhowever i have one question： in Postprocess stage,how do you determine the decrease value<br>\nThank you for your response in advance.</p>",
      "rawMarkdown": "Thank you for your sharing，\nhowever i have one question： in Postprocess stage,how do you determine the decrease value\nThank you for your response in advance.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2840903,
      "author_name": "andreasbis",
      "author_url": "",
      "post_date": "05/28/2024 09:49:18",
      "content": "<p>Nice writeup! However, I have some questions:</p>\n<ol>\n<li>Did you use pseudo-labeling/meta-features?</li>\n<li>For the LightGBM models, was 'dart' better than any other boosting method (gbdt, rf)? For how long did training with dart last?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2841052,
          "author_name": "uplus26e7",
          "author_url": "",
          "post_date": "05/28/2024 11:28:28",
          "content": "<p>Thank you for your comment.</p>\n<ol>\n<li>Did you use pseudo-labeling/meta-features?<br>\nWe tried both, but since they didn't work on the public lb, we didn't include them in the final submission.</li>\n<li>For the LightGBM models, was 'dart' better than any other boosting method (gbdt, rf)? For how long did training with dart last?<br>\nThe best LightGBM model in local cv, public lb was gbdt with extra_tree set to true.<br>\nIn dart training, we fixed n_estimators to 1000.</li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2841113,
              "author_name": "andreasbis",
              "author_url": "",
              "post_date": "05/28/2024 12:12:42",
              "content": "<p><a href=\"https://www.kaggle.com/uplus26e7\" target=\"_blank\">@uplus26e7</a> <br>\nHow did you implement pseudo-labels? As OOF predictions on the validation or train set?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2841192,
                  "author_name": "uplus26e7",
                  "author_url": "",
                  "post_date": "05/28/2024 12:44:18",
                  "content": "<p>In each fold, we used the predictions of the base model to the validation data as targets.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2841344,
      "author_name": "varuniraothumsi",
      "author_url": "",
      "post_date": "05/28/2024 14:37:14",
      "content": "<p>Glad to see you have used Calibrated results. My mistake was that I did not select the calibrated response for the final submission, else it would have boosted me up by at least 50. I learned a good lesson in this first competition - to trust my instincts. I manually implemented the isotonic calibration on the final results of my LGB and CatBoost ensemble and it gave me a 0.09 point boost on the private LB but did not select it as final submission because it did not give me any boost in public LB. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2841741,
      "author_name": "nicholasyong216",
      "author_url": "",
      "post_date": "05/28/2024 17:15:21",
      "content": "<p>Thank you for your sharing.  I learned a lot from your post.<br>\nHowever, apart from using pseudo-labeling, are there any techniques I can use? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2842147,
          "author_name": "uplus26e7",
          "author_url": "",
          "post_date": "05/28/2024 22:14:28",
          "content": "<p>In my opinion, this competition is not suitable for learning materials because the postprocess effects are too big.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2844419,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "05/30/2024 04:25:25",
      "content": "<p>Congratulations on securing 14th place in this competition. Thanks for sharing details of your solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2885338,
      "author_name": "cauchemare",
      "author_url": "",
      "post_date": "06/23/2024 03:55:04",
      "content": "<p>Thank you for your sharing，<br>\nhowever i have one question： in Postprocess stage,how do you determine the decrease value<br>\nThank you for your response in advance.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2840860": "First of all, thanks to Kaggle and the host of the competition. To be honest, we don't fully understand what was effective, let me briefly describe our solution.\n\nThanks @kentookumura @pegasus27 for the collaboration.\n\n## Context\n\nBusiness context: [https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview)\n\nData context: [https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data)\n\n## **Overview of the approach**\n\nIn feature engineering, we created 772 handcrafted features using tables other than credit_bureau_b1, other_1, deposit_1, debitcard_1. We performed manual and correlation-based feature selection and reduced the number of features to 411. \n\nIn modeling, we built 10 models using LightGBM, XGBoost, CatBoost, and HistGradientBoostingClassifier. I stacked these outputs with a RidgeClassifier and applied probability calibration to create predicted values. Then, we created the final predicted values using random seed averaging.\n\nIn postprocess, we added a simple process to take the maximum value of pmts_year_1139T in the credit_bureau_a_2 table and apply a negative correction to the score for each year.\n\n## **Details of the submission**\n\n### Feature Engineering\n\nThe feature engineering was mainly done by @kentookumura. Initially, we decided not to use the credit_bureau_b_2, other_1, deposit_1, debitcard_1 tables which had a high rate of missing case_id. We created handcrafted features based on the results of EDA and the solutions from past competitions. We reduced the number of features from 772 to 411 through manual and correlation-based feature selection. Below, we will describe the points that we believe were effective.\n\n**Era**\n\n```python\ndf_base = df_base.with_columns(\n\t\t((pl.col(\"first_birth_259D\") / 10).floor() * 10).alias(\"era\").cast(pl.Int32),\n)\n```\n\n**Age at Start of Employment**\n\n```python\ndf = df.with_columns(\n        ((pl.col(\"empl_employedfrom_271D\") - pl.col(\"birth_259D\")).dt.total_days() // 365).cast(pl.Int32).alias(\"agestartofemploymentA\"), \n)\n```\n\n**Employment Period**\n\n```python\ndf_base = df_base.with_columns(\n        (pl.col(\"first_birth_259D\") - pl.col(\"first_agestartofemploymentA\")).alias(\"durationofemploymentA\"),\n)\n```\n\n**Date processing other than suffix D**\n\n```python\ndef handle_dates(df):\n    for colin df.columns:\n        if col[-1]in (\"D\",):\n            df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n            df = df.with_columns(pl.col(col).dt.total_days())\n            df = df.with_columns(pl.col(col).cast(pl.Float32))\n\n\t\t\t\telif \"year\" in col:\n            df = df.with_columns(pl.col(col) - pl.col(\"date_decision\").dt.year())\n            df = df.with_columns(pl.col(col).cast(pl.Int32))\n```\n\n**Merge tax_registry tables**\n\nFrom some case_id with multiple provider information, we inferred the correspondence of each table column and made it into one table.\n\n**Aggregation of String type (mode and n_unique)**\nWe used the process `pl.col(col).drop_nans().drop_nulls().mode().sort().first()` for reproducibility in polars.\n\n**Removal of features that fluctuate greatly during the training data period**\n\nWe manually checked and removed features that fluctuate greatly with each WEEK_NUM.\n\n### Modeling\n\n@uplus26e7 mainly handled the modeling. We used StratifiedGroupKFold(k=5) based on WEEK_NUM for CV. We tried multiple GBDT models that do not require scaling of features or missing value completion, as these did not go well. In order to create diverse models, we created 10 models with multiple parameters and stacked them with RidgeClassifier. Finally, we corrected the predicted values using Scikit-Learn's CalibratedClassifierCV.\n\nThe above model performed random seed averaging (5 seeds) and used it for the final inference.\n\n| Model | Local CV AUC (average 5 seeds) | Main Parameters |\n| --- | --- | --- |\n| XGBoost | 0.8569388045 |  |\n| CatBoost | 0.8543810988 |  |\n| LightGBM | 0.8576141787 | boosting=”gbdt”, extra_tree=True |\n| LightGBM | 0.8569003931 | boosting=”gbdt” |\n| LightGBM | 0.8068982316 | boosting=”rf” |\n| LightGBM | 0.7993733596 | boosting=”rf”, extra_tree=True |\n| LightGBM | 0.8546620954 | boosting=”dart” |\n| LightGBM | 0.8517172791 | boosting=”dart”, extra_tree=True |\n| HistGradientBoostingClassifier | 0.8496922895 |  |\n| LightGBM | 0.8574351195 | boosting=”gbdt”, extra_tree=True, data_sample_strategy=”goss” |\n| CalibratedClassifierCV (RidgeClassifier) | 0.859322 |  |\n\n### Postprocess\n\n@kentookumura's thorough EDA and experiments revealed that the 'pmts_year_1139T' in the 'credit_bureau_a_2' table is likely the most recent 'date_decision' year. It was also observed that no date column transformations were added in data changes. Based on these findings, we implemented post-processing to decrease the predicted value based on the maximum 'pmts_year_1139T' value.\n\n```python\nsubmission = pd.read_csv(\"submission.csv\")\npmts_year = ... # max pmts_year_1139T group by case_id\n    submission.loc[pmts_year == 2020, \"score\"] = (submission.loc[pmts_year == 2020, \"score\"] - 0.07).clip(0)\n    submission.loc[pmts_year == 2021, \"score\"] = (submission.loc[pmts_year == 2021, \"score\"] - 0.06).clip(0)\n    submission.loc[pmts_year == 2022, \"score\"] = (submission.loc[pmts_year==2022, \"score\"] - 0.02).clip(0)\n    submission.to_csv(\"submission.csv\", index=False)\n```",
    "2840903": "Nice writeup! However, I have some questions:\n\n1. Did you use pseudo-labeling/meta-features?\n2. For the LightGBM models, was 'dart' better than any other boosting method (gbdt, rf)? For how long did training with dart last?",
    "2841052": "Thank you for your comment.\n\n1. Did you use pseudo-labeling/meta-features?\n    We tried both, but since they didn't work on the public lb, we didn't include them in the final submission.\n2. For the LightGBM models, was 'dart' better than any other boosting method (gbdt, rf)? For how long did training with dart last?\n    The best LightGBM model in local cv, public lb was gbdt with extra_tree set to true.\n    In dart training, we fixed n_estimators to 1000.",
    "2841113": "uplus26e7 \nHow did you implement pseudo-labels? As OOF predictions on the validation or train set?",
    "2841192": "In each fold, we used the predictions of the base model to the validation data as targets.",
    "2841344": "Glad to see you have used Calibrated results. My mistake was that I did not select the calibrated response for the final submission, else it would have boosted me up by at least 50. I learned a good lesson in this first competition - to trust my instincts. I manually implemented the isotonic calibration on the final results of my LGB and CatBoost ensemble and it gave me a 0.09 point boost on the private LB but did not select it as final submission because it did not give me any boost in public LB.",
    "2841741": "Thank you for your sharing.  I learned a lot from your post.\nHowever, apart from using pseudo-labeling, are there any techniques I can use?",
    "2842147": "In my opinion, this competition is not suitable for learning materials because the postprocess effects are too big.",
    "2844419": "Congratulations on securing 14th place in this competition. Thanks for sharing details of your solution.",
    "2885338": "Thank you for your sharing，\nhowever i have one question： in Postprocess stage,how do you determine the decrease value\nThank you for your response in advance."
  },
  "source": "meta"
}