{
  "id": 508202,
  "title": "59th Place Solution for the Home Credit - Credit Risk Model Stability Competition",
  "url": "/competitions/home-credit-credit-risk-model-stability/writeups/nguyenosaurus-59th-place-solution-for-the-home-cre",
  "author_name": "",
  "post_date": "2024-06-11T08:50:16.547Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview</a></p>\n<p>Data context: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data</a></p>\n<h1>Overview of the Approach</h1>\n<p>Our final model was a combination of 4 models of me and 2 models of <a href=\"https://www.kaggle.com/htkhiem2000\" target=\"_blank\">@htkhiem2000</a> . My models are CatBoost (Public/Private LB of 0.586/0.508), LGBM (0.579/0.504), XGB (0.558/0.492) and LightAutoML. Khiem also built CatBoost and LGBM. . And we engineered features independently. One final submission was an ensemble with highest weight for CatBoost model. </p>\n<h1>Details of the submission</h1>\n<p>For each of the categorical features of appl_prev_1, we counted the number of occurrences of each categories, which increase the LB by 0.013. Khiem also perform count encoding on the categorical features.</p>\n<pre><code>   df.select([pl.(pl.String), pl.(pl.Boolean)]).:\n     = df[].().to_list()\n     len() &lt;=   df[].is_null().() &lt; :\n         value  :\n            agg_cols += [pl.().filter(pl.() == value).count().(f)]\n</code></pre>\n<p>We got the idea to use the riskassesment_302T from <a href=\"https://www.kaggle.com/code/pereradulina/credit-risk-prediction-with-lightgbm-and-catboost\" target=\"_blank\">this notebook</a>. This increased LB by 0.001.<br>\nWe computed the number of months of each candidate on previous loan.</p>\n<pre><code>df()(\n        pl()(),\n        pl()(),\n    )()(\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()()\n)\n</code></pre>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-feature-engineering\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-feature-engineering</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-catboost\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-catboost</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-lgbm\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-lgbm</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-xgb\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-xgb</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-lightautoml\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-lightautoml</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-inference\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-inference</a></li>\n</ul>",
  "messages": [
    {
      "id": "2841593",
      "postDate": "05/28/2024 16:16:42",
      "content": "<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview</a></p>\n<p>Data context: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data</a></p>\n<h1>Overview of the Approach</h1>\n<p>Our final model was a combination of 4 models of me and 2 models of <a href=\"https://www.kaggle.com/htkhiem2000\" target=\"_blank\">@htkhiem2000</a> . My models are CatBoost (Public/Private LB of 0.586/0.508), LGBM (0.579/0.504), XGB (0.558/0.492) and LightAutoML. Khiem also built CatBoost and LGBM. . And we engineered features independently. One final submission was an ensemble with highest weight for CatBoost model. </p>\n<h1>Details of the submission</h1>\n<p>For each of the categorical features of appl_prev_1, we counted the number of occurrences of each categories, which increase the LB by 0.013. Khiem also perform count encoding on the categorical features.</p>\n<pre><code>   df.select([pl.(pl.String), pl.(pl.Boolean)]).:\n     = df[].().to_list()\n     len() &lt;=   df[].is_null().() &lt; :\n         value  :\n            agg_cols += [pl.().filter(pl.() == value).count().(f)]\n</code></pre>\n<p>We got the idea to use the riskassesment_302T from <a href=\"https://www.kaggle.com/code/pereradulina/credit-risk-prediction-with-lightgbm-and-catboost\" target=\"_blank\">this notebook</a>. This increased LB by 0.001.<br>\nWe computed the number of months of each candidate on previous loan.</p>\n<pre><code>df()(\n        pl()(),\n        pl()(),\n    )()(\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()(),\n        pl(pl() &gt; )(pl())(None)()()\n)\n</code></pre>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-feature-engineering\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-feature-engineering</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-catboost\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-catboost</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-lgbm\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-lgbm</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-xgb\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-xgb</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-lightautoml\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-lightautoml</a></li>\n<li><a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-inference\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-inference</a></li>\n</ul>",
      "rawMarkdown": "# Context\nBusiness context: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview\n\nData context: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\n# Overview of the Approach\nOur final model was a combination of 4 models of me and 2 models of @htkhiem2000 . My models are CatBoost (Public/Private LB of 0.586/0.508), LGBM (0.579/0.504), XGB (0.558/0.492) and LightAutoML. Khiem also built CatBoost and LGBM. . And we engineered features independently. One final submission was an ensemble with highest weight for CatBoost model. \n# Details of the submission\nFor each of the categorical features of appl_prev_1, we counted the number of occurrences of each categories, which increase the LB by 0.013. Khiem also perform count encoding on the categorical features.\n```\nfor col in df.select([pl.col(pl.String), pl.col(pl.Boolean)]).columns:\n    values = df[col].unique().to_list()\n    if len(values) <= 10 and df[col].is_null().mean() < 0.9:\n        for value in values:\n            agg_cols += [pl.col(col).filter(pl.col(col) == value).count().alias(f\"{col}_{value}_C\")]\n```\nWe got the idea to use the riskassesment_302T from [this notebook](https://www.kaggle.com/code/pereradulina/credit-risk-prediction-with-lightgbm-and-catboost). This increased LB by 0.001.\nWe computed the number of months of each candidate on previous loan.\n```\ndf.group_by([\"case_id\", \"num_group1\"]).agg(\n        pl.count(\"pmts_month_158T\").alias(\"count_pmts_month_158T\"),\n        pl.count(\"pmts_month_706T\").alias(\"count_pmts_month_706T\"),\n    ).group_by(\"case_id\").agg(\n        pl.when(pl.col(\"count_pmts_month_158T\") > 0).then(pl.col(\"count_pmts_month_158T\")).otherwise(None).min().alias(\"min_count_pmts_month_158T\"),\n        pl.when(pl.col(\"count_pmts_month_158T\") > 0).then(pl.col(\"count_pmts_month_158T\")).otherwise(None).max().alias(\"max_count_pmts_month_158T\"),\n        pl.when(pl.col(\"count_pmts_month_158T\") > 0).then(pl.col(\"count_pmts_month_158T\")).otherwise(None).mean().alias(\"mean_count_pmts_month_158T\"),\n        pl.when(pl.col(\"count_pmts_month_158T\") > 0).then(pl.col(\"count_pmts_month_158T\")).otherwise(None).std().alias(\"std_count_pmts_month_158T\"),\n        pl.when(pl.col(\"count_pmts_month_706T\") > 0).then(pl.col(\"count_pmts_month_706T\")).otherwise(None).min().alias(\"min_count_pmts_month_706T\"),\n        pl.when(pl.col(\"count_pmts_month_706T\") > 0).then(pl.col(\"count_pmts_month_706T\")).otherwise(None).max().alias(\"max_count_pmts_month_706T\"),\n        pl.when(pl.col(\"count_pmts_month_706T\") > 0).then(pl.col(\"count_pmts_month_706T\")).otherwise(None).mean().alias(\"mean_count_pmts_month_706T\"),\n        pl.when(pl.col(\"count_pmts_month_706T\") > 0).then(pl.col(\"count_pmts_month_706T\")).otherwise(None).std().alias(\"std_count_pmts_month_706T\")\n)\n```\n# Sources\n- https://www.kaggle.com/nguyenosaurus/home-credit-feature-engineering\n- https://www.kaggle.com/nguyenosaurus/home-credit-catboost\n- https://www.kaggle.com/nguyenosaurus/home-credit-lgbm\n- https://www.kaggle.com/nguyenosaurus/home-credit-xgb\n- https://www.kaggle.com/nguyenosaurus/home-credit-lightautoml\n- https://www.kaggle.com/nguyenosaurus/home-credit-inference",
      "votes": null
    },
    {
      "id": "2841924",
      "postDate": "05/28/2024 18:53:55",
      "content": "<p>Thanks for your solution! Could you explain what you mean by \"ensemble with highest weight for CatBoost model\"?</p>",
      "rawMarkdown": "Thanks for your solution! Could you explain what you mean by \"ensemble with highest weight for CatBoost model\"?",
      "votes": null
    },
    {
      "id": "2842233",
      "postDate": "05/29/2024 01:18:43",
      "content": "<p>The final prediction is a weighted average of all model’s predictions and I gave a larger weight to CatBoost prediction because it has the highest score on LB. You can find more information at this notebook <a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-inference\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-inference</a></p>",
      "rawMarkdown": "The final prediction is a weighted average of all model’s predictions and I gave a larger weight to CatBoost prediction because it has the highest score on LB. You can find more information at this notebook https://www.kaggle.com/nguyenosaurus/home-credit-inference",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2841924,
      "author_name": "lorryzouzelun",
      "author_url": "",
      "post_date": "05/28/2024 18:53:55",
      "content": "<p>Thanks for your solution! Could you explain what you mean by \"ensemble with highest weight for CatBoost model\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2842233,
          "author_name": "nguyenosaurus",
          "author_url": "",
          "post_date": "05/29/2024 01:18:43",
          "content": "<p>The final prediction is a weighted average of all model’s predictions and I gave a larger weight to CatBoost prediction because it has the highest score on LB. You can find more information at this notebook <a href=\"https://www.kaggle.com/nguyenosaurus/home-credit-inference\" target=\"_blank\">https://www.kaggle.com/nguyenosaurus/home-credit-inference</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2841593": "# Context\nBusiness context: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/overview\n\nData context: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\n# Overview of the Approach\nOur final model was a combination of 4 models of me and 2 models of @htkhiem2000 . My models are CatBoost (Public/Private LB of 0.586/0.508), LGBM (0.579/0.504), XGB (0.558/0.492) and LightAutoML. Khiem also built CatBoost and LGBM. . And we engineered features independently. One final submission was an ensemble with highest weight for CatBoost model. \n# Details of the submission\nFor each of the categorical features of appl_prev_1, we counted the number of occurrences of each categories, which increase the LB by 0.013. Khiem also perform count encoding on the categorical features.\n```\nfor col in df.select([pl.col(pl.String), pl.col(pl.Boolean)]).columns:\n    values = df[col].unique().to_list()\n    if len(values) <= 10 and df[col].is_null().mean() < 0.9:\n        for value in values:\n            agg_cols += [pl.col(col).filter(pl.col(col) == value).count().alias(f\"{col}_{value}_C\")]\n```\nWe got the idea to use the riskassesment_302T from [this notebook](https://www.kaggle.com/code/pereradulina/credit-risk-prediction-with-lightgbm-and-catboost). This increased LB by 0.001.\nWe computed the number of months of each candidate on previous loan.\n```\ndf.group_by([\"case_id\", \"num_group1\"]).agg(\n        pl.count(\"pmts_month_158T\").alias(\"count_pmts_month_158T\"),\n        pl.count(\"pmts_month_706T\").alias(\"count_pmts_month_706T\"),\n    ).group_by(\"case_id\").agg(\n        pl.when(pl.col(\"count_pmts_month_158T\") > 0).then(pl.col(\"count_pmts_month_158T\")).otherwise(None).min().alias(\"min_count_pmts_month_158T\"),\n        pl.when(pl.col(\"count_pmts_month_158T\") > 0).then(pl.col(\"count_pmts_month_158T\")).otherwise(None).max().alias(\"max_count_pmts_month_158T\"),\n        pl.when(pl.col(\"count_pmts_month_158T\") > 0).then(pl.col(\"count_pmts_month_158T\")).otherwise(None).mean().alias(\"mean_count_pmts_month_158T\"),\n        pl.when(pl.col(\"count_pmts_month_158T\") > 0).then(pl.col(\"count_pmts_month_158T\")).otherwise(None).std().alias(\"std_count_pmts_month_158T\"),\n        pl.when(pl.col(\"count_pmts_month_706T\") > 0).then(pl.col(\"count_pmts_month_706T\")).otherwise(None).min().alias(\"min_count_pmts_month_706T\"),\n        pl.when(pl.col(\"count_pmts_month_706T\") > 0).then(pl.col(\"count_pmts_month_706T\")).otherwise(None).max().alias(\"max_count_pmts_month_706T\"),\n        pl.when(pl.col(\"count_pmts_month_706T\") > 0).then(pl.col(\"count_pmts_month_706T\")).otherwise(None).mean().alias(\"mean_count_pmts_month_706T\"),\n        pl.when(pl.col(\"count_pmts_month_706T\") > 0).then(pl.col(\"count_pmts_month_706T\")).otherwise(None).std().alias(\"std_count_pmts_month_706T\")\n)\n```\n# Sources\n- https://www.kaggle.com/nguyenosaurus/home-credit-feature-engineering\n- https://www.kaggle.com/nguyenosaurus/home-credit-catboost\n- https://www.kaggle.com/nguyenosaurus/home-credit-lgbm\n- https://www.kaggle.com/nguyenosaurus/home-credit-xgb\n- https://www.kaggle.com/nguyenosaurus/home-credit-lightautoml\n- https://www.kaggle.com/nguyenosaurus/home-credit-inference",
    "2841924": "Thanks for your solution! Could you explain what you mean by \"ensemble with highest weight for CatBoost model\"?",
    "2842233": "The final prediction is a weighted average of all model’s predictions and I gave a larger weight to CatBoost prediction because it has the highest score on LB. You can find more information at this notebook https://www.kaggle.com/nguyenosaurus/home-credit-inference"
  },
  "source": "meta"
}