{
  "id": 508242,
  "title": "53th Place Solution (without metric hack)",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508242",
  "author_name": "Sercan Yeşilöz",
  "post_date": "2024-05-28T18:02:27.485000",
  "votes": 29,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, my teammate <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> and I would like to thank the organizers for hosting this competition and congratulate all the winners. We focused on improving our CV before the metric hack was announced to be legal and It wasn't easy to hold on to the public leaderboard after this announcement. We've tried metric hack techniques like linear shifting or restoring week numbers but reached a better private score without using any metric hack. Here is a summary of the solution that secured the 53th place without the metric hack.</p>\n<h2>Feature Engineering</h2>\n<p>We've mostly used public codes for feature engineering and made some modifications to them. For example, we scaled the total day difference between the date features and the <code>date_decision</code> by dividing them by -365. This minor adjustment affected all date-based features and enhanced our cross-validation performance.</p>\n<pre><code>def handle:\n     col  df.columns:\n         col  (,):\n            df = df. - pl.col())\n            df = df..dt.total-)\n    df = df.drop(, )\n    return df\n</code></pre>\n<p>We examined the most important features of the previous Home Credit competition and tried to extract some features that had characteristics similar to them. Here are some features that improved our CV.</p>\n<pre><code>(df_base / df_base)()\n\n(df_base / df_base)()\n\n(df_base / df_base)()\n\n(df_base / ( + df_base))()\n</code></pre>\n<h2>Modeling</h2>\n<p>We have trained many models but have chosen to use four of them in our ensemble model, including two CatBoost models, one LightGBM model, and one Neural Network model. We determined their weights by comparing their cross-validation scores to obtain the ensemble predictions. One of the CatBoost models was trained using 5 stratified folds by grouping week numbers, while the other CatBoost model and the LightGBM model were trained using 5 stratified folds with sample weights that were created based on <code>WEEK_NUM</code></p>\n<p><br></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Folds</th>\n<th>Sample Weights</th>\n<th>AVG AUC Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LightGBM</td>\n<td>StratifiedKFold</td>\n<td>5</td>\n<td>TRUE</td>\n<td>0,85766</td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>StratifiedGroupKFold</td>\n<td>5</td>\n<td>FALSE</td>\n<td>0,85324</td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>StratifiedKFold</td>\n<td>5</td>\n<td>TRUE</td>\n<td>0,85180</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2841851,
      "postDate": "2024-05-28T18:02:27.487Z",
      "content": "<p>First of all, my teammate <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> and I would like to thank the organizers for hosting this competition and congratulate all the winners. We focused on improving our CV before the metric hack was announced to be legal and It wasn't easy to hold on to the public leaderboard after this announcement. We've tried metric hack techniques like linear shifting or restoring week numbers but reached a better private score without using any metric hack. Here is a summary of the solution that secured the 53th place without the metric hack.</p>\n<h2>Feature Engineering</h2>\n<p>We've mostly used public codes for feature engineering and made some modifications to them. For example, we scaled the total day difference between the date features and the <code>date_decision</code> by dividing them by -365. This minor adjustment affected all date-based features and enhanced our cross-validation performance.</p>\n<pre><code>def handle:\n     col  df.columns:\n         col  (,):\n            df = df. - pl.col())\n            df = df..dt.total-)\n    df = df.drop(, )\n    return df\n</code></pre>\n<p>We examined the most important features of the previous Home Credit competition and tried to extract some features that had characteristics similar to them. Here are some features that improved our CV.</p>\n<pre><code>(df_base / df_base)()\n\n(df_base / df_base)()\n\n(df_base / df_base)()\n\n(df_base / ( + df_base))()\n</code></pre>\n<h2>Modeling</h2>\n<p>We have trained many models but have chosen to use four of them in our ensemble model, including two CatBoost models, one LightGBM model, and one Neural Network model. We determined their weights by comparing their cross-validation scores to obtain the ensemble predictions. One of the CatBoost models was trained using 5 stratified folds by grouping week numbers, while the other CatBoost model and the LightGBM model were trained using 5 stratified folds with sample weights that were created based on <code>WEEK_NUM</code></p>\n<p><br></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Folds</th>\n<th>Sample Weights</th>\n<th>AVG AUC Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LightGBM</td>\n<td>StratifiedKFold</td>\n<td>5</td>\n<td>TRUE</td>\n<td>0,85766</td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>StratifiedGroupKFold</td>\n<td>5</td>\n<td>FALSE</td>\n<td>0,85324</td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>StratifiedKFold</td>\n<td>5</td>\n<td>TRUE</td>\n<td>0,85180</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "First of all, my teammate @snnclsr and I would like to thank the organizers for hosting this competition and congratulate all the winners. We focused on improving our CV before the metric hack was announced to be legal and It wasn't easy to hold on to the public leaderboard after this announcement. We've tried metric hack techniques like linear shifting or restoring week numbers but reached a better private score without using any metric hack. Here is a summary of the solution that secured the 53th place without the metric hack.\n\n## Feature Engineering\n\nWe've mostly used public codes for feature engineering and made some modifications to them. For example, we scaled the total day difference between the date features and the <code>date_decision</code> by dividing them by -365. This minor adjustment affected all date-based features and enhanced our cross-validation performance.\n\n```\ndef handle_dates(df):\n    for col in df.columns:\n        if col[-1] in (\"D\",):\n            df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n            df = df.with_columns(pl.col(col).dt.total_days() / -365)\n    df = df.drop(\"date_decision\", \"MONTH\")\n    return df\n```\n\nWe examined the most important features of the previous Home Credit competition and tried to extract some features that had characteristics similar to them. Here are some features that improved our CV.\n\n```\n(df_base['price_1097A'] / df_base['annuity_780A']).alias('credit_annuity_ratio')\n\n(df_base['eir_270L'] / df_base['price_1097A']).alias('interest_share')\n\n(df_base['disbursedcredamount_1113A'] / df_base['credamount_770A']).alias('cred_disbursed_ratio')\n\n(df_base['totaldebt_9A'] / (1 + df_base['credamount_770A'])).alias('totaldebt_credamount_ratio')\n```\n\n## Modeling\n\nWe have trained many models but have chosen to use four of them in our ensemble model, including two CatBoost models, one LightGBM model, and one Neural Network model. We determined their weights by comparing their cross-validation scores to obtain the ensemble predictions. One of the CatBoost models was trained using 5 stratified folds by grouping week numbers, while the other CatBoost model and the LightGBM model were trained using 5 stratified folds with sample weights that were created based on <code>WEEK_NUM</code>\n\n<br>\n\n\n| Model    | CV                   | Folds | Sample Weights | AVG AUC Score |\n| -------- | -------------------- | ----- | -------------- | ------------- |\n| LightGBM | StratifiedKFold      | 5     | TRUE           | 0,85766       |\n| CatBoost | StratifiedGroupKFold | 5     | FALSE          | 0,85324       |\n| CatBoost | StratifiedKFold      | 5     | TRUE           | 0,85180       |",
      "votes": 29
    },
    {
      "id": 2844417,
      "postDate": "2024-05-30T04:23:53.200Z",
      "content": "<p>Congratulations on scoring 57th rank in this competition. Thanks for sharing details of your solution. </p>",
      "rawMarkdown": "Congratulations on scoring 57th rank in this competition. Thanks for sharing details of your solution. ",
      "votes": 1
    },
    {
      "id": 2841885,
      "postDate": "2024-05-28T18:21:10.893Z",
      "content": "<p>nice job, first top 100 notebook ive seen without cheating. well done!</p>",
      "rawMarkdown": "nice job, first top 100 notebook ive seen without cheating. well done!",
      "votes": -2
    },
    {
      "id": 2841900,
      "postDate": "2024-05-28T18:35:49.883Z",
      "content": "<p>congratulations, If possible, share your notebook with us <a href=\"https://www.kaggle.com/sercanyesiloz\" target=\"_blank\">@sercanyesiloz</a> </p>",
      "rawMarkdown": "congratulations, If possible, share your notebook with us @sercanyesiloz "
    },
    {
      "id": 2841910,
      "postDate": "2024-05-28T18:46:14.460Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2842284,
      "postDate": "2024-05-29T02:38:24.177Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2844417,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-05-30T04:23:53.200000",
      "content": "<p>Congratulations on scoring 57th rank in this competition. Thanks for sharing details of your solution. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2841885,
      "author_name": "Gerald R",
      "author_url": "",
      "post_date": "2024-05-28T18:21:10.893000",
      "content": "<p>nice job, first top 100 notebook ive seen without cheating. well done!</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 2841900,
      "author_name": "Thiago Mantuani",
      "author_url": "",
      "post_date": "2024-05-28T18:35:49.883000",
      "content": "<p>congratulations, If possible, share your notebook with us <a href=\"https://www.kaggle.com/sercanyesiloz\" target=\"_blank\">@sercanyesiloz</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2841910,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-28T18:46:14.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2842284,
      "author_name": "Logos",
      "author_url": "",
      "post_date": "2024-05-29T02:38:24.177000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2841851": "First of all, my teammate @snnclsr and I would like to thank the organizers for hosting this competition and congratulate all the winners. We focused on improving our CV before the metric hack was announced to be legal and It wasn't easy to hold on to the public leaderboard after this announcement. We've tried metric hack techniques like linear shifting or restoring week numbers but reached a better private score without using any metric hack. Here is a summary of the solution that secured the 53th place without the metric hack.\n\n## Feature Engineering\n\nWe've mostly used public codes for feature engineering and made some modifications to them. For example, we scaled the total day difference between the date features and the <code>date_decision</code> by dividing them by -365. This minor adjustment affected all date-based features and enhanced our cross-validation performance.\n\n```\ndef handle_dates(df):\n    for col in df.columns:\n        if col[-1] in (\"D\",):\n            df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n            df = df.with_columns(pl.col(col).dt.total_days() / -365)\n    df = df.drop(\"date_decision\", \"MONTH\")\n    return df\n```\n\nWe examined the most important features of the previous Home Credit competition and tried to extract some features that had characteristics similar to them. Here are some features that improved our CV.\n\n```\n(df_base['price_1097A'] / df_base['annuity_780A']).alias('credit_annuity_ratio')\n\n(df_base['eir_270L'] / df_base['price_1097A']).alias('interest_share')\n\n(df_base['disbursedcredamount_1113A'] / df_base['credamount_770A']).alias('cred_disbursed_ratio')\n\n(df_base['totaldebt_9A'] / (1 + df_base['credamount_770A'])).alias('totaldebt_credamount_ratio')\n```\n\n## Modeling\n\nWe have trained many models but have chosen to use four of them in our ensemble model, including two CatBoost models, one LightGBM model, and one Neural Network model. We determined their weights by comparing their cross-validation scores to obtain the ensemble predictions. One of the CatBoost models was trained using 5 stratified folds by grouping week numbers, while the other CatBoost model and the LightGBM model were trained using 5 stratified folds with sample weights that were created based on <code>WEEK_NUM</code>\n\n<br>\n\n\n| Model    | CV                   | Folds | Sample Weights | AVG AUC Score |\n| -------- | -------------------- | ----- | -------------- | ------------- |\n| LightGBM | StratifiedKFold      | 5     | TRUE           | 0,85766       |\n| CatBoost | StratifiedGroupKFold | 5     | FALSE          | 0,85324       |\n| CatBoost | StratifiedKFold      | 5     | TRUE           | 0,85180       |",
    "2844417": "Congratulations on scoring 57th rank in this competition. Thanks for sharing details of your solution. ",
    "2841885": "nice job, first top 100 notebook ive seen without cheating. well done!",
    "2841900": "congratulations, If possible, share your notebook with us @sercanyesiloz ",
    "2841910": "",
    "2842284": "Thanks for sharing!"
  }
}