{
  "id": 509069,
  "title": "This was not the way  (public 18th private 83rd solution)",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/509069",
  "author_name": "",
  "post_date": "2024-06-01T03:57:51.707494700Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello, this is the AITC team.<br>\nWe would like to thank the contest organizers Home Credit and their staff, Kaggle staff, and all those involved.</p>\n<p>The participation period was after the competition resumed.<br>\nInitially, we focused on improving the model, but metric hacking was allowed, so we worked on replicating week_num.<br>\nSince metric hacking was allowed, successful metric hacking became the minimum requirement.<br>\nOf course, this is not a required task for a data scientist, but it was necessary to win. Regardless of the situation, as a Kaggler, it is a fun experience to build models and try different approaches to improve your score.<br>\nWhile we pursued the week_num strategy, we also continued to improve the client defaulted prediction model, since accuracy is what matters in the end.</p>\n<h1>Overview of  the client defaulted prediction model</h1>\n<p>Our team created two models.</p>\n<h2>tacorice model</h2>\n<p>We added quartiles and unique counts to basic aggregation features (max, mean, median, final, standard deviation, etc.) for a total of 524 features. Since CV and LB were not correlated and the hyperparameters were sensitive to LB, I adopted a model with a good balance between CV and LB. (AUC CV 0.861 public LB 0.590)<br>\nlgbm + cat (xgb did not work)<br>\nThe ensemble consists of 5 LightGBM models, 4 CatBoost models.<br>\nFor stability, lgbm and cat used different learning rates and max_depth. We improved the score by slightly ensembling with depth=8 and depth=5, using depth=7 as the base.</p>\n<h2>Cookie model</h2>\n<p>We employed 484 basic aggregation functions (max, mean, median, last, std, etc.).<br>\nBy imputing missing values ​​of birthdate_XX and dateofbirth_XX from the filled columns, the leaderboard score improved slightly.<br>\nThe CV and LB were adjusted to be balanced. (AUC CV 0.860 public LB 0.592)<br>\nlgbm + cat + xgb<br>\nThe ensemble consists of 7 LightGBM models, 4 CatBoost models, and 1 XGBoost model. The blending ratio was manually adjusted based on the leaderboard.</p>\n<p>LightGBM and CatBoost were trained on the entire dataset, but the XGBoost model was only effective when trained on a reduced dataset of 50,000 samples. Inference time is about 2.5 hours.</p>\n<p>For stability, lgbm and cat used different learning rates and max_depth and num_leaves. For each model, depth was varied from 3 to 8, learning_rate was adjusted between 0.03 and 0.1, and num_of_leaves was varied from 16 to 64.</p>\n<h1>week_num strategy</h1>\n<p>For the strategy of replicating week_num, we created a regression model. (Replication from features did not work well.)</p>\n<p>To train the week_num predictor, we used  the client defaulted prediction model features, but only a few features actually worked. The week_num features were very simple and quickly saturated, but they performed well within the training data.</p>\n<p>Based on this strategy, I observed that by making about 15%-20% of week_num worse, the score started to improve (public LB above 0.6).<br>\nThis was about 2 weeks before the end.<br>\nSoon after, the problematic \"This is the way\" notebook was published.<br>\nCompared to the week_num predictor, the latter improved the score, so I abandoned the week_num strategy because there was little time left (it was difficult to advance the week_num predictor due to the limited information available).<br>\nFortunately or unfortunately, the scores in the public notebooks were adjusted by impression seekers and quickly saturated, allowing me to avoid unnecessary parameter adjustments.<br>\nAt this point, I was optimistic that the pure accuracy of  the client defaulted prediction model could determine the final ranking (if everyone used metric hacking).</p>\n<p>In the final week, I focused again on improving the client defaulted prediction model. It was fluctuating between the gold zones, but it wasn't a bad experience.</p>\n<h1>In the end</h1>\n<p>I was ranked 83rd. Not bad, but disappointing.<br>\nAnother big disappointment is that if we had gone ahead with our original strategy with the week_num predictor, we might have won the gold medal with public 0.649, private 0.580 (only with the tacorice model, so we could have improved it further with team model blending and model tuning).</p>\n<p>Of course, this is hypothetical. When the public notebook were published, we were defeated.<br>\nHas anyone else felt like us and thought, \"This is <strong>not</strong> the way\" ?</p>",
  "messages": [
    {
      "id": "2848476",
      "postDate": "06/01/2024 03:57:51",
      "content": "<p>Hello, this is the AITC team.<br>\nWe would like to thank the contest organizers Home Credit and their staff, Kaggle staff, and all those involved.</p>\n<p>The participation period was after the competition resumed.<br>\nInitially, we focused on improving the model, but metric hacking was allowed, so we worked on replicating week_num.<br>\nSince metric hacking was allowed, successful metric hacking became the minimum requirement.<br>\nOf course, this is not a required task for a data scientist, but it was necessary to win. Regardless of the situation, as a Kaggler, it is a fun experience to build models and try different approaches to improve your score.<br>\nWhile we pursued the week_num strategy, we also continued to improve the client defaulted prediction model, since accuracy is what matters in the end.</p>\n<h1>Overview of  the client defaulted prediction model</h1>\n<p>Our team created two models.</p>\n<h2>tacorice model</h2>\n<p>We added quartiles and unique counts to basic aggregation features (max, mean, median, final, standard deviation, etc.) for a total of 524 features. Since CV and LB were not correlated and the hyperparameters were sensitive to LB, I adopted a model with a good balance between CV and LB. (AUC CV 0.861 public LB 0.590)<br>\nlgbm + cat (xgb did not work)<br>\nThe ensemble consists of 5 LightGBM models, 4 CatBoost models.<br>\nFor stability, lgbm and cat used different learning rates and max_depth. We improved the score by slightly ensembling with depth=8 and depth=5, using depth=7 as the base.</p>\n<h2>Cookie model</h2>\n<p>We employed 484 basic aggregation functions (max, mean, median, last, std, etc.).<br>\nBy imputing missing values ​​of birthdate_XX and dateofbirth_XX from the filled columns, the leaderboard score improved slightly.<br>\nThe CV and LB were adjusted to be balanced. (AUC CV 0.860 public LB 0.592)<br>\nlgbm + cat + xgb<br>\nThe ensemble consists of 7 LightGBM models, 4 CatBoost models, and 1 XGBoost model. The blending ratio was manually adjusted based on the leaderboard.</p>\n<p>LightGBM and CatBoost were trained on the entire dataset, but the XGBoost model was only effective when trained on a reduced dataset of 50,000 samples. Inference time is about 2.5 hours.</p>\n<p>For stability, lgbm and cat used different learning rates and max_depth and num_leaves. For each model, depth was varied from 3 to 8, learning_rate was adjusted between 0.03 and 0.1, and num_of_leaves was varied from 16 to 64.</p>\n<h1>week_num strategy</h1>\n<p>For the strategy of replicating week_num, we created a regression model. (Replication from features did not work well.)</p>\n<p>To train the week_num predictor, we used  the client defaulted prediction model features, but only a few features actually worked. The week_num features were very simple and quickly saturated, but they performed well within the training data.</p>\n<p>Based on this strategy, I observed that by making about 15%-20% of week_num worse, the score started to improve (public LB above 0.6).<br>\nThis was about 2 weeks before the end.<br>\nSoon after, the problematic \"This is the way\" notebook was published.<br>\nCompared to the week_num predictor, the latter improved the score, so I abandoned the week_num strategy because there was little time left (it was difficult to advance the week_num predictor due to the limited information available).<br>\nFortunately or unfortunately, the scores in the public notebooks were adjusted by impression seekers and quickly saturated, allowing me to avoid unnecessary parameter adjustments.<br>\nAt this point, I was optimistic that the pure accuracy of  the client defaulted prediction model could determine the final ranking (if everyone used metric hacking).</p>\n<p>In the final week, I focused again on improving the client defaulted prediction model. It was fluctuating between the gold zones, but it wasn't a bad experience.</p>\n<h1>In the end</h1>\n<p>I was ranked 83rd. Not bad, but disappointing.<br>\nAnother big disappointment is that if we had gone ahead with our original strategy with the week_num predictor, we might have won the gold medal with public 0.649, private 0.580 (only with the tacorice model, so we could have improved it further with team model blending and model tuning).</p>\n<p>Of course, this is hypothetical. When the public notebook were published, we were defeated.<br>\nHas anyone else felt like us and thought, \"This is <strong>not</strong> the way\" ?</p>",
      "rawMarkdown": "Hello, this is the AITC team.\nWe would like to thank the contest organizers Home Credit and their staff, Kaggle staff, and all those involved.\n\nThe participation period was after the competition resumed.\nInitially, we focused on improving the model, but metric hacking was allowed, so we worked on replicating week_num.\nSince metric hacking was allowed, successful metric hacking became the minimum requirement.\nOf course, this is not a required task for a data scientist, but it was necessary to win. Regardless of the situation, as a Kaggler, it is a fun experience to build models and try different approaches to improve your score.\nWhile we pursued the week_num strategy, we also continued to improve the client defaulted prediction model, since accuracy is what matters in the end.\n\n\n# Overview of  the client defaulted prediction model\nOur team created two models.\n\n## tacorice model\nWe added quartiles and unique counts to basic aggregation features (max, mean, median, final, standard deviation, etc.) for a total of 524 features. Since CV and LB were not correlated and the hyperparameters were sensitive to LB, I adopted a model with a good balance between CV and LB. (AUC CV 0.861 public LB 0.590)\nlgbm + cat (xgb did not work)\nThe ensemble consists of 5 LightGBM models, 4 CatBoost models.\nFor stability, lgbm and cat used different learning rates and max_depth. We improved the score by slightly ensembling with depth=8 and depth=5, using depth=7 as the base.\n\n## Cookie model\nWe employed 484 basic aggregation functions (max, mean, median, last, std, etc.).\nBy imputing missing values ​​of birthdate_XX and dateofbirth_XX from the filled columns, the leaderboard score improved slightly.\nThe CV and LB were adjusted to be balanced. (AUC CV 0.860 public LB 0.592)\nlgbm + cat + xgb\nThe ensemble consists of 7 LightGBM models, 4 CatBoost models, and 1 XGBoost model. The blending ratio was manually adjusted based on the leaderboard.\n\nLightGBM and CatBoost were trained on the entire dataset, but the XGBoost model was only effective when trained on a reduced dataset of 50,000 samples. Inference time is about 2.5 hours.\n\nFor stability, lgbm and cat used different learning rates and max_depth and num_leaves. For each model, depth was varied from 3 to 8, learning_rate was adjusted between 0.03 and 0.1, and num_of_leaves was varied from 16 to 64.\n\n# week_num strategy\nFor the strategy of replicating week_num, we created a regression model. (Replication from features did not work well.)\n\nTo train the week_num predictor, we used  the client defaulted prediction model features, but only a few features actually worked. The week_num features were very simple and quickly saturated, but they performed well within the training data.\n\nBased on this strategy, I observed that by making about 15%-20% of week_num worse, the score started to improve (public LB above 0.6).\nThis was about 2 weeks before the end.\nSoon after, the problematic \"This is the way\" notebook was published.\nCompared to the week_num predictor, the latter improved the score, so I abandoned the week_num strategy because there was little time left (it was difficult to advance the week_num predictor due to the limited information available).\nFortunately or unfortunately, the scores in the public notebooks were adjusted by impression seekers and quickly saturated, allowing me to avoid unnecessary parameter adjustments.\nAt this point, I was optimistic that the pure accuracy of  the client defaulted prediction model could determine the final ranking (if everyone used metric hacking).\n\nIn the final week, I focused again on improving the client defaulted prediction model. It was fluctuating between the gold zones, but it wasn't a bad experience.\n\n# In the end\nI was ranked 83rd. Not bad, but disappointing.\nAnother big disappointment is that if we had gone ahead with our original strategy with the week_num predictor, we might have won the gold medal with public 0.649, private 0.580 (only with the tacorice model, so we could have improved it further with team model blending and model tuning).\n\nOf course, this is hypothetical. When the public notebook were published, we were defeated.\nHas anyone else felt like us and thought, \"This is **not** the way\" ?",
      "votes": null
    },
    {
      "id": "2848772",
      "postDate": "06/01/2024 07:36:31",
      "content": "<p>The entire \"hack thing\" is NOT the way, it could be a much more interesting challenge with the banned hacking, competing between healthy models and a normal leaderboard (not overflooded with synthetic scores). There should be many good models that are out of medals range that are better than \"hacking the way model\" types on top.</p>",
      "rawMarkdown": "The entire \"hack thing\" is NOT the way, it could be a much more interesting challenge with the banned hacking, competing between healthy models and a normal leaderboard (not overflooded with synthetic scores). There should be many good models that are out of medals range that are better than \"hacking the way model\" types on top.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2848772,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "06/01/2024 07:36:31",
      "content": "<p>The entire \"hack thing\" is NOT the way, it could be a much more interesting challenge with the banned hacking, competing between healthy models and a normal leaderboard (not overflooded with synthetic scores). There should be many good models that are out of medals range that are better than \"hacking the way model\" types on top.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2848476": "Hello, this is the AITC team.\nWe would like to thank the contest organizers Home Credit and their staff, Kaggle staff, and all those involved.\n\nThe participation period was after the competition resumed.\nInitially, we focused on improving the model, but metric hacking was allowed, so we worked on replicating week_num.\nSince metric hacking was allowed, successful metric hacking became the minimum requirement.\nOf course, this is not a required task for a data scientist, but it was necessary to win. Regardless of the situation, as a Kaggler, it is a fun experience to build models and try different approaches to improve your score.\nWhile we pursued the week_num strategy, we also continued to improve the client defaulted prediction model, since accuracy is what matters in the end.\n\n\n# Overview of  the client defaulted prediction model\nOur team created two models.\n\n## tacorice model\nWe added quartiles and unique counts to basic aggregation features (max, mean, median, final, standard deviation, etc.) for a total of 524 features. Since CV and LB were not correlated and the hyperparameters were sensitive to LB, I adopted a model with a good balance between CV and LB. (AUC CV 0.861 public LB 0.590)\nlgbm + cat (xgb did not work)\nThe ensemble consists of 5 LightGBM models, 4 CatBoost models.\nFor stability, lgbm and cat used different learning rates and max_depth. We improved the score by slightly ensembling with depth=8 and depth=5, using depth=7 as the base.\n\n## Cookie model\nWe employed 484 basic aggregation functions (max, mean, median, last, std, etc.).\nBy imputing missing values ​​of birthdate_XX and dateofbirth_XX from the filled columns, the leaderboard score improved slightly.\nThe CV and LB were adjusted to be balanced. (AUC CV 0.860 public LB 0.592)\nlgbm + cat + xgb\nThe ensemble consists of 7 LightGBM models, 4 CatBoost models, and 1 XGBoost model. The blending ratio was manually adjusted based on the leaderboard.\n\nLightGBM and CatBoost were trained on the entire dataset, but the XGBoost model was only effective when trained on a reduced dataset of 50,000 samples. Inference time is about 2.5 hours.\n\nFor stability, lgbm and cat used different learning rates and max_depth and num_leaves. For each model, depth was varied from 3 to 8, learning_rate was adjusted between 0.03 and 0.1, and num_of_leaves was varied from 16 to 64.\n\n# week_num strategy\nFor the strategy of replicating week_num, we created a regression model. (Replication from features did not work well.)\n\nTo train the week_num predictor, we used  the client defaulted prediction model features, but only a few features actually worked. The week_num features were very simple and quickly saturated, but they performed well within the training data.\n\nBased on this strategy, I observed that by making about 15%-20% of week_num worse, the score started to improve (public LB above 0.6).\nThis was about 2 weeks before the end.\nSoon after, the problematic \"This is the way\" notebook was published.\nCompared to the week_num predictor, the latter improved the score, so I abandoned the week_num strategy because there was little time left (it was difficult to advance the week_num predictor due to the limited information available).\nFortunately or unfortunately, the scores in the public notebooks were adjusted by impression seekers and quickly saturated, allowing me to avoid unnecessary parameter adjustments.\nAt this point, I was optimistic that the pure accuracy of  the client defaulted prediction model could determine the final ranking (if everyone used metric hacking).\n\nIn the final week, I focused again on improving the client defaulted prediction model. It was fluctuating between the gold zones, but it wasn't a bad experience.\n\n# In the end\nI was ranked 83rd. Not bad, but disappointing.\nAnother big disappointment is that if we had gone ahead with our original strategy with the week_num predictor, we might have won the gold medal with public 0.649, private 0.580 (only with the tacorice model, so we could have improved it further with team model blending and model tuning).\n\nOf course, this is hypothetical. When the public notebook were published, we were defeated.\nHas anyone else felt like us and thought, \"This is **not** the way\" ?",
    "2848772": "The entire \"hack thing\" is NOT the way, it could be a much more interesting challenge with the banned hacking, competing between healthy models and a normal leaderboard (not overflooded with synthetic scores). There should be many good models that are out of medals range that are better than \"hacking the way model\" types on top."
  },
  "source": "meta"
}