{
  "id": 347996,
  "title": "135th place solution : +1483 shake-up",
  "url": "/competitions/amex-default-prediction/writeups/hyeongchan-kim-135th-place-solution-1483-shake-up",
  "author_name": "",
  "post_date": "2022-08-27T04:42:00.197Z",
  "votes": 19,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n<p>First, thank you Amex for hosting a fun competition! Also, congratulations to all the winners!</p>\n<h1>TL;DR</h1>\n<p>I couldn't spend lots of time on the competition (only made 30 submissions :(). In the meantime, the competition metric is kinda noisy and we also expected a shake-up/down (not a planet-scale, but for some cases). So, my strategy is focused on protecting a shake-down as possible i can (instead of bulding new features).</p>\n<h1>Overview</h1>\n<p>My strategy is <code>building various datasets, folds, seeds, models</code>. I'll explain them one by one.</p>\n<h2>Data (Pre-Processing)</h2>\n<p>My base dataset is based on the raddar's dataset (huge thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>). Also, most of the pre-processing logic can be found in the <code>Code</code> section.</p>\n<p>The differences are </p>\n<ol>\n<li>using more lagging features (to 3 months)</li>\n<li>not just using a single dataset, but multiple datasets (I just added features incrementally) for the variousity.<ul>\n<li>A dataset</li>\n<li>B dataset = A dataset + (features)</li>\n<li>C dataset = B dataset + (another features)</li></ul></li>\n</ol>\n<p>I didn't check the exact effectiveness of using the datasets on multiple models, however, it seems that positive effects when ensembling in my experiments.</p>\n<h2>Model</h2>\n<p>I built 6 models (3 gbtm, 3 nn) to secure the variousity and roboustness. Also, a few models (LightGBM, CatBoost) are trained on multiple seeds (1, 42, 1337) with the same training recipe. Lastly, some models are trained with 10, 20 folds.</p>\n<ul>\n<li>Xgboost</li>\n<li>CatBoost</li>\n<li>LightGBM (w/ dart, w/o dart)</li>\n<li>5-layers NN</li>\n<li>stacked bi-GRU</li>\n<li>Transformer</li>\n</ul>\n<p>Here's the best CV by the model (sorry for the LB, PB scores, I rarely submitted a single model)</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>bi-GRU</td>\n<td>0.787006</td>\n<td></td>\n</tr>\n<tr>\n<td>Transformer</td>\n<td>0.785647</td>\n<td></td>\n</tr>\n<tr>\n<td>NN</td>\n<td>0.789874</td>\n<td></td>\n</tr>\n<tr>\n<td>Xgboost</td>\n<td>0.795940</td>\n<td>only using the given(?) cat features as <code>cat_features</code></td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>0.797058</td>\n<td>using all <code>np.int8</code> features as <code>cat_features</code></td>\n</tr>\n<tr>\n<td>LighGBM</td>\n<td>0.798410</td>\n<td>w/ dart</td>\n</tr>\n</tbody>\n</table>\n<p>The CV score of the single neural network model isn't good. Nevertheless, when ensembling, It works good with the tree-based models.</p>\n<h2>Blend (Ensemble)</h2>\n<p>Inspired by the discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329103\" target=\"_blank\">log-odds</a>, I found weighted ensemble with log-odds probability is better than a normal weighted ensemble (I tuned the weights with <code>Optuna</code> library based on the OOF). But, one difference is not <code>ln</code>, but <code>log10</code>. In my experiments, It's better to optimize the weights with <code>log10</code>. However, It brings little boost (4th digit difference).</p>\n<p>I ensembled about 50 models, and there's no post-processing logic.</p>\n<h1>Summary</h1>\n<p>The final score is</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Ensemble</td>\n<td><code>0.8009</code></td>\n<td><code>0.7992</code></td>\n<td><code>0.8075</code></td>\n</tr>\n</tbody>\n</table>\n<p>Last day of the competition, I selected about 1600th Public LB solution (my best CV solution). Luckily, <code>Trust CV score</code> wins again :) (Actually, my best CV is also my best LB, and when the cv score increases, lb score increases, so there's little difference between best CV &amp; LB for my cases)</p>\n<p>After the competition, I checked the correlation among the scores (CV vs Private LB, CV vs Public LB). then, I found the CV score is more correlated with Private LB than Public LB in my case.</p>\n<h2>Works</h2>\n<ul>\n<li>blending various models (gbtm + nn), even if there're huge CV gaps <ul>\n<li>e.g. nn 0.790, lgbm 0.798</li></ul></li>\n<li>(maybe) various datasets, models, seeds bring a robust prediction I guess</li>\n</ul>\n<h2>Didn't work</h2>\n<ul>\n<li>pseudo labeling (w/ hard label)<ul>\n<li>maybe <code>soft-label</code> or <code>hard label</code> with a more strict threshold could be worked i guess.</li></ul></li>\n<li>deeper NN models<ul>\n<li>5-layers nn is enough</li></ul></li>\n<li>num of folds doesn't matter (5 folds are enough)<ul>\n<li>there's no significant difference between 5 folds vs 20 folds</li></ul></li>\n<li>rank weighted ensemble</li>\n</ul>\n<p>I hope this you could help :) Thank you! </p>",
  "messages": [
    {
      "id": "1914650",
      "postDate": "08/26/2022 09:17:15",
      "content": "<p>Hello everyone!</p>\n<p>First, thank you Amex for hosting a fun competition! Also, congratulations to all the winners!</p>\n<h1>TL;DR</h1>\n<p>I couldn't spend lots of time on the competition (only made 30 submissions :(). In the meantime, the competition metric is kinda noisy and we also expected a shake-up/down (not a planet-scale, but for some cases). So, my strategy is focused on protecting a shake-down as possible i can (instead of bulding new features).</p>\n<h1>Overview</h1>\n<p>My strategy is <code>building various datasets, folds, seeds, models</code>. I'll explain them one by one.</p>\n<h2>Data (Pre-Processing)</h2>\n<p>My base dataset is based on the raddar's dataset (huge thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>). Also, most of the pre-processing logic can be found in the <code>Code</code> section.</p>\n<p>The differences are </p>\n<ol>\n<li>using more lagging features (to 3 months)</li>\n<li>not just using a single dataset, but multiple datasets (I just added features incrementally) for the variousity.<ul>\n<li>A dataset</li>\n<li>B dataset = A dataset + (features)</li>\n<li>C dataset = B dataset + (another features)</li></ul></li>\n</ol>\n<p>I didn't check the exact effectiveness of using the datasets on multiple models, however, it seems that positive effects when ensembling in my experiments.</p>\n<h2>Model</h2>\n<p>I built 6 models (3 gbtm, 3 nn) to secure the variousity and roboustness. Also, a few models (LightGBM, CatBoost) are trained on multiple seeds (1, 42, 1337) with the same training recipe. Lastly, some models are trained with 10, 20 folds.</p>\n<ul>\n<li>Xgboost</li>\n<li>CatBoost</li>\n<li>LightGBM (w/ dart, w/o dart)</li>\n<li>5-layers NN</li>\n<li>stacked bi-GRU</li>\n<li>Transformer</li>\n</ul>\n<p>Here's the best CV by the model (sorry for the LB, PB scores, I rarely submitted a single model)</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>bi-GRU</td>\n<td>0.787006</td>\n<td></td>\n</tr>\n<tr>\n<td>Transformer</td>\n<td>0.785647</td>\n<td></td>\n</tr>\n<tr>\n<td>NN</td>\n<td>0.789874</td>\n<td></td>\n</tr>\n<tr>\n<td>Xgboost</td>\n<td>0.795940</td>\n<td>only using the given(?) cat features as <code>cat_features</code></td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>0.797058</td>\n<td>using all <code>np.int8</code> features as <code>cat_features</code></td>\n</tr>\n<tr>\n<td>LighGBM</td>\n<td>0.798410</td>\n<td>w/ dart</td>\n</tr>\n</tbody>\n</table>\n<p>The CV score of the single neural network model isn't good. Nevertheless, when ensembling, It works good with the tree-based models.</p>\n<h2>Blend (Ensemble)</h2>\n<p>Inspired by the discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/329103\" target=\"_blank\">log-odds</a>, I found weighted ensemble with log-odds probability is better than a normal weighted ensemble (I tuned the weights with <code>Optuna</code> library based on the OOF). But, one difference is not <code>ln</code>, but <code>log10</code>. In my experiments, It's better to optimize the weights with <code>log10</code>. However, It brings little boost (4th digit difference).</p>\n<p>I ensembled about 50 models, and there's no post-processing logic.</p>\n<h1>Summary</h1>\n<p>The final score is</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Ensemble</td>\n<td><code>0.8009</code></td>\n<td><code>0.7992</code></td>\n<td><code>0.8075</code></td>\n</tr>\n</tbody>\n</table>\n<p>Last day of the competition, I selected about 1600th Public LB solution (my best CV solution). Luckily, <code>Trust CV score</code> wins again :) (Actually, my best CV is also my best LB, and when the cv score increases, lb score increases, so there's little difference between best CV &amp; LB for my cases)</p>\n<p>After the competition, I checked the correlation among the scores (CV vs Private LB, CV vs Public LB). then, I found the CV score is more correlated with Private LB than Public LB in my case.</p>\n<h2>Works</h2>\n<ul>\n<li>blending various models (gbtm + nn), even if there're huge CV gaps <ul>\n<li>e.g. nn 0.790, lgbm 0.798</li></ul></li>\n<li>(maybe) various datasets, models, seeds bring a robust prediction I guess</li>\n</ul>\n<h2>Didn't work</h2>\n<ul>\n<li>pseudo labeling (w/ hard label)<ul>\n<li>maybe <code>soft-label</code> or <code>hard label</code> with a more strict threshold could be worked i guess.</li></ul></li>\n<li>deeper NN models<ul>\n<li>5-layers nn is enough</li></ul></li>\n<li>num of folds doesn't matter (5 folds are enough)<ul>\n<li>there's no significant difference between 5 folds vs 20 folds</li></ul></li>\n<li>rank weighted ensemble</li>\n</ul>\n<p>I hope this you could help :) Thank you! </p>",
      "rawMarkdown": "Hello everyone!\n\nFirst, thank you Amex for hosting a fun competition! Also, congratulations to all the winners!\n\n# TL;DR\n\nI couldn't spend lots of time on the competition (only made 30 submissions :(). In the meantime, the competition metric is kinda noisy and we also expected a shake-up/down (not a planet-scale, but for some cases). So, my strategy is focused on protecting a shake-down as possible i can (instead of bulding new features).\n\n# Overview\n\nMy strategy is `building various datasets, folds, seeds, models`. I'll explain them one by one.\n\n## Data (Pre-Processing)\n\nMy base dataset is based on the raddar's dataset (huge thanks to @raddar). Also, most of the pre-processing logic can be found in the `Code` section.\n\nThe differences are \n1. using more lagging features (to 3 months)\n2. not just using a single dataset, but multiple datasets (I just added features incrementally) for the variousity.\n  * A dataset\n  * B dataset = A dataset + (features)\n  * C dataset = B dataset + (another features)\n\nI didn't check the exact effectiveness of using the datasets on multiple models, however, it seems that positive effects when ensembling in my experiments.\n\n## Model\n\nI built 6 models (3 gbtm, 3 nn) to secure the variousity and roboustness. Also, a few models (LightGBM, CatBoost) are trained on multiple seeds (1, 42, 1337) with the same training recipe. Lastly, some models are trained with 10, 20 folds.\n\n* Xgboost\n* CatBoost\n* LightGBM (w/ dart, w/o dart)\n* 5-layers NN\n* stacked bi-GRU\n* Transformer\n\nHere's the best CV by the model (sorry for the LB, PB scores, I rarely submitted a single model)\n\n| Model | CV | Note |\n| :---: | :---: | :---: |\n| bi-GRU | 0.787006 | |\n| Transformer | 0.785647 | |\n| NN | 0.789874 | |\n| Xgboost | 0.795940 | only using the given(?) cat features as `cat_features` |\n| CatBoost | 0.797058 | using all `np.int8` features as `cat_features` |\n| LighGBM | 0.798410 | w/ dart |\n\nThe CV score of the single neural network model isn't good. Nevertheless, when ensembling, It works good with the tree-based models.\n\n## Blend (Ensemble)\n\nInspired by the discussion [log-odds](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329103), I found weighted ensemble with log-odds probability is better than a normal weighted ensemble (I tuned the weights with `Optuna` library based on the OOF). But, one difference is not `ln`, but `log10`. In my experiments, It's better to optimize the weights with `log10`. However, It brings little boost (4th digit difference).\n\nI ensembled about 50 models, and there's no post-processing logic.\n\n# Summary\n\nThe final score is\n| Model | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: |\n| Ensemble | `0.8009` | `0.7992` | `0.8075` |\n\nLast day of the competition, I selected about 1600th Public LB solution (my best CV solution). Luckily, `Trust CV score` wins again :) (Actually, my best CV is also my best LB, and when the cv score increases, lb score increases, so there's little difference between best CV & LB for my cases)\n\nAfter the competition, I checked the correlation among the scores (CV vs Private LB, CV vs Public LB). then, I found the CV score is more correlated with Private LB than Public LB in my case.\n\n## Works\n\n* blending various models (gbtm + nn), even if there're huge CV gaps \n  * e.g. nn 0.790, lgbm 0.798\n* (maybe) various datasets, models, seeds bring a robust prediction I guess\n\n## Didn't work\n\n* pseudo labeling (w/ hard label)\n  * maybe `soft-label` or `hard label` with a more strict threshold could be worked i guess.\n* deeper NN models\n  * 5-layers nn is enough\n* num of folds doesn't matter (5 folds are enough)\n  * there's no significant difference between 5 folds vs 20 folds\n* rank weighted ensemble\n\nI hope this you could help :) Thank you!",
      "votes": null
    },
    {
      "id": "1914727",
      "postDate": "08/26/2022 11:03:07",
      "content": "<p>Very helpful. Thank you <a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a> for sharing. Only one dumb question. You said your strategy focused on protecting a shake-down as possible. Is there any specific part of your strategy that you notice to be important in pursuing this shake-down protection? Thanks once more for sharing your great work and congrats for the silver. </p>",
      "rawMarkdown": "Very helpful. Thank you @kozistr for sharing. Only one dumb question. You said your strategy focused on protecting a shake-down as possible. Is there any specific part of your strategy that you notice to be important in pursuing this shake-down protection? Thanks once more for sharing your great work and congrats for the silver.",
      "votes": null
    },
    {
      "id": "1914736",
      "postDate": "08/26/2022 11:22:07",
      "content": "<p>great question! I didn't check and compare all of my experiments, but i guess that ensembling NN &amp; tree-based models (various types of models) is most important in my cases.</p>\n<p>When ensembling, similar type of models doesn't give a significant boost on some levels, but mixing various types of models gives an improvment.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "great question! I didn't check and compare all of my experiments, but i guess that ensembling NN & tree-based models (various types of models) is most important in my cases.\n\nWhen ensembling, similar type of models doesn't give a significant boost on some levels, but mixing various types of models gives an improvment.\n\nThank you!",
      "votes": null
    },
    {
      "id": "1914764",
      "postDate": "08/26/2022 11:53:40",
      "content": "<p>Thanks a lot. I agree with your point and will try to increase the variety of models when ensembling in future competitions. Wish you the best. </p>",
      "rawMarkdown": "Thanks a lot. I agree with your point and will try to increase the variety of models when ensembling in future competitions. Wish you the best.",
      "votes": null
    },
    {
      "id": "1915550",
      "postDate": "08/27/2022 04:42:47",
      "content": "<p>Congrats for the medal and shakeup!</p>\n<p><code>My strategy is building various datasets, folds, seeds, models.</code></p>\n<p>I went down the same path for basically the same reasons. Lots of similarities with my strategy and approach to this comp. But in the end I did not select my submissions that contained NN models.</p>",
      "rawMarkdown": "Congrats for the medal and shakeup!\n\n`My strategy is building various datasets, folds, seeds, models.`\n\nI went down the same path for basically the same reasons. Lots of similarities with my strategy and approach to this comp. But in the end I did not select my submissions that contained NN models.",
      "votes": null
    },
    {
      "id": "1922112",
      "postDate": "09/01/2022 09:20:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @kozistr, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    },
    {
      "id": "1933099",
      "postDate": "09/10/2022 09:05:25",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a> and thanks for sharing. As you said, your CV is more correlated with Private LB than Public. That means your final model is robust to the time period in Private LB, but not the Public. Do you think that is caused by features you created or by the model you ensembled?</p>",
      "rawMarkdown": "Congratulations @kozistr and thanks for sharing. As you said, your CV is more correlated with Private LB than Public. That means your final model is robust to the time period in Private LB, but not the Public. Do you think that is caused by features you created or by the model you ensembled?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1914727,
      "author_name": "clacflores",
      "author_url": "",
      "post_date": "08/26/2022 11:03:07",
      "content": "<p>Very helpful. Thank you <a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a> for sharing. Only one dumb question. You said your strategy focused on protecting a shake-down as possible. Is there any specific part of your strategy that you notice to be important in pursuing this shake-down protection? Thanks once more for sharing your great work and congrats for the silver. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1914736,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "08/26/2022 11:22:07",
          "content": "<p>great question! I didn't check and compare all of my experiments, but i guess that ensembling NN &amp; tree-based models (various types of models) is most important in my cases.</p>\n<p>When ensembling, similar type of models doesn't give a significant boost on some levels, but mixing various types of models gives an improvment.</p>\n<p>Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914764,
          "author_name": "clacflores",
          "author_url": "",
          "post_date": "08/26/2022 11:53:40",
          "content": "<p>Thanks a lot. I agree with your point and will try to increase the variety of models when ensembling in future competitions. Wish you the best. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1915550,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "08/27/2022 04:42:47",
      "content": "<p>Congrats for the medal and shakeup!</p>\n<p><code>My strategy is building various datasets, folds, seeds, models.</code></p>\n<p>I went down the same path for basically the same reasons. Lots of similarities with my strategy and approach to this comp. But in the end I did not select my submissions that contained NN models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1922112,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 09:20:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a>, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1933099,
      "author_name": "jacksonyou",
      "author_url": "",
      "post_date": "09/10/2022 09:05:25",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a> and thanks for sharing. As you said, your CV is more correlated with Private LB than Public. That means your final model is robust to the time period in Private LB, but not the Public. Do you think that is caused by features you created or by the model you ensembled?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914650": "Hello everyone!\n\nFirst, thank you Amex for hosting a fun competition! Also, congratulations to all the winners!\n\n# TL;DR\n\nI couldn't spend lots of time on the competition (only made 30 submissions :(). In the meantime, the competition metric is kinda noisy and we also expected a shake-up/down (not a planet-scale, but for some cases). So, my strategy is focused on protecting a shake-down as possible i can (instead of bulding new features).\n\n# Overview\n\nMy strategy is `building various datasets, folds, seeds, models`. I'll explain them one by one.\n\n## Data (Pre-Processing)\n\nMy base dataset is based on the raddar's dataset (huge thanks to @raddar). Also, most of the pre-processing logic can be found in the `Code` section.\n\nThe differences are \n1. using more lagging features (to 3 months)\n2. not just using a single dataset, but multiple datasets (I just added features incrementally) for the variousity.\n  * A dataset\n  * B dataset = A dataset + (features)\n  * C dataset = B dataset + (another features)\n\nI didn't check the exact effectiveness of using the datasets on multiple models, however, it seems that positive effects when ensembling in my experiments.\n\n## Model\n\nI built 6 models (3 gbtm, 3 nn) to secure the variousity and roboustness. Also, a few models (LightGBM, CatBoost) are trained on multiple seeds (1, 42, 1337) with the same training recipe. Lastly, some models are trained with 10, 20 folds.\n\n* Xgboost\n* CatBoost\n* LightGBM (w/ dart, w/o dart)\n* 5-layers NN\n* stacked bi-GRU\n* Transformer\n\nHere's the best CV by the model (sorry for the LB, PB scores, I rarely submitted a single model)\n\n| Model | CV | Note |\n| :---: | :---: | :---: |\n| bi-GRU | 0.787006 | |\n| Transformer | 0.785647 | |\n| NN | 0.789874 | |\n| Xgboost | 0.795940 | only using the given(?) cat features as `cat_features` |\n| CatBoost | 0.797058 | using all `np.int8` features as `cat_features` |\n| LighGBM | 0.798410 | w/ dart |\n\nThe CV score of the single neural network model isn't good. Nevertheless, when ensembling, It works good with the tree-based models.\n\n## Blend (Ensemble)\n\nInspired by the discussion [log-odds](https://www.kaggle.com/competitions/amex-default-prediction/discussion/329103), I found weighted ensemble with log-odds probability is better than a normal weighted ensemble (I tuned the weights with `Optuna` library based on the OOF). But, one difference is not `ln`, but `log10`. In my experiments, It's better to optimize the weights with `log10`. However, It brings little boost (4th digit difference).\n\nI ensembled about 50 models, and there's no post-processing logic.\n\n# Summary\n\nThe final score is\n| Model | CV | Public LB | Private LB |\n| :---: | :---: | :---: | :---: |\n| Ensemble | `0.8009` | `0.7992` | `0.8075` |\n\nLast day of the competition, I selected about 1600th Public LB solution (my best CV solution). Luckily, `Trust CV score` wins again :) (Actually, my best CV is also my best LB, and when the cv score increases, lb score increases, so there's little difference between best CV & LB for my cases)\n\nAfter the competition, I checked the correlation among the scores (CV vs Private LB, CV vs Public LB). then, I found the CV score is more correlated with Private LB than Public LB in my case.\n\n## Works\n\n* blending various models (gbtm + nn), even if there're huge CV gaps \n  * e.g. nn 0.790, lgbm 0.798\n* (maybe) various datasets, models, seeds bring a robust prediction I guess\n\n## Didn't work\n\n* pseudo labeling (w/ hard label)\n  * maybe `soft-label` or `hard label` with a more strict threshold could be worked i guess.\n* deeper NN models\n  * 5-layers nn is enough\n* num of folds doesn't matter (5 folds are enough)\n  * there's no significant difference between 5 folds vs 20 folds\n* rank weighted ensemble\n\nI hope this you could help :) Thank you!",
    "1914727": "Very helpful. Thank you @kozistr for sharing. Only one dumb question. You said your strategy focused on protecting a shake-down as possible. Is there any specific part of your strategy that you notice to be important in pursuing this shake-down protection? Thanks once more for sharing your great work and congrats for the silver.",
    "1914736": "great question! I didn't check and compare all of my experiments, but i guess that ensembling NN & tree-based models (various types of models) is most important in my cases.\n\nWhen ensembling, similar type of models doesn't give a significant boost on some levels, but mixing various types of models gives an improvment.\n\nThank you!",
    "1914764": "Thanks a lot. I agree with your point and will try to increase the variety of models when ensembling in future competitions. Wish you the best.",
    "1915550": "Congrats for the medal and shakeup!\n\n`My strategy is building various datasets, folds, seeds, models.`\n\nI went down the same path for basically the same reasons. Lots of similarities with my strategy and approach to this comp. But in the end I did not select my submissions that contained NN models.",
    "1922112": "Hi @kozistr, May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1933099": "Congratulations @kozistr and thanks for sharing. As you said, your CV is more correlated with Private LB than Public. That means your final model is robust to the time period in Private LB, but not the Public. Do you think that is caused by features you created or by the model you ensembled?"
  },
  "source": "meta"
}