{
  "id": 508337,
  "title": "[1st Place Solution] My Betting Strategy",
  "url": "/competitions/home-credit-credit-risk-model-stability/writeups/yuuniee-1st-place-solution-my-betting-strategy",
  "author_name": "",
  "post_date": "2024-05-31T02:41:29.600Z",
  "votes": 175,
  "comment_count": 65,
  "views": 0,
  "content": "<h2>Foreword</h2>\n<p>Hello, this is yuuniee.<br>\nThank you to the competition host Home Credit and staff, Kaggle staff and everyone involved.</p>\n<p>The image below is a summary of my solution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6960421%2F92721e389d44b29d04f931d244fed316%2F.png?generation=1716967303955938&amp;alt=media\" alt=\"image_\"></p>\n<p>The maximum Private LB Score for my model pool without Metric Hacking is close to about 0.53.</p>\n<p>As everyone knows, this competition takes place in two phases.<br>\nPhase 1 was ML (Machine Learning), and Phase 2 was MH (Metric Hack).<br>\nI will briefly introduce each part.<br>\n(If you are only curious about ML, please read only Phase 1.)</p>\n<h2>I. Phase 1 - ML (Machine Learning)</h2>\n<p>Thank you for writing a great notebook for EDA and solution construction at the beginning of the competition.<br>\n<a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> - <a href=\"https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission\" target=\"_blank\">https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission</a><br>\n<a href=\"https://www.kaggle.com/greysky\" target=\"_blank\">@greysky</a> - <a href=\"https://www.kaggle.com/code/greysky/home-credit-baseline\" target=\"_blank\">https://www.kaggle.com/code/greysky/home-credit-baseline</a><br>\nAdditionally, I would like to thank the many participants who provided various insights.</p>\n<ol>\n<li><p>The CV strategy is StratifiedGroupKFold, and has been tested and trained in various forms: No Shuffle and Shuffle. In my case, differences in CV of around 0.001 to 0.005 had a low correlation with LB, while differences above 0.01 had a high correlation with LB. In addition, CV improvement through model parameter adjustment had a relatively low correlation with LB, and CV improvement through FE had a relatively high correlation with LB.</p></li>\n<li><p>There's nothing special about Feature Engineering. Max, Min, Avg, Var, First, Last and Max-Min Difference of each item were utilized. (I will release the code after cleaning it up.)</p></li>\n<li><p>LGBM was a bit lacking compared to Catboost, but was good for the ensemble. The reason why the number of features is smaller is because some items (mainly categorical types) were excluded because they caused performance degradation and overfitting in LGBM, and other slightly different types of features were added and deleted.</p></li>\n<li><p>DNN is a Denselight model and utilizes the LightAutoML library. There were some bugs, so I had to fix some of them myself, but overall, I think it is a good library. At first, I planned to lightly test with Denselight and then build the final model with a larger model (FT-Transformer, etc.). Surprisingly, I haven't created or found a model that beats Denselight performance. Perhaps I made a mistake, or the nature of the data is sensitive to overfitting, so I guess a simpler model worked better.</p></li>\n<li><p>Catboost was the model that performed best in this competition and I guess it performed well on a large number of categorical features (about 117).</p></li>\n</ol>\n<p>During this period, my public LB was close to top 5 and the competition seemed to be going very smoothly.<br>\nBut…</p>\n<blockquote>\n  <h3>+ What didn't work</h3>\n  <ol>\n  <li>Transformer type models such as Tabnet, TabTransformer, FT-Transformer, etc.</li>\n  <li>Income, expenditure, and tax statistics by period (month, week, etc.).</li>\n  <li>Differences between individual income, expenses, and taxes.</li>\n  <li>K-means clustering</li>\n  </ol>\n</blockquote>\n<h2>II. Phase 2 - MH(Metric Hack)</h2>\n<p>Pointed out problems with MH at the beginning of the competition,<br>\n<a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> - <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449</a><br>\nAnnounced the revival of MH in the latter half of the competition,<br>\n<a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> - <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167</a></p>\n<p>Also, thanks to the many participants who suggested various other comments and solutions. Thanks to you, my knowledge has expanded.<br>\nIn particular, the two people above are true scientists and journalists who could have hidden this discovery and profited from it, but pursued the public interest by exposing it to everyone. (Why not consider Kaggle giving a special award for this sharing?)<br>\nAs a result, the problem was not solved, but we learned something(?) from it and at least I was able to take minimal measures.</p>\n<p>After <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a>'s \"Metric hack again, sorry\" post, I tried Data Reversing and found that the difference between date_decision and min_refreshdate_3813885D had a high correlation (about 0.9 or more) with WEEK_NUM.<br>\n(More information can be found at <a href=\"https://www.kaggle.com/code/eivolkova/how-to-restore-the-dates?scriptVersionId=180157891\" target=\"_blank\">https://www.kaggle.com/code/eivolkova/how-to-restore-the-dates?scriptVersionId=180157891</a>, published by <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> immediately after the competition.)<br>\nAnd the way suggested by <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a>:</p>\n<pre><code>DEVIDE = /\nREDUCE = \ncondition = df[] &lt; (df[].()-df[].())*DEVIDE+df[].()\ndf.loc[condition, ] = (df.loc[condition, ] - REDUCE).clip()\n</code></pre>\n<p>Based on this, the results of examining the score details submitted by changing the REDUCE and DEVIDE values ​​were as follows.</p>\n<ol>\n<li>The public/private division of test data spans the entire period and it is difficult to identify a specific distribution, and the optimal value in private will probably be distributed around 1/4 to 3/4.</li>\n<li>The optimal public LB value of REDUCE is 0.03 (at least in my model), and the optimal value in private is probably distributed around 0.02 to 0.04.</li>\n</ol>\n<p>The difference in pure model performance between top competitors was 0.00X,<br>\nSince the difference due to the above adjustment value was 0.0X,<br>\nAt this stage it became gambling for me.</p>\n<p>Some will bet low, others will bet high.<br>\nHere, I trusted my model and chose neutral, and chose DEVIDE=1/2, REDUCE=0.03, which is the median of the expected distribution.<br>\nI expected someone who bet high or low to take the prize, and I expected about a 50% chance of finishing in the top 10.</p>\n<h2>III. Conclusion</h2>\n<ul>\n<li><p>The first selected submission is the No Hack model with CV Best Score created in Phase 1, and the second selected submission is the application of Metric Hack to it. (I thought there was a high probability that Metric Hack would work, but I also submitted No Hack just in case)</p></li>\n<li><p>Rather than simply sharing solutions, I thought it would be better to share my experience in detail. To improve the bad parts that happened in this competition.</p></li>\n<li><p>My ranking was due to luck, and although it is the Best Private Score submitted, it probably won't be Best Solution.</p></li>\n<li><p>Although the competition ended on a gloomy note due to the metric hacking issue, the host and staff worked hard to make it a meaningful and successful competition. </p></li>\n<li><p>Based on this, we expect a more developed system in the future.</p></li>\n</ul>\n<p>thank you</p>\n<blockquote>\n  <p>Edit : I posted the FE code here.<br>\n  644(for lgbm) - <a href=\"https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-risk-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-risk-lightgbm</a><br>\n  661(for cat &amp; dnn) - <a href=\"https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-catboost-inference\" target=\"_blank\">https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-catboost-inference</a></p>\n</blockquote>",
  "messages": [
    {
      "id": "2842729",
      "postDate": "05/29/2024 08:24:57",
      "content": "<h2>Foreword</h2>\n<p>Hello, this is yuuniee.<br>\nThank you to the competition host Home Credit and staff, Kaggle staff and everyone involved.</p>\n<p>The image below is a summary of my solution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6960421%2F92721e389d44b29d04f931d244fed316%2F.png?generation=1716967303955938&amp;alt=media\" alt=\"image_\"></p>\n<p>The maximum Private LB Score for my model pool without Metric Hacking is close to about 0.53.</p>\n<p>As everyone knows, this competition takes place in two phases.<br>\nPhase 1 was ML (Machine Learning), and Phase 2 was MH (Metric Hack).<br>\nI will briefly introduce each part.<br>\n(If you are only curious about ML, please read only Phase 1.)</p>\n<h2>I. Phase 1 - ML (Machine Learning)</h2>\n<p>Thank you for writing a great notebook for EDA and solution construction at the beginning of the competition.<br>\n<a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> - <a href=\"https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission\" target=\"_blank\">https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission</a><br>\n<a href=\"https://www.kaggle.com/greysky\" target=\"_blank\">@greysky</a> - <a href=\"https://www.kaggle.com/code/greysky/home-credit-baseline\" target=\"_blank\">https://www.kaggle.com/code/greysky/home-credit-baseline</a><br>\nAdditionally, I would like to thank the many participants who provided various insights.</p>\n<ol>\n<li><p>The CV strategy is StratifiedGroupKFold, and has been tested and trained in various forms: No Shuffle and Shuffle. In my case, differences in CV of around 0.001 to 0.005 had a low correlation with LB, while differences above 0.01 had a high correlation with LB. In addition, CV improvement through model parameter adjustment had a relatively low correlation with LB, and CV improvement through FE had a relatively high correlation with LB.</p></li>\n<li><p>There's nothing special about Feature Engineering. Max, Min, Avg, Var, First, Last and Max-Min Difference of each item were utilized. (I will release the code after cleaning it up.)</p></li>\n<li><p>LGBM was a bit lacking compared to Catboost, but was good for the ensemble. The reason why the number of features is smaller is because some items (mainly categorical types) were excluded because they caused performance degradation and overfitting in LGBM, and other slightly different types of features were added and deleted.</p></li>\n<li><p>DNN is a Denselight model and utilizes the LightAutoML library. There were some bugs, so I had to fix some of them myself, but overall, I think it is a good library. At first, I planned to lightly test with Denselight and then build the final model with a larger model (FT-Transformer, etc.). Surprisingly, I haven't created or found a model that beats Denselight performance. Perhaps I made a mistake, or the nature of the data is sensitive to overfitting, so I guess a simpler model worked better.</p></li>\n<li><p>Catboost was the model that performed best in this competition and I guess it performed well on a large number of categorical features (about 117).</p></li>\n</ol>\n<p>During this period, my public LB was close to top 5 and the competition seemed to be going very smoothly.<br>\nBut…</p>\n<blockquote>\n  <h3>+ What didn't work</h3>\n  <ol>\n  <li>Transformer type models such as Tabnet, TabTransformer, FT-Transformer, etc.</li>\n  <li>Income, expenditure, and tax statistics by period (month, week, etc.).</li>\n  <li>Differences between individual income, expenses, and taxes.</li>\n  <li>K-means clustering</li>\n  </ol>\n</blockquote>\n<h2>II. Phase 2 - MH(Metric Hack)</h2>\n<p>Pointed out problems with MH at the beginning of the competition,<br>\n<a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> - <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449</a><br>\nAnnounced the revival of MH in the latter half of the competition,<br>\n<a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> - <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167</a></p>\n<p>Also, thanks to the many participants who suggested various other comments and solutions. Thanks to you, my knowledge has expanded.<br>\nIn particular, the two people above are true scientists and journalists who could have hidden this discovery and profited from it, but pursued the public interest by exposing it to everyone. (Why not consider Kaggle giving a special award for this sharing?)<br>\nAs a result, the problem was not solved, but we learned something(?) from it and at least I was able to take minimal measures.</p>\n<p>After <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a>'s \"Metric hack again, sorry\" post, I tried Data Reversing and found that the difference between date_decision and min_refreshdate_3813885D had a high correlation (about 0.9 or more) with WEEK_NUM.<br>\n(More information can be found at <a href=\"https://www.kaggle.com/code/eivolkova/how-to-restore-the-dates?scriptVersionId=180157891\" target=\"_blank\">https://www.kaggle.com/code/eivolkova/how-to-restore-the-dates?scriptVersionId=180157891</a>, published by <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> immediately after the competition.)<br>\nAnd the way suggested by <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a>:</p>\n<pre><code>DEVIDE = /\nREDUCE = \ncondition = df[] &lt; (df[].()-df[].())*DEVIDE+df[].()\ndf.loc[condition, ] = (df.loc[condition, ] - REDUCE).clip()\n</code></pre>\n<p>Based on this, the results of examining the score details submitted by changing the REDUCE and DEVIDE values ​​were as follows.</p>\n<ol>\n<li>The public/private division of test data spans the entire period and it is difficult to identify a specific distribution, and the optimal value in private will probably be distributed around 1/4 to 3/4.</li>\n<li>The optimal public LB value of REDUCE is 0.03 (at least in my model), and the optimal value in private is probably distributed around 0.02 to 0.04.</li>\n</ol>\n<p>The difference in pure model performance between top competitors was 0.00X,<br>\nSince the difference due to the above adjustment value was 0.0X,<br>\nAt this stage it became gambling for me.</p>\n<p>Some will bet low, others will bet high.<br>\nHere, I trusted my model and chose neutral, and chose DEVIDE=1/2, REDUCE=0.03, which is the median of the expected distribution.<br>\nI expected someone who bet high or low to take the prize, and I expected about a 50% chance of finishing in the top 10.</p>\n<h2>III. Conclusion</h2>\n<ul>\n<li><p>The first selected submission is the No Hack model with CV Best Score created in Phase 1, and the second selected submission is the application of Metric Hack to it. (I thought there was a high probability that Metric Hack would work, but I also submitted No Hack just in case)</p></li>\n<li><p>Rather than simply sharing solutions, I thought it would be better to share my experience in detail. To improve the bad parts that happened in this competition.</p></li>\n<li><p>My ranking was due to luck, and although it is the Best Private Score submitted, it probably won't be Best Solution.</p></li>\n<li><p>Although the competition ended on a gloomy note due to the metric hacking issue, the host and staff worked hard to make it a meaningful and successful competition. </p></li>\n<li><p>Based on this, we expect a more developed system in the future.</p></li>\n</ul>\n<p>thank you</p>\n<blockquote>\n  <p>Edit : I posted the FE code here.<br>\n  644(for lgbm) - <a href=\"https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-risk-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-risk-lightgbm</a><br>\n  661(for cat &amp; dnn) - <a href=\"https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-catboost-inference\" target=\"_blank\">https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-catboost-inference</a></p>\n</blockquote>",
      "rawMarkdown": "##Foreword\n\nHello, this is yuuniee.\nThank you to the competition host Home Credit and staff, Kaggle staff and everyone involved.\n\nThe image below is a summary of my solution.\n\n![image_](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6960421%2F92721e389d44b29d04f931d244fed316%2F.png?generation=1716967303955938&alt=media)\n\nThe maximum Private LB Score for my model pool without Metric Hacking is close to about 0.53.\n\nAs everyone knows, this competition takes place in two phases.\nPhase 1 was ML (Machine Learning), and Phase 2 was MH (Metric Hack).\nI will briefly introduce each part.\n(If you are only curious about ML, please read only Phase 1.)\n\n\n## I. Phase 1 - ML (Machine Learning)\n\nThank you for writing a great notebook for EDA and solution construction at the beginning of the competition.\n@sergiosaharovskiy - https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission\n@greysky - https://www.kaggle.com/code/greysky/home-credit-baseline\nAdditionally, I would like to thank the many participants who provided various insights.\n\n1. The CV strategy is StratifiedGroupKFold, and has been tested and trained in various forms: No Shuffle and Shuffle. In my case, differences in CV of around 0.001 to 0.005 had a low correlation with LB, while differences above 0.01 had a high correlation with LB. In addition, CV improvement through model parameter adjustment had a relatively low correlation with LB, and CV improvement through FE had a relatively high correlation with LB.\n\n2. There's nothing special about Feature Engineering. Max, Min, Avg, Var, First, Last and Max-Min Difference of each item were utilized. (I will release the code after cleaning it up.)\n\n3. LGBM was a bit lacking compared to Catboost, but was good for the ensemble. The reason why the number of features is smaller is because some items (mainly categorical types) were excluded because they caused performance degradation and overfitting in LGBM, and other slightly different types of features were added and deleted.\n\n4. DNN is a Denselight model and utilizes the LightAutoML library. There were some bugs, so I had to fix some of them myself, but overall, I think it is a good library. At first, I planned to lightly test with Denselight and then build the final model with a larger model (FT-Transformer, etc.). Surprisingly, I haven't created or found a model that beats Denselight performance. Perhaps I made a mistake, or the nature of the data is sensitive to overfitting, so I guess a simpler model worked better.\n\n5. Catboost was the model that performed best in this competition and I guess it performed well on a large number of categorical features (about 117).\n\nDuring this period, my public LB was close to top 5 and the competition seemed to be going very smoothly.\nBut...\n\n\n>### + What didn't work\n>1. Transformer type models such as Tabnet, TabTransformer, FT-Transformer, etc.\n>2. Income, expenditure, and tax statistics by period (month, week, etc.).\n>3. Differences between individual income, expenses, and taxes.\n>4. K-means clustering\n\n\n## II. Phase 2 - MH(Metric Hack)\n\nPointed out problems with MH at the beginning of the competition,\n@at7459 - https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449\nAnnounced the revival of MH in the latter half of the competition,\n@johnpateha - https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167\n\nAlso, thanks to the many participants who suggested various other comments and solutions. Thanks to you, my knowledge has expanded.\nIn particular, the two people above are true scientists and journalists who could have hidden this discovery and profited from it, but pursued the public interest by exposing it to everyone. (Why not consider Kaggle giving a special award for this sharing?)\nAs a result, the problem was not solved, but we learned something(?) from it and at least I was able to take minimal measures.\n\nAfter @johnpateha's \"Metric hack again, sorry\" post, I tried Data Reversing and found that the difference between date_decision and min_refreshdate_3813885D had a high correlation (about 0.9 or more) with WEEK_NUM.\n(More information can be found at https://www.kaggle.com/code/eivolkova/how-to-restore-the-dates?scriptVersionId=180157891, published by @eivolkova immediately after the competition.)\nAnd the way suggested by @at7459:\n\n```python\nDEVIDE = 1/2\nREDUCE = 0.02\ncondition = df['WEEK_NUM'] < (df['WEEK_NUM'].max()-df['WEEK_NUM'].min())*DEVIDE+df['WEEK_NUM'].min()\ndf.loc[condition, 'score'] = (df.loc[condition, 'score'] - REDUCE).clip(0)\n```\n\nBased on this, the results of examining the score details submitted by changing the REDUCE and DEVIDE values ​​were as follows.\n\n1. The public/private division of test data spans the entire period and it is difficult to identify a specific distribution, and the optimal value in private will probably be distributed around 1/4 to 3/4.\n2. The optimal public LB value of REDUCE is 0.03 (at least in my model), and the optimal value in private is probably distributed around 0.02 to 0.04.\n\nThe difference in pure model performance between top competitors was 0.00X,\nSince the difference due to the above adjustment value was 0.0X,\nAt this stage it became gambling for me.\n\nSome will bet low, others will bet high.\nHere, I trusted my model and chose neutral, and chose DEVIDE=1/2, REDUCE=0.03, which is the median of the expected distribution.\nI expected someone who bet high or low to take the prize, and I expected about a 50% chance of finishing in the top 10.\n\n\n## III. Conclusion\n\n+ The first selected submission is the No Hack model with CV Best Score created in Phase 1, and the second selected submission is the application of Metric Hack to it. (I thought there was a high probability that Metric Hack would work, but I also submitted No Hack just in case)\n\n+ Rather than simply sharing solutions, I thought it would be better to share my experience in detail. To improve the bad parts that happened in this competition.\n\n+ My ranking was due to luck, and although it is the Best Private Score submitted, it probably won't be Best Solution.\n\n+ Although the competition ended on a gloomy note due to the metric hacking issue, the host and staff worked hard to make it a meaningful and successful competition. \n\n+ Based on this, we expect a more developed system in the future.\n\n\nthank you\n\n\n>Edit : I posted the FE code here.\n>644(for lgbm) - https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-risk-lightgbm\n>661(for cat & dnn) - https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-catboost-inference",
      "votes": null
    },
    {
      "id": "2842917",
      "postDate": "05/29/2024 10:37:11",
      "content": "<p>Congratulations on the placement <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> , how did you choose the features for each model?<br>\nIf possible, share your notebook with us.</p>",
      "rawMarkdown": "Congratulations on the placement @yuuniekiri , how did you choose the features for each model?\nIf possible, share your notebook with us.",
      "votes": null
    },
    {
      "id": "2842974",
      "postDate": "05/29/2024 11:24:24",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/thiagomantuani\" target=\"_blank\">@thiagomantuani</a>. A notebook link has been added.</p>",
      "rawMarkdown": "Thank you, @thiagomantuani. A notebook link has been added.",
      "votes": null
    },
    {
      "id": "2843263",
      "postDate": "05/29/2024 13:35:53",
      "content": "<p><strong>Congratulations on your second #1 rank competition <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>!</strong> Even though I didn't participate in the competition, I was following along since the beginning after my initial post some months ago. After reading your straightforward write-up and both notebooks, I totally agree with your choice of models, such as CatBoost, LGBM, and DNN (eventually NaiveBayes would've helped a bit). My only question that I have left is, did you try to use <code>rankdata()</code> after making the predictions? That gave me a huge boost in my first competition with almost exactly the same ensemble models. I'm just curious, if you don't mind, whenever you have time, could you please try a \"Late submission\" with the following updated part of your \"Fork of Home Credit Risk (LightGBM)\" at the very end:</p>\n<p><strong>Edit after <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> comment:</strong></p>\n<pre><code> scipy.stats  rankdata  \n\n reader\ngc.collect()\n\nranked_preds = rankdata(y_pred) / (y_pred)  \n\ndf_subm = pd.read_csv()\ndf_subm = df_subm.set_index()\ndf_subm[] = ranked_preds  \ndf_subm.to_csv()\n(df_subm)\n</code></pre>\n<p>Thanks a lot for your try-out and response. Enjoy your preliminary win!<br>\nBest regards,<br>\nEtienne</p>",
      "rawMarkdown": "**Congratulations on your second #1 rank competition @yuuniekiri!** Even though I didn't participate in the competition, I was following along since the beginning after my initial post some months ago. After reading your straightforward write-up and both notebooks, I totally agree with your choice of models, such as CatBoost, LGBM, and DNN (eventually NaiveBayes would've helped a bit). My only question that I have left is, did you try to use `rankdata()` after making the predictions? That gave me a huge boost in my first competition with almost exactly the same ensemble models. I'm just curious, if you don't mind, whenever you have time, could you please try a \"Late submission\" with the following updated part of your \"Fork of Home Credit Risk (LightGBM)\" at the very end:\n\n**Edit after @yuuniekiri comment:**\n\n```python\nfrom scipy.stats import rankdata  # <- newly added, import rankdata\n\ndel reader\ngc.collect()\n\nranked_preds = rankdata(y_pred) / len(y_pred)  # <- newly added\n\ndf_subm = pd.read_csv(f\"{ROOT}/sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm[\"score\"] = ranked_preds  # <- newly adjusted\ndf_subm.to_csv(\"submission.csv\")\nprint(df_subm)\n```\n\nThanks a lot for your try-out and response. Enjoy your preliminary win!\nBest regards,\nEtienne",
      "votes": null
    },
    {
      "id": "2843306",
      "postDate": "05/29/2024 14:02:29",
      "content": "<p>Thank you for your congratulations. <br>\nHowever, the submissions resulting from the code above are different from the competition format.<br>\n(Only two columns are allowed: case_id and score.)<br>\nIt will probably cause an error.</p>",
      "rawMarkdown": "Thank you for your congratulations. \nHowever, the submissions resulting from the code above are different from the competition format.\n(Only two columns are allowed: case_id and score.)\nIt will probably cause an error.",
      "votes": null
    },
    {
      "id": "2843340",
      "postDate": "05/29/2024 14:22:38",
      "content": "<p>You're welcome. That's right, I had a logical mistake. We don't need an additional column, I updated my code above (Edit). I looked into my old competition and that's the way how I made it there. Now the predictions should be averaged ranked. Give it a shot, if it doesn't work it's also fine - just for quick testing purpose :) Thanks!</p>",
      "rawMarkdown": "You're welcome. That's right, I had a logical mistake. We don't need an additional column, I updated my code above (Edit). I looked into my old competition and that's the way how I made it there. Now the predictions should be averaged ranked. Give it a shot, if it doesn't work it's also fine - just for quick testing purpose :) Thanks!",
      "votes": null
    },
    {
      "id": "2844027",
      "postDate": "05/29/2024 20:57:16",
      "content": "<p>Congratulations! Thanks for the detailed explanation. Finally, the winning solution used the metric hack as was expected for some Kaggle experts.</p>",
      "rawMarkdown": "Congratulations! Thanks for the detailed explanation. Finally, the winning solution used the metric hack as was expected for some Kaggle experts.",
      "votes": null
    },
    {
      "id": "2844034",
      "postDate": "05/29/2024 21:02:01",
      "content": "<p>This is awesome! I'll be reading this in detail soon! :) Great work <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>!</p>",
      "rawMarkdown": "This is awesome! I'll be reading this in detail soon! :) Great work @yuuniekiri!",
      "votes": null
    },
    {
      "id": "2844081",
      "postDate": "05/29/2024 22:38:52",
      "content": "<p>Hi! May i know what's the logic of it? Isn't all the scores equal to the mean df_subm[\"score\"] = final_preds_ranked? Thank you!</p>",
      "rawMarkdown": "Hi! May i know what's the logic of it? Isn't all the scores equal to the mean df_subm[\"score\"] = final_preds_ranked? Thank you!",
      "votes": null
    },
    {
      "id": "2844111",
      "postDate": "05/29/2024 23:59:54",
      "content": "<p>Congratulations! This notebook is very impressive for me. </p>",
      "rawMarkdown": "Congratulations! This notebook is very impressive for me.",
      "votes": null
    },
    {
      "id": "2844391",
      "postDate": "05/30/2024 04:19:02",
      "content": "<p>Congratulations on winning this competition. Thanks for detailed explanation of your solution and sharing valuable comments on topic of Metric Hacking. Can you comment on execution time of your notebook and share best practices for the competition participants. </p>",
      "rawMarkdown": "Congratulations on winning this competition. Thanks for detailed explanation of your solution and sharing valuable comments on topic of Metric Hacking. Can you comment on execution time of your notebook and share best practices for the competition participants.",
      "votes": null
    },
    {
      "id": "2844427",
      "postDate": "05/30/2024 04:30:15",
      "content": "<p>Congratulations! I learn a lot~</p>",
      "rawMarkdown": "Congratulations! I learn a lot~",
      "votes": null
    },
    {
      "id": "2844545",
      "postDate": "05/30/2024 06:03:01",
      "content": "<p>Thank you very much. </p>\n<ol>\n<li><p>The execution time of my notes is approximately 2 to 3 hours. (I have never measured it accurately.)</p></li>\n<li><p>Best practice - I don't know if I'm qualified to comment on this. It's difficult to pinpoint something, but since financial data has a lot of noise and is vulnerable to overfitting, I think we need to pay special attention to generalization. For example, you should be very careful about randomly inserting unexplained data (but this may not always be accurate). In addition, we explore various methods to prevent overfitting.</p></li>\n</ol>",
      "rawMarkdown": "Thank you very much. \n\n1. The execution time of my notes is approximately 2 to 3 hours. (I have never measured it accurately.)\n\n2. Best practice - I don't know if I'm qualified to comment on this. It's difficult to pinpoint something, but since financial data has a lot of noise and is vulnerable to overfitting, I think we need to pay special attention to generalization. For example, you should be very careful about randomly inserting unexplained data (but this may not always be accurate). In addition, we explore various methods to prevent overfitting.",
      "votes": null
    },
    {
      "id": "2844603",
      "postDate": "05/30/2024 06:39:27",
      "content": "<p>Your idea is so amazing. Thank you for sharing awesome idea. </p>",
      "rawMarkdown": "Your idea is so amazing. Thank you for sharing awesome idea.",
      "votes": null
    },
    {
      "id": "2844613",
      "postDate": "05/30/2024 06:49:12",
      "content": "<p><a href=\"https://www.kaggle.com/mengjingyang\" target=\"_blank\">@mengjingyang</a> Thanks for your question. Before using <code>rankdata()</code>, the score for each test case is just the raw probability. These probabilities are averaged across all models, resulting in a single score for each test case. However, without <code>rankdata()</code>, these scores can still be skewed or affected by outliers.</p>\n<p>By applying <code>rankdata()</code>, the predictions are converted into ranks and then normalized. This process makes the scores more robust by reducing the influence of outliers and ensures a uniform distribution (relative) instead of an absolute distribution (which can be skewed). Hope that clarifies to understand why I thought of that.</p>",
      "rawMarkdown": "mengjingyang Thanks for your question. Before using `rankdata()`, the score for each test case is just the raw probability. These probabilities are averaged across all models, resulting in a single score for each test case. However, without `rankdata()`, these scores can still be skewed or affected by outliers.\n\nBy applying `rankdata()`, the predictions are converted into ranks and then normalized. This process makes the scores more robust by reducing the influence of outliers and ensures a uniform distribution (relative) instead of an absolute distribution (which can be skewed). Hope that clarifies to understand why I thought of that.",
      "votes": null
    },
    {
      "id": "2844674",
      "postDate": "05/30/2024 07:26:14",
      "content": "<p><code>In addition, CV improvement through model parameter adjustment had a relatively low correlation with LB, and CV improvement through FE had a relatively high correlation with LB.</code></p>\n<p>This is consistently my experience as well on various problems (not only this competition).</p>",
      "rawMarkdown": "```In addition, CV improvement through model parameter adjustment had a relatively low correlation with LB, and CV improvement through FE had a relatively high correlation with LB.```\n\nThis is consistently my experience as well on various problems (not only this competition).",
      "votes": null
    },
    {
      "id": "2844679",
      "postDate": "05/30/2024 07:29:12",
      "content": "<p>Thanks for sharing - this is great. One question out of curiosity: have you managed to use credit bureau A data successfully?<br>\n(for context - I haven't - my post with details here <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360</a>)</p>",
      "rawMarkdown": "Thanks for sharing - this is great. One question out of curiosity: have you managed to use credit bureau A data successfully?\n(for context - I haven't - my post with details here https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360)",
      "votes": null
    },
    {
      "id": "2844882",
      "postDate": "05/30/2024 09:13:23",
      "content": "<p>Congratulations , thanks for sharing your wonderfull journay!!</p>",
      "rawMarkdown": "Congratulations , thanks for sharing your wonderfull journay!!",
      "votes": null
    },
    {
      "id": "2844897",
      "postDate": "05/30/2024 09:28:31",
      "content": "<p>Congrats Yuuniee!! From your shared code I assume you used several models for each model category (LGBM, DNN, Catboost). Were those models the result of each folod of the CV process or were models result of training on the whole train dataset but with different params?</p>",
      "rawMarkdown": "Congrats Yuuniee!! From your shared code I assume you used several models for each model category (LGBM, DNN, Catboost). Were those models the result of each folod of the CV process or were models result of training on the whole train dataset but with different params?",
      "votes": null
    },
    {
      "id": "2844941",
      "postDate": "05/30/2024 09:55:52",
      "content": "<p>Thanks a lot. <br>\nThe CV score above is the score for all out of fold(OOF) of 5folds at StratifiedGroupKFold (no shuffle).</p>",
      "rawMarkdown": "Thanks a lot. \nThe CV score above is the score for all out of fold(OOF) of 5folds at StratifiedGroupKFold (no shuffle).",
      "votes": null
    },
    {
      "id": "2844961",
      "postDate": "05/30/2024 10:05:57",
      "content": "<p>This is quite a mystery to me too. <br>\nAs you can see from the code and comments for LGBM posted above, I had quite a bit of trouble with it.<br>\nAs a result, many features were deleted from bureau A and only some were retained. <br>\n(And I experienced something similar with person data.)<br>\nBecause test data cannot be seen, it is difficult to determine the cause.</p>",
      "rawMarkdown": "This is quite a mystery to me too. \nAs you can see from the code and comments for LGBM posted above, I had quite a bit of trouble with it.\nAs a result, many features were deleted from bureau A and only some were retained. \n(And I experienced something similar with person data.)\nBecause test data cannot be seen, it is difficult to determine the cause.",
      "votes": null
    },
    {
      "id": "2845079",
      "postDate": "05/30/2024 11:42:27",
      "content": "<p>From Credit Bureau data I just used the feature related to the financial institution that approved the loan. Other features degraded the stability metric. I also thought it was better not to use debit card data. On the other hand, the best way I  found to use tax data, was the sum of all deduction tax of each client, it makes sense because it represents the tax deductions in the last year. That feature boosted my model's performance.</p>",
      "rawMarkdown": "From Credit Bureau data I just used the feature related to the financial institution that approved the loan. Other features degraded the stability metric. I also thought it was better not to use debit card data. On the other hand, the best way I  found to use tax data, was the sum of all deduction tax of each client, it makes sense because it represents the tax deductions in the last year. That feature boosted my model's performance.",
      "votes": null
    },
    {
      "id": "2845269",
      "postDate": "05/30/2024 13:42:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>,</p>\n<p>Congrats with the great performance and amazing solution! Also I would like to thank you for using LightAutoML models in your solution 😎</p>\n<blockquote>\n  <p>DNN is a Denselight model and utilizes the Lightautoml library. There were some bugs, so I had to fix some of them myself, but overall, I think it is a good library. </p>\n</blockquote>\n<p>Could you please share the problems you had with our library using <a href=\"https://github.com/sb-ai-lab/LightAutoML/issues\" target=\"_blank\">GitHub issue mechanism</a>? It will help us to make LightAutoML better 😊</p>\n<p>P.S. If you already have fixes for the problems and you would like to share them, please use the <a href=\"https://github.com/sb-ai-lab/LightAutoML/pulls\" target=\"_blank\">Pull Request mechanism</a></p>\n<p>Alex<br>\n(Head of LightAutoML team)</p>",
      "rawMarkdown": "Hi @yuuniekiri,\n\nCongrats with the great performance and amazing solution! Also I would like to thank you for using LightAutoML models in your solution 😎\n\n>DNN is a Denselight model and utilizes the Lightautoml library. There were some bugs, so I had to fix some of them myself, but overall, I think it is a good library. \n\nCould you please share the problems you had with our library using [GitHub issue mechanism](https://github.com/sb-ai-lab/LightAutoML/issues)? It will help us to make LightAutoML better 😊\n\nP.S. If you already have fixes for the problems and you would like to share them, please use the [Pull Request mechanism](https://github.com/sb-ai-lab/LightAutoML/pulls)\n\nAlex\n(Head of LightAutoML team)",
      "votes": null
    },
    {
      "id": "2845308",
      "postDate": "05/30/2024 13:57:36",
      "content": "<p>Amazing <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> ! Congrats!</p>\n<p>How did you manage OOM errors?</p>",
      "rawMarkdown": "Amazing @yuuniekiri ! Congrats!\n\nHow did you manage OOM errors?",
      "votes": null
    },
    {
      "id": "2845622",
      "postDate": "05/30/2024 15:52:55",
      "content": "<p>Wow! Quite a nice job! Congrats!!!</p>",
      "rawMarkdown": "Wow! Quite a nice job! Congrats!!!",
      "votes": null
    },
    {
      "id": "2845692",
      "postDate": "05/30/2024 16:28:55",
      "content": "<p>My interpretation is that hosts downsampled this source of data heavily for the test set, and hence it led to overfitting.</p>",
      "rawMarkdown": "My interpretation is that hosts downsampled this source of data heavily for the test set, and hence it led to overfitting.",
      "votes": null
    },
    {
      "id": "2845807",
      "postDate": "05/30/2024 17:34:24",
      "content": "<p>Let me add my 5 cents. With MLP model I also had problems with overfitting, so I used only top 20 features (by gain importance of LGBM). But GBDT models worked pretty well with cred_b_a_* tables, especially one version of CatBoost with a private score of 0.520. Of course some features were filtered, but nothing special (maybe with the exception of additional aggregators that are applied to some num cols). In ensemble, these models (using cred_b_a_* tables) worked well.</p>\n<p>And also in early stages I filtered out some rows with null cols greater then some threshold, but then I think I just found out combination of aggregators that worked well with all tables and rows of cred_b_a_* data.</p>",
      "rawMarkdown": "Let me add my 5 cents. With MLP model I also had problems with overfitting, so I used only top 20 features (by gain importance of LGBM). But GBDT models worked pretty well with cred_b_a_* tables, especially one version of CatBoost with a private score of 0.520. Of course some features were filtered, but nothing special (maybe with the exception of additional aggregators that are applied to some num cols). In ensemble, these models (using cred_b_a_* tables) worked well.\n\nAnd also in early stages I filtered out some rows with null cols greater then some threshold, but then I think I just found out combination of aggregators that worked well with all tables and rows of cred_b_a_* data.",
      "votes": null
    },
    {
      "id": "2845822",
      "postDate": "05/30/2024 17:48:55",
      "content": "<p>I use credit_bureau_a_1_3 for training instead of credit_bureau_a_1_*.</p>",
      "rawMarkdown": "I use credit_bureau_a_1_3 for training instead of credit_bureau_a_1_*.",
      "votes": null
    },
    {
      "id": "2845992",
      "postDate": "05/30/2024 19:45:05",
      "content": "<p>Can you elaborate?</p>",
      "rawMarkdown": "Can you elaborate?",
      "votes": null
    },
    {
      "id": "2846139",
      "postDate": "05/31/2024 00:03:40",
      "content": "<p>Congratulations on your award! <br>\nThank you for the solution!  I will keep this in mind for future reference.</p>",
      "rawMarkdown": "Congratulations on your award! \nThank you for the solution!  I will keep this in mind for future reference.",
      "votes": null
    },
    {
      "id": "2846187",
      "postDate": "05/31/2024 02:21:03",
      "content": "<p>Congratulations, thanks for sharing. <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> </p>",
      "rawMarkdown": "Congratulations, thanks for sharing. @yuuniekiri",
      "votes": null
    },
    {
      "id": "2846194",
      "postDate": "05/31/2024 02:27:56",
      "content": "<p>Thank you so much.</p>\n<p>LightAutoML is a really convenient and useful library.<br>\nThe problem I had was not fatal, but I will share it there soon. :)</p>",
      "rawMarkdown": "Thank you so much.\n\nLightAutoML is a really convenient and useful library.\nThe problem I had was not fatal, but I will share it there soon. :)",
      "votes": null
    },
    {
      "id": "2846201",
      "postDate": "05/31/2024 02:40:09",
      "content": "<p>Thank you for your congrats. </p>\n<p>Except when there are more than 1000 features, I have not experienced OOM (Out of memory). But if you're having trouble, the following methods may help:</p>\n<blockquote>\n  <ol>\n  <li>Capture as many useless features as possible and remove them all.</li>\n  <li>Use reduce_memory_useage in the step of reading each file.</li>\n  <li>In addition to RAM, GPU RAM is used together. (A representative example is the cuDF library created by NVIDIA's RAPIDS team.)</li>\n  <li>When data processing is complete, save the data locally as a file. Then, delete the DataFrame and reset the memory with commands such as gc.collect() and reset.</li>\n  <li>In the prediction step, locally stored files are read in chunks and each is predicted.</li>\n  </ol>\n</blockquote>",
      "rawMarkdown": "Thank you for your congrats. \n\nExcept when there are more than 1000 features, I have not experienced OOM (Out of memory). But if you're having trouble, the following methods may help:\n\n>1. Capture as many useless features as possible and remove them all.\n>2. Use reduce_memory_useage in the step of reading each file.\n>3. In addition to RAM, GPU RAM is used together. (A representative example is the cuDF library created by NVIDIA's RAPIDS team.)\n>4. When data processing is complete, save the data locally as a file. Then, delete the DataFrame and reset the memory with commands such as gc.collect() and reset.\n>5. In the prediction step, locally stored files are read in chunks and each is predicted.",
      "votes": null
    },
    {
      "id": "2846220",
      "postDate": "05/31/2024 03:03:17",
      "content": "<p>Congratulations. Thank you for sharing.</p>",
      "rawMarkdown": "Congratulations. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "2846286",
      "postDate": "05/31/2024 04:01:16",
      "content": "<p>Thank you for clear cut explanation</p>",
      "rawMarkdown": "Thank you for clear cut explanation",
      "votes": null
    },
    {
      "id": "2846477",
      "postDate": "05/31/2024 06:00:01",
      "content": "<p>great hacking man</p>",
      "rawMarkdown": "great hacking man",
      "votes": null
    },
    {
      "id": "2846618",
      "postDate": "05/31/2024 07:08:27",
      "content": "<p>Hi, we will have to look into it, in production models credit bureau contributes quite significantly…<br>\nin dataset should be data as they are in our real processes (except for anonymization/masking), and there were no adjustments like downsampling in test set.</p>",
      "rawMarkdown": "Hi, we will have to look into it, in production models credit bureau contributes quite significantly...\nin dataset should be data as they are in our real processes (except for anonymization/masking), and there were no adjustments like downsampling in test set.",
      "votes": null
    },
    {
      "id": "2846640",
      "postDate": "05/31/2024 07:23:09",
      "content": "<p>Thank you! This logic makes sense.  So ranked_preds is a 2d array(N * model_number) here? </p>",
      "rawMarkdown": "Thank you! This logic makes sense.  So ranked_preds is a 2d array(N * model_number) here?",
      "votes": null
    },
    {
      "id": "2846673",
      "postDate": "05/31/2024 07:39:50",
      "content": "<p>You're welcome! No, <code>ranked_preds</code> is not a 2D array. It's a 1D array that represents the normalized ranks of the combined predictions from <strong>all</strong> models.</p>",
      "rawMarkdown": "You're welcome! No, `ranked_preds` is not a 2D array. It's a 1D array that represents the normalized ranks of the combined predictions from **all** models.",
      "votes": null
    },
    {
      "id": "2847445",
      "postDate": "05/31/2024 14:39:11",
      "content": "<p>Interesting</p>",
      "rawMarkdown": "Interesting",
      "votes": null
    },
    {
      "id": "2847765",
      "postDate": "05/31/2024 17:14:49",
      "content": "<p>Congratulations! My team had similar idea - Ensemble model of LightGBM, CatBoost and Tensorflow. We were really close to developing the ensemble model. However after the change of rules and allowing the metric hack, our focus shifted on the metric hack.</p>",
      "rawMarkdown": "Congratulations! My team had similar idea - Ensemble model of LightGBM, CatBoost and Tensorflow. We were really close to developing the ensemble model. However after the change of rules and allowing the metric hack, our focus shifted on the metric hack.",
      "votes": null
    },
    {
      "id": "2847982",
      "postDate": "05/31/2024 18:37:21",
      "content": "<p>I see. So final_preds_ranked is a number. I am still confused as to why you set the entire score column to the same number. Thank you again! </p>",
      "rawMarkdown": "I see. So final_preds_ranked is a number. I am still confused as to why you set the entire score column to the same number. Thank you again!",
      "votes": null
    },
    {
      "id": "2848059",
      "postDate": "05/31/2024 19:05:04",
      "content": "<p>Great approach</p>",
      "rawMarkdown": "Great approach",
      "votes": null
    },
    {
      "id": "2848245",
      "postDate": "05/31/2024 23:36:47",
      "content": "<p>helpful for me!</p>",
      "rawMarkdown": "helpful for me!",
      "votes": null
    },
    {
      "id": "2849145",
      "postDate": "06/01/2024 12:33:17",
      "content": "<p>Congratulations on your award!<br>\nThank you for the solution! Very interesting approach</p>",
      "rawMarkdown": "Congratulations on your award!\nThank you for the solution! Very interesting approach",
      "votes": null
    },
    {
      "id": "2849248",
      "postDate": "06/01/2024 13:28:13",
      "content": "<p>Great work man!</p>",
      "rawMarkdown": "Great work man!",
      "votes": null
    },
    {
      "id": "2850116",
      "postDate": "06/02/2024 00:10:08",
      "content": "<p>Smart way of using the hacks!</p>",
      "rawMarkdown": "Smart way of using the hacks!",
      "votes": null
    },
    {
      "id": "2851303",
      "postDate": "06/02/2024 17:02:24",
      "content": "<p>Thanks! Very useful article for fresh man</p>",
      "rawMarkdown": "Thanks! Very useful article for fresh man",
      "votes": null
    },
    {
      "id": "2851969",
      "postDate": "06/03/2024 04:24:18",
      "content": "<p>a very good guide ,thanks</p>",
      "rawMarkdown": "a very good guide ,thanks",
      "votes": null
    },
    {
      "id": "2852381",
      "postDate": "06/03/2024 08:58:11",
      "content": "<p>Thanks! Very useful article for fresh man.</p>",
      "rawMarkdown": "Thanks! Very useful article for fresh man.",
      "votes": null
    },
    {
      "id": "2852723",
      "postDate": "06/03/2024 12:29:02",
      "content": "<p>선생님처럼 할려면 얼마나 해야할까요..</p>",
      "rawMarkdown": "선생님처럼 할려면 얼마나 해야할까요..",
      "votes": null
    },
    {
      "id": "2852941",
      "postDate": "06/03/2024 14:24:27",
      "content": "<p>Participate actively in the competition.<br>\nI think it's best to have fun!</p>",
      "rawMarkdown": "Participate actively in the competition.\nI think it's best to have fun!",
      "votes": null
    },
    {
      "id": "2852944",
      "postDate": "06/03/2024 14:26:11",
      "content": "<p>Keep it in English, please.</p>",
      "rawMarkdown": "Keep it in English, please.",
      "votes": null
    },
    {
      "id": "2853669",
      "postDate": "06/03/2024 21:18:45",
      "content": "<p>thank you for the summary. Very inspiring </p>",
      "rawMarkdown": "thank you for the summary. Very inspiring",
      "votes": null
    },
    {
      "id": "2856429",
      "postDate": "06/05/2024 09:14:15",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> for the winning. May I ask you what approach did you choose to find the best ensemble weights? </p>",
      "rawMarkdown": "Congrats @yuuniekiri for the winning. May I ask you what approach did you choose to find the best ensemble weights?",
      "votes": null
    },
    {
      "id": "2856467",
      "postDate": "06/05/2024 10:01:55",
      "content": "<p>Thank you for your congrats.<br>\nMy approach is linear regression and weighted testing in OOF.<br>\nInterestingly, the optimal weights of OOF score was close to the optimal weights of LB.</p>",
      "rawMarkdown": "Thank you for your congrats.\nMy approach is linear regression and weighted testing in OOF.\nInterestingly, the optimal weights of OOF score was close to the optimal weights of LB.",
      "votes": null
    },
    {
      "id": "2858305",
      "postDate": "06/06/2024 12:04:12",
      "content": "<p>Thank you for sharing such a detailed and insightful breakdown of your solution, yuuniee! Your approach to combining LGBM, DNN, and Catboost models with a weighted ensemble clearly highlights the importance of diversity in modeling techniques.</p>\n<p>I particularly appreciate the transparency in discussing both phases of the competition—Machine Learning and Metric Hacking. Your use of StratifiedGroupKFold and various feature engineering techniques, along with your observations about the correlation between CV improvements and LB performance, are incredibly valuable for the community.</p>\n<p>The post-processing strategy and the thoughtful discussion on the impact of WEEK_NUM adjustments and the choices around DEVIDE and REDUCE values offer practical insights into handling similar challenges.</p>\n<p>It's also commendable how you acknowledged the contributions and insights from other participants like <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> and <a href=\"https://www.kaggle.com/greysky\" target=\"_blank\">@greysky</a> for their foundational notebooks, and <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> and <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> for their critical discussions around Metric Hacking. The collaboration and knowledge sharing within this community are what make these competitions so enriching.</p>\n<p>Despite the challenges posed by Metric Hacking, your balanced approach to selecting submission strategies demonstrates a clear understanding of risk management in competition settings.</p>",
      "rawMarkdown": "Thank you for sharing such a detailed and insightful breakdown of your solution, yuuniee! Your approach to combining LGBM, DNN, and Catboost models with a weighted ensemble clearly highlights the importance of diversity in modeling techniques.\n\nI particularly appreciate the transparency in discussing both phases of the competition—Machine Learning and Metric Hacking. Your use of StratifiedGroupKFold and various feature engineering techniques, along with your observations about the correlation between CV improvements and LB performance, are incredibly valuable for the community.\n\nThe post-processing strategy and the thoughtful discussion on the impact of WEEK_NUM adjustments and the choices around DEVIDE and REDUCE values offer practical insights into handling similar challenges.\n\nIt's also commendable how you acknowledged the contributions and insights from other participants like @sergiosaharovskiy and @greysky for their foundational notebooks, and @at7459 and @johnpateha for their critical discussions around Metric Hacking. The collaboration and knowledge sharing within this community are what make these competitions so enriching.\n\nDespite the challenges posed by Metric Hacking, your balanced approach to selecting submission strategies demonstrates a clear understanding of risk management in competition settings.",
      "votes": null
    },
    {
      "id": "2865334",
      "postDate": "06/10/2024 17:01:02",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": null
    },
    {
      "id": "2875243",
      "postDate": "06/17/2024 01:10:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> for the winning. How do we filter out effective features and toxic features?Thank you for your response in advance.</p>",
      "rawMarkdown": "Congrats @yuuniekiri for the winning. How do we filter out effective features and toxic features?Thank you for your response in advance.",
      "votes": null
    },
    {
      "id": "2875270",
      "postDate": "06/17/2024 02:01:37",
      "content": "<p>As you can see in the article above, features that have a significant advantage in the CV score are kept, and features that have no advantage are removed.<br>\nAdditionally, even if there is no significant gain in score, logically meaningful and stable features can be kept. This can also apply in the opposite case.</p>",
      "rawMarkdown": "As you can see in the article above, features that have a significant advantage in the CV score are kept, and features that have no advantage are removed.\nAdditionally, even if there is no significant gain in score, logically meaningful and stable features can be kept. This can also apply in the opposite case.",
      "votes": null
    },
    {
      "id": "2880567",
      "postDate": "06/20/2024 08:49:52",
      "content": "<p>Nice share, Congratulations for your first solo gold.</p>",
      "rawMarkdown": "Nice share, Congratulations for your first solo gold.",
      "votes": null
    },
    {
      "id": "2889766",
      "postDate": "06/25/2024 16:53:07",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": null
    },
    {
      "id": "2904468",
      "postDate": "07/04/2024 11:27:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>, I cannot find the final submission notebook for the first place. Can you please share it?</p>",
      "rawMarkdown": "Hi @yuuniekiri, I cannot find the final submission notebook for the first place. Can you please share it?",
      "votes": null
    },
    {
      "id": "3012991",
      "postDate": "10/09/2024 14:54:37",
      "content": "<p>Thanks for your share!</p>",
      "rawMarkdown": "Thanks for your share!",
      "votes": null
    },
    {
      "id": "3135086",
      "postDate": "02/27/2025 03:49:04",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>, Congratulations on your win. I love the way you approach Feature Engineering, especially Aggregation. I was wondering—if all feature names were hidden, how would you do in Aggregation?</p>",
      "rawMarkdown": "Hi @yuuniekiri, Congratulations on your win. I love the way you approach Feature Engineering, especially Aggregation. I was wondering—if all feature names were hidden, how would you do in Aggregation?",
      "votes": null
    },
    {
      "id": "3312920",
      "postDate": "11/08/2025 08:47:50",
      "content": "<p>Nice, keep doing well</p>",
      "rawMarkdown": "Nice, keep doing well",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2842917,
      "author_name": "thiagomantuani",
      "author_url": "",
      "post_date": "05/29/2024 10:37:11",
      "content": "<p>Congratulations on the placement <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> , how did you choose the features for each model?<br>\nIf possible, share your notebook with us.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2842974,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "05/29/2024 11:24:24",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/thiagomantuani\" target=\"_blank\">@thiagomantuani</a>. A notebook link has been added.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2843263,
      "author_name": "etiennekaiser",
      "author_url": "",
      "post_date": "05/29/2024 13:35:53",
      "content": "<p><strong>Congratulations on your second #1 rank competition <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>!</strong> Even though I didn't participate in the competition, I was following along since the beginning after my initial post some months ago. After reading your straightforward write-up and both notebooks, I totally agree with your choice of models, such as CatBoost, LGBM, and DNN (eventually NaiveBayes would've helped a bit). My only question that I have left is, did you try to use <code>rankdata()</code> after making the predictions? That gave me a huge boost in my first competition with almost exactly the same ensemble models. I'm just curious, if you don't mind, whenever you have time, could you please try a \"Late submission\" with the following updated part of your \"Fork of Home Credit Risk (LightGBM)\" at the very end:</p>\n<p><strong>Edit after <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> comment:</strong></p>\n<pre><code> scipy.stats  rankdata  \n\n reader\ngc.collect()\n\nranked_preds = rankdata(y_pred) / (y_pred)  \n\ndf_subm = pd.read_csv()\ndf_subm = df_subm.set_index()\ndf_subm[] = ranked_preds  \ndf_subm.to_csv()\n(df_subm)\n</code></pre>\n<p>Thanks a lot for your try-out and response. Enjoy your preliminary win!<br>\nBest regards,<br>\nEtienne</p>",
      "votes": null,
      "replies": [
        {
          "id": 2843306,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "05/29/2024 14:02:29",
          "content": "<p>Thank you for your congratulations. <br>\nHowever, the submissions resulting from the code above are different from the competition format.<br>\n(Only two columns are allowed: case_id and score.)<br>\nIt will probably cause an error.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2843340,
              "author_name": "etiennekaiser",
              "author_url": "",
              "post_date": "05/29/2024 14:22:38",
              "content": "<p>You're welcome. That's right, I had a logical mistake. We don't need an additional column, I updated my code above (Edit). I looked into my old competition and that's the way how I made it there. Now the predictions should be averaged ranked. Give it a shot, if it doesn't work it's also fine - just for quick testing purpose :) Thanks!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2844081,
                  "author_name": "mengjingyang",
                  "author_url": "",
                  "post_date": "05/29/2024 22:38:52",
                  "content": "<p>Hi! May i know what's the logic of it? Isn't all the scores equal to the mean df_subm[\"score\"] = final_preds_ranked? Thank you!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2844613,
                      "author_name": "etiennekaiser",
                      "author_url": "",
                      "post_date": "05/30/2024 06:49:12",
                      "content": "<p><a href=\"https://www.kaggle.com/mengjingyang\" target=\"_blank\">@mengjingyang</a> Thanks for your question. Before using <code>rankdata()</code>, the score for each test case is just the raw probability. These probabilities are averaged across all models, resulting in a single score for each test case. However, without <code>rankdata()</code>, these scores can still be skewed or affected by outliers.</p>\n<p>By applying <code>rankdata()</code>, the predictions are converted into ranks and then normalized. This process makes the scores more robust by reducing the influence of outliers and ensures a uniform distribution (relative) instead of an absolute distribution (which can be skewed). Hope that clarifies to understand why I thought of that.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2846640,
                          "author_name": "mengjingyang",
                          "author_url": "",
                          "post_date": "05/31/2024 07:23:09",
                          "content": "<p>Thank you! This logic makes sense.  So ranked_preds is a 2d array(N * model_number) here? </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2846673,
                              "author_name": "etiennekaiser",
                              "author_url": "",
                              "post_date": "05/31/2024 07:39:50",
                              "content": "<p>You're welcome! No, <code>ranked_preds</code> is not a 2D array. It's a 1D array that represents the normalized ranks of the combined predictions from <strong>all</strong> models.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2847982,
                                  "author_name": "mengjingyang",
                                  "author_url": "",
                                  "post_date": "05/31/2024 18:37:21",
                                  "content": "<p>I see. So final_preds_ranked is a number. I am still confused as to why you set the entire score column to the same number. Thank you again! </p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2844027,
      "author_name": "eduardotoloza",
      "author_url": "",
      "post_date": "05/29/2024 20:57:16",
      "content": "<p>Congratulations! Thanks for the detailed explanation. Finally, the winning solution used the metric hack as was expected for some Kaggle experts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844034,
      "author_name": "vishrutgrover",
      "author_url": "",
      "post_date": "05/29/2024 21:02:01",
      "content": "<p>This is awesome! I'll be reading this in detail soon! :) Great work <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844111,
      "author_name": "chrischan0204",
      "author_url": "",
      "post_date": "05/29/2024 23:59:54",
      "content": "<p>Congratulations! This notebook is very impressive for me. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844391,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "05/30/2024 04:19:02",
      "content": "<p>Congratulations on winning this competition. Thanks for detailed explanation of your solution and sharing valuable comments on topic of Metric Hacking. Can you comment on execution time of your notebook and share best practices for the competition participants. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2844545,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "05/30/2024 06:03:01",
          "content": "<p>Thank you very much. </p>\n<ol>\n<li><p>The execution time of my notes is approximately 2 to 3 hours. (I have never measured it accurately.)</p></li>\n<li><p>Best practice - I don't know if I'm qualified to comment on this. It's difficult to pinpoint something, but since financial data has a lot of noise and is vulnerable to overfitting, I think we need to pay special attention to generalization. For example, you should be very careful about randomly inserting unexplained data (but this may not always be accurate). In addition, we explore various methods to prevent overfitting.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2844427,
      "author_name": "ddooyeong",
      "author_url": "",
      "post_date": "05/30/2024 04:30:15",
      "content": "<p>Congratulations! I learn a lot~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844603,
      "author_name": "ishlove7",
      "author_url": "",
      "post_date": "05/30/2024 06:39:27",
      "content": "<p>Your idea is so amazing. Thank you for sharing awesome idea. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844674,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "05/30/2024 07:26:14",
      "content": "<p><code>In addition, CV improvement through model parameter adjustment had a relatively low correlation with LB, and CV improvement through FE had a relatively high correlation with LB.</code></p>\n<p>This is consistently my experience as well on various problems (not only this competition).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844679,
      "author_name": "narsil",
      "author_url": "",
      "post_date": "05/30/2024 07:29:12",
      "content": "<p>Thanks for sharing - this is great. One question out of curiosity: have you managed to use credit bureau A data successfully?<br>\n(for context - I haven't - my post with details here <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360</a>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2844961,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "05/30/2024 10:05:57",
          "content": "<p>This is quite a mystery to me too. <br>\nAs you can see from the code and comments for LGBM posted above, I had quite a bit of trouble with it.<br>\nAs a result, many features were deleted from bureau A and only some were retained. <br>\n(And I experienced something similar with person data.)<br>\nBecause test data cannot be seen, it is difficult to determine the cause.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2845079,
              "author_name": "eduardotoloza",
              "author_url": "",
              "post_date": "05/30/2024 11:42:27",
              "content": "<p>From Credit Bureau data I just used the feature related to the financial institution that approved the loan. Other features degraded the stability metric. I also thought it was better not to use debit card data. On the other hand, the best way I  found to use tax data, was the sum of all deduction tax of each client, it makes sense because it represents the tax deductions in the last year. That feature boosted my model's performance.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2845692,
              "author_name": "narsil",
              "author_url": "",
              "post_date": "05/30/2024 16:28:55",
              "content": "<p>My interpretation is that hosts downsampled this source of data heavily for the test set, and hence it led to overfitting.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2846618,
                  "author_name": "tomasjeline2",
                  "author_url": "",
                  "post_date": "05/31/2024 07:08:27",
                  "content": "<p>Hi, we will have to look into it, in production models credit bureau contributes quite significantly…<br>\nin dataset should be data as they are in our real processes (except for anonymization/masking), and there were no adjustments like downsampling in test set.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2845807,
          "author_name": "andreynesterov",
          "author_url": "",
          "post_date": "05/30/2024 17:34:24",
          "content": "<p>Let me add my 5 cents. With MLP model I also had problems with overfitting, so I used only top 20 features (by gain importance of LGBM). But GBDT models worked pretty well with cred_b_a_* tables, especially one version of CatBoost with a private score of 0.520. Of course some features were filtered, but nothing special (maybe with the exception of additional aggregators that are applied to some num cols). In ensemble, these models (using cred_b_a_* tables) worked well.</p>\n<p>And also in early stages I filtered out some rows with null cols greater then some threshold, but then I think I just found out combination of aggregators that worked well with all tables and rows of cred_b_a_* data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2845822,
          "author_name": "thunderthunder",
          "author_url": "",
          "post_date": "05/30/2024 17:48:55",
          "content": "<p>I use credit_bureau_a_1_3 for training instead of credit_bureau_a_1_*.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2845992,
              "author_name": "eduardotoloza",
              "author_url": "",
              "post_date": "05/30/2024 19:45:05",
              "content": "<p>Can you elaborate?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2844882,
      "author_name": "kunduruanil",
      "author_url": "",
      "post_date": "05/30/2024 09:13:23",
      "content": "<p>Congratulations , thanks for sharing your wonderfull journay!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2844897,
      "author_name": "diegoiglesias",
      "author_url": "",
      "post_date": "05/30/2024 09:28:31",
      "content": "<p>Congrats Yuuniee!! From your shared code I assume you used several models for each model category (LGBM, DNN, Catboost). Were those models the result of each folod of the CV process or were models result of training on the whole train dataset but with different params?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2844941,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "05/30/2024 09:55:52",
          "content": "<p>Thanks a lot. <br>\nThe CV score above is the score for all out of fold(OOF) of 5folds at StratifiedGroupKFold (no shuffle).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2845269,
      "author_name": "alexryzhkov",
      "author_url": "",
      "post_date": "05/30/2024 13:42:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>,</p>\n<p>Congrats with the great performance and amazing solution! Also I would like to thank you for using LightAutoML models in your solution 😎</p>\n<blockquote>\n  <p>DNN is a Denselight model and utilizes the Lightautoml library. There were some bugs, so I had to fix some of them myself, but overall, I think it is a good library. </p>\n</blockquote>\n<p>Could you please share the problems you had with our library using <a href=\"https://github.com/sb-ai-lab/LightAutoML/issues\" target=\"_blank\">GitHub issue mechanism</a>? It will help us to make LightAutoML better 😊</p>\n<p>P.S. If you already have fixes for the problems and you would like to share them, please use the <a href=\"https://github.com/sb-ai-lab/LightAutoML/pulls\" target=\"_blank\">Pull Request mechanism</a></p>\n<p>Alex<br>\n(Head of LightAutoML team)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2846194,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "05/31/2024 02:27:56",
          "content": "<p>Thank you so much.</p>\n<p>LightAutoML is a really convenient and useful library.<br>\nThe problem I had was not fatal, but I will share it there soon. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2845308,
      "author_name": "octaviograu",
      "author_url": "",
      "post_date": "05/30/2024 13:57:36",
      "content": "<p>Amazing <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> ! Congrats!</p>\n<p>How did you manage OOM errors?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2846201,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "05/31/2024 02:40:09",
          "content": "<p>Thank you for your congrats. </p>\n<p>Except when there are more than 1000 features, I have not experienced OOM (Out of memory). But if you're having trouble, the following methods may help:</p>\n<blockquote>\n  <ol>\n  <li>Capture as many useless features as possible and remove them all.</li>\n  <li>Use reduce_memory_useage in the step of reading each file.</li>\n  <li>In addition to RAM, GPU RAM is used together. (A representative example is the cuDF library created by NVIDIA's RAPIDS team.)</li>\n  <li>When data processing is complete, save the data locally as a file. Then, delete the DataFrame and reset the memory with commands such as gc.collect() and reset.</li>\n  <li>In the prediction step, locally stored files are read in chunks and each is predicted.</li>\n  </ol>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2845622,
      "author_name": "barbosajaf",
      "author_url": "",
      "post_date": "05/30/2024 15:52:55",
      "content": "<p>Wow! Quite a nice job! Congrats!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2846139,
      "author_name": "hoshinoyumemi",
      "author_url": "",
      "post_date": "05/31/2024 00:03:40",
      "content": "<p>Congratulations on your award! <br>\nThank you for the solution!  I will keep this in mind for future reference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2846187,
      "author_name": "muhammedtausif",
      "author_url": "",
      "post_date": "05/31/2024 02:21:03",
      "content": "<p>Congratulations, thanks for sharing. <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2846220,
      "author_name": "clovisdalmolinvieira",
      "author_url": "",
      "post_date": "05/31/2024 03:03:17",
      "content": "<p>Congratulations. Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2846286,
      "author_name": "masopf",
      "author_url": "",
      "post_date": "05/31/2024 04:01:16",
      "content": "<p>Thank you for clear cut explanation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2846477,
      "author_name": "antonygithinji",
      "author_url": "",
      "post_date": "05/31/2024 06:00:01",
      "content": "<p>great hacking man</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2847445,
      "author_name": "ineerav",
      "author_url": "",
      "post_date": "05/31/2024 14:39:11",
      "content": "<p>Interesting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2847765,
      "author_name": "matousfamera",
      "author_url": "",
      "post_date": "05/31/2024 17:14:49",
      "content": "<p>Congratulations! My team had similar idea - Ensemble model of LightGBM, CatBoost and Tensorflow. We were really close to developing the ensemble model. However after the change of rules and allowing the metric hack, our focus shifted on the metric hack.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2848059,
      "author_name": "tomaraman",
      "author_url": "",
      "post_date": "05/31/2024 19:05:04",
      "content": "<p>Great approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2848245,
      "author_name": "aa000098",
      "author_url": "",
      "post_date": "05/31/2024 23:36:47",
      "content": "<p>helpful for me!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2849145,
      "author_name": "arjunm97",
      "author_url": "",
      "post_date": "06/01/2024 12:33:17",
      "content": "<p>Congratulations on your award!<br>\nThank you for the solution! Very interesting approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2849248,
      "author_name": "wuuthraad",
      "author_url": "",
      "post_date": "06/01/2024 13:28:13",
      "content": "<p>Great work man!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2850116,
      "author_name": "manavtrivedi",
      "author_url": "",
      "post_date": "06/02/2024 00:10:08",
      "content": "<p>Smart way of using the hacks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2851303,
      "author_name": "larrylin666",
      "author_url": "",
      "post_date": "06/02/2024 17:02:24",
      "content": "<p>Thanks! Very useful article for fresh man</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2851969,
      "author_name": "mubashirsidiki",
      "author_url": "",
      "post_date": "06/03/2024 04:24:18",
      "content": "<p>a very good guide ,thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2852381,
      "author_name": "kerenshang",
      "author_url": "",
      "post_date": "06/03/2024 08:58:11",
      "content": "<p>Thanks! Very useful article for fresh man.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2852723,
      "author_name": "jidaengwarrior",
      "author_url": "",
      "post_date": "06/03/2024 12:29:02",
      "content": "<p>선생님처럼 할려면 얼마나 해야할까요..</p>",
      "votes": null,
      "replies": [
        {
          "id": 2852941,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "06/03/2024 14:24:27",
          "content": "<p>Participate actively in the competition.<br>\nI think it's best to have fun!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2852944,
          "author_name": "matousfamera",
          "author_url": "",
          "post_date": "06/03/2024 14:26:11",
          "content": "<p>Keep it in English, please.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2853669,
      "author_name": "milkyway0809",
      "author_url": "",
      "post_date": "06/03/2024 21:18:45",
      "content": "<p>thank you for the summary. Very inspiring </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2856429,
      "author_name": "rram12",
      "author_url": "",
      "post_date": "06/05/2024 09:14:15",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> for the winning. May I ask you what approach did you choose to find the best ensemble weights? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2856467,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "06/05/2024 10:01:55",
          "content": "<p>Thank you for your congrats.<br>\nMy approach is linear regression and weighted testing in OOF.<br>\nInterestingly, the optimal weights of OOF score was close to the optimal weights of LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2858305,
      "author_name": "jickymen",
      "author_url": "",
      "post_date": "06/06/2024 12:04:12",
      "content": "<p>Thank you for sharing such a detailed and insightful breakdown of your solution, yuuniee! Your approach to combining LGBM, DNN, and Catboost models with a weighted ensemble clearly highlights the importance of diversity in modeling techniques.</p>\n<p>I particularly appreciate the transparency in discussing both phases of the competition—Machine Learning and Metric Hacking. Your use of StratifiedGroupKFold and various feature engineering techniques, along with your observations about the correlation between CV improvements and LB performance, are incredibly valuable for the community.</p>\n<p>The post-processing strategy and the thoughtful discussion on the impact of WEEK_NUM adjustments and the choices around DEVIDE and REDUCE values offer practical insights into handling similar challenges.</p>\n<p>It's also commendable how you acknowledged the contributions and insights from other participants like <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> and <a href=\"https://www.kaggle.com/greysky\" target=\"_blank\">@greysky</a> for their foundational notebooks, and <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> and <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> for their critical discussions around Metric Hacking. The collaboration and knowledge sharing within this community are what make these competitions so enriching.</p>\n<p>Despite the challenges posed by Metric Hacking, your balanced approach to selecting submission strategies demonstrates a clear understanding of risk management in competition settings.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2865334,
      "author_name": "troostiticgh",
      "author_url": "",
      "post_date": "06/10/2024 17:01:02",
      "content": "<p>Congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2875243,
      "author_name": "wajyjpku",
      "author_url": "",
      "post_date": "06/17/2024 01:10:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> for the winning. How do we filter out effective features and toxic features?Thank you for your response in advance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2875270,
          "author_name": "yuuniekiri",
          "author_url": "",
          "post_date": "06/17/2024 02:01:37",
          "content": "<p>As you can see in the article above, features that have a significant advantage in the CV score are kept, and features that have no advantage are removed.<br>\nAdditionally, even if there is no significant gain in score, logically meaningful and stable features can be kept. This can also apply in the opposite case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2880567,
      "author_name": "xianzwaikato",
      "author_url": "",
      "post_date": "06/20/2024 08:49:52",
      "content": "<p>Nice share, Congratulations for your first solo gold.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2889766,
      "author_name": "nonchokablexd",
      "author_url": "",
      "post_date": "06/25/2024 16:53:07",
      "content": "<p>Congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2904468,
      "author_name": "jeannkouagou",
      "author_url": "",
      "post_date": "07/04/2024 11:27:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>, I cannot find the final submission notebook for the first place. Can you please share it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3012991,
      "author_name": "outlierrr",
      "author_url": "",
      "post_date": "10/09/2024 14:54:37",
      "content": "<p>Thanks for your share!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3135086,
      "author_name": "trannhattrinhk16hcm",
      "author_url": "",
      "post_date": "02/27/2025 03:49:04",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a>, Congratulations on your win. I love the way you approach Feature Engineering, especially Aggregation. I was wondering—if all feature names were hidden, how would you do in Aggregation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3312920,
      "author_name": "afeedahamed",
      "author_url": "",
      "post_date": "11/08/2025 08:47:50",
      "content": "<p>Nice, keep doing well</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2842729": "##Foreword\n\nHello, this is yuuniee.\nThank you to the competition host Home Credit and staff, Kaggle staff and everyone involved.\n\nThe image below is a summary of my solution.\n\n![image_](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6960421%2F92721e389d44b29d04f931d244fed316%2F.png?generation=1716967303955938&alt=media)\n\nThe maximum Private LB Score for my model pool without Metric Hacking is close to about 0.53.\n\nAs everyone knows, this competition takes place in two phases.\nPhase 1 was ML (Machine Learning), and Phase 2 was MH (Metric Hack).\nI will briefly introduce each part.\n(If you are only curious about ML, please read only Phase 1.)\n\n\n## I. Phase 1 - ML (Machine Learning)\n\nThank you for writing a great notebook for EDA and solution construction at the beginning of the competition.\n@sergiosaharovskiy - https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission\n@greysky - https://www.kaggle.com/code/greysky/home-credit-baseline\nAdditionally, I would like to thank the many participants who provided various insights.\n\n1. The CV strategy is StratifiedGroupKFold, and has been tested and trained in various forms: No Shuffle and Shuffle. In my case, differences in CV of around 0.001 to 0.005 had a low correlation with LB, while differences above 0.01 had a high correlation with LB. In addition, CV improvement through model parameter adjustment had a relatively low correlation with LB, and CV improvement through FE had a relatively high correlation with LB.\n\n2. There's nothing special about Feature Engineering. Max, Min, Avg, Var, First, Last and Max-Min Difference of each item were utilized. (I will release the code after cleaning it up.)\n\n3. LGBM was a bit lacking compared to Catboost, but was good for the ensemble. The reason why the number of features is smaller is because some items (mainly categorical types) were excluded because they caused performance degradation and overfitting in LGBM, and other slightly different types of features were added and deleted.\n\n4. DNN is a Denselight model and utilizes the LightAutoML library. There were some bugs, so I had to fix some of them myself, but overall, I think it is a good library. At first, I planned to lightly test with Denselight and then build the final model with a larger model (FT-Transformer, etc.). Surprisingly, I haven't created or found a model that beats Denselight performance. Perhaps I made a mistake, or the nature of the data is sensitive to overfitting, so I guess a simpler model worked better.\n\n5. Catboost was the model that performed best in this competition and I guess it performed well on a large number of categorical features (about 117).\n\nDuring this period, my public LB was close to top 5 and the competition seemed to be going very smoothly.\nBut...\n\n\n>### + What didn't work\n>1. Transformer type models such as Tabnet, TabTransformer, FT-Transformer, etc.\n>2. Income, expenditure, and tax statistics by period (month, week, etc.).\n>3. Differences between individual income, expenses, and taxes.\n>4. K-means clustering\n\n\n## II. Phase 2 - MH(Metric Hack)\n\nPointed out problems with MH at the beginning of the competition,\n@at7459 - https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449\nAnnounced the revival of MH in the latter half of the competition,\n@johnpateha - https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167\n\nAlso, thanks to the many participants who suggested various other comments and solutions. Thanks to you, my knowledge has expanded.\nIn particular, the two people above are true scientists and journalists who could have hidden this discovery and profited from it, but pursued the public interest by exposing it to everyone. (Why not consider Kaggle giving a special award for this sharing?)\nAs a result, the problem was not solved, but we learned something(?) from it and at least I was able to take minimal measures.\n\nAfter @johnpateha's \"Metric hack again, sorry\" post, I tried Data Reversing and found that the difference between date_decision and min_refreshdate_3813885D had a high correlation (about 0.9 or more) with WEEK_NUM.\n(More information can be found at https://www.kaggle.com/code/eivolkova/how-to-restore-the-dates?scriptVersionId=180157891, published by @eivolkova immediately after the competition.)\nAnd the way suggested by @at7459:\n\n```python\nDEVIDE = 1/2\nREDUCE = 0.02\ncondition = df['WEEK_NUM'] < (df['WEEK_NUM'].max()-df['WEEK_NUM'].min())*DEVIDE+df['WEEK_NUM'].min()\ndf.loc[condition, 'score'] = (df.loc[condition, 'score'] - REDUCE).clip(0)\n```\n\nBased on this, the results of examining the score details submitted by changing the REDUCE and DEVIDE values ​​were as follows.\n\n1. The public/private division of test data spans the entire period and it is difficult to identify a specific distribution, and the optimal value in private will probably be distributed around 1/4 to 3/4.\n2. The optimal public LB value of REDUCE is 0.03 (at least in my model), and the optimal value in private is probably distributed around 0.02 to 0.04.\n\nThe difference in pure model performance between top competitors was 0.00X,\nSince the difference due to the above adjustment value was 0.0X,\nAt this stage it became gambling for me.\n\nSome will bet low, others will bet high.\nHere, I trusted my model and chose neutral, and chose DEVIDE=1/2, REDUCE=0.03, which is the median of the expected distribution.\nI expected someone who bet high or low to take the prize, and I expected about a 50% chance of finishing in the top 10.\n\n\n## III. Conclusion\n\n+ The first selected submission is the No Hack model with CV Best Score created in Phase 1, and the second selected submission is the application of Metric Hack to it. (I thought there was a high probability that Metric Hack would work, but I also submitted No Hack just in case)\n\n+ Rather than simply sharing solutions, I thought it would be better to share my experience in detail. To improve the bad parts that happened in this competition.\n\n+ My ranking was due to luck, and although it is the Best Private Score submitted, it probably won't be Best Solution.\n\n+ Although the competition ended on a gloomy note due to the metric hacking issue, the host and staff worked hard to make it a meaningful and successful competition. \n\n+ Based on this, we expect a more developed system in the future.\n\n\nthank you\n\n\n>Edit : I posted the FE code here.\n>644(for lgbm) - https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-risk-lightgbm\n>661(for cat & dnn) - https://www.kaggle.com/code/yuuniekiri/fork-of-home-credit-catboost-inference",
    "2842917": "Congratulations on the placement @yuuniekiri , how did you choose the features for each model?\nIf possible, share your notebook with us.",
    "2842974": "Thank you, @thiagomantuani. A notebook link has been added.",
    "2843263": "**Congratulations on your second #1 rank competition @yuuniekiri!** Even though I didn't participate in the competition, I was following along since the beginning after my initial post some months ago. After reading your straightforward write-up and both notebooks, I totally agree with your choice of models, such as CatBoost, LGBM, and DNN (eventually NaiveBayes would've helped a bit). My only question that I have left is, did you try to use `rankdata()` after making the predictions? That gave me a huge boost in my first competition with almost exactly the same ensemble models. I'm just curious, if you don't mind, whenever you have time, could you please try a \"Late submission\" with the following updated part of your \"Fork of Home Credit Risk (LightGBM)\" at the very end:\n\n**Edit after @yuuniekiri comment:**\n\n```python\nfrom scipy.stats import rankdata  # <- newly added, import rankdata\n\ndel reader\ngc.collect()\n\nranked_preds = rankdata(y_pred) / len(y_pred)  # <- newly added\n\ndf_subm = pd.read_csv(f\"{ROOT}/sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm[\"score\"] = ranked_preds  # <- newly adjusted\ndf_subm.to_csv(\"submission.csv\")\nprint(df_subm)\n```\n\nThanks a lot for your try-out and response. Enjoy your preliminary win!\nBest regards,\nEtienne",
    "2843306": "Thank you for your congratulations. \nHowever, the submissions resulting from the code above are different from the competition format.\n(Only two columns are allowed: case_id and score.)\nIt will probably cause an error.",
    "2843340": "You're welcome. That's right, I had a logical mistake. We don't need an additional column, I updated my code above (Edit). I looked into my old competition and that's the way how I made it there. Now the predictions should be averaged ranked. Give it a shot, if it doesn't work it's also fine - just for quick testing purpose :) Thanks!",
    "2844027": "Congratulations! Thanks for the detailed explanation. Finally, the winning solution used the metric hack as was expected for some Kaggle experts.",
    "2844034": "This is awesome! I'll be reading this in detail soon! :) Great work @yuuniekiri!",
    "2844081": "Hi! May i know what's the logic of it? Isn't all the scores equal to the mean df_subm[\"score\"] = final_preds_ranked? Thank you!",
    "2844111": "Congratulations! This notebook is very impressive for me.",
    "2844391": "Congratulations on winning this competition. Thanks for detailed explanation of your solution and sharing valuable comments on topic of Metric Hacking. Can you comment on execution time of your notebook and share best practices for the competition participants.",
    "2844427": "Congratulations! I learn a lot~",
    "2844545": "Thank you very much. \n\n1. The execution time of my notes is approximately 2 to 3 hours. (I have never measured it accurately.)\n\n2. Best practice - I don't know if I'm qualified to comment on this. It's difficult to pinpoint something, but since financial data has a lot of noise and is vulnerable to overfitting, I think we need to pay special attention to generalization. For example, you should be very careful about randomly inserting unexplained data (but this may not always be accurate). In addition, we explore various methods to prevent overfitting.",
    "2844603": "Your idea is so amazing. Thank you for sharing awesome idea.",
    "2844613": "mengjingyang Thanks for your question. Before using `rankdata()`, the score for each test case is just the raw probability. These probabilities are averaged across all models, resulting in a single score for each test case. However, without `rankdata()`, these scores can still be skewed or affected by outliers.\n\nBy applying `rankdata()`, the predictions are converted into ranks and then normalized. This process makes the scores more robust by reducing the influence of outliers and ensures a uniform distribution (relative) instead of an absolute distribution (which can be skewed). Hope that clarifies to understand why I thought of that.",
    "2844674": "```In addition, CV improvement through model parameter adjustment had a relatively low correlation with LB, and CV improvement through FE had a relatively high correlation with LB.```\n\nThis is consistently my experience as well on various problems (not only this competition).",
    "2844679": "Thanks for sharing - this is great. One question out of curiosity: have you managed to use credit bureau A data successfully?\n(for context - I haven't - my post with details here https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478360)",
    "2844882": "Congratulations , thanks for sharing your wonderfull journay!!",
    "2844897": "Congrats Yuuniee!! From your shared code I assume you used several models for each model category (LGBM, DNN, Catboost). Were those models the result of each folod of the CV process or were models result of training on the whole train dataset but with different params?",
    "2844941": "Thanks a lot. \nThe CV score above is the score for all out of fold(OOF) of 5folds at StratifiedGroupKFold (no shuffle).",
    "2844961": "This is quite a mystery to me too. \nAs you can see from the code and comments for LGBM posted above, I had quite a bit of trouble with it.\nAs a result, many features were deleted from bureau A and only some were retained. \n(And I experienced something similar with person data.)\nBecause test data cannot be seen, it is difficult to determine the cause.",
    "2845079": "From Credit Bureau data I just used the feature related to the financial institution that approved the loan. Other features degraded the stability metric. I also thought it was better not to use debit card data. On the other hand, the best way I  found to use tax data, was the sum of all deduction tax of each client, it makes sense because it represents the tax deductions in the last year. That feature boosted my model's performance.",
    "2845269": "Hi @yuuniekiri,\n\nCongrats with the great performance and amazing solution! Also I would like to thank you for using LightAutoML models in your solution 😎\n\n>DNN is a Denselight model and utilizes the Lightautoml library. There were some bugs, so I had to fix some of them myself, but overall, I think it is a good library. \n\nCould you please share the problems you had with our library using [GitHub issue mechanism](https://github.com/sb-ai-lab/LightAutoML/issues)? It will help us to make LightAutoML better 😊\n\nP.S. If you already have fixes for the problems and you would like to share them, please use the [Pull Request mechanism](https://github.com/sb-ai-lab/LightAutoML/pulls)\n\nAlex\n(Head of LightAutoML team)",
    "2845308": "Amazing @yuuniekiri ! Congrats!\n\nHow did you manage OOM errors?",
    "2845622": "Wow! Quite a nice job! Congrats!!!",
    "2845692": "My interpretation is that hosts downsampled this source of data heavily for the test set, and hence it led to overfitting.",
    "2845807": "Let me add my 5 cents. With MLP model I also had problems with overfitting, so I used only top 20 features (by gain importance of LGBM). But GBDT models worked pretty well with cred_b_a_* tables, especially one version of CatBoost with a private score of 0.520. Of course some features were filtered, but nothing special (maybe with the exception of additional aggregators that are applied to some num cols). In ensemble, these models (using cred_b_a_* tables) worked well.\n\nAnd also in early stages I filtered out some rows with null cols greater then some threshold, but then I think I just found out combination of aggregators that worked well with all tables and rows of cred_b_a_* data.",
    "2845822": "I use credit_bureau_a_1_3 for training instead of credit_bureau_a_1_*.",
    "2845992": "Can you elaborate?",
    "2846139": "Congratulations on your award! \nThank you for the solution!  I will keep this in mind for future reference.",
    "2846187": "Congratulations, thanks for sharing. @yuuniekiri",
    "2846194": "Thank you so much.\n\nLightAutoML is a really convenient and useful library.\nThe problem I had was not fatal, but I will share it there soon. :)",
    "2846201": "Thank you for your congrats. \n\nExcept when there are more than 1000 features, I have not experienced OOM (Out of memory). But if you're having trouble, the following methods may help:\n\n>1. Capture as many useless features as possible and remove them all.\n>2. Use reduce_memory_useage in the step of reading each file.\n>3. In addition to RAM, GPU RAM is used together. (A representative example is the cuDF library created by NVIDIA's RAPIDS team.)\n>4. When data processing is complete, save the data locally as a file. Then, delete the DataFrame and reset the memory with commands such as gc.collect() and reset.\n>5. In the prediction step, locally stored files are read in chunks and each is predicted.",
    "2846220": "Congratulations. Thank you for sharing.",
    "2846286": "Thank you for clear cut explanation",
    "2846477": "great hacking man",
    "2846618": "Hi, we will have to look into it, in production models credit bureau contributes quite significantly...\nin dataset should be data as they are in our real processes (except for anonymization/masking), and there were no adjustments like downsampling in test set.",
    "2846640": "Thank you! This logic makes sense.  So ranked_preds is a 2d array(N * model_number) here?",
    "2846673": "You're welcome! No, `ranked_preds` is not a 2D array. It's a 1D array that represents the normalized ranks of the combined predictions from **all** models.",
    "2847445": "Interesting",
    "2847765": "Congratulations! My team had similar idea - Ensemble model of LightGBM, CatBoost and Tensorflow. We were really close to developing the ensemble model. However after the change of rules and allowing the metric hack, our focus shifted on the metric hack.",
    "2847982": "I see. So final_preds_ranked is a number. I am still confused as to why you set the entire score column to the same number. Thank you again!",
    "2848059": "Great approach",
    "2848245": "helpful for me!",
    "2849145": "Congratulations on your award!\nThank you for the solution! Very interesting approach",
    "2849248": "Great work man!",
    "2850116": "Smart way of using the hacks!",
    "2851303": "Thanks! Very useful article for fresh man",
    "2851969": "a very good guide ,thanks",
    "2852381": "Thanks! Very useful article for fresh man.",
    "2852723": "선생님처럼 할려면 얼마나 해야할까요..",
    "2852941": "Participate actively in the competition.\nI think it's best to have fun!",
    "2852944": "Keep it in English, please.",
    "2853669": "thank you for the summary. Very inspiring",
    "2856429": "Congrats @yuuniekiri for the winning. May I ask you what approach did you choose to find the best ensemble weights?",
    "2856467": "Thank you for your congrats.\nMy approach is linear regression and weighted testing in OOF.\nInterestingly, the optimal weights of OOF score was close to the optimal weights of LB.",
    "2858305": "Thank you for sharing such a detailed and insightful breakdown of your solution, yuuniee! Your approach to combining LGBM, DNN, and Catboost models with a weighted ensemble clearly highlights the importance of diversity in modeling techniques.\n\nI particularly appreciate the transparency in discussing both phases of the competition—Machine Learning and Metric Hacking. Your use of StratifiedGroupKFold and various feature engineering techniques, along with your observations about the correlation between CV improvements and LB performance, are incredibly valuable for the community.\n\nThe post-processing strategy and the thoughtful discussion on the impact of WEEK_NUM adjustments and the choices around DEVIDE and REDUCE values offer practical insights into handling similar challenges.\n\nIt's also commendable how you acknowledged the contributions and insights from other participants like @sergiosaharovskiy and @greysky for their foundational notebooks, and @at7459 and @johnpateha for their critical discussions around Metric Hacking. The collaboration and knowledge sharing within this community are what make these competitions so enriching.\n\nDespite the challenges posed by Metric Hacking, your balanced approach to selecting submission strategies demonstrates a clear understanding of risk management in competition settings.",
    "2865334": "Congratulations",
    "2875243": "Congrats @yuuniekiri for the winning. How do we filter out effective features and toxic features?Thank you for your response in advance.",
    "2875270": "As you can see in the article above, features that have a significant advantage in the CV score are kept, and features that have no advantage are removed.\nAdditionally, even if there is no significant gain in score, logically meaningful and stable features can be kept. This can also apply in the opposite case.",
    "2880567": "Nice share, Congratulations for your first solo gold.",
    "2889766": "Congratulations",
    "2904468": "Hi @yuuniekiri, I cannot find the final submission notebook for the first place. Can you please share it?",
    "3012991": "Thanks for your share!",
    "3135086": "Hi @yuuniekiri, Congratulations on your win. I love the way you approach Feature Engineering, especially Aggregation. I was wondering—if all feature names were hidden, how would you do in Aggregation?",
    "3312920": "Nice, keep doing well"
  },
  "source": "meta"
}