{
  "id": 508588,
  "title": "10th Place Solution",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508588",
  "author_name": "",
  "post_date": "2024-05-30T07:44:49.959135800Z",
  "votes": 29,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>Thank the hosts for organizing this competition.<br>\nOn behalf of my team, I'd like to share my team's approaches.</p>\n<p><strong>CV:</strong><br>\nWe tried many CV strategies, including StratifiedGroupKfold, MultilabelStratifiedKFold, RollOut, and HoldOut, but we couldn't find a good CV that aligns with LB. Therefore, we decided to use StratifiedGroupKfold (group=WEEK_NUM) which we trust the best.</p>\n<p><strong>Feature engineering:</strong></p>\n<ul>\n<li>Hand-crafted features: we spent almost three-quarters of our time on feature engineering. All new handcrafted features come from tables with depths 0 and 1. We generated and tested only one new feature at a time to make sure this feature aligned on both CV and LB. Then, we combined all worked features, which improved by 1e-2 on both CV and LB.</li>\n<li>Future information: while we did error analysis, we found that the train data contained some future information on some features (e.g., feature lastapplicationdate_877D). At that time, we remove all leaked rows before training because we expect that the test data will not contain any future information. However, at the end of the competition, the test data seems not to exclude future information.</li>\n<li>Category encode: additionally, we conducted some experiments about the category encoders. We found that using CountEncoder on the whole train data is better than using any encoders on each fold. It's about 3e-3 boosts up on CV.</li>\n<li>High correlation features: <a href=\"https://www.kaggle.com/code/meloncc/notebook\" target=\"_blank\">We used this code</a></li>\n<li>Collapse: The final key we used is to collapse categorical features whose distinct number of values is more than 200. Instead of removing these features, we merge some values with low frequencies into a group. We got a little jump, about 0.003.</li>\n<li>Process null: We filtered out columns with the percent of null values higher than 95%, built a model on those features, chose the top 10 important features, and put it back into the main model. LB increases about 2e-3.</li>\n</ul>\n<p><strong>Modeling:</strong></p>\n<ul>\n<li>Single model and ensemble: during the competition, ensemble models significantly improve the LB. Unfortunately, Catboost consumed a lot of time to do pseudo label training so we left it out. Our final submission is just a single XGB model.</li>\n<li>NN models: we also applied the tab transformer and distillation models, but their performance dropped unexpectedly on the LB even though they boosted the CV.</li>\n<li>Pseudo-label: about the pseudo-labeling technique, it worked! However, we had to filter users that we’re confident about, so we tuned the threshold for label 1 based on precision (i.e., we chose precision to be 90% because it's the best in CV) and that for label 0 based on the absolute number of users. We climbed up to 6e-3 on CV with this strategy.</li>\n</ul>\n<p><strong>Post-process (metric hack):</strong></p>\n<ul>\n<li>We just paid attention to hacking in 2 last weeks because its impact was much greater than that of traditional modeling. There are 2 available hacking methods: by reversing WEEK_NUM and adversarial models.</li>\n<li>Reversing WEEK_NUM: we found that the WEEK_NUM can be restored by the feature 'min_refreshdate_3813885D' due to the correlation up to 0.99. We built a simple linear regression for this feature to restore WEEK_NUM in the test data. The reversed WEEK_NUM can be divided into 3 groups:<br>\n(1) null<br>\n(2) From 92 -&gt; 142<br>\n(3) From 0 -&gt; 53~55<br>\nHacking on group (2) significantly boosts the LB score. The group (3) is kinda weird but the host confirmed that test data is the future of training data. We decided to only hack on group (2) for safety.</li>\n<li>Adversarial model: another way to hack is to use the adversarial model to detect the continuous WEEK_NUM with the train data in the test data. This method seems to work best on public LB. We decided not to use this method at the last minute because the adversarial model only detects 3 or 4 weeks near the train data (e.i., we tested it on the RollOut CV). It's not even longer than 30% of the test data, then if the private data is shifted it won't affect the final score.</li>\n</ul>\n<p><strong>Probing:</strong><br>\nI think probing the test data is the key to help us get a gold medal. We probed a lot to understand the test dataset. The table below shows some of our probing results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18133706%2F161b73e3cb01073cfc8bc1f0998648d7%2Fhomecredit_probe_results.png?generation=1717054678434516&amp;alt=media\" alt=\"Probing results\"></p>\n<p><strong>Final submission:</strong><br>\nOn the last day of the competition, we decided to select 2 hacking submissions. The problem of this competition is the metric itself, not about the data. The metric hack will work regardless of the test data. If you can know/guess the starting weeks of the test data, you can hack it. We also concluded that the adversarial hacking method is not robust because it only edited some first few weeks, if the host changed the starting week of the private data (and they did it) the adversarial hack will not work. At the end, we choose to hack for a longer period (from week 92 to 130). We also applied a linear reduction (e.g., on week 92 we subtracted the score by 0.04, 93 by 0.0375, etc) but we don’t think it improves much.</p>\n<p>This is the first time I put more than 100% effort into a competition which was fun and challenging. Besides that, I’ve learned a lot from my teammates, it’s such a wonderful opportunity. Thank Mr. Duc <a href=\"https://www.kaggle.com/mathormad\" target=\"_blank\">@mathormad</a> and Mr. Tu <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> for caring and teaching me a lot. I’m so grateful for that.</p>\n<p>HARD WORK PAYS OFF</p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": "2844702",
      "postDate": "05/30/2024 07:44:49",
      "content": "<p>Hi all,</p>\n<p>Thank the hosts for organizing this competition.<br>\nOn behalf of my team, I'd like to share my team's approaches.</p>\n<p><strong>CV:</strong><br>\nWe tried many CV strategies, including StratifiedGroupKfold, MultilabelStratifiedKFold, RollOut, and HoldOut, but we couldn't find a good CV that aligns with LB. Therefore, we decided to use StratifiedGroupKfold (group=WEEK_NUM) which we trust the best.</p>\n<p><strong>Feature engineering:</strong></p>\n<ul>\n<li>Hand-crafted features: we spent almost three-quarters of our time on feature engineering. All new handcrafted features come from tables with depths 0 and 1. We generated and tested only one new feature at a time to make sure this feature aligned on both CV and LB. Then, we combined all worked features, which improved by 1e-2 on both CV and LB.</li>\n<li>Future information: while we did error analysis, we found that the train data contained some future information on some features (e.g., feature lastapplicationdate_877D). At that time, we remove all leaked rows before training because we expect that the test data will not contain any future information. However, at the end of the competition, the test data seems not to exclude future information.</li>\n<li>Category encode: additionally, we conducted some experiments about the category encoders. We found that using CountEncoder on the whole train data is better than using any encoders on each fold. It's about 3e-3 boosts up on CV.</li>\n<li>High correlation features: <a href=\"https://www.kaggle.com/code/meloncc/notebook\" target=\"_blank\">We used this code</a></li>\n<li>Collapse: The final key we used is to collapse categorical features whose distinct number of values is more than 200. Instead of removing these features, we merge some values with low frequencies into a group. We got a little jump, about 0.003.</li>\n<li>Process null: We filtered out columns with the percent of null values higher than 95%, built a model on those features, chose the top 10 important features, and put it back into the main model. LB increases about 2e-3.</li>\n</ul>\n<p><strong>Modeling:</strong></p>\n<ul>\n<li>Single model and ensemble: during the competition, ensemble models significantly improve the LB. Unfortunately, Catboost consumed a lot of time to do pseudo label training so we left it out. Our final submission is just a single XGB model.</li>\n<li>NN models: we also applied the tab transformer and distillation models, but their performance dropped unexpectedly on the LB even though they boosted the CV.</li>\n<li>Pseudo-label: about the pseudo-labeling technique, it worked! However, we had to filter users that we’re confident about, so we tuned the threshold for label 1 based on precision (i.e., we chose precision to be 90% because it's the best in CV) and that for label 0 based on the absolute number of users. We climbed up to 6e-3 on CV with this strategy.</li>\n</ul>\n<p><strong>Post-process (metric hack):</strong></p>\n<ul>\n<li>We just paid attention to hacking in 2 last weeks because its impact was much greater than that of traditional modeling. There are 2 available hacking methods: by reversing WEEK_NUM and adversarial models.</li>\n<li>Reversing WEEK_NUM: we found that the WEEK_NUM can be restored by the feature 'min_refreshdate_3813885D' due to the correlation up to 0.99. We built a simple linear regression for this feature to restore WEEK_NUM in the test data. The reversed WEEK_NUM can be divided into 3 groups:<br>\n(1) null<br>\n(2) From 92 -&gt; 142<br>\n(3) From 0 -&gt; 53~55<br>\nHacking on group (2) significantly boosts the LB score. The group (3) is kinda weird but the host confirmed that test data is the future of training data. We decided to only hack on group (2) for safety.</li>\n<li>Adversarial model: another way to hack is to use the adversarial model to detect the continuous WEEK_NUM with the train data in the test data. This method seems to work best on public LB. We decided not to use this method at the last minute because the adversarial model only detects 3 or 4 weeks near the train data (e.i., we tested it on the RollOut CV). It's not even longer than 30% of the test data, then if the private data is shifted it won't affect the final score.</li>\n</ul>\n<p><strong>Probing:</strong><br>\nI think probing the test data is the key to help us get a gold medal. We probed a lot to understand the test dataset. The table below shows some of our probing results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18133706%2F161b73e3cb01073cfc8bc1f0998648d7%2Fhomecredit_probe_results.png?generation=1717054678434516&amp;alt=media\" alt=\"Probing results\"></p>\n<p><strong>Final submission:</strong><br>\nOn the last day of the competition, we decided to select 2 hacking submissions. The problem of this competition is the metric itself, not about the data. The metric hack will work regardless of the test data. If you can know/guess the starting weeks of the test data, you can hack it. We also concluded that the adversarial hacking method is not robust because it only edited some first few weeks, if the host changed the starting week of the private data (and they did it) the adversarial hack will not work. At the end, we choose to hack for a longer period (from week 92 to 130). We also applied a linear reduction (e.g., on week 92 we subtracted the score by 0.04, 93 by 0.0375, etc) but we don’t think it improves much.</p>\n<p>This is the first time I put more than 100% effort into a competition which was fun and challenging. Besides that, I’ve learned a lot from my teammates, it’s such a wonderful opportunity. Thank Mr. Duc <a href=\"https://www.kaggle.com/mathormad\" target=\"_blank\">@mathormad</a> and Mr. Tu <a href=\"https://www.kaggle.com/minhtu123\" target=\"_blank\">@minhtu123</a> for caring and teaching me a lot. I’m so grateful for that.</p>\n<p>HARD WORK PAYS OFF</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Hi all,\n\nThank the hosts for organizing this competition.\nOn behalf of my team, I'd like to share my team's approaches.\n\n**CV:**\nWe tried many CV strategies, including StratifiedGroupKfold, MultilabelStratifiedKFold, RollOut, and HoldOut, but we couldn't find a good CV that aligns with LB. Therefore, we decided to use StratifiedGroupKfold (group=WEEK_NUM) which we trust the best.\n\n**Feature engineering:**\n- Hand-crafted features: we spent almost three-quarters of our time on feature engineering. All new handcrafted features come from tables with depths 0 and 1. We generated and tested only one new feature at a time to make sure this feature aligned on both CV and LB. Then, we combined all worked features, which improved by 1e-2 on both CV and LB.\n- Future information: while we did error analysis, we found that the train data contained some future information on some features (e.g., feature lastapplicationdate_877D). At that time, we remove all leaked rows before training because we expect that the test data will not contain any future information. However, at the end of the competition, the test data seems not to exclude future information.\n- Category encode: additionally, we conducted some experiments about the category encoders. We found that using CountEncoder on the whole train data is better than using any encoders on each fold. It's about 3e-3 boosts up on CV.\n- High correlation features: [We used this code](https://www.kaggle.com/code/meloncc/notebook)\n- Collapse: The final key we used is to collapse categorical features whose distinct number of values is more than 200. Instead of removing these features, we merge some values with low frequencies into a group. We got a little jump, about 0.003.\n- Process null: We filtered out columns with the percent of null values higher than 95%, built a model on those features, chose the top 10 important features, and put it back into the main model. LB increases about 2e-3.\n\n**Modeling:**\n- Single model and ensemble: during the competition, ensemble models significantly improve the LB. Unfortunately, Catboost consumed a lot of time to do pseudo label training so we left it out. Our final submission is just a single XGB model.\n- NN models: we also applied the tab transformer and distillation models, but their performance dropped unexpectedly on the LB even though they boosted the CV.\n- Pseudo-label: about the pseudo-labeling technique, it worked! However, we had to filter users that we’re confident about, so we tuned the threshold for label 1 based on precision (i.e., we chose precision to be 90% because it's the best in CV) and that for label 0 based on the absolute number of users. We climbed up to 6e-3 on CV with this strategy.\n\n**Post-process (metric hack):**\n- We just paid attention to hacking in 2 last weeks because its impact was much greater than that of traditional modeling. There are 2 available hacking methods: by reversing WEEK_NUM and adversarial models.\n- Reversing WEEK_NUM: we found that the WEEK_NUM can be restored by the feature 'min_refreshdate_3813885D' due to the correlation up to 0.99. We built a simple linear regression for this feature to restore WEEK_NUM in the test data. The reversed WEEK_NUM can be divided into 3 groups:\n(1) null\n(2) From 92 -> 142\n(3) From 0 -> 53~55\nHacking on group (2) significantly boosts the LB score. The group (3) is kinda weird but the host confirmed that test data is the future of training data. We decided to only hack on group (2) for safety.\n- Adversarial model: another way to hack is to use the adversarial model to detect the continuous WEEK_NUM with the train data in the test data. This method seems to work best on public LB. We decided not to use this method at the last minute because the adversarial model only detects 3 or 4 weeks near the train data (e.i., we tested it on the RollOut CV). It's not even longer than 30% of the test data, then if the private data is shifted it won't affect the final score.\n\n**Probing:**\nI think probing the test data is the key to help us get a gold medal. We probed a lot to understand the test dataset. The table below shows some of our probing results:\n![Probing results](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18133706%2F161b73e3cb01073cfc8bc1f0998648d7%2Fhomecredit_probe_results.png?generation=1717054678434516&alt=media)\n\n**Final submission:**\nOn the last day of the competition, we decided to select 2 hacking submissions. The problem of this competition is the metric itself, not about the data. The metric hack will work regardless of the test data. If you can know/guess the starting weeks of the test data, you can hack it. We also concluded that the adversarial hacking method is not robust because it only edited some first few weeks, if the host changed the starting week of the private data (and they did it) the adversarial hack will not work. At the end, we choose to hack for a longer period (from week 92 to 130). We also applied a linear reduction (e.g., on week 92 we subtracted the score by 0.04, 93 by 0.0375, etc) but we don’t think it improves much.\n\n\n\n\nThis is the first time I put more than 100% effort into a competition which was fun and challenging. Besides that, I’ve learned a lot from my teammates, it’s such a wonderful opportunity. Thank Mr. Duc @mathormad and Mr. Tu @minhtu123 for caring and teaching me a lot. I’m so grateful for that.\n\nHARD WORK PAYS OFF\n\nThank you.",
      "votes": null
    },
    {
      "id": "2846709",
      "postDate": "05/31/2024 08:11:45",
      "content": "<p>Congratulation!!! 🎉🎉🎉</p>",
      "rawMarkdown": "Congratulation!!! 🎉🎉🎉",
      "votes": null
    },
    {
      "id": "2853241",
      "postDate": "06/03/2024 17:20:35",
      "content": "<p>This is a such great method to learn from! I joined this competition as a learning experience, and I indeed learn so much from the competitors!</p>",
      "rawMarkdown": "This is a such great method to learn from! I joined this competition as a learning experience, and I indeed learn so much from the competitors!",
      "votes": null
    },
    {
      "id": "2854645",
      "postDate": "06/04/2024 11:39:14",
      "content": "<p>Such a good idea! Congratulations!</p>",
      "rawMarkdown": "Such a good idea! Congratulations!",
      "votes": null
    },
    {
      "id": "2858072",
      "postDate": "06/06/2024 08:53:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> ,</p>\n<p>Can you please attach my team's solution on the leaderboard?</p>\n<p>Thank you.</p>",
      "rawMarkdown": "Hi @addisonhoward ,\n\nCan you please attach my team's solution on the leaderboard?\n\nThank you.",
      "votes": null
    },
    {
      "id": "2858271",
      "postDate": "06/06/2024 11:39:10",
      "content": "<p>Hi there - you actually need to add it yourself! Here's a <a href=\"https://www.kaggle.com/discussions/product-feedback/373153\" target=\"_blank\">post to learn more</a>.</p>",
      "rawMarkdown": "Hi there - you actually need to add it yourself! Here's a [post to learn more](https://www.kaggle.com/discussions/product-feedback/373153).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2846709,
      "author_name": "linhlethuy",
      "author_url": "",
      "post_date": "05/31/2024 08:11:45",
      "content": "<p>Congratulation!!! 🎉🎉🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2853241,
      "author_name": "suehuynh",
      "author_url": "",
      "post_date": "06/03/2024 17:20:35",
      "content": "<p>This is a such great method to learn from! I joined this competition as a learning experience, and I indeed learn so much from the competitors!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2854645,
      "author_name": "xianminwang",
      "author_url": "",
      "post_date": "06/04/2024 11:39:14",
      "content": "<p>Such a good idea! Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2858072,
      "author_name": "pntan17",
      "author_url": "",
      "post_date": "06/06/2024 08:53:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> ,</p>\n<p>Can you please attach my team's solution on the leaderboard?</p>\n<p>Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2858271,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "06/06/2024 11:39:10",
          "content": "<p>Hi there - you actually need to add it yourself! Here's a <a href=\"https://www.kaggle.com/discussions/product-feedback/373153\" target=\"_blank\">post to learn more</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2844702": "Hi all,\n\nThank the hosts for organizing this competition.\nOn behalf of my team, I'd like to share my team's approaches.\n\n**CV:**\nWe tried many CV strategies, including StratifiedGroupKfold, MultilabelStratifiedKFold, RollOut, and HoldOut, but we couldn't find a good CV that aligns with LB. Therefore, we decided to use StratifiedGroupKfold (group=WEEK_NUM) which we trust the best.\n\n**Feature engineering:**\n- Hand-crafted features: we spent almost three-quarters of our time on feature engineering. All new handcrafted features come from tables with depths 0 and 1. We generated and tested only one new feature at a time to make sure this feature aligned on both CV and LB. Then, we combined all worked features, which improved by 1e-2 on both CV and LB.\n- Future information: while we did error analysis, we found that the train data contained some future information on some features (e.g., feature lastapplicationdate_877D). At that time, we remove all leaked rows before training because we expect that the test data will not contain any future information. However, at the end of the competition, the test data seems not to exclude future information.\n- Category encode: additionally, we conducted some experiments about the category encoders. We found that using CountEncoder on the whole train data is better than using any encoders on each fold. It's about 3e-3 boosts up on CV.\n- High correlation features: [We used this code](https://www.kaggle.com/code/meloncc/notebook)\n- Collapse: The final key we used is to collapse categorical features whose distinct number of values is more than 200. Instead of removing these features, we merge some values with low frequencies into a group. We got a little jump, about 0.003.\n- Process null: We filtered out columns with the percent of null values higher than 95%, built a model on those features, chose the top 10 important features, and put it back into the main model. LB increases about 2e-3.\n\n**Modeling:**\n- Single model and ensemble: during the competition, ensemble models significantly improve the LB. Unfortunately, Catboost consumed a lot of time to do pseudo label training so we left it out. Our final submission is just a single XGB model.\n- NN models: we also applied the tab transformer and distillation models, but their performance dropped unexpectedly on the LB even though they boosted the CV.\n- Pseudo-label: about the pseudo-labeling technique, it worked! However, we had to filter users that we’re confident about, so we tuned the threshold for label 1 based on precision (i.e., we chose precision to be 90% because it's the best in CV) and that for label 0 based on the absolute number of users. We climbed up to 6e-3 on CV with this strategy.\n\n**Post-process (metric hack):**\n- We just paid attention to hacking in 2 last weeks because its impact was much greater than that of traditional modeling. There are 2 available hacking methods: by reversing WEEK_NUM and adversarial models.\n- Reversing WEEK_NUM: we found that the WEEK_NUM can be restored by the feature 'min_refreshdate_3813885D' due to the correlation up to 0.99. We built a simple linear regression for this feature to restore WEEK_NUM in the test data. The reversed WEEK_NUM can be divided into 3 groups:\n(1) null\n(2) From 92 -> 142\n(3) From 0 -> 53~55\nHacking on group (2) significantly boosts the LB score. The group (3) is kinda weird but the host confirmed that test data is the future of training data. We decided to only hack on group (2) for safety.\n- Adversarial model: another way to hack is to use the adversarial model to detect the continuous WEEK_NUM with the train data in the test data. This method seems to work best on public LB. We decided not to use this method at the last minute because the adversarial model only detects 3 or 4 weeks near the train data (e.i., we tested it on the RollOut CV). It's not even longer than 30% of the test data, then if the private data is shifted it won't affect the final score.\n\n**Probing:**\nI think probing the test data is the key to help us get a gold medal. We probed a lot to understand the test dataset. The table below shows some of our probing results:\n![Probing results](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18133706%2F161b73e3cb01073cfc8bc1f0998648d7%2Fhomecredit_probe_results.png?generation=1717054678434516&alt=media)\n\n**Final submission:**\nOn the last day of the competition, we decided to select 2 hacking submissions. The problem of this competition is the metric itself, not about the data. The metric hack will work regardless of the test data. If you can know/guess the starting weeks of the test data, you can hack it. We also concluded that the adversarial hacking method is not robust because it only edited some first few weeks, if the host changed the starting week of the private data (and they did it) the adversarial hack will not work. At the end, we choose to hack for a longer period (from week 92 to 130). We also applied a linear reduction (e.g., on week 92 we subtracted the score by 0.04, 93 by 0.0375, etc) but we don’t think it improves much.\n\n\n\n\nThis is the first time I put more than 100% effort into a competition which was fun and challenging. Besides that, I’ve learned a lot from my teammates, it’s such a wonderful opportunity. Thank Mr. Duc @mathormad and Mr. Tu @minhtu123 for caring and teaching me a lot. I’m so grateful for that.\n\nHARD WORK PAYS OFF\n\nThank you.",
    "2846709": "Congratulation!!! 🎉🎉🎉",
    "2853241": "This is a such great method to learn from! I joined this competition as a learning experience, and I indeed learn so much from the competitors!",
    "2854645": "Such a good idea! Congratulations!",
    "2858072": "Hi @addisonhoward ,\n\nCan you please attach my team's solution on the leaderboard?\n\nThank you.",
    "2858271": "Hi there - you actually need to add it yourself! Here's a [post to learn more](https://www.kaggle.com/discussions/product-feedback/373153)."
  },
  "source": "meta"
}