{
  "id": 507982,
  "title": "All you need is to hack further weeks",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/507982",
  "author_name": "minhtu.mt.mt",
  "post_date": "2024-05-28T04:14:23.357000",
  "votes": 24,
  "comment_count": 12,
  "views": 0,
  "content": "<p>We have some extremely high private score submissions in some of our experience. Those submissions are all hacked-versions using reverse engineering to reproduce the WEEK_NUM. The special thing is that instead of hacking the first-half weeks, we hacked 3/4.</p>\n<p>Our 10th solution is also a hacked version, hack 3/4 weeks with linear decay (penalty first weeks more than further ones). We will post our solution in detail and the reason why we decided to choose the 3/4 version later.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023686%2Fbb522a08fffe3913afe47b034b1ce93c%2FScreenshot%202024-05-28%20at%2011.13.40.png?generation=1716869633788833&amp;alt=media\"></p>",
  "messages": [
    {
      "id": 2840272,
      "postDate": "2024-05-28T04:14:23.357Z",
      "content": "<p>We have some extremely high private score submissions in some of our experience. Those submissions are all hacked-versions using reverse engineering to reproduce the WEEK_NUM. The special thing is that instead of hacking the first-half weeks, we hacked 3/4.</p>\n<p>Our 10th solution is also a hacked version, hack 3/4 weeks with linear decay (penalty first weeks more than further ones). We will post our solution in detail and the reason why we decided to choose the 3/4 version later.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023686%2Fbb522a08fffe3913afe47b034b1ce93c%2FScreenshot%202024-05-28%20at%2011.13.40.png?generation=1716869633788833&amp;alt=media\"></p>",
      "rawMarkdown": "We have some extremely high private score submissions in some of our experience. Those submissions are all hacked-versions using reverse engineering to reproduce the WEEK_NUM. The special thing is that instead of hacking the first-half weeks, we hacked 3/4.\n\nOur 10th solution is also a hacked version, hack 3/4 weeks with linear decay (penalty first weeks more than further ones). We will post our solution in detail and the reason why we decided to choose the 3/4 version later.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023686%2Fbb522a08fffe3913afe47b034b1ce93c%2FScreenshot%202024-05-28%20at%2011.13.40.png?generation=1716869633788833&alt=media)\n",
      "votes": 23
    },
    {
      "id": 2840518,
      "postDate": "2024-05-28T06:01:48.870Z",
      "content": "<p>i was giving the hosts the benefit of the doubt at first, but their decisions just got worse and worse throughout the competition</p>\n<p>and a bizarre split like that just to make sure nobody knows what they should select for submission makes me think the hosts really are just incompetent beyond repair</p>\n<p>edit: from what i can gather from my submissions they had 2021 and maybe 1/3 of 2022 as the public lb and the rest from 2021 onward as private lb</p>\n<p>which sucks for me because i just hacked 2021, if i just sorted by the dates and hacked the first half i would've scored something around 0.63</p>\n<p>i feel like the only reasonable split to have the best feedback from the leaderboard is to take 30/70 of EVERY week_num or date_decision, the way they did it is just misleading and damaging for competitors</p>\n<p>edit2: actually looking at it further the split they did must've been even more bizarre than that, they must've taken all of 2021 (because all my submissions which hack 2021 score the same) a bit of 2022 and then also bit near the end of 2024 and/or 2023 (because originally if i hack the first half of the week_nums it is possible to increase the score significantly, which i think wouldn't be possible if they just took the first 30% of the data as the public leaderboard)</p>\n<p>all just to make things as misleading as possible… they should really make it standard to reveal how the data was split in the competition description.</p>",
      "rawMarkdown": "i was giving the hosts the benefit of the doubt at first, but their decisions just got worse and worse throughout the competition\n\nand a bizarre split like that just to make sure nobody knows what they should select for submission makes me think the hosts really are just incompetent beyond repair\n\nedit: from what i can gather from my submissions they had 2021 and maybe 1/3 of 2022 as the public lb and the rest from 2021 onward as private lb\n\nwhich sucks for me because i just hacked 2021, if i just sorted by the dates and hacked the first half i would've scored something around 0.63\n\ni feel like the only reasonable split to have the best feedback from the leaderboard is to take 30/70 of EVERY week_num or date_decision, the way they did it is just misleading and damaging for competitors\n\nedit2: actually looking at it further the split they did must've been even more bizarre than that, they must've taken all of 2021 (because all my submissions which hack 2021 score the same) a bit of 2022 and then also bit near the end of 2024 and/or 2023 (because originally if i hack the first half of the week_nums it is possible to increase the score significantly, which i think wouldn't be possible if they just took the first 30% of the data as the public leaderboard)\n\nall just to make things as misleading as possible... they should really make it standard to reveal how the data was split in the competition description.",
      "votes": 20,
      "replies": [
        {
          "id": 2840690,
          "postDate": "2024-05-28T08:02:23.717Z",
          "content": "<p>I totally agree with you. The metric hack is inevitable.<br>\nAt first, I thought that the hack wont work in private LB because the host can shift the first weeks of the private dataset so that overfit-hacking on public LB will shake up.<br>\nBut on the last day, I realized that the hack will always works. They key is how to tune a stable hacking method so that it can generalize well in both public and private LB. At this point, we had to gamble about the distribution of private dataset.<br>\nFortunately (and thank to a lot of probing results), we won the bet.</p>",
          "rawMarkdown": "I totally agree with you. The metric hack is inevitable.\nAt first, I thought that the hack wont work in private LB because the host can shift the first weeks of the private dataset so that overfit-hacking on public LB will shake up.\nBut on the last day, I realized that the hack will always works. They key is how to tune a stable hacking method so that it can generalize well in both public and private LB. At this point, we had to gamble about the distribution of private dataset.\nFortunately (and thank to a lot of probing results), we won the bet.",
          "votes": 1,
          "replies": [
            {
              "id": 2840700,
              "postDate": "2024-05-28T08:09:11.857Z",
              "content": "<p>yeah, unlucky for me i just stopped as soon as i figured out you can get the year back, and jsut assumed they would do aany type of reasonable split</p>\n<p>i should've just done more probing and at least reconstructed my score from before the dataset got transformed… it would've been obvious then that they didn't do a reasonable split and i would've made a completely different submission…</p>",
              "rawMarkdown": "yeah, unlucky for me i just stopped as soon as i figured out you can get the year back, and jsut assumed they would do aany type of reasonable split\n\ni should've just done more probing and at least reconstructed my score from before the dataset got transformed... it would've been obvious then that they didn't do a reasonable split and i would've made a completely different submission...",
              "votes": 2
            }
          ]
        },
        {
          "id": 2840858,
          "postDate": "2024-05-28T09:27:29.113Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> , I remember that before everybody started to hack the metric when the host said they won't do anything about this, you had the first place and the score about 0.620. Was it also the result of hacking or did you manage to find a way to reduce the slope withouth the hack?</p>",
          "rawMarkdown": "Hey @at7459 , I remember that before everybody started to hack the metric when the host said they won't do anything about this, you had the first place and the score about 0.620. Was it also the result of hacking or did you manage to find a way to reduce the slope withouth the hack?",
          "votes": 2,
          "replies": [
            {
              "id": 2842050,
              "postDate": "2024-05-28T20:29:39.137Z",
              "content": "<p>well, i figured there are really only two ways to deal with stability</p>\n<ol>\n<li>retrain the model frequently</li>\n<li>make sure your data sources are consistent and reliable across time</li>\n</ol>\n<p>in the data there are columns called \"riskassessment\" which seem to be models from credit bureaus, which take care of one or both of these problems over the long time frame of the test set.</p>\n<p>they are more stable and likely use more first hand information than our model, which is not retrained frequently and uses a lot of data from various different data sources which can become unreliable and inconsistent.</p>\n<p>the columns are over 90% nan, but even then, just ensembling them with your model makes it a lot more stable under this metric. in my validation setup - which looked at the longest periods of time where the slope is negative in the train set - this has about a 80% chance of increasing your score by an avg ov 0.015 under the competition metric, the longer the time period the more beneficial it becomes.</p>\n<p>i had one submission, which was overfit on the leaderboard and doesn't preserve the meaning of a probabilty, that scores 0.617 on the public lb and 0.549 on the private leaderboard</p>\n<p>and another which had the best validation score for periods of negative slope and also preserve the meaning of a probabilty which scored 0.611 on the public lb and 0.544 on the private lb</p>\n<p>it is roughly a 0.04 increase in score in both the public and private lb without hacking the metric</p>\n<p>now, i wouldn't use any of this in actual practice because it comes at the cost of making your predictions worse during periods of positive slope; it is just a side effect of a terrible asymmetric metric and knowing that the testing period has a negative slope</p>\n<p>but the thing to learn there is still the same, that stability comes from retraining your model frequently and making sure your data sources are correct</p>",
              "rawMarkdown": "well, i figured there are really only two ways to deal with stability\n\n1. retrain the model frequently\n2. make sure your data sources are consistent and reliable across time\n\nin the data there are columns called \"riskassessment\" which seem to be models from credit bureaus, which take care of one or both of these problems over the long time frame of the test set.\n\nthey are more stable and likely use more first hand information than our model, which is not retrained frequently and uses a lot of data from various different data sources which can become unreliable and inconsistent.\n\nthe columns are over 90% nan, but even then, just ensembling them with your model makes it a lot more stable under this metric. in my validation setup - which looked at the longest periods of time where the slope is negative in the train set - this has about a 80% chance of increasing your score by an avg ov 0.015 under the competition metric, the longer the time period the more beneficial it becomes.\n\ni had one submission, which was overfit on the leaderboard and doesn't preserve the meaning of a probabilty, that scores 0.617 on the public lb and 0.549 on the private leaderboard\n\nand another which had the best validation score for periods of negative slope and also preserve the meaning of a probabilty which scored 0.611 on the public lb and 0.544 on the private lb\n\nit is roughly a 0.04 increase in score in both the public and private lb without hacking the metric\n\nnow, i wouldn't use any of this in actual practice because it comes at the cost of making your predictions worse during periods of positive slope; it is just a side effect of a terrible asymmetric metric and knowing that the testing period has a negative slope\n\nbut the thing to learn there is still the same, that stability comes from retraining your model frequently and making sure your data sources are correct\n\n",
              "votes": 2
            }
          ]
        },
        {
          "id": 2841797,
          "postDate": "2024-05-28T17:37:28.823Z",
          "content": "<p>hypothesis in edit2 is completely incorrect.</p>",
          "rawMarkdown": "hypothesis in edit2 is completely incorrect.",
          "votes": -4,
          "replies": [
            {
              "id": 2842056,
              "postDate": "2024-05-28T20:33:46.407Z",
              "content": "<p>yeah, obviously as you revealed in another thread now you decided to change the test set mid competition by removing the first 30 weeks of only the private leaderboard for seemingly no reason at all other than to screw over anyone who assumed everything stayed the same.</p>",
              "rawMarkdown": "yeah, obviously as you revealed in another thread now you decided to change the test set mid competition by removing the first 30 weeks of only the private leaderboard for seemingly no reason at all other than to screw over anyone who assumed everything stayed the same.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2840401,
      "postDate": "2024-05-28T04:59:05.943Z",
      "content": "<p>I stopped trying \"hard\" a while ago in this competition, and if I selected a hacked-based sub I would have made it to the medal area (bronze most likely). </p>\n<p>But I am very disappointed with the competition.</p>\n<p>Best of luck to you all! </p>",
      "rawMarkdown": "I stopped trying \"hard\" a while ago in this competition, and if I selected a hacked-based sub I would have made it to the medal area (bronze most likely). \n\nBut I am very disappointed with the competition.\n\nBest of luck to you all! \n",
      "votes": 5
    },
    {
      "id": 2841508,
      "postDate": "2024-05-28T15:43:04.303Z",
      "content": "<p>I cant wait for HomeCredit to implement these hacks in production for improved decisions making for their customers!</p>",
      "rawMarkdown": "I cant wait for HomeCredit to implement these hacks in production for improved decisions making for their customers!",
      "votes": 4,
      "replies": [
        {
          "id": 2841937,
          "postDate": "2024-05-28T19:01:29.897Z",
          "content": "<p>Bet that next year, they will launch a new competition with the same objective, after they found out that an over-engineered metric and solutions that make the best exploitation of it do not guarantee success!</p>",
          "rawMarkdown": "Bet that next year, they will launch a new competition with the same objective, after they found out that an over-engineered metric and solutions that make the best exploitation of it do not guarantee success!",
          "votes": 3
        }
      ]
    },
    {
      "id": 2846467,
      "postDate": "2024-05-31T05:57:12.227Z",
      "content": "<p>Could you please share your hacked version? Thank you!</p>",
      "rawMarkdown": "Could you please share your hacked version? Thank you!"
    },
    {
      "id": 2840423,
      "postDate": "2024-05-28T05:05:30.450Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2840518,
      "author_name": "at7459",
      "author_url": "",
      "post_date": "2024-05-28T06:01:48.870000",
      "content": "<p>i was giving the hosts the benefit of the doubt at first, but their decisions just got worse and worse throughout the competition</p>\n<p>and a bizarre split like that just to make sure nobody knows what they should select for submission makes me think the hosts really are just incompetent beyond repair</p>\n<p>edit: from what i can gather from my submissions they had 2021 and maybe 1/3 of 2022 as the public lb and the rest from 2021 onward as private lb</p>\n<p>which sucks for me because i just hacked 2021, if i just sorted by the dates and hacked the first half i would've scored something around 0.63</p>\n<p>i feel like the only reasonable split to have the best feedback from the leaderboard is to take 30/70 of EVERY week_num or date_decision, the way they did it is just misleading and damaging for competitors</p>\n<p>edit2: actually looking at it further the split they did must've been even more bizarre than that, they must've taken all of 2021 (because all my submissions which hack 2021 score the same) a bit of 2022 and then also bit near the end of 2024 and/or 2023 (because originally if i hack the first half of the week_nums it is possible to increase the score significantly, which i think wouldn't be possible if they just took the first 30% of the data as the public leaderboard)</p>\n<p>all just to make things as misleading as possible… they should really make it standard to reveal how the data was split in the competition description.</p>",
      "votes": 20,
      "replies": [
        {
          "id": 2840690,
          "author_name": "minhtu.mt.mt",
          "author_url": "",
          "post_date": "2024-05-28T08:02:23.717000",
          "content": "<p>I totally agree with you. The metric hack is inevitable.<br>\nAt first, I thought that the hack wont work in private LB because the host can shift the first weeks of the private dataset so that overfit-hacking on public LB will shake up.<br>\nBut on the last day, I realized that the hack will always works. They key is how to tune a stable hacking method so that it can generalize well in both public and private LB. At this point, we had to gamble about the distribution of private dataset.<br>\nFortunately (and thank to a lot of probing results), we won the bet.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2840700,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-05-28T08:09:11.857000",
              "content": "<p>yeah, unlucky for me i just stopped as soon as i figured out you can get the year back, and jsut assumed they would do aany type of reasonable split</p>\n<p>i should've just done more probing and at least reconstructed my score from before the dataset got transformed… it would've been obvious then that they didn't do a reasonable split and i would've made a completely different submission…</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2840858,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2024-05-28T09:27:29.113000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> , I remember that before everybody started to hack the metric when the host said they won't do anything about this, you had the first place and the score about 0.620. Was it also the result of hacking or did you manage to find a way to reduce the slope withouth the hack?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2842050,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-05-28T20:29:39.137000",
              "content": "<p>well, i figured there are really only two ways to deal with stability</p>\n<ol>\n<li>retrain the model frequently</li>\n<li>make sure your data sources are consistent and reliable across time</li>\n</ol>\n<p>in the data there are columns called \"riskassessment\" which seem to be models from credit bureaus, which take care of one or both of these problems over the long time frame of the test set.</p>\n<p>they are more stable and likely use more first hand information than our model, which is not retrained frequently and uses a lot of data from various different data sources which can become unreliable and inconsistent.</p>\n<p>the columns are over 90% nan, but even then, just ensembling them with your model makes it a lot more stable under this metric. in my validation setup - which looked at the longest periods of time where the slope is negative in the train set - this has about a 80% chance of increasing your score by an avg ov 0.015 under the competition metric, the longer the time period the more beneficial it becomes.</p>\n<p>i had one submission, which was overfit on the leaderboard and doesn't preserve the meaning of a probabilty, that scores 0.617 on the public lb and 0.549 on the private leaderboard</p>\n<p>and another which had the best validation score for periods of negative slope and also preserve the meaning of a probabilty which scored 0.611 on the public lb and 0.544 on the private lb</p>\n<p>it is roughly a 0.04 increase in score in both the public and private lb without hacking the metric</p>\n<p>now, i wouldn't use any of this in actual practice because it comes at the cost of making your predictions worse during periods of positive slope; it is just a side effect of a terrible asymmetric metric and knowing that the testing period has a negative slope</p>\n<p>but the thing to learn there is still the same, that stability comes from retraining your model frequently and making sure your data sources are correct</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2841797,
          "author_name": "Tomas Jelinek",
          "author_url": "",
          "post_date": "2024-05-28T17:37:28.823000",
          "content": "<p>hypothesis in edit2 is completely incorrect.</p>",
          "votes": -4,
          "replies": [
            {
              "id": 2842056,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-05-28T20:33:46.407000",
              "content": "<p>yeah, obviously as you revealed in another thread now you decided to change the test set mid competition by removing the first 30 weeks of only the private leaderboard for seemingly no reason at all other than to screw over anyone who assumed everything stayed the same.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2840401,
      "author_name": "Ali",
      "author_url": "",
      "post_date": "2024-05-28T04:59:05.943000",
      "content": "<p>I stopped trying \"hard\" a while ago in this competition, and if I selected a hacked-based sub I would have made it to the medal area (bronze most likely). </p>\n<p>But I am very disappointed with the competition.</p>\n<p>Best of luck to you all! </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2841508,
      "author_name": "NxGTR",
      "author_url": "",
      "post_date": "2024-05-28T15:43:04.303000",
      "content": "<p>I cant wait for HomeCredit to implement these hacks in production for improved decisions making for their customers!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2841937,
          "author_name": "Andreas Bisiadis",
          "author_url": "",
          "post_date": "2024-05-28T19:01:29.897000",
          "content": "<p>Bet that next year, they will launch a new competition with the same objective, after they found out that an over-engineered metric and solutions that make the best exploitation of it do not guarantee success!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2846467,
      "author_name": "MrsHope",
      "author_url": "",
      "post_date": "2024-05-31T05:57:12.227000",
      "content": "<p>Could you please share your hacked version? Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2840423,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-28T05:05:30.450000",
      "content": "",
      "votes": -4,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2840272": "We have some extremely high private score submissions in some of our experience. Those submissions are all hacked-versions using reverse engineering to reproduce the WEEK_NUM. The special thing is that instead of hacking the first-half weeks, we hacked 3/4.\n\nOur 10th solution is also a hacked version, hack 3/4 weeks with linear decay (penalty first weeks more than further ones). We will post our solution in detail and the reason why we decided to choose the 3/4 version later.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1023686%2Fbb522a08fffe3913afe47b034b1ce93c%2FScreenshot%202024-05-28%20at%2011.13.40.png?generation=1716869633788833&alt=media)\n",
    "2840518": "i was giving the hosts the benefit of the doubt at first, but their decisions just got worse and worse throughout the competition\n\nand a bizarre split like that just to make sure nobody knows what they should select for submission makes me think the hosts really are just incompetent beyond repair\n\nedit: from what i can gather from my submissions they had 2021 and maybe 1/3 of 2022 as the public lb and the rest from 2021 onward as private lb\n\nwhich sucks for me because i just hacked 2021, if i just sorted by the dates and hacked the first half i would've scored something around 0.63\n\ni feel like the only reasonable split to have the best feedback from the leaderboard is to take 30/70 of EVERY week_num or date_decision, the way they did it is just misleading and damaging for competitors\n\nedit2: actually looking at it further the split they did must've been even more bizarre than that, they must've taken all of 2021 (because all my submissions which hack 2021 score the same) a bit of 2022 and then also bit near the end of 2024 and/or 2023 (because originally if i hack the first half of the week_nums it is possible to increase the score significantly, which i think wouldn't be possible if they just took the first 30% of the data as the public leaderboard)\n\nall just to make things as misleading as possible... they should really make it standard to reveal how the data was split in the competition description.",
    "2840401": "I stopped trying \"hard\" a while ago in this competition, and if I selected a hacked-based sub I would have made it to the medal area (bronze most likely). \n\nBut I am very disappointed with the competition.\n\nBest of luck to you all! \n",
    "2841508": "I cant wait for HomeCredit to implement these hacks in production for improved decisions making for their customers!",
    "2846467": "Could you please share your hacked version? Thank you!",
    "2840423": ""
  }
}