{
  "id": 507556,
  "title": "Shakeup is all you need!!",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/507556",
  "author_name": "",
  "post_date": "2024-05-26T10:01:21.871904100Z",
  "votes": 14,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Well, if u take from 60s to 91s week_num from train dataset and make from that the holdout test dataset, the adversarial validation will show u this type of probability distributions by week_num's. So, the idea is too simple to exploit the metric hack, adversarial validation is worste to predict first weeks from test dataset. And simple probability shift on this samples is good for boosting ur public LB score. But it seems this type of simple metric hack exploiting doesn't work for later periods, and probably host knows that more data in private LB is later week_num's. Also the public notebooks with metric hacking with simple mask of probability &lt; threshold found about 5 percents of total test size and this also proves observation above.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F502737%2F094f322cf5eb45ed1f6d8e82cae3b6a0%2Fdownload.png?generation=1716717063194160&amp;alt=media\"></p>\n<p>Hope host make private LB with only laters week_num's. Just wanna look at huge shakeup!!  🤟</p>",
  "messages": [
    {
      "id": "2837147",
      "postDate": "05/26/2024 10:01:21",
      "content": "<p>Well, if u take from 60s to 91s week_num from train dataset and make from that the holdout test dataset, the adversarial validation will show u this type of probability distributions by week_num's. So, the idea is too simple to exploit the metric hack, adversarial validation is worste to predict first weeks from test dataset. And simple probability shift on this samples is good for boosting ur public LB score. But it seems this type of simple metric hack exploiting doesn't work for later periods, and probably host knows that more data in private LB is later week_num's. Also the public notebooks with metric hacking with simple mask of probability &lt; threshold found about 5 percents of total test size and this also proves observation above.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F502737%2F094f322cf5eb45ed1f6d8e82cae3b6a0%2Fdownload.png?generation=1716717063194160&amp;alt=media\"></p>\n<p>Hope host make private LB with only laters week_num's. Just wanna look at huge shakeup!!  🤟</p>",
      "rawMarkdown": "Well, if u take from 60s to 91s week_num from train dataset and make from that the holdout test dataset, the adversarial validation will show u this type of probability distributions by week_num's. So, the idea is too simple to exploit the metric hack, adversarial validation is worste to predict first weeks from test dataset. And simple probability shift on this samples is good for boosting ur public LB score. But it seems this type of simple metric hack exploiting doesn't work for later periods, and probably host knows that more data in private LB is later week_num's. Also the public notebooks with metric hacking with simple mask of probability < threshold found about 5 percents of total test size and this also proves observation above.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F502737%2F094f322cf5eb45ed1f6d8e82cae3b6a0%2Fdownload.png?generation=1716717063194160&alt=media)\n\nHope host make private LB with only laters week_num's. Just wanna look at huge shakeup!!  🤟",
      "votes": null
    },
    {
      "id": "2837213",
      "postDate": "05/26/2024 11:17:18",
      "content": "<p>I also align with you on this, I think we are all collectively in for a gargantuan shakeup 2 days later <a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a> </p>",
      "rawMarkdown": "I also align with you on this, I think we are all collectively in for a gargantuan shakeup 2 days later @sggpls",
      "votes": null
    },
    {
      "id": "2838211",
      "postDate": "05/26/2024 23:44:41",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a>. It will be interesting to see what the maximum score is in the private LB. My best ensemble  without the hack could not score better than 0.594. Perhaps the limit might be below 6.00 which is way lower than all those useless submissions with the hack😀</p>",
      "rawMarkdown": "Thanks for sharing @sggpls. It will be interesting to see what the maximum score is in the private LB. My best ensemble  without the hack could not score better than 0.594. Perhaps the limit might be below 6.00 which is way lower than all those useless submissions with the hack😀",
      "votes": null
    },
    {
      "id": "2838402",
      "postDate": "05/27/2024 02:49:57",
      "content": "<p>Our best ensemble without hacking failed to score above 0.599+. It is possible that the limit could be lower than 0.612 which is much lower than all applications with the hack.</p>",
      "rawMarkdown": "Our best ensemble without hacking failed to score above 0.599+. It is possible that the limit could be lower than 0.612 which is much lower than all applications with the hack.",
      "votes": null
    },
    {
      "id": "2838418",
      "postDate": "05/27/2024 03:12:56",
      "content": "<p>If we lower the accuracy in the first week, wouldn’t the score improve to some extent, whether it’s public or private? Or is the data from the week immediately following the training data excluded in the private dataset?</p>",
      "rawMarkdown": "If we lower the accuracy in the first week, wouldn’t the score improve to some extent, whether it’s public or private? Or is the data from the week immediately following the training data excluded in the private dataset?",
      "votes": null
    },
    {
      "id": "2838482",
      "postDate": "05/27/2024 04:12:04",
      "content": "<p>Maybe a dumb question, but why do the public notebook threshold hack actually work? I have given some thoughts but could not figure out. You mentioned \"simple mask of probability &lt; threshold found about 5 percents of total test size\", where does the 5 percents come from?<br>\nCan someone kindly explain it? Really interested in the math behind it.🤔</p>",
      "rawMarkdown": "Maybe a dumb question, but why do the public notebook threshold hack actually work? I have given some thoughts but could not figure out. You mentioned \"simple mask of probability < threshold found about 5 percents of total test size\", where does the 5 percents come from?\nCan someone kindly explain it? Really interested in the math behind it.🤔",
      "votes": null
    },
    {
      "id": "2838555",
      "postDate": "05/27/2024 05:28:38",
      "content": "<p>I hope so ! It would be funny and the revenge of the organizers…🤣</p>",
      "rawMarkdown": "I hope so ! It would be funny and the revenge of the organizers…🤣",
      "votes": null
    },
    {
      "id": "2838598",
      "postDate": "05/27/2024 06:06:18",
      "content": "<p>I bet solutions based on the 2ndary model trained online for the hack will end up with OOM.. and it would be kind of sweet.. ,)</p>",
      "rawMarkdown": "I bet solutions based on the 2ndary model trained online for the hack will end up with OOM.. and it would be kind of sweet.. ,)",
      "votes": null
    },
    {
      "id": "2838869",
      "postDate": "05/27/2024 08:38:42",
      "content": "<p>Somewhere one of the hosts mentioned that the public and private split, both contain data from near covid to post covid. This was to make the public leaderboard reliable. Can't seem to find it now. That being said I still believe that the winners will have strong base model (0.60+) and a little thresholding on top. However, thresholds also have to be chosen very carefully based on some logic :) so that you are not just guessing. Anyway All the best to all of us who have worked hard to build a good base model. May the best model win (at least be in the top 10). Let's choose our two submissions carefully.</p>",
      "rawMarkdown": "Somewhere one of the hosts mentioned that the public and private split, both contain data from near covid to post covid. This was to make the public leaderboard reliable. Can't seem to find it now. That being said I still believe that the winners will have strong base model (0.60+) and a little thresholding on top. However, thresholds also have to be chosen very carefully based on some logic :) so that you are not just guessing. Anyway All the best to all of us who have worked hard to build a good base model. May the best model win (at least be in the top 10). Let's choose our two submissions carefully.",
      "votes": null
    },
    {
      "id": "2838874",
      "postDate": "05/27/2024 08:43:53",
      "content": "<p>if its scored its already has been computed for private data too, afaik</p>",
      "rawMarkdown": "if its scored its already has been computed for private data too, afaik",
      "votes": null
    },
    {
      "id": "2838891",
      "postDate": "05/27/2024 08:53:44",
      "content": "<p>That's a critical thing to know! May I ask how do you know it?</p>",
      "rawMarkdown": "That's a critical thing to know! May I ask how do you know it?",
      "votes": null
    },
    {
      "id": "2838907",
      "postDate": "05/27/2024 09:05:12",
      "content": "<p>If yes is the answer for the second ur question or it is most of the samples from first week in public than there is a shake up ;)</p>",
      "rawMarkdown": "If yes is the answer for the second ur question or it is most of the samples from first week in public than there is a shake up ;)",
      "votes": null
    },
    {
      "id": "2838980",
      "postDate": "05/27/2024 10:12:05",
      "content": "<p>I believe that latest public hack affected only first 10-20 weeks. If hosts used all weeks for  random  public/private split, hack will work on private too, may be less effective (within .001-.005). Otherwise it will be a huge shake up :)<br>\nBig chances that public hack not very far from hacking limit  (based on my simulation) and this gives us some equality and now we are competing based on initial strength of our models. </p>\n<p>My model without any attempts to lower score has around .605 on LB and I believe it could be higher</p>",
      "rawMarkdown": "I believe that latest public hack affected only first 10-20 weeks. If hosts used all weeks for  random  public/private split, hack will work on private too, may be less effective (within .001-.005). Otherwise it will be a huge shake up :)\nBig chances that public hack not very far from hacking limit  (based on my simulation) and this gives us some equality and now we are competing based on initial strength of our models. \n\nMy model without any attempts to lower score has around .605 on LB and I believe it could be higher",
      "votes": null
    },
    {
      "id": "2838986",
      "postDate": "05/27/2024 10:18:54",
      "content": "<p>As I know, hosts can see both public and private scores before the deadline, which means scores are prepared for full test, but you can only see the score for public part as a participant</p>",
      "rawMarkdown": "As I know, hosts can see both public and private scores before the deadline, which means scores are prepared for full test, but you can only see the score for public part as a participant",
      "votes": null
    },
    {
      "id": "2839009",
      "postDate": "05/27/2024 10:37:42",
      "content": "<p>Really? Wanna see that comment. On the contrary, I saw the host's comment that they wouldn't disclose how they divided the test dataset into public and private. </p>",
      "rawMarkdown": "Really? Wanna see that comment. On the contrary, I saw the host's comment that they wouldn't disclose how they divided the test dataset into public and private.",
      "votes": null
    },
    {
      "id": "2839027",
      "postDate": "05/27/2024 10:48:16",
      "content": "<p>That's true. Changing the ratio of the samples from the first week is a simple and effective way to inactivate the popular hacking method (the one with adversarial validation). Let's see how this competition ends :D</p>",
      "rawMarkdown": "That's true. Changing the ratio of the samples from the first week is a simple and effective way to inactivate the popular hacking method (the one with adversarial validation). Let's see how this competition ends :D",
      "votes": null
    },
    {
      "id": "2839041",
      "postDate": "05/27/2024 11:03:49",
      "content": "<p>Found it. I hope this is the case. Otherwise, the shakeup will be huge :) . Not only in terms of thresholding but also in terms of base model variance of prediction in time.<br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764</a></p>",
      "rawMarkdown": "Found it. I hope this is the case. Otherwise, the shakeup will be huge :) . Not only in terms of thresholding but also in terms of base model variance of prediction in time.\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764",
      "votes": null
    },
    {
      "id": "2839064",
      "postDate": "05/27/2024 11:20:21",
      "content": "<p>Thanks a ton! This info is so important. </p>",
      "rawMarkdown": "Thanks a ton! This info is so important.",
      "votes": null
    },
    {
      "id": "2839082",
      "postDate": "05/27/2024 11:29:54",
      "content": "<p>More likely that in this answer Tomas wrote about train dataset as it was open to public</p>",
      "rawMarkdown": "More likely that in this answer Tomas wrote about train dataset as it was open to public",
      "votes": null
    },
    {
      "id": "2839090",
      "postDate": "05/27/2024 11:34:43",
      "content": "<p><a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> Read Daniel's comment later in the discussion :)</p>",
      "rawMarkdown": "johnpateha Read Daniel's comment later in the discussion :)",
      "votes": null
    },
    {
      "id": "2839111",
      "postDate": "05/27/2024 11:46:50",
      "content": "<p>Thank you, you are right. It could be different weeks too, but likely they used random split for each of week</p>",
      "rawMarkdown": "Thank you, you are right. It could be different weeks too, but likely they used random split for each of week",
      "votes": null
    },
    {
      "id": "2839117",
      "postDate": "05/27/2024 11:52:07",
      "content": "<p>yes, otherwise I might have chosen the wrong competition as my first tabular on Kaggle 🤣</p>",
      "rawMarkdown": "yes, otherwise I might have chosen the wrong competition as my first tabular on Kaggle 🤣",
      "votes": null
    },
    {
      "id": "2839123",
      "postDate": "05/27/2024 11:53:41",
      "content": "<p>And how about the hack of retrieving WEEK_NUM and using them to lower score for initial weeks? I remember you made a post about it.</p>",
      "rawMarkdown": "And how about the hack of retrieving WEEK_NUM and using them to lower score for initial weeks? I remember you made a post about it.",
      "votes": null
    },
    {
      "id": "2839158",
      "postDate": "05/27/2024 12:11:27",
      "content": "<p>There many ways to restore weeks for hack, but final results were not very different for my models, I prefer to build better model </p>",
      "rawMarkdown": "There many ways to restore weeks for hack, but final results were not very different for my models, I prefer to build better model",
      "votes": null
    },
    {
      "id": "2839165",
      "postDate": "05/27/2024 12:14:35",
      "content": "<p>Every competition teaches us something. No matter how this competition ends, in my ranking, it is far from the worst. It’s hard to knock the M5 off the podium.</p>",
      "rawMarkdown": "Every competition teaches us something. No matter how this competition ends, in my ranking, it is far from the worst. It’s hard to knock the M5 off the podium.",
      "votes": null
    },
    {
      "id": "2839192",
      "postDate": "05/27/2024 12:22:35",
      "content": "<p>I'm not sure if a 'better model' in Public LB is worth exploring. If the public/private used all weeks for both then yes a better model is worth building provided a good hack is implemented. Otherwise, if the private/public were split based on week_num (private is later weeks) then I think the results will be random. The train set has a positive slope (and that's why we all are getting good stability in our CVs), and the public set has a negative one (hence the need for hacking), so in that case will depend on the starting period chosen for the private set.</p>",
      "rawMarkdown": "I'm not sure if a 'better model' in Public LB is worth exploring. If the public/private used all weeks for both then yes a better model is worth building provided a good hack is implemented. Otherwise, if the private/public were split based on week_num (private is later weeks) then I think the results will be random. The train set has a positive slope (and that's why we all are getting good stability in our CVs), and the public set has a negative one (hence the need for hacking), so in that case will depend on the starting period chosen for the private set.",
      "votes": null
    },
    {
      "id": "2839209",
      "postDate": "05/27/2024 12:34:11",
      "content": "<p>if private starts after public the impact of hack would be not so big. cutting 5% from scores lead to decrease Gini by around 0.1 It's too much for shorter period. <br>\nIt's possible that split was done by weeks, for example 1st week public, 2-3 - private, 4 - public, etc. In that case hack will work on private too with some fluctuation. But fluctuation is possible even if each weeks was splitted 30 x 70</p>",
      "rawMarkdown": "if private starts after public the impact of hack would be not so big. cutting 5% from scores lead to decrease Gini by around 0.1 It's too much for shorter period. \nIt's possible that split was done by weeks, for example 1st week public, 2-3 - private, 4 - public, etc. In that case hack will work on private too with some fluctuation. But fluctuation is possible even if each weeks was splitted 30 x 70",
      "votes": null
    },
    {
      "id": "2839230",
      "postDate": "05/27/2024 12:46:28",
      "content": "<p>Agreed, sir.</p>",
      "rawMarkdown": "Agreed, sir.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2837213,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/26/2024 11:17:18",
      "content": "<p>I also align with you on this, I think we are all collectively in for a gargantuan shakeup 2 days later <a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2838211,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/26/2024 23:44:41",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a>. It will be interesting to see what the maximum score is in the private LB. My best ensemble  without the hack could not score better than 0.594. Perhaps the limit might be below 6.00 which is way lower than all those useless submissions with the hack😀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2838402,
      "author_name": "alexxanderlarko",
      "author_url": "",
      "post_date": "05/27/2024 02:49:57",
      "content": "<p>Our best ensemble without hacking failed to score above 0.599+. It is possible that the limit could be lower than 0.612 which is much lower than all applications with the hack.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2838418,
      "author_name": "atsuno",
      "author_url": "",
      "post_date": "05/27/2024 03:12:56",
      "content": "<p>If we lower the accuracy in the first week, wouldn’t the score improve to some extent, whether it’s public or private? Or is the data from the week immediately following the training data excluded in the private dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2838907,
          "author_name": "sggpls",
          "author_url": "",
          "post_date": "05/27/2024 09:05:12",
          "content": "<p>If yes is the answer for the second ur question or it is most of the samples from first week in public than there is a shake up ;)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2839027,
              "author_name": "atsuno",
              "author_url": "",
              "post_date": "05/27/2024 10:48:16",
              "content": "<p>That's true. Changing the ratio of the samples from the first week is a simple and effective way to inactivate the popular hacking method (the one with adversarial validation). Let's see how this competition ends :D</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2838482,
      "author_name": "shyhandsome",
      "author_url": "",
      "post_date": "05/27/2024 04:12:04",
      "content": "<p>Maybe a dumb question, but why do the public notebook threshold hack actually work? I have given some thoughts but could not figure out. You mentioned \"simple mask of probability &lt; threshold found about 5 percents of total test size\", where does the 5 percents come from?<br>\nCan someone kindly explain it? Really interested in the math behind it.🤔</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2838555,
      "author_name": "pourchot",
      "author_url": "",
      "post_date": "05/27/2024 05:28:38",
      "content": "<p>I hope so ! It would be funny and the revenge of the organizers…🤣</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2838598,
      "author_name": "lohmaa",
      "author_url": "",
      "post_date": "05/27/2024 06:06:18",
      "content": "<p>I bet solutions based on the 2ndary model trained online for the hack will end up with OOM.. and it would be kind of sweet.. ,)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2838874,
          "author_name": "bluepill",
          "author_url": "",
          "post_date": "05/27/2024 08:43:53",
          "content": "<p>if its scored its already has been computed for private data too, afaik</p>",
          "votes": null,
          "replies": [
            {
              "id": 2838891,
              "author_name": "lohmaa",
              "author_url": "",
              "post_date": "05/27/2024 08:53:44",
              "content": "<p>That's a critical thing to know! May I ask how do you know it?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2838986,
                  "author_name": "bluepill",
                  "author_url": "",
                  "post_date": "05/27/2024 10:18:54",
                  "content": "<p>As I know, hosts can see both public and private scores before the deadline, which means scores are prepared for full test, but you can only see the score for public part as a participant</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2838869,
      "author_name": "krishnapriya18",
      "author_url": "",
      "post_date": "05/27/2024 08:38:42",
      "content": "<p>Somewhere one of the hosts mentioned that the public and private split, both contain data from near covid to post covid. This was to make the public leaderboard reliable. Can't seem to find it now. That being said I still believe that the winners will have strong base model (0.60+) and a little thresholding on top. However, thresholds also have to be chosen very carefully based on some logic :) so that you are not just guessing. Anyway All the best to all of us who have worked hard to build a good base model. May the best model win (at least be in the top 10). Let's choose our two submissions carefully.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2839009,
          "author_name": "atsuno",
          "author_url": "",
          "post_date": "05/27/2024 10:37:42",
          "content": "<p>Really? Wanna see that comment. On the contrary, I saw the host's comment that they wouldn't disclose how they divided the test dataset into public and private. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2839041,
              "author_name": "krishnapriya18",
              "author_url": "",
              "post_date": "05/27/2024 11:03:49",
              "content": "<p>Found it. I hope this is the case. Otherwise, the shakeup will be huge :) . Not only in terms of thresholding but also in terms of base model variance of prediction in time.<br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2839064,
                  "author_name": "atsuno",
                  "author_url": "",
                  "post_date": "05/27/2024 11:20:21",
                  "content": "<p>Thanks a ton! This info is so important. </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2839082,
                  "author_name": "johnpateha",
                  "author_url": "",
                  "post_date": "05/27/2024 11:29:54",
                  "content": "<p>More likely that in this answer Tomas wrote about train dataset as it was open to public</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2839090,
                      "author_name": "krishnapriya18",
                      "author_url": "",
                      "post_date": "05/27/2024 11:34:43",
                      "content": "<p><a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> Read Daniel's comment later in the discussion :)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2839111,
                          "author_name": "johnpateha",
                          "author_url": "",
                          "post_date": "05/27/2024 11:46:50",
                          "content": "<p>Thank you, you are right. It could be different weeks too, but likely they used random split for each of week</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2839117,
                              "author_name": "krishnapriya18",
                              "author_url": "",
                              "post_date": "05/27/2024 11:52:07",
                              "content": "<p>yes, otherwise I might have chosen the wrong competition as my first tabular on Kaggle 🤣</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2839165,
                                  "author_name": "johnpateha",
                                  "author_url": "",
                                  "post_date": "05/27/2024 12:14:35",
                                  "content": "<p>Every competition teaches us something. No matter how this competition ends, in my ranking, it is far from the worst. It’s hard to knock the M5 off the podium.</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2839230,
                                      "author_name": "krishnapriya18",
                                      "author_url": "",
                                      "post_date": "05/27/2024 12:46:28",
                                      "content": "<p>Agreed, sir.</p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2838980,
      "author_name": "johnpateha",
      "author_url": "",
      "post_date": "05/27/2024 10:12:05",
      "content": "<p>I believe that latest public hack affected only first 10-20 weeks. If hosts used all weeks for  random  public/private split, hack will work on private too, may be less effective (within .001-.005). Otherwise it will be a huge shake up :)<br>\nBig chances that public hack not very far from hacking limit  (based on my simulation) and this gives us some equality and now we are competing based on initial strength of our models. </p>\n<p>My model without any attempts to lower score has around .605 on LB and I believe it could be higher</p>",
      "votes": null,
      "replies": [
        {
          "id": 2839123,
          "author_name": "simoelm",
          "author_url": "",
          "post_date": "05/27/2024 11:53:41",
          "content": "<p>And how about the hack of retrieving WEEK_NUM and using them to lower score for initial weeks? I remember you made a post about it.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2839158,
              "author_name": "johnpateha",
              "author_url": "",
              "post_date": "05/27/2024 12:11:27",
              "content": "<p>There many ways to restore weeks for hack, but final results were not very different for my models, I prefer to build better model </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2839192,
                  "author_name": "simoelm",
                  "author_url": "",
                  "post_date": "05/27/2024 12:22:35",
                  "content": "<p>I'm not sure if a 'better model' in Public LB is worth exploring. If the public/private used all weeks for both then yes a better model is worth building provided a good hack is implemented. Otherwise, if the private/public were split based on week_num (private is later weeks) then I think the results will be random. The train set has a positive slope (and that's why we all are getting good stability in our CVs), and the public set has a negative one (hence the need for hacking), so in that case will depend on the starting period chosen for the private set.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2839209,
                      "author_name": "johnpateha",
                      "author_url": "",
                      "post_date": "05/27/2024 12:34:11",
                      "content": "<p>if private starts after public the impact of hack would be not so big. cutting 5% from scores lead to decrease Gini by around 0.1 It's too much for shorter period. <br>\nIt's possible that split was done by weeks, for example 1st week public, 2-3 - private, 4 - public, etc. In that case hack will work on private too with some fluctuation. But fluctuation is possible even if each weeks was splitted 30 x 70</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2837147": "Well, if u take from 60s to 91s week_num from train dataset and make from that the holdout test dataset, the adversarial validation will show u this type of probability distributions by week_num's. So, the idea is too simple to exploit the metric hack, adversarial validation is worste to predict first weeks from test dataset. And simple probability shift on this samples is good for boosting ur public LB score. But it seems this type of simple metric hack exploiting doesn't work for later periods, and probably host knows that more data in private LB is later week_num's. Also the public notebooks with metric hacking with simple mask of probability < threshold found about 5 percents of total test size and this also proves observation above.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F502737%2F094f322cf5eb45ed1f6d8e82cae3b6a0%2Fdownload.png?generation=1716717063194160&alt=media)\n\nHope host make private LB with only laters week_num's. Just wanna look at huge shakeup!!  🤟",
    "2837213": "I also align with you on this, I think we are all collectively in for a gargantuan shakeup 2 days later @sggpls",
    "2838211": "Thanks for sharing @sggpls. It will be interesting to see what the maximum score is in the private LB. My best ensemble  without the hack could not score better than 0.594. Perhaps the limit might be below 6.00 which is way lower than all those useless submissions with the hack😀",
    "2838402": "Our best ensemble without hacking failed to score above 0.599+. It is possible that the limit could be lower than 0.612 which is much lower than all applications with the hack.",
    "2838418": "If we lower the accuracy in the first week, wouldn’t the score improve to some extent, whether it’s public or private? Or is the data from the week immediately following the training data excluded in the private dataset?",
    "2838482": "Maybe a dumb question, but why do the public notebook threshold hack actually work? I have given some thoughts but could not figure out. You mentioned \"simple mask of probability < threshold found about 5 percents of total test size\", where does the 5 percents come from?\nCan someone kindly explain it? Really interested in the math behind it.🤔",
    "2838555": "I hope so ! It would be funny and the revenge of the organizers…🤣",
    "2838598": "I bet solutions based on the 2ndary model trained online for the hack will end up with OOM.. and it would be kind of sweet.. ,)",
    "2838869": "Somewhere one of the hosts mentioned that the public and private split, both contain data from near covid to post covid. This was to make the public leaderboard reliable. Can't seem to find it now. That being said I still believe that the winners will have strong base model (0.60+) and a little thresholding on top. However, thresholds also have to be chosen very carefully based on some logic :) so that you are not just guessing. Anyway All the best to all of us who have worked hard to build a good base model. May the best model win (at least be in the top 10). Let's choose our two submissions carefully.",
    "2838874": "if its scored its already has been computed for private data too, afaik",
    "2838891": "That's a critical thing to know! May I ask how do you know it?",
    "2838907": "If yes is the answer for the second ur question or it is most of the samples from first week in public than there is a shake up ;)",
    "2838980": "I believe that latest public hack affected only first 10-20 weeks. If hosts used all weeks for  random  public/private split, hack will work on private too, may be less effective (within .001-.005). Otherwise it will be a huge shake up :)\nBig chances that public hack not very far from hacking limit  (based on my simulation) and this gives us some equality and now we are competing based on initial strength of our models. \n\nMy model without any attempts to lower score has around .605 on LB and I believe it could be higher",
    "2838986": "As I know, hosts can see both public and private scores before the deadline, which means scores are prepared for full test, but you can only see the score for public part as a participant",
    "2839009": "Really? Wanna see that comment. On the contrary, I saw the host's comment that they wouldn't disclose how they divided the test dataset into public and private.",
    "2839027": "That's true. Changing the ratio of the samples from the first week is a simple and effective way to inactivate the popular hacking method (the one with adversarial validation). Let's see how this competition ends :D",
    "2839041": "Found it. I hope this is the case. Otherwise, the shakeup will be huge :) . Not only in terms of thresholding but also in terms of base model variance of prediction in time.\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764",
    "2839064": "Thanks a ton! This info is so important.",
    "2839082": "More likely that in this answer Tomas wrote about train dataset as it was open to public",
    "2839090": "johnpateha Read Daniel's comment later in the discussion :)",
    "2839111": "Thank you, you are right. It could be different weeks too, but likely they used random split for each of week",
    "2839117": "yes, otherwise I might have chosen the wrong competition as my first tabular on Kaggle 🤣",
    "2839123": "And how about the hack of retrieving WEEK_NUM and using them to lower score for initial weeks? I remember you made a post about it.",
    "2839158": "There many ways to restore weeks for hack, but final results were not very different for my models, I prefer to build better model",
    "2839165": "Every competition teaches us something. No matter how this competition ends, in my ranking, it is far from the worst. It’s hard to knock the M5 off the podium.",
    "2839192": "I'm not sure if a 'better model' in Public LB is worth exploring. If the public/private used all weeks for both then yes a better model is worth building provided a good hack is implemented. Otherwise, if the private/public were split based on week_num (private is later weeks) then I think the results will be random. The train set has a positive slope (and that's why we all are getting good stability in our CVs), and the public set has a negative one (hence the need for hacking), so in that case will depend on the starting period chosen for the private set.",
    "2839209": "if private starts after public the impact of hack would be not so big. cutting 5% from scores lead to decrease Gini by around 0.1 It's too much for shorter period. \nIt's possible that split was done by weeks, for example 1st week public, 2-3 - private, 4 - public, etc. In that case hack will work on private too with some fluctuation. But fluctuation is possible even if each weeks was splitted 30 x 70",
    "2839230": "Agreed, sir."
  },
  "source": "meta"
}