{
  "id": 501654,
  "title": "Is the explosion of good scores related to restoring of WEEK_NUM?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501654",
  "author_name": "narsil (jobs-in-data.com)",
  "post_date": "2024-05-10T07:48:57.164000",
  "votes": 22,
  "comment_count": 24,
  "views": 0,
  "content": "<p>We have seen for a long time a stagnation of scores around 0.600 and now we have an explosion of scores between .620- .650. Especially considering that:</p>\n<ul>\n<li>it happened after <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167\" target=\"_blank\">this</a> great post from <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a></li>\n<li>we know that metric hacking based on <code>WEEK_NUM</code> <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449\" target=\"_blank\">leads to great improvements in scores</a></li>\n</ul>\n<p>I am very curious if this is indeed related to restoring <code>WEEK_NUM</code> and thus my prophecy is being fulfilled that I posted as an immediate reaction to the organizers' way of handling the metric hacking problem - that once you let the Genie out of the bottle, there is no coming back.</p>",
  "messages": [
    {
      "id": 2804775,
      "postDate": "2024-05-10T07:48:57.163Z",
      "content": "<p>We have seen for a long time a stagnation of scores around 0.600 and now we have an explosion of scores between .620- .650. Especially considering that:</p>\n<ul>\n<li>it happened after <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167\" target=\"_blank\">this</a> great post from <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a></li>\n<li>we know that metric hacking based on <code>WEEK_NUM</code> <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449\" target=\"_blank\">leads to great improvements in scores</a></li>\n</ul>\n<p>I am very curious if this is indeed related to restoring <code>WEEK_NUM</code> and thus my prophecy is being fulfilled that I posted as an immediate reaction to the organizers' way of handling the metric hacking problem - that once you let the Genie out of the bottle, there is no coming back.</p>",
      "rawMarkdown": "We have seen for a long time a stagnation of scores around 0.600 and now we have an explosion of scores between .620- .650. Especially considering that:\n- it happened after [this](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167) great post from @johnpateha\n- we know that metric hacking based on `WEEK_NUM` [leads to great improvements in scores](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449)\n\nI am very curious if this is indeed related to restoring `WEEK_NUM` and thus my prophecy is being fulfilled that I posted as an immediate reaction to the organizers' way of handling the metric hacking problem - that once you let the Genie out of the bottle, there is no coming back.",
      "votes": 22
    },
    {
      "id": 2805290,
      "postDate": "2024-05-10T13:35:04Z",
      "content": "<p>yes, explosion of scores definitely related to the metric hack, after clarification that it not breaks Kaggle rules. For me it gave 5+% up on LB<br>\nIt's sad, but similar situations were many times previously here, so nothing new. </p>\n<p>Metric's hack has some limit, so big chance that strong models will win anyway.<br>\nSome participants will find it by themselves, other will join a teams with someone who found it and finally we will have competition between strong models on top level. Previously Kaggle worked this way. </p>",
      "rawMarkdown": "yes, explosion of scores definitely related to the metric hack, after clarification that it not breaks Kaggle rules. For me it gave 5+% up on LB\nIt's sad, but similar situations were many times previously here, so nothing new. \n\nMetric's hack has some limit, so big chance that strong models will win anyway.\nSome participants will find it by themselves, other will join a teams with someone who found it and finally we will have competition between strong models on top level. Previously Kaggle worked this way. \n",
      "votes": 6,
      "replies": [
        {
          "id": 2805326,
          "postDate": "2024-05-10T13:58:27.763Z",
          "content": "<p>Unfortunately, one also needs to finetune the hack to maximize the score and hope that the chosen parameters will also be optimal for the hidden dataset. This introduces an element of randomness to the final scores, so the winner might not be the person with the best AUC, but rather the one who best randomly guessed how to optimally hack the metric. Even if it can be probed, the winner could be the one who discovered the hack earlier and therefore had more submissions left than others for probing.</p>",
          "rawMarkdown": "Unfortunately, one also needs to finetune the hack to maximize the score and hope that the chosen parameters will also be optimal for the hidden dataset. This introduces an element of randomness to the final scores, so the winner might not be the person with the best AUC, but rather the one who best randomly guessed how to optimally hack the metric. Even if it can be probed, the winner could be the one who discovered the hack earlier and therefore had more submissions left than others for probing.",
          "votes": 7,
          "replies": [
            {
              "id": 2805410,
              "postDate": "2024-05-10T14:34:49.820Z",
              "content": "<p>You are right that super tune of hack on LB could lead to overfit on private, so the quality of base model is still very important. <br>\nWe don't know how exactly test data was splitted, so shake up is possible. Will see it in three weeks. <br>\nI remember M5 competition, where hosts chose as private data period with high anomaly, compared with train and public data and very specific custom metric too. There was a huge shake up and I'm not sure that host got best solutions, but may be their goal was different. </p>\n<p>This time we also have specific metric which could be good for business but bad for competitions.  May be next competitions will avoid the same problem.</p>",
              "rawMarkdown": "You are right that super tune of hack on LB could lead to overfit on private, so the quality of base model is still very important. \nWe don't know how exactly test data was splitted, so shake up is possible. Will see it in three weeks. \nI remember M5 competition, where hosts chose as private data period with high anomaly, compared with train and public data and very specific custom metric too. There was a huge shake up and I'm not sure that host got best solutions, but may be their goal was different. \n\nThis time we also have specific metric which could be good for business but bad for competitions.  May be next competitions will avoid the same problem.",
              "votes": 3
            },
            {
              "id": 2805427,
              "postDate": "2024-05-10T14:42:46.237Z",
              "content": "<p>and of course you have two submission to choose for private scoring, one could be with aggressive hack and another with light correction or without it at all</p>",
              "rawMarkdown": "and of course you have two submission to choose for private scoring, one could be with aggressive hack and another with light correction or without it at all",
              "votes": 2
            },
            {
              "id": 2805725,
              "postDate": "2024-05-10T18:04:04.963Z",
              "content": "<p>For anyone who doesn't remember (or doesn't want to remember):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2F442140b962675a0bc1e7b009ce63c74c%2FM5_shakeup_plot.png?generation=1715364197906310&amp;alt=media\"><br>\n(Source: <a href=\"https://www.kaggle.com/code/carlmcbrideellis/shakeup-scatterplots-boxes-strings-and-things\" target=\"_blank\">\"<em>Shakeup scatterplots: Boxes, strings and things…</em>\"</a>)</p>",
              "rawMarkdown": "For anyone who doesn't remember (or doesn't want to remember):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2F442140b962675a0bc1e7b009ce63c74c%2FM5_shakeup_plot.png?generation=1715364197906310&alt=media)\n(Source: [\"*Shakeup scatterplots: Boxes, strings and things...*\"](https://www.kaggle.com/code/carlmcbrideellis/shakeup-scatterplots-boxes-strings-and-things))",
              "votes": 7
            },
            {
              "id": 2806100,
              "postDate": "2024-05-10T21:45:18.740Z",
              "content": "<p>The problem here is that the variance in the score related to the hack can be similar to or greater than the variance in the score related to the quality of the model. In the previous Home Credit competition, the top solution scores ranged from 0.6 to 0.61, whereas the impact of the hack here can be up to 0.4 (?). And it's unclear how this impact might change on the private leaderboard.</p>\n<p>But I'm not suggesting any solutions; just expressing concern about an upcoming shakeup:) I actually think that the choice of the metric makes this competition more reflective of real life, as we often must work with strange KPIs that we need to optimize blindly.</p>",
              "rawMarkdown": "The problem here is that the variance in the score related to the hack can be similar to or greater than the variance in the score related to the quality of the model. In the previous Home Credit competition, the top solution scores ranged from 0.6 to 0.61, whereas the impact of the hack here can be up to 0.4 (?). And it's unclear how this impact might change on the private leaderboard.\n\nBut I'm not suggesting any solutions; just expressing concern about an upcoming shakeup:) I actually think that the choice of the metric makes this competition more reflective of real life, as we often must work with strange KPIs that we need to optimize blindly.",
              "votes": 3
            },
            {
              "id": 2806166,
              "postDate": "2024-05-10T23:49:07.217Z",
              "content": "<p><a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> Yeah, I agree. Optimal recovery in a blackbox is a little weird, but it is still a stimulating challenge.</p>",
              "rawMarkdown": "@eivolkova Yeah, I agree. Optimal recovery in a blackbox is a little weird, but it is still a stimulating challenge."
            },
            {
              "id": 2812669,
              "postDate": "2024-05-14T10:52:04.407Z",
              "content": "<p>Maybe! Just  so</p>",
              "rawMarkdown": "Maybe! Just  so"
            }
          ]
        },
        {
          "id": 2805409,
          "postDate": "2024-05-10T14:34:48.730Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2804900,
      "postDate": "2024-05-10T08:55:37.670Z",
      "content": "<p>I had a good model without hacking and now it falls in the bottom of the PLB 🤣<br>\nThe ironic of the situation is the organizer who looked for getting stable models and after hiding the dates and now changing the rules again, he will get the best hacked metrics,and not necessarily with good models because the hacking is so efficient… I wonder if his company wanted to spend money for that ? 🤔</p>",
      "rawMarkdown": "I had a good model without hacking and now it falls in the bottom of the PLB 🤣\nThe ironic of the situation is the organizer who looked for getting stable models and after hiding the dates and now changing the rules again, he will get the best hacked metrics,and not necessarily with good models because the hacking is so efficient... I wonder if his company wanted to spend money for that ? 🤔",
      "votes": 5,
      "replies": [
        {
          "id": 2804939,
          "postDate": "2024-05-10T09:09:12.203Z",
          "content": "<p>I guess even without intentionally hacking it, that metric will still lead to models that happens to be worse (worse than they could be) on the initial test period. I wonder if that's really desirable? 🤷‍♂️</p>",
          "rawMarkdown": "I guess even without intentionally hacking it, that metric will still lead to models that happens to be worse (worse than they could be) on the initial test period. I wonder if that's really desirable? 🤷‍♂️",
          "votes": 2
        },
        {
          "id": 2805207,
          "postDate": "2024-05-10T12:39:49.613Z",
          "content": "<p>My guess is they will still get a lot of value if they read Top 100 -200 solutions - just because of the tons of work community did there.<br>\nBut - the value of a solution from the host perspective may be only loosely correlated with its LB score.</p>",
          "rawMarkdown": "My guess is they will still get a lot of value if they read Top 100 -200 solutions - just because of the tons of work community did there.\nBut - the value of a solution from the host perspective may be only loosely correlated with its LB score."
        },
        {
          "id": 2809540,
          "postDate": "2024-05-12T20:20:57.843Z",
          "content": "<p>I see you caught up <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> ! :) </p>",
          "rawMarkdown": "I see you caught up @pourchot ! :) ",
          "replies": [
            {
              "id": 2809959,
              "postDate": "2024-05-13T04:47:42.740Z",
              "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> : Good model + metric hacking 🤣 </p>",
              "rawMarkdown": "@narsil : Good model + metric hacking 🤣 ",
              "votes": 2
            },
            {
              "id": 2813706,
              "postDate": "2024-05-15T01:58:15.547Z",
              "content": "<p>Have you ever use model stacking or something? Or spent more time on the feature engineering part? I found myself with no improvement for a relatively long period of time since I'm really just a beginner without ideas. Could you give me some advice to start again? 🥹</p>",
              "rawMarkdown": "Have you ever use model stacking or something? Or spent more time on the feature engineering part? I found myself with no improvement for a relatively long period of time since I'm really just a beginner without ideas. Could you give me some advice to start again? 🥹"
            }
          ]
        }
      ]
    },
    {
      "id": 2804781,
      "postDate": "2024-05-10T07:53:11.423Z",
      "content": "<p>Yes, it actually happened after the host stated they are not going to disqualify those who hack the metric, contrary to what was said earlier. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412</a></p>",
      "rawMarkdown": "Yes, it actually happened after the host stated they are not going to disqualify those who hack the metric, contrary to what was said earlier. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412",
      "votes": 5,
      "replies": [
        {
          "id": 2805197,
          "postDate": "2024-05-10T12:35:54.363Z",
          "content": "<p>Yes, I saw this…</p>\n<p>Well - what can I say - I had the intuition this would happen the moment when hosts declined to confirm that metrick hacks based on restored <code>WEEK_NUM</code> would be ineligible for Kaggle points.</p>\n<p>We have in Poland the Kaggle meetup group with the official name \"Leaks &amp; Shakeups\". Maybe we should change our name to \"Leaks, Shakeups and Rule Changes\" 😉</p>",
          "rawMarkdown": "Yes, I saw this...\n\nWell - what can I say - I had the intuition this would happen the moment when hosts declined to confirm that metrick hacks based on restored `WEEK_NUM` would be ineligible for Kaggle points.\n\nWe have in Poland the Kaggle meetup group with the official name \"Leaks & Shakeups\". Maybe we should change our name to \"Leaks, Shakeups and Rule Changes\" 😉",
          "votes": 2
        }
      ]
    },
    {
      "id": 2804930,
      "postDate": "2024-05-10T09:04:42.130Z",
      "content": "<p>It seems that everyone is able to get improvement from hacking the metric. Maybe the way to restore WEEK_NUM is posted somewhere I do not have access to? 😯</p>",
      "rawMarkdown": "It seems that everyone is able to get improvement from hacking the metric. Maybe the way to restore WEEK_NUM is posted somewhere I do not have access to? 😯",
      "votes": 3,
      "replies": [
        {
          "id": 2805014,
          "postDate": "2024-05-10T10:13:20.250Z",
          "content": "<p>No, this is top secret. I can't restore the weeks too. </p>",
          "rawMarkdown": "No, this is top secret. I can't restore the weeks too. "
        },
        {
          "id": 2805831,
          "postDate": "2024-05-10T18:53:11.477Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuanzhezhou\" target=\"_blank\">@yuanzhezhou</a> feel free not to answer: is it true that you've reached your current public LB position without any WEEK_NUM / metric hacking?</p>",
          "rawMarkdown": "@yuanzhezhou feel free not to answer: is it true that you've reached your current public LB position without any WEEK_NUM / metric hacking?"
        },
        {
          "id": 2806046,
          "postDate": "2024-05-10T20:39:08.977Z",
          "content": "<p>Actually, I am finding it difficult to get any benefit with my models. If you look at the training data, it is possible to recover WEEK_NUM from date_decision. However, I’m not sure that the aggregation tracks in the test set. Another option is perhaps MONTH which is less precise but surely tracks the aggregation properly. Not sure if this is how everyone else is doing it or not.</p>",
          "rawMarkdown": "Actually, I am finding it difficult to get any benefit with my models. If you look at the training data, it is possible to recover WEEK_NUM from date_decision. However, I’m not sure that the aggregation tracks in the test set. Another option is perhaps MONTH which is less precise but surely tracks the aggregation properly. Not sure if this is how everyone else is doing it or not."
        }
      ]
    },
    {
      "id": 2805001,
      "postDate": "2024-05-10T10:01:35.753Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2805066,
          "postDate": "2024-05-10T10:45:29.843Z",
          "content": "<p>I'm afraid it will work on private lb too, but let's hope I'm mistaken. Anyway, that sounds like a good idea. </p>\n<p>If I will do anything more on this I will try creating one \"good\" solution (meaning aiming for an ensemble with high AUC in general (like high lowest AUC), without reducing any performance on purpose), and one \"hacking\" solution more out of curiosity. </p>",
          "rawMarkdown": "I'm afraid it will work on private lb too, but let's hope I'm mistaken. Anyway, that sounds like a good idea. \n\nIf I will do anything more on this I will try creating one \"good\" solution (meaning aiming for an ensemble with high AUC in general (like high lowest AUC), without reducing any performance on purpose), and one \"hacking\" solution more out of curiosity. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2806238,
      "postDate": "2024-05-11T01:24:33.553Z",
      "content": "<p>thanks for disccusions</p>",
      "rawMarkdown": "thanks for disccusions",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 2805290,
      "author_name": "Evgeny Patekha",
      "author_url": "",
      "post_date": "2024-05-10T13:35:04",
      "content": "<p>yes, explosion of scores definitely related to the metric hack, after clarification that it not breaks Kaggle rules. For me it gave 5+% up on LB<br>\nIt's sad, but similar situations were many times previously here, so nothing new. </p>\n<p>Metric's hack has some limit, so big chance that strong models will win anyway.<br>\nSome participants will find it by themselves, other will join a teams with someone who found it and finally we will have competition between strong models on top level. Previously Kaggle worked this way. </p>",
      "votes": 6,
      "replies": [
        {
          "id": 2805326,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2024-05-10T13:58:27.763000",
          "content": "<p>Unfortunately, one also needs to finetune the hack to maximize the score and hope that the chosen parameters will also be optimal for the hidden dataset. This introduces an element of randomness to the final scores, so the winner might not be the person with the best AUC, but rather the one who best randomly guessed how to optimally hack the metric. Even if it can be probed, the winner could be the one who discovered the hack earlier and therefore had more submissions left than others for probing.</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2805410,
              "author_name": "Evgeny Patekha",
              "author_url": "",
              "post_date": "2024-05-10T14:34:49.820000",
              "content": "<p>You are right that super tune of hack on LB could lead to overfit on private, so the quality of base model is still very important. <br>\nWe don't know how exactly test data was splitted, so shake up is possible. Will see it in three weeks. <br>\nI remember M5 competition, where hosts chose as private data period with high anomaly, compared with train and public data and very specific custom metric too. There was a huge shake up and I'm not sure that host got best solutions, but may be their goal was different. </p>\n<p>This time we also have specific metric which could be good for business but bad for competitions.  May be next competitions will avoid the same problem.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2805427,
              "author_name": "Evgeny Patekha",
              "author_url": "",
              "post_date": "2024-05-10T14:42:46.237000",
              "content": "<p>and of course you have two submission to choose for private scoring, one could be with aggressive hack and another with light correction or without it at all</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2805725,
              "author_name": "Carl McBride Ellis",
              "author_url": "",
              "post_date": "2024-05-10T18:04:04.963000",
              "content": "<p>For anyone who doesn't remember (or doesn't want to remember):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2F442140b962675a0bc1e7b009ce63c74c%2FM5_shakeup_plot.png?generation=1715364197906310&amp;alt=media\"><br>\n(Source: <a href=\"https://www.kaggle.com/code/carlmcbrideellis/shakeup-scatterplots-boxes-strings-and-things\" target=\"_blank\">\"<em>Shakeup scatterplots: Boxes, strings and things…</em>\"</a>)</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2806100,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-05-10T21:45:18.740000",
              "content": "<p>The problem here is that the variance in the score related to the hack can be similar to or greater than the variance in the score related to the quality of the model. In the previous Home Credit competition, the top solution scores ranged from 0.6 to 0.61, whereas the impact of the hack here can be up to 0.4 (?). And it's unclear how this impact might change on the private leaderboard.</p>\n<p>But I'm not suggesting any solutions; just expressing concern about an upcoming shakeup:) I actually think that the choice of the metric makes this competition more reflective of real life, as we often must work with strange KPIs that we need to optimize blindly.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2806166,
              "author_name": "ejwlk",
              "author_url": "",
              "post_date": "2024-05-10T23:49:07.217000",
              "content": "<p><a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> Yeah, I agree. Optimal recovery in a blackbox is a little weird, but it is still a stimulating challenge.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2812669,
              "author_name": "May FY",
              "author_url": "",
              "post_date": "2024-05-14T10:52:04.407000",
              "content": "<p>Maybe! Just  so</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2805409,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-10T14:34:48.730000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2804900,
      "author_name": "Laurent Pourchot",
      "author_url": "",
      "post_date": "2024-05-10T08:55:37.670000",
      "content": "<p>I had a good model without hacking and now it falls in the bottom of the PLB 🤣<br>\nThe ironic of the situation is the organizer who looked for getting stable models and after hiding the dates and now changing the rules again, he will get the best hacked metrics,and not necessarily with good models because the hacking is so efficient… I wonder if his company wanted to spend money for that ? 🤔</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2804939,
          "author_name": "Ern711",
          "author_url": "",
          "post_date": "2024-05-10T09:09:12.203000",
          "content": "<p>I guess even without intentionally hacking it, that metric will still lead to models that happens to be worse (worse than they could be) on the initial test period. I wonder if that's really desirable? 🤷‍♂️</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2805207,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-05-10T12:39:49.613000",
          "content": "<p>My guess is they will still get a lot of value if they read Top 100 -200 solutions - just because of the tons of work community did there.<br>\nBut - the value of a solution from the host perspective may be only loosely correlated with its LB score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2809540,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-05-12T20:20:57.843000",
          "content": "<p>I see you caught up <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> ! :) </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2809959,
              "author_name": "Laurent Pourchot",
              "author_url": "",
              "post_date": "2024-05-13T04:47:42.740000",
              "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> : Good model + metric hacking 🤣 </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2813706,
              "author_name": "Harry_Chan123",
              "author_url": "",
              "post_date": "2024-05-15T01:58:15.547000",
              "content": "<p>Have you ever use model stacking or something? Or spent more time on the feature engineering part? I found myself with no improvement for a relatively long period of time since I'm really just a beginner without ideas. Could you give me some advice to start again? 🥹</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2804781,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2024-05-10T07:53:11.423000",
      "content": "<p>Yes, it actually happened after the host stated they are not going to disqualify those who hack the metric, contrary to what was said earlier. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2805197,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-05-10T12:35:54.363000",
          "content": "<p>Yes, I saw this…</p>\n<p>Well - what can I say - I had the intuition this would happen the moment when hosts declined to confirm that metrick hacks based on restored <code>WEEK_NUM</code> would be ineligible for Kaggle points.</p>\n<p>We have in Poland the Kaggle meetup group with the official name \"Leaks &amp; Shakeups\". Maybe we should change our name to \"Leaks, Shakeups and Rule Changes\" 😉</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2804930,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-05-10T09:04:42.130000",
      "content": "<p>It seems that everyone is able to get improvement from hacking the metric. Maybe the way to restore WEEK_NUM is posted somewhere I do not have access to? 😯</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2805014,
          "author_name": "Rafał Pawłowski",
          "author_url": "",
          "post_date": "2024-05-10T10:13:20.250000",
          "content": "<p>No, this is top secret. I can't restore the weeks too. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2805831,
          "author_name": "Taizhuo Tang",
          "author_url": "",
          "post_date": "2024-05-10T18:53:11.477000",
          "content": "<p><a href=\"https://www.kaggle.com/yuanzhezhou\" target=\"_blank\">@yuanzhezhou</a> feel free not to answer: is it true that you've reached your current public LB position without any WEEK_NUM / metric hacking?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2806046,
          "author_name": "Rob Freeman",
          "author_url": "",
          "post_date": "2024-05-10T20:39:08.977000",
          "content": "<p>Actually, I am finding it difficult to get any benefit with my models. If you look at the training data, it is possible to recover WEEK_NUM from date_decision. However, I’m not sure that the aggregation tracks in the test set. Another option is perhaps MONTH which is less precise but surely tracks the aggregation properly. Not sure if this is how everyone else is doing it or not.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2805001,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-10T10:01:35.753000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2805066,
          "author_name": "Ern711",
          "author_url": "",
          "post_date": "2024-05-10T10:45:29.843000",
          "content": "<p>I'm afraid it will work on private lb too, but let's hope I'm mistaken. Anyway, that sounds like a good idea. </p>\n<p>If I will do anything more on this I will try creating one \"good\" solution (meaning aiming for an ensemble with high AUC in general (like high lowest AUC), without reducing any performance on purpose), and one \"hacking\" solution more out of curiosity. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2806238,
      "author_name": "Mark Hu",
      "author_url": "",
      "post_date": "2024-05-11T01:24:33.553000",
      "content": "<p>thanks for disccusions</p>",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2804775": "We have seen for a long time a stagnation of scores around 0.600 and now we have an explosion of scores between .620- .650. Especially considering that:\n- it happened after [this](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167) great post from @johnpateha\n- we know that metric hacking based on `WEEK_NUM` [leads to great improvements in scores](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449)\n\nI am very curious if this is indeed related to restoring `WEEK_NUM` and thus my prophecy is being fulfilled that I posted as an immediate reaction to the organizers' way of handling the metric hacking problem - that once you let the Genie out of the bottle, there is no coming back.",
    "2805290": "yes, explosion of scores definitely related to the metric hack, after clarification that it not breaks Kaggle rules. For me it gave 5+% up on LB\nIt's sad, but similar situations were many times previously here, so nothing new. \n\nMetric's hack has some limit, so big chance that strong models will win anyway.\nSome participants will find it by themselves, other will join a teams with someone who found it and finally we will have competition between strong models on top level. Previously Kaggle worked this way. \n",
    "2804900": "I had a good model without hacking and now it falls in the bottom of the PLB 🤣\nThe ironic of the situation is the organizer who looked for getting stable models and after hiding the dates and now changing the rules again, he will get the best hacked metrics,and not necessarily with good models because the hacking is so efficient... I wonder if his company wanted to spend money for that ? 🤔",
    "2804781": "Yes, it actually happened after the host stated they are not going to disqualify those who hack the metric, contrary to what was said earlier. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497337#2796412",
    "2804930": "It seems that everyone is able to get improvement from hacking the metric. Maybe the way to restore WEEK_NUM is posted somewhere I do not have access to? 😯",
    "2805001": "",
    "2806238": "thanks for disccusions"
  }
}