{
  "id": 501744,
  "title": "My team's viewpoint about metric hacking",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501744",
  "author_name": "minhtu.mt.mt",
  "post_date": "2024-05-10T14:52:13.559000",
  "votes": 32,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Our team successfully hacked the LB. As you can see, our LB scores before and after hacking are 0.599 and 0.652. <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> had reported the same amount of improvement, ~5%. We don't want to spend time optimizing the hack so the number can be larger.<br>\nHowever, in my opinion, this hacking method is unlikely to work in the Private dataset. I don't know whether my assumption is correct or not, but I really hope so.<br>\nLook on the bright side, after many trials, we still have good models which work well on both CV and LB. Our team members also decide to focus on improving the CV, not to continue to exploit the hack. Let's wait until the end before judging the competition. I think we all learnt a lot in this competition.<br>\nHappy kaggling!</p>",
  "messages": [
    {
      "id": 2805440,
      "postDate": "2024-05-10T14:52:13.560Z",
      "content": "<p>Our team successfully hacked the LB. As you can see, our LB scores before and after hacking are 0.599 and 0.652. <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> had reported the same amount of improvement, ~5%. We don't want to spend time optimizing the hack so the number can be larger.<br>\nHowever, in my opinion, this hacking method is unlikely to work in the Private dataset. I don't know whether my assumption is correct or not, but I really hope so.<br>\nLook on the bright side, after many trials, we still have good models which work well on both CV and LB. Our team members also decide to focus on improving the CV, not to continue to exploit the hack. Let's wait until the end before judging the competition. I think we all learnt a lot in this competition.<br>\nHappy kaggling!</p>",
      "rawMarkdown": "Our team successfully hacked the LB. As you can see, our LB scores before and after hacking are 0.599 and 0.652. @johnpateha had reported the same amount of improvement, ~5%. We don't want to spend time optimizing the hack so the number can be larger.\nHowever, in my opinion, this hacking method is unlikely to work in the Private dataset. I don't know whether my assumption is correct or not, but I really hope so.\nLook on the bright side, after many trials, we still have good models which work well on both CV and LB. Our team members also decide to focus on improving the CV, not to continue to exploit the hack. Let's wait until the end before judging the competition. I think we all learnt a lot in this competition.\nHappy kaggling!",
      "votes": 32
    },
    {
      "id": 2816001,
      "postDate": "2024-05-16T06:26:24.270Z",
      "content": "<p>\"How to fine tune your llm to make it look like it's NOT generated by llm\" </p>",
      "rawMarkdown": "\"How to fine tune your llm to make it look like it's NOT generated by llm\" ",
      "votes": 3
    },
    {
      "id": 2805711,
      "postDate": "2024-05-10T17:57:14.737Z",
      "content": "<p>I think it will work because they stopped the competition for quite a while after seeing this issue. They would just say, oh it doesn't matter because you will be overfitting to the public LB or something. And from the previous experiences, I think it's also pointing towards that.</p>\n<p>Of course I might be wrong as well but let's see. </p>",
      "rawMarkdown": "I think it will work because they stopped the competition for quite a while after seeing this issue. They would just say, oh it doesn't matter because you will be overfitting to the public LB or something. And from the previous experiences, I think it's also pointing towards that.\n\nOf course I might be wrong as well but let's see. ",
      "votes": 4,
      "replies": [
        {
          "id": 2812842,
          "postDate": "2024-05-14T12:56:19.907Z",
          "content": "<p>Agree. There seems to be a deep meaning behind their recent silence. Let's watch over it…</p>",
          "rawMarkdown": "Agree. There seems to be a deep meaning behind their recent silence. Let's watch over it..."
        },
        {
          "id": 2816277,
          "postDate": "2024-05-16T08:22:29.730Z",
          "content": "<p>I think so…</p>",
          "rawMarkdown": "I think so..."
        }
      ]
    },
    {
      "id": 2805497,
      "postDate": "2024-05-10T15:21:02.833Z",
      "content": "<p>I also share the opinion that all metric hacks apply only to the public test set, and for \"different\" data - aka the private test - may result to entirely false predictions. My two cents is that we should develop a model that does well - but not great on the public leaderboard - perhaps <code>stability</code> means NOT doing great on a specific set of data but decently on any set of data. </p>",
      "rawMarkdown": "I also share the opinion that all metric hacks apply only to the public test set, and for \"different\" data - aka the private test - may result to entirely false predictions. My two cents is that we should develop a model that does well - but not great on the public leaderboard - perhaps `stability` means NOT doing great on a specific set of data but decently on any set of data. ",
      "votes": 1,
      "replies": [
        {
          "id": 2809933,
          "postDate": "2024-05-13T04:23:29.030Z",
          "content": "<p>I also hope so..</p>",
          "rawMarkdown": "I also hope so.."
        }
      ]
    },
    {
      "id": 2805460,
      "postDate": "2024-05-10T15:02:07.363Z",
      "content": "<p>Why do you think the hack won't work on the private LB? I think it will, but maybe our assumptions differ</p>",
      "rawMarkdown": "Why do you think the hack won't work on the private LB? I think it will, but maybe our assumptions differ",
      "votes": 2,
      "replies": [
        {
          "id": 2805550,
          "postDate": "2024-05-10T15:46:40.693Z",
          "content": "<p>I agree that the hack is more likely to work on the private LB, but I also see some options/set-up that can efficiently help reducing the impact of the hack</p>",
          "rawMarkdown": "I agree that the hack is more likely to work on the private LB, but I also see some options/set-up that can efficiently help reducing the impact of the hack",
          "votes": 1,
          "replies": [
            {
              "id": 2807985,
              "postDate": "2024-05-12T03:17:12.153Z",
              "content": "<p>Hmm…just a curious. What kind of options/set-up could be made by the competition host on the private LB?</p>",
              "rawMarkdown": "Hmm...just a curious. What kind of options/set-up could be made by the competition host on the private LB?",
              "votes": 1
            },
            {
              "id": 2809074,
              "postDate": "2024-05-12T14:03:27.967Z",
              "content": "<p><a href=\"https://www.kaggle.com/ttkuma\" target=\"_blank\">@ttkuma</a> 1 option is that the private LB can begin from some later weeks. For example, we know the whole test data starts from week 92 (train ends in 91). If the host set the public to start from 92, but the private start from a further week (e.g. 130), the hack will only work in public LB</p>",
              "rawMarkdown": "@ttkuma 1 option is that the private LB can begin from some later weeks. For example, we know the whole test data starts from week 92 (train ends in 91). If the host set the public to start from 92, but the private start from a further week (e.g. 130), the hack will only work in public LB",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2805445,
      "postDate": "2024-05-10T14:55:15.597Z",
      "content": "<p>I believe that it's like 90% chance that the hack work in private LB from the past competitions.</p>",
      "rawMarkdown": "I believe that it's like 90% chance that the hack work in private LB from the past competitions.",
      "votes": 2,
      "replies": [
        {
          "id": 2805476,
          "postDate": "2024-05-10T15:06:50.233Z",
          "content": "<p>So let's pray for the remaining 10%</p>",
          "rawMarkdown": "So let's pray for the remaining 10%"
        }
      ]
    },
    {
      "id": 2806218,
      "postDate": "2024-05-11T01:09:37.073Z",
      "content": "<p>Congratulations to your team for successfully hacking!</p>",
      "rawMarkdown": "Congratulations to your team for successfully hacking!",
      "votes": -7
    },
    {
      "id": 2823257,
      "postDate": "2024-05-19T05:21:35.983Z",
      "content": "<p>For what its worth, my model ensemble that scores LB=0.592, score much lower i.e. LB=0.515 with the latest one liner trick😀.</p>",
      "rawMarkdown": "For what its worth, my model ensemble that scores LB=0.592, score much lower i.e. LB=0.515 with the latest one liner trick😀."
    },
    {
      "id": 2809842,
      "postDate": "2024-05-13T03:02:49.197Z",
      "content": "<p>good work!</p>",
      "rawMarkdown": "good work!"
    },
    {
      "id": 2812763,
      "postDate": "2024-05-14T12:05:23.610Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2816001,
      "author_name": "hongyi shui",
      "author_url": "",
      "post_date": "2024-05-16T06:26:24.270000",
      "content": "<p>\"How to fine tune your llm to make it look like it's NOT generated by llm\" </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2805711,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2024-05-10T17:57:14.737000",
      "content": "<p>I think it will work because they stopped the competition for quite a while after seeing this issue. They would just say, oh it doesn't matter because you will be overfitting to the public LB or something. And from the previous experiences, I think it's also pointing towards that.</p>\n<p>Of course I might be wrong as well but let's see. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2812842,
          "author_name": "mh",
          "author_url": "",
          "post_date": "2024-05-14T12:56:19.907000",
          "content": "<p>Agree. There seems to be a deep meaning behind their recent silence. Let's watch over it…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2816277,
          "author_name": "Stephen Wuzh",
          "author_url": "",
          "post_date": "2024-05-16T08:22:29.730000",
          "content": "<p>I think so…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2805497,
      "author_name": "Andreas Bisiadis",
      "author_url": "",
      "post_date": "2024-05-10T15:21:02.833000",
      "content": "<p>I also share the opinion that all metric hacks apply only to the public test set, and for \"different\" data - aka the private test - may result to entirely false predictions. My two cents is that we should develop a model that does well - but not great on the public leaderboard - perhaps <code>stability</code> means NOT doing great on a specific set of data but decently on any set of data. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2809933,
          "author_name": "ArcherPan",
          "author_url": "",
          "post_date": "2024-05-13T04:23:29.030000",
          "content": "<p>I also hope so..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2805460,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2024-05-10T15:02:07.363000",
      "content": "<p>Why do you think the hack won't work on the private LB? I think it will, but maybe our assumptions differ</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2805550,
          "author_name": "minhtu.mt.mt",
          "author_url": "",
          "post_date": "2024-05-10T15:46:40.693000",
          "content": "<p>I agree that the hack is more likely to work on the private LB, but I also see some options/set-up that can efficiently help reducing the impact of the hack</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2807985,
              "author_name": "Ttkuma",
              "author_url": "",
              "post_date": "2024-05-12T03:17:12.153000",
              "content": "<p>Hmm…just a curious. What kind of options/set-up could be made by the competition host on the private LB?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2809074,
              "author_name": "minhtu.mt.mt",
              "author_url": "",
              "post_date": "2024-05-12T14:03:27.967000",
              "content": "<p><a href=\"https://www.kaggle.com/ttkuma\" target=\"_blank\">@ttkuma</a> 1 option is that the private LB can begin from some later weeks. For example, we know the whole test data starts from week 92 (train ends in 91). If the host set the public to start from 92, but the private start from a further week (e.g. 130), the hack will only work in public LB</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2805445,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-05-10T14:55:15.597000",
      "content": "<p>I believe that it's like 90% chance that the hack work in private LB from the past competitions.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2805476,
          "author_name": "minhtu.mt.mt",
          "author_url": "",
          "post_date": "2024-05-10T15:06:50.233000",
          "content": "<p>So let's pray for the remaining 10%</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2806218,
      "author_name": "暗黑AGI",
      "author_url": "",
      "post_date": "2024-05-11T01:09:37.073000",
      "content": "<p>Congratulations to your team for successfully hacking!</p>",
      "votes": -7,
      "replies": []
    },
    {
      "id": 2823257,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2024-05-19T05:21:35.983000",
      "content": "<p>For what its worth, my model ensemble that scores LB=0.592, score much lower i.e. LB=0.515 with the latest one liner trick😀.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2809842,
      "author_name": "Cheng Zhao",
      "author_url": "",
      "post_date": "2024-05-13T03:02:49.197000",
      "content": "<p>good work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2812763,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-14T12:05:23.610000",
      "content": "",
      "votes": -2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2805440": "Our team successfully hacked the LB. As you can see, our LB scores before and after hacking are 0.599 and 0.652. @johnpateha had reported the same amount of improvement, ~5%. We don't want to spend time optimizing the hack so the number can be larger.\nHowever, in my opinion, this hacking method is unlikely to work in the Private dataset. I don't know whether my assumption is correct or not, but I really hope so.\nLook on the bright side, after many trials, we still have good models which work well on both CV and LB. Our team members also decide to focus on improving the CV, not to continue to exploit the hack. Let's wait until the end before judging the competition. I think we all learnt a lot in this competition.\nHappy kaggling!",
    "2816001": "\"How to fine tune your llm to make it look like it's NOT generated by llm\" ",
    "2805711": "I think it will work because they stopped the competition for quite a while after seeing this issue. They would just say, oh it doesn't matter because you will be overfitting to the public LB or something. And from the previous experiences, I think it's also pointing towards that.\n\nOf course I might be wrong as well but let's see. ",
    "2805497": "I also share the opinion that all metric hacks apply only to the public test set, and for \"different\" data - aka the private test - may result to entirely false predictions. My two cents is that we should develop a model that does well - but not great on the public leaderboard - perhaps `stability` means NOT doing great on a specific set of data but decently on any set of data. ",
    "2805460": "Why do you think the hack won't work on the private LB? I think it will, but maybe our assumptions differ",
    "2805445": "I believe that it's like 90% chance that the hack work in private LB from the past competitions.",
    "2806218": "Congratulations to your team for successfully hacking!",
    "2823257": "For what its worth, my model ensemble that scores LB=0.592, score much lower i.e. LB=0.515 with the latest one liner trick😀.",
    "2809842": "good work!",
    "2812763": ""
  }
}