{
  "id": 496898,
  "title": "How to cheat, I mean \"improve\" your score",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/496898",
  "author_name": "NxGTR",
  "post_date": "2024-04-22T21:29:09.018000",
  "votes": 49,
  "comment_count": 28,
  "views": 0,
  "content": "<p>If you are looking for , improvement ideas, I came across to <a href=\"https://www.kaggle.com/code/carloshuertas/cheatingtime/\" target=\"_blank\">this</a> code that I made a private copy in case its gone.</p>\n<p>I am way too lazy to spend time on this, and I could be wrong here, but when I see stuff like:</p>\n<pre><code> i, most_similar_row  (most_similar_indices):\n        df2.at[most_similar_row, ] -= \n        df2.at[most_similar_row, ] = \n</code></pre>\n<p>mm… I smell foul play. As you should already know, the metric is very <em>hackable</em>, to the point that hacking the metric is far better than trying to do something clever, if your goal is to score high.</p>\n<p>The metric was not changed, but without knowing WEEK_NUM the problem just turns into, let me guess the WEEK_NUM, then back to original hacking.</p>\n<p>I hope to be wrong and this to be some sort of magic, either way, you have it!.</p>",
  "messages": [
    {
      "id": 2768481,
      "postDate": "2024-04-22T21:29:09.020Z",
      "content": "<p>If you are looking for , improvement ideas, I came across to <a href=\"https://www.kaggle.com/code/carloshuertas/cheatingtime/\" target=\"_blank\">this</a> code that I made a private copy in case its gone.</p>\n<p>I am way too lazy to spend time on this, and I could be wrong here, but when I see stuff like:</p>\n<pre><code> i, most_similar_row  (most_similar_indices):\n        df2.at[most_similar_row, ] -= \n        df2.at[most_similar_row, ] = \n</code></pre>\n<p>mm… I smell foul play. As you should already know, the metric is very <em>hackable</em>, to the point that hacking the metric is far better than trying to do something clever, if your goal is to score high.</p>\n<p>The metric was not changed, but without knowing WEEK_NUM the problem just turns into, let me guess the WEEK_NUM, then back to original hacking.</p>\n<p>I hope to be wrong and this to be some sort of magic, either way, you have it!.</p>",
      "rawMarkdown": "If you are looking for ~~cheating~~, improvement ideas, I came across to [this](https://www.kaggle.com/code/carloshuertas/cheatingtime/) code that I made a private copy in case its gone.\n\nI am way too lazy to spend time on this, and I could be wrong here, but when I see stuff like:\n\n```python\nfor i, most_similar_row in enumerate(most_similar_indices):\n        df2.at[most_similar_row, 'score'] -= 0.04\n        df2.at[most_similar_row, 'matched'] = True\n```\n\nmm... I smell foul play. As you should already know, the metric is very *hackable*, to the point that hacking the metric is far better than trying to do something clever, if your goal is to score high.\n\nThe metric was not changed, but without knowing WEEK_NUM the problem just turns into, let me guess the WEEK_NUM, then back to original hacking.\n\nI hope to be wrong and this to be some sort of magic, either way, you have it!.",
      "votes": 49
    },
    {
      "id": 2769207,
      "postDate": "2024-04-23T08:02:58.580Z",
      "content": "<p>as far as i understood, any submissions, which manually alter some fine tuned row selection, will be removed from the private leaderboard; so this is not a real problem.  sadly though, the public lb will become useless if this works. i haven't tried hacking anything myself, but something clearly works or scores close to 0.64 wouldn't be possible.</p>\n<p>anyway, i won't be wasting any time on this and just trust the hosts that they will be able to detect hacking like that without problems and remove them from the final leaderboard.</p>",
      "rawMarkdown": "as far as i understood, any submissions, which manually alter some fine tuned row selection, will be removed from the private leaderboard; so this is not a real problem.  sadly though, the public lb will become useless if this works. i haven't tried hacking anything myself, but something clearly works or scores close to 0.64 wouldn't be possible.\n\nanyway, i won't be wasting any time on this and just trust the hosts that they will be able to detect hacking like that without problems and remove them from the final leaderboard.\n",
      "votes": 10,
      "replies": [
        {
          "id": 2769552,
          "postDate": "2024-04-23T12:14:44.577Z",
          "content": "<p>I am not planning to hack the metrix eather. Measure, detect and tackle datadrift is an interesting subject and it is wrose to be analyed</p>",
          "rawMarkdown": "I am not planning to hack the metrix eather. Measure, detect and tackle datadrift is an interesting subject and it is wrose to be analyed"
        },
        {
          "id": 2769628,
          "postDate": "2024-04-23T13:21:37.873Z",
          "content": "<p>I hope the competition hosts are going to manually check the submissions and exclude the cheating solutions from the leaderboards. Otherwise it would be a huge disadvantage for those who do not cheat and huge waste of time and effort. </p>",
          "rawMarkdown": "I hope the competition hosts are going to manually check the submissions and exclude the cheating solutions from the leaderboards. Otherwise it would be a huge disadvantage for those who do not cheat and huge waste of time and effort. ",
          "votes": 1,
          "replies": [
            {
              "id": 2769668,
              "postDate": "2024-04-23T13:39:49.050Z",
              "content": "<p>Firstly, manually checking the code requires a lot of time, and at the same time, some tricks may not necessarily be good on the public list or the private list.</p>",
              "rawMarkdown": "Firstly, manually checking the code requires a lot of time, and at the same time, some tricks may not necessarily be good on the public list or the private list.",
              "votes": 1
            },
            {
              "id": 2769700,
              "postDate": "2024-04-23T14:04:29.497Z",
              "content": "<p>Yes. Checking every single submission would take a lot of time. But checking the top submissions on the private leaderboard after the deadline is manageable.</p>",
              "rawMarkdown": "Yes. Checking every single submission would take a lot of time. But checking the top submissions on the private leaderboard after the deadline is manageable."
            },
            {
              "id": 2770014,
              "postDate": "2024-04-23T16:29:41.227Z",
              "content": "<p>That only works if a few people cheat, if someone publish a notebook doing it, a lot of people will just copy &amp; submit, that's what happened the last time, the top 200 on the plb were  using performance enhancing tricks.</p>",
              "rawMarkdown": "That only works if a few people cheat, if someone publish a notebook doing it, a lot of people will just copy & submit, that's what happened the last time, the top 200 on the plb were ~~cheating~~ using performance enhancing tricks.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2774034,
          "postDate": "2024-04-25T03:22:20.017Z",
          "content": "<p>I don't think the hosts' track record so far should inspire much trust. </p>\n<p>The problem with manual review is that there a thousand ways to disguise the metric hack in post-processing that appears legitimate or would be legitimate if it were done for different reasons. You essentially have to rely on the evaluators being mind readers with this as your criteria. It makes the whole evaluation subjective which is ridiculous for a competition that so easily could have been objectively evaluated under a different metric. Kaggle really ought to have some standard here and not allow hackable metrics or manual evaluation on competitions that give medal and points.</p>",
          "rawMarkdown": "I don't think the hosts' track record so far should inspire much trust. \n\nThe problem with manual review is that there a thousand ways to disguise the metric hack in post-processing that appears legitimate or would be legitimate if it were done for different reasons. You essentially have to rely on the evaluators being mind readers with this as your criteria. It makes the whole evaluation subjective which is ridiculous for a competition that so easily could have been objectively evaluated under a different metric. Kaggle really ought to have some standard here and not allow hackable metrics or manual evaluation on competitions that give medal and points.",
          "votes": 6
        },
        {
          "id": 2774353,
          "postDate": "2024-04-25T06:50:51.667Z",
          "content": "<p>What will they do in case that all the processing and training is made externally and on kaggle  is loaded just the pretrained model for prediction? In this case they can't verify anything even manually, or I'm missing something?</p>",
          "rawMarkdown": "What will they do in case that all the processing and training is made externally and on kaggle  is loaded just the pretrained model for prediction? In this case they can't verify anything even manually, or I'm missing something?",
          "replies": [
            {
              "id": 2774602,
              "postDate": "2024-04-25T09:05:04.063Z",
              "content": "<p>i think there are only two ways hacks are possible</p>\n<ol>\n<li>changing rows early in the timeseries of the predictions by adding noise</li>\n<li>changing rows early in the timeseries of the test dataframes by adding noise</li>\n</ol>\n<p>both can be detected manually by looking at the submission, it would probably just take quite a while because it can take many different forms, but i think these are the only two ways; i can't think of any way that changing the train dataset could result in hacking.</p>\n<p>my understanding is that the only thing the hosts don't want is that you deliberatly make your predictions worse based on some knowledge of the date of the prediction in the test dataframe.</p>\n<p>if you just use what's given to you, without altering anything the test dataframes or your predictions by adding noise to some selection of rows, you're good.</p>",
              "rawMarkdown": "i think there are only two ways hacks are possible\n\n1. changing rows early in the timeseries of the predictions by adding noise\n2. changing rows early in the timeseries of the test dataframes by adding noise\n\nboth can be detected manually by looking at the submission, it would probably just take quite a while because it can take many different forms, but i think these are the only two ways; i can't think of any way that changing the train dataset could result in hacking.\n\nmy understanding is that the only thing the hosts don't want is that you deliberatly make your predictions worse based on some knowledge of the date of the prediction in the test dataframe.\n\nif you just use what's given to you, without altering anything the test dataframes or your predictions by adding noise to some selection of rows, you're good.",
              "votes": 2
            },
            {
              "id": 2774681,
              "postDate": "2024-04-25T09:33:48.313Z",
              "content": "<p>That makes sens - they can check how the model is applied on the test set and doesn't matter what are you doing during training, thanks for clarity</p>",
              "rawMarkdown": "That makes sens - they can check how the model is applied on the test set and doesn't matter what are you doing during training, thanks for clarity"
            },
            {
              "id": 2774758,
              "postDate": "2024-04-25T10:12:32.923Z",
              "content": "<p>You can use a weaker model for the first rows of the test data. Then, how can one prove it was a deliberate worsening of predictions? You might argue there was no intention to do so; perhaps you believed the first rows of test data were from the COVID period and decided to use a different model for that period because it differs significantly from others.</p>\n<p>In other words, the rules should be transparent, clear, and explicit. Perhaps the difficulty in formulating this rule was the reason it hasn't been introduced yet, even though the problem was revealed more than a month ago and it was clear back then that hacking the metric is still possible.</p>",
              "rawMarkdown": "You can use a weaker model for the first rows of the test data. Then, how can one prove it was a deliberate worsening of predictions? You might argue there was no intention to do so; perhaps you believed the first rows of test data were from the COVID period and decided to use a different model for that period because it differs significantly from others.\n\nIn other words, the rules should be transparent, clear, and explicit. Perhaps the difficulty in formulating this rule was the reason it hasn't been introduced yet, even though the problem was revealed more than a month ago and it was clear back then that hacking the metric is still possible.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2768526,
      "postDate": "2024-04-22T22:46:48.070Z",
      "content": "<p>That was expected, and that's why the host said they are going to manually check top solutions. However, they still haven't told us whether the cheaters will get medals…</p>\n<p>The main problem here is that it's possible to make the cheating less obvious. Instead of directly changing predictions, one could use a poorer model, making it appear like a legitimate modeling decision. </p>",
      "rawMarkdown": "That was expected, and that's why the host said they are going to manually check top solutions. However, they still haven't told us whether the cheaters will get medals...\n\nThe main problem here is that it's possible to make the cheating less obvious. Instead of directly changing predictions, one could use a poorer model, making it appear like a legitimate modeling decision. ",
      "votes": 7,
      "replies": [
        {
          "id": 2774361,
          "postDate": "2024-04-25T06:54:30.097Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2787515,
          "postDate": "2024-05-01T19:08:50.733Z",
          "content": "<p>How do you define \"top solutions\"? is it just gold, or silver and bronze also?</p>",
          "rawMarkdown": "How do you define \"top solutions\"? is it just gold, or silver and bronze also?",
          "replies": [
            {
              "id": 2787556,
              "postDate": "2024-05-01T19:44:58.970Z",
              "content": "<p>I think they've mentioned top-100 solutions</p>",
              "rawMarkdown": "I think they've mentioned top-100 solutions",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2769650,
      "postDate": "2024-04-23T13:33:15.590Z",
      "content": "<p>Does this work in LB?</p>",
      "rawMarkdown": "Does this work in LB?",
      "votes": 1
    },
    {
      "id": 2769819,
      "postDate": "2024-04-23T15:05:21.323Z",
      "content": "<p>One other concerning part is how to detect if people are legitimately doing post-processing or hacking the metric. Post-processing has been done in every other competition to boost your final rank a little further (sometimes even more). I know competitions where post-processing put you into the gold zone from the bronze zone (i.e. <a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection\" target=\"_blank\">https://www.kaggle.com/competitions/rfcx-species-audio-detection</a>) solution by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220389</a>. </p>",
      "rawMarkdown": "One other concerning part is how to detect if people are legitimately doing post-processing or hacking the metric. Post-processing has been done in every other competition to boost your final rank a little further (sometimes even more). I know competitions where post-processing put you into the gold zone from the bronze zone (i.e. https://www.kaggle.com/competitions/rfcx-species-audio-detection) solution by @cdeotte https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220389. ",
      "votes": 2
    },
    {
      "id": 2768780,
      "postDate": "2024-04-23T03:10:42.980Z",
      "content": "<p>Does this work in LB? I believe the host should simply use AUC rather than a metric that can be hacked. (Extension?) They will not get what they want with current metric.</p>",
      "rawMarkdown": "Does this work in LB? I believe the host should simply use AUC rather than a metric that can be hacked. (Extension?) They will not get what they want with current metric.",
      "votes": 2,
      "replies": [
        {
          "id": 2769347,
          "postDate": "2024-04-23T09:46:51.393Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2770996,
          "postDate": "2024-04-24T05:14:37.497Z",
          "content": "<p>at beginning of competition multiple people pointed out that the metric is very hackable and it will lead to problems.<br>\nthe host didn't want to change it.</p>",
          "rawMarkdown": "at beginning of competition multiple people pointed out that the metric is very hackable and it will lead to problems.\nthe host didn't want to change it.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2772400,
      "postDate": "2024-04-24T17:48:17.397Z",
      "content": "<p>What about include some train samples around 09-2020 in test and these samples would not be scored, will this solve the problem? </p>",
      "rawMarkdown": "What about include some train samples around 09-2020 in test and these samples would not be scored, will this solve the problem? "
    },
    {
      "id": 2768564,
      "postDate": "2024-04-22T23:52:32.583Z",
      "content": "<p>Is the public and private ranking of test data divided by time? Can this method be used for private ranking?</p>",
      "rawMarkdown": "Is the public and private ranking of test data divided by time? Can this method be used for private ranking?"
    },
    {
      "id": 2768495,
      "postDate": "2024-04-22T22:07:14.097Z",
      "content": "<p>I wonder if this \"trick\" really helped to improve someone's score, because similar data doesn't automatically mean closer <code>WEEK_NUM</code>. However, I agree that the decision to keep the same metric by transforming the data indeed doesn't look like a robust solution.</p>",
      "rawMarkdown": "I wonder if this \"trick\" really helped to improve someone's score, because similar data doesn't automatically mean closer `WEEK_NUM`. However, I agree that the decision to keep the same metric by transforming the data indeed doesn't look like a robust solution.",
      "replies": [
        {
          "id": 2768497,
          "postDate": "2024-04-22T22:10:11.787Z",
          "content": "<p>Maybe… if I find the similar weeks right after train-end?, the rationale, if there is distribution shift, its more likely to happen over the long-term, so, whatever is similar to the end-of-train, has potential to be the next WEEK_NUM, remember, I dont need to guess them all, all I need is to guess the first ones.</p>",
          "rawMarkdown": "Maybe... if I find the similar weeks right after train-end?, the rationale, if there is distribution shift, its more likely to happen over the long-term, so, whatever is similar to the end-of-train, has potential to be the next WEEK_NUM, remember, I dont need to guess them all, all I need is to guess the first ones.",
          "replies": [
            {
              "id": 2768503,
              "postDate": "2024-04-22T22:24:01.503Z",
              "content": "<p>Right, but I'm not sure that end-of-train and start-of-test are going to be similar. From what hosts <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764\" target=\"_blank\">said</a> test data also contain some pretty old samples.</p>",
              "rawMarkdown": "Right, but I'm not sure that end-of-train and start-of-test are going to be similar. From what hosts [said](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764) test data also contain some pretty old samples.",
              "votes": 1
            },
            {
              "id": 2768529,
              "postDate": "2024-04-22T22:56:49.667Z",
              "content": "<p>I'm not sure they meant that the test sample contains data before train-end. It seems that by 'old data,' he meant the covid period. </p>",
              "rawMarkdown": "I'm not sure they meant that the test sample contains data before train-end. It seems that by 'old data,' he meant the covid period. \n",
              "votes": 2
            },
            {
              "id": 2769267,
              "postDate": "2024-04-23T08:51:07.657Z",
              "content": "<p>It could probably be easily checked with a few submissions, looking at the current LB status I feel like this trick does work.</p>",
              "rawMarkdown": "It could probably be easily checked with a few submissions, looking at the current LB status I feel like this trick does work.",
              "votes": 1
            },
            {
              "id": 2773963,
              "postDate": "2024-04-25T02:01:14.447Z",
              "content": "<blockquote>\n  <p>In the test sample, WEEK_NUM continues sequentially from the last training value of WEEK_NUM.</p>\n</blockquote>\n<p>I found this in the Data description. Not sure whether they have changed the test data.</p>",
              "rawMarkdown": ">In the test sample, WEEK_NUM continues sequentially from the last training value of WEEK_NUM.\n\nI found this in the Data description. Not sure whether they have changed the test data.",
              "votes": 2
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2769207,
      "author_name": "at7459",
      "author_url": "",
      "post_date": "2024-04-23T08:02:58.580000",
      "content": "<p>as far as i understood, any submissions, which manually alter some fine tuned row selection, will be removed from the private leaderboard; so this is not a real problem.  sadly though, the public lb will become useless if this works. i haven't tried hacking anything myself, but something clearly works or scores close to 0.64 wouldn't be possible.</p>\n<p>anyway, i won't be wasting any time on this and just trust the hosts that they will be able to detect hacking like that without problems and remove them from the final leaderboard.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 2769552,
          "author_name": "LucaMTB",
          "author_url": "",
          "post_date": "2024-04-23T12:14:44.577000",
          "content": "<p>I am not planning to hack the metrix eather. Measure, detect and tackle datadrift is an interesting subject and it is wrose to be analyed</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2769628,
          "author_name": "Matous Famera",
          "author_url": "",
          "post_date": "2024-04-23T13:21:37.873000",
          "content": "<p>I hope the competition hosts are going to manually check the submissions and exclude the cheating solutions from the leaderboards. Otherwise it would be a huge disadvantage for those who do not cheat and huge waste of time and effort. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2769668,
              "author_name": "yunsuxiaozi",
              "author_url": "",
              "post_date": "2024-04-23T13:39:49.050000",
              "content": "<p>Firstly, manually checking the code requires a lot of time, and at the same time, some tricks may not necessarily be good on the public list or the private list.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2769700,
              "author_name": "Matous Famera",
              "author_url": "",
              "post_date": "2024-04-23T14:04:29.497000",
              "content": "<p>Yes. Checking every single submission would take a lot of time. But checking the top submissions on the private leaderboard after the deadline is manageable.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2770014,
              "author_name": "Antonio Félix",
              "author_url": "",
              "post_date": "2024-04-23T16:29:41.227000",
              "content": "<p>That only works if a few people cheat, if someone publish a notebook doing it, a lot of people will just copy &amp; submit, that's what happened the last time, the top 200 on the plb were  using performance enhancing tricks.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2774034,
          "author_name": "Jacoby Jaeger",
          "author_url": "",
          "post_date": "2024-04-25T03:22:20.017000",
          "content": "<p>I don't think the hosts' track record so far should inspire much trust. </p>\n<p>The problem with manual review is that there a thousand ways to disguise the metric hack in post-processing that appears legitimate or would be legitimate if it were done for different reasons. You essentially have to rely on the evaluators being mind readers with this as your criteria. It makes the whole evaluation subjective which is ridiculous for a competition that so easily could have been objectively evaluated under a different metric. Kaggle really ought to have some standard here and not allow hackable metrics or manual evaluation on competitions that give medal and points.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2774353,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-04-25T06:50:51.667000",
          "content": "<p>What will they do in case that all the processing and training is made externally and on kaggle  is loaded just the pretrained model for prediction? In this case they can't verify anything even manually, or I'm missing something?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2774602,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-04-25T09:05:04.063000",
              "content": "<p>i think there are only two ways hacks are possible</p>\n<ol>\n<li>changing rows early in the timeseries of the predictions by adding noise</li>\n<li>changing rows early in the timeseries of the test dataframes by adding noise</li>\n</ol>\n<p>both can be detected manually by looking at the submission, it would probably just take quite a while because it can take many different forms, but i think these are the only two ways; i can't think of any way that changing the train dataset could result in hacking.</p>\n<p>my understanding is that the only thing the hosts don't want is that you deliberatly make your predictions worse based on some knowledge of the date of the prediction in the test dataframe.</p>\n<p>if you just use what's given to you, without altering anything the test dataframes or your predictions by adding noise to some selection of rows, you're good.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2774681,
              "author_name": "Danu A.",
              "author_url": "",
              "post_date": "2024-04-25T09:33:48.313000",
              "content": "<p>That makes sens - they can check how the model is applied on the test set and doesn't matter what are you doing during training, thanks for clarity</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2774758,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-04-25T10:12:32.923000",
              "content": "<p>You can use a weaker model for the first rows of the test data. Then, how can one prove it was a deliberate worsening of predictions? You might argue there was no intention to do so; perhaps you believed the first rows of test data were from the COVID period and decided to use a different model for that period because it differs significantly from others.</p>\n<p>In other words, the rules should be transparent, clear, and explicit. Perhaps the difficulty in formulating this rule was the reason it hasn't been introduced yet, even though the problem was revealed more than a month ago and it was clear back then that hacking the metric is still possible.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2768526,
      "author_name": "Evgeniia Grigoreva",
      "author_url": "",
      "post_date": "2024-04-22T22:46:48.070000",
      "content": "<p>That was expected, and that's why the host said they are going to manually check top solutions. However, they still haven't told us whether the cheaters will get medals…</p>\n<p>The main problem here is that it's possible to make the cheating less obvious. Instead of directly changing predictions, one could use a poorer model, making it appear like a legitimate modeling decision. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 2774361,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-25T06:54:30.097000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2787515,
          "author_name": "Eduard Stefanescu",
          "author_url": "",
          "post_date": "2024-05-01T19:08:50.733000",
          "content": "<p>How do you define \"top solutions\"? is it just gold, or silver and bronze also?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2787556,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-05-01T19:44:58.970000",
              "content": "<p>I think they've mentioned top-100 solutions</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2769650,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-04-23T13:33:15.590000",
      "content": "<p>Does this work in LB?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2769819,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2024-04-23T15:05:21.323000",
      "content": "<p>One other concerning part is how to detect if people are legitimately doing post-processing or hacking the metric. Post-processing has been done in every other competition to boost your final rank a little further (sometimes even more). I know competitions where post-processing put you into the gold zone from the bronze zone (i.e. <a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection\" target=\"_blank\">https://www.kaggle.com/competitions/rfcx-species-audio-detection</a>) solution by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220389</a>. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2768780,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-04-23T03:10:42.980000",
      "content": "<p>Does this work in LB? I believe the host should simply use AUC rather than a metric that can be hacked. (Extension?) They will not get what they want with current metric.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2769347,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-04-23T09:46:51.393000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2770996,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2024-04-24T05:14:37.497000",
          "content": "<p>at beginning of competition multiple people pointed out that the metric is very hackable and it will lead to problems.<br>\nthe host didn't want to change it.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2772400,
      "author_name": "Evan",
      "author_url": "",
      "post_date": "2024-04-24T17:48:17.397000",
      "content": "<p>What about include some train samples around 09-2020 in test and these samples would not be scored, will this solve the problem? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2768564,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-04-22T23:52:32.583000",
      "content": "<p>Is the public and private ranking of test data divided by time? Can this method be used for private ranking?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2768495,
      "author_name": "Oleksiy Kononenko",
      "author_url": "",
      "post_date": "2024-04-22T22:07:14.097000",
      "content": "<p>I wonder if this \"trick\" really helped to improve someone's score, because similar data doesn't automatically mean closer <code>WEEK_NUM</code>. However, I agree that the decision to keep the same metric by transforming the data indeed doesn't look like a robust solution.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2768497,
          "author_name": "NxGTR",
          "author_url": "",
          "post_date": "2024-04-22T22:10:11.787000",
          "content": "<p>Maybe… if I find the similar weeks right after train-end?, the rationale, if there is distribution shift, its more likely to happen over the long-term, so, whatever is similar to the end-of-train, has potential to be the next WEEK_NUM, remember, I dont need to guess them all, all I need is to guess the first ones.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2768503,
              "author_name": "Oleksiy Kononenko",
              "author_url": "",
              "post_date": "2024-04-22T22:24:01.503000",
              "content": "<p>Right, but I'm not sure that end-of-train and start-of-test are going to be similar. From what hosts <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/473765#2638764\" target=\"_blank\">said</a> test data also contain some pretty old samples.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2768529,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-04-22T22:56:49.667000",
              "content": "<p>I'm not sure they meant that the test sample contains data before train-end. It seems that by 'old data,' he meant the covid period. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2769267,
              "author_name": "Oleksiy Kononenko",
              "author_url": "",
              "post_date": "2024-04-23T08:51:07.657000",
              "content": "<p>It could probably be easily checked with a few submissions, looking at the current LB status I feel like this trick does work.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2773963,
              "author_name": "Wisp Vale",
              "author_url": "",
              "post_date": "2024-04-25T02:01:14.447000",
              "content": "<blockquote>\n  <p>In the test sample, WEEK_NUM continues sequentially from the last training value of WEEK_NUM.</p>\n</blockquote>\n<p>I found this in the Data description. Not sure whether they have changed the test data.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2768481": "If you are looking for ~~cheating~~, improvement ideas, I came across to [this](https://www.kaggle.com/code/carloshuertas/cheatingtime/) code that I made a private copy in case its gone.\n\nI am way too lazy to spend time on this, and I could be wrong here, but when I see stuff like:\n\n```python\nfor i, most_similar_row in enumerate(most_similar_indices):\n        df2.at[most_similar_row, 'score'] -= 0.04\n        df2.at[most_similar_row, 'matched'] = True\n```\n\nmm... I smell foul play. As you should already know, the metric is very *hackable*, to the point that hacking the metric is far better than trying to do something clever, if your goal is to score high.\n\nThe metric was not changed, but without knowing WEEK_NUM the problem just turns into, let me guess the WEEK_NUM, then back to original hacking.\n\nI hope to be wrong and this to be some sort of magic, either way, you have it!.",
    "2769207": "as far as i understood, any submissions, which manually alter some fine tuned row selection, will be removed from the private leaderboard; so this is not a real problem.  sadly though, the public lb will become useless if this works. i haven't tried hacking anything myself, but something clearly works or scores close to 0.64 wouldn't be possible.\n\nanyway, i won't be wasting any time on this and just trust the hosts that they will be able to detect hacking like that without problems and remove them from the final leaderboard.\n",
    "2768526": "That was expected, and that's why the host said they are going to manually check top solutions. However, they still haven't told us whether the cheaters will get medals...\n\nThe main problem here is that it's possible to make the cheating less obvious. Instead of directly changing predictions, one could use a poorer model, making it appear like a legitimate modeling decision. ",
    "2769650": "Does this work in LB?",
    "2769819": "One other concerning part is how to detect if people are legitimately doing post-processing or hacking the metric. Post-processing has been done in every other competition to boost your final rank a little further (sometimes even more). I know competitions where post-processing put you into the gold zone from the bronze zone (i.e. https://www.kaggle.com/competitions/rfcx-species-audio-detection) solution by @cdeotte https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/220389. ",
    "2768780": "Does this work in LB? I believe the host should simply use AUC rather than a metric that can be hacked. (Extension?) They will not get what they want with current metric.",
    "2772400": "What about include some train samples around 09-2020 in test and these samples would not be scored, will this solve the problem? ",
    "2768564": "Is the public and private ranking of test data divided by time? Can this method be used for private ranking?",
    "2768495": "I wonder if this \"trick\" really helped to improve someone's score, because similar data doesn't automatically mean closer `WEEK_NUM`. However, I agree that the decision to keep the same metric by transforming the data indeed doesn't look like a robust solution."
  }
}