{
  "id": 497167,
  "title": " Metric hack again, sorry",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/497167",
  "author_name": "Evgeny Patekha",
  "post_date": "2024-04-23T22:14:34.859000",
  "votes": 73,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Today, looking at the leaderboard, I decided to test one idea and my submission showed that the dates in the test could be accurately restored to initial values. This discovery allows for a metric hack that boosts the leaderboard score by several percentage points.</p>\n<p>I re-read the rules and found no explicit prohibition of these actions. I've seen organizer's responses indicating they wouldn't be happy about metric hacking, but I found no restrictions that would lead to disqualification of solutions. Maybe I didn’t search thoroughly, but clear guidelines would be appreciated.</p>\n<p>This is Kaggle after all, and ultimately, it's the rankings that count, not how much you adhere to fair play. If a clean result doesn't top the leaderboard because others used a hack, then ultimately, it doesn't mean anything. Sadly, that's life :(</p>\n<p>After a few years' break, I find some time to Kaggle again, and I would like to achieve good results that are practically useful, but going against reality seems foolish.</p>\n<p>It would be great if rule-based strict prohibitions against metric hacks were implemented, making such methods lead to disqualification. However, there might be complexities that have prevented this so far.</p>\n<p>Another option is to fix the leak, which could be done with minimal impact on the \"pure\" model. If the organizers are interested in the nature of the leak, I'm ready to explain it privately. Or they could check my last submission—it’s all quite transparent there.</p>\n<p>Otherwise, things will stay as they are, and we will compete to see who can exploit this hack in addition to having a strong model. I believe many will be able to find this leak.</p>\n<p>I think this news might disappoint many, but it seemed right to share it while there’s still a month ahead to make corrections.</p>",
  "messages": [
    {
      "id": 2770542,
      "postDate": "2024-04-23T22:14:34.860Z",
      "content": "<p>Today, looking at the leaderboard, I decided to test one idea and my submission showed that the dates in the test could be accurately restored to initial values. This discovery allows for a metric hack that boosts the leaderboard score by several percentage points.</p>\n<p>I re-read the rules and found no explicit prohibition of these actions. I've seen organizer's responses indicating they wouldn't be happy about metric hacking, but I found no restrictions that would lead to disqualification of solutions. Maybe I didn’t search thoroughly, but clear guidelines would be appreciated.</p>\n<p>This is Kaggle after all, and ultimately, it's the rankings that count, not how much you adhere to fair play. If a clean result doesn't top the leaderboard because others used a hack, then ultimately, it doesn't mean anything. Sadly, that's life :(</p>\n<p>After a few years' break, I find some time to Kaggle again, and I would like to achieve good results that are practically useful, but going against reality seems foolish.</p>\n<p>It would be great if rule-based strict prohibitions against metric hacks were implemented, making such methods lead to disqualification. However, there might be complexities that have prevented this so far.</p>\n<p>Another option is to fix the leak, which could be done with minimal impact on the \"pure\" model. If the organizers are interested in the nature of the leak, I'm ready to explain it privately. Or they could check my last submission—it’s all quite transparent there.</p>\n<p>Otherwise, things will stay as they are, and we will compete to see who can exploit this hack in addition to having a strong model. I believe many will be able to find this leak.</p>\n<p>I think this news might disappoint many, but it seemed right to share it while there’s still a month ahead to make corrections.</p>",
      "rawMarkdown": "Today, looking at the leaderboard, I decided to test one idea and my submission showed that the dates in the test could be accurately restored to initial values. This discovery allows for a metric hack that boosts the leaderboard score by several percentage points.\n\nI re-read the rules and found no explicit prohibition of these actions. I've seen organizer's responses indicating they wouldn't be happy about metric hacking, but I found no restrictions that would lead to disqualification of solutions. Maybe I didn’t search thoroughly, but clear guidelines would be appreciated.\n\nThis is Kaggle after all, and ultimately, it's the rankings that count, not how much you adhere to fair play. If a clean result doesn't top the leaderboard because others used a hack, then ultimately, it doesn't mean anything. Sadly, that's life :(\n\nAfter a few years' break, I find some time to Kaggle again, and I would like to achieve good results that are practically useful, but going against reality seems foolish.\n\nIt would be great if rule-based strict prohibitions against metric hacks were implemented, making such methods lead to disqualification. However, there might be complexities that have prevented this so far.\n\nAnother option is to fix the leak, which could be done with minimal impact on the \"pure\" model. If the organizers are interested in the nature of the leak, I'm ready to explain it privately. Or they could check my last submission—it’s all quite transparent there.\n\nOtherwise, things will stay as they are, and we will compete to see who can exploit this hack in addition to having a strong model. I believe many will be able to find this leak.\n\nI think this news might disappoint many, but it seemed right to share it while there’s still a month ahead to make corrections.",
      "votes": 73
    },
    {
      "id": 2801572,
      "postDate": "2024-05-08T17:03:03.807Z",
      "content": "<p>It seems to be better just reveal real date/week than hiding it to benefit minor persons …</p>",
      "rawMarkdown": "It seems to be better just reveal real date/week than hiding it to benefit minor persons ...",
      "votes": 5
    },
    {
      "id": 2801299,
      "postDate": "2024-05-08T15:13:17.980Z",
      "content": "<p>Just curious, your current score of 0.641 was achieved with or without your metric hack? Just to be clear, I'm not pointing fingers… I'm just curious if such scores can be achieved by simply training models and using their direct predictions. Thanks.</p>",
      "rawMarkdown": "Just curious, your current score of 0.641 was achieved with or without your metric hack? Just to be clear, I'm not pointing fingers... I'm just curious if such scores can be achieved by simply training models and using their direct predictions. Thanks.",
      "votes": 3
    },
    {
      "id": 2770665,
      "postDate": "2024-04-24T01:09:28.860Z",
      "content": "<p>Perhaps AUC is a viable evaluation metric</p>",
      "rawMarkdown": "Perhaps AUC is a viable evaluation metric",
      "votes": 2,
      "replies": [
        {
          "id": 2770672,
          "postDate": "2024-04-24T01:19:23.337Z",
          "content": "<p>Based on my current experiments and discussions, I have come to the conclusion that if it is AUC, everyone will try to construct as many features as possible. If it is the current indicator, a balance must be struck between AUC and stability. If there are too few features, AUC is too low, and too many features can easily affect stability.</p>",
          "rawMarkdown": "Based on my current experiments and discussions, I have come to the conclusion that if it is AUC, everyone will try to construct as many features as possible. If it is the current indicator, a balance must be struck between AUC and stability. If there are too few features, AUC is too low, and too many features can easily affect stability.",
          "votes": 3,
          "replies": [
            {
              "id": 2770678,
              "postDate": "2024-04-24T01:29:28.100Z",
              "content": "<p>According to me, GINI should have been used for the competition and the host could have examined/ executed the top n solutions and assessed stability and handed over a special prize (like the additional stability prize in the existing competition) <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
              "rawMarkdown": "According to me, GINI should have been used for the competition and the host could have examined/ executed the top n solutions and assessed stability and handed over a special prize (like the additional stability prize in the existing competition) @yunsuxiaozi ",
              "votes": 4
            },
            {
              "id": 2770699,
              "postDate": "2024-04-24T01:41:17.553Z",
              "content": "<p>I agree with your statement.</p>",
              "rawMarkdown": "I agree with your statement.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2772390,
          "postDate": "2024-04-24T17:39:53.933Z",
          "content": "<p>An alternative approach could be to evaluate performance using AUC, but assessing it on a test dataset representing data occurring X months into the future from the latest training data. For example, if the training data spans from 2018 to 2020, and the original test data covers 2021 to 2023, one could exclude the 2021 data entirely and evaluate the model's performance on the 2022-2023 data subset.</p>\n<p>Then it will not measure stability over time as defined now, but still measure the models performance on data points occurring far into the future, which (IMO) may be more relevant in many cases. Just because a model begins with an AUC of 0.8 and it decreases to 0.7 after 6 months, it does not inherently suggest a continual decline in performance. Similarly, a model starting with 0.7 and maintaining this level after 6 months does not guarantee greater stability over the subsequent 6 months.</p>\n<p>In addition, while metric hack is possible with the current stability metric, it could still be the case that the current LB score is achieved just by doing something completely different nobody else has thought of (even though metric hacking might be the more plausible guess, given the current facts)</p>",
          "rawMarkdown": "An alternative approach could be to evaluate performance using AUC, but assessing it on a test dataset representing data occurring X months into the future from the latest training data. For example, if the training data spans from 2018 to 2020, and the original test data covers 2021 to 2023, one could exclude the 2021 data entirely and evaluate the model's performance on the 2022-2023 data subset.\n\nThen it will not measure stability over time as defined now, but still measure the models performance on data points occurring far into the future, which (IMO) may be more relevant in many cases. Just because a model begins with an AUC of 0.8 and it decreases to 0.7 after 6 months, it does not inherently suggest a continual decline in performance. Similarly, a model starting with 0.7 and maintaining this level after 6 months does not guarantee greater stability over the subsequent 6 months.\n\nIn addition, while metric hack is possible with the current stability metric, it could still be the case that the current LB score is achieved just by doing something completely different nobody else has thought of (even though metric hacking might be the more plausible guess, given the current facts)\n",
          "votes": 3
        },
        {
          "id": 2813064,
          "postDate": "2024-05-14T14:47:57.910Z",
          "content": "<p>Unfortunately, in this competition, AUC does not represent everything. I once submitted a model with a lower AUC, but its score was much higher compared to a model with a higher AUC after tuning. I think this might be due to some issues where a high AUC score makes the model lack stability or an issue with the competition metric.</p>",
          "rawMarkdown": "Unfortunately, in this competition, AUC does not represent everything. I once submitted a model with a lower AUC, but its score was much higher compared to a model with a higher AUC after tuning. I think this might be due to some issues where a high AUC score makes the model lack stability or an issue with the competition metric.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2770663,
      "postDate": "2024-04-24T01:06:17.863Z",
      "content": "<p>I think one simple rule that could be added for this competition would be that the submission must be  based on the pure predictions of the models (or their weighted average of ensembles) without any other post processing. </p>\n<p>So for example  y_submission =  0.5 * model1.predict() + 0.5 * model2.predict()</p>",
      "rawMarkdown": "I think one simple rule that could be added for this competition would be that the submission must be  based on the pure predictions of the models (or their weighted average of ensembles) without any other post processing. \n\nSo for example  y_submission =  0.5 * model1.predict() + 0.5 * model2.predict()",
      "replies": [
        {
          "id": 2770670,
          "postDate": "2024-04-24T01:16:23.833Z",
          "content": "<p>Does Kaggle endorse such rules? AFAIK Kaggle does not have any restrictions for hacking metrics. As mentioned by <a href=\"https://www.kaggle.com/bruceqdu\" target=\"_blank\">@bruceqdu</a> using the ROC-AUC score would have been a better choice <a href=\"https://www.kaggle.com/fadylabib\" target=\"_blank\">@fadylabib</a> </p>",
          "rawMarkdown": "Does Kaggle endorse such rules? AFAIK Kaggle does not have any restrictions for hacking metrics. As mentioned by @bruceqdu using the ROC-AUC score would have been a better choice @fadylabib ",
          "votes": 1,
          "replies": [
            {
              "id": 2770709,
              "postDate": "2024-04-24T01:50:42Z",
              "content": "<p>My proposal was that the host adds this rule as a competition specific rule, the same way they indicated that they will review and disqualify any solution with hacking in the top 100 places…. It would be a competition specific rule.</p>",
              "rawMarkdown": "My proposal was that the host adds this rule as a competition specific rule, the same way they indicated that they will review and disqualify any solution with hacking in the top 100 places.... It would be a competition specific rule."
            },
            {
              "id": 2770712,
              "postDate": "2024-04-24T01:53:25.353Z",
              "content": "<p>However, then the public leaderboard will not work at all…</p>",
              "rawMarkdown": "However, then the public leaderboard will not work at all..."
            }
          ]
        },
        {
          "id": 2771373,
          "postDate": "2024-04-24T08:07:13.313Z",
          "content": "<p>What if someone wants to use one model for one part of the observations and a different model for another part? For example, what if they split the sample by income size and fit different models accordingly? What if the split is based on row similarity instead? Or by distinguishing between the COVID and non-COVID periods?</p>\n<p>Is this considered postprocessing, or not? Where is the formal line drawn between a legitimate modeling decision and hacking?</p>",
          "rawMarkdown": "What if someone wants to use one model for one part of the observations and a different model for another part? For example, what if they split the sample by income size and fit different models accordingly? What if the split is based on row similarity instead? Or by distinguishing between the COVID and non-COVID periods?\n\nIs this considered postprocessing, or not? Where is the formal line drawn between a legitimate modeling decision and hacking?",
          "votes": 9,
          "replies": [
            {
              "id": 2783717,
              "postDate": "2024-04-29T22:39:20.577Z",
              "content": "<p>I also had this question! Please try to answer, it will bring about much more clarity.</p>",
              "rawMarkdown": "I also had this question! Please try to answer, it will bring about much more clarity."
            },
            {
              "id": 2783738,
              "postDate": "2024-04-29T23:22:49.633Z",
              "content": "<p>There is a normal practice for Kaggle to use many models, including some segment models. </p>\n<p>Host could eliminate leak and make it difficult to hack metric with high precision. <br>\nOr other option - to decrease weight of slope component of metric - if we could not eliminate hacking for 100%, we could do it less attractive. In this case some strong models have chance to compete against hacking. </p>",
              "rawMarkdown": "There is a normal practice for Kaggle to use many models, including some segment models. \n\nHost could eliminate leak and make it difficult to hack metric with high precision. \nOr other option - to decrease weight of slope component of metric - if we could not eliminate hacking for 100%, we could do it less attractive. In this case some strong models have chance to compete against hacking. \n"
            },
            {
              "id": 2783767,
              "postDate": "2024-04-30T00:09:01.223Z",
              "content": "<p>When the hack was first discovered a few months ago, the host stated they were not going to change the metric; maybe this time they will reconsider their decision.</p>\n<blockquote>\n  <p>The host could eliminate the leak and make it difficult to hack the metric with high precision. </p>\n</blockquote>\n<p>Do you think that the fix for the leak you've found will make this useless <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898</a> ? It seems that even if the obvious sources of date information are eliminated, it would still be possible to identify rows similar to those at the end of the training sample. However, I haven't checked this yet.</p>",
              "rawMarkdown": "When the hack was first discovered a few months ago, the host stated they were not going to change the metric; maybe this time they will reconsider their decision.\n\n>The host could eliminate the leak and make it difficult to hack the metric with high precision. \n\nDo you think that the fix for the leak you've found will make this useless https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898 ? It seems that even if the obvious sources of date information are eliminated, it would still be possible to identify rows similar to those at the end of the training sample. However, I haven't checked this yet."
            },
            {
              "id": 2784309,
              "postDate": "2024-04-30T07:40:15.210Z",
              "content": "<p>I didn't try to hack metric intentionally, my discovery was side effect of data analysis. So there is a chance that other information which helps to  identify rows still exists. <br>\nHost have much more info about data and I'll hope they'll find a way to close this issue</p>",
              "rawMarkdown": "I didn't try to hack metric intentionally, my discovery was side effect of data analysis. So there is a chance that other information which helps to  identify rows still exists. \nHost have much more info about data and I'll hope they'll find a way to close this issue",
              "votes": -1
            }
          ]
        }
      ]
    },
    {
      "id": 2774976,
      "postDate": "2024-04-25T12:26:05.247Z",
      "content": "<p>I agree with your progress!!</p>",
      "rawMarkdown": "I agree with your progress!!",
      "votes": -5
    },
    {
      "id": 2803032,
      "postDate": "2024-05-09T09:51:12.530Z",
      "content": "<p><a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> Looking at your score it seems it works ;)</p>",
      "rawMarkdown": "@johnpateha Looking at your score it seems it works ;)",
      "replies": [
        {
          "id": 2805276,
          "postDate": "2024-05-10T13:17:49.760Z",
          "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>  yes,  score up more than 5%. </p>",
          "rawMarkdown": "@narsil  yes,  score up more than 5%. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2785670,
      "postDate": "2024-05-01T01:18:45.200Z",
      "content": "<p>Perhaps this is the price of achieving a ‘stability’ model.</p>",
      "rawMarkdown": "Perhaps this is the price of achieving a ‘stability’ model."
    },
    {
      "id": 2800992,
      "postDate": "2024-05-08T12:44:10.540Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2799962,
      "postDate": "2024-05-08T04:37:38.453Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2801572,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-05-08T17:03:03.807000",
      "content": "<p>It seems to be better just reveal real date/week than hiding it to benefit minor persons …</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2801299,
      "author_name": "Eduard Stefanescu",
      "author_url": "",
      "post_date": "2024-05-08T15:13:17.980000",
      "content": "<p>Just curious, your current score of 0.641 was achieved with or without your metric hack? Just to be clear, I'm not pointing fingers… I'm just curious if such scores can be achieved by simply training models and using their direct predictions. Thanks.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2770665,
      "author_name": "Bruce",
      "author_url": "",
      "post_date": "2024-04-24T01:09:28.860000",
      "content": "<p>Perhaps AUC is a viable evaluation metric</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2770672,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "2024-04-24T01:19:23.337000",
          "content": "<p>Based on my current experiments and discussions, I have come to the conclusion that if it is AUC, everyone will try to construct as many features as possible. If it is the current indicator, a balance must be struck between AUC and stability. If there are too few features, AUC is too low, and too many features can easily affect stability.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2770678,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-04-24T01:29:28.100000",
              "content": "<p>According to me, GINI should have been used for the competition and the host could have examined/ executed the top n solutions and assessed stability and handed over a special prize (like the additional stability prize in the existing competition) <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a> </p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2770699,
              "author_name": "yunsuxiaozi",
              "author_url": "",
              "post_date": "2024-04-24T01:41:17.553000",
              "content": "<p>I agree with your statement.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2772390,
          "author_name": "Ern711",
          "author_url": "",
          "post_date": "2024-04-24T17:39:53.933000",
          "content": "<p>An alternative approach could be to evaluate performance using AUC, but assessing it on a test dataset representing data occurring X months into the future from the latest training data. For example, if the training data spans from 2018 to 2020, and the original test data covers 2021 to 2023, one could exclude the 2021 data entirely and evaluate the model's performance on the 2022-2023 data subset.</p>\n<p>Then it will not measure stability over time as defined now, but still measure the models performance on data points occurring far into the future, which (IMO) may be more relevant in many cases. Just because a model begins with an AUC of 0.8 and it decreases to 0.7 after 6 months, it does not inherently suggest a continual decline in performance. Similarly, a model starting with 0.7 and maintaining this level after 6 months does not guarantee greater stability over the subsequent 6 months.</p>\n<p>In addition, while metric hack is possible with the current stability metric, it could still be the case that the current LB score is achieved just by doing something completely different nobody else has thought of (even though metric hacking might be the more plausible guess, given the current facts)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2813064,
          "author_name": "SStats L",
          "author_url": "",
          "post_date": "2024-05-14T14:47:57.910000",
          "content": "<p>Unfortunately, in this competition, AUC does not represent everything. I once submitted a model with a lower AUC, but its score was much higher compared to a model with a higher AUC after tuning. I think this might be due to some issues where a high AUC score makes the model lack stability or an issue with the competition metric.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2770663,
      "author_name": "ARMADA",
      "author_url": "",
      "post_date": "2024-04-24T01:06:17.863000",
      "content": "<p>I think one simple rule that could be added for this competition would be that the submission must be  based on the pure predictions of the models (or their weighted average of ensembles) without any other post processing. </p>\n<p>So for example  y_submission =  0.5 * model1.predict() + 0.5 * model2.predict()</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2770670,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-04-24T01:16:23.833000",
          "content": "<p>Does Kaggle endorse such rules? AFAIK Kaggle does not have any restrictions for hacking metrics. As mentioned by <a href=\"https://www.kaggle.com/bruceqdu\" target=\"_blank\">@bruceqdu</a> using the ROC-AUC score would have been a better choice <a href=\"https://www.kaggle.com/fadylabib\" target=\"_blank\">@fadylabib</a> </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2770709,
              "author_name": "ARMADA",
              "author_url": "",
              "post_date": "2024-04-24T01:50:42",
              "content": "<p>My proposal was that the host adds this rule as a competition specific rule, the same way they indicated that they will review and disqualify any solution with hacking in the top 100 places…. It would be a competition specific rule.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2770712,
              "author_name": "Wisp Vale",
              "author_url": "",
              "post_date": "2024-04-24T01:53:25.353000",
              "content": "<p>However, then the public leaderboard will not work at all…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2771373,
          "author_name": "Evgeniia Grigoreva",
          "author_url": "",
          "post_date": "2024-04-24T08:07:13.313000",
          "content": "<p>What if someone wants to use one model for one part of the observations and a different model for another part? For example, what if they split the sample by income size and fit different models accordingly? What if the split is based on row similarity instead? Or by distinguishing between the COVID and non-COVID periods?</p>\n<p>Is this considered postprocessing, or not? Where is the formal line drawn between a legitimate modeling decision and hacking?</p>",
          "votes": 9,
          "replies": [
            {
              "id": 2783717,
              "author_name": "YashvardhanSingh11",
              "author_url": "",
              "post_date": "2024-04-29T22:39:20.577000",
              "content": "<p>I also had this question! Please try to answer, it will bring about much more clarity.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2783738,
              "author_name": "Evgeny Patekha",
              "author_url": "",
              "post_date": "2024-04-29T23:22:49.633000",
              "content": "<p>There is a normal practice for Kaggle to use many models, including some segment models. </p>\n<p>Host could eliminate leak and make it difficult to hack metric with high precision. <br>\nOr other option - to decrease weight of slope component of metric - if we could not eliminate hacking for 100%, we could do it less attractive. In this case some strong models have chance to compete against hacking. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2783767,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-04-30T00:09:01.223000",
              "content": "<p>When the hack was first discovered a few months ago, the host stated they were not going to change the metric; maybe this time they will reconsider their decision.</p>\n<blockquote>\n  <p>The host could eliminate the leak and make it difficult to hack the metric with high precision. </p>\n</blockquote>\n<p>Do you think that the fix for the leak you've found will make this useless <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898</a> ? It seems that even if the obvious sources of date information are eliminated, it would still be possible to identify rows similar to those at the end of the training sample. However, I haven't checked this yet.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2784309,
              "author_name": "Evgeny Patekha",
              "author_url": "",
              "post_date": "2024-04-30T07:40:15.210000",
              "content": "<p>I didn't try to hack metric intentionally, my discovery was side effect of data analysis. So there is a chance that other information which helps to  identify rows still exists. <br>\nHost have much more info about data and I'll hope they'll find a way to close this issue</p>",
              "votes": -1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2774976,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-04-25T12:26:05.247000",
      "content": "<p>I agree with your progress!!</p>",
      "votes": -5,
      "replies": []
    },
    {
      "id": 2803032,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2024-05-09T09:51:12.530000",
      "content": "<p><a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> Looking at your score it seems it works ;)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2805276,
          "author_name": "Evgeny Patekha",
          "author_url": "",
          "post_date": "2024-05-10T13:17:49.760000",
          "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>  yes,  score up more than 5%. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2785670,
      "author_name": "暗黑AGI",
      "author_url": "",
      "post_date": "2024-05-01T01:18:45.200000",
      "content": "<p>Perhaps this is the price of achieving a ‘stability’ model.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2800992,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-08T12:44:10.540000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2799962,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-08T04:37:38.453000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2770542": "Today, looking at the leaderboard, I decided to test one idea and my submission showed that the dates in the test could be accurately restored to initial values. This discovery allows for a metric hack that boosts the leaderboard score by several percentage points.\n\nI re-read the rules and found no explicit prohibition of these actions. I've seen organizer's responses indicating they wouldn't be happy about metric hacking, but I found no restrictions that would lead to disqualification of solutions. Maybe I didn’t search thoroughly, but clear guidelines would be appreciated.\n\nThis is Kaggle after all, and ultimately, it's the rankings that count, not how much you adhere to fair play. If a clean result doesn't top the leaderboard because others used a hack, then ultimately, it doesn't mean anything. Sadly, that's life :(\n\nAfter a few years' break, I find some time to Kaggle again, and I would like to achieve good results that are practically useful, but going against reality seems foolish.\n\nIt would be great if rule-based strict prohibitions against metric hacks were implemented, making such methods lead to disqualification. However, there might be complexities that have prevented this so far.\n\nAnother option is to fix the leak, which could be done with minimal impact on the \"pure\" model. If the organizers are interested in the nature of the leak, I'm ready to explain it privately. Or they could check my last submission—it’s all quite transparent there.\n\nOtherwise, things will stay as they are, and we will compete to see who can exploit this hack in addition to having a strong model. I believe many will be able to find this leak.\n\nI think this news might disappoint many, but it seemed right to share it while there’s still a month ahead to make corrections.",
    "2801572": "It seems to be better just reveal real date/week than hiding it to benefit minor persons ...",
    "2801299": "Just curious, your current score of 0.641 was achieved with or without your metric hack? Just to be clear, I'm not pointing fingers... I'm just curious if such scores can be achieved by simply training models and using their direct predictions. Thanks.",
    "2770665": "Perhaps AUC is a viable evaluation metric",
    "2770663": "I think one simple rule that could be added for this competition would be that the submission must be  based on the pure predictions of the models (or their weighted average of ensembles) without any other post processing. \n\nSo for example  y_submission =  0.5 * model1.predict() + 0.5 * model2.predict()",
    "2774976": "I agree with your progress!!",
    "2803032": "@johnpateha Looking at your score it seems it works ;)",
    "2785670": "Perhaps this is the price of achieving a ‘stability’ model.",
    "2800992": "",
    "2799962": ""
  }
}