{
  "id": 477014,
  "title": "Lottery, no thanks",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/477014",
  "author_name": "José A. Guerrero",
  "post_date": "2024-02-14T11:43:21.417000",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, I was thinking in joining this competition after a time out of Kaggle, but a few days ago, based in my experience in other competitions, I though, is better wait the first data leaks (in this case metrics leaks) appears.<br>\nAnd here we go! I read the comments and seems the competition became in a new lottery deciding if test set will performance better or worse, after tuning the predictions. <br>\nI hate lottery competitions.<br>\nAs much complex is the metric more risk it can be hacked.<br>\nPlease, use mean(weekly_gini) - k * sd(weekly_gini) for a k value. This is the lower level of the confidence interval.<br>\nIt is simple an I think is robust respecting hacking attack.<br>\nPlease, give us the opporunity to work with this interesting dataset.</p>\n<p>EDIT: <br>\nAn alternative if you want give more weight to a permormance almost constant over time instead to taking into account the worse periods is to use something similar to sharpe ratio:<br>\nmean(weekly_gini) / (sd(weekly_gini) + K)<br>\nI put K to avoid extremely constant but bad predictions have a very high score.<br>\nI like more the first idea (lower extrem of confidence interval), but this is an alternative.</p>",
  "messages": [
    {
      "id": 2651818,
      "postDate": "2024-02-14T11:43:21.417Z",
      "content": "<p>Hi, I was thinking in joining this competition after a time out of Kaggle, but a few days ago, based in my experience in other competitions, I though, is better wait the first data leaks (in this case metrics leaks) appears.<br>\nAnd here we go! I read the comments and seems the competition became in a new lottery deciding if test set will performance better or worse, after tuning the predictions. <br>\nI hate lottery competitions.<br>\nAs much complex is the metric more risk it can be hacked.<br>\nPlease, use mean(weekly_gini) - k * sd(weekly_gini) for a k value. This is the lower level of the confidence interval.<br>\nIt is simple an I think is robust respecting hacking attack.<br>\nPlease, give us the opporunity to work with this interesting dataset.</p>\n<p>EDIT: <br>\nAn alternative if you want give more weight to a permormance almost constant over time instead to taking into account the worse periods is to use something similar to sharpe ratio:<br>\nmean(weekly_gini) / (sd(weekly_gini) + K)<br>\nI put K to avoid extremely constant but bad predictions have a very high score.<br>\nI like more the first idea (lower extrem of confidence interval), but this is an alternative.</p>",
      "rawMarkdown": "Hi, I was thinking in joining this competition after a time out of Kaggle, but a few days ago, based in my experience in other competitions, I though, is better wait the first data leaks (in this case metrics leaks) appears.\nAnd here we go! I read the comments and seems the competition became in a new lottery deciding if test set will performance better or worse, after tuning the predictions. \nI hate lottery competitions.\nAs much complex is the metric more risk it can be hacked.\nPlease, use mean(weekly_gini) - k * sd(weekly_gini) for a k value. This is the lower level of the confidence interval.\nIt is simple an I think is robust respecting hacking attack.\nPlease, give us the opporunity to work with this interesting dataset.\n\nEDIT: \nAn alternative if you want give more weight to a permormance almost constant over time instead to taking into account the worse periods is to use something similar to sharpe ratio:\nmean(weekly_gini) / (sd(weekly_gini) + K)\nI put K to avoid extremely constant but bad predictions have a very high score.\nI like more the first idea (lower extrem of confidence interval), but this is an alternative.",
      "votes": 14
    },
    {
      "id": 2651857,
      "postDate": "2024-02-14T12:09:19.193Z",
      "content": "<p>You should read <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867\" target=\"_blank\">this announcement</a>. Hosts are really active in this competition. Considering that we KNOW that test data comes from major economic and social disaster which was COVID gives us more information, than anybody would normally have.</p>",
      "rawMarkdown": "You should read [this announcement](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867). Hosts are really active in this competition. Considering that we KNOW that test data comes from major economic and social disaster which was COVID gives us more information, than anybody would normally have.",
      "votes": 5
    },
    {
      "id": 2652887,
      "postDate": "2024-02-15T04:10:03.903Z",
      "content": "<p>If the metric can be exploited in <em>any</em> way. It will be exploited, and at that point many people will be out of interest. No longer an ML challenge.</p>\n<p>The problem is interesting, and makes sense, but the metric is not the quality needed. It is not well designed and there is math to defend it. There is nothing wrong with the competition goal, its the metric that does not reflect the intention.</p>\n<p>In the search for quality models a quality metric must exist.</p>",
      "rawMarkdown": "If the metric can be exploited in *any* way. It will be exploited, and at that point many people will be out of interest. No longer an ML challenge.\n\nThe problem is interesting, and makes sense, but the metric is not the quality needed. It is not well designed and there is math to defend it. There is nothing wrong with the competition goal, its the metric that does not reflect the intention.\n\nIn the search for quality models a quality metric must exist.",
      "votes": 3
    },
    {
      "id": 2652056,
      "postDate": "2024-02-14T14:40:58.613Z",
      "content": "<p>I am sorry to hear that you won't be joining due to the fact that you are unhappy with the metric. You could give it a try anyway while we are working on the fix.</p>\n<p>I must say even though the \"hack\" can give you some additional points, winning the competition would not be definitely just about \"hacking\" the metric. Sure, as it is done now, the top (let's say) 5% would need to use it, but winning will still be based also on well designed stable model. So, I don't think it is correct to say that the whole competition is now about hacking the metric, that is a little bit of an overstatement. </p>\n<p>As I mentioned in Discussion several times already. Omitting the falling rate would not be a satisfactory solution. Without the falling rate we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk manager in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. </p>\n<p>I am afraid there is no simple solution, let's have a look at the extremes. 1) we could omit mean gini or 2) we could omit the part with falling rate. When going with 2) the issue becomes that the models are pushed towards excellent short-term performance, but yet not stable in time. When going with 1) the issue becomes tweaking of the score so that we have as stable gini in time as possible. Changing the weights will not result in solving the problem, but rather transformation from one issue into another.</p>\n<p>As mentioned <a href=\"https://www.kaggle.com/marekgp\" target=\"_blank\">@marekgp</a> have a look at our official statement <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867</a> we definitely want to make the competition more fair for everyone and penalize hacking the metric.</p>",
      "rawMarkdown": "I am sorry to hear that you won't be joining due to the fact that you are unhappy with the metric. You could give it a try anyway while we are working on the fix.\n\nI must say even though the \"hack\" can give you some additional points, winning the competition would not be definitely just about \"hacking\" the metric. Sure, as it is done now, the top (let's say) 5% would need to use it, but winning will still be based also on well designed stable model. So, I don't think it is correct to say that the whole competition is now about hacking the metric, that is a little bit of an overstatement. \n\nAs I mentioned in Discussion several times already. Omitting the falling rate would not be a satisfactory solution. Without the falling rate we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk manager in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. \n\nI am afraid there is no simple solution, let's have a look at the extremes. 1) we could omit mean gini or 2) we could omit the part with falling rate. When going with 2) the issue becomes that the models are pushed towards excellent short-term performance, but yet not stable in time. When going with 1) the issue becomes tweaking of the score so that we have as stable gini in time as possible. Changing the weights will not result in solving the problem, but rather transformation from one issue into another.\n\nAs mentioned @marekgp have a look at our official statement https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867 we definitely want to make the competition more fair for everyone and penalize hacking the metric.",
      "votes": 1,
      "replies": [
        {
          "id": 2653140,
          "postDate": "2024-02-15T08:32:37.470Z",
          "content": "<blockquote>\n  <p>I must say even though the \"hack\" can give you some additional points, winning the competition would not be definitely just about \"hacking\" the metric. </p>\n</blockquote>\n<p>I have a different perspective on that. At the moment, metric hacking yields so much more gain to public LB compared to feature engineering and/or ensembling (<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/477228\" target=\"_blank\">as i show here</a>) that it incentivizes people to focus the majority of their work on metric hacking.</p>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> I think the changes you are discussing should eliminate any gain in public LB from purposefully decreasing the model score. For example, consider the following fix: assign higher weights to weekly gini scores that are further-in time. It would promote models that are not degrading, but at the same time one would gain nothing by decreasing their scores.</p>",
          "rawMarkdown": "> I must say even though the \"hack\" can give you some additional points, winning the competition would not be definitely just about \"hacking\" the metric. \n\nI have a different perspective on that. At the moment, metric hacking yields so much more gain to public LB compared to feature engineering and/or ensembling ([as i show here](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/477228)) that it incentivizes people to focus the majority of their work on metric hacking.\n\n@jetakow I think the changes you are discussing should eliminate any gain in public LB from purposefully decreasing the model score. For example, consider the following fix: assign higher weights to weekly gini scores that are further-in time. It would promote models that are not degrading, but at the same time one would gain nothing by decreasing their scores.",
          "votes": 4
        },
        {
          "id": 2655484,
          "postDate": "2024-02-16T23:06:17.720Z",
          "content": "<p>Thanks for the response, Daniel. However, it feels like déjà vu from many other competitions where organizers insisted on using an inappropriate metric against multiple well-founded opinions, ultimately resulting in an unusable model. Currently, I already have several ideas to optimize the metric (let's call optimization what it really is, hacking), and after seeing comments from people claiming to achieve an order of magnitude improvement with optimization compared to variable engineering, I am more interested in giving it a try</p>",
          "rawMarkdown": "Thanks for the response, Daniel. However, it feels like déjà vu from many other competitions where organizers insisted on using an inappropriate metric against multiple well-founded opinions, ultimately resulting in an unusable model. Currently, I already have several ideas to optimize the metric (let's call optimization what it really is, hacking), and after seeing comments from people claiming to achieve an order of magnitude improvement with optimization compared to variable engineering, I am more interested in giving it a try",
          "votes": 4
        }
      ]
    },
    {
      "id": 2651859,
      "postDate": "2024-02-14T12:11:33.860Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2651857,
      "author_name": "Marek Przybyłowicz",
      "author_url": "",
      "post_date": "2024-02-14T12:09:19.193000",
      "content": "<p>You should read <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867\" target=\"_blank\">this announcement</a>. Hosts are really active in this competition. Considering that we KNOW that test data comes from major economic and social disaster which was COVID gives us more information, than anybody would normally have.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2652887,
      "author_name": "NxGTR",
      "author_url": "",
      "post_date": "2024-02-15T04:10:03.903000",
      "content": "<p>If the metric can be exploited in <em>any</em> way. It will be exploited, and at that point many people will be out of interest. No longer an ML challenge.</p>\n<p>The problem is interesting, and makes sense, but the metric is not the quality needed. It is not well designed and there is math to defend it. There is nothing wrong with the competition goal, its the metric that does not reflect the intention.</p>\n<p>In the search for quality models a quality metric must exist.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2652056,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-14T14:40:58.613000",
      "content": "<p>I am sorry to hear that you won't be joining due to the fact that you are unhappy with the metric. You could give it a try anyway while we are working on the fix.</p>\n<p>I must say even though the \"hack\" can give you some additional points, winning the competition would not be definitely just about \"hacking\" the metric. Sure, as it is done now, the top (let's say) 5% would need to use it, but winning will still be based also on well designed stable model. So, I don't think it is correct to say that the whole competition is now about hacking the metric, that is a little bit of an overstatement. </p>\n<p>As I mentioned in Discussion several times already. Omitting the falling rate would not be a satisfactory solution. Without the falling rate we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk manager in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. </p>\n<p>I am afraid there is no simple solution, let's have a look at the extremes. 1) we could omit mean gini or 2) we could omit the part with falling rate. When going with 2) the issue becomes that the models are pushed towards excellent short-term performance, but yet not stable in time. When going with 1) the issue becomes tweaking of the score so that we have as stable gini in time as possible. Changing the weights will not result in solving the problem, but rather transformation from one issue into another.</p>\n<p>As mentioned <a href=\"https://www.kaggle.com/marekgp\" target=\"_blank\">@marekgp</a> have a look at our official statement <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867</a> we definitely want to make the competition more fair for everyone and penalize hacking the metric.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2653140,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-02-15T08:32:37.470000",
          "content": "<blockquote>\n  <p>I must say even though the \"hack\" can give you some additional points, winning the competition would not be definitely just about \"hacking\" the metric. </p>\n</blockquote>\n<p>I have a different perspective on that. At the moment, metric hacking yields so much more gain to public LB compared to feature engineering and/or ensembling (<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/477228\" target=\"_blank\">as i show here</a>) that it incentivizes people to focus the majority of their work on metric hacking.</p>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> I think the changes you are discussing should eliminate any gain in public LB from purposefully decreasing the model score. For example, consider the following fix: assign higher weights to weekly gini scores that are further-in time. It would promote models that are not degrading, but at the same time one would gain nothing by decreasing their scores.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2655484,
          "author_name": "José A. Guerrero",
          "author_url": "",
          "post_date": "2024-02-16T23:06:17.720000",
          "content": "<p>Thanks for the response, Daniel. However, it feels like déjà vu from many other competitions where organizers insisted on using an inappropriate metric against multiple well-founded opinions, ultimately resulting in an unusable model. Currently, I already have several ideas to optimize the metric (let's call optimization what it really is, hacking), and after seeing comments from people claiming to achieve an order of magnitude improvement with optimization compared to variable engineering, I am more interested in giving it a try</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2651859,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-14T12:11:33.860000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2651818": "Hi, I was thinking in joining this competition after a time out of Kaggle, but a few days ago, based in my experience in other competitions, I though, is better wait the first data leaks (in this case metrics leaks) appears.\nAnd here we go! I read the comments and seems the competition became in a new lottery deciding if test set will performance better or worse, after tuning the predictions. \nI hate lottery competitions.\nAs much complex is the metric more risk it can be hacked.\nPlease, use mean(weekly_gini) - k * sd(weekly_gini) for a k value. This is the lower level of the confidence interval.\nIt is simple an I think is robust respecting hacking attack.\nPlease, give us the opporunity to work with this interesting dataset.\n\nEDIT: \nAn alternative if you want give more weight to a permormance almost constant over time instead to taking into account the worse periods is to use something similar to sharpe ratio:\nmean(weekly_gini) / (sd(weekly_gini) + K)\nI put K to avoid extremely constant but bad predictions have a very high score.\nI like more the first idea (lower extrem of confidence interval), but this is an alternative.",
    "2651857": "You should read [this announcement](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867). Hosts are really active in this competition. Considering that we KNOW that test data comes from major economic and social disaster which was COVID gives us more information, than anybody would normally have.",
    "2652887": "If the metric can be exploited in *any* way. It will be exploited, and at that point many people will be out of interest. No longer an ML challenge.\n\nThe problem is interesting, and makes sense, but the metric is not the quality needed. It is not well designed and there is math to defend it. There is nothing wrong with the competition goal, its the metric that does not reflect the intention.\n\nIn the search for quality models a quality metric must exist.",
    "2652056": "I am sorry to hear that you won't be joining due to the fact that you are unhappy with the metric. You could give it a try anyway while we are working on the fix.\n\nI must say even though the \"hack\" can give you some additional points, winning the competition would not be definitely just about \"hacking\" the metric. Sure, as it is done now, the top (let's say) 5% would need to use it, but winning will still be based also on well designed stable model. So, I don't think it is correct to say that the whole competition is now about hacking the metric, that is a little bit of an overstatement. \n\nAs I mentioned in Discussion several times already. Omitting the falling rate would not be a satisfactory solution. Without the falling rate we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk manager in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. \n\nI am afraid there is no simple solution, let's have a look at the extremes. 1) we could omit mean gini or 2) we could omit the part with falling rate. When going with 2) the issue becomes that the models are pushed towards excellent short-term performance, but yet not stable in time. When going with 1) the issue becomes tweaking of the score so that we have as stable gini in time as possible. Changing the weights will not result in solving the problem, but rather transformation from one issue into another.\n\nAs mentioned @marekgp have a look at our official statement https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867 we definitely want to make the competition more fair for everyone and penalize hacking the metric.",
    "2651859": ""
  }
}