{
  "id": 476867,
  "title": "Announcement regarding competition metric - UPDATED!",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/476867",
  "author_name": "Tomas Jelinek",
  "post_date": "2024-02-13T19:02:47.900000",
  "votes": 83,
  "comment_count": 65,
  "views": 0,
  "content": "<p><strong><em><em></em></em></strong>Dear Kagglers!</p>\n<p>As you have probably noticed in the discussion already, there is issue with our competition metric. In short, it is possible to gain extra score by artificially reducing predictive power of model for first part of evaluation sample, thus smoothing gini over time and reduce penalty for instability.</p>\n<p>This is something we are not looking for. Our goal is to maximize both predictive ability and stability. At the same time our implicit expectation was that for all clients will be applied same rules/model/approach, in other words the scoring would not explicitly depend on date decision. From the discussion on Kaggle I have strong feeling that you see it similarly, and you don’t want competition where winning strategy is to game scoring metric. However, we don’t want to remove stability altogether, from the beginning we wanted to present you some new problem, not just repeat previous Home Credit competition.</p>\n<p>Therefore, we have agreed we have to react. Together with Kaggle team we are analyzing several ideas how to prevent possibility to improve score by artificially changing models for certain time periods. At the same time we fully understand that it will be change in on-going competition, which should not be done under normal circumstances. We want to make sure that our fix will be correct and final, so we want to thoroughly analyze and test our ideas – it will take us few days. But we want to inform you we are working on solution and be transparent with you.</p>\n<p>Last but not least, I want to thank many of you who were pointing at this issue, namely <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> for sharing openly his solution and for proving that our metric is not working as intended.</p>\n<p>Will keep you updated.</p>\n<p>Best,</p>\n<p>Tomas &amp; Daniel</p>",
  "messages": [
    {
      "id": 2650939,
      "postDate": "2024-02-13T19:02:47.900Z",
      "content": "<p><strong><em><em></em></em></strong>Dear Kagglers!</p>\n<p>As you have probably noticed in the discussion already, there is issue with our competition metric. In short, it is possible to gain extra score by artificially reducing predictive power of model for first part of evaluation sample, thus smoothing gini over time and reduce penalty for instability.</p>\n<p>This is something we are not looking for. Our goal is to maximize both predictive ability and stability. At the same time our implicit expectation was that for all clients will be applied same rules/model/approach, in other words the scoring would not explicitly depend on date decision. From the discussion on Kaggle I have strong feeling that you see it similarly, and you don’t want competition where winning strategy is to game scoring metric. However, we don’t want to remove stability altogether, from the beginning we wanted to present you some new problem, not just repeat previous Home Credit competition.</p>\n<p>Therefore, we have agreed we have to react. Together with Kaggle team we are analyzing several ideas how to prevent possibility to improve score by artificially changing models for certain time periods. At the same time we fully understand that it will be change in on-going competition, which should not be done under normal circumstances. We want to make sure that our fix will be correct and final, so we want to thoroughly analyze and test our ideas – it will take us few days. But we want to inform you we are working on solution and be transparent with you.</p>\n<p>Last but not least, I want to thank many of you who were pointing at this issue, namely <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> for sharing openly his solution and for proving that our metric is not working as intended.</p>\n<p>Will keep you updated.</p>\n<p>Best,</p>\n<p>Tomas &amp; Daniel</p>",
      "rawMarkdown": "********Dear Kagglers!\n\nAs you have probably noticed in the discussion already, there is issue with our competition metric. In short, it is possible to gain extra score by artificially reducing predictive power of model for first part of evaluation sample, thus smoothing gini over time and reduce penalty for instability.\n\nThis is something we are not looking for. Our goal is to maximize both predictive ability and stability. At the same time our implicit expectation was that for all clients will be applied same rules/model/approach, in other words the scoring would not explicitly depend on date decision. From the discussion on Kaggle I have strong feeling that you see it similarly, and you don’t want competition where winning strategy is to game scoring metric. However, we don’t want to remove stability altogether, from the beginning we wanted to present you some new problem, not just repeat previous Home Credit competition.\n\nTherefore, we have agreed we have to react. Together with Kaggle team we are analyzing several ideas how to prevent possibility to improve score by artificially changing models for certain time periods. At the same time we fully understand that it will be change in on-going competition, which should not be done under normal circumstances. We want to make sure that our fix will be correct and final, so we want to thoroughly analyze and test our ideas – it will take us few days. But we want to inform you we are working on solution and be transparent with you.\n\nLast but not least, I want to thank many of you who were pointing at this issue, namely @at7459 for sharing openly his solution and for proving that our metric is not working as intended.\n\nWill keep you updated.\n\nBest,\n\nTomas & Daniel\n",
      "votes": 82
    },
    {
      "id": 2651259,
      "postDate": "2024-02-14T03:45:58.743Z",
      "content": "<p>First, <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> thank you for the report. <br>\nAnd thank you to the host <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> , <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>   for sincerely considering a revision of the metric for this topic.</p>\n<p>My thought is that, perhaps, if we consider the reverse, isn't there a potential issue like in Case 3, where improving the Gini value in parts could worsen the overall score? (Even though the average value of Gini might increase, the slope becomes significantly negative, and its weight is large). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F4f77e7c26d8d69e6c9a3e2897368bc0e%2FClipboard03.jpg?generation=1707882775177471&amp;alt=media\"></p>\n<p>※ The green represents a fixed Gini value for explanation purposes. The blue is intentionally shifted for explanation purposes.</p>\n<p>And, this potential issue is incompatible with the description in the data tab below and could significantly impact the overall score.</p>\n<pre><code>It\n</code></pre>\n<p>I would appreciate it if you could also consider this potential issue. But, if that's what the host desires, I won't say anything.</p>\n<p>Personally, using polyfit makes it difficult to balance the mean value of Gini with the term of the slope, making it challenging to excel in this competition.</p>\n<p>Perhaps, instead of using polyfit, for example, weighting the Gini function for each week number, taking the sum, and then normalizing it at the end, might also be a good metric.</p>\n<p>However, this is just my opinion, so I leave the decision to the host. I would appreciate it if you could take this into consideration. Please feel free to point out any mistakes.</p>",
      "rawMarkdown": "First, @at7459 thank you for the report. \nAnd thank you to the host @tomasjeline2 , @jetakow   for sincerely considering a revision of the metric for this topic.\n\nMy thought is that, perhaps, if we consider the reverse, isn't there a potential issue like in Case 3, where improving the Gini value in parts could worsen the overall score? (Even though the average value of Gini might increase, the slope becomes significantly negative, and its weight is large). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F4f77e7c26d8d69e6c9a3e2897368bc0e%2FClipboard03.jpg?generation=1707882775177471&alt=media)\n\n※ The green represents a fixed Gini value for explanation purposes. The blue is intentionally shifted for explanation purposes.\n\n\nAnd, this potential issue is incompatible with the description in the data tab below and could significantly impact the overall score.\n~~~\nIt's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated.\n~~~\n\nI would appreciate it if you could also consider this potential issue. But, if that's what the host desires, I won't say anything.\n\nPersonally, using polyfit makes it difficult to balance the mean value of Gini with the term of the slope, making it challenging to excel in this competition.\n\nPerhaps, instead of using polyfit, for example, weighting the Gini function for each week number, taking the sum, and then normalizing it at the end, might also be a good metric.\n\nHowever, this is just my opinion, so I leave the decision to the host. I would appreciate it if you could take this into consideration. Please feel free to point out any mistakes.",
      "votes": 20
    },
    {
      "id": 2662222,
      "postDate": "2024-02-21T19:13:01.210Z",
      "content": "<p>Dear Kagglers, </p>\n<p>We prepared this year’s competition with an intention to present you something else and more challenging than the previous Home Credit kaggle competition. We decided to select stability of models in production as the main objective of this competition. Therefore, we will keep stability and won’t remove it from the competition.  </p>\n<p>We would like to further clarify the meaning of stability. Our ideal stable model has constant performance by weeks in time and as high of a mean gini over weeks as possible. That is the reason why we have three terms in the stability metric as of now since this is desirable in business and not only to maximize AUC. The metric itself is a good approximation of the decisions made in-house when assessing the performance of models in production. Despite different opinions here within the competition, removing the terms in the stability metric would significantly deviate from the way we rate performance internally. When people leverage the second term and nullify it by carefully choosing scores on the test sample, it does not represent how we deploy models in the production. This is undesirable and we are forced to make steps that will discourage hacking of the metric.  </p>\n<p>We will make a change to the test dataset in such a way that it will represent better client scoring in production. When the model is running in production, it can only see its past performance and predictions. This would be quite difficult to reproduce in a Kaggle competition, but we have an alternative solution. We will remove the date decision, month and week number from the test base table. This part will take two weeks to amend, and we will prolong the competition by this period. Because of this, it is very likely that we won’t be able to rescore the old submissions. We apologize in advance for this inconvenience. </p>\n<p>Besides the change of test data, we have a second plan regarding metric. At this moment we are still considering metric change, and we welcome suggestions that would be aligned with how we view the value of model stability and in a way that would prevent further hacking. We are unlikely to go with a new metric that would prevent this type of hacking and would significantly deviate from the stability part. There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.</p>\n<p>We would like to apologize for the delay and inconvenience caused by the issues connected to the stability metric. We also want to thank the community for many proposals on how to fix the issue. We have tested and analyzed many of them, thank you for all the suggestions! We will disable submissions for two weeks from now on and prolong the competition by two weeks.  </p>\n<p>Tomas &amp; Daniel</p>",
      "rawMarkdown": "Dear Kagglers, \n \nWe prepared this year’s competition with an intention to present you something else and more challenging than the previous Home Credit kaggle competition. We decided to select stability of models in production as the main objective of this competition. Therefore, we will keep stability and won’t remove it from the competition.  \n\nWe would like to further clarify the meaning of stability. Our ideal stable model has constant performance by weeks in time and as high of a mean gini over weeks as possible. That is the reason why we have three terms in the stability metric as of now since this is desirable in business and not only to maximize AUC. The metric itself is a good approximation of the decisions made in-house when assessing the performance of models in production. Despite different opinions here within the competition, removing the terms in the stability metric would significantly deviate from the way we rate performance internally. When people leverage the second term and nullify it by carefully choosing scores on the test sample, it does not represent how we deploy models in the production. This is undesirable and we are forced to make steps that will discourage hacking of the metric.  \n \nWe will make a change to the test dataset in such a way that it will represent better client scoring in production. When the model is running in production, it can only see its past performance and predictions. This would be quite difficult to reproduce in a Kaggle competition, but we have an alternative solution. We will remove the date decision, month and week number from the test base table. This part will take two weeks to amend, and we will prolong the competition by this period. Because of this, it is very likely that we won’t be able to rescore the old submissions. We apologize in advance for this inconvenience. \n \nBesides the change of test data, we have a second plan regarding metric. At this moment we are still considering metric change, and we welcome suggestions that would be aligned with how we view the value of model stability and in a way that would prevent further hacking. We are unlikely to go with a new metric that would prevent this type of hacking and would significantly deviate from the stability part. There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.\n \nWe would like to apologize for the delay and inconvenience caused by the issues connected to the stability metric. We also want to thank the community for many proposals on how to fix the issue. We have tested and analyzed many of them, thank you for all the suggestions! We will disable submissions for two weeks from now on and prolong the competition by two weeks.  \n\nTomas & Daniel",
      "votes": 15,
      "replies": [
        {
          "id": 2662256,
          "postDate": "2024-02-21T19:38:03.710Z",
          "content": "<p>This will likely lead to a separate set of models, focused on predicting the <code>week_number</code> based on available features. Once you let the Genie out of the bottle there is no coming back :) But I get it - there are no ideal solutions. Thank you for the effort.</p>",
          "rawMarkdown": "This will likely lead to a separate set of models, focused on predicting the `week_number` based on available features. Once you let the Genie out of the bottle there is no coming back :) But I get it - there are no ideal solutions. Thank you for the effort.",
          "votes": 5,
          "replies": [
            {
              "id": 2662265,
              "postDate": "2024-02-21T19:42:13.943Z",
              "content": "<p>I am pretty sure as long as the metric won’t be strictly decreasing in gini scores for every time point it will lead to people hacking the leaderboard.</p>",
              "rawMarkdown": "I am pretty sure as long as the metric won’t be strictly decreasing in gini scores for every time point it will lead to people hacking the leaderboard."
            },
            {
              "id": 2662269,
              "postDate": "2024-02-21T19:44:38.400Z",
              "content": "<p>The winning models will be reviewed manually :) We want stable model/models in time, not models that predict what is the week number. This approach wouldn't be the winning one. </p>",
              "rawMarkdown": "The winning models will be reviewed manually :) We want stable model/models in time, not models that predict what is the week number. This approach wouldn't be the winning one. "
            },
            {
              "id": 2662280,
              "postDate": "2024-02-21T19:52:00.407Z",
              "content": "<p>I think it will be enormous amount of people submitting such stuff. do you mean you would manually review hundreds of submissions?<br>\nAlso it’s frustrating for people who want to build better models if one never can get a real benchmark of how good the model performs on the LB due to hundreds of hacked solutions.<br>\nI don’t get why not chose a mathematical „safe“ metric which is provably not hackable by decreasing model performance.</p>",
              "rawMarkdown": "I think it will be enormous amount of people submitting such stuff. do you mean you would manually review hundreds of submissions?\nAlso it’s frustrating for people who want to build better models if one never can get a real benchmark of how good the model performs on the LB due to hundreds of hacked solutions.\nI don’t get why not chose a mathematical „safe“ metric which is provably not hackable by decreasing model performance.",
              "votes": 2
            },
            {
              "id": 2662295,
              "postDate": "2024-02-21T20:06:14.907Z",
              "content": "<p>I can assure you that we will do our best to detect such hacking of the metric even if someone fortunate passes the private test set. If the top 100 solutions will be using some way of week num prediction, yes we will go over those hundred solutions in the top private leaderboard until we find a solution that doesn't have this approach :) </p>\n<p>We discourage any effort that will be put into a prediction of date within the test sample. </p>\n<p>If you have any metric on your mind, don't hesitate to suggest it and I will review it personally. The main issue is that those mathematically safe metrics don't asses very well decreasing performance. I am quite optimistic nd I believe such metric exists, however, we are limited by time. </p>",
              "rawMarkdown": "I can assure you that we will do our best to detect such hacking of the metric even if someone fortunate passes the private test set. If the top 100 solutions will be using some way of week num prediction, yes we will go over those hundred solutions in the top private leaderboard until we find a solution that doesn't have this approach :) \n\nWe discourage any effort that will be put into a prediction of date within the test sample. \n\nIf you have any metric on your mind, don't hesitate to suggest it and I will review it personally. The main issue is that those mathematically safe metrics don't asses very well decreasing performance. I am quite optimistic nd I believe such metric exists, however, we are limited by time. "
            },
            {
              "id": 2662570,
              "postDate": "2024-02-22T02:06:24.570Z",
              "content": "<p>That only covers top submissions. Most people will likely be forced to do hacking just to keep a reasonable ranking.</p>",
              "rawMarkdown": "That only covers top submissions. Most people will likely be forced to do hacking just to keep a reasonable ranking.",
              "votes": 4
            },
            {
              "id": 2663000,
              "postDate": "2024-02-22T09:18:30.477Z",
              "content": "<p>However, if you delete 'date_decision', it will cause most of the data on the date is not useful. In terms of my model, it may reduce my score about 0.01</p>",
              "rawMarkdown": "However, if you delete 'date_decision', it will cause most of the data on the date is not useful. In terms of my model, it may reduce my score about 0.01",
              "votes": 1
            },
            {
              "id": 2663018,
              "postDate": "2024-02-22T09:26:05.837Z",
              "content": "<p>If you decide to delete \"date_decision\", I think I need other feather, like:<br>\nif col[-1] == \"D\":<br>\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))<br>\n                df = df.with_columns(pl.col(col).dt.total_days())<br>\nI hope that  original data about date changes like this.</p>",
              "rawMarkdown": "If you decide to delete \"date_decision\", I think I need other feather, like:\nif col[-1] == \"D\":\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n                df = df.with_columns(pl.col(col).dt.total_days())\nI hope that  original data about date changes like this."
            },
            {
              "id": 2664509,
              "postDate": "2024-02-23T04:43:31.757Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2662268,
          "postDate": "2024-02-21T19:44:21.597Z",
          "content": "<p>Hi Tomas and Daniel,</p>\n<p>To aid with the alternative metric proposals, can you publish the \"units tests\" that the metric must pass?  Can you publish multiple examples with two hypothetical week-by-week Gini behaviors, where you judge one behavior to be more desirable than the other?</p>\n<p>I think that it would be hard to propose metrics that align with your vision without something quantifiable, especially since there is probably a disconnect between what Home Credit considers desirable and what competitors here would assume to be desirable.  For better or worse, the units tests would establish a hard ground truth that we would have to respect with our proposals.</p>",
          "rawMarkdown": "Hi Tomas and Daniel,\n\nTo aid with the alternative metric proposals, can you publish the \"units tests\" that the metric must pass?  Can you publish multiple examples with two hypothetical week-by-week Gini behaviors, where you judge one behavior to be more desirable than the other?\n\nI think that it would be hard to propose metrics that align with your vision without something quantifiable, especially since there is probably a disconnect between what Home Credit considers desirable and what competitors here would assume to be desirable.  For better or worse, the units tests would establish a hard ground truth that we would have to respect with our proposals.",
          "votes": 1,
          "replies": [
            {
              "id": 2662273,
              "postDate": "2024-02-21T19:48:09.383Z",
              "content": "<p>I am afraid we can't post such unit tests that we have in-house. However, I am willing to score any metric you will send. After scoring I could also issue examples where the proposed metric is not better when I feel it would help. What would you say to that?</p>",
              "rawMarkdown": "I am afraid we can't post such unit tests that we have in-house. However, I am willing to score any metric you will send. After scoring I could also issue examples where the proposed metric is not better when I feel it would help. What would you say to that?"
            },
            {
              "id": 2662285,
              "postDate": "2024-02-21T19:56:58.513Z",
              "content": "<p>I think it won't lead to effective iteration when you're trying to hit the target you can't see, and don't have instant feedback as to whether you're on the right track.</p>\n<p>Can there be a compromise solution here?  Can you publish some preliminary units tests so that we have something tangible to go on, but still make the final determination with your private unit tests?</p>",
              "rawMarkdown": "I think it won't lead to effective iteration when you're trying to hit the target you can't see, and don't have instant feedback as to whether you're on the right track.\n\nCan there be a compromise solution here?  Can you publish some preliminary units tests so that we have something tangible to go on, but still make the final determination with your private unit tests?",
              "votes": 1
            },
            {
              "id": 2662297,
              "postDate": "2024-02-21T20:10:19.613Z",
              "content": "<p>It doesn't have to be effective, in the end we will be selecting the metric. I am afraid we simply can not give you the same data we used for determining the current metric. We can however use them to score any suggested metric as I said. We could derive such unit tests, but in the sake of using the time wisely, I am afraid we won't be able to spend much time on it. It would be better to spend a considerable amount of time on this activity so there is no scope for inconsistency and at the same time, we won't leak the data used for such a decision during the competition itself. </p>",
              "rawMarkdown": "It doesn't have to be effective, in the end we will be selecting the metric. I am afraid we simply can not give you the same data we used for determining the current metric. We can however use them to score any suggested metric as I said. We could derive such unit tests, but in the sake of using the time wisely, I am afraid we won't be able to spend much time on it. It would be better to spend a considerable amount of time on this activity so there is no scope for inconsistency and at the same time, we won't leak the data used for such a decision during the competition itself. "
            }
          ]
        },
        {
          "id": 2662374,
          "postDate": "2024-02-21T21:05:27.063Z",
          "content": "<p>\"We will remove the date decision…\": but \"date decision\" is subtracted from all the \"date\" inputs to turn them into time periods; without \"date decision\" all the date inputs immediately become less usable. And they can be used to impute the missing \"date decision\" anyway.</p>",
          "rawMarkdown": "\"We will remove the date decision...\": but \"date decision\" is subtracted from all the \"date\" inputs to turn them into time periods; without \"date decision\" all the date inputs immediately become less usable. And they can be used to impute the missing \"date decision\" anyway.",
          "votes": 4,
          "replies": [
            {
              "id": 2662420,
              "postDate": "2024-02-21T21:53:13.897Z",
              "content": "<blockquote>\n  <p>\"We will remove the date decision…\": but \"date decision\" is subtracted from all the \"date\" inputs to turn them into time periods; without \"date decision\" all the date inputs immediately become less usable. </p>\n</blockquote>\n<p>We are aware of that, it is acceptable to us at this moment. Unless we of course in a week come up with a resonable new metric.</p>\n<blockquote>\n  <p>And they can be used to impute the missing \"date decision\" anyway.</p>\n</blockquote>\n<p>Don't worry about this part. Data will be transformed again :)</p>",
              "rawMarkdown": ">\"We will remove the date decision…\": but \"date decision\" is subtracted from all the \"date\" inputs to turn them into time periods; without \"date decision\" all the date inputs immediately become less usable. \n\nWe are aware of that, it is acceptable to us at this moment. Unless we of course in a week come up with a resonable new metric.\n\n>And they can be used to impute the missing \"date decision\" anyway.\n\nDon't worry about this part. Data will be transformed again :)",
              "votes": 1
            }
          ]
        },
        {
          "id": 2662551,
          "postDate": "2024-02-22T01:21:52.080Z",
          "content": "<p>很抱歉我英文不好，我还是用中文表达吧。我已经观察到了测试数据的问题，并且利用这个问题拿到了public榜单的第一名。我认为修改公式并不能很好的解决这个问题。建议修改测试数据。只保留未来的部分。</p>",
          "rawMarkdown": "很抱歉我英文不好，我还是用中文表达吧。我已经观察到了测试数据的问题，并且利用这个问题拿到了public榜单的第一名。我认为修改公式并不能很好的解决这个问题。建议修改测试数据。只保留未来的部分。",
          "replies": [
            {
              "id": 2663040,
              "postDate": "2024-02-22T09:33:20.530Z",
              "content": "<p><a href=\"https://www.kaggle.com/boristown\" target=\"_blank\">@boristown</a> I used google translate so I hope I got it correctly. We appreciate your feedback. I also agree that the change of metric is unlikely to solve the problem. That is why we see it as secondary at this point. This doesn't mean that there is a metric that is monotonic wrt to sample scores and also aligned with our internal business views. This would be win-win for Home Credit and majority of the kagglers in this competition as they would feel that the new metric is safe to work with.</p>",
              "rawMarkdown": "@boristown I used google translate so I hope I got it correctly. We appreciate your feedback. I also agree that the change of metric is unlikely to solve the problem. That is why we see it as secondary at this point. This doesn't mean that there is a metric that is monotonic wrt to sample scores and also aligned with our internal business views. This would be win-win for Home Credit and majority of the kagglers in this competition as they would feel that the new metric is safe to work with."
            },
            {
              "id": 2663311,
              "postDate": "2024-02-22T12:16:35.877Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2663313,
              "postDate": "2024-02-22T12:17:18.097Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2652701,
      "postDate": "2024-02-15T00:23:13.247Z",
      "content": "<p>Whatever the outcome is, please make sure you are able to re-score previous solutions. Otherwise a lot of experiments need unnecessary repetition and its a waste of effort.</p>",
      "rawMarkdown": "Whatever the outcome is, please make sure you are able to re-score previous solutions. Otherwise a lot of experiments need unnecessary repetition and its a waste of effort.",
      "votes": 13,
      "replies": [
        {
          "id": 2653063,
          "postDate": "2024-02-15T06:53:59.573Z",
          "content": "<p>Yep, super important</p>",
          "rawMarkdown": "Yep, super important",
          "votes": 2
        },
        {
          "id": 2653996,
          "postDate": "2024-02-15T19:00:34.940Z",
          "content": "<p>Agree, despite I have none yet</p>",
          "rawMarkdown": "Agree, despite I have none yet",
          "votes": 1
        },
        {
          "id": 2657228,
          "postDate": "2024-02-18T11:43:41.767Z",
          "content": "<p>I agree. I was disheartened that my experiments and models could not improve beyond 0.5. After seeing this, my model may have steeper slope and I hope the revised metric will show their true capabilities.</p>",
          "rawMarkdown": "I agree. I was disheartened that my experiments and models could not improve beyond 0.5. After seeing this, my model may have steeper slope and I hope the revised metric will show their true capabilities.",
          "votes": 2
        },
        {
          "id": 2658288,
          "postDate": "2024-02-19T05:30:52.720Z",
          "content": "<p>Agreed!!!!</p>",
          "rawMarkdown": "Agreed!!!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2651381,
      "postDate": "2024-02-14T05:59:21.363Z",
      "content": "<h1>Covid, Leaderboard &amp; the Metric</h1>\n<p>I wonder if this is really a metric issue, or a fundamental fact that hosts cannot probe the future, while we can probe the LB, and that one cannot model COVID default rates with pre-COVID data.</p>\n<p>Some thoughts:</p>\n<ul>\n<li>COVID was a macro event that changed the data generation process &amp; ideal credit scoring function so much, that no amount of historical data would give us a stable cross-validation. The best way to check the actual model performance on the Private LB is to probe the Public LB, no matter how small % it constitutes (in this context I think it would be in hosts' interest to reveal the split method of private/public LB - just to let Kagglers know if probing LB makes sense or not. Without clear indication, most people will use LB above CV - and maybe waste their time maybe not - but hosts will not benefit from it).</li>\n<li>While we can probe LB, hosts cannot probe the future while building a model, so models which benefit from LB probing are of no use to hosts.</li>\n<li>Also, COVID chances or happening again are rather small (hopefully). Even if a next pandemic happens, it will be again such a table-turning event, that data generation process &amp; ideal credit scoring function will change so much, that models built in this competition will be of no use.</li>\n<li>Will the winning models built with the current setup of train set (pre-covid and early covid) and eval set (late covid &amp; post covid) be useful at all for the 'new normal'?</li>\n</ul>\n<p><strong>Maybe it is better to give us access to the post Covid data, because models built with it will be useful for longer to the hosts?</strong> I understand it takes a lot of time for a credit to verify the default, but we can validate with short-term loans (like 3-6m), or we can change the target definition (default within first 3-6 months), or we can just wait for the new data to come - there have been competitions on Kaggle which took really long - look at the <a href=\"https://www.kaggle.com/competitions/zillow-prize-1\" target=\"_blank\">Zillow competition</a> for example, which lasted for almost 2 years.</p>\n<p>I don't have a clear solution yet, just throwing my 2 cents, but given the hosts engagement I would like this competition to produce something really useful for them.</p>",
      "rawMarkdown": "# Covid, Leaderboard & the Metric\n\nI wonder if this is really a metric issue, or a fundamental fact that hosts cannot probe the future, while we can probe the LB, and that one cannot model COVID default rates with pre-COVID data.\n\nSome thoughts:\n- COVID was a macro event that changed the data generation process & ideal credit scoring function so much, that no amount of historical data would give us a stable cross-validation. The best way to check the actual model performance on the Private LB is to probe the Public LB, no matter how small % it constitutes (in this context I think it would be in hosts' interest to reveal the split method of private/public LB - just to let Kagglers know if probing LB makes sense or not. Without clear indication, most people will use LB above CV - and maybe waste their time maybe not - but hosts will not benefit from it).\n- While we can probe LB, hosts cannot probe the future while building a model, so models which benefit from LB probing are of no use to hosts.\n- Also, COVID chances or happening again are rather small (hopefully). Even if a next pandemic happens, it will be again such a table-turning event, that data generation process & ideal credit scoring function will change so much, that models built in this competition will be of no use.\n- Will the winning models built with the current setup of train set (pre-covid and early covid) and eval set (late covid & post covid) be useful at all for the 'new normal'?\n\n**Maybe it is better to give us access to the post Covid data, because models built with it will be useful for longer to the hosts?** I understand it takes a lot of time for a credit to verify the default, but we can validate with short-term loans (like 3-6m), or we can change the target definition (default within first 3-6 months), or we can just wait for the new data to come - there have been competitions on Kaggle which took really long - look at the [Zillow competition](https://www.kaggle.com/competitions/zillow-prize-1) for example, which lasted for almost 2 years.\n\nI don't have a clear solution yet, just throwing my 2 cents, but given the hosts engagement I would like this competition to produce something really useful for them.",
      "votes": 12,
      "replies": [
        {
          "id": 2657834,
          "postDate": "2024-02-18T19:07:20.470Z",
          "content": "<p>The broader problem you mention is the more interesting thing to solve in 2024, IMO. How do you maximize the useful signal from pre-anomaly data but mute the part that doesn't apply going forward? It's a big deal in healthcare right now for time-series problems. I'm guessing the good folks at Home Credit have thought a lot about it and how the scoring metric relates.</p>\n<p>It would be cool if a competition could somehow include metric development so that the winning model depended on a team's choice of objective and eval functions. But what metric would you use for the leaderboard?🤨 For now I guess iteration on metrics and ultimate evaluation has to be done outside of Kaggle World. </p>",
          "rawMarkdown": "The broader problem you mention is the more interesting thing to solve in 2024, IMO. How do you maximize the useful signal from pre-anomaly data but mute the part that doesn't apply going forward? It's a big deal in healthcare right now for time-series problems. I'm guessing the good folks at Home Credit have thought a lot about it and how the scoring metric relates.\n\nIt would be cool if a competition could somehow include metric development so that the winning model depended on a team's choice of objective and eval functions. But what metric would you use for the leaderboard?🤨 For now I guess iteration on metrics and ultimate evaluation has to be done outside of Kaggle World. ",
          "votes": 3,
          "replies": [
            {
              "id": 2657865,
              "postDate": "2024-02-18T19:40:42Z",
              "content": "<p>To your point on time intervals <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>:  If I understand correctly, we're trying to predict defaults post-covid without post-covid data in the public test set? That would be consistent with HC's focus on stable models (for the next big bad thing). I would think then a winning model pays a lot more attention to conservative estimation than to accuracy and leads to the metric problem we see now.</p>\n<p>That's actually OK in the real world. Models geared toward a worst-case scenario are often chosen over the most-likely scenario when volatility increases. It doesn't really work for Kaggle though.</p>\n<p>I could be wrong, of course. I still plan though to protect myself financially assuming HC is predicting a recession as that next big disruption. 😁</p>",
              "rawMarkdown": "To your point on time intervals @narsil:  If I understand correctly, we're trying to predict defaults post-covid without post-covid data in the public test set? That would be consistent with HC's focus on stable models (for the next big bad thing). I would think then a winning model pays a lot more attention to conservative estimation than to accuracy and leads to the metric problem we see now.\n\nThat's actually OK in the real world. Models geared toward a worst-case scenario are often chosen over the most-likely scenario when volatility increases. It doesn't really work for Kaggle though.\n\nI could be wrong, of course. I still plan though to protect myself financially assuming HC is predicting a recession as that next big disruption. 😁",
              "votes": 3
            },
            {
              "id": 2657866,
              "postDate": "2024-02-18T19:41:44.697Z",
              "content": "<p><a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> Thanks for appreciating the idea. The initial idea was to come up with model+metric design that would be scored by profitability in the long run. There were just too many cons for the initial idea so we had to go with the simplified version of it. The stability metric itself is an approximation of the decisions made when observing deployed models, but it is sufficient for the sake of this competition. People who say there is a disconnect between the metric and the competition idea perhaps don't see the importance of stable performance.</p>\n<p>We have a special prize for better implementation of stability in the development. I can imagine if someone comes up with a better stability metric that you can directly optimize in some kind of workflow to come up with a way of extracting the useful signal, he/she will be a good candidate for winning this special prize. </p>",
              "rawMarkdown": "@jpmiller Thanks for appreciating the idea. The initial idea was to come up with model+metric design that would be scored by profitability in the long run. There were just too many cons for the initial idea so we had to go with the simplified version of it. The stability metric itself is an approximation of the decisions made when observing deployed models, but it is sufficient for the sake of this competition. People who say there is a disconnect between the metric and the competition idea perhaps don't see the importance of stable performance.\n\nWe have a special prize for better implementation of stability in the development. I can imagine if someone comes up with a better stability metric that you can directly optimize in some kind of workflow to come up with a way of extracting the useful signal, he/she will be a good candidate for winning this special prize. ",
              "votes": 3
            },
            {
              "id": 2658174,
              "postDate": "2024-02-19T03:33:19.347Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 2658508,
              "postDate": "2024-02-19T08:36:05.107Z",
              "content": "<p>I have essentially two points:</p>\n<ul>\n<li>Asking to predict post-covid default rates with pre-covid data is a form of a very strong stress-testing. I am curious (but not sure) that models which best survive such stress-testing will be the most useful to the hosts, but it is their call, as they know their business best.</li>\n<li>Not revealing the public vs private test split + knowledge that LB score allows some access (albeit limited) to the future will practically mean that people will make many decisions based on LB score. We all know what is the consequence of that. Also - this is impractical since hosts don't have limited access to the future when building their models. So the process leading to the best model could likely not be reproduced by hosts when building their models.</li>\n</ul>",
              "rawMarkdown": "I have essentially two points:\n- Asking to predict post-covid default rates with pre-covid data is a form of a very strong stress-testing. I am curious (but not sure) that models which best survive such stress-testing will be the most useful to the hosts, but it is their call, as they know their business best.\n- Not revealing the public vs private test split + knowledge that LB score allows some access (albeit limited) to the future will practically mean that people will make many decisions based on LB score. We all know what is the consequence of that. Also - this is impractical since hosts don't have limited access to the future when building their models. So the process leading to the best model could likely not be reproduced by hosts when building their models.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2651091,
      "postDate": "2024-02-13T22:51:04.597Z",
      "content": "<p>It shouldn't be a surprise that a metric that gives an incentive to make some predictions worse, intentionally or otherwise, is problematic. Unless there is some reason to want worse performance on those samples, it is an obvious indication that the metric does not align with one's actual goals. </p>\n<p>Kaggle really ought to have standards in place for the competition metrics that it uses. A good rule that would avoid this kind of problem would be that the gradient of user scores wrt to sample scores must be strictly non-negative.</p>\n<p>Punishing score variability without breaking this rule can easily be done by using a loss term that grows rapidly on bad sample scores ie: mean( -log( gini  ) ) or even mean( (-log( gini))**exponent ) with exponent&gt;1. This has the benefit of punishing outliers with very bad performance without punishing those with very good performance. </p>\n<p>Punishing longer periods of poor scores can also be done without breaking the rule. For example one could use mean( -log( MovingAverage(gini) ) ). This metric is punishing to periods of consistent poor performance without providing false incentives or punishing good scores. </p>\n<p>The choice of initial metric for this competition was careless to a degree that should not be accepted. Kaggle should adopt rules that do not allow for these kinds of mistakes in the future.</p>",
      "rawMarkdown": "It shouldn't be a surprise that a metric that gives an incentive to make some predictions worse, intentionally or otherwise, is problematic. Unless there is some reason to want worse performance on those samples, it is an obvious indication that the metric does not align with one's actual goals. \n\nKaggle really ought to have standards in place for the competition metrics that it uses. A good rule that would avoid this kind of problem would be that the gradient of user scores wrt to sample scores must be strictly non-negative.\n\nPunishing score variability without breaking this rule can easily be done by using a loss term that grows rapidly on bad sample scores ie: mean( -log( gini  ) ) or even mean( (-log( gini))**exponent ) with exponent>1. This has the benefit of punishing outliers with very bad performance without punishing those with very good performance. \n\nPunishing longer periods of poor scores can also be done without breaking the rule. For example one could use mean( -log( MovingAverage(gini) ) ). This metric is punishing to periods of consistent poor performance without providing false incentives or punishing good scores. \n\nThe choice of initial metric for this competition was careless to a degree that should not be accepted. Kaggle should adopt rules that do not allow for these kinds of mistakes in the future.",
      "votes": 9,
      "replies": [
        {
          "id": 2652299,
          "postDate": "2024-02-14T17:07:13.180Z",
          "content": "<blockquote>\n  <p>The choice of initial metric for this competition was careless to a degree that should not be accepted. Kaggle should adopt rules that do not allow for these kinds of mistakes in the future.</p>\n</blockquote>\n<p>A warmup period (eg 1-2 weeks) prior the official competition lunch, where kagglers check the dataset/metric might help for spotting such issues (for \"non-trivial\" challenges).</p>",
          "rawMarkdown": "> The choice of initial metric for this competition was careless to a degree that should not be accepted. Kaggle should adopt rules that do not allow for these kinds of mistakes in the future.\n\nA warmup period (eg 1-2 weeks) prior the official competition lunch, where kagglers check the dataset/metric might help for spotting such issues (for \"non-trivial\" challenges).",
          "votes": 1
        }
      ]
    },
    {
      "id": 2660882,
      "postDate": "2024-02-20T21:31:21.227Z",
      "content": "<p>One of the issues I commented on before is that you can have a case where Model A has a higher Gini than Model B for every week, and yet would score worse.  No matter how much you desire stability, surely you shouldn't desire Model B over Model A, so any metric that prefers Model B cannot fit the business case well.</p>\n<p>To prevent this failure mode, how about we simply use the average of N worst weekly Ginis as the score (with N chosen from some experimentation)?  Evaluating the model on its worst weeks will implicitly penalize lack of stability, but not to the point that a strictly worse model could come out better solely on the strength of being very consistently awful at predicting things.</p>",
      "rawMarkdown": "One of the issues I commented on before is that you can have a case where Model A has a higher Gini than Model B for every week, and yet would score worse.  No matter how much you desire stability, surely you shouldn't desire Model B over Model A, so any metric that prefers Model B cannot fit the business case well.\n\nTo prevent this failure mode, how about we simply use the average of N worst weekly Ginis as the score (with N chosen from some experimentation)?  Evaluating the model on its worst weeks will implicitly penalize lack of stability, but not to the point that a strictly worse model could come out better solely on the strength of being very consistently awful at predicting things.",
      "votes": 5,
      "replies": [
        {
          "id": 2663049,
          "postDate": "2024-02-22T09:35:39.567Z",
          "content": "<p>Actually, it is possible that we would desire such model. Examples of gini in time A: [0.3, 0.3, 0.3], B: [0.8, 0.6, 0.4], we would still prefer model A here. </p>",
          "rawMarkdown": "Actually, it is possible that we would desire such model. Examples of gini in time A: [0.3, 0.3, 0.3], B: [0.8, 0.6, 0.4], we would still prefer model A here. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2651109,
      "postDate": "2024-02-13T23:52:56.703Z",
      "content": "<p>Why not remove this penalty metric and just use the auc to evaluate scores.</p>",
      "rawMarkdown": "Why not remove this penalty metric and just use the auc to evaluate scores.",
      "votes": 5,
      "replies": [
        {
          "id": 2652168,
          "postDate": "2024-02-14T15:44:47.697Z",
          "content": "<p>Or assign higher weights to weekly auc scores more distant in time.</p>",
          "rawMarkdown": "Or assign higher weights to weekly auc scores more distant in time.",
          "votes": 5
        },
        {
          "id": 2655278,
          "postDate": "2024-02-16T19:11:52.747Z",
          "content": "<p>Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/477074\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/477074</a></p>",
          "rawMarkdown": "Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/477074"
        }
      ]
    },
    {
      "id": 2675153,
      "postDate": "2024-02-29T18:23:04.680Z",
      "content": "<p>I find it generally weird that you would remove 'weeks' a part of what is supposed to be optimised. Or conversely that you would use a metric including weeks if you remove that info. </p>\n<p>More generally it seems rather difficult to remove all 'time' info from the test. You can spend a lot of time reworking the data without garantee. I would suggest not doing this and keeping a simple metric like avg(gini) (or even a <a href=\"https://en.wikipedia.org/wiki/Scoring_rule\" target=\"_blank\">proper scoring rule</a> ?). </p>",
      "rawMarkdown": "I find it generally weird that you would remove 'weeks' a part of what is supposed to be optimised. Or conversely that you would use a metric including weeks if you remove that info. \n\nMore generally it seems rather difficult to remove all 'time' info from the test. You can spend a lot of time reworking the data without garantee. I would suggest not doing this and keeping a simple metric like avg(gini) (or even a [proper scoring rule](https://en.wikipedia.org/wiki/Scoring_rule) ?). ",
      "votes": 1
    },
    {
      "id": 2658216,
      "postDate": "2024-02-19T04:40:22.917Z",
      "content": "<p>why dont you just use a mean with more weight on the recent values for mean gini</p>",
      "rawMarkdown": "why dont you just use a mean with more weight on the recent values for mean gini",
      "votes": 1
    },
    {
      "id": 2653082,
      "postDate": "2024-02-15T07:18:46.117Z",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>  I can suggest that using Gini as evaluation metric on LB and stability score for stability prize.  HC can restrict stability prize like one can win it if she/he in top %1-2 on private LB and stability score should be calculated with the submission which has  maximum gini score among contender's 2 final submissions (people still could hack the metric with two submission, 1 for gini score other one is for stability score).  I believe that will be beneficial for both parties (HC and kagglers ).</p>",
      "rawMarkdown": "@tomasjeline2 @jetakow  I can suggest that using Gini as evaluation metric on LB and stability score for stability prize.  HC can restrict stability prize like one can win it if she/he in top %1-2 on private LB and stability score should be calculated with the submission which has  maximum gini score among contender's 2 final submissions (people still could hack the metric with two submission, 1 for gini score other one is for stability score).  I believe that will be beneficial for both parties (HC and kagglers ).",
      "votes": 1,
      "replies": [
        {
          "id": 2655270,
          "postDate": "2024-02-16T19:08:26.230Z",
          "content": "<p>I am afraid that would not be a satisfactory solution on behalf of HC. This year the competition is about stability, without having the stability directly in the evaluation metric we would have little control over the stability of the submitted solutions. </p>",
          "rawMarkdown": "I am afraid that would not be a satisfactory solution on behalf of HC. This year the competition is about stability, without having the stability directly in the evaluation metric we would have little control over the stability of the submitted solutions. ",
          "votes": 1,
          "replies": [
            {
              "id": 2655404,
              "postDate": "2024-02-16T20:38:58.970Z",
              "content": "<p>without changing to an usual metric (i.e. a metric that is always worse when an individual gini score get's worse) the result will be solutions that are not stable but just carefully chosen hacks or maybe even more possible purely lucky hacks.<br>\nthere are already some suggestions for metric whcih satisfy the above condition and still encourage stability over time.</p>",
              "rawMarkdown": "without changing to an usual metric (i.e. a metric that is always worse when an individual gini score get's worse) the result will be solutions that are not stable but just carefully chosen hacks or maybe even more possible purely lucky hacks.\nthere are already some suggestions for metric whcih satisfy the above condition and still encourage stability over time."
            },
            {
              "id": 2655447,
              "postDate": "2024-02-16T21:44:32.010Z",
              "content": "<p>We are working on a fix that will discourage such hacking. We are less likely to directly change the metric as now it is a good approximation of how risk managers \"rate\" our models in production. I must ask you for patience. </p>",
              "rawMarkdown": "We are working on a fix that will discourage such hacking. We are less likely to directly change the metric as now it is a good approximation of how risk managers \"rate\" our models in production. I must ask you for patience. "
            }
          ]
        }
      ]
    },
    {
      "id": 2650940,
      "postDate": "2024-02-13T19:04:24.210Z",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> kindly inform us in advance as this is likely to impact our plans/ ideas/ models for the competition. <br>\nThanks for being so proactive!</p>",
      "rawMarkdown": "@tomasjeline2 kindly inform us in advance as this is likely to impact our plans/ ideas/ models for the competition. \nThanks for being so proactive!",
      "votes": 1
    },
    {
      "id": 2658425,
      "postDate": "2024-02-19T07:17:28.027Z",
      "content": "<p>$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1}\\sqrt{ \\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{peak} - GiniScore_i)\\right)^2}<br>\n$$</p>\n<p>GiniScore_peak : is the highest Gini score observed up to week i <br>\nGiniScore_i : is the Gini score for week i </p>\n<p>I modified the  ulcer index from finance domain (<a href=\"https://en.wikipedia.org/wiki/Ulcer_index\" target=\"_blank\">https://en.wikipedia.org/wiki/Ulcer_index</a>)</p>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> </p>",
      "rawMarkdown": "$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1}\\sqrt{ \\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{peak} - GiniScore_i)\\right)^2}\n$$\n\nGiniScore_peak : is the highest Gini score observed up to week i \nGiniScore_i : is the Gini score for week i \n\nI modified the  ulcer index from finance domain (https://en.wikipedia.org/wiki/Ulcer_index)\n\n@jetakow @tomasjeline2 ",
      "votes": 2,
      "replies": [
        {
          "id": 2658652,
          "postDate": "2024-02-19T10:02:30.473Z",
          "content": "<p>Thanks for the idea. We wanted to make the metric as robust as possible and at the same time as simple as possible. Since you have GinniScore_peak in the metric it would be susceptible in different type of hacking. That would result in residues (of weekly gini in respect to y=ax+b) with a skew whereas we don't want to penalize less the weeks where the performance is worse. In a sense the modified Ulcer index is similar to the term with std(gini) except that it doesn't penalize equally the negative drops in performance and positive changes in performance (in weekly gini). There is no overall trend as a term, we definitely don't want to promote performance that is on a stable decline.  </p>",
          "rawMarkdown": "Thanks for the idea. We wanted to make the metric as robust as possible and at the same time as simple as possible. Since you have GinniScore_peak in the metric it would be susceptible in different type of hacking. That would result in residues (of weekly gini in respect to y=ax+b) with a skew whereas we don't want to penalize less the weeks where the performance is worse. In a sense the modified Ulcer index is similar to the term with std(gini) except that it doesn't penalize equally the negative drops in performance and positive changes in performance (in weekly gini). There is no overall trend as a term, we definitely don't want to promote performance that is on a stable decline.  ",
          "replies": [
            {
              "id": 2658670,
              "postDate": "2024-02-19T10:24:08.583Z",
              "content": "<p>Note that GiniScore_peak is rolling max not absolute max  and  when someone tries to hack it, he/she will end up lowering the mean(GiniScore) at almost same degree I think that makes it more robust than current one.  HC used 88.0⋅𝑚𝑖𝑛(0,𝑎) term, this also does not penalize equally for positive and negative changes. The term in the sqrt behaves like a trend  because of comparing with rolling max of GiniScore value, Thus,  it gives penalty for negative gini performance slope, 0 penalty for positive (and steady performance) slopes.  You can test the metric with several scenarios. </p>",
              "rawMarkdown": "Note that GiniScore_peak is rolling max not absolute max  and  when someone tries to hack it, he/she will end up lowering the mean(GiniScore) at almost same degree I think that makes it more robust than current one.  HC used 88.0⋅𝑚𝑖𝑛(0,𝑎) term, this also does not penalize equally for positive and negative changes. The term in the sqrt behaves like a trend  because of comparing with rolling max of GiniScore value, Thus,  it gives penalty for negative gini performance slope, 0 penalty for positive (and steady performance) slopes.  You can test the metric with several scenarios. ",
              "votes": 3
            },
            {
              "id": 2658716,
              "postDate": "2024-02-19T11:00:00.710Z",
              "content": "<p>I will test it, thanks for the suggestion.</p>",
              "rawMarkdown": "I will test it, thanks for the suggestion."
            },
            {
              "id": 2658882,
              "postDate": "2024-02-19T13:25:04.013Z",
              "content": "<p>i have modified the formula (noticed an error) and tried a couple of scenarios, it seems OK so far</p>\n<p>$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1} \\sqrt{\\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{rollingMax} - GiniScore_i)\\right)^2}<br>\n$$</p>\n<p>GiniScore_rollingMax : is the highest Gini score observed up to week i <br>\nGiniScore_i : is the Gini score for week i </p>\n<p>a = [0.60, 0.66, 0.65, 0.70, 0.64, 0.61, 0.64, 0.65, 0.64]    -- benchmark--<br>\nb = [0.55, 0.60, 0.60, 0.65, 0.64, 0.61, 0.64, 0.65, 0.64]    worsened first 4 weeks by 0.05<br>\nc = [0.60, 0.61, 0.64, 0.64, 0.64, 0.65, 0.65, 0.66, 0.70]    max stable scenario with benchmark gini scores <br>\nd = [0.59, 0.65, 0.64, 0.69, 0.63, 0.60, 0.63, 0.64, 0.63]   worsened all weeks by 0.01<br>\ne = [0.70, 0.66, 0.65, 0.65, 0.64, 0.64, 0.64, 0.61, 0.60]    worst scenario with benchmark gini scores <br>\nf = [0.58, 0.64, 0.63, 0.67, 0.64, 0.61, 0.64, 0.65, 0.64]    worsened first 4 weeks by 0.02  <br>\ng = [0.56, 0.62, 0.62, 0.62, 0.64, 0.61, 0.64, 0.65, 0.64]  worsened first 4 weeks with random values <br>\nh = [0.56, 0.62, 0.62, 0.62, 0.63, 0.63, 0.63, 0.63, 0.63]  worsened all weeks (monotonic increasing)<br>\ni = [0.62, 0.61, 0.60, 0.65, 0.64, 0.64, 0.64, 0.61, 0.60]    worsened first 4 weeks for scenario e <br>\nj = [0.65, 0.62, 0.61, 0.63, 0.64, 0.64, 0.64, 0.61, 0.60]   worsened first 4 weeks for scenario e</p>\n<p>a : 0.6250<br>\nb: 0.6145<br>\nc: 0.6433<br>\nd: 0.6150<br>\ne: 0.6197<br>\nf: 0.6230<br>\ng: 0.6182<br>\nh: 0.6188<br>\ni: 0.6145<br>\nj: 0.6159</p>",
              "rawMarkdown": "i have modified the formula (noticed an error) and tried a couple of scenarios, it seems OK so far\n\n$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1} \\sqrt{\\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{rollingMax} - GiniScore_i)\\right)^2}\n$$\n\nGiniScore_rollingMax : is the highest Gini score observed up to week i \nGiniScore_i : is the Gini score for week i \n\na = [0.60, 0.66, 0.65, 0.70, 0.64, 0.61, 0.64, 0.65, 0.64]    -- benchmark--\nb = [0.55, 0.60, 0.60, 0.65, 0.64, 0.61, 0.64, 0.65, 0.64]    worsened first 4 weeks by 0.05\nc = [0.60, 0.61, 0.64, 0.64, 0.64, 0.65, 0.65, 0.66, 0.70]    max stable scenario with benchmark gini scores \nd = [0.59, 0.65, 0.64, 0.69, 0.63, 0.60, 0.63, 0.64, 0.63]   worsened all weeks by 0.01\ne = [0.70, 0.66, 0.65, 0.65, 0.64, 0.64, 0.64, 0.61, 0.60]    worst scenario with benchmark gini scores \nf = [0.58, 0.64, 0.63, 0.67, 0.64, 0.61, 0.64, 0.65, 0.64]    worsened first 4 weeks by 0.02  \ng = [0.56, 0.62, 0.62, 0.62, 0.64, 0.61, 0.64, 0.65, 0.64]  worsened first 4 weeks with random values \nh = [0.56, 0.62, 0.62, 0.62, 0.63, 0.63, 0.63, 0.63, 0.63]  worsened all weeks (monotonic increasing)\ni = [0.62, 0.61, 0.60, 0.65, 0.64, 0.64, 0.64, 0.61, 0.60]    worsened first 4 weeks for scenario e \nj = [0.65, 0.62, 0.61, 0.63, 0.64, 0.64, 0.64, 0.61, 0.60]   worsened first 4 weeks for scenario e\n\n\na : 0.6250\nb: 0.6145\nc: 0.6433\nd: 0.6150\ne: 0.6197\nf: 0.6230\ng: 0.6182\nh: 0.6188\ni: 0.6145\nj: 0.6159\n",
              "votes": 3
            },
            {
              "id": 2660103,
              "postDate": "2024-02-20T10:47:34Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> i have tested the metric I suggested  for 100M  random scenarios, </p>\n<p>number of weeks : 18<br>\ninitialized the gini scores by week between [0.55, 0.65]<br>\ncreated a mask which has random size between [3,9] -- refers to which first N weeks to worsen the gini scores. mask has values  between [0.9 , 1.0]  If    mask[4] = 0.95 that mean I worsen 5th week' gini by %5.  </p>\n<p>here is the result I get:</p>\n<ul>\n<li>for only %0.27 cases, one can improve the metric by hacking </li>\n<li>for rest of the cases, score is distorted </li>\n<li>maximum improvement one can get by hacking is 0.0070 for 18 weeks, but when I increase the number of weeks it is getting harder to hack so  maximum improvement one can get  decreases.</li>\n<li>median improvement among %0.27 cases  is  0.0005 (for 18 weeks)</li>\n</ul>\n<p>improvement quantiles (for %0.27 of  cases)</p>\n<p>Quantile 10 8.898260757736014e-05<br>\nQuantile 20 0.00018749605249894266<br>\nQuantile 30 0.00029993356066040623<br>\nQuantile 40 0.0004269924733614792<br>\nQuantile 50 0.0005784240749482758<br>\nQuantile 60 0.0007601426463899664<br>\nQuantile 70 0.0009925642541695821<br>\nQuantile 80 0.0013097591959578612<br>\nQuantile 90 0.0018370136698832628<br>\nQuantile 100 0.006960994642879724</p>\n<p>now I am running 1B scenarios for 50 weeks. I will share the results and the code here. </p>\n<p>To hack the metric I suggested, one should know which week has  the maximum gini score otherwise the score will be distorted, I think no one would try to hack it.  </p>",
              "rawMarkdown": "@jetakow i have tested the metric I suggested  for 100M  random scenarios, \n\n\nnumber of weeks : 18\ninitialized the gini scores by week between [0.55, 0.65]\ncreated a mask which has random size between [3,9] -- refers to which first N weeks to worsen the gini scores. mask has values  between [0.9 , 1.0]  If    mask[4] = 0.95 that mean I worsen 5th week' gini by %5.  \n\n\nhere is the result I get:\n\n- for only %0.27 cases, one can improve the metric by hacking \n- for rest of the cases, score is distorted \n- maximum improvement one can get by hacking is 0.0070 for 18 weeks, but when I increase the number of weeks it is getting harder to hack so  maximum improvement one can get  decreases.\n- median improvement among %0.27 cases  is  0.0005 (for 18 weeks)\n\nimprovement quantiles (for %0.27 of  cases)\n\nQuantile 10 8.898260757736014e-05\nQuantile 20 0.00018749605249894266\nQuantile 30 0.00029993356066040623\nQuantile 40 0.0004269924733614792\nQuantile 50 0.0005784240749482758\nQuantile 60 0.0007601426463899664\nQuantile 70 0.0009925642541695821\nQuantile 80 0.0013097591959578612\nQuantile 90 0.0018370136698832628\nQuantile 100 0.006960994642879724\n\nnow I am running 1B scenarios for 50 weeks. I will share the results and the code here. \n\nTo hack the metric I suggested, one should know which week has  the maximum gini score otherwise the score will be distorted, I think no one would try to hack it.  \n",
              "votes": 3
            },
            {
              "id": 2660141,
              "postDate": "2024-02-20T11:25:42.563Z",
              "content": "<p>Why choose this metric if <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a> suggested already a metric which is provably non hackable.</p>",
              "rawMarkdown": "Why choose this metric if @jacobyjaeger suggested already a metric which is provably non hackable."
            },
            {
              "id": 2660169,
              "postDate": "2024-02-20T12:00:24.467Z",
              "content": "<p>now I have tested  Jacob's  metric with 1M scenarios, it seems  it is hackable and one can get up to  0.038  improvement (final score range is like 0.25-0.30.  0.038  means  ~ %13 improvement)    </p>\n<p>score range : (0.25-0.30)<br>\nImprovement Quantiles<br>\nQuantile 5 0.0018123009184825712<br>\nQuantile 10 0.002954071577671077<br>\nQuantile 15 0.0042528432747884854<br>\nQuantile 20 0.005571776912732721<br>\nQuantile 25 0.006885412084717121<br>\nQuantile 30 0.008219234721494929<br>\nQuantile 35 0.009542760995397888<br>\nQuantile 40 0.01087649085647983<br>\nQuantile 45 0.012200174484067327<br>\nQuantile 50 0.01351322138801158<br>\nQuantile 55 0.014840927605872758<br>\nQuantile 60 0.016168883974780883<br>\nQuantile 65 0.017497915477009143<br>\nQuantile 70 0.018820113281295393<br>\nQuantile 75 0.02017423466541196<br>\nQuantile 80 0.021564749464102084<br>\nQuantile 85 0.023043430204338983<br>\nQuantile 90 0.024716932816428373<br>\nQuantile 95 0.02689879395363132<br>\nQuantile 100 0.03894322510151606</p>\n<p>i have tested mine with 200M scenarios and  50 weeks:  <br>\nfor  %0.021 cases it is hackable and most improvement one can get  : 0.0021<br>\nmedian improvement (for %0.021 cases) : 0.0001<br>\nscore range : (0.57- 0.64)</p>\n<p>Improvement Quantiles<br>\nQuantile 5 1.2699755546674734e-05<br>\nQuantile 10 2.6765098840468564e-05<br>\nQuantile 15 4.102476662966947e-05<br>\nQuantile 20 5.6175799148184025e-05<br>\nQuantile 25 7.213998299378261e-05<br>\nQuantile 30 8.979772687923974e-05<br>\nQuantile 35 0.00010814053350019279<br>\nQuantile 40 0.00012823791281313051<br>\nQuantile 45 0.00015023664735298696<br>\nQuantile 50 0.0001735378974960051<br>\nQuantile 55 0.00020047277886452565<br>\nQuantile 60 0.0002300174182662148<br>\nQuantile 65 0.00026180227673274555<br>\nQuantile 70 0.00029811598482456093<br>\nQuantile 75 0.0003424203629695818<br>\nQuantile 80 0.00039457860089038663<br>\nQuantile 85 0.00046220233188403397<br>\nQuantile 90 0.0005567642604209829<br>\nQuantile 95 0.0007121168562655653<br>\nQuantile 100 0.0021852690646599017</p>\n<p>if we increase the week num it is getting harder to hack.<br>\ni have attached code for mine and Jacob's metric. </p>",
              "rawMarkdown": "now I have tested  Jacob's  metric with 1M scenarios, it seems  it is hackable and one can get up to  0.038  improvement (final score range is like 0.25-0.30.  0.038  means  ~ %13 improvement)    \n\nscore range : (0.25-0.30)\nImprovement Quantiles\nQuantile 5 0.0018123009184825712\nQuantile 10 0.002954071577671077\nQuantile 15 0.0042528432747884854\nQuantile 20 0.005571776912732721\nQuantile 25 0.006885412084717121\nQuantile 30 0.008219234721494929\nQuantile 35 0.009542760995397888\nQuantile 40 0.01087649085647983\nQuantile 45 0.012200174484067327\nQuantile 50 0.01351322138801158\nQuantile 55 0.014840927605872758\nQuantile 60 0.016168883974780883\nQuantile 65 0.017497915477009143\nQuantile 70 0.018820113281295393\nQuantile 75 0.02017423466541196\nQuantile 80 0.021564749464102084\nQuantile 85 0.023043430204338983\nQuantile 90 0.024716932816428373\nQuantile 95 0.02689879395363132\nQuantile 100 0.03894322510151606\n\n\n\n\ni have tested mine with 200M scenarios and  50 weeks:  \nfor  %0.021 cases it is hackable and most improvement one can get  : 0.0021\nmedian improvement (for %0.021 cases) : 0.0001\nscore range : (0.57- 0.64)\n\nImprovement Quantiles\nQuantile 5 1.2699755546674734e-05\nQuantile 10 2.6765098840468564e-05\nQuantile 15 4.102476662966947e-05\nQuantile 20 5.6175799148184025e-05\nQuantile 25 7.213998299378261e-05\nQuantile 30 8.979772687923974e-05\nQuantile 35 0.00010814053350019279\nQuantile 40 0.00012823791281313051\nQuantile 45 0.00015023664735298696\nQuantile 50 0.0001735378974960051\nQuantile 55 0.00020047277886452565\nQuantile 60 0.0002300174182662148\nQuantile 65 0.00026180227673274555\nQuantile 70 0.00029811598482456093\nQuantile 75 0.0003424203629695818\nQuantile 80 0.00039457860089038663\nQuantile 85 0.00046220233188403397\nQuantile 90 0.0005567642604209829\nQuantile 95 0.0007121168562655653\nQuantile 100 0.0021852690646599017\n\nif we increase the week num it is getting harder to hack.\ni have attached code for mine and Jacob's metric. \n\n",
              "votes": 2
            },
            {
              "id": 2660319,
              "postDate": "2024-02-20T14:30:50.200Z",
              "content": "<p>It is probably non hackable by decreasing gini score. Please see the other discussion for explanation. If you look carefully at the formula you will see that every decreasing of gini score will make the metric worse.</p>",
              "rawMarkdown": "It is probably non hackable by decreasing gini score. Please see the other discussion for explanation. If you look carefully at the formula you will see that every decreasing of gini score will make the metric worse."
            },
            {
              "id": 2660498,
              "postDate": "2024-02-20T16:29:56.317Z",
              "content": "<p>Hacking the metric is done by decreasing gini score for half weeks. See the main thread here:</p>\n<p><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449</a></p>",
              "rawMarkdown": "Hacking the metric is done by decreasing gini score for half weeks. See the main thread here:\n\nhttps://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449"
            },
            {
              "id": 2660563,
              "postDate": "2024-02-20T17:11:59.497Z",
              "content": "<p>I was talking about metric proposed by above user in another thread:<br>\nmean( ( -log( MovingAverage(gini) ) )**exponent ) with exponent&gt;=1</p>\n<p>I don’t see how this metric can be improved by artificially lowering any gini scores. If I am wrong, please correct me.</p>",
              "rawMarkdown": "I was talking about metric proposed by above user in another thread:\nmean( ( -log( MovingAverage(gini) ) )**exponent ) with exponent>=1\n\nI don’t see how this metric can be improved by artificially lowering any gini scores. If I am wrong, please correct me."
            },
            {
              "id": 2660616,
              "postDate": "2024-02-20T17:50:56.267Z",
              "content": "<p>See the test code i attached.  i generated 1M scenario which artificially lowering gini scores. All cases improved the stability scores (up to %13 improvement ) </p>",
              "rawMarkdown": "See the test code i attached.  i generated 1M scenario which artificially lowering gini scores. All cases improved the stability scores (up to %13 improvement ) "
            },
            {
              "id": 2660621,
              "postDate": "2024-02-20T17:59:19.417Z",
              "content": "<p>Are you aware of the fact, that in case of this metric lower is better? I can’t think of a reason why any of the moving averages should get higher by lowering gini scores. If you decrease the ginis the averages will get smaller, that means the -log of that will get larger and the same total mean over all time windows will increase as well which means the score  is worse.  Or am I wrong?</p>",
              "rawMarkdown": "Are you aware of the fact, that in case of this metric lower is better? I can’t think of a reason why any of the moving averages should get higher by lowering gini scores. If you decrease the ginis the averages will get smaller, that means the -log of that will get larger and the same total mean over all time windows will increase as well which means the score  is worse.  Or am I wrong?"
            },
            {
              "id": 2660632,
              "postDate": "2024-02-20T18:16:24.090Z",
              "content": "<p>if lower is better, then Jacob’s metric is fine.  %100 of cases distorted if you lower the gini scores. </p>",
              "rawMarkdown": "if lower is better, then Jacob’s metric is fine.  %100 of cases distorted if you lower the gini scores. ",
              "votes": 1
            },
            {
              "id": 2660672,
              "postDate": "2024-02-20T18:47:02.077Z",
              "content": "<p>Yes. The best possible outcome is that every Gini = 1 (for every timestep) then every -log(1)=0 and the overall score is =0.</p>",
              "rawMarkdown": "Yes. The best possible outcome is that every Gini = 1 (for every timestep) then every -log(1)=0 and the overall score is =0."
            }
          ]
        }
      ]
    },
    {
      "id": 2671312,
      "postDate": "2024-02-27T12:42:48.937Z",
      "content": "<p>When is it expected that submissions are possible again?</p>",
      "rawMarkdown": "When is it expected that submissions are possible again?"
    },
    {
      "id": 2662874,
      "postDate": "2024-02-22T07:10:16.423Z",
      "content": "<p>Astrologers lied even by chance, data science determines the optimal direction according to the data and not without it Therefore I think that it is difficult to provide stability metric that skips to the accuracy of the near or far future alike Therefore it is possible to manipulate any stability metric according to its function The best solution is to separate the evaluation into two or three parts so that the data has to be involved in it For example evaluating the person in terms of whether he is able to pay or not based on what is available in the data available in The beginning, otherwise if someone can deduce the future by more than half and succeeds in that, he will not give jealousy because in this case he has the power of a god</p>",
      "rawMarkdown": "Astrologers lied even by chance, data science determines the optimal direction according to the data and not without it Therefore I think that it is difficult to provide stability metric that skips to the accuracy of the near or far future alike Therefore it is possible to manipulate any stability metric according to its function The best solution is to separate the evaluation into two or three parts so that the data has to be involved in it For example evaluating the person in terms of whether he is able to pay or not based on what is available in the data available in The beginning, otherwise if someone can deduce the future by more than half and succeeds in that, he will not give jealousy because in this case he has the power of a god"
    },
    {
      "id": 2658036,
      "postDate": "2024-02-18T23:22:17.487Z",
      "content": "<p>If the main concern is that teams are gaming the stability metric by lowering the score of the first few weeks, one idea to keep the stability metric and also discourage lowering the score of the first weeks would be to add an additional weighted term to the current metric e.g.  0.3*mean(gini1, gini2,… giniN) for the first N weeks. This means that we want solutions that can score highly on the first N weeks and also decline very slowly overtime (which was the original intent). </p>\n<p>or perhaps replace the first term in the current metric:  mean(gini) by mean(gini1, gini2,… giniN) for N weeks.</p>",
      "rawMarkdown": "If the main concern is that teams are gaming the stability metric by lowering the score of the first few weeks, one idea to keep the stability metric and also discourage lowering the score of the first weeks would be to add an additional weighted term to the current metric e.g.  0.3*mean(gini1, gini2,... giniN) for the first N weeks. This means that we want solutions that can score highly on the first N weeks and also decline very slowly overtime (which was the original intent). \n\nor perhaps replace the first term in the current metric:  mean(gini) by mean(gini1, gini2,... giniN) for N weeks."
    },
    {
      "id": 2653519,
      "postDate": "2024-02-15T13:39:48.923Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2651259,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "2024-02-14T03:45:58.743000",
      "content": "<p>First, <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> thank you for the report. <br>\nAnd thank you to the host <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> , <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>   for sincerely considering a revision of the metric for this topic.</p>\n<p>My thought is that, perhaps, if we consider the reverse, isn't there a potential issue like in Case 3, where improving the Gini value in parts could worsen the overall score? (Even though the average value of Gini might increase, the slope becomes significantly negative, and its weight is large). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F4f77e7c26d8d69e6c9a3e2897368bc0e%2FClipboard03.jpg?generation=1707882775177471&amp;alt=media\"></p>\n<p>※ The green represents a fixed Gini value for explanation purposes. The blue is intentionally shifted for explanation purposes.</p>\n<p>And, this potential issue is incompatible with the description in the data tab below and could significantly impact the overall score.</p>\n<pre><code>It\n</code></pre>\n<p>I would appreciate it if you could also consider this potential issue. But, if that's what the host desires, I won't say anything.</p>\n<p>Personally, using polyfit makes it difficult to balance the mean value of Gini with the term of the slope, making it challenging to excel in this competition.</p>\n<p>Perhaps, instead of using polyfit, for example, weighting the Gini function for each week number, taking the sum, and then normalizing it at the end, might also be a good metric.</p>\n<p>However, this is just my opinion, so I leave the decision to the host. I would appreciate it if you could take this into consideration. Please feel free to point out any mistakes.</p>",
      "votes": 20,
      "replies": []
    },
    {
      "id": 2662222,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-21T19:13:01.210000",
      "content": "<p>Dear Kagglers, </p>\n<p>We prepared this year’s competition with an intention to present you something else and more challenging than the previous Home Credit kaggle competition. We decided to select stability of models in production as the main objective of this competition. Therefore, we will keep stability and won’t remove it from the competition.  </p>\n<p>We would like to further clarify the meaning of stability. Our ideal stable model has constant performance by weeks in time and as high of a mean gini over weeks as possible. That is the reason why we have three terms in the stability metric as of now since this is desirable in business and not only to maximize AUC. The metric itself is a good approximation of the decisions made in-house when assessing the performance of models in production. Despite different opinions here within the competition, removing the terms in the stability metric would significantly deviate from the way we rate performance internally. When people leverage the second term and nullify it by carefully choosing scores on the test sample, it does not represent how we deploy models in the production. This is undesirable and we are forced to make steps that will discourage hacking of the metric.  </p>\n<p>We will make a change to the test dataset in such a way that it will represent better client scoring in production. When the model is running in production, it can only see its past performance and predictions. This would be quite difficult to reproduce in a Kaggle competition, but we have an alternative solution. We will remove the date decision, month and week number from the test base table. This part will take two weeks to amend, and we will prolong the competition by this period. Because of this, it is very likely that we won’t be able to rescore the old submissions. We apologize in advance for this inconvenience. </p>\n<p>Besides the change of test data, we have a second plan regarding metric. At this moment we are still considering metric change, and we welcome suggestions that would be aligned with how we view the value of model stability and in a way that would prevent further hacking. We are unlikely to go with a new metric that would prevent this type of hacking and would significantly deviate from the stability part. There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.</p>\n<p>We would like to apologize for the delay and inconvenience caused by the issues connected to the stability metric. We also want to thank the community for many proposals on how to fix the issue. We have tested and analyzed many of them, thank you for all the suggestions! We will disable submissions for two weeks from now on and prolong the competition by two weeks.  </p>\n<p>Tomas &amp; Daniel</p>",
      "votes": 15,
      "replies": [
        {
          "id": 2662256,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-02-21T19:38:03.710000",
          "content": "<p>This will likely lead to a separate set of models, focused on predicting the <code>week_number</code> based on available features. Once you let the Genie out of the bottle there is no coming back :) But I get it - there are no ideal solutions. Thank you for the effort.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2662265,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-21T19:42:13.943000",
              "content": "<p>I am pretty sure as long as the metric won’t be strictly decreasing in gini scores for every time point it will lead to people hacking the leaderboard.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2662269,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-21T19:44:38.400000",
              "content": "<p>The winning models will be reviewed manually :) We want stable model/models in time, not models that predict what is the week number. This approach wouldn't be the winning one. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2662280,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-21T19:52:00.407000",
              "content": "<p>I think it will be enormous amount of people submitting such stuff. do you mean you would manually review hundreds of submissions?<br>\nAlso it’s frustrating for people who want to build better models if one never can get a real benchmark of how good the model performs on the LB due to hundreds of hacked solutions.<br>\nI don’t get why not chose a mathematical „safe“ metric which is provably not hackable by decreasing model performance.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2662295,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-21T20:06:14.907000",
              "content": "<p>I can assure you that we will do our best to detect such hacking of the metric even if someone fortunate passes the private test set. If the top 100 solutions will be using some way of week num prediction, yes we will go over those hundred solutions in the top private leaderboard until we find a solution that doesn't have this approach :) </p>\n<p>We discourage any effort that will be put into a prediction of date within the test sample. </p>\n<p>If you have any metric on your mind, don't hesitate to suggest it and I will review it personally. The main issue is that those mathematically safe metrics don't asses very well decreasing performance. I am quite optimistic nd I believe such metric exists, however, we are limited by time. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2662570,
              "author_name": "silver",
              "author_url": "",
              "post_date": "2024-02-22T02:06:24.570000",
              "content": "<p>That only covers top submissions. Most people will likely be forced to do hacking just to keep a reasonable ranking.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2663000,
              "author_name": "Ye Daotian",
              "author_url": "",
              "post_date": "2024-02-22T09:18:30.477000",
              "content": "<p>However, if you delete 'date_decision', it will cause most of the data on the date is not useful. In terms of my model, it may reduce my score about 0.01</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2663018,
              "author_name": "Ye Daotian",
              "author_url": "",
              "post_date": "2024-02-22T09:26:05.837000",
              "content": "<p>If you decide to delete \"date_decision\", I think I need other feather, like:<br>\nif col[-1] == \"D\":<br>\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))<br>\n                df = df.with_columns(pl.col(col).dt.total_days())<br>\nI hope that  original data about date changes like this.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2664509,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-23T04:43:31.757000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2662268,
          "author_name": "Dmitriy Guller",
          "author_url": "",
          "post_date": "2024-02-21T19:44:21.597000",
          "content": "<p>Hi Tomas and Daniel,</p>\n<p>To aid with the alternative metric proposals, can you publish the \"units tests\" that the metric must pass?  Can you publish multiple examples with two hypothetical week-by-week Gini behaviors, where you judge one behavior to be more desirable than the other?</p>\n<p>I think that it would be hard to propose metrics that align with your vision without something quantifiable, especially since there is probably a disconnect between what Home Credit considers desirable and what competitors here would assume to be desirable.  For better or worse, the units tests would establish a hard ground truth that we would have to respect with our proposals.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2662273,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-21T19:48:09.383000",
              "content": "<p>I am afraid we can't post such unit tests that we have in-house. However, I am willing to score any metric you will send. After scoring I could also issue examples where the proposed metric is not better when I feel it would help. What would you say to that?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2662285,
              "author_name": "Dmitriy Guller",
              "author_url": "",
              "post_date": "2024-02-21T19:56:58.513000",
              "content": "<p>I think it won't lead to effective iteration when you're trying to hit the target you can't see, and don't have instant feedback as to whether you're on the right track.</p>\n<p>Can there be a compromise solution here?  Can you publish some preliminary units tests so that we have something tangible to go on, but still make the final determination with your private unit tests?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2662297,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-21T20:10:19.613000",
              "content": "<p>It doesn't have to be effective, in the end we will be selecting the metric. I am afraid we simply can not give you the same data we used for determining the current metric. We can however use them to score any suggested metric as I said. We could derive such unit tests, but in the sake of using the time wisely, I am afraid we won't be able to spend much time on it. It would be better to spend a considerable amount of time on this activity so there is no scope for inconsistency and at the same time, we won't leak the data used for such a decision during the competition itself. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2662374,
          "author_name": "Youri Matiounine",
          "author_url": "",
          "post_date": "2024-02-21T21:05:27.063000",
          "content": "<p>\"We will remove the date decision…\": but \"date decision\" is subtracted from all the \"date\" inputs to turn them into time periods; without \"date decision\" all the date inputs immediately become less usable. And they can be used to impute the missing \"date decision\" anyway.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2662420,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-21T21:53:13.897000",
              "content": "<blockquote>\n  <p>\"We will remove the date decision…\": but \"date decision\" is subtracted from all the \"date\" inputs to turn them into time periods; without \"date decision\" all the date inputs immediately become less usable. </p>\n</blockquote>\n<p>We are aware of that, it is acceptable to us at this moment. Unless we of course in a week come up with a resonable new metric.</p>\n<blockquote>\n  <p>And they can be used to impute the missing \"date decision\" anyway.</p>\n</blockquote>\n<p>Don't worry about this part. Data will be transformed again :)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2662551,
          "author_name": "暗黑AGI",
          "author_url": "",
          "post_date": "2024-02-22T01:21:52.080000",
          "content": "<p>很抱歉我英文不好，我还是用中文表达吧。我已经观察到了测试数据的问题，并且利用这个问题拿到了public榜单的第一名。我认为修改公式并不能很好的解决这个问题。建议修改测试数据。只保留未来的部分。</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2663040,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-22T09:33:20.530000",
              "content": "<p><a href=\"https://www.kaggle.com/boristown\" target=\"_blank\">@boristown</a> I used google translate so I hope I got it correctly. We appreciate your feedback. I also agree that the change of metric is unlikely to solve the problem. That is why we see it as secondary at this point. This doesn't mean that there is a metric that is monotonic wrt to sample scores and also aligned with our internal business views. This would be win-win for Home Credit and majority of the kagglers in this competition as they would feel that the new metric is safe to work with.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2663311,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-22T12:16:35.877000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2663313,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-22T12:17:18.097000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2652701,
      "author_name": "NxGTR",
      "author_url": "",
      "post_date": "2024-02-15T00:23:13.247000",
      "content": "<p>Whatever the outcome is, please make sure you are able to re-score previous solutions. Otherwise a lot of experiments need unnecessary repetition and its a waste of effort.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 2653063,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-02-15T06:53:59.573000",
          "content": "<p>Yep, super important</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2653996,
          "author_name": "Liubomyr Klymiuk",
          "author_url": "",
          "post_date": "2024-02-15T19:00:34.940000",
          "content": "<p>Agree, despite I have none yet</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2657228,
          "author_name": "Vallian Sayoga",
          "author_url": "",
          "post_date": "2024-02-18T11:43:41.767000",
          "content": "<p>I agree. I was disheartened that my experiments and models could not improve beyond 0.5. After seeing this, my model may have steeper slope and I hope the revised metric will show their true capabilities.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2658288,
          "author_name": "Riya Raizada",
          "author_url": "",
          "post_date": "2024-02-19T05:30:52.720000",
          "content": "<p>Agreed!!!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2651381,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2024-02-14T05:59:21.363000",
      "content": "<h1>Covid, Leaderboard &amp; the Metric</h1>\n<p>I wonder if this is really a metric issue, or a fundamental fact that hosts cannot probe the future, while we can probe the LB, and that one cannot model COVID default rates with pre-COVID data.</p>\n<p>Some thoughts:</p>\n<ul>\n<li>COVID was a macro event that changed the data generation process &amp; ideal credit scoring function so much, that no amount of historical data would give us a stable cross-validation. The best way to check the actual model performance on the Private LB is to probe the Public LB, no matter how small % it constitutes (in this context I think it would be in hosts' interest to reveal the split method of private/public LB - just to let Kagglers know if probing LB makes sense or not. Without clear indication, most people will use LB above CV - and maybe waste their time maybe not - but hosts will not benefit from it).</li>\n<li>While we can probe LB, hosts cannot probe the future while building a model, so models which benefit from LB probing are of no use to hosts.</li>\n<li>Also, COVID chances or happening again are rather small (hopefully). Even if a next pandemic happens, it will be again such a table-turning event, that data generation process &amp; ideal credit scoring function will change so much, that models built in this competition will be of no use.</li>\n<li>Will the winning models built with the current setup of train set (pre-covid and early covid) and eval set (late covid &amp; post covid) be useful at all for the 'new normal'?</li>\n</ul>\n<p><strong>Maybe it is better to give us access to the post Covid data, because models built with it will be useful for longer to the hosts?</strong> I understand it takes a lot of time for a credit to verify the default, but we can validate with short-term loans (like 3-6m), or we can change the target definition (default within first 3-6 months), or we can just wait for the new data to come - there have been competitions on Kaggle which took really long - look at the <a href=\"https://www.kaggle.com/competitions/zillow-prize-1\" target=\"_blank\">Zillow competition</a> for example, which lasted for almost 2 years.</p>\n<p>I don't have a clear solution yet, just throwing my 2 cents, but given the hosts engagement I would like this competition to produce something really useful for them.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2657834,
          "author_name": "JohnM",
          "author_url": "",
          "post_date": "2024-02-18T19:07:20.470000",
          "content": "<p>The broader problem you mention is the more interesting thing to solve in 2024, IMO. How do you maximize the useful signal from pre-anomaly data but mute the part that doesn't apply going forward? It's a big deal in healthcare right now for time-series problems. I'm guessing the good folks at Home Credit have thought a lot about it and how the scoring metric relates.</p>\n<p>It would be cool if a competition could somehow include metric development so that the winning model depended on a team's choice of objective and eval functions. But what metric would you use for the leaderboard?🤨 For now I guess iteration on metrics and ultimate evaluation has to be done outside of Kaggle World. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2657865,
              "author_name": "JohnM",
              "author_url": "",
              "post_date": "2024-02-18T19:40:42",
              "content": "<p>To your point on time intervals <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>:  If I understand correctly, we're trying to predict defaults post-covid without post-covid data in the public test set? That would be consistent with HC's focus on stable models (for the next big bad thing). I would think then a winning model pays a lot more attention to conservative estimation than to accuracy and leads to the metric problem we see now.</p>\n<p>That's actually OK in the real world. Models geared toward a worst-case scenario are often chosen over the most-likely scenario when volatility increases. It doesn't really work for Kaggle though.</p>\n<p>I could be wrong, of course. I still plan though to protect myself financially assuming HC is predicting a recession as that next big disruption. 😁</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2657866,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-18T19:41:44.697000",
              "content": "<p><a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> Thanks for appreciating the idea. The initial idea was to come up with model+metric design that would be scored by profitability in the long run. There were just too many cons for the initial idea so we had to go with the simplified version of it. The stability metric itself is an approximation of the decisions made when observing deployed models, but it is sufficient for the sake of this competition. People who say there is a disconnect between the metric and the competition idea perhaps don't see the importance of stable performance.</p>\n<p>We have a special prize for better implementation of stability in the development. I can imagine if someone comes up with a better stability metric that you can directly optimize in some kind of workflow to come up with a way of extracting the useful signal, he/she will be a good candidate for winning this special prize. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2658174,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-19T03:33:19.347000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2658508,
              "author_name": "narsil (jobs-in-data.com)",
              "author_url": "",
              "post_date": "2024-02-19T08:36:05.107000",
              "content": "<p>I have essentially two points:</p>\n<ul>\n<li>Asking to predict post-covid default rates with pre-covid data is a form of a very strong stress-testing. I am curious (but not sure) that models which best survive such stress-testing will be the most useful to the hosts, but it is their call, as they know their business best.</li>\n<li>Not revealing the public vs private test split + knowledge that LB score allows some access (albeit limited) to the future will practically mean that people will make many decisions based on LB score. We all know what is the consequence of that. Also - this is impractical since hosts don't have limited access to the future when building their models. So the process leading to the best model could likely not be reproduced by hosts when building their models.</li>\n</ul>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2651091,
      "author_name": "Jacoby Jaeger",
      "author_url": "",
      "post_date": "2024-02-13T22:51:04.597000",
      "content": "<p>It shouldn't be a surprise that a metric that gives an incentive to make some predictions worse, intentionally or otherwise, is problematic. Unless there is some reason to want worse performance on those samples, it is an obvious indication that the metric does not align with one's actual goals. </p>\n<p>Kaggle really ought to have standards in place for the competition metrics that it uses. A good rule that would avoid this kind of problem would be that the gradient of user scores wrt to sample scores must be strictly non-negative.</p>\n<p>Punishing score variability without breaking this rule can easily be done by using a loss term that grows rapidly on bad sample scores ie: mean( -log( gini  ) ) or even mean( (-log( gini))**exponent ) with exponent&gt;1. This has the benefit of punishing outliers with very bad performance without punishing those with very good performance. </p>\n<p>Punishing longer periods of poor scores can also be done without breaking the rule. For example one could use mean( -log( MovingAverage(gini) ) ). This metric is punishing to periods of consistent poor performance without providing false incentives or punishing good scores. </p>\n<p>The choice of initial metric for this competition was careless to a degree that should not be accepted. Kaggle should adopt rules that do not allow for these kinds of mistakes in the future.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 2652299,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2024-02-14T17:07:13.180000",
          "content": "<blockquote>\n  <p>The choice of initial metric for this competition was careless to a degree that should not be accepted. Kaggle should adopt rules that do not allow for these kinds of mistakes in the future.</p>\n</blockquote>\n<p>A warmup period (eg 1-2 weeks) prior the official competition lunch, where kagglers check the dataset/metric might help for spotting such issues (for \"non-trivial\" challenges).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2660882,
      "author_name": "Dmitriy Guller",
      "author_url": "",
      "post_date": "2024-02-20T21:31:21.227000",
      "content": "<p>One of the issues I commented on before is that you can have a case where Model A has a higher Gini than Model B for every week, and yet would score worse.  No matter how much you desire stability, surely you shouldn't desire Model B over Model A, so any metric that prefers Model B cannot fit the business case well.</p>\n<p>To prevent this failure mode, how about we simply use the average of N worst weekly Ginis as the score (with N chosen from some experimentation)?  Evaluating the model on its worst weeks will implicitly penalize lack of stability, but not to the point that a strictly worse model could come out better solely on the strength of being very consistently awful at predicting things.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2663049,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-22T09:35:39.567000",
          "content": "<p>Actually, it is possible that we would desire such model. Examples of gini in time A: [0.3, 0.3, 0.3], B: [0.8, 0.6, 0.4], we would still prefer model A here. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2651109,
      "author_name": "暗黑AGI",
      "author_url": "",
      "post_date": "2024-02-13T23:52:56.703000",
      "content": "<p>Why not remove this penalty metric and just use the auc to evaluate scores.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2652168,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-02-14T15:44:47.697000",
          "content": "<p>Or assign higher weights to weekly auc scores more distant in time.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2655278,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-16T19:11:52.747000",
          "content": "<p>Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/477074\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/477074</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2675153,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2024-02-29T18:23:04.680000",
      "content": "<p>I find it generally weird that you would remove 'weeks' a part of what is supposed to be optimised. Or conversely that you would use a metric including weeks if you remove that info. </p>\n<p>More generally it seems rather difficult to remove all 'time' info from the test. You can spend a lot of time reworking the data without garantee. I would suggest not doing this and keeping a simple metric like avg(gini) (or even a <a href=\"https://en.wikipedia.org/wiki/Scoring_rule\" target=\"_blank\">proper scoring rule</a> ?). </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2658216,
      "author_name": "random_abnormal",
      "author_url": "",
      "post_date": "2024-02-19T04:40:22.917000",
      "content": "<p>why dont you just use a mean with more weight on the recent values for mean gini</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2653082,
      "author_name": "Davut Polat",
      "author_url": "",
      "post_date": "2024-02-15T07:18:46.117000",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>  I can suggest that using Gini as evaluation metric on LB and stability score for stability prize.  HC can restrict stability prize like one can win it if she/he in top %1-2 on private LB and stability score should be calculated with the submission which has  maximum gini score among contender's 2 final submissions (people still could hack the metric with two submission, 1 for gini score other one is for stability score).  I believe that will be beneficial for both parties (HC and kagglers ).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2655270,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-16T19:08:26.230000",
          "content": "<p>I am afraid that would not be a satisfactory solution on behalf of HC. This year the competition is about stability, without having the stability directly in the evaluation metric we would have little control over the stability of the submitted solutions. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2655404,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-16T20:38:58.970000",
              "content": "<p>without changing to an usual metric (i.e. a metric that is always worse when an individual gini score get's worse) the result will be solutions that are not stable but just carefully chosen hacks or maybe even more possible purely lucky hacks.<br>\nthere are already some suggestions for metric whcih satisfy the above condition and still encourage stability over time.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2655447,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-16T21:44:32.010000",
              "content": "<p>We are working on a fix that will discourage such hacking. We are less likely to directly change the metric as now it is a good approximation of how risk managers \"rate\" our models in production. I must ask you for patience. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2650940,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-02-13T19:04:24.210000",
      "content": "<p><a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> kindly inform us in advance as this is likely to impact our plans/ ideas/ models for the competition. <br>\nThanks for being so proactive!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2658425,
      "author_name": "Davut Polat",
      "author_url": "",
      "post_date": "2024-02-19T07:17:28.027000",
      "content": "<p>$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1}\\sqrt{ \\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{peak} - GiniScore_i)\\right)^2}<br>\n$$</p>\n<p>GiniScore_peak : is the highest Gini score observed up to week i <br>\nGiniScore_i : is the Gini score for week i </p>\n<p>I modified the  ulcer index from finance domain (<a href=\"https://en.wikipedia.org/wiki/Ulcer_index\" target=\"_blank\">https://en.wikipedia.org/wiki/Ulcer_index</a>)</p>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2658652,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-19T10:02:30.473000",
          "content": "<p>Thanks for the idea. We wanted to make the metric as robust as possible and at the same time as simple as possible. Since you have GinniScore_peak in the metric it would be susceptible in different type of hacking. That would result in residues (of weekly gini in respect to y=ax+b) with a skew whereas we don't want to penalize less the weeks where the performance is worse. In a sense the modified Ulcer index is similar to the term with std(gini) except that it doesn't penalize equally the negative drops in performance and positive changes in performance (in weekly gini). There is no overall trend as a term, we definitely don't want to promote performance that is on a stable decline.  </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2658670,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-19T10:24:08.583000",
              "content": "<p>Note that GiniScore_peak is rolling max not absolute max  and  when someone tries to hack it, he/she will end up lowering the mean(GiniScore) at almost same degree I think that makes it more robust than current one.  HC used 88.0⋅𝑚𝑖𝑛(0,𝑎) term, this also does not penalize equally for positive and negative changes. The term in the sqrt behaves like a trend  because of comparing with rolling max of GiniScore value, Thus,  it gives penalty for negative gini performance slope, 0 penalty for positive (and steady performance) slopes.  You can test the metric with several scenarios. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2658716,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-19T11:00:00.710000",
              "content": "<p>I will test it, thanks for the suggestion.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2658882,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-19T13:25:04.013000",
              "content": "<p>i have modified the formula (noticed an error) and tried a couple of scenarios, it seems OK so far</p>\n<p>$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1} \\sqrt{\\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{rollingMax} - GiniScore_i)\\right)^2}<br>\n$$</p>\n<p>GiniScore_rollingMax : is the highest Gini score observed up to week i <br>\nGiniScore_i : is the Gini score for week i </p>\n<p>a = [0.60, 0.66, 0.65, 0.70, 0.64, 0.61, 0.64, 0.65, 0.64]    -- benchmark--<br>\nb = [0.55, 0.60, 0.60, 0.65, 0.64, 0.61, 0.64, 0.65, 0.64]    worsened first 4 weeks by 0.05<br>\nc = [0.60, 0.61, 0.64, 0.64, 0.64, 0.65, 0.65, 0.66, 0.70]    max stable scenario with benchmark gini scores <br>\nd = [0.59, 0.65, 0.64, 0.69, 0.63, 0.60, 0.63, 0.64, 0.63]   worsened all weeks by 0.01<br>\ne = [0.70, 0.66, 0.65, 0.65, 0.64, 0.64, 0.64, 0.61, 0.60]    worst scenario with benchmark gini scores <br>\nf = [0.58, 0.64, 0.63, 0.67, 0.64, 0.61, 0.64, 0.65, 0.64]    worsened first 4 weeks by 0.02  <br>\ng = [0.56, 0.62, 0.62, 0.62, 0.64, 0.61, 0.64, 0.65, 0.64]  worsened first 4 weeks with random values <br>\nh = [0.56, 0.62, 0.62, 0.62, 0.63, 0.63, 0.63, 0.63, 0.63]  worsened all weeks (monotonic increasing)<br>\ni = [0.62, 0.61, 0.60, 0.65, 0.64, 0.64, 0.64, 0.61, 0.60]    worsened first 4 weeks for scenario e <br>\nj = [0.65, 0.62, 0.61, 0.63, 0.64, 0.64, 0.64, 0.61, 0.60]   worsened first 4 weeks for scenario e</p>\n<p>a : 0.6250<br>\nb: 0.6145<br>\nc: 0.6433<br>\nd: 0.6150<br>\ne: 0.6197<br>\nf: 0.6230<br>\ng: 0.6182<br>\nh: 0.6188<br>\ni: 0.6145<br>\nj: 0.6159</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2660103,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-20T10:47:34",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> i have tested the metric I suggested  for 100M  random scenarios, </p>\n<p>number of weeks : 18<br>\ninitialized the gini scores by week between [0.55, 0.65]<br>\ncreated a mask which has random size between [3,9] -- refers to which first N weeks to worsen the gini scores. mask has values  between [0.9 , 1.0]  If    mask[4] = 0.95 that mean I worsen 5th week' gini by %5.  </p>\n<p>here is the result I get:</p>\n<ul>\n<li>for only %0.27 cases, one can improve the metric by hacking </li>\n<li>for rest of the cases, score is distorted </li>\n<li>maximum improvement one can get by hacking is 0.0070 for 18 weeks, but when I increase the number of weeks it is getting harder to hack so  maximum improvement one can get  decreases.</li>\n<li>median improvement among %0.27 cases  is  0.0005 (for 18 weeks)</li>\n</ul>\n<p>improvement quantiles (for %0.27 of  cases)</p>\n<p>Quantile 10 8.898260757736014e-05<br>\nQuantile 20 0.00018749605249894266<br>\nQuantile 30 0.00029993356066040623<br>\nQuantile 40 0.0004269924733614792<br>\nQuantile 50 0.0005784240749482758<br>\nQuantile 60 0.0007601426463899664<br>\nQuantile 70 0.0009925642541695821<br>\nQuantile 80 0.0013097591959578612<br>\nQuantile 90 0.0018370136698832628<br>\nQuantile 100 0.006960994642879724</p>\n<p>now I am running 1B scenarios for 50 weeks. I will share the results and the code here. </p>\n<p>To hack the metric I suggested, one should know which week has  the maximum gini score otherwise the score will be distorted, I think no one would try to hack it.  </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2660141,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-20T11:25:42.563000",
              "content": "<p>Why choose this metric if <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a> suggested already a metric which is provably non hackable.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2660169,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-20T12:00:24.467000",
              "content": "<p>now I have tested  Jacob's  metric with 1M scenarios, it seems  it is hackable and one can get up to  0.038  improvement (final score range is like 0.25-0.30.  0.038  means  ~ %13 improvement)    </p>\n<p>score range : (0.25-0.30)<br>\nImprovement Quantiles<br>\nQuantile 5 0.0018123009184825712<br>\nQuantile 10 0.002954071577671077<br>\nQuantile 15 0.0042528432747884854<br>\nQuantile 20 0.005571776912732721<br>\nQuantile 25 0.006885412084717121<br>\nQuantile 30 0.008219234721494929<br>\nQuantile 35 0.009542760995397888<br>\nQuantile 40 0.01087649085647983<br>\nQuantile 45 0.012200174484067327<br>\nQuantile 50 0.01351322138801158<br>\nQuantile 55 0.014840927605872758<br>\nQuantile 60 0.016168883974780883<br>\nQuantile 65 0.017497915477009143<br>\nQuantile 70 0.018820113281295393<br>\nQuantile 75 0.02017423466541196<br>\nQuantile 80 0.021564749464102084<br>\nQuantile 85 0.023043430204338983<br>\nQuantile 90 0.024716932816428373<br>\nQuantile 95 0.02689879395363132<br>\nQuantile 100 0.03894322510151606</p>\n<p>i have tested mine with 200M scenarios and  50 weeks:  <br>\nfor  %0.021 cases it is hackable and most improvement one can get  : 0.0021<br>\nmedian improvement (for %0.021 cases) : 0.0001<br>\nscore range : (0.57- 0.64)</p>\n<p>Improvement Quantiles<br>\nQuantile 5 1.2699755546674734e-05<br>\nQuantile 10 2.6765098840468564e-05<br>\nQuantile 15 4.102476662966947e-05<br>\nQuantile 20 5.6175799148184025e-05<br>\nQuantile 25 7.213998299378261e-05<br>\nQuantile 30 8.979772687923974e-05<br>\nQuantile 35 0.00010814053350019279<br>\nQuantile 40 0.00012823791281313051<br>\nQuantile 45 0.00015023664735298696<br>\nQuantile 50 0.0001735378974960051<br>\nQuantile 55 0.00020047277886452565<br>\nQuantile 60 0.0002300174182662148<br>\nQuantile 65 0.00026180227673274555<br>\nQuantile 70 0.00029811598482456093<br>\nQuantile 75 0.0003424203629695818<br>\nQuantile 80 0.00039457860089038663<br>\nQuantile 85 0.00046220233188403397<br>\nQuantile 90 0.0005567642604209829<br>\nQuantile 95 0.0007121168562655653<br>\nQuantile 100 0.0021852690646599017</p>\n<p>if we increase the week num it is getting harder to hack.<br>\ni have attached code for mine and Jacob's metric. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2660319,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-20T14:30:50.200000",
              "content": "<p>It is probably non hackable by decreasing gini score. Please see the other discussion for explanation. If you look carefully at the formula you will see that every decreasing of gini score will make the metric worse.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2660498,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-20T16:29:56.317000",
              "content": "<p>Hacking the metric is done by decreasing gini score for half weeks. See the main thread here:</p>\n<p><a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476449</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2660563,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-20T17:11:59.497000",
              "content": "<p>I was talking about metric proposed by above user in another thread:<br>\nmean( ( -log( MovingAverage(gini) ) )**exponent ) with exponent&gt;=1</p>\n<p>I don’t see how this metric can be improved by artificially lowering any gini scores. If I am wrong, please correct me.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2660616,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-20T17:50:56.267000",
              "content": "<p>See the test code i attached.  i generated 1M scenario which artificially lowering gini scores. All cases improved the stability scores (up to %13 improvement ) </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2660621,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-20T17:59:19.417000",
              "content": "<p>Are you aware of the fact, that in case of this metric lower is better? I can’t think of a reason why any of the moving averages should get higher by lowering gini scores. If you decrease the ginis the averages will get smaller, that means the -log of that will get larger and the same total mean over all time windows will increase as well which means the score  is worse.  Or am I wrong?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2660632,
              "author_name": "Davut Polat",
              "author_url": "",
              "post_date": "2024-02-20T18:16:24.090000",
              "content": "<p>if lower is better, then Jacob’s metric is fine.  %100 of cases distorted if you lower the gini scores. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2660672,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-20T18:47:02.077000",
              "content": "<p>Yes. The best possible outcome is that every Gini = 1 (for every timestep) then every -log(1)=0 and the overall score is =0.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2671312,
      "author_name": "Thomas Meißner",
      "author_url": "",
      "post_date": "2024-02-27T12:42:48.937000",
      "content": "<p>When is it expected that submissions are possible again?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2662874,
      "author_name": "mohamed el nashar",
      "author_url": "",
      "post_date": "2024-02-22T07:10:16.423000",
      "content": "<p>Astrologers lied even by chance, data science determines the optimal direction according to the data and not without it Therefore I think that it is difficult to provide stability metric that skips to the accuracy of the near or far future alike Therefore it is possible to manipulate any stability metric according to its function The best solution is to separate the evaluation into two or three parts so that the data has to be involved in it For example evaluating the person in terms of whether he is able to pay or not based on what is available in the data available in The beginning, otherwise if someone can deduce the future by more than half and succeeds in that, he will not give jealousy because in this case he has the power of a god</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2658036,
      "author_name": "ARMADA",
      "author_url": "",
      "post_date": "2024-02-18T23:22:17.487000",
      "content": "<p>If the main concern is that teams are gaming the stability metric by lowering the score of the first few weeks, one idea to keep the stability metric and also discourage lowering the score of the first weeks would be to add an additional weighted term to the current metric e.g.  0.3*mean(gini1, gini2,… giniN) for the first N weeks. This means that we want solutions that can score highly on the first N weeks and also decline very slowly overtime (which was the original intent). </p>\n<p>or perhaps replace the first term in the current metric:  mean(gini) by mean(gini1, gini2,… giniN) for N weeks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2653519,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-15T13:39:48.923000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2650939": "********Dear Kagglers!\n\nAs you have probably noticed in the discussion already, there is issue with our competition metric. In short, it is possible to gain extra score by artificially reducing predictive power of model for first part of evaluation sample, thus smoothing gini over time and reduce penalty for instability.\n\nThis is something we are not looking for. Our goal is to maximize both predictive ability and stability. At the same time our implicit expectation was that for all clients will be applied same rules/model/approach, in other words the scoring would not explicitly depend on date decision. From the discussion on Kaggle I have strong feeling that you see it similarly, and you don’t want competition where winning strategy is to game scoring metric. However, we don’t want to remove stability altogether, from the beginning we wanted to present you some new problem, not just repeat previous Home Credit competition.\n\nTherefore, we have agreed we have to react. Together with Kaggle team we are analyzing several ideas how to prevent possibility to improve score by artificially changing models for certain time periods. At the same time we fully understand that it will be change in on-going competition, which should not be done under normal circumstances. We want to make sure that our fix will be correct and final, so we want to thoroughly analyze and test our ideas – it will take us few days. But we want to inform you we are working on solution and be transparent with you.\n\nLast but not least, I want to thank many of you who were pointing at this issue, namely @at7459 for sharing openly his solution and for proving that our metric is not working as intended.\n\nWill keep you updated.\n\nBest,\n\nTomas & Daniel\n",
    "2651259": "First, @at7459 thank you for the report. \nAnd thank you to the host @tomasjeline2 , @jetakow   for sincerely considering a revision of the metric for this topic.\n\nMy thought is that, perhaps, if we consider the reverse, isn't there a potential issue like in Case 3, where improving the Gini value in parts could worsen the overall score? (Even though the average value of Gini might increase, the slope becomes significantly negative, and its weight is large). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F4f77e7c26d8d69e6c9a3e2897368bc0e%2FClipboard03.jpg?generation=1707882775177471&alt=media)\n\n※ The green represents a fixed Gini value for explanation purposes. The blue is intentionally shifted for explanation purposes.\n\n\nAnd, this potential issue is incompatible with the description in the data tab below and could significantly impact the overall score.\n~~~\nIt's worth noting that some external data providers might not be available for future (test) evaluations, which is anticipated.\n~~~\n\nI would appreciate it if you could also consider this potential issue. But, if that's what the host desires, I won't say anything.\n\nPersonally, using polyfit makes it difficult to balance the mean value of Gini with the term of the slope, making it challenging to excel in this competition.\n\nPerhaps, instead of using polyfit, for example, weighting the Gini function for each week number, taking the sum, and then normalizing it at the end, might also be a good metric.\n\nHowever, this is just my opinion, so I leave the decision to the host. I would appreciate it if you could take this into consideration. Please feel free to point out any mistakes.",
    "2662222": "Dear Kagglers, \n \nWe prepared this year’s competition with an intention to present you something else and more challenging than the previous Home Credit kaggle competition. We decided to select stability of models in production as the main objective of this competition. Therefore, we will keep stability and won’t remove it from the competition.  \n\nWe would like to further clarify the meaning of stability. Our ideal stable model has constant performance by weeks in time and as high of a mean gini over weeks as possible. That is the reason why we have three terms in the stability metric as of now since this is desirable in business and not only to maximize AUC. The metric itself is a good approximation of the decisions made in-house when assessing the performance of models in production. Despite different opinions here within the competition, removing the terms in the stability metric would significantly deviate from the way we rate performance internally. When people leverage the second term and nullify it by carefully choosing scores on the test sample, it does not represent how we deploy models in the production. This is undesirable and we are forced to make steps that will discourage hacking of the metric.  \n \nWe will make a change to the test dataset in such a way that it will represent better client scoring in production. When the model is running in production, it can only see its past performance and predictions. This would be quite difficult to reproduce in a Kaggle competition, but we have an alternative solution. We will remove the date decision, month and week number from the test base table. This part will take two weeks to amend, and we will prolong the competition by this period. Because of this, it is very likely that we won’t be able to rescore the old submissions. We apologize in advance for this inconvenience. \n \nBesides the change of test data, we have a second plan regarding metric. At this moment we are still considering metric change, and we welcome suggestions that would be aligned with how we view the value of model stability and in a way that would prevent further hacking. We are unlikely to go with a new metric that would prevent this type of hacking and would significantly deviate from the stability part. There is very little guarantee that a new metric would not open doors to a new type of hacking. This part will take at least a week.\n \nWe would like to apologize for the delay and inconvenience caused by the issues connected to the stability metric. We also want to thank the community for many proposals on how to fix the issue. We have tested and analyzed many of them, thank you for all the suggestions! We will disable submissions for two weeks from now on and prolong the competition by two weeks.  \n\nTomas & Daniel",
    "2652701": "Whatever the outcome is, please make sure you are able to re-score previous solutions. Otherwise a lot of experiments need unnecessary repetition and its a waste of effort.",
    "2651381": "# Covid, Leaderboard & the Metric\n\nI wonder if this is really a metric issue, or a fundamental fact that hosts cannot probe the future, while we can probe the LB, and that one cannot model COVID default rates with pre-COVID data.\n\nSome thoughts:\n- COVID was a macro event that changed the data generation process & ideal credit scoring function so much, that no amount of historical data would give us a stable cross-validation. The best way to check the actual model performance on the Private LB is to probe the Public LB, no matter how small % it constitutes (in this context I think it would be in hosts' interest to reveal the split method of private/public LB - just to let Kagglers know if probing LB makes sense or not. Without clear indication, most people will use LB above CV - and maybe waste their time maybe not - but hosts will not benefit from it).\n- While we can probe LB, hosts cannot probe the future while building a model, so models which benefit from LB probing are of no use to hosts.\n- Also, COVID chances or happening again are rather small (hopefully). Even if a next pandemic happens, it will be again such a table-turning event, that data generation process & ideal credit scoring function will change so much, that models built in this competition will be of no use.\n- Will the winning models built with the current setup of train set (pre-covid and early covid) and eval set (late covid & post covid) be useful at all for the 'new normal'?\n\n**Maybe it is better to give us access to the post Covid data, because models built with it will be useful for longer to the hosts?** I understand it takes a lot of time for a credit to verify the default, but we can validate with short-term loans (like 3-6m), or we can change the target definition (default within first 3-6 months), or we can just wait for the new data to come - there have been competitions on Kaggle which took really long - look at the [Zillow competition](https://www.kaggle.com/competitions/zillow-prize-1) for example, which lasted for almost 2 years.\n\nI don't have a clear solution yet, just throwing my 2 cents, but given the hosts engagement I would like this competition to produce something really useful for them.",
    "2651091": "It shouldn't be a surprise that a metric that gives an incentive to make some predictions worse, intentionally or otherwise, is problematic. Unless there is some reason to want worse performance on those samples, it is an obvious indication that the metric does not align with one's actual goals. \n\nKaggle really ought to have standards in place for the competition metrics that it uses. A good rule that would avoid this kind of problem would be that the gradient of user scores wrt to sample scores must be strictly non-negative.\n\nPunishing score variability without breaking this rule can easily be done by using a loss term that grows rapidly on bad sample scores ie: mean( -log( gini  ) ) or even mean( (-log( gini))**exponent ) with exponent>1. This has the benefit of punishing outliers with very bad performance without punishing those with very good performance. \n\nPunishing longer periods of poor scores can also be done without breaking the rule. For example one could use mean( -log( MovingAverage(gini) ) ). This metric is punishing to periods of consistent poor performance without providing false incentives or punishing good scores. \n\nThe choice of initial metric for this competition was careless to a degree that should not be accepted. Kaggle should adopt rules that do not allow for these kinds of mistakes in the future.",
    "2660882": "One of the issues I commented on before is that you can have a case where Model A has a higher Gini than Model B for every week, and yet would score worse.  No matter how much you desire stability, surely you shouldn't desire Model B over Model A, so any metric that prefers Model B cannot fit the business case well.\n\nTo prevent this failure mode, how about we simply use the average of N worst weekly Ginis as the score (with N chosen from some experimentation)?  Evaluating the model on its worst weeks will implicitly penalize lack of stability, but not to the point that a strictly worse model could come out better solely on the strength of being very consistently awful at predicting things.",
    "2651109": "Why not remove this penalty metric and just use the auc to evaluate scores.",
    "2675153": "I find it generally weird that you would remove 'weeks' a part of what is supposed to be optimised. Or conversely that you would use a metric including weeks if you remove that info. \n\nMore generally it seems rather difficult to remove all 'time' info from the test. You can spend a lot of time reworking the data without garantee. I would suggest not doing this and keeping a simple metric like avg(gini) (or even a [proper scoring rule](https://en.wikipedia.org/wiki/Scoring_rule) ?). ",
    "2658216": "why dont you just use a mean with more weight on the recent values for mean gini",
    "2653082": "@tomasjeline2 @jetakow  I can suggest that using Gini as evaluation metric on LB and stability score for stability prize.  HC can restrict stability prize like one can win it if she/he in top %1-2 on private LB and stability score should be calculated with the submission which has  maximum gini score among contender's 2 final submissions (people still could hack the metric with two submission, 1 for gini score other one is for stability score).  I believe that will be beneficial for both parties (HC and kagglers ).",
    "2650940": "@tomasjeline2 kindly inform us in advance as this is likely to impact our plans/ ideas/ models for the competition. \nThanks for being so proactive!",
    "2658425": "$$StabilityScore =  mean(GiniScore) - \\frac{1}{n-1}\\sqrt{ \\sum_{i=2}^{n} \\left(\\max(0, GiniScore_{peak} - GiniScore_i)\\right)^2}\n$$\n\nGiniScore_peak : is the highest Gini score observed up to week i \nGiniScore_i : is the Gini score for week i \n\nI modified the  ulcer index from finance domain (https://en.wikipedia.org/wiki/Ulcer_index)\n\n@jetakow @tomasjeline2 ",
    "2671312": "When is it expected that submissions are possible again?",
    "2662874": "Astrologers lied even by chance, data science determines the optimal direction according to the data and not without it Therefore I think that it is difficult to provide stability metric that skips to the accuracy of the near or far future alike Therefore it is possible to manipulate any stability metric according to its function The best solution is to separate the evaluation into two or three parts so that the data has to be involved in it For example evaluating the person in terms of whether he is able to pay or not based on what is available in the data available in The beginning, otherwise if someone can deduce the future by more than half and succeeds in that, he will not give jealousy because in this case he has the power of a god",
    "2658036": "If the main concern is that teams are gaming the stability metric by lowering the score of the first few weeks, one idea to keep the stability metric and also discourage lowering the score of the first weeks would be to add an additional weighted term to the current metric e.g.  0.3*mean(gini1, gini2,... giniN) for the first N weeks. This means that we want solutions that can score highly on the first N weeks and also decline very slowly overtime (which was the original intent). \n\nor perhaps replace the first term in the current metric:  mean(gini) by mean(gini1, gini2,... giniN) for N weeks.",
    "2653519": ""
  }
}