{
  "id": 476449,
  "title": "Problem with competition metric",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/476449",
  "author_name": "at7459",
  "post_date": "2024-02-12T12:12:13.027000",
  "votes": 130,
  "comment_count": 47,
  "views": 0,
  "content": "<p>i could abuse this, but i guess by the end most people will be doing something like this ,and it doesn't really have anything to do with data science. also, in the end, people who didn't see this will just get pissed off.</p>\n<p>The problem is that the 88 * min(0,a) term is too large, so you can just make your model worse for the first weeks and improve your score a lot. i like the idea of the metric in itself but this is a competition after all and people will just pull the dirtiest tricks out of their bag to come out ahead.</p>\n<p>for example something like this,</p>\n<p><code>condition = df_subm['WEEK_NUM'] &lt; (df_subm['WEEK_NUM'].max()-df_subm['WEEK_NUM'].min())/2+df_subm['WEEK_NUM'].min()\ndf_subm.loc[condition, 'score'] = (df_subm.loc[condition, 'score'] - 0.02).clip(0)\ndf_subm=df_subm[[\"case_id\",\"score\"]]\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm.to_csv(\"submission.csv\")\nprint(df_subm)</code></p>\n<p>where i make my model worse for the first half of the testing period, improves my score by almost 0.03</p>\n<p>possible solution:</p>\n<p>as i said the 88 is too large, when you look at the contributions from the metric for a model which gets worse over time, this term is about 4-6 times larger than the other term. maybe 18 or something would be more fitting, it takes some experimenting i guess</p>",
  "messages": [
    {
      "id": 2648775,
      "postDate": "2024-02-12T12:12:13.027Z",
      "content": "<p>i could abuse this, but i guess by the end most people will be doing something like this ,and it doesn't really have anything to do with data science. also, in the end, people who didn't see this will just get pissed off.</p>\n<p>The problem is that the 88 * min(0,a) term is too large, so you can just make your model worse for the first weeks and improve your score a lot. i like the idea of the metric in itself but this is a competition after all and people will just pull the dirtiest tricks out of their bag to come out ahead.</p>\n<p>for example something like this,</p>\n<p><code>condition = df_subm['WEEK_NUM'] &lt; (df_subm['WEEK_NUM'].max()-df_subm['WEEK_NUM'].min())/2+df_subm['WEEK_NUM'].min()\ndf_subm.loc[condition, 'score'] = (df_subm.loc[condition, 'score'] - 0.02).clip(0)\ndf_subm=df_subm[[\"case_id\",\"score\"]]\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm.to_csv(\"submission.csv\")\nprint(df_subm)</code></p>\n<p>where i make my model worse for the first half of the testing period, improves my score by almost 0.03</p>\n<p>possible solution:</p>\n<p>as i said the 88 is too large, when you look at the contributions from the metric for a model which gets worse over time, this term is about 4-6 times larger than the other term. maybe 18 or something would be more fitting, it takes some experimenting i guess</p>",
      "rawMarkdown": "i could abuse this, but i guess by the end most people will be doing something like this ,and it doesn't really have anything to do with data science. also, in the end, people who didn't see this will just get pissed off.\n\nThe problem is that the 88 * min(0,a) term is too large, so you can just make your model worse for the first weeks and improve your score a lot. i like the idea of the metric in itself but this is a competition after all and people will just pull the dirtiest tricks out of their bag to come out ahead.\n\nfor example something like this,\n\n`condition = df_subm['WEEK_NUM'] < (df_subm['WEEK_NUM'].max()-df_subm['WEEK_NUM'].min())/2+df_subm['WEEK_NUM'].min()\ndf_subm.loc[condition, 'score'] = (df_subm.loc[condition, 'score'] - 0.02).clip(0)\ndf_subm=df_subm[[\"case_id\",\"score\"]]\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm.to_csv(\"submission.csv\")\nprint(df_subm)`\n\nwhere i make my model worse for the first half of the testing period, improves my score by almost 0.03\n\npossible solution:\n\nas i said the 88 is too large, when you look at the contributions from the metric for a model which gets worse over time, this term is about 4-6 times larger than the other term. maybe 18 or something would be more fitting, it takes some experimenting i guess",
      "votes": 128
    },
    {
      "id": 2649090,
      "postDate": "2024-02-12T15:53:38.777Z",
      "content": "<p>We greatly appreciate your finding and that you came forward with. </p>\n<p>I am afraid there is no simple solution, let's have a look at the extremes. 1) we could omit mean gini or 2) we could omit the part with falling rate. When going with 2) the issue becomes that the models are pushed towards excellent short-term performance, but yet not stable in time. When going with 1) the issue becomes tweaking of the score so that we have as stable gini in time as possible. Changing the weights will not result in solving the problem, but rather transformation from one issue into another. </p>\n<p>We will discuss this also with Kaggle. Tweaking the score manually based on time is not desired in this competition and we are already thinking about steps how to partially mitigate this issue. We will make an official statement later.</p>",
      "rawMarkdown": "We greatly appreciate your finding and that you came forward with. \n\nI am afraid there is no simple solution, let's have a look at the extremes. 1) we could omit mean gini or 2) we could omit the part with falling rate. When going with 2) the issue becomes that the models are pushed towards excellent short-term performance, but yet not stable in time. When going with 1) the issue becomes tweaking of the score so that we have as stable gini in time as possible. Changing the weights will not result in solving the problem, but rather transformation from one issue into another. \n\nWe will discuss this also with Kaggle. Tweaking the score manually based on time is not desired in this competition and we are already thinking about steps how to partially mitigate this issue. We will make an official statement later.",
      "votes": 16,
      "replies": [
        {
          "id": 2649151,
          "postDate": "2024-02-12T16:36:19.870Z",
          "content": "<p>alright you can do it however you like but i think changing the weights could definitely solve the problem.</p>\n<p>for example, currently the score for the test period probably looks something like <br>\nmean(gini) : <strong>+0.7</strong><br>\nmin(0,a) : <strong>-0.09</strong><br>\nstd(residuals: <strong>-0.01</strong><br>\nwhich sums to <strong>0.6</strong><br>\nwhen i make my model worse for the first half it turns into<br>\nmean(gini) : <strong>+0.69</strong><br>\nmin(0,a) : <strong>-0.05</strong><br>\nstd(residuals: <strong>-0.02</strong><br>\nwhich sums to <strong>0.62</strong></p>\n<p>changing the weights of the min(0,a) term to say 44 would turn this into a net negative strategy for me</p>\n<p>granted it would put less emphasis on stability and take some experimenting but it could still work out overall</p>",
          "rawMarkdown": "alright you can do it however you like but i think changing the weights could definitely solve the problem.\n\nfor example, currently the score for the test period probably looks something like \nmean(gini) : **+0.7**\nmin(0,a) : **-0.09**\nstd(residuals: **-0.01**\nwhich sums to **0.6**\nwhen i make my model worse for the first half it turns into\nmean(gini) : **+0.69**\nmin(0,a) : **-0.05**\nstd(residuals: **-0.02**\nwhich sums to **0.62**\n\nchanging the weights of the min(0,a) term to say 44 would turn this into a net negative strategy for me\n\ngranted it would put less emphasis on stability and take some experimenting but it could still work out overall\n",
          "votes": 7,
          "replies": [
            {
              "id": 2649189,
              "postDate": "2024-02-12T17:05:23.350Z",
              "content": "<p>Of course this would make this particular strategy less interesting, but another may rise with new set of weights. Besides that smaller weight for falling rate would encourage less stable models, which is something we ultimately don't aim for. Most likely the countermeasure will not contain change in weights. </p>",
              "rawMarkdown": "Of course this would make this particular strategy less interesting, but another may rise with new set of weights. Besides that smaller weight for falling rate would encourage less stable models, which is something we ultimately don't aim for. Most likely the countermeasure will not contain change in weights. ",
              "votes": 4
            },
            {
              "id": 2649202,
              "postDate": "2024-02-12T17:14:10.707Z",
              "content": "<p>I think a valid reason for not changing the current metric would be this post-processing is risky; the public lb is only 30 % of the whole test data, and you mentioned this method doesn't work well in the training set. 0.03 change is huge but also not so huge.</p>\n<ul>\n<li>If the weekly gini in the hidden set keep falling, then this post-processing might work well</li>\n<li>if the weekly gini rises again, this post-processing might not work.</li>\n</ul>\n<p>The actual number of weeks and the percentage of predictions to force 0 are hard to validate. If someone can have a trick that works well in training and testing data, this might also show some proper stability.</p>\n<p>But I think some teams will use 1 of their final submissions to take this risk.</p>",
              "rawMarkdown": "I think a valid reason for not changing the current metric would be this post-processing is risky; the public lb is only 30 % of the whole test data, and you mentioned this method doesn't work well in the training set. 0.03 change is huge but also not so huge.\n\n- If the weekly gini in the hidden set keep falling, then this post-processing might work well\n- if the weekly gini rises again, this post-processing might not work.\n\nThe actual number of weeks and the percentage of predictions to force 0 are hard to validate. If someone can have a trick that works well in training and testing data, this might also show some proper stability.\n\nBut I think some teams will use 1 of their final submissions to take this risk.",
              "votes": 6
            },
            {
              "id": 2650343,
              "postDate": "2024-02-13T11:56:25.520Z",
              "content": "<p>All of those are valid points. We are discussing the issue internally. </p>",
              "rawMarkdown": "All of those are valid points. We are discussing the issue internally. ",
              "votes": 1
            },
            {
              "id": 2651850,
              "postDate": "2024-02-14T12:05:25.470Z",
              "content": "<p>Another alternative (draft idea) for changing the weights could be to change them dynamically, e.g. by multiplying the penalty by K/gini_mean (K - some constant). So higher gini_mean would have less penalty. But it seems too complex and needs more rigorous mathematical proofs. And instead of using linear regression, Huber regression could be used (slight optimization).</p>",
              "rawMarkdown": "Another alternative (draft idea) for changing the weights could be to change them dynamically, e.g. by multiplying the penalty by K/gini_mean (K - some constant). So higher gini_mean would have less penalty. But it seems too complex and needs more rigorous mathematical proofs. And instead of using linear regression, Huber regression could be used (slight optimization)."
            }
          ]
        },
        {
          "id": 2649405,
          "postDate": "2024-02-12T20:01:42.303Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>, I understand that the point of this competition is to come up with a stable nicely performing model. However, I’m not sure how the penalizing term helps. What is the advantage of having a model that performs worse on the first weeks and better on the last ones? As we can see from this post, this requirement can actually result in an overall worse model.</p>\n<p>Is it because you assume the data after the last weeks will stay similar and will never be close to the first weeks data? I would appreciate it if you could elaborate a little bit on this.</p>",
          "rawMarkdown": "Dear @jetakow, I understand that the point of this competition is to come up with a stable nicely performing model. However, I’m not sure how the penalizing term helps. What is the advantage of having a model that performs worse on the first weeks and better on the last ones? As we can see from this post, this requirement can actually result in an overall worse model.\n\nIs it because you assume the data after the last weeks will stay similar and will never be close to the first weeks data? I would appreciate it if you could elaborate a little bit on this.",
          "votes": 1,
          "replies": [
            {
              "id": 2649463,
              "postDate": "2024-02-12T20:46:25.787Z",
              "content": "<p>We determined that this year's problem we wanted to offer to kagglers is stability. Using no penalizing terms would lead to only unstable models. We will not leak any information on purpose about test sample. I can only ask you for patience while we will try to make things more fair for everyone. </p>",
              "rawMarkdown": "We determined that this year's problem we wanted to offer to kagglers is stability. Using no penalizing terms would lead to only unstable models. We will not leak any information on purpose about test sample. I can only ask you for patience while we will try to make things more fair for everyone. ",
              "votes": 5
            },
            {
              "id": 2650082,
              "postDate": "2024-02-13T08:34:27.997Z",
              "content": "<p>imo the best way to deal with stability is one of those competition structures where you're continously fed new data and can retrain your model. obviously, as you stated, due to how your company works and how long a model gets used, this is not really possible.</p>\n<p>but in this competition structure right now - due to the fact that the testing period is so different from the training period - whether you fix the problem or not, it will mostly just be about leaderboard probing and checking what works well on the leaderboard. ive had very bad experience with these types of competitions so far, honestly the only reason i joined is due to the added stability thing.</p>",
              "rawMarkdown": "imo the best way to deal with stability is one of those competition structures where you're continously fed new data and can retrain your model. obviously, as you stated, due to how your company works and how long a model gets used, this is not really possible.\n\nbut in this competition structure right now - due to the fact that the testing period is so different from the training period - whether you fix the problem or not, it will mostly just be about leaderboard probing and checking what works well on the leaderboard. ive had very bad experience with these types of competitions so far, honestly the only reason i joined is due to the added stability thing.",
              "votes": 4
            },
            {
              "id": 2650338,
              "postDate": "2024-02-13T11:54:28.227Z",
              "content": "<blockquote>\n  <p>honestly the only reason i joined is due to the added stability thing.</p>\n</blockquote>\n<p>Then I hope it will be still challenging for you enough that you remain in the competition. I ask you for patience and please understand that we have only limited options at the moment. </p>",
              "rawMarkdown": ">honestly the only reason i joined is due to the added stability thing.\n\nThen I hope it will be still challenging for you enough that you remain in the competition. I ask you for patience and please understand that we have only limited options at the moment. "
            },
            {
              "id": 2650578,
              "postDate": "2024-02-13T14:41:28.907Z",
              "content": "<p>I'm not sure the point being made here is completely congruent with the stated intent of the model and the information included for modelling purposes (vintage). </p>\n<p>There's a big disconnect between the metric, the stated intent and the data being used for modelling.</p>",
              "rawMarkdown": "I'm not sure the point being made here is completely congruent with the stated intent of the model and the information included for modelling purposes (vintage). \n\nThere's a big disconnect between the metric, the stated intent and the data being used for modelling."
            },
            {
              "id": 2650753,
              "postDate": "2024-02-13T17:25:47.313Z",
              "content": "<p>I'm not suggesting anything here. But from a metric design perspective, I am curious: Why can't we use std(weekly gini) here?</p>\n<pre><code>(weekly gini) - (weekly gini)\n</code></pre>\n<p>This keeps the idea of weekly gini and stability and doesn't need weight.</p>\n<p>I agree with <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> ; lb probing doesn't feel good. The 30% public leader board becomes tricky to interpret.  I am also interested in the stability concept here. Online learning is a way to handle stability, but it also adds extra costs. I agree that Online learning could be an easier method to handle this,  but I am also excited about the concept of training a model from pre-COVID weeks and surviving the COVID weeks. </p>",
              "rawMarkdown": "I'm not suggesting anything here. But from a metric design perspective, I am curious: Why can't we use std(weekly gini) here?\n```\nmean(weekly gini) - std(weekly gini)\n```\nThis keeps the idea of weekly gini and stability and doesn't need weight.\n\nI agree with @at7459 ; lb probing doesn't feel good. The 30% public leader board becomes tricky to interpret.  I am also interested in the stability concept here. Online learning is a way to handle stability, but it also adds extra costs. I agree that Online learning could be an easier method to handle this,  but I am also excited about the concept of training a model from pre-COVID weeks and surviving the COVID weeks. ",
              "votes": 3
            },
            {
              "id": 2650795,
              "postDate": "2024-02-13T17:47:26.793Z",
              "content": "<blockquote>\n  <p>Using no penalizing terms would lead to only unstable models.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> I only meant the second term. I understand you don’t want the model’s performance to deteriorate with time, but since credit models are only valid for one year (as you say in description), doesn’t mean and std address stability issue good enough if the evaluation period is similar to the model’s lifetime?</p>",
              "rawMarkdown": ">Using no penalizing terms would lead to only unstable models.\n\n@jetakow I only meant the second term. I understand you don’t want the model’s performance to deteriorate with time, but since credit models are only valid for one year (as you say in description), doesn’t mean and std address stability issue good enough if the evaluation period is similar to the model’s lifetime?"
            },
            {
              "id": 2650849,
              "postDate": "2024-02-13T18:12:11.017Z",
              "content": "<p>At this very moment, we consider also this scenario. It may be right that most of the credit models are made to last only one year, but that is no universal rule. They can be deployed for much longer than one year. Removing the falling rate term would promote unstable models, that perform well shortly, but quickly degrade. We need to penalize this behaviour and that is the reason why such term is present. </p>",
              "rawMarkdown": "At this very moment, we consider also this scenario. It may be right that most of the credit models are made to last only one year, but that is no universal rule. They can be deployed for much longer than one year. Removing the falling rate term would promote unstable models, that perform well shortly, but quickly degrade. We need to penalize this behaviour and that is the reason why such term is present. ",
              "votes": 1
            },
            {
              "id": 2650868,
              "postDate": "2024-02-13T18:18:44.670Z",
              "content": "<p>I see, may be that term could be compensated by making timeframe of test data longer, i.e. similar to whatever the typical period of model’s deployment is. It could be already the case, not sure, but from my opinion testing on real data is better than trying to extrapolate performance with a linear function.</p>",
              "rawMarkdown": "I see, may be that term could be compensated by making timeframe of test data longer, i.e. similar to whatever the typical period of model’s deployment is. It could be already the case, not sure, but from my opinion testing on real data is better than trying to extrapolate performance with a linear function.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2657889,
          "postDate": "2024-02-18T20:13:36.810Z",
          "content": "<p>I agree Daniel that the solution is not so simple (at least for my brain). IME, weightings like these are usually a business decision, and data science decisions are made to support the business needs. Here though we have competing goals: HC wants stability; Kagglers want to win the competition. That leads to a hard balancing act. I made an assumption in <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2657865\" target=\"_blank\">this thread</a> that HC's desire for stable models is geared toward managing downside risk for the upcoming recession. Thoughts on that? </p>",
          "rawMarkdown": "I agree Daniel that the solution is not so simple (at least for my brain). IME, weightings like these are usually a business decision, and data science decisions are made to support the business needs. Here though we have competing goals: HC wants stability; Kagglers want to win the competition. That leads to a hard balancing act. I made an assumption in [this thread](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2657865) that HC's desire for stable models is geared toward managing downside risk for the upcoming recession. Thoughts on that? ",
          "replies": [
            {
              "id": 2657952,
              "postDate": "2024-02-18T20:58:04.680Z",
              "content": "<p>I can't comment on HC internal plans for models. There are many challenges to business not only recession. The ability to develop more stable/predictable models would have a positive impact on business in the long run. </p>",
              "rawMarkdown": "I can't comment on HC internal plans for models. There are many challenges to business not only recession. The ability to develop more stable/predictable models would have a positive impact on business in the long run. "
            },
            {
              "id": 2658605,
              "postDate": "2024-02-19T09:39:36.330Z",
              "content": "<p>I will be more open in my answer than Daniel :-)</p>\n<p>Home Credit's objective is stability of model. From business point of view it means also \"reliability\" of model as a risk management tool during turbulent times, meaning if something bad happens we have tools how to react.</p>\n<p>It has nothing to do with our expectation of recession - if we expect it (I am not saying we do), we would reduce our risk appetite, but not change or models.</p>\n<p>What I see, many of Kagglers focus on Covid, which is nice and obvious example of turbulent times. But I can assure you it's not the only one we have experienced in recent years - Home Credit is operating in several countries and something is happening all the time, so we need to be ready all the time…</p>",
              "rawMarkdown": "I will be more open in my answer than Daniel :-)\n\nHome Credit's objective is stability of model. From business point of view it means also \"reliability\" of model as a risk management tool during turbulent times, meaning if something bad happens we have tools how to react.\n\nIt has nothing to do with our expectation of recession - if we expect it (I am not saying we do), we would reduce our risk appetite, but not change or models.\n\nWhat I see, many of Kagglers focus on Covid, which is nice and obvious example of turbulent times. But I can assure you it's not the only one we have experienced in recent years - Home Credit is operating in several countries and something is happening all the time, so we need to be ready all the time...",
              "votes": 3
            }
          ]
        },
        {
          "id": 2658857,
          "postDate": "2024-02-19T13:04:32.277Z",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> why isnt  the solution simple, use traditional metrics like mean(weakly gini)/ std( weekly gini) as the metric for finding stable nicely performing models</p>",
          "rawMarkdown": "@jetakow why isnt  the solution simple, use traditional metrics like mean(weakly gini)/ std( weekly gini) as the metric for finding stable nicely performing models"
        },
        {
          "id": 2660680,
          "postDate": "2024-02-20T18:58:00.850Z",
          "content": "<p>One suggestion, if it hasn't been tested yet, is that instead of relying on a single metric, the organizers could consider using multiple metrics to assess different aspects of the models' performance. For example, one metric could focus on overall accuracy, while another could concentrate on temporal stability.</p>",
          "rawMarkdown": "One suggestion, if it hasn't been tested yet, is that instead of relying on a single metric, the organizers could consider using multiple metrics to assess different aspects of the models' performance. For example, one metric could focus on overall accuracy, while another could concentrate on temporal stability."
        }
      ]
    },
    {
      "id": 2649222,
      "postDate": "2024-02-12T17:22:17.293Z",
      "content": "<p>Thanks for sharing!. Unlike previous topics, this one actually proves the metric does not work as it should, making a model intentionally worse should not work.</p>",
      "rawMarkdown": "Thanks for sharing!. Unlike previous topics, this one actually proves the metric does not work as it should, making a model intentionally worse should not work.",
      "votes": 12,
      "replies": [
        {
          "id": 2650198,
          "postDate": "2024-02-13T09:58:42.400Z",
          "content": "<p>The metric is resistant to attacks based on imputing external data based on state of economy at date X - that was my claim. It comes from the weekly-based gini calculation.</p>\n<p>At the same time, as was discovered here, the metric is vulnerable to attacks based on manipulation of <code>a</code>.</p>",
          "rawMarkdown": "The metric is resistant to attacks based on imputing external data based on state of economy at date X - that was my claim. It comes from the weekly-based gini calculation.\n\nAt the same time, as was discovered here, the metric is vulnerable to attacks based on manipulation of `a`.",
          "votes": 2,
          "replies": [
            {
              "id": 2652891,
              "postDate": "2024-02-15T04:21:20.517Z",
              "content": "<p>you are right about this, I guess.</p>",
              "rawMarkdown": "you are right about this, I guess.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2648783,
      "postDate": "2024-02-12T12:20:49.847Z",
      "content": "<p>It is probably worth to tag organizers <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> and <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>, so that they could take a look. I have to admit, that it is very fair of you to report this issue. </p>\n<p><strong>UPD:</strong> I did a small <a href=\"https://www.kaggle.com/code/kononenko/metric-s-trick-home-credit-baseline-inference\" target=\"_blank\">experiment</a> to see the effect of the metric trick on one of the top-scoring public notebooks and indeed got a similar 0.03 boost for LB score.</p>",
      "rawMarkdown": "It is probably worth to tag organizers @tomasjeline2 and @jetakow, so that they could take a look. I have to admit, that it is very fair of you to report this issue. \n\n**UPD:** I did a small [experiment](https://www.kaggle.com/code/kononenko/metric-s-trick-home-credit-baseline-inference) to see the effect of the metric trick on one of the top-scoring public notebooks and indeed got a similar 0.03 boost for LB score.",
      "votes": 5,
      "replies": [
        {
          "id": 2648797,
          "postDate": "2024-02-12T12:37:20.760Z",
          "content": "<p>yea thanks</p>",
          "rawMarkdown": "yea thanks",
          "votes": 2
        }
      ]
    },
    {
      "id": 2649054,
      "postDate": "2024-02-12T15:36:39.410Z",
      "content": "<p>Thanks for the important remark!</p>\n<p>I tried experimenting, and I was able to repeat it quite easily on synthetic data.</p>\n<p>You can watch it on my <a href=\"https://www.kaggle.com/code/neibyr/visualise-metric-problem/notebook\" target=\"_blank\">notebook</a>.</p>",
      "rawMarkdown": "Thanks for the important remark!\n\nI tried experimenting, and I was able to repeat it quite easily on synthetic data.\n\nYou can watch it on my [notebook](https://www.kaggle.com/code/neibyr/visualise-metric-problem/notebook).",
      "votes": 3,
      "replies": [
        {
          "id": 2650498,
          "postDate": "2024-02-13T13:42:36.517Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2648861,
      "postDate": "2024-02-12T13:24:24.090Z",
      "content": "<p>Thanks a lot for heads up, let us discuss internally…</p>",
      "rawMarkdown": "Thanks a lot for heads up, let us discuss internally...",
      "votes": 4,
      "replies": [
        {
          "id": 2648895,
          "postDate": "2024-02-12T13:37:53.520Z",
          "content": "<p>ok, as i said, it's only really a problem when the model gets a lot worse over time (which the testing period seems to be) then the term with the 88 is too punishing, so that it becomes beneficial score-wise to decrease the model performance in the beginning.</p>\n<p>In the real world this would not matter and the 88 might be appropiate for your circumstances, but in a competition i think it is too punishing</p>",
          "rawMarkdown": "ok, as i said, it's only really a problem when the model gets a lot worse over time (which the testing period seems to be) then the term with the 88 is too punishing, so that it becomes beneficial score-wise to decrease the model performance in the beginning.\n\n In the real world this would not matter and the 88 might be appropiate for your circumstances, but in a competition i think it is too punishing"
        },
        {
          "id": 2648973,
          "postDate": "2024-02-12T14:39:43.247Z",
          "content": "<p>It is a reasonable/common trade-off to take for model generalization that</p>\n<ul>\n<li>perform poor in recent/\"similar\" data </li>\n<li>but become better (less worsen) at future/\"out of distribution\" data</li>\n</ul>\n<p>But what is happening in this trick is that</p>\n<ul>\n<li>perform poor in recent/\"similar\" data </li>\n<li>but stay the same in future/\"out of distribution\" data</li>\n</ul>\n<p>Actually, what is the problem of dropping the linear regression logic, simply using mean(weekly gini) - std(weekly gini)? <br>\nmean and std function give the same unit + scale, the subtraction should make sense here?</p>",
          "rawMarkdown": "It is a reasonable/common trade-off to take for model generalization that\n- perform poor in recent/\"similar\" data \n- but become better (less worsen) at future/\"out of distribution\" data\n\nBut what is happening in this trick is that\n- perform poor in recent/\"similar\" data \n- but stay the same in future/\"out of distribution\" data\n\nActually, what is the problem of dropping the linear regression logic, simply using mean(weekly gini) - std(weekly gini)? \nmean and std function give the same unit + scale, the subtraction should make sense here?"
        }
      ]
    },
    {
      "id": 2649016,
      "postDate": "2024-02-12T15:00:18.107Z",
      "content": "<blockquote>\n  <p>where i make my model worse for the first half of the testing period improves my score by almost 0.3</p>\n</blockquote>\n<p>you mean 0.03? from 0.576 to 0.606?</p>",
      "rawMarkdown": ">where i make my model worse for the first half of the testing period improves my score by almost 0.3\n\nyou mean 0.03? from 0.576 to 0.606?",
      "votes": 1,
      "replies": [
        {
          "id": 2649038,
          "postDate": "2024-02-12T15:23:44.500Z",
          "content": "<p>yes, thanks, fixed it, from 0.582 to 0.606</p>",
          "rawMarkdown": "yes, thanks, fixed it, from 0.582 to 0.606",
          "votes": 3,
          "replies": [
            {
              "id": 2649048,
              "postDate": "2024-02-12T15:32:12.017Z",
              "content": "<p>Cool. Anyway, thanks for sharing - it is most generous of you.</p>",
              "rawMarkdown": "Cool. Anyway, thanks for sharing - it is most generous of you.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2659254,
      "postDate": "2024-02-19T17:47:50.170Z",
      "content": "<p>the main thing is not to retrain on validation. Otherwise it will become useless.  I still think it's a dead end to use the offset. Maybe it's worth correcting the offset somehow, making it more dynamic? Do something like: </p>\n<pre><code>N =   #   weeks\nsub[] = sub.groupby()[].(lambda x: x.rolling(=N, min_periods=).mean())\n</code></pre>",
      "rawMarkdown": "the main thing is not to retrain on validation. Otherwise it will become useless.  I still think it's a dead end to use the offset. Maybe it's worth correcting the offset somehow, making it more dynamic? Do something like: \n```\nN = 4  # offset by weeks\nsub['dynamic_SHIFT'] = sub.groupby('WEEK_NUM')['score'].transform(lambda x: x.rolling(window=N, min_periods=1).mean())\n```",
      "votes": 2
    },
    {
      "id": 2648946,
      "postDate": "2024-02-12T14:15:15.737Z",
      "content": "<p>Btw, for those who are confused what exactly has been done here.<br>\nOnly -0.02 part should not change the metric value, it -0.02 + clip 0</p>\n<pre><code> sklearn.metrics  roc_auc_score\n numpy  np\nt = (np.random.rand() &gt; ).astype()\ny = np.random.rand()\n(roc_auc_score(t, y)) \n(roc_auc_score(t, y - ))\n(roc_auc_score(t, (y - ).clip()))\n\n\n\n</code></pre>\n<p>Effectively, it \"floors\" some predictions to 0 for the first half of the testing period.<br>\nIf this makes the model performance in the first half get worse + the second half is poor,<br>\nthe slope of the fall becomes smaller.</p>\n<p><a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a>  Have you tried running the same post processing to other week nums configurations? like only the first week or applied it to all weeks?</p>",
      "rawMarkdown": "Btw, for those who are confused what exactly has been done here.\nOnly -0.02 part should not change the metric value, it -0.02 + clip 0\n\n```python\nfrom sklearn.metrics import roc_auc_score\nimport numpy as np\nt = (np.random.rand(100) > .5).astype(int)\ny = np.random.rand(100)\nprint(roc_auc_score(t, y)) \nprint(roc_auc_score(t, y - 0.02))\nprint(roc_auc_score(t, (y - 0.02).clip(0)))\n# 0.5232371794871795\n# 0.5232371794871795\n# 0.5240384615384616\n```\n\nEffectively, it \"floors\" some predictions to 0 for the first half of the testing period.\nIf this makes the model performance in the first half get worse + the second half is poor,\nthe slope of the fall becomes smaller.\n\n@at7459  Have you tried running the same post processing to other week nums configurations? like only the first week or applied it to all weeks?",
      "votes": 2,
      "replies": [
        {
          "id": 2648957,
          "postDate": "2024-02-12T14:29:07.140Z",
          "content": "<p>nah i didnt realize that about roc_auc, but yeah the point is to make the performance worse in the beginning to make the slope smaller.</p>\n<p>i didnt really find any periods where the model gets a lot worse over time during training, what i did instead to test something like this is to make the end period a lot worse and the beginning period slightly worse to roughly see the contributions from each term to get a good feel for the metric</p>",
          "rawMarkdown": "nah i didnt realize that about roc_auc, but yeah the point is to make the performance worse in the beginning to make the slope smaller.\n\ni didnt really find any periods where the model gets a lot worse over time during training, what i did instead to test something like this is to make the end period a lot worse and the beginning period slightly worse to roughly see the contributions from each term to get a good feel for the metric",
          "votes": 1
        },
        {
          "id": 2649063,
          "postDate": "2024-02-12T15:42:43.027Z",
          "content": "<p>I'm trying to replicate this PP in train data (using my OOF predictions, SGKF on week).  </p>\n<p><code>WEEK min/max/cutoff: 0/91, 45  (applied same logic as above)</code></p>\n<pre><code>Baseline CV score: \navg_gini:  \nres_std:  \na:  \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fe9f52cdf8ee7e58326855716b0ec5fcb%2Fgini_pp.png?generation=1707752081526627&amp;alt=media\"></p>\n<p>Legend:<br>\nblue: baseline no pp<br>\norange: pp with -0.02<br>\ngreen: pp with -0.01 </p>\n<p>Overall score reduced in both cases. Need to test other week cutoffs.. </p>\n<p>PS: Not sure if the problem is that my CV scheme is not time-series, i.e. valid only on <strong>future</strong> weeks.. </p>\n<p>[EDIT]: do you also get similar plots for gini in time?</p>",
          "rawMarkdown": "I'm trying to replicate this PP in train data (using my OOF predictions, SGKF on week).  \n\n`WEEK min/max/cutoff: 0/91, 45  (applied same logic as above)`\n\n```python\nBaseline CV score: 0.6465\navg_gini: 0.66375 \nres_std: 0.03445 \na: 0.0006534 \n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fe9f52cdf8ee7e58326855716b0ec5fcb%2Fgini_pp.png?generation=1707752081526627&alt=media)\n\n\nLegend:\nblue: baseline no pp\norange: pp with -0.02\ngreen: pp with -0.01 \n\nOverall score reduced in both cases. Need to test other week cutoffs.. \n\nPS: Not sure if the problem is that my CV scheme is not time-series, i.e. valid only on **future** weeks.. \n\n[EDIT]: do you also get similar plots for gini in time?",
          "votes": 3,
          "replies": [
            {
              "id": 2649086,
              "postDate": "2024-02-12T15:52:29.423Z",
              "content": "<p>well, the thing is is that the data we're training on is almost purely non covid, and the data during the testing period is covid and post-covid (and also a lot longer into the future than you would be able to make with any split in the training data)</p>\n<p>so our models will fall off hard during testing (they obviously chose this split on purpose due to their stability agenda)</p>\n<p>i dont think it's something you will be able to experiment well with only with the training data because it's mostly just non covid</p>",
              "rawMarkdown": "well, the thing is is that the data we're training on is almost purely non covid, and the data during the testing period is covid and post-covid (and also a lot longer into the future than you would be able to make with any split in the training data)\n\nso our models will fall off hard during testing (they obviously chose this split on purpose due to their stability agenda)\n\ni dont think it's something you will be able to experiment well with only with the training data because it's mostly just non covid",
              "votes": 7
            }
          ]
        }
      ]
    },
    {
      "id": 2650496,
      "postDate": "2024-02-13T13:40:16.223Z",
      "content": "<p>and what is so early =))) it was necessary the day before the end of the competition =))</p>",
      "rawMarkdown": "and what is so early =))) it was necessary the day before the end of the competition =))"
    },
    {
      "id": 2728558,
      "postDate": "2024-04-02T10:04:59.510Z",
      "content": "<p>I didn't think about the possibility of hacking the metrics, so thanks for the help.<br>\nThank you so much.</p>",
      "rawMarkdown": "I didn't think about the possibility of hacking the metrics, so thanks for the help.\nThank you so much."
    },
    {
      "id": 2661659,
      "postDate": "2024-02-21T12:11:31.640Z",
      "content": "<p>would it get big shake up?im pretty skepital,wasnt it a kind of post process?<br>\ni dont really onderstand.<br>\nthank you very much</p>",
      "rawMarkdown": "would it get big shake up?im pretty skepital,wasnt it a kind of post process?\ni dont really onderstand.\nthank you very much"
    },
    {
      "id": 2657526,
      "postDate": "2024-02-18T14:59:14.413Z",
      "content": "<p>It is a really helpful one !! So late for me to see it!  </p>",
      "rawMarkdown": "It is a really helpful one !! So late for me to see it!  "
    },
    {
      "id": 2651929,
      "postDate": "2024-02-14T13:01:37.663Z",
      "content": "<p><a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a> What would you think about non-uniform weighting of predictions, so that predictions on samples in the more distant future had more weight? This would fix the hack in question, while preserving the objective, roughly. I don't know about side effects, though… just an idea.</p>",
      "rawMarkdown": "@at7459 @jacobyjaeger What would you think about non-uniform weighting of predictions, so that predictions on samples in the more distant future had more weight? This would fix the hack in question, while preserving the objective, roughly. I don't know about side effects, though... just an idea.",
      "replies": [
        {
          "id": 2652025,
          "postDate": "2024-02-14T14:20:13.253Z",
          "content": "<p>yeah sure something like that could work, you definitely couldn't abuse the metric then. but you might aswell just cut out the first half of the data and i dont know if that would ultimately achieve the desired result that the most stable models would win.</p>\n<p>i think it would just shift the focus more on doing better on a shorter, more restriced time frame, but i guess that's not that different from the current metric overall.</p>",
          "rawMarkdown": "yeah sure something like that could work, you definitely couldn't abuse the metric then. but you might aswell just cut out the first half of the data and i dont know if that would ultimately achieve the desired result that the most stable models would win.\n\ni think it would just shift the focus more on doing better on a shorter, more restriced time frame, but i guess that's not that different from the current metric overall.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2650361,
      "postDate": "2024-02-13T12:15:24.400Z",
      "content": "<p>Thanks for sharing!. Unlike previous topics, this one actually proves the metric does not work as it should, making a model intentionally worse should not work.</p>\n<p>Great work and Keep it up!</p>",
      "rawMarkdown": "Thanks for sharing!. Unlike previous topics, this one actually proves the metric does not work as it should, making a model intentionally worse should not work.\n\nGreat work and Keep it up!"
    },
    {
      "id": 2649414,
      "postDate": "2024-02-12T20:14:09.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2649080,
      "postDate": "2024-02-12T15:50:31.020Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2649090,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-12T15:53:38.777000",
      "content": "<p>We greatly appreciate your finding and that you came forward with. </p>\n<p>I am afraid there is no simple solution, let's have a look at the extremes. 1) we could omit mean gini or 2) we could omit the part with falling rate. When going with 2) the issue becomes that the models are pushed towards excellent short-term performance, but yet not stable in time. When going with 1) the issue becomes tweaking of the score so that we have as stable gini in time as possible. Changing the weights will not result in solving the problem, but rather transformation from one issue into another. </p>\n<p>We will discuss this also with Kaggle. Tweaking the score manually based on time is not desired in this competition and we are already thinking about steps how to partially mitigate this issue. We will make an official statement later.</p>",
      "votes": 16,
      "replies": [
        {
          "id": 2649151,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-02-12T16:36:19.870000",
          "content": "<p>alright you can do it however you like but i think changing the weights could definitely solve the problem.</p>\n<p>for example, currently the score for the test period probably looks something like <br>\nmean(gini) : <strong>+0.7</strong><br>\nmin(0,a) : <strong>-0.09</strong><br>\nstd(residuals: <strong>-0.01</strong><br>\nwhich sums to <strong>0.6</strong><br>\nwhen i make my model worse for the first half it turns into<br>\nmean(gini) : <strong>+0.69</strong><br>\nmin(0,a) : <strong>-0.05</strong><br>\nstd(residuals: <strong>-0.02</strong><br>\nwhich sums to <strong>0.62</strong></p>\n<p>changing the weights of the min(0,a) term to say 44 would turn this into a net negative strategy for me</p>\n<p>granted it would put less emphasis on stability and take some experimenting but it could still work out overall</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2649189,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-12T17:05:23.350000",
              "content": "<p>Of course this would make this particular strategy less interesting, but another may rise with new set of weights. Besides that smaller weight for falling rate would encourage less stable models, which is something we ultimately don't aim for. Most likely the countermeasure will not contain change in weights. </p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2649202,
              "author_name": "Anthony Chiu",
              "author_url": "",
              "post_date": "2024-02-12T17:14:10.707000",
              "content": "<p>I think a valid reason for not changing the current metric would be this post-processing is risky; the public lb is only 30 % of the whole test data, and you mentioned this method doesn't work well in the training set. 0.03 change is huge but also not so huge.</p>\n<ul>\n<li>If the weekly gini in the hidden set keep falling, then this post-processing might work well</li>\n<li>if the weekly gini rises again, this post-processing might not work.</li>\n</ul>\n<p>The actual number of weeks and the percentage of predictions to force 0 are hard to validate. If someone can have a trick that works well in training and testing data, this might also show some proper stability.</p>\n<p>But I think some teams will use 1 of their final submissions to take this risk.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2650343,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-13T11:56:25.520000",
              "content": "<p>All of those are valid points. We are discussing the issue internally. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2651850,
              "author_name": "Andrey Nesterov",
              "author_url": "",
              "post_date": "2024-02-14T12:05:25.470000",
              "content": "<p>Another alternative (draft idea) for changing the weights could be to change them dynamically, e.g. by multiplying the penalty by K/gini_mean (K - some constant). So higher gini_mean would have less penalty. But it seems too complex and needs more rigorous mathematical proofs. And instead of using linear regression, Huber regression could be used (slight optimization).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2649405,
          "author_name": "Oleksiy Kononenko",
          "author_url": "",
          "post_date": "2024-02-12T20:01:42.303000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>, I understand that the point of this competition is to come up with a stable nicely performing model. However, I’m not sure how the penalizing term helps. What is the advantage of having a model that performs worse on the first weeks and better on the last ones? As we can see from this post, this requirement can actually result in an overall worse model.</p>\n<p>Is it because you assume the data after the last weeks will stay similar and will never be close to the first weeks data? I would appreciate it if you could elaborate a little bit on this.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2649463,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-12T20:46:25.787000",
              "content": "<p>We determined that this year's problem we wanted to offer to kagglers is stability. Using no penalizing terms would lead to only unstable models. We will not leak any information on purpose about test sample. I can only ask you for patience while we will try to make things more fair for everyone. </p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2650082,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-02-13T08:34:27.997000",
              "content": "<p>imo the best way to deal with stability is one of those competition structures where you're continously fed new data and can retrain your model. obviously, as you stated, due to how your company works and how long a model gets used, this is not really possible.</p>\n<p>but in this competition structure right now - due to the fact that the testing period is so different from the training period - whether you fix the problem or not, it will mostly just be about leaderboard probing and checking what works well on the leaderboard. ive had very bad experience with these types of competitions so far, honestly the only reason i joined is due to the added stability thing.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2650338,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-13T11:54:28.227000",
              "content": "<blockquote>\n  <p>honestly the only reason i joined is due to the added stability thing.</p>\n</blockquote>\n<p>Then I hope it will be still challenging for you enough that you remain in the competition. I ask you for patience and please understand that we have only limited options at the moment. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2650578,
              "author_name": "Sean McManus",
              "author_url": "",
              "post_date": "2024-02-13T14:41:28.907000",
              "content": "<p>I'm not sure the point being made here is completely congruent with the stated intent of the model and the information included for modelling purposes (vintage). </p>\n<p>There's a big disconnect between the metric, the stated intent and the data being used for modelling.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2650753,
              "author_name": "Anthony Chiu",
              "author_url": "",
              "post_date": "2024-02-13T17:25:47.313000",
              "content": "<p>I'm not suggesting anything here. But from a metric design perspective, I am curious: Why can't we use std(weekly gini) here?</p>\n<pre><code>(weekly gini) - (weekly gini)\n</code></pre>\n<p>This keeps the idea of weekly gini and stability and doesn't need weight.</p>\n<p>I agree with <a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> ; lb probing doesn't feel good. The 30% public leader board becomes tricky to interpret.  I am also interested in the stability concept here. Online learning is a way to handle stability, but it also adds extra costs. I agree that Online learning could be an easier method to handle this,  but I am also excited about the concept of training a model from pre-COVID weeks and surviving the COVID weeks. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2650795,
              "author_name": "Oleksiy Kononenko",
              "author_url": "",
              "post_date": "2024-02-13T17:47:26.793000",
              "content": "<blockquote>\n  <p>Using no penalizing terms would lead to only unstable models.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> I only meant the second term. I understand you don’t want the model’s performance to deteriorate with time, but since credit models are only valid for one year (as you say in description), doesn’t mean and std address stability issue good enough if the evaluation period is similar to the model’s lifetime?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2650849,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-13T18:12:11.017000",
              "content": "<p>At this very moment, we consider also this scenario. It may be right that most of the credit models are made to last only one year, but that is no universal rule. They can be deployed for much longer than one year. Removing the falling rate term would promote unstable models, that perform well shortly, but quickly degrade. We need to penalize this behaviour and that is the reason why such term is present. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2650868,
              "author_name": "Oleksiy Kononenko",
              "author_url": "",
              "post_date": "2024-02-13T18:18:44.670000",
              "content": "<p>I see, may be that term could be compensated by making timeframe of test data longer, i.e. similar to whatever the typical period of model’s deployment is. It could be already the case, not sure, but from my opinion testing on real data is better than trying to extrapolate performance with a linear function.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2657889,
          "author_name": "JohnM",
          "author_url": "",
          "post_date": "2024-02-18T20:13:36.810000",
          "content": "<p>I agree Daniel that the solution is not so simple (at least for my brain). IME, weightings like these are usually a business decision, and data science decisions are made to support the business needs. Here though we have competing goals: HC wants stability; Kagglers want to win the competition. That leads to a hard balancing act. I made an assumption in <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867#2657865\" target=\"_blank\">this thread</a> that HC's desire for stable models is geared toward managing downside risk for the upcoming recession. Thoughts on that? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2657952,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-18T20:58:04.680000",
              "content": "<p>I can't comment on HC internal plans for models. There are many challenges to business not only recession. The ability to develop more stable/predictable models would have a positive impact on business in the long run. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2658605,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-02-19T09:39:36.330000",
              "content": "<p>I will be more open in my answer than Daniel :-)</p>\n<p>Home Credit's objective is stability of model. From business point of view it means also \"reliability\" of model as a risk management tool during turbulent times, meaning if something bad happens we have tools how to react.</p>\n<p>It has nothing to do with our expectation of recession - if we expect it (I am not saying we do), we would reduce our risk appetite, but not change or models.</p>\n<p>What I see, many of Kagglers focus on Covid, which is nice and obvious example of turbulent times. But I can assure you it's not the only one we have experienced in recent years - Home Credit is operating in several countries and something is happening all the time, so we need to be ready all the time…</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2658857,
          "author_name": "pranavpoduval",
          "author_url": "",
          "post_date": "2024-02-19T13:04:32.277000",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> why isnt  the solution simple, use traditional metrics like mean(weakly gini)/ std( weekly gini) as the metric for finding stable nicely performing models</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2660680,
          "author_name": "Fernando Rodrigues",
          "author_url": "",
          "post_date": "2024-02-20T18:58:00.850000",
          "content": "<p>One suggestion, if it hasn't been tested yet, is that instead of relying on a single metric, the organizers could consider using multiple metrics to assess different aspects of the models' performance. For example, one metric could focus on overall accuracy, while another could concentrate on temporal stability.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2649222,
      "author_name": "NxGTR",
      "author_url": "",
      "post_date": "2024-02-12T17:22:17.293000",
      "content": "<p>Thanks for sharing!. Unlike previous topics, this one actually proves the metric does not work as it should, making a model intentionally worse should not work.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2650198,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2024-02-13T09:58:42.400000",
          "content": "<p>The metric is resistant to attacks based on imputing external data based on state of economy at date X - that was my claim. It comes from the weekly-based gini calculation.</p>\n<p>At the same time, as was discovered here, the metric is vulnerable to attacks based on manipulation of <code>a</code>.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2652891,
              "author_name": "Malik Muhammad Noman Ghafoor",
              "author_url": "",
              "post_date": "2024-02-15T04:21:20.517000",
              "content": "<p>you are right about this, I guess.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2648783,
      "author_name": "Oleksiy Kononenko",
      "author_url": "",
      "post_date": "2024-02-12T12:20:49.847000",
      "content": "<p>It is probably worth to tag organizers <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> and <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>, so that they could take a look. I have to admit, that it is very fair of you to report this issue. </p>\n<p><strong>UPD:</strong> I did a small <a href=\"https://www.kaggle.com/code/kononenko/metric-s-trick-home-credit-baseline-inference\" target=\"_blank\">experiment</a> to see the effect of the metric trick on one of the top-scoring public notebooks and indeed got a similar 0.03 boost for LB score.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2648797,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-02-12T12:37:20.760000",
          "content": "<p>yea thanks</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2649054,
      "author_name": "Kirill Tushin",
      "author_url": "",
      "post_date": "2024-02-12T15:36:39.410000",
      "content": "<p>Thanks for the important remark!</p>\n<p>I tried experimenting, and I was able to repeat it quite easily on synthetic data.</p>\n<p>You can watch it on my <a href=\"https://www.kaggle.com/code/neibyr/visualise-metric-problem/notebook\" target=\"_blank\">notebook</a>.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2650498,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-13T13:42:36.517000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2648861,
      "author_name": "Tomas Jelinek",
      "author_url": "",
      "post_date": "2024-02-12T13:24:24.090000",
      "content": "<p>Thanks a lot for heads up, let us discuss internally…</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2648895,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-02-12T13:37:53.520000",
          "content": "<p>ok, as i said, it's only really a problem when the model gets a lot worse over time (which the testing period seems to be) then the term with the 88 is too punishing, so that it becomes beneficial score-wise to decrease the model performance in the beginning.</p>\n<p>In the real world this would not matter and the 88 might be appropiate for your circumstances, but in a competition i think it is too punishing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2648973,
          "author_name": "Anthony Chiu",
          "author_url": "",
          "post_date": "2024-02-12T14:39:43.247000",
          "content": "<p>It is a reasonable/common trade-off to take for model generalization that</p>\n<ul>\n<li>perform poor in recent/\"similar\" data </li>\n<li>but become better (less worsen) at future/\"out of distribution\" data</li>\n</ul>\n<p>But what is happening in this trick is that</p>\n<ul>\n<li>perform poor in recent/\"similar\" data </li>\n<li>but stay the same in future/\"out of distribution\" data</li>\n</ul>\n<p>Actually, what is the problem of dropping the linear regression logic, simply using mean(weekly gini) - std(weekly gini)? <br>\nmean and std function give the same unit + scale, the subtraction should make sense here?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2649016,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2024-02-12T15:00:18.107000",
      "content": "<blockquote>\n  <p>where i make my model worse for the first half of the testing period improves my score by almost 0.3</p>\n</blockquote>\n<p>you mean 0.03? from 0.576 to 0.606?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2649038,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-02-12T15:23:44.500000",
          "content": "<p>yes, thanks, fixed it, from 0.582 to 0.606</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2649048,
              "author_name": "narsil (jobs-in-data.com)",
              "author_url": "",
              "post_date": "2024-02-12T15:32:12.017000",
              "content": "<p>Cool. Anyway, thanks for sharing - it is most generous of you.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2659254,
      "author_name": "wasjaip",
      "author_url": "",
      "post_date": "2024-02-19T17:47:50.170000",
      "content": "<p>the main thing is not to retrain on validation. Otherwise it will become useless.  I still think it's a dead end to use the offset. Maybe it's worth correcting the offset somehow, making it more dynamic? Do something like: </p>\n<pre><code>N =   #   weeks\nsub[] = sub.groupby()[].(lambda x: x.rolling(=N, min_periods=).mean())\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2648946,
      "author_name": "Anthony Chiu",
      "author_url": "",
      "post_date": "2024-02-12T14:15:15.737000",
      "content": "<p>Btw, for those who are confused what exactly has been done here.<br>\nOnly -0.02 part should not change the metric value, it -0.02 + clip 0</p>\n<pre><code> sklearn.metrics  roc_auc_score\n numpy  np\nt = (np.random.rand() &gt; ).astype()\ny = np.random.rand()\n(roc_auc_score(t, y)) \n(roc_auc_score(t, y - ))\n(roc_auc_score(t, (y - ).clip()))\n\n\n\n</code></pre>\n<p>Effectively, it \"floors\" some predictions to 0 for the first half of the testing period.<br>\nIf this makes the model performance in the first half get worse + the second half is poor,<br>\nthe slope of the fall becomes smaller.</p>\n<p><a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a>  Have you tried running the same post processing to other week nums configurations? like only the first week or applied it to all weeks?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2648957,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-02-12T14:29:07.140000",
          "content": "<p>nah i didnt realize that about roc_auc, but yeah the point is to make the performance worse in the beginning to make the slope smaller.</p>\n<p>i didnt really find any periods where the model gets a lot worse over time during training, what i did instead to test something like this is to make the end period a lot worse and the beginning period slightly worse to roughly see the contributions from each term to get a good feel for the metric</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2649063,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2024-02-12T15:42:43.027000",
          "content": "<p>I'm trying to replicate this PP in train data (using my OOF predictions, SGKF on week).  </p>\n<p><code>WEEK min/max/cutoff: 0/91, 45  (applied same logic as above)</code></p>\n<pre><code>Baseline CV score: \navg_gini:  \nres_std:  \na:  \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fe9f52cdf8ee7e58326855716b0ec5fcb%2Fgini_pp.png?generation=1707752081526627&amp;alt=media\"></p>\n<p>Legend:<br>\nblue: baseline no pp<br>\norange: pp with -0.02<br>\ngreen: pp with -0.01 </p>\n<p>Overall score reduced in both cases. Need to test other week cutoffs.. </p>\n<p>PS: Not sure if the problem is that my CV scheme is not time-series, i.e. valid only on <strong>future</strong> weeks.. </p>\n<p>[EDIT]: do you also get similar plots for gini in time?</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2649086,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-02-12T15:52:29.423000",
              "content": "<p>well, the thing is is that the data we're training on is almost purely non covid, and the data during the testing period is covid and post-covid (and also a lot longer into the future than you would be able to make with any split in the training data)</p>\n<p>so our models will fall off hard during testing (they obviously chose this split on purpose due to their stability agenda)</p>\n<p>i dont think it's something you will be able to experiment well with only with the training data because it's mostly just non covid</p>",
              "votes": 7,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2650496,
      "author_name": "wasjaip",
      "author_url": "",
      "post_date": "2024-02-13T13:40:16.223000",
      "content": "<p>and what is so early =))) it was necessary the day before the end of the competition =))</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2728558,
      "author_name": "Ryo trust CV",
      "author_url": "",
      "post_date": "2024-04-02T10:04:59.510000",
      "content": "<p>I didn't think about the possibility of hacking the metrics, so thanks for the help.<br>\nThank you so much.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2661659,
      "author_name": "OBS$E$SED",
      "author_url": "",
      "post_date": "2024-02-21T12:11:31.640000",
      "content": "<p>would it get big shake up?im pretty skepital,wasnt it a kind of post process?<br>\ni dont really onderstand.<br>\nthank you very much</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2657526,
      "author_name": "Vicccc_L",
      "author_url": "",
      "post_date": "2024-02-18T14:59:14.413000",
      "content": "<p>It is a really helpful one !! So late for me to see it!  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2651929,
      "author_name": "loh-maa",
      "author_url": "",
      "post_date": "2024-02-14T13:01:37.663000",
      "content": "<p><a href=\"https://www.kaggle.com/at7459\" target=\"_blank\">@at7459</a> <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a> What would you think about non-uniform weighting of predictions, so that predictions on samples in the more distant future had more weight? This would fix the hack in question, while preserving the objective, roughly. I don't know about side effects, though… just an idea.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2652025,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-02-14T14:20:13.253000",
          "content": "<p>yeah sure something like that could work, you definitely couldn't abuse the metric then. but you might aswell just cut out the first half of the data and i dont know if that would ultimately achieve the desired result that the most stable models would win.</p>\n<p>i think it would just shift the focus more on doing better on a shorter, more restriced time frame, but i guess that's not that different from the current metric overall.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2650361,
      "author_name": "Tanishq dublish",
      "author_url": "",
      "post_date": "2024-02-13T12:15:24.400000",
      "content": "<p>Thanks for sharing!. Unlike previous topics, this one actually proves the metric does not work as it should, making a model intentionally worse should not work.</p>\n<p>Great work and Keep it up!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2649414,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-12T20:14:09.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2649080,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-12T15:50:31.020000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2648775": "i could abuse this, but i guess by the end most people will be doing something like this ,and it doesn't really have anything to do with data science. also, in the end, people who didn't see this will just get pissed off.\n\nThe problem is that the 88 * min(0,a) term is too large, so you can just make your model worse for the first weeks and improve your score a lot. i like the idea of the metric in itself but this is a competition after all and people will just pull the dirtiest tricks out of their bag to come out ahead.\n\nfor example something like this,\n\n`condition = df_subm['WEEK_NUM'] < (df_subm['WEEK_NUM'].max()-df_subm['WEEK_NUM'].min())/2+df_subm['WEEK_NUM'].min()\ndf_subm.loc[condition, 'score'] = (df_subm.loc[condition, 'score'] - 0.02).clip(0)\ndf_subm=df_subm[[\"case_id\",\"score\"]]\ndf_subm = df_subm.set_index(\"case_id\")\ndf_subm.to_csv(\"submission.csv\")\nprint(df_subm)`\n\nwhere i make my model worse for the first half of the testing period, improves my score by almost 0.03\n\npossible solution:\n\nas i said the 88 is too large, when you look at the contributions from the metric for a model which gets worse over time, this term is about 4-6 times larger than the other term. maybe 18 or something would be more fitting, it takes some experimenting i guess",
    "2649090": "We greatly appreciate your finding and that you came forward with. \n\nI am afraid there is no simple solution, let's have a look at the extremes. 1) we could omit mean gini or 2) we could omit the part with falling rate. When going with 2) the issue becomes that the models are pushed towards excellent short-term performance, but yet not stable in time. When going with 1) the issue becomes tweaking of the score so that we have as stable gini in time as possible. Changing the weights will not result in solving the problem, but rather transformation from one issue into another. \n\nWe will discuss this also with Kaggle. Tweaking the score manually based on time is not desired in this competition and we are already thinking about steps how to partially mitigate this issue. We will make an official statement later.",
    "2649222": "Thanks for sharing!. Unlike previous topics, this one actually proves the metric does not work as it should, making a model intentionally worse should not work.",
    "2648783": "It is probably worth to tag organizers @tomasjeline2 and @jetakow, so that they could take a look. I have to admit, that it is very fair of you to report this issue. \n\n**UPD:** I did a small [experiment](https://www.kaggle.com/code/kononenko/metric-s-trick-home-credit-baseline-inference) to see the effect of the metric trick on one of the top-scoring public notebooks and indeed got a similar 0.03 boost for LB score.",
    "2649054": "Thanks for the important remark!\n\nI tried experimenting, and I was able to repeat it quite easily on synthetic data.\n\nYou can watch it on my [notebook](https://www.kaggle.com/code/neibyr/visualise-metric-problem/notebook).",
    "2648861": "Thanks a lot for heads up, let us discuss internally...",
    "2649016": ">where i make my model worse for the first half of the testing period improves my score by almost 0.3\n\nyou mean 0.03? from 0.576 to 0.606?",
    "2659254": "the main thing is not to retrain on validation. Otherwise it will become useless.  I still think it's a dead end to use the offset. Maybe it's worth correcting the offset somehow, making it more dynamic? Do something like: \n```\nN = 4  # offset by weeks\nsub['dynamic_SHIFT'] = sub.groupby('WEEK_NUM')['score'].transform(lambda x: x.rolling(window=N, min_periods=1).mean())\n```",
    "2648946": "Btw, for those who are confused what exactly has been done here.\nOnly -0.02 part should not change the metric value, it -0.02 + clip 0\n\n```python\nfrom sklearn.metrics import roc_auc_score\nimport numpy as np\nt = (np.random.rand(100) > .5).astype(int)\ny = np.random.rand(100)\nprint(roc_auc_score(t, y)) \nprint(roc_auc_score(t, y - 0.02))\nprint(roc_auc_score(t, (y - 0.02).clip(0)))\n# 0.5232371794871795\n# 0.5232371794871795\n# 0.5240384615384616\n```\n\nEffectively, it \"floors\" some predictions to 0 for the first half of the testing period.\nIf this makes the model performance in the first half get worse + the second half is poor,\nthe slope of the fall becomes smaller.\n\n@at7459  Have you tried running the same post processing to other week nums configurations? like only the first week or applied it to all weeks?",
    "2650496": "and what is so early =))) it was necessary the day before the end of the competition =))",
    "2728558": "I didn't think about the possibility of hacking the metrics, so thanks for the help.\nThank you so much.",
    "2661659": "would it get big shake up?im pretty skepital,wasnt it a kind of post process?\ni dont really onderstand.\nthank you very much",
    "2657526": "It is a really helpful one !! So late for me to see it!  ",
    "2651929": "@at7459 @jacobyjaeger What would you think about non-uniform weighting of predictions, so that predictions on samples in the more distant future had more weight? This would fix the hack in question, while preserving the objective, roughly. I don't know about side effects, though... just an idea.",
    "2650361": "Thanks for sharing!. Unlike previous topics, this one actually proves the metric does not work as it should, making a model intentionally worse should not work.\n\nGreat work and Keep it up!",
    "2649414": "",
    "2649080": ""
  }
}