{
  "id": 475118,
  "title": "strange metric?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/475118",
  "author_name": "",
  "post_date": "2024-02-07T06:32:36.608983600Z",
  "votes": 18,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Consider gini score decrease linearly from score_max to score_min in k weeks, then three terms in metric will be:<br>\n(score_max+score_min)/2, -(score_max-score_min)*88/k, 0<br>\nIn case score_max==score_min they sum to score_min<br>\nif we subtract it from the general one we will get:<br>\n(score_max-score_min)(0.5-88/k)<br>\nif 88/k &gt; 0.5, then metric is better for the second case where all gini score is smaller .<br>\nAnd we have 92 weeks in training set (I think it will be similar in test set?) so it lies in this situation.<br>\nIs this an expected behaviour? Personally I expect the metric will be better with better gini score for every week and it may avoid some meaningless postprocessing. And metric property changes with evaluation period length is also a bit counter intuitive.</p>",
  "messages": [
    {
      "id": "2640903",
      "postDate": "02/07/2024 06:32:36",
      "content": "<p>Consider gini score decrease linearly from score_max to score_min in k weeks, then three terms in metric will be:<br>\n(score_max+score_min)/2, -(score_max-score_min)*88/k, 0<br>\nIn case score_max==score_min they sum to score_min<br>\nif we subtract it from the general one we will get:<br>\n(score_max-score_min)(0.5-88/k)<br>\nif 88/k &gt; 0.5, then metric is better for the second case where all gini score is smaller .<br>\nAnd we have 92 weeks in training set (I think it will be similar in test set?) so it lies in this situation.<br>\nIs this an expected behaviour? Personally I expect the metric will be better with better gini score for every week and it may avoid some meaningless postprocessing. And metric property changes with evaluation period length is also a bit counter intuitive.</p>",
      "rawMarkdown": "Consider gini score decrease linearly from score_max to score_min in k weeks, then three terms in metric will be:\n(score_max+score_min)/2, -(score_max-score_min)*88/k, 0\nIn case score_max==score_min they sum to score_min\nif we subtract it from the general one we will get:\n(score_max-score_min)(0.5-88/k)\nif 88/k > 0.5, then metric is better for the second case where all gini score is smaller .\nAnd we have 92 weeks in training set (I think it will be similar in test set?) so it lies in this situation.\nIs this an expected behaviour? Personally I expect the metric will be better with better gini score for every week and it may avoid some meaningless postprocessing. And metric property changes with evaluation period length is also a bit counter intuitive.",
      "votes": null
    },
    {
      "id": "2641076",
      "postDate": "02/07/2024 09:13:47",
      "content": "<p>The point of the metric is to simulate how we rate the performance of production models. The high gini performance just after deployment is pointless if over a period of 6 months, the performance will drop to half of that. Instead of integrating the gini in time we decided to settle on this metric which takes into account not only gini, but also the decrease of performance and variance of performance. </p>",
      "rawMarkdown": "The point of the metric is to simulate how we rate the performance of production models. The high gini performance just after deployment is pointless if over a period of 6 months, the performance will drop to half of that. Instead of integrating the gini in time we decided to settle on this metric which takes into account not only gini, but also the decrease of performance and variance of performance.",
      "votes": null
    },
    {
      "id": "2641090",
      "postDate": "02/07/2024 09:32:03",
      "content": "<p>So we shall consider model with constant gini score 0.5 in 6 months better than model with gini score decreased from 1 to 0.5 within the same period?</p>",
      "rawMarkdown": "So we shall consider model with constant gini score 0.5 in 6 months better than model with gini score decreased from 1 to 0.5 within the same period?",
      "votes": null
    },
    {
      "id": "2641121",
      "postDate": "02/07/2024 09:50:23",
      "content": "<p>Hi Ruby,<br>\nYes. Because if this trend continues, we can expect that in another 6 months gini will drop further to 0. Such model with so unstable performance is not suitable for usage in production. That's why we penalize instability in our metric.</p>",
      "rawMarkdown": "Hi Ruby,\nYes. Because if this trend continues, we can expect that in another 6 months gini will drop further to 0. Such model with so unstable performance is not suitable for usage in production. That's why we penalize instability in our metric.",
      "votes": null
    },
    {
      "id": "2641136",
      "postDate": "02/07/2024 10:03:19",
      "content": "<p>As explained in my post the similar conclusion extends to period of 88*2 weeks (3.4 years), I guess this far beyond your model lifecycle? I think the metric design is reasonable, but 88 might be too large? Is there any reason to choose 88 as weight?</p>",
      "rawMarkdown": "As explained in my post the similar conclusion extends to period of 88*2 weeks (3.4 years), I guess this far beyond your model lifecycle? I think the metric design is reasonable, but 88 might be too large? Is there any reason to choose 88 as weight?",
      "votes": null
    },
    {
      "id": "2641153",
      "postDate": "02/07/2024 10:11:49",
      "content": "<p>We have presented models with different performance profiles to our risk managers and from their answers we estimated sensitivity to instability. The metric is outcome of this exercise.</p>",
      "rawMarkdown": "We have presented models with different performance profiles to our risk managers and from their answers we estimated sensitivity to instability. The metric is outcome of this exercise.",
      "votes": null
    },
    {
      "id": "2641190",
      "postDate": "02/07/2024 10:35:53",
      "content": "<p>Exactly as said Tomas, the weights are not random and should simulate how we rate the model performance. The metric itself is a crude simplification of the actual rating of performance profiles. Nevertheless, it will suffice for the sake of this competition.  </p>",
      "rawMarkdown": "Exactly as said Tomas, the weights are not random and should simulate how we rate the model performance. The metric itself is a crude simplification of the actual rating of performance profiles. Nevertheless, it will suffice for the sake of this competition.",
      "votes": null
    },
    {
      "id": "2641254",
      "postDate": "02/07/2024 11:22:52",
      "content": "<p>Honestly, this is the first time I've seen this approach to scoring model metrics. Wouldn't it be better to focus on ROC-AUC (or Gini) and Kolmogorov-Smirnov statistics for primary validation and look at PSI (and optionally Herfindahl-Hirschman index) to monitor model drift (during secondary validation) and set up an automated pipeline for model retraining (using Airflow, Mlflow, DVC) when metrics fall below a certain threshold? </p>",
      "rawMarkdown": "Honestly, this is the first time I've seen this approach to scoring model metrics. Wouldn't it be better to focus on ROC-AUC (or Gini) and Kolmogorov-Smirnov statistics for primary validation and look at PSI (and optionally Herfindahl-Hirschman index) to monitor model drift (during secondary validation) and set up an automated pipeline for model retraining (using Airflow, Mlflow, DVC) when metrics fall below a certain threshold?",
      "votes": null
    },
    {
      "id": "2641272",
      "postDate": "02/07/2024 11:38:20",
      "content": "<p>Thank you for your reply. It is clear now.</p>",
      "rawMarkdown": "Thank you for your reply. It is clear now.",
      "votes": null
    },
    {
      "id": "2641276",
      "postDate": "02/07/2024 11:45:55",
      "content": "<p>Some models can not be automatically retrained or even if they could be the approval process takes time. The motivation here is simple. We would like to see methods that maximize performance yet can be stable in future at least for a while. </p>",
      "rawMarkdown": "Some models can not be automatically retrained or even if they could be the approval process takes time. The motivation here is simple. We would like to see methods that maximize performance yet can be stable in future at least for a while.",
      "votes": null
    },
    {
      "id": "2641290",
      "postDate": "02/07/2024 11:59:12",
      "content": "<p>Well, we wanted to present Kagglers some new challenge :-)</p>\n<p>PSI - that's something we definitely look at, however it measures changes in structure of customer population, not changes in relationship between score / feature value and credit risk. Moreover, there is significant delay in observability of targets, so automation/redevelopment is not always option - we would anyway use many months old data. That's why we want to incorporate stability into evaluation metric so our models are more durable.</p>",
      "rawMarkdown": "Well, we wanted to present Kagglers some new challenge :-)\n\nPSI - that's something we definitely look at, however it measures changes in structure of customer population, not changes in relationship between score / feature value and credit risk. Moreover, there is significant delay in observability of targets, so automation/redevelopment is not always option - we would anyway use many months old data. That's why we want to incorporate stability into evaluation metric so our models are more durable.",
      "votes": null
    },
    {
      "id": "2641311",
      "postDate": "02/07/2024 12:21:26",
      "content": "<p>I agree that each bank has its own approach to scoring models. Some require high interpretability and prefer logistic regression models with feature categorization (binning), while others prioritize the F1 metric over the algorithm used. It is beneficial to have a platform like Kaggle where we can exchange opinions and experiences.  </p>",
      "rawMarkdown": "I agree that each bank has its own approach to scoring models. Some require high interpretability and prefer logistic regression models with feature categorization (binning), while others prioritize the F1 metric over the algorithm used. It is beneficial to have a platform like Kaggle where we can exchange opinions and experiences.",
      "votes": null
    },
    {
      "id": "2641545",
      "postDate": "02/07/2024 14:53:01",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> </p>\n<p>for stability prize, it is stated<br>\n\"2) incorporate stability directly into your training model / loss function (not on the level of feature preparation or feature selection)\"</p>\n<p>and you mention here \"The point of the metric is to simulate how we rate the performance of production models …\"</p>\n<p>I would suggest if you can  add a third criteria for  stability prize:<br>\n3)  (optional) design other validation metric function that can test the stability of inference model ….</p>\n<p>usually, loss and metric are not quite the same thing, although they are related </p>",
      "rawMarkdown": "jetakow \n\nfor stability prize, it is stated\n\"2) incorporate stability directly into your training model / loss function (not on the level of feature preparation or feature selection)\"\n\nand you mention here \"The point of the metric is to simulate how we rate the performance of production models ...\"\n\nI would suggest if you can  add a third criteria for  stability prize:\n3)  (optional) design other validation metric function that can test the stability of inference model ....\n\nusually, loss and metric are not quite the same thing, although they are related",
      "votes": null
    },
    {
      "id": "2641567",
      "postDate": "02/07/2024 15:03:34",
      "content": "<p>Thanks for the suggestion. I don't think we will be changing the criteria for now. If you come up with such a validation metric and it will help you produce a more stable model (in respect to our metric) it will be definitely a plus. You don't have to use our rating metric at all despite I encourage you to do so. </p>\n<p>\"usually, loss and metric are not quite the same thing, although they are related\"<br>\nWe know. If you would be able to alter the loss function/training in such a way that it takes into account the stability it will increase your chances of winning this price. </p>",
      "rawMarkdown": "Thanks for the suggestion. I don't think we will be changing the criteria for now. If you come up with such a validation metric and it will help you produce a more stable model (in respect to our metric) it will be definitely a plus. You don't have to use our rating metric at all despite I encourage you to do so. \n\n\"usually, loss and metric are not quite the same thing, although they are related\"\nWe know. If you would be able to alter the loss function/training in such a way that it takes into account the stability it will increase your chances of winning this price.",
      "votes": null
    },
    {
      "id": "2643276",
      "postDate": "02/08/2024 18:25:15",
      "content": "<p>Thanks for the insightful analysis.  It does seem like stability penalty plays an excessively large role in the metric.</p>\n<p>Another possible flaw with the metric is that the Gini of the model may change for reasons that have nothing to do with the quality of the model.  The population that is fed through the model itself can change, which may make it less predictable.  For example, let's say that predictor A is extremely predictive, raising your Gini from 75 to 80.  However, over time, the applicants have less and less variation in that predictor, until every application has exactly the same value for it.  Since there is no more variance in the predictor, you go back from 80 to 75.</p>\n<p>Would you have been better off without that predictor?  I would argue emphatically that you were better of making use of that predictor when it was predictive.  It's a shame that the predictor lost its value for your use case over time, but you were making better decisions for a while, and your model is not any worse now that it would've been had you never included it.  If variance in predictor A picks up again, the model will continue using that predictor properly.  However, this metric would punish you for building a model that by every reasonable measure is strictly superior.</p>\n<p>In my opinion, too many people in the industry vastly overemphasize the value of model stability, often without even properly quantifying and putting in proper context its effect.  Being consistently bad at predicting is rarely better than being inconsistently good at it.</p>",
      "rawMarkdown": "Thanks for the insightful analysis.  It does seem like stability penalty plays an excessively large role in the metric.\n\nAnother possible flaw with the metric is that the Gini of the model may change for reasons that have nothing to do with the quality of the model.  The population that is fed through the model itself can change, which may make it less predictable.  For example, let's say that predictor A is extremely predictive, raising your Gini from 75 to 80.  However, over time, the applicants have less and less variation in that predictor, until every application has exactly the same value for it.  Since there is no more variance in the predictor, you go back from 80 to 75.\n\nWould you have been better off without that predictor?  I would argue emphatically that you were better of making use of that predictor when it was predictive.  It's a shame that the predictor lost its value for your use case over time, but you were making better decisions for a while, and your model is not any worse now that it would've been had you never included it.  If variance in predictor A picks up again, the model will continue using that predictor properly.  However, this metric would punish you for building a model that by every reasonable measure is strictly superior.\n\nIn my opinion, too many people in the industry vastly overemphasize the value of model stability, often without even properly quantifying and putting in proper context its effect.  Being consistently bad at predicting is rarely better than being inconsistently good at it.",
      "votes": null
    },
    {
      "id": "2643395",
      "postDate": "02/08/2024 19:52:32",
      "content": "<blockquote>\n  <p>It does seem like stability penalty plays an excessively large role in the metric.</p>\n</blockquote>\n<p>Indeed, that is intended.</p>\n<blockquote>\n  <p>Another possible flaw with the metric is that the Gini of the model may change for reasons that have nothing to do with the quality of the model. The population that is fed through the model itself can change, which may make it less predictable. For example, let's say that predictor A is extremely predictive, raising your Gini from 75 to 80. However, over time, the applicants have less and less variation in that predictor, until every application has exactly the same value for it. Since there is no more variance in the predictor, you go back from 80 to 75.</p>\n</blockquote>\n<p>We are well aware of that, that is also the reason for instability penalization. </p>\n<blockquote>\n  <p>In my opinion, too many people in the industry vastly overemphasize the value of model stability, often without even properly quantifying and putting in proper context its effect. Being consistently bad at predicting is rarely better than being inconsistently good at it.</p>\n</blockquote>\n<p>Model stability is an issue that troubles us at Home Credit. We have some means of dealing with this issue, yet we believe this part of modelling can be still improved. There were also other ideas for this year's competition, but stability was selected for this round. </p>",
      "rawMarkdown": ">It does seem like stability penalty plays an excessively large role in the metric.\n\nIndeed, that is intended.\n\n>Another possible flaw with the metric is that the Gini of the model may change for reasons that have nothing to do with the quality of the model. The population that is fed through the model itself can change, which may make it less predictable. For example, let's say that predictor A is extremely predictive, raising your Gini from 75 to 80. However, over time, the applicants have less and less variation in that predictor, until every application has exactly the same value for it. Since there is no more variance in the predictor, you go back from 80 to 75.\n\nWe are well aware of that, that is also the reason for instability penalization. \n\n>In my opinion, too many people in the industry vastly overemphasize the value of model stability, often without even properly quantifying and putting in proper context its effect. Being consistently bad at predicting is rarely better than being inconsistently good at it.\n\nModel stability is an issue that troubles us at Home Credit. We have some means of dealing with this issue, yet we believe this part of modelling can be still improved. There were also other ideas for this year's competition, but stability was selected for this round.",
      "votes": null
    },
    {
      "id": "2643416",
      "postDate": "02/08/2024 20:12:18",
      "content": "<blockquote>\n  <p>We are well aware of that, that is also the reason for instability penalization.</p>\n</blockquote>\n<p>Should that be a reason for instability penalization, or actually a reason against it?  As I argued, I think that should be a reason against it, at least against that way of measuring it.  </p>\n<p>I can see why you want to penalize the model for making use of relationships between predictors that don't stay stable over time.  I don't see why it useful in any business context to penalize the model for finding relationships that are stable over time, but whose effect on predictive power changes due to a change in population distribution.  Both effects look like a decline in Gini over time, but they have very different implications on business value gained from the model.</p>",
      "rawMarkdown": ">We are well aware of that, that is also the reason for instability penalization.\n\nShould that be a reason for instability penalization, or actually a reason against it?  As I argued, I think that should be a reason against it, at least against that way of measuring it.  \n\nI can see why you want to penalize the model for making use of relationships between predictors that don't stay stable over time.  I don't see why it useful in any business context to penalize the model for finding relationships that are stable over time, but whose effect on predictive power changes due to a change in population distribution.  Both effects look like a decline in Gini over time, but they have very different implications on business value gained from the model.",
      "votes": null
    },
    {
      "id": "2643568",
      "postDate": "02/08/2024 22:50:58",
      "content": "<blockquote>\n  <p>I don't see why it useful in any business context to penalize the model for finding relationships that are stable over time, but whose effect on predictive power changes due to a change in population distribution.</p>\n</blockquote>\n<p>The result of this competition won't be implemented directly in production, we seek new ways of implementing models that would be stable over time. In other words, whether the metric is perfectly aligned with the business perfectly or not is not relevant for this competition. </p>\n<p>The metric is for sure more strict in a way you described, but the subset of models (according to the more strict metric) we will see in this competition, will be stable over time. We considered also A) creating different metric that would rather track drift and different properties of the predictions or we could B) alter the data so that we have the change of underlying distributions under control. For A) we didn't want to overcomplicate the metric and we wanted to keep it simple and as robust as possible. For B) we didn't want to use synthetic data. </p>\n<p>I hope this explanation helps.</p>",
      "rawMarkdown": ">I don't see why it useful in any business context to penalize the model for finding relationships that are stable over time, but whose effect on predictive power changes due to a change in population distribution.\n\nThe result of this competition won't be implemented directly in production, we seek new ways of implementing models that would be stable over time. In other words, whether the metric is perfectly aligned with the business perfectly or not is not relevant for this competition. \n\nThe metric is for sure more strict in a way you described, but the subset of models (according to the more strict metric) we will see in this competition, will be stable over time. We considered also A) creating different metric that would rather track drift and different properties of the predictions or we could B) alter the data so that we have the change of underlying distributions under control. For A) we didn't want to overcomplicate the metric and we wanted to keep it simple and as robust as possible. For B) we didn't want to use synthetic data. \n\nI hope this explanation helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2641076,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/07/2024 09:13:47",
      "content": "<p>The point of the metric is to simulate how we rate the performance of production models. The high gini performance just after deployment is pointless if over a period of 6 months, the performance will drop to half of that. Instead of integrating the gini in time we decided to settle on this metric which takes into account not only gini, but also the decrease of performance and variance of performance. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2641090,
          "author_name": "w5833946",
          "author_url": "",
          "post_date": "02/07/2024 09:32:03",
          "content": "<p>So we shall consider model with constant gini score 0.5 in 6 months better than model with gini score decreased from 1 to 0.5 within the same period?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2641121,
              "author_name": "tomasjeline2",
              "author_url": "",
              "post_date": "02/07/2024 09:50:23",
              "content": "<p>Hi Ruby,<br>\nYes. Because if this trend continues, we can expect that in another 6 months gini will drop further to 0. Such model with so unstable performance is not suitable for usage in production. That's why we penalize instability in our metric.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2641136,
                  "author_name": "w5833946",
                  "author_url": "",
                  "post_date": "02/07/2024 10:03:19",
                  "content": "<p>As explained in my post the similar conclusion extends to period of 88*2 weeks (3.4 years), I guess this far beyond your model lifecycle? I think the metric design is reasonable, but 88 might be too large? Is there any reason to choose 88 as weight?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2641153,
                      "author_name": "tomasjeline2",
                      "author_url": "",
                      "post_date": "02/07/2024 10:11:49",
                      "content": "<p>We have presented models with different performance profiles to our risk managers and from their answers we estimated sensitivity to instability. The metric is outcome of this exercise.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 2641190,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "02/07/2024 10:35:53",
              "content": "<p>Exactly as said Tomas, the weights are not random and should simulate how we rate the model performance. The metric itself is a crude simplification of the actual rating of performance profiles. Nevertheless, it will suffice for the sake of this competition.  </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2641272,
                  "author_name": "w5833946",
                  "author_url": "",
                  "post_date": "02/07/2024 11:38:20",
                  "content": "<p>Thank you for your reply. It is clear now.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2641254,
          "author_name": "bratkovskyevgeny",
          "author_url": "",
          "post_date": "02/07/2024 11:22:52",
          "content": "<p>Honestly, this is the first time I've seen this approach to scoring model metrics. Wouldn't it be better to focus on ROC-AUC (or Gini) and Kolmogorov-Smirnov statistics for primary validation and look at PSI (and optionally Herfindahl-Hirschman index) to monitor model drift (during secondary validation) and set up an automated pipeline for model retraining (using Airflow, Mlflow, DVC) when metrics fall below a certain threshold? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2641276,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "02/07/2024 11:45:55",
              "content": "<p>Some models can not be automatically retrained or even if they could be the approval process takes time. The motivation here is simple. We would like to see methods that maximize performance yet can be stable in future at least for a while. </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2641290,
              "author_name": "tomasjeline2",
              "author_url": "",
              "post_date": "02/07/2024 11:59:12",
              "content": "<p>Well, we wanted to present Kagglers some new challenge :-)</p>\n<p>PSI - that's something we definitely look at, however it measures changes in structure of customer population, not changes in relationship between score / feature value and credit risk. Moreover, there is significant delay in observability of targets, so automation/redevelopment is not always option - we would anyway use many months old data. That's why we want to incorporate stability into evaluation metric so our models are more durable.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2641311,
                  "author_name": "bratkovskyevgeny",
                  "author_url": "",
                  "post_date": "02/07/2024 12:21:26",
                  "content": "<p>I agree that each bank has its own approach to scoring models. Some require high interpretability and prefer logistic regression models with feature categorization (binning), while others prioritize the F1 metric over the algorithm used. It is beneficial to have a platform like Kaggle where we can exchange opinions and experiences.  </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2641545,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/07/2024 14:53:01",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> </p>\n<p>for stability prize, it is stated<br>\n\"2) incorporate stability directly into your training model / loss function (not on the level of feature preparation or feature selection)\"</p>\n<p>and you mention here \"The point of the metric is to simulate how we rate the performance of production models …\"</p>\n<p>I would suggest if you can  add a third criteria for  stability prize:<br>\n3)  (optional) design other validation metric function that can test the stability of inference model ….</p>\n<p>usually, loss and metric are not quite the same thing, although they are related </p>",
          "votes": null,
          "replies": [
            {
              "id": 2641567,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "02/07/2024 15:03:34",
              "content": "<p>Thanks for the suggestion. I don't think we will be changing the criteria for now. If you come up with such a validation metric and it will help you produce a more stable model (in respect to our metric) it will be definitely a plus. You don't have to use our rating metric at all despite I encourage you to do so. </p>\n<p>\"usually, loss and metric are not quite the same thing, although they are related\"<br>\nWe know. If you would be able to alter the loss function/training in such a way that it takes into account the stability it will increase your chances of winning this price. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2643276,
      "author_name": "dmitriyguller",
      "author_url": "",
      "post_date": "02/08/2024 18:25:15",
      "content": "<p>Thanks for the insightful analysis.  It does seem like stability penalty plays an excessively large role in the metric.</p>\n<p>Another possible flaw with the metric is that the Gini of the model may change for reasons that have nothing to do with the quality of the model.  The population that is fed through the model itself can change, which may make it less predictable.  For example, let's say that predictor A is extremely predictive, raising your Gini from 75 to 80.  However, over time, the applicants have less and less variation in that predictor, until every application has exactly the same value for it.  Since there is no more variance in the predictor, you go back from 80 to 75.</p>\n<p>Would you have been better off without that predictor?  I would argue emphatically that you were better of making use of that predictor when it was predictive.  It's a shame that the predictor lost its value for your use case over time, but you were making better decisions for a while, and your model is not any worse now that it would've been had you never included it.  If variance in predictor A picks up again, the model will continue using that predictor properly.  However, this metric would punish you for building a model that by every reasonable measure is strictly superior.</p>\n<p>In my opinion, too many people in the industry vastly overemphasize the value of model stability, often without even properly quantifying and putting in proper context its effect.  Being consistently bad at predicting is rarely better than being inconsistently good at it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2643395,
          "author_name": "jetakow",
          "author_url": "",
          "post_date": "02/08/2024 19:52:32",
          "content": "<blockquote>\n  <p>It does seem like stability penalty plays an excessively large role in the metric.</p>\n</blockquote>\n<p>Indeed, that is intended.</p>\n<blockquote>\n  <p>Another possible flaw with the metric is that the Gini of the model may change for reasons that have nothing to do with the quality of the model. The population that is fed through the model itself can change, which may make it less predictable. For example, let's say that predictor A is extremely predictive, raising your Gini from 75 to 80. However, over time, the applicants have less and less variation in that predictor, until every application has exactly the same value for it. Since there is no more variance in the predictor, you go back from 80 to 75.</p>\n</blockquote>\n<p>We are well aware of that, that is also the reason for instability penalization. </p>\n<blockquote>\n  <p>In my opinion, too many people in the industry vastly overemphasize the value of model stability, often without even properly quantifying and putting in proper context its effect. Being consistently bad at predicting is rarely better than being inconsistently good at it.</p>\n</blockquote>\n<p>Model stability is an issue that troubles us at Home Credit. We have some means of dealing with this issue, yet we believe this part of modelling can be still improved. There were also other ideas for this year's competition, but stability was selected for this round. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2643416,
              "author_name": "dmitriyguller",
              "author_url": "",
              "post_date": "02/08/2024 20:12:18",
              "content": "<blockquote>\n  <p>We are well aware of that, that is also the reason for instability penalization.</p>\n</blockquote>\n<p>Should that be a reason for instability penalization, or actually a reason against it?  As I argued, I think that should be a reason against it, at least against that way of measuring it.  </p>\n<p>I can see why you want to penalize the model for making use of relationships between predictors that don't stay stable over time.  I don't see why it useful in any business context to penalize the model for finding relationships that are stable over time, but whose effect on predictive power changes due to a change in population distribution.  Both effects look like a decline in Gini over time, but they have very different implications on business value gained from the model.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2643568,
                  "author_name": "jetakow",
                  "author_url": "",
                  "post_date": "02/08/2024 22:50:58",
                  "content": "<blockquote>\n  <p>I don't see why it useful in any business context to penalize the model for finding relationships that are stable over time, but whose effect on predictive power changes due to a change in population distribution.</p>\n</blockquote>\n<p>The result of this competition won't be implemented directly in production, we seek new ways of implementing models that would be stable over time. In other words, whether the metric is perfectly aligned with the business perfectly or not is not relevant for this competition. </p>\n<p>The metric is for sure more strict in a way you described, but the subset of models (according to the more strict metric) we will see in this competition, will be stable over time. We considered also A) creating different metric that would rather track drift and different properties of the predictions or we could B) alter the data so that we have the change of underlying distributions under control. For A) we didn't want to overcomplicate the metric and we wanted to keep it simple and as robust as possible. For B) we didn't want to use synthetic data. </p>\n<p>I hope this explanation helps.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2640903": "Consider gini score decrease linearly from score_max to score_min in k weeks, then three terms in metric will be:\n(score_max+score_min)/2, -(score_max-score_min)*88/k, 0\nIn case score_max==score_min they sum to score_min\nif we subtract it from the general one we will get:\n(score_max-score_min)(0.5-88/k)\nif 88/k > 0.5, then metric is better for the second case where all gini score is smaller .\nAnd we have 92 weeks in training set (I think it will be similar in test set?) so it lies in this situation.\nIs this an expected behaviour? Personally I expect the metric will be better with better gini score for every week and it may avoid some meaningless postprocessing. And metric property changes with evaluation period length is also a bit counter intuitive.",
    "2641076": "The point of the metric is to simulate how we rate the performance of production models. The high gini performance just after deployment is pointless if over a period of 6 months, the performance will drop to half of that. Instead of integrating the gini in time we decided to settle on this metric which takes into account not only gini, but also the decrease of performance and variance of performance.",
    "2641090": "So we shall consider model with constant gini score 0.5 in 6 months better than model with gini score decreased from 1 to 0.5 within the same period?",
    "2641121": "Hi Ruby,\nYes. Because if this trend continues, we can expect that in another 6 months gini will drop further to 0. Such model with so unstable performance is not suitable for usage in production. That's why we penalize instability in our metric.",
    "2641136": "As explained in my post the similar conclusion extends to period of 88*2 weeks (3.4 years), I guess this far beyond your model lifecycle? I think the metric design is reasonable, but 88 might be too large? Is there any reason to choose 88 as weight?",
    "2641153": "We have presented models with different performance profiles to our risk managers and from their answers we estimated sensitivity to instability. The metric is outcome of this exercise.",
    "2641190": "Exactly as said Tomas, the weights are not random and should simulate how we rate the model performance. The metric itself is a crude simplification of the actual rating of performance profiles. Nevertheless, it will suffice for the sake of this competition.",
    "2641254": "Honestly, this is the first time I've seen this approach to scoring model metrics. Wouldn't it be better to focus on ROC-AUC (or Gini) and Kolmogorov-Smirnov statistics for primary validation and look at PSI (and optionally Herfindahl-Hirschman index) to monitor model drift (during secondary validation) and set up an automated pipeline for model retraining (using Airflow, Mlflow, DVC) when metrics fall below a certain threshold?",
    "2641272": "Thank you for your reply. It is clear now.",
    "2641276": "Some models can not be automatically retrained or even if they could be the approval process takes time. The motivation here is simple. We would like to see methods that maximize performance yet can be stable in future at least for a while.",
    "2641290": "Well, we wanted to present Kagglers some new challenge :-)\n\nPSI - that's something we definitely look at, however it measures changes in structure of customer population, not changes in relationship between score / feature value and credit risk. Moreover, there is significant delay in observability of targets, so automation/redevelopment is not always option - we would anyway use many months old data. That's why we want to incorporate stability into evaluation metric so our models are more durable.",
    "2641311": "I agree that each bank has its own approach to scoring models. Some require high interpretability and prefer logistic regression models with feature categorization (binning), while others prioritize the F1 metric over the algorithm used. It is beneficial to have a platform like Kaggle where we can exchange opinions and experiences.",
    "2641545": "jetakow \n\nfor stability prize, it is stated\n\"2) incorporate stability directly into your training model / loss function (not on the level of feature preparation or feature selection)\"\n\nand you mention here \"The point of the metric is to simulate how we rate the performance of production models ...\"\n\nI would suggest if you can  add a third criteria for  stability prize:\n3)  (optional) design other validation metric function that can test the stability of inference model ....\n\nusually, loss and metric are not quite the same thing, although they are related",
    "2641567": "Thanks for the suggestion. I don't think we will be changing the criteria for now. If you come up with such a validation metric and it will help you produce a more stable model (in respect to our metric) it will be definitely a plus. You don't have to use our rating metric at all despite I encourage you to do so. \n\n\"usually, loss and metric are not quite the same thing, although they are related\"\nWe know. If you would be able to alter the loss function/training in such a way that it takes into account the stability it will increase your chances of winning this price.",
    "2643276": "Thanks for the insightful analysis.  It does seem like stability penalty plays an excessively large role in the metric.\n\nAnother possible flaw with the metric is that the Gini of the model may change for reasons that have nothing to do with the quality of the model.  The population that is fed through the model itself can change, which may make it less predictable.  For example, let's say that predictor A is extremely predictive, raising your Gini from 75 to 80.  However, over time, the applicants have less and less variation in that predictor, until every application has exactly the same value for it.  Since there is no more variance in the predictor, you go back from 80 to 75.\n\nWould you have been better off without that predictor?  I would argue emphatically that you were better of making use of that predictor when it was predictive.  It's a shame that the predictor lost its value for your use case over time, but you were making better decisions for a while, and your model is not any worse now that it would've been had you never included it.  If variance in predictor A picks up again, the model will continue using that predictor properly.  However, this metric would punish you for building a model that by every reasonable measure is strictly superior.\n\nIn my opinion, too many people in the industry vastly overemphasize the value of model stability, often without even properly quantifying and putting in proper context its effect.  Being consistently bad at predicting is rarely better than being inconsistently good at it.",
    "2643395": ">It does seem like stability penalty plays an excessively large role in the metric.\n\nIndeed, that is intended.\n\n>Another possible flaw with the metric is that the Gini of the model may change for reasons that have nothing to do with the quality of the model. The population that is fed through the model itself can change, which may make it less predictable. For example, let's say that predictor A is extremely predictive, raising your Gini from 75 to 80. However, over time, the applicants have less and less variation in that predictor, until every application has exactly the same value for it. Since there is no more variance in the predictor, you go back from 80 to 75.\n\nWe are well aware of that, that is also the reason for instability penalization. \n\n>In my opinion, too many people in the industry vastly overemphasize the value of model stability, often without even properly quantifying and putting in proper context its effect. Being consistently bad at predicting is rarely better than being inconsistently good at it.\n\nModel stability is an issue that troubles us at Home Credit. We have some means of dealing with this issue, yet we believe this part of modelling can be still improved. There were also other ideas for this year's competition, but stability was selected for this round.",
    "2643416": ">We are well aware of that, that is also the reason for instability penalization.\n\nShould that be a reason for instability penalization, or actually a reason against it?  As I argued, I think that should be a reason against it, at least against that way of measuring it.  \n\nI can see why you want to penalize the model for making use of relationships between predictors that don't stay stable over time.  I don't see why it useful in any business context to penalize the model for finding relationships that are stable over time, but whose effect on predictive power changes due to a change in population distribution.  Both effects look like a decline in Gini over time, but they have very different implications on business value gained from the model.",
    "2643568": ">I don't see why it useful in any business context to penalize the model for finding relationships that are stable over time, but whose effect on predictive power changes due to a change in population distribution.\n\nThe result of this competition won't be implemented directly in production, we seek new ways of implementing models that would be stable over time. In other words, whether the metric is perfectly aligned with the business perfectly or not is not relevant for this competition. \n\nThe metric is for sure more strict in a way you described, but the subset of models (according to the more strict metric) we will see in this competition, will be stable over time. We considered also A) creating different metric that would rather track drift and different properties of the predictions or we could B) alter the data so that we have the change of underlying distributions under control. For A) we didn't want to overcomplicate the metric and we wanted to keep it simple and as robust as possible. For B) we didn't want to use synthetic data. \n\nI hope this explanation helps."
  },
  "source": "meta"
}