{
  "id": 505852,
  "title": "The big(er) elephant: is stability regime dependent?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/505852",
  "author_name": "",
  "post_date": "2024-05-19T10:56:13.200488400Z",
  "votes": 18,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This competition is not over yet, but it is true it doesn't look too good. In my opinion, worse than the metric choice is the repeated change of mind regarding the evaluation.</p>\n<p>I stopped Kaggling over two years ago, lacking the confidence that, conditioned on the host, my effort of many weeks wouldn't be dealt with by an RNG for leaderboard. On the sidelines since, Home Credit Stability looked like an interesting premise, at first sight, maybe time to come back to the arena… and here we are.</p>\n<p>Still, I see more than one \"elephant in the room\" here. We are supposed to achieve \"stability\" in contrast to \"model degradation\" over time. The question is, <strong>why does a model degrade \"over time\" in general, and why does a model degrade in this domain and task in particular?</strong></p>\n<p>In a nutshell, in general, I would give three main reasons for a model to degrade over time:</p>\n<p><strong>1. Concept Drift:</strong> the answers to the same questions are now different.</p>\n<p><strong>2. Covariate Shift:</strong> the answers are the same, but the questions are worded differently and harder to relate to previous questions.</p>\n<p><strong>3. User adaptation:</strong> somehow you influenced the test you are seeing, i.e. you changed the future with your previous answers and that future sample is now different, so possibly both points 1 and 2.</p>\n<p><strong>So, regarding this competition, which one applies?</strong></p>\n<p>For sure in the \"real world\" point 3 applies as credit companies will use their models to select which credits to give, hence influencing the future sample their models will be seeing. This is a well-known problem of causal/ interactive/ reinforcement learning when you change your future with your model. In this particular task, models can be thought of as learning to make inferences on a sample of approved credits. <strong>Is this relevant for this competition? I would bet possibly not because we know to be working with heavily ad-hoc subsampled data.</strong> One less thing to worry about for us, but possibly something that skips one of the interesting (and solvable) sources of instability.</p>\n<p><strong>But there is the problem number 2, the big one. The question would go like this: are the statistical properties of this task time/regime dependent in a non-stationary way?</strong> Or, in other words, how predictable is model degradation in this use case? Because it looks to me like a financial-regime dependent, non-stationary problem. This means that the infamous slope of the weekly gini regression can be up or down, depending on conditions that most companies struggle to forecast at all. It also would need much wider data time span for any attempt to model. <strong>So here goes this critic to the metric, not from its hackability but because of it pursuing stability by measuring something that is possibly regime-dependent, non stationary, and pointless to attempt to model at this time scale.</strong></p>\n<p>I hope that all this made some sense. Anyway, I do find this competition interesting in spite of its flaws. I also understand that the hosts can still obtain insights from the solutions to the problem in its current form,  no matter how much sense it makes from a Kaggler point of view.  Quite curious to witness how it will unfold. </p>\n<p>Good luck to all!</p>",
  "messages": [
    {
      "id": "2823667",
      "postDate": "05/19/2024 10:56:13",
      "content": "<p>This competition is not over yet, but it is true it doesn't look too good. In my opinion, worse than the metric choice is the repeated change of mind regarding the evaluation.</p>\n<p>I stopped Kaggling over two years ago, lacking the confidence that, conditioned on the host, my effort of many weeks wouldn't be dealt with by an RNG for leaderboard. On the sidelines since, Home Credit Stability looked like an interesting premise, at first sight, maybe time to come back to the arena… and here we are.</p>\n<p>Still, I see more than one \"elephant in the room\" here. We are supposed to achieve \"stability\" in contrast to \"model degradation\" over time. The question is, <strong>why does a model degrade \"over time\" in general, and why does a model degrade in this domain and task in particular?</strong></p>\n<p>In a nutshell, in general, I would give three main reasons for a model to degrade over time:</p>\n<p><strong>1. Concept Drift:</strong> the answers to the same questions are now different.</p>\n<p><strong>2. Covariate Shift:</strong> the answers are the same, but the questions are worded differently and harder to relate to previous questions.</p>\n<p><strong>3. User adaptation:</strong> somehow you influenced the test you are seeing, i.e. you changed the future with your previous answers and that future sample is now different, so possibly both points 1 and 2.</p>\n<p><strong>So, regarding this competition, which one applies?</strong></p>\n<p>For sure in the \"real world\" point 3 applies as credit companies will use their models to select which credits to give, hence influencing the future sample their models will be seeing. This is a well-known problem of causal/ interactive/ reinforcement learning when you change your future with your model. In this particular task, models can be thought of as learning to make inferences on a sample of approved credits. <strong>Is this relevant for this competition? I would bet possibly not because we know to be working with heavily ad-hoc subsampled data.</strong> One less thing to worry about for us, but possibly something that skips one of the interesting (and solvable) sources of instability.</p>\n<p><strong>But there is the problem number 2, the big one. The question would go like this: are the statistical properties of this task time/regime dependent in a non-stationary way?</strong> Or, in other words, how predictable is model degradation in this use case? Because it looks to me like a financial-regime dependent, non-stationary problem. This means that the infamous slope of the weekly gini regression can be up or down, depending on conditions that most companies struggle to forecast at all. It also would need much wider data time span for any attempt to model. <strong>So here goes this critic to the metric, not from its hackability but because of it pursuing stability by measuring something that is possibly regime-dependent, non stationary, and pointless to attempt to model at this time scale.</strong></p>\n<p>I hope that all this made some sense. Anyway, I do find this competition interesting in spite of its flaws. I also understand that the hosts can still obtain insights from the solutions to the problem in its current form,  no matter how much sense it makes from a Kaggler point of view.  Quite curious to witness how it will unfold. </p>\n<p>Good luck to all!</p>",
      "rawMarkdown": "This competition is not over yet, but it is true it doesn't look too good. In my opinion, worse than the metric choice is the repeated change of mind regarding the evaluation.\n\nI stopped Kaggling over two years ago, lacking the confidence that, conditioned on the host, my effort of many weeks wouldn't be dealt with by an RNG for leaderboard. On the sidelines since, Home Credit Stability looked like an interesting premise, at first sight, maybe time to come back to the arena... and here we are.\n\nStill, I see more than one \"elephant in the room\" here. We are supposed to achieve \"stability\" in contrast to \"model degradation\" over time. The question is, **why does a model degrade \"over time\" in general, and why does a model degrade in this domain and task in particular?**\n\nIn a nutshell, in general, I would give three main reasons for a model to degrade over time:\n\n**1. Concept Drift:** the answers to the same questions are now different.\n\n**2. Covariate Shift:** the answers are the same, but the questions are worded differently and harder to relate to previous questions.\n\n**3. User adaptation:** somehow you influenced the test you are seeing, i.e. you changed the future with your previous answers and that future sample is now different, so possibly both points 1 and 2.\n\n**So, regarding this competition, which one applies?**\n\nFor sure in the \"real world\" point 3 applies as credit companies will use their models to select which credits to give, hence influencing the future sample their models will be seeing. This is a well-known problem of causal/ interactive/ reinforcement learning when you change your future with your model. In this particular task, models can be thought of as learning to make inferences on a sample of approved credits. **Is this relevant for this competition? I would bet possibly not because we know to be working with heavily ad-hoc subsampled data.** One less thing to worry about for us, but possibly something that skips one of the interesting (and solvable) sources of instability.\n\n**But there is the problem number 2, the big one. The question would go like this: are the statistical properties of this task time/regime dependent in a non-stationary way?** Or, in other words, how predictable is model degradation in this use case? Because it looks to me like a financial-regime dependent, non-stationary problem. This means that the infamous slope of the weekly gini regression can be up or down, depending on conditions that most companies struggle to forecast at all. It also would need much wider data time span for any attempt to model. **So here goes this critic to the metric, not from its hackability but because of it pursuing stability by measuring something that is possibly regime-dependent, non stationary, and pointless to attempt to model at this time scale.**\n\nI hope that all this made some sense. Anyway, I do find this competition interesting in spite of its flaws. I also understand that the hosts can still obtain insights from the solutions to the problem in its current form,  no matter how much sense it makes from a Kaggler point of view.  Quite curious to witness how it will unfold. \n\nGood luck to all!",
      "votes": null
    },
    {
      "id": "2823692",
      "postDate": "05/19/2024 11:21:24",
      "content": "<blockquote>\n  <p>Because it looks to me like a financial-regime dependent, non-stationary problem</p>\n</blockquote>\n<p>I agree, based on the training data we have observed.</p>\n<p>I think the current metric could be modified to better capture predictable changes in performance while reducing the probability of false alarms on stochastic trends. I've discussed this here <a href=\"https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach\" target=\"_blank\">https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach</a>. </p>\n<p>However, I believe there are no predictable changes in performance, so including a volatility penalty should be sufficient. </p>\n<p>Perhaps formulating the problem in terms of the financial losses of the company would be a better approach to capture the costs of dropping performance.</p>",
      "rawMarkdown": ">Because it looks to me like a financial-regime dependent, non-stationary problem\n\nI agree, based on the training data we have observed.\n\nI think the current metric could be modified to better capture predictable changes in performance while reducing the probability of false alarms on stochastic trends. I've discussed this here https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach. \n\nHowever, I believe there are no predictable changes in performance, so including a volatility penalty should be sufficient. \n\nPerhaps formulating the problem in terms of the financial losses of the company would be a better approach to capture the costs of dropping performance.",
      "votes": null
    },
    {
      "id": "2823729",
      "postDate": "05/19/2024 11:53:35",
      "content": "<p>That's interesting. Regarding the financial P&amp;L of a company, I'm afraid it would imply adding many more layers of complexity over the uncertainty problem. I'm not an expert in this domain but if today I had to draw a tentative evaluation of the stability of a model for this context, I would attempt to normalize performance for each time point's by each current global credit market-performance.</p>",
      "rawMarkdown": "That's interesting. Regarding the financial P&L of a company, I'm afraid it would imply adding many more layers of complexity over the uncertainty problem. I'm not an expert in this domain but if today I had to draw a tentative evaluation of the stability of a model for this context, I would attempt to normalize performance for each time point's by each current global credit market-performance.",
      "votes": null
    },
    {
      "id": "2823905",
      "postDate": "05/19/2024 13:22:03",
      "content": "<p>Interested in your comment regarding user adaptation which I take it to include issues like reject inference: <em>one of the interesting (and solvable) sources of instability</em><br>\nWhat solutions did you have in mind?</p>",
      "rawMarkdown": "Interested in your comment regarding user adaptation which I take it to include issues like reject inference: *one of the interesting (and solvable) sources of instability*\nWhat solutions did you have in mind?",
      "votes": null
    },
    {
      "id": "2823910",
      "postDate": "05/19/2024 13:25:21",
      "content": "<p>I've seen a few papers on this, e.g. <a href=\"https://www.sciencedirect.com/science/article/pii/S0305048323001688#:~:text=The%20total%20profit%20of%20a,the%20MP%20measure%20%5B5%5D\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S0305048323001688#:~:text=The%20total%20profit%20of%20a,the%20MP%20measure%20%5B5%5D</a>. </p>\n<blockquote>\n  <p>Models that were profitable during economic recovery might bring losses during turbulent times because of changes in the benefits associated with the correct classification and losses from misclassification. Hence, there has been a recent surge in demand for estimating AI systems’ profit in credit scoring under changing market conditions, especially during the COVID-19 pandemic</p>\n</blockquote>\n<p>.</p>\n<blockquote>\n  <p>I would attempt to normalize performance for each time point's by each current global credit market-performance.</p>\n</blockquote>\n<p>That's an interesting idea, though it may be difficult to find an appropriate benchmark</p>",
      "rawMarkdown": "I've seen a few papers on this, e.g. https://www.sciencedirect.com/science/article/pii/S0305048323001688#:~:text=The%20total%20profit%20of%20a,the%20MP%20measure%20%5B5%5D. \n>Models that were profitable during economic recovery might bring losses during turbulent times because of changes in the benefits associated with the correct classification and losses from misclassification. Hence, there has been a recent surge in demand for estimating AI systems’ profit in credit scoring under changing market conditions, especially during the COVID-19 pandemic\n\n.\n\n>I would attempt to normalize performance for each time point's by each current global credit market-performance.\n\nThat's an interesting idea, though it may be difficult to find an appropriate benchmark",
      "votes": null
    },
    {
      "id": "2823989",
      "postDate": "05/19/2024 14:14:36",
      "content": "<p>I agree. This is definitely regime dependent.<br>\nThere are striking similarities between this problem and the quant strategies evaluation in Finance. The expected returns are the AUC which we dont want them to decrease over time or have huge downtrends, and the volatility/dowdrown are what we're trying to eliminate here. An equivalent of the sharpe ratio (mean divided by std) is probably best suited to evaluate this, as it is used in quant firms. See: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501708\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501708</a></p>",
      "rawMarkdown": "I agree. This is definitely regime dependent.\nThere are striking similarities between this problem and the quant strategies evaluation in Finance. The expected returns are the AUC which we dont want them to decrease over time or have huge downtrends, and the volatility/dowdrown are what we're trying to eliminate here. An equivalent of the sharpe ratio (mean divided by std) is probably best suited to evaluate this, as it is used in quant firms. See: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501708",
      "votes": null
    },
    {
      "id": "2824262",
      "postDate": "05/19/2024 17:00:28",
      "content": "<p>Just had a first look at the paper. As I get it, the approach does add some complexity and some assumptions while optimizing for a different (more interesting) risk dependent objective.</p>\n<p>In this competition setup a multi-stage use of the models is implied, being our task essentially a ranking by week of candidates. Cutoff optimization and portfolio risk tunning probably coming afterwards through non automated human decision. The paper expands the scoring scope to profit and risk hence feature selection will optimize for that. Not sure that is equivalent to reduce the uncertainty of the ranking, more I would say the uncertainty is circumvented while optimizing for profit, quite similarly to how money allocation accounts for a given trading system uncertainty.</p>\n<p>Anyway the paper is an interesting finding, also makes more clear the meaning of your point about using P&amp;L for optimization that I had not fully understood in your first mention. Now it makes sense, it might be definitely a way to explore.</p>",
      "rawMarkdown": "Just had a first look at the paper. As I get it, the approach does add some complexity and some assumptions while optimizing for a different (more interesting) risk dependent objective.\n\nIn this competition setup a multi-stage use of the models is implied, being our task essentially a ranking by week of candidates. Cutoff optimization and portfolio risk tunning probably coming afterwards through non automated human decision. The paper expands the scoring scope to profit and risk hence feature selection will optimize for that. Not sure that is equivalent to reduce the uncertainty of the ranking, more I would say the uncertainty is circumvented while optimizing for profit, quite similarly to how money allocation accounts for a given trading system uncertainty.\n\nAnyway the paper is an interesting finding, also makes more clear the meaning of your point about using P&L for optimization that I had not fully understood in your first mention. Now it makes sense, it might be definitely a way to explore.",
      "votes": null
    },
    {
      "id": "2824268",
      "postDate": "05/19/2024 17:08:53",
      "content": "<p>Yes, reject inference definitely belongs to the core of user adaptation problem.  </p>\n<p>About being solvable, of course I meant partially solvable. Actual solutions are both problem and model dependent.<br>\nI simply expressed my impression that it is not a relevant source of degradation in this competition while alternative competition designs might have included it as something to mitigate/account for.</p>",
      "rawMarkdown": "Yes, reject inference definitely belongs to the core of user adaptation problem.  \n\nAbout being solvable, of course I meant partially solvable. Actual solutions are both problem and model dependent.\nI simply expressed my impression that it is not a relevant source of degradation in this competition while alternative competition designs might have included it as something to mitigate/account for.",
      "votes": null
    },
    {
      "id": "2824455",
      "postDate": "05/19/2024 19:16:24",
      "content": "<p>For credit ranking problems I have not seen good methods. Some methods in the literature or common practice are intuitive but with no objective improvement metric (unless one randomizes acceptance criteria and accepts the cost for a small random sample, not realistic).</p>",
      "rawMarkdown": "For credit ranking problems I have not seen good methods. Some methods in the literature or common practice are intuitive but with no objective improvement metric (unless one randomizes acceptance criteria and accepts the cost for a small random sample, not realistic).",
      "votes": null
    },
    {
      "id": "2825956",
      "postDate": "05/20/2024 16:41:00",
      "content": "<p>I agree. This is definitely regime dependent!</p>",
      "rawMarkdown": "I agree. This is definitely regime dependent!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2823692,
      "author_name": "eivolkova",
      "author_url": "",
      "post_date": "05/19/2024 11:21:24",
      "content": "<blockquote>\n  <p>Because it looks to me like a financial-regime dependent, non-stationary problem</p>\n</blockquote>\n<p>I agree, based on the training data we have observed.</p>\n<p>I think the current metric could be modified to better capture predictable changes in performance while reducing the probability of false alarms on stochastic trends. I've discussed this here <a href=\"https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach\" target=\"_blank\">https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach</a>. </p>\n<p>However, I believe there are no predictable changes in performance, so including a volatility penalty should be sufficient. </p>\n<p>Perhaps formulating the problem in terms of the financial losses of the company would be a better approach to capture the costs of dropping performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2823729,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "05/19/2024 11:53:35",
          "content": "<p>That's interesting. Regarding the financial P&amp;L of a company, I'm afraid it would imply adding many more layers of complexity over the uncertainty problem. I'm not an expert in this domain but if today I had to draw a tentative evaluation of the stability of a model for this context, I would attempt to normalize performance for each time point's by each current global credit market-performance.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2823910,
              "author_name": "eivolkova",
              "author_url": "",
              "post_date": "05/19/2024 13:25:21",
              "content": "<p>I've seen a few papers on this, e.g. <a href=\"https://www.sciencedirect.com/science/article/pii/S0305048323001688#:~:text=The%20total%20profit%20of%20a,the%20MP%20measure%20%5B5%5D\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S0305048323001688#:~:text=The%20total%20profit%20of%20a,the%20MP%20measure%20%5B5%5D</a>. </p>\n<blockquote>\n  <p>Models that were profitable during economic recovery might bring losses during turbulent times because of changes in the benefits associated with the correct classification and losses from misclassification. Hence, there has been a recent surge in demand for estimating AI systems’ profit in credit scoring under changing market conditions, especially during the COVID-19 pandemic</p>\n</blockquote>\n<p>.</p>\n<blockquote>\n  <p>I would attempt to normalize performance for each time point's by each current global credit market-performance.</p>\n</blockquote>\n<p>That's an interesting idea, though it may be difficult to find an appropriate benchmark</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2824262,
                  "author_name": "miguelpm",
                  "author_url": "",
                  "post_date": "05/19/2024 17:00:28",
                  "content": "<p>Just had a first look at the paper. As I get it, the approach does add some complexity and some assumptions while optimizing for a different (more interesting) risk dependent objective.</p>\n<p>In this competition setup a multi-stage use of the models is implied, being our task essentially a ranking by week of candidates. Cutoff optimization and portfolio risk tunning probably coming afterwards through non automated human decision. The paper expands the scoring scope to profit and risk hence feature selection will optimize for that. Not sure that is equivalent to reduce the uncertainty of the ranking, more I would say the uncertainty is circumvented while optimizing for profit, quite similarly to how money allocation accounts for a given trading system uncertainty.</p>\n<p>Anyway the paper is an interesting finding, also makes more clear the meaning of your point about using P&amp;L for optimization that I had not fully understood in your first mention. Now it makes sense, it might be definitely a way to explore.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2823905,
      "author_name": "casras",
      "author_url": "",
      "post_date": "05/19/2024 13:22:03",
      "content": "<p>Interested in your comment regarding user adaptation which I take it to include issues like reject inference: <em>one of the interesting (and solvable) sources of instability</em><br>\nWhat solutions did you have in mind?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2824268,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "05/19/2024 17:08:53",
          "content": "<p>Yes, reject inference definitely belongs to the core of user adaptation problem.  </p>\n<p>About being solvable, of course I meant partially solvable. Actual solutions are both problem and model dependent.<br>\nI simply expressed my impression that it is not a relevant source of degradation in this competition while alternative competition designs might have included it as something to mitigate/account for.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2824455,
              "author_name": "casras",
              "author_url": "",
              "post_date": "05/19/2024 19:16:24",
              "content": "<p>For credit ranking problems I have not seen good methods. Some methods in the literature or common practice are intuitive but with no objective improvement metric (unless one randomizes acceptance criteria and accepts the cost for a small random sample, not realistic).</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2823989,
      "author_name": "simoelm",
      "author_url": "",
      "post_date": "05/19/2024 14:14:36",
      "content": "<p>I agree. This is definitely regime dependent.<br>\nThere are striking similarities between this problem and the quant strategies evaluation in Finance. The expected returns are the AUC which we dont want them to decrease over time or have huge downtrends, and the volatility/dowdrown are what we're trying to eliminate here. An equivalent of the sharpe ratio (mean divided by std) is probably best suited to evaluate this, as it is used in quant firms. See: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501708\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501708</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2825956,
      "author_name": "sheemazain",
      "author_url": "",
      "post_date": "05/20/2024 16:41:00",
      "content": "<p>I agree. This is definitely regime dependent!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2823667": "This competition is not over yet, but it is true it doesn't look too good. In my opinion, worse than the metric choice is the repeated change of mind regarding the evaluation.\n\nI stopped Kaggling over two years ago, lacking the confidence that, conditioned on the host, my effort of many weeks wouldn't be dealt with by an RNG for leaderboard. On the sidelines since, Home Credit Stability looked like an interesting premise, at first sight, maybe time to come back to the arena... and here we are.\n\nStill, I see more than one \"elephant in the room\" here. We are supposed to achieve \"stability\" in contrast to \"model degradation\" over time. The question is, **why does a model degrade \"over time\" in general, and why does a model degrade in this domain and task in particular?**\n\nIn a nutshell, in general, I would give three main reasons for a model to degrade over time:\n\n**1. Concept Drift:** the answers to the same questions are now different.\n\n**2. Covariate Shift:** the answers are the same, but the questions are worded differently and harder to relate to previous questions.\n\n**3. User adaptation:** somehow you influenced the test you are seeing, i.e. you changed the future with your previous answers and that future sample is now different, so possibly both points 1 and 2.\n\n**So, regarding this competition, which one applies?**\n\nFor sure in the \"real world\" point 3 applies as credit companies will use their models to select which credits to give, hence influencing the future sample their models will be seeing. This is a well-known problem of causal/ interactive/ reinforcement learning when you change your future with your model. In this particular task, models can be thought of as learning to make inferences on a sample of approved credits. **Is this relevant for this competition? I would bet possibly not because we know to be working with heavily ad-hoc subsampled data.** One less thing to worry about for us, but possibly something that skips one of the interesting (and solvable) sources of instability.\n\n**But there is the problem number 2, the big one. The question would go like this: are the statistical properties of this task time/regime dependent in a non-stationary way?** Or, in other words, how predictable is model degradation in this use case? Because it looks to me like a financial-regime dependent, non-stationary problem. This means that the infamous slope of the weekly gini regression can be up or down, depending on conditions that most companies struggle to forecast at all. It also would need much wider data time span for any attempt to model. **So here goes this critic to the metric, not from its hackability but because of it pursuing stability by measuring something that is possibly regime-dependent, non stationary, and pointless to attempt to model at this time scale.**\n\nI hope that all this made some sense. Anyway, I do find this competition interesting in spite of its flaws. I also understand that the hosts can still obtain insights from the solutions to the problem in its current form,  no matter how much sense it makes from a Kaggler point of view.  Quite curious to witness how it will unfold. \n\nGood luck to all!",
    "2823692": ">Because it looks to me like a financial-regime dependent, non-stationary problem\n\nI agree, based on the training data we have observed.\n\nI think the current metric could be modified to better capture predictable changes in performance while reducing the probability of false alarms on stochastic trends. I've discussed this here https://www.kaggle.com/code/eivolkova/stability-metric-issue-probabilistic-approach. \n\nHowever, I believe there are no predictable changes in performance, so including a volatility penalty should be sufficient. \n\nPerhaps formulating the problem in terms of the financial losses of the company would be a better approach to capture the costs of dropping performance.",
    "2823729": "That's interesting. Regarding the financial P&L of a company, I'm afraid it would imply adding many more layers of complexity over the uncertainty problem. I'm not an expert in this domain but if today I had to draw a tentative evaluation of the stability of a model for this context, I would attempt to normalize performance for each time point's by each current global credit market-performance.",
    "2823905": "Interested in your comment regarding user adaptation which I take it to include issues like reject inference: *one of the interesting (and solvable) sources of instability*\nWhat solutions did you have in mind?",
    "2823910": "I've seen a few papers on this, e.g. https://www.sciencedirect.com/science/article/pii/S0305048323001688#:~:text=The%20total%20profit%20of%20a,the%20MP%20measure%20%5B5%5D. \n>Models that were profitable during economic recovery might bring losses during turbulent times because of changes in the benefits associated with the correct classification and losses from misclassification. Hence, there has been a recent surge in demand for estimating AI systems’ profit in credit scoring under changing market conditions, especially during the COVID-19 pandemic\n\n.\n\n>I would attempt to normalize performance for each time point's by each current global credit market-performance.\n\nThat's an interesting idea, though it may be difficult to find an appropriate benchmark",
    "2823989": "I agree. This is definitely regime dependent.\nThere are striking similarities between this problem and the quant strategies evaluation in Finance. The expected returns are the AUC which we dont want them to decrease over time or have huge downtrends, and the volatility/dowdrown are what we're trying to eliminate here. An equivalent of the sharpe ratio (mean divided by std) is probably best suited to evaluate this, as it is used in quant firms. See: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/501708",
    "2824262": "Just had a first look at the paper. As I get it, the approach does add some complexity and some assumptions while optimizing for a different (more interesting) risk dependent objective.\n\nIn this competition setup a multi-stage use of the models is implied, being our task essentially a ranking by week of candidates. Cutoff optimization and portfolio risk tunning probably coming afterwards through non automated human decision. The paper expands the scoring scope to profit and risk hence feature selection will optimize for that. Not sure that is equivalent to reduce the uncertainty of the ranking, more I would say the uncertainty is circumvented while optimizing for profit, quite similarly to how money allocation accounts for a given trading system uncertainty.\n\nAnyway the paper is an interesting finding, also makes more clear the meaning of your point about using P&L for optimization that I had not fully understood in your first mention. Now it makes sense, it might be definitely a way to explore.",
    "2824268": "Yes, reject inference definitely belongs to the core of user adaptation problem.  \n\nAbout being solvable, of course I meant partially solvable. Actual solutions are both problem and model dependent.\nI simply expressed my impression that it is not a relevant source of degradation in this competition while alternative competition designs might have included it as something to mitigate/account for.",
    "2824455": "For credit ranking problems I have not seen good methods. Some methods in the literature or common practice are intuitive but with no objective improvement metric (unless one randomizes acceptance criteria and accepts the cost for a small random sample, not realistic).",
    "2825956": "I agree. This is definitely regime dependent!"
  },
  "source": "meta"
}