{
  "id": 477074,
  "title": "FAQ - Please read me before posting",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/477074",
  "author_name": "Daniel Herman",
  "post_date": "2024-02-14T15:34:50.855000",
  "votes": 37,
  "comment_count": 33,
  "views": 0,
  "content": "<p><strong>Q: Why can't I submit a solution in my country?</strong><br>\nA: We disabled the submissions as a temporary measure. It will be enabled again around 6.3.2024. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478716\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478716</a></p>\n<p><strong>Q: Questions regarding transformation of test data.</strong><br>\nA: The transformation was done to train and test before the start of the competition and now we will transform test data in different manner. We won't specify what is the transformation exactly. We will keep column names, dtypes and the change is not going to be drastic. </p>\n<p><strong>Q: Any question related to private test set</strong><br>\nA: We will not disclose any information about private test set. There is no point in asking questions about the private test set. We made the split in such a way that it will be fair for final evaluation.</p>\n<p><strong>Q: How long is the test set?</strong><br>\nA: For the sake of stability evaluation it would not make sense to disclose you that information. This information would lead even to more \"hacks\" of the metric. Think about it this way. The model will be deployed and you don't know for how long it is going to be.</p>\n<p><strong>Q: What is the meaning of num_group1 and num_group2?</strong><br>\nA: Those are indices that will help you to aggregate tables with depth=1,2. Please have a look at Data section.</p>\n<p><strong>Q: Any questions regarding stability metric and hacking.</strong><br>\nPlease read <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867</a> first. We are working on a solution that would minimize such hacking by artificially worsening the score in the near future. </p>\n<p><strong>Q: Can't we just remove the falling rate from metric?</strong><br>\nA:  Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic.</p>\n<p><strong>Q: What is the reason for having those weights in the stability metric?</strong><br>\nA: Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models.</p>\n<p><strong>Q: Why is my notebook failing?</strong><br>\nA: We don't have capacity to review every notebook, unfortunately. You can post it in discussion and maybe someone will help. In future, we can collect those errors into this thread. </p>\n<p><strong>Q: Dates are inconsistent and I see dates from the future in respect to date decision.</strong><br>\nA: This might happen as all dates were transformed, but there might be some inconsistencies. Also, the underlying data from our primary sources can be flawed, sometimes it happens. </p>\n<p><strong>Q: What is the definition of the used target?</strong><br>\nA: It's unpaid payment (one is enough) in certain time period. There is also some time tolerance (e.g. one day late is not default) and amount tolerance (if client paid $100 instead of $100.10). We can't disclose further details. </p>",
  "messages": [
    {
      "id": 2652153,
      "postDate": "2024-02-14T15:34:50.857Z",
      "content": "<p><strong>Q: Why can't I submit a solution in my country?</strong><br>\nA: We disabled the submissions as a temporary measure. It will be enabled again around 6.3.2024. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478716\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478716</a></p>\n<p><strong>Q: Questions regarding transformation of test data.</strong><br>\nA: The transformation was done to train and test before the start of the competition and now we will transform test data in different manner. We won't specify what is the transformation exactly. We will keep column names, dtypes and the change is not going to be drastic. </p>\n<p><strong>Q: Any question related to private test set</strong><br>\nA: We will not disclose any information about private test set. There is no point in asking questions about the private test set. We made the split in such a way that it will be fair for final evaluation.</p>\n<p><strong>Q: How long is the test set?</strong><br>\nA: For the sake of stability evaluation it would not make sense to disclose you that information. This information would lead even to more \"hacks\" of the metric. Think about it this way. The model will be deployed and you don't know for how long it is going to be.</p>\n<p><strong>Q: What is the meaning of num_group1 and num_group2?</strong><br>\nA: Those are indices that will help you to aggregate tables with depth=1,2. Please have a look at Data section.</p>\n<p><strong>Q: Any questions regarding stability metric and hacking.</strong><br>\nPlease read <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867</a> first. We are working on a solution that would minimize such hacking by artificially worsening the score in the near future. </p>\n<p><strong>Q: Can't we just remove the falling rate from metric?</strong><br>\nA:  Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic.</p>\n<p><strong>Q: What is the reason for having those weights in the stability metric?</strong><br>\nA: Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models.</p>\n<p><strong>Q: Why is my notebook failing?</strong><br>\nA: We don't have capacity to review every notebook, unfortunately. You can post it in discussion and maybe someone will help. In future, we can collect those errors into this thread. </p>\n<p><strong>Q: Dates are inconsistent and I see dates from the future in respect to date decision.</strong><br>\nA: This might happen as all dates were transformed, but there might be some inconsistencies. Also, the underlying data from our primary sources can be flawed, sometimes it happens. </p>\n<p><strong>Q: What is the definition of the used target?</strong><br>\nA: It's unpaid payment (one is enough) in certain time period. There is also some time tolerance (e.g. one day late is not default) and amount tolerance (if client paid $100 instead of $100.10). We can't disclose further details. </p>",
      "rawMarkdown": "**Q: Why can't I submit a solution in my country?**\nA: We disabled the submissions as a temporary measure. It will be enabled again around 6.3.2024. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478716\n\n**Q: Questions regarding transformation of test data.**\nA: The transformation was done to train and test before the start of the competition and now we will transform test data in different manner. We won't specify what is the transformation exactly. We will keep column names, dtypes and the change is not going to be drastic. \n\n**Q: Any question related to private test set**\nA: We will not disclose any information about private test set. There is no point in asking questions about the private test set. We made the split in such a way that it will be fair for final evaluation.\n\n**Q: How long is the test set?**\nA: For the sake of stability evaluation it would not make sense to disclose you that information. This information would lead even to more \"hacks\" of the metric. Think about it this way. The model will be deployed and you don't know for how long it is going to be.\n\n**Q: What is the meaning of num_group1 and num_group2?**\nA: Those are indices that will help you to aggregate tables with depth=1,2. Please have a look at Data section.\n\n**Q: Any questions regarding stability metric and hacking.**\nPlease read https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867 first. We are working on a solution that would minimize such hacking by artificially worsening the score in the near future. \n\n**Q: Can't we just remove the falling rate from metric?**\nA:  Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic.\n\n**Q: What is the reason for having those weights in the stability metric?**\nA: Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models.\n\n**Q: Why is my notebook failing?**\nA: We don't have capacity to review every notebook, unfortunately. You can post it in discussion and maybe someone will help. In future, we can collect those errors into this thread. \n\n**Q: Dates are inconsistent and I see dates from the future in respect to date decision.**\nA: This might happen as all dates were transformed, but there might be some inconsistencies. Also, the underlying data from our primary sources can be flawed, sometimes it happens. \n\n**Q: What is the definition of the used target?**\nA: It's unpaid payment (one is enough) in certain time period. There is also some time tolerance (e.g. one day late is not default) and amount tolerance (if client paid $100 instead of $100.10). We can't disclose further details. ",
      "votes": 37
    },
    {
      "id": 2652661,
      "postDate": "2024-02-14T23:10:47.477Z",
      "content": "<blockquote>\n  <p>Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk manager in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic.</p>\n</blockquote>\n<p>The falling rate was ill conceived. There should be no scenario where a model that has consistently poor performance is deemed better than a model that sometimes has good performance and sometimes poor performance. To use such a metric is to misunderstand what the value of stability is. That is, the value of stability is never in being consistently bad, but in avoiding periods that are particularly bad. Once this is understood, it becomes obvious that the best way to encourage stability is by aggressively punishing periods of particularly bad performance. We can do this with a metric that rapidly become worse when very bad samples are present. As I suggested in the other thread, this can be done with a metric of the form: </p>\n<p>mean(  ( -log( MovingAverage(gini) ) )**exponent   ) with exponent&gt;=1</p>\n<p>Under this metric, periods of poor performance are heavily penalized without the silliness of punishing good rounds. Stability should require nothing more than this.</p>",
      "rawMarkdown": ">Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk manager in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic.\n\nThe falling rate was ill conceived. There should be no scenario where a model that has consistently poor performance is deemed better than a model that sometimes has good performance and sometimes poor performance. To use such a metric is to misunderstand what the value of stability is. That is, the value of stability is never in being consistently bad, but in avoiding periods that are particularly bad. Once this is understood, it becomes obvious that the best way to encourage stability is by aggressively punishing periods of particularly bad performance. We can do this with a metric that rapidly become worse when very bad samples are present. As I suggested in the other thread, this can be done with a metric of the form: \n\nmean(  ( -log( MovingAverage(gini) ) )**exponent   ) with exponent>=1\n\nUnder this metric, periods of poor performance are heavily penalized without the silliness of punishing good rounds. Stability should require nothing more than this.",
      "votes": 8,
      "replies": [
        {
          "id": 2653310,
          "postDate": "2024-02-15T10:49:34.977Z",
          "content": "<blockquote>\n  <p>The falling rate was ill conceived. There should be no scenario where a model that has consistently poor performance is deemed better than a model that sometimes has good performance and sometimes poor performance. To use such a metric is to misunderstand what the value of stability is. </p>\n</blockquote>\n<p>Frankly, I think there is also some misunderstanding on your side. Consider this. From the perspective of risk managers at HC it is undesirable to have a model that is performing well let's say first 6 months and then it levels on for example 60%-80% of that performance. Even if the model performed better in terms of integral of weekly gini over time the model was unpredictable in sense that its performance lowered after the first 6 months. What is desirable is a model that has the same exact performance every week and if there is some decrease it is well predictable. </p>\n<p>Omitting the falling rate term will allow models that will be performing well on some portion of test sample and then they will have decrease in performance that might not be even linear of exponential. </p>\n<p>Regarding proposed metric, you could try to rewrite it using LaTeX with proper sums so it is clear what is applied to what. As it is now I am not sure if I understand it right. </p>\n<p>Regarding the moving average idea, I believe that this might lower the amount of hacking, but it will not solve the problem completely. Consider following model that is performing well the first N weeks and during next 3 weeks the performance drops to 60% of that. By lowering the score artificially on the first N weeks you will improve the stability metric. That is aligned with risk managers' view of performance. It is not aligned with metric being fair and fun for everyone, we are working on that and I ask you for patience. The score from the metric that you suggested will improve when decreasing the score artificially first N weeks. If this is not the case, then I don't understand it perhaps correctly and you could explain it in finer detail. </p>",
          "rawMarkdown": ">The falling rate was ill conceived. There should be no scenario where a model that has consistently poor performance is deemed better than a model that sometimes has good performance and sometimes poor performance. To use such a metric is to misunderstand what the value of stability is. \n\nFrankly, I think there is also some misunderstanding on your side. Consider this. From the perspective of risk managers at HC it is undesirable to have a model that is performing well let's say first 6 months and then it levels on for example 60%-80% of that performance. Even if the model performed better in terms of integral of weekly gini over time the model was unpredictable in sense that its performance lowered after the first 6 months. What is desirable is a model that has the same exact performance every week and if there is some decrease it is well predictable. \n\nOmitting the falling rate term will allow models that will be performing well on some portion of test sample and then they will have decrease in performance that might not be even linear of exponential. \n\nRegarding proposed metric, you could try to rewrite it using LaTeX with proper sums so it is clear what is applied to what. As it is now I am not sure if I understand it right. \n\nRegarding the moving average idea, I believe that this might lower the amount of hacking, but it will not solve the problem completely. Consider following model that is performing well the first N weeks and during next 3 weeks the performance drops to 60% of that. By lowering the score artificially on the first N weeks you will improve the stability metric. That is aligned with risk managers' view of performance. It is not aligned with metric being fair and fun for everyone, we are working on that and I ask you for patience. The score from the metric that you suggested will improve when decreasing the score artificially first N weeks. If this is not the case, then I don't understand it perhaps correctly and you could explain it in finer detail. ",
          "votes": 3,
          "replies": [
            {
              "id": 2653528,
              "postDate": "2024-02-15T13:58:10.207Z",
              "content": "<p>I don’t think it will improve by lowering the score in first few weeks. Every lowering of gini score will result in a lower moving average. If moving average gets smaller, -log will get larger. The best performance is reached when all ginis in all relevant time windows are 1, then the above metric is 0 as log(1)=0.<br>\nThat badscores are punished heavily is also clear as the log decreases rapidly for arguments smaller than 1 (and even more so the closer we get to zero.)<br>\nPlease correct if I’m wrong <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a> </p>",
              "rawMarkdown": "I don’t think it will improve by lowering the score in first few weeks. Every lowering of gini score will result in a lower moving average. If moving average gets smaller, -log will get larger. The best performance is reached when all ginis in all relevant time windows are 1, then the above metric is 0 as log(1)=0.\nThat badscores are punished heavily is also clear as the log decreases rapidly for arguments smaller than 1 (and even more so the closer we get to zero.)\nPlease correct if I’m wrong @jacobyjaeger ",
              "votes": 2
            },
            {
              "id": 2654189,
              "postDate": "2024-02-16T00:21:26.750Z",
              "content": "<blockquote>\n  <p>Frankly, I think there is also some misunderstanding on your side. Consider this. From the perspective of risk managers at HC it is undesirable to have a model that is performing well let's say first 6 months and then it levels on for example 60%-80% of that performance. </p>\n</blockquote>\n<p>What you need is not to be able to predict the performance but to be able to put a good lower bound on what the performance could be. This is because the event of the model performing better than you might have prepared for is never a problem and as such does not need to be penalized. The set of models that can be given a high lower bound on their expected performance are those that have demonstrated a high lower bound during the validation period, which are those that perform well under a metric that heavily penalizes a given model's worst rounds as mine does.</p>\n<blockquote>\n  <p>Regarding proposed metric, you could try to rewrite it using LaTeX with proper sums so it is clear what is applied to what. As it is now I am not sure if I understand it right.</p>\n</blockquote>\n<p>Once predictions have been converted to ginis, there is only one axis that we need to be concerned with, the time axis. Here is a simple implementation of the metric. Higher scores are considered worse under this metric. </p>\n<pre><code> numpy  np\n matplotlib  pyplot  plt\n\n\n ():\n    (axis==)\n\n    x = np.cumsum( x, axis= )\n\n    x = np.concatenate([ *x[:], x ],)\n\n     (x[n:] - x[:-n])/n\n\n\n ():\n    \n\n     ():\n\n        scores = ( - np.log(  MA( ginis, ma_len, axis= ) ) )**exponent\n\n          scores\n\n     metric \n\n\n\nN_GINIS = \nN_MA = \nEXPONENT = \n\nginis = np.random.rand(N_GINIS)\n\ntime = np.arange(N_GINIS)\n\n\nmetric = get_metric(EXPONENT, N_MA)\n\n\nscores = metric(ginis)\n\n\nplt.plot(time, ginis)\nplt.plot(time[N_MA//: N_GINIS-(N_MA-)//],  scores)\n\n\nscore = np.mean(scores)\nscore\n</code></pre>\n<blockquote>\n  <p>Omitting the falling rate term will allow models that will be performing well on some portion of test sample and then they will have decrease in performance that might not be even linear of exponential.</p>\n</blockquote>\n<p>Any period of poor performance is heavily punished under my metric. But the falling rate can only detect periods of poor performance if they come at the end of the test period.</p>\n<blockquote>\n  <p>Regarding the moving average idea, I believe that this might lower the amount of hacking, but it will not solve the problem completely. Consider following model that is performing well the first N weeks and during next 3 weeks the performance drops to 60% of that. By lowering the score artificially on the first N weeks you will improve the stability metric. That is aligned with risk managers' view of performance.</p>\n</blockquote>\n<p>I'm not sure what you are trying to say here. Are you suggesting that performance on my metric would be improved by such hacking? That is not the case, it can be seen from the fact that the gradient wrt sample ginis is strictly negative that making any round's performance worse can only hurt performance under my metric. </p>\n<p>Or are you saying that it is desirable that such hacking results in a higher score? That would be rather silly, you can make a declining performance look flat during the test period by making the early test rounds artificially worse, but that does not change the fact that true performance was declining which will become evident after the test period when you no longer have artificially deflated rounds left to compare the performance to. </p>",
              "rawMarkdown": ">Frankly, I think there is also some misunderstanding on your side. Consider this. From the perspective of risk managers at HC it is undesirable to have a model that is performing well let's say first 6 months and then it levels on for example 60%-80% of that performance. \n\nWhat you need is not to be able to predict the performance but to be able to put a good lower bound on what the performance could be. This is because the event of the model performing better than you might have prepared for is never a problem and as such does not need to be penalized. The set of models that can be given a high lower bound on their expected performance are those that have demonstrated a high lower bound during the validation period, which are those that perform well under a metric that heavily penalizes a given model's worst rounds as mine does.\n\n>Regarding proposed metric, you could try to rewrite it using LaTeX with proper sums so it is clear what is applied to what. As it is now I am not sure if I understand it right.\n\n\nOnce predictions have been converted to ginis, there is only one axis that we need to be concerned with, the time axis. Here is a simple implementation of the metric. Higher scores are considered worse under this metric. \n\n```python\n\nimport numpy as np\nfrom matplotlib import pyplot as plt\n\n\ndef MA( x, n, axis=0 ):\n    assert(axis==0)\n    \n    x = np.cumsum( x, axis=0 )\n    \n    x = np.concatenate([ 0*x[:1], x ],0)\n    \n    return (x[n:] - x[:-n])/n\n\n\ndef get_metric(exponent=1, ma_len=5):\n    '''\n    A metric that heavily punishes rounds of poor performance.\n    Higher exponent values punish bad rounds more disproportionately.\n    Larger MA len's place focus on longer term performance.\n    \n    Lower mean scores mean better performance.\n    '''\n    \n    def metric( ginis ):\n        \n        scores = ( - np.log(  MA( ginis, ma_len, axis=0 ) ) )**exponent\n        \n        return  scores\n    \n    return metric \n\n\n\nN_GINIS = 100\nN_MA = 5\nEXPONENT = 2\n\nginis = np.random.rand(N_GINIS)\n\ntime = np.arange(N_GINIS)\n\n\nmetric = get_metric(EXPONENT, N_MA)\n\n\nscores = metric(ginis)\n\n\nplt.plot(time, ginis)\nplt.plot(time[N_MA//2: N_GINIS-(N_MA-1)//2],  scores)\n\n\nscore = np.mean(scores)\nscore\n```\n\n>Omitting the falling rate term will allow models that will be performing well on some portion of test sample and then they will have decrease in performance that might not be even linear of exponential.\n\nAny period of poor performance is heavily punished under my metric. But the falling rate can only detect periods of poor performance if they come at the end of the test period.\n\n\n>Regarding the moving average idea, I believe that this might lower the amount of hacking, but it will not solve the problem completely. Consider following model that is performing well the first N weeks and during next 3 weeks the performance drops to 60% of that. By lowering the score artificially on the first N weeks you will improve the stability metric. That is aligned with risk managers' view of performance.\n\nI'm not sure what you are trying to say here. Are you suggesting that performance on my metric would be improved by such hacking? That is not the case, it can be seen from the fact that the gradient wrt sample ginis is strictly negative that making any round's performance worse can only hurt performance under my metric. \n\nOr are you saying that it is desirable that such hacking results in a higher score? That would be rather silly, you can make a declining performance look flat during the test period by making the early test rounds artificially worse, but that does not change the fact that true performance was declining which will become evident after the test period when you no longer have artificially deflated rounds left to compare the performance to. "
            },
            {
              "id": 2654192,
              "postDate": "2024-02-16T00:26:12.940Z",
              "content": "<p>You're exactly right, Simon. We can check the sign of the gradient of the final score wrt the per round scores to prove that it is never desirable to make any round score worse. I suggest that kaggle should perform such a check on every competition metric so that we don't see hackable metrics again in the future.</p>",
              "rawMarkdown": "You're exactly right, Simon. We can check the sign of the gradient of the final score wrt the per round scores to prove that it is never desirable to make any round score worse. I suggest that kaggle should perform such a check on every competition metric so that we don't see hackable metrics again in the future.",
              "votes": 2
            },
            {
              "id": 2654518,
              "postDate": "2024-02-16T07:24:44.077Z",
              "content": "<blockquote>\n  <p>What you need is not to be able to predict the performance but to be able to put a good lower bound on what the performance could be. This is because the event of the model performing better than you might have prepared for is never a problem and as such does not need to be penalized. </p>\n</blockquote>\n<p>From business point of view this statement is not true. It's not only about the credit risk model, but whole underwriting setup, product parameters, profitability models and many other things. We believe we know why we defined stability as we did.</p>",
              "rawMarkdown": ">What you need is not to be able to predict the performance but to be able to put a good lower bound on what the performance could be. This is because the event of the model performing better than you might have prepared for is never a problem and as such does not need to be penalized. \n\nFrom business point of view this statement is not true. It's not only about the credit risk model, but whole underwriting setup, product parameters, profitability models and many other things. We believe we know why we defined stability as we did.",
              "votes": 3
            },
            {
              "id": 2654536,
              "postDate": "2024-02-16T07:47:46.097Z",
              "content": "<p>In that case if a model performs \"too\" well in a given period, what is the negative consequence? What is the problem with it being \"too profitable\" one month?</p>",
              "rawMarkdown": "In that case if a model performs \"too\" well in a given period, what is the negative consequence? What is the problem with it being \"too profitable\" one month?",
              "votes": 3
            },
            {
              "id": 2654853,
              "postDate": "2024-02-16T13:36:07.927Z",
              "content": "<p>The problem is not being too profitable, the problem is being unpredictable in production. What we want is to get a model that this predictable in performance for as long period as possible and ideally stable in time (in terms of performance). The score of a model is processed further with different processes. In other words, the score is not used directly to approve loans. </p>\n<p>Regarding your previous post, I will test the proposed metric and reply to that later, I am working on the announced fix. </p>",
              "rawMarkdown": "The problem is not being too profitable, the problem is being unpredictable in production. What we want is to get a model that this predictable in performance for as long period as possible and ideally stable in time (in terms of performance). The score of a model is processed further with different processes. In other words, the score is not used directly to approve loans. \n\nRegarding your previous post, I will test the proposed metric and reply to that later, I am working on the announced fix. ",
              "votes": 1
            },
            {
              "id": 2654915,
              "postDate": "2024-02-16T14:52:13.080Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2662077,
              "postDate": "2024-02-21T17:28:55.827Z",
              "content": "<p>Okay <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a> I am back after testing it. I am here to provide you few examples where the current stability metric performs as we want it to perform, but your proposed metric does not select the correct option. Here are 5 pairs of simulated performance and A is always the preferred choice. The stability metric correctly says that the A is more preferred (with higher value of metric) and your metric says the opposite. Please have a look at this.</p>\n<pre><code> numpy  np\n matplotlib  pyplot  plt\n\n ():\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n     avg_gini + w_fallingrate * (, a) + w_resstd * res_std\n\n ():\n    x = np.cumsum(gini_in_time, axis=)\n    x = np.concatenate([*x[:], x], )\n    scores = -np.mean(-np.log(np.maximum((x[ma_len:] - x[:-ma_len])/ma_len, ))**exponent)\n     scores \n\n\n(proposed_metric([, , , , ]), proposed_metric([, , , , ]))\n\n\n\ndata = [\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n]\n row  data:\n    target = row[]\n    stability_metric_A = gini_stability(row[], w_fallingrate=, w_resstd=-)\n    stability_metric_B = gini_stability(row[], w_fallingrate=, w_resstd=-)\n    proposed_metric_A = proposed_metric(row[], exponent=, ma_len=)\n    proposed_metric_B = proposed_metric(row[], exponent=, ma_len=)\n     target ==   stability_metric_A &gt; stability_metric_B  proposed_metric_A &lt; proposed_metric_B:\n        ()\n        ()\n        plt.plot(row[], label=)\n        plt.plot(row[], label=)\n        plt.ylabel()\n        plt.xlabel()\n        plt.legend()\n        plt.show()\n</code></pre>",
              "rawMarkdown": "Okay @jacobyjaeger I am back after testing it. I am here to provide you few examples where the current stability metric performs as we want it to perform, but your proposed metric does not select the correct option. Here are 5 pairs of simulated performance and A is always the preferred choice. The stability metric correctly says that the A is more preferred (with higher value of metric) and your metric says the opposite. Please have a look at this.\n\n```python\nimport numpy as np\nfrom matplotlib import pyplot as plt\n\ndef gini_stability(gini_in_time, w_fallingrate=88.0, w_resstd=-0.5):\n    x = np.arange(len(gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, 1)\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n    return avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std\n\ndef proposed_metric(gini_in_time, exponent=1, ma_len=5):\n    x = np.cumsum(gini_in_time, axis=0)\n    x = np.concatenate([0*x[:1], x], 0)\n    scores = -np.mean(-np.log(np.maximum((x[ma_len:] - x[:-ma_len])/ma_len, 1e-5))**exponent)\n    return scores \n\n# what is the meaning of higher score?\nprint(proposed_metric([0.9, 0.9, 0.9, 0.9, 0.9]), proposed_metric([0.9, 0.9, 0.9, 0.9, 0.8]))\n# -0.10536051565782628 -0.12783337150988477\n# higher is better\n\ndata = [\n    {'timeseries_A': [0.45931853897247277, 0.5567067540723954, 0.4825901177734739, 0.512594840757675, 0.5213442065673528, 0.5135029054410122, 0.4943336324321425, 0.5152234525848737, 0.5146463547382705, 0.5038023813007937, 0.5397269895234323, 0.5266249170508955, 0.4921021327544269, 0.48378519753277266, 0.4822064844476577, 0.5043200548192804, 0.4917529149366844, 0.5279542186854929, 0.5312255516354792, 0.49067434612077354, 0.5272594478007019, 0.4757360669629569, 0.4781711160803342, 0.5386923816576479, 0.5110302894041262, 0.4639983930246085, 0.4930094020868559, 0.5132064930156196, 0.4808508989691657, 0.42406569232376445, 0.506688297601109, 0.4893717267144967, 0.4468218834679482, 0.48818171974084273, 0.48886184076865924, 0.5300270237737726, 0.46358870277229225, 0.4891250255256556, 0.49466748406586813, 0.5179960577137955, 0.503486257333416, 0.47193092429499284, 0.490661118730109, 0.4735919949685026, 0.45169056674732805, 0.5078023284980676, 0.4446048993565578, 0.46562273113525227, 0.46320328044161835, 0.45381385090335985, 0.4877709870401442, 0.5009434622402376, 0.4442606940494906, 0.44835045036753335, 0.415673871747707, 0.4872443945432088, 0.4816445401411326, 0.48343370136827457, 0.4573879128043373, 0.4409649975935981, 0.47543186314679936, 0.4770095693838089, 0.44085806859971166, 0.4775226304738473, 0.4773773917097745, 0.48405348833055517, 0.42428918158907125, 0.4212659355199594, 0.4376028510073539, 0.49328204971140704, 0.4267081415260231, 0.45883453677208114, 0.41397316007240414, 0.46972198330207715, 0.48225935899984423, 0.46274111318728006, 0.4240172855770951, 0.46215584767072365, 0.4514716790994658, 0.4593814759428128, 0.4518050578434685, 0.4685243577064591, 0.43803906175282525, 0.4620023371152737, 0.4393578049863158, 0.46599791571812066, 0.46237610301217386, 0.45446742219601777, 0.4377169920230219, 0.4211815014455423, 0.3991602617067614, 0.44824171296247134, 0.4610268770993053, 0.4243746444677177, 0.4248544840760521, 0.4572760605096456, 0.4255960555365488, 0.39644047098573837, 0.4343217659342772, 0.45324334728114857, 0.43953211031260486, 0.46328365151682577], 'timeseries_B': [0.5680099987408721, 0.5884270580234222, 0.5898876074432621, 0.6003138685157438, 0.551228714999531, 0.5349790537323957, 0.5656571404451941, 0.5944802938798136, 0.5779805268295671, 0.581862855213639, 0.5343184692567533, 0.5955247349192402, 0.5389691428884885, 0.5415294805103056, 0.5862649316006328, 0.5697536342248863, 0.554349682757325, 0.5336471062394862, 0.5572602631305275, 0.5533107900913204, 0.5504162486594195, 0.5105567182157634, 0.5159309378197687, 0.5536228998721683, 0.5028986348537446, 0.5179791876278644, 0.5382504758153427, 0.5641923428496285, 0.49597113867830817, 0.5646953569782098, 0.5165616033513313, 0.5168699500287747, 0.5187389976292371, 0.5256125744883205, 0.5368723641108188, 0.5017694891794096, 0.4860979738169708, 0.5253362003459279, 0.5123437279231432, 0.5484175623873022, 0.5052839490937887, 0.48150707831461814, 0.48091773956347983, 0.484696691723045, 0.487279353526466, 0.49839306889474994, 0.4850468879627133, 0.4984101241249189, 0.48307679885912164, 0.5234584251535545, 0.47656636050148155, 0.4952015310113332, 0.5279890269332961, 0.5111752455047976, 0.484260580046016, 0.4733799092708589, 0.4613318746373804, 0.4812666110109198, 0.5039152343900071, 0.42789856771355606, 0.4938043628101529, 0.48437544971844537, 0.4455690671766333, 0.4869412694074684, 0.42976251032235097, 0.4337192019547387, 0.4462709308706712, 0.4813780440534373, 0.4289591521542533, 0.44000823833392827, 0.49758904135162313, 0.4494426129326302, 0.45684668285615954, 0.40102329662438224, 0.4497459115785034, 0.4451294954750487, 0.4240119311470284, 0.42085644871902606, 0.4418773512799505, 0.41652055832623475, 0.4599043963276754, 0.4225341862689644, 0.4403486933309671, 0.41556175615032304, 0.40935046448215295, 0.433138660302346, 0.4225141536453778, 0.3770137589894561, 0.38942837176841844, 0.44119762428550213, 0.42565196780255543, 0.422840302272181, 0.39975892682757513, 0.3951348393288056, 0.3511627699105334, 0.4080825397831498, 0.40911990171909773, 0.4020680356510034, 0.3866797342897758, 0.36483843121041143, 0.35840410925834054, 0.39095415259603034], 'target': 'A'},\n    {'timeseries_A': [0.38277105177821946, 0.4298153459114027, 0.4555700800535953, 0.49978357733725964, 0.5120119357572293, 0.5014759728509398, 0.42886969443172707, 0.6239415331700883, 0.5104019745688886, 0.4385300637353016, 0.1865617367176554, 0.536227986578527, 0.5546562723683358, 0.627280868864281, 0.4675918721715071, 0.4550186895621498, 0.4969487286750642, 0.49876778019721535, 0.3887648937850893, 0.553054715485712, 0.4321418536576932, 0.5672033427323433, 0.42035451718693206, 0.4435663596050058, 0.5647498347815504, 0.41451105296252416, 0.39241995242761635, 0.5521707472452356, 0.5166664603144179, 0.5141158273711287, 0.4750506169283815, 0.25067114513919275, 0.4649908791178805, 0.42382415828524705, 0.430893600765377, 0.44669977528242405, 0.44002267862385763, 0.29461542789608797, 0.31911957798874313, 0.5766629645256923, 0.7086370083485398, 0.6071466046220485, 0.4410141848652491, 0.38349620514362176, 0.46778946811189465, 0.580738676918999, 0.43446598818733656, 0.4737676086919768, 0.35759795575399506, 0.43730343584945025, 0.47114589726857903, 0.5810386827998576, 0.6190314059729546, 0.38443884731150124, 0.5643003663220149, 0.39972213437349496, 0.5094329989161557, 0.5603519818103472, 0.47361600297609086, 0.39217730194580436, 0.4062027938645795, 0.37366747729430355, 0.31788656596435866, 0.4498987139662878, 0.39060289149163313, 0.39762601154312205, 0.5839523596231831, 0.43761599918984606, 0.3170675939769194, 0.5194855688378123, 0.48183283006824834, 0.39752161332184116, 0.34396597166193527, 0.4061004923572341, 0.42547037697171103, 0.3524249879465508, 0.3374836438650132, 0.49636695833173455, 0.5163471592579683, 0.4614449065149918, 0.37633373729388003, 0.3889473218054684, 0.37542364391049204, 0.4474402381844267, 0.4342172044656317, 0.4643518656614464, 0.5247468171089611, 0.2658205249534029, 0.4348094424093696, 0.45408663431042534, 0.36073016113961354, 0.3736412214061613, 0.38090603679837864, 0.5567927496749354, 0.3620428138914703, 0.40682295721953204, 0.4245521616048502, 0.25020727434461554, 0.4415705653379547, 0.18045756246173553, 0.540110131734328, 0.3268584063566512], 'timeseries_B': [0.5649290098858608, 0.5604061207474527, 0.5999928579099665, 0.5911852582901042, 0.5829027994764711, 0.5672210507264529, 0.5729483984699694, 0.5705795432935264, 0.5223803966343707, 0.5741702012451493, 0.5645631988585531, 0.5365698806534049, 0.5611480200208953, 0.5678444744988338, 0.563756158703405, 0.5257265590994212, 0.5660756885174074, 0.5548845953510743, 0.5545760801756416, 0.5116631849776518, 0.5331221914504484, 0.544739660084671, 0.5345521070019253, 0.4994706928489658, 0.49045195814773085, 0.5244101139627537, 0.5431879070199751, 0.4969432120579304, 0.48817988285590985, 0.4774937415457261, 0.5099676210392469, 0.5048371245277136, 0.5390626354426095, 0.48932810535486987, 0.5205855091397935, 0.5035558243684439, 0.5067774680753907, 0.5098100106561956, 0.5059451312137639, 0.4851446172650819, 0.4741171556475985, 0.45423802076276804, 0.48490161940618826, 0.47503971367838976, 0.474225266023418, 0.460057073566178, 0.484955930957043, 0.4537182718943504, 0.46011776827874146, 0.5037157013530582, 0.41573127611183097, 0.42575854794463974, 0.41106009809244454, 0.43680454284553777, 0.47801890813637515, 0.39515906684994695, 0.4079883631059276, 0.43756872816061243, 0.45083493396929647, 0.4274635426962311, 0.46897746002766805, 0.4269580524683769, 0.42172898466595754, 0.44034176266892394, 0.4505840343406947, 0.4253240988151345, 0.4058941086590048, 0.39168969928097325, 0.4265219147075367, 0.44057864836064453, 0.3727412304335596, 0.4077042440658397, 0.3695207196317026, 0.4236600764122127, 0.40933736175942353, 0.40405616681967876, 0.3934353234281606, 0.3916703899740906, 0.3781587329997265, 0.3475798440504509, 0.3923555301893651, 0.35015533737540777, 0.3751782140741693, 0.3664029719731482, 0.39764384992769697, 0.3955332349193155, 0.3955334521023883, 0.3710524928924976, 0.3487070552498921, 0.3839920269700164, 0.37769099005545503, 0.3072866195619903, 0.37529462750674236, 0.3835438911945801, 0.3559786219884512, 0.3634521148072316, 0.36096491503918215, 0.3511807986486451, 0.3349316631545844, 0.3342886963029886, 0.35203569798351136, 0.342491151972554], 'target': 'A'},\n    {'timeseries_A': [0.4615033714847938, 0.430480721970956, 0.44208643685501214, 0.4302268934888613, 0.407246130251733, 0.4534614552044678, 0.4647500988403568, 0.4077374718517324, 0.4901057286480192, 0.44640404488693664, 0.4820873903652818, 0.4695438309385333, 0.4171452442392155, 0.40657530304369444, 0.46594878544637286, 0.46001107550227627, 0.4936073922006132, 0.43062082161040016, 0.4236090623502089, 0.42014379393918355, 0.4208917440287419, 0.41898763803192335, 0.39068352236484505, 0.4366474778421588, 0.42588213326885005, 0.412745585894296, 0.4255446413976233, 0.4243048141667452, 0.44089312320025237, 0.44248136576992697, 0.4162072689113627, 0.4432580556198079, 0.43909724415758267, 0.38819754419266617, 0.33822104570275835, 0.3716047551771261, 0.43012553387851854, 0.42156125208105144, 0.41779622211384837, 0.43422006716273087, 0.43167175396647467, 0.47019100382660667, 0.4061855590789723, 0.41983084954640354, 0.4239865963336714, 0.368593307909171, 0.4474135237929123, 0.3851351790685198, 0.41294130867900575, 0.3952452832048591, 0.4136972603077734, 0.4237490875310168, 0.44402267665304335, 0.4560349941646884, 0.38939836596253696, 0.4479332458273756, 0.44389034205336064, 0.43919315776857876, 0.4832592139946704, 0.4282388302497168, 0.3820059657002886, 0.4104138521040905, 0.3812100974205045, 0.38758745673391665, 0.4471273294525189, 0.37075830012528244, 0.37542360924147933, 0.40268379169670726, 0.3591545628677168, 0.36297892409050325, 0.3824134559838041, 0.3900891778490022, 0.4597496279368515, 0.37333888503866974, 0.41331153437171925, 0.4485308914491443, 0.4290545821426362, 0.39965490823744476, 0.40339336410532434, 0.3935262737570563, 0.38471404599571246, 0.37730106633856686, 0.3845799043680018, 0.4810262191436202, 0.40098866322993765, 0.3329390737350509, 0.39804883655093143, 0.41713809272826535, 0.36496622672781753, 0.4055079850315084, 0.4169923078007224, 0.4038904609887994, 0.44007445914966853, 0.3985308737390522, 0.44516805563083667, 0.37801580999689066, 0.3898193061963585, 0.41712899758227845, 0.45971520101683017, 0.404043305651106, 0.4869961231471868, 0.4385210618126883], 'timeseries_B': [0.5082494660968147, 0.52760185149098, 0.47838773309901567, 0.4721836946278241, 0.486944756110326, 0.46441316310266473, 0.49281502879890837, 0.48561054084376226, 0.4739583386559003, 0.49271434175151074, 0.5041227298633644, 0.4636240598441501, 0.4582209717451241, 0.49260343863667805, 0.4540601827717916, 0.45190832000632214, 0.4880975260225627, 0.4711430228621206, 0.4622345224839258, 0.4546221654532595, 0.5118219720217506, 0.47617030485790174, 0.47533283046418673, 0.5203244230710035, 0.4560889328265441, 0.44782774047429846, 0.44332442333931904, 0.43288392484614324, 0.4860581383381998, 0.49181882417798384, 0.45504544636937055, 0.4457903187262269, 0.4422057007375859, 0.45513150343260644, 0.4642559875661327, 0.499578184588642, 0.413611136087428, 0.44070287156589527, 0.48340680517351853, 0.41286915174099853, 0.44007039047734364, 0.4854103911637223, 0.44117086093499447, 0.4961870870493389, 0.4602946037472932, 0.4385859145659286, 0.4180126416640228, 0.46271563793378795, 0.4670333471859822, 0.46521977943892034, 0.443311098648369, 0.4340616974892226, 0.458921657452505, 0.43064892462008386, 0.46034534936834176, 0.4878953429720048, 0.4327342485959988, 0.4502245197560584, 0.4614610543122852, 0.39752605588183865, 0.3615189084673878, 0.42226654610396114, 0.39635153520187216, 0.4328746006373983, 0.41442524484767596, 0.4261671704688302, 0.39328059696005657, 0.4118038126428938, 0.43316904267631895, 0.42351180046941733, 0.4002619743463075, 0.38731815845905726, 0.3721042600590473, 0.4261937424052764, 0.4228183014013086, 0.3881210897099942, 0.39009446840809253, 0.4239572784579651, 0.4362301600487653, 0.42632499165327437, 0.3732806394796505, 0.372966121994878, 0.39896447246118094, 0.37820021355641154, 0.36590854079444596, 0.3673451474567097, 0.37583221796036326, 0.35669982286372054, 0.35074536585266164, 0.38919471409398226, 0.31754372329242747, 0.41087855484167823, 0.345072364633625, 0.34309076784808756, 0.34722161478427754, 0.40424789248276094, 0.353508439481339, 0.39950637073022693, 0.38244917446908294, 0.40667363400538137, 0.3954837357567307, 0.3598082255052257], 'target': 'A'},\n    {'timeseries_A': [0.3695193437677706, 0.5800089636747562, 0.5479263500595669, 0.3744096830544058, 0.35719547680732205, 0.3989931021524135, 0.30675724645262614, 0.36736093944798276, 0.3286390493077659, 0.238035373460873, 0.31422415437855755, 0.25026833287810196, 0.2522897798439113, 0.22443520072510906, 0.40476510646711894, 0.4424716473568172, 0.3135489153397458, 0.29310668767098363, 0.39704137031864567, 0.3700473006878307, 0.3600654354578455, 0.28875664256465694, 0.219486663494644, 0.3241041265613376, 0.39472124092875327, 0.40065302143185566, 0.3720271330570407, 0.3586718414149773, 0.40324183132995983, 0.4059140281629289, 0.3580355614947487, 0.36382599717579567, 0.14984067860008324, 0.37352363447092507, 0.2946553904935494, 0.4030895689122731, 0.40066575391230924, 0.4059063990562177, 0.43866486246637515, 0.36268473550374386, 0.2382094925565018, 0.3393993094140351, 0.3759015478382634, 0.3922744248069139, 0.21229480461250913, 0.2882614681747582, 0.4939459025784824, 0.40943804697003405, 0.18704739610036342, 0.35085281199824114, 0.2053515955338319, 0.30053596632354407, 0.3777670836077214, 0.4199592569997364, 0.3866944725207437, 0.37150976310833406, 0.39048570087940837, 0.2616730131758571, 0.2666063743112383, 0.47969925512768885, 0.4994118468821412, 0.2636007224911605, 0.33041616966755905, 0.48133358969742573, 0.38457531428114794, 0.2882245387397079, 0.29845635950294136, 0.3829511366599437, 0.3733744682225122, 0.28752371314041053, 0.4403898365272012, 0.3628563684985745, 0.2723897010964237, 0.3175166186459605, 0.3529321229929289, 0.46205461662952174, 0.43077760357364236, 0.28968694761326774, 0.24182134979857306, 0.2763041914104344, 0.40874419808204043, 0.412096181080196, 0.47400084009637333, 0.3040568424431095, 0.35295478297907074, 0.42335996739039566, 0.4173576177457297, 0.37782691932746454, 0.2772854277483117, 0.3176967047757471, 0.4033299378975606, 0.3650275632755715, 0.31148680370934656, 0.33632360090892544, 0.15576993027039682, 0.33169007194014183, 0.34363721096833166, 0.32312889509183995, 0.18809907739549062, 0.39112488085690245, 0.4917023586601301, 0.3103147662751032], 'timeseries_B': [0.532580680143779, 0.5145676661323622, 0.49351082439775007, 0.5387979379974333, 0.4965366226927702, 0.5276327232237269, 0.5560893980180144, 0.5034984543303929, 0.47309170448270665, 0.4680894917870857, 0.5211980069986586, 0.4444000620784131, 0.5035519723040783, 0.4701194223925642, 0.4311133474337341, 0.4921176692223041, 0.5023452970566123, 0.3914316009827064, 0.5206094194756454, 0.4315881488664204, 0.4680373476002116, 0.4419106363048167, 0.39692952335028736, 0.4355600075980052, 0.5098930425446738, 0.4716921032935254, 0.4305488935137656, 0.41727418581255554, 0.44638026538066433, 0.3786758020296681, 0.43541673080073506, 0.4578601150075236, 0.43142856073452474, 0.3263429374716187, 0.4095622654484418, 0.41754349834693627, 0.49719664029709126, 0.4460795327687775, 0.41869136342794233, 0.3764462780099209, 0.38290924226274925, 0.48707635054780696, 0.38304345006014945, 0.46650987904778807, 0.413166992537095, 0.45583743843808444, 0.3870875443520592, 0.3333957312847632, 0.37139622522746873, 0.4438745410530312, 0.40056758615400045, 0.39097738006126187, 0.37788540892081407, 0.4118908803143205, 0.4696783097627302, 0.3993362845988085, 0.38566404705181434, 0.4197991668365185, 0.3988967671006601, 0.2839307239541484, 0.4463319040915371, 0.4022498454452814, 0.4581405466511147, 0.42730437674886257, 0.3993511521883696, 0.3510445304938805, 0.3710259243217678, 0.41652419942049523, 0.37615917624937223, 0.3640925090895095, 0.3767475921732916, 0.38965379245466664, 0.3604338495830648, 0.39849322469071524, 0.3733341917349463, 0.4107493874294734, 0.37031038085418894, 0.32770641631474345, 0.3657719680339455, 0.3692698642257253, 0.34287241094035437, 0.3498519458266234, 0.32393176346966346, 0.37013160928177347, 0.36327401661328423, 0.3619455668410678, 0.271363367381472, 0.36141695733302376, 0.369082876878374, 0.3064132526602471, 0.30480764443115493, 0.30799652348347006, 0.2767386850111629, 0.3399303424082364, 0.36706776081021114, 0.37683074919160486, 0.3377764478288203, 0.3834174101656307, 0.3410218029451091, 0.2692658751175656, 0.2982030891320605, 0.24134906170516945], 'target': 'A'},\n    {'timeseries_A': [0.7324539392032496, 0.7315324255819039, 0.7406013306995103, 0.6919432316419213, 0.7043290173804724, 0.6840703819733793, 0.7100559547093143, 0.7411591643737254, 0.6939994247931566, 0.7267144356130599, 0.6735220198262422, 0.6941802061991229, 0.7485985844073486, 0.7147801718806396, 0.6714758901735557, 0.6827686842447191, 0.6934636389749588, 0.6870769245953977, 0.6807198957527497, 0.664243912584791, 0.7147294200104243, 0.6525046456568223, 0.6382095826550254, 0.6378802573960717, 0.6967064312081331, 0.6329731946514959, 0.6860623206397027, 0.6062906771321356, 0.6772946017019581, 0.6619426010142747, 0.6542444061570126, 0.6438299602485235, 0.6782019943668309, 0.649673736447479, 0.6062575275735259, 0.6226336457431877, 0.6710299648266628, 0.6090108527959578, 0.6470922539980055, 0.6246431324445965, 0.6261370819836367, 0.6319421617953249, 0.6593296301243167, 0.6371993357676435, 0.6244872978906821, 0.6144078696307285, 0.6760466729668273, 0.6599087573865529, 0.6145704341892096, 0.6433286967578029, 0.6348137316960929, 0.6381252008507567, 0.6669175983595702, 0.6165649713639787, 0.6287621712412056, 0.6205856973341815, 0.6145728555317658, 0.64906970134065, 0.5825544479579965, 0.6365567004404068, 0.5759981942238256, 0.5921234822686602, 0.5758981926189904, 0.6401941747003246, 0.5740132182311531, 0.6134881468379683, 0.6610538039159286, 0.5981869474622106, 0.5960918282559386, 0.5797831213686314, 0.5767533119430479, 0.5765156539810978, 0.555399489001275, 0.6096506320068641, 0.5827108495686911, 0.5702832892542127, 0.5851727496621798, 0.5590912032947655, 0.5853744376088734, 0.6102555508491666, 0.5251946955985918, 0.5909924192482054, 0.571172750762496, 0.5407846426324003, 0.586297866793387, 0.5403286441662387, 0.5759049646768166, 0.5602388555652903, 0.5518506986813858, 0.5232706399192941, 0.5682837117905885, 0.5693550213051449, 0.5198871050404484, 0.5778895081167977, 0.5416479185278099, 0.5450576864471701, 0.5483655536149011, 0.5475213984373535, 0.5551996881834433, 0.5434836547691035, 0.539910246376495, 0.5443754261504585], 'timeseries_B': [0.744296800545373, 0.7492832915467056, 0.7426809485509256, 0.7085206068932708, 0.7313996045196183, 0.7246429078508698, 0.7160342433337276, 0.7586484466251815, 0.6991151179620149, 0.7258164141009468, 0.7219767028039261, 0.6959687792743458, 0.734264522502458, 0.750345584693445, 0.7041301939328065, 0.7391988904746676, 0.7225433659450554, 0.7423814035070792, 0.7041685967568511, 0.6830812680925529, 0.6742501938426577, 0.7397506890442799, 0.6964919496781319, 0.6746827614344237, 0.6797886078725782, 0.6854415317635011, 0.6893753835233, 0.694512196340604, 0.657024750349058, 0.6855125568055386, 0.6645312907673777, 0.6969653898290525, 0.67405638038444, 0.7199011207673212, 0.68116010612748, 0.6604356318775635, 0.7026242161517433, 0.6391538141786133, 0.6415388055814865, 0.6479043384418115, 0.6410399333718384, 0.6619191485521101, 0.6776722084525457, 0.6544071630902605, 0.6426920241383717, 0.6721550672872098, 0.6448599807546069, 0.6157098333667348, 0.6486499066631287, 0.6477672138618698, 0.6423572641902929, 0.6348219134671025, 0.638607546664917, 0.6278063260339439, 0.6255440378476963, 0.6346983556282139, 0.6007719616164944, 0.6453971091180846, 0.5986394809097678, 0.6014347986553429, 0.632114504971426, 0.6219000374574222, 0.6584210628195133, 0.6412176508503323, 0.593304087893014, 0.5494225096554454, 0.6184936406902766, 0.5994587241354398, 0.6246339462897043, 0.5894706043553081, 0.5893562400701743, 0.6072519863609487, 0.6065041860351612, 0.6325832823303527, 0.5688963359824745, 0.5781813750324756, 0.6067868356213674, 0.5618826892225179, 0.6026634669740552, 0.5568236691003231, 0.5935109920714875, 0.6199912239814926, 0.5589893524844813, 0.5701988772256402, 0.5818669155047405, 0.6020634158797262, 0.5611225910392237, 0.5284808216440328, 0.5471118882378313, 0.5576504625293679, 0.545546425516719, 0.5489933497633878, 0.548423606060742, 0.5722956039828841, 0.5416767993725854, 0.5315576306786028, 0.5084196569012209, 0.5357866121194605, 0.5310648008188599, 0.5471574468589391, 0.5378799189423918, 0.5295660016732836], 'target': 'A'},\n]\nfor row in data:\n    target = row[\"target\"]\n    stability_metric_A = gini_stability(row[\"timeseries_A\"], w_fallingrate=88.0, w_resstd=-0.5)\n    stability_metric_B = gini_stability(row[\"timeseries_B\"], w_fallingrate=88.0, w_resstd=-0.5)\n    proposed_metric_A = proposed_metric(row[\"timeseries_A\"], exponent=1, ma_len=5)\n    proposed_metric_B = proposed_metric(row[\"timeseries_B\"], exponent=1, ma_len=5)\n    if target == \"A\" and stability_metric_A > stability_metric_B and proposed_metric_A < proposed_metric_B:\n        print(f\"stability_metric_A: {stability_metric_A}, stability_metric_B: {stability_metric_B}\")\n        print(f\"proposed_metric_A: {proposed_metric_A}, proposed_metric_B: {proposed_metric_B}\")\n        plt.plot(row[\"timeseries_A\"], label=\"A\")\n        plt.plot(row[\"timeseries_B\"], label=\"B\")\n        plt.ylabel(\"gini\")\n        plt.xlabel(\"week number\")\n        plt.legend()\n        plt.show()\n```",
              "votes": 1
            },
            {
              "id": 2662899,
              "postDate": "2024-02-22T07:45:39.983Z",
              "content": "<p>The problem is that you're trying to extrapolate future performance beyond the test period which requires assigning negative weighting to early rounds. This is not suited to a competition because it is both hackable and noisy. The solution is to use a longer test period so you don't need to extrapolate.</p>\n<p>As I understand it, the test period already is not strictly future data, so why not just make it longer?</p>",
              "rawMarkdown": "The problem is that you're trying to extrapolate future performance beyond the test period which requires assigning negative weighting to early rounds. This is not suited to a competition because it is both hackable and noisy. The solution is to use a longer test period so you don't need to extrapolate.\n\nAs I understand it, the test period already is not strictly future data, so why not just make it longer?"
            },
            {
              "id": 2662986,
              "postDate": "2024-02-22T09:11:07.990Z",
              "content": "<p>I understand your points and I agree with the point that it is not perfectly suited for competition. You need to understand that The metric is done in a way that it very well approximates the decisions made in-house. This has the apparent benefit of directly considering the outcomes of the competition in our business. While the metric would only consider the timeframe on which kagglers operate it would not be directly applicable internally. This can be ofc disputed, but we have reasons for doing it in this way. Since we have problems with ppl hacking the LB we consider sacrificing this property if the new metric will be sufficiently similar. So far your metric is the closest one. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699</a></p>",
              "rawMarkdown": "I understand your points and I agree with the point that it is not perfectly suited for competition. You need to understand that The metric is done in a way that it very well approximates the decisions made in-house. This has the apparent benefit of directly considering the outcomes of the competition in our business. While the metric would only consider the timeframe on which kagglers operate it would not be directly applicable internally. This can be ofc disputed, but we have reasons for doing it in this way. Since we have problems with ppl hacking the LB we consider sacrificing this property if the new metric will be sufficiently similar. So far your metric is the closest one. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699"
            },
            {
              "id": 2663102,
              "postDate": "2024-02-22T09:57:57Z",
              "content": "<p>A metric with extrapolation should be about equal to the same metric without extrapolation taken over a longer test period. That is what extrapolation achieves, it compensates for a lack of future data.</p>\n<p>Using a longer test period should accomplish the same effect that extrapolating does, but better because you're using the actual future performance instead of guessing at it.</p>\n<p>There's nothing extrapolation does for you that extending the test period would not.</p>",
              "rawMarkdown": "A metric with extrapolation should be about equal to the same metric without extrapolation taken over a longer test period. That is what extrapolation achieves, it compensates for a lack of future data.\n\nUsing a longer test period should accomplish the same effect that extrapolating does, but better because you're using the actual future performance instead of guessing at it.\n\nThere's nothing extrapolation does for you that extending the test period would not."
            },
            {
              "id": 2663109,
              "postDate": "2024-02-22T10:05:02.073Z",
              "content": "<p>I agree only partially. For the sake of this competition, it would be ideal to have very long test sample so that we don't have to deal with the extrapolation anymore. Apart from the competition, the issue is that the decision of risk managers is usually about whether we should keep the model in production or not. Hence, they need to extrapolate the near future performance. How does it translate into the competition? Given that the metric approximates well in-house decisions winning model will be the one that is most likely kept in production after deployment on test set. If we knew beforehand that the deployment will be only for the period of test set, if would make sense to use only metric that takes the performance on test set without any kind of extrapolation. That is however rarely the case. </p>",
              "rawMarkdown": "I agree only partially. For the sake of this competition, it would be ideal to have very long test sample so that we don't have to deal with the extrapolation anymore. Apart from the competition, the issue is that the decision of risk managers is usually about whether we should keep the model in production or not. Hence, they need to extrapolate the near future performance. How does it translate into the competition? Given that the metric approximates well in-house decisions winning model will be the one that is most likely kept in production after deployment on test set. If we knew beforehand that the deployment will be only for the period of test set, if would make sense to use only metric that takes the performance on test set without any kind of extrapolation. That is however rarely the case. ",
              "votes": 1
            },
            {
              "id": 2663141,
              "postDate": "2024-02-22T10:38:35.467Z",
              "content": "<p>You may have to guess at the future performance of models to decide whether to keep them running, but competition scores should not be guesses. They should be an evaluation of proven historical performance on real data. </p>\n<p>It's better to score on 6 months of real data than to score on 3 months of real data and whatever the linear fit predicts the next three months will be. If this differs from the equation you use in house it's because it's purpose differs from your needs in house. A competition needs a fair evaluation of performance without excessive randomness. This shouldn't be sacrificed just to make the evaluation period as recent as possible.</p>",
              "rawMarkdown": "You may have to guess at the future performance of models to decide whether to keep them running, but competition scores should not be guesses. They should be an evaluation of proven historical performance on real data. \n\nIt's better to score on 6 months of real data than to score on 3 months of real data and whatever the linear fit predicts the next three months will be. If this differs from the equation you use in house it's because it's purpose differs from your needs in house. A competition needs a fair evaluation of performance without excessive randomness. This shouldn't be sacrificed just to make the evaluation period as recent as possible.\n\n\n\n",
              "votes": 2
            },
            {
              "id": 2663413,
              "postDate": "2024-02-22T13:28:53.570Z",
              "content": "<p>I don't think we could completely abandon the extrapolation nature of the current stability metric. It would indeed solve the hacking issue, but on the other hand, it would deviate from the intention of this competition. I can't see how we could make a metric that would have this extrapolation property and yet at the same it would be monotonic in respect to every prediction.</p>\n<p>Perhaps there could be a compromise. We could lower the penalization that stems from the extrapolation somehow in new metric and at the same time have a metric that is monotonic in respect to majority of samples.  We don't have to use gini by weeks as long as the alternative would be similar to gini. Think about it, we are willing to compromise, but we are unlikely to remove the intended extrapolation nature of the metric. Again the metric simulates rating of models when the decision is going to be made, whether it should be left in production or not. </p>",
              "rawMarkdown": "I don't think we could completely abandon the extrapolation nature of the current stability metric. It would indeed solve the hacking issue, but on the other hand, it would deviate from the intention of this competition. I can't see how we could make a metric that would have this extrapolation property and yet at the same it would be monotonic in respect to every prediction.\n\nPerhaps there could be a compromise. We could lower the penalization that stems from the extrapolation somehow in new metric and at the same time have a metric that is monotonic in respect to majority of samples.  We don't have to use gini by weeks as long as the alternative would be similar to gini. Think about it, we are willing to compromise, but we are unlikely to remove the intended extrapolation nature of the metric. Again the metric simulates rating of models when the decision is going to be made, whether it should be left in production or not. ",
              "votes": 1
            },
            {
              "id": 2667400,
              "postDate": "2024-02-25T05:24:38.100Z",
              "content": "<p>If the intention of the competition is to test stability, using a longer test period instead of extrapolation does not deviate from that intention, it better accomplishes it. To even call the metric a \"stability metric\" is misleading because it does not so much test stability as imperfectly simulate a longer test period. The metric I gave you is more of a stability metric in that it punishes models that experience periods of poor performance in a way that cannot be replicated by extending the test period. </p>\n<p>To say that the test set has to extrapolate because fund managers extrapolate seems like cargo cultism to me: a sort of naive copying without regard for the purpose or context of the original practice. While fund managers need to choose models that look like they will do well in the near future, competition scores should be based on real performance. At the end of the competition models can be retrained on new data and then projections can be made. A model that has done well for 3 months and is projected to do well for 3 months simply has not demonstrated better performance or stability than a model that has done well for 6 months.  </p>\n<p>The dataset we have is also particularly unsuitable for extrapolation. The fact the features can disappear from the dataset creates a risk that a one time drop due to a feature leaving will be wrongly extrapolated as if it were a downward trend when no such trend is likely to materialize in reality. Obfuscating dates is also problematic because it precludes the creation of lagged time features that may otherwise improve models. And then there is of course the risk of de-obfuscation. You may think you can just disqualify any models that use such a technique, but it is always possible that features that appear legitimate also leak time information and choices that accomplish the goal of de-obfuscation are hidden behind what appear to be legitimate designs or designs that are too borderline to make a decision that isn't arbitrary. Kaggle has a long history of hosts thinking they have a bulletproof setup only to see their entire leaderboard populated by exploits.</p>\n<p>I appreciate that you are willing to make changes, but if you are only willing change things unrelated to the problem nothing can be accomplished.</p>",
              "rawMarkdown": "If the intention of the competition is to test stability, using a longer test period instead of extrapolation does not deviate from that intention, it better accomplishes it. To even call the metric a \"stability metric\" is misleading because it does not so much test stability as imperfectly simulate a longer test period. The metric I gave you is more of a stability metric in that it punishes models that experience periods of poor performance in a way that cannot be replicated by extending the test period. \n\nTo say that the test set has to extrapolate because fund managers extrapolate seems like cargo cultism to me: a sort of naive copying without regard for the purpose or context of the original practice. While fund managers need to choose models that look like they will do well in the near future, competition scores should be based on real performance. At the end of the competition models can be retrained on new data and then projections can be made. A model that has done well for 3 months and is projected to do well for 3 months simply has not demonstrated better performance or stability than a model that has done well for 6 months.  \n\nThe dataset we have is also particularly unsuitable for extrapolation. The fact the features can disappear from the dataset creates a risk that a one time drop due to a feature leaving will be wrongly extrapolated as if it were a downward trend when no such trend is likely to materialize in reality. Obfuscating dates is also problematic because it precludes the creation of lagged time features that may otherwise improve models. And then there is of course the risk of de-obfuscation. You may think you can just disqualify any models that use such a technique, but it is always possible that features that appear legitimate also leak time information and choices that accomplish the goal of de-obfuscation are hidden behind what appear to be legitimate designs or designs that are too borderline to make a decision that isn't arbitrary. Kaggle has a long history of hosts thinking they have a bulletproof setup only to see their entire leaderboard populated by exploits.\n\nI appreciate that you are willing to make changes, but if you are only willing change things unrelated to the problem nothing can be accomplished.",
              "votes": 1
            },
            {
              "id": 2750951,
              "postDate": "2024-04-13T23:18:07.583Z",
              "content": "<p>This is worth looking at…</p>",
              "rawMarkdown": "This is worth looking at..."
            }
          ]
        }
      ]
    },
    {
      "id": 2652189,
      "postDate": "2024-02-14T15:57:05.190Z",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <br>\nThe dates in the future are errors?</p>\n<p>Or is it one of the following scenarios:</p>\n<ul>\n<li>The dates are correct for various reasons even if they are in the future, as they pertain to an event that has actually occurred.</li>\n<li>The dates are inconsistent because they result from calculations/extractions that sometimes lead to inconsistencies.</li>\n</ul>\n<p>Ty</p>",
      "rawMarkdown": "@jetakow \nThe dates in the future are errors?\n\nOr is it one of the following scenarios:\n\n- The dates are correct for various reasons even if they are in the future, as they pertain to an event that has actually occurred.\n- The dates are inconsistent because they result from calculations/extractions that sometimes lead to inconsistencies.\n\nTy",
      "votes": 1,
      "replies": [
        {
          "id": 2653249,
          "postDate": "2024-02-15T10:12:01.880Z",
          "content": "<p>We believe the majority will be the second case AND that we transformed all dates and there might be some inconsistencies as a result of that. There might be a case that the dates are from future in a sense that they indicate future date of something. Even if that is the case we don't think it will bring any significant information value.</p>",
          "rawMarkdown": "We believe the majority will be the second case AND that we transformed all dates and there might be some inconsistencies as a result of that. There might be a case that the dates are from future in a sense that they indicate future date of something. Even if that is the case we don't think it will bring any significant information value.",
          "votes": 2,
          "replies": [
            {
              "id": 2671059,
              "postDate": "2024-02-27T09:12:08.220Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <br>\nActually transforming the dates will mess up a lot of features. For example, changing decision date will change the correct age of the applicant at the time of decision,  which is a big discrepancy compared to train data. The same applies to FE around covid periods where people lost wages and were stood down…it will incorrectly apply these features to test data if the dates are altered.</p>",
              "rawMarkdown": "@jetakow \nActually transforming the dates will mess up a lot of features. For example, changing decision date will change the correct age of the applicant at the time of decision,  which is a big discrepancy compared to train data. The same applies to FE around covid periods where people lost wages and were stood down...it will incorrectly apply these features to test data if the dates are altered."
            },
            {
              "id": 2671256,
              "postDate": "2024-02-27T12:15:37.727Z",
              "content": "<p>I understand you concern and I agree that it will render a lot of features useless. We initially wanted to leave this information, but as you can see it was mainly exploited for metric hacking. Since the gain from metric hacking is about the same size as basic FE, we have to remove this information from dataset. It is still not certain, but very likely as of now. Kagglers will be given two additional weeks of time and they can use the time now to inspect train set further. Which is in my opinion very generous. The worst case scenario would be to leave it as it is now and just make the hack accessible to everyone. </p>",
              "rawMarkdown": "I understand you concern and I agree that it will render a lot of features useless. We initially wanted to leave this information, but as you can see it was mainly exploited for metric hacking. Since the gain from metric hacking is about the same size as basic FE, we have to remove this information from dataset. It is still not certain, but very likely as of now. Kagglers will be given two additional weeks of time and they can use the time now to inspect train set further. Which is in my opinion very generous. The worst case scenario would be to leave it as it is now and just make the hack accessible to everyone. ",
              "votes": 1
            },
            {
              "id": 2672017,
              "postDate": "2024-02-27T20:47:27.503Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> OK will you please at least publish the dates that are no longer representative of the actual dates and whether at least the year or the month are correct? so for example if the concern is predicting WEEK_NUM then if the modified date is still within +/- 2 months of the correct date we can keep some features like age and COVID related behaviour. I am not asking to disclose the new transformation but just to indicate the fields that got new transformation and a rough indication of +/- N moths from the actual date.</p>",
              "rawMarkdown": "@jetakow OK will you please at least publish the dates that are no longer representative of the actual dates and whether at least the year or the month are correct? so for example if the concern is predicting WEEK_NUM then if the modified date is still within +/- 2 months of the correct date we can keep some features like age and COVID related behaviour. I am not asking to disclose the new transformation but just to indicate the fields that got new transformation and a rough indication of +/- N moths from the actual date."
            },
            {
              "id": 2672103,
              "postDate": "2024-02-27T22:00:44.860Z",
              "content": "<p><a href=\"https://www.kaggle.com/fadylabib\" target=\"_blank\">@fadylabib</a> Thanks for asking, but no we will not do that. The point of transformation is to make it impossible to link to dates which there are now. So providing you this information would be the same as enabling you to exploit the metric hack again. Again, we know that by doing this we reduce the possible features that can be created. Also, please note that in production, the scoring is usually time-independent in that it doesn't matter at which time the scoring is done in the future. </p>",
              "rawMarkdown": "@fadylabib Thanks for asking, but no we will not do that. The point of transformation is to make it impossible to link to dates which there are now. So providing you this information would be the same as enabling you to exploit the metric hack again. Again, we know that by doing this we reduce the possible features that can be created. Also, please note that in production, the scoring is usually time-independent in that it doesn't matter at which time the scoring is done in the future. "
            },
            {
              "id": 2672150,
              "postDate": "2024-02-28T00:23:49.490Z",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>  thank you. Will you be publishing the test columns that are having a different transformation from the train set or would this be done without explicitly saying which columns have different transforms ?</p>\n<p>For this part \"please note that in production, the scoring is usually time-independent\". I would disagree (I already mentioned that age is a factor and age will depend on the date e.g. if someone has / has not turned 18 even if the difference is just a few months) . Anyway, I understand that we will not know the dates, so please at least let us know which columns are no longer reliable.</p>",
              "rawMarkdown": "@jetakow  thank you. Will you be publishing the test columns that are having a different transformation from the train set or would this be done without explicitly saying which columns have different transforms ?\n\nFor this part \"please note that in production, the scoring is usually time-independent\". I would disagree (I already mentioned that age is a factor and age will depend on the date e.g. if someone has / has not turned 18 even if the difference is just a few months) . Anyway, I understand that we will not know the dates, so please at least let us know which columns are no longer reliable.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2652301,
      "postDate": "2024-02-14T17:09:15.957Z",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> congrats to you for the expert title. Thanks for being so proactive!</p>",
      "rawMarkdown": "@jetakow congrats to you for the expert title. Thanks for being so proactive!",
      "votes": 2,
      "replies": [
        {
          "id": 2653245,
          "postDate": "2024-02-15T10:04:50.050Z",
          "content": "<p>Thanks, I would never guess a year ago that I will get this first title thanks to kaggle competition :)</p>",
          "rawMarkdown": "Thanks, I would never guess a year ago that I will get this first title thanks to kaggle competition :)",
          "replies": [
            {
              "id": 2821232,
              "postDate": "2024-05-18T02:50:51.567Z",
              "content": "<p>Congrats Ravi</p>",
              "rawMarkdown": "Congrats Ravi"
            }
          ]
        }
      ]
    },
    {
      "id": 2671055,
      "postDate": "2024-02-27T09:08:19.783Z",
      "content": "<p>Regarding this point:  <strong>The transformation was done to train and test before the start of the competition and now we will transform test data in different manner</strong>. </p>\n<p>This will create a big discrepancy between the two sets…for example if a solution is trying to do FE based on covid lock down periods (people were stood down, salaries were cut by 20%….etc). Having the dates processed differently in the test data will render all these creative FE ideas useless and they won't apply to test data. They will actually perform poorly if the dates were transformed.</p>\n<p>May I suggest that the we keep the focus on choosing the new metric that would avoid stability hacks without changing the test data. This proposed change is a big change that will make a lot of our efforts useless.</p>",
      "rawMarkdown": "Regarding this point:  **The transformation was done to train and test before the start of the competition and now we will transform test data in different manner**. \n\nThis will create a big discrepancy between the two sets...for example if a solution is trying to do FE based on covid lock down periods (people were stood down, salaries were cut by 20%....etc). Having the dates processed differently in the test data will render all these creative FE ideas useless and they won't apply to test data. They will actually perform poorly if the dates were transformed.\n\nMay I suggest that the we keep the focus on choosing the new metric that would avoid stability hacks without changing the test data. This proposed change is a big change that will make a lot of our efforts useless.\n",
      "replies": [
        {
          "id": 2671250,
          "postDate": "2024-02-27T12:09:08.233Z",
          "content": "<p>I think there is no reason to panic. There has to be done some transformation so the test data are not the same as it is now. It should not change meaning of those features, so unless the FE is overfitted to train set or to LB, it won't be useless.</p>",
          "rawMarkdown": "I think there is no reason to panic. There has to be done some transformation so the test data are not the same as it is now. It should not change meaning of those features, so unless the FE is overfitted to train set or to LB, it won't be useless.",
          "votes": -1
        }
      ]
    },
    {
      "id": 2840141,
      "postDate": "2024-05-28T01:45:29.843Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2838965,
      "postDate": "2024-05-27T09:58:39.157Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2809713,
      "postDate": "2024-05-13T00:10:56.107Z",
      "rawMarkdown": "",
      "votes": 6,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2652661,
      "author_name": "Jacoby Jaeger",
      "author_url": "",
      "post_date": "2024-02-14T23:10:47.477000",
      "content": "<blockquote>\n  <p>Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk manager in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic.</p>\n</blockquote>\n<p>The falling rate was ill conceived. There should be no scenario where a model that has consistently poor performance is deemed better than a model that sometimes has good performance and sometimes poor performance. To use such a metric is to misunderstand what the value of stability is. That is, the value of stability is never in being consistently bad, but in avoiding periods that are particularly bad. Once this is understood, it becomes obvious that the best way to encourage stability is by aggressively punishing periods of particularly bad performance. We can do this with a metric that rapidly become worse when very bad samples are present. As I suggested in the other thread, this can be done with a metric of the form: </p>\n<p>mean(  ( -log( MovingAverage(gini) ) )**exponent   ) with exponent&gt;=1</p>\n<p>Under this metric, periods of poor performance are heavily penalized without the silliness of punishing good rounds. Stability should require nothing more than this.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2653310,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-15T10:49:34.977000",
          "content": "<blockquote>\n  <p>The falling rate was ill conceived. There should be no scenario where a model that has consistently poor performance is deemed better than a model that sometimes has good performance and sometimes poor performance. To use such a metric is to misunderstand what the value of stability is. </p>\n</blockquote>\n<p>Frankly, I think there is also some misunderstanding on your side. Consider this. From the perspective of risk managers at HC it is undesirable to have a model that is performing well let's say first 6 months and then it levels on for example 60%-80% of that performance. Even if the model performed better in terms of integral of weekly gini over time the model was unpredictable in sense that its performance lowered after the first 6 months. What is desirable is a model that has the same exact performance every week and if there is some decrease it is well predictable. </p>\n<p>Omitting the falling rate term will allow models that will be performing well on some portion of test sample and then they will have decrease in performance that might not be even linear of exponential. </p>\n<p>Regarding proposed metric, you could try to rewrite it using LaTeX with proper sums so it is clear what is applied to what. As it is now I am not sure if I understand it right. </p>\n<p>Regarding the moving average idea, I believe that this might lower the amount of hacking, but it will not solve the problem completely. Consider following model that is performing well the first N weeks and during next 3 weeks the performance drops to 60% of that. By lowering the score artificially on the first N weeks you will improve the stability metric. That is aligned with risk managers' view of performance. It is not aligned with metric being fair and fun for everyone, we are working on that and I ask you for patience. The score from the metric that you suggested will improve when decreasing the score artificially first N weeks. If this is not the case, then I don't understand it perhaps correctly and you could explain it in finer detail. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2653528,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2024-02-15T13:58:10.207000",
              "content": "<p>I don’t think it will improve by lowering the score in first few weeks. Every lowering of gini score will result in a lower moving average. If moving average gets smaller, -log will get larger. The best performance is reached when all ginis in all relevant time windows are 1, then the above metric is 0 as log(1)=0.<br>\nThat badscores are punished heavily is also clear as the log decreases rapidly for arguments smaller than 1 (and even more so the closer we get to zero.)<br>\nPlease correct if I’m wrong <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a> </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2654189,
              "author_name": "Jacoby Jaeger",
              "author_url": "",
              "post_date": "2024-02-16T00:21:26.750000",
              "content": "<blockquote>\n  <p>Frankly, I think there is also some misunderstanding on your side. Consider this. From the perspective of risk managers at HC it is undesirable to have a model that is performing well let's say first 6 months and then it levels on for example 60%-80% of that performance. </p>\n</blockquote>\n<p>What you need is not to be able to predict the performance but to be able to put a good lower bound on what the performance could be. This is because the event of the model performing better than you might have prepared for is never a problem and as such does not need to be penalized. The set of models that can be given a high lower bound on their expected performance are those that have demonstrated a high lower bound during the validation period, which are those that perform well under a metric that heavily penalizes a given model's worst rounds as mine does.</p>\n<blockquote>\n  <p>Regarding proposed metric, you could try to rewrite it using LaTeX with proper sums so it is clear what is applied to what. As it is now I am not sure if I understand it right.</p>\n</blockquote>\n<p>Once predictions have been converted to ginis, there is only one axis that we need to be concerned with, the time axis. Here is a simple implementation of the metric. Higher scores are considered worse under this metric. </p>\n<pre><code> numpy  np\n matplotlib  pyplot  plt\n\n\n ():\n    (axis==)\n\n    x = np.cumsum( x, axis= )\n\n    x = np.concatenate([ *x[:], x ],)\n\n     (x[n:] - x[:-n])/n\n\n\n ():\n    \n\n     ():\n\n        scores = ( - np.log(  MA( ginis, ma_len, axis= ) ) )**exponent\n\n          scores\n\n     metric \n\n\n\nN_GINIS = \nN_MA = \nEXPONENT = \n\nginis = np.random.rand(N_GINIS)\n\ntime = np.arange(N_GINIS)\n\n\nmetric = get_metric(EXPONENT, N_MA)\n\n\nscores = metric(ginis)\n\n\nplt.plot(time, ginis)\nplt.plot(time[N_MA//: N_GINIS-(N_MA-)//],  scores)\n\n\nscore = np.mean(scores)\nscore\n</code></pre>\n<blockquote>\n  <p>Omitting the falling rate term will allow models that will be performing well on some portion of test sample and then they will have decrease in performance that might not be even linear of exponential.</p>\n</blockquote>\n<p>Any period of poor performance is heavily punished under my metric. But the falling rate can only detect periods of poor performance if they come at the end of the test period.</p>\n<blockquote>\n  <p>Regarding the moving average idea, I believe that this might lower the amount of hacking, but it will not solve the problem completely. Consider following model that is performing well the first N weeks and during next 3 weeks the performance drops to 60% of that. By lowering the score artificially on the first N weeks you will improve the stability metric. That is aligned with risk managers' view of performance.</p>\n</blockquote>\n<p>I'm not sure what you are trying to say here. Are you suggesting that performance on my metric would be improved by such hacking? That is not the case, it can be seen from the fact that the gradient wrt sample ginis is strictly negative that making any round's performance worse can only hurt performance under my metric. </p>\n<p>Or are you saying that it is desirable that such hacking results in a higher score? That would be rather silly, you can make a declining performance look flat during the test period by making the early test rounds artificially worse, but that does not change the fact that true performance was declining which will become evident after the test period when you no longer have artificially deflated rounds left to compare the performance to. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2654192,
              "author_name": "Jacoby Jaeger",
              "author_url": "",
              "post_date": "2024-02-16T00:26:12.940000",
              "content": "<p>You're exactly right, Simon. We can check the sign of the gradient of the final score wrt the per round scores to prove that it is never desirable to make any round score worse. I suggest that kaggle should perform such a check on every competition metric so that we don't see hackable metrics again in the future.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2654518,
              "author_name": "Tomas Jelinek",
              "author_url": "",
              "post_date": "2024-02-16T07:24:44.077000",
              "content": "<blockquote>\n  <p>What you need is not to be able to predict the performance but to be able to put a good lower bound on what the performance could be. This is because the event of the model performing better than you might have prepared for is never a problem and as such does not need to be penalized. </p>\n</blockquote>\n<p>From business point of view this statement is not true. It's not only about the credit risk model, but whole underwriting setup, product parameters, profitability models and many other things. We believe we know why we defined stability as we did.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2654536,
              "author_name": "Jacoby Jaeger",
              "author_url": "",
              "post_date": "2024-02-16T07:47:46.097000",
              "content": "<p>In that case if a model performs \"too\" well in a given period, what is the negative consequence? What is the problem with it being \"too profitable\" one month?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2654853,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-16T13:36:07.927000",
              "content": "<p>The problem is not being too profitable, the problem is being unpredictable in production. What we want is to get a model that this predictable in performance for as long period as possible and ideally stable in time (in terms of performance). The score of a model is processed further with different processes. In other words, the score is not used directly to approve loans. </p>\n<p>Regarding your previous post, I will test the proposed metric and reply to that later, I am working on the announced fix. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2654915,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-16T14:52:13.080000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2662077,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-21T17:28:55.827000",
              "content": "<p>Okay <a href=\"https://www.kaggle.com/jacobyjaeger\" target=\"_blank\">@jacobyjaeger</a> I am back after testing it. I am here to provide you few examples where the current stability metric performs as we want it to perform, but your proposed metric does not select the correct option. Here are 5 pairs of simulated performance and A is always the preferred choice. The stability metric correctly says that the A is more preferred (with higher value of metric) and your metric says the opposite. Please have a look at this.</p>\n<pre><code> numpy  np\n matplotlib  pyplot  plt\n\n ():\n    x = np.arange((gini_in_time))\n    y = gini_in_time\n    a, b = np.polyfit(x, y, )\n    y_hat = a*x + b\n    residuals = y - y_hat\n    res_std = np.std(residuals)\n    avg_gini = np.mean(gini_in_time)\n     avg_gini + w_fallingrate * (, a) + w_resstd * res_std\n\n ():\n    x = np.cumsum(gini_in_time, axis=)\n    x = np.concatenate([*x[:], x], )\n    scores = -np.mean(-np.log(np.maximum((x[ma_len:] - x[:-ma_len])/ma_len, ))**exponent)\n     scores \n\n\n(proposed_metric([, , , , ]), proposed_metric([, , , , ]))\n\n\n\ndata = [\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n    {: [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ], : },\n]\n row  data:\n    target = row[]\n    stability_metric_A = gini_stability(row[], w_fallingrate=, w_resstd=-)\n    stability_metric_B = gini_stability(row[], w_fallingrate=, w_resstd=-)\n    proposed_metric_A = proposed_metric(row[], exponent=, ma_len=)\n    proposed_metric_B = proposed_metric(row[], exponent=, ma_len=)\n     target ==   stability_metric_A &gt; stability_metric_B  proposed_metric_A &lt; proposed_metric_B:\n        ()\n        ()\n        plt.plot(row[], label=)\n        plt.plot(row[], label=)\n        plt.ylabel()\n        plt.xlabel()\n        plt.legend()\n        plt.show()\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2662899,
              "author_name": "Jacoby Jaeger",
              "author_url": "",
              "post_date": "2024-02-22T07:45:39.983000",
              "content": "<p>The problem is that you're trying to extrapolate future performance beyond the test period which requires assigning negative weighting to early rounds. This is not suited to a competition because it is both hackable and noisy. The solution is to use a longer test period so you don't need to extrapolate.</p>\n<p>As I understand it, the test period already is not strictly future data, so why not just make it longer?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2662986,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-22T09:11:07.990000",
              "content": "<p>I understand your points and I agree with the point that it is not perfectly suited for competition. You need to understand that The metric is done in a way that it very well approximates the decisions made in-house. This has the apparent benefit of directly considering the outcomes of the competition in our business. While the metric would only consider the timeframe on which kagglers operate it would not be directly applicable internally. This can be ofc disputed, but we have reasons for doing it in this way. Since we have problems with ppl hacking the LB we consider sacrificing this property if the new metric will be sufficiently similar. So far your metric is the closest one. <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478699</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2663102,
              "author_name": "Jacoby Jaeger",
              "author_url": "",
              "post_date": "2024-02-22T09:57:57",
              "content": "<p>A metric with extrapolation should be about equal to the same metric without extrapolation taken over a longer test period. That is what extrapolation achieves, it compensates for a lack of future data.</p>\n<p>Using a longer test period should accomplish the same effect that extrapolating does, but better because you're using the actual future performance instead of guessing at it.</p>\n<p>There's nothing extrapolation does for you that extending the test period would not.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2663109,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-22T10:05:02.073000",
              "content": "<p>I agree only partially. For the sake of this competition, it would be ideal to have very long test sample so that we don't have to deal with the extrapolation anymore. Apart from the competition, the issue is that the decision of risk managers is usually about whether we should keep the model in production or not. Hence, they need to extrapolate the near future performance. How does it translate into the competition? Given that the metric approximates well in-house decisions winning model will be the one that is most likely kept in production after deployment on test set. If we knew beforehand that the deployment will be only for the period of test set, if would make sense to use only metric that takes the performance on test set without any kind of extrapolation. That is however rarely the case. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2663141,
              "author_name": "Jacoby Jaeger",
              "author_url": "",
              "post_date": "2024-02-22T10:38:35.467000",
              "content": "<p>You may have to guess at the future performance of models to decide whether to keep them running, but competition scores should not be guesses. They should be an evaluation of proven historical performance on real data. </p>\n<p>It's better to score on 6 months of real data than to score on 3 months of real data and whatever the linear fit predicts the next three months will be. If this differs from the equation you use in house it's because it's purpose differs from your needs in house. A competition needs a fair evaluation of performance without excessive randomness. This shouldn't be sacrificed just to make the evaluation period as recent as possible.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2663413,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-22T13:28:53.570000",
              "content": "<p>I don't think we could completely abandon the extrapolation nature of the current stability metric. It would indeed solve the hacking issue, but on the other hand, it would deviate from the intention of this competition. I can't see how we could make a metric that would have this extrapolation property and yet at the same it would be monotonic in respect to every prediction.</p>\n<p>Perhaps there could be a compromise. We could lower the penalization that stems from the extrapolation somehow in new metric and at the same time have a metric that is monotonic in respect to majority of samples.  We don't have to use gini by weeks as long as the alternative would be similar to gini. Think about it, we are willing to compromise, but we are unlikely to remove the intended extrapolation nature of the metric. Again the metric simulates rating of models when the decision is going to be made, whether it should be left in production or not. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2667400,
              "author_name": "Jacoby Jaeger",
              "author_url": "",
              "post_date": "2024-02-25T05:24:38.100000",
              "content": "<p>If the intention of the competition is to test stability, using a longer test period instead of extrapolation does not deviate from that intention, it better accomplishes it. To even call the metric a \"stability metric\" is misleading because it does not so much test stability as imperfectly simulate a longer test period. The metric I gave you is more of a stability metric in that it punishes models that experience periods of poor performance in a way that cannot be replicated by extending the test period. </p>\n<p>To say that the test set has to extrapolate because fund managers extrapolate seems like cargo cultism to me: a sort of naive copying without regard for the purpose or context of the original practice. While fund managers need to choose models that look like they will do well in the near future, competition scores should be based on real performance. At the end of the competition models can be retrained on new data and then projections can be made. A model that has done well for 3 months and is projected to do well for 3 months simply has not demonstrated better performance or stability than a model that has done well for 6 months.  </p>\n<p>The dataset we have is also particularly unsuitable for extrapolation. The fact the features can disappear from the dataset creates a risk that a one time drop due to a feature leaving will be wrongly extrapolated as if it were a downward trend when no such trend is likely to materialize in reality. Obfuscating dates is also problematic because it precludes the creation of lagged time features that may otherwise improve models. And then there is of course the risk of de-obfuscation. You may think you can just disqualify any models that use such a technique, but it is always possible that features that appear legitimate also leak time information and choices that accomplish the goal of de-obfuscation are hidden behind what appear to be legitimate designs or designs that are too borderline to make a decision that isn't arbitrary. Kaggle has a long history of hosts thinking they have a bulletproof setup only to see their entire leaderboard populated by exploits.</p>\n<p>I appreciate that you are willing to make changes, but if you are only willing change things unrelated to the problem nothing can be accomplished.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2750951,
              "author_name": "Kabir Olawale Mohammed",
              "author_url": "",
              "post_date": "2024-04-13T23:18:07.583000",
              "content": "<p>This is worth looking at…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2652189,
      "author_name": "Davide Stenner",
      "author_url": "",
      "post_date": "2024-02-14T15:57:05.190000",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <br>\nThe dates in the future are errors?</p>\n<p>Or is it one of the following scenarios:</p>\n<ul>\n<li>The dates are correct for various reasons even if they are in the future, as they pertain to an event that has actually occurred.</li>\n<li>The dates are inconsistent because they result from calculations/extractions that sometimes lead to inconsistencies.</li>\n</ul>\n<p>Ty</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2653249,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-15T10:12:01.880000",
          "content": "<p>We believe the majority will be the second case AND that we transformed all dates and there might be some inconsistencies as a result of that. There might be a case that the dates are from future in a sense that they indicate future date of something. Even if that is the case we don't think it will bring any significant information value.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2671059,
              "author_name": "ARMADA",
              "author_url": "",
              "post_date": "2024-02-27T09:12:08.220000",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> <br>\nActually transforming the dates will mess up a lot of features. For example, changing decision date will change the correct age of the applicant at the time of decision,  which is a big discrepancy compared to train data. The same applies to FE around covid periods where people lost wages and were stood down…it will incorrectly apply these features to test data if the dates are altered.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2671256,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-27T12:15:37.727000",
              "content": "<p>I understand you concern and I agree that it will render a lot of features useless. We initially wanted to leave this information, but as you can see it was mainly exploited for metric hacking. Since the gain from metric hacking is about the same size as basic FE, we have to remove this information from dataset. It is still not certain, but very likely as of now. Kagglers will be given two additional weeks of time and they can use the time now to inspect train set further. Which is in my opinion very generous. The worst case scenario would be to leave it as it is now and just make the hack accessible to everyone. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2672017,
              "author_name": "ARMADA",
              "author_url": "",
              "post_date": "2024-02-27T20:47:27.503000",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> OK will you please at least publish the dates that are no longer representative of the actual dates and whether at least the year or the month are correct? so for example if the concern is predicting WEEK_NUM then if the modified date is still within +/- 2 months of the correct date we can keep some features like age and COVID related behaviour. I am not asking to disclose the new transformation but just to indicate the fields that got new transformation and a rough indication of +/- N moths from the actual date.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2672103,
              "author_name": "Daniel Herman",
              "author_url": "",
              "post_date": "2024-02-27T22:00:44.860000",
              "content": "<p><a href=\"https://www.kaggle.com/fadylabib\" target=\"_blank\">@fadylabib</a> Thanks for asking, but no we will not do that. The point of transformation is to make it impossible to link to dates which there are now. So providing you this information would be the same as enabling you to exploit the metric hack again. Again, we know that by doing this we reduce the possible features that can be created. Also, please note that in production, the scoring is usually time-independent in that it doesn't matter at which time the scoring is done in the future. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2672150,
              "author_name": "ARMADA",
              "author_url": "",
              "post_date": "2024-02-28T00:23:49.490000",
              "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>  thank you. Will you be publishing the test columns that are having a different transformation from the train set or would this be done without explicitly saying which columns have different transforms ?</p>\n<p>For this part \"please note that in production, the scoring is usually time-independent\". I would disagree (I already mentioned that age is a factor and age will depend on the date e.g. if someone has / has not turned 18 even if the difference is just a few months) . Anyway, I understand that we will not know the dates, so please at least let us know which columns are no longer reliable.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2652301,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-02-14T17:09:15.957000",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> congrats to you for the expert title. Thanks for being so proactive!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2653245,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-15T10:04:50.050000",
          "content": "<p>Thanks, I would never guess a year ago that I will get this first title thanks to kaggle competition :)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2821232,
              "author_name": "vinothkarunakaran",
              "author_url": "",
              "post_date": "2024-05-18T02:50:51.567000",
              "content": "<p>Congrats Ravi</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2671055,
      "author_name": "ARMADA",
      "author_url": "",
      "post_date": "2024-02-27T09:08:19.783000",
      "content": "<p>Regarding this point:  <strong>The transformation was done to train and test before the start of the competition and now we will transform test data in different manner</strong>. </p>\n<p>This will create a big discrepancy between the two sets…for example if a solution is trying to do FE based on covid lock down periods (people were stood down, salaries were cut by 20%….etc). Having the dates processed differently in the test data will render all these creative FE ideas useless and they won't apply to test data. They will actually perform poorly if the dates were transformed.</p>\n<p>May I suggest that the we keep the focus on choosing the new metric that would avoid stability hacks without changing the test data. This proposed change is a big change that will make a lot of our efforts useless.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2671250,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-27T12:09:08.233000",
          "content": "<p>I think there is no reason to panic. There has to be done some transformation so the test data are not the same as it is now. It should not change meaning of those features, so unless the FE is overfitted to train set or to LB, it won't be useless.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2840141,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-28T01:45:29.843000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2838965,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-27T09:58:39.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2809713,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-13T00:10:56.107000",
      "content": "",
      "votes": 6,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2652153": "**Q: Why can't I submit a solution in my country?**\nA: We disabled the submissions as a temporary measure. It will be enabled again around 6.3.2024. https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/478716\n\n**Q: Questions regarding transformation of test data.**\nA: The transformation was done to train and test before the start of the competition and now we will transform test data in different manner. We won't specify what is the transformation exactly. We will keep column names, dtypes and the change is not going to be drastic. \n\n**Q: Any question related to private test set**\nA: We will not disclose any information about private test set. There is no point in asking questions about the private test set. We made the split in such a way that it will be fair for final evaluation.\n\n**Q: How long is the test set?**\nA: For the sake of stability evaluation it would not make sense to disclose you that information. This information would lead even to more \"hacks\" of the metric. Think about it this way. The model will be deployed and you don't know for how long it is going to be.\n\n**Q: What is the meaning of num_group1 and num_group2?**\nA: Those are indices that will help you to aggregate tables with depth=1,2. Please have a look at Data section.\n\n**Q: Any questions regarding stability metric and hacking.**\nPlease read https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476867 first. We are working on a solution that would minimize such hacking by artificially worsening the score in the near future. \n\n**Q: Can't we just remove the falling rate from metric?**\nA:  Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic.\n\n**Q: What is the reason for having those weights in the stability metric?**\nA: Those particular weights were selected so that they would perfectly represent the average decision of risk managers in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models.\n\n**Q: Why is my notebook failing?**\nA: We don't have capacity to review every notebook, unfortunately. You can post it in discussion and maybe someone will help. In future, we can collect those errors into this thread. \n\n**Q: Dates are inconsistent and I see dates from the future in respect to date decision.**\nA: This might happen as all dates were transformed, but there might be some inconsistencies. Also, the underlying data from our primary sources can be flawed, sometimes it happens. \n\n**Q: What is the definition of the used target?**\nA: It's unpaid payment (one is enough) in certain time period. There is also some time tolerance (e.g. one day late is not default) and amount tolerance (if client paid $100 instead of $100.10). We can't disclose further details. ",
    "2652661": ">Omitting the falling rate would not be a satisfactory solution. Without the falling rate, we would greatly encourage unstable models in time. Those particular weights were selected so that they would perfectly represent the average decision of risk manager in HC. That is desirable as we will be able to use the knowledge gained from winners and compare it with our processes and models. Removing the falling rate would remove the stability component in time. This year we wanted to explore stability as the main topic.\n\nThe falling rate was ill conceived. There should be no scenario where a model that has consistently poor performance is deemed better than a model that sometimes has good performance and sometimes poor performance. To use such a metric is to misunderstand what the value of stability is. That is, the value of stability is never in being consistently bad, but in avoiding periods that are particularly bad. Once this is understood, it becomes obvious that the best way to encourage stability is by aggressively punishing periods of particularly bad performance. We can do this with a metric that rapidly become worse when very bad samples are present. As I suggested in the other thread, this can be done with a metric of the form: \n\nmean(  ( -log( MovingAverage(gini) ) )**exponent   ) with exponent>=1\n\nUnder this metric, periods of poor performance are heavily penalized without the silliness of punishing good rounds. Stability should require nothing more than this.",
    "2652189": "@jetakow \nThe dates in the future are errors?\n\nOr is it one of the following scenarios:\n\n- The dates are correct for various reasons even if they are in the future, as they pertain to an event that has actually occurred.\n- The dates are inconsistent because they result from calculations/extractions that sometimes lead to inconsistencies.\n\nTy",
    "2652301": "@jetakow congrats to you for the expert title. Thanks for being so proactive!",
    "2671055": "Regarding this point:  **The transformation was done to train and test before the start of the competition and now we will transform test data in different manner**. \n\nThis will create a big discrepancy between the two sets...for example if a solution is trying to do FE based on covid lock down periods (people were stood down, salaries were cut by 20%....etc). Having the dates processed differently in the test data will render all these creative FE ideas useless and they won't apply to test data. They will actually perform poorly if the dates were transformed.\n\nMay I suggest that the we keep the focus on choosing the new metric that would avoid stability hacks without changing the test data. This proposed change is a big change that will make a lot of our efforts useless.\n",
    "2840141": "",
    "2838965": "",
    "2809713": ""
  }
}