{
  "id": 498038,
  "title": "What about stability?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/498038",
  "author_name": "",
  "post_date": "2024-04-26T17:19:08.538692800Z",
  "votes": 18,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I've seen a lot of (legitimate) discussions about metric hacking, but surprisingly I have not seen many kagglers discussing about the actual topic of the competition: how to tackle the stability issue. <br>\nThe major problem I'm facing is the difficulty to elaborate a CV strategy to check for improvements / performances of my models for what concerns stability: even if I train a LGBM with only the first 30 WEEK_NUM, I see that the penalty term \"a\" is positive if computed on the remaining part of the data! This means that there's no decrease of performance and that I'm not able to see what works and what does not work in the local CV…I can just submit and infer it from the score I get back.<br>\nDid you find any idea on how to check for stability improvements locally?</p>",
  "messages": [
    {
      "id": "2777486",
      "postDate": "04/26/2024 17:19:08",
      "content": "<p>I've seen a lot of (legitimate) discussions about metric hacking, but surprisingly I have not seen many kagglers discussing about the actual topic of the competition: how to tackle the stability issue. <br>\nThe major problem I'm facing is the difficulty to elaborate a CV strategy to check for improvements / performances of my models for what concerns stability: even if I train a LGBM with only the first 30 WEEK_NUM, I see that the penalty term \"a\" is positive if computed on the remaining part of the data! This means that there's no decrease of performance and that I'm not able to see what works and what does not work in the local CV…I can just submit and infer it from the score I get back.<br>\nDid you find any idea on how to check for stability improvements locally?</p>",
      "rawMarkdown": "I've seen a lot of (legitimate) discussions about metric hacking, but surprisingly I have not seen many kagglers discussing about the actual topic of the competition: how to tackle the stability issue. \nThe major problem I'm facing is the difficulty to elaborate a CV strategy to check for improvements / performances of my models for what concerns stability: even if I train a LGBM with only the first 30 WEEK_NUM, I see that the penalty term \"a\" is positive if computed on the remaining part of the data! This means that there's no decrease of performance and that I'm not able to see what works and what does not work in the local CV...I can just submit and infer it from the score I get back.\nDid you find any idea on how to check for stability improvements locally?",
      "votes": null
    },
    {
      "id": "2777751",
      "postDate": "04/26/2024 19:40:05",
      "content": "<p><a href=\"https://www.kaggle.com/andrealunch\" target=\"_blank\">@andrealunch</a> this is the biggest challenge in this competition. The metric is new and we have no reasonable regime to assess stability locally, at least from our collective present work. <br>\nI am sure we will learn more in this regard when the challenge will end, but now we can try out various approaches and validate the results!</p>",
      "rawMarkdown": "andrealunch this is the biggest challenge in this competition. The metric is new and we have no reasonable regime to assess stability locally, at least from our collective present work. \nI am sure we will learn more in this regard when the challenge will end, but now we can try out various approaches and validate the results!",
      "votes": null
    },
    {
      "id": "2777991",
      "postDate": "04/26/2024 22:30:39",
      "content": "<p>TLDR: check if the model has a lower number of false positives.</p>\n<p>What I have found to correlate local scores to the public lb in individual models, I check if the model performs better with the hard examples of the negative class.</p>\n<p>The dataset is highly imbalanced and the base metric is a vanilla roc-auc, so for two models with similar auc, usually the one with a lower number of fp is the one that performs the best in this competition.</p>\n<p>A few examples:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F6f6f6237e63690c1ab67ad1ad4dbe8fe%2FScreenshot%202024-04-26%20at%2011.52.56PM.png?generation=1714168388965736&amp;alt=media\"></p>\n<p>In the best scenario you are able to predict the true positives and have a very low number of false positives, but is worse to predict the true positives with a high number of false positives than is to predict a low number of true positives with a low number of false positives.</p>\n<p>The competition dataset is even more imbalanced, so it gets even worse if you don't have good predictions for the negative class.</p>\n<p>if you visualise the predictions of this models, they performs worse at predicting who is going to default, which ironically make the models perform better due to the gini calculations.</p>\n<p>Here is an example of a transformer with decent local gini stability (0.7) but a super bad score on the public lb (0.36).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F1e499af58cc28d51f4da24c41cd5a966%2FScreenshot%202024-04-26%20at%2011.53.42PM.png?generation=1714168434981982&amp;alt=media\"></p>\n<p>As expected, the predictions of this transformer has tons of false positives.</p>\n<p>Another way to test this, apply weights to the criterion to boost the positive class, train the model to a similar auc, the better it gets at predicting the true positive class the worse it performs in the public lb, if you check the predictions, the number of false positives also increased.</p>\n<p>There is probably better ways to correlate local validation to the public lb, this is what I have found so far.</p>",
      "rawMarkdown": "TLDR: check if the model has a lower number of false positives.\n\nWhat I have found to correlate local scores to the public lb in individual models, I check if the model performs better with the hard examples of the negative class.\n\nThe dataset is highly imbalanced and the base metric is a vanilla roc-auc, so for two models with similar auc, usually the one with a lower number of fp is the one that performs the best in this competition.\n\nA few examples:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F6f6f6237e63690c1ab67ad1ad4dbe8fe%2FScreenshot%202024-04-26%20at%2011.52.56PM.png?generation=1714168388965736&alt=media)\n\nIn the best scenario you are able to predict the true positives and have a very low number of false positives, but is worse to predict the true positives with a high number of false positives than is to predict a low number of true positives with a low number of false positives.\n\nThe competition dataset is even more imbalanced, so it gets even worse if you don't have good predictions for the negative class.\n\nif you visualise the predictions of this models, they performs worse at predicting who is going to default, which ironically make the models perform better due to the gini calculations.\n\nHere is an example of a transformer with decent local gini stability (0.7) but a super bad score on the public lb (0.36).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F1e499af58cc28d51f4da24c41cd5a966%2FScreenshot%202024-04-26%20at%2011.53.42PM.png?generation=1714168434981982&alt=media)\n\nAs expected, the predictions of this transformer has tons of false positives.\n\nAnother way to test this, apply weights to the criterion to boost the positive class, train the model to a similar auc, the better it gets at predicting the true positive class the worse it performs in the public lb, if you check the predictions, the number of false positives also increased.\n\nThere is probably better ways to correlate local validation to the public lb, this is what I have found so far.",
      "votes": null
    },
    {
      "id": "2778815",
      "postDate": "04/27/2024 11:29:29",
      "content": "<p>Thanks Antonio, this is very useful!</p>",
      "rawMarkdown": "Thanks Antonio, this is very useful!",
      "votes": null
    },
    {
      "id": "2800620",
      "postDate": "05/08/2024 09:20:29",
      "content": "<p><a href=\"https://www.kaggle.com/enriquezaf\" target=\"_blank\">@enriquezaf</a>  Thanks for sharing !! can we use false positive rate as optimization metric in LightGBM , Catboosting to get good LB score ? </p>",
      "rawMarkdown": "enriquezaf  Thanks for sharing !! can we use false positive rate as optimization metric in LightGBM , Catboosting to get good LB score ?",
      "votes": null
    },
    {
      "id": "2803027",
      "postDate": "05/09/2024 09:50:08",
      "content": "<p>While I find the stability topic very interesting in general, I have one major problem with it, excluding hacking considerations.</p>\n<p>The way the metric is currently constructed, it appears that a model starting with an AUC of 0.7 and remaining there for 6 months is considered better than a model starting with 0.8 AUC and dropping to 0.7 after 6 months. The only reason this would make sense to me is if one assumes that 'the model that was more stable in the first 6 months will continue being more stable'. To me, no such implication exists, and it seems more like wishful thinking. I mean, we have no clue which model will perform better in the upcoming 6 months.</p>\n<p>I know that some people would argue something like 'but the model more stable in the first 6 months has shown itself robust to drift.' Equally well, you could argue that the model starting with 0.8 learned more complex patterns (allowing it to have better performance), and when these underlying patterns change, it's likely to deteriorate and stabilize around 0.7. </p>\n<p>I mean, each model's trajectory could be influenced by a myriad of factors, so we have no clue.</p>\n<p>Some things I think could be interesting in general for stability include:</p>\n<ul>\n<li>Ensemble of many models and variables (goes without saying)</li>\n<li>Discretizing variables (possibly less influcenced by drift?)</li>\n<li>Bin probabilities into categories based on something (and perhaps have this as new target)</li>\n<li>Train a classifier to detect drift (0 for old training data, 1 for new production data) and use that to find out if distribution has changed (and somehow make use of that information)</li>\n</ul>\n<p>If it would work though, I have no idea 😅</p>",
      "rawMarkdown": "While I find the stability topic very interesting in general, I have one major problem with it, excluding hacking considerations.\n\nThe way the metric is currently constructed, it appears that a model starting with an AUC of 0.7 and remaining there for 6 months is considered better than a model starting with 0.8 AUC and dropping to 0.7 after 6 months. The only reason this would make sense to me is if one assumes that 'the model that was more stable in the first 6 months will continue being more stable'. To me, no such implication exists, and it seems more like wishful thinking. I mean, we have no clue which model will perform better in the upcoming 6 months.\n\nI know that some people would argue something like 'but the model more stable in the first 6 months has shown itself robust to drift.' Equally well, you could argue that the model starting with 0.8 learned more complex patterns (allowing it to have better performance), and when these underlying patterns change, it's likely to deteriorate and stabilize around 0.7. \n\nI mean, each model's trajectory could be influenced by a myriad of factors, so we have no clue.\n\nSome things I think could be interesting in general for stability include:\n\n* Ensemble of many models and variables (goes without saying)\n* Discretizing variables (possibly less influcenced by drift?)\n* Bin probabilities into categories based on something (and perhaps have this as new target)\n* Train a classifier to detect drift (0 for old training data, 1 for new production data) and use that to find out if distribution has changed (and somehow make use of that information)\n\nIf it would work though, I have no idea 😅",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2777751,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "04/26/2024 19:40:05",
      "content": "<p><a href=\"https://www.kaggle.com/andrealunch\" target=\"_blank\">@andrealunch</a> this is the biggest challenge in this competition. The metric is new and we have no reasonable regime to assess stability locally, at least from our collective present work. <br>\nI am sure we will learn more in this regard when the challenge will end, but now we can try out various approaches and validate the results!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2777991,
      "author_name": "enriquezaf",
      "author_url": "",
      "post_date": "04/26/2024 22:30:39",
      "content": "<p>TLDR: check if the model has a lower number of false positives.</p>\n<p>What I have found to correlate local scores to the public lb in individual models, I check if the model performs better with the hard examples of the negative class.</p>\n<p>The dataset is highly imbalanced and the base metric is a vanilla roc-auc, so for two models with similar auc, usually the one with a lower number of fp is the one that performs the best in this competition.</p>\n<p>A few examples:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F6f6f6237e63690c1ab67ad1ad4dbe8fe%2FScreenshot%202024-04-26%20at%2011.52.56PM.png?generation=1714168388965736&amp;alt=media\"></p>\n<p>In the best scenario you are able to predict the true positives and have a very low number of false positives, but is worse to predict the true positives with a high number of false positives than is to predict a low number of true positives with a low number of false positives.</p>\n<p>The competition dataset is even more imbalanced, so it gets even worse if you don't have good predictions for the negative class.</p>\n<p>if you visualise the predictions of this models, they performs worse at predicting who is going to default, which ironically make the models perform better due to the gini calculations.</p>\n<p>Here is an example of a transformer with decent local gini stability (0.7) but a super bad score on the public lb (0.36).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F1e499af58cc28d51f4da24c41cd5a966%2FScreenshot%202024-04-26%20at%2011.53.42PM.png?generation=1714168434981982&amp;alt=media\"></p>\n<p>As expected, the predictions of this transformer has tons of false positives.</p>\n<p>Another way to test this, apply weights to the criterion to boost the positive class, train the model to a similar auc, the better it gets at predicting the true positive class the worse it performs in the public lb, if you check the predictions, the number of false positives also increased.</p>\n<p>There is probably better ways to correlate local validation to the public lb, this is what I have found so far.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2778815,
          "author_name": "andrealunch",
          "author_url": "",
          "post_date": "04/27/2024 11:29:29",
          "content": "<p>Thanks Antonio, this is very useful!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2800620,
          "author_name": "kunduruanil",
          "author_url": "",
          "post_date": "05/08/2024 09:20:29",
          "content": "<p><a href=\"https://www.kaggle.com/enriquezaf\" target=\"_blank\">@enriquezaf</a>  Thanks for sharing !! can we use false positive rate as optimization metric in LightGBM , Catboosting to get good LB score ? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2803027,
      "author_name": "ern711",
      "author_url": "",
      "post_date": "05/09/2024 09:50:08",
      "content": "<p>While I find the stability topic very interesting in general, I have one major problem with it, excluding hacking considerations.</p>\n<p>The way the metric is currently constructed, it appears that a model starting with an AUC of 0.7 and remaining there for 6 months is considered better than a model starting with 0.8 AUC and dropping to 0.7 after 6 months. The only reason this would make sense to me is if one assumes that 'the model that was more stable in the first 6 months will continue being more stable'. To me, no such implication exists, and it seems more like wishful thinking. I mean, we have no clue which model will perform better in the upcoming 6 months.</p>\n<p>I know that some people would argue something like 'but the model more stable in the first 6 months has shown itself robust to drift.' Equally well, you could argue that the model starting with 0.8 learned more complex patterns (allowing it to have better performance), and when these underlying patterns change, it's likely to deteriorate and stabilize around 0.7. </p>\n<p>I mean, each model's trajectory could be influenced by a myriad of factors, so we have no clue.</p>\n<p>Some things I think could be interesting in general for stability include:</p>\n<ul>\n<li>Ensemble of many models and variables (goes without saying)</li>\n<li>Discretizing variables (possibly less influcenced by drift?)</li>\n<li>Bin probabilities into categories based on something (and perhaps have this as new target)</li>\n<li>Train a classifier to detect drift (0 for old training data, 1 for new production data) and use that to find out if distribution has changed (and somehow make use of that information)</li>\n</ul>\n<p>If it would work though, I have no idea 😅</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2777486": "I've seen a lot of (legitimate) discussions about metric hacking, but surprisingly I have not seen many kagglers discussing about the actual topic of the competition: how to tackle the stability issue. \nThe major problem I'm facing is the difficulty to elaborate a CV strategy to check for improvements / performances of my models for what concerns stability: even if I train a LGBM with only the first 30 WEEK_NUM, I see that the penalty term \"a\" is positive if computed on the remaining part of the data! This means that there's no decrease of performance and that I'm not able to see what works and what does not work in the local CV...I can just submit and infer it from the score I get back.\nDid you find any idea on how to check for stability improvements locally?",
    "2777751": "andrealunch this is the biggest challenge in this competition. The metric is new and we have no reasonable regime to assess stability locally, at least from our collective present work. \nI am sure we will learn more in this regard when the challenge will end, but now we can try out various approaches and validate the results!",
    "2777991": "TLDR: check if the model has a lower number of false positives.\n\nWhat I have found to correlate local scores to the public lb in individual models, I check if the model performs better with the hard examples of the negative class.\n\nThe dataset is highly imbalanced and the base metric is a vanilla roc-auc, so for two models with similar auc, usually the one with a lower number of fp is the one that performs the best in this competition.\n\nA few examples:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F6f6f6237e63690c1ab67ad1ad4dbe8fe%2FScreenshot%202024-04-26%20at%2011.52.56PM.png?generation=1714168388965736&alt=media)\n\nIn the best scenario you are able to predict the true positives and have a very low number of false positives, but is worse to predict the true positives with a high number of false positives than is to predict a low number of true positives with a low number of false positives.\n\nThe competition dataset is even more imbalanced, so it gets even worse if you don't have good predictions for the negative class.\n\nif you visualise the predictions of this models, they performs worse at predicting who is going to default, which ironically make the models perform better due to the gini calculations.\n\nHere is an example of a transformer with decent local gini stability (0.7) but a super bad score on the public lb (0.36).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F1e499af58cc28d51f4da24c41cd5a966%2FScreenshot%202024-04-26%20at%2011.53.42PM.png?generation=1714168434981982&alt=media)\n\nAs expected, the predictions of this transformer has tons of false positives.\n\nAnother way to test this, apply weights to the criterion to boost the positive class, train the model to a similar auc, the better it gets at predicting the true positive class the worse it performs in the public lb, if you check the predictions, the number of false positives also increased.\n\nThere is probably better ways to correlate local validation to the public lb, this is what I have found so far.",
    "2778815": "Thanks Antonio, this is very useful!",
    "2800620": "enriquezaf  Thanks for sharing !! can we use false positive rate as optimization metric in LightGBM , Catboosting to get good LB score ?",
    "2803027": "While I find the stability topic very interesting in general, I have one major problem with it, excluding hacking considerations.\n\nThe way the metric is currently constructed, it appears that a model starting with an AUC of 0.7 and remaining there for 6 months is considered better than a model starting with 0.8 AUC and dropping to 0.7 after 6 months. The only reason this would make sense to me is if one assumes that 'the model that was more stable in the first 6 months will continue being more stable'. To me, no such implication exists, and it seems more like wishful thinking. I mean, we have no clue which model will perform better in the upcoming 6 months.\n\nI know that some people would argue something like 'but the model more stable in the first 6 months has shown itself robust to drift.' Equally well, you could argue that the model starting with 0.8 learned more complex patterns (allowing it to have better performance), and when these underlying patterns change, it's likely to deteriorate and stabilize around 0.7. \n\nI mean, each model's trajectory could be influenced by a myriad of factors, so we have no clue.\n\nSome things I think could be interesting in general for stability include:\n\n* Ensemble of many models and variables (goes without saying)\n* Discretizing variables (possibly less influcenced by drift?)\n* Bin probabilities into categories based on something (and perhaps have this as new target)\n* Train a classifier to detect drift (0 for old training data, 1 for new production data) and use that to find out if distribution has changed (and somehow make use of that information)\n\nIf it would work though, I have no idea 😅"
  },
  "source": "meta"
}