{
  "id": 338283,
  "title": "Experimenting with different Early Stopping strategies",
  "url": "/competitions/amex-default-prediction/discussion/338283",
  "author_name": "",
  "post_date": "2022-07-19T23:31:35.657776100Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The 4% metric is <em>very</em> noisy. And both competition metrics are relative rank metrics, so even after model training is finished, there's several options for which model iteration is the \"best\" one to use for prediction:</p>\n<ul>\n<li>Best AMEX score (since it's the final criteria)</li>\n<li>Best Gini (auc) score (since it's a high quality, more consistent, relative rank metric</li>\n<li>Best log-loss score (since it's a high quality true prediction metric and goes hand-in-hand with the objective function.)</li>\n</ul>\n<p>Further, when ensembling each fold together, there's at least a couple options. Even for CV, without true ensembling, there's options on whether/how to normalize model A and model B's predictions together.</p>\n<ul>\n<li>Prediction score (for holdout, this itself has two varieties, raw margin or the actual 0.0-1.0 range prediction value)</li>\n<li>Relative rank</li>\n</ul>\n<p>A good experiment would be to do - not sure the technical term - doubled k-fold validation+holdout, map results from a few different seeds on each holdout set, see which strategy best generalizes to unseen information, and also see if the experiment results have statistical significance.</p>\n<p>But even just looking at CV score, it's pretty interesting: Here are the 6 different scores for a single 5 fold model, using early stopping on auc, but then retroactively checking and finding best iteration for AMEX score, auc, or log-loss, and then taking the actual pred values and stitching the results together into a single dataset, or doing ranked results and stitching those together (tiebreaker using the pred value).</p>\n<p>You would expect taking the noisy AMEX score, cherry-picking the ones you already determined are best, and stitching them together would get the highest CV score (the best holdout score? remains to be tested). And you'd be right. But <em>only</em> if converting the raw scores to ranking, otherwise (due to less consistent scale between models?) it gets the worst score… at least on one data point.</p>\n<p>If it's not clear, amex vs gini vs logloss in labels below just means which iteration of the model I stopped at and used for predictions. For example:<br>\nFold0:<br>\nBest amex ntree limit: 8117   (0.7979915790671912)<br>\nBest gini ntree limit: 8341   (0.9258354899602064)<br>\nBest logloss ntree limit: 8328   (0.21375135626306707)</p>\n<p>And here's the different final CV scores, depending on the 6 (3x2) different methodologies:<br>\nAmex_pred:<br>\n  4%  : 0.66828525<br>\n  Gini: 0.925125257468853<br>\nKaggle: 0.7967052540663051</p>\n<p>Amex_rank:<br>\n  4%  : 0.6685882<br>\n  Gini: 0.9251309052658099<br>\nKaggle: 0.7968595631694803</p>\n<p>Gini_pred:<br>\n  4%  : 0.6684115<br>\n  Gini: 0.9251342153129204<br>\nKaggle: 0.7967728543071559</p>\n<p>Gini_rank:<br>\n  4%  : 0.6684367<br>\n  Gini: 0.9251394110109513<br>\nKaggle: 0.7967880585385414</p>\n<p>Logloss_pred:<br>\n  4%  : 0.66851246<br>\n  Gini: 0.9251324394821973<br>\nKaggle: 0.7968224515259192</p>\n<p>Logloss_rank:<br>\n  4%  : 0.6683778<br>\n  Gini: 0.9251367014363389<br>\nKaggle: 0.7967572590567162</p>",
  "messages": [
    {
      "id": "1862718",
      "postDate": "07/19/2022 23:31:35",
      "content": "<p>The 4% metric is <em>very</em> noisy. And both competition metrics are relative rank metrics, so even after model training is finished, there's several options for which model iteration is the \"best\" one to use for prediction:</p>\n<ul>\n<li>Best AMEX score (since it's the final criteria)</li>\n<li>Best Gini (auc) score (since it's a high quality, more consistent, relative rank metric</li>\n<li>Best log-loss score (since it's a high quality true prediction metric and goes hand-in-hand with the objective function.)</li>\n</ul>\n<p>Further, when ensembling each fold together, there's at least a couple options. Even for CV, without true ensembling, there's options on whether/how to normalize model A and model B's predictions together.</p>\n<ul>\n<li>Prediction score (for holdout, this itself has two varieties, raw margin or the actual 0.0-1.0 range prediction value)</li>\n<li>Relative rank</li>\n</ul>\n<p>A good experiment would be to do - not sure the technical term - doubled k-fold validation+holdout, map results from a few different seeds on each holdout set, see which strategy best generalizes to unseen information, and also see if the experiment results have statistical significance.</p>\n<p>But even just looking at CV score, it's pretty interesting: Here are the 6 different scores for a single 5 fold model, using early stopping on auc, but then retroactively checking and finding best iteration for AMEX score, auc, or log-loss, and then taking the actual pred values and stitching the results together into a single dataset, or doing ranked results and stitching those together (tiebreaker using the pred value).</p>\n<p>You would expect taking the noisy AMEX score, cherry-picking the ones you already determined are best, and stitching them together would get the highest CV score (the best holdout score? remains to be tested). And you'd be right. But <em>only</em> if converting the raw scores to ranking, otherwise (due to less consistent scale between models?) it gets the worst score… at least on one data point.</p>\n<p>If it's not clear, amex vs gini vs logloss in labels below just means which iteration of the model I stopped at and used for predictions. For example:<br>\nFold0:<br>\nBest amex ntree limit: 8117   (0.7979915790671912)<br>\nBest gini ntree limit: 8341   (0.9258354899602064)<br>\nBest logloss ntree limit: 8328   (0.21375135626306707)</p>\n<p>And here's the different final CV scores, depending on the 6 (3x2) different methodologies:<br>\nAmex_pred:<br>\n  4%  : 0.66828525<br>\n  Gini: 0.925125257468853<br>\nKaggle: 0.7967052540663051</p>\n<p>Amex_rank:<br>\n  4%  : 0.6685882<br>\n  Gini: 0.9251309052658099<br>\nKaggle: 0.7968595631694803</p>\n<p>Gini_pred:<br>\n  4%  : 0.6684115<br>\n  Gini: 0.9251342153129204<br>\nKaggle: 0.7967728543071559</p>\n<p>Gini_rank:<br>\n  4%  : 0.6684367<br>\n  Gini: 0.9251394110109513<br>\nKaggle: 0.7967880585385414</p>\n<p>Logloss_pred:<br>\n  4%  : 0.66851246<br>\n  Gini: 0.9251324394821973<br>\nKaggle: 0.7968224515259192</p>\n<p>Logloss_rank:<br>\n  4%  : 0.6683778<br>\n  Gini: 0.9251367014363389<br>\nKaggle: 0.7967572590567162</p>",
      "rawMarkdown": "The 4% metric is _very_ noisy. And both competition metrics are relative rank metrics, so even after model training is finished, there's several options for which model iteration is the \"best\" one to use for prediction:\n* Best AMEX score (since it's the final criteria)\n* Best Gini (auc) score (since it's a high quality, more consistent, relative rank metric\n* Best log-loss score (since it's a high quality true prediction metric and goes hand-in-hand with the objective function.)\n\nFurther, when ensembling each fold together, there's at least a couple options. Even for CV, without true ensembling, there's options on whether/how to normalize model A and model B's predictions together.\n* Prediction score (for holdout, this itself has two varieties, raw margin or the actual 0.0-1.0 range prediction value)\n* Relative rank\n\nA good experiment would be to do - not sure the technical term - doubled k-fold validation+holdout, map results from a few different seeds on each holdout set, see which strategy best generalizes to unseen information, and also see if the experiment results have statistical significance.\n\nBut even just looking at CV score, it's pretty interesting: Here are the 6 different scores for a single 5 fold model, using early stopping on auc, but then retroactively checking and finding best iteration for AMEX score, auc, or log-loss, and then taking the actual pred values and stitching the results together into a single dataset, or doing ranked results and stitching those together (tiebreaker using the pred value).\n\nYou would expect taking the noisy AMEX score, cherry-picking the ones you already determined are best, and stitching them together would get the highest CV score (the best holdout score? remains to be tested). And you'd be right. But *only* if converting the raw scores to ranking, otherwise (due to less consistent scale between models?) it gets the worst score... at least on one data point.\n\nIf it's not clear, amex vs gini vs logloss in labels below just means which iteration of the model I stopped at and used for predictions. For example:\nFold0:\nBest amex ntree limit: 8117   (0.7979915790671912)\nBest gini ntree limit: 8341   (0.9258354899602064)\nBest logloss ntree limit: 8328   (0.21375135626306707)\n\nAnd here's the different final CV scores, depending on the 6 (3x2) different methodologies:\nAmex_pred:\n  4%  : 0.66828525\n  Gini: 0.925125257468853\nKaggle: 0.7967052540663051\n\nAmex_rank:\n  4%  : 0.6685882\n  Gini: 0.9251309052658099\nKaggle: 0.7968595631694803\n\nGini_pred:\n  4%  : 0.6684115\n  Gini: 0.9251342153129204\nKaggle: 0.7967728543071559\n\nGini_rank:\n  4%  : 0.6684367\n  Gini: 0.9251394110109513\nKaggle: 0.7967880585385414\n\nLogloss_pred:\n  4%  : 0.66851246\n  Gini: 0.9251324394821973\nKaggle: 0.7968224515259192\n\nLogloss_rank:\n  4%  : 0.6683778\n  Gini: 0.9251367014363389\nKaggle: 0.7967572590567162",
      "votes": null
    },
    {
      "id": "1862760",
      "postDate": "07/20/2022 01:41:09",
      "content": "<p>I like the questions you are asking, and agree with your way of thinking. However, I am not sure we can ever figure this one out properly without test labels. Even full test scores rather than those truncated at 3rd decimal place would help somewhat.</p>\n<p>I think the biggest problem is that we don't know which of the metrics available to us is the best proxy for AmEx scores, as I discussed <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/338268\" target=\"_blank\"><strong>here</strong></a>. We can assume that early on in training any of the log-loss, AUC or Gini behave in a way that is consistent with AmEx scores (the first goes monotonically down as AmEx goes up, while the last two go along with AmEx). Let's keep in mind that our training machines, no matter what kind they are exactly, are minimizing the loss function (log-loss specifically, unless one is using a modified loss function). At some point when we get close to a log-loss minimum, the training will still be moving in a direction that further minimizes log-loss, which may or may not be maximizing either one of AUC, Gini or AmEx. So we can rely on our loss function up to a point, but sooner or later we hit the area where optimizing our loss function may not be optimizing our final scoring function. Yes, these differences are mostly on 3rd decimal place and beyond, but in a competition all of them matter.</p>\n<p>That's why luck plays such a prominent role when it comes to how we split our folds, and whether we do early stopping after 300 or 500 iterations without improvement. I think there is a good chance the LGB-DART submission is overfitting to the LB, and anything we base on that submission will be dropping on private LB. I like that script and it has many excellent parts, but if there are many drops by 1000+ places on the private LB it will likely be because of it.</p>",
      "rawMarkdown": "I like the questions you are asking, and agree with your way of thinking. However, I am not sure we can ever figure this one out properly without test labels. Even full test scores rather than those truncated at 3rd decimal place would help somewhat.\n\nI think the biggest problem is that we don't know which of the metrics available to us is the best proxy for AmEx scores, as I discussed [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/338268). We can assume that early on in training any of the log-loss, AUC or Gini behave in a way that is consistent with AmEx scores (the first goes monotonically down as AmEx goes up, while the last two go along with AmEx). Let's keep in mind that our training machines, no matter what kind they are exactly, are minimizing the loss function (log-loss specifically, unless one is using a modified loss function). At some point when we get close to a log-loss minimum, the training will still be moving in a direction that further minimizes log-loss, which may or may not be maximizing either one of AUC, Gini or AmEx. So we can rely on our loss function up to a point, but sooner or later we hit the area where optimizing our loss function may not be optimizing our final scoring function. Yes, these differences are mostly on 3rd decimal place and beyond, but in a competition all of them matter.\n\nThat's why luck plays such a prominent role when it comes to how we split our folds, and whether we do early stopping after 300 or 500 iterations without improvement. I think there is a good chance the LGB-DART submission is overfitting to the LB, and anything we base on that submission will be dropping on private LB. I like that script and it has many excellent parts, but if there are many drops by 1000+ places on the private LB it will likely be because of it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1862760,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "07/20/2022 01:41:09",
      "content": "<p>I like the questions you are asking, and agree with your way of thinking. However, I am not sure we can ever figure this one out properly without test labels. Even full test scores rather than those truncated at 3rd decimal place would help somewhat.</p>\n<p>I think the biggest problem is that we don't know which of the metrics available to us is the best proxy for AmEx scores, as I discussed <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/338268\" target=\"_blank\"><strong>here</strong></a>. We can assume that early on in training any of the log-loss, AUC or Gini behave in a way that is consistent with AmEx scores (the first goes monotonically down as AmEx goes up, while the last two go along with AmEx). Let's keep in mind that our training machines, no matter what kind they are exactly, are minimizing the loss function (log-loss specifically, unless one is using a modified loss function). At some point when we get close to a log-loss minimum, the training will still be moving in a direction that further minimizes log-loss, which may or may not be maximizing either one of AUC, Gini or AmEx. So we can rely on our loss function up to a point, but sooner or later we hit the area where optimizing our loss function may not be optimizing our final scoring function. Yes, these differences are mostly on 3rd decimal place and beyond, but in a competition all of them matter.</p>\n<p>That's why luck plays such a prominent role when it comes to how we split our folds, and whether we do early stopping after 300 or 500 iterations without improvement. I think there is a good chance the LGB-DART submission is overfitting to the LB, and anything we base on that submission will be dropping on private LB. I like that script and it has many excellent parts, but if there are many drops by 1000+ places on the private LB it will likely be because of it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1862718": "The 4% metric is _very_ noisy. And both competition metrics are relative rank metrics, so even after model training is finished, there's several options for which model iteration is the \"best\" one to use for prediction:\n* Best AMEX score (since it's the final criteria)\n* Best Gini (auc) score (since it's a high quality, more consistent, relative rank metric\n* Best log-loss score (since it's a high quality true prediction metric and goes hand-in-hand with the objective function.)\n\nFurther, when ensembling each fold together, there's at least a couple options. Even for CV, without true ensembling, there's options on whether/how to normalize model A and model B's predictions together.\n* Prediction score (for holdout, this itself has two varieties, raw margin or the actual 0.0-1.0 range prediction value)\n* Relative rank\n\nA good experiment would be to do - not sure the technical term - doubled k-fold validation+holdout, map results from a few different seeds on each holdout set, see which strategy best generalizes to unseen information, and also see if the experiment results have statistical significance.\n\nBut even just looking at CV score, it's pretty interesting: Here are the 6 different scores for a single 5 fold model, using early stopping on auc, but then retroactively checking and finding best iteration for AMEX score, auc, or log-loss, and then taking the actual pred values and stitching the results together into a single dataset, or doing ranked results and stitching those together (tiebreaker using the pred value).\n\nYou would expect taking the noisy AMEX score, cherry-picking the ones you already determined are best, and stitching them together would get the highest CV score (the best holdout score? remains to be tested). And you'd be right. But *only* if converting the raw scores to ranking, otherwise (due to less consistent scale between models?) it gets the worst score... at least on one data point.\n\nIf it's not clear, amex vs gini vs logloss in labels below just means which iteration of the model I stopped at and used for predictions. For example:\nFold0:\nBest amex ntree limit: 8117   (0.7979915790671912)\nBest gini ntree limit: 8341   (0.9258354899602064)\nBest logloss ntree limit: 8328   (0.21375135626306707)\n\nAnd here's the different final CV scores, depending on the 6 (3x2) different methodologies:\nAmex_pred:\n  4%  : 0.66828525\n  Gini: 0.925125257468853\nKaggle: 0.7967052540663051\n\nAmex_rank:\n  4%  : 0.6685882\n  Gini: 0.9251309052658099\nKaggle: 0.7968595631694803\n\nGini_pred:\n  4%  : 0.6684115\n  Gini: 0.9251342153129204\nKaggle: 0.7967728543071559\n\nGini_rank:\n  4%  : 0.6684367\n  Gini: 0.9251394110109513\nKaggle: 0.7967880585385414\n\nLogloss_pred:\n  4%  : 0.66851246\n  Gini: 0.9251324394821973\nKaggle: 0.7968224515259192\n\nLogloss_rank:\n  4%  : 0.6683778\n  Gini: 0.9251367014363389\nKaggle: 0.7967572590567162",
    "1862760": "I like the questions you are asking, and agree with your way of thinking. However, I am not sure we can ever figure this one out properly without test labels. Even full test scores rather than those truncated at 3rd decimal place would help somewhat.\n\nI think the biggest problem is that we don't know which of the metrics available to us is the best proxy for AmEx scores, as I discussed [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/338268). We can assume that early on in training any of the log-loss, AUC or Gini behave in a way that is consistent with AmEx scores (the first goes monotonically down as AmEx goes up, while the last two go along with AmEx). Let's keep in mind that our training machines, no matter what kind they are exactly, are minimizing the loss function (log-loss specifically, unless one is using a modified loss function). At some point when we get close to a log-loss minimum, the training will still be moving in a direction that further minimizes log-loss, which may or may not be maximizing either one of AUC, Gini or AmEx. So we can rely on our loss function up to a point, but sooner or later we hit the area where optimizing our loss function may not be optimizing our final scoring function. Yes, these differences are mostly on 3rd decimal place and beyond, but in a competition all of them matter.\n\nThat's why luck plays such a prominent role when it comes to how we split our folds, and whether we do early stopping after 300 or 500 iterations without improvement. I think there is a good chance the LGB-DART submission is overfitting to the LB, and anything we base on that submission will be dropping on private LB. I like that script and it has many excellent parts, but if there are many drops by 1000+ places on the private LB it will likely be because of it."
  },
  "source": "meta"
}