{
  "id": 337188,
  "title": "Model ensembling",
  "url": "/competitions/amex-default-prediction/discussion/337188",
  "author_name": "",
  "post_date": "2022-07-14T23:29:03.671196800Z",
  "votes": 27,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Sometimes model ensembling works, and other times it doesn't. <strong>There is always a possibility that you are making better models from an ensemble, but they have the same scores according to the public leaderboard (it shows only the first 3 decimal places).</strong> If we sort our submissions by public scores, it should rank them properly even if they have identical scores.</p>\n<p>If you are truly not getting better models by ensembling, there are 2 possibilities:</p>\n<ul>\n<li>Your best model is so much better than any other, and all your models are highly correlated</li>\n<li>You are not doing the ensembling properly</li>\n</ul>\n<p>Highly correlated models do not ensemble well, for what should be an obvious reason. If the two models have near-identical predictions, the outcome will be almost identical whether you do simple averaging, or linearly weigh their contributions in some other way (<code>[0.7*model1] + [0.3*model2]</code>, or whatever other weight combination). Experts with non-overlapping expertise (non-correlated models) are much better when combined.</p>\n<p>For the second possibility, many people try to guess model contributions such that weights add up to 1, or some variation of that. That's not ensembling - that's overfitting to the public leaderboard. A proper way to ensemble is by using out-of-fold (OOF) predictions (from train data) to gauge how much weight each model should be given, and then those weights are used with test data. But using OOF files alone is not enough, as our models also must have similar gaps between cross-validation (CV) and public leaderboard scores. I can elaborate further on this if needed, but generally speaking models whose CV scores underestimate the leaderboard scores will be underweighted in the final ensemble. Conversely, models whose CV scores overestimated the leaderboard will be given higher than optimal weights.</p>",
  "messages": [
    {
      "id": "1855795",
      "postDate": "07/14/2022 23:29:03",
      "content": "<p>Sometimes model ensembling works, and other times it doesn't. <strong>There is always a possibility that you are making better models from an ensemble, but they have the same scores according to the public leaderboard (it shows only the first 3 decimal places).</strong> If we sort our submissions by public scores, it should rank them properly even if they have identical scores.</p>\n<p>If you are truly not getting better models by ensembling, there are 2 possibilities:</p>\n<ul>\n<li>Your best model is so much better than any other, and all your models are highly correlated</li>\n<li>You are not doing the ensembling properly</li>\n</ul>\n<p>Highly correlated models do not ensemble well, for what should be an obvious reason. If the two models have near-identical predictions, the outcome will be almost identical whether you do simple averaging, or linearly weigh their contributions in some other way (<code>[0.7*model1] + [0.3*model2]</code>, or whatever other weight combination). Experts with non-overlapping expertise (non-correlated models) are much better when combined.</p>\n<p>For the second possibility, many people try to guess model contributions such that weights add up to 1, or some variation of that. That's not ensembling - that's overfitting to the public leaderboard. A proper way to ensemble is by using out-of-fold (OOF) predictions (from train data) to gauge how much weight each model should be given, and then those weights are used with test data. But using OOF files alone is not enough, as our models also must have similar gaps between cross-validation (CV) and public leaderboard scores. I can elaborate further on this if needed, but generally speaking models whose CV scores underestimate the leaderboard scores will be underweighted in the final ensemble. Conversely, models whose CV scores overestimated the leaderboard will be given higher than optimal weights.</p>",
      "rawMarkdown": "Sometimes model ensembling works, and other times it doesn't. **There is always a possibility that you are making better models from an ensemble, but they have the same scores according to the public leaderboard (it shows only the first 3 decimal places).** If we sort our submissions by public scores, it should rank them properly even if they have identical scores.\n\nIf you are truly not getting better models by ensembling, there are 2 possibilities:\n\n - Your best model is so much better than any other, and all your models are highly correlated\n - You are not doing the ensembling properly\n\nHighly correlated models do not ensemble well, for what should be an obvious reason. If the two models have near-identical predictions, the outcome will be almost identical whether you do simple averaging, or linearly weigh their contributions in some other way (`[0.7*model1] + [0.3*model2]`, or whatever other weight combination). Experts with non-overlapping expertise (non-correlated models) are much better when combined.\n\nFor the second possibility, many people try to guess model contributions such that weights add up to 1, or some variation of that. That's not ensembling - that's overfitting to the public leaderboard. A proper way to ensemble is by using out-of-fold (OOF) predictions (from train data) to gauge how much weight each model should be given, and then those weights are used with test data. But using OOF files alone is not enough, as our models also must have similar gaps between cross-validation (CV) and public leaderboard scores. I can elaborate further on this if needed, but generally speaking models whose CV scores underestimate the leaderboard scores will be underweighted in the final ensemble. Conversely, models whose CV scores overestimated the leaderboard will be given higher than optimal weights.",
      "votes": null
    },
    {
      "id": "1856009",
      "postDate": "07/15/2022 04:32:58",
      "content": "<p>Very well written! Thanks for sharing</p>",
      "rawMarkdown": "Very well written! Thanks for sharing",
      "votes": null
    },
    {
      "id": "1856170",
      "postDate": "07/15/2022 07:03:04",
      "content": "<p>Very good point.<br>\nI have a question.You state that when an oof CV underestimates the score it will be underweighted in final ensemble.<br>\nI think that means that a proper adjustment &gt;1 must be applied to them. Have I understood your point correctly??<br>\nThanks</p>",
      "rawMarkdown": "Very good point.\nI have a question.You state that when an oof CV underestimates the score it will be underweighted in final ensemble.\nI think that means that a proper adjustment >1 must be applied to them. Have I understood your point correctly??\nThanks",
      "votes": null
    },
    {
      "id": "1856997",
      "postDate": "07/15/2022 19:01:35",
      "content": "<p>Applying manual adjustments is not what I would do, because to me that is not much better than trying to guess weights. Still, if it can't be avoided, finding weights from OOF files and slightly adjusting them may be a more informed approach than completely guessing.</p>\n<p>I emphasize again that using models with similar CV-LB gap is the best approach. For most of my models the LB scores are better by 0.001 than my CV scores. That is difficult to ascertain exactly because we have more significant digits from CV than from LB scores, but they are in that ballpark. As long as we don't include a model where <code>[CV-LB]</code> is 0.003 or something like that, the ensembling should work well. That could mean excluding the very best model, and I know it sounds counterintuitive. Still, second and third best models may ensemble better with each other than the best one with either of them. Especially for a large collection of good models, the best model may be unnecessary.</p>\n<p>PS I don't think weights need to sum up to 1 and negative weights are OK. I have discussed that in one of my <a href=\"https://www.kaggle.com/code/tilii7/cross-validation-weighted-linear-blending-errors\" target=\"_blank\">old notebooks</a>.</p>",
      "rawMarkdown": "Applying manual adjustments is not what I would do, because to me that is not much better than trying to guess weights. Still, if it can't be avoided, finding weights from OOF files and slightly adjusting them may be a more informed approach than completely guessing.\n\nI emphasize again that using models with similar CV-LB gap is the best approach. For most of my models the LB scores are better by 0.001 than my CV scores. That is difficult to ascertain exactly because we have more significant digits from CV than from LB scores, but they are in that ballpark. As long as we don't include a model where `[CV-LB]` is 0.003 or something like that, the ensembling should work well. That could mean excluding the very best model, and I know it sounds counterintuitive. Still, second and third best models may ensemble better with each other than the best one with either of them. Especially for a large collection of good models, the best model may be unnecessary.\n\nPS I don't think weights need to sum up to 1 and negative weights are OK. I have discussed that in one of my [old notebooks](https://www.kaggle.com/code/tilii7/cross-validation-weighted-linear-blending-errors).",
      "votes": null
    },
    {
      "id": "1857122",
      "postDate": "07/15/2022 21:22:40",
      "content": "<p>Thanks for your answer.</p>",
      "rawMarkdown": "Thanks for your answer.",
      "votes": null
    },
    {
      "id": "1857904",
      "postDate": "07/16/2022 13:36:50",
      "content": "<p>I've just used different types of tree models in my ensemble. Should I include linear models as well?</p>",
      "rawMarkdown": "I've just used different types of tree models in my ensemble. Should I include linear models as well?",
      "votes": null
    },
    {
      "id": "1858293",
      "postDate": "07/16/2022 19:39:46",
      "content": "<p>Diverse models are essential for good ensembling, and it is important to include them no matter how they were obtained. That goes even for models that have relatively low scores.</p>\n<p>Linear models are very likely to be different from tree models, and are generally good candidates to improve ensemble diversity. But any two models can be different enough even if both of them are tree-based. For example, if they were made using different features and different hyperparameters, there is a good chance that even two LGB models could be different enough.</p>",
      "rawMarkdown": "Diverse models are essential for good ensembling, and it is important to include them no matter how they were obtained. That goes even for models that have relatively low scores.\n\nLinear models are very likely to be different from tree models, and are generally good candidates to improve ensemble diversity. But any two models can be different enough even if both of them are tree-based. For example, if they were made using different features and different hyperparameters, there is a good chance that even two LGB models could be different enough.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1856009,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/15/2022 04:32:58",
      "content": "<p>Very well written! Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1856170,
      "author_name": "georgem",
      "author_url": "",
      "post_date": "07/15/2022 07:03:04",
      "content": "<p>Very good point.<br>\nI have a question.You state that when an oof CV underestimates the score it will be underweighted in final ensemble.<br>\nI think that means that a proper adjustment &gt;1 must be applied to them. Have I understood your point correctly??<br>\nThanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1856997,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/15/2022 19:01:35",
          "content": "<p>Applying manual adjustments is not what I would do, because to me that is not much better than trying to guess weights. Still, if it can't be avoided, finding weights from OOF files and slightly adjusting them may be a more informed approach than completely guessing.</p>\n<p>I emphasize again that using models with similar CV-LB gap is the best approach. For most of my models the LB scores are better by 0.001 than my CV scores. That is difficult to ascertain exactly because we have more significant digits from CV than from LB scores, but they are in that ballpark. As long as we don't include a model where <code>[CV-LB]</code> is 0.003 or something like that, the ensembling should work well. That could mean excluding the very best model, and I know it sounds counterintuitive. Still, second and third best models may ensemble better with each other than the best one with either of them. Especially for a large collection of good models, the best model may be unnecessary.</p>\n<p>PS I don't think weights need to sum up to 1 and negative weights are OK. I have discussed that in one of my <a href=\"https://www.kaggle.com/code/tilii7/cross-validation-weighted-linear-blending-errors\" target=\"_blank\">old notebooks</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1857122,
          "author_name": "georgem",
          "author_url": "",
          "post_date": "07/15/2022 21:22:40",
          "content": "<p>Thanks for your answer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1857904,
      "author_name": "vivekjan14",
      "author_url": "",
      "post_date": "07/16/2022 13:36:50",
      "content": "<p>I've just used different types of tree models in my ensemble. Should I include linear models as well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1858293,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/16/2022 19:39:46",
          "content": "<p>Diverse models are essential for good ensembling, and it is important to include them no matter how they were obtained. That goes even for models that have relatively low scores.</p>\n<p>Linear models are very likely to be different from tree models, and are generally good candidates to improve ensemble diversity. But any two models can be different enough even if both of them are tree-based. For example, if they were made using different features and different hyperparameters, there is a good chance that even two LGB models could be different enough.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1855795": "Sometimes model ensembling works, and other times it doesn't. **There is always a possibility that you are making better models from an ensemble, but they have the same scores according to the public leaderboard (it shows only the first 3 decimal places).** If we sort our submissions by public scores, it should rank them properly even if they have identical scores.\n\nIf you are truly not getting better models by ensembling, there are 2 possibilities:\n\n - Your best model is so much better than any other, and all your models are highly correlated\n - You are not doing the ensembling properly\n\nHighly correlated models do not ensemble well, for what should be an obvious reason. If the two models have near-identical predictions, the outcome will be almost identical whether you do simple averaging, or linearly weigh their contributions in some other way (`[0.7*model1] + [0.3*model2]`, or whatever other weight combination). Experts with non-overlapping expertise (non-correlated models) are much better when combined.\n\nFor the second possibility, many people try to guess model contributions such that weights add up to 1, or some variation of that. That's not ensembling - that's overfitting to the public leaderboard. A proper way to ensemble is by using out-of-fold (OOF) predictions (from train data) to gauge how much weight each model should be given, and then those weights are used with test data. But using OOF files alone is not enough, as our models also must have similar gaps between cross-validation (CV) and public leaderboard scores. I can elaborate further on this if needed, but generally speaking models whose CV scores underestimate the leaderboard scores will be underweighted in the final ensemble. Conversely, models whose CV scores overestimated the leaderboard will be given higher than optimal weights.",
    "1856009": "Very well written! Thanks for sharing",
    "1856170": "Very good point.\nI have a question.You state that when an oof CV underestimates the score it will be underweighted in final ensemble.\nI think that means that a proper adjustment >1 must be applied to them. Have I understood your point correctly??\nThanks",
    "1856997": "Applying manual adjustments is not what I would do, because to me that is not much better than trying to guess weights. Still, if it can't be avoided, finding weights from OOF files and slightly adjusting them may be a more informed approach than completely guessing.\n\nI emphasize again that using models with similar CV-LB gap is the best approach. For most of my models the LB scores are better by 0.001 than my CV scores. That is difficult to ascertain exactly because we have more significant digits from CV than from LB scores, but they are in that ballpark. As long as we don't include a model where `[CV-LB]` is 0.003 or something like that, the ensembling should work well. That could mean excluding the very best model, and I know it sounds counterintuitive. Still, second and third best models may ensemble better with each other than the best one with either of them. Especially for a large collection of good models, the best model may be unnecessary.\n\nPS I don't think weights need to sum up to 1 and negative weights are OK. I have discussed that in one of my [old notebooks](https://www.kaggle.com/code/tilii7/cross-validation-weighted-linear-blending-errors).",
    "1857122": "Thanks for your answer.",
    "1857904": "I've just used different types of tree models in my ensemble. Should I include linear models as well?",
    "1858293": "Diverse models are essential for good ensembling, and it is important to include them no matter how they were obtained. That goes even for models that have relatively low scores.\n\nLinear models are very likely to be different from tree models, and are generally good candidates to improve ensemble diversity. But any two models can be different enough even if both of them are tree-based. For example, if they were made using different features and different hyperparameters, there is a good chance that even two LGB models could be different enough."
  },
  "source": "meta"
}