{
  "id": 341222,
  "title": "Better Ensembling Methods",
  "url": "/competitions/amex-default-prediction/discussion/341222",
  "author_name": "",
  "post_date": "2022-08-01T21:29:40.460025800Z",
  "votes": -1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hey everyone, I am relatively new to Kaggle and also many ensemble methods.</p>\n<p>A lot of what I have seen and read involves simply training many different diverse models and averaging<br>\nthe predictions. I get that this can reduce variance to some degree but in general it seems very<br>\nshallow and weak.</p>\n<p>I have also seen model stacking methods in which many diverse models are trained, (say models <br>\nA, B, and C on some train set), and then a meta model is trained using a holdout set on the predictions<br>\nof models A, B, and C (e.g. logistic regression).</p>\n<p>This seems equally stupid to me. The idea that you can identify which model to trust when making<br>\na prediction simply based of its confidence seems dubious.</p>\n<p>I know gradient boosting works very well for <em>many</em> weak learners, but I feel like there must be better ways to <br>\nensemble a small number of strong models. At the very least, a meta model should also have the original<br>\ninput included in it to determine under which circumstances to trust different predictors. (LMK if anyone has<br>\ntried attention mechanisms)</p>\n<p>I'm curious if any of you more experienced Kagglers have any insights on this.</p>\n<p>Thanks,</p>\n<p>Duck</p>",
  "messages": [
    {
      "id": "1880611",
      "postDate": "08/01/2022 21:29:40",
      "content": "<p>Hey everyone, I am relatively new to Kaggle and also many ensemble methods.</p>\n<p>A lot of what I have seen and read involves simply training many different diverse models and averaging<br>\nthe predictions. I get that this can reduce variance to some degree but in general it seems very<br>\nshallow and weak.</p>\n<p>I have also seen model stacking methods in which many diverse models are trained, (say models <br>\nA, B, and C on some train set), and then a meta model is trained using a holdout set on the predictions<br>\nof models A, B, and C (e.g. logistic regression).</p>\n<p>This seems equally stupid to me. The idea that you can identify which model to trust when making<br>\na prediction simply based of its confidence seems dubious.</p>\n<p>I know gradient boosting works very well for <em>many</em> weak learners, but I feel like there must be better ways to <br>\nensemble a small number of strong models. At the very least, a meta model should also have the original<br>\ninput included in it to determine under which circumstances to trust different predictors. (LMK if anyone has<br>\ntried attention mechanisms)</p>\n<p>I'm curious if any of you more experienced Kagglers have any insights on this.</p>\n<p>Thanks,</p>\n<p>Duck</p>",
      "rawMarkdown": "Hey everyone, I am relatively new to Kaggle and also many ensemble methods.\n\nA lot of what I have seen and read involves simply training many different diverse models and averaging\nthe predictions. I get that this can reduce variance to some degree but in general it seems very\nshallow and weak.\n\nI have also seen model stacking methods in which many diverse models are trained, (say models \nA, B, and C on some train set), and then a meta model is trained using a holdout set on the predictions\nof models A, B, and C (e.g. logistic regression).\n\nThis seems equally stupid to me. The idea that you can identify which model to trust when making\na prediction simply based of its confidence seems dubious.\n\nI know gradient boosting works very well for *many* weak learners, but I feel like there must be better ways to \nensemble a small number of strong models. At the very least, a meta model should also have the original\ninput included in it to determine under which circumstances to trust different predictors. (LMK if anyone has\ntried attention mechanisms)\n\nI'm curious if any of you more experienced Kagglers have any insights on this.\n\nThanks,\n\nDuck",
      "votes": null
    },
    {
      "id": "1881055",
      "postDate": "08/02/2022 08:10:34",
      "content": "<p>See this discussion where <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> links to the ultimate kaggle ensemble <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729\" target=\"_blank\">guide</a>. It is the best explanation I have seen.</p>",
      "rawMarkdown": "See this discussion where @carlmcbrideellis links to the ultimate kaggle ensemble [guide](https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729). It is the best explanation I have seen.",
      "votes": null
    },
    {
      "id": "1881776",
      "postDate": "08/02/2022 19:08:57",
      "content": "<p>There is a <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/341227\" target=\"_blank\"><strong>nice post</strong></a> that came out recently and explains some of the issues you raised.</p>\n<blockquote>\n  <p>This seems equally stupid to me. The idea that you can identify which model to trust when making a prediction simply based of its confidence seems dubious.</p>\n</blockquote>\n<p>There is a wealth of precedent that this idea is not stupid, and in fact most of Kaggle winners have used it in some shape or form. But it has to be done right, and I am not convinced from your explanation that you understand the right way of doing it. I suggest you read the linked posts that are already in this thread, and I'd be happy to revisit if you still need an explanation why this works.</p>\n<p>I could be wrong, but you are giving an impression that it is worth combining only strong models - at least that is how your inquiry is worded. This is in part a self-promotion, but I suggest that you read <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058\" target=\"_blank\"><strong>this post</strong></a> as it may help you understand why model diversity is more important that the strength of individual models. I will give you an example from this competition. I combine two 0.795 models (both from LGB runs and non-identical, though fairly similar) and get 0.795 on the LB again. Then I add two 0.790 models (both from neural networks, very diverse to the first two) and get 0.797.</p>",
      "rawMarkdown": "There is a [**nice post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/341227) that came out recently and explains some of the issues you raised.\n\n> This seems equally stupid to me. The idea that you can identify which model to trust when making a prediction simply based of its confidence seems dubious.\n\nThere is a wealth of precedent that this idea is not stupid, and in fact most of Kaggle winners have used it in some shape or form. But it has to be done right, and I am not convinced from your explanation that you understand the right way of doing it. I suggest you read the linked posts that are already in this thread, and I'd be happy to revisit if you still need an explanation why this works.\n\nI could be wrong, but you are giving an impression that it is worth combining only strong models - at least that is how your inquiry is worded. This is in part a self-promotion, but I suggest that you read [**this post**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058) as it may help you understand why model diversity is more important that the strength of individual models. I will give you an example from this competition. I combine two 0.795 models (both from LGB runs and non-identical, though fairly similar) and get 0.795 on the LB again. Then I add two 0.790 models (both from neural networks, very diverse to the first two) and get 0.797.",
      "votes": null
    },
    {
      "id": "1881873",
      "postDate": "08/02/2022 22:03:04",
      "content": "<p>Thanks I'll check it out!</p>",
      "rawMarkdown": "Thanks I'll check it out!",
      "votes": null
    },
    {
      "id": "1881875",
      "postDate": "08/02/2022 22:06:41",
      "content": "<p>Thanks for replying and thanks for the reading material!</p>\n<p>It is abundantly clear that ensembling works in various forms. I guess I am still struggling on model stacking and ensembling of stronger models. </p>\n<p>I understand that diversity is important, I just feel like there should be better ways of combining<br>\ndiverse models than simple weighted averaging. <br>\nLike trying to create a multiple experts type of model where the weighting is based off the inputs.</p>\n<p>Still working it out though. I'll read the materials you sent and think it over some more.</p>\n<p>FINAL NOTE: Based off the response it seems my post had some unintended hostility. I did not mean to criticize<br>\nany Kagglers at all. I just want to dive deeper and try to understand the current methods and see if there are better ones. :)</p>",
      "rawMarkdown": "Thanks for replying and thanks for the reading material!\n\nIt is abundantly clear that ensembling works in various forms. I guess I am still struggling on model stacking and ensembling of stronger models. \n\nI understand that diversity is important, I just feel like there should be better ways of combining\ndiverse models than simple weighted averaging. \nLike trying to create a multiple experts type of model where the weighting is based off the inputs.\n\nStill working it out though. I'll read the materials you sent and think it over some more.\n\nFINAL NOTE: Based off the response it seems my post had some unintended hostility. I did not mean to criticize\nany Kagglers at all. I just want to dive deeper and try to understand the current methods and see if there are better ones. :)",
      "votes": null
    },
    {
      "id": "1881922",
      "postDate": "08/02/2022 23:29:58",
      "content": "<blockquote>\n  <p>It is abundantly clear that ensembling works in various forms. I guess I am still struggling on model stacking and ensembling of stronger models.</p>\n</blockquote>\n<p>Again, not sure why you are stuck on this concept of combining stronger models. If two strong models are highly correlated, they will not ensemble well. See for example how my explanation <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337610#1864201\" target=\"_blank\"><strong>here</strong></a>. Just in this competition I have written 3 posts on the subject of ensembling - see <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337610\" target=\"_blank\"><strong>post 1</strong></a>, <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337813\" target=\"_blank\"><strong>post 2</strong></a>, and <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337188\" target=\"_blank\"><strong>post 3</strong></a>. There are exceptions to everything, but in general it is better to ensemble two diverse models of which one is strong and the other one is only decent, compared to two strong models that are very similar to each other. In other words, rather than focusing on making a bunch of similar models (similar in terms of quality and the underlying predictive methodology), it may be more productive to create a few diverse models, even if they are less strong compared to the top model.</p>\n<p>In the post 3 linked above I answer this question.</p>\n<blockquote>\n  <p>I understand that diversity is important, I just feel like there should be better ways of combining diverse models than simple weighted averaging. Like trying to create a multiple experts type of model where the weighting is based off the inputs.</p>\n</blockquote>\n<p>Yes, there is a better way, which is using out-of-fold (OOF) predictions. That concept is explained in several posts that have already been linked here, but this is the gist: during N-fold training, for each fold the training is made on a N-1/N subset of data, and the prediction is made on a 1/N subset of training data, plus on test data. If we do this N times, we will have a complete and unbiased prediction for the train data (N * 1/N folds equals one full prediction for train data) and N predictions for the test data (so we sum them up take the average). Now we can use these OOF files to inform how much weight is given to each model during ensembling. Generally speaking, models with higher OOF scores get larger weight during ensembling, though there are exceptions to that as well. The point is that there is no need for simple averaging (even though it works) or to guess model weights (also works sometimes) because we can do informed model weighting.</p>",
      "rawMarkdown": "> It is abundantly clear that ensembling works in various forms. I guess I am still struggling on model stacking and ensembling of stronger models.\n\nAgain, not sure why you are stuck on this concept of combining stronger models. If two strong models are highly correlated, they will not ensemble well. See for example how my explanation [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337610#1864201). Just in this competition I have written 3 posts on the subject of ensembling - see [**post 1**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337610), [**post 2**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337813), and [**post 3**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337188). There are exceptions to everything, but in general it is better to ensemble two diverse models of which one is strong and the other one is only decent, compared to two strong models that are very similar to each other. In other words, rather than focusing on making a bunch of similar models (similar in terms of quality and the underlying predictive methodology), it may be more productive to create a few diverse models, even if they are less strong compared to the top model.\n\nIn the post 3 linked above I answer this question.\n\n> I understand that diversity is important, I just feel like there should be better ways of combining diverse models than simple weighted averaging. Like trying to create a multiple experts type of model where the weighting is based off the inputs.\n\nYes, there is a better way, which is using out-of-fold (OOF) predictions. That concept is explained in several posts that have already been linked here, but this is the gist: during N-fold training, for each fold the training is made on a N-1/N subset of data, and the prediction is made on a 1/N subset of training data, plus on test data. If we do this N times, we will have a complete and unbiased prediction for the train data (N * 1/N folds equals one full prediction for train data) and N predictions for the test data (so we sum them up take the average). Now we can use these OOF files to inform how much weight is given to each model during ensembling. Generally speaking, models with higher OOF scores get larger weight during ensembling, though there are exceptions to that as well. The point is that there is no need for simple averaging (even though it works) or to guess model weights (also works sometimes) because we can do informed model weighting.",
      "votes": null
    },
    {
      "id": "1881931",
      "postDate": "08/02/2022 23:53:23",
      "content": "<p>Hmm okay. I guess in my mind diverse and strong were not mutually exclusive, but you are suggesting that this may not be the case. Also I have been using OOF predictions. </p>\n<p>My original question (probably not clearly phrased) was regarding a specific method of informed model weighting. </p>\n<p>The methods you are referring to sound a lot like the random forest algorithm and rely on model performance on the folds (or some variant of this) to weight the models, but I was wondering if there were a way to effectively learn which models performed better under different circumstances.</p>\n<p>A basic example of this might be to say \"the meta model learned that when the input data belongs to categories A, B, &amp; C and value D is greater than 2.4, then the NN model tends to outperform the gradient boosting models.\" This would require however that the meta model also takes in the original input data to learn when to trust each of the first level models.</p>",
      "rawMarkdown": "Hmm okay. I guess in my mind diverse and strong were not mutually exclusive, but you are suggesting that this may not be the case. Also I have been using OOF predictions. \n\nMy original question (probably not clearly phrased) was regarding a specific method of informed model weighting. \n\nThe methods you are referring to sound a lot like the random forest algorithm and rely on model performance on the folds (or some variant of this) to weight the models, but I was wondering if there were a way to effectively learn which models performed better under different circumstances.\n\nA basic example of this might be to say \"the meta model learned that when the input data belongs to categories A, B, & C and value D is greater than 2.4, then the NN model tends to outperform the gradient boosting models.\" This would require however that the meta model also takes in the original input data to learn when to trust each of the first level models.",
      "votes": null
    },
    {
      "id": "1881961",
      "postDate": "08/03/2022 01:15:11",
      "content": "<blockquote>\n  <p>I guess in my mind diverse and strong were not mutually exclusive, but you are suggesting that this may not be the case.</p>\n</blockquote>\n<p>Never said anything of the sort. If you can get two strong and also diverse models, that is definitely better than one strong and one decent model, even if diverse. The point is that diversity of models is more important than \"the strength,\" but more power to you if you can get models in that are both strong and diverse.</p>\n<p>I will refrain from commenting on the rest of your interpretations, as I am spending a lot of time and seemingly not making much or inroads. You have plenty of information in this thread and all the links to do proper model ensembling. This is a known and well studied area of ML, and in your shoes I would try to understand and learn to apply the existing approaches before inventing something new.</p>",
      "rawMarkdown": "> I guess in my mind diverse and strong were not mutually exclusive, but you are suggesting that this may not be the case.\n\nNever said anything of the sort. If you can get two strong and also diverse models, that is definitely better than one strong and one decent model, even if diverse. The point is that diversity of models is more important than \"the strength,\" but more power to you if you can get models in that are both strong and diverse.\n\nI will refrain from commenting on the rest of your interpretations, as I am spending a lot of time and seemingly not making much or inroads. You have plenty of information in this thread and all the links to do proper model ensembling. This is a known and well studied area of ML, and in your shoes I would try to understand and learn to apply the existing approaches before inventing something new.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1881055,
      "author_name": "jakelj",
      "author_url": "",
      "post_date": "08/02/2022 08:10:34",
      "content": "<p>See this discussion where <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> links to the ultimate kaggle ensemble <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729\" target=\"_blank\">guide</a>. It is the best explanation I have seen.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1881873,
          "author_name": "jacobmehlman",
          "author_url": "",
          "post_date": "08/02/2022 22:03:04",
          "content": "<p>Thanks I'll check it out!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1881776,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "08/02/2022 19:08:57",
      "content": "<p>There is a <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/341227\" target=\"_blank\"><strong>nice post</strong></a> that came out recently and explains some of the issues you raised.</p>\n<blockquote>\n  <p>This seems equally stupid to me. The idea that you can identify which model to trust when making a prediction simply based of its confidence seems dubious.</p>\n</blockquote>\n<p>There is a wealth of precedent that this idea is not stupid, and in fact most of Kaggle winners have used it in some shape or form. But it has to be done right, and I am not convinced from your explanation that you understand the right way of doing it. I suggest you read the linked posts that are already in this thread, and I'd be happy to revisit if you still need an explanation why this works.</p>\n<p>I could be wrong, but you are giving an impression that it is worth combining only strong models - at least that is how your inquiry is worded. This is in part a self-promotion, but I suggest that you read <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058\" target=\"_blank\"><strong>this post</strong></a> as it may help you understand why model diversity is more important that the strength of individual models. I will give you an example from this competition. I combine two 0.795 models (both from LGB runs and non-identical, though fairly similar) and get 0.795 on the LB again. Then I add two 0.790 models (both from neural networks, very diverse to the first two) and get 0.797.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1881875,
          "author_name": "jacobmehlman",
          "author_url": "",
          "post_date": "08/02/2022 22:06:41",
          "content": "<p>Thanks for replying and thanks for the reading material!</p>\n<p>It is abundantly clear that ensembling works in various forms. I guess I am still struggling on model stacking and ensembling of stronger models. </p>\n<p>I understand that diversity is important, I just feel like there should be better ways of combining<br>\ndiverse models than simple weighted averaging. <br>\nLike trying to create a multiple experts type of model where the weighting is based off the inputs.</p>\n<p>Still working it out though. I'll read the materials you sent and think it over some more.</p>\n<p>FINAL NOTE: Based off the response it seems my post had some unintended hostility. I did not mean to criticize<br>\nany Kagglers at all. I just want to dive deeper and try to understand the current methods and see if there are better ones. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1881922,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "08/02/2022 23:29:58",
          "content": "<blockquote>\n  <p>It is abundantly clear that ensembling works in various forms. I guess I am still struggling on model stacking and ensembling of stronger models.</p>\n</blockquote>\n<p>Again, not sure why you are stuck on this concept of combining stronger models. If two strong models are highly correlated, they will not ensemble well. See for example how my explanation <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337610#1864201\" target=\"_blank\"><strong>here</strong></a>. Just in this competition I have written 3 posts on the subject of ensembling - see <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337610\" target=\"_blank\"><strong>post 1</strong></a>, <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337813\" target=\"_blank\"><strong>post 2</strong></a>, and <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/337188\" target=\"_blank\"><strong>post 3</strong></a>. There are exceptions to everything, but in general it is better to ensemble two diverse models of which one is strong and the other one is only decent, compared to two strong models that are very similar to each other. In other words, rather than focusing on making a bunch of similar models (similar in terms of quality and the underlying predictive methodology), it may be more productive to create a few diverse models, even if they are less strong compared to the top model.</p>\n<p>In the post 3 linked above I answer this question.</p>\n<blockquote>\n  <p>I understand that diversity is important, I just feel like there should be better ways of combining diverse models than simple weighted averaging. Like trying to create a multiple experts type of model where the weighting is based off the inputs.</p>\n</blockquote>\n<p>Yes, there is a better way, which is using out-of-fold (OOF) predictions. That concept is explained in several posts that have already been linked here, but this is the gist: during N-fold training, for each fold the training is made on a N-1/N subset of data, and the prediction is made on a 1/N subset of training data, plus on test data. If we do this N times, we will have a complete and unbiased prediction for the train data (N * 1/N folds equals one full prediction for train data) and N predictions for the test data (so we sum them up take the average). Now we can use these OOF files to inform how much weight is given to each model during ensembling. Generally speaking, models with higher OOF scores get larger weight during ensembling, though there are exceptions to that as well. The point is that there is no need for simple averaging (even though it works) or to guess model weights (also works sometimes) because we can do informed model weighting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1881931,
          "author_name": "jacobmehlman",
          "author_url": "",
          "post_date": "08/02/2022 23:53:23",
          "content": "<p>Hmm okay. I guess in my mind diverse and strong were not mutually exclusive, but you are suggesting that this may not be the case. Also I have been using OOF predictions. </p>\n<p>My original question (probably not clearly phrased) was regarding a specific method of informed model weighting. </p>\n<p>The methods you are referring to sound a lot like the random forest algorithm and rely on model performance on the folds (or some variant of this) to weight the models, but I was wondering if there were a way to effectively learn which models performed better under different circumstances.</p>\n<p>A basic example of this might be to say \"the meta model learned that when the input data belongs to categories A, B, &amp; C and value D is greater than 2.4, then the NN model tends to outperform the gradient boosting models.\" This would require however that the meta model also takes in the original input data to learn when to trust each of the first level models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1881961,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "08/03/2022 01:15:11",
          "content": "<blockquote>\n  <p>I guess in my mind diverse and strong were not mutually exclusive, but you are suggesting that this may not be the case.</p>\n</blockquote>\n<p>Never said anything of the sort. If you can get two strong and also diverse models, that is definitely better than one strong and one decent model, even if diverse. The point is that diversity of models is more important than \"the strength,\" but more power to you if you can get models in that are both strong and diverse.</p>\n<p>I will refrain from commenting on the rest of your interpretations, as I am spending a lot of time and seemingly not making much or inroads. You have plenty of information in this thread and all the links to do proper model ensembling. This is a known and well studied area of ML, and in your shoes I would try to understand and learn to apply the existing approaches before inventing something new.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1880611": "Hey everyone, I am relatively new to Kaggle and also many ensemble methods.\n\nA lot of what I have seen and read involves simply training many different diverse models and averaging\nthe predictions. I get that this can reduce variance to some degree but in general it seems very\nshallow and weak.\n\nI have also seen model stacking methods in which many diverse models are trained, (say models \nA, B, and C on some train set), and then a meta model is trained using a holdout set on the predictions\nof models A, B, and C (e.g. logistic regression).\n\nThis seems equally stupid to me. The idea that you can identify which model to trust when making\na prediction simply based of its confidence seems dubious.\n\nI know gradient boosting works very well for *many* weak learners, but I feel like there must be better ways to \nensemble a small number of strong models. At the very least, a meta model should also have the original\ninput included in it to determine under which circumstances to trust different predictors. (LMK if anyone has\ntried attention mechanisms)\n\nI'm curious if any of you more experienced Kagglers have any insights on this.\n\nThanks,\n\nDuck",
    "1881055": "See this discussion where @carlmcbrideellis links to the ultimate kaggle ensemble [guide](https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729). It is the best explanation I have seen.",
    "1881776": "There is a [**nice post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/341227) that came out recently and explains some of the issues you raised.\n\n> This seems equally stupid to me. The idea that you can identify which model to trust when making a prediction simply based of its confidence seems dubious.\n\nThere is a wealth of precedent that this idea is not stupid, and in fact most of Kaggle winners have used it in some shape or form. But it has to be done right, and I am not convinced from your explanation that you understand the right way of doing it. I suggest you read the linked posts that are already in this thread, and I'd be happy to revisit if you still need an explanation why this works.\n\nI could be wrong, but you are giving an impression that it is worth combining only strong models - at least that is how your inquiry is worded. This is in part a self-promotion, but I suggest that you read [**this post**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058) as it may help you understand why model diversity is more important that the strength of individual models. I will give you an example from this competition. I combine two 0.795 models (both from LGB runs and non-identical, though fairly similar) and get 0.795 on the LB again. Then I add two 0.790 models (both from neural networks, very diverse to the first two) and get 0.797.",
    "1881873": "Thanks I'll check it out!",
    "1881875": "Thanks for replying and thanks for the reading material!\n\nIt is abundantly clear that ensembling works in various forms. I guess I am still struggling on model stacking and ensembling of stronger models. \n\nI understand that diversity is important, I just feel like there should be better ways of combining\ndiverse models than simple weighted averaging. \nLike trying to create a multiple experts type of model where the weighting is based off the inputs.\n\nStill working it out though. I'll read the materials you sent and think it over some more.\n\nFINAL NOTE: Based off the response it seems my post had some unintended hostility. I did not mean to criticize\nany Kagglers at all. I just want to dive deeper and try to understand the current methods and see if there are better ones. :)",
    "1881922": "> It is abundantly clear that ensembling works in various forms. I guess I am still struggling on model stacking and ensembling of stronger models.\n\nAgain, not sure why you are stuck on this concept of combining stronger models. If two strong models are highly correlated, they will not ensemble well. See for example how my explanation [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337610#1864201). Just in this competition I have written 3 posts on the subject of ensembling - see [**post 1**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337610), [**post 2**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337813), and [**post 3**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/337188). There are exceptions to everything, but in general it is better to ensemble two diverse models of which one is strong and the other one is only decent, compared to two strong models that are very similar to each other. In other words, rather than focusing on making a bunch of similar models (similar in terms of quality and the underlying predictive methodology), it may be more productive to create a few diverse models, even if they are less strong compared to the top model.\n\nIn the post 3 linked above I answer this question.\n\n> I understand that diversity is important, I just feel like there should be better ways of combining diverse models than simple weighted averaging. Like trying to create a multiple experts type of model where the weighting is based off the inputs.\n\nYes, there is a better way, which is using out-of-fold (OOF) predictions. That concept is explained in several posts that have already been linked here, but this is the gist: during N-fold training, for each fold the training is made on a N-1/N subset of data, and the prediction is made on a 1/N subset of training data, plus on test data. If we do this N times, we will have a complete and unbiased prediction for the train data (N * 1/N folds equals one full prediction for train data) and N predictions for the test data (so we sum them up take the average). Now we can use these OOF files to inform how much weight is given to each model during ensembling. Generally speaking, models with higher OOF scores get larger weight during ensembling, though there are exceptions to that as well. The point is that there is no need for simple averaging (even though it works) or to guess model weights (also works sometimes) because we can do informed model weighting.",
    "1881931": "Hmm okay. I guess in my mind diverse and strong were not mutually exclusive, but you are suggesting that this may not be the case. Also I have been using OOF predictions. \n\nMy original question (probably not clearly phrased) was regarding a specific method of informed model weighting. \n\nThe methods you are referring to sound a lot like the random forest algorithm and rely on model performance on the folds (or some variant of this) to weight the models, but I was wondering if there were a way to effectively learn which models performed better under different circumstances.\n\nA basic example of this might be to say \"the meta model learned that when the input data belongs to categories A, B, & C and value D is greater than 2.4, then the NN model tends to outperform the gradient boosting models.\" This would require however that the meta model also takes in the original input data to learn when to trust each of the first level models.",
    "1881961": "> I guess in my mind diverse and strong were not mutually exclusive, but you are suggesting that this may not be the case.\n\nNever said anything of the sort. If you can get two strong and also diverse models, that is definitely better than one strong and one decent model, even if diverse. The point is that diversity of models is more important than \"the strength,\" but more power to you if you can get models in that are both strong and diverse.\n\nI will refrain from commenting on the rest of your interpretations, as I am spending a lot of time and seemingly not making much or inroads. You have plenty of information in this thread and all the links to do proper model ensembling. This is a known and well studied area of ML, and in your shoes I would try to understand and learn to apply the existing approaches before inventing something new."
  },
  "source": "meta"
}