{
  "id": 220104,
  "title": "Noob question on ensemble model",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220104",
  "author_name": "",
  "post_date": "2021-02-17T13:25:57.377742Z",
  "votes": 3,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>My current model is an ensemble of 3 models and as a score of 0.903<br>\nI added a different 4th model with a score of 0.867 with uses a different technique from the other 3</p>\n<p>When I make an ensemble model with all 4 models, my score dips and is in the range of  0.602 to 0.796. This is worse than the individual  scores of any of the 4 models.</p>\n<p>Can someone help me understand why this may happen?</p>\n<p>Thanks for answering!</p>",
  "messages": [
    {
      "id": "1206617",
      "postDate": "02/17/2021 13:25:57",
      "content": "<p>Hi,</p>\n<p>My current model is an ensemble of 3 models and as a score of 0.903<br>\nI added a different 4th model with a score of 0.867 with uses a different technique from the other 3</p>\n<p>When I make an ensemble model with all 4 models, my score dips and is in the range of  0.602 to 0.796. This is worse than the individual  scores of any of the 4 models.</p>\n<p>Can someone help me understand why this may happen?</p>\n<p>Thanks for answering!</p>",
      "rawMarkdown": "Hi,\n\nMy current model is an ensemble of 3 models and as a score of 0.903\nI added a different 4th model with a score of 0.867 with uses a different technique from the other 3\n\nWhen I make an ensemble model with all 4 models, my score dips and is in the range of  0.602 to 0.796. This is worse than the individual  scores of any of the 4 models.\n\nCan someone help me understand why this may happen?\n\nThanks for answering!",
      "votes": null
    },
    {
      "id": "1206941",
      "postDate": "02/17/2021 16:13:18",
      "content": "<p>What type of ensemble method are you using?</p>",
      "rawMarkdown": "What type of ensemble method are you using?",
      "votes": null
    },
    {
      "id": "1207040",
      "postDate": "02/17/2021 17:37:37",
      "content": "<p>I take the predictions of each model and do a weighted average, and then apply softmax</p>",
      "rawMarkdown": "I take the predictions of each model and do a weighted average, and then apply softmax",
      "votes": null
    },
    {
      "id": "1207146",
      "postDate": "02/17/2021 18:23:38",
      "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> Did you try without weighting? Try it once with all having the same weights</p>",
      "rawMarkdown": "kmldas Did you try without weighting? Try it once with all having the same weights",
      "votes": null
    },
    {
      "id": "1207164",
      "postDate": "02/17/2021 18:27:30",
      "content": "<p>yes, tried a simple average as well… surprising to see… likely will end up being some stupid coding error on my part!</p>",
      "rawMarkdown": "yes, tried a simple average as well... surprising to see... likely will end up being some stupid coding error on my part!",
      "votes": null
    },
    {
      "id": "1207319",
      "postDate": "02/17/2021 19:51:52",
      "content": "<p><strong>François Chollet</strong>(Creator of Keras) explained it as : Ensembling relies on the assumption that different good models trained independently are likely to be good for different reasons: each model looks at slightly different aspects of the data to make its predictions, getting part of the “truth” but not all of it. You may be familiar with the ancient parable of the blind men and the elephant: a group of blind men come across an elephant for the first time and try to understand what the elephant is by touching it. Each man touches a different part of the elephant’s body—just one part, such as the trunk or a leg. Then the men describe to<br>\neach other what an elephant is: “It’s like a snake,” “Like a pillar or a tree,” and so on.The blind men are essentially machine-learning models trying to understand the manifold of the training data, each from its own perspective, using its own assumptions(provided by the unique architecture of the model and the unique random weight initialization). Each of them gets part of the truth of the data, but not the whole truth. By pooling their perspectives together, you can get a far more accurate description of the data. The elephant is a combination of parts: not any single blind man gets it quite<br>\nright, but, interviewed together, they can tell a fairly accurate story.</p>\n<p>Let’s use classification as an example. The easiest way to pool the predictions of a set<br>\nof classifiers (to ensemble the classifiers) is to average their predictions at inference time:<br>\nUse four different models to compute initial predictions.</p>\n<pre><code>preds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n</code></pre>\n<p>This new prediction array should be more accurate than any of the initial ones.</p>\n<p><code>final_preds = 0.25 * (preds_a + preds_b + preds_c + preds_d)</code></p>\n<p>This will work only if the classifiers are more or less equally good. If one of them is significantly worse than the others, the final predictions may not be as good as the best classifier of the group.<br>\nA smarter way to ensemble classifiers is to do a weighted average, where the weights are learned on the validation data typically, the better classifiers are given a higher weight, and the worse classifiers are given a lower weight. To search for a good set of ensembling weights, you can use random search or a simple optimization algorithm such as Nelder-Mead:</p>\n<pre><code>preds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n</code></pre>\n<p>These weights (0.5, 0.25,0.1, 0.15) are assumed to be learned empirically.<br>\n<code>final_preds = 0.5 * preds_a + 0.25 * preds_b + 0.1 * preds_c + 0.15 * preds_d</code></p>",
      "rawMarkdown": "**François Chollet**(Creator of Keras) explained it as : Ensembling relies on the assumption that different good models trained independently are likely to be good for different reasons: each model looks at slightly different aspects of the data to make its predictions, getting part of the “truth” but not all of it. You may be familiar with the ancient parable of the blind men and the elephant: a group of blind men come across an elephant for the first time and try to understand what the elephant is by touching it. Each man touches a different part of the elephant’s body—just one part, such as the trunk or a leg. Then the men describe to\neach other what an elephant is: “It’s like a snake,” “Like a pillar or a tree,” and so on.The blind men are essentially machine-learning models trying to understand the manifold of the training data, each from its own perspective, using its own assumptions(provided by the unique architecture of the model and the unique random weight initialization). Each of them gets part of the truth of the data, but not the whole truth. By pooling their perspectives together, you can get a far more accurate description of the data. The elephant is a combination of parts: not any single blind man gets it quite\nright, but, interviewed together, they can tell a fairly accurate story.\n\nLet’s use classification as an example. The easiest way to pool the predictions of a set\nof classifiers (to ensemble the classifiers) is to average their predictions at inference time:\nUse four different models to compute initial predictions.\n\n```\npreds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n```\n\nThis new prediction array should be more accurate than any of the initial ones.\n\n`final_preds = 0.25 * (preds_a + preds_b + preds_c + preds_d)`\n\nThis will work only if the classifiers are more or less equally good. If one of them is significantly worse than the others, the final predictions may not be as good as the best classifier of the group.\nA smarter way to ensemble classifiers is to do a weighted average, where the weights are learned on the validation data typically, the better classifiers are given a higher weight, and the worse classifiers are given a lower weight. To search for a good set of ensembling weights, you can use random search or a simple optimization algorithm such as Nelder-Mead:\n```\n\npreds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n\n```\nThese weights (0.5, 0.25,0.1, 0.15) are assumed to be learned empirically.\n`final_preds = 0.5 * preds_a + 0.25 * preds_b + 0.1 * preds_c + 0.15 * preds_d`",
      "votes": null
    },
    {
      "id": "1207580",
      "postDate": "02/17/2021 23:40:03",
      "content": "<p>You mean the order of models can affect predict accuracy?</p>",
      "rawMarkdown": "You mean the order of models can affect predict accuracy?",
      "votes": null
    },
    {
      "id": "1207803",
      "postDate": "02/18/2021 03:36:37",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/vickygoyal\" target=\"_blank\">@vickygoyal</a> for the detailed explanation! </p>",
      "rawMarkdown": "Thanks @vickygoyal for the detailed explanation!",
      "votes": null
    },
    {
      "id": "1207882",
      "postDate": "02/18/2021 04:35:07",
      "content": "<p>Did you sort all your predicitons before blending?</p>",
      "rawMarkdown": "Did you sort all your predicitons before blending?",
      "votes": null
    },
    {
      "id": "1208529",
      "postDate": "02/18/2021 10:18:23",
      "content": "<p>Wew. Very good explanation </p>",
      "rawMarkdown": "Wew. Very good explanation",
      "votes": null
    },
    {
      "id": "1208813",
      "postDate": "02/18/2021 13:52:04",
      "content": "<p>hmm I had faced the same, but I thought some bug was in the code.. maybe the ordering of rows matters (?) as <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> pointed out.. how we can verify this in the hidden test ? </p>",
      "rawMarkdown": "hmm I had faced the same, but I thought some bug was in the code.. maybe the ordering of rows matters (?) as @yosukeyama pointed out.. how we can verify this in the hidden test ?",
      "votes": null
    },
    {
      "id": "1208860",
      "postDate": "02/18/2021 14:17:41",
      "content": "<p>Try averaging the softmax predictions instead of the pre-softmax predictions. This is a bit more robust I think. Otherwise, you might accidentally give some models much more weight than others if you average pre-softmax predictions.</p>",
      "rawMarkdown": "Try averaging the softmax predictions instead of the pre-softmax predictions. This is a bit more robust I think. Otherwise, you might accidentally give some models much more weight than others if you average pre-softmax predictions.",
      "votes": null
    },
    {
      "id": "1208884",
      "postDate": "02/18/2021 14:37:01",
      "content": "<p>In my case I use the <code>softmax</code> preds (ie. probs from all classes sum to 1) from 3 models and have tried with weights and plain average as well.. LB 0.439/0.468 <br>\nps: the individual model scores: 0.898 - 0.90x</p>\n<p>one way to verify if it is bug or not is to try the same ensemble with CV (I'll get back to that..)</p>",
      "rawMarkdown": "In my case I use the `softmax` preds (ie. probs from all classes sum to 1) from 3 models and have tried with weights and plain average as well.. LB 0.439/0.468 \nps: the individual model scores: 0.898 - 0.90x\n\none way to verify if it is bug or not is to try the same ensemble with CV (I'll get back to that..)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1206941,
      "author_name": "mohneesh7",
      "author_url": "",
      "post_date": "02/17/2021 16:13:18",
      "content": "<p>What type of ensemble method are you using?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207040,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/17/2021 17:37:37",
          "content": "<p>I take the predictions of each model and do a weighted average, and then apply softmax</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207146,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "02/17/2021 18:23:38",
          "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> Did you try without weighting? Try it once with all having the same weights</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207164,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/17/2021 18:27:30",
          "content": "<p>yes, tried a simple average as well… surprising to see… likely will end up being some stupid coding error on my part!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208860,
          "author_name": "raivokoot",
          "author_url": "",
          "post_date": "02/18/2021 14:17:41",
          "content": "<p>Try averaging the softmax predictions instead of the pre-softmax predictions. This is a bit more robust I think. Otherwise, you might accidentally give some models much more weight than others if you average pre-softmax predictions.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1208884,
              "author_name": "imeintanis",
              "author_url": "",
              "post_date": "02/18/2021 14:37:01",
              "content": "<p>In my case I use the <code>softmax</code> preds (ie. probs from all classes sum to 1) from 3 models and have tried with weights and plain average as well.. LB 0.439/0.468 <br>\nps: the individual model scores: 0.898 - 0.90x</p>\n<p>one way to verify if it is bug or not is to try the same ensemble with CV (I'll get back to that..)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1207319,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/17/2021 19:51:52",
      "content": "<p><strong>François Chollet</strong>(Creator of Keras) explained it as : Ensembling relies on the assumption that different good models trained independently are likely to be good for different reasons: each model looks at slightly different aspects of the data to make its predictions, getting part of the “truth” but not all of it. You may be familiar with the ancient parable of the blind men and the elephant: a group of blind men come across an elephant for the first time and try to understand what the elephant is by touching it. Each man touches a different part of the elephant’s body—just one part, such as the trunk or a leg. Then the men describe to<br>\neach other what an elephant is: “It’s like a snake,” “Like a pillar or a tree,” and so on.The blind men are essentially machine-learning models trying to understand the manifold of the training data, each from its own perspective, using its own assumptions(provided by the unique architecture of the model and the unique random weight initialization). Each of them gets part of the truth of the data, but not the whole truth. By pooling their perspectives together, you can get a far more accurate description of the data. The elephant is a combination of parts: not any single blind man gets it quite<br>\nright, but, interviewed together, they can tell a fairly accurate story.</p>\n<p>Let’s use classification as an example. The easiest way to pool the predictions of a set<br>\nof classifiers (to ensemble the classifiers) is to average their predictions at inference time:<br>\nUse four different models to compute initial predictions.</p>\n<pre><code>preds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n</code></pre>\n<p>This new prediction array should be more accurate than any of the initial ones.</p>\n<p><code>final_preds = 0.25 * (preds_a + preds_b + preds_c + preds_d)</code></p>\n<p>This will work only if the classifiers are more or less equally good. If one of them is significantly worse than the others, the final predictions may not be as good as the best classifier of the group.<br>\nA smarter way to ensemble classifiers is to do a weighted average, where the weights are learned on the validation data typically, the better classifiers are given a higher weight, and the worse classifiers are given a lower weight. To search for a good set of ensembling weights, you can use random search or a simple optimization algorithm such as Nelder-Mead:</p>\n<pre><code>preds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n</code></pre>\n<p>These weights (0.5, 0.25,0.1, 0.15) are assumed to be learned empirically.<br>\n<code>final_preds = 0.5 * preds_a + 0.25 * preds_b + 0.1 * preds_c + 0.15 * preds_d</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1207580,
          "author_name": "leejingwan",
          "author_url": "",
          "post_date": "02/17/2021 23:40:03",
          "content": "<p>You mean the order of models can affect predict accuracy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207803,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/18/2021 03:36:37",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/vickygoyal\" target=\"_blank\">@vickygoyal</a> for the detailed explanation! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208529,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "02/18/2021 10:18:23",
          "content": "<p>Wew. Very good explanation </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207882,
      "author_name": "yosukeyama",
      "author_url": "",
      "post_date": "02/18/2021 04:35:07",
      "content": "<p>Did you sort all your predicitons before blending?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208813,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "02/18/2021 13:52:04",
      "content": "<p>hmm I had faced the same, but I thought some bug was in the code.. maybe the ordering of rows matters (?) as <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> pointed out.. how we can verify this in the hidden test ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1206617": "Hi,\n\nMy current model is an ensemble of 3 models and as a score of 0.903\nI added a different 4th model with a score of 0.867 with uses a different technique from the other 3\n\nWhen I make an ensemble model with all 4 models, my score dips and is in the range of  0.602 to 0.796. This is worse than the individual  scores of any of the 4 models.\n\nCan someone help me understand why this may happen?\n\nThanks for answering!",
    "1206941": "What type of ensemble method are you using?",
    "1207040": "I take the predictions of each model and do a weighted average, and then apply softmax",
    "1207146": "kmldas Did you try without weighting? Try it once with all having the same weights",
    "1207164": "yes, tried a simple average as well... surprising to see... likely will end up being some stupid coding error on my part!",
    "1207319": "**François Chollet**(Creator of Keras) explained it as : Ensembling relies on the assumption that different good models trained independently are likely to be good for different reasons: each model looks at slightly different aspects of the data to make its predictions, getting part of the “truth” but not all of it. You may be familiar with the ancient parable of the blind men and the elephant: a group of blind men come across an elephant for the first time and try to understand what the elephant is by touching it. Each man touches a different part of the elephant’s body—just one part, such as the trunk or a leg. Then the men describe to\neach other what an elephant is: “It’s like a snake,” “Like a pillar or a tree,” and so on.The blind men are essentially machine-learning models trying to understand the manifold of the training data, each from its own perspective, using its own assumptions(provided by the unique architecture of the model and the unique random weight initialization). Each of them gets part of the truth of the data, but not the whole truth. By pooling their perspectives together, you can get a far more accurate description of the data. The elephant is a combination of parts: not any single blind man gets it quite\nright, but, interviewed together, they can tell a fairly accurate story.\n\nLet’s use classification as an example. The easiest way to pool the predictions of a set\nof classifiers (to ensemble the classifiers) is to average their predictions at inference time:\nUse four different models to compute initial predictions.\n\n```\npreds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n```\n\nThis new prediction array should be more accurate than any of the initial ones.\n\n`final_preds = 0.25 * (preds_a + preds_b + preds_c + preds_d)`\n\nThis will work only if the classifiers are more or less equally good. If one of them is significantly worse than the others, the final predictions may not be as good as the best classifier of the group.\nA smarter way to ensemble classifiers is to do a weighted average, where the weights are learned on the validation data typically, the better classifiers are given a higher weight, and the worse classifiers are given a lower weight. To search for a good set of ensembling weights, you can use random search or a simple optimization algorithm such as Nelder-Mead:\n```\n\npreds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n\n```\nThese weights (0.5, 0.25,0.1, 0.15) are assumed to be learned empirically.\n`final_preds = 0.5 * preds_a + 0.25 * preds_b + 0.1 * preds_c + 0.15 * preds_d`",
    "1207580": "You mean the order of models can affect predict accuracy?",
    "1207803": "Thanks @vickygoyal for the detailed explanation!",
    "1207882": "Did you sort all your predicitons before blending?",
    "1208529": "Wew. Very good explanation",
    "1208813": "hmm I had faced the same, but I thought some bug was in the code.. maybe the ordering of rows matters (?) as @yosukeyama pointed out.. how we can verify this in the hidden test ?",
    "1208860": "Try averaging the softmax predictions instead of the pre-softmax predictions. This is a bit more robust I think. Otherwise, you might accidentally give some models much more weight than others if you average pre-softmax predictions.",
    "1208884": "In my case I use the `softmax` preds (ie. probs from all classes sum to 1) from 3 models and have tried with weights and plain average as well.. LB 0.439/0.468 \nps: the individual model scores: 0.898 - 0.90x\n\none way to verify if it is bug or not is to try the same ensemble with CV (I'll get back to that..)"
  },
  "source": "meta"
}