{
  "id": 219882,
  "title": "Model Ensemble",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/219882",
  "author_name": "",
  "post_date": "2021-02-16T17:15:41.075142900Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Can someone explain how to ensemble multiple models? Should you simply take the average of the outputs after TTA or should there be a weight associated with one model versus another?</p>",
  "messages": [
    {
      "id": "1205442",
      "postDate": "02/16/2021 17:15:41",
      "content": "<p>Can someone explain how to ensemble multiple models? Should you simply take the average of the outputs after TTA or should there be a weight associated with one model versus another?</p>",
      "rawMarkdown": "Can someone explain how to ensemble multiple models? Should you simply take the average of the outputs after TTA or should there be a weight associated with one model versus another?",
      "votes": null
    },
    {
      "id": "1205649",
      "postDate": "02/16/2021 20:57:08",
      "content": "<p>There are two common ways to ensemble models in a classification task:</p>\n<ul>\n<li>Soft voting: Average of probabilities of each model</li>\n<li>Hard voting: Majority voting which class gets predicted the most gets choosen</li>\n</ul>\n<p>Soft voting usually results in better accuracy and should be choosen if possible. </p>\n<p>You can ensemble your predictions after TTA together for that. You can also give them different weights, I usually give the more accurate model a little higher weight than the others.</p>\n<p>If you dont want to choose weights you could train another model on top which blends all your probs (so called Stacking see: <a href=\"https://medium.com/ml-research-lab/stacking-ensemble-meta-algorithms-for-improve-predictions-f4b4cf3b9237\" target=\"_blank\">https://medium.com/ml-research-lab/stacking-ensemble-meta-algorithms-for-improve-predictions-f4b4cf3b9237</a>)</p>",
      "rawMarkdown": "There are two common ways to ensemble models in a classification task:\n- Soft voting: Average of probabilities of each model\n- Hard voting: Majority voting which class gets predicted the most gets choosen\n\nSoft voting usually results in better accuracy and should be choosen if possible. \n\nYou can ensemble your predictions after TTA together for that. You can also give them different weights, I usually give the more accurate model a little higher weight than the others.\n\nIf you dont want to choose weights you could train another model on top which blends all your probs (so called Stacking see: https://medium.com/ml-research-lab/stacking-ensemble-meta-algorithms-for-improve-predictions-f4b4cf3b9237)",
      "votes": null
    },
    {
      "id": "1205709",
      "postDate": "02/16/2021 22:49:00",
      "content": "<p>Thanks for the suggestions.</p>",
      "rawMarkdown": "Thanks for the suggestions.",
      "votes": null
    },
    {
      "id": "1206239",
      "postDate": "02/17/2021 08:00:11",
      "content": "<p>Hi! It's not an unambiguous question like using SVC or LogitRegression. For sure, a weighted prediction of other models' predictions will perform better on the most problems. There are already well-made kit1, kit2 in sklearn for ensembling classic ML algorithms. But there are no or maybe few libraries for ensembling deep learning models. So you are free to create several different DL models in TF, Torch or Keras, fine-tune them, save weights and then build another model, which will get your trained models predictions and return a final prediction. Good luck!</p>",
      "rawMarkdown": "Hi! It's not an unambiguous question like using SVC or LogitRegression. For sure, a weighted prediction of other models' predictions will perform better on the most problems. There are already well-made kit1, kit2 in sklearn for ensembling classic ML algorithms. But there are no or maybe few libraries for ensembling deep learning models. So you are free to create several different DL models in TF, Torch or Keras, fine-tune them, save weights and then build another model, which will get your trained models predictions and return a final prediction. Good luck!",
      "votes": null
    },
    {
      "id": "1206248",
      "postDate": "02/17/2021 08:01:15",
      "content": "<p>If you are interested how to ensemble, you train your models according to a selected statical algorithm (stacking, bagging and combinations). All them are opened to be read in the internet.</p>",
      "rawMarkdown": "If you are interested how to ensemble, you train your models according to a selected statical algorithm (stacking, bagging and combinations). All them are opened to be read in the internet.",
      "votes": null
    },
    {
      "id": "1206767",
      "postDate": "02/17/2021 14:55:56",
      "content": "<p>Nice explanation.<br>\nI have used both Soft voting and hard voting, Soft voting is giving me a slightly better result. Tested it on my LB score as well, soft voting is better in my case at least.</p>",
      "rawMarkdown": "Nice explanation.\nI have used both Soft voting and hard voting, Soft voting is giving me a slightly better result. Tested it on my LB score as well, soft voting is better in my case at least.",
      "votes": null
    },
    {
      "id": "1207322",
      "postDate": "02/17/2021 19:52:51",
      "content": "<p><strong>François Chollet</strong>(Creator of Keras) explained it as : Ensembling relies on the assumption that different good models trained independently are likely to be good for different reasons: each model looks at slightly different aspects of the data to make its predictions, getting part of the “truth” but not all of it. You may be familiar with the ancient parable of the blind men and the elephant: a group of blind men come across an elephant for the first time and try to understand what the elephant is by touching it. Each man touches a different part of the elephant’s body—just one part, such as the trunk or a leg. Then the men describe to<br>\neach other what an elephant is: “It’s like a snake,” “Like a pillar or a tree,” and so on.The blind men are essentially machine-learning models trying to understand the manifold of the training data, each from its own perspective, using its own assumptions(provided by the unique architecture of the model and the unique random weight initialization). Each of them gets part of the truth of the data, but not the whole truth. By pooling their perspectives together, you can get a far more accurate description of the data. The elephant is a combination of parts: not any single blind man gets it quite<br>\nright, but, interviewed together, they can tell a fairly accurate story.</p>\n<p>Let’s use classification as an example. The easiest way to pool the predictions of a set<br>\nof classifiers (to ensemble the classifiers) is to average their predictions at inference time:<br>\nUse four different models to compute initial predictions.</p>\n<pre><code>preds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n</code></pre>\n<p>This new prediction array should be more accurate than any of the initial ones.</p>\n<p><code>final_preds = 0.25 * (preds_a + preds_b + preds_c + preds_d)</code></p>\n<p>This will work only if the classifiers are more or less equally good. If one of them is significantly worse than the others, the final predictions may not be as good as the best classifier of the group.<br>\nA smarter way to ensemble classifiers is to do a weighted average, where the weights are learned on the validation data typically, the better classifiers are given a higher weight, and the worse classifiers are given a lower weight. To search for a good set of ensembling weights, you can use random search or a simple optimization algorithm such as Nelder-Mead:</p>\n<pre><code>preds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n</code></pre>\n<p>These weights (0.5, 0.25,0.1, 0.15) are assumed to be learned empirically.<br>\n<code>final_preds = 0.5 * preds_a + 0.25 * preds_b + 0.1 * preds_c + 0.15 * preds_d</code></p>",
      "rawMarkdown": "**François Chollet**(Creator of Keras) explained it as : Ensembling relies on the assumption that different good models trained independently are likely to be good for different reasons: each model looks at slightly different aspects of the data to make its predictions, getting part of the “truth” but not all of it. You may be familiar with the ancient parable of the blind men and the elephant: a group of blind men come across an elephant for the first time and try to understand what the elephant is by touching it. Each man touches a different part of the elephant’s body—just one part, such as the trunk or a leg. Then the men describe to\neach other what an elephant is: “It’s like a snake,” “Like a pillar or a tree,” and so on.The blind men are essentially machine-learning models trying to understand the manifold of the training data, each from its own perspective, using its own assumptions(provided by the unique architecture of the model and the unique random weight initialization). Each of them gets part of the truth of the data, but not the whole truth. By pooling their perspectives together, you can get a far more accurate description of the data. The elephant is a combination of parts: not any single blind man gets it quite\nright, but, interviewed together, they can tell a fairly accurate story.\n\nLet’s use classification as an example. The easiest way to pool the predictions of a set\nof classifiers (to ensemble the classifiers) is to average their predictions at inference time:\nUse four different models to compute initial predictions.\n\n```\npreds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n```\n\nThis new prediction array should be more accurate than any of the initial ones.\n\n`final_preds = 0.25 * (preds_a + preds_b + preds_c + preds_d)`\n\nThis will work only if the classifiers are more or less equally good. If one of them is significantly worse than the others, the final predictions may not be as good as the best classifier of the group.\nA smarter way to ensemble classifiers is to do a weighted average, where the weights are learned on the validation data typically, the better classifiers are given a higher weight, and the worse classifiers are given a lower weight. To search for a good set of ensembling weights, you can use random search or a simple optimization algorithm such as Nelder-Mead:\n```\n\npreds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n\n```\nThese weights (0.5, 0.25,0.1, 0.15) are assumed to be learned empirically.\n`final_preds = 0.5 * preds_a + 0.25 * preds_b + 0.1 * preds_c + 0.15 * preds_d`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1205649,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "02/16/2021 20:57:08",
      "content": "<p>There are two common ways to ensemble models in a classification task:</p>\n<ul>\n<li>Soft voting: Average of probabilities of each model</li>\n<li>Hard voting: Majority voting which class gets predicted the most gets choosen</li>\n</ul>\n<p>Soft voting usually results in better accuracy and should be choosen if possible. </p>\n<p>You can ensemble your predictions after TTA together for that. You can also give them different weights, I usually give the more accurate model a little higher weight than the others.</p>\n<p>If you dont want to choose weights you could train another model on top which blends all your probs (so called Stacking see: <a href=\"https://medium.com/ml-research-lab/stacking-ensemble-meta-algorithms-for-improve-predictions-f4b4cf3b9237\" target=\"_blank\">https://medium.com/ml-research-lab/stacking-ensemble-meta-algorithms-for-improve-predictions-f4b4cf3b9237</a>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1205709,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "02/16/2021 22:49:00",
          "content": "<p>Thanks for the suggestions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1206767,
          "author_name": "mohneesh7",
          "author_url": "",
          "post_date": "02/17/2021 14:55:56",
          "content": "<p>Nice explanation.<br>\nI have used both Soft voting and hard voting, Soft voting is giving me a slightly better result. Tested it on my LB score as well, soft voting is better in my case at least.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1206239,
      "author_name": "ivankalinchuk",
      "author_url": "",
      "post_date": "02/17/2021 08:00:11",
      "content": "<p>Hi! It's not an unambiguous question like using SVC or LogitRegression. For sure, a weighted prediction of other models' predictions will perform better on the most problems. There are already well-made kit1, kit2 in sklearn for ensembling classic ML algorithms. But there are no or maybe few libraries for ensembling deep learning models. So you are free to create several different DL models in TF, Torch or Keras, fine-tune them, save weights and then build another model, which will get your trained models predictions and return a final prediction. Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1206248,
      "author_name": "ivankalinchuk",
      "author_url": "",
      "post_date": "02/17/2021 08:01:15",
      "content": "<p>If you are interested how to ensemble, you train your models according to a selected statical algorithm (stacking, bagging and combinations). All them are opened to be read in the internet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207322,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/17/2021 19:52:51",
      "content": "<p><strong>François Chollet</strong>(Creator of Keras) explained it as : Ensembling relies on the assumption that different good models trained independently are likely to be good for different reasons: each model looks at slightly different aspects of the data to make its predictions, getting part of the “truth” but not all of it. You may be familiar with the ancient parable of the blind men and the elephant: a group of blind men come across an elephant for the first time and try to understand what the elephant is by touching it. Each man touches a different part of the elephant’s body—just one part, such as the trunk or a leg. Then the men describe to<br>\neach other what an elephant is: “It’s like a snake,” “Like a pillar or a tree,” and so on.The blind men are essentially machine-learning models trying to understand the manifold of the training data, each from its own perspective, using its own assumptions(provided by the unique architecture of the model and the unique random weight initialization). Each of them gets part of the truth of the data, but not the whole truth. By pooling their perspectives together, you can get a far more accurate description of the data. The elephant is a combination of parts: not any single blind man gets it quite<br>\nright, but, interviewed together, they can tell a fairly accurate story.</p>\n<p>Let’s use classification as an example. The easiest way to pool the predictions of a set<br>\nof classifiers (to ensemble the classifiers) is to average their predictions at inference time:<br>\nUse four different models to compute initial predictions.</p>\n<pre><code>preds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n</code></pre>\n<p>This new prediction array should be more accurate than any of the initial ones.</p>\n<p><code>final_preds = 0.25 * (preds_a + preds_b + preds_c + preds_d)</code></p>\n<p>This will work only if the classifiers are more or less equally good. If one of them is significantly worse than the others, the final predictions may not be as good as the best classifier of the group.<br>\nA smarter way to ensemble classifiers is to do a weighted average, where the weights are learned on the validation data typically, the better classifiers are given a higher weight, and the worse classifiers are given a lower weight. To search for a good set of ensembling weights, you can use random search or a simple optimization algorithm such as Nelder-Mead:</p>\n<pre><code>preds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n</code></pre>\n<p>These weights (0.5, 0.25,0.1, 0.15) are assumed to be learned empirically.<br>\n<code>final_preds = 0.5 * preds_a + 0.25 * preds_b + 0.1 * preds_c + 0.15 * preds_d</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1205442": "Can someone explain how to ensemble multiple models? Should you simply take the average of the outputs after TTA or should there be a weight associated with one model versus another?",
    "1205649": "There are two common ways to ensemble models in a classification task:\n- Soft voting: Average of probabilities of each model\n- Hard voting: Majority voting which class gets predicted the most gets choosen\n\nSoft voting usually results in better accuracy and should be choosen if possible. \n\nYou can ensemble your predictions after TTA together for that. You can also give them different weights, I usually give the more accurate model a little higher weight than the others.\n\nIf you dont want to choose weights you could train another model on top which blends all your probs (so called Stacking see: https://medium.com/ml-research-lab/stacking-ensemble-meta-algorithms-for-improve-predictions-f4b4cf3b9237)",
    "1205709": "Thanks for the suggestions.",
    "1206239": "Hi! It's not an unambiguous question like using SVC or LogitRegression. For sure, a weighted prediction of other models' predictions will perform better on the most problems. There are already well-made kit1, kit2 in sklearn for ensembling classic ML algorithms. But there are no or maybe few libraries for ensembling deep learning models. So you are free to create several different DL models in TF, Torch or Keras, fine-tune them, save weights and then build another model, which will get your trained models predictions and return a final prediction. Good luck!",
    "1206248": "If you are interested how to ensemble, you train your models according to a selected statical algorithm (stacking, bagging and combinations). All them are opened to be read in the internet.",
    "1206767": "Nice explanation.\nI have used both Soft voting and hard voting, Soft voting is giving me a slightly better result. Tested it on my LB score as well, soft voting is better in my case at least.",
    "1207322": "**François Chollet**(Creator of Keras) explained it as : Ensembling relies on the assumption that different good models trained independently are likely to be good for different reasons: each model looks at slightly different aspects of the data to make its predictions, getting part of the “truth” but not all of it. You may be familiar with the ancient parable of the blind men and the elephant: a group of blind men come across an elephant for the first time and try to understand what the elephant is by touching it. Each man touches a different part of the elephant’s body—just one part, such as the trunk or a leg. Then the men describe to\neach other what an elephant is: “It’s like a snake,” “Like a pillar or a tree,” and so on.The blind men are essentially machine-learning models trying to understand the manifold of the training data, each from its own perspective, using its own assumptions(provided by the unique architecture of the model and the unique random weight initialization). Each of them gets part of the truth of the data, but not the whole truth. By pooling their perspectives together, you can get a far more accurate description of the data. The elephant is a combination of parts: not any single blind man gets it quite\nright, but, interviewed together, they can tell a fairly accurate story.\n\nLet’s use classification as an example. The easiest way to pool the predictions of a set\nof classifiers (to ensemble the classifiers) is to average their predictions at inference time:\nUse four different models to compute initial predictions.\n\n```\npreds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n```\n\nThis new prediction array should be more accurate than any of the initial ones.\n\n`final_preds = 0.25 * (preds_a + preds_b + preds_c + preds_d)`\n\nThis will work only if the classifiers are more or less equally good. If one of them is significantly worse than the others, the final predictions may not be as good as the best classifier of the group.\nA smarter way to ensemble classifiers is to do a weighted average, where the weights are learned on the validation data typically, the better classifiers are given a higher weight, and the worse classifiers are given a lower weight. To search for a good set of ensembling weights, you can use random search or a simple optimization algorithm such as Nelder-Mead:\n```\n\npreds_a = model_a.predict(x_val)\npreds_b = model_b.predict(x_val)\npreds_c = model_c.predict(x_val)\npreds_d = model_d.predict(x_val)\n\n```\nThese weights (0.5, 0.25,0.1, 0.15) are assumed to be learned empirically.\n`final_preds = 0.5 * preds_a + 0.25 * preds_b + 0.1 * preds_c + 0.15 * preds_d`"
  },
  "source": "meta"
}