{
  "id": 348366,
  "title": "Why ensemble model is more precise not the opposite?",
  "url": "/competitions/amex-default-prediction/discussion/348366",
  "author_name": "",
  "post_date": "2022-08-28T03:36:34.370153300Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Congrats and thank you all. It's a great competition. I leaned so much from all the great shared notebooks. </p>\n<p>While from the notebooks and discussions there has some common points, more and more features or more model to ensemble.</p>\n<p>I know that more features will be helpful for digging unseen feature. </p>\n<p><strong>But why ensemble model is more precise not the opposite?</strong><br>\n(Of course, the model to be ensembled is trained well)</p>\n<p>It seems that the unstable predicted samples from one model will be corrected by other model not the opposite.</p>\n<p>Can anyone explain this in depth?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1916627",
      "postDate": "08/28/2022 03:36:34",
      "content": "<p>Congrats and thank you all. It's a great competition. I leaned so much from all the great shared notebooks. </p>\n<p>While from the notebooks and discussions there has some common points, more and more features or more model to ensemble.</p>\n<p>I know that more features will be helpful for digging unseen feature. </p>\n<p><strong>But why ensemble model is more precise not the opposite?</strong><br>\n(Of course, the model to be ensembled is trained well)</p>\n<p>It seems that the unstable predicted samples from one model will be corrected by other model not the opposite.</p>\n<p>Can anyone explain this in depth?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Congrats and thank you all. It's a great competition. I leaned so much from all the great shared notebooks. \n\nWhile from the notebooks and discussions there has some common points, more and more features or more model to ensemble.\n\nI know that more features will be helpful for digging unseen feature. \n\n**But why ensemble model is more precise not the opposite?**\n(Of course, the model to be ensembled is trained well)\n\nIt seems that the unstable predicted samples from one model will be corrected by other model not the opposite.\n\nCan anyone explain this in depth?\n\nThanks!",
      "votes": null
    },
    {
      "id": "1916648",
      "postDate": "08/28/2022 04:05:35",
      "content": "<p>Awhile back, Carl linked to a great ensembling guide, it covers this question pretty well in the first section. <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729</a> </p>",
      "rawMarkdown": "Awhile back, Carl linked to a great ensembling guide, it covers this question pretty well in the first section. https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729",
      "votes": null
    },
    {
      "id": "1916688",
      "postDate": "08/28/2022 05:28:23",
      "content": "<p>Great! Thank you very much.</p>",
      "rawMarkdown": "Great! Thank you very much.",
      "votes": null
    },
    {
      "id": "1917410",
      "postDate": "08/28/2022 17:47:43",
      "content": "<p>A simple explanation is that models that complement each other will make a good ensemble, rather than models that agree with each other. That's why diversity is always emphasized when creating ensembles.</p>\n<p>A longer explanation is <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058\" target=\"_blank\"><strong>here</strong></a>, and beware that I am sending you to one of my previous posts.</p>",
      "rawMarkdown": "A simple explanation is that models that complement each other will make a good ensemble, rather than models that agree with each other. That's why diversity is always emphasized when creating ensembles.\n\nA longer explanation is [**here**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058), and beware that I am sending you to one of my previous posts.",
      "votes": null
    },
    {
      "id": "1917660",
      "postDate": "08/29/2022 00:47:14",
      "content": "<p>If two models are very similar then if the first model has score 0.798 and the second model has score 0.796. Then the ensemble will result in the middle at 0.797.</p>\n<p>When two models are different, then the result magically boosts both! If the first model has score 0.798 and the second model has score 0.796, then the ensemble will result in 0.799 or 0.800!</p>\n<p>Successful ensemble is about achieving model diversity!</p>",
      "rawMarkdown": "If two models are very similar then if the first model has score 0.798 and the second model has score 0.796. Then the ensemble will result in the middle at 0.797.\n\nWhen two models are different, then the result magically boosts both! If the first model has score 0.798 and the second model has score 0.796, then the ensemble will result in 0.799 or 0.800!\n\nSuccessful ensemble is about achieving model diversity!",
      "votes": null
    },
    {
      "id": "1920796",
      "postDate": "08/31/2022 12:15:08",
      "content": "<p>In short, It extends the hypotheses space.<br>\nThere is a requirement that is the models are diverse and independent models.<br>\nSo the generalization ability will be promoted.</p>",
      "rawMarkdown": "In short, It extends the hypotheses space.\nThere is a requirement that is the models are diverse and independent models.\nSo the generalization ability will be promoted.",
      "votes": null
    },
    {
      "id": "1920800",
      "postDate": "08/31/2022 12:20:06",
      "content": "<p>Think of any ensemble as a single random forest classifier with <code>n_estimators=N</code>, where <code>N</code> is the number of models you put into ensemble.</p>",
      "rawMarkdown": "Think of any ensemble as a single random forest classifier with `n_estimators=N`, where `N` is the number of models you put into ensemble.",
      "votes": null
    },
    {
      "id": "1922275",
      "postDate": "09/01/2022 11:36:05",
      "content": "<p>Many high profile Kagglers already answered (providing excelent insights) and, happily autoreferential, their answers complete each other, so that the ensemble of the answers has even better value than each of them separately.</p>",
      "rawMarkdown": "Many high profile Kagglers already answered (providing excelent insights) and, happily autoreferential, their answers complete each other, so that the ensemble of the answers has even better value than each of them separately.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1916648,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/28/2022 04:05:35",
      "content": "<p>Awhile back, Carl linked to a great ensembling guide, it covers this question pretty well in the first section. <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1916688,
          "author_name": "itisliukun",
          "author_url": "",
          "post_date": "08/28/2022 05:28:23",
          "content": "<p>Great! Thank you very much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1917410,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "08/28/2022 17:47:43",
      "content": "<p>A simple explanation is that models that complement each other will make a good ensemble, rather than models that agree with each other. That's why diversity is always emphasized when creating ensembles.</p>\n<p>A longer explanation is <a href=\"https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058\" target=\"_blank\"><strong>here</strong></a>, and beware that I am sending you to one of my previous posts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1917660,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/29/2022 00:47:14",
      "content": "<p>If two models are very similar then if the first model has score 0.798 and the second model has score 0.796. Then the ensemble will result in the middle at 0.797.</p>\n<p>When two models are different, then the result magically boosts both! If the first model has score 0.798 and the second model has score 0.796, then the ensemble will result in 0.799 or 0.800!</p>\n<p>Successful ensemble is about achieving model diversity!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1920796,
      "author_name": "zhehaoliang",
      "author_url": "",
      "post_date": "08/31/2022 12:15:08",
      "content": "<p>In short, It extends the hypotheses space.<br>\nThere is a requirement that is the models are diverse and independent models.<br>\nSo the generalization ability will be promoted.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1920800,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "08/31/2022 12:20:06",
      "content": "<p>Think of any ensemble as a single random forest classifier with <code>n_estimators=N</code>, where <code>N</code> is the number of models you put into ensemble.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1922275,
      "author_name": "gpreda",
      "author_url": "",
      "post_date": "09/01/2022 11:36:05",
      "content": "<p>Many high profile Kagglers already answered (providing excelent insights) and, happily autoreferential, their answers complete each other, so that the ensemble of the answers has even better value than each of them separately.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1916627": "Congrats and thank you all. It's a great competition. I leaned so much from all the great shared notebooks. \n\nWhile from the notebooks and discussions there has some common points, more and more features or more model to ensemble.\n\nI know that more features will be helpful for digging unseen feature. \n\n**But why ensemble model is more precise not the opposite?**\n(Of course, the model to be ensembled is trained well)\n\nIt seems that the unstable predicted samples from one model will be corrected by other model not the opposite.\n\nCan anyone explain this in depth?\n\nThanks!",
    "1916648": "Awhile back, Carl linked to a great ensembling guide, it covers this question pretty well in the first section. https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729",
    "1916688": "Great! Thank you very much.",
    "1917410": "A simple explanation is that models that complement each other will make a good ensemble, rather than models that agree with each other. That's why diversity is always emphasized when creating ensembles.\n\nA longer explanation is [**here**](https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/discussion/51058), and beware that I am sending you to one of my previous posts.",
    "1917660": "If two models are very similar then if the first model has score 0.798 and the second model has score 0.796. Then the ensemble will result in the middle at 0.797.\n\nWhen two models are different, then the result magically boosts both! If the first model has score 0.798 and the second model has score 0.796, then the ensemble will result in 0.799 or 0.800!\n\nSuccessful ensemble is about achieving model diversity!",
    "1920796": "In short, It extends the hypotheses space.\nThere is a requirement that is the models are diverse and independent models.\nSo the generalization ability will be promoted.",
    "1920800": "Think of any ensemble as a single random forest classifier with `n_estimators=N`, where `N` is the number of models you put into ensemble.",
    "1922275": "Many high profile Kagglers already answered (providing excelent insights) and, happily autoreferential, their answers complete each other, so that the ensemble of the answers has even better value than each of them separately."
  },
  "source": "meta"
}