{
  "id": 76372,
  "title": "What is proper ways to make ensemble models?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/76372",
  "author_name": "",
  "post_date": "2019-01-02T02:30:07.704521800Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I trained models with different cross-validation folds and architectures.</p>\n\n<p>But naive ensemble model using majority voting improve the performance about +0.02.</p>\n\n<p>Also I tried to use xgboost for outputs from multiple models, it is degradate the performance.</p>\n\n<p>Any suggestion?</p>",
  "messages": [
    {
      "id": "448752",
      "postDate": "01/02/2019 02:30:07",
      "content": "<p>I trained models with different cross-validation folds and architectures.</p>\n\n<p>But naive ensemble model using majority voting improve the performance about +0.02.</p>\n\n<p>Also I tried to use xgboost for outputs from multiple models, it is degradate the performance.</p>\n\n<p>Any suggestion?</p>",
      "rawMarkdown": "I trained models with different cross-validation folds and architectures.\n\nBut naive ensemble model using majority voting improve the performance about +0.02.\n\nAlso I tried to use xgboost for outputs from multiple models, it is degradate the performance.\n\nAny suggestion?",
      "votes": null
    },
    {
      "id": "449093",
      "postDate": "01/02/2019 16:20:42",
      "content": "<p>I have the same observation.\nI used lightGBM, but the results are worse than my best single model.</p>",
      "rawMarkdown": "I have the same observation.\nI used lightGBM, but the results are worse than my best single model.",
      "votes": null
    },
    {
      "id": "449247",
      "postDate": "01/02/2019 21:28:30",
      "content": "<p>Simple averaging for the moment yielded the best results in terms of LB score. Improvements are in the order of ~0.03 for me.</p>\n\n<p>Linear regression, lightgbm, xgboost... all score significantly less. </p>\n\n<p>Btw, if I'm not mistaken in lightgbm you can only use a one-vs-all approach, multi-label problems are not supported.</p>",
      "rawMarkdown": "Simple averaging for the moment yielded the best results in terms of LB score. Improvements are in the order of ~0.03 for me.\n\nLinear regression, lightgbm, xgboost... all score significantly less. \n\nBtw, if I'm not mistaken in lightgbm you can only use a one-vs-all approach, multi-label problems are not supported.",
      "votes": null
    },
    {
      "id": "449711",
      "postDate": "01/03/2019 16:29:05",
      "content": "<p>I used lightGBM for each label separately.\nMaybe I should try average or majority voting for ensembling.</p>",
      "rawMarkdown": "I used lightGBM for each label separately.\nMaybe I should try average or majority voting for ensembling.",
      "votes": null
    },
    {
      "id": "449794",
      "postDate": "01/03/2019 18:56:51",
      "content": "<p>I have been averaging the raw numerical predictions of multiple models, possibly weighting the models differently based on their relative performances, then thresholding the result to get the label predictions.  This method has been fairly effective.  However, as has been pointed out in another thread, averaging the numerical predictions changes the shape of their distribution and so may affect the choice of thresholds.</p>",
      "rawMarkdown": "I have been averaging the raw numerical predictions of multiple models, possibly weighting the models differently based on their relative performances, then thresholding the result to get the label predictions.  This method has been fairly effective.  However, as has been pointed out in another thread, averaging the numerical predictions changes the shape of their distribution and so may affect the choice of thresholds.",
      "votes": null
    },
    {
      "id": "451609",
      "postDate": "01/07/2019 10:49:02",
      "content": "<p>Simple averaging +1.  I even build a small ensemble network to do the ensemble but none of them beat averaging.</p>",
      "rawMarkdown": "Simple averaging +1.  I even build a small ensemble network to do the ensemble but none of them beat averaging.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 449093,
      "author_name": "zhijianli",
      "author_url": "",
      "post_date": "01/02/2019 16:20:42",
      "content": "<p>I have the same observation.\nI used lightGBM, but the results are worse than my best single model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449247,
      "author_name": "thundo",
      "author_url": "",
      "post_date": "01/02/2019 21:28:30",
      "content": "<p>Simple averaging for the moment yielded the best results in terms of LB score. Improvements are in the order of ~0.03 for me.</p>\n\n<p>Linear regression, lightgbm, xgboost... all score significantly less. </p>\n\n<p>Btw, if I'm not mistaken in lightgbm you can only use a one-vs-all approach, multi-label problems are not supported.</p>",
      "votes": null,
      "replies": [
        {
          "id": 449711,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "01/03/2019 16:29:05",
          "content": "<p>I used lightGBM for each label separately.\nMaybe I should try average or majority voting for ensembling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 449794,
      "author_name": "dslate",
      "author_url": "",
      "post_date": "01/03/2019 18:56:51",
      "content": "<p>I have been averaging the raw numerical predictions of multiple models, possibly weighting the models differently based on their relative performances, then thresholding the result to get the label predictions.  This method has been fairly effective.  However, as has been pointed out in another thread, averaging the numerical predictions changes the shape of their distribution and so may affect the choice of thresholds.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 451609,
      "author_name": "snaker",
      "author_url": "",
      "post_date": "01/07/2019 10:49:02",
      "content": "<p>Simple averaging +1.  I even build a small ensemble network to do the ensemble but none of them beat averaging.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "448752": "I trained models with different cross-validation folds and architectures.\n\nBut naive ensemble model using majority voting improve the performance about +0.02.\n\nAlso I tried to use xgboost for outputs from multiple models, it is degradate the performance.\n\nAny suggestion?",
    "449093": "I have the same observation.\nI used lightGBM, but the results are worse than my best single model.",
    "449247": "Simple averaging for the moment yielded the best results in terms of LB score. Improvements are in the order of ~0.03 for me.\n\nLinear regression, lightgbm, xgboost... all score significantly less. \n\nBtw, if I'm not mistaken in lightgbm you can only use a one-vs-all approach, multi-label problems are not supported.",
    "449711": "I used lightGBM for each label separately.\nMaybe I should try average or majority voting for ensembling.",
    "449794": "I have been averaging the raw numerical predictions of multiple models, possibly weighting the models differently based on their relative performances, then thresholding the result to get the label predictions.  This method has been fairly effective.  However, as has been pointed out in another thread, averaging the numerical predictions changes the shape of their distribution and so may affect the choice of thresholds.",
    "451609": "Simple averaging +1.  I even build a small ensemble network to do the ensemble but none of them beat averaging."
  },
  "source": "meta"
}