{
  "id": 332729,
  "title": "The \"Kaggle Ensembling Guide\"",
  "url": "/competitions/amex-default-prediction/discussion/332729",
  "author_name": "Carl McBride Ellis",
  "post_date": "2022-06-23T04:45:25.361000",
  "votes": 55,
  "comment_count": 14,
  "views": 0,
  "content": "<p>TL;DR: <a href=\"https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">\"Kaggle Ensembling Guide\"</a></p>\n<p>A single individual model will (almost certainly) not be the winning solution; an \"ensemble\" will have to be made. The how and why of this is outlined beautifully in the above link, and for anyone unfamiliar with the long tradition that ensembling has on kaggle  (and is not often seen in ML textbooks) I cannot recommend this guide enough.</p>\n<p>All the best,<br>\ncarl</p>",
  "messages": [
    {
      "id": 1829978,
      "postDate": "2022-06-23T04:45:25.360Z",
      "content": "<p>TL;DR: <a href=\"https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/\" target=\"_blank\">\"Kaggle Ensembling Guide\"</a></p>\n<p>A single individual model will (almost certainly) not be the winning solution; an \"ensemble\" will have to be made. The how and why of this is outlined beautifully in the above link, and for anyone unfamiliar with the long tradition that ensembling has on kaggle  (and is not often seen in ML textbooks) I cannot recommend this guide enough.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "TL;DR: [\"Kaggle Ensembling Guide\"](https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/)\n\nA single individual model will (almost certainly) not be the winning solution; an \"ensemble\" will have to be made. The how and why of this is outlined beautifully in the above link, and for anyone unfamiliar with the long tradition that ensembling has on kaggle  (and is not often seen in ML textbooks) I cannot recommend this guide enough.\n\nAll the best,\ncarl",
      "votes": 55
    },
    {
      "id": 1853471,
      "postDate": "2022-07-12T22:45:49.593Z",
      "content": "<blockquote>\n  <p>with the long tradition that ensembling has on kaggle (and is not often seen in ML textbooks)</p>\n</blockquote>\n<p>This surprised me the most when I joined Kaggle. I learned about many practical things on Kaggle that were not discussed in the literature I read on machine learning.</p>",
      "rawMarkdown": "> with the long tradition that ensembling has on kaggle (and is not often seen in ML textbooks)\n\nThis surprised me the most when I joined Kaggle. I learned about many practical things on Kaggle that were not discussed in the literature I read on machine learning.",
      "votes": 3,
      "replies": [
        {
          "id": 1903345,
          "postDate": "2022-08-17T10:41:13.820Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> </p>\n<p>I think one reason it is not often seen in ML textbooks is perhaps because the worst and most frequent kind of ensembling, that of combining public <code>submission.csv</code> files and iteratively playing with the weights to achieve the best possible Public Leaderboard score, is complete nonsense anywhere other than on kaggle. Indeed I have just made a suggestion that I think would help curtail such unsound approaches: <a href=\"https://www.kaggle.com/discussions/product-feedback/345940\" target=\"_blank\">\"Competitions: All test data should be hidden ⊗ all comps are code\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @fritzcremer \n\nI think one reason it is not often seen in ML textbooks is perhaps because the worst and most frequent kind of ensembling, that of combining public `submission.csv` files and iteratively playing with the weights to achieve the best possible Public Leaderboard score, is complete nonsense anywhere other than on kaggle. Indeed I have just made a suggestion that I think would help curtail such unsound approaches: [\"Competitions: All test data should be hidden ⊗ all comps are code\"](https://www.kaggle.com/discussions/product-feedback/345940).\n\nAll the best,\ncarl",
          "votes": 3
        }
      ]
    },
    {
      "id": 1829998,
      "postDate": "2022-06-23T05:09:33.667Z",
      "content": "<p>I would also suggest focusing more on building a single model (not sure if it is applicable here). Usually, a successful ensemble/team merge requires most models to be in the gold/top-silver zone.</p>",
      "rawMarkdown": "I would also suggest focusing more on building a single model (not sure if it is applicable here). Usually, a successful ensemble/team merge requires most models to be in the gold/top-silver zone.",
      "votes": 4,
      "replies": [
        {
          "id": 1830004,
          "postDate": "2022-06-23T05:20:44.937Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> </p>\n<blockquote>\n  <p>\"<em>I would also suggest focusing more on building a single model</em>\"</p>\n</blockquote>\n<p>Oh most definitely! I would hazard a guess that the top teams in the end will be those that have built several strong models via diverse \"low-correlation\" approaches, then in the  last week or so they produce an ensemble of those solutions to squeeze out that little extra. That said I do seem to remember that in perhaps one of the Santander (?) competitions one of the winning solutions was actually a <a href=\"https://www.ismll.uni-hildesheim.de/pub/pdfs/wistuba_et_al_SDM_2017.pdf\" target=\"_blank\"><em>\"Frankenstein\" ensemble</em></a>  of public notebooks.</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @kingychiu \n\n> \"*I would also suggest focusing more on building a single model*\"\n\nOh most definitely! I would hazard a guess that the top teams in the end will be those that have built several strong models via diverse \"low-correlation\" approaches, then in the  last week or so they produce an ensemble of those solutions to squeeze out that little extra. That said I do seem to remember that in perhaps one of the Santander (?) competitions one of the winning solutions was actually a [*\"Frankenstein\" ensemble*](https://www.ismll.uni-hildesheim.de/pub/pdfs/wistuba_et_al_SDM_2017.pdf)  of public notebooks.\n\nAll the best,\ncarl",
          "votes": 7
        },
        {
          "id": 1853426,
          "postDate": "2022-07-12T21:47:17.323Z",
          "content": "<p>In that case how did the guys winning get confidence without a way of knowing the oof cv or was the oof also available in the case you mentioned publicly.</p>",
          "rawMarkdown": "In that case how did the guys winning get confidence without a way of knowing the oof cv or was the oof also available in the case you mentioned publicly.",
          "votes": 1
        },
        {
          "id": 1861614,
          "postDate": "2022-07-19T05:55:18.117Z",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> </p>\n<p>In the big competitions, with 2k+ participants and a very compact LB, sometimes a lucky blend can win. <br>\n…Even a stopped clock is right twice a day.</p>\n<p>I personally wouldn't recommend creating such a \"<em>Frankenstein</em>\" ensemble; to think that doing so is actually a good strategy would be an example of <a href=\"https://en.wikipedia.org/wiki/Survivorship_bias\" target=\"_blank\">survivorship bias</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "rawMarkdown": "Dear @gauravbrills \n\nIn the big competitions, with 2k+ participants and a very compact LB, sometimes a lucky blend can win. \n...Even a stopped clock is right twice a day.\n\nI personally wouldn't recommend creating such a \"*Frankenstein*\" ensemble; to think that doing so is actually a good strategy would be an example of [survivorship bias](https://en.wikipedia.org/wiki/Survivorship_bias).\n\nAll the best,\ncarl",
          "votes": 2
        }
      ]
    },
    {
      "id": 1831287,
      "postDate": "2022-06-24T05:16:04.663Z",
      "content": "<p>This sounds amazing. Thanks for sharing!</p>",
      "rawMarkdown": "This sounds amazing. Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1830029,
      "postDate": "2022-06-23T05:39:55.320Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a>  for sharing this resource.</p>",
      "rawMarkdown": "Thanks @carlmcbrideellis  for sharing this resource.",
      "votes": 1
    },
    {
      "id": 1830368,
      "postDate": "2022-06-23T12:00:22.010Z",
      "content": "<p>Thanks for sharing!  I was looking for something like this</p>",
      "rawMarkdown": "Thanks for sharing!  I was looking for something like this",
      "votes": 2
    },
    {
      "id": 1829986,
      "postDate": "2022-06-23T04:56:14.563Z",
      "content": "<p>Thank you for sharing this Carl! I have always been curious about model ensemble strategies!</p>",
      "rawMarkdown": "Thank you for sharing this Carl! I have always been curious about model ensemble strategies!",
      "votes": 2
    },
    {
      "id": 1917203,
      "postDate": "2022-08-28T14:21:42.210Z",
      "content": "<p>I totally agree with  <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a>. Ensembling more than ten models is not cost effective strategy in practice that cost too much computing resource.</p>",
      "rawMarkdown": "I totally agree with  @carlmcbrideellis. Ensembling more than ten models is not cost effective strategy in practice that cost too much computing resource."
    },
    {
      "id": 1835746,
      "postDate": "2022-06-28T03:54:24.863Z",
      "content": "<p>Thanks for sharing … ! <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> </p>",
      "rawMarkdown": "Thanks for sharing ... ! @carlmcbrideellis "
    },
    {
      "id": 1835298,
      "postDate": "2022-06-27T16:29:13.187Z",
      "content": "<p>Thanks for sharing this awesome guide <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> </p>",
      "rawMarkdown": "Thanks for sharing this awesome guide @carlmcbrideellis "
    },
    {
      "id": 1831295,
      "postDate": "2022-06-24T05:17:30.883Z",
      "content": "<p>Really helpful!</p>",
      "rawMarkdown": "Really helpful!",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1853471,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2022-07-12T22:45:49.593000",
      "content": "<blockquote>\n  <p>with the long tradition that ensembling has on kaggle (and is not often seen in ML textbooks)</p>\n</blockquote>\n<p>This surprised me the most when I joined Kaggle. I learned about many practical things on Kaggle that were not discussed in the literature I read on machine learning.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1903345,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-08-17T10:41:13.820000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> </p>\n<p>I think one reason it is not often seen in ML textbooks is perhaps because the worst and most frequent kind of ensembling, that of combining public <code>submission.csv</code> files and iteratively playing with the weights to achieve the best possible Public Leaderboard score, is complete nonsense anywhere other than on kaggle. Indeed I have just made a suggestion that I think would help curtail such unsound approaches: <a href=\"https://www.kaggle.com/discussions/product-feedback/345940\" target=\"_blank\">\"Competitions: All test data should be hidden ⊗ all comps are code\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1829998,
      "author_name": "Anthony Chiu",
      "author_url": "",
      "post_date": "2022-06-23T05:09:33.667000",
      "content": "<p>I would also suggest focusing more on building a single model (not sure if it is applicable here). Usually, a successful ensemble/team merge requires most models to be in the gold/top-silver zone.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1830004,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-06-23T05:20:44.937000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> </p>\n<blockquote>\n  <p>\"<em>I would also suggest focusing more on building a single model</em>\"</p>\n</blockquote>\n<p>Oh most definitely! I would hazard a guess that the top teams in the end will be those that have built several strong models via diverse \"low-correlation\" approaches, then in the  last week or so they produce an ensemble of those solutions to squeeze out that little extra. That said I do seem to remember that in perhaps one of the Santander (?) competitions one of the winning solutions was actually a <a href=\"https://www.ismll.uni-hildesheim.de/pub/pdfs/wistuba_et_al_SDM_2017.pdf\" target=\"_blank\"><em>\"Frankenstein\" ensemble</em></a>  of public notebooks.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1853426,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2022-07-12T21:47:17.323000",
          "content": "<p>In that case how did the guys winning get confidence without a way of knowing the oof cv or was the oof also available in the case you mentioned publicly.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1861614,
          "author_name": "Carl McBride Ellis",
          "author_url": "",
          "post_date": "2022-07-19T05:55:18.117000",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> </p>\n<p>In the big competitions, with 2k+ participants and a very compact LB, sometimes a lucky blend can win. <br>\n…Even a stopped clock is right twice a day.</p>\n<p>I personally wouldn't recommend creating such a \"<em>Frankenstein</em>\" ensemble; to think that doing so is actually a good strategy would be an example of <a href=\"https://en.wikipedia.org/wiki/Survivorship_bias\" target=\"_blank\">survivorship bias</a>.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1831287,
      "author_name": "ANMOL BHATIA",
      "author_url": "",
      "post_date": "2022-06-24T05:16:04.663000",
      "content": "<p>This sounds amazing. Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1830029,
      "author_name": "VSB",
      "author_url": "",
      "post_date": "2022-06-23T05:39:55.320000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a>  for sharing this resource.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1830368,
      "author_name": "Paulo Junqueira",
      "author_url": "",
      "post_date": "2022-06-23T12:00:22.010000",
      "content": "<p>Thanks for sharing!  I was looking for something like this</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1829986,
      "author_name": "Tonghui Li",
      "author_url": "",
      "post_date": "2022-06-23T04:56:14.563000",
      "content": "<p>Thank you for sharing this Carl! I have always been curious about model ensemble strategies!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1917203,
      "author_name": "yiqin",
      "author_url": "",
      "post_date": "2022-08-28T14:21:42.210000",
      "content": "<p>I totally agree with  <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a>. Ensembling more than ten models is not cost effective strategy in practice that cost too much computing resource.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1835746,
      "author_name": "Pankaj Kumar",
      "author_url": "",
      "post_date": "2022-06-28T03:54:24.863000",
      "content": "<p>Thanks for sharing … ! <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1835298,
      "author_name": "Dhamu",
      "author_url": "",
      "post_date": "2022-06-27T16:29:13.187000",
      "content": "<p>Thanks for sharing this awesome guide <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1831295,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-24T05:17:30.883000",
      "content": "<p>Really helpful!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1829978": "TL;DR: [\"Kaggle Ensembling Guide\"](https://web.archive.org/web/20160304031055/http://mlwave.com/kaggle-ensembling-guide/)\n\nA single individual model will (almost certainly) not be the winning solution; an \"ensemble\" will have to be made. The how and why of this is outlined beautifully in the above link, and for anyone unfamiliar with the long tradition that ensembling has on kaggle  (and is not often seen in ML textbooks) I cannot recommend this guide enough.\n\nAll the best,\ncarl",
    "1853471": "> with the long tradition that ensembling has on kaggle (and is not often seen in ML textbooks)\n\nThis surprised me the most when I joined Kaggle. I learned about many practical things on Kaggle that were not discussed in the literature I read on machine learning.",
    "1829998": "I would also suggest focusing more on building a single model (not sure if it is applicable here). Usually, a successful ensemble/team merge requires most models to be in the gold/top-silver zone.",
    "1831287": "This sounds amazing. Thanks for sharing!",
    "1830029": "Thanks @carlmcbrideellis  for sharing this resource.",
    "1830368": "Thanks for sharing!  I was looking for something like this",
    "1829986": "Thank you for sharing this Carl! I have always been curious about model ensemble strategies!",
    "1917203": "I totally agree with  @carlmcbrideellis. Ensembling more than ten models is not cost effective strategy in practice that cost too much computing resource.",
    "1835746": "Thanks for sharing ... ! @carlmcbrideellis ",
    "1835298": "Thanks for sharing this awesome guide @carlmcbrideellis ",
    "1831295": "Really helpful!"
  }
}