{
  "id": 202344,
  "title": "Single Model Ensembling Guide | LightGBM Example",
  "url": "/competitions/riiid-test-answer-prediction/discussion/202344",
  "author_name": "",
  "post_date": "2020-12-09T14:27:25.695107800Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Tl;dr</strong> We can use individual models for ensembling. For Neural Networks, save models at regular intervals and use the weighted average as the prediction. For GBMs, create multiple outputs by limiting the number of trees used. Check <a href=\"https://arxiv.org/pdf/1912.02757.pdf\" target=\"_blank\">this paper</a> or (<a href=\"https://www.youtube.com/watch?v=5IRlUVrEVL8\" target=\"_blank\">this review video</a>) for more insight.</p>\n<hr>\n<h2>Single Model Ensembling</h2>\n<p>Everyone here agrees that it is important to build individual models but in order to get a medal, we have to do some kind of ensembling. But for competitions like this one, it is difficult to create a huge number of versions of the same model (KFold with Multiple Seeds situation). In this type of case, it would be useful to understand that individual models can also be used to create ensembles and improve the performance by some <em>basis points</em> depending upon the problem. </p>\n<p>I came across this paper (<a href=\"https://arxiv.org/pdf/1912.02757.pdf\" target=\"_blank\">Deep Ensembles: A Loss Landscape Perspective</a>) a while back after watching <a href=\"https://www.youtube.com/watch?v=5IRlUVrEVL8\" target=\"_blank\">its paper review by Yannic Kilcher</a>. I thought this would be a good chance to use this concept. The main idea is that even though different versions (NN: starting with different weights, using different optimizers, …) of the same model give the same accuracy, there will be some examples where they will give different results. Using ensembles of the outputs will help us ignore the local biases.</p>\n<p><img src=\"https://i.imgur.com/tEazOJl.png\" alt=\"Ensembles help use to ignore local biases\"> <br>\nsorry for the bad quality image, wasn't able to get any better one. </p>\n<p>Go through the paper if you are interested. The paper is very informative and interesting. </p>\n<p><strong>Main point</strong>: Save DNNs at different epochs and use them for ensembling. </p>\n<p><img src=\"https://i.imgur.com/aGCwn7R.png\" alt=\"Save DNNs at different epochs and use them for ensembling\"></p>\n<h2>LightGBM Example</h2>\n<p>The paper goes through in detail how to use NNs, but doesn't mention Tree-based models. I have few ideas:</p>\n<ol>\n<li><p>Limit the number of trees used while scoring to create different outputs for ensembling. As an example, I have created <a href=\"https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-blending-training\" target=\"_blank\">sample training</a> and <a href=\"https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-ensembling-scoring\" target=\"_blank\">scoring notebook</a>. In this case, I picked 400, 700, and 989 as the breakpoints. So if we have an overfitted GBM model, then this technique will be really useful to use this kind of technique. In my example, the performance degraded (Valid AUC). One way to explain it is to present how Boosting Machines work. <img src=\"https://i.imgur.com/qlQTh0b.png\" alt=\"Limit the number of trees used while scoring to create different outputs for ensembling\"></p></li>\n<li><p>Retrain the same model with a different dataset. GBM models allow us to retrain them with a different subset of data (Performance is problem-dependent). To train with a full dataset with more than 20 variables is difficult in the Kaggle environment. So we can split the data into 2 or more splits. Train on the first split, save the model, and retrain the same model with the other splits. This way we can create copies for ensembles. <img src=\"https://i.imgur.com/RJ2DRYq.png\" alt=\"Retrain the same model with a different dataset\"></p></li>\n</ol>\n<p>Links:</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1912.02757\" target=\"_blank\">Deep Ensembles: A Loss Landscape Perspective paper</a> and <a href=\"https://www.youtube.com/watch?v=5IRlUVrEVL8\" target=\"_blank\">it's review video by Yannic Kilcher on YouTube</a></li>\n<li><a href=\"https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-blending-training\" target=\"_blank\">Training Script</a> modified from <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering/\" target=\"_blank\">this popular public notebook</a>. </li>\n<li><a href=\"https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-ensembling-scoring\" target=\"_blank\">Scoring Script</a></li>\n</ul>\n<hr>\n<p><strong>Note</strong>: I in <a href=\"https://www.kaggle.com/c/lish-moa/discussion/202612\" target=\"_blank\">no way suggest/support beginners to just create ensembles of public notebooks</a>. If all you are doing is that, then understand that there is a better way to spend your time. Ensembling helps in scoring medals but understanding that it is only a small part of Data Science and is rarely used in real life is very important. I saw some threads asking for blending ideas, so I am just sharing them. </p>\n<p>If you know other ways of ensembling please let me know. </p>",
  "messages": [
    {
      "id": "1107254",
      "postDate": "12/09/2020 14:27:25",
      "content": "<p><strong>Tl;dr</strong> We can use individual models for ensembling. For Neural Networks, save models at regular intervals and use the weighted average as the prediction. For GBMs, create multiple outputs by limiting the number of trees used. Check <a href=\"https://arxiv.org/pdf/1912.02757.pdf\" target=\"_blank\">this paper</a> or (<a href=\"https://www.youtube.com/watch?v=5IRlUVrEVL8\" target=\"_blank\">this review video</a>) for more insight.</p>\n<hr>\n<h2>Single Model Ensembling</h2>\n<p>Everyone here agrees that it is important to build individual models but in order to get a medal, we have to do some kind of ensembling. But for competitions like this one, it is difficult to create a huge number of versions of the same model (KFold with Multiple Seeds situation). In this type of case, it would be useful to understand that individual models can also be used to create ensembles and improve the performance by some <em>basis points</em> depending upon the problem. </p>\n<p>I came across this paper (<a href=\"https://arxiv.org/pdf/1912.02757.pdf\" target=\"_blank\">Deep Ensembles: A Loss Landscape Perspective</a>) a while back after watching <a href=\"https://www.youtube.com/watch?v=5IRlUVrEVL8\" target=\"_blank\">its paper review by Yannic Kilcher</a>. I thought this would be a good chance to use this concept. The main idea is that even though different versions (NN: starting with different weights, using different optimizers, …) of the same model give the same accuracy, there will be some examples where they will give different results. Using ensembles of the outputs will help us ignore the local biases.</p>\n<p><img src=\"https://i.imgur.com/tEazOJl.png\" alt=\"Ensembles help use to ignore local biases\"> <br>\nsorry for the bad quality image, wasn't able to get any better one. </p>\n<p>Go through the paper if you are interested. The paper is very informative and interesting. </p>\n<p><strong>Main point</strong>: Save DNNs at different epochs and use them for ensembling. </p>\n<p><img src=\"https://i.imgur.com/aGCwn7R.png\" alt=\"Save DNNs at different epochs and use them for ensembling\"></p>\n<h2>LightGBM Example</h2>\n<p>The paper goes through in detail how to use NNs, but doesn't mention Tree-based models. I have few ideas:</p>\n<ol>\n<li><p>Limit the number of trees used while scoring to create different outputs for ensembling. As an example, I have created <a href=\"https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-blending-training\" target=\"_blank\">sample training</a> and <a href=\"https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-ensembling-scoring\" target=\"_blank\">scoring notebook</a>. In this case, I picked 400, 700, and 989 as the breakpoints. So if we have an overfitted GBM model, then this technique will be really useful to use this kind of technique. In my example, the performance degraded (Valid AUC). One way to explain it is to present how Boosting Machines work. <img src=\"https://i.imgur.com/qlQTh0b.png\" alt=\"Limit the number of trees used while scoring to create different outputs for ensembling\"></p></li>\n<li><p>Retrain the same model with a different dataset. GBM models allow us to retrain them with a different subset of data (Performance is problem-dependent). To train with a full dataset with more than 20 variables is difficult in the Kaggle environment. So we can split the data into 2 or more splits. Train on the first split, save the model, and retrain the same model with the other splits. This way we can create copies for ensembles. <img src=\"https://i.imgur.com/RJ2DRYq.png\" alt=\"Retrain the same model with a different dataset\"></p></li>\n</ol>\n<p>Links:</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/1912.02757\" target=\"_blank\">Deep Ensembles: A Loss Landscape Perspective paper</a> and <a href=\"https://www.youtube.com/watch?v=5IRlUVrEVL8\" target=\"_blank\">it's review video by Yannic Kilcher on YouTube</a></li>\n<li><a href=\"https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-blending-training\" target=\"_blank\">Training Script</a> modified from <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering/\" target=\"_blank\">this popular public notebook</a>. </li>\n<li><a href=\"https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-ensembling-scoring\" target=\"_blank\">Scoring Script</a></li>\n</ul>\n<hr>\n<p><strong>Note</strong>: I in <a href=\"https://www.kaggle.com/c/lish-moa/discussion/202612\" target=\"_blank\">no way suggest/support beginners to just create ensembles of public notebooks</a>. If all you are doing is that, then understand that there is a better way to spend your time. Ensembling helps in scoring medals but understanding that it is only a small part of Data Science and is rarely used in real life is very important. I saw some threads asking for blending ideas, so I am just sharing them. </p>\n<p>If you know other ways of ensembling please let me know. </p>",
      "rawMarkdown": "**Tl;dr** We can use individual models for ensembling. For Neural Networks, save models at regular intervals and use the weighted average as the prediction. For GBMs, create multiple outputs by limiting the number of trees used. Check [this paper](https://arxiv.org/pdf/1912.02757.pdf) or ([this review video](https://www.youtube.com/watch?v=5IRlUVrEVL8)) for more insight.\n \n---\n\n## Single Model Ensembling\nEveryone here agrees that it is important to build individual models but in order to get a medal, we have to do some kind of ensembling. But for competitions like this one, it is difficult to create a huge number of versions of the same model (KFold with Multiple Seeds situation). In this type of case, it would be useful to understand that individual models can also be used to create ensembles and improve the performance by some *basis points* depending upon the problem. \n\nI came across this paper ([Deep Ensembles: A Loss Landscape Perspective](https://arxiv.org/pdf/1912.02757.pdf)) a while back after watching [its paper review by Yannic Kilcher](https://www.youtube.com/watch?v=5IRlUVrEVL8). I thought this would be a good chance to use this concept. The main idea is that even though different versions (NN: starting with different weights, using different optimizers, ...) of the same model give the same accuracy, there will be some examples where they will give different results. Using ensembles of the outputs will help us ignore the local biases.\n\n![Ensembles help use to ignore local biases](https://i.imgur.com/tEazOJl.png) \nsorry for the bad quality image, wasn't able to get any better one. \n\nGo through the paper if you are interested. The paper is very informative and interesting. \n\n**Main point**: Save DNNs at different epochs and use them for ensembling. \n\n![Save DNNs at different epochs and use them for ensembling](https://i.imgur.com/aGCwn7R.png)\n\n## LightGBM Example\nThe paper goes through in detail how to use NNs, but doesn't mention Tree-based models. I have few ideas:\n1. Limit the number of trees used while scoring to create different outputs for ensembling. As an example, I have created [sample training](https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-blending-training) and [scoring notebook](https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-ensembling-scoring). In this case, I picked 400, 700, and 989 as the breakpoints. So if we have an overfitted GBM model, then this technique will be really useful to use this kind of technique. In my example, the performance degraded (Valid AUC). One way to explain it is to present how Boosting Machines work. ![Limit the number of trees used while scoring to create different outputs for ensembling](https://i.imgur.com/qlQTh0b.png)\n\n2. Retrain the same model with a different dataset. GBM models allow us to retrain them with a different subset of data (Performance is problem-dependent). To train with a full dataset with more than 20 variables is difficult in the Kaggle environment. So we can split the data into 2 or more splits. Train on the first split, save the model, and retrain the same model with the other splits. This way we can create copies for ensembles. ![Retrain the same model with a different dataset](https://i.imgur.com/RJ2DRYq.png)\n\nLinks:\n- [Deep Ensembles: A Loss Landscape Perspective paper](https://arxiv.org/abs/1912.02757) and [it's review video by Yannic Kilcher on YouTube](https://www.youtube.com/watch?v=5IRlUVrEVL8)\n- [Training Script](https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-blending-training) modified from [this popular public notebook](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering/). \n- [Scoring Script](https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-ensembling-scoring)\n\n---\n\n**Note**: I in [no way suggest/support beginners to just create ensembles of public notebooks](https://www.kaggle.com/c/lish-moa/discussion/202612). If all you are doing is that, then understand that there is a better way to spend your time. Ensembling helps in scoring medals but understanding that it is only a small part of Data Science and is rarely used in real life is very important. I saw some threads asking for blending ideas, so I am just sharing them. \n\nIf you know other ways of ensembling please let me know.",
      "votes": null
    },
    {
      "id": "1108445",
      "postDate": "12/10/2020 16:49:27",
      "content": "<p>Very helpful, thank you</p>",
      "rawMarkdown": "Very helpful, thank you",
      "votes": null
    },
    {
      "id": "1109749",
      "postDate": "12/12/2020 03:06:15",
      "content": "<p>Glad to be of help. </p>",
      "rawMarkdown": "Glad to be of help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1108445,
      "author_name": "ruslankupr",
      "author_url": "",
      "post_date": "12/10/2020 16:49:27",
      "content": "<p>Very helpful, thank you</p>",
      "votes": null,
      "replies": [
        {
          "id": 1109749,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "12/12/2020 03:06:15",
          "content": "<p>Glad to be of help. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1107254": "**Tl;dr** We can use individual models for ensembling. For Neural Networks, save models at regular intervals and use the weighted average as the prediction. For GBMs, create multiple outputs by limiting the number of trees used. Check [this paper](https://arxiv.org/pdf/1912.02757.pdf) or ([this review video](https://www.youtube.com/watch?v=5IRlUVrEVL8)) for more insight.\n \n---\n\n## Single Model Ensembling\nEveryone here agrees that it is important to build individual models but in order to get a medal, we have to do some kind of ensembling. But for competitions like this one, it is difficult to create a huge number of versions of the same model (KFold with Multiple Seeds situation). In this type of case, it would be useful to understand that individual models can also be used to create ensembles and improve the performance by some *basis points* depending upon the problem. \n\nI came across this paper ([Deep Ensembles: A Loss Landscape Perspective](https://arxiv.org/pdf/1912.02757.pdf)) a while back after watching [its paper review by Yannic Kilcher](https://www.youtube.com/watch?v=5IRlUVrEVL8). I thought this would be a good chance to use this concept. The main idea is that even though different versions (NN: starting with different weights, using different optimizers, ...) of the same model give the same accuracy, there will be some examples where they will give different results. Using ensembles of the outputs will help us ignore the local biases.\n\n![Ensembles help use to ignore local biases](https://i.imgur.com/tEazOJl.png) \nsorry for the bad quality image, wasn't able to get any better one. \n\nGo through the paper if you are interested. The paper is very informative and interesting. \n\n**Main point**: Save DNNs at different epochs and use them for ensembling. \n\n![Save DNNs at different epochs and use them for ensembling](https://i.imgur.com/aGCwn7R.png)\n\n## LightGBM Example\nThe paper goes through in detail how to use NNs, but doesn't mention Tree-based models. I have few ideas:\n1. Limit the number of trees used while scoring to create different outputs for ensembling. As an example, I have created [sample training](https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-blending-training) and [scoring notebook](https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-ensembling-scoring). In this case, I picked 400, 700, and 989 as the breakpoints. So if we have an overfitted GBM model, then this technique will be really useful to use this kind of technique. In my example, the performance degraded (Valid AUC). One way to explain it is to present how Boosting Machines work. ![Limit the number of trees used while scoring to create different outputs for ensembling](https://i.imgur.com/qlQTh0b.png)\n\n2. Retrain the same model with a different dataset. GBM models allow us to retrain them with a different subset of data (Performance is problem-dependent). To train with a full dataset with more than 20 variables is difficult in the Kaggle environment. So we can split the data into 2 or more splits. Train on the first split, save the model, and retrain the same model with the other splits. This way we can create copies for ensembles. ![Retrain the same model with a different dataset](https://i.imgur.com/RJ2DRYq.png)\n\nLinks:\n- [Deep Ensembles: A Loss Landscape Perspective paper](https://arxiv.org/abs/1912.02757) and [it's review video by Yannic Kilcher on YouTube](https://www.youtube.com/watch?v=5IRlUVrEVL8)\n- [Training Script](https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-blending-training) modified from [this popular public notebook](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering/). \n- [Scoring Script](https://www.kaggle.com/manikanthr5/riiid-lgbm-single-model-ensembling-scoring)\n\n---\n\n**Note**: I in [no way suggest/support beginners to just create ensembles of public notebooks](https://www.kaggle.com/c/lish-moa/discussion/202612). If all you are doing is that, then understand that there is a better way to spend your time. Ensembling helps in scoring medals but understanding that it is only a small part of Data Science and is rarely used in real life is very important. I saw some threads asking for blending ideas, so I am just sharing them. \n\nIf you know other ways of ensembling please let me know.",
    "1108445": "Very helpful, thank you",
    "1109749": "Glad to be of help."
  },
  "source": "meta"
}