{
  "id": 221669,
  "title": "How to Approach Ensemble // First competition medal [Bronze/missed silver(0.8999)]",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/221669",
  "author_name": "Vatsal Mavani",
  "post_date": "2021-02-23T17:51:36.194000",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have many ensembled models scoring 0.8999 on private which could lead me to the top 100 and on the public, they scored 0.8999 to 0.9012. But I chose 0.9042 and ends up having 363th.</p>\n<p>One of the best unusual things in this competition was no blending works. I just started working for this competition 10 days before the deadline. And trained a few models but didn't make it. Because I trust LB.</p>\n<h4>How to approach Ensembling?</h4>\n<p>I wanted to know more about different Ensembling approaches and how to evaluate the model for ensembling. If anyone has something it can be helpful for new kagglers like me. Share it here.</p>",
  "messages": [
    {
      "id": 1215536,
      "postDate": "2021-02-23T17:51:36.193Z",
      "content": "<p>I have many ensembled models scoring 0.8999 on private which could lead me to the top 100 and on the public, they scored 0.8999 to 0.9012. But I chose 0.9042 and ends up having 363th.</p>\n<p>One of the best unusual things in this competition was no blending works. I just started working for this competition 10 days before the deadline. And trained a few models but didn't make it. Because I trust LB.</p>\n<h4>How to approach Ensembling?</h4>\n<p>I wanted to know more about different Ensembling approaches and how to evaluate the model for ensembling. If anyone has something it can be helpful for new kagglers like me. Share it here.</p>",
      "rawMarkdown": "I have many ensembled models scoring 0.8999 on private which could lead me to the top 100 and on the public, they scored 0.8999 to 0.9012. But I chose 0.9042 and ends up having 363th.\n\nOne of the best unusual things in this competition was no blending works. I just started working for this competition 10 days before the deadline. And trained a few models but didn't make it. Because I trust LB.\n\n#### How to approach Ensembling?\nI wanted to know more about different Ensembling approaches and how to evaluate the model for ensembling. If anyone has something it can be helpful for new kagglers like me. Share it here.",
      "votes": 4
    },
    {
      "id": 1215722,
      "postDate": "2021-02-23T22:24:29.660Z",
      "content": "<p>Congratulations to your first medal!</p>\n<p>There are different ensembling methods:</p>\n<p><strong>- Bagging and Pasting:</strong></p>\n<ol>\n<li>Bagging trains N models on K randomly choosed train samples (with using the same train sample multiple times)</li>\n<li>Pasting trains N models on K randomly choosed train samples (without using the same train sample multiple times)</li>\n</ol>\n<p><strong>- Boosting</strong></p>\n<ol>\n<li>AdaBoosting trains N models on the trainset, models have to be trained sequentially because they are depended on each other. After training a model the model will be evaluated on the trainset, the training samples with the most errors get higher weighted and the next model learns these more</li>\n<li>GradientBoosting trains N models on the trainset, models have to be trained sequentially because they are depended on each other. After training a model the model will be evaluated on the trainset again and the differences will be used to train the next model. The ensemble is the sum of all predictions from each model.</li>\n</ol>\n<p><strong>- Stacking</strong></p>\n<ol>\n<li>Training N models on a subset of the trainset afterwards you predict on the rest of the trainset which you hold out. These predictions will be used to train another model, so called Blender, which tries to weight the prediction of your N models to the Groundtruth value.</li>\n</ol>\n<p>In general the more different your models within an ensemble is the better the accuracy. Because different models tend to do seperate errors and you can compress the error by using multiple models.</p>\n<p>On top of that there are two different ways in handling each output from each model within an ensemble:</p>\n<ul>\n<li>Soft voting: Average of probabilities of each model</li>\n<li>Hard voting: Majority voting which class gets predicted the most gets choosen</li>\n</ul>\n<p>Soft voting usually results in better accuracy and should be choosen if possible.</p>\n<p>You can also give them different weights, I usually give the more accurate model a little higher weight than the others.</p>",
      "rawMarkdown": "Congratulations to your first medal!\n\nThere are different ensembling methods:\n\n**- Bagging and Pasting:**\n1. Bagging trains N models on K randomly choosed train samples (with using the same train sample multiple times)\n2. Pasting trains N models on K randomly choosed train samples (without using the same train sample multiple times)\n\n**- Boosting**\n1. AdaBoosting trains N models on the trainset, models have to be trained sequentially because they are depended on each other. After training a model the model will be evaluated on the trainset, the training samples with the most errors get higher weighted and the next model learns these more\n2. GradientBoosting trains N models on the trainset, models have to be trained sequentially because they are depended on each other. After training a model the model will be evaluated on the trainset again and the differences will be used to train the next model. The ensemble is the sum of all predictions from each model.\n\n**- Stacking**\n1. Training N models on a subset of the trainset afterwards you predict on the rest of the trainset which you hold out. These predictions will be used to train another model, so called Blender, which tries to weight the prediction of your N models to the Groundtruth value.\n\nIn general the more different your models within an ensemble is the better the accuracy. Because different models tend to do seperate errors and you can compress the error by using multiple models.\n\nOn top of that there are two different ways in handling each output from each model within an ensemble:\n\n- Soft voting: Average of probabilities of each model\n- Hard voting: Majority voting which class gets predicted the most gets choosen\n\nSoft voting usually results in better accuracy and should be choosen if possible.\n\nYou can also give them different weights, I usually give the more accurate model a little higher weight than the others.",
      "votes": 2,
      "replies": [
        {
          "id": 1215893,
          "postDate": "2021-02-24T03:02:51.300Z",
          "content": "<p>Thank you for explaining different methods <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a>. For this competition, I trained 5 folds on two different architectures and then used Soft Voting. It works well.</p>\n<ol>\n<li>Is there an implementation notebook if you have one or know about one for these methods?</li>\n<li>Is Boosting method possible(in terms of training time) for the computer vision dataset?</li>\n</ol>",
          "rawMarkdown": "Thank you for explaining different methods @aliabdin1. For this competition, I trained 5 folds on two different architectures and then used Soft Voting. It works well.\n\n1. Is there an implementation notebook if you have one or know about one for these methods?\n2. Is Boosting method possible(in terms of training time) for the computer vision dataset?"
        },
        {
          "id": 1222406,
          "postDate": "2021-03-01T19:03:28.947Z",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a>, according to your description, am I correct that only hold out subset is used for training of Blender model for Stacking? (i.e. part that participated in CV is not used here)</p>",
          "rawMarkdown": "@aliabdin1, according to your description, am I correct that only hold out subset is used for training of Blender model for Stacking? (i.e. part that participated in CV is not used here)"
        },
        {
          "id": 1225405,
          "postDate": "2021-03-03T15:43:48.583Z",
          "content": "<p><a href=\"https://www.kaggle.com/dmitrynovikov\" target=\"_blank\">@dmitrynovikov</a> Yes that is right. You would use the hold out subset to evaluate your trained models on these and the predictions will be used to train your Blender model (Input: predictions from N trained models evaluated on holdout subset Target: ground truth data).</p>",
          "rawMarkdown": "@dmitrynovikov Yes that is right. You would use the hold out subset to evaluate your trained models on these and the predictions will be used to train your Blender model (Input: predictions from N trained models evaluated on holdout subset Target: ground truth data).",
          "votes": 1
        },
        {
          "id": 1225540,
          "postDate": "2021-03-03T17:13:26.313Z",
          "content": "<p>Thank you for your explanation</p>",
          "rawMarkdown": "Thank you for your explanation"
        }
      ]
    },
    {
      "id": 1216236,
      "postDate": "2021-02-24T06:47:18.247Z",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!"
    },
    {
      "id": 1215544,
      "postDate": "2021-02-23T18:03:58.957Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1225116,
      "postDate": "2021-03-03T10:32:26.910Z",
      "content": "<p>Thank you very much</p>",
      "rawMarkdown": "Thank you very much"
    }
  ],
  "comments": [
    {
      "id": 1215722,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2021-02-23T22:24:29.660000",
      "content": "<p>Congratulations to your first medal!</p>\n<p>There are different ensembling methods:</p>\n<p><strong>- Bagging and Pasting:</strong></p>\n<ol>\n<li>Bagging trains N models on K randomly choosed train samples (with using the same train sample multiple times)</li>\n<li>Pasting trains N models on K randomly choosed train samples (without using the same train sample multiple times)</li>\n</ol>\n<p><strong>- Boosting</strong></p>\n<ol>\n<li>AdaBoosting trains N models on the trainset, models have to be trained sequentially because they are depended on each other. After training a model the model will be evaluated on the trainset, the training samples with the most errors get higher weighted and the next model learns these more</li>\n<li>GradientBoosting trains N models on the trainset, models have to be trained sequentially because they are depended on each other. After training a model the model will be evaluated on the trainset again and the differences will be used to train the next model. The ensemble is the sum of all predictions from each model.</li>\n</ol>\n<p><strong>- Stacking</strong></p>\n<ol>\n<li>Training N models on a subset of the trainset afterwards you predict on the rest of the trainset which you hold out. These predictions will be used to train another model, so called Blender, which tries to weight the prediction of your N models to the Groundtruth value.</li>\n</ol>\n<p>In general the more different your models within an ensemble is the better the accuracy. Because different models tend to do seperate errors and you can compress the error by using multiple models.</p>\n<p>On top of that there are two different ways in handling each output from each model within an ensemble:</p>\n<ul>\n<li>Soft voting: Average of probabilities of each model</li>\n<li>Hard voting: Majority voting which class gets predicted the most gets choosen</li>\n</ul>\n<p>Soft voting usually results in better accuracy and should be choosen if possible.</p>\n<p>You can also give them different weights, I usually give the more accurate model a little higher weight than the others.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1215893,
          "author_name": "Vatsal Mavani",
          "author_url": "",
          "post_date": "2021-02-24T03:02:51.300000",
          "content": "<p>Thank you for explaining different methods <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a>. For this competition, I trained 5 folds on two different architectures and then used Soft Voting. It works well.</p>\n<ol>\n<li>Is there an implementation notebook if you have one or know about one for these methods?</li>\n<li>Is Boosting method possible(in terms of training time) for the computer vision dataset?</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222406,
          "author_name": "DN",
          "author_url": "",
          "post_date": "2021-03-01T19:03:28.947000",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a>, according to your description, am I correct that only hold out subset is used for training of Blender model for Stacking? (i.e. part that participated in CV is not used here)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1225405,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2021-03-03T15:43:48.583000",
          "content": "<p><a href=\"https://www.kaggle.com/dmitrynovikov\" target=\"_blank\">@dmitrynovikov</a> Yes that is right. You would use the hold out subset to evaluate your trained models on these and the predictions will be used to train your Blender model (Input: predictions from N trained models evaluated on holdout subset Target: ground truth data).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1225540,
          "author_name": "DN",
          "author_url": "",
          "post_date": "2021-03-03T17:13:26.313000",
          "content": "<p>Thank you for your explanation</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1216236,
      "author_name": "Himanshu Mehndiratta",
      "author_url": "",
      "post_date": "2021-02-24T06:47:18.247000",
      "content": "<p>Congratulations!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1215544,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-23T18:03:58.957000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1225116,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-03T10:32:26.910000",
      "content": "<p>Thank you very much</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1215536": "I have many ensembled models scoring 0.8999 on private which could lead me to the top 100 and on the public, they scored 0.8999 to 0.9012. But I chose 0.9042 and ends up having 363th.\n\nOne of the best unusual things in this competition was no blending works. I just started working for this competition 10 days before the deadline. And trained a few models but didn't make it. Because I trust LB.\n\n#### How to approach Ensembling?\nI wanted to know more about different Ensembling approaches and how to evaluate the model for ensembling. If anyone has something it can be helpful for new kagglers like me. Share it here.",
    "1215722": "Congratulations to your first medal!\n\nThere are different ensembling methods:\n\n**- Bagging and Pasting:**\n1. Bagging trains N models on K randomly choosed train samples (with using the same train sample multiple times)\n2. Pasting trains N models on K randomly choosed train samples (without using the same train sample multiple times)\n\n**- Boosting**\n1. AdaBoosting trains N models on the trainset, models have to be trained sequentially because they are depended on each other. After training a model the model will be evaluated on the trainset, the training samples with the most errors get higher weighted and the next model learns these more\n2. GradientBoosting trains N models on the trainset, models have to be trained sequentially because they are depended on each other. After training a model the model will be evaluated on the trainset again and the differences will be used to train the next model. The ensemble is the sum of all predictions from each model.\n\n**- Stacking**\n1. Training N models on a subset of the trainset afterwards you predict on the rest of the trainset which you hold out. These predictions will be used to train another model, so called Blender, which tries to weight the prediction of your N models to the Groundtruth value.\n\nIn general the more different your models within an ensemble is the better the accuracy. Because different models tend to do seperate errors and you can compress the error by using multiple models.\n\nOn top of that there are two different ways in handling each output from each model within an ensemble:\n\n- Soft voting: Average of probabilities of each model\n- Hard voting: Majority voting which class gets predicted the most gets choosen\n\nSoft voting usually results in better accuracy and should be choosen if possible.\n\nYou can also give them different weights, I usually give the more accurate model a little higher weight than the others.",
    "1216236": "Congratulations!!",
    "1215544": "",
    "1225116": "Thank you very much"
  }
}