{
  "id": 71947,
  "title": "How could I ensemble in this kind  of data output",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/71947",
  "author_name": "",
  "post_date": "2018-11-18T16:38:04.747989400Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello ,\nI am interested to find  some ideea  regarding  how could I ensemble in this kind  of competition with this   kind  of output .</p>\n\n<p>Tnks</p>",
  "messages": [
    {
      "id": "423593",
      "postDate": "11/18/2018 16:38:04",
      "content": "<p>Hello ,\nI am interested to find  some ideea  regarding  how could I ensemble in this kind  of competition with this   kind  of output .</p>\n\n<p>Tnks</p>",
      "rawMarkdown": "Hello ,\nI am interested to find  some ideea  regarding  how could I ensemble in this kind  of competition with this   kind  of output .\n\nTnks",
      "votes": null
    },
    {
      "id": "423691",
      "postDate": "11/18/2018 22:08:09",
      "content": "<p>I would suggest starting your ensemble by averaging the outputs of several models. Let's say you have 5 models you want to ensemble together. Each of the models has a prediction for each image in the test set, so you'd have 5 arrays of length 28 (for 28 classes) for each test image. To make a prediction using the ensemble, you'd take the average of those 5 arrays and end up with a single array of length 28.</p>\n\n<p>One thing you might want to test out is whether to do the averaging before thresholding or after thresholding, see what gives you a better result.</p>",
      "rawMarkdown": "I would suggest starting your ensemble by averaging the outputs of several models. Let's say you have 5 models you want to ensemble together. Each of the models has a prediction for each image in the test set, so you'd have 5 arrays of length 28 (for 28 classes) for each test image. To make a prediction using the ensemble, you'd take the average of those 5 arrays and end up with a single array of length 28.\n\nOne thing you might want to test out is whether to do the averaging before thresholding or after thresholding, see what gives you a better result.",
      "votes": null
    },
    {
      "id": "424047",
      "postDate": "11/19/2018 13:23:23",
      "content": "<p>I ensemble by averaging the raw numerical predictions of different models before thresholding, and that has produced significant score improvements over single models.  I haven't tried averaging after thresholding, but it seems like that would discard some useful information.  It would also produce fractional results, so you would have to threshold again to get final predictions.</p>",
      "rawMarkdown": "I ensemble by averaging the raw numerical predictions of different models before thresholding, and that has produced significant score improvements over single models.  I haven't tried averaging after thresholding, but it seems like that would discard some useful information.  It would also produce fractional results, so you would have to threshold again to get final predictions.",
      "votes": null
    },
    {
      "id": "425828",
      "postDate": "11/22/2018 07:14:58",
      "content": "<p>Contrary to what has been reported by others, I see larger improvements on the LB by hard-voting rather than averaging the probabilities. </p>",
      "rawMarkdown": "Contrary to what has been reported by others, I see larger improvements on the LB by hard-voting rather than averaging the probabilities.",
      "votes": null
    },
    {
      "id": "426311",
      "postDate": "11/23/2018 03:37:01",
      "content": "<p>@Chase the Trane\nCan you explain what you mean by \"hard-voting\"?  For example, what happens if you have two models and one votes for a particular class but the other votes against it?</p>",
      "rawMarkdown": "Chase the Trane\nCan you explain what you mean by \"hard-voting\"?  For example, what happens if you have two models and one votes for a particular class but the other votes against it?",
      "votes": null
    },
    {
      "id": "426349",
      "postDate": "11/23/2018 05:43:30",
      "content": "<p><a href=\"/dslate\">@dslate</a> People usually do use odd number of models on majority voting</p>",
      "rawMarkdown": "dslate People usually do use odd number of models on majority voting",
      "votes": null
    },
    {
      "id": "426381",
      "postDate": "11/23/2018 06:51:42",
      "content": "<p>I try to use odd number of models. However, I noticed that with even number of models when you have a draw it's better to vote in favour of the class, at least in average. So in your case (1-1) I would vote in favour of the class.</p>",
      "rawMarkdown": "I try to use odd number of models. However, I noticed that with even number of models when you have a draw it's better to vote in favour of the class, at least in average. So in your case (1-1) I would vote in favour of the class.",
      "votes": null
    },
    {
      "id": "426774",
      "postDate": "11/23/2018 20:15:57",
      "content": "<p><a href=\"/stecasasso\">@stecasasso</a> It probably depends on the loss you use too. When using even numbers of models, would it be helpful to select the decision based on the model that performs better in local CV when there is a (1-1)?</p>",
      "rawMarkdown": "stecasasso It probably depends on the loss you use too. When using even numbers of models, would it be helpful to select the decision based on the model that performs better in local CV when there is a (1-1)?",
      "votes": null
    },
    {
      "id": "426837",
      "postDate": "11/24/2018 01:02:03",
      "content": "<p>Thanks Hanke Chen and Chase the Trane .  After I asked my question it occurred to me that an odd number of models solves it.  I will try majority voting, although intuitively it doesn't seem like it should be as accurate as averaging the raw probabilities and then thresholding.</p>",
      "rawMarkdown": "Thanks Hanke Chen and Chase the Trane .  After I asked my question it occurred to me that an odd number of models solves it.  I will try majority voting, although intuitively it doesn't seem like it should be as accurate as averaging the raw probabilities and then thresholding.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 423691,
      "author_name": "hortonhearsafoo",
      "author_url": "",
      "post_date": "11/18/2018 22:08:09",
      "content": "<p>I would suggest starting your ensemble by averaging the outputs of several models. Let's say you have 5 models you want to ensemble together. Each of the models has a prediction for each image in the test set, so you'd have 5 arrays of length 28 (for 28 classes) for each test image. To make a prediction using the ensemble, you'd take the average of those 5 arrays and end up with a single array of length 28.</p>\n\n<p>One thing you might want to test out is whether to do the averaging before thresholding or after thresholding, see what gives you a better result.</p>",
      "votes": null,
      "replies": [
        {
          "id": 424047,
          "author_name": "dslate",
          "author_url": "",
          "post_date": "11/19/2018 13:23:23",
          "content": "<p>I ensemble by averaging the raw numerical predictions of different models before thresholding, and that has produced significant score improvements over single models.  I haven't tried averaging after thresholding, but it seems like that would discard some useful information.  It would also produce fractional results, so you would have to threshold again to get final predictions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 425828,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "11/22/2018 07:14:58",
      "content": "<p>Contrary to what has been reported by others, I see larger improvements on the LB by hard-voting rather than averaging the probabilities. </p>",
      "votes": null,
      "replies": [
        {
          "id": 426311,
          "author_name": "dslate",
          "author_url": "",
          "post_date": "11/23/2018 03:37:01",
          "content": "<p>@Chase the Trane\nCan you explain what you mean by \"hard-voting\"?  For example, what happens if you have two models and one votes for a particular class but the other votes against it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 426349,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/23/2018 05:43:30",
          "content": "<p><a href=\"/dslate\">@dslate</a> People usually do use odd number of models on majority voting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 426381,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "11/23/2018 06:51:42",
          "content": "<p>I try to use odd number of models. However, I noticed that with even number of models when you have a draw it's better to vote in favour of the class, at least in average. So in your case (1-1) I would vote in favour of the class.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 426774,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/23/2018 20:15:57",
          "content": "<p><a href=\"/stecasasso\">@stecasasso</a> It probably depends on the loss you use too. When using even numbers of models, would it be helpful to select the decision based on the model that performs better in local CV when there is a (1-1)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 426837,
          "author_name": "dslate",
          "author_url": "",
          "post_date": "11/24/2018 01:02:03",
          "content": "<p>Thanks Hanke Chen and Chase the Trane .  After I asked my question it occurred to me that an odd number of models solves it.  I will try majority voting, although intuitively it doesn't seem like it should be as accurate as averaging the raw probabilities and then thresholding.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "423593": "Hello ,\nI am interested to find  some ideea  regarding  how could I ensemble in this kind  of competition with this   kind  of output .\n\nTnks",
    "423691": "I would suggest starting your ensemble by averaging the outputs of several models. Let's say you have 5 models you want to ensemble together. Each of the models has a prediction for each image in the test set, so you'd have 5 arrays of length 28 (for 28 classes) for each test image. To make a prediction using the ensemble, you'd take the average of those 5 arrays and end up with a single array of length 28.\n\nOne thing you might want to test out is whether to do the averaging before thresholding or after thresholding, see what gives you a better result.",
    "424047": "I ensemble by averaging the raw numerical predictions of different models before thresholding, and that has produced significant score improvements over single models.  I haven't tried averaging after thresholding, but it seems like that would discard some useful information.  It would also produce fractional results, so you would have to threshold again to get final predictions.",
    "425828": "Contrary to what has been reported by others, I see larger improvements on the LB by hard-voting rather than averaging the probabilities.",
    "426311": "Chase the Trane\nCan you explain what you mean by \"hard-voting\"?  For example, what happens if you have two models and one votes for a particular class but the other votes against it?",
    "426349": "dslate People usually do use odd number of models on majority voting",
    "426381": "I try to use odd number of models. However, I noticed that with even number of models when you have a draw it's better to vote in favour of the class, at least in average. So in your case (1-1) I would vote in favour of the class.",
    "426774": "stecasasso It probably depends on the loss you use too. When using even numbers of models, would it be helpful to select the decision based on the model that performs better in local CV when there is a (1-1)?",
    "426837": "Thanks Hanke Chen and Chase the Trane .  After I asked my question it occurred to me that an odd number of models solves it.  I will try majority voting, although intuitively it doesn't seem like it should be as accurate as averaging the raw probabilities and then thresholding."
  },
  "source": "meta"
}