{
  "id": 206072,
  "title": "Voting vs Adding at Inference with multiple model or Augmentation",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/206072",
  "author_name": "",
  "post_date": "2020-12-23T05:38:40.533514700Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>When there are multiple models, or TTA are being used at inference time, there are two methods to get the final prediction results:</p>\n<p>Method 1:    Add up all the predictions  y = y1 + y2 + … yn,<br>\n                     predictions = y.argmax(axis=1)</p>\n<p>Method 2:   Take the argmax of each prediction:<br>\n                      p1 = y1.argmax(axis =1 )<br>\n                      p2= y2.argmax(axis =1 )<br>\n                        …………..<br>\n                      p2 = yn.argmax(axis =1 )<br>\n                     Then find the most common predictions among all individual prediction p1, p2, … pn, and use that for the results.</p>\n<p>Are there any research which method yield better results?  Has anyone compared them?</p>",
  "messages": [
    {
      "id": "1123294",
      "postDate": "12/23/2020 05:38:40",
      "content": "<p>When there are multiple models, or TTA are being used at inference time, there are two methods to get the final prediction results:</p>\n<p>Method 1:    Add up all the predictions  y = y1 + y2 + … yn,<br>\n                     predictions = y.argmax(axis=1)</p>\n<p>Method 2:   Take the argmax of each prediction:<br>\n                      p1 = y1.argmax(axis =1 )<br>\n                      p2= y2.argmax(axis =1 )<br>\n                        …………..<br>\n                      p2 = yn.argmax(axis =1 )<br>\n                     Then find the most common predictions among all individual prediction p1, p2, … pn, and use that for the results.</p>\n<p>Are there any research which method yield better results?  Has anyone compared them?</p>",
      "rawMarkdown": "When there are multiple models, or TTA are being used at inference time, there are two methods to get the final prediction results:\n\nMethod 1:    Add up all the predictions  y = y1 + y2 + ... yn,\n                     predictions = y.argmax(axis=1)\n\nMethod 2:   Take the argmax of each prediction:\n                      p1 = y1.argmax(axis =1 )\n                      p2= y2.argmax(axis =1 )\n                        ..............\n                      p2 = yn.argmax(axis =1 )\n                     Then find the most common predictions among all individual prediction p1, p2, ... pn, and use that for the results.\n\nAre there any research which method yield better results?  Has anyone compared them?",
      "votes": null
    },
    {
      "id": "1123533",
      "postDate": "12/23/2020 10:01:22",
      "content": "<p>I prefer the first one. When you make a vote system (your 2nd version), you will miss the information about confidences in the other categories except the biggest one<br>\nFor example: if you have the probabilities ( 0.40, 0.39, 0.07, 0.07, 0.07) you will be missing the information that the 2nd probability is almost as good as first. You will just go in with: this is 1<br>\nIf you sum the probabilities of all possible outputs (1st version), your system will have a better understanding of the phenomenon.<br>\nIf you really want to go deeper with the interpretation of the outputs you can go with a meta model</p>",
      "rawMarkdown": "I prefer the first one. When you make a vote system (your 2nd version), you will miss the information about confidences in the other categories except the biggest one\nFor example: if you have the probabilities ( 0.40, 0.39, 0.07, 0.07, 0.07) you will be missing the information that the 2nd probability is almost as good as first. You will just go in with: this is 1\nIf you sum the probabilities of all possible outputs (1st version), your system will have a better understanding of the phenomenon.\nIf you really want to go deeper with the interpretation of the outputs you can go with a meta model",
      "votes": null
    },
    {
      "id": "1124304",
      "postDate": "12/23/2020 20:02:03",
      "content": "<p>The first technique is the most used and consistent one. For the second method you will need to have a large amount of models available to get an \"accurate\" voting.</p>",
      "rawMarkdown": "The first technique is the most used and consistent one. For the second method you will need to have a large amount of models available to get an \"accurate\" voting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1123533,
      "author_name": "vladvdv",
      "author_url": "",
      "post_date": "12/23/2020 10:01:22",
      "content": "<p>I prefer the first one. When you make a vote system (your 2nd version), you will miss the information about confidences in the other categories except the biggest one<br>\nFor example: if you have the probabilities ( 0.40, 0.39, 0.07, 0.07, 0.07) you will be missing the information that the 2nd probability is almost as good as first. You will just go in with: this is 1<br>\nIf you sum the probabilities of all possible outputs (1st version), your system will have a better understanding of the phenomenon.<br>\nIf you really want to go deeper with the interpretation of the outputs you can go with a meta model</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1124304,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "12/23/2020 20:02:03",
      "content": "<p>The first technique is the most used and consistent one. For the second method you will need to have a large amount of models available to get an \"accurate\" voting.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1123294": "When there are multiple models, or TTA are being used at inference time, there are two methods to get the final prediction results:\n\nMethod 1:    Add up all the predictions  y = y1 + y2 + ... yn,\n                     predictions = y.argmax(axis=1)\n\nMethod 2:   Take the argmax of each prediction:\n                      p1 = y1.argmax(axis =1 )\n                      p2= y2.argmax(axis =1 )\n                        ..............\n                      p2 = yn.argmax(axis =1 )\n                     Then find the most common predictions among all individual prediction p1, p2, ... pn, and use that for the results.\n\nAre there any research which method yield better results?  Has anyone compared them?",
    "1123533": "I prefer the first one. When you make a vote system (your 2nd version), you will miss the information about confidences in the other categories except the biggest one\nFor example: if you have the probabilities ( 0.40, 0.39, 0.07, 0.07, 0.07) you will be missing the information that the 2nd probability is almost as good as first. You will just go in with: this is 1\nIf you sum the probabilities of all possible outputs (1st version), your system will have a better understanding of the phenomenon.\nIf you really want to go deeper with the interpretation of the outputs you can go with a meta model",
    "1124304": "The first technique is the most used and consistent one. For the second method you will need to have a large amount of models available to get an \"accurate\" voting."
  },
  "source": "meta"
}