{
  "id": 207281,
  "title": "Why do we take the mean during kfold/ensembling/tta?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207281",
  "author_name": "",
  "post_date": "2020-12-29T01:25:38.026944500Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone! Correct me if I'm wrong but I'm operating on the assumption that our model outputs one-hot predictions for the 5 classes like so: [0.1, 0.2, 0.5, 0.1, 0.1]</p>\n<p>In this case, when we have multiple predictions from kfold/ensembling/tta, why do we take the mean  values? If I have the previous one-hot prediction and another like [0.7, 0.05, 0.05, 0.1, 0.1], then the mean would give me [0.4, 0.125, 0.275, 0.1, 0.1], so the prediction overall would be label 1 instead of label 3. </p>\n<p>Does that mean we want to give more weight to more confident predictions?</p>",
  "messages": [
    {
      "id": "1130375",
      "postDate": "12/29/2020 01:25:38",
      "content": "<p>Hi everyone! Correct me if I'm wrong but I'm operating on the assumption that our model outputs one-hot predictions for the 5 classes like so: [0.1, 0.2, 0.5, 0.1, 0.1]</p>\n<p>In this case, when we have multiple predictions from kfold/ensembling/tta, why do we take the mean  values? If I have the previous one-hot prediction and another like [0.7, 0.05, 0.05, 0.1, 0.1], then the mean would give me [0.4, 0.125, 0.275, 0.1, 0.1], so the prediction overall would be label 1 instead of label 3. </p>\n<p>Does that mean we want to give more weight to more confident predictions?</p>",
      "rawMarkdown": "Hi everyone! Correct me if I'm wrong but I'm operating on the assumption that our model outputs one-hot predictions for the 5 classes like so: [0.1, 0.2, 0.5, 0.1, 0.1]\n\nIn this case, when we have multiple predictions from kfold/ensembling/tta, why do we take the mean  values? If I have the previous one-hot prediction and another like [0.7, 0.05, 0.05, 0.1, 0.1], then the mean would give me [0.4, 0.125, 0.275, 0.1, 0.1], so the prediction overall would be label 1 instead of label 3. \n\nDoes that mean we want to give more weight to more confident predictions?",
      "votes": null
    },
    {
      "id": "1130945",
      "postDate": "12/29/2020 12:47:01",
      "content": "<p>The mean is not your only choice, it just tends to often be a reasonable one. Additionally, there's the nice property that an (weighted) arithmetic mean ensures the probabilities will still add up to 1.</p>\n<p>For TTA or combining results from different folds some kind of unweighted average like mean, harmonic mean, geometric mean etc. will often make sense, because in a way there's no reason to believe that one prediction is better than another.</p>\n<p>Alternatives for stacking/ensembles include some kind of model on the predictions, or just weighted averages. When you combine very different models, you'd expect that some might truly in an identifiable way be better than the others.</p>",
      "rawMarkdown": "The mean is not your only choice, it just tends to often be a reasonable one. Additionally, there's the nice property that an (weighted) arithmetic mean ensures the probabilities will still add up to 1.\n\nFor TTA or combining results from different folds some kind of unweighted average like mean, harmonic mean, geometric mean etc. will often make sense, because in a way there's no reason to believe that one prediction is better than another.\n\nAlternatives for stacking/ensembles include some kind of model on the predictions, or just weighted averages. When you combine very different models, you'd expect that some might truly in an identifiable way be better than the others.",
      "votes": null
    },
    {
      "id": "1131016",
      "postDate": "12/29/2020 13:54:28",
      "content": "<p>Thanks a lot for the detailed reply! It helped me :)</p>",
      "rawMarkdown": "Thanks a lot for the detailed reply! It helped me :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1130945,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "12/29/2020 12:47:01",
      "content": "<p>The mean is not your only choice, it just tends to often be a reasonable one. Additionally, there's the nice property that an (weighted) arithmetic mean ensures the probabilities will still add up to 1.</p>\n<p>For TTA or combining results from different folds some kind of unweighted average like mean, harmonic mean, geometric mean etc. will often make sense, because in a way there's no reason to believe that one prediction is better than another.</p>\n<p>Alternatives for stacking/ensembles include some kind of model on the predictions, or just weighted averages. When you combine very different models, you'd expect that some might truly in an identifiable way be better than the others.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131016,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "12/29/2020 13:54:28",
          "content": "<p>Thanks a lot for the detailed reply! It helped me :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1130375": "Hi everyone! Correct me if I'm wrong but I'm operating on the assumption that our model outputs one-hot predictions for the 5 classes like so: [0.1, 0.2, 0.5, 0.1, 0.1]\n\nIn this case, when we have multiple predictions from kfold/ensembling/tta, why do we take the mean  values? If I have the previous one-hot prediction and another like [0.7, 0.05, 0.05, 0.1, 0.1], then the mean would give me [0.4, 0.125, 0.275, 0.1, 0.1], so the prediction overall would be label 1 instead of label 3. \n\nDoes that mean we want to give more weight to more confident predictions?",
    "1130945": "The mean is not your only choice, it just tends to often be a reasonable one. Additionally, there's the nice property that an (weighted) arithmetic mean ensures the probabilities will still add up to 1.\n\nFor TTA or combining results from different folds some kind of unweighted average like mean, harmonic mean, geometric mean etc. will often make sense, because in a way there's no reason to believe that one prediction is better than another.\n\nAlternatives for stacking/ensembles include some kind of model on the predictions, or just weighted averages. When you combine very different models, you'd expect that some might truly in an identifiable way be better than the others.",
    "1131016": "Thanks a lot for the detailed reply! It helped me :)"
  },
  "source": "meta"
}