{
  "id": 54395,
  "title": "Why was MAP@3 metric chosen for this competition?",
  "url": "/competitions/freesound-audio-tagging/discussion/54395",
  "author_name": "",
  "post_date": "2018-04-12T19:08:14.896602900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is my first time seeing this metric. I looked up the metric and it looks like something that would be good when order of correct prediction matters. I am not sure why would that be the case here. Can someone share why the metric is a good fit here? </p>",
  "messages": [
    {
      "id": "313091",
      "postDate": "04/12/2018 19:08:14",
      "content": "<p>This is my first time seeing this metric. I looked up the metric and it looks like something that would be good when order of correct prediction matters. I am not sure why would that be the case here. Can someone share why the metric is a good fit here? </p>",
      "rawMarkdown": "This is my first time seeing this metric. I looked up the metric and it looks like something that would be good when order of correct prediction matters. I am not sure why would that be the case here. Can someone share why the metric is a good fit here?",
      "votes": null
    },
    {
      "id": "313178",
      "postDate": "04/12/2018 21:32:38",
      "content": "<p>Since there's only one correct label per clip, my understanding is that the metric gives a score (per clip) of 1, 0.50, 0.33 or 0 depending on the rank the correct label was given in the predictions. The metric is a good fit if there is some merit to being <em>almost</em> correct. For example, the overview was trying to make the point that it's difficult, even for humans, to distinguish between chainsaw and blender.</p>",
      "rawMarkdown": "Since there's only one correct label per clip, my understanding is that the metric gives a score (per clip) of 1, 0.50, 0.33 or 0 depending on the rank the correct label was given in the predictions. The metric is a good fit if there is some merit to being *almost* correct. For example, the overview was trying to make the point that it's difficult, even for humans, to distinguish between chainsaw and blender.",
      "votes": null
    },
    {
      "id": "323667",
      "postDate": "05/05/2018 20:36:08",
      "content": "<p>Hi Aseem,</p>\n\n<p>the idea behind this metric is to give \"partial credit\" to predictions that are <em>almost</em> correct, as Turab pointed out. Specifically, the 3 (as in MAP@<strong>3</strong>) most probable guesses per sample can count for the score (with 1, 1/2 and 1/3, respectively).</p>\n\n<p>An alternative, for instance, would be to use % Accuracy as evaluation metric, i.e., \"all or nothing\". Which evaluation metric would you consider appropriate for this task?</p>",
      "rawMarkdown": "Hi Aseem,\n\nthe idea behind this metric is to give \"partial credit\" to predictions that are *almost* correct, as Turab pointed out. Specifically, the 3 (as in MAP@**3**) most probable guesses per sample can count for the score (with 1, 1/2 and 1/3, respectively).\n\nAn alternative, for instance, would be to use % Accuracy as evaluation metric, i.e., \"all or nothing\". Which evaluation metric would you consider appropriate for this task?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 313178,
      "author_name": "tqbl95",
      "author_url": "",
      "post_date": "04/12/2018 21:32:38",
      "content": "<p>Since there's only one correct label per clip, my understanding is that the metric gives a score (per clip) of 1, 0.50, 0.33 or 0 depending on the rank the correct label was given in the predictions. The metric is a good fit if there is some merit to being <em>almost</em> correct. For example, the overview was trying to make the point that it's difficult, even for humans, to distinguish between chainsaw and blender.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 323667,
      "author_name": "eduardofonseca",
      "author_url": "",
      "post_date": "05/05/2018 20:36:08",
      "content": "<p>Hi Aseem,</p>\n\n<p>the idea behind this metric is to give \"partial credit\" to predictions that are <em>almost</em> correct, as Turab pointed out. Specifically, the 3 (as in MAP@<strong>3</strong>) most probable guesses per sample can count for the score (with 1, 1/2 and 1/3, respectively).</p>\n\n<p>An alternative, for instance, would be to use % Accuracy as evaluation metric, i.e., \"all or nothing\". Which evaluation metric would you consider appropriate for this task?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "313091": "This is my first time seeing this metric. I looked up the metric and it looks like something that would be good when order of correct prediction matters. I am not sure why would that be the case here. Can someone share why the metric is a good fit here?",
    "313178": "Since there's only one correct label per clip, my understanding is that the metric gives a score (per clip) of 1, 0.50, 0.33 or 0 depending on the rank the correct label was given in the predictions. The metric is a good fit if there is some merit to being *almost* correct. For example, the overview was trying to make the point that it's difficult, even for humans, to distinguish between chainsaw and blender.",
    "323667": "Hi Aseem,\n\nthe idea behind this metric is to give \"partial credit\" to predictions that are *almost* correct, as Turab pointed out. Specifically, the 3 (as in MAP@**3**) most probable guesses per sample can count for the score (with 1, 1/2 and 1/3, respectively).\n\nAn alternative, for instance, would be to use % Accuracy as evaluation metric, i.e., \"all or nothing\". Which evaluation metric would you consider appropriate for this task?"
  },
  "source": "meta"
}