{
  "id": 44223,
  "title": "Evaluation metric",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44223",
  "author_name": "",
  "post_date": "2017-11-25T14:49:40.075748400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,\nI have a question regarding the evaluation metric. In the overview tab it says: \"Submissions are evaluated on Multiclass Accuracy, which is simply the average number of observations with the correct label.\" Is this the mean accuracy? E.g. for 3 classes: Accuracies of [0.0, 1.0, 0.0] results in mean Acc: 0.333. If this is the case, why does the \"Silence is Golden Benchmark\" give 0.09? After all: 1/12 = 0.083 and should therefore result in 0.08? What am I missing? Thanks!</p>",
  "messages": [
    {
      "id": "248285",
      "postDate": "11/25/2017 14:49:40",
      "content": "<p>Hello everyone,\nI have a question regarding the evaluation metric. In the overview tab it says: \"Submissions are evaluated on Multiclass Accuracy, which is simply the average number of observations with the correct label.\" Is this the mean accuracy? E.g. for 3 classes: Accuracies of [0.0, 1.0, 0.0] results in mean Acc: 0.333. If this is the case, why does the \"Silence is Golden Benchmark\" give 0.09? After all: 1/12 = 0.083 and should therefore result in 0.08? What am I missing? Thanks!</p>",
      "rawMarkdown": "Hello everyone,\nI have a question regarding the evaluation metric. In the overview tab it says: \"Submissions are evaluated on Multiclass Accuracy, which is simply the average number of observations with the correct label.\" Is this the mean accuracy? E.g. for 3 classes: Accuracies of [0.0, 1.0, 0.0] results in mean Acc: 0.333. If this is the case, why does the \"Silence is Golden Benchmark\" give 0.09? After all: 1/12 = 0.083 and should therefore result in 0.08? What am I missing? Thanks!",
      "votes": null
    },
    {
      "id": "248294",
      "postDate": "11/25/2017 15:31:44",
      "content": "<p>It could be that the class distribution is not perfectly balanced in the public test set.</p>",
      "rawMarkdown": "It could be that the class distribution is not perfectly balanced in the public test set.",
      "votes": null
    },
    {
      "id": "248300",
      "postDate": "11/25/2017 16:10:31",
      "content": "<p>Got it, thanks! I disregarded the private part.</p>",
      "rawMarkdown": "Got it, thanks! I disregarded the private part.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 248294,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "11/25/2017 15:31:44",
      "content": "<p>It could be that the class distribution is not perfectly balanced in the public test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 248300,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "11/25/2017 16:10:31",
          "content": "<p>Got it, thanks! I disregarded the private part.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "248285": "Hello everyone,\nI have a question regarding the evaluation metric. In the overview tab it says: \"Submissions are evaluated on Multiclass Accuracy, which is simply the average number of observations with the correct label.\" Is this the mean accuracy? E.g. for 3 classes: Accuracies of [0.0, 1.0, 0.0] results in mean Acc: 0.333. If this is the case, why does the \"Silence is Golden Benchmark\" give 0.09? After all: 1/12 = 0.083 and should therefore result in 0.08? What am I missing? Thanks!",
    "248294": "It could be that the class distribution is not perfectly balanced in the public test set.",
    "248300": "Got it, thanks! I disregarded the private part."
  },
  "source": "meta"
}