{
  "id": 46618,
  "title": "Which submission are used for the final score ? ",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/46618",
  "author_name": "",
  "post_date": "2017-12-30T19:13:01.883756200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Someone showed in a <a href=\"https://www.kaggle.com/ironbar/frequency-of-the-labels-in-the-public-test-set\">kernel</a> that the number of occurrences of each class is similar on LB. Yet, we can expect the final testing set to have more unknown samples than any other class because:</p>\n\n<ol>\n<li>The final goal of the dataset is probably to have 32 classes with approximately the same number of elements, not just 11+1 classes</li>\n<li>The training set has approximately 60% of elements of the class unknown</li>\n<li>Someone achieved 0.80 on LB while predicting 47% of unknown labels (<a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298\">here</a>)</li>\n</ol>\n\n<p>Then there could be a solution that has 11% of error for unknowns and ~1% for others on LB, that could perform worse than a solution with e.g. 11% of error for yes and ~1% for others (including unknown) because unknown should have many more samples.\nThus, is the best submission the only one used for the final score or are there more submissions taken into account ?</p>",
  "messages": [
    {
      "id": "263558",
      "postDate": "12/30/2017 19:13:01",
      "content": "<p>Someone showed in a <a href=\"https://www.kaggle.com/ironbar/frequency-of-the-labels-in-the-public-test-set\">kernel</a> that the number of occurrences of each class is similar on LB. Yet, we can expect the final testing set to have more unknown samples than any other class because:</p>\n\n<ol>\n<li>The final goal of the dataset is probably to have 32 classes with approximately the same number of elements, not just 11+1 classes</li>\n<li>The training set has approximately 60% of elements of the class unknown</li>\n<li>Someone achieved 0.80 on LB while predicting 47% of unknown labels (<a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298\">here</a>)</li>\n</ol>\n\n<p>Then there could be a solution that has 11% of error for unknowns and ~1% for others on LB, that could perform worse than a solution with e.g. 11% of error for yes and ~1% for others (including unknown) because unknown should have many more samples.\nThus, is the best submission the only one used for the final score or are there more submissions taken into account ?</p>",
      "rawMarkdown": "Someone showed in a [kernel][1] that the number of occurrences of each class is similar on LB. Yet, we can expect the final testing set to have more unknown samples than any other class because:\n\n 1. The final goal of the dataset is probably to have 32 classes with approximately the same number of elements, not just 11+1 classes\n 2. The training set has approximately 60% of elements of the class unknown\n 3. Someone achieved 0.80 on LB while predicting 47% of unknown labels ([here][2])\n\nThen there could be a solution that has 11% of error for unknowns and ~1% for others on LB, that could perform worse than a solution with e.g. 11% of error for yes and ~1% for others (including unknown) because unknown should have many more samples.\nThus, is the best submission the only one used for the final score or are there more submissions taken into account ?\n\n  [1]: https://www.kaggle.com/ironbar/frequency-of-the-labels-in-the-public-test-set\n  [2]: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298",
      "votes": null
    },
    {
      "id": "264302",
      "postDate": "01/02/2018 18:56:29",
      "content": "<p>You will need to select 2 submissions for final scoring.</p>",
      "rawMarkdown": "You will need to select 2 submissions for final scoring.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 264302,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "01/02/2018 18:56:29",
      "content": "<p>You will need to select 2 submissions for final scoring.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "263558": "Someone showed in a [kernel][1] that the number of occurrences of each class is similar on LB. Yet, we can expect the final testing set to have more unknown samples than any other class because:\n\n 1. The final goal of the dataset is probably to have 32 classes with approximately the same number of elements, not just 11+1 classes\n 2. The training set has approximately 60% of elements of the class unknown\n 3. Someone achieved 0.80 on LB while predicting 47% of unknown labels ([here][2])\n\nThen there could be a solution that has 11% of error for unknowns and ~1% for others on LB, that could perform worse than a solution with e.g. 11% of error for yes and ~1% for others (including unknown) because unknown should have many more samples.\nThus, is the best submission the only one used for the final score or are there more submissions taken into account ?\n\n  [1]: https://www.kaggle.com/ironbar/frequency-of-the-labels-in-the-public-test-set\n  [2]: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298",
    "264302": "You will need to select 2 submissions for final scoring."
  },
  "source": "meta"
}