{
  "id": 93111,
  "title": "Labels on Validation Test",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/93111",
  "author_name": "",
  "post_date": "2019-05-23T10:36:06.295744600Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi!</p>\n\n<p>We know that final results will be evaluated on the new test set which is 3 times bigger than the current one.</p>\n\n<p>Wll labels be the same as now? For now test set contains 80 distinct labels and only 74 of them are presented in train curated. I am worried that in new test could be new labels for test.</p>\n\n<p>And many public kernels are focused on the labels that are in the current test set might be wrong on the final valdiation.</p>\n\n<p>What do you think?</p>",
  "messages": [
    {
      "id": "535677",
      "postDate": "05/23/2019 10:36:06",
      "content": "<p>Hi!</p>\n\n<p>We know that final results will be evaluated on the new test set which is 3 times bigger than the current one.</p>\n\n<p>Wll labels be the same as now? For now test set contains 80 distinct labels and only 74 of them are presented in train curated. I am worried that in new test could be new labels for test.</p>\n\n<p>And many public kernels are focused on the labels that are in the current test set might be wrong on the final valdiation.</p>\n\n<p>What do you think?</p>",
      "rawMarkdown": "Hi!\n\nWe know that final results will be evaluated on the new test set which is 3 times bigger than the current one.\n\nWll labels be the same as now? For now test set contains 80 distinct labels and only 74 of them are presented in train curated. I am worried that in new test could be new labels for test.\n\nAnd many public kernels are focused on the labels that are in the current test set might be wrong on the final valdiation.\n\nWhat do you think?",
      "votes": null
    },
    {
      "id": "535693",
      "postDate": "05/23/2019 10:56:01",
      "content": "<blockquote>\n  <p>and only 74 of them are presented in train curated</p>\n</blockquote>\n\n<p>Really?</p>",
      "rawMarkdown": "&gt; and only 74 of them are presented in train curated\n\nReally?",
      "votes": null
    },
    {
      "id": "535817",
      "postDate": "05/23/2019 13:41:53",
      "content": "<p><a href=\"/pavetr\">@pavetr</a> All the labels appear in train curated data. I think you see single-label data only.  6 labels appear only in multi-label data.</p>",
      "rawMarkdown": "pavetr All the labels appear in train curated data. I think you see single-label data only.  6 labels appear only in multi-label data.",
      "votes": null
    },
    {
      "id": "535846",
      "postDate": "05/23/2019 14:10:16",
      "content": "<p><a href=\"/osciiart\">@osciiart</a> ok. But what do you think about private test data? Same labels as in sample submission file?</p>",
      "rawMarkdown": "osciiart ok. But what do you think about private test data? Same labels as in sample submission file?",
      "votes": null
    },
    {
      "id": "535909",
      "postDate": "05/23/2019 15:57:31",
      "content": "<p>All 80 classes are represented in each of our datasets: curated train, noisy train, public test, private test. We do not expect you to make predictions for classes that you have never seen :)</p>",
      "rawMarkdown": "All 80 classes are represented in each of our datasets: curated train, noisy train, public test, private test. We do not expect you to make predictions for classes that you have never seen :)",
      "votes": null
    },
    {
      "id": "535998",
      "postDate": "05/23/2019 19:08:45",
      "content": "<p><a href=\"/plakal\">@plakal</a> I meant if we have 1000+ classes in noisy and 80 classes in the current test set then Could there be new classes (from this 1000+ variants) in new test set which are not present in the current test set? Or new test classes == current test classes?</p>",
      "rawMarkdown": "plakal I meant if we have 1000+ classes in noisy and 80 classes in the current test set then Could there be new classes (from this 1000+ variants) in new test set which are not present in the current test set? Or new test classes == current test classes?",
      "votes": null
    },
    {
      "id": "536011",
      "postDate": "05/23/2019 19:49:57",
      "content": "<p>I don't understand what you mean by 1000+ classes (or variants) in the noisy dataset. All datasets are labeled using the same 80 classes, and no more. All classes are represented in each of the datasets. The public and private test sets were split from the same original test set, with roughly equal representation of all 80 classes. There are no new classes in the private test set.</p>",
      "rawMarkdown": "I don't understand what you mean by 1000+ classes (or variants) in the noisy dataset. All datasets are labeled using the same 80 classes, and no more. All classes are represented in each of the datasets. The public and private test sets were split from the same original test set, with roughly equal representation of all 80 classes. There are no new classes in the private test set.",
      "votes": null
    },
    {
      "id": "536412",
      "postDate": "05/24/2019 12:24:50",
      "content": "<p>No, we think there was a misunderstanding by <a href=\"/pavetr\">@pavetr</a> . The curated train data has  data for <strong>all</strong> the 80 class labels. Take a look at Manoj's comments above for good clarification.</p>\n\n<p><a href=\"/pavetr\">@pavetr</a> could you please edit your initial post to clarify this? (so that no more participants get confused) thanks! :)</p>",
      "rawMarkdown": "No, we think there was a misunderstanding by @pavetr . The curated train data has  data for **all** the 80 class labels. Take a look at Manoj's comments above for good clarification.\n\n@pavetr could you please edit your initial post to clarify this? (so that no more participants get confused) thanks! :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 535693,
      "author_name": "vzaguskin",
      "author_url": "",
      "post_date": "05/23/2019 10:56:01",
      "content": "<blockquote>\n  <p>and only 74 of them are presented in train curated</p>\n</blockquote>\n\n<p>Really?</p>",
      "votes": null,
      "replies": [
        {
          "id": 536412,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "05/24/2019 12:24:50",
          "content": "<p>No, we think there was a misunderstanding by <a href=\"/pavetr\">@pavetr</a> . The curated train data has  data for <strong>all</strong> the 80 class labels. Take a look at Manoj's comments above for good clarification.</p>\n\n<p><a href=\"/pavetr\">@pavetr</a> could you please edit your initial post to clarify this? (so that no more participants get confused) thanks! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 535817,
      "author_name": "osciiart",
      "author_url": "",
      "post_date": "05/23/2019 13:41:53",
      "content": "<p><a href=\"/pavetr\">@pavetr</a> All the labels appear in train curated data. I think you see single-label data only.  6 labels appear only in multi-label data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 535846,
          "author_name": "pavetr",
          "author_url": "",
          "post_date": "05/23/2019 14:10:16",
          "content": "<p><a href=\"/osciiart\">@osciiart</a> ok. But what do you think about private test data? Same labels as in sample submission file?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535909,
          "author_name": "plakal",
          "author_url": "",
          "post_date": "05/23/2019 15:57:31",
          "content": "<p>All 80 classes are represented in each of our datasets: curated train, noisy train, public test, private test. We do not expect you to make predictions for classes that you have never seen :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535998,
          "author_name": "pavetr",
          "author_url": "",
          "post_date": "05/23/2019 19:08:45",
          "content": "<p><a href=\"/plakal\">@plakal</a> I meant if we have 1000+ classes in noisy and 80 classes in the current test set then Could there be new classes (from this 1000+ variants) in new test set which are not present in the current test set? Or new test classes == current test classes?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536011,
          "author_name": "plakal",
          "author_url": "",
          "post_date": "05/23/2019 19:49:57",
          "content": "<p>I don't understand what you mean by 1000+ classes (or variants) in the noisy dataset. All datasets are labeled using the same 80 classes, and no more. All classes are represented in each of the datasets. The public and private test sets were split from the same original test set, with roughly equal representation of all 80 classes. There are no new classes in the private test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "535677": "Hi!\n\nWe know that final results will be evaluated on the new test set which is 3 times bigger than the current one.\n\nWll labels be the same as now? For now test set contains 80 distinct labels and only 74 of them are presented in train curated. I am worried that in new test could be new labels for test.\n\nAnd many public kernels are focused on the labels that are in the current test set might be wrong on the final valdiation.\n\nWhat do you think?",
    "535693": "&gt; and only 74 of them are presented in train curated\n\nReally?",
    "535817": "pavetr All the labels appear in train curated data. I think you see single-label data only.  6 labels appear only in multi-label data.",
    "535846": "osciiart ok. But what do you think about private test data? Same labels as in sample submission file?",
    "535909": "All 80 classes are represented in each of our datasets: curated train, noisy train, public test, private test. We do not expect you to make predictions for classes that you have never seen :)",
    "535998": "plakal I meant if we have 1000+ classes in noisy and 80 classes in the current test set then Could there be new classes (from this 1000+ variants) in new test set which are not present in the current test set? Or new test classes == current test classes?",
    "536011": "I don't understand what you mean by 1000+ classes (or variants) in the noisy dataset. All datasets are labeled using the same 80 classes, and no more. All classes are represented in each of the datasets. The public and private test sets were split from the same original test set, with roughly equal representation of all 80 classes. There are no new classes in the private test set.",
    "536412": "No, we think there was a misunderstanding by @pavetr . The curated train data has  data for **all** the 80 class labels. Take a look at Manoj's comments above for good clarification.\n\n@pavetr could you please edit your initial post to clarify this? (so that no more participants get confused) thanks! :)"
  },
  "source": "meta"
}