{
  "id": 171004,
  "title": "Squirrels, unknown birds, and scoring",
  "url": "/competitions/birdsong-recognition/discussion/171004",
  "author_name": "",
  "post_date": "2020-07-30T01:13:10.204711200Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Please pardon me if these are silly questions, or have already been answered.</p>\n\n<p>1) I see in file example_test_audio_summary.csv that \"squirrel\" is included as one of the calls. But \"squirrel\" does not appear in any of the training samples. Does this mean that \"squirrel\" might need to be identified and matched when computing the score for the test set?</p>\n\n<p>2) More generally, suppose a test sample contains a call that does not correspond to one of the 264 training species. The overall score will be penalized if an unknown call is incorrectly classified as one of the 264, right?</p>\n\n<p>3) Therefore if a test set call can be classified as NOT one of the 264, it is best to make no identification (just ignore it)?</p>\n\n<p>Thanks for clarifying.</p>",
  "messages": [
    {
      "id": "951182",
      "postDate": "07/30/2020 01:13:10",
      "content": "<p>Please pardon me if these are silly questions, or have already been answered.</p>\n\n<p>1) I see in file example_test_audio_summary.csv that \"squirrel\" is included as one of the calls. But \"squirrel\" does not appear in any of the training samples. Does this mean that \"squirrel\" might need to be identified and matched when computing the score for the test set?</p>\n\n<p>2) More generally, suppose a test sample contains a call that does not correspond to one of the 264 training species. The overall score will be penalized if an unknown call is incorrectly classified as one of the 264, right?</p>\n\n<p>3) Therefore if a test set call can be classified as NOT one of the 264, it is best to make no identification (just ignore it)?</p>\n\n<p>Thanks for clarifying.</p>",
      "rawMarkdown": "Please pardon me if these are silly questions, or have already been answered.\n\n1) I see in file example_test_audio_summary.csv that \"squirrel\" is included as one of the calls. But \"squirrel\" does not appear in any of the training samples. Does this mean that \"squirrel\" might need to be identified and matched when computing the score for the test set?\n\n2) More generally, suppose a test sample contains a call that does not correspond to one of the 264 training species. The overall score will be penalized if an unknown call is incorrectly classified as one of the 264, right?\n\n3) Therefore if a test set call can be classified as NOT one of the 264, it is best to make no identification (just ignore it)?\n\nThanks for clarifying.",
      "votes": null
    },
    {
      "id": "951607",
      "postDate": "07/30/2020 08:58:39",
      "content": "<p>If you check the sample submission file - use nocall when the prediction is not one of the 264. </p>",
      "rawMarkdown": "If you check the sample submission file - use nocall when the prediction is not one of the 264.",
      "votes": null
    },
    {
      "id": "952567",
      "postDate": "07/31/2020 03:50:55",
      "content": "<p>this is actually common in audio competitions.</p>\n\n<ul>\n<li>you are given N class as train data</li>\n<li>you are asked to annotate test clips using N+1 class (the 1 is for background=no events)</li>\n<li>there is additional class in the test clip,  i.e. unknown class  = other events</li>\n</ul>\n\n<p>so in reality, there are N+2 class! And you have to collect your own  data for the two missing class!</p>\n\n<p>In the previous tensorflow spotting keyword kaggle challenge, we exploit the fact that test data is given. Then we can detect non-event background noise interval in the test (pseudo labels). Then we mix this test noise with the train data (perfect simulation of test environment noise)</p>\n\n<p>but we cannot do this here.</p>",
      "rawMarkdown": "this is actually common in audio competitions.\n\n- you are given N class as train data\n- you are asked to annotate test clips using N+1 class (the 1 is for background=no events)\n- there is additional class in the test clip,  i.e. unknown class  = other events\n\nso in reality, there are N+2 class! And you have to collect your own  data for the two missing class!\n\nIn the previous tensorflow spotting keyword kaggle challenge, we exploit the fact that test data is given. Then we can detect non-event background noise interval in the test (pseudo labels). Then we mix this test noise with the train data (perfect simulation of test environment noise)\n\nbut we cannot do this here.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 951607,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "07/30/2020 08:58:39",
      "content": "<p>If you check the sample submission file - use nocall when the prediction is not one of the 264. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 952567,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/31/2020 03:50:55",
      "content": "<p>this is actually common in audio competitions.</p>\n\n<ul>\n<li>you are given N class as train data</li>\n<li>you are asked to annotate test clips using N+1 class (the 1 is for background=no events)</li>\n<li>there is additional class in the test clip,  i.e. unknown class  = other events</li>\n</ul>\n\n<p>so in reality, there are N+2 class! And you have to collect your own  data for the two missing class!</p>\n\n<p>In the previous tensorflow spotting keyword kaggle challenge, we exploit the fact that test data is given. Then we can detect non-event background noise interval in the test (pseudo labels). Then we mix this test noise with the train data (perfect simulation of test environment noise)</p>\n\n<p>but we cannot do this here.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "951182": "Please pardon me if these are silly questions, or have already been answered.\n\n1) I see in file example_test_audio_summary.csv that \"squirrel\" is included as one of the calls. But \"squirrel\" does not appear in any of the training samples. Does this mean that \"squirrel\" might need to be identified and matched when computing the score for the test set?\n\n2) More generally, suppose a test sample contains a call that does not correspond to one of the 264 training species. The overall score will be penalized if an unknown call is incorrectly classified as one of the 264, right?\n\n3) Therefore if a test set call can be classified as NOT one of the 264, it is best to make no identification (just ignore it)?\n\nThanks for clarifying.",
    "951607": "If you check the sample submission file - use nocall when the prediction is not one of the 264.",
    "952567": "this is actually common in audio competitions.\n\n- you are given N class as train data\n- you are asked to annotate test clips using N+1 class (the 1 is for background=no events)\n- there is additional class in the test clip,  i.e. unknown class  = other events\n\nso in reality, there are N+2 class! And you have to collect your own  data for the two missing class!\n\nIn the previous tensorflow spotting keyword kaggle challenge, we exploit the fact that test data is given. Then we can detect non-event background noise interval in the test (pseudo labels). Then we mix this test noise with the train data (perfect simulation of test environment noise)\n\nbut we cannot do this here."
  },
  "source": "meta"
}