{
  "id": 214862,
  "title": "[Q] Clarification what to submit / Test files",
  "url": "/competitions/rfcx-species-audio-detection/discussion/214862",
  "author_name": "",
  "post_date": "2021-01-27T22:50:01.460212900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Could somebody please help me to understand the rule?  </p>\n<p>At the Leaderboard I see the note, </p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 21% of the test data.</p>\n  <p>The final results will be based on the other 79%, so the final standings may be different.</p>\n</blockquote>\n<p>This means that </p>\n<p>1) the public LB is calculated on 21% of the 1992 entires in the<br>\n   submitted csv, which is about 418 rows. The private LB will be<br>\n   calculated on the rest of 1574 entries in submission.csv.</p>\n<p>2) or, there are <em>hidden</em> test files that is about 4 times larger than <br>\n   the test recordings in the 'rfcx-species-audio-detection/test' directory?</p>\n<p>My understanding is 1). 'submission.csv' is the only product we have<br>\nto submit. We do not need to submit the code itself, nor the code <br>\ndoes not need to deal with new recordings that are not in the public<br>\ndirectory. We do not need to do the inference on the kaggle notebook, <br>\neither. (do I understand correct?)</p>\n<p>.. probably these things are clear to experienced kagglers, but <br>\nnot to me, I need to ask… </p>\n<p>thank you so much for your time,<br>\nmegner</p>",
  "messages": [
    {
      "id": "1173458",
      "postDate": "01/27/2021 22:50:01",
      "content": "<p>Could somebody please help me to understand the rule?  </p>\n<p>At the Leaderboard I see the note, </p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 21% of the test data.</p>\n  <p>The final results will be based on the other 79%, so the final standings may be different.</p>\n</blockquote>\n<p>This means that </p>\n<p>1) the public LB is calculated on 21% of the 1992 entires in the<br>\n   submitted csv, which is about 418 rows. The private LB will be<br>\n   calculated on the rest of 1574 entries in submission.csv.</p>\n<p>2) or, there are <em>hidden</em> test files that is about 4 times larger than <br>\n   the test recordings in the 'rfcx-species-audio-detection/test' directory?</p>\n<p>My understanding is 1). 'submission.csv' is the only product we have<br>\nto submit. We do not need to submit the code itself, nor the code <br>\ndoes not need to deal with new recordings that are not in the public<br>\ndirectory. We do not need to do the inference on the kaggle notebook, <br>\neither. (do I understand correct?)</p>\n<p>.. probably these things are clear to experienced kagglers, but <br>\nnot to me, I need to ask… </p>\n<p>thank you so much for your time,<br>\nmegner</p>",
      "rawMarkdown": "Could somebody please help me to understand the rule?  \n\nAt the Leaderboard I see the note, \n\n> This leaderboard is calculated with approximately 21% of the test data.\n\n> The final results will be based on the other 79%, so the final standings may be different.\n\nThis means that \n\n1) the public LB is calculated on 21% of the 1992 entires in the\n   submitted csv, which is about 418 rows. The private LB will be\n   calculated on the rest of 1574 entries in submission.csv.\n\n2) or, there are _hidden_ test files that is about 4 times larger than \n   the test recordings in the 'rfcx-species-audio-detection/test' directory?\n\nMy understanding is 1). 'submission.csv' is the only product we have\nto submit. We do not need to submit the code itself, nor the code \ndoes not need to deal with new recordings that are not in the public\ndirectory. We do not need to do the inference on the kaggle notebook, \neither. (do I understand correct?)\n\n.. probably these things are clear to experienced kagglers, but \nnot to me, I need to ask... \n\nthank you so much for your time,\nmegner",
      "votes": null
    },
    {
      "id": "1174471",
      "postDate": "01/28/2021 13:49:21",
      "content": "<p>Yeah, I believe it's 1).<br>\nYour submitted csv is sufficient to compute score for both public and private set of data. <br>\nBut before the competition ends, only the public part of score will be showed.</p>",
      "rawMarkdown": "Yeah, I believe it's 1).\nYour submitted csv is sufficient to compute score for both public and private set of data. \nBut before the competition ends, only the public part of score will be showed.",
      "votes": null
    },
    {
      "id": "1174804",
      "postDate": "01/28/2021 18:01:53",
      "content": "<p>Dear Buffalo, </p>\n<p>thank you so much for your response. <br>\nI was really worried, in case the note above means 2), all <br>\nour efforts will be useless. Now it is clear to me. Thank you so much, </p>\n<p>megner</p>",
      "rawMarkdown": "Dear Buffalo, \n\nthank you so much for your response. \nI was really worried, in case the note above means 2), all \nour efforts will be useless. Now it is clear to me. Thank you so much, \n\nmegner",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1174471,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "01/28/2021 13:49:21",
      "content": "<p>Yeah, I believe it's 1).<br>\nYour submitted csv is sufficient to compute score for both public and private set of data. <br>\nBut before the competition ends, only the public part of score will be showed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1174804,
      "author_name": "megner",
      "author_url": "",
      "post_date": "01/28/2021 18:01:53",
      "content": "<p>Dear Buffalo, </p>\n<p>thank you so much for your response. <br>\nI was really worried, in case the note above means 2), all <br>\nour efforts will be useless. Now it is clear to me. Thank you so much, </p>\n<p>megner</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1173458": "Could somebody please help me to understand the rule?  \n\nAt the Leaderboard I see the note, \n\n> This leaderboard is calculated with approximately 21% of the test data.\n\n> The final results will be based on the other 79%, so the final standings may be different.\n\nThis means that \n\n1) the public LB is calculated on 21% of the 1992 entires in the\n   submitted csv, which is about 418 rows. The private LB will be\n   calculated on the rest of 1574 entries in submission.csv.\n\n2) or, there are _hidden_ test files that is about 4 times larger than \n   the test recordings in the 'rfcx-species-audio-detection/test' directory?\n\nMy understanding is 1). 'submission.csv' is the only product we have\nto submit. We do not need to submit the code itself, nor the code \ndoes not need to deal with new recordings that are not in the public\ndirectory. We do not need to do the inference on the kaggle notebook, \neither. (do I understand correct?)\n\n.. probably these things are clear to experienced kagglers, but \nnot to me, I need to ask... \n\nthank you so much for your time,\nmegner",
    "1174471": "Yeah, I believe it's 1).\nYour submitted csv is sufficient to compute score for both public and private set of data. \nBut before the competition ends, only the public part of score will be showed.",
    "1174804": "Dear Buffalo, \n\nthank you so much for your response. \nI was really worried, in case the note above means 2), all \nour efforts will be useless. Now it is clear to me. Thank you so much, \n\nmegner"
  },
  "source": "meta"
}