{
  "id": 158987,
  "title": "Confusion about test set",
  "url": "/competitions/birdsong-recognition/discussion/158987",
  "author_name": "Hidehisa Arai",
  "post_date": "2020-06-16T03:18:46.880000",
  "votes": 69,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I'm still not clear about the structure of test data.</p>\n\n<p>In data description, they say</p>\n\n<p>&gt; <strong>test_audio</strong> The hidden test set audio consists of approximately 150 recordings in mp3 format, each roughly 10 minutes long. The recordings were taken at three separate remote locations. Sites 1 and 2 were labeled in 5 second increments and need matching predictions, but due to the time consuming nature of the labeling process the site 3 files are only labeled at the file level. Accordingly, site 3 has relatively few rows in the test set and needs lower time resolution predictions.</p>\n\n<p>so we cannot see test set at all. Instead we have <code>example_test_audio</code> but we don't need to do prediction on audio files in this directory, am I right about this?</p>\n\n<p>What we need to do in the inference phase is</p>\n\n<ol>\n<li>Read <code>test.csv</code></li>\n<li>Repeat 3 - 5 for each row in <code>test.csv</code></li>\n<li>Open test audio file which is contained in <code>/kaggle/input/birdsong-recognition/test_audio</code> and named like <code>audio_id.mp3</code> .</li>\n<li>If <code>site</code> is <code>site_3</code> use the whole clip, otherwise cut 5 seconds short clip out of the long audio clip which we loaded in step3 that ends in the second specified in <code>seconds</code> column.</li>\n<li>Perform prediction on the short clip which we prepared in step 4. Note that each clip may have multiple labels.</li>\n</ol>\n\n<p>↑Am I right about this?</p>",
  "messages": [
    {
      "id": 887963,
      "postDate": "2020-06-16T03:18:46.880Z",
      "content": "<p>I'm still not clear about the structure of test data.</p>\n\n<p>In data description, they say</p>\n\n<p>&gt; <strong>test_audio</strong> The hidden test set audio consists of approximately 150 recordings in mp3 format, each roughly 10 minutes long. The recordings were taken at three separate remote locations. Sites 1 and 2 were labeled in 5 second increments and need matching predictions, but due to the time consuming nature of the labeling process the site 3 files are only labeled at the file level. Accordingly, site 3 has relatively few rows in the test set and needs lower time resolution predictions.</p>\n\n<p>so we cannot see test set at all. Instead we have <code>example_test_audio</code> but we don't need to do prediction on audio files in this directory, am I right about this?</p>\n\n<p>What we need to do in the inference phase is</p>\n\n<ol>\n<li>Read <code>test.csv</code></li>\n<li>Repeat 3 - 5 for each row in <code>test.csv</code></li>\n<li>Open test audio file which is contained in <code>/kaggle/input/birdsong-recognition/test_audio</code> and named like <code>audio_id.mp3</code> .</li>\n<li>If <code>site</code> is <code>site_3</code> use the whole clip, otherwise cut 5 seconds short clip out of the long audio clip which we loaded in step3 that ends in the second specified in <code>seconds</code> column.</li>\n<li>Perform prediction on the short clip which we prepared in step 4. Note that each clip may have multiple labels.</li>\n</ol>\n\n<p>↑Am I right about this?</p>",
      "rawMarkdown": "I'm still not clear about the structure of test data.\n\nIn data description, they say\n\n&gt; **test_audio** The hidden test set audio consists of approximately 150 recordings in mp3 format, each roughly 10 minutes long. The recordings were taken at three separate remote locations. Sites 1 and 2 were labeled in 5 second increments and need matching predictions, but due to the time consuming nature of the labeling process the site 3 files are only labeled at the file level. Accordingly, site 3 has relatively few rows in the test set and needs lower time resolution predictions.\n\nso we cannot see test set at all. Instead we have `example_test_audio` but we don't need to do prediction on audio files in this directory, am I right about this?\n\nWhat we need to do in the inference phase is\n\n1. Read `test.csv`\n2. Repeat 3 - 5 for each row in `test.csv`\n3. Open test audio file which is contained in `/kaggle/input/birdsong-recognition/test_audio` and named like `audio_id.mp3` .\n4. If `site` is `site_3` use the whole clip, otherwise cut 5 seconds short clip out of the long audio clip which we loaded in step3 that ends in the second specified in `seconds` column.\n5. Perform prediction on the short clip which we prepared in step 4. Note that each clip may have multiple labels.\n\n↑Am I right about this?",
      "votes": 68
    },
    {
      "id": 890937,
      "postDate": "2020-06-17T19:19:23.350Z",
      "content": "<p>For others having trouble, it looks like <a href=\"/cwthompson\">@cwthompson</a> put together a working skeletal submission example here, which you can refer to:</p>\n\n<p><a href=\"https://www.kaggle.com/cwthompson/birdsong-making-a-prediction\">https://www.kaggle.com/cwthompson/birdsong-making-a-prediction</a></p>",
      "rawMarkdown": "For others having trouble, it looks like @cwthompson put together a working skeletal submission example here, which you can refer to:\n\nhttps://www.kaggle.com/cwthompson/birdsong-making-a-prediction",
      "votes": 11,
      "replies": [
        {
          "id": 948368,
          "postDate": "2020-07-27T21:25:22.310Z",
          "content": "<p>Unless it is just me, but this skeleton sample - Errors out because it can't find the files is is referencing, throws an Exception (with no print output) and due to the exception is set to output the default submission file, rather than a submission file with a bunch of random predictions.  So this is (semi) broken code????</p>\n\n<p>[UPDATE] - The code breaks pre-submit, but works after Submit is done it seems.  I updated with a version that works pre-submit also with <a href=\"/shonenkov\">@shonenkov</a> 's test data, hopefully it helps people.  (I was confused by method, this is my making sense of it.)</p>",
          "rawMarkdown": "Unless it is just me, but this skeleton sample - Errors out because it can't find the files is is referencing, throws an Exception (with no print output) and due to the exception is set to output the default submission file, rather than a submission file with a bunch of random predictions.  So this is (semi) broken code????\n \n[UPDATE] - The code breaks pre-submit, but works after Submit is done it seems.  I updated with a version that works pre-submit also with @shonenkov 's test data, hopefully it helps people.  (I was confused by method, this is my making sense of it.)",
          "votes": 1
        }
      ]
    },
    {
      "id": 890473,
      "postDate": "2020-06-17T14:20:43.037Z",
      "content": "<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> your description is accurate.</p>",
      "rawMarkdown": "@hidehisaarai1213 your description is accurate.",
      "votes": 6,
      "replies": [
        {
          "id": 890485,
          "postDate": "2020-06-17T14:27:03.657Z",
          "content": "<p>Thank you for your confirmation !</p>",
          "rawMarkdown": "Thank you for your confirmation !",
          "votes": 1
        }
      ]
    },
    {
      "id": 888738,
      "postDate": "2020-06-16T14:38:21.400Z",
      "content": "<p>What I gather from looking at the dataset is:</p>\n\n<ol>\n<li><p>0/150 audio files from the test data are being shared, so inference will be done only during kernel re-run on submission of <code>sample_submission.csv</code></p></li>\n<li><p>For <strong>validation</strong>, we can use audio file: <code>example_test_audio</code> and rows <code>filename_seconds</code> and <code>birds</code> from <code>example_test_audio_summary.csv</code> , but this is <strong>not</strong> part of the test set.</p></li>\n<li><p>For <em>site1</em> and <em>site 2</em>, predictions of <code>birds</code> is done for every <strong>5 second interval</strong> of the audio-file.\nFor site3, prediction of <code>birds</code> is for entire audio-file</p></li>\n<li><p>For each row, there can be multiple <code>birds</code></p></li>\n</ol>",
      "rawMarkdown": "What I gather from looking at the dataset is:\n\n1. 0/150 audio files from the test data are being shared, so inference will be done only during kernel re-run on submission of `sample_submission.csv`\n\n2. For **validation**, we can use audio file: `example_test_audio` and rows `filename_seconds` and `birds` from `example_test_audio_summary.csv` , but this is **not** part of the test set.\n\n3. For *site1* and *site 2*, predictions of `birds` is done for every **5 second interval** of the audio-file.\n    For site3, prediction of `birds` is for entire audio-file\n\n4. For each row, there can be multiple `birds`",
      "votes": 4,
      "replies": [
        {
          "id": 888790,
          "postDate": "2020-06-16T15:04:51.710Z",
          "content": "<p>Thanks for the comment !\nI totally agree with you but still wants clarification from the host.</p>",
          "rawMarkdown": "Thanks for the comment !\nI totally agree with you but still wants clarification from the host.",
          "votes": 2
        }
      ]
    },
    {
      "id": 888513,
      "postDate": "2020-06-16T12:09:58.340Z",
      "content": "<p>I don't see a test audio folder.  I see an example test audio that goes with the example test metadata and example test summary.</p>\n\n<p>Plus test.csv is only 3 rows.  </p>",
      "rawMarkdown": "I don't see a test audio folder.  I see an example test audio that goes with the example test metadata and example test summary.\n\nPlus test.csv is only 3 rows.  ",
      "votes": 1,
      "replies": [
        {
          "id": 888530,
          "postDate": "2020-06-16T12:23:22.833Z",
          "content": "<blockquote>\n  <p>I don't see a test audio folder. I see an example test audio that goes with the example test metadata and example test summary.</p>\n</blockquote>\n\n<p>Yes, that's why I am confused. What I would like to know is whether it is safe to assume there is a <code>test_audio</code> directory at kernel re-running.</p>\n\n<p><a href=\"/stefankahl\">@stefankahl</a> </p>",
          "rawMarkdown": "&gt; I don't see a test audio folder. I see an example test audio that goes with the example test metadata and example test summary.\n\nYes, that's why I am confused. What I would like to know is whether it is safe to assume there is a `test_audio` directory at kernel re-running.\n\n@stefankahl ",
          "votes": 3
        },
        {
          "id": 892370,
          "postDate": "2020-06-18T20:42:52.730Z",
          "content": "<p>What about <code>test_audio</code>? You forgot this folder? Why are you ignoring us? <a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> </p>",
          "rawMarkdown": "What about `test_audio`? You forgot this folder? Why are you ignoring us? @stefankahl @sohier "
        },
        {
          "id": 892407,
          "postDate": "2020-06-18T21:15:18.730Z",
          "content": "<p>Could you add in test.csv cases with all sites? <a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> it really needs for check. Your tasks for site1(2) and site3 are different, I suppose competitors will prepare different models, need check. Also, why 2 submissions per day? It is not enough even for debugging your not clearly understanding structure of files, not to mention about own bugs... </p>",
          "rawMarkdown": "Could you add in test.csv cases with all sites? @stefankahl @sohier it really needs for check. Your tasks for site1(2) and site3 are different, I suppose competitors will prepare different models, need check. Also, why 2 submissions per day? It is not enough even for debugging your not clearly understanding structure of files, not to mention about own bugs... "
        },
        {
          "id": 892415,
          "postDate": "2020-06-18T21:21:03.630Z",
          "content": "<p>Hi, the test_audio folder is visible to the submitted notebook.\n<a href=\"/cwthompson\">@cwthompson</a> put together a nice example submission notebook here, which you can follow:\n<a href=\"https://www.kaggle.com/cwthompson/birdsong-making-a-prediction\">https://www.kaggle.com/cwthompson/birdsong-making-a-prediction</a></p>\n\n<p>Alex: It's up to the competitors to decide how to handle the third site. You could take a short-segment model and aggregate up to the 10 minute window, or use an entirely different model; your call.</p>",
          "rawMarkdown": "Hi, the test_audio folder is visible to the submitted notebook.\n@cwthompson put together a nice example submission notebook here, which you can follow:\nhttps://www.kaggle.com/cwthompson/birdsong-making-a-prediction\n\nAlex: It's up to the competitors to decide how to handle the third site. You could take a short-segment model and aggregate up to the 10 minute window, or use an entirely different model; your call."
        },
        {
          "id": 892424,
          "postDate": "2020-06-18T21:33:03.373Z",
          "content": "<p>Thank you for answer! if you look in last code competitions with hidden test you can see this folder visible in check phase too. It is right way, because it checks bugs with base structure of files without button submission. Also your approach breaks display public notebooks without this folder. </p>\n\n<p>You should add this folder and change content after button submission</p>\n\n<p>And you didn’t answer about 2 submissions per day.</p>\n\n<p><a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> <a href=\"/tomdenton\">@tomdenton</a> </p>",
          "rawMarkdown": "Thank you for answer! if you look in last code competitions with hidden test you can see this folder visible in check phase too. It is right way, because it checks bugs with base structure of files without button submission. Also your approach breaks display public notebooks without this folder. \n\nYou should add this folder and change content after button submission\n\nAnd you didn’t answer about 2 submissions per day.\n\n @stefankahl @sohier @tomdenton ",
          "votes": 11
        },
        {
          "id": 892582,
          "postDate": "2020-06-19T02:55:33.690Z",
          "content": "<p>I agree with <a href=\"/shonenkov\">@shonenkov</a> , it would be much easier for us to create a submission if we have <code>test_audio</code> folder not only in the submission phase but also <em>before</em> committing (I mean coding phase). I know you do not want to make it visible to avoid competitor trying nonsense stuffs, but as <a href=\"/shonenkov\">@shonenkov</a> says you can change the contents in <code>test_audio</code> between coding phase and submission phase.</p>\n\n<p>Right now we have to do without <code>test_audio</code> folder and that makes us unconfident about whether our commit works in submission phase.</p>",
          "rawMarkdown": "I agree with @shonenkov , it would be much easier for us to create a submission if we have `test_audio` folder not only in the submission phase but also *before* committing (I mean coding phase). I know you do not want to make it visible to avoid competitor trying nonsense stuffs, but as @shonenkov says you can change the contents in `test_audio` between coding phase and submission phase.\n\nRight now we have to do without `test_audio` folder and that makes us unconfident about whether our commit works in submission phase.",
          "votes": 6
        },
        {
          "id": 893460,
          "postDate": "2020-06-19T16:13:02.230Z",
          "content": "<p>The existing structure is sufficient to write reliable submission code; there are some examples of testing for the existence of the test_audio folder in some other competitor notebooks, for example. If you're very concerned, and you're doing any development outside the Kaggle editors, you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.) We're hesitant to change things after the competition has started unless really necessary.</p>\n\n<p>Alex: We limit to two submissions per day to avoid submission spamming to probe/overfit the test datasets. </p>",
          "rawMarkdown": "The existing structure is sufficient to write reliable submission code; there are some examples of testing for the existence of the test_audio folder in some other competitor notebooks, for example. If you're very concerned, and you're doing any development outside the Kaggle editors, you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.) We're hesitant to change things after the competition has started unless really necessary.\n\nAlex: We limit to two submissions per day to avoid submission spamming to probe/overfit the test datasets. ",
          "votes": -19
        },
        {
          "id": 893484,
          "postDate": "2020-06-19T16:39:03.040Z",
          "content": "<p><a href=\"/tomdenton\">@tomdenton</a> are you sure?</p>\n\n<p>this topics from Novice --&gt; GM about this problem.</p>\n\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159992\">https://www.kaggle.com/c/birdsong-recognition/discussion/159992</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160007\">https://www.kaggle.com/c/birdsong-recognition/discussion/160007</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159173\">https://www.kaggle.com/c/birdsong-recognition/discussion/159173</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159123\">https://www.kaggle.com/c/birdsong-recognition/discussion/159123</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158951\">https://www.kaggle.com/c/birdsong-recognition/discussion/158951</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159507\">https://www.kaggle.com/c/birdsong-recognition/discussion/159507</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160111\">https://www.kaggle.com/c/birdsong-recognition/discussion/160111</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160110\">https://www.kaggle.com/c/birdsong-recognition/discussion/160110</a></p>\n\n<p>It is disrespectful of you to ignore obvious problems. I spent last 6 submissions (3 days!?) for solving it for all competitors (I did your work):</p>\n\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159993\">https://www.kaggle.com/c/birdsong-recognition/discussion/159993</a></p>\n\n<p>But I wanted only to share my baseline for your competition...</p>\n\n<p>P.S. Dont forget to read <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#886135\">topic about DFDC</a> where hosts ignored obvious problems during competition. </p>",
          "rawMarkdown": "@tomdenton are you sure?\n\nthis topics from Novice --&gt; GM about this problem.\n\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159992\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/160007\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159173\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159123\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/158951\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159507\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/160111\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/160110\n\nIt is disrespectful of you to ignore obvious problems. I spent last 6 submissions (3 days!?) for solving it for all competitors (I did your work):\n\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/159993\n\nBut I wanted only to share my baseline for your competition...\n\nP.S. Dont forget to read [topic about DFDC](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#886135) where hosts ignored obvious problems during competition. ",
          "votes": 19
        },
        {
          "id": 894889,
          "postDate": "2020-06-20T21:33:30.313Z",
          "content": "<p>Hi <a href=\"/tomdenton\">@tomdenton</a>,\nYou said 'The existing structure is sufficient ...', see above for full post.\nTechnically your statement is probably correct, but the data structure is a bit of a headache.\nMany of us already have to deal with such complicated data structures during day-time at out jobs. \nMy fear is that more that one will end-up ignoring this competition in favor of others where they can directly dive into the fun stuff like ML and DNNs.</p>",
          "rawMarkdown": "Hi @tomdenton,\nYou said 'The existing structure is sufficient ...', see above for full post.\nTechnically your statement is probably correct, but the data structure is a bit of a headache.\nMany of us already have to deal with such complicated data structures during day-time at out jobs. \nMy fear is that more that one will end-up ignoring this competition in favor of others where they can directly dive into the fun stuff like ML and DNNs.",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 894900,
          "postDate": "2020-06-20T21:58:28.127Z",
          "content": "<p>&gt; The existing structure is sufficient to write reliable submission code; there are some examples of testing for the existence of the test_audio folder in some other competitor notebooks, for example. If you're very concerned, and you're doing any development outside the Kaggle editors, you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.)</p>\n\n<p>I agree it is <em>possible</em> to write unit tests and checks (like how <a href=\"/shonenkov\">@shonenkov</a> has graciously shared one such method for it) but it is not the standard way and not at all beginner-friendly. Especially considering Kaggle already has a feature and system that is designed to solve this problem as described by <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> </p>\n\n<p>&gt; We're hesitant to change things after the competition has started unless really necessary.</p>\n\n<p>The 3-month long competition began less than a week back and as of this point, no one has made any submission that beats the sample submission. If this is resolved / fixed soon (even if it requires changes or a restart), I can bet Kagglers will be much happier without any complaints.</p>",
          "rawMarkdown": "&gt; The existing structure is sufficient to write reliable submission code; there are some examples of testing for the existence of the test_audio folder in some other competitor notebooks, for example. If you're very concerned, and you're doing any development outside the Kaggle editors, you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.)\n\nI agree it is *possible* to write unit tests and checks (like how @shonenkov has graciously shared one such method for it) but it is not the standard way and not at all beginner-friendly. Especially considering Kaggle already has a feature and system that is designed to solve this problem as described by @hidehisaarai1213 \n\n&gt; We're hesitant to change things after the competition has started unless really necessary.\n\nThe 3-month long competition began less than a week back and as of this point, no one has made any submission that beats the sample submission. If this is resolved / fixed soon (even if it requires changes or a restart), I can bet Kagglers will be much happier without any complaints.",
          "votes": 8
        },
        {
          "id": 948306,
          "postDate": "2020-07-27T19:39:20.463Z",
          "content": "<p><a href=\"/shonenkov\">@shonenkov</a> - My goodness the Deep Fake scoring / data provided was atrocious!</p>",
          "rawMarkdown": "@shonenkov - My goodness the Deep Fake scoring / data provided was atrocious!",
          "votes": 1
        },
        {
          "id": 995480,
          "postDate": "2020-09-02T13:35:51.723Z",
          "content": "<p><a href=\"https://www.kaggle.com/shonenkov\" target=\"_blank\">@shonenkov</a> I am reading all your posts as it seems I am following your path with failed submissions that really look fine.  Thanks for the fake test data.</p>",
          "rawMarkdown": "@shonenkov I am reading all your posts as it seems I am following your path with failed submissions that really look fine.  Thanks for the fake test data.",
          "votes": 2
        },
        {
          "id": 995760,
          "postDate": "2020-09-02T18:41:28.567Z",
          "content": "<blockquote>\n  <p>you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.) </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> are you aware that what you call a unit test is a submission for us?  We have only two of them per day.  </p>\n<p>I am only beginning this journey but when I see a seasoned Kaggler GM being unable to get a successful submission in few days this week, then I believe there is a real issue with the setup.  I asked a very simple question in another thread: can one of you check that reading all actual test files with librosa gives files with sampling rates of 32000 and duration 10 minutes?</p>\n<p>Using the time of hundreds of Kagglers to debug by trial and error something that could be just documented properly baffles me to be honest.  Main reason is that while we are doing this we are not building better models for you.</p>",
          "rawMarkdown": ">  you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.) \n\n@tomdenton are you aware that what you call a unit test is a submission for us?  We have only two of them per day.  \n\nI am only beginning this journey but when I see a seasoned Kaggler GM being unable to get a successful submission in few days this week, then I believe there is a real issue with the setup.  I asked a very simple question in another thread: can one of you check that reading all actual test files with librosa gives files with sampling rates of 32000 and duration 10 minutes?\n\nUsing the time of hundreds of Kagglers to debug by trial and error something that could be just documented properly baffles me to be honest.  Main reason is that while we are doing this we are not building better models for you.",
          "votes": 3
        }
      ]
    },
    {
      "id": 946035,
      "postDate": "2020-07-26T09:55:04.557Z",
      "content": "<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> sorry for such a naive question but there is something that has been bugging me and I will be really happy if you answer this . All the metadata columns which we have in train are not available during test time right? and we have only and only audio data available ? \nThis also means everything other than the audio becomes useless during train time right?</p>",
      "rawMarkdown": "@hidehisaarai1213 sorry for such a naive question but there is something that has been bugging me and I will be really happy if you answer this . All the metadata columns which we have in train are not available during test time right? and we have only and only audio data available ? \nThis also means everything other than the audio becomes useless during train time right?",
      "votes": 2,
      "replies": [
        {
          "id": 946066,
          "postDate": "2020-07-26T10:25:54.830Z",
          "content": "<p>Yup. Metadata is not for training the model.</p>",
          "rawMarkdown": "Yup. Metadata is not for training the model.",
          "votes": 2
        },
        {
          "id": 946079,
          "postDate": "2020-07-26T10:41:52.320Z",
          "content": "<blockquote>\n  <p>This also means everything other than the audio becomes useless during train time right?</p>\n</blockquote>\n\n<p>Basically yes, but if you get luck in predicting some of the columns, they may be useful.</p>",
          "rawMarkdown": "&gt; This also means everything other than the audio becomes useless during train time right?\n\nBasically yes, but if you get luck in predicting some of the columns, they may be useful.",
          "votes": 4
        },
        {
          "id": 946086,
          "postDate": "2020-07-26T10:49:05.163Z",
          "content": "<p>Thanks for your answers Hidehisa and Dhruv</p>",
          "rawMarkdown": "Thanks for your answers Hidehisa and Dhruv",
          "votes": 1
        }
      ]
    },
    {
      "id": 890497,
      "postDate": "2020-06-17T14:32:36.913Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 888003,
      "postDate": "2020-06-16T04:03:38.637Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 890937,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2020-06-17T19:19:23.350000",
      "content": "<p>For others having trouble, it looks like <a href=\"/cwthompson\">@cwthompson</a> put together a working skeletal submission example here, which you can refer to:</p>\n\n<p><a href=\"https://www.kaggle.com/cwthompson/birdsong-making-a-prediction\">https://www.kaggle.com/cwthompson/birdsong-making-a-prediction</a></p>",
      "votes": 11,
      "replies": [
        {
          "id": 948368,
          "author_name": "Mark Eckdahl",
          "author_url": "",
          "post_date": "2020-07-27T21:25:22.310000",
          "content": "<p>Unless it is just me, but this skeleton sample - Errors out because it can't find the files is is referencing, throws an Exception (with no print output) and due to the exception is set to output the default submission file, rather than a submission file with a bunch of random predictions.  So this is (semi) broken code????</p>\n\n<p>[UPDATE] - The code breaks pre-submit, but works after Submit is done it seems.  I updated with a version that works pre-submit also with <a href=\"/shonenkov\">@shonenkov</a> 's test data, hopefully it helps people.  (I was confused by method, this is my making sense of it.)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 890473,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2020-06-17T14:20:43.037000",
      "content": "<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> your description is accurate.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 890485,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-06-17T14:27:03.657000",
          "content": "<p>Thank you for your confirmation !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 888738,
      "author_name": "Dhruv Naik",
      "author_url": "",
      "post_date": "2020-06-16T14:38:21.400000",
      "content": "<p>What I gather from looking at the dataset is:</p>\n\n<ol>\n<li><p>0/150 audio files from the test data are being shared, so inference will be done only during kernel re-run on submission of <code>sample_submission.csv</code></p></li>\n<li><p>For <strong>validation</strong>, we can use audio file: <code>example_test_audio</code> and rows <code>filename_seconds</code> and <code>birds</code> from <code>example_test_audio_summary.csv</code> , but this is <strong>not</strong> part of the test set.</p></li>\n<li><p>For <em>site1</em> and <em>site 2</em>, predictions of <code>birds</code> is done for every <strong>5 second interval</strong> of the audio-file.\nFor site3, prediction of <code>birds</code> is for entire audio-file</p></li>\n<li><p>For each row, there can be multiple <code>birds</code></p></li>\n</ol>",
      "votes": 4,
      "replies": [
        {
          "id": 888790,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-06-16T15:04:51.710000",
          "content": "<p>Thanks for the comment !\nI totally agree with you but still wants clarification from the host.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 888513,
      "author_name": "Eric Freeman",
      "author_url": "",
      "post_date": "2020-06-16T12:09:58.340000",
      "content": "<p>I don't see a test audio folder.  I see an example test audio that goes with the example test metadata and example test summary.</p>\n\n<p>Plus test.csv is only 3 rows.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 888530,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-06-16T12:23:22.833000",
          "content": "<blockquote>\n  <p>I don't see a test audio folder. I see an example test audio that goes with the example test metadata and example test summary.</p>\n</blockquote>\n\n<p>Yes, that's why I am confused. What I would like to know is whether it is safe to assume there is a <code>test_audio</code> directory at kernel re-running.</p>\n\n<p><a href=\"/stefankahl\">@stefankahl</a> </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 892370,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-18T20:42:52.730000",
          "content": "<p>What about <code>test_audio</code>? You forgot this folder? Why are you ignoring us? <a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 892407,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-18T21:15:18.730000",
          "content": "<p>Could you add in test.csv cases with all sites? <a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> it really needs for check. Your tasks for site1(2) and site3 are different, I suppose competitors will prepare different models, need check. Also, why 2 submissions per day? It is not enough even for debugging your not clearly understanding structure of files, not to mention about own bugs... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 892415,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2020-06-18T21:21:03.630000",
          "content": "<p>Hi, the test_audio folder is visible to the submitted notebook.\n<a href=\"/cwthompson\">@cwthompson</a> put together a nice example submission notebook here, which you can follow:\n<a href=\"https://www.kaggle.com/cwthompson/birdsong-making-a-prediction\">https://www.kaggle.com/cwthompson/birdsong-making-a-prediction</a></p>\n\n<p>Alex: It's up to the competitors to decide how to handle the third site. You could take a short-segment model and aggregate up to the 10 minute window, or use an entirely different model; your call.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 892424,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-18T21:33:03.373000",
          "content": "<p>Thank you for answer! if you look in last code competitions with hidden test you can see this folder visible in check phase too. It is right way, because it checks bugs with base structure of files without button submission. Also your approach breaks display public notebooks without this folder. </p>\n\n<p>You should add this folder and change content after button submission</p>\n\n<p>And you didn’t answer about 2 submissions per day.</p>\n\n<p><a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> <a href=\"/tomdenton\">@tomdenton</a> </p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 892582,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-06-19T02:55:33.690000",
          "content": "<p>I agree with <a href=\"/shonenkov\">@shonenkov</a> , it would be much easier for us to create a submission if we have <code>test_audio</code> folder not only in the submission phase but also <em>before</em> committing (I mean coding phase). I know you do not want to make it visible to avoid competitor trying nonsense stuffs, but as <a href=\"/shonenkov\">@shonenkov</a> says you can change the contents in <code>test_audio</code> between coding phase and submission phase.</p>\n\n<p>Right now we have to do without <code>test_audio</code> folder and that makes us unconfident about whether our commit works in submission phase.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 893460,
          "author_name": "Tom Denton",
          "author_url": "",
          "post_date": "2020-06-19T16:13:02.230000",
          "content": "<p>The existing structure is sufficient to write reliable submission code; there are some examples of testing for the existence of the test_audio folder in some other competitor notebooks, for example. If you're very concerned, and you're doing any development outside the Kaggle editors, you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.) We're hesitant to change things after the competition has started unless really necessary.</p>\n\n<p>Alex: We limit to two submissions per day to avoid submission spamming to probe/overfit the test datasets. </p>",
          "votes": -19,
          "replies": []
        },
        {
          "id": 893484,
          "author_name": "Alex Shonenkov",
          "author_url": "",
          "post_date": "2020-06-19T16:39:03.040000",
          "content": "<p><a href=\"/tomdenton\">@tomdenton</a> are you sure?</p>\n\n<p>this topics from Novice --&gt; GM about this problem.</p>\n\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159992\">https://www.kaggle.com/c/birdsong-recognition/discussion/159992</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160007\">https://www.kaggle.com/c/birdsong-recognition/discussion/160007</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159173\">https://www.kaggle.com/c/birdsong-recognition/discussion/159173</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159123\">https://www.kaggle.com/c/birdsong-recognition/discussion/159123</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158951\">https://www.kaggle.com/c/birdsong-recognition/discussion/158951</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159507\">https://www.kaggle.com/c/birdsong-recognition/discussion/159507</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160111\">https://www.kaggle.com/c/birdsong-recognition/discussion/160111</a>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/160110\">https://www.kaggle.com/c/birdsong-recognition/discussion/160110</a></p>\n\n<p>It is disrespectful of you to ignore obvious problems. I spent last 6 submissions (3 days!?) for solving it for all competitors (I did your work):</p>\n\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/159993\">https://www.kaggle.com/c/birdsong-recognition/discussion/159993</a></p>\n\n<p>But I wanted only to share my baseline for your competition...</p>\n\n<p>P.S. Dont forget to read <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#886135\">topic about DFDC</a> where hosts ignored obvious problems during competition. </p>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 894889,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-20T21:33:30.313000",
          "content": "<p>Hi <a href=\"/tomdenton\">@tomdenton</a>,\nYou said 'The existing structure is sufficient ...', see above for full post.\nTechnically your statement is probably correct, but the data structure is a bit of a headache.\nMany of us already have to deal with such complicated data structures during day-time at out jobs. \nMy fear is that more that one will end-up ignoring this competition in favor of others where they can directly dive into the fun stuff like ML and DNNs.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 894900,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-06-20T21:58:28.127000",
          "content": "<p>&gt; The existing structure is sufficient to write reliable submission code; there are some examples of testing for the existence of the test_audio folder in some other competitor notebooks, for example. If you're very concerned, and you're doing any development outside the Kaggle editors, you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.)</p>\n\n<p>I agree it is <em>possible</em> to write unit tests and checks (like how <a href=\"/shonenkov\">@shonenkov</a> has graciously shared one such method for it) but it is not the standard way and not at all beginner-friendly. Especially considering Kaggle already has a feature and system that is designed to solve this problem as described by <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> </p>\n\n<p>&gt; We're hesitant to change things after the competition has started unless really necessary.</p>\n\n<p>The 3-month long competition began less than a week back and as of this point, no one has made any submission that beats the sample submission. If this is resolved / fixed soon (even if it requires changes or a restart), I can bet Kagglers will be much happier without any complaints.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 948306,
          "author_name": "Mark Eckdahl",
          "author_url": "",
          "post_date": "2020-07-27T19:39:20.463000",
          "content": "<p><a href=\"/shonenkov\">@shonenkov</a> - My goodness the Deep Fake scoring / data provided was atrocious!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 995480,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-02T13:35:51.723000",
          "content": "<p><a href=\"https://www.kaggle.com/shonenkov\" target=\"_blank\">@shonenkov</a> I am reading all your posts as it seems I am following your path with failed submissions that really look fine.  Thanks for the fake test data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 995760,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-02T18:41:28.567000",
          "content": "<blockquote>\n  <p>you could write some unit tests to ensure that your code will work as expected at submission time. (This is what I would do in this situation.) </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> are you aware that what you call a unit test is a submission for us?  We have only two of them per day.  </p>\n<p>I am only beginning this journey but when I see a seasoned Kaggler GM being unable to get a successful submission in few days this week, then I believe there is a real issue with the setup.  I asked a very simple question in another thread: can one of you check that reading all actual test files with librosa gives files with sampling rates of 32000 and duration 10 minutes?</p>\n<p>Using the time of hundreds of Kagglers to debug by trial and error something that could be just documented properly baffles me to be honest.  Main reason is that while we are doing this we are not building better models for you.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 946035,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2020-07-26T09:55:04.557000",
      "content": "<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> sorry for such a naive question but there is something that has been bugging me and I will be really happy if you answer this . All the metadata columns which we have in train are not available during test time right? and we have only and only audio data available ? \nThis also means everything other than the audio becomes useless during train time right?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 946066,
          "author_name": "Dhruv Naik",
          "author_url": "",
          "post_date": "2020-07-26T10:25:54.830000",
          "content": "<p>Yup. Metadata is not for training the model.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 946079,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-07-26T10:41:52.320000",
          "content": "<blockquote>\n  <p>This also means everything other than the audio becomes useless during train time right?</p>\n</blockquote>\n\n<p>Basically yes, but if you get luck in predicting some of the columns, they may be useful.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 946086,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2020-07-26T10:49:05.163000",
          "content": "<p>Thanks for your answers Hidehisa and Dhruv</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 890497,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-17T14:32:36.913000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 888003,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-16T04:03:38.637000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "887963": "I'm still not clear about the structure of test data.\n\nIn data description, they say\n\n&gt; **test_audio** The hidden test set audio consists of approximately 150 recordings in mp3 format, each roughly 10 minutes long. The recordings were taken at three separate remote locations. Sites 1 and 2 were labeled in 5 second increments and need matching predictions, but due to the time consuming nature of the labeling process the site 3 files are only labeled at the file level. Accordingly, site 3 has relatively few rows in the test set and needs lower time resolution predictions.\n\nso we cannot see test set at all. Instead we have `example_test_audio` but we don't need to do prediction on audio files in this directory, am I right about this?\n\nWhat we need to do in the inference phase is\n\n1. Read `test.csv`\n2. Repeat 3 - 5 for each row in `test.csv`\n3. Open test audio file which is contained in `/kaggle/input/birdsong-recognition/test_audio` and named like `audio_id.mp3` .\n4. If `site` is `site_3` use the whole clip, otherwise cut 5 seconds short clip out of the long audio clip which we loaded in step3 that ends in the second specified in `seconds` column.\n5. Perform prediction on the short clip which we prepared in step 4. Note that each clip may have multiple labels.\n\n↑Am I right about this?",
    "890937": "For others having trouble, it looks like @cwthompson put together a working skeletal submission example here, which you can refer to:\n\nhttps://www.kaggle.com/cwthompson/birdsong-making-a-prediction",
    "890473": "@hidehisaarai1213 your description is accurate.",
    "888738": "What I gather from looking at the dataset is:\n\n1. 0/150 audio files from the test data are being shared, so inference will be done only during kernel re-run on submission of `sample_submission.csv`\n\n2. For **validation**, we can use audio file: `example_test_audio` and rows `filename_seconds` and `birds` from `example_test_audio_summary.csv` , but this is **not** part of the test set.\n\n3. For *site1* and *site 2*, predictions of `birds` is done for every **5 second interval** of the audio-file.\n    For site3, prediction of `birds` is for entire audio-file\n\n4. For each row, there can be multiple `birds`",
    "888513": "I don't see a test audio folder.  I see an example test audio that goes with the example test metadata and example test summary.\n\nPlus test.csv is only 3 rows.  ",
    "946035": "@hidehisaarai1213 sorry for such a naive question but there is something that has been bugging me and I will be really happy if you answer this . All the metadata columns which we have in train are not available during test time right? and we have only and only audio data available ? \nThis also means everything other than the audio becomes useless during train time right?",
    "890497": "",
    "888003": ""
  }
}