{
  "id": 309001,
  "title": "Found discrepancy between Test CSV and provided soundscape",
  "url": "/competitions/birdclef-2022/discussion/309001",
  "author_name": "",
  "post_date": "2022-02-21T11:22:24.985312200Z",
  "votes": 8,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hello, I have found some discrepancy between Test CSV and provided test soundscape which prevents me to run testing. The test CSV file includes column <code>file_id</code> but the particular name does not exist in <code>test_soundscapes</code> folder. The sample test table is:</p>\n<pre><code>row_id,file_id,bird,end_time\nsoundscape_1000170626_akiapo_5,soundscape_1000170626,akiapo,5\nsoundscape_1000170626_akiapo_10,soundscape_1000170626,akiapo,10\nsoundscape_1000170626_akiapo_15,soundscape_1000170626,akiapo,15\n</code></pre>\n<p>but the provided data/dataset is:</p>\n<pre><code>birdclef-2022\n |- test_soundscapes\n |   L soundscape_453028782.ogg\n |- train_audio\n |- eBird_Taxonomy_v2021.csv\n |- sample_submission.csv\n |- scored_birds.json\n |- test.csv\n L train_metadata.csv\n</code></pre>\n<p>So I am wondering if it is a bug or I miss some part of the test data mapping… <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>?</p>",
  "messages": [
    {
      "id": "1699701",
      "postDate": "02/21/2022 11:22:24",
      "content": "<p>Hello, I have found some discrepancy between Test CSV and provided test soundscape which prevents me to run testing. The test CSV file includes column <code>file_id</code> but the particular name does not exist in <code>test_soundscapes</code> folder. The sample test table is:</p>\n<pre><code>row_id,file_id,bird,end_time\nsoundscape_1000170626_akiapo_5,soundscape_1000170626,akiapo,5\nsoundscape_1000170626_akiapo_10,soundscape_1000170626,akiapo,10\nsoundscape_1000170626_akiapo_15,soundscape_1000170626,akiapo,15\n</code></pre>\n<p>but the provided data/dataset is:</p>\n<pre><code>birdclef-2022\n |- test_soundscapes\n |   L soundscape_453028782.ogg\n |- train_audio\n |- eBird_Taxonomy_v2021.csv\n |- sample_submission.csv\n |- scored_birds.json\n |- test.csv\n L train_metadata.csv\n</code></pre>\n<p>So I am wondering if it is a bug or I miss some part of the test data mapping… <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>?</p>",
      "rawMarkdown": "Hello, I have found some discrepancy between Test CSV and provided test soundscape which prevents me to run testing. The test CSV file includes column `file_id` but the particular name does not exist in `test_soundscapes` folder. The sample test table is:\n\n```\nrow_id,file_id,bird,end_time\nsoundscape_1000170626_akiapo_5,soundscape_1000170626,akiapo,5\nsoundscape_1000170626_akiapo_10,soundscape_1000170626,akiapo,10\nsoundscape_1000170626_akiapo_15,soundscape_1000170626,akiapo,15\n```\n\nbut the provided data/dataset is:\n```\nbirdclef-2022\n |- test_soundscapes\n |   L soundscape_453028782.ogg\n |- train_audio\n |- eBird_Taxonomy_v2021.csv\n |- sample_submission.csv\n |- scored_birds.json\n |- test.csv\n L train_metadata.csv\n```\n\nSo I am wondering if it is a bug or I miss some part of the test data mapping... @stefankahl?",
      "votes": null
    },
    {
      "id": "1700121",
      "postDate": "02/21/2022 17:08:02",
      "content": "<p>This is somewhat intentional and no bug, yet not a 100% ideal. For testing, you might need to parse the test_soundscape folder for files and then switch to the test.csv for a submission. Or stick with parsing the soundscape dir. Both should work fine when submitting.</p>",
      "rawMarkdown": "This is somewhat intentional and no bug, yet not a 100% ideal. For testing, you might need to parse the test_soundscape folder for files and then switch to the test.csv for a submission. Or stick with parsing the soundscape dir. Both should work fine when submitting.",
      "votes": null
    },
    {
      "id": "1700147",
      "postDate": "02/21/2022 17:38:10",
      "content": "<p>ok, maybe I miss something… to make a submission I need to have a valid kernel (which runs with the dummy test data). It means the kernel takes the test data and creates a submission file. So when I open the test.csv I have reference to an audio file named <code>soundscape_1000170626</code> but there is no such file in the dataset (in the test folder is <code>soundscape_453028782</code>) so based on what shall I make the predictions?</p>\n<p>My point is that the name in CSV does not match the audio file name…</p>",
      "rawMarkdown": "ok, maybe I miss something... to make a submission I need to have a valid kernel (which runs with the dummy test data). It means the kernel takes the test data and creates a submission file. So when I open the test.csv I have reference to an audio file named `soundscape_1000170626` but there is no such file in the dataset (in the test folder is `soundscape_453028782`) so based on what shall I make the predictions?\n\nMy point is that the name in CSV does not match the audio file name...",
      "votes": null
    },
    {
      "id": "1700186",
      "postDate": "02/21/2022 18:18:38",
      "content": "<p>I'd suggest to have a look on the <br>\n<a href=\"https://www.kaggle.com/stefankahl/how-to-submit-to-birdclef-2022\" target=\"_blank\">submission nb</a> provided by host - it should make clear how to parse test audio files</p>",
      "rawMarkdown": "I'd suggest to have a look on the \n[submission nb](https://www.kaggle.com/stefankahl/how-to-submit-to-birdclef-2022) provided by host - it should make clear how to parse test audio files",
      "votes": null
    },
    {
      "id": "1700217",
      "postDate": "02/21/2022 18:42:06",
      "content": "<p>I see, that the sample kernel runs prediction for all 5-secund frames for all birds… completely ignoring the test.csv so any idea why we should care about it for submission as descriptions states \"the full test.csv is provided in the hidden test set\"?</p>\n<p>interestingly, why do we need the submission.csv as mentioned \"the full submission.csv is provided in the hidden test set.\"</p>",
      "rawMarkdown": "I see, that the sample kernel runs prediction for all 5-secund frames for all birds... completely ignoring the test.csv so any idea why we should care about it for submission as descriptions states \"the full test.csv is provided in the hidden test set\"?\n\ninterestingly, why do we need the submission.csv as mentioned \"the full submission.csv is provided in the hidden test set.\"",
      "votes": null
    },
    {
      "id": "1700335",
      "postDate": "02/21/2022 20:57:25",
      "content": "<p>they mean that when you press the \"Submit\" they replace the sample submission csv with the full one that contain ALL audios (public and private a.k.a hidden test).. as usual in code competition format   </p>",
      "rawMarkdown": "they mean that when you press the \"Submit\" they replace the sample submission csv with the full one that contain ALL audios (public and private a.k.a hidden test).. as usual in code competition format",
      "votes": null
    },
    {
      "id": "1700345",
      "postDate": "02/21/2022 21:33:45",
      "content": "<p>Yeah, I see your point, but please see mine too.. with the actual setting you are not able to use their test as a mock to develop your submission… for illustration you have the task of result <code>a+b</code> for inpouts 2 and 4 and result is 42</p>",
      "rawMarkdown": "Yeah, I see your point, but please see mine too.. with the actual setting you are not able to use their test as a mock to develop your submission... for illustration you have the task of result `a+b` for inpouts 2 and 4 and result is 42",
      "votes": null
    },
    {
      "id": "1713755",
      "postDate": "03/06/2022 10:45:01",
      "content": "<p><a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> have you made any progress on the matter? I've just joined the competition and I'm kinda puzzled, too.</p>",
      "rawMarkdown": "jirkaborovec have you made any progress on the matter? I've just joined the competition and I'm kinda puzzled, too.",
      "votes": null
    },
    {
      "id": "1713806",
      "postDate": "03/06/2022 12:04:20",
      "content": "<p>seems there is no update on the obvious issue, so just continued with creating predictions for all birds and ignoring the test.csv file, see <a href=\"https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions</a></p>",
      "rawMarkdown": "seems there is no update on the obvious issue, so just continued with creating predictions for all birds and ignoring the test.csv file, see https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions",
      "votes": null
    },
    {
      "id": "1714862",
      "postDate": "03/07/2022 12:27:07",
      "content": "<p>Hey I did this to have an easy way around the problem :</p>\n<pre><code>test = pd.read_csv('../input/birdclef-2022/test.csv')\nif len(test) == 3:\n    test['file_id'] = 'soundscape_453028782'\n</code></pre>\n<p>Then, when I save the notebook I am sure it will run without problem and that it can work for the actual submission.</p>",
      "rawMarkdown": "Hey I did this to have an easy way around the problem :\n\n```\ntest = pd.read_csv('../input/birdclef-2022/test.csv')\nif len(test) == 3:\n    test['file_id'] = 'soundscape_453028782'\n```\n\n\nThen, when I save the notebook I am sure it will run without problem and that it can work for the actual submission.",
      "votes": null
    },
    {
      "id": "1714882",
      "postDate": "03/07/2022 12:55:10",
      "content": "<p>cool, but then you will run your predictions on one audio file instead of the expected test dataset…</p>",
      "rawMarkdown": "cool, but then you will run your predictions on one audio file instead of the expected test dataset...",
      "votes": null
    },
    {
      "id": "1714910",
      "postDate": "03/07/2022 13:26:01",
      "content": "<p>Yes, if the length of the dataset (test) is equal to 3.</p>\n<p>But when it comes to the actual submission the test dataframe will be longer than 3. Thus, you will use the real files*.</p>\n<p>*Because the 'len(test) == 3' condition will not be true, so you will not change the file_id names.</p>",
      "rawMarkdown": "Yes, if the length of the dataset (test) is equal to 3.\n\nBut when it comes to the actual submission the test dataframe will be longer than 3. Thus, you will use the real files*.\n\n*Because the 'len(test) == 3' condition will not be true, so you will not change the file_id names.",
      "votes": null
    },
    {
      "id": "1714946",
      "postDate": "03/07/2022 14:07:35",
      "content": "<p>I know already how to generate the submission without the test.csv, I was just asking the hosted could fix it and make the dataset part related to submission rather useful… :)</p>",
      "rawMarkdown": "I know already how to generate the submission without the test.csv, I was just asking the hosted could fix it and make the dataset part related to submission rather useful... :)",
      "votes": null
    },
    {
      "id": "1714949",
      "postDate": "03/07/2022 14:09:04",
      "content": "<p>Ha sorry ! I misunderstood the problem then :) !</p>",
      "rawMarkdown": "Ha sorry ! I misunderstood the problem then :) !",
      "votes": null
    },
    {
      "id": "1720691",
      "postDate": "03/13/2022 03:27:53",
      "content": "<blockquote>\n  <p>Or stick with parsing the soundscape dir. Both should work fine when submitting.</p>\n</blockquote>\n<p>Is this really correct?<br>\nI searched for soundscape dir and tried to create a submission file, <br>\nbut I got an error if I don't set the number of segments to 12 (equivalent to 60 seconds) for all files.</p>\n<p><strong>Are files with more than 60 seconds evaluated only for the first 60 seconds?</strong></p>\n<p>I haven't tried submission based on <code>test.csv</code> because I've exceeded today's submission limit.<br>\nBut I'm guessing that the maximum <code>end_time</code> of each file in the <code>test.csv</code> is \"_60\"  even though there are files longer than 60 seconds..</p>",
      "rawMarkdown": "> Or stick with parsing the soundscape dir. Both should work fine when submitting.\n\nIs this really correct?\nI searched for soundscape dir and tried to create a submission file, \nbut I got an error if I don't set the number of segments to 12 (equivalent to 60 seconds) for all files.\n\n**Are files with more than 60 seconds evaluated only for the first 60 seconds?**\n\nI haven't tried submission based on `test.csv` because I've exceeded today's submission limit.\nBut I'm guessing that the maximum `end_time` of each file in the `test.csv` is \"_60\"  even though there are files longer than 60 seconds..",
      "votes": null
    },
    {
      "id": "1720871",
      "postDate": "03/13/2022 08:13:21",
      "content": "<p>yes, generating the submission purely on the folder with soundtracks and <code>scored_birds.json</code> seems to be working for me in this dummy submission: <a href=\"https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions</a></p>",
      "rawMarkdown": "yes, generating the submission purely on the folder with soundtracks and `scored_birds.json` seems to be working for me in this dummy submission: https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions",
      "votes": null
    },
    {
      "id": "1720954",
      "postDate": "03/13/2022 09:37:24",
      "content": "<p>Thank you for your reply!<br>\nThe only difference between my failed code and your code seems to be 'ceil' (yours) or 'floor' (mine).<br>\nSo it seems that the cause of my failure is a lack of segments, but..<br>\nI'm confused by the success of my code with a fixed number of segments of 12.</p>\n<p>If all the test data are 60 seconds, there should be no difference between 'ceil' and 'floor'.<br>\nIf there is a test data that is not 60 seconds, I don't know why the submission is successful with the fixed 12 segments.</p>\n<p>I will try it tomorrow with reference to your code!<br>\nThanks again!!</p>\n<p>~~~<br>\nThis is my failed code. (Of course, sample submission works fine.)</p>\n<pre><code>for fpath in test_files:\n    sig, rate = librosa.load(fpath, sr=32000, mono=True)\n    num_segments = len(sig) // (rate*5)    \n    for seg_idx in range(num_segments): # &lt;---\n        ...\n</code></pre>\n<p>The following code was successful.</p>\n<pre><code>for fpath in test_files:\n    sig, rate = librosa.load(fpath, sr=32000, mono=True)\n    for seg_idx in range(12): # &lt;---\n        ...\n</code></pre>",
      "rawMarkdown": "Thank you for your reply!\nThe only difference between my failed code and your code seems to be 'ceil' (yours) or 'floor' (mine).\nSo it seems that the cause of my failure is a lack of segments, but..\nI'm confused by the success of my code with a fixed number of segments of 12.\n\nIf all the test data are 60 seconds, there should be no difference between 'ceil' and 'floor'.\nIf there is a test data that is not 60 seconds, I don't know why the submission is successful with the fixed 12 segments.\n\nI will try it tomorrow with reference to your code!\nThanks again!!\n\n~~~\nThis is my failed code. (Of course, sample submission works fine.)\n```\nfor fpath in test_files:\n    sig, rate = librosa.load(fpath, sr=32000, mono=True)\n    num_segments = len(sig) // (rate*5)    \n    for seg_idx in range(num_segments): # <---\n        ...\n```\n\nThe following code was successful.\n```\nfor fpath in test_files:\n    sig, rate = librosa.load(fpath, sr=32000, mono=True)\n    for seg_idx in range(12): # <---\n        ...\n```",
      "votes": null
    },
    {
      "id": "1721009",
      "postDate": "03/13/2022 10:48:33",
      "content": "<p>from data tab</p>\n<blockquote>\n  <p>the test_soundscapes directory will be populated with approximately 5,500 recordings to be used for scoring. These are each <strong>within a few milliseconds of 1 minute long</strong> and in the ogg audio format.</p>\n</blockquote>",
      "rawMarkdown": "from data tab\n\n> the test_soundscapes directory will be populated with approximately 5,500 recordings to be used for scoring. These are each **within a few milliseconds of 1 minute long** and in the ogg audio format.",
      "votes": null
    },
    {
      "id": "1721036",
      "postDate": "03/13/2022 11:06:51",
      "content": "<p>Thank you!<br>\nI completely missed it..<br>\nSince 'ceil' works well, the audio file is probably up to 1 minute long (not over 1 minute).</p>",
      "rawMarkdown": "Thank you!\nI completely missed it..\nSince 'ceil' works well, the audio file is probably up to 1 minute long (not over 1 minute).",
      "votes": null
    },
    {
      "id": "1721445",
      "postDate": "03/13/2022 17:28:42",
      "content": "<p>I tripped over the exact same thing but after some failed attempts it works fine with <code>ceil</code>. But good to know that they all seem to be <code>&lt;= 60 sec</code> - <em>within a few milliseconds of 1 minute long</em> could be <code>60k ms +- n ms</code>.</p>",
      "rawMarkdown": "I tripped over the exact same thing but after some failed attempts it works fine with `ceil`. But good to know that they all seem to be `<= 60 sec` - *within a few milliseconds of 1 minute long* could be `60k ms +- n ms`.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1700121,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "02/21/2022 17:08:02",
      "content": "<p>This is somewhat intentional and no bug, yet not a 100% ideal. For testing, you might need to parse the test_soundscape folder for files and then switch to the test.csv for a submission. Or stick with parsing the soundscape dir. Both should work fine when submitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1700147,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "02/21/2022 17:38:10",
          "content": "<p>ok, maybe I miss something… to make a submission I need to have a valid kernel (which runs with the dummy test data). It means the kernel takes the test data and creates a submission file. So when I open the test.csv I have reference to an audio file named <code>soundscape_1000170626</code> but there is no such file in the dataset (in the test folder is <code>soundscape_453028782</code>) so based on what shall I make the predictions?</p>\n<p>My point is that the name in CSV does not match the audio file name…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1700186,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "02/21/2022 18:18:38",
          "content": "<p>I'd suggest to have a look on the <br>\n<a href=\"https://www.kaggle.com/stefankahl/how-to-submit-to-birdclef-2022\" target=\"_blank\">submission nb</a> provided by host - it should make clear how to parse test audio files</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1700217,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "02/21/2022 18:42:06",
          "content": "<p>I see, that the sample kernel runs prediction for all 5-secund frames for all birds… completely ignoring the test.csv so any idea why we should care about it for submission as descriptions states \"the full test.csv is provided in the hidden test set\"?</p>\n<p>interestingly, why do we need the submission.csv as mentioned \"the full submission.csv is provided in the hidden test set.\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1700335,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "02/21/2022 20:57:25",
          "content": "<p>they mean that when you press the \"Submit\" they replace the sample submission csv with the full one that contain ALL audios (public and private a.k.a hidden test).. as usual in code competition format   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1700345,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "02/21/2022 21:33:45",
          "content": "<p>Yeah, I see your point, but please see mine too.. with the actual setting you are not able to use their test as a mock to develop your submission… for illustration you have the task of result <code>a+b</code> for inpouts 2 and 4 and result is 42</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1713755,
          "author_name": "m02ph3u5",
          "author_url": "",
          "post_date": "03/06/2022 10:45:01",
          "content": "<p><a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> have you made any progress on the matter? I've just joined the competition and I'm kinda puzzled, too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1713806,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "03/06/2022 12:04:20",
          "content": "<p>seems there is no update on the obvious issue, so just continued with creating predictions for all birds and ignoring the test.csv file, see <a href=\"https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1714862,
          "author_name": "gabiboubibel",
          "author_url": "",
          "post_date": "03/07/2022 12:27:07",
          "content": "<p>Hey I did this to have an easy way around the problem :</p>\n<pre><code>test = pd.read_csv('../input/birdclef-2022/test.csv')\nif len(test) == 3:\n    test['file_id'] = 'soundscape_453028782'\n</code></pre>\n<p>Then, when I save the notebook I am sure it will run without problem and that it can work for the actual submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1714882,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "03/07/2022 12:55:10",
          "content": "<p>cool, but then you will run your predictions on one audio file instead of the expected test dataset…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1714910,
          "author_name": "gabiboubibel",
          "author_url": "",
          "post_date": "03/07/2022 13:26:01",
          "content": "<p>Yes, if the length of the dataset (test) is equal to 3.</p>\n<p>But when it comes to the actual submission the test dataframe will be longer than 3. Thus, you will use the real files*.</p>\n<p>*Because the 'len(test) == 3' condition will not be true, so you will not change the file_id names.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1714946,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "03/07/2022 14:07:35",
          "content": "<p>I know already how to generate the submission without the test.csv, I was just asking the hosted could fix it and make the dataset part related to submission rather useful… :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1714949,
          "author_name": "gabiboubibel",
          "author_url": "",
          "post_date": "03/07/2022 14:09:04",
          "content": "<p>Ha sorry ! I misunderstood the problem then :) !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1720691,
          "author_name": "octpath0302",
          "author_url": "",
          "post_date": "03/13/2022 03:27:53",
          "content": "<blockquote>\n  <p>Or stick with parsing the soundscape dir. Both should work fine when submitting.</p>\n</blockquote>\n<p>Is this really correct?<br>\nI searched for soundscape dir and tried to create a submission file, <br>\nbut I got an error if I don't set the number of segments to 12 (equivalent to 60 seconds) for all files.</p>\n<p><strong>Are files with more than 60 seconds evaluated only for the first 60 seconds?</strong></p>\n<p>I haven't tried submission based on <code>test.csv</code> because I've exceeded today's submission limit.<br>\nBut I'm guessing that the maximum <code>end_time</code> of each file in the <code>test.csv</code> is \"_60\"  even though there are files longer than 60 seconds..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1720871,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "03/13/2022 08:13:21",
          "content": "<p>yes, generating the submission purely on the folder with soundtracks and <code>scored_birds.json</code> seems to be working for me in this dummy submission: <a href=\"https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1720954,
          "author_name": "octpath0302",
          "author_url": "",
          "post_date": "03/13/2022 09:37:24",
          "content": "<p>Thank you for your reply!<br>\nThe only difference between my failed code and your code seems to be 'ceil' (yours) or 'floor' (mine).<br>\nSo it seems that the cause of my failure is a lack of segments, but..<br>\nI'm confused by the success of my code with a fixed number of segments of 12.</p>\n<p>If all the test data are 60 seconds, there should be no difference between 'ceil' and 'floor'.<br>\nIf there is a test data that is not 60 seconds, I don't know why the submission is successful with the fixed 12 segments.</p>\n<p>I will try it tomorrow with reference to your code!<br>\nThanks again!!</p>\n<p>~~~<br>\nThis is my failed code. (Of course, sample submission works fine.)</p>\n<pre><code>for fpath in test_files:\n    sig, rate = librosa.load(fpath, sr=32000, mono=True)\n    num_segments = len(sig) // (rate*5)    \n    for seg_idx in range(num_segments): # &lt;---\n        ...\n</code></pre>\n<p>The following code was successful.</p>\n<pre><code>for fpath in test_files:\n    sig, rate = librosa.load(fpath, sr=32000, mono=True)\n    for seg_idx in range(12): # &lt;---\n        ...\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721009,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "03/13/2022 10:48:33",
          "content": "<p>from data tab</p>\n<blockquote>\n  <p>the test_soundscapes directory will be populated with approximately 5,500 recordings to be used for scoring. These are each <strong>within a few milliseconds of 1 minute long</strong> and in the ogg audio format.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721036,
          "author_name": "octpath0302",
          "author_url": "",
          "post_date": "03/13/2022 11:06:51",
          "content": "<p>Thank you!<br>\nI completely missed it..<br>\nSince 'ceil' works well, the audio file is probably up to 1 minute long (not over 1 minute).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721445,
          "author_name": "m02ph3u5",
          "author_url": "",
          "post_date": "03/13/2022 17:28:42",
          "content": "<p>I tripped over the exact same thing but after some failed attempts it works fine with <code>ceil</code>. But good to know that they all seem to be <code>&lt;= 60 sec</code> - <em>within a few milliseconds of 1 minute long</em> could be <code>60k ms +- n ms</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1699701": "Hello, I have found some discrepancy between Test CSV and provided test soundscape which prevents me to run testing. The test CSV file includes column `file_id` but the particular name does not exist in `test_soundscapes` folder. The sample test table is:\n\n```\nrow_id,file_id,bird,end_time\nsoundscape_1000170626_akiapo_5,soundscape_1000170626,akiapo,5\nsoundscape_1000170626_akiapo_10,soundscape_1000170626,akiapo,10\nsoundscape_1000170626_akiapo_15,soundscape_1000170626,akiapo,15\n```\n\nbut the provided data/dataset is:\n```\nbirdclef-2022\n |- test_soundscapes\n |   L soundscape_453028782.ogg\n |- train_audio\n |- eBird_Taxonomy_v2021.csv\n |- sample_submission.csv\n |- scored_birds.json\n |- test.csv\n L train_metadata.csv\n```\n\nSo I am wondering if it is a bug or I miss some part of the test data mapping... @stefankahl?",
    "1700121": "This is somewhat intentional and no bug, yet not a 100% ideal. For testing, you might need to parse the test_soundscape folder for files and then switch to the test.csv for a submission. Or stick with parsing the soundscape dir. Both should work fine when submitting.",
    "1700147": "ok, maybe I miss something... to make a submission I need to have a valid kernel (which runs with the dummy test data). It means the kernel takes the test data and creates a submission file. So when I open the test.csv I have reference to an audio file named `soundscape_1000170626` but there is no such file in the dataset (in the test folder is `soundscape_453028782`) so based on what shall I make the predictions?\n\nMy point is that the name in CSV does not match the audio file name...",
    "1700186": "I'd suggest to have a look on the \n[submission nb](https://www.kaggle.com/stefankahl/how-to-submit-to-birdclef-2022) provided by host - it should make clear how to parse test audio files",
    "1700217": "I see, that the sample kernel runs prediction for all 5-secund frames for all birds... completely ignoring the test.csv so any idea why we should care about it for submission as descriptions states \"the full test.csv is provided in the hidden test set\"?\n\ninterestingly, why do we need the submission.csv as mentioned \"the full submission.csv is provided in the hidden test set.\"",
    "1700335": "they mean that when you press the \"Submit\" they replace the sample submission csv with the full one that contain ALL audios (public and private a.k.a hidden test).. as usual in code competition format",
    "1700345": "Yeah, I see your point, but please see mine too.. with the actual setting you are not able to use their test as a mock to develop your submission... for illustration you have the task of result `a+b` for inpouts 2 and 4 and result is 42",
    "1713755": "jirkaborovec have you made any progress on the matter? I've just joined the competition and I'm kinda puzzled, too.",
    "1713806": "seems there is no update on the obvious issue, so just continued with creating predictions for all birds and ignoring the test.csv file, see https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions",
    "1714862": "Hey I did this to have an easy way around the problem :\n\n```\ntest = pd.read_csv('../input/birdclef-2022/test.csv')\nif len(test) == 3:\n    test['file_id'] = 'soundscape_453028782'\n```\n\n\nThen, when I save the notebook I am sure it will run without problem and that it can work for the actual submission.",
    "1714882": "cool, but then you will run your predictions on one audio file instead of the expected test dataset...",
    "1714910": "Yes, if the length of the dataset (test) is equal to 3.\n\nBut when it comes to the actual submission the test dataframe will be longer than 3. Thus, you will use the real files*.\n\n*Because the 'len(test) == 3' condition will not be true, so you will not change the file_id names.",
    "1714946": "I know already how to generate the submission without the test.csv, I was just asking the hosted could fix it and make the dataset part related to submission rather useful... :)",
    "1714949": "Ha sorry ! I misunderstood the problem then :) !",
    "1720691": "> Or stick with parsing the soundscape dir. Both should work fine when submitting.\n\nIs this really correct?\nI searched for soundscape dir and tried to create a submission file, \nbut I got an error if I don't set the number of segments to 12 (equivalent to 60 seconds) for all files.\n\n**Are files with more than 60 seconds evaluated only for the first 60 seconds?**\n\nI haven't tried submission based on `test.csv` because I've exceeded today's submission limit.\nBut I'm guessing that the maximum `end_time` of each file in the `test.csv` is \"_60\"  even though there are files longer than 60 seconds..",
    "1720871": "yes, generating the submission purely on the folder with soundtracks and `scored_birds.json` seems to be working for me in this dummy submission: https://www.kaggle.com/jirkaborovec/birdclef-fake-predictions",
    "1720954": "Thank you for your reply!\nThe only difference between my failed code and your code seems to be 'ceil' (yours) or 'floor' (mine).\nSo it seems that the cause of my failure is a lack of segments, but..\nI'm confused by the success of my code with a fixed number of segments of 12.\n\nIf all the test data are 60 seconds, there should be no difference between 'ceil' and 'floor'.\nIf there is a test data that is not 60 seconds, I don't know why the submission is successful with the fixed 12 segments.\n\nI will try it tomorrow with reference to your code!\nThanks again!!\n\n~~~\nThis is my failed code. (Of course, sample submission works fine.)\n```\nfor fpath in test_files:\n    sig, rate = librosa.load(fpath, sr=32000, mono=True)\n    num_segments = len(sig) // (rate*5)    \n    for seg_idx in range(num_segments): # <---\n        ...\n```\n\nThe following code was successful.\n```\nfor fpath in test_files:\n    sig, rate = librosa.load(fpath, sr=32000, mono=True)\n    for seg_idx in range(12): # <---\n        ...\n```",
    "1721009": "from data tab\n\n> the test_soundscapes directory will be populated with approximately 5,500 recordings to be used for scoring. These are each **within a few milliseconds of 1 minute long** and in the ogg audio format.",
    "1721036": "Thank you!\nI completely missed it..\nSince 'ceil' works well, the audio file is probably up to 1 minute long (not over 1 minute).",
    "1721445": "I tripped over the exact same thing but after some failed attempts it works fine with `ceil`. But good to know that they all seem to be `<= 60 sec` - *within a few milliseconds of 1 minute long* could be `60k ms +- n ms`."
  },
  "source": "meta"
}