{
  "id": 490964,
  "title": "[Solved] What is the `soundscape_id` in the submission.csv?",
  "url": "/competitions/birdclef-2024/discussion/490964",
  "author_name": "penguin46",
  "post_date": "2024-04-04T04:22:47.702000",
  "votes": 6,
  "comment_count": 14,
  "views": 0,
  "content": "<p>The <code>row_id</code> in sample_submission.csv is described as below, but what is <code>soundscape_id</code>?</p>\n<blockquote>\n  <p>row_id: A slug of [soundscape_id]_[end_time] for the prediction.</p>\n</blockquote>\n<p>Does the test_soundscapes directory has files named like \"soundscape_1446779.ogg\", for which the corresponding <code>soundscape_id</code> is \"soundscape_1446779\"? Or is the test file names like \"1446779.ogg\" and the <code>soundscape_id</code> is \"soundscape_[file_name]\"?</p>\n<hr>\n<p>Solved: The dataset description has been updated and confirmed that the submission passes correctly. Thank you.</p>",
  "messages": [
    {
      "id": 2734328,
      "postDate": "2024-04-04T04:22:47.703Z",
      "content": "<p>The <code>row_id</code> in sample_submission.csv is described as below, but what is <code>soundscape_id</code>?</p>\n<blockquote>\n  <p>row_id: A slug of [soundscape_id]_[end_time] for the prediction.</p>\n</blockquote>\n<p>Does the test_soundscapes directory has files named like \"soundscape_1446779.ogg\", for which the corresponding <code>soundscape_id</code> is \"soundscape_1446779\"? Or is the test file names like \"1446779.ogg\" and the <code>soundscape_id</code> is \"soundscape_[file_name]\"?</p>\n<hr>\n<p>Solved: The dataset description has been updated and confirmed that the submission passes correctly. Thank you.</p>",
      "rawMarkdown": "The `row_id` in sample_submission.csv is described as below, but what is `soundscape_id`?\n\n> row_id: A slug of [soundscape_id]_[end_time] for the prediction.\n\nDoes the test_soundscapes directory has files named like \"soundscape_1446779.ogg\", for which the corresponding `soundscape_id` is \"soundscape_1446779\"? Or is the test file names like \"1446779.ogg\" and the `soundscape_id` is \"soundscape_[file_name]\"?\n\n---\n\nSolved: The dataset description has been updated and confirmed that the submission passes correctly. Thank you.",
      "votes": 6
    },
    {
      "id": 2735579,
      "postDate": "2024-04-04T19:24:11.127Z",
      "content": "<p>I see that each file in \"unlabeled_soundscapes\" folder is in format \"xxxxxx.ogg\", </p>\n<p>Which format does the filenames in test_soundscapes folder in the hidden test use, \"soundscape_xxxxxx.ogg\" or \"xxxxxx.ogg\"?</p>\n<p>I am adapting to sample submissions but my notebook fails in the first minute of hidden test, the same notebook runs in birdclef 2023 without error, as I have run multiple debug runs using other data provided, my suspect is the filename format is changed?</p>",
      "rawMarkdown": "I see that each file in \"unlabeled_soundscapes\" folder is in format \"xxxxxx.ogg\", \n\nWhich format does the filenames in test_soundscapes folder in the hidden test use, \"soundscape_xxxxxx.ogg\" or \"xxxxxx.ogg\"?\n\nI am adapting to sample submissions but my notebook fails in the first minute of hidden test, the same notebook runs in birdclef 2023 without error, as I have run multiple debug runs using other data provided, my suspect is the filename format is changed?",
      "votes": 3,
      "replies": [
        {
          "id": 2735769,
          "postDate": "2024-04-04T22:29:58.043Z",
          "content": "<p>The hidden test set uses file names like <code>soundscape_xxxxxx.ogg</code>.</p>",
          "rawMarkdown": "The hidden test set uses file names like `soundscape_xxxxxx.ogg`.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2735271,
      "postDate": "2024-04-04T16:23:54.807Z",
      "content": "<p>Hello!<br>\nIt is best to adapt some sample submissions. The notebook from last year should help, as the data format is the same (though note that the species list is different):<br>\n<a href=\"https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2023\" target=\"_blank\">https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2023</a></p>",
      "rawMarkdown": "Hello!\nIt is best to adapt some sample submissions. The notebook from last year should help, as the data format is the same (though note that the species list is different):\nhttps://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2023",
      "votes": 3,
      "replies": [
        {
          "id": 2736105,
          "postDate": "2024-04-05T03:26:16.770Z",
          "content": "<p>I've modified the code to adapt to BirdCLEF2024 using the provided reference. However, I'm encountering a submission scoring error. Could you please check if there's something wrong with the <a href=\"https://www.kaggle.com/code/myso1987/submission-scoring-error-submit-to-birdclef-2024/notebook\" target=\"_blank\">code</a>?</p>\n<p><a href=\"https://www.kaggle.com/code/myso1987/submission-scoring-error-submit-to-birdclef-2024/notebook\" target=\"_blank\">https://www.kaggle.com/code/myso1987/submission-scoring-error-submit-to-birdclef-2024/notebook</a></p>",
          "rawMarkdown": "I've modified the code to adapt to BirdCLEF2024 using the provided reference. However, I'm encountering a submission scoring error. Could you please check if there's something wrong with the [code](https://www.kaggle.com/code/myso1987/submission-scoring-error-submit-to-birdclef-2024/notebook)?\n\nhttps://www.kaggle.com/code/myso1987/submission-scoring-error-submit-to-birdclef-2024/notebook",
          "replies": [
            {
              "id": 2736385,
              "postDate": "2024-04-05T06:59:24.540Z",
              "content": "<p>I've fixed the notebook and can submit it successfully. I plan to update the notebook tomorrow. Thanks for sharing the reference.</p>",
              "rawMarkdown": "I've fixed the notebook and can submit it successfully. I plan to update the notebook tomorrow. Thanks for sharing the reference."
            }
          ]
        }
      ]
    },
    {
      "id": 2735299,
      "postDate": "2024-04-04T16:44:08.943Z",
      "content": "<p>Sorry for the confusion. I've updated the data description.</p>",
      "rawMarkdown": "Sorry for the confusion. I've updated the data description.",
      "votes": 1,
      "replies": [
        {
          "id": 2735665,
          "postDate": "2024-04-04T20:33:24.037Z",
          "content": "<p></p>\n<pre><code> pandas  pd\n glob  glob\n os\n numpy  np\n pathlib  \n\n#   bird names\nall_birds = pd.read_csv(\"/kaggle/input/birdclef-2024/sample_submission.csv\").[:].tolist()\n\n#   soundscape_ids\nogg_dir = Path(\"/kaggle/input/birdclef-2024/test_soundscapes\")\n# ogg_dir = Path(\"/kaggle/input/birdclef-2024/unlabeled_soundscapes\")\nogg_files = glob(str(ogg_dir / \"*.ogg\"))\nsoundscape_ids = [Path(f).stem  f  ogg_files]\n\n# generate submission.csv\ndf = pd.DataFrame(np.zeros((len(soundscape_ids), len(all_birds))), =all_birds, )\ndf[\"soundscape_id\"] = soundscape_ids\n\ndfs = []\n end_time  range(, , ):\n    this_df = df.()\n    this_df[\"end_time\"] = end_time\n    this_df[\"row_id\"] = \"soundscape_\" + this_df[\"soundscape_id\"] + \"_\" + this_df[\"end_time\"].astype(str)\n    dfs.append(this_df)\nsub = pd.concat(dfs).sort_values([\"soundscape_id\", \"end_time\"])[[\"row_id\"] + all_birds].reset_index(=)\n\nsub.to_csv(\"submission.csv\", =)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F1856467ad226e02619619cf8fe25f650%2Fsub.png?generation=1712262706796504&amp;alt=media\"></p>",
          "rawMarkdown": "~~Thanks for the update! But I am still confused. The following code still throws a scoring error.\nI would appreciate it if you could check what is wrong.~~\n\n```\nimport pandas as pd\nfrom glob import glob\nimport os\nimport numpy as np\nfrom pathlib import Path\n\n# prepare all bird names\nall_birds = pd.read_csv(\"/kaggle/input/birdclef-2024/sample_submission.csv\").columns[1:].tolist()\n\n# prepare all soundscape_ids\nogg_dir = Path(\"/kaggle/input/birdclef-2024/test_soundscapes\")\n# ogg_dir = Path(\"/kaggle/input/birdclef-2024/unlabeled_soundscapes\")\nogg_files = glob(str(ogg_dir / \"*.ogg\"))\nsoundscape_ids = [Path(f).stem for f in ogg_files]\n\n# generate submission.csv\ndf = pd.DataFrame(np.zeros((len(soundscape_ids), len(all_birds))), columns=all_birds, )\ndf[\"soundscape_id\"] = soundscape_ids\n\ndfs = []\nfor end_time in range(5, 241, 5):\n    this_df = df.copy()\n    this_df[\"end_time\"] = end_time\n    this_df[\"row_id\"] = \"soundscape_\" + this_df[\"soundscape_id\"] + \"_\" + this_df[\"end_time\"].astype(str)\n    dfs.append(this_df)\nsub = pd.concat(dfs).sort_values([\"soundscape_id\", \"end_time\"])[[\"row_id\"] + all_birds].reset_index(drop=True)\n\nsub.to_csv(\"submission.csv\", index=False)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F1856467ad226e02619619cf8fe25f650%2Fsub.png?generation=1712262706796504&alt=media)",
          "replies": [
            {
              "id": 2735886,
              "postDate": "2024-04-05T00:04:02.877Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2736218,
              "postDate": "2024-04-05T05:14:13.320Z",
              "content": "<p>Could you add something like this in the code:</p>\n<pre><code>sample_sub = pd.read_csv()\n\n (sample_sub) != (sub):\n     (a_variable_that_does_not_exist) \n</code></pre>\n<p>the notebook will throw exception error if the submission you are creating is not the same in length as sample_submission, which would mean either there are extra .ogg files in test_soundscapes or maybe not every .ogg file is exactly 4 minutes</p>",
              "rawMarkdown": "Could you add something like this in the code:\n\n```python\nsample_sub = pd.read_csv(\"/kaggle/input/birdclef-2024/sample_submission.csv\")\n\nif len(sample_sub) != len(sub):\n     print(a_variable_that_does_not_exist) \n```\n\nthe notebook will throw exception error if the submission you are creating is not the same in length as sample_submission, which would mean either there are extra .ogg files in test_soundscapes or maybe not every .ogg file is exactly 4 minutes",
              "votes": 1
            },
            {
              "id": 2736252,
              "postDate": "2024-04-05T05:32:02.507Z",
              "content": "<p>Thanks. I added the following code and the error is now \"notebook threw exception\".<br>\nSo probably the number of our predictions is wrong.</p>\n<pre><code> = ()\n\n\n</code></pre>",
              "rawMarkdown": "Thanks. I added the following code and the error is now \"notebook threw exception\".\nSo probably the number of our predictions is wrong.\n\n```\nsample_sub = pd.read_csv(\"/kaggle/input/birdclef-2024/sample_submission.csv\")\n\nassert(len(sample_sub) == len(sub))\n```",
              "votes": 1
            },
            {
              "id": 2736305,
              "postDate": "2024-04-05T06:01:45.753Z",
              "content": "<p>This code threw \"notebook threw exception\", so the length of submission.csv is not a multiple of 48 (=4 x 60 // 5). </p>\n<pre><code> pandas  pd\n glob  glob\n os\n numpy  np\n pathlib  Path\n\n = pd.read_csv()\n len(sample_sub) %  != :\n    # exception\n    raise\n</code></pre>\n<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nCould you please check the length of sample_submission.csv and .ogg files?</p>",
              "rawMarkdown": "This code threw \"notebook threw exception\", so the length of submission.csv is not a multiple of 48 (=4 x 60 // 5). \n\n```\nimport pandas as pd\nfrom glob import glob\nimport os\nimport numpy as np\nfrom pathlib import Path\n\nsample_sub = pd.read_csv(\"/kaggle/input/birdclef-2024/sample_submission.csv\")\nif len(sample_sub) % 48 != 0:\n    # exception\n    raise\n\n```\n\nHi @sohier \nCould you please check the length of sample_submission.csv and .ogg files?\n",
              "votes": 1
            },
            {
              "id": 2736307,
              "postDate": "2024-04-05T06:02:32.600Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2736461,
              "postDate": "2024-04-05T08:03:37.470Z",
              "content": "<p>Thanks for investigating, folks. It seems like something is off with the number of rows the scoring system is expecting, which indeed is not a multiple of 48. The audio files however, are indeed 4min long with no exception. We'll take a closer look and update ASAP. Thanks for your patience.</p>",
              "rawMarkdown": "Thanks for investigating, folks. It seems like something is off with the number of rows the scoring system is expecting, which indeed is not a multiple of 48. The audio files however, are indeed 4min long with no exception. We'll take a closer look and update ASAP. Thanks for your patience.",
              "votes": 4
            },
            {
              "id": 2736943,
              "postDate": "2024-04-05T13:39:07.457Z",
              "content": "<p>I wonder why some can still submit successfully when the scoring system expects less or more than 48x1100 rows.</p>",
              "rawMarkdown": "I wonder why some can still submit successfully when the scoring system expects less or more than 48x1100 rows."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2735579,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2024-04-04T19:24:11.127000",
      "content": "<p>I see that each file in \"unlabeled_soundscapes\" folder is in format \"xxxxxx.ogg\", </p>\n<p>Which format does the filenames in test_soundscapes folder in the hidden test use, \"soundscape_xxxxxx.ogg\" or \"xxxxxx.ogg\"?</p>\n<p>I am adapting to sample submissions but my notebook fails in the first minute of hidden test, the same notebook runs in birdclef 2023 without error, as I have run multiple debug runs using other data provided, my suspect is the filename format is changed?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2735769,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2024-04-04T22:29:58.043000",
          "content": "<p>The hidden test set uses file names like <code>soundscape_xxxxxx.ogg</code>.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2735271,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2024-04-04T16:23:54.807000",
      "content": "<p>Hello!<br>\nIt is best to adapt some sample submissions. The notebook from last year should help, as the data format is the same (though note that the species list is different):<br>\n<a href=\"https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2023\" target=\"_blank\">https://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2023</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2736105,
          "author_name": "MYSO",
          "author_url": "",
          "post_date": "2024-04-05T03:26:16.770000",
          "content": "<p>I've modified the code to adapt to BirdCLEF2024 using the provided reference. However, I'm encountering a submission scoring error. Could you please check if there's something wrong with the <a href=\"https://www.kaggle.com/code/myso1987/submission-scoring-error-submit-to-birdclef-2024/notebook\" target=\"_blank\">code</a>?</p>\n<p><a href=\"https://www.kaggle.com/code/myso1987/submission-scoring-error-submit-to-birdclef-2024/notebook\" target=\"_blank\">https://www.kaggle.com/code/myso1987/submission-scoring-error-submit-to-birdclef-2024/notebook</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2736385,
              "author_name": "MYSO",
              "author_url": "",
              "post_date": "2024-04-05T06:59:24.540000",
              "content": "<p>I've fixed the notebook and can submit it successfully. I plan to update the notebook tomorrow. Thanks for sharing the reference.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2735299,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2024-04-04T16:44:08.943000",
      "content": "<p>Sorry for the confusion. I've updated the data description.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2735665,
          "author_name": "penguin46",
          "author_url": "",
          "post_date": "2024-04-04T20:33:24.037000",
          "content": "<p></p>\n<pre><code> pandas  pd\n glob  glob\n os\n numpy  np\n pathlib  \n\n#   bird names\nall_birds = pd.read_csv(\"/kaggle/input/birdclef-2024/sample_submission.csv\").[:].tolist()\n\n#   soundscape_ids\nogg_dir = Path(\"/kaggle/input/birdclef-2024/test_soundscapes\")\n# ogg_dir = Path(\"/kaggle/input/birdclef-2024/unlabeled_soundscapes\")\nogg_files = glob(str(ogg_dir / \"*.ogg\"))\nsoundscape_ids = [Path(f).stem  f  ogg_files]\n\n# generate submission.csv\ndf = pd.DataFrame(np.zeros((len(soundscape_ids), len(all_birds))), =all_birds, )\ndf[\"soundscape_id\"] = soundscape_ids\n\ndfs = []\n end_time  range(, , ):\n    this_df = df.()\n    this_df[\"end_time\"] = end_time\n    this_df[\"row_id\"] = \"soundscape_\" + this_df[\"soundscape_id\"] + \"_\" + this_df[\"end_time\"].astype(str)\n    dfs.append(this_df)\nsub = pd.concat(dfs).sort_values([\"soundscape_id\", \"end_time\"])[[\"row_id\"] + all_birds].reset_index(=)\n\nsub.to_csv(\"submission.csv\", =)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F1856467ad226e02619619cf8fe25f650%2Fsub.png?generation=1712262706796504&amp;alt=media\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2735886,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-05T00:04:02.877000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2736218,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2024-04-05T05:14:13.320000",
              "content": "<p>Could you add something like this in the code:</p>\n<pre><code>sample_sub = pd.read_csv()\n\n (sample_sub) != (sub):\n     (a_variable_that_does_not_exist) \n</code></pre>\n<p>the notebook will throw exception error if the submission you are creating is not the same in length as sample_submission, which would mean either there are extra .ogg files in test_soundscapes or maybe not every .ogg file is exactly 4 minutes</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2736252,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-04-05T05:32:02.507000",
              "content": "<p>Thanks. I added the following code and the error is now \"notebook threw exception\".<br>\nSo probably the number of our predictions is wrong.</p>\n<pre><code> = ()\n\n\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2736305,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-04-05T06:01:45.753000",
              "content": "<p>This code threw \"notebook threw exception\", so the length of submission.csv is not a multiple of 48 (=4 x 60 // 5). </p>\n<pre><code> pandas  pd\n glob  glob\n os\n numpy  np\n pathlib  Path\n\n = pd.read_csv()\n len(sample_sub) %  != :\n    # exception\n    raise\n</code></pre>\n<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nCould you please check the length of sample_submission.csv and .ogg files?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2736307,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-04-05T06:02:32.600000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2736461,
              "author_name": "Stefan Kahl",
              "author_url": "",
              "post_date": "2024-04-05T08:03:37.470000",
              "content": "<p>Thanks for investigating, folks. It seems like something is off with the number of rows the scoring system is expecting, which indeed is not a multiple of 48. The audio files however, are indeed 4min long with no exception. We'll take a closer look and update ASAP. Thanks for your patience.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2736943,
              "author_name": "Hicham Bellafkir",
              "author_url": "",
              "post_date": "2024-04-05T13:39:07.457000",
              "content": "<p>I wonder why some can still submit successfully when the scoring system expects less or more than 48x1100 rows.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2734328": "The `row_id` in sample_submission.csv is described as below, but what is `soundscape_id`?\n\n> row_id: A slug of [soundscape_id]_[end_time] for the prediction.\n\nDoes the test_soundscapes directory has files named like \"soundscape_1446779.ogg\", for which the corresponding `soundscape_id` is \"soundscape_1446779\"? Or is the test file names like \"1446779.ogg\" and the `soundscape_id` is \"soundscape_[file_name]\"?\n\n---\n\nSolved: The dataset description has been updated and confirmed that the submission passes correctly. Thank you.",
    "2735579": "I see that each file in \"unlabeled_soundscapes\" folder is in format \"xxxxxx.ogg\", \n\nWhich format does the filenames in test_soundscapes folder in the hidden test use, \"soundscape_xxxxxx.ogg\" or \"xxxxxx.ogg\"?\n\nI am adapting to sample submissions but my notebook fails in the first minute of hidden test, the same notebook runs in birdclef 2023 without error, as I have run multiple debug runs using other data provided, my suspect is the filename format is changed?",
    "2735271": "Hello!\nIt is best to adapt some sample submissions. The notebook from last year should help, as the data format is the same (though note that the species list is different):\nhttps://www.kaggle.com/code/stefankahl/how-to-submit-to-birdclef-2023",
    "2735299": "Sorry for the confusion. I've updated the data description."
  }
}