{
  "id": 323574,
  "title": "Competition dataset divided into sub clips",
  "url": "/competitions/birdclef-2022/discussion/323574",
  "author_name": "",
  "post_date": "2022-05-07T07:13:15.841203600Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I would like to present three datasets that I have recently released to the public[1-3].<br>\nThese data sets are a split of the main competition data set into sub-clips of 5, 10 and 60 seconds, respectively.</p>\n<p>I made this because the original dataset had very small sample sizes for several species, making it difficult to cut cross-validation folds.</p>\n<p>The created dataset contains metadata and audio files. The metadata column includes the file names of the original sound clips, so you can merge them with the metadata of this competition for further detailed data analysis and processing.</p>\n<p>These datasets are licensed under CC BY-NC-SA 4.0. Please note the license restrictions when reusing the data.</p>\n<h1>Reference</h1>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-5-sec\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-5-sec</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-10-sec\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-10-sec</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-60-sec\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-60-sec</a></li>\n</ul>",
  "messages": [
    {
      "id": "1780182",
      "postDate": "05/07/2022 07:13:15",
      "content": "<p>I would like to present three datasets that I have recently released to the public[1-3].<br>\nThese data sets are a split of the main competition data set into sub-clips of 5, 10 and 60 seconds, respectively.</p>\n<p>I made this because the original dataset had very small sample sizes for several species, making it difficult to cut cross-validation folds.</p>\n<p>The created dataset contains metadata and audio files. The metadata column includes the file names of the original sound clips, so you can merge them with the metadata of this competition for further detailed data analysis and processing.</p>\n<p>These datasets are licensed under CC BY-NC-SA 4.0. Please note the license restrictions when reusing the data.</p>\n<h1>Reference</h1>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-5-sec\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-5-sec</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-10-sec\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-10-sec</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-60-sec\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-60-sec</a></li>\n</ul>",
      "rawMarkdown": "I would like to present three datasets that I have recently released to the public[1-3].\nThese data sets are a split of the main competition data set into sub-clips of 5, 10 and 60 seconds, respectively.\n\nI made this because the original dataset had very small sample sizes for several species, making it difficult to cut cross-validation folds.\n\nThe created dataset contains metadata and audio files. The metadata column includes the file names of the original sound clips, so you can merge them with the metadata of this competition for further detailed data analysis and processing.\n\nThese datasets are licensed under CC BY-NC-SA 4.0. Please note the license restrictions when reusing the data.\n\n# Reference\n\n- [1] https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-5-sec\n- [2] https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-10-sec\n- [3] https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-60-sec",
      "votes": null
    },
    {
      "id": "1780190",
      "postDate": "05/07/2022 07:22:32",
      "content": "<p>Note: the number of sub-clips for the species to be scored in these data sets is as follows:</p>\n<h3>sub-clip in 5 sec or less</h3>\n<table>\n<thead>\n<tr>\n<th>primary_label</th>\n<th>num_subclips</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>skylar</td>\n<td>5796</td>\n</tr>\n<tr>\n<td>houfin</td>\n<td>3582</td>\n</tr>\n<tr>\n<td>jabwar</td>\n<td>881</td>\n</tr>\n<tr>\n<td>warwhe1</td>\n<td>571</td>\n</tr>\n<tr>\n<td>apapan</td>\n<td>540</td>\n</tr>\n<tr>\n<td>hawcre</td>\n<td>504</td>\n</tr>\n<tr>\n<td>yefcan</td>\n<td>494</td>\n</tr>\n<tr>\n<td>iiwi</td>\n<td>455</td>\n</tr>\n<tr>\n<td>omao</td>\n<td>245</td>\n</tr>\n<tr>\n<td>barpet</td>\n<td>210</td>\n</tr>\n<tr>\n<td>hawama</td>\n<td>161</td>\n</tr>\n<tr>\n<td>akiapo</td>\n<td>159</td>\n</tr>\n<tr>\n<td>elepai</td>\n<td>145</td>\n</tr>\n<tr>\n<td>aniani</td>\n<td>95</td>\n</tr>\n<tr>\n<td>hawhaw</td>\n<td>55</td>\n</tr>\n<tr>\n<td>hawgoo</td>\n<td>50</td>\n</tr>\n<tr>\n<td>maupar</td>\n<td>50</td>\n</tr>\n<tr>\n<td>hawpet1</td>\n<td>35</td>\n</tr>\n<tr>\n<td>crehon</td>\n<td>25</td>\n</tr>\n<tr>\n<td>ercfra</td>\n<td>17</td>\n</tr>\n<tr>\n<td>puaioh</td>\n<td>10</td>\n</tr>\n</tbody>\n</table>\n<h3>sub-clip in 60 sec or less</h3>\n<table>\n<thead>\n<tr>\n<th>primary_label</th>\n<th>num_subclips</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>skylar</td>\n<td>758</td>\n</tr>\n<tr>\n<td>houfin</td>\n<td>465</td>\n</tr>\n<tr>\n<td>jabwar</td>\n<td>114</td>\n</tr>\n<tr>\n<td>warwhe1</td>\n<td>87</td>\n</tr>\n<tr>\n<td>yefcan</td>\n<td>76</td>\n</tr>\n<tr>\n<td>apapan</td>\n<td>68</td>\n</tr>\n<tr>\n<td>iiwi</td>\n<td>55</td>\n</tr>\n<tr>\n<td>hawcre</td>\n<td>51</td>\n</tr>\n<tr>\n<td>omao</td>\n<td>31</td>\n</tr>\n<tr>\n<td>hawama</td>\n<td>26</td>\n</tr>\n<tr>\n<td>barpet</td>\n<td>25</td>\n</tr>\n<tr>\n<td>akiapo</td>\n<td>21</td>\n</tr>\n<tr>\n<td>elepai</td>\n<td>20</td>\n</tr>\n<tr>\n<td>aniani</td>\n<td>15</td>\n</tr>\n<tr>\n<td>hawgoo</td>\n<td>9</td>\n</tr>\n<tr>\n<td>hawhaw</td>\n<td>7</td>\n</tr>\n<tr>\n<td>ercfra</td>\n<td>6</td>\n</tr>\n<tr>\n<td>maupar</td>\n<td>5</td>\n</tr>\n<tr>\n<td>hawpet1</td>\n<td>4</td>\n</tr>\n<tr>\n<td>puaioh</td>\n<td>3</td>\n</tr>\n<tr>\n<td>crehon</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Note: the number of sub-clips for the species to be scored in these data sets is as follows:\n\n### sub-clip in 5 sec or less\n\n| primary_label | num_subclips |\n|---------------|--------------|\n| skylar        | 5796         |\n| houfin        | 3582         |\n| jabwar        | 881          |\n| warwhe1       | 571          |\n| apapan        | 540          |\n| hawcre        | 504          |\n| yefcan        | 494          |\n| iiwi          | 455          |\n| omao          | 245          |\n| barpet        | 210          |\n| hawama        | 161          |\n| akiapo        | 159          |\n| elepai        | 145          |\n| aniani        | 95           |\n| hawhaw        | 55           |\n| hawgoo        | 50           |\n| maupar        | 50           |\n| hawpet1       | 35           |\n| crehon        | 25           |\n| ercfra        | 17           |\n| puaioh        | 10           |\n\n### sub-clip in 60 sec or less\n\n| primary_label | num_subclips |\n|:-------------:|:------------:|\n| skylar        |          758 |\n| houfin        |          465 |\n| jabwar        |          114 |\n| warwhe1       |           87 |\n| yefcan        |           76 |\n| apapan        |           68 |\n| iiwi          |           55 |\n| hawcre        |           51 |\n| omao          |           31 |\n| hawama        |           26 |\n| barpet        |           25 |\n| akiapo        |           21 |\n| elepai        |           20 |\n| aniani        |           15 |\n| hawgoo        |            9 |\n| hawhaw        |            7 |\n| ercfra        |            6 |\n| maupar        |            5 |\n| hawpet1       |            4 |\n| puaioh        |            3 |\n| crehon        |            3 |",
      "votes": null
    },
    {
      "id": "1780246",
      "postDate": "05/07/2022 08:26:54",
      "content": "<p>These datasets look really useful! I'm glad you released them to the public. Thanks for sharing!</p>",
      "rawMarkdown": "These datasets look really useful! I'm glad you released them to the public. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1780393",
      "postDate": "05/07/2022 11:20:54",
      "content": "<p>Note: I also shared dataset of 10 sec sub-clips.</p>",
      "rawMarkdown": "Note: I also shared dataset of 10 sec sub-clips.",
      "votes": null
    },
    {
      "id": "1782800",
      "postDate": "05/09/2022 21:34:59",
      "content": "<p>Note that the same clip is split into multiple clips in this data set, so the score is likely to be higher in local validation. (Call features within the same recording tend to be very similar.)</p>",
      "rawMarkdown": "Note that the same clip is split into multiple clips in this data set, so the score is likely to be higher in local validation. (Call features within the same recording tend to be very similar.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1780190,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/07/2022 07:22:32",
      "content": "<p>Note: the number of sub-clips for the species to be scored in these data sets is as follows:</p>\n<h3>sub-clip in 5 sec or less</h3>\n<table>\n<thead>\n<tr>\n<th>primary_label</th>\n<th>num_subclips</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>skylar</td>\n<td>5796</td>\n</tr>\n<tr>\n<td>houfin</td>\n<td>3582</td>\n</tr>\n<tr>\n<td>jabwar</td>\n<td>881</td>\n</tr>\n<tr>\n<td>warwhe1</td>\n<td>571</td>\n</tr>\n<tr>\n<td>apapan</td>\n<td>540</td>\n</tr>\n<tr>\n<td>hawcre</td>\n<td>504</td>\n</tr>\n<tr>\n<td>yefcan</td>\n<td>494</td>\n</tr>\n<tr>\n<td>iiwi</td>\n<td>455</td>\n</tr>\n<tr>\n<td>omao</td>\n<td>245</td>\n</tr>\n<tr>\n<td>barpet</td>\n<td>210</td>\n</tr>\n<tr>\n<td>hawama</td>\n<td>161</td>\n</tr>\n<tr>\n<td>akiapo</td>\n<td>159</td>\n</tr>\n<tr>\n<td>elepai</td>\n<td>145</td>\n</tr>\n<tr>\n<td>aniani</td>\n<td>95</td>\n</tr>\n<tr>\n<td>hawhaw</td>\n<td>55</td>\n</tr>\n<tr>\n<td>hawgoo</td>\n<td>50</td>\n</tr>\n<tr>\n<td>maupar</td>\n<td>50</td>\n</tr>\n<tr>\n<td>hawpet1</td>\n<td>35</td>\n</tr>\n<tr>\n<td>crehon</td>\n<td>25</td>\n</tr>\n<tr>\n<td>ercfra</td>\n<td>17</td>\n</tr>\n<tr>\n<td>puaioh</td>\n<td>10</td>\n</tr>\n</tbody>\n</table>\n<h3>sub-clip in 60 sec or less</h3>\n<table>\n<thead>\n<tr>\n<th>primary_label</th>\n<th>num_subclips</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>skylar</td>\n<td>758</td>\n</tr>\n<tr>\n<td>houfin</td>\n<td>465</td>\n</tr>\n<tr>\n<td>jabwar</td>\n<td>114</td>\n</tr>\n<tr>\n<td>warwhe1</td>\n<td>87</td>\n</tr>\n<tr>\n<td>yefcan</td>\n<td>76</td>\n</tr>\n<tr>\n<td>apapan</td>\n<td>68</td>\n</tr>\n<tr>\n<td>iiwi</td>\n<td>55</td>\n</tr>\n<tr>\n<td>hawcre</td>\n<td>51</td>\n</tr>\n<tr>\n<td>omao</td>\n<td>31</td>\n</tr>\n<tr>\n<td>hawama</td>\n<td>26</td>\n</tr>\n<tr>\n<td>barpet</td>\n<td>25</td>\n</tr>\n<tr>\n<td>akiapo</td>\n<td>21</td>\n</tr>\n<tr>\n<td>elepai</td>\n<td>20</td>\n</tr>\n<tr>\n<td>aniani</td>\n<td>15</td>\n</tr>\n<tr>\n<td>hawgoo</td>\n<td>9</td>\n</tr>\n<tr>\n<td>hawhaw</td>\n<td>7</td>\n</tr>\n<tr>\n<td>ercfra</td>\n<td>6</td>\n</tr>\n<tr>\n<td>maupar</td>\n<td>5</td>\n</tr>\n<tr>\n<td>hawpet1</td>\n<td>4</td>\n</tr>\n<tr>\n<td>puaioh</td>\n<td>3</td>\n</tr>\n<tr>\n<td>crehon</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1780246,
      "author_name": "",
      "author_url": "",
      "post_date": "05/07/2022 08:26:54",
      "content": "<p>These datasets look really useful! I'm glad you released them to the public. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1780393,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/07/2022 11:20:54",
      "content": "<p>Note: I also shared dataset of 10 sec sub-clips.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1782800,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/09/2022 21:34:59",
      "content": "<p>Note that the same clip is split into multiple clips in this data set, so the score is likely to be higher in local validation. (Call features within the same recording tend to be very similar.)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1780182": "I would like to present three datasets that I have recently released to the public[1-3].\nThese data sets are a split of the main competition data set into sub-clips of 5, 10 and 60 seconds, respectively.\n\nI made this because the original dataset had very small sample sizes for several species, making it difficult to cut cross-validation folds.\n\nThe created dataset contains metadata and audio files. The metadata column includes the file names of the original sound clips, so you can merge them with the metadata of this competition for further detailed data analysis and processing.\n\nThese datasets are licensed under CC BY-NC-SA 4.0. Please note the license restrictions when reusing the data.\n\n# Reference\n\n- [1] https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-5-sec\n- [2] https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-10-sec\n- [3] https://www.kaggle.com/datasets/tatamikenn/birdclef-2022-subclip-60-sec",
    "1780190": "Note: the number of sub-clips for the species to be scored in these data sets is as follows:\n\n### sub-clip in 5 sec or less\n\n| primary_label | num_subclips |\n|---------------|--------------|\n| skylar        | 5796         |\n| houfin        | 3582         |\n| jabwar        | 881          |\n| warwhe1       | 571          |\n| apapan        | 540          |\n| hawcre        | 504          |\n| yefcan        | 494          |\n| iiwi          | 455          |\n| omao          | 245          |\n| barpet        | 210          |\n| hawama        | 161          |\n| akiapo        | 159          |\n| elepai        | 145          |\n| aniani        | 95           |\n| hawhaw        | 55           |\n| hawgoo        | 50           |\n| maupar        | 50           |\n| hawpet1       | 35           |\n| crehon        | 25           |\n| ercfra        | 17           |\n| puaioh        | 10           |\n\n### sub-clip in 60 sec or less\n\n| primary_label | num_subclips |\n|:-------------:|:------------:|\n| skylar        |          758 |\n| houfin        |          465 |\n| jabwar        |          114 |\n| warwhe1       |           87 |\n| yefcan        |           76 |\n| apapan        |           68 |\n| iiwi          |           55 |\n| hawcre        |           51 |\n| omao          |           31 |\n| hawama        |           26 |\n| barpet        |           25 |\n| akiapo        |           21 |\n| elepai        |           20 |\n| aniani        |           15 |\n| hawgoo        |            9 |\n| hawhaw        |            7 |\n| ercfra        |            6 |\n| maupar        |            5 |\n| hawpet1       |            4 |\n| puaioh        |            3 |\n| crehon        |            3 |",
    "1780246": "These datasets look really useful! I'm glad you released them to the public. Thanks for sharing!",
    "1780393": "Note: I also shared dataset of 10 sec sub-clips.",
    "1782800": "Note that the same clip is split into multiple clips in this data set, so the score is likely to be higher in local validation. (Call features within the same recording tend to be very similar.)"
  },
  "source": "meta"
}