{
  "id": 236799,
  "title": "Some of the short audio isn't actually short",
  "url": "/competitions/birdclef-2021/discussion/236799",
  "author_name": "",
  "post_date": "2021-05-05T21:26:39.884584400Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm just getting into this competition and was looking at the clip lengths and gathered the following statistics:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>count</th>\n<th>mean</th>\n<th>std</th>\n<th>min</th>\n<th>25%</th>\n<th>50%</th>\n<th>75%</th>\n<th>max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>clip_length</td>\n<td>62874</td>\n<td>56.2553</td>\n<td>74.0424</td>\n<td>5.95812</td>\n<td>18.3783</td>\n<td>34.2607</td>\n<td>66.205</td>\n<td>2745.35</td>\n</tr>\n</tbody>\n</table>\n<p>It turns out that some of the clips are quite long with the longest being up to 2745 seconds (45 minutes?)</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>file</th>\n<th>folder</th>\n<th>sample_rate</th>\n<th>clip_length</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>45169</td>\n<td>XC478859.ogg</td>\n<td>rudtur</td>\n<td>32000</td>\n<td>1976.91</td>\n</tr>\n<tr>\n<td>25013</td>\n<td>XC246425.ogg</td>\n<td>comrav</td>\n<td>32000</td>\n<td>2283.9</td>\n</tr>\n<tr>\n<td>19108</td>\n<td>XC310358.ogg</td>\n<td>eursta</td>\n<td>32000</td>\n<td>2354.64</td>\n</tr>\n<tr>\n<td>32153</td>\n<td>XC147860.ogg</td>\n<td>blbthr1</td>\n<td>32000</td>\n<td>2739.66</td>\n</tr>\n<tr>\n<td>44232</td>\n<td>XC244537.ogg</td>\n<td>whevir</td>\n<td>32000</td>\n<td>2745.35</td>\n</tr>\n</tbody>\n</table>\n<p>Obviously, the longer clips are more likely to have multiple bird calls in them. There are 17881 clips longer than 60 seconds.</p>\n<p>What sort of methods are people using to handle the differing clip lengths? Should we ignore the unusually long clips since we don't know which part is the target sound? Or should we randomly chop these clips into smaller ones?</p>\n<p>Let's say I limit to using clips between 0 - 60 seconds. What other methods are available other than padding the short clips to match the 60-second clips?</p>",
  "messages": [
    {
      "id": "1294741",
      "postDate": "05/05/2021 21:26:39",
      "content": "<p>I'm just getting into this competition and was looking at the clip lengths and gathered the following statistics:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>count</th>\n<th>mean</th>\n<th>std</th>\n<th>min</th>\n<th>25%</th>\n<th>50%</th>\n<th>75%</th>\n<th>max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>clip_length</td>\n<td>62874</td>\n<td>56.2553</td>\n<td>74.0424</td>\n<td>5.95812</td>\n<td>18.3783</td>\n<td>34.2607</td>\n<td>66.205</td>\n<td>2745.35</td>\n</tr>\n</tbody>\n</table>\n<p>It turns out that some of the clips are quite long with the longest being up to 2745 seconds (45 minutes?)</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>file</th>\n<th>folder</th>\n<th>sample_rate</th>\n<th>clip_length</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>45169</td>\n<td>XC478859.ogg</td>\n<td>rudtur</td>\n<td>32000</td>\n<td>1976.91</td>\n</tr>\n<tr>\n<td>25013</td>\n<td>XC246425.ogg</td>\n<td>comrav</td>\n<td>32000</td>\n<td>2283.9</td>\n</tr>\n<tr>\n<td>19108</td>\n<td>XC310358.ogg</td>\n<td>eursta</td>\n<td>32000</td>\n<td>2354.64</td>\n</tr>\n<tr>\n<td>32153</td>\n<td>XC147860.ogg</td>\n<td>blbthr1</td>\n<td>32000</td>\n<td>2739.66</td>\n</tr>\n<tr>\n<td>44232</td>\n<td>XC244537.ogg</td>\n<td>whevir</td>\n<td>32000</td>\n<td>2745.35</td>\n</tr>\n</tbody>\n</table>\n<p>Obviously, the longer clips are more likely to have multiple bird calls in them. There are 17881 clips longer than 60 seconds.</p>\n<p>What sort of methods are people using to handle the differing clip lengths? Should we ignore the unusually long clips since we don't know which part is the target sound? Or should we randomly chop these clips into smaller ones?</p>\n<p>Let's say I limit to using clips between 0 - 60 seconds. What other methods are available other than padding the short clips to match the 60-second clips?</p>",
      "rawMarkdown": "I'm just getting into this competition and was looking at the clip lengths and gathered the following statistics:\n\n|             |   count |    mean |     std |     min |     25% |     50% |    75% |     max |\n|:------------|--------:|--------:|--------:|--------:|--------:|--------:|-------:|--------:|\n| clip_length |   62874 | 56.2553 | 74.0424 | 5.95812 | 18.3783 | 34.2607 | 66.205 | 2745.35 |\n\nIt turns out that some of the clips are quite long with the longest being up to 2745 seconds (45 minutes?)\n\n|       | file         | folder   |   sample_rate |   clip_length |\n|------:|:-------------|:---------|--------------:|--------------:|\n| 45169 | XC478859.ogg | rudtur   |         32000 |       1976.91 |\n| 25013 | XC246425.ogg | comrav   |         32000 |       2283.9  |\n| 19108 | XC310358.ogg | eursta   |         32000 |       2354.64 |\n| 32153 | XC147860.ogg | blbthr1  |         32000 |       2739.66 |\n| 44232 | XC244537.ogg | whevir   |         32000 |       2745.35 |\n\nObviously, the longer clips are more likely to have multiple bird calls in them. There are 17881 clips longer than 60 seconds.\n\nWhat sort of methods are people using to handle the differing clip lengths? Should we ignore the unusually long clips since we don't know which part is the target sound? Or should we randomly chop these clips into smaller ones?\n\nLet's say I limit to using clips between 0 - 60 seconds. What other methods are available other than padding the short clips to match the 60-second clips?",
      "votes": null
    },
    {
      "id": "1294750",
      "postDate": "05/05/2021 21:45:41",
      "content": "<blockquote>\n  <p>Should we ignore the unusually long clips since we don't know which part is the target sound?</p>\n</blockquote>\n<p>I think it becomes quite obvious once you compare recordings of the same primary_label 🙂</p>",
      "rawMarkdown": "> Should we ignore the unusually long clips since we don't know which part is the target sound?\n\nI think it becomes quite obvious once you compare recordings of the same primary_label 🙂",
      "votes": null
    },
    {
      "id": "1295955",
      "postDate": "05/06/2021 20:34:46",
      "content": "<p>By the kernel i study.<br>\nit can be one solution that, for example, split audio by 7sec and make list of it.<br>\nso if you have 100sec audio file, it can be a list of 15 items.<br>\nand you can pick items from list by your own methods.</p>",
      "rawMarkdown": "By the kernel i study.\nit can be one solution that, for example, split audio by 7sec and make list of it.\nso if you have 100sec audio file, it can be a list of 15 items.\nand you can pick items from list by your own methods.",
      "votes": null
    },
    {
      "id": "1296463",
      "postDate": "05/07/2021 09:20:20",
      "content": "<p>I think you're right and I'm totally overthinking this. Time for less thinking and more coding! 🙂</p>",
      "rawMarkdown": "I think you're right and I'm totally overthinking this. Time for less thinking and more coding! 🙂",
      "votes": null
    },
    {
      "id": "1309318",
      "postDate": "05/15/2021 20:13:43",
      "content": "<p>I haven't gotten there yet, but my (kind of dumb) solution to this was just going to be using random cropping resize on the spectrogram images in my dataloader.</p>",
      "rawMarkdown": "I haven't gotten there yet, but my (kind of dumb) solution to this was just going to be using random cropping resize on the spectrogram images in my dataloader.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1294750,
      "author_name": "shtrausslearning",
      "author_url": "",
      "post_date": "05/05/2021 21:45:41",
      "content": "<blockquote>\n  <p>Should we ignore the unusually long clips since we don't know which part is the target sound?</p>\n</blockquote>\n<p>I think it becomes quite obvious once you compare recordings of the same primary_label 🙂</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1295955,
      "author_name": "kangyunho",
      "author_url": "",
      "post_date": "05/06/2021 20:34:46",
      "content": "<p>By the kernel i study.<br>\nit can be one solution that, for example, split audio by 7sec and make list of it.<br>\nso if you have 100sec audio file, it can be a list of 15 items.<br>\nand you can pick items from list by your own methods.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1296463,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "05/07/2021 09:20:20",
          "content": "<p>I think you're right and I'm totally overthinking this. Time for less thinking and more coding! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1309318,
      "author_name": "aaronmead",
      "author_url": "",
      "post_date": "05/15/2021 20:13:43",
      "content": "<p>I haven't gotten there yet, but my (kind of dumb) solution to this was just going to be using random cropping resize on the spectrogram images in my dataloader.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1294741": "I'm just getting into this competition and was looking at the clip lengths and gathered the following statistics:\n\n|             |   count |    mean |     std |     min |     25% |     50% |    75% |     max |\n|:------------|--------:|--------:|--------:|--------:|--------:|--------:|-------:|--------:|\n| clip_length |   62874 | 56.2553 | 74.0424 | 5.95812 | 18.3783 | 34.2607 | 66.205 | 2745.35 |\n\nIt turns out that some of the clips are quite long with the longest being up to 2745 seconds (45 minutes?)\n\n|       | file         | folder   |   sample_rate |   clip_length |\n|------:|:-------------|:---------|--------------:|--------------:|\n| 45169 | XC478859.ogg | rudtur   |         32000 |       1976.91 |\n| 25013 | XC246425.ogg | comrav   |         32000 |       2283.9  |\n| 19108 | XC310358.ogg | eursta   |         32000 |       2354.64 |\n| 32153 | XC147860.ogg | blbthr1  |         32000 |       2739.66 |\n| 44232 | XC244537.ogg | whevir   |         32000 |       2745.35 |\n\nObviously, the longer clips are more likely to have multiple bird calls in them. There are 17881 clips longer than 60 seconds.\n\nWhat sort of methods are people using to handle the differing clip lengths? Should we ignore the unusually long clips since we don't know which part is the target sound? Or should we randomly chop these clips into smaller ones?\n\nLet's say I limit to using clips between 0 - 60 seconds. What other methods are available other than padding the short clips to match the 60-second clips?",
    "1294750": "> Should we ignore the unusually long clips since we don't know which part is the target sound?\n\nI think it becomes quite obvious once you compare recordings of the same primary_label 🙂",
    "1295955": "By the kernel i study.\nit can be one solution that, for example, split audio by 7sec and make list of it.\nso if you have 100sec audio file, it can be a list of 15 items.\nand you can pick items from list by your own methods.",
    "1296463": "I think you're right and I'm totally overthinking this. Time for less thinking and more coding! 🙂",
    "1309318": "I haven't gotten there yet, but my (kind of dumb) solution to this was just going to be using random cropping resize on the spectrogram images in my dataloader."
  },
  "source": "meta"
}