{
  "id": 312355,
  "title": "These files are probably corrupted",
  "url": "/competitions/birdclef-2022/discussion/312355",
  "author_name": "",
  "post_date": "2022-03-11T15:06:39.676905200Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The files are:</p>\n<table>\n<thead>\n<tr>\n<th>filename</th>\n<th>length</th>\n<th>link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>blkfra/XC649198.ogg</td>\n<td>0.034844s</td>\n<td><a href=\"https://www.xeno-canto.org/649198\" target=\"_blank\">https://www.xeno-canto.org/649198</a></td>\n</tr>\n<tr>\n<td>normoc/XC150238.ogg</td>\n<td>0.104500s</td>\n<td><a href=\"https://www.xeno-canto.org/150238\" target=\"_blank\">https://www.xeno-canto.org/150238</a></td>\n</tr>\n</tbody>\n</table>\n<p>I listened to them, and they don't make any sense. But the recordings are good on the xeno-canto website. So probably they are only corrupted here?</p>\n<p>By the way, I have removed the data folder and re-added, things didn't change.</p>",
  "messages": [
    {
      "id": "1719207",
      "postDate": "03/11/2022 15:06:39",
      "content": "<p>The files are:</p>\n<table>\n<thead>\n<tr>\n<th>filename</th>\n<th>length</th>\n<th>link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>blkfra/XC649198.ogg</td>\n<td>0.034844s</td>\n<td><a href=\"https://www.xeno-canto.org/649198\" target=\"_blank\">https://www.xeno-canto.org/649198</a></td>\n</tr>\n<tr>\n<td>normoc/XC150238.ogg</td>\n<td>0.104500s</td>\n<td><a href=\"https://www.xeno-canto.org/150238\" target=\"_blank\">https://www.xeno-canto.org/150238</a></td>\n</tr>\n</tbody>\n</table>\n<p>I listened to them, and they don't make any sense. But the recordings are good on the xeno-canto website. So probably they are only corrupted here?</p>\n<p>By the way, I have removed the data folder and re-added, things didn't change.</p>",
      "rawMarkdown": "The files are:\n| filename | length  | link|\n| --- | --- | --- |\n| blkfra/XC649198.ogg |0.034844s  |https://www.xeno-canto.org/649198 |\n| normoc/XC150238.ogg |0.104500s |https://www.xeno-canto.org/150238 |\n\nI listened to them, and they don't make any sense. But the recordings are good on the xeno-canto website. So probably they are only corrupted here?\n\nBy the way, I have removed the data folder and re-added, things didn't change.",
      "votes": null
    },
    {
      "id": "1721913",
      "postDate": "03/14/2022 05:49:14",
      "content": "<p>Thanks for pointing this out, I noticed the same issue on those two files in my data too. I'm now trying to go through the 14k audio files to find such anomalies so I can exclude them from the list. The CPU / memory seem to be maxing out when I run my code so trying to work on that. Once I have that in, happy to post it for others to use :)  </p>",
      "rawMarkdown": "Thanks for pointing this out, I noticed the same issue on those two files in my data too. I'm now trying to go through the 14k audio files to find such anomalies so I can exclude them from the list. The CPU / memory seem to be maxing out when I run my code so trying to work on that. Once I have that in, happy to post it for others to use :)",
      "votes": null
    },
    {
      "id": "1723222",
      "postDate": "03/15/2022 08:33:40",
      "content": "<p>that's nice, looking forward to it</p>",
      "rawMarkdown": "that's nice, looking forward to it",
      "votes": null
    },
    {
      "id": "1723868",
      "postDate": "03/15/2022 20:06:31",
      "content": "<p>Okay, so scanned all 14,852 files and the only two files corrupt (lengths of under 1 sec) where the ones you pointed out already 😀 <br>\nSo here goes the code for parallel execution</p>\n<pre><code># Import libraries\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.auto import tqdm\nimport pandas as pd\n\n# Load the data into the dataframe\ndf_birdclef = pd.read_csv('/kaggle/input/birdclef-2022/train_metadata.csv')\n\n# Function to get size of the audio\ndef get_length(fn):\n    fp = os.path.join(path, \"train_audio\", fn)\n    waveform, sample_rate = torchaudio.load(fp)\n    return waveform.size()[-1]\n\n# Run the jobs in parallel to get length and add to list called 'sizes'\nsizes = Parallel(n_jobs=os.cpu_count())(delayed(get_length)(fn) for fn in tqdm(df_birdclef[\"filename\"]))\n\n# Check if any sizes of audio less than 1 sec\ndf_sizes = pd.DataFrame(sizes)\ndf_sizes[df_sizes[0]&lt;10000]\n</code></pre>\n<p>Note - the code above doesn't give the filename but it can be modified to give you that. I just left it at this.</p>\n<p>Credit: The code snippet above was heavily used from <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> 's notebook mentioned below<br>\n<a href=\"https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce</a></p>",
      "rawMarkdown": "Okay, so scanned all 14,852 files and the only two files corrupt (lengths of under 1 sec) where the ones you pointed out already 😀 \nSo here goes the code for parallel execution\n\n```\n# Import libraries\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.auto import tqdm\nimport pandas as pd\n\n# Load the data into the dataframe\ndf_birdclef = pd.read_csv('/kaggle/input/birdclef-2022/train_metadata.csv')\n\n# Function to get size of the audio\ndef get_length(fn):\n    fp = os.path.join(path, \"train_audio\", fn)\n    waveform, sample_rate = torchaudio.load(fp)\n    return waveform.size()[-1]\n\n# Run the jobs in parallel to get length and add to list called 'sizes'\nsizes = Parallel(n_jobs=os.cpu_count())(delayed(get_length)(fn) for fn in tqdm(df_birdclef[\"filename\"]))\n\n# Check if any sizes of audio less than 1 sec\ndf_sizes = pd.DataFrame(sizes)\ndf_sizes[df_sizes[0]<10000]\n```\n\nNote - the code above doesn't give the filename but it can be modified to give you that. I just left it at this.\n\nCredit: The code snippet above was heavily used from @jirkaborovec 's notebook mentioned below\nhttps://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce",
      "votes": null
    },
    {
      "id": "1725397",
      "postDate": "03/17/2022 04:15:55",
      "content": "<p>Thanks for sharing this.</p>",
      "rawMarkdown": "Thanks for sharing this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1721913,
      "author_name": "hadicurtay",
      "author_url": "",
      "post_date": "03/14/2022 05:49:14",
      "content": "<p>Thanks for pointing this out, I noticed the same issue on those two files in my data too. I'm now trying to go through the 14k audio files to find such anomalies so I can exclude them from the list. The CPU / memory seem to be maxing out when I run my code so trying to work on that. Once I have that in, happy to post it for others to use :)  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1723222,
          "author_name": "chrisqiu",
          "author_url": "",
          "post_date": "03/15/2022 08:33:40",
          "content": "<p>that's nice, looking forward to it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1723868,
          "author_name": "hadicurtay",
          "author_url": "",
          "post_date": "03/15/2022 20:06:31",
          "content": "<p>Okay, so scanned all 14,852 files and the only two files corrupt (lengths of under 1 sec) where the ones you pointed out already 😀 <br>\nSo here goes the code for parallel execution</p>\n<pre><code># Import libraries\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.auto import tqdm\nimport pandas as pd\n\n# Load the data into the dataframe\ndf_birdclef = pd.read_csv('/kaggle/input/birdclef-2022/train_metadata.csv')\n\n# Function to get size of the audio\ndef get_length(fn):\n    fp = os.path.join(path, \"train_audio\", fn)\n    waveform, sample_rate = torchaudio.load(fp)\n    return waveform.size()[-1]\n\n# Run the jobs in parallel to get length and add to list called 'sizes'\nsizes = Parallel(n_jobs=os.cpu_count())(delayed(get_length)(fn) for fn in tqdm(df_birdclef[\"filename\"]))\n\n# Check if any sizes of audio less than 1 sec\ndf_sizes = pd.DataFrame(sizes)\ndf_sizes[df_sizes[0]&lt;10000]\n</code></pre>\n<p>Note - the code above doesn't give the filename but it can be modified to give you that. I just left it at this.</p>\n<p>Credit: The code snippet above was heavily used from <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> 's notebook mentioned below<br>\n<a href=\"https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce\" target=\"_blank\">https://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725397,
          "author_name": "chrisqiu",
          "author_url": "",
          "post_date": "03/17/2022 04:15:55",
          "content": "<p>Thanks for sharing this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1719207": "The files are:\n| filename | length  | link|\n| --- | --- | --- |\n| blkfra/XC649198.ogg |0.034844s  |https://www.xeno-canto.org/649198 |\n| normoc/XC150238.ogg |0.104500s |https://www.xeno-canto.org/150238 |\n\nI listened to them, and they don't make any sense. But the recordings are good on the xeno-canto website. So probably they are only corrupted here?\n\nBy the way, I have removed the data folder and re-added, things didn't change.",
    "1721913": "Thanks for pointing this out, I noticed the same issue on those two files in my data too. I'm now trying to go through the 14k audio files to find such anomalies so I can exclude them from the list. The CPU / memory seem to be maxing out when I run my code so trying to work on that. Once I have that in, happy to post it for others to use :)",
    "1723222": "that's nice, looking forward to it",
    "1723868": "Okay, so scanned all 14,852 files and the only two files corrupt (lengths of under 1 sec) where the ones you pointed out already 😀 \nSo here goes the code for parallel execution\n\n```\n# Import libraries\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.auto import tqdm\nimport pandas as pd\n\n# Load the data into the dataframe\ndf_birdclef = pd.read_csv('/kaggle/input/birdclef-2022/train_metadata.csv')\n\n# Function to get size of the audio\ndef get_length(fn):\n    fp = os.path.join(path, \"train_audio\", fn)\n    waveform, sample_rate = torchaudio.load(fp)\n    return waveform.size()[-1]\n\n# Run the jobs in parallel to get length and add to list called 'sizes'\nsizes = Parallel(n_jobs=os.cpu_count())(delayed(get_length)(fn) for fn in tqdm(df_birdclef[\"filename\"]))\n\n# Check if any sizes of audio less than 1 sec\ndf_sizes = pd.DataFrame(sizes)\ndf_sizes[df_sizes[0]<10000]\n```\n\nNote - the code above doesn't give the filename but it can be modified to give you that. I just left it at this.\n\nCredit: The code snippet above was heavily used from @jirkaborovec 's notebook mentioned below\nhttps://www.kaggle.com/jirkaborovec/birdclef-convert-spectrograms-noise-reduce",
    "1725397": "Thanks for sharing this."
  },
  "source": "meta"
}