{
  "id": 445547,
  "title": "torchaudio.load() is failing to load '.mp3' files",
  "url": "/competitions/bengaliai-speech/discussion/445547",
  "author_name": "UDOY DAS",
  "post_date": "2023-10-07T14:19:27.829000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am getting this error for quite a long time. I was trying to train a model but this error is occurring while preparing the dataset. I have used,</p>\n<p><code>train_data_d = train_data_d.cast_column('path', Audio(sampling_rate=16000))</code></p>\n<p>When I called,</p>\n<p><code>train_data_d[0]</code></p>\n<p>Below error occured,</p>\n<p><code>RuntimeError: Failed to load audio from /kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3</code></p>\n<p>This error occurs in <strong>Kaggle</strong> only</p>",
  "messages": [
    {
      "id": 2474438,
      "postDate": "2023-10-09T08:17:55.410Z",
      "content": "<p>Installing tag 2.14.15 of <a href=\"https://github.com/huggingface/datasets/tags\" target=\"_blank\">datasets</a> fix this problem</p>",
      "rawMarkdown": "Installing tag 2.14.15 of [datasets](https://github.com/huggingface/datasets/tags) fix this problem",
      "votes": 1
    },
    {
      "id": 2473570,
      "postDate": "2023-10-08T11:48:43.283Z",
      "content": "<p>Yes. Not working, I've replaced with librosa as below.  May be not a best option</p>\n<pre><code>fin = (, )\ndata = fin.read()\ndata = data.replace(\n    , \n    )\nfin.close()\nfin = (, )\nfin.write(data)\nfin.close()\n</code></pre>",
      "rawMarkdown": "Yes. Not working, I've replaced with librosa as below.  May be not a best option\n\n```python\nfin = open('/opt/conda/lib/python3.10/site-packages/datasets/features/audio.py', \"rt\")\ndata = fin.read()\ndata = data.replace(\n    'array, sampling_rate = torchaudio.load(path_or_file, format=\"mp3\")', \n    'import librosa;import torch;array, sampling_rate = librosa.load(path_or_file, sr=None);array = array.reshape(((1, array.shape[0])));array = torch.tensor(array)')\nfin.close()\nfin = open('/opt/conda/lib/python3.10/site-packages/datasets/features/audio.py', \"wt\")\nfin.write(data)\nfin.close()\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2485037,
          "postDate": "2023-10-16T21:54:52.207Z",
          "content": "<p>That's a very clever solution actually, huge help! Thanks a lot.</p>",
          "rawMarkdown": "That's a very clever solution actually, huge help! Thanks a lot."
        }
      ]
    },
    {
      "id": 2472979,
      "postDate": "2023-10-07T18:57:59.997Z",
      "content": "<p>This is probably due to the version mismatch of torchaudio. One plausible fix is to convert mp3s to wavs and then load them. </p>\n<pre><code> pydub import AudioSegment\naudio = AudioSegment.from_file(,=)\naudio.(,=)\n</code></pre>",
      "rawMarkdown": "This is probably due to the version mismatch of torchaudio. One plausible fix is to convert mp3s to wavs and then load them. \n```\nfrom pydub import AudioSegment\naudio = AudioSegment.from_file(\"mp3_file\",format=\"mp3\")\naudio.export(\"temp.wav\",format=\"wav\")\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2473558,
          "postDate": "2023-10-08T11:41:00.577Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2472692,
      "postDate": "2023-10-07T14:19:27.830Z",
      "content": "<p>I am getting this error for quite a long time. I was trying to train a model but this error is occurring while preparing the dataset. I have used,</p>\n<p><code>train_data_d = train_data_d.cast_column('path', Audio(sampling_rate=16000))</code></p>\n<p>When I called,</p>\n<p><code>train_data_d[0]</code></p>\n<p>Below error occured,</p>\n<p><code>RuntimeError: Failed to load audio from /kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3</code></p>\n<p>This error occurs in <strong>Kaggle</strong> only</p>",
      "rawMarkdown": "I am getting this error for quite a long time. I was trying to train a model but this error is occurring while preparing the dataset. I have used,\n\n`train_data_d = train_data_d.cast_column('path', Audio(sampling_rate=16000))`\n\nWhen I called,\n\n`train_data_d[0]`\n\nBelow error occured,\n\n`RuntimeError: Failed to load audio from /kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3`\n\nThis error occurs in **Kaggle** only",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2474438,
      "author_name": "Mikhail Edukov",
      "author_url": "",
      "post_date": "2023-10-09T08:17:55.410000",
      "content": "<p>Installing tag 2.14.15 of <a href=\"https://github.com/huggingface/datasets/tags\" target=\"_blank\">datasets</a> fix this problem</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2473570,
      "author_name": "Dmitriy Gerasimov",
      "author_url": "",
      "post_date": "2023-10-08T11:48:43.283000",
      "content": "<p>Yes. Not working, I've replaced with librosa as below.  May be not a best option</p>\n<pre><code>fin = (, )\ndata = fin.read()\ndata = data.replace(\n    , \n    )\nfin.close()\nfin = (, )\nfin.write(data)\nfin.close()\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2485037,
          "author_name": "Abdur Rahim",
          "author_url": "",
          "post_date": "2023-10-16T21:54:52.207000",
          "content": "<p>That's a very clever solution actually, huge help! Thanks a lot.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2472979,
      "author_name": "Md Boktiar Mahbub Murad",
      "author_url": "",
      "post_date": "2023-10-07T18:57:59.997000",
      "content": "<p>This is probably due to the version mismatch of torchaudio. One plausible fix is to convert mp3s to wavs and then load them. </p>\n<pre><code> pydub import AudioSegment\naudio = AudioSegment.from_file(,=)\naudio.(,=)\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2473558,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-10-08T11:41:00.577000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2474438": "Installing tag 2.14.15 of [datasets](https://github.com/huggingface/datasets/tags) fix this problem",
    "2473570": "Yes. Not working, I've replaced with librosa as below.  May be not a best option\n\n```python\nfin = open('/opt/conda/lib/python3.10/site-packages/datasets/features/audio.py', \"rt\")\ndata = fin.read()\ndata = data.replace(\n    'array, sampling_rate = torchaudio.load(path_or_file, format=\"mp3\")', \n    'import librosa;import torch;array, sampling_rate = librosa.load(path_or_file, sr=None);array = array.reshape(((1, array.shape[0])));array = torch.tensor(array)')\nfin.close()\nfin = open('/opt/conda/lib/python3.10/site-packages/datasets/features/audio.py', \"wt\")\nfin.write(data)\nfin.close()\n```",
    "2472979": "This is probably due to the version mismatch of torchaudio. One plausible fix is to convert mp3s to wavs and then load them. \n```\nfrom pydub import AudioSegment\naudio = AudioSegment.from_file(\"mp3_file\",format=\"mp3\")\naudio.export(\"temp.wav\",format=\"wav\")\n```",
    "2472692": "I am getting this error for quite a long time. I was trying to train a model but this error is occurring while preparing the dataset. I have used,\n\n`train_data_d = train_data_d.cast_column('path', Audio(sampling_rate=16000))`\n\nWhen I called,\n\n`train_data_d[0]`\n\nBelow error occured,\n\n`RuntimeError: Failed to load audio from /kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3`\n\nThis error occurs in **Kaggle** only"
  }
}