{
  "id": 431049,
  "title": "How to fix torchaudio error when using GPU?",
  "url": "/competitions/bengaliai-speech/discussion/431049",
  "author_name": "",
  "post_date": "2023-08-11T21:00:00.830129400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi! I am trying to finetune Whisper using the hugging face Dataset library. But when I use GPU, it says error in loading the mp3 audio file. But it works perfectly when using the CPU. How do I fix the issue?</p>\n<pre><code>DatasetDict({\n    train: Dataset({\n        features: [, ],\n        num_rows: \n    })\n    valid: Dataset({\n        features: [, ],\n        num_rows: \n    })\n})\n\ndataset[][]\n</code></pre>\n<p>RuntimeError: Failed to load audio from /kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3</p>\n<p>A similar error is coming when I am using the <code>map</code> function.</p>\n<pre><code>common_voice = dataset.(prepare_dataset, remove_columns=dataset.column_names[])\n</code></pre>\n<p>How do I use GPU without the error?</p>",
  "messages": [
    {
      "id": "2386296",
      "postDate": "08/11/2023 21:00:00",
      "content": "<p>Hi! I am trying to finetune Whisper using the hugging face Dataset library. But when I use GPU, it says error in loading the mp3 audio file. But it works perfectly when using the CPU. How do I fix the issue?</p>\n<pre><code>DatasetDict({\n    train: Dataset({\n        features: [, ],\n        num_rows: \n    })\n    valid: Dataset({\n        features: [, ],\n        num_rows: \n    })\n})\n\ndataset[][]\n</code></pre>\n<p>RuntimeError: Failed to load audio from /kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3</p>\n<p>A similar error is coming when I am using the <code>map</code> function.</p>\n<pre><code>common_voice = dataset.(prepare_dataset, remove_columns=dataset.column_names[])\n</code></pre>\n<p>How do I use GPU without the error?</p>",
      "rawMarkdown": "Hi! I am trying to finetune Whisper using the hugging face Dataset library. But when I use GPU, it says error in loading the mp3 audio file. But it works perfectly when using the CPU. How do I fix the issue?\n\n```python\nDatasetDict({\n    train: Dataset({\n        features: ['audio', 'sentence'],\n        num_rows: 934065\n    })\n    valid: Dataset({\n        features: ['audio', 'sentence'],\n        num_rows: 29588\n    })\n})\n\ndataset['train'][0]\n```\n\nRuntimeError: Failed to load audio from /kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3\n\nA similar error is coming when I am using the `map` function.\n\n```python\ncommon_voice = dataset.map(prepare_dataset, remove_columns=dataset.column_names[\"train\"])\n```\n\nHow do I use GPU without the error?",
      "votes": null
    },
    {
      "id": "2386477",
      "postDate": "08/12/2023 01:38:43",
      "content": "<p>I also met that problem. This is my way to solve it :<br>\nJust use another way to load the file.</p>\n<pre><code> wav, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n wav = np.expand\n wav = torch.from\n</code></pre>\n<p>change the \"wav, sr = torchaudio.load(example_audio)\" into the above one. </p>",
      "rawMarkdown": "I also met that problem. This is my way to solve it :\nJust use another way to load the file.\n```   \n wav, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n wav = np.expand_dims(wav,axis=0)\n wav = torch.from_numpy(wav)\n```\n\nchange the \"wav, sr = torchaudio.load(example_audio)\" into the above one.",
      "votes": null
    },
    {
      "id": "2387002",
      "postDate": "08/12/2023 10:20:19",
      "content": "<p>Thanks. This helped.</p>",
      "rawMarkdown": "Thanks. This helped.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2386477,
      "author_name": "hongori",
      "author_url": "",
      "post_date": "08/12/2023 01:38:43",
      "content": "<p>I also met that problem. This is my way to solve it :<br>\nJust use another way to load the file.</p>\n<pre><code> wav, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n wav = np.expand\n wav = torch.from\n</code></pre>\n<p>change the \"wav, sr = torchaudio.load(example_audio)\" into the above one. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2387002,
          "author_name": "pritamsinha23",
          "author_url": "",
          "post_date": "08/12/2023 10:20:19",
          "content": "<p>Thanks. This helped.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2386296": "Hi! I am trying to finetune Whisper using the hugging face Dataset library. But when I use GPU, it says error in loading the mp3 audio file. But it works perfectly when using the CPU. How do I fix the issue?\n\n```python\nDatasetDict({\n    train: Dataset({\n        features: ['audio', 'sentence'],\n        num_rows: 934065\n    })\n    valid: Dataset({\n        features: ['audio', 'sentence'],\n        num_rows: 29588\n    })\n})\n\ndataset['train'][0]\n```\n\nRuntimeError: Failed to load audio from /kaggle/input/bengaliai-speech/train_mp3s/000005f3362c.mp3\n\nA similar error is coming when I am using the `map` function.\n\n```python\ncommon_voice = dataset.map(prepare_dataset, remove_columns=dataset.column_names[\"train\"])\n```\n\nHow do I use GPU without the error?",
    "2386477": "I also met that problem. This is my way to solve it :\nJust use another way to load the file.\n```   \n wav, sr = librosa.load(audio_path, sr=processor.feature_extractor.sampling_rate)\n wav = np.expand_dims(wav,axis=0)\n wav = torch.from_numpy(wav)\n```\n\nchange the \"wav, sr = torchaudio.load(example_audio)\" into the above one.",
    "2387002": "Thanks. This helped."
  },
  "source": "meta"
}