{
  "id": 430400,
  "title": "Seeking for help about voice extraction which takes too long to run",
  "url": "/competitions/bengaliai-speech/discussion/430400",
  "author_name": "Mutian Hong",
  "post_date": "2023-08-09T15:37:00.435000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Considering that the test set includes a lot of noice and back ground music, which can disturb the inference, I wanted to extract the human voice. So I tried to use Ultimate Vocal Remover (UVR) to preprocess the audio. I used the CIL virsion from github <a href=\"url\" target=\"_blank\">https://github.com/seanghay/uvr</a> . And here is the same kaggle dataset:<a href=\"url\" target=\"_blank\">https://www.kaggle.com/datasets/hongori/uvr-cil-v3</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13325950%2Fbc803a3b1eb56e8565da9ce9042c6240%2FScreenshot%202023-08-09%20232722.png?generation=1691594947047429&amp;alt=media\" alt=\"\"><br>\n I used it to infer as following:</p>\n<pre><code> ():\n    \n    results = []\n    command = \n    os.system(command)\n    new_paths = []\n     path  audio_paths:\n        new_path = \n        new_paths.append(new_path)\n     path  new_paths:\n        pred = \n        :\n            pred = infer(path)\n        :\n            pred = \n         (pred)==:\n            pred = \n        results.append(pred)\n     results\n</code></pre>\n<p>it works well, but takes too long. It takes average about 10s to run a iter. And finally it ran over 9 hours and was canceled.</p>\n<p>So I posted this topic to brainstorm if is there some ways that works faster.</p>\n<p>I have not tried to use different models, and I think that might not improve the speed a lot because the process takes about 6 seconds, and loading the data and other steps takes another 4 seconds. And that did not take the infer step into consideration. (now the pure infer takes about 2 hours for me). The lighter model maybe can shorten the process to 3 seconds, but still not enough. </p>\n<p>There might be some trivial ways to do it like using equalizer or filter, I haven't tried yet. Did anyone used them and got imiprovement? And do anyone got some ideas about accelerating UVR? Thanks a lot!</p>",
  "messages": [
    {
      "id": 2382144,
      "postDate": "2023-08-09T15:37:00.437Z",
      "content": "<p>Considering that the test set includes a lot of noice and back ground music, which can disturb the inference, I wanted to extract the human voice. So I tried to use Ultimate Vocal Remover (UVR) to preprocess the audio. I used the CIL virsion from github <a href=\"url\" target=\"_blank\">https://github.com/seanghay/uvr</a> . And here is the same kaggle dataset:<a href=\"url\" target=\"_blank\">https://www.kaggle.com/datasets/hongori/uvr-cil-v3</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13325950%2Fbc803a3b1eb56e8565da9ce9042c6240%2FScreenshot%202023-08-09%20232722.png?generation=1691594947047429&amp;alt=media\" alt=\"\"><br>\n I used it to infer as following:</p>\n<pre><code> ():\n    \n    results = []\n    command = \n    os.system(command)\n    new_paths = []\n     path  audio_paths:\n        new_path = \n        new_paths.append(new_path)\n     path  new_paths:\n        pred = \n        :\n            pred = infer(path)\n        :\n            pred = \n         (pred)==:\n            pred = \n        results.append(pred)\n     results\n</code></pre>\n<p>it works well, but takes too long. It takes average about 10s to run a iter. And finally it ran over 9 hours and was canceled.</p>\n<p>So I posted this topic to brainstorm if is there some ways that works faster.</p>\n<p>I have not tried to use different models, and I think that might not improve the speed a lot because the process takes about 6 seconds, and loading the data and other steps takes another 4 seconds. And that did not take the infer step into consideration. (now the pure infer takes about 2 hours for me). The lighter model maybe can shorten the process to 3 seconds, but still not enough. </p>\n<p>There might be some trivial ways to do it like using equalizer or filter, I haven't tried yet. Did anyone used them and got imiprovement? And do anyone got some ideas about accelerating UVR? Thanks a lot!</p>",
      "rawMarkdown": "Considering that the test set includes a lot of noice and back ground music, which can disturb the inference, I wanted to extract the human voice. So I tried to use Ultimate Vocal Remover (UVR) to preprocess the audio. I used the CIL virsion from github [https://github.com/seanghay/uvr](url) . And here is the same kaggle dataset:[https://www.kaggle.com/datasets/hongori/uvr-cil-v3](url)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13325950%2Fbc803a3b1eb56e8565da9ce9042c6240%2FScreenshot%202023-08-09%20232722.png?generation=1691594947047429&alt=media)\n I used it to infer as following:\n\n```\ndef batch_infer(audio_paths, batch_size=BATCH_SIZE):\n    '''\n    infers on a batch of audio\n    args:\n      audio_paths  : list of path to audio files <list of string>\n    returns:\n      bangla predicted texts <list of string>\n    '''\n    results = []\n    command = f\"! python /kaggle/input/uvr-cil-v3/vocal-remover-v5.0.4/vocal-remover/inference.py {' '.join(['-i %s' % p for p in audio_paths])} -P /kaggle/input/uvr-cil-v3/vocal-remover-v5.0.4/vocal-remover/models/baseline.pth -o /kaggle/working/processed -r 36000\"\n    os.system(command)\n    new_paths = []\n    for path in audio_paths:\n        new_path = f'/kaggle/working/processed/{path.split(\"/\")[-1].split(\".\")[0]}_Vocals.wav'\n        new_paths.append(new_path)\n    for path in new_paths:\n        pred = \"\"\n        try:\n            pred = infer(path)\n        except:\n            pred = \"এ\"\n        if len(pred)==0:\n            pred = \"এ\"\n        results.append(pred)\n    return results\n```\n\n\nit works well, but takes too long. It takes average about 10s to run a iter. And finally it ran over 9 hours and was canceled.\n\nSo I posted this topic to brainstorm if is there some ways that works faster.\n\nI have not tried to use different models, and I think that might not improve the speed a lot because the process takes about 6 seconds, and loading the data and other steps takes another 4 seconds. And that did not take the infer step into consideration. (now the pure infer takes about 2 hours for me). The lighter model maybe can shorten the process to 3 seconds, but still not enough. \n\nThere might be some trivial ways to do it like using equalizer or filter, I haven't tried yet. Did anyone used them and got imiprovement? And do anyone got some ideas about accelerating UVR? Thanks a lot!",
      "votes": 2
    },
    {
      "id": 2383409,
      "postDate": "2023-08-10T10:38:37.077Z",
      "content": "<p>It seems that inference.py from <a href=\"https://www.kaggle.com/datasets/hongori/uvr-cil-v3\" target=\"_blank\">this dataset</a> doesn't support multiple files as input and only one per batch can be processed. Did you adjust code to handle multiple files?</p>",
      "rawMarkdown": "It seems that inference.py from [this dataset](https://www.kaggle.com/datasets/hongori/uvr-cil-v3) doesn't support multiple files as input and only one per batch can be processed. Did you adjust code to handle multiple files?",
      "isDeleted": true,
      "replies": [
        {
          "id": 2383557,
          "postDate": "2023-08-10T12:48:42.887Z",
          "content": "<p>No. I failed to do batch infer. Cuz my UVR   version do not support batch infer. I am still trying :)</p>",
          "rawMarkdown": "No. I failed to do batch infer. Cuz my UVR   version do not support batch infer. I am still trying :)",
          "replies": [
            {
              "id": 2383847,
              "postDate": "2023-08-10T15:36:24.817Z",
              "content": "<p>I think use a pretrained <a href=\"https://github.com/facebookresearch/denoiser\" target=\"_blank\">facebook-research denoiser</a> is overall better solution.</p>\n<pre><code> denoiser  pretrained\n denoiser.dsp  convert_audio\n\nmodel = pretrained.dns64()\nwav, sr = torchaudio.load(example_audio)\nwav = convert_audio(wav, sr, model.sample_rate, model.chin)\n\n torch.no_grad():\n    denoised = model(wav.unsqueeze())[]\n</code></pre>",
              "rawMarkdown": "I think use a pretrained [facebook-research denoiser](https://github.com/facebookresearch/denoiser) is overall better solution.\n\n```python\nfrom denoiser import pretrained\nfrom denoiser.dsp import convert_audio\n\nmodel = pretrained.dns64()\nwav, sr = torchaudio.load(example_audio)\nwav = convert_audio(wav, sr, model.sample_rate, model.chin)\n\nwith torch.no_grad():\n    denoised = model(wav.unsqueeze(0))[0]\n```\n",
              "votes": 4,
              "isDeleted": true
            },
            {
              "id": 2383961,
              "postDate": "2023-08-10T17:11:02.407Z",
              "content": "<p>Thanks! I'll give it a try!</p>",
              "rawMarkdown": "Thanks! I'll give it a try!"
            },
            {
              "id": 2385189,
              "postDate": "2023-08-11T07:54:16.223Z",
              "content": "<p>In the submission, the internet is off. And I don't know how to load the model from local. Do you have any ideas to solve this?🥲</p>",
              "rawMarkdown": "In the submission, the internet is off. And I don't know how to load the model from local. Do you have any ideas to solve this?🥲"
            },
            {
              "id": 2385394,
              "postDate": "2023-08-11T10:03:54.283Z",
              "content": "<p>Yes, you can simply save parameters and load model directly.</p>\n<pre><code> denoiser  pretrained\n denoiser.demucs  Demucs\n\nmodel = pretrained.dns64()\ntorch.save(model.state_dict(), )\nmodel = Demucs(hidden=, sample_rate=)\nmodel.load_state_dict(torch.load())\n</code></pre>",
              "rawMarkdown": "Yes, you can simply save parameters and load model directly.\n\n```python\nfrom denoiser import pretrained\nfrom denoiser.demucs import Demucs\n\nmodel = pretrained.dns64()\ntorch.save(model.state_dict(), './params.pt')\nmodel = Demucs(hidden=64, sample_rate=16000)\nmodel.load_state_dict(torch.load('./params.pt'))\n```",
              "votes": 2,
              "isDeleted": true
            },
            {
              "id": 2387206,
              "postDate": "2023-08-12T13:43:50.267Z",
              "content": "<p>Thanks a lot!</p>",
              "rawMarkdown": "Thanks a lot!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2383409,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-10T10:38:37.077000",
      "content": "<p>It seems that inference.py from <a href=\"https://www.kaggle.com/datasets/hongori/uvr-cil-v3\" target=\"_blank\">this dataset</a> doesn't support multiple files as input and only one per batch can be processed. Did you adjust code to handle multiple files?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2383557,
          "author_name": "Mutian Hong",
          "author_url": "",
          "post_date": "2023-08-10T12:48:42.887000",
          "content": "<p>No. I failed to do batch infer. Cuz my UVR   version do not support batch infer. I am still trying :)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2383847,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-10T15:36:24.817000",
              "content": "<p>I think use a pretrained <a href=\"https://github.com/facebookresearch/denoiser\" target=\"_blank\">facebook-research denoiser</a> is overall better solution.</p>\n<pre><code> denoiser  pretrained\n denoiser.dsp  convert_audio\n\nmodel = pretrained.dns64()\nwav, sr = torchaudio.load(example_audio)\nwav = convert_audio(wav, sr, model.sample_rate, model.chin)\n\n torch.no_grad():\n    denoised = model(wav.unsqueeze())[]\n</code></pre>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2383961,
              "author_name": "Mutian Hong",
              "author_url": "",
              "post_date": "2023-08-10T17:11:02.407000",
              "content": "<p>Thanks! I'll give it a try!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2385189,
              "author_name": "Mutian Hong",
              "author_url": "",
              "post_date": "2023-08-11T07:54:16.223000",
              "content": "<p>In the submission, the internet is off. And I don't know how to load the model from local. Do you have any ideas to solve this?🥲</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2385394,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-11T10:03:54.283000",
              "content": "<p>Yes, you can simply save parameters and load model directly.</p>\n<pre><code> denoiser  pretrained\n denoiser.demucs  Demucs\n\nmodel = pretrained.dns64()\ntorch.save(model.state_dict(), )\nmodel = Demucs(hidden=, sample_rate=)\nmodel.load_state_dict(torch.load())\n</code></pre>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2387206,
              "author_name": "Mutian Hong",
              "author_url": "",
              "post_date": "2023-08-12T13:43:50.267000",
              "content": "<p>Thanks a lot!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2382144": "Considering that the test set includes a lot of noice and back ground music, which can disturb the inference, I wanted to extract the human voice. So I tried to use Ultimate Vocal Remover (UVR) to preprocess the audio. I used the CIL virsion from github [https://github.com/seanghay/uvr](url) . And here is the same kaggle dataset:[https://www.kaggle.com/datasets/hongori/uvr-cil-v3](url)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13325950%2Fbc803a3b1eb56e8565da9ce9042c6240%2FScreenshot%202023-08-09%20232722.png?generation=1691594947047429&alt=media)\n I used it to infer as following:\n\n```\ndef batch_infer(audio_paths, batch_size=BATCH_SIZE):\n    '''\n    infers on a batch of audio\n    args:\n      audio_paths  : list of path to audio files <list of string>\n    returns:\n      bangla predicted texts <list of string>\n    '''\n    results = []\n    command = f\"! python /kaggle/input/uvr-cil-v3/vocal-remover-v5.0.4/vocal-remover/inference.py {' '.join(['-i %s' % p for p in audio_paths])} -P /kaggle/input/uvr-cil-v3/vocal-remover-v5.0.4/vocal-remover/models/baseline.pth -o /kaggle/working/processed -r 36000\"\n    os.system(command)\n    new_paths = []\n    for path in audio_paths:\n        new_path = f'/kaggle/working/processed/{path.split(\"/\")[-1].split(\".\")[0]}_Vocals.wav'\n        new_paths.append(new_path)\n    for path in new_paths:\n        pred = \"\"\n        try:\n            pred = infer(path)\n        except:\n            pred = \"এ\"\n        if len(pred)==0:\n            pred = \"এ\"\n        results.append(pred)\n    return results\n```\n\n\nit works well, but takes too long. It takes average about 10s to run a iter. And finally it ran over 9 hours and was canceled.\n\nSo I posted this topic to brainstorm if is there some ways that works faster.\n\nI have not tried to use different models, and I think that might not improve the speed a lot because the process takes about 6 seconds, and loading the data and other steps takes another 4 seconds. And that did not take the infer step into consideration. (now the pure infer takes about 2 hours for me). The lighter model maybe can shorten the process to 3 seconds, but still not enough. \n\nThere might be some trivial ways to do it like using equalizer or filter, I haven't tried yet. Did anyone used them and got imiprovement? And do anyone got some ideas about accelerating UVR? Thanks a lot!",
    "2383409": "It seems that inference.py from [this dataset](https://www.kaggle.com/datasets/hongori/uvr-cil-v3) doesn't support multiple files as input and only one per batch can be processed. Did you adjust code to handle multiple files?"
  }
}