{
  "id": 499048,
  "title": "OpenVino not working with torchaudio MelSpectrogram?",
  "url": "/competitions/birdclef-2024/discussion/499048",
  "author_name": "",
  "post_date": "2024-04-30T12:53:38.517665200Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I was following birdclef 2023 solutions, and quite a few of them used torchaudio's MelSpectrogram as part of feature extractor on GPU (locally). For kaggle inference, when I try to convert this model (MelSpectrogram-&gt;timm-&gt;fc), I get weird OpenVino error. </p>\n<p>If anyone has helpful hints/solution to solve this it would be great. Currently, I've ditched openvino and inferencing using torch.jit (which i suspect is slower). </p>",
  "messages": [
    {
      "id": "2784706",
      "postDate": "04/30/2024 12:53:38",
      "content": "<p>I was following birdclef 2023 solutions, and quite a few of them used torchaudio's MelSpectrogram as part of feature extractor on GPU (locally). For kaggle inference, when I try to convert this model (MelSpectrogram-&gt;timm-&gt;fc), I get weird OpenVino error. </p>\n<p>If anyone has helpful hints/solution to solve this it would be great. Currently, I've ditched openvino and inferencing using torch.jit (which i suspect is slower). </p>",
      "rawMarkdown": "I was following birdclef 2023 solutions, and quite a few of them used torchaudio's MelSpectrogram as part of feature extractor on GPU (locally). For kaggle inference, when I try to convert this model (MelSpectrogram->timm->fc), I get weird OpenVino error. \n\nIf anyone has helpful hints/solution to solve this it would be great. Currently, I've ditched openvino and inferencing using torch.jit (which i suspect is slower).",
      "votes": null
    },
    {
      "id": "2784749",
      "postDate": "04/30/2024 13:25:42",
      "content": "<p>I don't know about the openvino issue - but you can pull the melspectrogram pipeline outside your model and do it separate  from the openvino. </p>",
      "rawMarkdown": "I don't know about the openvino issue - but you can pull the melspectrogram pipeline outside your model and do it separate  from the openvino.",
      "votes": null
    },
    {
      "id": "2784754",
      "postDate": "04/30/2024 13:29:11",
      "content": "<p>Yep, I know that. Just wanted additional flexibility in my train pipeline, hence MelSpec inside nnet. </p>",
      "rawMarkdown": "Yep, I know that. Just wanted additional flexibility in my train pipeline, hence MelSpec inside nnet.",
      "votes": null
    },
    {
      "id": "2784875",
      "postDate": "04/30/2024 14:25:41",
      "content": "<p>This is how I did it and used the output as an input to my inference notebook.</p>\n<ul>\n<li>First install the openvino related libraries</li>\n</ul>\n<pre><code>TMP_DIR = \n! -p {TMP_DIR}\n!pip wheel openvino --wheel-dir={TMP_DIR} --no-cache-dir\n!pip install wheels/openvino-2024.0.0-14509-cp310-cp310-manylinux2014_x86_64.whl --no-index --find-links wheels/\n!pip install openvino-dev[onnx]\n</code></pre>\n<ul>\n<li>Convert your model to onnx.</li>\n</ul>\n<pre><code>rand_input = torch.randn(input_shape_of_your_model)\nmodel = ... \n\ninput_names = ... \noutput_names = ... \ntorch.onnx.export(model, rand_input, , input_names=input_names, output_names=output_names)\n</code></pre>\n<ul>\n<li>Convert your model to openvino</li>\n</ul>\n<pre><code>!mo --input_model model.onnx --compress_to_fp16 \n</code></pre>",
      "rawMarkdown": "This is how I did it and used the output as an input to my inference notebook.\n\n- First install the openvino related libraries\n```bash\nTMP_DIR = \"wheels\"\n!mkdir -p {TMP_DIR}\n!pip wheel openvino --wheel-dir={TMP_DIR} --no-cache-dir\n!pip install wheels/openvino-2024.0.0-14509-cp310-cp310-manylinux2014_x86_64.whl --no-index --find-links wheels/\n!pip install openvino-dev[onnx]\n```\n- Convert your model to onnx.\n```python\nrand_input = torch.randn(input_shape_of_your_model)\nmodel = ... # your model\n\ninput_names = ... # the names of the params in your forward function\noutput_names = ... # the output variable name in your forward function\ntorch.onnx.export(model, rand_input, \"model.onnx\", input_names=input_names, output_names=output_names)\n```\n- Convert your model to openvino\n```bash\n!mo --input_model model.onnx --compress_to_fp16 # fp16 is optional\n```",
      "votes": null
    },
    {
      "id": "2784933",
      "postDate": "04/30/2024 14:59:21",
      "content": "<p>Does your model contain a torchaudio layer? </p>\n<pre><code>import torchaudio as T\n\nself.spect_layer= nn(\n    T(),\n    T(top_db = ),\n)\n</code></pre>\n<p>Like <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>, when I try onnx conversion without removing the torchaudio layer, there are errors due to operations not implemented in onnx.</p>",
      "rawMarkdown": "Does your model contain a torchaudio layer? \n```\nimport torchaudio.transforms as T\n\nself.spect_layer= nn.Sequential(\n    T.MelSpectrogram(),\n    T.AmplitudeToDB(top_db = 80),\n)\n```\nLike @pheadrus, when I try onnx conversion without removing the torchaudio layer, there are errors due to operations not implemented in onnx.",
      "votes": null
    },
    {
      "id": "2785286",
      "postDate": "04/30/2024 17:48:03",
      "content": "<p>Yes I am using torchaudio, I have the following ones:</p>\n<pre><code>torchaudio.transforms.MelSpectrogram\ntorchaudio.transforms.TimeMasking\ntorchaudio.transforms.FrequencyMasking\n</code></pre>",
      "rawMarkdown": "Yes I am using torchaudio, I have the following ones:\n\n```python\ntorchaudio.transforms.MelSpectrogram\ntorchaudio.transforms.TimeMasking\ntorchaudio.transforms.FrequencyMasking\n```",
      "votes": null
    },
    {
      "id": "2809309",
      "postDate": "05/12/2024 17:00:45",
      "content": "<p>Yes, because of the STFT implementation in pytorch.<br>\nYou  can write STFT by yourself with Conv1D proposed by 12th place last year.<br>\nI tried, which makes inference super slow… maybe I did not use it correctly</p>",
      "rawMarkdown": "Yes, because of the STFT implementation in pytorch.\nYou  can write STFT by yourself with Conv1D proposed by 12th place last year.\nI tried, which makes inference super slow... maybe I did not use it correctly",
      "votes": null
    },
    {
      "id": "2812594",
      "postDate": "05/14/2024 10:12:18",
      "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> is it still a dead end converting eca_nfnet model for openvino ? I've read that in your past solution write up (I'm stuck with weird error on my side) </p>",
      "rawMarkdown": "honglihang is it still a dead end converting eca_nfnet model for openvino ? I've read that in your past solution write up (I'm stuck with weird error on my side)",
      "votes": null
    },
    {
      "id": "2812627",
      "postDate": "05/14/2024 10:33:22",
      "content": "<p>I remember last year's 9th place manually modified the eca_nfnet and made it work. He or She left a link in last year's 9th place solution discussion.</p>",
      "rawMarkdown": "I remember last year's 9th place manually modified the eca_nfnet and made it work. He or She left a link in last year's 9th place solution discussion.",
      "votes": null
    },
    {
      "id": "2824018",
      "postDate": "05/19/2024 14:44:27",
      "content": "<p>You can use <a href=\"https://github.com/adobe-research/convmelspec\" target=\"_blank\">https://github.com/adobe-research/convmelspec</a></p>",
      "rawMarkdown": "You can use https://github.com/adobe-research/convmelspec",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2784749,
      "author_name": "llleeeoooh",
      "author_url": "",
      "post_date": "04/30/2024 13:25:42",
      "content": "<p>I don't know about the openvino issue - but you can pull the melspectrogram pipeline outside your model and do it separate  from the openvino. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2784754,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "04/30/2024 13:29:11",
          "content": "<p>Yep, I know that. Just wanted additional flexibility in my train pipeline, hence MelSpec inside nnet. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2784875,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "04/30/2024 14:25:41",
      "content": "<p>This is how I did it and used the output as an input to my inference notebook.</p>\n<ul>\n<li>First install the openvino related libraries</li>\n</ul>\n<pre><code>TMP_DIR = \n! -p {TMP_DIR}\n!pip wheel openvino --wheel-dir={TMP_DIR} --no-cache-dir\n!pip install wheels/openvino-2024.0.0-14509-cp310-cp310-manylinux2014_x86_64.whl --no-index --find-links wheels/\n!pip install openvino-dev[onnx]\n</code></pre>\n<ul>\n<li>Convert your model to onnx.</li>\n</ul>\n<pre><code>rand_input = torch.randn(input_shape_of_your_model)\nmodel = ... \n\ninput_names = ... \noutput_names = ... \ntorch.onnx.export(model, rand_input, , input_names=input_names, output_names=output_names)\n</code></pre>\n<ul>\n<li>Convert your model to openvino</li>\n</ul>\n<pre><code>!mo --input_model model.onnx --compress_to_fp16 \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2784933,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "04/30/2024 14:59:21",
          "content": "<p>Does your model contain a torchaudio layer? </p>\n<pre><code>import torchaudio as T\n\nself.spect_layer= nn(\n    T(),\n    T(top_db = ),\n)\n</code></pre>\n<p>Like <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>, when I try onnx conversion without removing the torchaudio layer, there are errors due to operations not implemented in onnx.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2785286,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "04/30/2024 17:48:03",
              "content": "<p>Yes I am using torchaudio, I have the following ones:</p>\n<pre><code>torchaudio.transforms.MelSpectrogram\ntorchaudio.transforms.TimeMasking\ntorchaudio.transforms.FrequencyMasking\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2809309,
      "author_name": "honglihang",
      "author_url": "",
      "post_date": "05/12/2024 17:00:45",
      "content": "<p>Yes, because of the STFT implementation in pytorch.<br>\nYou  can write STFT by yourself with Conv1D proposed by 12th place last year.<br>\nI tried, which makes inference super slow… maybe I did not use it correctly</p>",
      "votes": null,
      "replies": [
        {
          "id": 2812594,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "05/14/2024 10:12:18",
          "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> is it still a dead end converting eca_nfnet model for openvino ? I've read that in your past solution write up (I'm stuck with weird error on my side) </p>",
          "votes": null,
          "replies": [
            {
              "id": 2812627,
              "author_name": "honglihang",
              "author_url": "",
              "post_date": "05/14/2024 10:33:22",
              "content": "<p>I remember last year's 9th place manually modified the eca_nfnet and made it work. He or She left a link in last year's 9th place solution discussion.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2824018,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "05/19/2024 14:44:27",
      "content": "<p>You can use <a href=\"https://github.com/adobe-research/convmelspec\" target=\"_blank\">https://github.com/adobe-research/convmelspec</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2784706": "I was following birdclef 2023 solutions, and quite a few of them used torchaudio's MelSpectrogram as part of feature extractor on GPU (locally). For kaggle inference, when I try to convert this model (MelSpectrogram->timm->fc), I get weird OpenVino error. \n\nIf anyone has helpful hints/solution to solve this it would be great. Currently, I've ditched openvino and inferencing using torch.jit (which i suspect is slower).",
    "2784749": "I don't know about the openvino issue - but you can pull the melspectrogram pipeline outside your model and do it separate  from the openvino.",
    "2784754": "Yep, I know that. Just wanted additional flexibility in my train pipeline, hence MelSpec inside nnet.",
    "2784875": "This is how I did it and used the output as an input to my inference notebook.\n\n- First install the openvino related libraries\n```bash\nTMP_DIR = \"wheels\"\n!mkdir -p {TMP_DIR}\n!pip wheel openvino --wheel-dir={TMP_DIR} --no-cache-dir\n!pip install wheels/openvino-2024.0.0-14509-cp310-cp310-manylinux2014_x86_64.whl --no-index --find-links wheels/\n!pip install openvino-dev[onnx]\n```\n- Convert your model to onnx.\n```python\nrand_input = torch.randn(input_shape_of_your_model)\nmodel = ... # your model\n\ninput_names = ... # the names of the params in your forward function\noutput_names = ... # the output variable name in your forward function\ntorch.onnx.export(model, rand_input, \"model.onnx\", input_names=input_names, output_names=output_names)\n```\n- Convert your model to openvino\n```bash\n!mo --input_model model.onnx --compress_to_fp16 # fp16 is optional\n```",
    "2784933": "Does your model contain a torchaudio layer? \n```\nimport torchaudio.transforms as T\n\nself.spect_layer= nn.Sequential(\n    T.MelSpectrogram(),\n    T.AmplitudeToDB(top_db = 80),\n)\n```\nLike @pheadrus, when I try onnx conversion without removing the torchaudio layer, there are errors due to operations not implemented in onnx.",
    "2785286": "Yes I am using torchaudio, I have the following ones:\n\n```python\ntorchaudio.transforms.MelSpectrogram\ntorchaudio.transforms.TimeMasking\ntorchaudio.transforms.FrequencyMasking\n```",
    "2809309": "Yes, because of the STFT implementation in pytorch.\nYou  can write STFT by yourself with Conv1D proposed by 12th place last year.\nI tried, which makes inference super slow... maybe I did not use it correctly",
    "2812594": "honglihang is it still a dead end converting eca_nfnet model for openvino ? I've read that in your past solution write up (I'm stuck with weird error on my side)",
    "2812627": "I remember last year's 9th place manually modified the eca_nfnet and made it work. He or She left a link in last year's 9th place solution discussion.",
    "2824018": "You can use https://github.com/adobe-research/convmelspec"
  },
  "source": "meta"
}