{
  "id": 501619,
  "title": "No speed up from jit to onnx",
  "url": "/competitions/birdclef-2024/discussion/501619",
  "author_name": "",
  "post_date": "2024-05-10T04:36:24.668906Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I did transformations offline with my computer. It seems that I did not get improvement of inference speed from jit <br>\n to onnx/openvino. It can be even slower. What can be the reason for it?</p>\n<p>I used the following code for inference</p>\n<pre><code> concurrent.futures.ThreadPoolExecutor(max_workers=)  executor:\n        dicts = (executor.(prediction_for_clip, all_audios))\n</code></pre>",
  "messages": [
    {
      "id": "2804446",
      "postDate": "05/10/2024 04:36:24",
      "content": "<p>I did transformations offline with my computer. It seems that I did not get improvement of inference speed from jit <br>\n to onnx/openvino. It can be even slower. What can be the reason for it?</p>\n<p>I used the following code for inference</p>\n<pre><code> concurrent.futures.ThreadPoolExecutor(max_workers=)  executor:\n        dicts = (executor.(prediction_for_clip, all_audios))\n</code></pre>",
      "rawMarkdown": "I did transformations offline with my computer. It seems that I did not get improvement of inference speed from jit \n to onnx/openvino. It can be even slower. What can be the reason for it?\n\nI used the following code for inference\n\n```python\nwith concurrent.futures.ThreadPoolExecutor(max_workers=4) as executor:\n        dicts = list(executor.map(prediction_for_clip, all_audios))\n```",
      "votes": null
    },
    {
      "id": "2804552",
      "postDate": "05/10/2024 05:30:09",
      "content": "<p>Probably predictor itself is using all cores already, so no reason to use <code>concurrent.futures</code> here</p>",
      "rawMarkdown": "Probably predictor itself is using all cores already, so no reason to use `concurrent.futures` here",
      "votes": null
    },
    {
      "id": "2804661",
      "postDate": "05/10/2024 06:28:03",
      "content": "<p>That's true, interestingly increasing batchsize decrease inference speed for jit by a lot.</p>",
      "rawMarkdown": "That's true, interestingly increasing batchsize decrease inference speed for jit by a lot.",
      "votes": null
    },
    {
      "id": "2805202",
      "postDate": "05/10/2024 12:36:46",
      "content": "<p>What is the acceleration from base to onnx inference? My base reasoning is the same speed as onnx</p>",
      "rawMarkdown": "What is the acceleration from base to onnx inference? My base reasoning is the same speed as onnx",
      "votes": null
    },
    {
      "id": "2805860",
      "postDate": "05/10/2024 19:10:31",
      "content": "<p>have you tried with <code>torch.jit.optimized_execution(False)</code>. I've noticed for some of the backbones this speeds up inference.</p>",
      "rawMarkdown": "have you tried with `torch.jit.optimized_execution(False)`. I've noticed for some of the backbones this speeds up inference.",
      "votes": null
    },
    {
      "id": "2806300",
      "postDate": "05/11/2024 02:52:38",
      "content": "<p>I used it. From my local test, onnx is twice faster than jit. But not this case in kaggle notebook. It's very weird.</p>",
      "rawMarkdown": "I used it. From my local test, onnx is twice faster than jit. But not this case in kaggle notebook. It's very weird.",
      "votes": null
    },
    {
      "id": "2806301",
      "postDate": "05/11/2024 02:53:07",
      "content": "<p>Local test shows that onnx  is twice (at most) faster than jit.</p>",
      "rawMarkdown": "Local test shows that onnx  is twice (at most) faster than jit.",
      "votes": null
    },
    {
      "id": "2806948",
      "postDate": "05/11/2024 12:19:25",
      "content": "<p>I get some speedup form onnx, but not a lot. Openvino is much faster, like 2x fatser,  but I lose some accuracy on LB.</p>\n<p>I am not sure why given others claim openvino works just fine.</p>",
      "rawMarkdown": "I get some speedup form onnx, but not a lot. Openvino is much faster, like 2x fatser,  but I lose some accuracy on LB.\n\nI am not sure why given others claim openvino works just fine.",
      "votes": null
    },
    {
      "id": "2807160",
      "postDate": "05/11/2024 14:57:22",
      "content": "<p>yeah I think there is quantization going on under the hood with openvino </p>",
      "rawMarkdown": "yeah I think there is quantization going on under the hood with openvino",
      "votes": null
    },
    {
      "id": "2811919",
      "postDate": "05/14/2024 02:01:33",
      "content": "<p>What's your batch_size and num_of_workers? BTW are you open to teaming up (unable to contact you by email)?</p>",
      "rawMarkdown": "What's your batch_size and num_of_workers? BTW are you open to teaming up (unable to contact you by email)?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2804552,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "05/10/2024 05:30:09",
      "content": "<p>Probably predictor itself is using all cores already, so no reason to use <code>concurrent.futures</code> here</p>",
      "votes": null,
      "replies": [
        {
          "id": 2804661,
          "author_name": "yuanzhezhou",
          "author_url": "",
          "post_date": "05/10/2024 06:28:03",
          "content": "<p>That's true, interestingly increasing batchsize decrease inference speed for jit by a lot.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2805860,
              "author_name": "willrice",
              "author_url": "",
              "post_date": "05/10/2024 19:10:31",
              "content": "<p>have you tried with <code>torch.jit.optimized_execution(False)</code>. I've noticed for some of the backbones this speeds up inference.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2806300,
                  "author_name": "yuanzhezhou",
                  "author_url": "",
                  "post_date": "05/11/2024 02:52:38",
                  "content": "<p>I used it. From my local test, onnx is twice faster than jit. But not this case in kaggle notebook. It's very weird.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2805202,
      "author_name": "pingfan",
      "author_url": "",
      "post_date": "05/10/2024 12:36:46",
      "content": "<p>What is the acceleration from base to onnx inference? My base reasoning is the same speed as onnx</p>",
      "votes": null,
      "replies": [
        {
          "id": 2806301,
          "author_name": "yuanzhezhou",
          "author_url": "",
          "post_date": "05/11/2024 02:53:07",
          "content": "<p>Local test shows that onnx  is twice (at most) faster than jit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2806948,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/11/2024 12:19:25",
      "content": "<p>I get some speedup form onnx, but not a lot. Openvino is much faster, like 2x fatser,  but I lose some accuracy on LB.</p>\n<p>I am not sure why given others claim openvino works just fine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2807160,
          "author_name": "willrice",
          "author_url": "",
          "post_date": "05/11/2024 14:57:22",
          "content": "<p>yeah I think there is quantization going on under the hood with openvino </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2811919,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "05/14/2024 02:01:33",
      "content": "<p>What's your batch_size and num_of_workers? BTW are you open to teaming up (unable to contact you by email)?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2804446": "I did transformations offline with my computer. It seems that I did not get improvement of inference speed from jit \n to onnx/openvino. It can be even slower. What can be the reason for it?\n\nI used the following code for inference\n\n```python\nwith concurrent.futures.ThreadPoolExecutor(max_workers=4) as executor:\n        dicts = list(executor.map(prediction_for_clip, all_audios))\n```",
    "2804552": "Probably predictor itself is using all cores already, so no reason to use `concurrent.futures` here",
    "2804661": "That's true, interestingly increasing batchsize decrease inference speed for jit by a lot.",
    "2805202": "What is the acceleration from base to onnx inference? My base reasoning is the same speed as onnx",
    "2805860": "have you tried with `torch.jit.optimized_execution(False)`. I've noticed for some of the backbones this speeds up inference.",
    "2806300": "I used it. From my local test, onnx is twice faster than jit. But not this case in kaggle notebook. It's very weird.",
    "2806301": "Local test shows that onnx  is twice (at most) faster than jit.",
    "2806948": "I get some speedup form onnx, but not a lot. Openvino is much faster, like 2x fatser,  but I lose some accuracy on LB.\n\nI am not sure why given others claim openvino works just fine.",
    "2807160": "yeah I think there is quantization going on under the hood with openvino",
    "2811919": "What's your batch_size and num_of_workers? BTW are you open to teaming up (unable to contact you by email)?"
  },
  "source": "meta"
}