{
  "id": 492649,
  "title": "Optimize Torch Model For Inference | 2X with 1 LOC!",
  "url": "/competitions/birdclef-2024/discussion/492649",
  "author_name": "",
  "post_date": "2024-04-10T12:47:20.199904600Z",
  "votes": 28,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This competition has strict rules when it comes to inference time.<br>\nWith only 120 minutes on CPU for 1100 samples of 4 minutes efficiency is key.</p>\n<p>Luckily Torch provides an easy solution which speeds up models for inference:</p>\n<pre><code>\n = torch.load(MODEL_PATH, map_location=torch.device())\n\n = torch.jit.optimize_for_inference(torch.jit.script(model.eval()))\n</code></pre>\n<p>On EfficientVit-B1 the speedup is over 2x on a subset of 100 samples, reducing the time from 612s → 301s (-51%)</p>",
  "messages": [
    {
      "id": "2745175",
      "postDate": "04/10/2024 12:47:20",
      "content": "<p>This competition has strict rules when it comes to inference time.<br>\nWith only 120 minutes on CPU for 1100 samples of 4 minutes efficiency is key.</p>\n<p>Luckily Torch provides an easy solution which speeds up models for inference:</p>\n<pre><code>\n = torch.load(MODEL_PATH, map_location=torch.device())\n\n = torch.jit.optimize_for_inference(torch.jit.script(model.eval()))\n</code></pre>\n<p>On EfficientVit-B1 the speedup is over 2x on a subset of 100 samples, reducing the time from 612s → 301s (-51%)</p>",
      "rawMarkdown": "This competition has strict rules when it comes to inference time.\nWith only 120 minutes on CPU for 1100 samples of 4 minutes efficiency is key.\n\nLuckily Torch provides an easy solution which speeds up models for inference:\n\n```\n# Load Your Model\nmodel = torch.load(MODEL_PATH, map_location=torch.device('cpu'))\n# Magic Line Of Code That Optimizes The Model For Inference\nmodel = torch.jit.optimize_for_inference(torch.jit.script(model.eval()))\n```\n\nOn EfficientVit-B1 the speedup is over 2x on a subset of 100 samples, reducing the time from 612s → 301s (-51%)",
      "votes": null
    },
    {
      "id": "2745185",
      "postDate": "04/10/2024 12:56:19",
      "content": "<p>That sounds great, I didn't know this at all.</p>\n<p>You can also take a look at ONNX:<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412996\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412996</a></p>",
      "rawMarkdown": "That sounds great, I didn't know this at all.\n\nYou can also take a look at ONNX:\nhttps://www.kaggle.com/competitions/birdclef-2023/discussion/412996",
      "votes": null
    },
    {
      "id": "2745190",
      "postDate": "04/10/2024 13:02:08",
      "content": "<p>I just tested with a efficientnet-b0 on a subset of 100 of the unlabelled soundscapes and I don't get a speed boost (in kaggle env). I'll look into why that is and update here later</p>",
      "rawMarkdown": "I just tested with a efficientnet-b0 on a subset of 100 of the unlabelled soundscapes and I don't get a speed boost (in kaggle env). I'll look into why that is and update here later",
      "votes": null
    },
    {
      "id": "2745266",
      "postDate": "04/10/2024 14:09:16",
      "content": "<p>Are you performing inference on CPU or GPU?<br>\nWhat might also play a role is that EfficientNet is a CNN and EfficientViT is a Transformer.</p>",
      "rawMarkdown": "Are you performing inference on CPU or GPU?\nWhat might also play a role is that EfficientNet is a CNN and EfficientViT is a Transformer.",
      "votes": null
    },
    {
      "id": "2745332",
      "postDate": "04/10/2024 14:54:10",
      "content": "<p>Inference on CPU, Yes that is what I was thinking as well. I don't have time to look up what difference this function makes but I suspect this is something along those lines.</p>",
      "rawMarkdown": "Inference on CPU, Yes that is what I was thinking as well. I don't have time to look up what difference this function makes but I suspect this is something along those lines.",
      "votes": null
    },
    {
      "id": "2747030",
      "postDate": "04/11/2024 16:00:42",
      "content": "<p>+3s every 10 samples on my server😂</p>",
      "rawMarkdown": "3s every 10 samples on my server😂",
      "votes": null
    },
    {
      "id": "2747342",
      "postDate": "04/11/2024 20:05:41",
      "content": "<p>Are we allowed to do inference on GPU?</p>",
      "rawMarkdown": "Are we allowed to do inference on GPU?",
      "votes": null
    },
    {
      "id": "2747449",
      "postDate": "04/11/2024 22:19:21",
      "content": "<p><a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> i havee tried using your notebook for inference of my efficientViT model and have faced timeouts could there be a exact reason as to why they are occuring</p>",
      "rawMarkdown": "markwijkhuizen i havee tried using your notebook for inference of my efficientViT model and have faced timeouts could there be a exact reason as to why they are occuring",
      "votes": null
    },
    {
      "id": "2747456",
      "postDate": "04/11/2024 22:36:23",
      "content": "<p>no, you can't</p>",
      "rawMarkdown": "no, you can't",
      "votes": null
    },
    {
      "id": "2759885",
      "postDate": "04/19/2024 01:12:31",
      "content": "<p>Thank you for the interesting share. I compared inference times using EfficientNet-B3, and in my environment, ONNX was faster than using jit.optimize_for_inference.</p>",
      "rawMarkdown": "Thank you for the interesting share. I compared inference times using EfficientNet-B3, and in my environment, ONNX was faster than using jit.optimize_for_inference.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2745185,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "04/10/2024 12:56:19",
      "content": "<p>That sounds great, I didn't know this at all.</p>\n<p>You can also take a look at ONNX:<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/412996\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2023/discussion/412996</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2745190,
          "author_name": "janmpia",
          "author_url": "",
          "post_date": "04/10/2024 13:02:08",
          "content": "<p>I just tested with a efficientnet-b0 on a subset of 100 of the unlabelled soundscapes and I don't get a speed boost (in kaggle env). I'll look into why that is and update here later</p>",
          "votes": null,
          "replies": [
            {
              "id": 2745266,
              "author_name": "markwijkhuizen",
              "author_url": "",
              "post_date": "04/10/2024 14:09:16",
              "content": "<p>Are you performing inference on CPU or GPU?<br>\nWhat might also play a role is that EfficientNet is a CNN and EfficientViT is a Transformer.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2745332,
                  "author_name": "janmpia",
                  "author_url": "",
                  "post_date": "04/10/2024 14:54:10",
                  "content": "<p>Inference on CPU, Yes that is what I was thinking as well. I don't have time to look up what difference this function makes but I suspect this is something along those lines.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2747342,
                  "author_name": "jose091",
                  "author_url": "",
                  "post_date": "04/11/2024 20:05:41",
                  "content": "<p>Are we allowed to do inference on GPU?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2747456,
                      "author_name": "janmpia",
                      "author_url": "",
                      "post_date": "04/11/2024 22:36:23",
                      "content": "<p>no, you can't</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2747030,
      "author_name": "seeingtimes",
      "author_url": "",
      "post_date": "04/11/2024 16:00:42",
      "content": "<p>+3s every 10 samples on my server😂</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2747449,
      "author_name": "manav2805",
      "author_url": "",
      "post_date": "04/11/2024 22:19:21",
      "content": "<p><a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> i havee tried using your notebook for inference of my efficientViT model and have faced timeouts could there be a exact reason as to why they are occuring</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2759885,
      "author_name": "kmatsu01",
      "author_url": "",
      "post_date": "04/19/2024 01:12:31",
      "content": "<p>Thank you for the interesting share. I compared inference times using EfficientNet-B3, and in my environment, ONNX was faster than using jit.optimize_for_inference.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2745175": "This competition has strict rules when it comes to inference time.\nWith only 120 minutes on CPU for 1100 samples of 4 minutes efficiency is key.\n\nLuckily Torch provides an easy solution which speeds up models for inference:\n\n```\n# Load Your Model\nmodel = torch.load(MODEL_PATH, map_location=torch.device('cpu'))\n# Magic Line Of Code That Optimizes The Model For Inference\nmodel = torch.jit.optimize_for_inference(torch.jit.script(model.eval()))\n```\n\nOn EfficientVit-B1 the speedup is over 2x on a subset of 100 samples, reducing the time from 612s → 301s (-51%)",
    "2745185": "That sounds great, I didn't know this at all.\n\nYou can also take a look at ONNX:\nhttps://www.kaggle.com/competitions/birdclef-2023/discussion/412996",
    "2745190": "I just tested with a efficientnet-b0 on a subset of 100 of the unlabelled soundscapes and I don't get a speed boost (in kaggle env). I'll look into why that is and update here later",
    "2745266": "Are you performing inference on CPU or GPU?\nWhat might also play a role is that EfficientNet is a CNN and EfficientViT is a Transformer.",
    "2745332": "Inference on CPU, Yes that is what I was thinking as well. I don't have time to look up what difference this function makes but I suspect this is something along those lines.",
    "2747030": "3s every 10 samples on my server😂",
    "2747342": "Are we allowed to do inference on GPU?",
    "2747449": "markwijkhuizen i havee tried using your notebook for inference of my efficientViT model and have faced timeouts could there be a exact reason as to why they are occuring",
    "2747456": "no, you can't",
    "2759885": "Thank you for the interesting share. I compared inference times using EfficientNet-B3, and in my environment, ONNX was faster than using jit.optimize_for_inference."
  },
  "source": "meta"
}