{
  "id": 556561,
  "title": "Converting PyTorch Checkpoints to TensorRT Models",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/556561",
  "author_name": "",
  "post_date": "2025-01-14T03:29:41.954188400Z",
  "votes": 23,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Converting PyTorch checkpoints to TensorRT models can be time-consuming in Kaggle, especially when building the necessary wheels and finding the most effective conversion method. After some trial and error, I’ve found a solution that works, and I wanted to share it with the community in the hope that it will help others.</p>\n<p>This approach can accelerate model inference by around 30%,  with no  loss in public lb score. Use it with caution and ensure it aligns with your specific needs.</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models\" target=\"_blank\">https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models</a></p>\n</blockquote>",
  "messages": [
    {
      "id": "3096051",
      "postDate": "01/14/2025 03:29:41",
      "content": "<p>Converting PyTorch checkpoints to TensorRT models can be time-consuming in Kaggle, especially when building the necessary wheels and finding the most effective conversion method. After some trial and error, I’ve found a solution that works, and I wanted to share it with the community in the hope that it will help others.</p>\n<p>This approach can accelerate model inference by around 30%,  with no  loss in public lb score. Use it with caution and ensure it aligns with your specific needs.</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models\" target=\"_blank\">https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models</a></p>\n</blockquote>",
      "rawMarkdown": "Converting PyTorch checkpoints to TensorRT models can be time-consuming in Kaggle, especially when building the necessary wheels and finding the most effective conversion method. After some trial and error, I’ve found a solution that works, and I wanted to share it with the community in the hope that it will help others.\n\nThis approach can accelerate model inference by around 30%, ~~it may slightly reduce your public score by approximately 0.002~~ with no  loss in public lb score. Use it with caution and ensure it aligns with your specific needs.\n\n>https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models",
      "votes": null
    },
    {
      "id": "3096059",
      "postDate": "01/14/2025 03:49:26",
      "content": "<p>Thank you very much for sharing the great code! 😀</p>\n<p>By the way, I was wondering if the <code>BSD-3-Clause</code> license used by torch-tensorrt is allowed on this competition.</p>",
      "rawMarkdown": "Thank you very much for sharing the great code! 😀\n\nBy the way, I was wondering if the `BSD-3-Clause` license used by torch-tensorrt is allowed on this competition.",
      "votes": null
    },
    {
      "id": "3096936",
      "postDate": "01/14/2025 20:44:50",
      "content": "<p>Thank you for sharing! The link seems to be broken though</p>",
      "rawMarkdown": "Thank you for sharing! The link seems to be broken though",
      "votes": null
    },
    {
      "id": "3096955",
      "postDate": "01/14/2025 21:06:48",
      "content": "<p>I encountered this problem as well. However, when I tried again, I didn't encounter it.😀 Maybe copying the link is a good idea.</p>",
      "rawMarkdown": "I encountered this problem as well. However, when I tried again, I didn't encounter it.😀 Maybe copying the link is a good idea.",
      "votes": null
    },
    {
      "id": "3096966",
      "postDate": "01/14/2025 21:27:11",
      "content": "<p>Interesting, it works for me too now :D </p>",
      "rawMarkdown": "Interesting, it works for me too now :D",
      "votes": null
    },
    {
      "id": "3097454",
      "postDate": "01/15/2025 11:39:48",
      "content": "<p>Apologies for the confusion. We use TensorRT, licensed under the Apache license, and Torch2TRT, licensed under the MIT license, instead of Torch-TensorRT.</p>",
      "rawMarkdown": "Apologies for the confusion. We use TensorRT, licensed under the Apache license, and Torch2TRT, licensed under the MIT license, instead of Torch-TensorRT.",
      "votes": null
    },
    {
      "id": "3097521",
      "postDate": "01/15/2025 13:20:01",
      "content": "<p>Thank you very much for your sharing. Could you please share the dataset <code>/kaggle/input/tensorrt-10-1-0?</code> I would really appreciate it.</p>",
      "rawMarkdown": "Thank you very much for your sharing. Could you please share the dataset `/kaggle/input/tensorrt-10-1-0?` I would really appreciate it.",
      "votes": null
    },
    {
      "id": "3097528",
      "postDate": "01/15/2025 13:25:53",
      "content": "<p>Sorry for not making the dataset publicly available. You can now access it.</p>",
      "rawMarkdown": "Sorry for not making the dataset publicly available. You can now access it.",
      "votes": null
    },
    {
      "id": "3098111",
      "postDate": "01/16/2025 05:09:59",
      "content": "<p>Great Thanks!</p>",
      "rawMarkdown": "Great Thanks!",
      "votes": null
    },
    {
      "id": "3099568",
      "postDate": "01/17/2025 23:10:11",
      "content": "<p>Thank you for sharing this solution—it’s incredibly helpful! A 30% boost in inference speed is impressive.</p>",
      "rawMarkdown": "Thank you for sharing this solution—it’s incredibly helpful! A 30% boost in inference speed is impressive.",
      "votes": null
    },
    {
      "id": "3102546",
      "postDate": "01/22/2025 09:50:01",
      "content": "<p>Thank you very much for sharing the great code!</p>\n<p>May I ask two questions?<br>\nFirst, have you successfully submitted a competition in this notebook environment (GPU T4)? I am not sure why I am getting errors when submitting and am unable to submit.<br>\nSecondly, is it possible to create a sensorrt wheel file for the P100 or can you tell me how to make one?</p>",
      "rawMarkdown": "Thank you very much for sharing the great code!\n\nMay I ask two questions?\nFirst, have you successfully submitted a competition in this notebook environment (GPU T4)? I am not sure why I am getting errors when submitting and am unable to submit.\nSecondly, is it possible to create a sensorrt wheel file for the P100 or can you tell me how to make one?",
      "votes": null
    },
    {
      "id": "3102560",
      "postDate": "01/22/2025 10:19:12",
      "content": "<p>Yes, I have successfully submitted TensorRT models using a setup with two T4 GPUs and did not encounter any errors during the process. If possible, could you share the specific error message you encountered? That would help diagnose the issue more effectively.</p>\n<p>May I ask why you prefer the P100 over the T4? The T4 is considerably faster for FP16 inference, while the P100 performs better with FP32 inference. Nonetheless, the process of creating and using TensorRT wheel files for the P100 should be similar to that for the T4.</p>",
      "rawMarkdown": "Yes, I have successfully submitted TensorRT models using a setup with two T4 GPUs and did not encounter any errors during the process. If possible, could you share the specific error message you encountered? That would help diagnose the issue more effectively.\n\nMay I ask why you prefer the P100 over the T4? The T4 is considerably faster for FP16 inference, while the P100 performs better with FP32 inference. Nonetheless, the process of creating and using TensorRT wheel files for the P100 should be similar to that for the T4.",
      "votes": null
    },
    {
      "id": "3102583",
      "postDate": "01/22/2025 10:58:13",
      "content": "<p>Thank you for your reply!<br>\nI'm glad to confirm that submissions are working in your environment. In my case, it runs without any issues when executing the notebook, but it stops during submission with a \"Notebook Threw Exception\" error. The only change made was switching the GPU from P100 to T4, and even submissions that previously worked are now failing. Do you have any insights on this issue?</p>\n<p>Also, does the P100 not benefit much from float16? I'm considering converting to TensorRT and performing inference with float16 on the P100, but is it unlikely to see speed improvements in that case?</p>",
      "rawMarkdown": "Thank you for your reply!\nI'm glad to confirm that submissions are working in your environment. In my case, it runs without any issues when executing the notebook, but it stops during submission with a \"Notebook Threw Exception\" error. The only change made was switching the GPU from P100 to T4, and even submissions that previously worked are now failing. Do you have any insights on this issue?\n\nAlso, does the P100 not benefit much from float16? I'm considering converting to TensorRT and performing inference with float16 on the P100, but is it unlikely to see speed improvements in that case?",
      "votes": null
    },
    {
      "id": "3102591",
      "postDate": "01/22/2025 11:09:31",
      "content": "<p>I also encountered the same issue where even <strong>submissions that previously worked are now failing</strong>. I’m unsure what might be causing this, and it does seem quite unusual.</p>\n<p>Since you have access to 2xT4 GPUs, you could configure your program to run in parallel. With TensorRT acceleration, this should significantly enhance processing speed.</p>\n<p>As for the P100, while it can benefit from FP16 inference, the performance improvement is not as pronounced as with the T4.</p>",
      "rawMarkdown": "I also encountered the same issue where even **submissions that previously worked are now failing**. I’m unsure what might be causing this, and it does seem quite unusual.\n\nSince you have access to 2xT4 GPUs, you could configure your program to run in parallel. With TensorRT acceleration, this should significantly enhance processing speed.\n\nAs for the P100, while it can benefit from FP16 inference, the performance improvement is not as pronounced as with the T4.",
      "votes": null
    },
    {
      "id": "3102593",
      "postDate": "01/22/2025 11:14:35",
      "content": "<p>Thank you for your detailed response! I will investigate the issue a bit further. Just knowing that submissions work in your environment is already helpful. If you notice anything else, even just a small detail, I would greatly appreciate it if you could share it with me.</p>",
      "rawMarkdown": "Thank you for your detailed response! I will investigate the issue a bit further. Just knowing that submissions work in your environment is already helpful. If you notice anything else, even just a small detail, I would greatly appreciate it if you could share it with me.",
      "votes": null
    },
    {
      "id": "3102595",
      "postDate": "01/22/2025 11:24:34",
      "content": "<p>Moreover, in your case, it might be helpful to debug whether the issue is related to GPU memory limitations. The P100 has 16 GB of GPU memory, while the T4 has only 15 GB, which could potentially cause memory-related issues during processing.</p>",
      "rawMarkdown": "Moreover, in your case, it might be helpful to debug whether the issue is related to GPU memory limitations. The P100 has 16 GB of GPU memory, while the T4 has only 15 GB, which could potentially cause memory-related issues during processing.",
      "votes": null
    },
    {
      "id": "3102904",
      "postDate": "01/22/2025 19:10:41",
      "content": "<p>A test that has helped me debug these things is to run the script without submitting, but make a test set of 500 tomograms by repeating the original test set several times. You can then run the script and see the error that pops up (in my case I found a memory leak this way).</p>",
      "rawMarkdown": "A test that has helped me debug these things is to run the script without submitting, but make a test set of 500 tomograms by repeating the original test set several times. You can then run the script and see the error that pops up (in my case I found a memory leak this way).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3096059,
      "author_name": "siwooyong",
      "author_url": "",
      "post_date": "01/14/2025 03:49:26",
      "content": "<p>Thank you very much for sharing the great code! 😀</p>\n<p>By the way, I was wondering if the <code>BSD-3-Clause</code> license used by torch-tensorrt is allowed on this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3097454,
          "author_name": "sjtuwangshuo",
          "author_url": "",
          "post_date": "01/15/2025 11:39:48",
          "content": "<p>Apologies for the confusion. We use TensorRT, licensed under the Apache license, and Torch2TRT, licensed under the MIT license, instead of Torch-TensorRT.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3096936,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "01/14/2025 20:44:50",
      "content": "<p>Thank you for sharing! The link seems to be broken though</p>",
      "votes": null,
      "replies": [
        {
          "id": 3096955,
          "author_name": "luoziqian",
          "author_url": "",
          "post_date": "01/14/2025 21:06:48",
          "content": "<p>I encountered this problem as well. However, when I tried again, I didn't encounter it.😀 Maybe copying the link is a good idea.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3096966,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "01/14/2025 21:27:11",
              "content": "<p>Interesting, it works for me too now :D </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3097521,
      "author_name": "peilwang",
      "author_url": "",
      "post_date": "01/15/2025 13:20:01",
      "content": "<p>Thank you very much for your sharing. Could you please share the dataset <code>/kaggle/input/tensorrt-10-1-0?</code> I would really appreciate it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3097528,
          "author_name": "sjtuwangshuo",
          "author_url": "",
          "post_date": "01/15/2025 13:25:53",
          "content": "<p>Sorry for not making the dataset publicly available. You can now access it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3098111,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "01/16/2025 05:09:59",
      "content": "<p>Great Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3099568,
      "author_name": "iamramzanai",
      "author_url": "",
      "post_date": "01/17/2025 23:10:11",
      "content": "<p>Thank you for sharing this solution—it’s incredibly helpful! A 30% boost in inference speed is impressive.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3102546,
      "author_name": "tkitagawa19999",
      "author_url": "",
      "post_date": "01/22/2025 09:50:01",
      "content": "<p>Thank you very much for sharing the great code!</p>\n<p>May I ask two questions?<br>\nFirst, have you successfully submitted a competition in this notebook environment (GPU T4)? I am not sure why I am getting errors when submitting and am unable to submit.<br>\nSecondly, is it possible to create a sensorrt wheel file for the P100 or can you tell me how to make one?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3102560,
          "author_name": "sjtuwangshuo",
          "author_url": "",
          "post_date": "01/22/2025 10:19:12",
          "content": "<p>Yes, I have successfully submitted TensorRT models using a setup with two T4 GPUs and did not encounter any errors during the process. If possible, could you share the specific error message you encountered? That would help diagnose the issue more effectively.</p>\n<p>May I ask why you prefer the P100 over the T4? The T4 is considerably faster for FP16 inference, while the P100 performs better with FP32 inference. Nonetheless, the process of creating and using TensorRT wheel files for the P100 should be similar to that for the T4.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3102583,
              "author_name": "tkitagawa19999",
              "author_url": "",
              "post_date": "01/22/2025 10:58:13",
              "content": "<p>Thank you for your reply!<br>\nI'm glad to confirm that submissions are working in your environment. In my case, it runs without any issues when executing the notebook, but it stops during submission with a \"Notebook Threw Exception\" error. The only change made was switching the GPU from P100 to T4, and even submissions that previously worked are now failing. Do you have any insights on this issue?</p>\n<p>Also, does the P100 not benefit much from float16? I'm considering converting to TensorRT and performing inference with float16 on the P100, but is it unlikely to see speed improvements in that case?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3102591,
                  "author_name": "sjtuwangshuo",
                  "author_url": "",
                  "post_date": "01/22/2025 11:09:31",
                  "content": "<p>I also encountered the same issue where even <strong>submissions that previously worked are now failing</strong>. I’m unsure what might be causing this, and it does seem quite unusual.</p>\n<p>Since you have access to 2xT4 GPUs, you could configure your program to run in parallel. With TensorRT acceleration, this should significantly enhance processing speed.</p>\n<p>As for the P100, while it can benefit from FP16 inference, the performance improvement is not as pronounced as with the T4.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3102593,
                      "author_name": "tkitagawa19999",
                      "author_url": "",
                      "post_date": "01/22/2025 11:14:35",
                      "content": "<p>Thank you for your detailed response! I will investigate the issue a bit further. Just knowing that submissions work in your environment is already helpful. If you notice anything else, even just a small detail, I would greatly appreciate it if you could share it with me.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3102595,
                          "author_name": "sjtuwangshuo",
                          "author_url": "",
                          "post_date": "01/22/2025 11:24:34",
                          "content": "<p>Moreover, in your case, it might be helpful to debug whether the issue is related to GPU memory limitations. The P100 has 16 GB of GPU memory, while the T4 has only 15 GB, which could potentially cause memory-related issues during processing.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3102904,
                              "author_name": "jeroencottaar",
                              "author_url": "",
                              "post_date": "01/22/2025 19:10:41",
                              "content": "<p>A test that has helped me debug these things is to run the script without submitting, but make a test set of 500 tomograms by repeating the original test set several times. You can then run the script and see the error that pops up (in my case I found a memory leak this way).</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3096051": "Converting PyTorch checkpoints to TensorRT models can be time-consuming in Kaggle, especially when building the necessary wheels and finding the most effective conversion method. After some trial and error, I’ve found a solution that works, and I wanted to share it with the community in the hope that it will help others.\n\nThis approach can accelerate model inference by around 30%, ~~it may slightly reduce your public score by approximately 0.002~~ with no  loss in public lb score. Use it with caution and ensure it aligns with your specific needs.\n\n>https://www.kaggle.com/code/sjtuwangshuo/converting-pytorch-checkpoints-to-tensorrt-models",
    "3096059": "Thank you very much for sharing the great code! 😀\n\nBy the way, I was wondering if the `BSD-3-Clause` license used by torch-tensorrt is allowed on this competition.",
    "3096936": "Thank you for sharing! The link seems to be broken though",
    "3096955": "I encountered this problem as well. However, when I tried again, I didn't encounter it.😀 Maybe copying the link is a good idea.",
    "3096966": "Interesting, it works for me too now :D",
    "3097454": "Apologies for the confusion. We use TensorRT, licensed under the Apache license, and Torch2TRT, licensed under the MIT license, instead of Torch-TensorRT.",
    "3097521": "Thank you very much for your sharing. Could you please share the dataset `/kaggle/input/tensorrt-10-1-0?` I would really appreciate it.",
    "3097528": "Sorry for not making the dataset publicly available. You can now access it.",
    "3098111": "Great Thanks!",
    "3099568": "Thank you for sharing this solution—it’s incredibly helpful! A 30% boost in inference speed is impressive.",
    "3102546": "Thank you very much for sharing the great code!\n\nMay I ask two questions?\nFirst, have you successfully submitted a competition in this notebook environment (GPU T4)? I am not sure why I am getting errors when submitting and am unable to submit.\nSecondly, is it possible to create a sensorrt wheel file for the P100 or can you tell me how to make one?",
    "3102560": "Yes, I have successfully submitted TensorRT models using a setup with two T4 GPUs and did not encounter any errors during the process. If possible, could you share the specific error message you encountered? That would help diagnose the issue more effectively.\n\nMay I ask why you prefer the P100 over the T4? The T4 is considerably faster for FP16 inference, while the P100 performs better with FP32 inference. Nonetheless, the process of creating and using TensorRT wheel files for the P100 should be similar to that for the T4.",
    "3102583": "Thank you for your reply!\nI'm glad to confirm that submissions are working in your environment. In my case, it runs without any issues when executing the notebook, but it stops during submission with a \"Notebook Threw Exception\" error. The only change made was switching the GPU from P100 to T4, and even submissions that previously worked are now failing. Do you have any insights on this issue?\n\nAlso, does the P100 not benefit much from float16? I'm considering converting to TensorRT and performing inference with float16 on the P100, but is it unlikely to see speed improvements in that case?",
    "3102591": "I also encountered the same issue where even **submissions that previously worked are now failing**. I’m unsure what might be causing this, and it does seem quite unusual.\n\nSince you have access to 2xT4 GPUs, you could configure your program to run in parallel. With TensorRT acceleration, this should significantly enhance processing speed.\n\nAs for the P100, while it can benefit from FP16 inference, the performance improvement is not as pronounced as with the T4.",
    "3102593": "Thank you for your detailed response! I will investigate the issue a bit further. Just knowing that submissions work in your environment is already helpful. If you notice anything else, even just a small detail, I would greatly appreciate it if you could share it with me.",
    "3102595": "Moreover, in your case, it might be helpful to debug whether the issue is related to GPU memory limitations. The P100 has 16 GB of GPU memory, while the T4 has only 15 GB, which could potentially cause memory-related issues during processing.",
    "3102904": "A test that has helped me debug these things is to run the script without submitting, but make a test set of 500 tomograms by repeating the original test set several times. You can then run the script and see the error that pops up (in my case I found a memory leak this way)."
  },
  "source": "meta"
}