{
  "id": 556741,
  "title": "How to reduce runtime?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/556741",
  "author_name": "",
  "post_date": "2025-01-14T21:28:02.881297500Z",
  "votes": null,
  "comment_count": 14,
  "views": 0,
  "content": "<p>The competition requires that runtime &lt;= 12 hours. I am trying to reduce the runtime of a submission notebook. <strong>How minimal can a submission code be</strong>? Can I use pre trained weights instead of training in the notebook? Do I also need to make inference in the same notebook?</p>",
  "messages": [
    {
      "id": "3096967",
      "postDate": "01/14/2025 21:28:02",
      "content": "<p>The competition requires that runtime &lt;= 12 hours. I am trying to reduce the runtime of a submission notebook. <strong>How minimal can a submission code be</strong>? Can I use pre trained weights instead of training in the notebook? Do I also need to make inference in the same notebook?</p>",
      "rawMarkdown": "The competition requires that runtime <= 12 hours. I am trying to reduce the runtime of a submission notebook. **How minimal can a submission code be**? Can I use pre trained weights instead of training in the notebook? Do I also need to make inference in the same notebook?",
      "votes": null
    },
    {
      "id": "3096997",
      "postDate": "01/14/2025 22:56:34",
      "content": "<p>I do two things:</p>\n<ul>\n<li>Train in one notebook and do inference in another.</li>\n<li>Use the T4 x 2 GPU.  Although this takes some extra code to setup, but it's pretty easy to google.</li>\n</ul>\n<p>Note, this is a pretty good discussion post if you're just getting started:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715</a></p>",
      "rawMarkdown": "I do two things:\n- Train in one notebook and do inference in another.\n- Use the T4 x 2 GPU.  Although this takes some extra code to setup, but it's pretty easy to google.\n\nNote, this is a pretty good discussion post if you're just getting started:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715)",
      "votes": null
    },
    {
      "id": "3097459",
      "postDate": "01/15/2025 11:46:13",
      "content": "<p>Question, as I’m using P100 and I lack a coding/software background, T4 x 2 makes use of extra parallelization, thus if I’m running one batch at a time, it wouldn’t really help me, unless I’d run two batches at a time, is my understanding correct?</p>",
      "rawMarkdown": "Question, as I’m using P100 and I lack a coding/software background, T4 x 2 makes use of extra parallelization, thus if I’m running one batch at a time, it wouldn’t really help me, unless I’d run two batches at a time, is my understanding correct?",
      "votes": null
    },
    {
      "id": "3097533",
      "postDate": "01/15/2025 13:36:00",
      "content": "<p>I supposed you chose to work on high resolution with shape (184,630,630). Your code will be faster to execute on a medium resolution with shape (92,135,135), but I supposed it comes with drawback, as loss of information, and so, less effective prediction. </p>",
      "rawMarkdown": "I supposed you chose to work on high resolution with shape (184,630,630). Your code will be faster to execute on a medium resolution with shape (92,135,135), but I supposed it comes with drawback, as loss of information, and so, less effective prediction.",
      "votes": null
    },
    {
      "id": "3097544",
      "postDate": "01/15/2025 13:43:34",
      "content": "<p>Yes, you need to run 2 batches of data each time. You can run 4 batches during inference, which is convenient for using TTA during inference. For example, one batch is the original image, one batch is horizontally flipped, one batch is rotated 90 degrees, and so on.</p>",
      "rawMarkdown": "Yes, you need to run 2 batches of data each time. You can run 4 batches during inference, which is convenient for using TTA during inference. For example, one batch is the original image, one batch is horizontally flipped, one batch is rotated 90 degrees, and so on.",
      "votes": null
    },
    {
      "id": "3097553",
      "postDate": "01/15/2025 13:51:14",
      "content": "<p><a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a> We also use it to be able to do inference with big volume sizes in somehow big models. </p>",
      "rawMarkdown": "andreizamfir We also use it to be able to do inference with big volume sizes in somehow big models.",
      "votes": null
    },
    {
      "id": "3099268",
      "postDate": "01/17/2025 14:23:42",
      "content": "<p>how do you setup T4 x 2 GPU</p>",
      "rawMarkdown": "how do you setup T4 x 2 GPU",
      "votes": null
    },
    {
      "id": "3109228",
      "postDate": "01/28/2025 16:18:31",
      "content": "<p>I highly recommend training in a different notebook than what you run inference on.  Train in notebook A, save the weights to a new kaggle dataset, load those weights in the inference notebook and run only inference.  Also make sure you are making use of the GPUs and that your data and model on are on the cuda device.  With no data parallelization or other optimizations this alone should get your runtime low enough to make a submission.  My notebook takes around 8-9 hours with absolutely no optimizations.  Make sure you aren't doing lots of preprocessing that could take a long time when the dataset size goes up to 500.  I recommend benchmarking on the 3 samples and multiply the runtime out so you can estimate the total submission time.</p>\n<p>It could also help to look at current submission notebooks and see what they do to run it within the allowed time.</p>",
      "rawMarkdown": "I highly recommend training in a different notebook than what you run inference on.  Train in notebook A, save the weights to a new kaggle dataset, load those weights in the inference notebook and run only inference.  Also make sure you are making use of the GPUs and that your data and model on are on the cuda device.  With no data parallelization or other optimizations this alone should get your runtime low enough to make a submission.  My notebook takes around 8-9 hours with absolutely no optimizations.  Make sure you aren't doing lots of preprocessing that could take a long time when the dataset size goes up to 500.  I recommend benchmarking on the 3 samples and multiply the runtime out so you can estimate the total submission time.\n\nIt could also help to look at current submission notebooks and see what they do to run it within the allowed time.",
      "votes": null
    },
    {
      "id": "3116437",
      "postDate": "02/06/2025 00:53:29",
      "content": "<p><a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a> If you use sliding_window_inference() and wrap your model with DataParallel() it runs the patches in parallel if you have more than one GPU:</p>\n<pre><code>model = torch.nn.DataParallel(model, device_ids=[,])\nmodel.to()\n\n ():\n     ():\n         sliding_window_inference(\n            inputs=,\n            roi_size=(, , ),\n            sw_batch_size=,  \n            predictor=model,\n            overlap=,\n            \n            roi_weight_map=weight_map\n        )\n\n    \n     torch.amp.autocast():\n         _compute()\n</code></pre>\n<p>Note the sliding_window_inference() is used in one of the example notebooks from the CZII Data Portal.  <a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/inference.ipynb\" target=\"_blank\">https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/inference.ipynb</a></p>\n<p>P.S.  Sorry for the delayed response, but you were way ahead of me at the time.  :)</p>",
      "rawMarkdown": "andreizamfir If you use sliding_window_inference() and wrap your model with DataParallel() it runs the patches in parallel if you have more than one GPU:\n\n```\nmodel = torch.nn.DataParallel(model, device_ids=[0,1])\nmodel.to('cuda:0')\n\ndef inference(model, input):\n    def _compute(input):\n        return sliding_window_inference(\n            inputs=input,\n            roi_size=(96, 96, 96),\n            sw_batch_size=4,  # one window is proecessed at a time\n            predictor=model,\n            overlap=0.125,\n            #mode='gaussian',\n            roi_weight_map=weight_map\n        )\n\n    # with torch.cuda.amp.autocast():\n    with torch.amp.autocast('cuda'):\n        return _compute(input)\n```\n\nNote the sliding_window_inference() is used in one of the example notebooks from the CZII Data Portal.  [https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/inference.ipynb](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/inference.ipynb)\n\nP.S.  Sorry for the delayed response, but you were way ahead of me at the time.  :)",
      "votes": null
    },
    {
      "id": "3116456",
      "postDate": "02/06/2025 01:10:20",
      "content": "<p><a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> Haha, no worries at all, I've since looked it up online and implemented the DataParallel and ended up using 2xT4s. Congratulations on Gold!! Well deserved for all the efforts and help you've provided on the discussion forums 🥳</p>",
      "rawMarkdown": "davidlist Haha, no worries at all, I've since looked it up online and implemented the DataParallel and ended up using 2xT4s. Congratulations on Gold!! Well deserved for all the efforts and help you've provided on the discussion forums 🥳",
      "votes": null
    },
    {
      "id": "3116671",
      "postDate": "02/06/2025 07:28:03",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>!  Yeah, I was pretty sure you'd be able to figure it out.  :)</p>",
      "rawMarkdown": "Thanks @andreizamfir!  Yeah, I was pretty sure you'd be able to figure it out.  :)",
      "votes": null
    },
    {
      "id": "3117772",
      "postDate": "02/07/2025 08:34:49",
      "content": "<p>Interesting. I tried this approach and my inference time increased 4x compared to not using DataParallel. I wonder what's the problem here.</p>",
      "rawMarkdown": "Interesting. I tried this approach and my inference time increased 4x compared to not using DataParallel. I wonder what's the problem here.",
      "votes": null
    },
    {
      "id": "3117787",
      "postDate": "02/07/2025 08:51:30",
      "content": "<p>I think I had to try a couple of things to get it to work.  The 4x inference time sounds familiar.  If I recall, the solution that worked was a lot easier than what was on the web.</p>",
      "rawMarkdown": "I think I had to try a couple of things to get it to work.  The 4x inference time sounds familiar.  If I recall, the solution that worked was a lot easier than what was on the web.",
      "votes": null
    },
    {
      "id": "3117801",
      "postDate": "02/07/2025 09:01:51",
      "content": "<p>Is it probably related to this part? I use mode gaussian.</p>\n<h1>mode='gaussian',</h1>\n<p>roi_weight_map=weight_map<br>\nI also set <br>\nsw_device=torch.device('cuda'),<br>\ndevice=torch.device('cuda'),<br>\nor else the sliding_window will throw an exception thatw_t (weight map) is on a different device (CPU vs cuda).</p>",
      "rawMarkdown": "Is it probably related to this part? I use mode gaussian.\n#mode='gaussian',\nroi_weight_map=weight_map\nI also set \nsw_device=torch.device('cuda'),\ndevice=torch.device('cuda'),\nor else the sliding_window will throw an exception thatw_t (weight map) is on a different device (CPU vs cuda).",
      "votes": null
    },
    {
      "id": "3117815",
      "postDate": "02/07/2025 09:09:01",
      "content": "<p>Yep.  That looks familiar.  I think the with torch.amp.autocast('cuda') takes care of all that.  If you do both, I suspect it gets confused.  Oh, and, yeah, I think mine fails with cpu only, and maybe a single gpu.</p>",
      "rawMarkdown": "Yep.  That looks familiar.  I think the with torch.amp.autocast('cuda') takes care of all that.  If you do both, I suspect it gets confused.  Oh, and, yeah, I think mine fails with cpu only, and maybe a single gpu.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3096997,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "01/14/2025 22:56:34",
      "content": "<p>I do two things:</p>\n<ul>\n<li>Train in one notebook and do inference in another.</li>\n<li>Use the T4 x 2 GPU.  Although this takes some extra code to setup, but it's pretty easy to google.</li>\n</ul>\n<p>Note, this is a pretty good discussion post if you're just getting started:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3097459,
          "author_name": "andreizamfir",
          "author_url": "",
          "post_date": "01/15/2025 11:46:13",
          "content": "<p>Question, as I’m using P100 and I lack a coding/software background, T4 x 2 makes use of extra parallelization, thus if I’m running one batch at a time, it wouldn’t really help me, unless I’d run two batches at a time, is my understanding correct?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3097544,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "01/15/2025 13:43:34",
              "content": "<p>Yes, you need to run 2 batches of data each time. You can run 4 batches during inference, which is convenient for using TTA during inference. For example, one batch is the original image, one batch is horizontally flipped, one batch is rotated 90 degrees, and so on.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3097553,
                  "author_name": "carlosperez97",
                  "author_url": "",
                  "post_date": "01/15/2025 13:51:14",
                  "content": "<p><a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a> We also use it to be able to do inference with big volume sizes in somehow big models. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3116437,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "02/06/2025 00:53:29",
              "content": "<p><a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a> If you use sliding_window_inference() and wrap your model with DataParallel() it runs the patches in parallel if you have more than one GPU:</p>\n<pre><code>model = torch.nn.DataParallel(model, device_ids=[,])\nmodel.to()\n\n ():\n     ():\n         sliding_window_inference(\n            inputs=,\n            roi_size=(, , ),\n            sw_batch_size=,  \n            predictor=model,\n            overlap=,\n            \n            roi_weight_map=weight_map\n        )\n\n    \n     torch.amp.autocast():\n         _compute()\n</code></pre>\n<p>Note the sliding_window_inference() is used in one of the example notebooks from the CZII Data Portal.  <a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/inference.ipynb\" target=\"_blank\">https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/inference.ipynb</a></p>\n<p>P.S.  Sorry for the delayed response, but you were way ahead of me at the time.  :)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3116456,
                  "author_name": "andreizamfir",
                  "author_url": "",
                  "post_date": "02/06/2025 01:10:20",
                  "content": "<p><a href=\"https://www.kaggle.com/davidlist\" target=\"_blank\">@davidlist</a> Haha, no worries at all, I've since looked it up online and implemented the DataParallel and ended up using 2xT4s. Congratulations on Gold!! Well deserved for all the efforts and help you've provided on the discussion forums 🥳</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3116671,
                      "author_name": "davidlist",
                      "author_url": "",
                      "post_date": "02/06/2025 07:28:03",
                      "content": "<p>Thanks <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>!  Yeah, I was pretty sure you'd be able to figure it out.  :)</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                },
                {
                  "id": 3117772,
                  "author_name": "sakvaua",
                  "author_url": "",
                  "post_date": "02/07/2025 08:34:49",
                  "content": "<p>Interesting. I tried this approach and my inference time increased 4x compared to not using DataParallel. I wonder what's the problem here.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3117787,
                      "author_name": "davidlist",
                      "author_url": "",
                      "post_date": "02/07/2025 08:51:30",
                      "content": "<p>I think I had to try a couple of things to get it to work.  The 4x inference time sounds familiar.  If I recall, the solution that worked was a lot easier than what was on the web.</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 3117801,
                      "author_name": "sakvaua",
                      "author_url": "",
                      "post_date": "02/07/2025 09:01:51",
                      "content": "<p>Is it probably related to this part? I use mode gaussian.</p>\n<h1>mode='gaussian',</h1>\n<p>roi_weight_map=weight_map<br>\nI also set <br>\nsw_device=torch.device('cuda'),<br>\ndevice=torch.device('cuda'),<br>\nor else the sliding_window will throw an exception thatw_t (weight map) is on a different device (CPU vs cuda).</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3117815,
                          "author_name": "davidlist",
                          "author_url": "",
                          "post_date": "02/07/2025 09:09:01",
                          "content": "<p>Yep.  That looks familiar.  I think the with torch.amp.autocast('cuda') takes care of all that.  If you do both, I suspect it gets confused.  Oh, and, yeah, I think mine fails with cpu only, and maybe a single gpu.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3099268,
          "author_name": "successmoses",
          "author_url": "",
          "post_date": "01/17/2025 14:23:42",
          "content": "<p>how do you setup T4 x 2 GPU</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3097533,
      "author_name": "moonirm",
      "author_url": "",
      "post_date": "01/15/2025 13:36:00",
      "content": "<p>I supposed you chose to work on high resolution with shape (184,630,630). Your code will be faster to execute on a medium resolution with shape (92,135,135), but I supposed it comes with drawback, as loss of information, and so, less effective prediction. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3109228,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "01/28/2025 16:18:31",
      "content": "<p>I highly recommend training in a different notebook than what you run inference on.  Train in notebook A, save the weights to a new kaggle dataset, load those weights in the inference notebook and run only inference.  Also make sure you are making use of the GPUs and that your data and model on are on the cuda device.  With no data parallelization or other optimizations this alone should get your runtime low enough to make a submission.  My notebook takes around 8-9 hours with absolutely no optimizations.  Make sure you aren't doing lots of preprocessing that could take a long time when the dataset size goes up to 500.  I recommend benchmarking on the 3 samples and multiply the runtime out so you can estimate the total submission time.</p>\n<p>It could also help to look at current submission notebooks and see what they do to run it within the allowed time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3096967": "The competition requires that runtime <= 12 hours. I am trying to reduce the runtime of a submission notebook. **How minimal can a submission code be**? Can I use pre trained weights instead of training in the notebook? Do I also need to make inference in the same notebook?",
    "3096997": "I do two things:\n- Train in one notebook and do inference in another.\n- Use the T4 x 2 GPU.  Although this takes some extra code to setup, but it's pretty easy to google.\n\nNote, this is a pretty good discussion post if you're just getting started:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549715)",
    "3097459": "Question, as I’m using P100 and I lack a coding/software background, T4 x 2 makes use of extra parallelization, thus if I’m running one batch at a time, it wouldn’t really help me, unless I’d run two batches at a time, is my understanding correct?",
    "3097533": "I supposed you chose to work on high resolution with shape (184,630,630). Your code will be faster to execute on a medium resolution with shape (92,135,135), but I supposed it comes with drawback, as loss of information, and so, less effective prediction.",
    "3097544": "Yes, you need to run 2 batches of data each time. You can run 4 batches during inference, which is convenient for using TTA during inference. For example, one batch is the original image, one batch is horizontally flipped, one batch is rotated 90 degrees, and so on.",
    "3097553": "andreizamfir We also use it to be able to do inference with big volume sizes in somehow big models.",
    "3099268": "how do you setup T4 x 2 GPU",
    "3109228": "I highly recommend training in a different notebook than what you run inference on.  Train in notebook A, save the weights to a new kaggle dataset, load those weights in the inference notebook and run only inference.  Also make sure you are making use of the GPUs and that your data and model on are on the cuda device.  With no data parallelization or other optimizations this alone should get your runtime low enough to make a submission.  My notebook takes around 8-9 hours with absolutely no optimizations.  Make sure you aren't doing lots of preprocessing that could take a long time when the dataset size goes up to 500.  I recommend benchmarking on the 3 samples and multiply the runtime out so you can estimate the total submission time.\n\nIt could also help to look at current submission notebooks and see what they do to run it within the allowed time.",
    "3116437": "andreizamfir If you use sliding_window_inference() and wrap your model with DataParallel() it runs the patches in parallel if you have more than one GPU:\n\n```\nmodel = torch.nn.DataParallel(model, device_ids=[0,1])\nmodel.to('cuda:0')\n\ndef inference(model, input):\n    def _compute(input):\n        return sliding_window_inference(\n            inputs=input,\n            roi_size=(96, 96, 96),\n            sw_batch_size=4,  # one window is proecessed at a time\n            predictor=model,\n            overlap=0.125,\n            #mode='gaussian',\n            roi_weight_map=weight_map\n        )\n\n    # with torch.cuda.amp.autocast():\n    with torch.amp.autocast('cuda'):\n        return _compute(input)\n```\n\nNote the sliding_window_inference() is used in one of the example notebooks from the CZII Data Portal.  [https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/inference.ipynb](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/3d_unet_monai/inference.ipynb)\n\nP.S.  Sorry for the delayed response, but you were way ahead of me at the time.  :)",
    "3116456": "davidlist Haha, no worries at all, I've since looked it up online and implemented the DataParallel and ended up using 2xT4s. Congratulations on Gold!! Well deserved for all the efforts and help you've provided on the discussion forums 🥳",
    "3116671": "Thanks @andreizamfir!  Yeah, I was pretty sure you'd be able to figure it out.  :)",
    "3117772": "Interesting. I tried this approach and my inference time increased 4x compared to not using DataParallel. I wonder what's the problem here.",
    "3117787": "I think I had to try a couple of things to get it to work.  The 4x inference time sounds familiar.  If I recall, the solution that worked was a lot easier than what was on the web.",
    "3117801": "Is it probably related to this part? I use mode gaussian.\n#mode='gaussian',\nroi_weight_map=weight_map\nI also set \nsw_device=torch.device('cuda'),\ndevice=torch.device('cuda'),\nor else the sliding_window will throw an exception thatw_t (weight map) is on a different device (CPU vs cuda).",
    "3117815": "Yep.  That looks familiar.  I think the with torch.amp.autocast('cuda') takes care of all that.  If you do both, I suspect it gets confused.  Oh, and, yeah, I think mine fails with cpu only, and maybe a single gpu."
  },
  "source": "meta"
}