{
  "id": 297282,
  "title": "CUDA out of memory when training YOLOX-l and YOLOX-m",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/297282",
  "author_name": "",
  "post_date": "2021-12-26T12:13:08.325320Z",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p>When I'm training YOLOX-s, as guided by <a href=\"https://www.kaggle.com/remekkinas/yolox-training-pipeline-cots-dataset-lb-0-507?scriptVersionId=81353936\" target=\"_blank\">this amazing notebook</a> by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, everything worked as expected, but public LB is only 0.155 even after training for 8 hours over 50 epochs (on Kaggle Kernel).</p>\n<p>However, when I try to train YOLOX-l and YOLOX-m using the same notebook, I swap out the <code>Exp</code> class and change the pretrained weights, but I run into <code>RuntimeError: CUDA out of memory</code>. At the moment, I can't train any model larger than YOLOX-s, which means I cannot improve my public LB. I'm sorry if this is just a stupid bug, I'm relatively new to Kaggle.</p>\n<p>The full error message: <code>RuntimeError: CUDA out of memory. Tried to allocate 86.00 MiB (GPU 0; 15.90 GiB total capacity; 14.87 GiB already allocated; 45.75 MiB free; 14.92 GiB reserved in total by PyTorch)</code></p>\n<p>Does anyone know how to fix this? <strong>Here is <a href=\"https://www.kaggle.com/coderrexe/tensorflow-great-barrier-reef-training?scriptVersionId=83362594\" target=\"_blank\">my notebook</a> for reference, I would appreciate any help from anyone. Thank you.</strong></p>",
  "messages": [
    {
      "id": "1629647",
      "postDate": "12/26/2021 12:13:08",
      "content": "<p>When I'm training YOLOX-s, as guided by <a href=\"https://www.kaggle.com/remekkinas/yolox-training-pipeline-cots-dataset-lb-0-507?scriptVersionId=81353936\" target=\"_blank\">this amazing notebook</a> by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, everything worked as expected, but public LB is only 0.155 even after training for 8 hours over 50 epochs (on Kaggle Kernel).</p>\n<p>However, when I try to train YOLOX-l and YOLOX-m using the same notebook, I swap out the <code>Exp</code> class and change the pretrained weights, but I run into <code>RuntimeError: CUDA out of memory</code>. At the moment, I can't train any model larger than YOLOX-s, which means I cannot improve my public LB. I'm sorry if this is just a stupid bug, I'm relatively new to Kaggle.</p>\n<p>The full error message: <code>RuntimeError: CUDA out of memory. Tried to allocate 86.00 MiB (GPU 0; 15.90 GiB total capacity; 14.87 GiB already allocated; 45.75 MiB free; 14.92 GiB reserved in total by PyTorch)</code></p>\n<p>Does anyone know how to fix this? <strong>Here is <a href=\"https://www.kaggle.com/coderrexe/tensorflow-great-barrier-reef-training?scriptVersionId=83362594\" target=\"_blank\">my notebook</a> for reference, I would appreciate any help from anyone. Thank you.</strong></p>",
      "rawMarkdown": "When I'm training YOLOX-s, as guided by [this amazing notebook] (https://www.kaggle.com/remekkinas/yolox-training-pipeline-cots-dataset-lb-0-507?scriptVersionId=81353936) by @remekkinas, everything worked as expected, but public LB is only 0.155 even after training for 8 hours over 50 epochs (on Kaggle Kernel).\n\nHowever, when I try to train YOLOX-l and YOLOX-m using the same notebook, I swap out the `Exp` class and change the pretrained weights, but I run into `RuntimeError: CUDA out of memory`. At the moment, I can't train any model larger than YOLOX-s, which means I cannot improve my public LB. I'm sorry if this is just a stupid bug, I'm relatively new to Kaggle.\n\nThe full error message: `RuntimeError: CUDA out of memory. Tried to allocate 86.00 MiB (GPU 0; 15.90 GiB total capacity; 14.87 GiB already allocated; 45.75 MiB free; 14.92 GiB reserved in total by PyTorch)`\n\nDoes anyone know how to fix this? **Here is [my notebook](https://www.kaggle.com/coderrexe/tensorflow-great-barrier-reef-training?scriptVersionId=83362594) for reference, I would appreciate any help from anyone. Thank you.**",
      "votes": null
    },
    {
      "id": "1629654",
      "postDate": "12/26/2021 12:22:15",
      "content": "<p>Perhaps try to reduce the batch size. <br>\nYou can see from your error message that there is 45.75 Mib free but your training needs 86Mib. FYI, I think reducing batch size will just reduce the training speed, no harm for accuracy.</p>",
      "rawMarkdown": "Perhaps try to reduce the batch size. \nYou can see from your error message that there is 45.75 Mib free but your training needs 86Mib. FYI, I think reducing batch size will just reduce the training speed, no harm for accuracy.",
      "votes": null
    },
    {
      "id": "1629663",
      "postDate": "12/26/2021 12:33:17",
      "content": "<p>Thank you for replying, I think reducing the batch size evades the error message. I'll run through the full training pipeline just to be sure.</p>",
      "rawMarkdown": "Thank you for replying, I think reducing the batch size evades the error message. I'll run through the full training pipeline just to be sure.",
      "votes": null
    },
    {
      "id": "1629670",
      "postDate": "12/26/2021 12:49:33",
      "content": "<p>Hi, thank you for mentioning my notebook and work. I saw your notebook:</p>\n<ol>\n<li>Batch size 32 is too big - check 6-8-12 batch size for such image resolution.</li>\n<li>To speed up training process use --cache option. But …. check this: <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/296434\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/296434</a></li>\n<li>LB is only 0.155 - probably your inference code base on RGB (I am guessing). </li>\n</ol>",
      "rawMarkdown": "Hi, thank you for mentioning my notebook and work. I saw your notebook:\n1. Batch size 32 is too big - check 6-8-12 batch size for such image resolution.\n2. To speed up training process use --cache option. But .... check this: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/296434\n3.  LB is only 0.155 - probably your inference code base on RGB (I am guessing).",
      "votes": null
    },
    {
      "id": "1629679",
      "postDate": "12/26/2021 12:58:54",
      "content": "<p>Thank you so much for the suggestions! I'll definitely give them a try.</p>",
      "rawMarkdown": "Thank you so much for the suggestions! I'll definitely give them a try.",
      "votes": null
    },
    {
      "id": "1629704",
      "postDate": "12/26/2021 13:21:18",
      "content": "<p>If you have any question let me know … I will supportu you. No problem.</p>",
      "rawMarkdown": "If you have any question let me know ... I will supportu you. No problem.",
      "votes": null
    },
    {
      "id": "1630401",
      "postDate": "12/27/2021 09:02:21",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I confused about why the RGB will cause low score on LB?</p>",
      "rawMarkdown": "Hi @remekkinas I confused about why the RGB will cause low score on LB?",
      "votes": null
    },
    {
      "id": "1630414",
      "postDate": "12/27/2021 09:21:08",
      "content": "<p>Because YoloX inference and training is on BGR as I understand :)</p>",
      "rawMarkdown": "Because YoloX inference and training is on BGR as I understand :)",
      "votes": null
    },
    {
      "id": "1630851",
      "postDate": "12/27/2021 19:20:50",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> thanks for sharing that great notebook!  I’m having the same problem… I successfully trained a yolox_l with 20 epochs by reducing the batch size (solved the memory issue) however I am only getting a LB score of 0.28.  Do you think it’s just a matter of needing to train more epochs or the RGB to BGR conversion?  How should I adjust for BGR at inference time?  </p>\n<p>Note  I am using both your training and inference baselines with 960x960 size.</p>\n<p>Thanks for the help!</p>",
      "rawMarkdown": "Hi @remekkinas thanks for sharing that great notebook!  I’m having the same problem… I successfully trained a yolox_l with 20 epochs by reducing the batch size (solved the memory issue) however I am only getting a LB score of 0.28.  Do you think it’s just a matter of needing to train more epochs or the RGB to BGR conversion?  How should I adjust for BGR at inference time?  \n\nNote  I am using both your training and inference baselines with 960x960 size.\n\nThanks for the help!",
      "votes": null
    },
    {
      "id": "1630857",
      "postDate": "12/27/2021 19:26:20",
      "content": "<p>If you are using my both notebooks score should be higher (I am able to cross 0.53x without no problem now … and today I crossed 0.55x …). Use just infer notebook for inference (in training notebook inference is on RGB as fas as I remember). We spend a lot of time watching on dataset … and thinking how to feed NN with data. So some experiment we ended up with low score and some … with quite high (not such high like TOP people but … ok).</p>\n<p>Certainly our team score is a result of many many experiments.</p>",
      "rawMarkdown": "If you are using my both notebooks score should be higher (I am able to cross 0.53x without no problem now ... and today I crossed 0.55x ...). Use just infer notebook for inference (in training notebook inference is on RGB as fas as I remember). We spend a lot of time watching on dataset ... and thinking how to feed NN with data. So some experiment we ended up with low score and some ... with quite high (not such high like TOP people but ... ok).\n\nCertainly our team score is a result of many many experiments.",
      "votes": null
    },
    {
      "id": "1631100",
      "postDate": "12/28/2021 03:39:12",
      "content": "<p>Hey!<br>\nDecrease the batch size into the 16 it works for me! And the Same Time change the image size into 640  <br>\n<a href=\"https://www.kaggle.com/coderrexe\" target=\"_blank\">@coderrexe</a> what about your Learning rate Brother ?</p>",
      "rawMarkdown": "Hey!\nDecrease the batch size into the 16 it works for me! And the Same Time change the image size into 640  \n@coderrexe what about your Learning rate Brother ?",
      "votes": null
    },
    {
      "id": "1633293",
      "postDate": "12/30/2021 16:33:23",
      "content": "<p>I had the same issue multiple times either you can try what others mentioned before or you could just hit up the reset button and it works seamlessly </p>\n<p>P.S if you are trying the reset option just remove any reduntant piece of code, try keeping up code that's only required for running up the model and not displaying etc if your images, labels are placed inside your cuda cores since even that consumes memory</p>",
      "rawMarkdown": "I had the same issue multiple times either you can try what others mentioned before or you could just hit up the reset button and it works seamlessly \n\nP.S if you are trying the reset option just remove any reduntant piece of code, try keeping up code that's only required for running up the model and not displaying etc if your images, labels are placed inside your cuda cores since even that consumes memory",
      "votes": null
    },
    {
      "id": "2922226",
      "postDate": "07/15/2024 02:54:37",
      "content": "<p>You can set the num_workers variable to 1 or 0, and pin_memory to true</p>",
      "rawMarkdown": "You can set the num_workers variable to 1 or 0, and pin_memory to true",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1629654,
      "author_name": "owenxing",
      "author_url": "",
      "post_date": "12/26/2021 12:22:15",
      "content": "<p>Perhaps try to reduce the batch size. <br>\nYou can see from your error message that there is 45.75 Mib free but your training needs 86Mib. FYI, I think reducing batch size will just reduce the training speed, no harm for accuracy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1629663,
          "author_name": "coderrexe",
          "author_url": "",
          "post_date": "12/26/2021 12:33:17",
          "content": "<p>Thank you for replying, I think reducing the batch size evades the error message. I'll run through the full training pipeline just to be sure.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1629670,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "12/26/2021 12:49:33",
      "content": "<p>Hi, thank you for mentioning my notebook and work. I saw your notebook:</p>\n<ol>\n<li>Batch size 32 is too big - check 6-8-12 batch size for such image resolution.</li>\n<li>To speed up training process use --cache option. But …. check this: <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/296434\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/296434</a></li>\n<li>LB is only 0.155 - probably your inference code base on RGB (I am guessing). </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1629679,
          "author_name": "coderrexe",
          "author_url": "",
          "post_date": "12/26/2021 12:58:54",
          "content": "<p>Thank you so much for the suggestions! I'll definitely give them a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1629704,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/26/2021 13:21:18",
          "content": "<p>If you have any question let me know … I will supportu you. No problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1630401,
          "author_name": "jarviskevin",
          "author_url": "",
          "post_date": "12/27/2021 09:02:21",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I confused about why the RGB will cause low score on LB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1630414,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/27/2021 09:21:08",
          "content": "<p>Because YoloX inference and training is on BGR as I understand :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1630851,
          "author_name": "bartmaciszewski",
          "author_url": "",
          "post_date": "12/27/2021 19:20:50",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> thanks for sharing that great notebook!  I’m having the same problem… I successfully trained a yolox_l with 20 epochs by reducing the batch size (solved the memory issue) however I am only getting a LB score of 0.28.  Do you think it’s just a matter of needing to train more epochs or the RGB to BGR conversion?  How should I adjust for BGR at inference time?  </p>\n<p>Note  I am using both your training and inference baselines with 960x960 size.</p>\n<p>Thanks for the help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1630857,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "12/27/2021 19:26:20",
          "content": "<p>If you are using my both notebooks score should be higher (I am able to cross 0.53x without no problem now … and today I crossed 0.55x …). Use just infer notebook for inference (in training notebook inference is on RGB as fas as I remember). We spend a lot of time watching on dataset … and thinking how to feed NN with data. So some experiment we ended up with low score and some … with quite high (not such high like TOP people but … ok).</p>\n<p>Certainly our team score is a result of many many experiments.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1631100,
      "author_name": "balasubramaniamv",
      "author_url": "",
      "post_date": "12/28/2021 03:39:12",
      "content": "<p>Hey!<br>\nDecrease the batch size into the 16 it works for me! And the Same Time change the image size into 640  <br>\n<a href=\"https://www.kaggle.com/coderrexe\" target=\"_blank\">@coderrexe</a> what about your Learning rate Brother ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1633293,
      "author_name": "harshris21",
      "author_url": "",
      "post_date": "12/30/2021 16:33:23",
      "content": "<p>I had the same issue multiple times either you can try what others mentioned before or you could just hit up the reset button and it works seamlessly </p>\n<p>P.S if you are trying the reset option just remove any reduntant piece of code, try keeping up code that's only required for running up the model and not displaying etc if your images, labels are placed inside your cuda cores since even that consumes memory</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2922226,
      "author_name": "niuniuhahaha",
      "author_url": "",
      "post_date": "07/15/2024 02:54:37",
      "content": "<p>You can set the num_workers variable to 1 or 0, and pin_memory to true</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1629647": "When I'm training YOLOX-s, as guided by [this amazing notebook] (https://www.kaggle.com/remekkinas/yolox-training-pipeline-cots-dataset-lb-0-507?scriptVersionId=81353936) by @remekkinas, everything worked as expected, but public LB is only 0.155 even after training for 8 hours over 50 epochs (on Kaggle Kernel).\n\nHowever, when I try to train YOLOX-l and YOLOX-m using the same notebook, I swap out the `Exp` class and change the pretrained weights, but I run into `RuntimeError: CUDA out of memory`. At the moment, I can't train any model larger than YOLOX-s, which means I cannot improve my public LB. I'm sorry if this is just a stupid bug, I'm relatively new to Kaggle.\n\nThe full error message: `RuntimeError: CUDA out of memory. Tried to allocate 86.00 MiB (GPU 0; 15.90 GiB total capacity; 14.87 GiB already allocated; 45.75 MiB free; 14.92 GiB reserved in total by PyTorch)`\n\nDoes anyone know how to fix this? **Here is [my notebook](https://www.kaggle.com/coderrexe/tensorflow-great-barrier-reef-training?scriptVersionId=83362594) for reference, I would appreciate any help from anyone. Thank you.**",
    "1629654": "Perhaps try to reduce the batch size. \nYou can see from your error message that there is 45.75 Mib free but your training needs 86Mib. FYI, I think reducing batch size will just reduce the training speed, no harm for accuracy.",
    "1629663": "Thank you for replying, I think reducing the batch size evades the error message. I'll run through the full training pipeline just to be sure.",
    "1629670": "Hi, thank you for mentioning my notebook and work. I saw your notebook:\n1. Batch size 32 is too big - check 6-8-12 batch size for such image resolution.\n2. To speed up training process use --cache option. But .... check this: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/296434\n3.  LB is only 0.155 - probably your inference code base on RGB (I am guessing).",
    "1629679": "Thank you so much for the suggestions! I'll definitely give them a try.",
    "1629704": "If you have any question let me know ... I will supportu you. No problem.",
    "1630401": "Hi @remekkinas I confused about why the RGB will cause low score on LB?",
    "1630414": "Because YoloX inference and training is on BGR as I understand :)",
    "1630851": "Hi @remekkinas thanks for sharing that great notebook!  I’m having the same problem… I successfully trained a yolox_l with 20 epochs by reducing the batch size (solved the memory issue) however I am only getting a LB score of 0.28.  Do you think it’s just a matter of needing to train more epochs or the RGB to BGR conversion?  How should I adjust for BGR at inference time?  \n\nNote  I am using both your training and inference baselines with 960x960 size.\n\nThanks for the help!",
    "1630857": "If you are using my both notebooks score should be higher (I am able to cross 0.53x without no problem now ... and today I crossed 0.55x ...). Use just infer notebook for inference (in training notebook inference is on RGB as fas as I remember). We spend a lot of time watching on dataset ... and thinking how to feed NN with data. So some experiment we ended up with low score and some ... with quite high (not such high like TOP people but ... ok).\n\nCertainly our team score is a result of many many experiments.",
    "1631100": "Hey!\nDecrease the batch size into the 16 it works for me! And the Same Time change the image size into 640  \n@coderrexe what about your Learning rate Brother ?",
    "1633293": "I had the same issue multiple times either you can try what others mentioned before or you could just hit up the reset button and it works seamlessly \n\nP.S if you are trying the reset option just remove any reduntant piece of code, try keeping up code that's only required for running up the model and not displaying etc if your images, labels are placed inside your cuda cores since even that consumes memory",
    "2922226": "You can set the num_workers variable to 1 or 0, and pin_memory to true"
  },
  "source": "meta"
}