{
  "id": 293894,
  "title": "Inference every epoch is not suitable (and how to avoid CUDA/CPU OOM)",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/293894",
  "author_name": "",
  "post_date": "2021-12-07T14:12:10.022303600Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The most popular detectron2 notebook has very helpfully provided a custom MAPIoU Evaluator class that can be <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/287023\" target=\"_blank\">manipulated</a> to automatically save the best model with a few modifications. However for memory intensive models, this practise of inference every epoch should be avoided and we should use seperate validation script on all model checkpoints AFTER training is complete, and then select the best model</p>\n<p>I say this because in my own training experience with all the LiveCell Data on Colab Pro, inference took up a big chunk of memory (and eventually gave OOM, on both CUDA and CPU after ~2 epochs) which could be saved by just postponing the inference step.</p>\n<p><a href=\"https://www.kaggle.com/ferlockx/validation-score-for-multiple-files\" target=\"_blank\">Simple code</a> to do this with your private detectron datasets</p>",
  "messages": [
    {
      "id": "1610781",
      "postDate": "12/07/2021 14:12:10",
      "content": "<p>The most popular detectron2 notebook has very helpfully provided a custom MAPIoU Evaluator class that can be <a href=\"https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/287023\" target=\"_blank\">manipulated</a> to automatically save the best model with a few modifications. However for memory intensive models, this practise of inference every epoch should be avoided and we should use seperate validation script on all model checkpoints AFTER training is complete, and then select the best model</p>\n<p>I say this because in my own training experience with all the LiveCell Data on Colab Pro, inference took up a big chunk of memory (and eventually gave OOM, on both CUDA and CPU after ~2 epochs) which could be saved by just postponing the inference step.</p>\n<p><a href=\"https://www.kaggle.com/ferlockx/validation-score-for-multiple-files\" target=\"_blank\">Simple code</a> to do this with your private detectron datasets</p>",
      "rawMarkdown": "The most popular detectron2 notebook has very helpfully provided a custom MAPIoU Evaluator class that can be [manipulated](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/287023) to automatically save the best model with a few modifications. However for memory intensive models, this practise of inference every epoch should be avoided and we should use seperate validation script on all model checkpoints AFTER training is complete, and then select the best model\n\nI say this because in my own training experience with all the LiveCell Data on Colab Pro, inference took up a big chunk of memory (and eventually gave OOM, on both CUDA and CPU after ~2 epochs) which could be saved by just postponing the inference step.\n\n[Simple code](https://www.kaggle.com/ferlockx/validation-score-for-multiple-files) to do this with your private detectron datasets",
      "votes": null
    },
    {
      "id": "1616556",
      "postDate": "12/13/2021 13:51:50",
      "content": "<p>Thnak you for sharing <a href=\"https://www.kaggle.com/ferlockx\" target=\"_blank\">@ferlockx</a> </p>\n<p>I am using Google Colab Pro + :)</p>\n<p>I still get this </p>\n<p><code>[12/13 13:42:44 d2.utils.memory]: Attempting to copy inputs of &lt;function pairwise_iou at 0x7f04fbb97560&gt; to CPU due to CUDA OOM</code></p>\n<p>Even when my code is :</p>\n<pre><code>cfg.DATASETS.TRAIN = (\"sartorius_train\",\"sartorius_test\")\ncfg.DATASETS.TEST = (\"sartorius_val\",)\ncfg.SOLVER.MAX_ITER = 500000 \ncfg.TEST.EVAL_PERIOD = 499999   \ncfg.SOLVER.CHECKPOINT_PERIOD = 10000   \n</code></pre>",
      "rawMarkdown": "Thnak you for sharing @ferlockx \n\nI am using Google Colab Pro + :)\n\nI still get this \n\n`[12/13 13:42:44 d2.utils.memory]: Attempting to copy inputs of <function pairwise_iou at 0x7f04fbb97560> to CPU due to CUDA OOM`\n\nEven when my code is :\n\n```\ncfg.DATASETS.TRAIN = (\"sartorius_train\",\"sartorius_test\")\ncfg.DATASETS.TEST = (\"sartorius_val\",)\ncfg.SOLVER.MAX_ITER = 500000 \ncfg.TEST.EVAL_PERIOD = 499999   \ncfg.SOLVER.CHECKPOINT_PERIOD = 10000   \n```",
      "votes": null
    },
    {
      "id": "1616571",
      "postDate": "12/13/2021 14:04:24",
      "content": "<p>yes I was also getting this during training but it just loses us a bit of time, training should proceed even if you get that error :) I got it multiple times during training and even though I had set inference to take place every epoch again, training completed in 7ish hours (100k iters) </p>\n<p>500k seems a little excessive thats ~200 epochs, I think you shouldn't go beyond 30-40</p>",
      "rawMarkdown": "yes I was also getting this during training but it just loses us a bit of time, training should proceed even if you get that error :) I got it multiple times during training and even though I had set inference to take place every epoch again, training completed in 7ish hours (100k iters) \n\n500k seems a little excessive thats ~200 epochs, I think you shouldn't go beyond 30-40",
      "votes": null
    },
    {
      "id": "1616602",
      "postDate": "12/13/2021 14:28:46",
      "content": "<p>Thank you , I appreciate it.<br>\nAt this stage , it is LR game :)<br>\nAny hint about setting LR for transfer learning?</p>",
      "rawMarkdown": "Thank you , I appreciate it.\nAt this stage , it is LR game :)\nAny hint about setting LR for transfer learning?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1616556,
      "author_name": "faisalalsrheed",
      "author_url": "",
      "post_date": "12/13/2021 13:51:50",
      "content": "<p>Thnak you for sharing <a href=\"https://www.kaggle.com/ferlockx\" target=\"_blank\">@ferlockx</a> </p>\n<p>I am using Google Colab Pro + :)</p>\n<p>I still get this </p>\n<p><code>[12/13 13:42:44 d2.utils.memory]: Attempting to copy inputs of &lt;function pairwise_iou at 0x7f04fbb97560&gt; to CPU due to CUDA OOM</code></p>\n<p>Even when my code is :</p>\n<pre><code>cfg.DATASETS.TRAIN = (\"sartorius_train\",\"sartorius_test\")\ncfg.DATASETS.TEST = (\"sartorius_val\",)\ncfg.SOLVER.MAX_ITER = 500000 \ncfg.TEST.EVAL_PERIOD = 499999   \ncfg.SOLVER.CHECKPOINT_PERIOD = 10000   \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1616571,
          "author_name": "ferlockx",
          "author_url": "",
          "post_date": "12/13/2021 14:04:24",
          "content": "<p>yes I was also getting this during training but it just loses us a bit of time, training should proceed even if you get that error :) I got it multiple times during training and even though I had set inference to take place every epoch again, training completed in 7ish hours (100k iters) </p>\n<p>500k seems a little excessive thats ~200 epochs, I think you shouldn't go beyond 30-40</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1616602,
          "author_name": "faisalalsrheed",
          "author_url": "",
          "post_date": "12/13/2021 14:28:46",
          "content": "<p>Thank you , I appreciate it.<br>\nAt this stage , it is LR game :)<br>\nAny hint about setting LR for transfer learning?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1610781": "The most popular detectron2 notebook has very helpfully provided a custom MAPIoU Evaluator class that can be [manipulated](https://www.kaggle.com/c/sartorius-cell-instance-segmentation/discussion/287023) to automatically save the best model with a few modifications. However for memory intensive models, this practise of inference every epoch should be avoided and we should use seperate validation script on all model checkpoints AFTER training is complete, and then select the best model\n\nI say this because in my own training experience with all the LiveCell Data on Colab Pro, inference took up a big chunk of memory (and eventually gave OOM, on both CUDA and CPU after ~2 epochs) which could be saved by just postponing the inference step.\n\n[Simple code](https://www.kaggle.com/ferlockx/validation-score-for-multiple-files) to do this with your private detectron datasets",
    "1616556": "Thnak you for sharing @ferlockx \n\nI am using Google Colab Pro + :)\n\nI still get this \n\n`[12/13 13:42:44 d2.utils.memory]: Attempting to copy inputs of <function pairwise_iou at 0x7f04fbb97560> to CPU due to CUDA OOM`\n\nEven when my code is :\n\n```\ncfg.DATASETS.TRAIN = (\"sartorius_train\",\"sartorius_test\")\ncfg.DATASETS.TEST = (\"sartorius_val\",)\ncfg.SOLVER.MAX_ITER = 500000 \ncfg.TEST.EVAL_PERIOD = 499999   \ncfg.SOLVER.CHECKPOINT_PERIOD = 10000   \n```",
    "1616571": "yes I was also getting this during training but it just loses us a bit of time, training should proceed even if you get that error :) I got it multiple times during training and even though I had set inference to take place every epoch again, training completed in 7ish hours (100k iters) \n\n500k seems a little excessive thats ~200 epochs, I think you shouldn't go beyond 30-40",
    "1616602": "Thank you , I appreciate it.\nAt this stage , it is LR game :)\nAny hint about setting LR for transfer learning?"
  },
  "source": "meta"
}