{
  "id": 287023,
  "title": "Saving the Best checkpoint in Detectron2",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/287023",
  "author_name": "Yerram Varun",
  "post_date": "2021-11-11T18:34:49.735000",
  "votes": 86,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Thank you <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> for your excellent notebooks. They helped me learn a lot about Detectron2 and its working.<br>\nI found a way to save the best checkpoints from the Detectron2, This might help out someone in need. So sharing here :)</p>\n<p>Import some extra classes</p>\n<pre><code>from detectron2.engine import BestCheckpointer\nfrom detectron2.checkpoint import DetectionCheckpointer\n</code></pre>\n<p>Modify the trainer class to add a BestCheckpointer hook</p>\n<pre><code>class Trainer(DefaultTrainer):\n    @classmethod\n    def build_evaluator(cls, cfg, dataset_name, output_folder=None):\n        return MAPIOUEvaluator(dataset_name)\n\n    def build_hooks(self):\n\n        # copy of cfg\n        cfg = self.cfg.clone()\n\n        # build the original model hooks\n        hooks = super().build_hooks()\n\n        # add the best checkpointer hook\n        hooks.insert(-1, BestCheckpointer(cfg.TEST.EVAL_PERIOD, \n                                         DetectionCheckpointer(self.model, cfg.OUTPUT_DIR),\n                                         \"MaP IoU\",\n                                         \"max\",\n                                         ))\n        return hooks\n</code></pre>\n<p>This will save a <code>model_best.pth</code> in the OUTPUT directory. These weights will save the weights with Maximum MaP IoU and you can directly use them for inference.</p>",
  "messages": [
    {
      "id": 1579329,
      "postDate": "2021-11-11T18:34:49.737Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/slawekbiel\" target=\"_blank\">@slawekbiel</a> for your excellent notebooks. They helped me learn a lot about Detectron2 and its working.<br>\nI found a way to save the best checkpoints from the Detectron2, This might help out someone in need. So sharing here :)</p>\n<p>Import some extra classes</p>\n<pre><code>from detectron2.engine import BestCheckpointer\nfrom detectron2.checkpoint import DetectionCheckpointer\n</code></pre>\n<p>Modify the trainer class to add a BestCheckpointer hook</p>\n<pre><code>class Trainer(DefaultTrainer):\n    @classmethod\n    def build_evaluator(cls, cfg, dataset_name, output_folder=None):\n        return MAPIOUEvaluator(dataset_name)\n\n    def build_hooks(self):\n\n        # copy of cfg\n        cfg = self.cfg.clone()\n\n        # build the original model hooks\n        hooks = super().build_hooks()\n\n        # add the best checkpointer hook\n        hooks.insert(-1, BestCheckpointer(cfg.TEST.EVAL_PERIOD, \n                                         DetectionCheckpointer(self.model, cfg.OUTPUT_DIR),\n                                         \"MaP IoU\",\n                                         \"max\",\n                                         ))\n        return hooks\n</code></pre>\n<p>This will save a <code>model_best.pth</code> in the OUTPUT directory. These weights will save the weights with Maximum MaP IoU and you can directly use them for inference.</p>",
      "rawMarkdown": "Thank you @slawekbiel for your excellent notebooks. They helped me learn a lot about Detectron2 and its working.\nI found a way to save the best checkpoints from the Detectron2, This might help out someone in need. So sharing here :)\n\nImport some extra classes\n```\nfrom detectron2.engine import BestCheckpointer\nfrom detectron2.checkpoint import DetectionCheckpointer\n```\n\nModify the trainer class to add a BestCheckpointer hook\n\n```\nclass Trainer(DefaultTrainer):\n    @classmethod\n    def build_evaluator(cls, cfg, dataset_name, output_folder=None):\n        return MAPIOUEvaluator(dataset_name)\n    \n    def build_hooks(self):\n        \n        # copy of cfg\n        cfg = self.cfg.clone()\n        \n        # build the original model hooks\n        hooks = super().build_hooks()\n        \n        # add the best checkpointer hook\n        hooks.insert(-1, BestCheckpointer(cfg.TEST.EVAL_PERIOD, \n                                         DetectionCheckpointer(self.model, cfg.OUTPUT_DIR),\n                                         \"MaP IoU\",\n                                         \"max\",\n                                         ))\n        return hooks\n```\n\nThis will save a `model_best.pth` in the OUTPUT directory. These weights will save the weights with Maximum MaP IoU and you can directly use them for inference.\n",
      "votes": 86
    },
    {
      "id": 1581013,
      "postDate": "2021-11-13T10:19:56.297Z",
      "content": "<p>And do this to skip saving all weights, which is the default setting (In case you're still using the trainer's DefaultCheckpointer)<br>\n<code>cfg.SOLVER.CHECKPOINT_PERIOD = cfg.SOLVER.MAX_ITER+1</code></p>\n<p>This one setting can save lots and lots of space.</p>",
      "rawMarkdown": "And do this to skip saving all weights, which is the default setting (In case you're still using the trainer's DefaultCheckpointer)\n`cfg.SOLVER.CHECKPOINT_PERIOD = cfg.SOLVER.MAX_ITER+1`\n\nThis one setting can save lots and lots of space.",
      "votes": 3
    },
    {
      "id": 1588791,
      "postDate": "2021-11-19T16:26:55.763Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> thanks for that tip! +1</p>\n<p></p>\n<p></p>\n<p>EDIT: sort it out - works fine! </p>",
      "rawMarkdown": "Hi @yerramvarun thanks for that tip! +1\n\n~~Do you have any idea why in my case does not save the `best_model.pth`? \nI can only see `model_final.pth` (and the intermediate chekp if I don't suppress them) ~~\n\n~~anyone with same issue maybe ? ~~\n\nEDIT: sort it out - works fine! ",
      "votes": 1
    },
    {
      "id": 1588298,
      "postDate": "2021-11-19T10:33:36.910Z",
      "content": "<p>You are sooooooooooooooooo good!!!!!</p>",
      "rawMarkdown": "You are sooooooooooooooooo good!!!!!",
      "votes": 1
    },
    {
      "id": 1587823,
      "postDate": "2021-11-19T00:46:58.860Z",
      "content": "<p>Great work. It helps me a lot!</p>",
      "rawMarkdown": "Great work. It helps me a lot!",
      "votes": 1
    },
    {
      "id": 1587757,
      "postDate": "2021-11-18T23:39:52.197Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> </p>\n<p>Do you know why the size of the best model is different from others?</p>\n<p><img src=\"https://i.postimg.cc/x8dcWqch/sartorius.jpg\" alt=\"\"></p>",
      "rawMarkdown": "Hi @yerramvarun \n\nDo you know why the size of the best model is different from others?\n\n![](https://i.postimg.cc/x8dcWqch/sartorius.jpg)",
      "votes": 1,
      "replies": [
        {
          "id": 1587993,
          "postDate": "2021-11-19T05:41:59.923Z",
          "content": "<p>It might be because Bestcheckpointer doesn't save the optimizer states and param groups.</p>",
          "rawMarkdown": "It might be because Bestcheckpointer doesn't save the optimizer states and param groups.",
          "votes": 6
        }
      ]
    },
    {
      "id": 1580766,
      "postDate": "2021-11-13T05:21:41.140Z",
      "content": "<p>Much needed <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> . Thanks for it.</p>",
      "rawMarkdown": "Much needed @yerramvarun . Thanks for it.",
      "votes": 1
    },
    {
      "id": 1579370,
      "postDate": "2021-11-11T20:15:26.613Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> for sharing, I needed it. </p>\n<p>to resume training I am just changing this part of the code ( it seems like working)</p>\n<pre><code>cfg.MODEL.WEIGHTS = '../input/my-Sartorius-v1/output/model_final.pth'\ntrainer.resume_or_load(resume=True)\n</code></pre>",
      "rawMarkdown": "Thank you @yerramvarun for sharing, I needed it. \n\nto resume training I am just changing this part of the code ( it seems like working)\n```\ncfg.MODEL.WEIGHTS = '../input/my-Sartorius-v1/output/model_final.pth'\ntrainer.resume_or_load(resume=True)\n```\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1579514,
          "postDate": "2021-11-12T00:22:10.307Z",
          "content": "<p>You are welcome ;)<br>\nThanks for this! <a href=\"https://www.kaggle.com/faisalalsrheed\" target=\"_blank\">@faisalalsrheed</a> </p>",
          "rawMarkdown": "You are welcome ;)\nThanks for this! @faisalalsrheed ",
          "votes": 1
        },
        {
          "id": 1589687,
          "postDate": "2021-11-20T13:59:48.133Z",
          "content": "<p>you might need to use this as well for better resume training</p>\n<p><code>cfg.SOLVER.WARMUP_ITERS = 0</code></p>",
          "rawMarkdown": "you might need to use this as well for better resume training\n\n`cfg.SOLVER.WARMUP_ITERS = 0`\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1582510,
      "postDate": "2021-11-15T01:10:26.327Z",
      "content": "<p>May i ask a question? Do u know how to apply mutil-gpu training on detectrion2?</p>",
      "rawMarkdown": "May i ask a question? Do u know how to apply mutil-gpu training on detectrion2?",
      "replies": [
        {
          "id": 1587994,
          "postDate": "2021-11-19T05:43:08.770Z",
          "content": "<p><a href=\"https://www.kaggle.com/manyuli\" target=\"_blank\">@manyuli</a> Hey does multi-gpu argument in <a href=\"https://github.com/facebookresearch/detectron2/blob/main/tools/train_net.py\" target=\"_blank\">train_net.py</a> help?</p>",
          "rawMarkdown": "@manyuli Hey does multi-gpu argument in [train_net.py](https://github.com/facebookresearch/detectron2/blob/main/tools/train_net.py) help?"
        },
        {
          "id": 1587996,
          "postDate": "2021-11-19T05:44:51.607Z",
          "content": "<p>I use <code>launch</code> method to solve that problem</p>",
          "rawMarkdown": "I use `launch` method to solve that problem",
          "votes": 2
        },
        {
          "id": 1614728,
          "postDate": "2021-12-11T12:34:28.283Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1594883,
      "postDate": "2021-11-25T07:40:52Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1594879,
      "postDate": "2021-11-25T07:38:25.490Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1626799,
      "postDate": "2021-12-23T08:16:00.377Z",
      "content": "<p>Great tip, thanks!!</p>",
      "rawMarkdown": "Great tip, thanks!!"
    }
  ],
  "comments": [
    {
      "id": 1581013,
      "author_name": "Mighty Rains",
      "author_url": "",
      "post_date": "2021-11-13T10:19:56.297000",
      "content": "<p>And do this to skip saving all weights, which is the default setting (In case you're still using the trainer's DefaultCheckpointer)<br>\n<code>cfg.SOLVER.CHECKPOINT_PERIOD = cfg.SOLVER.MAX_ITER+1</code></p>\n<p>This one setting can save lots and lots of space.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1588791,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-11-19T16:26:55.763000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> thanks for that tip! +1</p>\n<p></p>\n<p></p>\n<p>EDIT: sort it out - works fine! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1588298,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2021-11-19T10:33:36.910000",
      "content": "<p>You are sooooooooooooooooo good!!!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1587823,
      "author_name": "Roc",
      "author_url": "",
      "post_date": "2021-11-19T00:46:58.860000",
      "content": "<p>Great work. It helps me a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1587757,
      "author_name": "Faisal Alsrheed",
      "author_url": "",
      "post_date": "2021-11-18T23:39:52.197000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> </p>\n<p>Do you know why the size of the best model is different from others?</p>\n<p><img src=\"https://i.postimg.cc/x8dcWqch/sartorius.jpg\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1587993,
          "author_name": "Yerram Varun",
          "author_url": "",
          "post_date": "2021-11-19T05:41:59.923000",
          "content": "<p>It might be because Bestcheckpointer doesn't save the optimizer states and param groups.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1580766,
      "author_name": "Sanchit Vijay",
      "author_url": "",
      "post_date": "2021-11-13T05:21:41.140000",
      "content": "<p>Much needed <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> . Thanks for it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1579370,
      "author_name": "Faisal Alsrheed",
      "author_url": "",
      "post_date": "2021-11-11T20:15:26.613000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/yerramvarun\" target=\"_blank\">@yerramvarun</a> for sharing, I needed it. </p>\n<p>to resume training I am just changing this part of the code ( it seems like working)</p>\n<pre><code>cfg.MODEL.WEIGHTS = '../input/my-Sartorius-v1/output/model_final.pth'\ntrainer.resume_or_load(resume=True)\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 1579514,
          "author_name": "Yerram Varun",
          "author_url": "",
          "post_date": "2021-11-12T00:22:10.307000",
          "content": "<p>You are welcome ;)<br>\nThanks for this! <a href=\"https://www.kaggle.com/faisalalsrheed\" target=\"_blank\">@faisalalsrheed</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1589687,
          "author_name": "Faisal Alsrheed",
          "author_url": "",
          "post_date": "2021-11-20T13:59:48.133000",
          "content": "<p>you might need to use this as well for better resume training</p>\n<p><code>cfg.SOLVER.WARMUP_ITERS = 0</code></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1582510,
      "author_name": "Manyu Li",
      "author_url": "",
      "post_date": "2021-11-15T01:10:26.327000",
      "content": "<p>May i ask a question? Do u know how to apply mutil-gpu training on detectrion2?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1587994,
          "author_name": "Yerram Varun",
          "author_url": "",
          "post_date": "2021-11-19T05:43:08.770000",
          "content": "<p><a href=\"https://www.kaggle.com/manyuli\" target=\"_blank\">@manyuli</a> Hey does multi-gpu argument in <a href=\"https://github.com/facebookresearch/detectron2/blob/main/tools/train_net.py\" target=\"_blank\">train_net.py</a> help?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1587996,
          "author_name": "Manyu Li",
          "author_url": "",
          "post_date": "2021-11-19T05:44:51.607000",
          "content": "<p>I use <code>launch</code> method to solve that problem</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1614728,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-12-11T12:34:28.283000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1594883,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-25T07:40:52",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1594879,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-25T07:38:25.490000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1626799,
      "author_name": "Boris Polishchuk",
      "author_url": "",
      "post_date": "2021-12-23T08:16:00.377000",
      "content": "<p>Great tip, thanks!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1579329": "Thank you @slawekbiel for your excellent notebooks. They helped me learn a lot about Detectron2 and its working.\nI found a way to save the best checkpoints from the Detectron2, This might help out someone in need. So sharing here :)\n\nImport some extra classes\n```\nfrom detectron2.engine import BestCheckpointer\nfrom detectron2.checkpoint import DetectionCheckpointer\n```\n\nModify the trainer class to add a BestCheckpointer hook\n\n```\nclass Trainer(DefaultTrainer):\n    @classmethod\n    def build_evaluator(cls, cfg, dataset_name, output_folder=None):\n        return MAPIOUEvaluator(dataset_name)\n    \n    def build_hooks(self):\n        \n        # copy of cfg\n        cfg = self.cfg.clone()\n        \n        # build the original model hooks\n        hooks = super().build_hooks()\n        \n        # add the best checkpointer hook\n        hooks.insert(-1, BestCheckpointer(cfg.TEST.EVAL_PERIOD, \n                                         DetectionCheckpointer(self.model, cfg.OUTPUT_DIR),\n                                         \"MaP IoU\",\n                                         \"max\",\n                                         ))\n        return hooks\n```\n\nThis will save a `model_best.pth` in the OUTPUT directory. These weights will save the weights with Maximum MaP IoU and you can directly use them for inference.\n",
    "1581013": "And do this to skip saving all weights, which is the default setting (In case you're still using the trainer's DefaultCheckpointer)\n`cfg.SOLVER.CHECKPOINT_PERIOD = cfg.SOLVER.MAX_ITER+1`\n\nThis one setting can save lots and lots of space.",
    "1588791": "Hi @yerramvarun thanks for that tip! +1\n\n~~Do you have any idea why in my case does not save the `best_model.pth`? \nI can only see `model_final.pth` (and the intermediate chekp if I don't suppress them) ~~\n\n~~anyone with same issue maybe ? ~~\n\nEDIT: sort it out - works fine! ",
    "1588298": "You are sooooooooooooooooo good!!!!!",
    "1587823": "Great work. It helps me a lot!",
    "1587757": "Hi @yerramvarun \n\nDo you know why the size of the best model is different from others?\n\n![](https://i.postimg.cc/x8dcWqch/sartorius.jpg)",
    "1580766": "Much needed @yerramvarun . Thanks for it.",
    "1579370": "Thank you @yerramvarun for sharing, I needed it. \n\nto resume training I am just changing this part of the code ( it seems like working)\n```\ncfg.MODEL.WEIGHTS = '../input/my-Sartorius-v1/output/model_final.pth'\ntrainer.resume_or_load(resume=True)\n```\n\n",
    "1582510": "May i ask a question? Do u know how to apply mutil-gpu training on detectrion2?",
    "1594883": "",
    "1594879": "",
    "1626799": "Great tip, thanks!!"
  }
}