{
  "id": 66337,
  "title": "Kernel timeout alternative",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/66337",
  "author_name": "",
  "post_date": "2018-09-20T18:51:41.999292700Z",
  "votes": null,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I am not able to train my model fully as the kernel dies after 6 hours. I don't have a high end pc. Can somebody tell me what setup or platforms they are using for training? </p>",
  "messages": [
    {
      "id": "390743",
      "postDate": "09/20/2018 18:51:42",
      "content": "<p>I am not able to train my model fully as the kernel dies after 6 hours. I don't have a high end pc. Can somebody tell me what setup or platforms they are using for training? </p>",
      "rawMarkdown": "I am not able to train my model fully as the kernel dies after 6 hours. I don't have a high end pc. Can somebody tell me what setup or platforms they are using for training?",
      "votes": null
    },
    {
      "id": "390768",
      "postDate": "09/20/2018 19:34:09",
      "content": "<p>With a bit of effort you can restart your training from more or less where you stopped. You can also use google colab and get 12h. Otherwise, you have to pay for Google cloud or Amazon (unless you never signed up with them - in this case Google gives you 300$ credit)</p>",
      "rawMarkdown": "With a bit of effort you can restart your training from more or less where you stopped. You can also use google colab and get 12h. Otherwise, you have to pay for Google cloud or Amazon (unless you never signed up with them - in this case Google gives you 300$ credit)",
      "votes": null
    },
    {
      "id": "390913",
      "postDate": "09/21/2018 03:06:31",
      "content": "<p>How can I stop my training and restart from where I stopped?</p>",
      "rawMarkdown": "How can I stop my training and restart from where I stopped?",
      "votes": null
    },
    {
      "id": "390920",
      "postDate": "09/21/2018 03:16:54",
      "content": "<p>Depends on what framework you use. If you make sure your snapshots are saved  and there are not too many files in the root, your snapshot will appear in the output of the kernel and can be used as dataset for another one (a fork)</p>",
      "rawMarkdown": "Depends on what framework you use. If you make sure your snapshots are saved  and there are not too many files in the root, your snapshot will appear in the output of the kernel and can be used as dataset for another one (a fork)",
      "votes": null
    },
    {
      "id": "391033",
      "postDate": "09/21/2018 06:55:35",
      "content": "<p>Can you guide me to a example?</p>",
      "rawMarkdown": "Can you guide me to a example?",
      "votes": null
    },
    {
      "id": "391053",
      "postDate": "09/21/2018 07:26:12",
      "content": "<p>Ummm all my examples qre private kernels but... Here are some more retails:\nWhen you commit a kernel, after it finish successfuly, it has \"output\" tab. If you kept your working directory tidy and removed all unnecessary things, everything you left in it will show on this tab,so make sure your backup/snapshot /checkpoint is left there.\nNext, fork the notebook, click \"add data source\", choose \"kernel output\" and choose the kernel that just finished. When starting the training again, load the pre trained weights from the proper directory</p>",
      "rawMarkdown": "Ummm all my examples qre private kernels but... Here are some more retails:\nWhen you commit a kernel, after it finish successfuly, it has \"output\" tab. If you kept your working directory tidy and removed all unnecessary things, everything you left in it will show on this tab,so make sure your backup/snapshot /checkpoint is left there.\nNext, fork the notebook, click \"add data source\", choose \"kernel output\" and choose the kernel that just finished. When starting the training again, load the pre trained weights from the proper directory",
      "votes": null
    },
    {
      "id": "391067",
      "postDate": "09/21/2018 08:20:48",
      "content": "<p>Ok thanks</p>",
      "rawMarkdown": "Ok thanks",
      "votes": null
    },
    {
      "id": "391727",
      "postDate": "09/22/2018 09:26:40",
      "content": "<p>I was thinking of training the model and save it using keras along with architecture and state of optimizer as .h5. Downloading it and starting a new kernel , loading the previous model and training it again. will this workout? will my model by initialized by previous weights and optimizer state will start from where i saved?</p>",
      "rawMarkdown": "I was thinking of training the model and save it using keras along with architecture and state of optimizer as .h5. Downloading it and starting a new kernel , loading the previous model and training it again. will this workout? will my model by initialized by previous weights and optimizer state will start from where i saved?",
      "votes": null
    },
    {
      "id": "391729",
      "postDate": "09/22/2018 09:46:14",
      "content": "<p>How can I save a mrcnn model?</p>",
      "rawMarkdown": "How can I save a mrcnn model?",
      "votes": null
    },
    {
      "id": "391742",
      "postDate": "09/22/2018 10:25:18",
      "content": "<p>You can do that. The h5 loads tge complete state (almost). My way is much more efficient but yours will work as well</p>",
      "rawMarkdown": "You can do that. The h5 loads tge complete state (almost). My way is much more efficient but yours will work as well",
      "votes": null
    },
    {
      "id": "391784",
      "postDate": "09/22/2018 12:28:09",
      "content": "<p>not able to load saved file. following error occurs:\nOSError: Unable to open file (truncated file: eof = 38932208, sblock-&gt;base_addr = 0, stored_eof = 179179456)</p>\n\n<p>Can you explain your way a little deeper?</p>",
      "rawMarkdown": "not able to load saved file. following error occurs:\nOSError: Unable to open file (truncated file: eof = 38932208, sblock-&gt;base_addr = 0, stored_eof = 179179456)\n\nCan you explain your way a little deeper?",
      "votes": null
    },
    {
      "id": "392007",
      "postDate": "09/22/2018 20:45:58",
      "content": "<p>This is a bug in your code. Try to iron out the kinks by writing a kernel that saves and load the model successfuly in the same code. When you get this right, you can progress to doing it across kernels</p>",
      "rawMarkdown": "This is a bug in your code. Try to iron out the kinks by writing a kernel that saves and load the model successfuly in the same code. When you get this right, you can progress to doing it across kernels",
      "votes": null
    },
    {
      "id": "392195",
      "postDate": "09/23/2018 08:50:44",
      "content": "<p>i figured it out. now this error arises when i try to load my model:\n'utf8' codec can't decode byte 0x9c in position 131: invalid start byte</p>",
      "rawMarkdown": "i figured it out. now this error arises when i try to load my model:\n'utf8' codec can't decode byte 0x9c in position 131: invalid start byte",
      "votes": null
    },
    {
      "id": "392216",
      "postDate": "09/23/2018 09:25:27",
      "content": "<p>Until you figure out how to save and load a model, there is very little to be done. </p>",
      "rawMarkdown": "Until you figure out how to save and load a model, there is very little to be done.",
      "votes": null
    },
    {
      "id": "392252",
      "postDate": "09/23/2018 11:08:38",
      "content": "<p>done with importingthe saved model to new kernel, but not able to retrain it. while loading the model following error occurs: \nOSError: Unable to open file (file signature not found)</p>",
      "rawMarkdown": "done with importingthe saved model to new kernel, but not able to retrain it. while loading the model following error occurs: \nOSError: Unable to open file (file signature not found)",
      "votes": null
    },
    {
      "id": "392271",
      "postDate": "09/23/2018 12:37:58",
      "content": "<p>@Moshel: Done with that error to. Thanks for help.\n I am storing my model to google drive and retrieving it from their. \nAlso i wanted to know in  following code :</p>\n\n<p>model_path='/content/mask_rcnn_pneumonia_0001.h5'\nmodel = modellib.MaskRCNN(mode='training', config=config, model_dir=ROOT_DIR)\nmodel.load_weights(model_path, by_name=True)</p>\n\n<p>Now if i train the model will it take care of the optimizer state and other stuff which was present during when i last trained the model?</p>",
      "rawMarkdown": "Moshel: Done with that error to. Thanks for help.\n I am storing my model to google drive and retrieving it from their. \nAlso i wanted to know in  following code :\n\nmodel_path='/content/mask_rcnn_pneumonia_0001.h5'\nmodel = modellib.MaskRCNN(mode='training', config=config, model_dir=ROOT_DIR)\nmodel.load_weights(model_path, by_name=True)\n\nNow if i train the model will it take care of the optimizer state and other stuff which was present during when i last trained the model?",
      "votes": null
    },
    {
      "id": "392558",
      "postDate": "09/24/2018 01:19:40",
      "content": "<p>It should. I have noticed it takes about ten min to get back to where it was, probably because of the running averages. This does not effect the training much</p>",
      "rawMarkdown": "It should. I have noticed it takes about ten min to get back to where it was, probably because of the running averages. This does not effect the training much",
      "votes": null
    },
    {
      "id": "392839",
      "postDate": "09/24/2018 13:12:54",
      "content": "<p>Ok. Thanks</p>",
      "rawMarkdown": "Ok. Thanks",
      "votes": null
    },
    {
      "id": "392843",
      "postDate": "09/24/2018 13:20:41",
      "content": "<p>Can you tell me how to combine two different model prediction to score better in this task? Ensemble by taking average for bounding box ?</p>",
      "rawMarkdown": "Can you tell me how to combine two different model prediction to score better in this task? Ensemble by taking average for bounding box ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 390768,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "09/20/2018 19:34:09",
      "content": "<p>With a bit of effort you can restart your training from more or less where you stopped. You can also use google colab and get 12h. Otherwise, you have to pay for Google cloud or Amazon (unless you never signed up with them - in this case Google gives you 300$ credit)</p>",
      "votes": null,
      "replies": [
        {
          "id": 390913,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/21/2018 03:06:31",
          "content": "<p>How can I stop my training and restart from where I stopped?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 390920,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 03:16:54",
          "content": "<p>Depends on what framework you use. If you make sure your snapshots are saved  and there are not too many files in the root, your snapshot will appear in the output of the kernel and can be used as dataset for another one (a fork)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391033,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/21/2018 06:55:35",
          "content": "<p>Can you guide me to a example?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391053,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/21/2018 07:26:12",
          "content": "<p>Ummm all my examples qre private kernels but... Here are some more retails:\nWhen you commit a kernel, after it finish successfuly, it has \"output\" tab. If you kept your working directory tidy and removed all unnecessary things, everything you left in it will show on this tab,so make sure your backup/snapshot /checkpoint is left there.\nNext, fork the notebook, click \"add data source\", choose \"kernel output\" and choose the kernel that just finished. When starting the training again, load the pre trained weights from the proper directory</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391067,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/21/2018 08:20:48",
          "content": "<p>Ok thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391727,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/22/2018 09:26:40",
          "content": "<p>I was thinking of training the model and save it using keras along with architecture and state of optimizer as .h5. Downloading it and starting a new kernel , loading the previous model and training it again. will this workout? will my model by initialized by previous weights and optimizer state will start from where i saved?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391729,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/22/2018 09:46:14",
          "content": "<p>How can I save a mrcnn model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391742,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/22/2018 10:25:18",
          "content": "<p>You can do that. The h5 loads tge complete state (almost). My way is much more efficient but yours will work as well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 391784,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/22/2018 12:28:09",
          "content": "<p>not able to load saved file. following error occurs:\nOSError: Unable to open file (truncated file: eof = 38932208, sblock-&gt;base_addr = 0, stored_eof = 179179456)</p>\n\n<p>Can you explain your way a little deeper?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 392007,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/22/2018 20:45:58",
          "content": "<p>This is a bug in your code. Try to iron out the kinks by writing a kernel that saves and load the model successfuly in the same code. When you get this right, you can progress to doing it across kernels</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 392195,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/23/2018 08:50:44",
          "content": "<p>i figured it out. now this error arises when i try to load my model:\n'utf8' codec can't decode byte 0x9c in position 131: invalid start byte</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 392216,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/23/2018 09:25:27",
          "content": "<p>Until you figure out how to save and load a model, there is very little to be done. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 392252,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/23/2018 11:08:38",
          "content": "<p>done with importingthe saved model to new kernel, but not able to retrain it. while loading the model following error occurs: \nOSError: Unable to open file (file signature not found)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 392271,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/23/2018 12:37:58",
          "content": "<p>@Moshel: Done with that error to. Thanks for help.\n I am storing my model to google drive and retrieving it from their. \nAlso i wanted to know in  following code :</p>\n\n<p>model_path='/content/mask_rcnn_pneumonia_0001.h5'\nmodel = modellib.MaskRCNN(mode='training', config=config, model_dir=ROOT_DIR)\nmodel.load_weights(model_path, by_name=True)</p>\n\n<p>Now if i train the model will it take care of the optimizer state and other stuff which was present during when i last trained the model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 392558,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "09/24/2018 01:19:40",
          "content": "<p>It should. I have noticed it takes about ten min to get back to where it was, probably because of the running averages. This does not effect the training much</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 392839,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/24/2018 13:12:54",
          "content": "<p>Ok. Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 392843,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "09/24/2018 13:20:41",
          "content": "<p>Can you tell me how to combine two different model prediction to score better in this task? Ensemble by taking average for bounding box ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "390743": "I am not able to train my model fully as the kernel dies after 6 hours. I don't have a high end pc. Can somebody tell me what setup or platforms they are using for training?",
    "390768": "With a bit of effort you can restart your training from more or less where you stopped. You can also use google colab and get 12h. Otherwise, you have to pay for Google cloud or Amazon (unless you never signed up with them - in this case Google gives you 300$ credit)",
    "390913": "How can I stop my training and restart from where I stopped?",
    "390920": "Depends on what framework you use. If you make sure your snapshots are saved  and there are not too many files in the root, your snapshot will appear in the output of the kernel and can be used as dataset for another one (a fork)",
    "391033": "Can you guide me to a example?",
    "391053": "Ummm all my examples qre private kernels but... Here are some more retails:\nWhen you commit a kernel, after it finish successfuly, it has \"output\" tab. If you kept your working directory tidy and removed all unnecessary things, everything you left in it will show on this tab,so make sure your backup/snapshot /checkpoint is left there.\nNext, fork the notebook, click \"add data source\", choose \"kernel output\" and choose the kernel that just finished. When starting the training again, load the pre trained weights from the proper directory",
    "391067": "Ok thanks",
    "391727": "I was thinking of training the model and save it using keras along with architecture and state of optimizer as .h5. Downloading it and starting a new kernel , loading the previous model and training it again. will this workout? will my model by initialized by previous weights and optimizer state will start from where i saved?",
    "391729": "How can I save a mrcnn model?",
    "391742": "You can do that. The h5 loads tge complete state (almost). My way is much more efficient but yours will work as well",
    "391784": "not able to load saved file. following error occurs:\nOSError: Unable to open file (truncated file: eof = 38932208, sblock-&gt;base_addr = 0, stored_eof = 179179456)\n\nCan you explain your way a little deeper?",
    "392007": "This is a bug in your code. Try to iron out the kinks by writing a kernel that saves and load the model successfuly in the same code. When you get this right, you can progress to doing it across kernels",
    "392195": "i figured it out. now this error arises when i try to load my model:\n'utf8' codec can't decode byte 0x9c in position 131: invalid start byte",
    "392216": "Until you figure out how to save and load a model, there is very little to be done.",
    "392252": "done with importingthe saved model to new kernel, but not able to retrain it. while loading the model following error occurs: \nOSError: Unable to open file (file signature not found)",
    "392271": "Moshel: Done with that error to. Thanks for help.\n I am storing my model to google drive and retrieving it from their. \nAlso i wanted to know in  following code :\n\nmodel_path='/content/mask_rcnn_pneumonia_0001.h5'\nmodel = modellib.MaskRCNN(mode='training', config=config, model_dir=ROOT_DIR)\nmodel.load_weights(model_path, by_name=True)\n\nNow if i train the model will it take care of the optimizer state and other stuff which was present during when i last trained the model?",
    "392558": "It should. I have noticed it takes about ten min to get back to where it was, probably because of the running averages. This does not effect the training much",
    "392839": "Ok. Thanks",
    "392843": "Can you tell me how to combine two different model prediction to score better in this task? Ensemble by taking average for bounding box ?"
  },
  "source": "meta"
}