{
  "id": 130099,
  "title": "\"Notebook Exceeded Allowed Compute\" cause",
  "url": "/competitions/bengaliai-cv19/discussion/130099",
  "author_name": "",
  "post_date": "2020-02-12T06:46:51.510418900Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p><strong>I have a notebook with vgg16 (with my own weights) performed on the data.</strong> I ran and commited it successfully, but got \"Notebook Exceeded Allowed Compute\". \n<strong>How to fix this?</strong> By the way, IMAGE SIZE I use is 128 and BATCH SIZE is 128 too. Can these values cause out of memory errors?</p>",
  "messages": [
    {
      "id": "743640",
      "postDate": "02/12/2020 06:46:51",
      "content": "<p><strong>I have a notebook with vgg16 (with my own weights) performed on the data.</strong> I ran and commited it successfully, but got \"Notebook Exceeded Allowed Compute\". \n<strong>How to fix this?</strong> By the way, IMAGE SIZE I use is 128 and BATCH SIZE is 128 too. Can these values cause out of memory errors?</p>",
      "rawMarkdown": "**I have a notebook with vgg16 (with my own weights) performed on the data.** I ran and commited it successfully, but got \"Notebook Exceeded Allowed Compute\". \n**How to fix this?** By the way, IMAGE SIZE I use is 128 and BATCH SIZE is 128 too. Can these values cause out of memory errors?",
      "votes": null
    },
    {
      "id": "743694",
      "postDate": "02/12/2020 07:49:05",
      "content": "<p>Maximum time allowed for the notebook to run with GPU is two hours. It mostly crossed that limit. Check the discussions for optimizing your inference kernel and reduce runtime. </p>",
      "rawMarkdown": "Maximum time allowed for the notebook to run with GPU is two hours. It mostly crossed that limit. Check the discussions for optimizing your inference kernel and reduce runtime.",
      "votes": null
    },
    {
      "id": "743735",
      "postDate": "02/12/2020 08:43:51",
      "content": "<p>Since, this is a kernel based submission, it will run the whole training process again when running on test data. I would suggest saving the model and then running the inference from saved model for final submission. But even then the memory use has to be optimized, as in actual run the test parquet files will have many thousands of images.</p>",
      "rawMarkdown": "Since, this is a kernel based submission, it will run the whole training process again when running on test data. I would suggest saving the model and then running the inference from saved model for final submission. But even then the memory use has to be optimized, as in actual run the test parquet files will have many thousands of images.",
      "votes": null
    },
    {
      "id": "743802",
      "postDate": "02/12/2020 10:19:04",
      "content": "<p>Try reducing batch size.</p>",
      "rawMarkdown": "Try reducing batch size.",
      "votes": null
    },
    {
      "id": "744268",
      "postDate": "02/12/2020 17:30:59",
      "content": "<p>I've been having huge issues with this and I couldn't even submit 1's.  So far I found that when I go to reshape the data, that seems to be the breaking point for me.  I created a public notebook related to my findings.\n<a href=\"https://www.kaggle.com/yeayates21/cant-submit-1s-whats-wrong-please-help\">https://www.kaggle.com/yeayates21/cant-submit-1s-whats-wrong-please-help</a></p>",
      "rawMarkdown": "I've been having huge issues with this and I couldn't even submit 1's.  So far I found that when I go to reshape the data, that seems to be the breaking point for me.  I created a public notebook related to my findings.\nhttps://www.kaggle.com/yeayates21/cant-submit-1s-whats-wrong-please-help",
      "votes": null
    },
    {
      "id": "746042",
      "postDate": "02/14/2020 14:33:12",
      "content": "<p><a href=\"/dwdkills\">@dwdkills</a> this error is because you are running out of RAM during the submission process. The submission will make your kernel run with a larger number of test images. Reduce the batch size to 32 for 128x128. To test if your kernel is running out of memory during submission, replace the test data with train data and watch the RAM usage and you will see the kernel dies and restarts.</p>\n\n<p>You will also need to optimize your resizing function. Also if you are normalizing the images between 0 and 1 by dividing by 255, eg. (<strong>images = X_test/255</strong>), this will definitely cause the kernel to crash during submission. You can avoid this by normalizing using the cv2.normalize and changing the dtype to float. Hope this helps.</p>",
      "rawMarkdown": "dwdkills this error is because you are running out of RAM during the submission process. The submission will make your kernel run with a larger number of test images. Reduce the batch size to 32 for 128x128. To test if your kernel is running out of memory during submission, replace the test data with train data and watch the RAM usage and you will see the kernel dies and restarts.\n\nYou will also need to optimize your resizing function. Also if you are normalizing the images between 0 and 1 by dividing by 255, eg. (**images = X_test/255**), this will definitely cause the kernel to crash during submission. You can avoid this by normalizing using the cv2.normalize and changing the dtype to float. Hope this helps.",
      "votes": null
    },
    {
      "id": "746117",
      "postDate": "02/14/2020 16:39:46",
      "content": "<p>Thanks for the help. I've already added additional resizing, but without changing the batch size. Notebook is running for 3 hours and it seems to me that the problem is gone (this one at least)</p>",
      "rawMarkdown": "Thanks for the help. I've already added additional resizing, but without changing the batch size. Notebook is running for 3 hours and it seems to me that the problem is gone (this one at least)",
      "votes": null
    },
    {
      "id": "771815",
      "postDate": "03/14/2020 16:56:49",
      "content": "<p>UPDATE:  My issue was that I was trying to use 3 models instead of 1.  Even loading 1 model at a time was causing issues.  Once I switched to using just 1 model, everything worked fine :-/.</p>",
      "rawMarkdown": "UPDATE:  My issue was that I was trying to use 3 models instead of 1.  Even loading 1 model at a time was causing issues.  Once I switched to using just 1 model, everything worked fine :-/.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 743694,
      "author_name": "p4rallax",
      "author_url": "",
      "post_date": "02/12/2020 07:49:05",
      "content": "<p>Maximum time allowed for the notebook to run with GPU is two hours. It mostly crossed that limit. Check the discussions for optimizing your inference kernel and reduce runtime. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 743735,
      "author_name": "anirbank",
      "author_url": "",
      "post_date": "02/12/2020 08:43:51",
      "content": "<p>Since, this is a kernel based submission, it will run the whole training process again when running on test data. I would suggest saving the model and then running the inference from saved model for final submission. But even then the memory use has to be optimized, as in actual run the test parquet files will have many thousands of images.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 743802,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "02/12/2020 10:19:04",
      "content": "<p>Try reducing batch size.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 744268,
      "author_name": "yeayates21",
      "author_url": "",
      "post_date": "02/12/2020 17:30:59",
      "content": "<p>I've been having huge issues with this and I couldn't even submit 1's.  So far I found that when I go to reshape the data, that seems to be the breaking point for me.  I created a public notebook related to my findings.\n<a href=\"https://www.kaggle.com/yeayates21/cant-submit-1s-whats-wrong-please-help\">https://www.kaggle.com/yeayates21/cant-submit-1s-whats-wrong-please-help</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 771815,
          "author_name": "yeayates21",
          "author_url": "",
          "post_date": "03/14/2020 16:56:49",
          "content": "<p>UPDATE:  My issue was that I was trying to use 3 models instead of 1.  Even loading 1 model at a time was causing issues.  Once I switched to using just 1 model, everything worked fine :-/.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 746042,
      "author_name": "yovinyahathugoda",
      "author_url": "",
      "post_date": "02/14/2020 14:33:12",
      "content": "<p><a href=\"/dwdkills\">@dwdkills</a> this error is because you are running out of RAM during the submission process. The submission will make your kernel run with a larger number of test images. Reduce the batch size to 32 for 128x128. To test if your kernel is running out of memory during submission, replace the test data with train data and watch the RAM usage and you will see the kernel dies and restarts.</p>\n\n<p>You will also need to optimize your resizing function. Also if you are normalizing the images between 0 and 1 by dividing by 255, eg. (<strong>images = X_test/255</strong>), this will definitely cause the kernel to crash during submission. You can avoid this by normalizing using the cv2.normalize and changing the dtype to float. Hope this helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 746117,
          "author_name": "dwdkills",
          "author_url": "",
          "post_date": "02/14/2020 16:39:46",
          "content": "<p>Thanks for the help. I've already added additional resizing, but without changing the batch size. Notebook is running for 3 hours and it seems to me that the problem is gone (this one at least)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "743640": "**I have a notebook with vgg16 (with my own weights) performed on the data.** I ran and commited it successfully, but got \"Notebook Exceeded Allowed Compute\". \n**How to fix this?** By the way, IMAGE SIZE I use is 128 and BATCH SIZE is 128 too. Can these values cause out of memory errors?",
    "743694": "Maximum time allowed for the notebook to run with GPU is two hours. It mostly crossed that limit. Check the discussions for optimizing your inference kernel and reduce runtime.",
    "743735": "Since, this is a kernel based submission, it will run the whole training process again when running on test data. I would suggest saving the model and then running the inference from saved model for final submission. But even then the memory use has to be optimized, as in actual run the test parquet files will have many thousands of images.",
    "743802": "Try reducing batch size.",
    "744268": "I've been having huge issues with this and I couldn't even submit 1's.  So far I found that when I go to reshape the data, that seems to be the breaking point for me.  I created a public notebook related to my findings.\nhttps://www.kaggle.com/yeayates21/cant-submit-1s-whats-wrong-please-help",
    "746042": "dwdkills this error is because you are running out of RAM during the submission process. The submission will make your kernel run with a larger number of test images. Reduce the batch size to 32 for 128x128. To test if your kernel is running out of memory during submission, replace the test data with train data and watch the RAM usage and you will see the kernel dies and restarts.\n\nYou will also need to optimize your resizing function. Also if you are normalizing the images between 0 and 1 by dividing by 255, eg. (**images = X_test/255**), this will definitely cause the kernel to crash during submission. You can avoid this by normalizing using the cv2.normalize and changing the dtype to float. Hope this helps.",
    "746117": "Thanks for the help. I've already added additional resizing, but without changing the batch size. Notebook is running for 3 hours and it seems to me that the problem is gone (this one at least)",
    "771815": "UPDATE:  My issue was that I was trying to use 3 models instead of 1.  Even loading 1 model at a time was causing issues.  Once I switched to using just 1 model, everything worked fine :-/."
  },
  "source": "meta"
}