{
  "id": 131844,
  "title": "Notebook Exceeded Allowed Compute..",
  "url": "/competitions/bengaliai-cv19/discussion/131844",
  "author_name": "",
  "post_date": "2020-02-22T05:07:41.428967600Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi guys,\n I'm quite new to Kaggle and this is my first time submitting with a kernel. Due to how big the data is, I pre-trained my model in Google Colab and then uploaded the model as a dataset in kernel. The Kernel submission is ok. But when I try to submit the output, <code>submission.csv</code> to the competition, it shows 'Notebook Exceeded Allowed Compute....' after an hour or so of running.\n Can anyone tell me what's going on? And it is allowed to predict using a pretrained model here? Many thanks.</p>",
  "messages": [
    {
      "id": "753371",
      "postDate": "02/22/2020 05:07:41",
      "content": "<p>Hi guys,\n I'm quite new to Kaggle and this is my first time submitting with a kernel. Due to how big the data is, I pre-trained my model in Google Colab and then uploaded the model as a dataset in kernel. The Kernel submission is ok. But when I try to submit the output, <code>submission.csv</code> to the competition, it shows 'Notebook Exceeded Allowed Compute....' after an hour or so of running.\n Can anyone tell me what's going on? And it is allowed to predict using a pretrained model here? Many thanks.</p>",
      "rawMarkdown": "Hi guys,\n I'm quite new to Kaggle and this is my first time submitting with a kernel. Due to how big the data is, I pre-trained my model in Google Colab and then uploaded the model as a dataset in kernel. The Kernel submission is ok. But when I try to submit the output, `submission.csv` to the competition, it shows 'Notebook Exceeded Allowed Compute....' after an hour or so of running.\n Can anyone tell me what's going on? And it is allowed to predict using a pretrained model here? Many thanks.",
      "votes": null
    },
    {
      "id": "753430",
      "postDate": "02/22/2020 06:57:27",
      "content": "<p>Most people are uploading pretrained models.</p>\n\n<p>My guess is that you used more than 13GB of RAM. Note that your submission is inferring 200,000 test images so you probably need to infer the test by batches. If you try to infer them all at once you'll probably overflow memory.</p>\n\n<p>Try predicting all the 200,000 train images in an interactive notebook and watch it run. If it runs successfully then when you predict the test images that should run successfully too.</p>",
      "rawMarkdown": "Most people are uploading pretrained models.\n\nMy guess is that you used more than 13GB of RAM. Note that your submission is inferring 200,000 test images so you probably need to infer the test by batches. If you try to infer them all at once you'll probably overflow memory.\n\nTry predicting all the 200,000 train images in an interactive notebook and watch it run. If it runs successfully then when you predict the test images that should run successfully too.",
      "votes": null
    },
    {
      "id": "753434",
      "postDate": "02/22/2020 07:16:15",
      "content": "<p>It is allowed to use a pre-trained model. If you use keras, you can try using predict_generator.</p>",
      "rawMarkdown": "It is allowed to use a pre-trained model. If you use keras, you can try using predict_generator.",
      "votes": null
    },
    {
      "id": "753662",
      "postDate": "02/22/2020 14:15:24",
      "content": "<p>Great advice. Thanks!</p>",
      "rawMarkdown": "Great advice. Thanks!",
      "votes": null
    },
    {
      "id": "753758",
      "postDate": "02/22/2020 16:20:47",
      "content": "<p>I am getting same error. I was able to validate a previous submit, but I change the preprocessing methods and I am not able to get it to work.\nI tried dividing each .parquet in 4 batches, but still no luck.\nThe strange think is that I am able to run the submission in kaggle when using the train_ file with 200.000 images without memory problems.</p>",
      "rawMarkdown": "I am getting same error. I was able to validate a previous submit, but I change the preprocessing methods and I am not able to get it to work.\nI tried dividing each .parquet in 4 batches, but still no luck.\nThe strange think is that I am able to run the submission in kaggle when using the train_ file with 200.000 images without memory problems.",
      "votes": null
    },
    {
      "id": "759173",
      "postDate": "02/28/2020 16:56:04",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Chris, can you expand on your comments about inferring 200,000 images? If it's trained externally and then uploaded, shouldn't it only be performing inference on the &lt; 20 test images when committed and the same &lt; 20 test images when the solutions to the commit are committed?</p>\n\n<p>I ask because I am running out of memory when submitting predictions, but not when when performing my commits.</p>",
      "rawMarkdown": "cdeotte Chris, can you expand on your comments about inferring 200,000 images? If it's trained externally and then uploaded, shouldn't it only be performing inference on the &lt; 20 test images when committed and the same &lt; 20 test images when the solutions to the commit are committed?\n\nI ask because I am running out of memory when submitting predictions, but not when when performing my commits.",
      "votes": null
    },
    {
      "id": "759177",
      "postDate": "02/28/2020 17:06:05",
      "content": "<p>The 12 test images that you can download are neither part of the public nor private dataset. They are only there to help you write code that handles test parquets. When you <strong>commit</strong> your notebook it uses those 4 parquets which contain 12 images. When you <strong>submit</strong> your notebook, it uses the real test data parquets which contain 200,000 images.</p>\n\n<p>Therefore, in an interactive notebook, change \"test parquet\" to \"train parquet\" and see if your model can make predictions on the 200,000 training images. If that works without memory error and under 2 hours, then your <strong>submit</strong> will work too (when you change \"train\" back to \"test\").</p>",
      "rawMarkdown": "The 12 test images that you can download are neither part of the public nor private dataset. They are only there to help you write code that handles test parquets. When you **commit** your notebook it uses those 4 parquets which contain 12 images. When you **submit** your notebook, it uses the real test data parquets which contain 200,000 images.\n\nTherefore, in an interactive notebook, change \"test parquet\" to \"train parquet\" and see if your model can make predictions on the 200,000 training images. If that works without memory error and under 2 hours, then your **submit** will work too (when you change \"train\" back to \"test\").",
      "votes": null
    },
    {
      "id": "759180",
      "postDate": "02/28/2020 17:13:47",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks so much for your reply. I see now what you're talking about -- after reading the whole data description again, I see that they say the test and sample submission they include for convenience contains only a few rows of the data, and that the full set is much closer to the size of the test set like you are saying. Thanks for the help!</p>",
      "rawMarkdown": "cdeotte Thanks so much for your reply. I see now what you're talking about -- after reading the whole data description again, I see that they say the test and sample submission they include for convenience contains only a few rows of the data, and that the full set is much closer to the size of the test set like you are saying. Thanks for the help!",
      "votes": null
    },
    {
      "id": "765460",
      "postDate": "03/06/2020 16:41:44",
      "content": "<p>Hi.. <a href=\"/xiaoyuez\">@xiaoyuez</a> . Were you able to solve this problem? I am facing the same issue now and dn't know what to do</p>",
      "rawMarkdown": "Hi.. @xiaoyuez . Were you able to solve this problem? I am facing the same issue now and dn't know what to do",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 753430,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/22/2020 06:57:27",
      "content": "<p>Most people are uploading pretrained models.</p>\n\n<p>My guess is that you used more than 13GB of RAM. Note that your submission is inferring 200,000 test images so you probably need to infer the test by batches. If you try to infer them all at once you'll probably overflow memory.</p>\n\n<p>Try predicting all the 200,000 train images in an interactive notebook and watch it run. If it runs successfully then when you predict the test images that should run successfully too.</p>",
      "votes": null,
      "replies": [
        {
          "id": 753662,
          "author_name": "xiaoyuez",
          "author_url": "",
          "post_date": "02/22/2020 14:15:24",
          "content": "<p>Great advice. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759173,
          "author_name": "egrimley",
          "author_url": "",
          "post_date": "02/28/2020 16:56:04",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Chris, can you expand on your comments about inferring 200,000 images? If it's trained externally and then uploaded, shouldn't it only be performing inference on the &lt; 20 test images when committed and the same &lt; 20 test images when the solutions to the commit are committed?</p>\n\n<p>I ask because I am running out of memory when submitting predictions, but not when when performing my commits.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759177,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/28/2020 17:06:05",
          "content": "<p>The 12 test images that you can download are neither part of the public nor private dataset. They are only there to help you write code that handles test parquets. When you <strong>commit</strong> your notebook it uses those 4 parquets which contain 12 images. When you <strong>submit</strong> your notebook, it uses the real test data parquets which contain 200,000 images.</p>\n\n<p>Therefore, in an interactive notebook, change \"test parquet\" to \"train parquet\" and see if your model can make predictions on the 200,000 training images. If that works without memory error and under 2 hours, then your <strong>submit</strong> will work too (when you change \"train\" back to \"test\").</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759180,
          "author_name": "egrimley",
          "author_url": "",
          "post_date": "02/28/2020 17:13:47",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Thanks so much for your reply. I see now what you're talking about -- after reading the whole data description again, I see that they say the test and sample submission they include for convenience contains only a few rows of the data, and that the full set is much closer to the size of the test set like you are saying. Thanks for the help!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 753434,
      "author_name": "kalinaxl",
      "author_url": "",
      "post_date": "02/22/2020 07:16:15",
      "content": "<p>It is allowed to use a pre-trained model. If you use keras, you can try using predict_generator.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 753758,
      "author_name": "macarrony00",
      "author_url": "",
      "post_date": "02/22/2020 16:20:47",
      "content": "<p>I am getting same error. I was able to validate a previous submit, but I change the preprocessing methods and I am not able to get it to work.\nI tried dividing each .parquet in 4 batches, but still no luck.\nThe strange think is that I am able to run the submission in kaggle when using the train_ file with 200.000 images without memory problems.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 765460,
      "author_name": "kvsnoufal",
      "author_url": "",
      "post_date": "03/06/2020 16:41:44",
      "content": "<p>Hi.. <a href=\"/xiaoyuez\">@xiaoyuez</a> . Were you able to solve this problem? I am facing the same issue now and dn't know what to do</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "753371": "Hi guys,\n I'm quite new to Kaggle and this is my first time submitting with a kernel. Due to how big the data is, I pre-trained my model in Google Colab and then uploaded the model as a dataset in kernel. The Kernel submission is ok. But when I try to submit the output, `submission.csv` to the competition, it shows 'Notebook Exceeded Allowed Compute....' after an hour or so of running.\n Can anyone tell me what's going on? And it is allowed to predict using a pretrained model here? Many thanks.",
    "753430": "Most people are uploading pretrained models.\n\nMy guess is that you used more than 13GB of RAM. Note that your submission is inferring 200,000 test images so you probably need to infer the test by batches. If you try to infer them all at once you'll probably overflow memory.\n\nTry predicting all the 200,000 train images in an interactive notebook and watch it run. If it runs successfully then when you predict the test images that should run successfully too.",
    "753434": "It is allowed to use a pre-trained model. If you use keras, you can try using predict_generator.",
    "753662": "Great advice. Thanks!",
    "753758": "I am getting same error. I was able to validate a previous submit, but I change the preprocessing methods and I am not able to get it to work.\nI tried dividing each .parquet in 4 batches, but still no luck.\nThe strange think is that I am able to run the submission in kaggle when using the train_ file with 200.000 images without memory problems.",
    "759173": "cdeotte Chris, can you expand on your comments about inferring 200,000 images? If it's trained externally and then uploaded, shouldn't it only be performing inference on the &lt; 20 test images when committed and the same &lt; 20 test images when the solutions to the commit are committed?\n\nI ask because I am running out of memory when submitting predictions, but not when when performing my commits.",
    "759177": "The 12 test images that you can download are neither part of the public nor private dataset. They are only there to help you write code that handles test parquets. When you **commit** your notebook it uses those 4 parquets which contain 12 images. When you **submit** your notebook, it uses the real test data parquets which contain 200,000 images.\n\nTherefore, in an interactive notebook, change \"test parquet\" to \"train parquet\" and see if your model can make predictions on the 200,000 training images. If that works without memory error and under 2 hours, then your **submit** will work too (when you change \"train\" back to \"test\").",
    "759180": "cdeotte Thanks so much for your reply. I see now what you're talking about -- after reading the whole data description again, I see that they say the test and sample submission they include for convenience contains only a few rows of the data, and that the full set is much closer to the size of the test set like you are saying. Thanks for the help!",
    "765460": "Hi.. @xiaoyuez . Were you able to solve this problem? I am facing the same issue now and dn't know what to do"
  },
  "source": "meta"
}