{
  "id": 213006,
  "title": "can't submit prediction time exceeded error",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/213006",
  "author_name": "",
  "post_date": "2021-01-21T06:53:56.978273300Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nI created separate  submission notebook and generated submission csv but when I submits notebook its taking too long and ends with an error.</p>",
  "messages": [
    {
      "id": "1162467",
      "postDate": "01/21/2021 06:53:56",
      "content": "<p>Hello everyone,<br>\nI created separate  submission notebook and generated submission csv but when I submits notebook its taking too long and ends with an error.</p>",
      "rawMarkdown": "Hello everyone,\nI created separate  submission notebook and generated submission csv but when I submits notebook its taking too long and ends with an error.",
      "votes": null
    },
    {
      "id": "1163424",
      "postDate": "01/21/2021 17:04:37",
      "content": "<p>Pretty hard for folks to provide an answer for this question - there are a whole long list of things that can lead to time out errors.  If you get no \"solutions\" in the next day or so - make your kernel public and come back with a new post that includes a link.  Also make any datasets (with the model file) public - than folks can run your code and give you much better feedback.</p>\n<p>In most completions Kaggle provides a larger public test set than 1 image - so you can get an idea of timing issues pretty quick.  In this competion a single test image is of marginal help.  Regardless - put a %%time on your prediction cell - RunAll within the kernel and see how fast your code predicts 1 image - multiple that by 15,000. (remove the %%time when you save the code later)</p>\n<p>My quick list of things to change when I get a timeout error (I have had a decent number of them for this competition :)</p>\n<p>If TTA - reduce the number to 3 or less.  If that works - slowly creep up the number with more submissions trys.<br>\nMake sure your batch_size is decent if using a generator - 16 to 32 has worked for me depending on the model size.<br>\nMy timeouts came with an assembly of models - hoped to use a dozen but had to drop back to 5 to avoid timeout.<br>\nComment out any fluff you might have - for example, don't display 20 images to show the effect of augmentation, etc.<br>\nBe sure that you have read the log - non fatal errors can cause time issues.<br>\nCheck the time of execution for your initial save - anything longer than a minute when your code did the single public test image is indication you have some bottleneck.<br>\nTurn on GPU accelerator.</p>",
      "rawMarkdown": "Pretty hard for folks to provide an answer for this question - there are a whole long list of things that can lead to time out errors.  If you get no \"solutions\" in the next day or so - make your kernel public and come back with a new post that includes a link.  Also make any datasets (with the model file) public - than folks can run your code and give you much better feedback.\n\nIn most completions Kaggle provides a larger public test set than 1 image - so you can get an idea of timing issues pretty quick.  In this competion a single test image is of marginal help.  Regardless - put a %%time on your prediction cell - RunAll within the kernel and see how fast your code predicts 1 image - multiple that by 15,000. (remove the %%time when you save the code later)\n\nMy quick list of things to change when I get a timeout error (I have had a decent number of them for this competition :)\n\nIf TTA - reduce the number to 3 or less.  If that works - slowly creep up the number with more submissions trys.\nMake sure your batch_size is decent if using a generator - 16 to 32 has worked for me depending on the model size.\nMy timeouts came with an assembly of models - hoped to use a dozen but had to drop back to 5 to avoid timeout.\nComment out any fluff you might have - for example, don't display 20 images to show the effect of augmentation, etc.\nBe sure that you have read the log - non fatal errors can cause time issues.\nCheck the time of execution for your initial save - anything longer than a minute when your code did the single public test image is indication you have some bottleneck.\nTurn on GPU accelerator.",
      "votes": null
    },
    {
      "id": "1163888",
      "postDate": "01/22/2021 02:54:45",
      "content": "<p>I would add to try submitting your submission under a GPU accelerator, it surely sped up my submission runtime.</p>",
      "rawMarkdown": "I would add to try submitting your submission under a GPU accelerator, it surely sped up my submission runtime.",
      "votes": null
    },
    {
      "id": "1166425",
      "postDate": "01/23/2021 16:06:57",
      "content": "<p>The runtime of your notebook is highly dependent on (1) if your accelerator is on (GPU since TPU isn't allowed in the submission portion of the competition) and (2) your model design. If your GPU accelerator is on, there's a high chance that your model is too slow (or too big) and goes over the 9 hour runtime limit as specified in the competition overview.</p>\n<p>I've previously tried running an ensemble of 3 models + TTA and was still able to get around ~4 hrs runtime. That said, it may be better to explore smaller models (e.g. smaller variations of EfficientNet/ResNet/ResNext) and try running them individually. Since the final public scoring will be based on the remaining hidden dataset, it would be better to not go over the 5 hour mark (you can monitor the runtime of your model programmatically or by regularly checking when your submission finishes.)</p>\n<p>Good luck!</p>",
      "rawMarkdown": "The runtime of your notebook is highly dependent on (1) if your accelerator is on (GPU since TPU isn't allowed in the submission portion of the competition) and (2) your model design. If your GPU accelerator is on, there's a high chance that your model is too slow (or too big) and goes over the 9 hour runtime limit as specified in the competition overview.\n\nI've previously tried running an ensemble of 3 models + TTA and was still able to get around ~4 hrs runtime. That said, it may be better to explore smaller models (e.g. smaller variations of EfficientNet/ResNet/ResNext) and try running them individually. Since the final public scoring will be based on the remaining hidden dataset, it would be better to not go over the 5 hour mark (you can monitor the runtime of your model programmatically or by regularly checking when your submission finishes.)\n\nGood luck!",
      "votes": null
    },
    {
      "id": "1166551",
      "postDate": "01/23/2021 17:18:37",
      "content": "<p><a href=\"https://www.kaggle.com/amielle\" target=\"_blank\">aim</a></p>\n<p>It is my understanding that the full test set is run when we make a submission, but only a percentage is scored and reported to us.  There is a post or two in this competition that talks about a leak that kaggle had were folks could look at the final LB score for a submission.</p>\n<p>So if your submission makes it under the 9 hours limit your good to go for that submission.</p>",
      "rawMarkdown": "[aim](https://www.kaggle.com/amielle)\n\nIt is my understanding that the full test set is run when we make a submission, but only a percentage is scored and reported to us.  There is a post or two in this competition that talks about a leak that kaggle had were folks could look at the final LB score for a submission.\n\nSo if your submission makes it under the 9 hours limit your good to go for that submission.",
      "votes": null
    },
    {
      "id": "1166725",
      "postDate": "01/23/2021 19:40:03",
      "content": "<p>Ohh that would make sense. Thank you for the clarification :)</p>",
      "rawMarkdown": "Ohh that would make sense. Thank you for the clarification :)",
      "votes": null
    },
    {
      "id": "1168265",
      "postDate": "01/24/2021 19:56:13",
      "content": "<p>I too am facing issues in this regard. I am using a single resnet based model (5 fold, no tta) with bs 32 and image size of 512. During training for each fold,  1 epoch runs at about 5 minutes(train set and validation set). But during submission it's taking hours. For 4-5 submissions it was completed within the 9 hr limit, but the same inference pipeline (taken mostly from <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> wonderful notebook) is now showing timeout error. I will make the notebook public soon as <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> suggested. But would like to know if others are facing such issues. </p>",
      "rawMarkdown": "I too am facing issues in this regard. I am using a single resnet based model (5 fold, no tta) with bs 32 and image size of 512. During training for each fold,  1 epoch runs at about 5 minutes(train set and validation set). But during submission it's taking hours. For 4-5 submissions it was completed within the 9 hr limit, but the same inference pipeline (taken mostly from @pestipeti wonderful notebook) is now showing timeout error. I will make the notebook public soon as @pcjimmmy suggested. But would like to know if others are facing such issues.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1163424,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/21/2021 17:04:37",
      "content": "<p>Pretty hard for folks to provide an answer for this question - there are a whole long list of things that can lead to time out errors.  If you get no \"solutions\" in the next day or so - make your kernel public and come back with a new post that includes a link.  Also make any datasets (with the model file) public - than folks can run your code and give you much better feedback.</p>\n<p>In most completions Kaggle provides a larger public test set than 1 image - so you can get an idea of timing issues pretty quick.  In this competion a single test image is of marginal help.  Regardless - put a %%time on your prediction cell - RunAll within the kernel and see how fast your code predicts 1 image - multiple that by 15,000. (remove the %%time when you save the code later)</p>\n<p>My quick list of things to change when I get a timeout error (I have had a decent number of them for this competition :)</p>\n<p>If TTA - reduce the number to 3 or less.  If that works - slowly creep up the number with more submissions trys.<br>\nMake sure your batch_size is decent if using a generator - 16 to 32 has worked for me depending on the model size.<br>\nMy timeouts came with an assembly of models - hoped to use a dozen but had to drop back to 5 to avoid timeout.<br>\nComment out any fluff you might have - for example, don't display 20 images to show the effect of augmentation, etc.<br>\nBe sure that you have read the log - non fatal errors can cause time issues.<br>\nCheck the time of execution for your initial save - anything longer than a minute when your code did the single public test image is indication you have some bottleneck.<br>\nTurn on GPU accelerator.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1163888,
      "author_name": "capiru",
      "author_url": "",
      "post_date": "01/22/2021 02:54:45",
      "content": "<p>I would add to try submitting your submission under a GPU accelerator, it surely sped up my submission runtime.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1166425,
      "author_name": "amielle",
      "author_url": "",
      "post_date": "01/23/2021 16:06:57",
      "content": "<p>The runtime of your notebook is highly dependent on (1) if your accelerator is on (GPU since TPU isn't allowed in the submission portion of the competition) and (2) your model design. If your GPU accelerator is on, there's a high chance that your model is too slow (or too big) and goes over the 9 hour runtime limit as specified in the competition overview.</p>\n<p>I've previously tried running an ensemble of 3 models + TTA and was still able to get around ~4 hrs runtime. That said, it may be better to explore smaller models (e.g. smaller variations of EfficientNet/ResNet/ResNext) and try running them individually. Since the final public scoring will be based on the remaining hidden dataset, it would be better to not go over the 5 hour mark (you can monitor the runtime of your model programmatically or by regularly checking when your submission finishes.)</p>\n<p>Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1166551,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/23/2021 17:18:37",
          "content": "<p><a href=\"https://www.kaggle.com/amielle\" target=\"_blank\">aim</a></p>\n<p>It is my understanding that the full test set is run when we make a submission, but only a percentage is scored and reported to us.  There is a post or two in this competition that talks about a leak that kaggle had were folks could look at the final LB score for a submission.</p>\n<p>So if your submission makes it under the 9 hours limit your good to go for that submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1166725,
          "author_name": "amielle",
          "author_url": "",
          "post_date": "01/23/2021 19:40:03",
          "content": "<p>Ohh that would make sense. Thank you for the clarification :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1168265,
      "author_name": "suryajrrafl",
      "author_url": "",
      "post_date": "01/24/2021 19:56:13",
      "content": "<p>I too am facing issues in this regard. I am using a single resnet based model (5 fold, no tta) with bs 32 and image size of 512. During training for each fold,  1 epoch runs at about 5 minutes(train set and validation set). But during submission it's taking hours. For 4-5 submissions it was completed within the 9 hr limit, but the same inference pipeline (taken mostly from <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> wonderful notebook) is now showing timeout error. I will make the notebook public soon as <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> suggested. But would like to know if others are facing such issues. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1162467": "Hello everyone,\nI created separate  submission notebook and generated submission csv but when I submits notebook its taking too long and ends with an error.",
    "1163424": "Pretty hard for folks to provide an answer for this question - there are a whole long list of things that can lead to time out errors.  If you get no \"solutions\" in the next day or so - make your kernel public and come back with a new post that includes a link.  Also make any datasets (with the model file) public - than folks can run your code and give you much better feedback.\n\nIn most completions Kaggle provides a larger public test set than 1 image - so you can get an idea of timing issues pretty quick.  In this competion a single test image is of marginal help.  Regardless - put a %%time on your prediction cell - RunAll within the kernel and see how fast your code predicts 1 image - multiple that by 15,000. (remove the %%time when you save the code later)\n\nMy quick list of things to change when I get a timeout error (I have had a decent number of them for this competition :)\n\nIf TTA - reduce the number to 3 or less.  If that works - slowly creep up the number with more submissions trys.\nMake sure your batch_size is decent if using a generator - 16 to 32 has worked for me depending on the model size.\nMy timeouts came with an assembly of models - hoped to use a dozen but had to drop back to 5 to avoid timeout.\nComment out any fluff you might have - for example, don't display 20 images to show the effect of augmentation, etc.\nBe sure that you have read the log - non fatal errors can cause time issues.\nCheck the time of execution for your initial save - anything longer than a minute when your code did the single public test image is indication you have some bottleneck.\nTurn on GPU accelerator.",
    "1163888": "I would add to try submitting your submission under a GPU accelerator, it surely sped up my submission runtime.",
    "1166425": "The runtime of your notebook is highly dependent on (1) if your accelerator is on (GPU since TPU isn't allowed in the submission portion of the competition) and (2) your model design. If your GPU accelerator is on, there's a high chance that your model is too slow (or too big) and goes over the 9 hour runtime limit as specified in the competition overview.\n\nI've previously tried running an ensemble of 3 models + TTA and was still able to get around ~4 hrs runtime. That said, it may be better to explore smaller models (e.g. smaller variations of EfficientNet/ResNet/ResNext) and try running them individually. Since the final public scoring will be based on the remaining hidden dataset, it would be better to not go over the 5 hour mark (you can monitor the runtime of your model programmatically or by regularly checking when your submission finishes.)\n\nGood luck!",
    "1166551": "[aim](https://www.kaggle.com/amielle)\n\nIt is my understanding that the full test set is run when we make a submission, but only a percentage is scored and reported to us.  There is a post or two in this competition that talks about a leak that kaggle had were folks could look at the final LB score for a submission.\n\nSo if your submission makes it under the 9 hours limit your good to go for that submission.",
    "1166725": "Ohh that would make sense. Thank you for the clarification :)",
    "1168265": "I too am facing issues in this regard. I am using a single resnet based model (5 fold, no tta) with bs 32 and image size of 512. During training for each fold,  1 epoch runs at about 5 minutes(train set and validation set). But during submission it's taking hours. For 4-5 submissions it was completed within the 9 hr limit, but the same inference pipeline (taken mostly from @pestipeti wonderful notebook) is now showing timeout error. I will make the notebook public soon as @pcjimmmy suggested. But would like to know if others are facing such issues."
  },
  "source": "meta"
}