{
  "id": 206068,
  "title": "Help needed to resolve submissions error",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/206068",
  "author_name": "",
  "post_date": "2020-12-23T05:18:59.339090500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello Kaggle Community,</p>\n<p>I have gone through a few discussions on the same topic but it has not helped yet.<br>\nFirst, let me share what I am doing. I have trained a Deeplab v3 model in PyTorch and doing inference for test images. The whole image is divided into tiles of the patch size and then weighted stitching is done prediction to recreate the prediction image of the test image. This prediction image is used to get RLE. Also, there is an overlap of 6.25% of the patch size.<br>\nTimings-</p>\n<ol>\n<li>using Pytorch data loader/dataset - patch size 512- num worker 2- it takes ~350s /image</li>\n<li>using Pytorch data loader/dataset - patch size 1024 (this reduces no of patches to be evaluated, later resize to 512)- it takes ~180s/image</li>\n<li>using TensorFlow Dataset - patch size 512- it takes ~340s/image<br>\nAvg memory usage is 10-11GB.</li>\n</ol>\n<p>All my submissions are giving Notebook timeout error.</p>",
  "messages": [
    {
      "id": "1123281",
      "postDate": "12/23/2020 05:18:59",
      "content": "<p>Hello Kaggle Community,</p>\n<p>I have gone through a few discussions on the same topic but it has not helped yet.<br>\nFirst, let me share what I am doing. I have trained a Deeplab v3 model in PyTorch and doing inference for test images. The whole image is divided into tiles of the patch size and then weighted stitching is done prediction to recreate the prediction image of the test image. This prediction image is used to get RLE. Also, there is an overlap of 6.25% of the patch size.<br>\nTimings-</p>\n<ol>\n<li>using Pytorch data loader/dataset - patch size 512- num worker 2- it takes ~350s /image</li>\n<li>using Pytorch data loader/dataset - patch size 1024 (this reduces no of patches to be evaluated, later resize to 512)- it takes ~180s/image</li>\n<li>using TensorFlow Dataset - patch size 512- it takes ~340s/image<br>\nAvg memory usage is 10-11GB.</li>\n</ol>\n<p>All my submissions are giving Notebook timeout error.</p>",
      "rawMarkdown": "Hello Kaggle Community,\n\nI have gone through a few discussions on the same topic but it has not helped yet.\nFirst, let me share what I am doing. I have trained a Deeplab v3 model in PyTorch and doing inference for test images. The whole image is divided into tiles of the patch size and then weighted stitching is done prediction to recreate the prediction image of the test image. This prediction image is used to get RLE. Also, there is an overlap of 6.25% of the patch size.\nTimings-\n1. using Pytorch data loader/dataset - patch size 512- num worker 2- it takes ~350s /image\n2. using Pytorch data loader/dataset - patch size 1024 (this reduces no of patches to be evaluated, later resize to 512)- it takes ~180s/image\n3. using TensorFlow Dataset - patch size 512- it takes ~340s/image\nAvg memory usage is 10-11GB.\n\nAll my submissions are giving Notebook timeout error.",
      "votes": null
    },
    {
      "id": "1123766",
      "postDate": "12/23/2020 13:40:32",
      "content": "<p>I have taken a notebook from the community with successful submission.  Its timings for the first 5 images are like this -</p>\n<pre><code>Time taken 1013.639083147049\n2 Predicting b9a3865fc\nTime taken 822.0375480651855\n3 Predicting c68fe75ea\nTime taken 855.8196265697479\n4 Predicting b2dc8411c\nTime taken 279.4383599758148\n5 Predicting 26dc41664\nTime taken 984.3189227581024\n</code></pre>\n<p><strong>While my code timeing is better than this but it got failed in submission-</strong></p>\n<pre><code>1 Predicting /kaggle/input/hubmap-kidney-segmentation/test/afa5e8098.tiff\n1303\nInference done....\nDone.. afa5e8098\nTotal Time Taken: 170.7640745639801\n2 Predicting /kaggle/input/hubmap-kidney-segmentation/test/b9a3865fc.tiff\n948\nInference done....\nDone.. b9a3865fc\nTotal Time Taken: 145.51686358451843\n3 Predicting /kaggle/input/hubmap-kidney-segmentation/test/c68fe75ea.tiff\n1315\nInference done....\nDone.. c68fe75ea\nTotal Time Taken: 136.53821325302124\n4 Predicting /kaggle/input/hubmap-kidney-segmentation/test/b2dc8411c.tiff\n252\nInference done....\nDone.. b2dc8411c\nTotal Time Taken: 43.90383338928223\n5 Predicting /kaggle/input/hubmap-kidney-segmentation/test/26dc41664.tiff\n979\nInference done....\nDone.. 26dc41664\nTotal Time Taken: 120.4419457912445\n</code></pre>",
      "rawMarkdown": "I have taken a notebook from the community with successful submission.  Its timings for the first 5 images are like this -\n```\nTime taken 1013.639083147049\n2 Predicting b9a3865fc\nTime taken 822.0375480651855\n3 Predicting c68fe75ea\nTime taken 855.8196265697479\n4 Predicting b2dc8411c\nTime taken 279.4383599758148\n5 Predicting 26dc41664\nTime taken 984.3189227581024\n```\n**While my code timeing is better than this but it got failed in submission-**\n\n\n```\n1 Predicting /kaggle/input/hubmap-kidney-segmentation/test/afa5e8098.tiff\n1303\nInference done....\nDone.. afa5e8098\nTotal Time Taken: 170.7640745639801\n2 Predicting /kaggle/input/hubmap-kidney-segmentation/test/b9a3865fc.tiff\n948\nInference done....\nDone.. b9a3865fc\nTotal Time Taken: 145.51686358451843\n3 Predicting /kaggle/input/hubmap-kidney-segmentation/test/c68fe75ea.tiff\n1315\nInference done....\nDone.. c68fe75ea\nTotal Time Taken: 136.53821325302124\n4 Predicting /kaggle/input/hubmap-kidney-segmentation/test/b2dc8411c.tiff\n252\nInference done....\nDone.. b2dc8411c\nTotal Time Taken: 43.90383338928223\n5 Predicting /kaggle/input/hubmap-kidney-segmentation/test/26dc41664.tiff\n979\nInference done....\nDone.. 26dc41664\nTotal Time Taken: 120.4419457912445\n```",
      "votes": null
    },
    {
      "id": "1124369",
      "postDate": "12/23/2020 21:27:44",
      "content": "<p>it has been discussed in several topics already. I ran into the same issue.<br>\nThe images in the private set are larger than the ones in the public test set. If the public set image processing is taking<br>\n10-11Gb of RAM the chances are high the private images will give out of memory error which will show up as timeout on submission. The RAM available is only 16Gb for cpu-only instances and 13Gb for gpu-enabled instances.</p>",
      "rawMarkdown": "it has been discussed in several topics already. I ran into the same issue.\nThe images in the private set are larger than the ones in the public test set. If the public set image processing is taking\n10-11Gb of RAM the chances are high the private images will give out of memory error which will show up as timeout on submission. The RAM available is only 16Gb for cpu-only instances and 13Gb for gpu-enabled instances.",
      "votes": null
    },
    {
      "id": "1125810",
      "postDate": "12/25/2020 05:08:39",
      "content": "<p>I'm also having the same trouble…<br>\nIs there any information about the size or amount of private test set?</p>",
      "rawMarkdown": "I'm also having the same trouble...\nIs there any information about the size or amount of private test set?",
      "votes": null
    },
    {
      "id": "1135422",
      "postDate": "01/02/2021 08:32:51",
      "content": "<p>Adding an observation-<br>\nif your kernel takes more than expected error on \"run &amp; commit\" then there might be an error that Kaggle doesn't give you direct and it is hogging your GPU quota. So better stop the commit and see the log error. Most likely it will be an error code 137 which implies-<br>\n             Exit Code 137: Indicates failure as container received SIGKILL (Manual intervention or 'oom-killer' [OUT-OF-MEMORY]).<br>\nSince you are forcing to stop so code 137 might come due to that also but in this Kideney case, I see an OOM issue.</p>",
      "rawMarkdown": "Adding an observation-\nif your kernel takes more than expected error on \"run & commit\" then there might be an error that Kaggle doesn't give you direct and it is hogging your GPU quota. So better stop the commit and see the log error. Most likely it will be an error code 137 which implies-\n             Exit Code 137: Indicates failure as container received SIGKILL (Manual intervention or 'oom-killer' [OUT-OF-MEMORY]).\nSince you are forcing to stop so code 137 might come due to that also but in this Kideney case, I see an OOM issue.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1123766,
      "author_name": "sumitjha19",
      "author_url": "",
      "post_date": "12/23/2020 13:40:32",
      "content": "<p>I have taken a notebook from the community with successful submission.  Its timings for the first 5 images are like this -</p>\n<pre><code>Time taken 1013.639083147049\n2 Predicting b9a3865fc\nTime taken 822.0375480651855\n3 Predicting c68fe75ea\nTime taken 855.8196265697479\n4 Predicting b2dc8411c\nTime taken 279.4383599758148\n5 Predicting 26dc41664\nTime taken 984.3189227581024\n</code></pre>\n<p><strong>While my code timeing is better than this but it got failed in submission-</strong></p>\n<pre><code>1 Predicting /kaggle/input/hubmap-kidney-segmentation/test/afa5e8098.tiff\n1303\nInference done....\nDone.. afa5e8098\nTotal Time Taken: 170.7640745639801\n2 Predicting /kaggle/input/hubmap-kidney-segmentation/test/b9a3865fc.tiff\n948\nInference done....\nDone.. b9a3865fc\nTotal Time Taken: 145.51686358451843\n3 Predicting /kaggle/input/hubmap-kidney-segmentation/test/c68fe75ea.tiff\n1315\nInference done....\nDone.. c68fe75ea\nTotal Time Taken: 136.53821325302124\n4 Predicting /kaggle/input/hubmap-kidney-segmentation/test/b2dc8411c.tiff\n252\nInference done....\nDone.. b2dc8411c\nTotal Time Taken: 43.90383338928223\n5 Predicting /kaggle/input/hubmap-kidney-segmentation/test/26dc41664.tiff\n979\nInference done....\nDone.. 26dc41664\nTotal Time Taken: 120.4419457912445\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1124369,
      "author_name": "igor14497",
      "author_url": "",
      "post_date": "12/23/2020 21:27:44",
      "content": "<p>it has been discussed in several topics already. I ran into the same issue.<br>\nThe images in the private set are larger than the ones in the public test set. If the public set image processing is taking<br>\n10-11Gb of RAM the chances are high the private images will give out of memory error which will show up as timeout on submission. The RAM available is only 16Gb for cpu-only instances and 13Gb for gpu-enabled instances.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1125810,
      "author_name": "hiromichikamata",
      "author_url": "",
      "post_date": "12/25/2020 05:08:39",
      "content": "<p>I'm also having the same trouble…<br>\nIs there any information about the size or amount of private test set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1135422,
      "author_name": "sumitjha19",
      "author_url": "",
      "post_date": "01/02/2021 08:32:51",
      "content": "<p>Adding an observation-<br>\nif your kernel takes more than expected error on \"run &amp; commit\" then there might be an error that Kaggle doesn't give you direct and it is hogging your GPU quota. So better stop the commit and see the log error. Most likely it will be an error code 137 which implies-<br>\n             Exit Code 137: Indicates failure as container received SIGKILL (Manual intervention or 'oom-killer' [OUT-OF-MEMORY]).<br>\nSince you are forcing to stop so code 137 might come due to that also but in this Kideney case, I see an OOM issue.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1123281": "Hello Kaggle Community,\n\nI have gone through a few discussions on the same topic but it has not helped yet.\nFirst, let me share what I am doing. I have trained a Deeplab v3 model in PyTorch and doing inference for test images. The whole image is divided into tiles of the patch size and then weighted stitching is done prediction to recreate the prediction image of the test image. This prediction image is used to get RLE. Also, there is an overlap of 6.25% of the patch size.\nTimings-\n1. using Pytorch data loader/dataset - patch size 512- num worker 2- it takes ~350s /image\n2. using Pytorch data loader/dataset - patch size 1024 (this reduces no of patches to be evaluated, later resize to 512)- it takes ~180s/image\n3. using TensorFlow Dataset - patch size 512- it takes ~340s/image\nAvg memory usage is 10-11GB.\n\nAll my submissions are giving Notebook timeout error.",
    "1123766": "I have taken a notebook from the community with successful submission.  Its timings for the first 5 images are like this -\n```\nTime taken 1013.639083147049\n2 Predicting b9a3865fc\nTime taken 822.0375480651855\n3 Predicting c68fe75ea\nTime taken 855.8196265697479\n4 Predicting b2dc8411c\nTime taken 279.4383599758148\n5 Predicting 26dc41664\nTime taken 984.3189227581024\n```\n**While my code timeing is better than this but it got failed in submission-**\n\n\n```\n1 Predicting /kaggle/input/hubmap-kidney-segmentation/test/afa5e8098.tiff\n1303\nInference done....\nDone.. afa5e8098\nTotal Time Taken: 170.7640745639801\n2 Predicting /kaggle/input/hubmap-kidney-segmentation/test/b9a3865fc.tiff\n948\nInference done....\nDone.. b9a3865fc\nTotal Time Taken: 145.51686358451843\n3 Predicting /kaggle/input/hubmap-kidney-segmentation/test/c68fe75ea.tiff\n1315\nInference done....\nDone.. c68fe75ea\nTotal Time Taken: 136.53821325302124\n4 Predicting /kaggle/input/hubmap-kidney-segmentation/test/b2dc8411c.tiff\n252\nInference done....\nDone.. b2dc8411c\nTotal Time Taken: 43.90383338928223\n5 Predicting /kaggle/input/hubmap-kidney-segmentation/test/26dc41664.tiff\n979\nInference done....\nDone.. 26dc41664\nTotal Time Taken: 120.4419457912445\n```",
    "1124369": "it has been discussed in several topics already. I ran into the same issue.\nThe images in the private set are larger than the ones in the public test set. If the public set image processing is taking\n10-11Gb of RAM the chances are high the private images will give out of memory error which will show up as timeout on submission. The RAM available is only 16Gb for cpu-only instances and 13Gb for gpu-enabled instances.",
    "1125810": "I'm also having the same trouble...\nIs there any information about the size or amount of private test set?",
    "1135422": "Adding an observation-\nif your kernel takes more than expected error on \"run & commit\" then there might be an error that Kaggle doesn't give you direct and it is hogging your GPU quota. So better stop the commit and see the log error. Most likely it will be an error code 137 which implies-\n             Exit Code 137: Indicates failure as container received SIGKILL (Manual intervention or 'oom-killer' [OUT-OF-MEMORY]).\nSince you are forcing to stop so code 137 might come due to that also but in this Kideney case, I see an OOM issue."
  },
  "source": "meta"
}