{
  "id": 215095,
  "title": "Why is my Submission timing out?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/215095",
  "author_name": "Mark A Lavin",
  "post_date": "2021-01-28T16:13:56.873000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I have made five submissions for the HuBMAP competition.  The first<br>\none timed out, the next two got \"Submission Scoring Error\", and the<br>\nlast two timed out again.  In all five of the cases, I was able to run<br>\nsuccessfully using the public Test set and [Save Version] ; the runs<br>\ncompleted in an hour and a half, and I was able to visualize the<br>\nresults by reading back the RLEs and decoding them, and the results<br>\nlook reasonable, with glom markers following the structure of the<br>\ninput images.</p>\n<p>So, why the timeout?  I'm going to try listing possible explanations,<br>\nand would appreciate any comments you might have.</p>\n<p>(1)  There is a bug in my code that's causing it to loop forever.  The<br>\nfact that I can submit Test cases that complete successfully contra-<br>\ndicts this.</p>\n<p>(2)  The inputs in the Private Test set are MUCH bigger, hence my<br>\nalgorithms are too slow to process them.  No way to know this, but<br>\nothers have submitted successfully.</p>\n<p>(3)  There is some algorithm inefficiency that causes <em>somewhat</em><br>\nbigger inputs to fail due to poor scaling.  The algorithm processes<br>\neach input image independently (except for non-recovered storage),<br>\nand within each image processes 512-pixel wide columns indepently,<br>\nconcatenating together the suitably-offset RLEs for each column.</p>\n<p>(4)  There are regions in the Private Test images that cause runtime<br>\nproblems for the model prediction.  However, since the prediction is<br>\na finite, determistic process (unlike fitting) this seems implausible.</p>\n<p>(5)  The run exhausts GPU allotment.  I have no visibility to this,<br>\nbut if it happens, it would certainly cause excessive runtimes.</p>\n<p>(6)  The RLE process blows up:  Given a random binary image as<br>\ninput, RLE will produce a very large string output.  However, examining<br>\nthe glom marker outputs for the Public Test set, the output images are<br>\nsparse collections of compact glom markers, far from random images.</p>",
  "messages": [
    {
      "id": 1177999,
      "postDate": "2021-01-30T15:41:55.240Z",
      "content": "<p>OK, some progress (finally!):</p>\n<p>I modified my code to only generate a trivial RLE ( \"0 1\") for one out of the N images, no output for the other N - 1.   I got a \"Submission Scoring Error\".   I then modified the code further so that it generates the trivial RLE for <em>all</em> N images, and the run actually succeeded (!), giving me a score of 0.0</p>\n<p>Several conclusions:<br>\n(1)   The overall <code>submission.csv</code> structure I'm generating is correct<br>\n(2)   The submitted <code>submission.csv</code> must contain RLE's for <em>all</em> input images</p>\n<p>My next step is to modify the code so it outputs complete RLE for one of the inputs and the trivial output for the other N-1 inputs.   We'll see whether this manages to get past the Submission Timed Out with a non-zero score.   If so, I will then start adding more images one at a time.</p>\n<p>EDIT:   Next step completed quickly, generating submission that Succeeded with a Leader Board (LB) score of 0.135  So, next step, try generating <em>two</em> \"real\" RLEs.</p>",
      "rawMarkdown": "OK, some progress (finally!):\n\nI modified my code to only generate a trivial RLE ( \"0 1\") for one out of the N images, no output for the other N - 1.   I got a \"Submission Scoring Error\".   I then modified the code further so that it generates the trivial RLE for *all* N images, and the run actually succeeded (!), giving me a score of 0.0\n\nSeveral conclusions:\n(1)   The overall ```submission.csv``` structure I'm generating is correct\n(2)   The submitted ```submission.csv``` must contain RLE's for *all* input images\n\nMy next step is to modify the code so it outputs complete RLE for one of the inputs and the trivial output for the other N-1 inputs.   We'll see whether this manages to get past the Submission Timed Out with a non-zero score.   If so, I will then start adding more images one at a time.\n\nEDIT:   Next step completed quickly, generating submission that Succeeded with a Leader Board (LB) score of 0.135  So, next step, try generating *two* \"real\" RLEs.",
      "votes": 1,
      "replies": [
        {
          "id": 1178911,
          "postDate": "2021-01-31T07:07:06.950Z",
          "content": "<p>This is a supplement.<br>\nUse <code>np.NaN</code> if you want to input empty mask RLE.</p>",
          "rawMarkdown": "This is a supplement.\nUse `np.NaN` if you want to input empty mask RLE."
        }
      ]
    },
    {
      "id": 1174643,
      "postDate": "2021-01-28T16:13:56.873Z",
      "content": "<p>I have made five submissions for the HuBMAP competition.  The first<br>\none timed out, the next two got \"Submission Scoring Error\", and the<br>\nlast two timed out again.  In all five of the cases, I was able to run<br>\nsuccessfully using the public Test set and [Save Version] ; the runs<br>\ncompleted in an hour and a half, and I was able to visualize the<br>\nresults by reading back the RLEs and decoding them, and the results<br>\nlook reasonable, with glom markers following the structure of the<br>\ninput images.</p>\n<p>So, why the timeout?  I'm going to try listing possible explanations,<br>\nand would appreciate any comments you might have.</p>\n<p>(1)  There is a bug in my code that's causing it to loop forever.  The<br>\nfact that I can submit Test cases that complete successfully contra-<br>\ndicts this.</p>\n<p>(2)  The inputs in the Private Test set are MUCH bigger, hence my<br>\nalgorithms are too slow to process them.  No way to know this, but<br>\nothers have submitted successfully.</p>\n<p>(3)  There is some algorithm inefficiency that causes <em>somewhat</em><br>\nbigger inputs to fail due to poor scaling.  The algorithm processes<br>\neach input image independently (except for non-recovered storage),<br>\nand within each image processes 512-pixel wide columns indepently,<br>\nconcatenating together the suitably-offset RLEs for each column.</p>\n<p>(4)  There are regions in the Private Test images that cause runtime<br>\nproblems for the model prediction.  However, since the prediction is<br>\na finite, determistic process (unlike fitting) this seems implausible.</p>\n<p>(5)  The run exhausts GPU allotment.  I have no visibility to this,<br>\nbut if it happens, it would certainly cause excessive runtimes.</p>\n<p>(6)  The RLE process blows up:  Given a random binary image as<br>\ninput, RLE will produce a very large string output.  However, examining<br>\nthe glom marker outputs for the Public Test set, the output images are<br>\nsparse collections of compact glom markers, far from random images.</p>",
      "rawMarkdown": "I have made five submissions for the HuBMAP competition.  The first\none timed out, the next two got \"Submission Scoring Error\", and the\nlast two timed out again.  In all five of the cases, I was able to run\nsuccessfully using the public Test set and [Save Version] ; the runs\ncompleted in an hour and a half, and I was able to visualize the\nresults by reading back the RLEs and decoding them, and the results\nlook reasonable, with glom markers following the structure of the\ninput images.\n\nSo, why the timeout?  I'm going to try listing possible explanations,\nand would appreciate any comments you might have.\n\n(1)  There is a bug in my code that's causing it to loop forever.  The\nfact that I can submit Test cases that complete successfully contra-\ndicts this.\n\n(2)  The inputs in the Private Test set are MUCH bigger, hence my\nalgorithms are too slow to process them.  No way to know this, but\nothers have submitted successfully.\n\n(3)  There is some algorithm inefficiency that causes *somewhat*\nbigger inputs to fail due to poor scaling.  The algorithm processes\neach input image independently (except for non-recovered storage),\nand within each image processes 512-pixel wide columns indepently,\nconcatenating together the suitably-offset RLEs for each column.\n\n(4)  There are regions in the Private Test images that cause runtime\nproblems for the model prediction.  However, since the prediction is\na finite, determistic process (unlike fitting) this seems implausible.\n\n(5)  The run exhausts GPU allotment.  I have no visibility to this,\nbut if it happens, it would certainly cause excessive runtimes.\n\n(6)  The RLE process blows up:  Given a random binary image as\ninput, RLE will produce a very large string output.  However, examining\nthe glom marker outputs for the Public Test set, the output images are\nsparse collections of compact glom markers, far from random images.",
      "votes": 2
    },
    {
      "id": 1179005,
      "postDate": "2021-01-31T08:56:31.473Z",
      "content": "<p>you could just singlely submit the result file without relevant execute code.</p>",
      "rawMarkdown": "you could just singlely submit the result file without relevant execute code."
    },
    {
      "id": 1176092,
      "postDate": "2021-01-29T13:59:09.350Z",
      "content": "<p>nice<br>\nI am upvoted<br>\nsee this notebook<br>\n<a href=\"https://www.kaggle.com/ashokkumarbibbab/digit-recognizer-machine-learning-model\" target=\"_blank\">https://www.kaggle.com/ashokkumarbibbab/digit-recognizer-machine-learning-model</a></p>",
      "rawMarkdown": "\nnice\nI am upvoted\nsee this notebook\nhttps://www.kaggle.com/ashokkumarbibbab/digit-recognizer-machine-learning-model"
    },
    {
      "id": 1175818,
      "postDate": "2021-01-29T11:12:10.510Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/markalavin\" target=\"_blank\">@markalavin</a> without looking at the code it is always difficult to troubleshoot. But may'be you could use below tips to narrow down your issue.</p>\n<p>One thing you could try todo to eliminate any issues with timing… modify your code to take only the first test image from the list. Only predict on that image… your score will be low of course as the other predictions are not there. But you can validate that the actual predictions on private set are made and submitted.</p>\n<p>Another thing you could try to do to catch any errors is to add some error handling and handle according to possible issues you suspect. For example if you think certain test images cause the issues then wrap the prediction inside some error handling. If there is an issue with that specific issue that causes an exception then just move on to the next test image.</p>\n<p>Hope this helps a bit.</p>",
      "rawMarkdown": "Hi @markalavin without looking at the code it is always difficult to troubleshoot. But may'be you could use below tips to narrow down your issue.\n\nOne thing you could try todo to eliminate any issues with timing... modify your code to take only the first test image from the list. Only predict on that image... your score will be low of course as the other predictions are not there. But you can validate that the actual predictions on private set are made and submitted.\n\nAnother thing you could try to do to catch any errors is to add some error handling and handle according to possible issues you suspect. For example if you think certain test images cause the issues then wrap the prediction inside some error handling. If there is an issue with that specific issue that causes an exception then just move on to the next test image.\n\nHope this helps a bit."
    },
    {
      "id": 1175114,
      "postDate": "2021-01-28T22:54:12.443Z",
      "content": "<p>I have a bunch of timeouts on my submission history too,<br>\nAnd for me at least, it did not have anything to do strictly with a timeout per se, but poor RAM management on my part. I was naively reading whole images. Didn't have an issue with the public test set but it seems that one of the images in the private test set is bigger than the biggest in the public test set. Fixed the issue by using rasterio/tifffile to read the tiles from disk and only load N tiles at a time to reduce the memory footprint. Also the time difference between the test prediction and the private one it's around ~2.5x slower, which seems in line with the 7 extra images added on the private test set. So if your prediction_time_for_the_test_set * 2.5 is faster than 9 hours, I would bet that your submission is exhausting the RAM. </p>\n<p>I hope my painful findings are useful for your endeavors!</p>",
      "rawMarkdown": "I have a bunch of timeouts on my submission history too,\nAnd for me at least, it did not have anything to do strictly with a timeout per se, but poor RAM management on my part. I was naively reading whole images. Didn't have an issue with the public test set but it seems that one of the images in the private test set is bigger than the biggest in the public test set. Fixed the issue by using rasterio/tifffile to read the tiles from disk and only load N tiles at a time to reduce the memory footprint. Also the time difference between the test prediction and the private one it's around ~2.5x slower, which seems in line with the 7 extra images added on the private test set. So if your prediction_time_for_the_test_set * 2.5 is faster than 9 hours, I would bet that your submission is exhausting the RAM. \n\nI hope my painful findings are useful for your endeavors!"
    },
    {
      "id": 1179288,
      "postDate": "2021-01-31T13:36:47.187Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1177999,
      "author_name": "Mark A Lavin",
      "author_url": "",
      "post_date": "2021-01-30T15:41:55.240000",
      "content": "<p>OK, some progress (finally!):</p>\n<p>I modified my code to only generate a trivial RLE ( \"0 1\") for one out of the N images, no output for the other N - 1.   I got a \"Submission Scoring Error\".   I then modified the code further so that it generates the trivial RLE for <em>all</em> N images, and the run actually succeeded (!), giving me a score of 0.0</p>\n<p>Several conclusions:<br>\n(1)   The overall <code>submission.csv</code> structure I'm generating is correct<br>\n(2)   The submitted <code>submission.csv</code> must contain RLE's for <em>all</em> input images</p>\n<p>My next step is to modify the code so it outputs complete RLE for one of the inputs and the trivial output for the other N-1 inputs.   We'll see whether this manages to get past the Submission Timed Out with a non-zero score.   If so, I will then start adding more images one at a time.</p>\n<p>EDIT:   Next step completed quickly, generating submission that Succeeded with a Leader Board (LB) score of 0.135  So, next step, try generating <em>two</em> \"real\" RLEs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1178911,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2021-01-31T07:07:06.950000",
          "content": "<p>This is a supplement.<br>\nUse <code>np.NaN</code> if you want to input empty mask RLE.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1179005,
      "author_name": "朴大福",
      "author_url": "",
      "post_date": "2021-01-31T08:56:31.473000",
      "content": "<p>you could just singlely submit the result file without relevant execute code.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1176092,
      "author_name": "Ashok kumar",
      "author_url": "",
      "post_date": "2021-01-29T13:59:09.350000",
      "content": "<p>nice<br>\nI am upvoted<br>\nsee this notebook<br>\n<a href=\"https://www.kaggle.com/ashokkumarbibbab/digit-recognizer-machine-learning-model\" target=\"_blank\">https://www.kaggle.com/ashokkumarbibbab/digit-recognizer-machine-learning-model</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1175818,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2021-01-29T11:12:10.510000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/markalavin\" target=\"_blank\">@markalavin</a> without looking at the code it is always difficult to troubleshoot. But may'be you could use below tips to narrow down your issue.</p>\n<p>One thing you could try todo to eliminate any issues with timing… modify your code to take only the first test image from the list. Only predict on that image… your score will be low of course as the other predictions are not there. But you can validate that the actual predictions on private set are made and submitted.</p>\n<p>Another thing you could try to do to catch any errors is to add some error handling and handle according to possible issues you suspect. For example if you think certain test images cause the issues then wrap the prediction inside some error handling. If there is an issue with that specific issue that causes an exception then just move on to the next test image.</p>\n<p>Hope this helps a bit.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1175114,
      "author_name": "Arslaan",
      "author_url": "",
      "post_date": "2021-01-28T22:54:12.443000",
      "content": "<p>I have a bunch of timeouts on my submission history too,<br>\nAnd for me at least, it did not have anything to do strictly with a timeout per se, but poor RAM management on my part. I was naively reading whole images. Didn't have an issue with the public test set but it seems that one of the images in the private test set is bigger than the biggest in the public test set. Fixed the issue by using rasterio/tifffile to read the tiles from disk and only load N tiles at a time to reduce the memory footprint. Also the time difference between the test prediction and the private one it's around ~2.5x slower, which seems in line with the 7 extra images added on the private test set. So if your prediction_time_for_the_test_set * 2.5 is faster than 9 hours, I would bet that your submission is exhausting the RAM. </p>\n<p>I hope my painful findings are useful for your endeavors!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1179288,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-31T13:36:47.187000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1177999": "OK, some progress (finally!):\n\nI modified my code to only generate a trivial RLE ( \"0 1\") for one out of the N images, no output for the other N - 1.   I got a \"Submission Scoring Error\".   I then modified the code further so that it generates the trivial RLE for *all* N images, and the run actually succeeded (!), giving me a score of 0.0\n\nSeveral conclusions:\n(1)   The overall ```submission.csv``` structure I'm generating is correct\n(2)   The submitted ```submission.csv``` must contain RLE's for *all* input images\n\nMy next step is to modify the code so it outputs complete RLE for one of the inputs and the trivial output for the other N-1 inputs.   We'll see whether this manages to get past the Submission Timed Out with a non-zero score.   If so, I will then start adding more images one at a time.\n\nEDIT:   Next step completed quickly, generating submission that Succeeded with a Leader Board (LB) score of 0.135  So, next step, try generating *two* \"real\" RLEs.",
    "1174643": "I have made five submissions for the HuBMAP competition.  The first\none timed out, the next two got \"Submission Scoring Error\", and the\nlast two timed out again.  In all five of the cases, I was able to run\nsuccessfully using the public Test set and [Save Version] ; the runs\ncompleted in an hour and a half, and I was able to visualize the\nresults by reading back the RLEs and decoding them, and the results\nlook reasonable, with glom markers following the structure of the\ninput images.\n\nSo, why the timeout?  I'm going to try listing possible explanations,\nand would appreciate any comments you might have.\n\n(1)  There is a bug in my code that's causing it to loop forever.  The\nfact that I can submit Test cases that complete successfully contra-\ndicts this.\n\n(2)  The inputs in the Private Test set are MUCH bigger, hence my\nalgorithms are too slow to process them.  No way to know this, but\nothers have submitted successfully.\n\n(3)  There is some algorithm inefficiency that causes *somewhat*\nbigger inputs to fail due to poor scaling.  The algorithm processes\neach input image independently (except for non-recovered storage),\nand within each image processes 512-pixel wide columns indepently,\nconcatenating together the suitably-offset RLEs for each column.\n\n(4)  There are regions in the Private Test images that cause runtime\nproblems for the model prediction.  However, since the prediction is\na finite, determistic process (unlike fitting) this seems implausible.\n\n(5)  The run exhausts GPU allotment.  I have no visibility to this,\nbut if it happens, it would certainly cause excessive runtimes.\n\n(6)  The RLE process blows up:  Given a random binary image as\ninput, RLE will produce a very large string output.  However, examining\nthe glom marker outputs for the Public Test set, the output images are\nsparse collections of compact glom markers, far from random images.",
    "1179005": "you could just singlely submit the result file without relevant execute code.",
    "1176092": "\nnice\nI am upvoted\nsee this notebook\nhttps://www.kaggle.com/ashokkumarbibbab/digit-recognizer-machine-learning-model",
    "1175818": "Hi @markalavin without looking at the code it is always difficult to troubleshoot. But may'be you could use below tips to narrow down your issue.\n\nOne thing you could try todo to eliminate any issues with timing... modify your code to take only the first test image from the list. Only predict on that image... your score will be low of course as the other predictions are not there. But you can validate that the actual predictions on private set are made and submitted.\n\nAnother thing you could try to do to catch any errors is to add some error handling and handle according to possible issues you suspect. For example if you think certain test images cause the issues then wrap the prediction inside some error handling. If there is an issue with that specific issue that causes an exception then just move on to the next test image.\n\nHope this helps a bit.",
    "1175114": "I have a bunch of timeouts on my submission history too,\nAnd for me at least, it did not have anything to do strictly with a timeout per se, but poor RAM management on my part. I was naively reading whole images. Didn't have an issue with the public test set but it seems that one of the images in the private test set is bigger than the biggest in the public test set. Fixed the issue by using rasterio/tifffile to read the tiles from disk and only load N tiles at a time to reduce the memory footprint. Also the time difference between the test prediction and the private one it's around ~2.5x slower, which seems in line with the 7 extra images added on the private test set. So if your prediction_time_for_the_test_set * 2.5 is faster than 9 hours, I would bet that your submission is exhausting the RAM. \n\nI hope my painful findings are useful for your endeavors!",
    "1179288": ""
  }
}