{
  "id": 215598,
  "title": "Main challenges",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/215598",
  "author_name": "",
  "post_date": "2021-01-30T13:50:58.238546500Z",
  "votes": 27,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I've joined lately but the competition is going to be restarted with updated data, so I'm not so late.</p>\n<p>I would like to share the main challenges I consider for this competition:</p>\n<ul>\n<li><p><strong>Large data</strong>:<br>\nCurrent train dataset includes 8 large RGB images and test dataset is 5 (public) + 7 (private) images.<br>\nThese images are also large and one could notice that one (or maybe more) image in private dataset seems larger than all in train/public.<br>\nI've spent a few submissions before realizing I got an OOM on a private image.<br>\nThe new train dataset planned will include 20 images which will make all our models much longer to train. Our models should be quite better however.</p></li>\n<li><p><strong>Resources management</strong>:<br>\nDid you get any <code>\"Notebook Exceeded Allowed Compute\"</code>, <code>\"Notebook Timeout\"</code> or <code>\"Submission Scoring Error\"</code>? I'm sure yes and you got frustrated to debug them.<br>\nI got <code>\"Notebook Exceeded Allowed Compute\"</code> when I exceeded the disk size limit. But I got <code>\"Notebook Timeout\"</code> or <code>\"Submission Scoring Error\"</code> mainly on OOM.<br>\nAssuming your code is designed to generate correct submission (correct RLE format, no missing id), you can still get <code>\"Submission Scoring Error\"</code> if the submission file<br>\nis not generated due to GPU or RAM crash or pressure. Generating it correctly for public data is required of course but not enough as you're blind on private data.</p>\n<p>Some numbers that could help you to see how tight are our Kaggle resources:</p>\n<p>float16 numpy array (36800, 43780, 1) is around is 3.0GB - Image <code>afa5e8098</code> HeightxWidth is 36800x43780<br>\nfloat16 numpy array (45000, 45000, 1) is around 3.7GB<br>\nfloat16 numpy array (50000, 50000, 1) is around 4.6GB<br>\nfloat16 numpy array (60000, 60000, 1) is around 6.0GB<br>\nfloat32 numpy array (60000, 60000, 1) is around 12.0GB<br>\nuint8   numpy array (60000, 60000, 1) is around 3.4GB</p>\n<p>So how many arrays can you maintain in memory without crashing? How many arrays do you need for final prediction?<br>\nKernel with GPU comes with <strong>13GB RAM</strong> and <strong>16GB GPU</strong>, you have to deal with that. <br>\nDon't try to load full TIFF image in memory, it will crash for some. Use <code>rasterio</code> or similar library to load slices from disk to feed your model.<br>\nKeep your predictions memory friendly, <code>float16</code> if you need to maintain probabilities, <code>boolean</code> for mask.<br>\nPersist data on disk if needed and if it's acceptable from runtime limit.</p></li>\n<li><p><strong>Data quality</strong>:<br>\nMany threads around report issues on masks. That's the main root cause for the gaps between CV and LB. If you apply the global shift from <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>  you will get around +0.02 (crazy!) on LB, so the ground truth is noisy. Fortunately, it should be fixed by the organizers with the updated data.<br>\nIt does not mean that the labels will be perfect but it should be more consistent.  </p>\n<p>Training some good models is also a challenge 😁</p></li>\n</ul>\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": "1177782",
      "postDate": "01/30/2021 13:50:58",
      "content": "<p>Hi,</p>\n<p>I've joined lately but the competition is going to be restarted with updated data, so I'm not so late.</p>\n<p>I would like to share the main challenges I consider for this competition:</p>\n<ul>\n<li><p><strong>Large data</strong>:<br>\nCurrent train dataset includes 8 large RGB images and test dataset is 5 (public) + 7 (private) images.<br>\nThese images are also large and one could notice that one (or maybe more) image in private dataset seems larger than all in train/public.<br>\nI've spent a few submissions before realizing I got an OOM on a private image.<br>\nThe new train dataset planned will include 20 images which will make all our models much longer to train. Our models should be quite better however.</p></li>\n<li><p><strong>Resources management</strong>:<br>\nDid you get any <code>\"Notebook Exceeded Allowed Compute\"</code>, <code>\"Notebook Timeout\"</code> or <code>\"Submission Scoring Error\"</code>? I'm sure yes and you got frustrated to debug them.<br>\nI got <code>\"Notebook Exceeded Allowed Compute\"</code> when I exceeded the disk size limit. But I got <code>\"Notebook Timeout\"</code> or <code>\"Submission Scoring Error\"</code> mainly on OOM.<br>\nAssuming your code is designed to generate correct submission (correct RLE format, no missing id), you can still get <code>\"Submission Scoring Error\"</code> if the submission file<br>\nis not generated due to GPU or RAM crash or pressure. Generating it correctly for public data is required of course but not enough as you're blind on private data.</p>\n<p>Some numbers that could help you to see how tight are our Kaggle resources:</p>\n<p>float16 numpy array (36800, 43780, 1) is around is 3.0GB - Image <code>afa5e8098</code> HeightxWidth is 36800x43780<br>\nfloat16 numpy array (45000, 45000, 1) is around 3.7GB<br>\nfloat16 numpy array (50000, 50000, 1) is around 4.6GB<br>\nfloat16 numpy array (60000, 60000, 1) is around 6.0GB<br>\nfloat32 numpy array (60000, 60000, 1) is around 12.0GB<br>\nuint8   numpy array (60000, 60000, 1) is around 3.4GB</p>\n<p>So how many arrays can you maintain in memory without crashing? How many arrays do you need for final prediction?<br>\nKernel with GPU comes with <strong>13GB RAM</strong> and <strong>16GB GPU</strong>, you have to deal with that. <br>\nDon't try to load full TIFF image in memory, it will crash for some. Use <code>rasterio</code> or similar library to load slices from disk to feed your model.<br>\nKeep your predictions memory friendly, <code>float16</code> if you need to maintain probabilities, <code>boolean</code> for mask.<br>\nPersist data on disk if needed and if it's acceptable from runtime limit.</p></li>\n<li><p><strong>Data quality</strong>:<br>\nMany threads around report issues on masks. That's the main root cause for the gaps between CV and LB. If you apply the global shift from <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>  you will get around +0.02 (crazy!) on LB, so the ground truth is noisy. Fortunately, it should be fixed by the organizers with the updated data.<br>\nIt does not mean that the labels will be perfect but it should be more consistent.  </p>\n<p>Training some good models is also a challenge 😁</p></li>\n</ul>\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "Hi,\n\nI've joined lately but the competition is going to be restarted with updated data, so I'm not so late.\n\nI would like to share the main challenges I consider for this competition:\n\n- **Large data**:\n  Current train dataset includes 8 large RGB images and test dataset is 5 (public) + 7 (private) images.\n  These images are also large and one could notice that one (or maybe more) image in private dataset seems larger than all in train/public.\n  I've spent a few submissions before realizing I got an OOM on a private image.\n  The new train dataset planned will include 20 images which will make all our models much longer to train. Our models should be quite better however.\n  \n- **Resources management**:\n  Did you get any `\"Notebook Exceeded Allowed Compute\"`, `\"Notebook Timeout\"` or `\"Submission Scoring Error\"`? I'm sure yes and you got frustrated to debug them.\n  I got `\"Notebook Exceeded Allowed Compute\"` when I exceeded the disk size limit. But I got `\"Notebook Timeout\"` or `\"Submission Scoring Error\"` mainly on OOM.\n  Assuming your code is designed to generate correct submission (correct RLE format, no missing id), you can still get `\"Submission Scoring Error\"` if the submission file\n  is not generated due to GPU or RAM crash or pressure. Generating it correctly for public data is required of course but not enough as you're blind on private data.\n  \n  Some numbers that could help you to see how tight are our Kaggle resources:\n  \n  float16 numpy array (36800, 43780, 1) is around is 3.0GB - Image `afa5e8098` HeightxWidth is 36800x43780\n  float16 numpy array (45000, 45000, 1) is around 3.7GB\n  float16 numpy array (50000, 50000, 1) is around 4.6GB\n  float16 numpy array (60000, 60000, 1) is around 6.0GB\n  float32 numpy array (60000, 60000, 1) is around 12.0GB\n  uint8   numpy array (60000, 60000, 1) is around 3.4GB\n  \n  So how many arrays can you maintain in memory without crashing? How many arrays do you need for final prediction?\n  Kernel with GPU comes with **13GB RAM** and **16GB GPU**, you have to deal with that. \n  Don't try to load full TIFF image in memory, it will crash for some. Use `rasterio` or similar library to load slices from disk to feed your model.\n  Keep your predictions memory friendly, `float16` if you need to maintain probabilities, `boolean` for mask.\n  Persist data on disk if needed and if it's acceptable from runtime limit.\n\n- **Data quality**:\n  Many threads around report issues on masks. That's the main root cause for the gaps between CV and LB. If you apply the global shift from @tivfrvqhs5  you will get around +0.02 (crazy!) on LB, so the ground truth is noisy. Fortunately, it should be fixed by the organizers with the updated data.\n  It does not mean that the labels will be perfect but it should be more consistent.  \n  \n  Training some good models is also a challenge 😁\n\nHappy Kaggling!",
      "votes": null
    },
    {
      "id": "1177902",
      "postDate": "01/30/2021 15:15:20",
      "content": "<p>Thank you for sharing.<br>\nI feel the same challenge with this competition.</p>\n<ul>\n<li>Resource Management<ul>\n<li>Errors in kaggle notebook are different from real ones<ul>\n<li>Time out is shown even when RAM is over</li></ul></li></ul></li>\n<li>Large size(height, width) image</li>\n<li>Quality of annotations (labels)<ul>\n<li>Annotation rules are ambiguous</li>\n<li>Boundaries of each object are ambiguous</li>\n<li>Shift problem</li></ul></li>\n</ul>",
      "rawMarkdown": "Thank you for sharing.\nI feel the same challenge with this competition.\n\n- Resource Management\n    - Errors in kaggle notebook are different from real ones\n        - Time out is shown even when RAM is over\n- Large size(height, width) image\n- Quality of annotations (labels)\n    - Annotation rules are ambiguous\n    - Boundaries of each object are ambiguous\n    - Shift problem",
      "votes": null
    },
    {
      "id": "1177956",
      "postDate": "01/30/2021 15:31:43",
      "content": "<blockquote>\n  <p>Errors in kaggle notebook are different from real ones</p>\n</blockquote>\n<p>Yes and it's not specific this competition. I don't know how they catch errors but sometimes it could be hidden to them indirectly when we use try/except clause. </p>",
      "rawMarkdown": "> Errors in kaggle notebook are different from real ones\n\nYes and it's not specific this competition. I don't know how they catch errors but sometimes it could be hidden to them indirectly when we use try/except clause.",
      "votes": null
    },
    {
      "id": "1179716",
      "postDate": "01/31/2021 18:50:30",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    },
    {
      "id": "1211184",
      "postDate": "02/20/2021 03:24:37",
      "content": "<p>Thanks for sharing.!</p>",
      "rawMarkdown": "Thanks for sharing.!",
      "votes": null
    },
    {
      "id": "1255138",
      "postDate": "03/28/2021 13:33:25",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> if we were to do some prep processing  can we use rasterio . even though i use gc collect , del tiff image after its use is over then also kernel not releasing all memory , so eventually upon subsequent tiff loads memory gets full. </p>\n<p>Any suggestion would be helpful. I load my tiff using pytorch ds  ,even i tried Yield</p>",
      "rawMarkdown": "mpware if we were to do some prep processing  can we use rasterio . even though i use gc collect , del tiff image after its use is over then also kernel not releasing all memory , so eventually upon subsequent tiff loads memory gets full. \n\nAny suggestion would be helpful. I load my tiff using pytorch ds  ,even i tried Yield",
      "votes": null
    },
    {
      "id": "1255161",
      "postDate": "03/28/2021 14:10:12",
      "content": "<p>Preprocessing should be done on small tiles loaded through RasterIO. If you want to do preprocessing on full image in one shot, then it means that you will need to load it in memory which may not work on Kaggle due to 13GB limit (for GPU kernel) and large images for this competition.</p>",
      "rawMarkdown": "Preprocessing should be done on small tiles loaded through RasterIO. If you want to do preprocessing on full image in one shot, then it means that you will need to load it in memory which may not work on Kaggle due to 13GB limit (for GPU kernel) and large images for this competition.",
      "votes": null
    },
    {
      "id": "1255367",
      "postDate": "03/28/2021 18:12:07",
      "content": "<p>hmm.. yes i m beginning to understand it, that is bit set back for my method  of preprocessing that requires loading of image which otherwise is accurately pointing to tissues, however with some better memory management process i can successfully complete 9 tiffs , 10th one which is biggest of all is trouble some. working on that one,if dsnt works then fall back option is read region only.</p>",
      "rawMarkdown": "hmm.. yes i m beginning to understand it, that is bit set back for my method  of preprocessing that requires loading of image which otherwise is accurately pointing to tissues, however with some better memory management process i can successfully complete 9 tiffs , 10th one which is biggest of all is trouble some. working on that one,if dsnt works then fall back option is read region only.",
      "votes": null
    },
    {
      "id": "1255974",
      "postDate": "03/29/2021 12:37:34",
      "content": "<p>Update:<br>\nFinally able to complete the preprocessing kernel  (full res tiff images) that required full load of Tiff images. <br>\nWith some tricks :) </p>",
      "rawMarkdown": "Update:\nFinally able to complete the preprocessing kernel  (full res tiff images) that required full load of Tiff images. \nWith some tricks :)",
      "votes": null
    },
    {
      "id": "1283601",
      "postDate": "04/25/2021 04:47:43",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1286780",
      "postDate": "04/28/2021 12:00:08",
      "content": "<p>Thank you, Submission error reason is 13GB RAM. It is really helpful.</p>",
      "rawMarkdown": "Thank you, Submission error reason is 13GB RAM. It is really helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1177902,
      "author_name": "yukkyo",
      "author_url": "",
      "post_date": "01/30/2021 15:15:20",
      "content": "<p>Thank you for sharing.<br>\nI feel the same challenge with this competition.</p>\n<ul>\n<li>Resource Management<ul>\n<li>Errors in kaggle notebook are different from real ones<ul>\n<li>Time out is shown even when RAM is over</li></ul></li></ul></li>\n<li>Large size(height, width) image</li>\n<li>Quality of annotations (labels)<ul>\n<li>Annotation rules are ambiguous</li>\n<li>Boundaries of each object are ambiguous</li>\n<li>Shift problem</li></ul></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1177956,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "01/30/2021 15:31:43",
          "content": "<blockquote>\n  <p>Errors in kaggle notebook are different from real ones</p>\n</blockquote>\n<p>Yes and it's not specific this competition. I don't know how they catch errors but sometimes it could be hidden to them indirectly when we use try/except clause. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255138,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "03/28/2021 13:33:25",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> if we were to do some prep processing  can we use rasterio . even though i use gc collect , del tiff image after its use is over then also kernel not releasing all memory , so eventually upon subsequent tiff loads memory gets full. </p>\n<p>Any suggestion would be helpful. I load my tiff using pytorch ds  ,even i tried Yield</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255161,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/28/2021 14:10:12",
          "content": "<p>Preprocessing should be done on small tiles loaded through RasterIO. If you want to do preprocessing on full image in one shot, then it means that you will need to load it in memory which may not work on Kaggle due to 13GB limit (for GPU kernel) and large images for this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255367,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "03/28/2021 18:12:07",
          "content": "<p>hmm.. yes i m beginning to understand it, that is bit set back for my method  of preprocessing that requires loading of image which otherwise is accurately pointing to tissues, however with some better memory management process i can successfully complete 9 tiffs , 10th one which is biggest of all is trouble some. working on that one,if dsnt works then fall back option is read region only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255974,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "03/29/2021 12:37:34",
          "content": "<p>Update:<br>\nFinally able to complete the preprocessing kernel  (full res tiff images) that required full load of Tiff images. <br>\nWith some tricks :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1179716,
      "author_name": "abhiest",
      "author_url": "",
      "post_date": "01/31/2021 18:50:30",
      "content": "<p>Thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211184,
      "author_name": "creeezy",
      "author_url": "",
      "post_date": "02/20/2021 03:24:37",
      "content": "<p>Thanks for sharing.!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1283601,
      "author_name": "",
      "author_url": "",
      "post_date": "04/25/2021 04:47:43",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1286780,
      "author_name": "shigengtian",
      "author_url": "",
      "post_date": "04/28/2021 12:00:08",
      "content": "<p>Thank you, Submission error reason is 13GB RAM. It is really helpful.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1177782": "Hi,\n\nI've joined lately but the competition is going to be restarted with updated data, so I'm not so late.\n\nI would like to share the main challenges I consider for this competition:\n\n- **Large data**:\n  Current train dataset includes 8 large RGB images and test dataset is 5 (public) + 7 (private) images.\n  These images are also large and one could notice that one (or maybe more) image in private dataset seems larger than all in train/public.\n  I've spent a few submissions before realizing I got an OOM on a private image.\n  The new train dataset planned will include 20 images which will make all our models much longer to train. Our models should be quite better however.\n  \n- **Resources management**:\n  Did you get any `\"Notebook Exceeded Allowed Compute\"`, `\"Notebook Timeout\"` or `\"Submission Scoring Error\"`? I'm sure yes and you got frustrated to debug them.\n  I got `\"Notebook Exceeded Allowed Compute\"` when I exceeded the disk size limit. But I got `\"Notebook Timeout\"` or `\"Submission Scoring Error\"` mainly on OOM.\n  Assuming your code is designed to generate correct submission (correct RLE format, no missing id), you can still get `\"Submission Scoring Error\"` if the submission file\n  is not generated due to GPU or RAM crash or pressure. Generating it correctly for public data is required of course but not enough as you're blind on private data.\n  \n  Some numbers that could help you to see how tight are our Kaggle resources:\n  \n  float16 numpy array (36800, 43780, 1) is around is 3.0GB - Image `afa5e8098` HeightxWidth is 36800x43780\n  float16 numpy array (45000, 45000, 1) is around 3.7GB\n  float16 numpy array (50000, 50000, 1) is around 4.6GB\n  float16 numpy array (60000, 60000, 1) is around 6.0GB\n  float32 numpy array (60000, 60000, 1) is around 12.0GB\n  uint8   numpy array (60000, 60000, 1) is around 3.4GB\n  \n  So how many arrays can you maintain in memory without crashing? How many arrays do you need for final prediction?\n  Kernel with GPU comes with **13GB RAM** and **16GB GPU**, you have to deal with that. \n  Don't try to load full TIFF image in memory, it will crash for some. Use `rasterio` or similar library to load slices from disk to feed your model.\n  Keep your predictions memory friendly, `float16` if you need to maintain probabilities, `boolean` for mask.\n  Persist data on disk if needed and if it's acceptable from runtime limit.\n\n- **Data quality**:\n  Many threads around report issues on masks. That's the main root cause for the gaps between CV and LB. If you apply the global shift from @tivfrvqhs5  you will get around +0.02 (crazy!) on LB, so the ground truth is noisy. Fortunately, it should be fixed by the organizers with the updated data.\n  It does not mean that the labels will be perfect but it should be more consistent.  \n  \n  Training some good models is also a challenge 😁\n\nHappy Kaggling!",
    "1177902": "Thank you for sharing.\nI feel the same challenge with this competition.\n\n- Resource Management\n    - Errors in kaggle notebook are different from real ones\n        - Time out is shown even when RAM is over\n- Large size(height, width) image\n- Quality of annotations (labels)\n    - Annotation rules are ambiguous\n    - Boundaries of each object are ambiguous\n    - Shift problem",
    "1177956": "> Errors in kaggle notebook are different from real ones\n\nYes and it's not specific this competition. I don't know how they catch errors but sometimes it could be hidden to them indirectly when we use try/except clause.",
    "1179716": "Thank you for sharing",
    "1211184": "Thanks for sharing.!",
    "1255138": "mpware if we were to do some prep processing  can we use rasterio . even though i use gc collect , del tiff image after its use is over then also kernel not releasing all memory , so eventually upon subsequent tiff loads memory gets full. \n\nAny suggestion would be helpful. I load my tiff using pytorch ds  ,even i tried Yield",
    "1255161": "Preprocessing should be done on small tiles loaded through RasterIO. If you want to do preprocessing on full image in one shot, then it means that you will need to load it in memory which may not work on Kaggle due to 13GB limit (for GPU kernel) and large images for this competition.",
    "1255367": "hmm.. yes i m beginning to understand it, that is bit set back for my method  of preprocessing that requires loading of image which otherwise is accurately pointing to tissues, however with some better memory management process i can successfully complete 9 tiffs , 10th one which is biggest of all is trouble some. working on that one,if dsnt works then fall back option is read region only.",
    "1255974": "Update:\nFinally able to complete the preprocessing kernel  (full res tiff images) that required full load of Tiff images. \nWith some tricks :)",
    "1283601": "Thanks for sharing",
    "1286780": "Thank you, Submission error reason is 13GB RAM. It is really helpful."
  },
  "source": "meta"
}