{
  "id": 210416,
  "title": "Some insights",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/210416",
  "author_name": "",
  "post_date": "2021-01-10T17:11:27.267572900Z",
  "votes": 37,
  "comment_count": 14,
  "views": 0,
  "content": "<p>It has been some time since I made one of these posts, so here is one to kick-start the new year. Let's go!</p>\n<h2>General</h2>\n<ul>\n<li>HuBMAP is short for Human BioMolecular Atlas Program.</li>\n<li>Sponsored by the NIH (National Institues of Health)</li>\n<li>Detect functional tissue units (FTUs in short) across different tissue preparation pipelines.</li>\n<li>An FTU is a 3D block of cells</li>\n</ul>\n<h2>Metric</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/S%C3%B8rensen%E2%80%93Dice_coefficient\" target=\"_blank\">Dice coefficient</a> between predicted and ground truth mask pixels.</p>\n<h2>Dataset</h2>\n<ul>\n<li>Metadata: HuBMAP-20-dataset_information.csv</li>\n<li>Various slices at different sizes of the train dataset can be found here (thanks to <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a> for providing these): <ul>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-512x512\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-512x512</a></li>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-256x256\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-256x256</a></li>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-1024x1024\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-1024x1024</a></li></ul></li>\n<li>The provided data consists of <code>11</code> fresh frozen (FF) and <code>9</code> Formalin Fixed Paraffin Embedded (FFPE) PAS kidney images. <br>\nHere is how the data is divided between train and test (this is for the <strong>previous</strong> dataset)</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Train</th>\n<th>Public test</th>\n<th>Private test</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>8 TIFF images + masks</td>\n<td>5 TIFF images + masks</td>\n<td>7 TIFF images + masks found on data portal</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Notice that this is the previous dataset and new data will be collected for the private test part.</li>\n<li>This was due to the discovery of the previous private dataset on the NIH HuBMAP data portal. Another reason is the bad<br>\nquality of the provided masks and the organizers are taking additional steps to insure better quality.</li>\n<li>The new private dataset will be comprised of <code>5</code> fresh frozen and <code>5</code> FFPE PAS images. This will be available around mid-January. So far, here is the grouping break of the updated dataset:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Train</th>\n<th>Public test</th>\n<th>Private test</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>20 TIFF images + masks</td>\n<td>Unknown</td>\n<td>10 TIFF images (5 FF + 5 FFPE)</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>For more details, read this <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884\" target=\"_blank\">discussion</a>. </li>\n<li>Here is what one of the original TIFF images looks like: </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Fdd80ff596b1cde86d22d76545e2c77e9%2Ftiff_sample.png?generation=1610298681198226&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Some metadata about the <strong>old</strong> dataset</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Ffd60353c1cc8d35b28ec1c418af20eea%2Fmetadata_old_hubmap.png?generation=1610479642195706&amp;alt=media\" alt=\"\"></p>\n<h2>Models</h2>\n<ul>\n<li><p>You can start with a ready to use model (thanks to <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a>). To train: </p></li>\n<li><p>clone the repo: <a href=\"https://github.com/ternaus/cloths_segmentation/tree/main/cloths_segmentation\" target=\"_blank\">https://github.com/ternaus/cloths_segmentation/tree/main/cloths_segmentation</a></p></li>\n<li><p>set env variables: <code>IMAGE_PATH</code> and <code>MASK_PATH</code></p></li>\n<li><p>run: <code>python -m cloths_segmentation.train -c config.yml</code>. You only need to provide a config.yml. </p></li>\n</ul>\n<p>Here is an example: </p>\n<pre><code>seed: 1984\n\nnum_workers: 4\nexperiment_name: \"best_experiment\"\n\nval_split: 0.1\n\nmodel:\n  type: segmentation_models_pytorch.Unet\n  encoder_name: timm-efficientnet-b3\n  classes: 1\n  encoder_weights: noisy-student\n\ntrainer:\n  type: pytorch_lightning.Trainer\n  gpus: 1\n  max_epochs: 70\n  distributed_backend: ddp\n  progress_bar_refresh_rate: 1\n  benchmark: True\n  precision: 16\n  gradient_clip_val: 5.0\n  num_sanity_val_steps: 2\n  sync_batchnorm: True\n\n\n\nscheduler:\n  type: torch.optim.lr_scheduler.CosineAnnealingWarmRestarts\n  T_0: 10\n  T_mult: 2\n\ntrain_parameters:\n  batch_size: 8\n\ncheckpoint_callback:\n  type: pytorch_lightning.callbacks.ModelCheckpoint\n  filepath: \"best_experiment\"\n  monitor: val_iou\n  verbose: True\n  mode: max\n  save_top_k: -1\n\nval_parameters:\n  batch_size: 2\n\noptimizer:\n  type: adamp.AdamP\n  lr: 0.0001\n\n\ntrain_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.PadIfNeeded\n        always_apply: False\n        min_height: 800\n        min_width: 800\n        border_mode: 0 # cv2.BORDER_CONSTANT\n        value: 0\n        mask_value: 0\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.RandomCrop\n        always_apply: False\n        height: 512\n        width: 512\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.HorizontalFlip\n        always_apply: False\n        p: 0.5\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n\nval_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.PadIfNeeded\n        always_apply: False\n        min_height: 800\n        min_width: 800\n        border_mode: 0 # cv2.BORDER_CONSTANT\n        value: 0\n        mask_value: 0\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n\ntest_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n</code></pre>\n<h2>Some medical terms</h2>\n<ul>\n<li><p>Glomerulus: <a href=\"https://en.wikipedia.org/wiki/Glomerulus_(kidney\" target=\"_blank\">https://en.wikipedia.org/wiki/Glomerulus_(kidney</a>)</p></li>\n<li><p>PAS (short for periodic acid-Schiff): <a href=\"https://en.wikipedia.org/wiki/Periodic_acid%E2%80%93Schiff_stain\" target=\"_blank\">https://en.wikipedia.org/wiki/Periodic_acid%E2%80%93Schiff_stain</a></p></li>\n<li><p>FFPE (short for Formalin Fixed Paraffin Embedded): <a href=\"https://www.biochain.com/general/what-is-ffpe-tissue/\" target=\"_blank\">https://www.biochain.com/general/what-is-ffpe-tissue/</a></p></li>\n<li><p>From the HuBMAP assays <a href=\"https://portal.hubmapconsortium.org/docs/assays\" target=\"_blank\">page</a>:</p></li>\n<li><p>Stained Microscopy:</p>\n<blockquote>\n  <p>Stained microscopy employs histological stains such as H&amp;E or PAS to improve resolution and contrast for     visualization  of anatomical structures such as tubules or glomeruli. </p>\n</blockquote></li>\n</ul>\n<h2>Previous similar challenges</h2>\n<p>Here is a list of previous similar challenges: i.e. either binary segmentation, medical imaging, or both.</p>\n<ul>\n<li>Human Protein Atlas Image Classification: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\" target=\"_blank\">https://www.kaggle.com/c/human-protein-atlas-image-classification</a> </li>\n<li>TGS Salt Identification Challenge: <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge\" target=\"_blank\">https://www.kaggle.com/c/tgs-salt-identification-challenge</a></li>\n<li>Ultrasound Nerve Segmentation: <a href=\"https://www.kaggle.com/c/ultrasound-nerve-segmentation\" target=\"_blank\">https://www.kaggle.com/c/ultrasound-nerve-segmentation</a> </li>\n</ul>\n<p>I will add more challenges in the list above with more details. In the meantime, use the great <a href=\"https://kagoole.herokuapp.com/\" target=\"_blank\">kagoole</a> <br>\nwebsite. Here is the result using \"segmentation\" as a search term:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F382bfbfbff76f9d42e8f7c11ca7855de%2Fsegmentation_list.png?generation=1610346191759100&amp;alt=media\" alt=\"\"></p>\n<h2>Some useful code snippets</h2>\n<ul>\n<li><p>Reading a TIFF image: </p>\n<pre><code>from osgeo import gdal\nimport numpy as np\nimport matplotlib.image as mpimg\npath_to_tiff = \"/hdd/hubmap-kidney-segmentation/train/2f6ecfcdf.tiff\"\ntiff = gdal.Open(path_to_tiff)\nimage = np.array(tiff.ReadAsArray())\nimage = image.transpose((1,2,0))\nfig, ax = plt.subplots(1, 1, figsize=(20, 20))\nax.imshow(image)\n</code></pre></li>\n</ul>\n<h2>References</h2>\n<ul>\n<li>Dataset details: <a href=\"https://www.kaggle.com/leahscherschel/dataset-details\" target=\"_blank\">https://www.kaggle.com/leahscherschel/dataset-details</a> </li>\n<li>Run-length explanation: <a href=\"https://www.kaggle.com/leahscherschel/run-length-encoding\" target=\"_blank\">https://www.kaggle.com/leahscherschel/run-length-encoding</a></li>\n<li>Segmentation library for Pytorch: <a href=\"https://github.com/qubvel/segmentation_models.pytorch\" target=\"_blank\">https://github.com/qubvel/segmentation_models.pytorch</a> </li>\n<li>List of tips from previous segmentation Kaggle competitions: <a href=\"https://neptune.ai/blog/image-segmentation-tips-and-tricks-from-kaggle-competitions\" target=\"_blank\">https://neptune.ai/blog/image-segmentation-tips-and-tricks-from-kaggle-competitions</a></li>\n<li>A good streamlit app for clothes segmentation: <a href=\"https://github.com/ternaus/cloths_segmentation_demo\" target=\"_blank\">https://github.com/ternaus/cloths_segmentation_demo</a></li>\n<li>TIFF: <a href=\"https://en.wikipedia.org/wiki/TIFF\" target=\"_blank\">https://en.wikipedia.org/wiki/TIFF</a></li>\n<li>HuBMAP website: <a href=\"https://hubmapconsortium.org/\" target=\"_blank\">https://hubmapconsortium.org/</a></li>\n<li>HuBMAP data portal: <a href=\"https://portal.hubmapconsortium.org/\" target=\"_blank\">https://portal.hubmapconsortium.org/</a></li>\n<li>A good discussion about wrong/missing annotations and how to correct them: <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201569\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201569</a></li>\n<li>Ground truth shift: <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816#1116263\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816#1116263</a> </li>\n<li>Annotation issues: <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664</a> </li>\n</ul>",
  "messages": [
    {
      "id": "1147751",
      "postDate": "01/10/2021 17:11:27",
      "content": "<p>It has been some time since I made one of these posts, so here is one to kick-start the new year. Let's go!</p>\n<h2>General</h2>\n<ul>\n<li>HuBMAP is short for Human BioMolecular Atlas Program.</li>\n<li>Sponsored by the NIH (National Institues of Health)</li>\n<li>Detect functional tissue units (FTUs in short) across different tissue preparation pipelines.</li>\n<li>An FTU is a 3D block of cells</li>\n</ul>\n<h2>Metric</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/S%C3%B8rensen%E2%80%93Dice_coefficient\" target=\"_blank\">Dice coefficient</a> between predicted and ground truth mask pixels.</p>\n<h2>Dataset</h2>\n<ul>\n<li>Metadata: HuBMAP-20-dataset_information.csv</li>\n<li>Various slices at different sizes of the train dataset can be found here (thanks to <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a> for providing these): <ul>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-512x512\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-512x512</a></li>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-256x256\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-256x256</a></li>\n<li><a href=\"https://www.kaggle.com/iafoss/hubmap-1024x1024\" target=\"_blank\">https://www.kaggle.com/iafoss/hubmap-1024x1024</a></li></ul></li>\n<li>The provided data consists of <code>11</code> fresh frozen (FF) and <code>9</code> Formalin Fixed Paraffin Embedded (FFPE) PAS kidney images. <br>\nHere is how the data is divided between train and test (this is for the <strong>previous</strong> dataset)</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Train</th>\n<th>Public test</th>\n<th>Private test</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>8 TIFF images + masks</td>\n<td>5 TIFF images + masks</td>\n<td>7 TIFF images + masks found on data portal</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>Notice that this is the previous dataset and new data will be collected for the private test part.</li>\n<li>This was due to the discovery of the previous private dataset on the NIH HuBMAP data portal. Another reason is the bad<br>\nquality of the provided masks and the organizers are taking additional steps to insure better quality.</li>\n<li>The new private dataset will be comprised of <code>5</code> fresh frozen and <code>5</code> FFPE PAS images. This will be available around mid-January. So far, here is the grouping break of the updated dataset:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Train</th>\n<th>Public test</th>\n<th>Private test</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>20 TIFF images + masks</td>\n<td>Unknown</td>\n<td>10 TIFF images (5 FF + 5 FFPE)</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>For more details, read this <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884\" target=\"_blank\">discussion</a>. </li>\n<li>Here is what one of the original TIFF images looks like: </li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Fdd80ff596b1cde86d22d76545e2c77e9%2Ftiff_sample.png?generation=1610298681198226&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Some metadata about the <strong>old</strong> dataset</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Ffd60353c1cc8d35b28ec1c418af20eea%2Fmetadata_old_hubmap.png?generation=1610479642195706&amp;alt=media\" alt=\"\"></p>\n<h2>Models</h2>\n<ul>\n<li><p>You can start with a ready to use model (thanks to <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a>). To train: </p></li>\n<li><p>clone the repo: <a href=\"https://github.com/ternaus/cloths_segmentation/tree/main/cloths_segmentation\" target=\"_blank\">https://github.com/ternaus/cloths_segmentation/tree/main/cloths_segmentation</a></p></li>\n<li><p>set env variables: <code>IMAGE_PATH</code> and <code>MASK_PATH</code></p></li>\n<li><p>run: <code>python -m cloths_segmentation.train -c config.yml</code>. You only need to provide a config.yml. </p></li>\n</ul>\n<p>Here is an example: </p>\n<pre><code>seed: 1984\n\nnum_workers: 4\nexperiment_name: \"best_experiment\"\n\nval_split: 0.1\n\nmodel:\n  type: segmentation_models_pytorch.Unet\n  encoder_name: timm-efficientnet-b3\n  classes: 1\n  encoder_weights: noisy-student\n\ntrainer:\n  type: pytorch_lightning.Trainer\n  gpus: 1\n  max_epochs: 70\n  distributed_backend: ddp\n  progress_bar_refresh_rate: 1\n  benchmark: True\n  precision: 16\n  gradient_clip_val: 5.0\n  num_sanity_val_steps: 2\n  sync_batchnorm: True\n\n\n\nscheduler:\n  type: torch.optim.lr_scheduler.CosineAnnealingWarmRestarts\n  T_0: 10\n  T_mult: 2\n\ntrain_parameters:\n  batch_size: 8\n\ncheckpoint_callback:\n  type: pytorch_lightning.callbacks.ModelCheckpoint\n  filepath: \"best_experiment\"\n  monitor: val_iou\n  verbose: True\n  mode: max\n  save_top_k: -1\n\nval_parameters:\n  batch_size: 2\n\noptimizer:\n  type: adamp.AdamP\n  lr: 0.0001\n\n\ntrain_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.PadIfNeeded\n        always_apply: False\n        min_height: 800\n        min_width: 800\n        border_mode: 0 # cv2.BORDER_CONSTANT\n        value: 0\n        mask_value: 0\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.RandomCrop\n        always_apply: False\n        height: 512\n        width: 512\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.HorizontalFlip\n        always_apply: False\n        p: 0.5\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n\nval_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.PadIfNeeded\n        always_apply: False\n        min_height: 800\n        min_width: 800\n        border_mode: 0 # cv2.BORDER_CONSTANT\n        value: 0\n        mask_value: 0\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n\ntest_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n</code></pre>\n<h2>Some medical terms</h2>\n<ul>\n<li><p>Glomerulus: <a href=\"https://en.wikipedia.org/wiki/Glomerulus_(kidney\" target=\"_blank\">https://en.wikipedia.org/wiki/Glomerulus_(kidney</a>)</p></li>\n<li><p>PAS (short for periodic acid-Schiff): <a href=\"https://en.wikipedia.org/wiki/Periodic_acid%E2%80%93Schiff_stain\" target=\"_blank\">https://en.wikipedia.org/wiki/Periodic_acid%E2%80%93Schiff_stain</a></p></li>\n<li><p>FFPE (short for Formalin Fixed Paraffin Embedded): <a href=\"https://www.biochain.com/general/what-is-ffpe-tissue/\" target=\"_blank\">https://www.biochain.com/general/what-is-ffpe-tissue/</a></p></li>\n<li><p>From the HuBMAP assays <a href=\"https://portal.hubmapconsortium.org/docs/assays\" target=\"_blank\">page</a>:</p></li>\n<li><p>Stained Microscopy:</p>\n<blockquote>\n  <p>Stained microscopy employs histological stains such as H&amp;E or PAS to improve resolution and contrast for     visualization  of anatomical structures such as tubules or glomeruli. </p>\n</blockquote></li>\n</ul>\n<h2>Previous similar challenges</h2>\n<p>Here is a list of previous similar challenges: i.e. either binary segmentation, medical imaging, or both.</p>\n<ul>\n<li>Human Protein Atlas Image Classification: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\" target=\"_blank\">https://www.kaggle.com/c/human-protein-atlas-image-classification</a> </li>\n<li>TGS Salt Identification Challenge: <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge\" target=\"_blank\">https://www.kaggle.com/c/tgs-salt-identification-challenge</a></li>\n<li>Ultrasound Nerve Segmentation: <a href=\"https://www.kaggle.com/c/ultrasound-nerve-segmentation\" target=\"_blank\">https://www.kaggle.com/c/ultrasound-nerve-segmentation</a> </li>\n</ul>\n<p>I will add more challenges in the list above with more details. In the meantime, use the great <a href=\"https://kagoole.herokuapp.com/\" target=\"_blank\">kagoole</a> <br>\nwebsite. Here is the result using \"segmentation\" as a search term:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F382bfbfbff76f9d42e8f7c11ca7855de%2Fsegmentation_list.png?generation=1610346191759100&amp;alt=media\" alt=\"\"></p>\n<h2>Some useful code snippets</h2>\n<ul>\n<li><p>Reading a TIFF image: </p>\n<pre><code>from osgeo import gdal\nimport numpy as np\nimport matplotlib.image as mpimg\npath_to_tiff = \"/hdd/hubmap-kidney-segmentation/train/2f6ecfcdf.tiff\"\ntiff = gdal.Open(path_to_tiff)\nimage = np.array(tiff.ReadAsArray())\nimage = image.transpose((1,2,0))\nfig, ax = plt.subplots(1, 1, figsize=(20, 20))\nax.imshow(image)\n</code></pre></li>\n</ul>\n<h2>References</h2>\n<ul>\n<li>Dataset details: <a href=\"https://www.kaggle.com/leahscherschel/dataset-details\" target=\"_blank\">https://www.kaggle.com/leahscherschel/dataset-details</a> </li>\n<li>Run-length explanation: <a href=\"https://www.kaggle.com/leahscherschel/run-length-encoding\" target=\"_blank\">https://www.kaggle.com/leahscherschel/run-length-encoding</a></li>\n<li>Segmentation library for Pytorch: <a href=\"https://github.com/qubvel/segmentation_models.pytorch\" target=\"_blank\">https://github.com/qubvel/segmentation_models.pytorch</a> </li>\n<li>List of tips from previous segmentation Kaggle competitions: <a href=\"https://neptune.ai/blog/image-segmentation-tips-and-tricks-from-kaggle-competitions\" target=\"_blank\">https://neptune.ai/blog/image-segmentation-tips-and-tricks-from-kaggle-competitions</a></li>\n<li>A good streamlit app for clothes segmentation: <a href=\"https://github.com/ternaus/cloths_segmentation_demo\" target=\"_blank\">https://github.com/ternaus/cloths_segmentation_demo</a></li>\n<li>TIFF: <a href=\"https://en.wikipedia.org/wiki/TIFF\" target=\"_blank\">https://en.wikipedia.org/wiki/TIFF</a></li>\n<li>HuBMAP website: <a href=\"https://hubmapconsortium.org/\" target=\"_blank\">https://hubmapconsortium.org/</a></li>\n<li>HuBMAP data portal: <a href=\"https://portal.hubmapconsortium.org/\" target=\"_blank\">https://portal.hubmapconsortium.org/</a></li>\n<li>A good discussion about wrong/missing annotations and how to correct them: <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201569\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201569</a></li>\n<li>Ground truth shift: <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816#1116263\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816#1116263</a> </li>\n<li>Annotation issues: <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664</a> </li>\n</ul>",
      "rawMarkdown": "It has been some time since I made one of these posts, so here is one to kick-start the new year. Let's go!\n\n\n\n## General \n\n* HuBMAP is short for Human BioMolecular Atlas Program.\n* Sponsored by the NIH (National Institues of Health)\n* Detect functional tissue units (FTUs in short) across different tissue preparation pipelines.\n* An FTU is a 3D block of cells\n\n\n\n## Metric\n\n[Dice coefficient](https://en.wikipedia.org/wiki/S%C3%B8rensen%E2%80%93Dice_coefficient) between predicted and ground truth mask pixels.\n\n\n\n## Dataset \n\n\n* Metadata: HuBMAP-20-dataset_information.csv\n* Various slices at different sizes of the train dataset can be found here (thanks to @Iafoss for providing these): \n    * https://www.kaggle.com/iafoss/hubmap-512x512\n    * https://www.kaggle.com/iafoss/hubmap-256x256\n    * https://www.kaggle.com/iafoss/hubmap-1024x1024\n* The provided data consists of `11` fresh frozen (FF) and `9` Formalin Fixed Paraffin Embedded (FFPE) PAS kidney images. \nHere is how the data is divided between train and test (this is for the **previous** dataset)\n\n\n\n| Train        | Public test  | Private test |\n| -------------|------------- |------------- |\n| 8 TIFF images + masks|   5 TIFF images + masks|              7 TIFF images + masks found on data portal|\n\n\n* Notice that this is the previous dataset and new data will be collected for the private test part.\n* This was due to the discovery of the previous private dataset on the NIH HuBMAP data portal. Another reason is the bad\nquality of the provided masks and the organizers are taking additional steps to insure better quality.\n* The new private dataset will be comprised of `5` fresh frozen and `5` FFPE PAS images. This will be available around mid-January. So far, here is the grouping break of the updated dataset:\n\n| Train        | Public test  | Private test |\n| -------------|------------- |------------- |\n| 20 TIFF images + masks|   Unknown|              10 TIFF images (5 FF + 5 FFPE)|\n\n* For more details, read this [discussion](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884). \n* Here is what one of the original TIFF images looks like: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Fdd80ff596b1cde86d22d76545e2c77e9%2Ftiff_sample.png?generation=1610298681198226&alt=media)\n\n* Some metadata about the **old** dataset\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Ffd60353c1cc8d35b28ec1c418af20eea%2Fmetadata_old_hubmap.png?generation=1610479642195706&alt=media)\n\n\n## Models\n\n\n* You can start with a ready to use model (thanks to @iglovikov). To train: \n\n- clone the repo: https://github.com/ternaus/cloths_segmentation/tree/main/cloths_segmentation\n- set env variables: `IMAGE_PATH` and `MASK_PATH`\n- run: `python -m cloths_segmentation.train -c config.yml`. You only need to provide a config.yml. \n\nHere is an example: \n\n\n```\nseed: 1984\n\nnum_workers: 4\nexperiment_name: \"best_experiment\"\n\nval_split: 0.1\n\nmodel:\n  type: segmentation_models_pytorch.Unet\n  encoder_name: timm-efficientnet-b3\n  classes: 1\n  encoder_weights: noisy-student\n\ntrainer:\n  type: pytorch_lightning.Trainer\n  gpus: 1\n  max_epochs: 70\n  distributed_backend: ddp\n  progress_bar_refresh_rate: 1\n  benchmark: True\n  precision: 16\n  gradient_clip_val: 5.0\n  num_sanity_val_steps: 2\n  sync_batchnorm: True\n\n\n\nscheduler:\n  type: torch.optim.lr_scheduler.CosineAnnealingWarmRestarts\n  T_0: 10\n  T_mult: 2\n\ntrain_parameters:\n  batch_size: 8\n\ncheckpoint_callback:\n  type: pytorch_lightning.callbacks.ModelCheckpoint\n  filepath: \"best_experiment\"\n  monitor: val_iou\n  verbose: True\n  mode: max\n  save_top_k: -1\n\nval_parameters:\n  batch_size: 2\n\noptimizer:\n  type: adamp.AdamP\n  lr: 0.0001\n\n\ntrain_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.PadIfNeeded\n        always_apply: False\n        min_height: 800\n        min_width: 800\n        border_mode: 0 # cv2.BORDER_CONSTANT\n        value: 0\n        mask_value: 0\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.RandomCrop\n        always_apply: False\n        height: 512\n        width: 512\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.HorizontalFlip\n        always_apply: False\n        p: 0.5\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n\nval_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.PadIfNeeded\n        always_apply: False\n        min_height: 800\n        min_width: 800\n        border_mode: 0 # cv2.BORDER_CONSTANT\n        value: 0\n        mask_value: 0\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n\ntest_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n```\n\n## Some medical terms\n\n\n* Glomerulus: https://en.wikipedia.org/wiki/Glomerulus_(kidney)\n* PAS (short for periodic acid-Schiff): https://en.wikipedia.org/wiki/Periodic_acid%E2%80%93Schiff_stain\n* FFPE (short for Formalin Fixed Paraffin Embedded): https://www.biochain.com/general/what-is-ffpe-tissue/\n* From the HuBMAP assays [page](https://portal.hubmapconsortium.org/docs/assays):\n* Stained Microscopy:\n\n    > Stained microscopy employs histological stains such as H&E or PAS to improve resolution and contrast for     visualization  of anatomical structures such as tubules or glomeruli. \n\n## Previous similar challenges\n\nHere is a list of previous similar challenges: i.e. either binary segmentation, medical imaging, or both.\n\n* Human Protein Atlas Image Classification: https://www.kaggle.com/c/human-protein-atlas-image-classification \n* TGS Salt Identification Challenge: https://www.kaggle.com/c/tgs-salt-identification-challenge\n* Ultrasound Nerve Segmentation: https://www.kaggle.com/c/ultrasound-nerve-segmentation \n\nI will add more challenges in the list above with more details. In the meantime, use the great [kagoole](https://kagoole.herokuapp.com/) \nwebsite. Here is the result using \"segmentation\" as a search term:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F382bfbfbff76f9d42e8f7c11ca7855de%2Fsegmentation_list.png?generation=1610346191759100&alt=media)\n\n\n## Some useful code snippets\n\n\n* Reading a TIFF image: \n\n    ```\n    from osgeo import gdal\n    import numpy as np\n    import matplotlib.image as mpimg\n    path_to_tiff = \"/hdd/hubmap-kidney-segmentation/train/2f6ecfcdf.tiff\"\n    tiff = gdal.Open(path_to_tiff)\n    image = np.array(tiff.ReadAsArray())\n    image = image.transpose((1,2,0))\n    fig, ax = plt.subplots(1, 1, figsize=(20, 20))\n    ax.imshow(image)\n    ```\n\n\n## References\n\n\n* Dataset details: https://www.kaggle.com/leahscherschel/dataset-details \n* Run-length explanation: https://www.kaggle.com/leahscherschel/run-length-encoding\n* Segmentation library for Pytorch: https://github.com/qubvel/segmentation_models.pytorch \n* List of tips from previous segmentation Kaggle competitions: https://neptune.ai/blog/image-segmentation-tips-and-tricks-from-kaggle-competitions\n* A good streamlit app for clothes segmentation: https://github.com/ternaus/cloths_segmentation_demo\n* TIFF: https://en.wikipedia.org/wiki/TIFF\n* HuBMAP website: https://hubmapconsortium.org/\n* HuBMAP data portal: https://portal.hubmapconsortium.org/\n* A good discussion about wrong/missing annotations and how to correct them: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201569\n* Ground truth shift: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816#1116263 \n* Annotation issues: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664",
      "votes": null
    },
    {
      "id": "1148192",
      "postDate": "01/11/2021 01:33:18",
      "content": "<p>Thanks for the elegant summary.<br>\nLooking forward to data updated soon…</p>",
      "rawMarkdown": "Thanks for the elegant summary.\nLooking forward to data updated soon...",
      "votes": null
    },
    {
      "id": "1148327",
      "postDate": "01/11/2021 04:42:51",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> Awesome !! Thanks for sharing. This post deserves an upvote. Keep posting more of such resources. The code snippets are such a bonus.</p>",
      "rawMarkdown": "yassinealouini Awesome !! Thanks for sharing. This post deserves an upvote. Keep posting more of such resources. The code snippets are such a bonus.",
      "votes": null
    },
    {
      "id": "1148336",
      "postDate": "01/11/2021 04:53:30",
      "content": "<p>Very insightful. Thank you for the effort.</p>",
      "rawMarkdown": "Very insightful. Thank you for the effort.",
      "votes": null
    },
    {
      "id": "1148407",
      "postDate": "01/11/2021 05:35:02",
      "content": "<p>You are welcome. I will keep updating the post, so come again later for more. ;)</p>",
      "rawMarkdown": "You are welcome. I will keep updating the post, so come again later for more. ;)",
      "votes": null
    },
    {
      "id": "1148408",
      "postDate": "01/11/2021 05:35:19",
      "content": "<p>Yes, I will update it once the data is released. </p>",
      "rawMarkdown": "Yes, I will update it once the data is released.",
      "votes": null
    },
    {
      "id": "1152023",
      "postDate": "01/13/2021 17:53:00",
      "content": "<p>You are welcome, glad it helps. </p>",
      "rawMarkdown": "You are welcome, glad it helps.",
      "votes": null
    },
    {
      "id": "1154342",
      "postDate": "01/15/2021 15:35:38",
      "content": "<p>Very insightful！Great works! Thank you!</p>",
      "rawMarkdown": "Very insightful！Great works! Thank you!",
      "votes": null
    },
    {
      "id": "1154465",
      "postDate": "01/15/2021 16:40:12",
      "content": "<p>One of the best resources in Kaggle. Appreciated . Upvoted</p>",
      "rawMarkdown": "One of the best resources in Kaggle. Appreciated . Upvoted",
      "votes": null
    },
    {
      "id": "1154635",
      "postDate": "01/15/2021 18:30:04",
      "content": "<p>Probably not the best but I appreciate the praise. :p</p>",
      "rawMarkdown": "Probably not the best but I appreciate the praise. :p",
      "votes": null
    },
    {
      "id": "1156759",
      "postDate": "01/17/2021 11:44:28",
      "content": "<p>Here is how the predicted mask might look like (done for the old dataset by the way):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F5ebf21ff800463d8239ad147ab1dae01%2F1e2425f28_637_preds_exploration.png?generation=1610883867486834&amp;alt=media\" alt=\"\"> </p>",
      "rawMarkdown": "Here is how the predicted mask might look like (done for the old dataset by the way):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F5ebf21ff800463d8239ad147ab1dae01%2F1e2425f28_637_preds_exploration.png?generation=1610883867486834&alt=media)",
      "votes": null
    },
    {
      "id": "1170971",
      "postDate": "01/26/2021 15:04:46",
      "content": "<p>thx for sharing</p>",
      "rawMarkdown": "thx for sharing",
      "votes": null
    },
    {
      "id": "1171211",
      "postDate": "01/26/2021 17:30:29",
      "content": "<p>You are welcome. I hope we get the new dataset soon so that I can update some of the insights. :)</p>",
      "rawMarkdown": "You are welcome. I hope we get the new dataset soon so that I can update some of the insights. :)",
      "votes": null
    },
    {
      "id": "1238157",
      "postDate": "03/14/2021 17:46:17",
      "content": "<p>UPDATE: For those that haven't noticed yet, the new dataset has been <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/224826\" target=\"_blank\">released</a>. I will update the post to reflect this. </p>",
      "rawMarkdown": "UPDATE: For those that haven't noticed yet, the new dataset has been [released](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/224826). I will update the post to reflect this.",
      "votes": null
    },
    {
      "id": "1280782",
      "postDate": "04/22/2021 10:41:13",
      "content": "<p>Great works! Thank you!</p>",
      "rawMarkdown": "Great works! Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1148192,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "01/11/2021 01:33:18",
      "content": "<p>Thanks for the elegant summary.<br>\nLooking forward to data updated soon…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148408,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "01/11/2021 05:35:19",
          "content": "<p>Yes, I will update it once the data is released. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148327,
      "author_name": "balaramk",
      "author_url": "",
      "post_date": "01/11/2021 04:42:51",
      "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> Awesome !! Thanks for sharing. This post deserves an upvote. Keep posting more of such resources. The code snippets are such a bonus.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148407,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "01/11/2021 05:35:02",
          "content": "<p>You are welcome. I will keep updating the post, so come again later for more. ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148336,
      "author_name": "datawarriors",
      "author_url": "",
      "post_date": "01/11/2021 04:53:30",
      "content": "<p>Very insightful. Thank you for the effort.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1152023,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "01/13/2021 17:53:00",
          "content": "<p>You are welcome, glad it helps. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1154342,
      "author_name": "aikeyz",
      "author_url": "",
      "post_date": "01/15/2021 15:35:38",
      "content": "<p>Very insightful！Great works! Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1154465,
      "author_name": "mragpavank",
      "author_url": "",
      "post_date": "01/15/2021 16:40:12",
      "content": "<p>One of the best resources in Kaggle. Appreciated . Upvoted</p>",
      "votes": null,
      "replies": [
        {
          "id": 1154635,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "01/15/2021 18:30:04",
          "content": "<p>Probably not the best but I appreciate the praise. :p</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1156759,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "01/17/2021 11:44:28",
      "content": "<p>Here is how the predicted mask might look like (done for the old dataset by the way):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F5ebf21ff800463d8239ad147ab1dae01%2F1e2425f28_637_preds_exploration.png?generation=1610883867486834&amp;alt=media\" alt=\"\"> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1170971,
      "author_name": "adammayor",
      "author_url": "",
      "post_date": "01/26/2021 15:04:46",
      "content": "<p>thx for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 1171211,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "01/26/2021 17:30:29",
          "content": "<p>You are welcome. I hope we get the new dataset soon so that I can update some of the insights. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1238157,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/14/2021 17:46:17",
      "content": "<p>UPDATE: For those that haven't noticed yet, the new dataset has been <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/224826\" target=\"_blank\">released</a>. I will update the post to reflect this. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1280782,
      "author_name": "niceanthony",
      "author_url": "",
      "post_date": "04/22/2021 10:41:13",
      "content": "<p>Great works! Thank you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1147751": "It has been some time since I made one of these posts, so here is one to kick-start the new year. Let's go!\n\n\n\n## General \n\n* HuBMAP is short for Human BioMolecular Atlas Program.\n* Sponsored by the NIH (National Institues of Health)\n* Detect functional tissue units (FTUs in short) across different tissue preparation pipelines.\n* An FTU is a 3D block of cells\n\n\n\n## Metric\n\n[Dice coefficient](https://en.wikipedia.org/wiki/S%C3%B8rensen%E2%80%93Dice_coefficient) between predicted and ground truth mask pixels.\n\n\n\n## Dataset \n\n\n* Metadata: HuBMAP-20-dataset_information.csv\n* Various slices at different sizes of the train dataset can be found here (thanks to @Iafoss for providing these): \n    * https://www.kaggle.com/iafoss/hubmap-512x512\n    * https://www.kaggle.com/iafoss/hubmap-256x256\n    * https://www.kaggle.com/iafoss/hubmap-1024x1024\n* The provided data consists of `11` fresh frozen (FF) and `9` Formalin Fixed Paraffin Embedded (FFPE) PAS kidney images. \nHere is how the data is divided between train and test (this is for the **previous** dataset)\n\n\n\n| Train        | Public test  | Private test |\n| -------------|------------- |------------- |\n| 8 TIFF images + masks|   5 TIFF images + masks|              7 TIFF images + masks found on data portal|\n\n\n* Notice that this is the previous dataset and new data will be collected for the private test part.\n* This was due to the discovery of the previous private dataset on the NIH HuBMAP data portal. Another reason is the bad\nquality of the provided masks and the organizers are taking additional steps to insure better quality.\n* The new private dataset will be comprised of `5` fresh frozen and `5` FFPE PAS images. This will be available around mid-January. So far, here is the grouping break of the updated dataset:\n\n| Train        | Public test  | Private test |\n| -------------|------------- |------------- |\n| 20 TIFF images + masks|   Unknown|              10 TIFF images (5 FF + 5 FFPE)|\n\n* For more details, read this [discussion](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884). \n* Here is what one of the original TIFF images looks like: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Fdd80ff596b1cde86d22d76545e2c77e9%2Ftiff_sample.png?generation=1610298681198226&alt=media)\n\n* Some metadata about the **old** dataset\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2Ffd60353c1cc8d35b28ec1c418af20eea%2Fmetadata_old_hubmap.png?generation=1610479642195706&alt=media)\n\n\n## Models\n\n\n* You can start with a ready to use model (thanks to @iglovikov). To train: \n\n- clone the repo: https://github.com/ternaus/cloths_segmentation/tree/main/cloths_segmentation\n- set env variables: `IMAGE_PATH` and `MASK_PATH`\n- run: `python -m cloths_segmentation.train -c config.yml`. You only need to provide a config.yml. \n\nHere is an example: \n\n\n```\nseed: 1984\n\nnum_workers: 4\nexperiment_name: \"best_experiment\"\n\nval_split: 0.1\n\nmodel:\n  type: segmentation_models_pytorch.Unet\n  encoder_name: timm-efficientnet-b3\n  classes: 1\n  encoder_weights: noisy-student\n\ntrainer:\n  type: pytorch_lightning.Trainer\n  gpus: 1\n  max_epochs: 70\n  distributed_backend: ddp\n  progress_bar_refresh_rate: 1\n  benchmark: True\n  precision: 16\n  gradient_clip_val: 5.0\n  num_sanity_val_steps: 2\n  sync_batchnorm: True\n\n\n\nscheduler:\n  type: torch.optim.lr_scheduler.CosineAnnealingWarmRestarts\n  T_0: 10\n  T_mult: 2\n\ntrain_parameters:\n  batch_size: 8\n\ncheckpoint_callback:\n  type: pytorch_lightning.callbacks.ModelCheckpoint\n  filepath: \"best_experiment\"\n  monitor: val_iou\n  verbose: True\n  mode: max\n  save_top_k: -1\n\nval_parameters:\n  batch_size: 2\n\noptimizer:\n  type: adamp.AdamP\n  lr: 0.0001\n\n\ntrain_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.PadIfNeeded\n        always_apply: False\n        min_height: 800\n        min_width: 800\n        border_mode: 0 # cv2.BORDER_CONSTANT\n        value: 0\n        mask_value: 0\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.RandomCrop\n        always_apply: False\n        height: 512\n        width: 512\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.HorizontalFlip\n        always_apply: False\n        p: 0.5\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n\nval_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.PadIfNeeded\n        always_apply: False\n        min_height: 800\n        min_width: 800\n        border_mode: 0 # cv2.BORDER_CONSTANT\n        value: 0\n        mask_value: 0\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n\ntest_aug:\n  transform:\n    __class_fullname__: albumentations.core.composition.Compose\n    bbox_params: null\n    keypoint_params: null\n    p: 1\n    transforms:\n      - __class_fullname__: albumentations.augmentations.transforms.LongestMaxSize\n        always_apply: False\n        max_size: 800\n        p: 1\n      - __class_fullname__: albumentations.augmentations.transforms.Normalize\n        always_apply: false\n        max_pixel_value: 255.0\n        mean:\n          - 0.485\n          - 0.456\n          - 0.406\n        p: 1\n        std:\n          - 0.229\n          - 0.224\n          - 0.225\n```\n\n## Some medical terms\n\n\n* Glomerulus: https://en.wikipedia.org/wiki/Glomerulus_(kidney)\n* PAS (short for periodic acid-Schiff): https://en.wikipedia.org/wiki/Periodic_acid%E2%80%93Schiff_stain\n* FFPE (short for Formalin Fixed Paraffin Embedded): https://www.biochain.com/general/what-is-ffpe-tissue/\n* From the HuBMAP assays [page](https://portal.hubmapconsortium.org/docs/assays):\n* Stained Microscopy:\n\n    > Stained microscopy employs histological stains such as H&E or PAS to improve resolution and contrast for     visualization  of anatomical structures such as tubules or glomeruli. \n\n## Previous similar challenges\n\nHere is a list of previous similar challenges: i.e. either binary segmentation, medical imaging, or both.\n\n* Human Protein Atlas Image Classification: https://www.kaggle.com/c/human-protein-atlas-image-classification \n* TGS Salt Identification Challenge: https://www.kaggle.com/c/tgs-salt-identification-challenge\n* Ultrasound Nerve Segmentation: https://www.kaggle.com/c/ultrasound-nerve-segmentation \n\nI will add more challenges in the list above with more details. In the meantime, use the great [kagoole](https://kagoole.herokuapp.com/) \nwebsite. Here is the result using \"segmentation\" as a search term:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F382bfbfbff76f9d42e8f7c11ca7855de%2Fsegmentation_list.png?generation=1610346191759100&alt=media)\n\n\n## Some useful code snippets\n\n\n* Reading a TIFF image: \n\n    ```\n    from osgeo import gdal\n    import numpy as np\n    import matplotlib.image as mpimg\n    path_to_tiff = \"/hdd/hubmap-kidney-segmentation/train/2f6ecfcdf.tiff\"\n    tiff = gdal.Open(path_to_tiff)\n    image = np.array(tiff.ReadAsArray())\n    image = image.transpose((1,2,0))\n    fig, ax = plt.subplots(1, 1, figsize=(20, 20))\n    ax.imshow(image)\n    ```\n\n\n## References\n\n\n* Dataset details: https://www.kaggle.com/leahscherschel/dataset-details \n* Run-length explanation: https://www.kaggle.com/leahscherschel/run-length-encoding\n* Segmentation library for Pytorch: https://github.com/qubvel/segmentation_models.pytorch \n* List of tips from previous segmentation Kaggle competitions: https://neptune.ai/blog/image-segmentation-tips-and-tricks-from-kaggle-competitions\n* A good streamlit app for clothes segmentation: https://github.com/ternaus/cloths_segmentation_demo\n* TIFF: https://en.wikipedia.org/wiki/TIFF\n* HuBMAP website: https://hubmapconsortium.org/\n* HuBMAP data portal: https://portal.hubmapconsortium.org/\n* A good discussion about wrong/missing annotations and how to correct them: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201569\n* Ground truth shift: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/201816#1116263 \n* Annotation issues: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/197664",
    "1148192": "Thanks for the elegant summary.\nLooking forward to data updated soon...",
    "1148327": "yassinealouini Awesome !! Thanks for sharing. This post deserves an upvote. Keep posting more of such resources. The code snippets are such a bonus.",
    "1148336": "Very insightful. Thank you for the effort.",
    "1148407": "You are welcome. I will keep updating the post, so come again later for more. ;)",
    "1148408": "Yes, I will update it once the data is released.",
    "1152023": "You are welcome, glad it helps.",
    "1154342": "Very insightful！Great works! Thank you!",
    "1154465": "One of the best resources in Kaggle. Appreciated . Upvoted",
    "1154635": "Probably not the best but I appreciate the praise. :p",
    "1156759": "Here is how the predicted mask might look like (done for the old dataset by the way):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F5ebf21ff800463d8239ad147ab1dae01%2F1e2425f28_637_preds_exploration.png?generation=1610883867486834&alt=media)",
    "1170971": "thx for sharing",
    "1171211": "You are welcome. I hope we get the new dataset soon so that I can update some of the insights. :)",
    "1238157": "UPDATE: For those that haven't noticed yet, the new dataset has been [released](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/224826). I will update the post to reflect this.",
    "1280782": "Great works! Thank you!"
  },
  "source": "meta"
}