{
  "id": 238090,
  "title": "28th Place Solution with inference code",
  "url": "/competitions/hubmap-kidney-segmentation/writeups/datasaurus-28th-place-solution-with-inference-code",
  "author_name": "",
  "post_date": "2021-05-11T07:49:32.863Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congratulations to all the winners and also to the hosts for such a successful competition which will hopefully benefit medical professionals in their work going forward. Commiserations for those who were stung by the shakeup, I'm sure you'll all be at the top of another LB again soon :)</p>\n<h1>CV Strategy</h1>\n<p>I knew early on that <code>d488c759a</code> was causing the public LB to be misaligned with CV, so I generally ignored the public LB other than loosely treating it as a 6th fold. My other 5 folds were simply set up like this to prevent leakage between images during training.</p>\n<pre><code>folds = [\n    [\"aaa6a05cc\", \"26dc41664\", \"b9a3865fc\"],\n    [\"2f6ecfcdf\", \"c68fe75ea\", \"e79de561c\"],\n    [\"1e2425f28\", \"b2dc8411c\", \"afa5e8098\"],\n    [\"0486052bb\", \"cb2d976f4\", \"54f2eec69\"],\n    [\"4ef6695ce\", \"8242609fa\", \"095bf7a1f\"],\n]\n</code></pre>\n<p>Although I did track per-tile dice during training, I calculated my final CV score by calculating the dice on the full resolution out-of-fold predictions (exactly as it's done in the submission notebook).</p>\n<h1>Preprocessing</h1>\n<p>I used a tile size of 2048 which was resized to 1024 before feeding to the model. Overlap of 32 (although I don't think this made much difference).</p>\n<p>I also used the JSON files to create a polygon defining the edge of the mask (since getting the edges right was important in this task). This allowed me to treat this as a two-channel segmentation problem, even though I'd only be using one of them for prediction</p>\n<h1>Augmentation</h1>\n<pre><code>trfm = A.Compose(\n    [\n        A.Resize(img_size_model, img_size_model),\n        A.Flip(),\n        A.RandomRotate90(),\n        A.ColorJitter(p=1),\n        A.OneOf(\n            [\n                A.ElasticTransform(),\n                A.GridDistortion(),\n                A.Blur(blur_limit=(3, 5)),\n            ],\n            p=0.5,\n        ),\n        A.ShiftScaleRotate(),\n        A.Normalize(),\n        ToTensorV2(),\n    ]\n)\n</code></pre>\n<h1>Model</h1>\n<p>UNet with ResNet34d &amp; ResNet50d. That's it! Deeper architectures didn't seem to bring many gains.</p>\n<ul>\n<li>Loss = 0.9 Focal Loss and 0.1 Lovasz</li>\n<li>MixUp with alpha=0.25</li>\n<li>No CutMix. When I was looking at predictions when using CutOut augmentation, the cut boundary between a glomerulus and a non-glom area would cause some mask bleeding, which made me uncomfortable using CutMix. CV changes were in the noise.</li>\n<li>Stochastic weight averaging</li>\n<li>AdamW(lr=0.0001)</li>\n</ul>\n<h1>Submission/postprocessing</h1>\n<p>For the predictions, I used an overlap of 256. This meant that in the corners of some images, you could potentially get 4 predictions which needed to be weighted accordingly. I did this by creating a weight matrix and dividing the full prediction by this. The RAM in the notebooks was too small to keep a prediction array and a weight array in memory, so I did this on the GPU. The mask threshold was 0.5. For TTA I only used flip transforms.</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet34D</td>\n<td>0.94037</td>\n<td>0.917</td>\n<td>0.947</td>\n</tr>\n<tr>\n<td>ResNet50D</td>\n<td>0.93859</td>\n<td>0.916</td>\n<td>0.946</td>\n</tr>\n</tbody>\n</table>\n<p>0.7:0.3 weighting = 0.947</p>\n<p>I did have some pseudo-labelled models using the provided test images which got into 0.948, but the CV wasn't as good so I didn't select those given how important trusting your CV was in this competition.</p>\n<h1>Inference code</h1>\n<p><a href=\"https://www.kaggle.com/anjum48/28th-place-inference-code\" target=\"_blank\">https://www.kaggle.com/anjum48/28th-place-inference-code</a></p>",
  "messages": [
    {
      "id": "1301663",
      "postDate": "05/11/2021 07:22:27",
      "content": "<p>Congratulations to all the winners and also to the hosts for such a successful competition which will hopefully benefit medical professionals in their work going forward. Commiserations for those who were stung by the shakeup, I'm sure you'll all be at the top of another LB again soon :)</p>\n<h1>CV Strategy</h1>\n<p>I knew early on that <code>d488c759a</code> was causing the public LB to be misaligned with CV, so I generally ignored the public LB other than loosely treating it as a 6th fold. My other 5 folds were simply set up like this to prevent leakage between images during training.</p>\n<pre><code>folds = [\n    [\"aaa6a05cc\", \"26dc41664\", \"b9a3865fc\"],\n    [\"2f6ecfcdf\", \"c68fe75ea\", \"e79de561c\"],\n    [\"1e2425f28\", \"b2dc8411c\", \"afa5e8098\"],\n    [\"0486052bb\", \"cb2d976f4\", \"54f2eec69\"],\n    [\"4ef6695ce\", \"8242609fa\", \"095bf7a1f\"],\n]\n</code></pre>\n<p>Although I did track per-tile dice during training, I calculated my final CV score by calculating the dice on the full resolution out-of-fold predictions (exactly as it's done in the submission notebook).</p>\n<h1>Preprocessing</h1>\n<p>I used a tile size of 2048 which was resized to 1024 before feeding to the model. Overlap of 32 (although I don't think this made much difference).</p>\n<p>I also used the JSON files to create a polygon defining the edge of the mask (since getting the edges right was important in this task). This allowed me to treat this as a two-channel segmentation problem, even though I'd only be using one of them for prediction</p>\n<h1>Augmentation</h1>\n<pre><code>trfm = A.Compose(\n    [\n        A.Resize(img_size_model, img_size_model),\n        A.Flip(),\n        A.RandomRotate90(),\n        A.ColorJitter(p=1),\n        A.OneOf(\n            [\n                A.ElasticTransform(),\n                A.GridDistortion(),\n                A.Blur(blur_limit=(3, 5)),\n            ],\n            p=0.5,\n        ),\n        A.ShiftScaleRotate(),\n        A.Normalize(),\n        ToTensorV2(),\n    ]\n)\n</code></pre>\n<h1>Model</h1>\n<p>UNet with ResNet34d &amp; ResNet50d. That's it! Deeper architectures didn't seem to bring many gains.</p>\n<ul>\n<li>Loss = 0.9 Focal Loss and 0.1 Lovasz</li>\n<li>MixUp with alpha=0.25</li>\n<li>No CutMix. When I was looking at predictions when using CutOut augmentation, the cut boundary between a glomerulus and a non-glom area would cause some mask bleeding, which made me uncomfortable using CutMix. CV changes were in the noise.</li>\n<li>Stochastic weight averaging</li>\n<li>AdamW(lr=0.0001)</li>\n</ul>\n<h1>Submission/postprocessing</h1>\n<p>For the predictions, I used an overlap of 256. This meant that in the corners of some images, you could potentially get 4 predictions which needed to be weighted accordingly. I did this by creating a weight matrix and dividing the full prediction by this. The RAM in the notebooks was too small to keep a prediction array and a weight array in memory, so I did this on the GPU. The mask threshold was 0.5. For TTA I only used flip transforms.</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet34D</td>\n<td>0.94037</td>\n<td>0.917</td>\n<td>0.947</td>\n</tr>\n<tr>\n<td>ResNet50D</td>\n<td>0.93859</td>\n<td>0.916</td>\n<td>0.946</td>\n</tr>\n</tbody>\n</table>\n<p>0.7:0.3 weighting = 0.947</p>\n<p>I did have some pseudo-labelled models using the provided test images which got into 0.948, but the CV wasn't as good so I didn't select those given how important trusting your CV was in this competition.</p>\n<h1>Inference code</h1>\n<p><a href=\"https://www.kaggle.com/anjum48/28th-place-inference-code\" target=\"_blank\">https://www.kaggle.com/anjum48/28th-place-inference-code</a></p>",
      "rawMarkdown": "Congratulations to all the winners and also to the hosts for such a successful competition which will hopefully benefit medical professionals in their work going forward. Commiserations for those who were stung by the shakeup, I'm sure you'll all be at the top of another LB again soon :)\n\n# CV Strategy\nI knew early on that `d488c759a` was causing the public LB to be misaligned with CV, so I generally ignored the public LB other than loosely treating it as a 6th fold. My other 5 folds were simply set up like this to prevent leakage between images during training.\n\n```\nfolds = [\n    [\"aaa6a05cc\", \"26dc41664\", \"b9a3865fc\"],\n    [\"2f6ecfcdf\", \"c68fe75ea\", \"e79de561c\"],\n    [\"1e2425f28\", \"b2dc8411c\", \"afa5e8098\"],\n    [\"0486052bb\", \"cb2d976f4\", \"54f2eec69\"],\n    [\"4ef6695ce\", \"8242609fa\", \"095bf7a1f\"],\n]\n```\nAlthough I did track per-tile dice during training, I calculated my final CV score by calculating the dice on the full resolution out-of-fold predictions (exactly as it's done in the submission notebook).\n\n# Preprocessing\nI used a tile size of 2048 which was resized to 1024 before feeding to the model. Overlap of 32 (although I don't think this made much difference).\n\nI also used the JSON files to create a polygon defining the edge of the mask (since getting the edges right was important in this task). This allowed me to treat this as a two-channel segmentation problem, even though I'd only be using one of them for prediction\n\n# Augmentation\n```\ntrfm = A.Compose(\n    [\n        A.Resize(img_size_model, img_size_model),\n        A.Flip(),\n        A.RandomRotate90(),\n        A.ColorJitter(p=1),\n        A.OneOf(\n            [\n                A.ElasticTransform(),\n                A.GridDistortion(),\n                A.Blur(blur_limit=(3, 5)),\n            ],\n            p=0.5,\n        ),\n        A.ShiftScaleRotate(),\n        A.Normalize(),\n        ToTensorV2(),\n    ]\n)\n```\n\n# Model\nUNet with ResNet34d & ResNet50d. That's it! Deeper architectures didn't seem to bring many gains.\n\n- Loss = 0.9 Focal Loss and 0.1 Lovasz\n- MixUp with alpha=0.25\n- No CutMix. When I was looking at predictions when using CutOut augmentation, the cut boundary between a glomerulus and a non-glom area would cause some mask bleeding, which made me uncomfortable using CutMix. CV changes were in the noise.\n- Stochastic weight averaging\n- AdamW(lr=0.0001)\n\n# Submission/postprocessing\nFor the predictions, I used an overlap of 256. This meant that in the corners of some images, you could potentially get 4 predictions which needed to be weighted accordingly. I did this by creating a weight matrix and dividing the full prediction by this. The RAM in the notebooks was too small to keep a prediction array and a weight array in memory, so I did this on the GPU. The mask threshold was 0.5. For TTA I only used flip transforms.\n\n| Backbone     |   CV   | Public | Private |\n| ------------- | ------ | ------ | ------- |\n| ResNet34D  | 0.94037 | 0.917  | 0.947   |\n| ResNet50D  | 0.93859| 0.916  | 0.946   |\n\n0.7:0.3 weighting = 0.947\n\nI did have some pseudo-labelled models using the provided test images which got into 0.948, but the CV wasn't as good so I didn't select those given how important trusting your CV was in this competition.\n\n# Inference code\nhttps://www.kaggle.com/anjum48/28th-place-inference-code",
      "votes": null
    },
    {
      "id": "1301715",
      "postDate": "05/11/2021 07:55:20",
      "content": "<p>Congrats on strongly finish <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> and thanks for sharing solution. </p>",
      "rawMarkdown": "Congrats on strongly finish @anjum48 and thanks for sharing solution.",
      "votes": null
    },
    {
      "id": "1302031",
      "postDate": "05/11/2021 11:12:36",
      "content": "<p>Thanks for sharing and congratulations for your medal , <br>\ndid you tried other Backbones ? </p>",
      "rawMarkdown": "Thanks for sharing and congratulations for your medal , \ndid you tried other Backbones ?",
      "votes": null
    },
    {
      "id": "1302091",
      "postDate": "05/11/2021 11:40:45",
      "content": "<p>I tried ResNet-200D, Densenet-121, EfficientNet-B0, but the mighty ResNet-34D beat them all.</p>\n<p>I also tried different decoder architectures like PAN, MANet, DeepNetV3+, UNet++ but the base UNet seemed to work best for me.</p>",
      "rawMarkdown": "I tried ResNet-200D, Densenet-121, EfficientNet-B0, but the mighty ResNet-34D beat them all.\n\nI also tried different decoder architectures like PAN, MANet, DeepNetV3+, UNet++ but the base UNet seemed to work best for me.",
      "votes": null
    },
    {
      "id": "1302601",
      "postDate": "05/11/2021 16:13:27",
      "content": "<p>Congratulations on the awesome result!</p>",
      "rawMarkdown": "Congratulations on the awesome result!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1301715,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "05/11/2021 07:55:20",
      "content": "<p>Congrats on strongly finish <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> and thanks for sharing solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1302031,
      "author_name": "salimkhazem",
      "author_url": "",
      "post_date": "05/11/2021 11:12:36",
      "content": "<p>Thanks for sharing and congratulations for your medal , <br>\ndid you tried other Backbones ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1302091,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "05/11/2021 11:40:45",
          "content": "<p>I tried ResNet-200D, Densenet-121, EfficientNet-B0, but the mighty ResNet-34D beat them all.</p>\n<p>I also tried different decoder architectures like PAN, MANet, DeepNetV3+, UNet++ but the base UNet seemed to work best for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1302601,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "05/11/2021 16:13:27",
      "content": "<p>Congratulations on the awesome result!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1301663": "Congratulations to all the winners and also to the hosts for such a successful competition which will hopefully benefit medical professionals in their work going forward. Commiserations for those who were stung by the shakeup, I'm sure you'll all be at the top of another LB again soon :)\n\n# CV Strategy\nI knew early on that `d488c759a` was causing the public LB to be misaligned with CV, so I generally ignored the public LB other than loosely treating it as a 6th fold. My other 5 folds were simply set up like this to prevent leakage between images during training.\n\n```\nfolds = [\n    [\"aaa6a05cc\", \"26dc41664\", \"b9a3865fc\"],\n    [\"2f6ecfcdf\", \"c68fe75ea\", \"e79de561c\"],\n    [\"1e2425f28\", \"b2dc8411c\", \"afa5e8098\"],\n    [\"0486052bb\", \"cb2d976f4\", \"54f2eec69\"],\n    [\"4ef6695ce\", \"8242609fa\", \"095bf7a1f\"],\n]\n```\nAlthough I did track per-tile dice during training, I calculated my final CV score by calculating the dice on the full resolution out-of-fold predictions (exactly as it's done in the submission notebook).\n\n# Preprocessing\nI used a tile size of 2048 which was resized to 1024 before feeding to the model. Overlap of 32 (although I don't think this made much difference).\n\nI also used the JSON files to create a polygon defining the edge of the mask (since getting the edges right was important in this task). This allowed me to treat this as a two-channel segmentation problem, even though I'd only be using one of them for prediction\n\n# Augmentation\n```\ntrfm = A.Compose(\n    [\n        A.Resize(img_size_model, img_size_model),\n        A.Flip(),\n        A.RandomRotate90(),\n        A.ColorJitter(p=1),\n        A.OneOf(\n            [\n                A.ElasticTransform(),\n                A.GridDistortion(),\n                A.Blur(blur_limit=(3, 5)),\n            ],\n            p=0.5,\n        ),\n        A.ShiftScaleRotate(),\n        A.Normalize(),\n        ToTensorV2(),\n    ]\n)\n```\n\n# Model\nUNet with ResNet34d & ResNet50d. That's it! Deeper architectures didn't seem to bring many gains.\n\n- Loss = 0.9 Focal Loss and 0.1 Lovasz\n- MixUp with alpha=0.25\n- No CutMix. When I was looking at predictions when using CutOut augmentation, the cut boundary between a glomerulus and a non-glom area would cause some mask bleeding, which made me uncomfortable using CutMix. CV changes were in the noise.\n- Stochastic weight averaging\n- AdamW(lr=0.0001)\n\n# Submission/postprocessing\nFor the predictions, I used an overlap of 256. This meant that in the corners of some images, you could potentially get 4 predictions which needed to be weighted accordingly. I did this by creating a weight matrix and dividing the full prediction by this. The RAM in the notebooks was too small to keep a prediction array and a weight array in memory, so I did this on the GPU. The mask threshold was 0.5. For TTA I only used flip transforms.\n\n| Backbone     |   CV   | Public | Private |\n| ------------- | ------ | ------ | ------- |\n| ResNet34D  | 0.94037 | 0.917  | 0.947   |\n| ResNet50D  | 0.93859| 0.916  | 0.946   |\n\n0.7:0.3 weighting = 0.947\n\nI did have some pseudo-labelled models using the provided test images which got into 0.948, but the CV wasn't as good so I didn't select those given how important trusting your CV was in this competition.\n\n# Inference code\nhttps://www.kaggle.com/anjum48/28th-place-inference-code",
    "1301715": "Congrats on strongly finish @anjum48 and thanks for sharing solution.",
    "1302031": "Thanks for sharing and congratulations for your medal , \ndid you tried other Backbones ?",
    "1302091": "I tried ResNet-200D, Densenet-121, EfficientNet-B0, but the mighty ResNet-34D beat them all.\n\nI also tried different decoder architectures like PAN, MANet, DeepNetV3+, UNet++ but the base UNet seemed to work best for me.",
    "1302601": "Congratulations on the awesome result!"
  },
  "source": "meta"
}