{
  "id": 199490,
  "title": "What I learned so far",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/199490",
  "author_name": "Marcos Novaes",
  "post_date": "2020-11-26T00:01:25.933000",
  "votes": 36,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello Kagglers!</p>\n<p>Wow, it's amazing how many high quality Notebooks have been created in this competition so far! I have not had the chance to read them all, but they all look great! Thanks for the hard work everyone!</p>\n<p>I would like to share a few things that I have myself learned in this competition</p>\n<p>1) Thanks <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for the <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198382\" target=\"_blank\">great set of notebooks</a> that leverage the lessons learned during previous segmentation challenges such as the <a href=\"https://www.kaggle.com/c/ultrasound-nerve-segmentation\" target=\"_blank\">Ultrasound Nerve Segmentation</a>.<br>\n I learned several tricks from this contribution, such as the trick of tiling the image with a \"reshape+transpose trick\". This is quite a trick indeed. But more importantly, <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> shared an important aspect of his approach: downsizing the image resolution by  factor of 4 to get to the score  of 0.836.  This is a huge insight for this problem. I think that this proves that in general we have to reduce the image such that we get complete glomeruli in them. If we tile the image at full resolution we get parts of glomeruli but not whole ones, so ML models would have difficulty learning shape. My initial models at full resolution confirm this.  Here is the section of code from the tiling section:</p>\n<p>#split image and mask into tiles using the reshape+transpose trick<br>\n        img = cv2.resize(img,(img.shape[1]//reduce,img.shape[0]//reduce),<br>\n                         interpolation = cv2.INTER_AREA)<br>\n        img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)<br>\n        img = img.transpose(0,2,1,3,4).reshape(-1,sz,sz,3)</p>\n<pre><code>    mask = cv2.resize(mask,(mask.shape[1]//reduce,mask.shape[0]//reduce),\n                      interpolation = cv2.INTER_NEAREST)\n</code></pre>\n<p>2) Second on my list of tricks would be<a href=\"https://www.kaggle.com/mistag/data-hubmap-sharded-tfrecords-512x512\" target=\"_blank\"> this one </a>by <a href=\"https://www.kaggle.com/mistag\" target=\"_blank\">@mistag</a> that demonstrates how to tile the images in 512x512 tiles in TFRecord format. I have also contributed a <a href=\"https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords\" target=\"_blank\">notebook</a> with an equivalent intent but I think this implementation is much more concise and the sharding scheme is very cool. This basically reduce my 500 lines of code to about 20, but wow, it is really dense python!</p>\n<p>3) Number three trick is the way to <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198440\" target=\"_blank\">create Kaggle dataset from a zip file</a> contributed by <a href=\"https://www.kaggle.com/sreevishnudamodaran\" target=\"_blank\">@sreevishnudamodaran</a> . Now we can save all the datasets with a directory structure! </p>\n<p>4) And finally, the <a href=\"https://www.kaggle.com/joshi98kishan/hubmap-pytorch-tpu-segmentation\" target=\"_blank\">recent notebook</a> by <a href=\"https://www.kaggle.com/joshi98kishan\" target=\"_blank\">@joshi98kishan</a> which explains how to leverage TPUs using PyTorch!!! Now we can all use TPUs with either Tensorflow or Pythorch, so there is no excuse that we should not soon reach 0.999 accuracy!</p>\n<p>It's been great to learn from all of you! Hope my snippets have also helped! See you all around and Happy Thanksgiving!</p>\n<p>Cheers,</p>\n<p>--Marcos</p>",
  "messages": [
    {
      "id": 1091323,
      "postDate": "2020-11-26T00:01:25.933Z",
      "content": "<p>Hello Kagglers!</p>\n<p>Wow, it's amazing how many high quality Notebooks have been created in this competition so far! I have not had the chance to read them all, but they all look great! Thanks for the hard work everyone!</p>\n<p>I would like to share a few things that I have myself learned in this competition</p>\n<p>1) Thanks <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> for the <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198382\" target=\"_blank\">great set of notebooks</a> that leverage the lessons learned during previous segmentation challenges such as the <a href=\"https://www.kaggle.com/c/ultrasound-nerve-segmentation\" target=\"_blank\">Ultrasound Nerve Segmentation</a>.<br>\n I learned several tricks from this contribution, such as the trick of tiling the image with a \"reshape+transpose trick\". This is quite a trick indeed. But more importantly, <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> shared an important aspect of his approach: downsizing the image resolution by  factor of 4 to get to the score  of 0.836.  This is a huge insight for this problem. I think that this proves that in general we have to reduce the image such that we get complete glomeruli in them. If we tile the image at full resolution we get parts of glomeruli but not whole ones, so ML models would have difficulty learning shape. My initial models at full resolution confirm this.  Here is the section of code from the tiling section:</p>\n<p>#split image and mask into tiles using the reshape+transpose trick<br>\n        img = cv2.resize(img,(img.shape[1]//reduce,img.shape[0]//reduce),<br>\n                         interpolation = cv2.INTER_AREA)<br>\n        img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)<br>\n        img = img.transpose(0,2,1,3,4).reshape(-1,sz,sz,3)</p>\n<pre><code>    mask = cv2.resize(mask,(mask.shape[1]//reduce,mask.shape[0]//reduce),\n                      interpolation = cv2.INTER_NEAREST)\n</code></pre>\n<p>2) Second on my list of tricks would be<a href=\"https://www.kaggle.com/mistag/data-hubmap-sharded-tfrecords-512x512\" target=\"_blank\"> this one </a>by <a href=\"https://www.kaggle.com/mistag\" target=\"_blank\">@mistag</a> that demonstrates how to tile the images in 512x512 tiles in TFRecord format. I have also contributed a <a href=\"https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords\" target=\"_blank\">notebook</a> with an equivalent intent but I think this implementation is much more concise and the sharding scheme is very cool. This basically reduce my 500 lines of code to about 20, but wow, it is really dense python!</p>\n<p>3) Number three trick is the way to <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198440\" target=\"_blank\">create Kaggle dataset from a zip file</a> contributed by <a href=\"https://www.kaggle.com/sreevishnudamodaran\" target=\"_blank\">@sreevishnudamodaran</a> . Now we can save all the datasets with a directory structure! </p>\n<p>4) And finally, the <a href=\"https://www.kaggle.com/joshi98kishan/hubmap-pytorch-tpu-segmentation\" target=\"_blank\">recent notebook</a> by <a href=\"https://www.kaggle.com/joshi98kishan\" target=\"_blank\">@joshi98kishan</a> which explains how to leverage TPUs using PyTorch!!! Now we can all use TPUs with either Tensorflow or Pythorch, so there is no excuse that we should not soon reach 0.999 accuracy!</p>\n<p>It's been great to learn from all of you! Hope my snippets have also helped! See you all around and Happy Thanksgiving!</p>\n<p>Cheers,</p>\n<p>--Marcos</p>",
      "rawMarkdown": "Hello Kagglers!\n\nWow, it's amazing how many high quality Notebooks have been created in this competition so far! I have not had the chance to read them all, but they all look great! Thanks for the hard work everyone!\n\nI would like to share a few things that I have myself learned in this competition\n\n1) Thanks @iafoss for the [great set of notebooks](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198382) that leverage the lessons learned during previous segmentation challenges such as the [Ultrasound Nerve Segmentation](https://www.kaggle.com/c/ultrasound-nerve-segmentation).\n I learned several tricks from this contribution, such as the trick of tiling the image with a \"reshape+transpose trick\". This is quite a trick indeed. But more importantly, @iafoss shared an important aspect of his approach: downsizing the image resolution by  factor of 4 to get to the score  of 0.836.  This is a huge insight for this problem. I think that this proves that in general we have to reduce the image such that we get complete glomeruli in them. If we tile the image at full resolution we get parts of glomeruli but not whole ones, so ML models would have difficulty learning shape. My initial models at full resolution confirm this.  Here is the section of code from the tiling section:\n\n #split image and mask into tiles using the reshape+transpose trick\n        img = cv2.resize(img,(img.shape[1]//reduce,img.shape[0]//reduce),\n                         interpolation = cv2.INTER_AREA)\n        img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)\n        img = img.transpose(0,2,1,3,4).reshape(-1,sz,sz,3)\n\n        mask = cv2.resize(mask,(mask.shape[1]//reduce,mask.shape[0]//reduce),\n                          interpolation = cv2.INTER_NEAREST)\n\n2) Second on my list of tricks would be[ this one ](https://www.kaggle.com/mistag/data-hubmap-sharded-tfrecords-512x512)by @mistag that demonstrates how to tile the images in 512x512 tiles in TFRecord format. I have also contributed a [notebook](https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords) with an equivalent intent but I think this implementation is much more concise and the sharding scheme is very cool. This basically reduce my 500 lines of code to about 20, but wow, it is really dense python!\n\n3) Number three trick is the way to [create Kaggle dataset from a zip file](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198440) contributed by @sreevishnudamodaran . Now we can save all the datasets with a directory structure! \n\n4) And finally, the [recent notebook](https://www.kaggle.com/joshi98kishan/hubmap-pytorch-tpu-segmentation) by @joshi98kishan which explains how to leverage TPUs using PyTorch!!! Now we can all use TPUs with either Tensorflow or Pythorch, so there is no excuse that we should not soon reach 0.999 accuracy!\n\nIt's been great to learn from all of you! Hope my snippets have also helped! See you all around and Happy Thanksgiving!\n\nCheers,\n\n--Marcos\n\n",
      "votes": 36
    },
    {
      "id": 1091581,
      "postDate": "2020-11-26T05:46:17.990Z",
      "content": "<p>But there are only a few notebooks that have a successful submission which is unusual comparing other competitions in Kaggle. I wonder why…</p>",
      "rawMarkdown": "But there are only a few notebooks that have a successful submission which is unusual comparing other competitions in Kaggle. I wonder why...",
      "votes": 1,
      "replies": [
        {
          "id": 1091705,
          "postDate": "2020-11-26T08:09:51.207Z",
          "content": "<p>This is possibly due to the fact that a bunch of Kaggle competitions just launched and people are spread out between them. The <a href=\"https://www.kaggle.com/c/nfl-impact-detection/overview\" target=\"_blank\">NFL impact detection</a> computer vision challenge started right before this one and has a similar prize. You'll see only 17 teams have entered.</p>",
          "rawMarkdown": "This is possibly due to the fact that a bunch of Kaggle competitions just launched and people are spread out between them. The [NFL impact detection](https://www.kaggle.com/c/nfl-impact-detection/overview) computer vision challenge started right before this one and has a similar prize. You'll see only 17 teams have entered.",
          "votes": 1
        },
        {
          "id": 1091777,
          "postDate": "2020-11-26T09:23:06.300Z",
          "content": "<p>But in <code>Casava</code> competition there are already 400+ teams. It was also launched in the same period.</p>",
          "rawMarkdown": "But in `Casava` competition there are already 400+ teams. It was also launched in the same period.\n"
        },
        {
          "id": 1091950,
          "postDate": "2020-11-26T12:14:25.897Z",
          "content": "<p>I think for image classification competitions like Casava most kagglers already have some code lying around so it is easy to submit baseline submissions quickly. This competition needs a bit more data processing</p>",
          "rawMarkdown": "I think for image classification competitions like Casava most kagglers already have some code lying around so it is easy to submit baseline submissions quickly. This competition needs a bit more data processing",
          "votes": 2
        },
        {
          "id": 1092354,
          "postDate": "2020-11-26T18:16:21.630Z",
          "content": "<p>I imagine <a href=\"https://www.youtube.com/watch?v=hBvUrj0FUiw&amp;t=2136s\" target=\"_blank\">this video</a> and the smaller dataset size drove a lot of Kagglers to Casava too ;)</p>",
          "rawMarkdown": "I imagine [this video](https://www.youtube.com/watch?v=hBvUrj0FUiw&t=2136s) and the smaller dataset size drove a lot of Kagglers to Casava too ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1094597,
      "postDate": "2020-11-28T18:45:49.917Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1092709,
      "postDate": "2020-11-27T05:29:41.277Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1092766,
          "postDate": "2020-11-27T06:37:00.450Z",
          "content": "<p>I don't think so, we have such data pipeline for the PyTorch. </p>\n<p>PyTorch XLA is still developing to leverage most of the power of the TPU.</p>",
          "rawMarkdown": "I don't think so, we have such data pipeline for the PyTorch. \n\nPyTorch XLA is still developing to leverage most of the power of the TPU."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1091581,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2020-11-26T05:46:17.990000",
      "content": "<p>But there are only a few notebooks that have a successful submission which is unusual comparing other competitions in Kaggle. I wonder why…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1091705,
          "author_name": "Matt",
          "author_url": "",
          "post_date": "2020-11-26T08:09:51.207000",
          "content": "<p>This is possibly due to the fact that a bunch of Kaggle competitions just launched and people are spread out between them. The <a href=\"https://www.kaggle.com/c/nfl-impact-detection/overview\" target=\"_blank\">NFL impact detection</a> computer vision challenge started right before this one and has a similar prize. You'll see only 17 teams have entered.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1091777,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2020-11-26T09:23:06.300000",
          "content": "<p>But in <code>Casava</code> competition there are already 400+ teams. It was also launched in the same period.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1091950,
          "author_name": "quantumgeek",
          "author_url": "",
          "post_date": "2020-11-26T12:14:25.897000",
          "content": "<p>I think for image classification competitions like Casava most kagglers already have some code lying around so it is easy to submit baseline submissions quickly. This competition needs a bit more data processing</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1092354,
          "author_name": "Matt",
          "author_url": "",
          "post_date": "2020-11-26T18:16:21.630000",
          "content": "<p>I imagine <a href=\"https://www.youtube.com/watch?v=hBvUrj0FUiw&amp;t=2136s\" target=\"_blank\">this video</a> and the smaller dataset size drove a lot of Kagglers to Casava too ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1094597,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-28T18:45:49.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1092709,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-27T05:29:41.277000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1092766,
          "author_name": "Kishan Joshi",
          "author_url": "",
          "post_date": "2020-11-27T06:37:00.450000",
          "content": "<p>I don't think so, we have such data pipeline for the PyTorch. </p>\n<p>PyTorch XLA is still developing to leverage most of the power of the TPU.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1091323": "Hello Kagglers!\n\nWow, it's amazing how many high quality Notebooks have been created in this competition so far! I have not had the chance to read them all, but they all look great! Thanks for the hard work everyone!\n\nI would like to share a few things that I have myself learned in this competition\n\n1) Thanks @iafoss for the [great set of notebooks](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198382) that leverage the lessons learned during previous segmentation challenges such as the [Ultrasound Nerve Segmentation](https://www.kaggle.com/c/ultrasound-nerve-segmentation).\n I learned several tricks from this contribution, such as the trick of tiling the image with a \"reshape+transpose trick\". This is quite a trick indeed. But more importantly, @iafoss shared an important aspect of his approach: downsizing the image resolution by  factor of 4 to get to the score  of 0.836.  This is a huge insight for this problem. I think that this proves that in general we have to reduce the image such that we get complete glomeruli in them. If we tile the image at full resolution we get parts of glomeruli but not whole ones, so ML models would have difficulty learning shape. My initial models at full resolution confirm this.  Here is the section of code from the tiling section:\n\n #split image and mask into tiles using the reshape+transpose trick\n        img = cv2.resize(img,(img.shape[1]//reduce,img.shape[0]//reduce),\n                         interpolation = cv2.INTER_AREA)\n        img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)\n        img = img.transpose(0,2,1,3,4).reshape(-1,sz,sz,3)\n\n        mask = cv2.resize(mask,(mask.shape[1]//reduce,mask.shape[0]//reduce),\n                          interpolation = cv2.INTER_NEAREST)\n\n2) Second on my list of tricks would be[ this one ](https://www.kaggle.com/mistag/data-hubmap-sharded-tfrecords-512x512)by @mistag that demonstrates how to tile the images in 512x512 tiles in TFRecord format. I have also contributed a [notebook](https://www.kaggle.com/marcosnovaes/hubmap-read-data-and-build-tfrecords) with an equivalent intent but I think this implementation is much more concise and the sharding scheme is very cool. This basically reduce my 500 lines of code to about 20, but wow, it is really dense python!\n\n3) Number three trick is the way to [create Kaggle dataset from a zip file](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198440) contributed by @sreevishnudamodaran . Now we can save all the datasets with a directory structure! \n\n4) And finally, the [recent notebook](https://www.kaggle.com/joshi98kishan/hubmap-pytorch-tpu-segmentation) by @joshi98kishan which explains how to leverage TPUs using PyTorch!!! Now we can all use TPUs with either Tensorflow or Pythorch, so there is no excuse that we should not soon reach 0.999 accuracy!\n\nIt's been great to learn from all of you! Hope my snippets have also helped! See you all around and Happy Thanksgiving!\n\nCheers,\n\n--Marcos\n\n",
    "1091581": "But there are only a few notebooks that have a successful submission which is unusual comparing other competitions in Kaggle. I wonder why...",
    "1094597": "",
    "1092709": ""
  }
}