{
  "id": 238365,
  "title": "What a ride! Parts of 6th place solution",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238365",
  "author_name": "Darek Kłeczek",
  "post_date": "2021-05-12T02:23:08.434000",
  "votes": 42,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I put all of my heart into this competition, and I need to start by thanking my wife for putting up with me in the last couple of months - buying a DL rig, waking up or going to sleep at crazy hours, putting my rig on car front seat on a weekend trip, and many more… all while going through Covid (fortunately mildly) and trying to entertain our Kaggle-famous 4yo future scientist while quarantined at home :) </p>\n<p>Thanks to the team members - <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> and <a href=\"https://www.kaggle.com/felipebihaiek\" target=\"_blank\">@felipebihaiek</a> for partnering early on, experimenting together and discussing hundreds of ideas, <a href=\"https://www.kaggle.com/scusywxy\" target=\"_blank\">@scusywxy</a> and <a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a> for the partnership and your lead to figure out the engineering details so that our combined solution finally worked. We had diverse approaches which gave us a boost (2 minutes late for the money zone!). I will only share my contributions here, a more comprehensive summary of our total approach will follow. </p>\n<p>Thanks to Kaggle and the hosts for this competition!</p>\n<h1>Model or Data-Centric Approach?</h1>\n<p>We started the competition with single-cell classification combined with HPA CellSegmentator, and I believed initially this will be the key to the competition. I planned a data centric approach, including pseudo labelling (it gave me a small boost early on) and iterative fixing of most-confused labels. I wanted to use a fastai widget for that, but afaik it doesn’t support multilabel classification unfortunately, so experimented with some other tools, but finally gave up when we realized labelling the cells is not very straightforward. </p>\n<p>We had accurate data to train image level models, so the question was how we can use those models to predict cells? After many experiments, we figured out two approaches that led us first to .549 public lb score, and gave us a 0.02 boost after blending with <a href=\"https://www.kaggle.com/scusywxy\" target=\"_blank\">@scusywxy</a> and <a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a> models. </p>\n<h1>GAPMASK Inference</h1>\n<p>We train a regular image-level model. At inference, we modify the model architecture so that it takes two inputs - image and cell-mask. The image goes through the network alone until the GAP layer. At that point, we do element wise multiplication of the image activations and cell mask. From that point, an image gets expanded into a batch of single-cell images, and we apply the model head to that entire batch. </p>\n<p><img src=\"https://pbs.twimg.com/media/E2SRV0qWQAM7kju?format=jpg&amp;name=medium\" alt=\"Gap-Mask-1\"></p>\n<p>We had two variations of this architecture - one for global average and max concat pooling, one for attention pooling. </p>\n<p>This is the code for our densenet model: </p>\n<pre><code>    def forward(self, images, masks):\n        …\n        e5 = F.relu(e5,inplace=True)\n        # GAP MASK STARTS HERE:\n        e5 = F.interpolate(e5, scale_factor=2, mode='bilinear') # increase grid size to 48x48\n        x = e5.permute(0,2,3,1) # [1, 48, 48, 1024]\n        x = x.squeeze() # [48, 48, 1024]\n        msk = masks.squeeze() # [15, 48, 48]\n        res = msk[...,None] * x[None,...] # [15, 48, 48, 1024]\n        res = res.permute(0,3,1,2)\n        # NOW APPLY THE MODEL HEAD WITH THE MASK DIMENSION AS THE BATCH DIMENSION\n        x = torch.cat((nn.AdaptiveAvgPool2d(1)(res), nn.AdaptiveMaxPool2d(1)(res)), dim=1)\n        x = x.view(x.size(0), -1)\n    …\n        x = self.logit(x)\n        return x\n</code></pre>\n<p>And this is the implementation for inception with attention: <br>\n<img src=\"https://pbs.twimg.com/media/E2SRK5IWUAMli02?format=jpg&amp;name=medium\" alt=\"Gap-Mask-2\"></p>\n<pre><code>    def forward(self, x, masks):\n        …\n    # GAP MASK STARTS HERE (dimensions reflect an image with 15 cells):\n        logits = self.last_linear(features_b)\n        logits = F.interpolate(logits, scale_factor=3, mode='nearest')\n        logits_attention = self.attention(features_b)\n        logits_attention = F.interpolate(logits_attention, scale_factor=3, mode='nearest')\n        logits_attention = logits_attention.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3)\n        attention = F.softmax(logits_attention, dim=2)\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3, self.attention_size * 3)\n        attention = attention.permute(0,2,3,1) # [1,66,66,19])\n        msk = masks.squeeze() # [15,66,66]       \n        attention = msk[...,None] * attention # [15,66,66,19]\n        attention = attention.permute(0,3,1,2) # [15,19,66,66]\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3) #[15, 19, 66*66]\n        attention = attention / (attention.sum(2).unsqueeze(-1) + 1e-7)\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3, self.attention_size * 3) #[15,19,66,66]\n        logits = logits * attention\n        return logits.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3).sum(2).view(-1, self.num_classes) # torch.Size([1, 19])\n</code></pre>\n<h1>GRIDIFY Inference</h1>\n<p>When we started with single cell tiles, we resized each cell to the same size, e.g. 128x128. The problem is that this changed cell resolutions, some<br>\ngot shrunk and some expanded. Inspired by <a href=\"https://arxiv.org/pdf/1906.06423.pdf\" target=\"_blank\">this paper</a>, I looked for an approach that would line up the features model sees while training on full images, with features seen by the model when running inference on single cells. </p>\n<p>The initial gridify approach was to take a single cell crop, copy-paste that crop with multiple augmentations into a 2048x2048 template, and then resize in the same way as during training. Finally, we converged on the following approach: </p>\n<ul>\n<li>Take 4x 512x512 crop from full image (independent of size) from each corner of the single cell</li>\n<li>Put those 4 crops into a single 1024x1024 image</li>\n<li>Resize to maintain same resolution as training. Eg. for the model trained on 768 size (from 2048), we resize into 384x384 cell tile</li>\n</ul>\n<p><img src=\"https://pbs.twimg.com/media/E1J0N0ZXEAYyffs?format=jpg&amp;name=large\" alt=\"gridify\"></p>\n<h1>Models</h1>\n<p>We used three models from HPA 2018 winning solutions: </p>\n<ul>\n<li>Densenet trained on size 768 from bestfitting</li>\n<li>Densenet trained on size 1536 from bestfitting </li>\n<li>Inceptionv3 trained on size 1024 from pudae</li>\n</ul>\n<p>All models were fine-tuned with original loss functions (focal-lovasz-logloss for densenet, focal for inception), starting with the 2018 weights. We froze the model head for 1 epoch, then trained for 3-6 epochs on train + public data (excl. classes 0 and 16) with cosine annealing. </p>\n<h1>Validation</h1>\n<p>Due to potential leakage from similar images (same plate) across folds, we run all images through metric learning model (from bestfitting’s 2018 solution) and cluster similar images together with UMAP/DBSCAN. We treat the clusters as groups and divide data in folds based on group stratified multilabel approach. </p>\n<h1>Postprocessing</h1>\n<p>We reduce probabilities for cells based on negative signal, this gave us a small boost (~0.003): preds[:,:18] = preds[:,:18] * (1 - preds[:,18]).unsqueeze(-1)</p>\n<h1>Tried but didn't work</h1>\n<ul>\n<li>Pseudolabels</li>\n<li>Multi-instance learning (treat each image as a bag of cells, train with bag of cells, predict on single cells). This was the same architecture as shared by Guanshuo Xu in his #8 solution. After analyzing what I did wrong, my conclusion is that I used too many cells per image. After reducing number of cells from 48 to 8, the results for my previously failed experiments improved significantly.</li>\n<li>Using rectangle-shape images during inference</li>\n<li>BagNet <a href=\"https://arxiv.org/pdf/1904.00760.pdf\" target=\"_blank\">paper</a></li>\n<li>…</li>\n</ul>",
  "messages": [
    {
      "id": 1303303,
      "postDate": "2021-05-12T02:23:08.433Z",
      "content": "<p>I put all of my heart into this competition, and I need to start by thanking my wife for putting up with me in the last couple of months - buying a DL rig, waking up or going to sleep at crazy hours, putting my rig on car front seat on a weekend trip, and many more… all while going through Covid (fortunately mildly) and trying to entertain our Kaggle-famous 4yo future scientist while quarantined at home :) </p>\n<p>Thanks to the team members - <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> and <a href=\"https://www.kaggle.com/felipebihaiek\" target=\"_blank\">@felipebihaiek</a> for partnering early on, experimenting together and discussing hundreds of ideas, <a href=\"https://www.kaggle.com/scusywxy\" target=\"_blank\">@scusywxy</a> and <a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a> for the partnership and your lead to figure out the engineering details so that our combined solution finally worked. We had diverse approaches which gave us a boost (2 minutes late for the money zone!). I will only share my contributions here, a more comprehensive summary of our total approach will follow. </p>\n<p>Thanks to Kaggle and the hosts for this competition!</p>\n<h1>Model or Data-Centric Approach?</h1>\n<p>We started the competition with single-cell classification combined with HPA CellSegmentator, and I believed initially this will be the key to the competition. I planned a data centric approach, including pseudo labelling (it gave me a small boost early on) and iterative fixing of most-confused labels. I wanted to use a fastai widget for that, but afaik it doesn’t support multilabel classification unfortunately, so experimented with some other tools, but finally gave up when we realized labelling the cells is not very straightforward. </p>\n<p>We had accurate data to train image level models, so the question was how we can use those models to predict cells? After many experiments, we figured out two approaches that led us first to .549 public lb score, and gave us a 0.02 boost after blending with <a href=\"https://www.kaggle.com/scusywxy\" target=\"_blank\">@scusywxy</a> and <a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a> models. </p>\n<h1>GAPMASK Inference</h1>\n<p>We train a regular image-level model. At inference, we modify the model architecture so that it takes two inputs - image and cell-mask. The image goes through the network alone until the GAP layer. At that point, we do element wise multiplication of the image activations and cell mask. From that point, an image gets expanded into a batch of single-cell images, and we apply the model head to that entire batch. </p>\n<p><img src=\"https://pbs.twimg.com/media/E2SRV0qWQAM7kju?format=jpg&amp;name=medium\" alt=\"Gap-Mask-1\"></p>\n<p>We had two variations of this architecture - one for global average and max concat pooling, one for attention pooling. </p>\n<p>This is the code for our densenet model: </p>\n<pre><code>    def forward(self, images, masks):\n        …\n        e5 = F.relu(e5,inplace=True)\n        # GAP MASK STARTS HERE:\n        e5 = F.interpolate(e5, scale_factor=2, mode='bilinear') # increase grid size to 48x48\n        x = e5.permute(0,2,3,1) # [1, 48, 48, 1024]\n        x = x.squeeze() # [48, 48, 1024]\n        msk = masks.squeeze() # [15, 48, 48]\n        res = msk[...,None] * x[None,...] # [15, 48, 48, 1024]\n        res = res.permute(0,3,1,2)\n        # NOW APPLY THE MODEL HEAD WITH THE MASK DIMENSION AS THE BATCH DIMENSION\n        x = torch.cat((nn.AdaptiveAvgPool2d(1)(res), nn.AdaptiveMaxPool2d(1)(res)), dim=1)\n        x = x.view(x.size(0), -1)\n    …\n        x = self.logit(x)\n        return x\n</code></pre>\n<p>And this is the implementation for inception with attention: <br>\n<img src=\"https://pbs.twimg.com/media/E2SRK5IWUAMli02?format=jpg&amp;name=medium\" alt=\"Gap-Mask-2\"></p>\n<pre><code>    def forward(self, x, masks):\n        …\n    # GAP MASK STARTS HERE (dimensions reflect an image with 15 cells):\n        logits = self.last_linear(features_b)\n        logits = F.interpolate(logits, scale_factor=3, mode='nearest')\n        logits_attention = self.attention(features_b)\n        logits_attention = F.interpolate(logits_attention, scale_factor=3, mode='nearest')\n        logits_attention = logits_attention.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3)\n        attention = F.softmax(logits_attention, dim=2)\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3, self.attention_size * 3)\n        attention = attention.permute(0,2,3,1) # [1,66,66,19])\n        msk = masks.squeeze() # [15,66,66]       \n        attention = msk[...,None] * attention # [15,66,66,19]\n        attention = attention.permute(0,3,1,2) # [15,19,66,66]\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3) #[15, 19, 66*66]\n        attention = attention / (attention.sum(2).unsqueeze(-1) + 1e-7)\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3, self.attention_size * 3) #[15,19,66,66]\n        logits = logits * attention\n        return logits.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3).sum(2).view(-1, self.num_classes) # torch.Size([1, 19])\n</code></pre>\n<h1>GRIDIFY Inference</h1>\n<p>When we started with single cell tiles, we resized each cell to the same size, e.g. 128x128. The problem is that this changed cell resolutions, some<br>\ngot shrunk and some expanded. Inspired by <a href=\"https://arxiv.org/pdf/1906.06423.pdf\" target=\"_blank\">this paper</a>, I looked for an approach that would line up the features model sees while training on full images, with features seen by the model when running inference on single cells. </p>\n<p>The initial gridify approach was to take a single cell crop, copy-paste that crop with multiple augmentations into a 2048x2048 template, and then resize in the same way as during training. Finally, we converged on the following approach: </p>\n<ul>\n<li>Take 4x 512x512 crop from full image (independent of size) from each corner of the single cell</li>\n<li>Put those 4 crops into a single 1024x1024 image</li>\n<li>Resize to maintain same resolution as training. Eg. for the model trained on 768 size (from 2048), we resize into 384x384 cell tile</li>\n</ul>\n<p><img src=\"https://pbs.twimg.com/media/E1J0N0ZXEAYyffs?format=jpg&amp;name=large\" alt=\"gridify\"></p>\n<h1>Models</h1>\n<p>We used three models from HPA 2018 winning solutions: </p>\n<ul>\n<li>Densenet trained on size 768 from bestfitting</li>\n<li>Densenet trained on size 1536 from bestfitting </li>\n<li>Inceptionv3 trained on size 1024 from pudae</li>\n</ul>\n<p>All models were fine-tuned with original loss functions (focal-lovasz-logloss for densenet, focal for inception), starting with the 2018 weights. We froze the model head for 1 epoch, then trained for 3-6 epochs on train + public data (excl. classes 0 and 16) with cosine annealing. </p>\n<h1>Validation</h1>\n<p>Due to potential leakage from similar images (same plate) across folds, we run all images through metric learning model (from bestfitting’s 2018 solution) and cluster similar images together with UMAP/DBSCAN. We treat the clusters as groups and divide data in folds based on group stratified multilabel approach. </p>\n<h1>Postprocessing</h1>\n<p>We reduce probabilities for cells based on negative signal, this gave us a small boost (~0.003): preds[:,:18] = preds[:,:18] * (1 - preds[:,18]).unsqueeze(-1)</p>\n<h1>Tried but didn't work</h1>\n<ul>\n<li>Pseudolabels</li>\n<li>Multi-instance learning (treat each image as a bag of cells, train with bag of cells, predict on single cells). This was the same architecture as shared by Guanshuo Xu in his #8 solution. After analyzing what I did wrong, my conclusion is that I used too many cells per image. After reducing number of cells from 48 to 8, the results for my previously failed experiments improved significantly.</li>\n<li>Using rectangle-shape images during inference</li>\n<li>BagNet <a href=\"https://arxiv.org/pdf/1904.00760.pdf\" target=\"_blank\">paper</a></li>\n<li>…</li>\n</ul>",
      "rawMarkdown": "I put all of my heart into this competition, and I need to start by thanking my wife for putting up with me in the last couple of months - buying a DL rig, waking up or going to sleep at crazy hours, putting my rig on car front seat on a weekend trip, and many more… all while going through Covid (fortunately mildly) and trying to entertain our Kaggle-famous 4yo future scientist while quarantined at home :) \n\nThanks to the team members - @dschettler8845 and @felipebihaiek for partnering early on, experimenting together and discussing hundreds of ideas, @scusywxy and @zehuigong for the partnership and your lead to figure out the engineering details so that our combined solution finally worked. We had diverse approaches which gave us a boost (2 minutes late for the money zone!). I will only share my contributions here, a more comprehensive summary of our total approach will follow. \n\nThanks to Kaggle and the hosts for this competition!\n\n# Model or Data-Centric Approach?\n\nWe started the competition with single-cell classification combined with HPA CellSegmentator, and I believed initially this will be the key to the competition. I planned a data centric approach, including pseudo labelling (it gave me a small boost early on) and iterative fixing of most-confused labels. I wanted to use a fastai widget for that, but afaik it doesn’t support multilabel classification unfortunately, so experimented with some other tools, but finally gave up when we realized labelling the cells is not very straightforward. \n\nWe had accurate data to train image level models, so the question was how we can use those models to predict cells? After many experiments, we figured out two approaches that led us first to .549 public lb score, and gave us a 0.02 boost after blending with @scusywxy and @zehuigong models. \n\n# GAPMASK Inference\n\nWe train a regular image-level model. At inference, we modify the model architecture so that it takes two inputs - image and cell-mask. The image goes through the network alone until the GAP layer. At that point, we do element wise multiplication of the image activations and cell mask. From that point, an image gets expanded into a batch of single-cell images, and we apply the model head to that entire batch. \n\n![Gap-Mask-1](https://pbs.twimg.com/media/E2SRV0qWQAM7kju?format=jpg&name=medium)\n\nWe had two variations of this architecture - one for global average and max concat pooling, one for attention pooling. \n\nThis is the code for our densenet model: \n\n```\n    def forward(self, images, masks):\n        …\n        e5 = F.relu(e5,inplace=True)\n        # GAP MASK STARTS HERE:\n        e5 = F.interpolate(e5, scale_factor=2, mode='bilinear') # increase grid size to 48x48\n        x = e5.permute(0,2,3,1) # [1, 48, 48, 1024]\n        x = x.squeeze() # [48, 48, 1024]\n        msk = masks.squeeze() # [15, 48, 48]\n        res = msk[...,None] * x[None,...] # [15, 48, 48, 1024]\n        res = res.permute(0,3,1,2)\n        # NOW APPLY THE MODEL HEAD WITH THE MASK DIMENSION AS THE BATCH DIMENSION\n        x = torch.cat((nn.AdaptiveAvgPool2d(1)(res), nn.AdaptiveMaxPool2d(1)(res)), dim=1)\n        x = x.view(x.size(0), -1)\n\t…\n        x = self.logit(x)\n        return x\n\n```\nAnd this is the implementation for inception with attention: \n![Gap-Mask-2](https://pbs.twimg.com/media/E2SRK5IWUAMli02?format=jpg&name=medium)\n\n```\n    def forward(self, x, masks):\n        …\n\t# GAP MASK STARTS HERE (dimensions reflect an image with 15 cells):\n        logits = self.last_linear(features_b)\n        logits = F.interpolate(logits, scale_factor=3, mode='nearest')\n        logits_attention = self.attention(features_b)\n        logits_attention = F.interpolate(logits_attention, scale_factor=3, mode='nearest')\n        logits_attention = logits_attention.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3)\n        attention = F.softmax(logits_attention, dim=2)\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3, self.attention_size * 3)\n        attention = attention.permute(0,2,3,1) # [1,66,66,19])\n        msk = masks.squeeze() # [15,66,66]       \n        attention = msk[...,None] * attention # [15,66,66,19]\n        attention = attention.permute(0,3,1,2) # [15,19,66,66]\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3) #[15, 19, 66*66]\n        attention = attention / (attention.sum(2).unsqueeze(-1) + 1e-7)\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3, self.attention_size * 3) #[15,19,66,66]\n        logits = logits * attention\n        return logits.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3).sum(2).view(-1, self.num_classes) # torch.Size([1, 19])\n```\n\n# GRIDIFY Inference\n\nWhen we started with single cell tiles, we resized each cell to the same size, e.g. 128x128. The problem is that this changed cell resolutions, some\ngot shrunk and some expanded. Inspired by [this paper](https://arxiv.org/pdf/1906.06423.pdf), I looked for an approach that would line up the features model sees while training on full images, with features seen by the model when running inference on single cells. \n\nThe initial gridify approach was to take a single cell crop, copy-paste that crop with multiple augmentations into a 2048x2048 template, and then resize in the same way as during training. Finally, we converged on the following approach: \n- Take 4x 512x512 crop from full image (independent of size) from each corner of the single cell\n- Put those 4 crops into a single 1024x1024 image\n- Resize to maintain same resolution as training. Eg. for the model trained on 768 size (from 2048), we resize into 384x384 cell tile\n\n![gridify](https://pbs.twimg.com/media/E1J0N0ZXEAYyffs?format=jpg&name=large)\n\n# Models\n\nWe used three models from HPA 2018 winning solutions: \n- Densenet trained on size 768 from bestfitting\n- Densenet trained on size 1536 from bestfitting \n- Inceptionv3 trained on size 1024 from pudae\n\nAll models were fine-tuned with original loss functions (focal-lovasz-logloss for densenet, focal for inception), starting with the 2018 weights. We froze the model head for 1 epoch, then trained for 3-6 epochs on train + public data (excl. classes 0 and 16) with cosine annealing. \n\n# Validation\n\nDue to potential leakage from similar images (same plate) across folds, we run all images through metric learning model (from bestfitting’s 2018 solution) and cluster similar images together with UMAP/DBSCAN. We treat the clusters as groups and divide data in folds based on group stratified multilabel approach. \n\n# Postprocessing\n\nWe reduce probabilities for cells based on negative signal, this gave us a small boost (~0.003): preds[:,:18] = preds[:,:18] * (1 - preds[:,18]).unsqueeze(-1)\n\n# Tried but didn't work\n\n- Pseudolabels\n- Multi-instance learning (treat each image as a bag of cells, train with bag of cells, predict on single cells). This was the same architecture as shared by Guanshuo Xu in his #8 solution. After analyzing what I did wrong, my conclusion is that I used too many cells per image. After reducing number of cells from 48 to 8, the results for my previously failed experiments improved significantly.\n- Using rectangle-shape images during inference\n- BagNet [paper](https://arxiv.org/pdf/1904.00760.pdf)\n- ...\n",
      "votes": 42
    },
    {
      "id": 1303833,
      "postDate": "2021-05-12T09:15:45.623Z",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> Congrats on your Gold Medal and your great work! 👍</p>",
      "rawMarkdown": "@thedrcat Congrats on your Gold Medal and your great work! 👍",
      "votes": 1,
      "replies": [
        {
          "id": 1304233,
          "postDate": "2021-05-12T13:54:28.297Z",
          "content": "<p>Thank you Jason!</p>",
          "rawMarkdown": "Thank you Jason!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1303691,
      "postDate": "2021-05-12T07:41:49.250Z",
      "content": "<p>Nice approach, thanks for the clarification, and congratulation </p>",
      "rawMarkdown": "Nice approach, thanks for the clarification, and congratulation ",
      "votes": 1,
      "replies": [
        {
          "id": 1304234,
          "postDate": "2021-05-12T13:54:47.897Z",
          "content": "<p>Thank you Salim!</p>",
          "rawMarkdown": "Thank you Salim!"
        }
      ]
    },
    {
      "id": 1303897,
      "postDate": "2021-05-12T10:23:34.743Z",
      "content": "<p>Awesome write up! Thank you so much for all your incredible work Darek! I learned a lot from working with you.</p>",
      "rawMarkdown": "Awesome write up! Thank you so much for all your incredible work Darek! I learned a lot from working with you.",
      "votes": 2,
      "replies": [
        {
          "id": 1304232,
          "postDate": "2021-05-12T13:54:13.277Z",
          "content": "<p>Your experiments pushed us in the direction of Densenet Darien, and you discovered the power of image + cell level blend before it was revealed on the forums. I'm glad we teamed up early, thanks for the collaboration!</p>",
          "rawMarkdown": "Your experiments pushed us in the direction of Densenet Darien, and you discovered the power of image + cell level blend before it was revealed on the forums. I'm glad we teamed up early, thanks for the collaboration!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1303491,
      "postDate": "2021-05-12T05:05:51.667Z",
      "content": "<p>Gridify inference is genius. Also, your notebooks helped me quite a bit in this competition. Thank you so much for sharing your solutions and notebooks.</p>",
      "rawMarkdown": "Gridify inference is genius. Also, your notebooks helped me quite a bit in this competition. Thank you so much for sharing your solutions and notebooks.",
      "votes": 2,
      "replies": [
        {
          "id": 1304235,
          "postDate": "2021-05-12T13:55:28.400Z",
          "content": "<p>I'm glad my public notebooks were helpful! Thank you!!</p>",
          "rawMarkdown": "I'm glad my public notebooks were helpful! Thank you!!"
        }
      ]
    },
    {
      "id": 1303464,
      "postDate": "2021-05-12T04:44:57.167Z",
      "content": "<p>Congratulations, Darek! 😊 And thank you for the great write-up! And thanks for your active participation in forums and notebooks in the first place! For me, your public contributions are a crucial piece of this competition's awesomeness! </p>",
      "rawMarkdown": "Congratulations, Darek! 😊 And thank you for the great write-up! And thanks for your active participation in forums and notebooks in the first place! For me, your public contributions are a crucial piece of this competition's awesomeness! ",
      "votes": 2,
      "replies": [
        {
          "id": 1304236,
          "postDate": "2021-05-12T13:56:49.413Z",
          "content": "<p>Thanks Raman, congratulations to you as well! Also thank you for your sharing, it helped our team a lot as well! True power of Kaggle community 🙏</p>",
          "rawMarkdown": "Thanks Raman, congratulations to you as well! Also thank you for your sharing, it helped our team a lot as well! True power of Kaggle community 🙏"
        }
      ]
    },
    {
      "id": 1303340,
      "postDate": "2021-05-12T03:04:11.873Z",
      "content": "<p>Congratulations!  Kaggle is fascinating.</p>",
      "rawMarkdown": "Congratulations!  Kaggle is fascinating.",
      "votes": 2,
      "replies": [
        {
          "id": 1304242,
          "postDate": "2021-05-12T14:00:28.663Z",
          "content": "<p>Thank you daishu! Your humble posts on the forum were very inspiring for me 🙏 Like <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> famously said, Kaggle is a legal drug 😂I need to go on a detox now and give back to my family ❤️❤️❤️</p>",
          "rawMarkdown": "Thank you daishu! Your humble posts on the forum were very inspiring for me 🙏 Like @cpmpml famously said, Kaggle is a legal drug 😂I need to go on a detox now and give back to my family ❤️❤️❤️"
        }
      ]
    },
    {
      "id": 1303346,
      "postDate": "2021-05-12T03:09:45.590Z",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> Congratulations and Thanks for sharing your approach </p>",
      "rawMarkdown": "@thedrcat Congratulations and Thanks for sharing your approach ",
      "replies": [
        {
          "id": 1304237,
          "postDate": "2021-05-12T13:57:10.257Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> !</p>",
          "rawMarkdown": "Thank you @usharengaraju !"
        }
      ]
    },
    {
      "id": 1937241,
      "postDate": "2022-09-13T11:14:46.683Z",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> Congratulations!</p>",
      "rawMarkdown": "@thedrcat Congratulations!"
    }
  ],
  "comments": [
    {
      "id": 1303833,
      "author_name": "Wang Xing",
      "author_url": "",
      "post_date": "2021-05-12T09:15:45.623000",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> Congrats on your Gold Medal and your great work! 👍</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304233,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-12T13:54:28.297000",
          "content": "<p>Thank you Jason!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1303691,
      "author_name": "Salim Khazem",
      "author_url": "",
      "post_date": "2021-05-12T07:41:49.250000",
      "content": "<p>Nice approach, thanks for the clarification, and congratulation </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304234,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-12T13:54:47.897000",
          "content": "<p>Thank you Salim!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1303897,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-05-12T10:23:34.743000",
      "content": "<p>Awesome write up! Thank you so much for all your incredible work Darek! I learned a lot from working with you.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1304232,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-12T13:54:13.277000",
          "content": "<p>Your experiments pushed us in the direction of Densenet Darien, and you discovered the power of image + cell level blend before it was revealed on the forums. I'm glad we teamed up early, thanks for the collaboration!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1303491,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "2021-05-12T05:05:51.667000",
      "content": "<p>Gridify inference is genius. Also, your notebooks helped me quite a bit in this competition. Thank you so much for sharing your solutions and notebooks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1304235,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-12T13:55:28.400000",
          "content": "<p>I'm glad my public notebooks were helpful! Thank you!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1303464,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-12T04:44:57.167000",
      "content": "<p>Congratulations, Darek! 😊 And thank you for the great write-up! And thanks for your active participation in forums and notebooks in the first place! For me, your public contributions are a crucial piece of this competition's awesomeness! </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1304236,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-12T13:56:49.413000",
          "content": "<p>Thanks Raman, congratulations to you as well! Also thank you for your sharing, it helped our team a lot as well! True power of Kaggle community 🙏</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1303340,
      "author_name": "Correlation",
      "author_url": "",
      "post_date": "2021-05-12T03:04:11.873000",
      "content": "<p>Congratulations!  Kaggle is fascinating.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1304242,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-12T14:00:28.663000",
          "content": "<p>Thank you daishu! Your humble posts on the forum were very inspiring for me 🙏 Like <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> famously said, Kaggle is a legal drug 😂I need to go on a detox now and give back to my family ❤️❤️❤️</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1303346,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-12T03:09:45.590000",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> Congratulations and Thanks for sharing your approach </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1304237,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-12T13:57:10.257000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1937241,
      "author_name": "Him@nshu Kandpal",
      "author_url": "",
      "post_date": "2022-09-13T11:14:46.683000",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> Congratulations!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1303303": "I put all of my heart into this competition, and I need to start by thanking my wife for putting up with me in the last couple of months - buying a DL rig, waking up or going to sleep at crazy hours, putting my rig on car front seat on a weekend trip, and many more… all while going through Covid (fortunately mildly) and trying to entertain our Kaggle-famous 4yo future scientist while quarantined at home :) \n\nThanks to the team members - @dschettler8845 and @felipebihaiek for partnering early on, experimenting together and discussing hundreds of ideas, @scusywxy and @zehuigong for the partnership and your lead to figure out the engineering details so that our combined solution finally worked. We had diverse approaches which gave us a boost (2 minutes late for the money zone!). I will only share my contributions here, a more comprehensive summary of our total approach will follow. \n\nThanks to Kaggle and the hosts for this competition!\n\n# Model or Data-Centric Approach?\n\nWe started the competition with single-cell classification combined with HPA CellSegmentator, and I believed initially this will be the key to the competition. I planned a data centric approach, including pseudo labelling (it gave me a small boost early on) and iterative fixing of most-confused labels. I wanted to use a fastai widget for that, but afaik it doesn’t support multilabel classification unfortunately, so experimented with some other tools, but finally gave up when we realized labelling the cells is not very straightforward. \n\nWe had accurate data to train image level models, so the question was how we can use those models to predict cells? After many experiments, we figured out two approaches that led us first to .549 public lb score, and gave us a 0.02 boost after blending with @scusywxy and @zehuigong models. \n\n# GAPMASK Inference\n\nWe train a regular image-level model. At inference, we modify the model architecture so that it takes two inputs - image and cell-mask. The image goes through the network alone until the GAP layer. At that point, we do element wise multiplication of the image activations and cell mask. From that point, an image gets expanded into a batch of single-cell images, and we apply the model head to that entire batch. \n\n![Gap-Mask-1](https://pbs.twimg.com/media/E2SRV0qWQAM7kju?format=jpg&name=medium)\n\nWe had two variations of this architecture - one for global average and max concat pooling, one for attention pooling. \n\nThis is the code for our densenet model: \n\n```\n    def forward(self, images, masks):\n        …\n        e5 = F.relu(e5,inplace=True)\n        # GAP MASK STARTS HERE:\n        e5 = F.interpolate(e5, scale_factor=2, mode='bilinear') # increase grid size to 48x48\n        x = e5.permute(0,2,3,1) # [1, 48, 48, 1024]\n        x = x.squeeze() # [48, 48, 1024]\n        msk = masks.squeeze() # [15, 48, 48]\n        res = msk[...,None] * x[None,...] # [15, 48, 48, 1024]\n        res = res.permute(0,3,1,2)\n        # NOW APPLY THE MODEL HEAD WITH THE MASK DIMENSION AS THE BATCH DIMENSION\n        x = torch.cat((nn.AdaptiveAvgPool2d(1)(res), nn.AdaptiveMaxPool2d(1)(res)), dim=1)\n        x = x.view(x.size(0), -1)\n\t…\n        x = self.logit(x)\n        return x\n\n```\nAnd this is the implementation for inception with attention: \n![Gap-Mask-2](https://pbs.twimg.com/media/E2SRK5IWUAMli02?format=jpg&name=medium)\n\n```\n    def forward(self, x, masks):\n        …\n\t# GAP MASK STARTS HERE (dimensions reflect an image with 15 cells):\n        logits = self.last_linear(features_b)\n        logits = F.interpolate(logits, scale_factor=3, mode='nearest')\n        logits_attention = self.attention(features_b)\n        logits_attention = F.interpolate(logits_attention, scale_factor=3, mode='nearest')\n        logits_attention = logits_attention.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3)\n        attention = F.softmax(logits_attention, dim=2)\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3, self.attention_size * 3)\n        attention = attention.permute(0,2,3,1) # [1,66,66,19])\n        msk = masks.squeeze() # [15,66,66]       \n        attention = msk[...,None] * attention # [15,66,66,19]\n        attention = attention.permute(0,3,1,2) # [15,19,66,66]\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3) #[15, 19, 66*66]\n        attention = attention / (attention.sum(2).unsqueeze(-1) + 1e-7)\n        attention = attention.view(-1, self.num_classes, self.attention_size * 3, self.attention_size * 3) #[15,19,66,66]\n        logits = logits * attention\n        return logits.view(-1, self.num_classes, self.attention_size * 3 * self.attention_size * 3).sum(2).view(-1, self.num_classes) # torch.Size([1, 19])\n```\n\n# GRIDIFY Inference\n\nWhen we started with single cell tiles, we resized each cell to the same size, e.g. 128x128. The problem is that this changed cell resolutions, some\ngot shrunk and some expanded. Inspired by [this paper](https://arxiv.org/pdf/1906.06423.pdf), I looked for an approach that would line up the features model sees while training on full images, with features seen by the model when running inference on single cells. \n\nThe initial gridify approach was to take a single cell crop, copy-paste that crop with multiple augmentations into a 2048x2048 template, and then resize in the same way as during training. Finally, we converged on the following approach: \n- Take 4x 512x512 crop from full image (independent of size) from each corner of the single cell\n- Put those 4 crops into a single 1024x1024 image\n- Resize to maintain same resolution as training. Eg. for the model trained on 768 size (from 2048), we resize into 384x384 cell tile\n\n![gridify](https://pbs.twimg.com/media/E1J0N0ZXEAYyffs?format=jpg&name=large)\n\n# Models\n\nWe used three models from HPA 2018 winning solutions: \n- Densenet trained on size 768 from bestfitting\n- Densenet trained on size 1536 from bestfitting \n- Inceptionv3 trained on size 1024 from pudae\n\nAll models were fine-tuned with original loss functions (focal-lovasz-logloss for densenet, focal for inception), starting with the 2018 weights. We froze the model head for 1 epoch, then trained for 3-6 epochs on train + public data (excl. classes 0 and 16) with cosine annealing. \n\n# Validation\n\nDue to potential leakage from similar images (same plate) across folds, we run all images through metric learning model (from bestfitting’s 2018 solution) and cluster similar images together with UMAP/DBSCAN. We treat the clusters as groups and divide data in folds based on group stratified multilabel approach. \n\n# Postprocessing\n\nWe reduce probabilities for cells based on negative signal, this gave us a small boost (~0.003): preds[:,:18] = preds[:,:18] * (1 - preds[:,18]).unsqueeze(-1)\n\n# Tried but didn't work\n\n- Pseudolabels\n- Multi-instance learning (treat each image as a bag of cells, train with bag of cells, predict on single cells). This was the same architecture as shared by Guanshuo Xu in his #8 solution. After analyzing what I did wrong, my conclusion is that I used too many cells per image. After reducing number of cells from 48 to 8, the results for my previously failed experiments improved significantly.\n- Using rectangle-shape images during inference\n- BagNet [paper](https://arxiv.org/pdf/1904.00760.pdf)\n- ...\n",
    "1303833": "@thedrcat Congrats on your Gold Medal and your great work! 👍",
    "1303691": "Nice approach, thanks for the clarification, and congratulation ",
    "1303897": "Awesome write up! Thank you so much for all your incredible work Darek! I learned a lot from working with you.",
    "1303491": "Gridify inference is genius. Also, your notebooks helped me quite a bit in this competition. Thank you so much for sharing your solutions and notebooks.",
    "1303464": "Congratulations, Darek! 😊 And thank you for the great write-up! And thanks for your active participation in forums and notebooks in the first place! For me, your public contributions are a crucial piece of this competition's awesomeness! ",
    "1303340": "Congratulations!  Kaggle is fascinating.",
    "1303346": "@thedrcat Congratulations and Thanks for sharing your approach ",
    "1937241": "@thedrcat Congratulations!"
  }
}