{
  "id": 238898,
  "title": "3rd place Dieter part",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238898",
  "author_name": "Dieter",
  "post_date": "2021-05-13T19:46:15.804000",
  "votes": 41,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks to Kaggle and hosts for this very interesting competition with a tricky setup. Also huge props to my team mates <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> and <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> who carried me to 3rd place. Since we worked quite independently on the models we thought it helps readability to split the work in 3 posts.  In the following, I want to give a rough overview of my part of our solution.</p>\n<h3>TLDR</h3>\n<p>I did an ensemble of several ideas to tackle the weak label problem. The first model type works on cropped single cells which are weighted using trainable attention to derive a bag-of-cell prediction on which the loss wrt to the weak image label is calculated. The second model type is on image level. Here cell masks derived using HPA Cell segmentor are used to extract local features and derive a single cell feature vector. All single cell feature vectors are weighted to derive with an image level label. The third model type is an image level model which does not account for single cells, but helps the overall prediction ranking.</p>\n<h3>Data setup &amp; CV</h3>\n<p>I settled early on the hyperparameters for the HPA cell segmenter, which I slightly modified to run faster. I generated masks for all images and used those and fixed single cell ids through the competition. For cross-validation I identified clusters of similar/ duplicate images and used GroupKfold to create a robust validation scheme. I used Public HPA data and images from the first HPA competition, which were not included here.</p>\n<h3>Models</h3>\n<h4>Single Cell Model</h4>\n<p>I tried different ways how to create a bag-of-cell model to leverage the weak labels. The following worked very well and is quite similar to <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> solution:</p>\n<ol>\n<li>Use HPA Cell Segmenter to crop single cells</li>\n<li>Feed batches x n_cells into the model</li>\n<li>Use attention pooling to train weighting of cells within a bag and derive bag-prediction</li>\n<li>BCE loss between image label and bag-predictiction</li>\n</ol>\n<p>The following gives a brief illustration (e.g. bs =1):</p>\n<p><img src=\"https://i.imgur.com/ilk8KrD.png\" alt=\"model1\"></p>\n<p>I want to note that the single cell model turned out to be especially robust on private LB as it can better handle high SCV.</p>\n<h4>Image model with aggregated local features</h4>\n<p>The second architecture uses the complete image but selects local features using a previously generated cell mask </p>\n<p>Cell mask is resized to 16x16 to map local feature vectors of backbone output (which has size (bs,1024,16,16)) to single cells and hence create an embedding vector for each cell. I weight the embedding vectors using trainable attention to derive the image level label. </p>\n<p><img src=\"https://i.imgur.com/vOCEPqD.png\" alt=\"model2\"></p>\n<p>I used backbones SE-ResNeXt26 and SE-ResNeXt101 from timm repository for these 2 model types.</p>\n<p>The third model is an efficientnet-b7 trained on image level and predicts the same label for each cell within an image.</p>\n<h4>Compressing into one</h4>\n<p>I want to note that by generating single cell out-of-fold predictions using this ensemble and training a single cell model with that results in a good but lightweight single model (Public LB 0.524)</p>\n<p>Thanks for reading. Happy to answer any questions.</p>",
  "messages": [
    {
      "id": 1306484,
      "postDate": "2021-05-13T19:46:15.803Z",
      "content": "<p>Thanks to Kaggle and hosts for this very interesting competition with a tricky setup. Also huge props to my team mates <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> and <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> who carried me to 3rd place. Since we worked quite independently on the models we thought it helps readability to split the work in 3 posts.  In the following, I want to give a rough overview of my part of our solution.</p>\n<h3>TLDR</h3>\n<p>I did an ensemble of several ideas to tackle the weak label problem. The first model type works on cropped single cells which are weighted using trainable attention to derive a bag-of-cell prediction on which the loss wrt to the weak image label is calculated. The second model type is on image level. Here cell masks derived using HPA Cell segmentor are used to extract local features and derive a single cell feature vector. All single cell feature vectors are weighted to derive with an image level label. The third model type is an image level model which does not account for single cells, but helps the overall prediction ranking.</p>\n<h3>Data setup &amp; CV</h3>\n<p>I settled early on the hyperparameters for the HPA cell segmenter, which I slightly modified to run faster. I generated masks for all images and used those and fixed single cell ids through the competition. For cross-validation I identified clusters of similar/ duplicate images and used GroupKfold to create a robust validation scheme. I used Public HPA data and images from the first HPA competition, which were not included here.</p>\n<h3>Models</h3>\n<h4>Single Cell Model</h4>\n<p>I tried different ways how to create a bag-of-cell model to leverage the weak labels. The following worked very well and is quite similar to <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> solution:</p>\n<ol>\n<li>Use HPA Cell Segmenter to crop single cells</li>\n<li>Feed batches x n_cells into the model</li>\n<li>Use attention pooling to train weighting of cells within a bag and derive bag-prediction</li>\n<li>BCE loss between image label and bag-predictiction</li>\n</ol>\n<p>The following gives a brief illustration (e.g. bs =1):</p>\n<p><img src=\"https://i.imgur.com/ilk8KrD.png\" alt=\"model1\"></p>\n<p>I want to note that the single cell model turned out to be especially robust on private LB as it can better handle high SCV.</p>\n<h4>Image model with aggregated local features</h4>\n<p>The second architecture uses the complete image but selects local features using a previously generated cell mask </p>\n<p>Cell mask is resized to 16x16 to map local feature vectors of backbone output (which has size (bs,1024,16,16)) to single cells and hence create an embedding vector for each cell. I weight the embedding vectors using trainable attention to derive the image level label. </p>\n<p><img src=\"https://i.imgur.com/vOCEPqD.png\" alt=\"model2\"></p>\n<p>I used backbones SE-ResNeXt26 and SE-ResNeXt101 from timm repository for these 2 model types.</p>\n<p>The third model is an efficientnet-b7 trained on image level and predicts the same label for each cell within an image.</p>\n<h4>Compressing into one</h4>\n<p>I want to note that by generating single cell out-of-fold predictions using this ensemble and training a single cell model with that results in a good but lightweight single model (Public LB 0.524)</p>\n<p>Thanks for reading. Happy to answer any questions.</p>",
      "rawMarkdown": "Thanks to Kaggle and hosts for this very interesting competition with a tricky setup. Also huge props to my team mates @zfturbo and @mpware who carried me to 3rd place. Since we worked quite independently on the models we thought it helps readability to split the work in 3 posts.  In the following, I want to give a rough overview of my part of our solution.\n\n### TLDR\nI did an ensemble of several ideas to tackle the weak label problem. The first model type works on cropped single cells which are weighted using trainable attention to derive a bag-of-cell prediction on which the loss wrt to the weak image label is calculated. The second model type is on image level. Here cell masks derived using HPA Cell segmentor are used to extract local features and derive a single cell feature vector. All single cell feature vectors are weighted to derive with an image level label. The third model type is an image level model which does not account for single cells, but helps the overall prediction ranking.\n\n### Data setup & CV\nI settled early on the hyperparameters for the HPA cell segmenter, which I slightly modified to run faster. I generated masks for all images and used those and fixed single cell ids through the competition. For cross-validation I identified clusters of similar/ duplicate images and used GroupKfold to create a robust validation scheme. I used Public HPA data and images from the first HPA competition, which were not included here.\n\n### Models\n\n\n#### Single Cell Model\n\nI tried different ways how to create a bag-of-cell model to leverage the weak labels. The following worked very well and is quite similar to @wowfattie solution:\n\n1. Use HPA Cell Segmenter to crop single cells\n2. Feed batches x n_cells into the model\n3. Use attention pooling to train weighting of cells within a bag and derive bag-prediction\n4. BCE loss between image label and bag-predictiction\n\nThe following gives a brief illustration (e.g. bs =1):\n\n![model1](https://i.imgur.com/ilk8KrD.png)\n\nI want to note that the single cell model turned out to be especially robust on private LB as it can better handle high SCV.\n\n#### Image model with aggregated local features\nThe second architecture uses the complete image but selects local features using a previously generated cell mask \n\nCell mask is resized to 16x16 to map local feature vectors of backbone output (which has size (bs,1024,16,16)) to single cells and hence create an embedding vector for each cell. I weight the embedding vectors using trainable attention to derive the image level label. \n\n![model2](https://i.imgur.com/vOCEPqD.png)\n\nI used backbones SE-ResNeXt26 and SE-ResNeXt101 from timm repository for these 2 model types.\n\nThe third model is an efficientnet-b7 trained on image level and predicts the same label for each cell within an image.\n\n\n#### Compressing into one\n\nI want to note that by generating single cell out-of-fold predictions using this ensemble and training a single cell model with that results in a good but lightweight single model (Public LB 0.524)\n\n\nThanks for reading. Happy to answer any questions.\n",
      "votes": 41
    },
    {
      "id": 1308074,
      "postDate": "2021-05-15T00:42:30.473Z",
      "content": "<p>Congratulations Dieter and team. Well done!</p>",
      "rawMarkdown": "Congratulations Dieter and team. Well done!",
      "votes": 1
    },
    {
      "id": 1308351,
      "postDate": "2021-05-15T06:21:32.337Z",
      "content": "<p>Congratulations on your 3rd position,<br>\nCan you share any code or paper referencing the \" attention pooling\" layer you have used? (I never heard of this type of pooling before)<br>\nThanks</p>",
      "rawMarkdown": "Congratulations on your 3rd position,\nCan you share any code or paper referencing the \" attention pooling\" layer you have used? (I never heard of this type of pooling before)\nThanks",
      "replies": [
        {
          "id": 1312543,
          "postDate": "2021-05-18T05:31:07.937Z",
          "content": "<pre><code>from torch import nn \nimport torch\n\nclass AttentionPool(nn.Module):\n    def __init__(self, in_units, hidden = 512):\n        super(AttentionPool, self).__init__()\n\n        self.att = nn.Sequential(nn.Linear(in_units, hidden),\n                                  nn.ReLU(),\n                                  nn.Linear(hidden, 1))\n\n\n    def forward(self, x):\n\n        # x has shape (batch_size, n_cell, feats)\n        att_weights = torch.softmax(self.att(x),dim=1)\n        x_pooled = (x * att_weights).sum(1)\n\n        return x_pooled\n</code></pre>\n<p>What I used here is a trainable layer which outputs weights for each cell. You then pool by weighted sum.</p>",
          "rawMarkdown": "```\nfrom torch import nn \nimport torch\n\nclass AttentionPool(nn.Module):\n    def __init__(self, in_units, hidden = 512):\n        super(AttentionPool, self).__init__()\n        \n        self.att = nn.Sequential(nn.Linear(in_units, hidden),\n                                  nn.ReLU(),\n                                  nn.Linear(hidden, 1))\n        \n\n    def forward(self, x):\n\n        # x has shape (batch_size, n_cell, feats)\n        att_weights = torch.softmax(self.att(x),dim=1)\n        x_pooled = (x * att_weights).sum(1)\n\n        return x_pooled\n\n```\n\nWhat I used here is a trainable layer which outputs weights for each cell. You then pool by weighted sum.",
          "votes": 6
        }
      ]
    },
    {
      "id": 1306909,
      "postDate": "2021-05-14T06:25:05.190Z",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> , congrats on your amazing performance! Mainly, thank you so much for this write-up! Reading about your image model with aggregated local features made my day! Such an elegant idea! </p>\n<p>*p.s. the bag-of-cells model is amazing, I was blown away by it when reading the <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> 's solution. <br>\nMay I have a quick technical question, please? For images containing less than n_cells, did you somehow mask out the padded (empty) cells in the attention pool, or what was your approach?<br>\nI'd be also grateful if you can briefly comment on handling different numbers of cells when pooling by attention in the image model with aggregated local features. Thank you!</p>",
      "rawMarkdown": "@christofhenkel , congrats on your amazing performance! Mainly, thank you so much for this write-up! Reading about your image model with aggregated local features made my day! Such an elegant idea! \n\n*p.s. the bag-of-cells model is amazing, I was blown away by it when reading the @wowfattie 's solution. \nMay I have a quick technical question, please? For images containing less than n_cells, did you somehow mask out the padded (empty) cells in the attention pool, or what was your approach?\nI'd be also grateful if you can briefly comment on handling different numbers of cells when pooling by attention in the image model with aggregated local features. Thank you!",
      "replies": [
        {
          "id": 1306925,
          "postDate": "2021-05-14T06:35:39.460Z",
          "content": "<blockquote>\n  <p>handling different numbers of cells when pooling by attention in the image model with aggregated local features. </p>\n</blockquote>\n<p>I predefine a max_len and then take randomly max_len number of cell vectors. If an image has less than max_len, I pad with zero-vectors</p>",
          "rawMarkdown": "> handling different numbers of cells when pooling by attention in the image model with aggregated local features. \n\nI predefine a max_len and then take randomly max_len number of cell vectors. If an image has less than max_len, I pad with zero-vectors",
          "votes": 1
        },
        {
          "id": 1306944,
          "postDate": "2021-05-14T06:50:20.393Z",
          "content": "<p>Thank you for the prompt detailed reply! </p>\n<p>It's a privilege to learn from you.</p>",
          "rawMarkdown": "Thank you for the prompt detailed reply! \n\nIt's a privilege to learn from you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1307362,
      "postDate": "2021-05-14T12:02:39.490Z",
      "content": "<p>Congrats! Thank you for sharing.</p>",
      "rawMarkdown": "Congrats! Thank you for sharing."
    }
  ],
  "comments": [
    {
      "id": 1308074,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-05-15T00:42:30.473000",
      "content": "<p>Congratulations Dieter and team. Well done!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1308351,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-05-15T06:21:32.337000",
      "content": "<p>Congratulations on your 3rd position,<br>\nCan you share any code or paper referencing the \" attention pooling\" layer you have used? (I never heard of this type of pooling before)<br>\nThanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1312543,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-05-18T05:31:07.937000",
          "content": "<pre><code>from torch import nn \nimport torch\n\nclass AttentionPool(nn.Module):\n    def __init__(self, in_units, hidden = 512):\n        super(AttentionPool, self).__init__()\n\n        self.att = nn.Sequential(nn.Linear(in_units, hidden),\n                                  nn.ReLU(),\n                                  nn.Linear(hidden, 1))\n\n\n    def forward(self, x):\n\n        # x has shape (batch_size, n_cell, feats)\n        att_weights = torch.softmax(self.att(x),dim=1)\n        x_pooled = (x * att_weights).sum(1)\n\n        return x_pooled\n</code></pre>\n<p>What I used here is a trainable layer which outputs weights for each cell. You then pool by weighted sum.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1306909,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-14T06:25:05.190000",
      "content": "<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> , congrats on your amazing performance! Mainly, thank you so much for this write-up! Reading about your image model with aggregated local features made my day! Such an elegant idea! </p>\n<p>*p.s. the bag-of-cells model is amazing, I was blown away by it when reading the <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> 's solution. <br>\nMay I have a quick technical question, please? For images containing less than n_cells, did you somehow mask out the padded (empty) cells in the attention pool, or what was your approach?<br>\nI'd be also grateful if you can briefly comment on handling different numbers of cells when pooling by attention in the image model with aggregated local features. Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1306925,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-05-14T06:35:39.460000",
          "content": "<blockquote>\n  <p>handling different numbers of cells when pooling by attention in the image model with aggregated local features. </p>\n</blockquote>\n<p>I predefine a max_len and then take randomly max_len number of cell vectors. If an image has less than max_len, I pad with zero-vectors</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1306944,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-14T06:50:20.393000",
          "content": "<p>Thank you for the prompt detailed reply! </p>\n<p>It's a privilege to learn from you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1307362,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-14T12:02:39.490000",
      "content": "<p>Congrats! Thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1306484": "Thanks to Kaggle and hosts for this very interesting competition with a tricky setup. Also huge props to my team mates @zfturbo and @mpware who carried me to 3rd place. Since we worked quite independently on the models we thought it helps readability to split the work in 3 posts.  In the following, I want to give a rough overview of my part of our solution.\n\n### TLDR\nI did an ensemble of several ideas to tackle the weak label problem. The first model type works on cropped single cells which are weighted using trainable attention to derive a bag-of-cell prediction on which the loss wrt to the weak image label is calculated. The second model type is on image level. Here cell masks derived using HPA Cell segmentor are used to extract local features and derive a single cell feature vector. All single cell feature vectors are weighted to derive with an image level label. The third model type is an image level model which does not account for single cells, but helps the overall prediction ranking.\n\n### Data setup & CV\nI settled early on the hyperparameters for the HPA cell segmenter, which I slightly modified to run faster. I generated masks for all images and used those and fixed single cell ids through the competition. For cross-validation I identified clusters of similar/ duplicate images and used GroupKfold to create a robust validation scheme. I used Public HPA data and images from the first HPA competition, which were not included here.\n\n### Models\n\n\n#### Single Cell Model\n\nI tried different ways how to create a bag-of-cell model to leverage the weak labels. The following worked very well and is quite similar to @wowfattie solution:\n\n1. Use HPA Cell Segmenter to crop single cells\n2. Feed batches x n_cells into the model\n3. Use attention pooling to train weighting of cells within a bag and derive bag-prediction\n4. BCE loss between image label and bag-predictiction\n\nThe following gives a brief illustration (e.g. bs =1):\n\n![model1](https://i.imgur.com/ilk8KrD.png)\n\nI want to note that the single cell model turned out to be especially robust on private LB as it can better handle high SCV.\n\n#### Image model with aggregated local features\nThe second architecture uses the complete image but selects local features using a previously generated cell mask \n\nCell mask is resized to 16x16 to map local feature vectors of backbone output (which has size (bs,1024,16,16)) to single cells and hence create an embedding vector for each cell. I weight the embedding vectors using trainable attention to derive the image level label. \n\n![model2](https://i.imgur.com/vOCEPqD.png)\n\nI used backbones SE-ResNeXt26 and SE-ResNeXt101 from timm repository for these 2 model types.\n\nThe third model is an efficientnet-b7 trained on image level and predicts the same label for each cell within an image.\n\n\n#### Compressing into one\n\nI want to note that by generating single cell out-of-fold predictions using this ensemble and training a single cell model with that results in a good but lightweight single model (Public LB 0.524)\n\n\nThanks for reading. Happy to answer any questions.\n",
    "1308074": "Congratulations Dieter and team. Well done!",
    "1308351": "Congratulations on your 3rd position,\nCan you share any code or paper referencing the \" attention pooling\" layer you have used? (I never heard of this type of pooling before)\nThanks",
    "1306909": "@christofhenkel , congrats on your amazing performance! Mainly, thank you so much for this write-up! Reading about your image model with aggregated local features made my day! Such an elegant idea! \n\n*p.s. the bag-of-cells model is amazing, I was blown away by it when reading the @wowfattie 's solution. \nMay I have a quick technical question, please? For images containing less than n_cells, did you somehow mask out the padded (empty) cells in the attention pool, or what was your approach?\nI'd be also grateful if you can briefly comment on handling different numbers of cells when pooling by attention in the image model with aggregated local features. Thank you!",
    "1307362": "Congrats! Thank you for sharing."
  }
}