{
  "id": 238624,
  "title": "41st Place Solution: Cell-based RoI Pooling + Transformer Encoder",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238624",
  "author_name": "Martin Chobanyan",
  "post_date": "2021-05-12T20:25:32.684000",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey fellow Kaggle competitors! </p>\n<p>I had a lot of fun with this unique competition and I'm looking forward to reading your own solutions to this problem. Here is an overview of my final approach + link to my code (open images in a new tab for higher res):</p>\n<p><a href=\"https://github.com/martin-chobanyan/hpa-single-cell\" target=\"_blank\">Github Repo</a></p>\n<p><img src=\"https://raw.githubusercontent.com/martin-chobanyan/hpa-single-cell/main/resources/cell-transformer-overview.png\" alt=\"cell transformer\"></p>\n<p>First off, in all of my models I used the pre-trained backbone CNN of <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> 's winning DenseNet-121 models from the previous competition (3 different pre-trained versions). I stacked all of the color stains into 4-channel images and resized them to 1536x1536 and fed them through the backbone CNN, resulting in 1024 feature maps.</p>\n<p>I then extracted a feature vector for each cell in the image by performing Region-of-Interest (RoI) pooling over the feature maps using the cell segmentation masks from the HPA Cell Segmentator. The pooling layer consisted of a concatenation of both avg-pool and max-pool features (resulting in a 2048-dim vector per cell). I also applied an adaptive average pool over the cell masks to reduce them to 8x8, which I then flattened. These downsampled, flatten masks served as position encoding for the cell regions (see diagram below):</p>\n<p><img src=\"https://raw.githubusercontent.com/martin-chobanyan/hpa-single-cell/main/resources/cell-roi-pool-overview.png\" alt=\"roi pool\"></p>\n<p>At this stage, we have a sequence of cell vectors, which allows us to explore the use of a Transformer model. The cell feature vectors are passed through a feed-forward layer which maps them to a 1024-dim embedding. Then, four encoder layers are applied as they were defined in the original Transformer paper. A final feed forward layer then maps each cell to an 18-dimensional logit.</p>\n<p>To predict the image-level labels, a log-sum-exponential layer is applied which is defined as f(x) = (1/r) * ln(mean(e^(r*x))) with r=5 in our case. This serves as a mechanism to interpolate between average pooling and max pooling with different values for r.</p>\n<p>For training I used Focal Loss with sigmoid activation and flip/rotate/crop augmentations. During inference, I extracted the cell-level logits and applied a sigmoid activation before their reduction to image-level labels. The final class prediction for a cell was averaged across all three models (three different pre-trained backbones).</p>\n<p>I also explored the use of <a href=\"https://arxiv.org/abs/1804.00880\" target=\"_blank\">Peak Response Maps</a> as a CAM based approach to the problem. Though the localizations looked good, the model was a bit too sensitive to classes which were not present in the image.</p>\n<p>Things I would have tried with more time:</p>\n<ul>\n<li>Ensembling with a cell-level classifier</li>\n<li>Explore more ways of merging the peak-response-map results with the cell transformer model</li>\n<li>Domain adaptation to the public test set (since we know the test set has higher variability in protein locations within a given image compared to the training set).</li>\n</ul>",
  "messages": [
    {
      "id": 1304735,
      "postDate": "2021-05-12T20:25:32.683Z",
      "content": "<p>Hey fellow Kaggle competitors! </p>\n<p>I had a lot of fun with this unique competition and I'm looking forward to reading your own solutions to this problem. Here is an overview of my final approach + link to my code (open images in a new tab for higher res):</p>\n<p><a href=\"https://github.com/martin-chobanyan/hpa-single-cell\" target=\"_blank\">Github Repo</a></p>\n<p><img src=\"https://raw.githubusercontent.com/martin-chobanyan/hpa-single-cell/main/resources/cell-transformer-overview.png\" alt=\"cell transformer\"></p>\n<p>First off, in all of my models I used the pre-trained backbone CNN of <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> 's winning DenseNet-121 models from the previous competition (3 different pre-trained versions). I stacked all of the color stains into 4-channel images and resized them to 1536x1536 and fed them through the backbone CNN, resulting in 1024 feature maps.</p>\n<p>I then extracted a feature vector for each cell in the image by performing Region-of-Interest (RoI) pooling over the feature maps using the cell segmentation masks from the HPA Cell Segmentator. The pooling layer consisted of a concatenation of both avg-pool and max-pool features (resulting in a 2048-dim vector per cell). I also applied an adaptive average pool over the cell masks to reduce them to 8x8, which I then flattened. These downsampled, flatten masks served as position encoding for the cell regions (see diagram below):</p>\n<p><img src=\"https://raw.githubusercontent.com/martin-chobanyan/hpa-single-cell/main/resources/cell-roi-pool-overview.png\" alt=\"roi pool\"></p>\n<p>At this stage, we have a sequence of cell vectors, which allows us to explore the use of a Transformer model. The cell feature vectors are passed through a feed-forward layer which maps them to a 1024-dim embedding. Then, four encoder layers are applied as they were defined in the original Transformer paper. A final feed forward layer then maps each cell to an 18-dimensional logit.</p>\n<p>To predict the image-level labels, a log-sum-exponential layer is applied which is defined as f(x) = (1/r) * ln(mean(e^(r*x))) with r=5 in our case. This serves as a mechanism to interpolate between average pooling and max pooling with different values for r.</p>\n<p>For training I used Focal Loss with sigmoid activation and flip/rotate/crop augmentations. During inference, I extracted the cell-level logits and applied a sigmoid activation before their reduction to image-level labels. The final class prediction for a cell was averaged across all three models (three different pre-trained backbones).</p>\n<p>I also explored the use of <a href=\"https://arxiv.org/abs/1804.00880\" target=\"_blank\">Peak Response Maps</a> as a CAM based approach to the problem. Though the localizations looked good, the model was a bit too sensitive to classes which were not present in the image.</p>\n<p>Things I would have tried with more time:</p>\n<ul>\n<li>Ensembling with a cell-level classifier</li>\n<li>Explore more ways of merging the peak-response-map results with the cell transformer model</li>\n<li>Domain adaptation to the public test set (since we know the test set has higher variability in protein locations within a given image compared to the training set).</li>\n</ul>",
      "rawMarkdown": "Hey fellow Kaggle competitors! \n\nI had a lot of fun with this unique competition and I'm looking forward to reading your own solutions to this problem. Here is an overview of my final approach + link to my code (open images in a new tab for higher res):\n\n[Github Repo](https://github.com/martin-chobanyan/hpa-single-cell)\n\n![cell transformer](https://raw.githubusercontent.com/martin-chobanyan/hpa-single-cell/main/resources/cell-transformer-overview.png)\n\nFirst off, in all of my models I used the pre-trained backbone CNN of @bestfitting 's winning DenseNet-121 models from the previous competition (3 different pre-trained versions). I stacked all of the color stains into 4-channel images and resized them to 1536x1536 and fed them through the backbone CNN, resulting in 1024 feature maps.\n\nI then extracted a feature vector for each cell in the image by performing Region-of-Interest (RoI) pooling over the feature maps using the cell segmentation masks from the HPA Cell Segmentator. The pooling layer consisted of a concatenation of both avg-pool and max-pool features (resulting in a 2048-dim vector per cell). I also applied an adaptive average pool over the cell masks to reduce them to 8x8, which I then flattened. These downsampled, flatten masks served as position encoding for the cell regions (see diagram below):\n\n![roi pool](https://raw.githubusercontent.com/martin-chobanyan/hpa-single-cell/main/resources/cell-roi-pool-overview.png)\n\nAt this stage, we have a sequence of cell vectors, which allows us to explore the use of a Transformer model. The cell feature vectors are passed through a feed-forward layer which maps them to a 1024-dim embedding. Then, four encoder layers are applied as they were defined in the original Transformer paper. A final feed forward layer then maps each cell to an 18-dimensional logit.\n\nTo predict the image-level labels, a log-sum-exponential layer is applied which is defined as f(x) = (1/r) * ln(mean(e^(r*x))) with r=5 in our case. This serves as a mechanism to interpolate between average pooling and max pooling with different values for r.\n\nFor training I used Focal Loss with sigmoid activation and flip/rotate/crop augmentations. During inference, I extracted the cell-level logits and applied a sigmoid activation before their reduction to image-level labels. The final class prediction for a cell was averaged across all three models (three different pre-trained backbones).\n\nI also explored the use of [Peak Response Maps](https://arxiv.org/abs/1804.00880) as a CAM based approach to the problem. Though the localizations looked good, the model was a bit too sensitive to classes which were not present in the image.\n\nThings I would have tried with more time:\n\n- Ensembling with a cell-level classifier\n- Explore more ways of merging the peak-response-map results with the cell transformer model\n- Domain adaptation to the public test set (since we know the test set has higher variability in protein locations within a given image compared to the training set).",
      "votes": 11
    },
    {
      "id": 1309793,
      "postDate": "2021-05-16T09:57:34.680Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mch123\" target=\"_blank\">@mch123</a> <br>\nI'm a newbie, I want to learn more and practice your source code, do you have any version above Kaggle notebook. If yes I hope you can share it so I can be more accessible. many thanks.</p>",
      "rawMarkdown": "Hi @mch123 \nI'm a newbie, I want to learn more and practice your source code, do you have any version above Kaggle notebook. If yes I hope you can share it so I can be more accessible. many thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 1310735,
          "postDate": "2021-05-16T22:34:29.927Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/nguyenhuudat\" target=\"_blank\">@nguyenhuudat</a>,</p>\n<p>I just made my submission kernel public: <a href=\"https://www.kaggle.com/mch123/ensemble-cell-transformer-2\" target=\"_blank\">https://www.kaggle.com/mch123/ensemble-cell-transformer-2</a></p>\n<p>You'll also have to use the model states in this dataset I made: <a href=\"https://www.kaggle.com/mch123/hparoimodels\" target=\"_blank\">https://www.kaggle.com/mch123/hparoimodels</a></p>",
          "rawMarkdown": "Hi @nguyenhuudat,\n\nI just made my submission kernel public: https://www.kaggle.com/mch123/ensemble-cell-transformer-2\n\nYou'll also have to use the model states in this dataset I made: https://www.kaggle.com/mch123/hparoimodels"
        }
      ]
    },
    {
      "id": 1305022,
      "postDate": "2021-05-13T04:25:12.803Z",
      "content": "<p><a href=\"https://www.kaggle.com/mch123\" target=\"_blank\">@mch123</a> Congratulations  and Thanks for sharing the approach</p>",
      "rawMarkdown": "@mch123 Congratulations  and Thanks for sharing the approach",
      "votes": 1,
      "replies": [
        {
          "id": 1307633,
          "postDate": "2021-05-14T15:09:47.800Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1309793,
      "author_name": "Dat Nguyen",
      "author_url": "",
      "post_date": "2021-05-16T09:57:34.680000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mch123\" target=\"_blank\">@mch123</a> <br>\nI'm a newbie, I want to learn more and practice your source code, do you have any version above Kaggle notebook. If yes I hope you can share it so I can be more accessible. many thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1310735,
          "author_name": "Martin Chobanyan",
          "author_url": "",
          "post_date": "2021-05-16T22:34:29.927000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/nguyenhuudat\" target=\"_blank\">@nguyenhuudat</a>,</p>\n<p>I just made my submission kernel public: <a href=\"https://www.kaggle.com/mch123/ensemble-cell-transformer-2\" target=\"_blank\">https://www.kaggle.com/mch123/ensemble-cell-transformer-2</a></p>\n<p>You'll also have to use the model states in this dataset I made: <a href=\"https://www.kaggle.com/mch123/hparoimodels\" target=\"_blank\">https://www.kaggle.com/mch123/hparoimodels</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1305022,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:25:12.803000",
      "content": "<p><a href=\"https://www.kaggle.com/mch123\" target=\"_blank\">@mch123</a> Congratulations  and Thanks for sharing the approach</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1307633,
          "author_name": "Martin Chobanyan",
          "author_url": "",
          "post_date": "2021-05-14T15:09:47.800000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1304735": "Hey fellow Kaggle competitors! \n\nI had a lot of fun with this unique competition and I'm looking forward to reading your own solutions to this problem. Here is an overview of my final approach + link to my code (open images in a new tab for higher res):\n\n[Github Repo](https://github.com/martin-chobanyan/hpa-single-cell)\n\n![cell transformer](https://raw.githubusercontent.com/martin-chobanyan/hpa-single-cell/main/resources/cell-transformer-overview.png)\n\nFirst off, in all of my models I used the pre-trained backbone CNN of @bestfitting 's winning DenseNet-121 models from the previous competition (3 different pre-trained versions). I stacked all of the color stains into 4-channel images and resized them to 1536x1536 and fed them through the backbone CNN, resulting in 1024 feature maps.\n\nI then extracted a feature vector for each cell in the image by performing Region-of-Interest (RoI) pooling over the feature maps using the cell segmentation masks from the HPA Cell Segmentator. The pooling layer consisted of a concatenation of both avg-pool and max-pool features (resulting in a 2048-dim vector per cell). I also applied an adaptive average pool over the cell masks to reduce them to 8x8, which I then flattened. These downsampled, flatten masks served as position encoding for the cell regions (see diagram below):\n\n![roi pool](https://raw.githubusercontent.com/martin-chobanyan/hpa-single-cell/main/resources/cell-roi-pool-overview.png)\n\nAt this stage, we have a sequence of cell vectors, which allows us to explore the use of a Transformer model. The cell feature vectors are passed through a feed-forward layer which maps them to a 1024-dim embedding. Then, four encoder layers are applied as they were defined in the original Transformer paper. A final feed forward layer then maps each cell to an 18-dimensional logit.\n\nTo predict the image-level labels, a log-sum-exponential layer is applied which is defined as f(x) = (1/r) * ln(mean(e^(r*x))) with r=5 in our case. This serves as a mechanism to interpolate between average pooling and max pooling with different values for r.\n\nFor training I used Focal Loss with sigmoid activation and flip/rotate/crop augmentations. During inference, I extracted the cell-level logits and applied a sigmoid activation before their reduction to image-level labels. The final class prediction for a cell was averaged across all three models (three different pre-trained backbones).\n\nI also explored the use of [Peak Response Maps](https://arxiv.org/abs/1804.00880) as a CAM based approach to the problem. Though the localizations looked good, the model was a bit too sensitive to classes which were not present in the image.\n\nThings I would have tried with more time:\n\n- Ensembling with a cell-level classifier\n- Explore more ways of merging the peak-response-map results with the cell transformer model\n- Domain adaptation to the public test set (since we know the test set has higher variability in protein locations within a given image compared to the training set).",
    "1309793": "Hi @mch123 \nI'm a newbie, I want to learn more and practice your source code, do you have any version above Kaggle notebook. If yes I hope you can share it so I can be more accessible. many thanks.",
    "1305022": "@mch123 Congratulations  and Thanks for sharing the approach"
  }
}