{
  "id": 238474,
  "title": "21st Place Solution: You don't need cell tiles",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238474",
  "author_name": "Alexander Riedel",
  "post_date": "2021-05-12T09:15:40.425000",
  "votes": 27,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hey  everyone, thanks for that awesome challenge! Congrats to everyone :)<br>\nI'm verry happy to get my first silver medal on Kaggle and want to share my approach with you.<br>\nAnd thanks to phalanx and his <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217395\" target=\"_blank\">post </a>on Puzzle-CAM for good inspiration</p>\n<p><strong>General Approach</strong><br>\nLike many of you, I read a lot a weakly-labeled instance segmentation and eventually wanted to go for a image-level training and inferencing method. For this to achieve, a model producing good Class-Activation-Maps was needed so I decided to try <a href=\"https://arxiv.org/abs/2101.11253\" target=\"_blank\">Puzzle-CAM</a> and do some mapping magic for inferencing to get probabilities from my CAMs. </p>\n<p><strong>Training</strong><br>\nI trained according to the Puzzle-CAM paper with each images being tiled to four single images and considering the full-image CAMs versus the tiled-image CAMs in a loss function. I used a ResNest-101 and an EfficientNet-B4 with the according GAP Layers added and Focal Loss function. </p>\n<p><strong>Inferencing</strong><br>\nHere's the interesting part. I'm simply multiplying the CAM of each class with the cell mask of each cell and the class probability the model produces <em>(using a Swish-Activation to obtain the CAMs gives slightly better results than raw CAMs or ReLU)</em>. This gives very large class activated values for each class for each cell, which have to be mapped to real class probabilities and I used two approaches for this:</p>\n<ol>\n<li>standardize the values of each image using <code>sklearn.preprocessing.StandardScaler</code> and applying a sigmoid function to these values (works surprisingly good)</li>\n<li>Do the inferencing on the single-class labeled train data to get the raw values and train a gradient boosting regressor to learn the according label (0..1) for each class (to make sure, that the right mapping function, that might be different from the sigmoid function, is found)</li>\n</ol>\n<p>In the end I combined both approaches. </p>\n<p>Now enjoy some nice graphics showing my approaches (click links for higher res) :)</p>\n<p><img src=\"https://images2.imgbox.com/e7/85/HVh20eFe_o.jpg\" alt=\"\"><br>\n<a href=\"https://images2.imgbox.com/e7/85/HVh20eFe_o.jpg\" target=\"_blank\">TRAIN</a></p>\n<p><img src=\"https://images2.imgbox.com/fd/b1/4bEYABtz_o.jpg\" alt=\"\"><br>\n<a href=\"https://images2.imgbox.com/fd/b1/4bEYABtz_o.jpg\" target=\"_blank\">INFERENCE</a></p>",
  "messages": [
    {
      "id": 1303832,
      "postDate": "2021-05-12T09:15:40.427Z",
      "content": "<p>Hey  everyone, thanks for that awesome challenge! Congrats to everyone :)<br>\nI'm verry happy to get my first silver medal on Kaggle and want to share my approach with you.<br>\nAnd thanks to phalanx and his <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217395\" target=\"_blank\">post </a>on Puzzle-CAM for good inspiration</p>\n<p><strong>General Approach</strong><br>\nLike many of you, I read a lot a weakly-labeled instance segmentation and eventually wanted to go for a image-level training and inferencing method. For this to achieve, a model producing good Class-Activation-Maps was needed so I decided to try <a href=\"https://arxiv.org/abs/2101.11253\" target=\"_blank\">Puzzle-CAM</a> and do some mapping magic for inferencing to get probabilities from my CAMs. </p>\n<p><strong>Training</strong><br>\nI trained according to the Puzzle-CAM paper with each images being tiled to four single images and considering the full-image CAMs versus the tiled-image CAMs in a loss function. I used a ResNest-101 and an EfficientNet-B4 with the according GAP Layers added and Focal Loss function. </p>\n<p><strong>Inferencing</strong><br>\nHere's the interesting part. I'm simply multiplying the CAM of each class with the cell mask of each cell and the class probability the model produces <em>(using a Swish-Activation to obtain the CAMs gives slightly better results than raw CAMs or ReLU)</em>. This gives very large class activated values for each class for each cell, which have to be mapped to real class probabilities and I used two approaches for this:</p>\n<ol>\n<li>standardize the values of each image using <code>sklearn.preprocessing.StandardScaler</code> and applying a sigmoid function to these values (works surprisingly good)</li>\n<li>Do the inferencing on the single-class labeled train data to get the raw values and train a gradient boosting regressor to learn the according label (0..1) for each class (to make sure, that the right mapping function, that might be different from the sigmoid function, is found)</li>\n</ol>\n<p>In the end I combined both approaches. </p>\n<p>Now enjoy some nice graphics showing my approaches (click links for higher res) :)</p>\n<p><img src=\"https://images2.imgbox.com/e7/85/HVh20eFe_o.jpg\" alt=\"\"><br>\n<a href=\"https://images2.imgbox.com/e7/85/HVh20eFe_o.jpg\" target=\"_blank\">TRAIN</a></p>\n<p><img src=\"https://images2.imgbox.com/fd/b1/4bEYABtz_o.jpg\" alt=\"\"><br>\n<a href=\"https://images2.imgbox.com/fd/b1/4bEYABtz_o.jpg\" target=\"_blank\">INFERENCE</a></p>",
      "rawMarkdown": "Hey  everyone, thanks for that awesome challenge! Congrats to everyone :)\nI'm verry happy to get my first silver medal on Kaggle and want to share my approach with you.\nAnd thanks to phalanx and his [post ](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217395)on Puzzle-CAM for good inspiration\n\n**General Approach**\nLike many of you, I read a lot a weakly-labeled instance segmentation and eventually wanted to go for a image-level training and inferencing method. For this to achieve, a model producing good Class-Activation-Maps was needed so I decided to try [Puzzle-CAM](https://arxiv.org/abs/2101.11253) and do some mapping magic for inferencing to get probabilities from my CAMs. \n\n**Training**\nI trained according to the Puzzle-CAM paper with each images being tiled to four single images and considering the full-image CAMs versus the tiled-image CAMs in a loss function. I used a ResNest-101 and an EfficientNet-B4 with the according GAP Layers added and Focal Loss function. \n\n**Inferencing**\nHere's the interesting part. I'm simply multiplying the CAM of each class with the cell mask of each cell and the class probability the model produces *(using a Swish-Activation to obtain the CAMs gives slightly better results than raw CAMs or ReLU)*. This gives very large class activated values for each class for each cell, which have to be mapped to real class probabilities and I used two approaches for this:\n1. standardize the values of each image using `sklearn.preprocessing.StandardScaler` and applying a sigmoid function to these values (works surprisingly good)\n2. Do the inferencing on the single-class labeled train data to get the raw values and train a gradient boosting regressor to learn the according label (0..1) for each class (to make sure, that the right mapping function, that might be different from the sigmoid function, is found)\n\nIn the end I combined both approaches. \n\nNow enjoy some nice graphics showing my approaches (click links for higher res) :)\n\n![](https://images2.imgbox.com/e7/85/HVh20eFe_o.jpg)\n[TRAIN](https://images2.imgbox.com/e7/85/HVh20eFe_o.jpg)\n\n![](https://images2.imgbox.com/fd/b1/4bEYABtz_o.jpg)\n[INFERENCE](https://images2.imgbox.com/fd/b1/4bEYABtz_o.jpg)",
      "votes": 27
    },
    {
      "id": 1304616,
      "postDate": "2021-05-12T18:29:06.987Z",
      "content": "<p>Hey, Alexander! Congrats! :) <br>\nAnd thanks for the write-up! Your solution is super impressive, with so many various model architectures, including transformers, and then combining it with Puzzle-CAM, simply WOW! :)</p>\n<p>p.s. also thanks a lot for your activity throughout the competition! :) </p>",
      "rawMarkdown": "Hey, Alexander! Congrats! :) \nAnd thanks for the write-up! Your solution is super impressive, with so many various model architectures, including transformers, and then combining it with Puzzle-CAM, simply WOW! :)\n\np.s. also thanks a lot for your activity throughout the competition! :) ",
      "votes": 1,
      "replies": [
        {
          "id": 1304640,
          "postDate": "2021-05-12T18:48:35.757Z",
          "content": "<p>Thanks Raman!</p>",
          "rawMarkdown": "Thanks Raman!"
        }
      ]
    },
    {
      "id": 1304445,
      "postDate": "2021-05-12T16:14:38.940Z",
      "content": "<p>Nice work! </p>\n<p>I also tried PuzzleCAM early on but I couldn't get it to work as good. I think that maybe because I was using Global Max Pooling at the end instead of Global Average Pooling. I was hesitant to use GAP because there were a lot of images where the labels did not apply to every cell in the image, especially in the multi-label images (and so I wanted to avoid forcing the model to try and learn features from those cells to make the prediction). It looks like that may not have been as big a problem as I thought.</p>\n<p>A few questions:</p>\n<ul>\n<li>Why did you decide to multiply the CAMs with the class probabilities? </li>\n<li>What is the bottom half of the inference diagram calculating (the branch with the ViT model)? Cellwise-probabilities as well?</li>\n</ul>\n<p>Thanks for the post.</p>",
      "rawMarkdown": "Nice work! \n\nI also tried PuzzleCAM early on but I couldn't get it to work as good. I think that maybe because I was using Global Max Pooling at the end instead of Global Average Pooling. I was hesitant to use GAP because there were a lot of images where the labels did not apply to every cell in the image, especially in the multi-label images (and so I wanted to avoid forcing the model to try and learn features from those cells to make the prediction). It looks like that may not have been as big a problem as I thought.\n\nA few questions:\n- Why did you decide to multiply the CAMs with the class probabilities? \n- What is the bottom half of the inference diagram calculating (the branch with the ViT model)? Cellwise-probabilities as well?\n\nThanks for the post.",
      "votes": 1,
      "replies": [
        {
          "id": 1304460,
          "postDate": "2021-05-12T16:25:54.207Z",
          "content": "<p>Thanks Martin!<br>\n1) It brought some stability to the CAM*mask-values because I had some very large class-activations for classes with very low probability (that presumeably evened out in the last fully connected layer of the model?)<br>\n2) The bottom half are image-level labels</p>",
          "rawMarkdown": "Thanks Martin!\n1) It brought some stability to the CAM*mask-values because I had some very large class-activations for classes with very low probability (that presumeably evened out in the last fully connected layer of the model?)\n2) The bottom half are image-level labels",
          "votes": 1
        }
      ]
    },
    {
      "id": 1304334,
      "postDate": "2021-05-12T15:02:51.440Z",
      "content": "<p>Thanks for sharing! I was very interested to see how you converted the CAMs into predictions.</p>",
      "rawMarkdown": "Thanks for sharing! I was very interested to see how you converted the CAMs into predictions.",
      "votes": 1
    },
    {
      "id": 1304263,
      "postDate": "2021-05-12T14:13:31.537Z",
      "content": "<p>Congratulations Alexander! Your posts and dataset were very helpful to us! Looks like a lot of interesting things can be done with CAM outputs, we also tested a bunch of different approaches to go from CAM to cell level predictions. I'm looking forward to study your solution in detail!</p>",
      "rawMarkdown": "Congratulations Alexander! Your posts and dataset were very helpful to us! Looks like a lot of interesting things can be done with CAM outputs, we also tested a bunch of different approaches to go from CAM to cell level predictions. I'm looking forward to study your solution in detail!",
      "votes": 1,
      "replies": [
        {
          "id": 1304433,
          "postDate": "2021-05-12T16:05:34.047Z",
          "content": "<p>Thanks Darek! That's really nice to hear, you're also an awesome contributor to Kaggle!</p>",
          "rawMarkdown": "Thanks Darek! That's really nice to hear, you're also an awesome contributor to Kaggle!"
        }
      ]
    },
    {
      "id": 1304142,
      "postDate": "2021-05-12T12:55:43.663Z",
      "content": "<p>Thanks for sharing, you tried various kinds of archs!</p>",
      "rawMarkdown": "Thanks for sharing, you tried various kinds of archs!",
      "votes": 1
    },
    {
      "id": 1304305,
      "postDate": "2021-05-12T14:46:31.147Z",
      "content": "<p>Thanks for sharing. Expecially your posted dataset were very helpfull.</p>",
      "rawMarkdown": "Thanks for sharing. Expecially your posted dataset were very helpfull.",
      "votes": 2
    },
    {
      "id": 1306386,
      "postDate": "2021-05-13T18:21:43.413Z",
      "content": "<p>Really impressive solution!<br>\ndid u use any app to draw this diagram?<br>\nalso, would you plan to release the source code of your solution?</p>",
      "rawMarkdown": "Really impressive solution!\ndid u use any app to draw this diagram?\nalso, would you plan to release the source code of your solution?",
      "replies": [
        {
          "id": 1307165,
          "postDate": "2021-05-14T09:41:54.183Z",
          "content": "<p>Thanks Alex! Yes I used Lucidchart: <a href=\"https://www.lucidchart.com/\" target=\"_blank\">https://www.lucidchart.com/</a><br>\nI made my inferencing code public: <a href=\"https://www.kaggle.com/alexanderriedel/hpa-inferencing?scriptVersionId=62555565\" target=\"_blank\">https://www.kaggle.com/alexanderriedel/hpa-inferencing?scriptVersionId=62555565</a></p>\n<p>the training is basically Puzze-CAM</p>",
          "rawMarkdown": "Thanks Alex! Yes I used Lucidchart: https://www.lucidchart.com/\nI made my inferencing code public: https://www.kaggle.com/alexanderriedel/hpa-inferencing?scriptVersionId=62555565\n\nthe training is basically Puzze-CAM\n"
        }
      ]
    },
    {
      "id": 1305034,
      "postDate": "2021-05-13T04:27:56.900Z",
      "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> Congratulations  and Thanks for sharing the approach</p>",
      "rawMarkdown": "@alexanderriedel Congratulations  and Thanks for sharing the approach"
    }
  ],
  "comments": [
    {
      "id": 1304616,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-12T18:29:06.987000",
      "content": "<p>Hey, Alexander! Congrats! :) <br>\nAnd thanks for the write-up! Your solution is super impressive, with so many various model architectures, including transformers, and then combining it with Puzzle-CAM, simply WOW! :)</p>\n<p>p.s. also thanks a lot for your activity throughout the competition! :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304640,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-05-12T18:48:35.757000",
          "content": "<p>Thanks Raman!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304445,
      "author_name": "Martin Chobanyan",
      "author_url": "",
      "post_date": "2021-05-12T16:14:38.940000",
      "content": "<p>Nice work! </p>\n<p>I also tried PuzzleCAM early on but I couldn't get it to work as good. I think that maybe because I was using Global Max Pooling at the end instead of Global Average Pooling. I was hesitant to use GAP because there were a lot of images where the labels did not apply to every cell in the image, especially in the multi-label images (and so I wanted to avoid forcing the model to try and learn features from those cells to make the prediction). It looks like that may not have been as big a problem as I thought.</p>\n<p>A few questions:</p>\n<ul>\n<li>Why did you decide to multiply the CAMs with the class probabilities? </li>\n<li>What is the bottom half of the inference diagram calculating (the branch with the ViT model)? Cellwise-probabilities as well?</li>\n</ul>\n<p>Thanks for the post.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304460,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-05-12T16:25:54.207000",
          "content": "<p>Thanks Martin!<br>\n1) It brought some stability to the CAM*mask-values because I had some very large class-activations for classes with very low probability (that presumeably evened out in the last fully connected layer of the model?)<br>\n2) The bottom half are image-level labels</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1304334,
      "author_name": "Andrew Tratz",
      "author_url": "",
      "post_date": "2021-05-12T15:02:51.440000",
      "content": "<p>Thanks for sharing! I was very interested to see how you converted the CAMs into predictions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1304263,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-05-12T14:13:31.537000",
      "content": "<p>Congratulations Alexander! Your posts and dataset were very helpful to us! Looks like a lot of interesting things can be done with CAM outputs, we also tested a bunch of different approaches to go from CAM to cell level predictions. I'm looking forward to study your solution in detail!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304433,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-05-12T16:05:34.047000",
          "content": "<p>Thanks Darek! That's really nice to hear, you're also an awesome contributor to Kaggle!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304142,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-12T12:55:43.663000",
      "content": "<p>Thanks for sharing, you tried various kinds of archs!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1304305,
      "author_name": "LucaMTB",
      "author_url": "",
      "post_date": "2021-05-12T14:46:31.147000",
      "content": "<p>Thanks for sharing. Expecially your posted dataset were very helpfull.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1306386,
      "author_name": "Alex Lau",
      "author_url": "",
      "post_date": "2021-05-13T18:21:43.413000",
      "content": "<p>Really impressive solution!<br>\ndid u use any app to draw this diagram?<br>\nalso, would you plan to release the source code of your solution?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1307165,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-05-14T09:41:54.183000",
          "content": "<p>Thanks Alex! Yes I used Lucidchart: <a href=\"https://www.lucidchart.com/\" target=\"_blank\">https://www.lucidchart.com/</a><br>\nI made my inferencing code public: <a href=\"https://www.kaggle.com/alexanderriedel/hpa-inferencing?scriptVersionId=62555565\" target=\"_blank\">https://www.kaggle.com/alexanderriedel/hpa-inferencing?scriptVersionId=62555565</a></p>\n<p>the training is basically Puzze-CAM</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1305034,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:27:56.900000",
      "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> Congratulations  and Thanks for sharing the approach</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1303832": "Hey  everyone, thanks for that awesome challenge! Congrats to everyone :)\nI'm verry happy to get my first silver medal on Kaggle and want to share my approach with you.\nAnd thanks to phalanx and his [post ](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217395)on Puzzle-CAM for good inspiration\n\n**General Approach**\nLike many of you, I read a lot a weakly-labeled instance segmentation and eventually wanted to go for a image-level training and inferencing method. For this to achieve, a model producing good Class-Activation-Maps was needed so I decided to try [Puzzle-CAM](https://arxiv.org/abs/2101.11253) and do some mapping magic for inferencing to get probabilities from my CAMs. \n\n**Training**\nI trained according to the Puzzle-CAM paper with each images being tiled to four single images and considering the full-image CAMs versus the tiled-image CAMs in a loss function. I used a ResNest-101 and an EfficientNet-B4 with the according GAP Layers added and Focal Loss function. \n\n**Inferencing**\nHere's the interesting part. I'm simply multiplying the CAM of each class with the cell mask of each cell and the class probability the model produces *(using a Swish-Activation to obtain the CAMs gives slightly better results than raw CAMs or ReLU)*. This gives very large class activated values for each class for each cell, which have to be mapped to real class probabilities and I used two approaches for this:\n1. standardize the values of each image using `sklearn.preprocessing.StandardScaler` and applying a sigmoid function to these values (works surprisingly good)\n2. Do the inferencing on the single-class labeled train data to get the raw values and train a gradient boosting regressor to learn the according label (0..1) for each class (to make sure, that the right mapping function, that might be different from the sigmoid function, is found)\n\nIn the end I combined both approaches. \n\nNow enjoy some nice graphics showing my approaches (click links for higher res) :)\n\n![](https://images2.imgbox.com/e7/85/HVh20eFe_o.jpg)\n[TRAIN](https://images2.imgbox.com/e7/85/HVh20eFe_o.jpg)\n\n![](https://images2.imgbox.com/fd/b1/4bEYABtz_o.jpg)\n[INFERENCE](https://images2.imgbox.com/fd/b1/4bEYABtz_o.jpg)",
    "1304616": "Hey, Alexander! Congrats! :) \nAnd thanks for the write-up! Your solution is super impressive, with so many various model architectures, including transformers, and then combining it with Puzzle-CAM, simply WOW! :)\n\np.s. also thanks a lot for your activity throughout the competition! :) ",
    "1304445": "Nice work! \n\nI also tried PuzzleCAM early on but I couldn't get it to work as good. I think that maybe because I was using Global Max Pooling at the end instead of Global Average Pooling. I was hesitant to use GAP because there were a lot of images where the labels did not apply to every cell in the image, especially in the multi-label images (and so I wanted to avoid forcing the model to try and learn features from those cells to make the prediction). It looks like that may not have been as big a problem as I thought.\n\nA few questions:\n- Why did you decide to multiply the CAMs with the class probabilities? \n- What is the bottom half of the inference diagram calculating (the branch with the ViT model)? Cellwise-probabilities as well?\n\nThanks for the post.",
    "1304334": "Thanks for sharing! I was very interested to see how you converted the CAMs into predictions.",
    "1304263": "Congratulations Alexander! Your posts and dataset were very helpful to us! Looks like a lot of interesting things can be done with CAM outputs, we also tested a bunch of different approaches to go from CAM to cell level predictions. I'm looking forward to study your solution in detail!",
    "1304142": "Thanks for sharing, you tried various kinds of archs!",
    "1304305": "Thanks for sharing. Expecially your posted dataset were very helpfull.",
    "1306386": "Really impressive solution!\ndid u use any app to draw this diagram?\nalso, would you plan to release the source code of your solution?",
    "1305034": "@alexanderriedel Congratulations  and Thanks for sharing the approach"
  }
}