{
  "id": 239001,
  "title": "Fair Cell Activation Network and Swin Transformer, the 1st place solution",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/239001",
  "author_name": "bestfitting",
  "post_date": "2021-05-14T08:54:36.111000",
  "votes": 221,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Congrats to all the winners, and thanks to the Human Protein Atlas team and kaggle hosted such an interesting competetion.</p>\n<p><strong>1. Introduction</strong></p>\n<p>The main challenge of this competition is to find a way to label every cell in a labeled image, it is a new type of weakly supervised challenge as we are provided with a cell segmentation model which means this is not a problem widely discussed like weakly supervised object detection or segmentation.</p>\n<p>The common method for this problem is to find CAM or attention on cells, but the activations of a CNN network is focus on most discriminative parts of an image which lead to a low recall rate, to solve this problem I developed a network called Fair Cell Activation Network(FCAN) based on Puzzle-CAM.</p>\n<p>After getting the prediction of each cell from FCAN, I relabeled the cells to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ] by rule and trained a Swin Transformer model to predict the cell label. </p>\n<p>Ensemble of this two models and post-processing by reducing the confidence of the cells on image border can achieve the first place with 0.555 on private LB, a more complex ensemble solution with 6 models can reach 0.566 on private LB.</p>\n<p><strong>2. Methods</strong></p>\n<p><strong>2.1 Fair Cell Activation Network</strong><br>\nThe activations of CNN on feature map of an image is focus on most descriminative instance of a class despite many instances exists. I call this phenomena unfair activation, to address this problem, a network was proposed based on Puzzle-CAM.</p>\n<p><strong>Training</strong><br>\n<img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png\" alt=\"FCAN train\"><br>\n<strong>Inference</strong><br>\n<img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png\" alt=\"FCAN inference\"><br>\nThe main difference to Puzzle-CAM in train part of this model is we can select cells instead of splitting the image to grid.<br>\nThe confidence of a cell is multiplication of image-level prediction and cell level prediction. </p>\n<p><strong>Details</strong></p>\n<p>Images are resized to 512x512 px</p>\n<p><strong>Backbone</strong>: EfficientNet-B0.</p>\n<p><strong>Losses</strong><br>\nLcls is FocalLoss + SymmetricLovaszLoss + HardLogLoss<br>\nLml is ArcFaceLoss metric learning supervised by antibody-id.<br>\nLre is MSELoss</p>\n<p><strong>Augmentation</strong><br>\nflip, transpose, scale, rotate, crop<br>\nAdding mitotic spindles with high confidence to other images to generate more positive samples of this type(lead to a boost with 0.02)<br>\nTest Time augmentation: default,flipud,fliplr,transpose.</p>\n<p><strong>Validation</strong><br>\nSelect 433 images in public test set which can be found in public-hpa dataset, and remove them from training set of the model.</p>\n<p><strong>Compare the score of models</strong><br>\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/CompareModels.png\" alt=\"FCAN compare results\"><br>\n<strong>Compare the models real images</strong><br>\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label1_compare.png\" alt=\"FCAN demo label1\"><br>\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label2_compare.png\" alt=\"FCAN demo label2\"><br>\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label15_compare.png\" alt=\"FCAN demo label15\"><br>\nThe left part of these images are the results from traditional CNN model,  the middle are the results from the puzzle-cam, the right part is the results from FCAN,.<br>\nThe top part of every figure is image with positive label. the bottom is negative image.<br>\nThe number on each cell is the confidence of this cell.</p>\n<p><strong>2.2 Swin Tranformer based cell classification model</strong></p>\n<p><strong>Data</strong><br>\nCrop the cells in an image by using Cell-Segmentaion model.<br>\nThe cells were labeled to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ], this is a rule based procedure, After getting the outputs of all cells of train set from FCAN introduced above, we can give higher label value if the image probability and cell probability are high, and the cells from an image with label A were given at least 0.25 of this label A. The thresholds of the rule were not sensitive according to my experiments.</p>\n<p><strong>Model</strong><br>\nSwin Transformer with pretrained weights small_patch4_window7. <br>\nThe cells are resized to 128x128 px to feed into the network</p>\n<p><strong>Loss</strong><br>\nFocalLoss</p>\n<p><strong>Augmentation</strong><br>\nFlip, transpose, scale, rotate, crop<br>\nTest Time augmentation:default,flipud,fliplr,transpose.</p>\n<p><strong>Validation Strategy</strong><br>\nGetting the max cell confidence in an image and use this confidence as the image confidence, calculate the MAP of the image level. <br>\nAlthough this is not a strategy always keep consistency with public LB, but it can reflect the capability of the model to some extend. </p>\n<p><strong>Inference</strong><br>\nThe confidence of a cell is multiplication of FCAN image-level prediction and cell level Swin Transformer prediction. </p>\n<p><strong>2.3 Ensemble</strong><br>\nWeighted average of the prediction of FCAN and Swin transformer. </p>\n<p><strong>2.4 Post-Processing</strong><br>\nAs the host did not label some cells on border, if we give the cell with high confidence, the Fasle Positive cells will increase, so I trained a model to predict the completeness of a cell. If the probability to be a whole cell is very low, the confidence of this cell is multiplied by low value such as 0.3.</p>\n<p><img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png\" alt=\"FCAN train\"><br>\nThe data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.</p>\n<p>The backbone of this model is EfficientNet-B0, 3 epochs is enough to get a quite good model.</p>\n<p>The score can improve 0.007 to 0.01 after this step.  </p>\n<p><strong>3. Results</strong></p>\n<p><strong>Results of simple solution</strong><br>\n<img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png\" alt=\"Simple-Solution\"></p>\n<p><strong>Results of final submission</strong><br>\n<img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Complex.png\" alt=\"Complex Solution\"></p>\n<p><strong>4. Conclusion</strong></p>\n<p>4.1 The Fair Cell Activation Network(FCAN) can increase cell level recall which is very important to this competition.</p>\n<p>4.2 The vision transformer models have shown promising capability.</p>\n<p>4.3 Larger model not always means better result as most pre-trained models are designed for ImageNet, our models should find relationship of relative position of pixels instead of abstract semantic. </p>\n<p>4.4 I found little differences among JPEG, PNG  and  8bit 16bit formats.</p>",
  "messages": [
    {
      "id": 1307108,
      "postDate": "2021-05-14T08:54:36.110Z",
      "content": "<p>Congrats to all the winners, and thanks to the Human Protein Atlas team and kaggle hosted such an interesting competetion.</p>\n<p><strong>1. Introduction</strong></p>\n<p>The main challenge of this competition is to find a way to label every cell in a labeled image, it is a new type of weakly supervised challenge as we are provided with a cell segmentation model which means this is not a problem widely discussed like weakly supervised object detection or segmentation.</p>\n<p>The common method for this problem is to find CAM or attention on cells, but the activations of a CNN network is focus on most discriminative parts of an image which lead to a low recall rate, to solve this problem I developed a network called Fair Cell Activation Network(FCAN) based on Puzzle-CAM.</p>\n<p>After getting the prediction of each cell from FCAN, I relabeled the cells to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ] by rule and trained a Swin Transformer model to predict the cell label. </p>\n<p>Ensemble of this two models and post-processing by reducing the confidence of the cells on image border can achieve the first place with 0.555 on private LB, a more complex ensemble solution with 6 models can reach 0.566 on private LB.</p>\n<p><strong>2. Methods</strong></p>\n<p><strong>2.1 Fair Cell Activation Network</strong><br>\nThe activations of CNN on feature map of an image is focus on most descriminative instance of a class despite many instances exists. I call this phenomena unfair activation, to address this problem, a network was proposed based on Puzzle-CAM.</p>\n<p><strong>Training</strong><br>\n<img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png\" alt=\"FCAN train\"><br>\n<strong>Inference</strong><br>\n<img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png\" alt=\"FCAN inference\"><br>\nThe main difference to Puzzle-CAM in train part of this model is we can select cells instead of splitting the image to grid.<br>\nThe confidence of a cell is multiplication of image-level prediction and cell level prediction. </p>\n<p><strong>Details</strong></p>\n<p>Images are resized to 512x512 px</p>\n<p><strong>Backbone</strong>: EfficientNet-B0.</p>\n<p><strong>Losses</strong><br>\nLcls is FocalLoss + SymmetricLovaszLoss + HardLogLoss<br>\nLml is ArcFaceLoss metric learning supervised by antibody-id.<br>\nLre is MSELoss</p>\n<p><strong>Augmentation</strong><br>\nflip, transpose, scale, rotate, crop<br>\nAdding mitotic spindles with high confidence to other images to generate more positive samples of this type(lead to a boost with 0.02)<br>\nTest Time augmentation: default,flipud,fliplr,transpose.</p>\n<p><strong>Validation</strong><br>\nSelect 433 images in public test set which can be found in public-hpa dataset, and remove them from training set of the model.</p>\n<p><strong>Compare the score of models</strong><br>\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/CompareModels.png\" alt=\"FCAN compare results\"><br>\n<strong>Compare the models real images</strong><br>\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label1_compare.png\" alt=\"FCAN demo label1\"><br>\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label2_compare.png\" alt=\"FCAN demo label2\"><br>\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label15_compare.png\" alt=\"FCAN demo label15\"><br>\nThe left part of these images are the results from traditional CNN model,  the middle are the results from the puzzle-cam, the right part is the results from FCAN,.<br>\nThe top part of every figure is image with positive label. the bottom is negative image.<br>\nThe number on each cell is the confidence of this cell.</p>\n<p><strong>2.2 Swin Tranformer based cell classification model</strong></p>\n<p><strong>Data</strong><br>\nCrop the cells in an image by using Cell-Segmentaion model.<br>\nThe cells were labeled to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ], this is a rule based procedure, After getting the outputs of all cells of train set from FCAN introduced above, we can give higher label value if the image probability and cell probability are high, and the cells from an image with label A were given at least 0.25 of this label A. The thresholds of the rule were not sensitive according to my experiments.</p>\n<p><strong>Model</strong><br>\nSwin Transformer with pretrained weights small_patch4_window7. <br>\nThe cells are resized to 128x128 px to feed into the network</p>\n<p><strong>Loss</strong><br>\nFocalLoss</p>\n<p><strong>Augmentation</strong><br>\nFlip, transpose, scale, rotate, crop<br>\nTest Time augmentation:default,flipud,fliplr,transpose.</p>\n<p><strong>Validation Strategy</strong><br>\nGetting the max cell confidence in an image and use this confidence as the image confidence, calculate the MAP of the image level. <br>\nAlthough this is not a strategy always keep consistency with public LB, but it can reflect the capability of the model to some extend. </p>\n<p><strong>Inference</strong><br>\nThe confidence of a cell is multiplication of FCAN image-level prediction and cell level Swin Transformer prediction. </p>\n<p><strong>2.3 Ensemble</strong><br>\nWeighted average of the prediction of FCAN and Swin transformer. </p>\n<p><strong>2.4 Post-Processing</strong><br>\nAs the host did not label some cells on border, if we give the cell with high confidence, the Fasle Positive cells will increase, so I trained a model to predict the completeness of a cell. If the probability to be a whole cell is very low, the confidence of this cell is multiplied by low value such as 0.3.</p>\n<p><img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png\" alt=\"FCAN train\"><br>\nThe data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.</p>\n<p>The backbone of this model is EfficientNet-B0, 3 epochs is enough to get a quite good model.</p>\n<p>The score can improve 0.007 to 0.01 after this step.  </p>\n<p><strong>3. Results</strong></p>\n<p><strong>Results of simple solution</strong><br>\n<img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png\" alt=\"Simple-Solution\"></p>\n<p><strong>Results of final submission</strong><br>\n<img src=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Complex.png\" alt=\"Complex Solution\"></p>\n<p><strong>4. Conclusion</strong></p>\n<p>4.1 The Fair Cell Activation Network(FCAN) can increase cell level recall which is very important to this competition.</p>\n<p>4.2 The vision transformer models have shown promising capability.</p>\n<p>4.3 Larger model not always means better result as most pre-trained models are designed for ImageNet, our models should find relationship of relative position of pixels instead of abstract semantic. </p>\n<p>4.4 I found little differences among JPEG, PNG  and  8bit 16bit formats.</p>",
      "rawMarkdown": "Congrats to all the winners, and thanks to the Human Protein Atlas team and kaggle hosted such an interesting competetion.\n\n**1. Introduction**\n\nThe main challenge of this competition is to find a way to label every cell in a labeled image, it is a new type of weakly supervised challenge as we are provided with a cell segmentation model which means this is not a problem widely discussed like weakly supervised object detection or segmentation.\n\nThe common method for this problem is to find CAM or attention on cells, but the activations of a CNN network is focus on most discriminative parts of an image which lead to a low recall rate, to solve this problem I developed a network called Fair Cell Activation Network(FCAN) based on Puzzle-CAM.\n\nAfter getting the prediction of each cell from FCAN, I relabeled the cells to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ] by rule and trained a Swin Transformer model to predict the cell label. \n\nEnsemble of this two models and post-processing by reducing the confidence of the cells on image border can achieve the first place with 0.555 on private LB, a more complex ensemble solution with 6 models can reach 0.566 on private LB.\n\n**2. Methods**\n\n**2.1 Fair Cell Activation Network**\nThe activations of CNN on feature map of an image is focus on most descriminative instance of a class despite many instances exists. I call this phenomena unfair activation, to address this problem, a network was proposed based on Puzzle-CAM.\n\n**Training**\n![FCAN train](https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png)\n**Inference**\n![FCAN inference](https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png)\nThe main difference to Puzzle-CAM in train part of this model is we can select cells instead of splitting the image to grid.\nThe confidence of a cell is multiplication of image-level prediction and cell level prediction. \n\n**Details**\n\nImages are resized to 512x512 px\n\n**Backbone**: EfficientNet-B0.\n\n**Losses**\nLcls is FocalLoss + SymmetricLovaszLoss + HardLogLoss\nLml is ArcFaceLoss metric learning supervised by antibody-id.\nLre is MSELoss\n\n**Augmentation**\nflip, transpose, scale, rotate, crop\nAdding mitotic spindles with high confidence to other images to generate more positive samples of this type(lead to a boost with 0.02)\nTest Time augmentation: default,flipud,fliplr,transpose.\n\n**Validation**\nSelect 433 images in public test set which can be found in public-hpa dataset, and remove them from training set of the model.\n\n**Compare the score of models**\n![FCAN compare results](https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/CompareModels.png)\n**Compare the models real images**\n![FCAN demo label1](https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label1_compare.png)\n![FCAN demo label2](https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label2_compare.png)\n![FCAN demo label15](https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label15_compare.png)\nThe left part of these images are the results from traditional CNN model,  the middle are the results from the puzzle-cam, the right part is the results from FCAN,.\nThe top part of every figure is image with positive label. the bottom is negative image.\nThe number on each cell is the confidence of this cell.\n\n\n**2.2 Swin Tranformer based cell classification model**\n\n**Data**\nCrop the cells in an image by using Cell-Segmentaion model.\nThe cells were labeled to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ], this is a rule based procedure, After getting the outputs of all cells of train set from FCAN introduced above, we can give higher label value if the image probability and cell probability are high, and the cells from an image with label A were given at least 0.25 of this label A. The thresholds of the rule were not sensitive according to my experiments.\n\n**Model**\nSwin Transformer with pretrained weights small_patch4_window7. \nThe cells are resized to 128x128 px to feed into the network\n\n**Loss**\nFocalLoss\n\n**Augmentation**\nFlip, transpose, scale, rotate, crop\nTest Time augmentation:default,flipud,fliplr,transpose.\n\n**Validation Strategy**\nGetting the max cell confidence in an image and use this confidence as the image confidence, calculate the MAP of the image level. \nAlthough this is not a strategy always keep consistency with public LB, but it can reflect the capability of the model to some extend. \n\n**Inference**\nThe confidence of a cell is multiplication of FCAN image-level prediction and cell level Swin Transformer prediction. \n\n\n**2.3 Ensemble**\nWeighted average of the prediction of FCAN and Swin transformer. \n\n**2.4 Post-Processing**\nAs the host did not label some cells on border, if we give the cell with high confidence, the Fasle Positive cells will increase, so I trained a model to predict the completeness of a cell. If the probability to be a whole cell is very low, the confidence of this cell is multiplied by low value such as 0.3.\n\n![FCAN train](https://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png)\nThe data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.\n\nThe backbone of this model is EfficientNet-B0, 3 epochs is enough to get a quite good model.\n\nThe score can improve 0.007 to 0.01 after this step.  \n\n**3. Results**\n\n**Results of simple solution**\n![Simple-Solution](https://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png)\n\n**Results of final submission**\n![Complex Solution](https://bestfitting.github.io/kaggle/hpa2021/figures/Complex.png)\n\n**4. Conclusion**\n\n4.1 The Fair Cell Activation Network(FCAN) can increase cell level recall which is very important to this competition.\n\n4.2 The vision transformer models have shown promising capability.\n\n4.3 Larger model not always means better result as most pre-trained models are designed for ImageNet, our models should find relationship of relative position of pixels instead of abstract semantic. \n\n4.4 I found little differences among JPEG, PNG  and  8bit 16bit formats.\n  \n\n",
      "votes": 221
    },
    {
      "id": 1307315,
      "postDate": "2021-05-14T11:25:24.897Z",
      "content": "<p>Awesome, congrats on being back to #1!</p>",
      "rawMarkdown": "Awesome, congrats on being back to #1!",
      "votes": 7,
      "replies": [
        {
          "id": 1308576,
          "postDate": "2021-05-15T10:05:57.060Z",
          "content": "<p>Thanks, I am sure you or guanshuo <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> will be back to #1 very soon, and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> is also on the way to this position :)</p>",
          "rawMarkdown": "Thanks, I am sure you or guanshuo @wowfattie will be back to #1 very soon, and @christofhenkel is also on the way to this position :)",
          "votes": 5
        }
      ]
    },
    {
      "id": 1308073,
      "postDate": "2021-05-15T00:41:49.230Z",
      "content": "<p>Congratulations winning HPA two years in a row even after posting your solution from the first win!</p>",
      "rawMarkdown": "Congratulations winning HPA two years in a row even after posting your solution from the first win!",
      "votes": 8
    },
    {
      "id": 1610714,
      "postDate": "2021-12-07T13:17:55.203Z",
      "content": "<p>日本語訳</p>\n<p>Congrats to all the winners, and thanks to the Human Protein Atlas team and kaggle hosted such an interesting competetion.</p>\n<ol>\n<li>Introduction</li>\n</ol>\n<p>The main challenge of this competition is to find a way to label every cell in a labeled image, it is a new type of weakly supervised challenge as we are provided with a cell segmentation model which means this is not a problem widely discussed like weakly supervised object detection or segmentation.</p>\n<p>The common method for this problem is to find CAM or attention on cells, but the activations of a CNN network is focus on most discriminative parts of an image which lead to a low recall rate, to solve this problem I developed a network called Fair Cell Activation Network(FCAN) based on Puzzle-CAM.</p>\n<p>After getting the prediction of each cell from FCAN, I relabeled the cells to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ] by rule and trained a Swin Transformer model to predict the cell label.</p>\n<p>Ensemble of this two models and post-processing by reducing the confidence of the cells on image border can achieve the first place with 0.555 on private LB, a more complex ensemble solution with 6 models can reach 0.566 on private LB.</p>\n<p>この競争の主な課題は、ラベル付けされた画像内のすべてのセルにラベルを付ける方法を見つけることです。これは、セルセグメンテーションモデルが提供されているため、新しいタイプの弱教師ありチャレンジです。これは、弱教師ありのように広く議論されている問題ではないことを意味します。オブジェクトの検出またはセグメンテーション。</p>\n<p>この問題の一般的な方法は、CAMまたは細胞への注意を見つけることですが、CNNネットワークのアクティブ化は、画像の最も識別可能な部分に焦点を当てているため、リコール率が低くなります。この問題を解決するために、FairCellというネットワークを開発しました。パズルCAMに基づくアクティベーションネットワーク（FCAN）。</p>\n<p>FCANから各セルの予測を取得した後、ルールによってセルをラベル[1.0、0.75、0.5、0.25、0]で5レベルに再ラベル付けし、セルラベルを予測するようにSwinTransformerモデルをトレーニングしました。</p>\n<p>この2つのモデルのアンサンブルと、画像境界のセルの信頼性を下げることによる後処理は、プライベートLBで0.555で最初の場所を達成でき、6つのモデルでのより複雑なアンサンブルソリューションは、プライベートLBで0.566に達することができます。</p>\n<ol>\n<li>Methods</li>\n</ol>\n<p>2.1 Fair Cell Activation Network<br>\nThe activations of CNN on feature map of an image is focus on most descriminative instance of a class despite many instances exists. I call this phenomena unfair activation, to address this problem, a network was proposed based on Puzzle-CAM.</p>\n<p>画像の機能マップでのCNNのアクティブ化は、多くのインスタンスが存在するにもかかわらず、クラスの最も識別力のあるインスタンスに焦点を合わせています。 私はこの現象を不公平な活性化と呼んでいますが、この問題に対処するために、Puzzle-CAMに基づいたネットワークが提案されました。</p>\n<p>Training<br>\n<a href=\"https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png\" target=\"_blank\">https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png</a><br>\nInference<br>\n<a href=\"https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png\" target=\"_blank\">https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png</a><br>\nThe main difference to Puzzle-CAM in train part of this model is we can select cells instead of splitting the image to grid.<br>\nThe confidence of a cell is multiplication of image-level prediction and cell level prediction.</p>\n<p>Details</p>\n<p>Images are resized to 512x512 px</p>\n<p>Backbone: EfficientNet-B0.</p>\n<p>Losses<br>\nLcls is FocalLoss + SymmetricLovaszLoss + HardLogLoss<br>\nLml is ArcFaceLoss metric learning supervised by antibody-id.<br>\nLre is MSELoss</p>\n<p>Augmentation<br>\nflip, transpose, scale, rotate, crop<br>\nAdding mitotic spindles with high confidence to other images to generate more positive samples of this type(lead to a boost with 0.02)<br>\nTest Time augmentation: default,flipud,fliplr,transpose.</p>\n<p>Validation<br>\nSelect 433 images in public test set which can be found in public-hpa dataset, and remove them from training set of the model.</p>\n<p>Compare the score of models<br>\nFCAN compare results<br>\nCompare the models real images<br>\nFCAN demo label1<br>\nFCAN demo label2<br>\nFCAN demo label15<br>\nThe left part of these images are the results from traditional CNN model, the middle are the results from the puzzle-cam, the right part is the results from FCAN,.<br>\nThe top part of every figure is image with positive label. the bottom is negative image.<br>\nThe number on each cell is the confidence of this cell.</p>\n<p>2.2 Swin Tranformer based cell classification model</p>\n<p>Data<br>\nCrop the cells in an image by using Cell-Segmentaion model.<br>\nThe cells were labeled to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ], this is a rule based procedure, After getting the outputs of all cells of train set from FCAN introduced above, we can give higher label value if the image probability and cell probability are high, and the cells from an image with label A were given at least 0.25 of this label A. The thresholds of the rule were not sensitive according to my experiments.</p>\n<p>Cell-Segmentaionモデルを使用して、画像内のセルをトリミングします。<br>\nセルはラベル[1.0、0.75、0.5、0.25、0]で5レベルにラベル付けされました。これはルールベースの手順です。上記で紹介したFCANからトレインセットのすべてのセルの出力を取得した後、次の場合に高いラベル値を与えることができます。 画像の確率とセルの確率は高く、ラベルAの画像のセルには、このラベルAの少なくとも0.25が与えられました。私の実験によると、ルールのしきい値は敏感ではありませんでした。</p>\n<p>Model<br>\nSwin Transformer with pretrained weights small_patch4_window7.<br>\nThe cells are resized to 128x128 px to feed into the network</p>\n<p>Loss<br>\nFocalLoss</p>\n<p>Augmentation<br>\nFlip, transpose, scale, rotate, crop<br>\nTest Time augmentation:default,flipud,fliplr,transpose.</p>\n<p>Validation Strategy<br>\nGetting the max cell confidence in an image and use this confidence as the image confidence, calculate the MAP of the image level.<br>\nAlthough this is not a strategy always keep consistency with public LB, but it can reflect the capability of the model to some extend.</p>\n<p>Inference<br>\nThe confidence of a cell is multiplication of FCAN image-level prediction and cell level Swin Transformer prediction.</p>\n<p>2.3 Ensemble<br>\nWeighted average of the prediction of FCAN and Swin transformer.</p>\n<p>2.4 Post-Processing<br>\nAs the host did not label some cells on border, if we give the cell with high confidence, the Fasle Positive cells will increase, so I trained a model to predict the completeness of a cell. If the probability to be a whole cell is very low, the confidence of this cell is multiplied by low value such as 0.3.</p>\n<p><a href=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png\" target=\"_blank\">https://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png</a><br>\nThe data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.</p>\n<p>The backbone of this model is EfficientNet-B0, 3 epochs is enough to get a quite good model.</p>\n<p>The score can improve 0.007 to 0.01 after this step.</p>\n<ol>\n<li>Results</li>\n</ol>\n<p>Results of simple solution<br>\n<a href=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png\" target=\"_blank\">https://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png</a></p>\n<p>Results of final submission<br>\nimagehttps://bestfitting.github.io/kaggle/hpa2021/figures/Complex.png</p>\n<ol>\n<li>Conclusion</li>\n</ol>\n<p>4.1 The Fair Cell Activation Network(FCAN) can increase cell level recall which is very important to this competition.</p>\n<p>4.2 The vision transformer models have shown promising capability.</p>\n<p>4.3 Larger model not always means better result as most pre-trained models are designed for ImageNet, our models should find relationship of relative position of pixels instead of abstract semantic.</p>\n<p>4.4 I found little differences among JPEG, PNG and 8bit 16bit formats.</p>",
      "rawMarkdown": "日本語訳\n\nCongrats to all the winners, and thanks to the Human Protein Atlas team and kaggle hosted such an interesting competetion.\n\n1. Introduction\n\nThe main challenge of this competition is to find a way to label every cell in a labeled image, it is a new type of weakly supervised challenge as we are provided with a cell segmentation model which means this is not a problem widely discussed like weakly supervised object detection or segmentation.\n\nThe common method for this problem is to find CAM or attention on cells, but the activations of a CNN network is focus on most discriminative parts of an image which lead to a low recall rate, to solve this problem I developed a network called Fair Cell Activation Network(FCAN) based on Puzzle-CAM.\n\nAfter getting the prediction of each cell from FCAN, I relabeled the cells to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ] by rule and trained a Swin Transformer model to predict the cell label.\n\nEnsemble of this two models and post-processing by reducing the confidence of the cells on image border can achieve the first place with 0.555 on private LB, a more complex ensemble solution with 6 models can reach 0.566 on private LB.\n\nこの競争の主な課題は、ラベル付けされた画像内のすべてのセルにラベルを付ける方法を見つけることです。これは、セルセグメンテーションモデルが提供されているため、新しいタイプの弱教師ありチャレンジです。これは、弱教師ありのように広く議論されている問題ではないことを意味します。オブジェクトの検出またはセグメンテーション。\n\nこの問題の一般的な方法は、CAMまたは細胞への注意を見つけることですが、CNNネットワークのアクティブ化は、画像の最も識別可能な部分に焦点を当てているため、リコール率が低くなります。この問題を解決するために、FairCellというネットワークを開発しました。パズルCAMに基づくアクティベーションネットワーク（FCAN）。\n\nFCANから各セルの予測を取得した後、ルールによってセルをラベル[1.0、0.75、0.5、0.25、0]で5レベルに再ラベル付けし、セルラベルを予測するようにSwinTransformerモデルをトレーニングしました。\n\nこの2つのモデルのアンサンブルと、画像境界のセルの信頼性を下げることによる後処理は、プライベートLBで0.555で最初の場所を達成でき、6つのモデルでのより複雑なアンサンブルソリューションは、プライベートLBで0.566に達することができます。\n\n2. Methods\n\n2.1 Fair Cell Activation Network\nThe activations of CNN on feature map of an image is focus on most descriminative instance of a class despite many instances exists. I call this phenomena unfair activation, to address this problem, a network was proposed based on Puzzle-CAM.\n\n画像の機能マップでのCNNのアクティブ化は、多くのインスタンスが存在するにもかかわらず、クラスの最も識別力のあるインスタンスに焦点を合わせています。 私はこの現象を不公平な活性化と呼んでいますが、この問題に対処するために、Puzzle-CAMに基づいたネットワークが提案されました。\n\nTraining\nhttps://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png\nInference\nhttps://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png\nThe main difference to Puzzle-CAM in train part of this model is we can select cells instead of splitting the image to grid.\nThe confidence of a cell is multiplication of image-level prediction and cell level prediction.\n\nDetails\n\nImages are resized to 512x512 px\n\nBackbone: EfficientNet-B0.\n\nLosses\nLcls is FocalLoss + SymmetricLovaszLoss + HardLogLoss\nLml is ArcFaceLoss metric learning supervised by antibody-id.\nLre is MSELoss\n\nAugmentation\nflip, transpose, scale, rotate, crop\nAdding mitotic spindles with high confidence to other images to generate more positive samples of this type(lead to a boost with 0.02)\nTest Time augmentation: default,flipud,fliplr,transpose.\n\nValidation\nSelect 433 images in public test set which can be found in public-hpa dataset, and remove them from training set of the model.\n\nCompare the score of models\nFCAN compare results\nCompare the models real images\nFCAN demo label1\nFCAN demo label2\nFCAN demo label15\nThe left part of these images are the results from traditional CNN model, the middle are the results from the puzzle-cam, the right part is the results from FCAN,.\nThe top part of every figure is image with positive label. the bottom is negative image.\nThe number on each cell is the confidence of this cell.\n\n2.2 Swin Tranformer based cell classification model\n\nData\nCrop the cells in an image by using Cell-Segmentaion model.\nThe cells were labeled to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ], this is a rule based procedure, After getting the outputs of all cells of train set from FCAN introduced above, we can give higher label value if the image probability and cell probability are high, and the cells from an image with label A were given at least 0.25 of this label A. The thresholds of the rule were not sensitive according to my experiments.\n\nCell-Segmentaionモデルを使用して、画像内のセルをトリミングします。\nセルはラベル[1.0、0.75、0.5、0.25、0]で5レベルにラベル付けされました。これはルールベースの手順です。上記で紹介したFCANからトレインセットのすべてのセルの出力を取得した後、次の場合に高いラベル値を与えることができます。 画像の確率とセルの確率は高く、ラベルAの画像のセルには、このラベルAの少なくとも0.25が与えられました。私の実験によると、ルールのしきい値は敏感ではありませんでした。\n\nModel\nSwin Transformer with pretrained weights small_patch4_window7.\nThe cells are resized to 128x128 px to feed into the network\n\nLoss\nFocalLoss\n\nAugmentation\nFlip, transpose, scale, rotate, crop\nTest Time augmentation:default,flipud,fliplr,transpose.\n\nValidation Strategy\nGetting the max cell confidence in an image and use this confidence as the image confidence, calculate the MAP of the image level.\nAlthough this is not a strategy always keep consistency with public LB, but it can reflect the capability of the model to some extend.\n\nInference\nThe confidence of a cell is multiplication of FCAN image-level prediction and cell level Swin Transformer prediction.\n\n2.3 Ensemble\nWeighted average of the prediction of FCAN and Swin transformer.\n\n2.4 Post-Processing\nAs the host did not label some cells on border, if we give the cell with high confidence, the Fasle Positive cells will increase, so I trained a model to predict the completeness of a cell. If the probability to be a whole cell is very low, the confidence of this cell is multiplied by low value such as 0.3.\n\nhttps://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png\nThe data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.\n\nThe backbone of this model is EfficientNet-B0, 3 epochs is enough to get a quite good model.\n\nThe score can improve 0.007 to 0.01 after this step.\n\n3. Results\n\nResults of simple solution\nhttps://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png\n\nResults of final submission\nimagehttps://bestfitting.github.io/kaggle/hpa2021/figures/Complex.png\n\n4. Conclusion\n\n4.1 The Fair Cell Activation Network(FCAN) can increase cell level recall which is very important to this competition.\n\n4.2 The vision transformer models have shown promising capability.\n\n4.3 Larger model not always means better result as most pre-trained models are designed for ImageNet, our models should find relationship of relative position of pixels instead of abstract semantic.\n\n4.4 I found little differences among JPEG, PNG and 8bit 16bit formats."
    },
    {
      "id": 1313179,
      "postDate": "2021-05-18T12:36:58.367Z",
      "content": "<p>Congratulations on your solo win!</p>",
      "rawMarkdown": "Congratulations on your solo win!",
      "votes": 1
    },
    {
      "id": 1312517,
      "postDate": "2021-05-18T05:14:51.160Z",
      "content": "<p>Congratulations and thank you for a very detailed solution summary, I'm learning so much!</p>\n<p>I have a question related to 'unfair activation'. I had assumed this was a problem for semantic segmentation, but for classification this should be less relevant - given that CNN's operate locally up until GAP layer, they should be able to find discriminative features in every cell in an image. This was the basis of my solution, and also some others like tito's. </p>\n<p>If I understand correctly what FCAN is doing, it does some regularisation, so that a model is pushed to focus on entire cells and activate them consistently across the entire image. Is my thinking correct?</p>",
      "rawMarkdown": "Congratulations and thank you for a very detailed solution summary, I'm learning so much!\n\nI have a question related to 'unfair activation'. I had assumed this was a problem for semantic segmentation, but for classification this should be less relevant - given that CNN's operate locally up until GAP layer, they should be able to find discriminative features in every cell in an image. This was the basis of my solution, and also some others like tito's. \n\nIf I understand correctly what FCAN is doing, it does some regularisation, so that a model is pushed to focus on entire cells and activate them consistently across the entire image. Is my thinking correct?",
      "votes": 1,
      "replies": [
        {
          "id": 1313975,
          "postDate": "2021-05-18T20:42:54.173Z",
          "content": "<p>Hi thedrcat,</p>\n<p>The activation maps/ CAMs can not find all the cells, they will focus on part of cells, this is the key problem with a CNN model.  What puzzle-cam and my network trying to do is force the network not so focus on the most discriminative part.</p>\n<p>You can visualize the activations of an image, you will find out the truth.</p>",
          "rawMarkdown": "Hi thedrcat,\n\nThe activation maps/ CAMs can not find all the cells, they will focus on part of cells, this is the key problem with a CNN model.  What puzzle-cam and my network trying to do is force the network not so focus on the most discriminative part.\n\nYou can visualize the activations of an image, you will find out the truth.\n\n",
          "votes": 1
        },
        {
          "id": 1333862,
          "postDate": "2021-06-03T05:24:40.043Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a>, I've done some visualizations as recommended. For a specific image and class, I plot the raw image, activation map, activation map when running the same model on individual cells and applying logits to cell mask, and the same activation map but taking rank order of specific class rather than raw logits. </p>\n<p>The unfair activation is definitely visible in the second column (raw activation map) but it is less salient when running inference on individual cells, especially when considering relative probabilities vs. other classes. For that reason, I thought the main benefit of FCAN model is regularization.</p>\n<p><img src=\"https://pbs.twimg.com/media/E27wfEBXIAEz0s4?format=jpg&amp;name=medium\" alt=\"visualization\"></p>",
          "rawMarkdown": "Hello @bestfitting, I've done some visualizations as recommended. For a specific image and class, I plot the raw image, activation map, activation map when running the same model on individual cells and applying logits to cell mask, and the same activation map but taking rank order of specific class rather than raw logits. \n\nThe unfair activation is definitely visible in the second column (raw activation map) but it is less salient when running inference on individual cells, especially when considering relative probabilities vs. other classes. For that reason, I thought the main benefit of FCAN model is regularization.\n\n![visualization](https://pbs.twimg.com/media/E27wfEBXIAEz0s4?format=jpg&name=medium)"
        },
        {
          "id": 1334043,
          "postDate": "2021-06-03T08:41:36.663Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a>,<br>\nNice to discuss with you!<br>\nI update my post with some visualizations, to get these cell probabilities, I forwarded every cell to the network. The confidences were ranked and then mix-max normalized to [0-1] on each class.</p>\n<p>By the way, there is a problem with grad-cam and other similar methods, the activations or the CAM can not compared between images.  </p>\n<p>And, we need not compare the prob with other class, the order of the confidences to be a class is important.</p>\n<p>As to the regularization, I think it play some role as we force the  network activate correct part of each cell.</p>",
          "rawMarkdown": "Hi @thedrcat,\nNice to discuss with you!\nI update my post with some visualizations, to get these cell probabilities, I forwarded every cell to the network. The confidences were ranked and then mix-max normalized to [0-1] on each class.\n\nBy the way, there is a problem with grad-cam and other similar methods, the activations or the CAM can not compared between images.  \n\nAnd, we need not compare the prob with other class, the order of the confidences to be a class is important.\n\nAs to the regularization, I think it play some role as we force the  network activate correct part of each cell.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1309164,
      "postDate": "2021-05-15T18:04:59.117Z",
      "content": "<p>Congrats for winning HPA twice with big margin</p>",
      "rawMarkdown": "Congrats for winning HPA twice with big margin",
      "votes": 1
    },
    {
      "id": 1307664,
      "postDate": "2021-05-14T15:35:00.640Z",
      "content": "<p>Congratulations on defending championship of HPA challenge. This solution is quite different than your previous one. Unlike last one, Densenet was not used here. But the loss function is reused. FCAN is an useful discovery!</p>\n<p>Two questions on the classification model:</p>\n<ol>\n<li>For Swin transformer model, did you assign class labels based on OOF set or the training set? </li>\n<li>Did you filter out the wrong predictions based on the true image level labels or kept everything?</li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratulations on defending championship of HPA challenge. This solution is quite different than your previous one. Unlike last one, Densenet was not used here. But the loss function is reused. FCAN is an useful discovery!\n\nTwo questions on the classification model:\n1. For Swin transformer model, did you assign class labels based on OOF set or the training set? \n2. Did you filter out the wrong predictions based on the true image level labels or kept everything?\n\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 1308581,
          "postDate": "2021-05-15T10:11:30.627Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sgalib\" target=\"_blank\">@sgalib</a> <br>\nThanks!<br>\nAs to you questions:</p>\n<ol>\n<li>Using OOF and training set perform similar but we should set different thresholds, I used training set.</li>\n<li>I trust image level labels.</li>\n</ol>",
          "rawMarkdown": "Hi @sgalib \nThanks!\nAs to you questions:\n1. Using OOF and training set perform similar but we should set different thresholds, I used training set.\n2. I trust image level labels.\n",
          "votes": 1
        },
        {
          "id": 1309077,
          "postDate": "2021-05-15T16:40:08.433Z",
          "content": "<p>Thank you very much!</p>",
          "rawMarkdown": "Thank you very much!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1307357,
      "postDate": "2021-05-14T11:59:10.343Z",
      "content": "<p>Congrats and thank you for posting solution. FCAM training procedure is clever! It's how we can utilize pre-computed segmentation information with puzzle cam idea!</p>",
      "rawMarkdown": "Congrats and thank you for posting solution. FCAM training procedure is clever! It's how we can utilize pre-computed segmentation information with puzzle cam idea!",
      "votes": 1
    },
    {
      "id": 1307286,
      "postDate": "2021-05-14T11:06:28.460Z",
      "content": "<p>Congratulations! Very valuable sharing! Sincere thanks!</p>",
      "rawMarkdown": "Congratulations! Very valuable sharing! Sincere thanks!",
      "votes": 1
    },
    {
      "id": 1313517,
      "postDate": "2021-05-18T15:40:23.903Z",
      "content": "<p>Congratulations on the win! I enjoyed your write-up/recap of your implementation. I've actually parsed through several of your write-ups and have found them all very informative. I appreciate you sharing your approaches and techniques, they're definitely helping to expand my knowledge.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratulations on the win! I enjoyed your write-up/recap of your implementation. I've actually parsed through several of your write-ups and have found them all very informative. I appreciate you sharing your approaches and techniques, they're definitely helping to expand my knowledge.\n\nThanks!",
      "votes": 2
    },
    {
      "id": 1308355,
      "postDate": "2021-05-15T06:26:28.410Z",
      "content": "<p>Congratulations for winning HPA again!</p>\n<p>I've got a question: where do you get the antibody-id information?</p>\n<p>I searched around in the <a href=\"https://www.proteinatlas.org/\" target=\"_blank\">https://www.proteinatlas.org/</a> but just cannot find it.</p>",
      "rawMarkdown": "Congratulations for winning HPA again!\n\nI've got a question: where do you get the antibody-id information?\n\nI searched around in the https://www.proteinatlas.org/ but just cannot find it.",
      "votes": 2,
      "replies": [
        {
          "id": 1308869,
          "postDate": "2021-05-15T14:11:37.347Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>,</p>\n<p>There is a link <a href=\"https://www.proteinatlas.org/about/download\" target=\"_blank\">https://www.proteinatlas.org/about/download</a>, you can download <a href=\"https://www.proteinatlas.org/download/subcellular_location.tsv.zip\" target=\"_blank\">https://www.proteinatlas.org/download/subcellular_location.tsv.zip</a> </p>\n<p>The first column is Gene, we can download related xml for every Gene, for example:<br>\n<a href=\"https://www.proteinatlas.org/ENSG00000134057.xml\" target=\"_blank\">https://www.proteinatlas.org/ENSG00000134057.xml</a></p>\n<p>There are a lot of information, we need  background knowledge to understand them all. But you can search antibody id= you will find one or many antibody for this gene. For exmaple:<br>\n<strong>antibody id=\"CAB000115\"</strong><br>\nFor every antibody, there are images, please search imageUrl, for example, <strong>http://images.proteinatlas.org/115/672_E2_1_blue_red_green.jpg</strong><br>\nyou will find the antibody-id(115) in the link, and you can also find the image_id(672_E2_1), then, we can add antibody-id attribute to this image, you can use other information, such as cell-line, age… </p>\n<p>Perhaps you did not enter the last HPA's competition, there are discussions on how to use information on <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860\" target=\"_blank\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860</a><br>\nfor example:<br>\n<a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/430860/10777/Parce%20XML%20and%20Download%20HPAv18%20Image.html\" target=\"_blank\">https://storage.googleapis.com/kaggle-forum-message-attachments/430860/10777/Parce%20XML%20and%20Download%20HPAv18%20Image.html</a></p>",
          "rawMarkdown": "Hi @haqishen,\n\nThere is a link https://www.proteinatlas.org/about/download, you can download https://www.proteinatlas.org/download/subcellular_location.tsv.zip \n\nThe first column is Gene, we can download related xml for every Gene, for example:\nhttps://www.proteinatlas.org/ENSG00000134057.xml\n\nThere are a lot of information, we need  background knowledge to understand them all. But you can search antibody id= you will find one or many antibody for this gene. For exmaple:\n**antibody id=\"CAB000115\"**\nFor every antibody, there are images, please search imageUrl, for example, <imageUrl>**http://images.proteinatlas.org/115/672_E2_1_blue_red_green.jpg**</imageUrl>\nyou will find the antibody-id(115) in the link, and you can also find the image_id(672_E2_1), then, we can add antibody-id attribute to this image, you can use other information, such as cell-line, age... \n\nPerhaps you did not enter the last HPA's competition, there are discussions on how to use information on https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860\nfor example:\nhttps://storage.googleapis.com/kaggle-forum-message-attachments/430860/10777/Parce%20XML%20and%20Download%20HPAv18%20Image.html\n\n\n\n",
          "votes": 5
        },
        {
          "id": 1308951,
          "postDate": "2021-05-15T14:55:46.553Z",
          "content": "<p>Thanks for the detail information, congratulation again for the great achievement! 👍</p>",
          "rawMarkdown": "Thanks for the detail information, congratulation again for the great achievement! 👍",
          "votes": 2
        }
      ]
    },
    {
      "id": 1307329,
      "postDate": "2021-05-14T11:32:29.310Z",
      "content": "<p>Waited last 2 weeks just for this write-up.<br>\nAmazing work again. <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> </p>",
      "rawMarkdown": "Waited last 2 weeks just for this write-up.\nAmazing work again. @bestfitting ",
      "votes": 2
    },
    {
      "id": 1307259,
      "postDate": "2021-05-14T10:53:12.157Z",
      "content": "<p>Congratulations. Your work is always so meticulous, sincerely admire.</p>",
      "rawMarkdown": "Congratulations. Your work is always so meticulous, sincerely admire.",
      "votes": 2
    },
    {
      "id": 1307256,
      "postDate": "2021-05-14T10:52:29.690Z",
      "content": "<p>Congratulations! My solution looks like a draft of yours (for initial steps) 🙂</p>",
      "rawMarkdown": "Congratulations! My solution looks like a draft of yours (for initial steps) 🙂",
      "votes": 2,
      "replies": [
        {
          "id": 1308510,
          "postDate": "2021-05-15T09:14:53.280Z",
          "content": "<p>Thanks, glad to find that our solutions are so similar. :)  </p>",
          "rawMarkdown": "Thanks, glad to find that our solutions are so similar. :)  ",
          "votes": 1
        },
        {
          "id": 1308519,
          "postDate": "2021-05-15T09:21:10.653Z",
          "content": "<p>Yes, but you've designed something really better and clever. <br>\nLessons learned for me: Spend more time on overall solution design at the beginning.</p>",
          "rawMarkdown": "Yes, but you've designed something really better and clever. \nLessons learned for me: Spend more time on overall solution design at the beginning.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1307177,
      "postDate": "2021-05-14T09:52:36.707Z",
      "content": "<p>Awesome solution! I also tried and increased the number of grid tiles for Puzzle-CAM but your way is probably the most intuitive to train for nice CAMs</p>",
      "rawMarkdown": "Awesome solution! I also tried and increased the number of grid tiles for Puzzle-CAM but your way is probably the most intuitive to train for nice CAMs",
      "votes": 2
    },
    {
      "id": 1534380,
      "postDate": "2021-10-04T20:47:20.880Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a>, Did you use one image at a time in training? or a batch?</p>",
      "rawMarkdown": "Hi @bestfitting, Did you use one image at a time in training? or a batch?",
      "replies": [
        {
          "id": 1545470,
          "postDate": "2021-10-15T08:42:28.173Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/micheomaano\" target=\"_blank\">@micheomaano</a>,<br>\nSorry for late reply, I did not open kaggle pages often..<br>\nWe can feed the images to the network in batch…As a training image-&gt;5 input images for FCAN, we can feed 2,3,4 training time to the network at same time, and the real batch-size will be 10,15,20.</p>",
          "rawMarkdown": "Hi @micheomaano,\nSorry for late reply, I did not open kaggle pages often..\nWe can feed the images to the network in batch...As a training image->5 input images for FCAN, we can feed 2,3,4 training time to the network at same time, and the real batch-size will be 10,15,20.\n\n"
        }
      ]
    },
    {
      "id": 1389078,
      "postDate": "2021-07-15T12:48:17.333Z",
      "content": "<p><a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> Congrats on winning the competition and thank you for a very detailed solution summary.<br>\nCould you please provide details about the relabelling of cells to 5 levels for Swin transformer. What exactly are the rules and thresholds for label assignments. <br>\nThanks very much.</p>",
      "rawMarkdown": "@bestfitting Congrats on winning the competition and thank you for a very detailed solution summary.\nCould you please provide details about the relabelling of cells to 5 levels for Swin transformer. What exactly are the rules and thresholds for label assignments. \nThanks very much.",
      "replies": [
        {
          "id": 1545469,
          "postDate": "2021-10-15T08:42:25.343Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sameedhusain\" target=\"_blank\">@sameedhusain</a>,<br>\nSorry for late reply!<br>\nIt's not easy to describe in detail, but the rule is quite easy to understanding, if the cell prediction from FCAN is high and the image level prediction is also high, then, the cell level label will be high…we can set the confidence of image level and cell level prediction to 2-3 levels and then set the cell level label accordingly. What's more, we should assign a label at least 0.25 if the label exists in the image level labels.</p>",
          "rawMarkdown": "Hi @sameedhusain,\nSorry for late reply!\nIt's not easy to describe in detail, but the rule is quite easy to understanding, if the cell prediction from FCAN is high and the image level prediction is also high, then, the cell level label will be high...we can set the confidence of image level and cell level prediction to 2-3 levels and then set the cell level label accordingly. What's more, we should assign a label at least 0.25 if the label exists in the image level labels."
        }
      ]
    },
    {
      "id": 1362472,
      "postDate": "2021-06-23T13:08:19.300Z",
      "content": "<p>Congratulations!  Thank you for sharing the solution to this challenge, and to the last one.</p>\n<p>Can you please give a bit more detail about how to identify a border cell?  </p>\n<blockquote>\n  <p>The data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.</p>\n</blockquote>\n<p>Do you mean that the model takes as its input an image with a cell, or part of a cell, inside it, and outputs the area (the number of pixels) occupied by the cell?  If so, how can you tell from this output whether the cell is a border cell or a whole cell?  Or, is the output a number between 0 and 1, where 1 means the cell is complete?</p>",
      "rawMarkdown": "Congratulations!  Thank you for sharing the solution to this challenge, and to the last one.\n\nCan you please give a bit more detail about how to identify a border cell?  \n> The data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.\n\nDo you mean that the model takes as its input an image with a cell, or part of a cell, inside it, and outputs the area (the number of pixels) occupied by the cell?  If so, how can you tell from this output whether the cell is a border cell or a whole cell?  Or, is the output a number between 0 and 1, where 1 means the cell is complete?",
      "replies": [
        {
          "id": 1365984,
          "postDate": "2021-06-26T10:32:08.143Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jackchungchiehyu\" target=\"_blank\">@jackchungchiehyu</a>, yes, the output is between 0 and 1, if the value&gt;0.98, it's a border cell, if the value &lt;0.1 or &lt;0.2, we should decrease the confidence of the cell.</p>",
          "rawMarkdown": "Hi @jackchungchiehyu, yes, the output is between 0 and 1, if the value>0.98, it's a border cell, if the value <0.1 or <0.2, we should decrease the confidence of the cell.\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1367673,
      "postDate": "2021-06-28T03:01:49.463Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1495841,
      "postDate": "2021-08-29T20:38:21.317Z",
      "content": "<p>Thanks for sharing!! very helpful</p>",
      "rawMarkdown": "Thanks for sharing!! very helpful\n\n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1307315,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-05-14T11:25:24.897000",
      "content": "<p>Awesome, congrats on being back to #1!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1308576,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-05-15T10:05:57.060000",
          "content": "<p>Thanks, I am sure you or guanshuo <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> will be back to #1 very soon, and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> is also on the way to this position :)</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1308073,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-05-15T00:41:49.230000",
      "content": "<p>Congratulations winning HPA two years in a row even after posting your solution from the first win!</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1610714,
      "author_name": "pixyz0130",
      "author_url": "",
      "post_date": "2021-12-07T13:17:55.203000",
      "content": "<p>日本語訳</p>\n<p>Congrats to all the winners, and thanks to the Human Protein Atlas team and kaggle hosted such an interesting competetion.</p>\n<ol>\n<li>Introduction</li>\n</ol>\n<p>The main challenge of this competition is to find a way to label every cell in a labeled image, it is a new type of weakly supervised challenge as we are provided with a cell segmentation model which means this is not a problem widely discussed like weakly supervised object detection or segmentation.</p>\n<p>The common method for this problem is to find CAM or attention on cells, but the activations of a CNN network is focus on most discriminative parts of an image which lead to a low recall rate, to solve this problem I developed a network called Fair Cell Activation Network(FCAN) based on Puzzle-CAM.</p>\n<p>After getting the prediction of each cell from FCAN, I relabeled the cells to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ] by rule and trained a Swin Transformer model to predict the cell label.</p>\n<p>Ensemble of this two models and post-processing by reducing the confidence of the cells on image border can achieve the first place with 0.555 on private LB, a more complex ensemble solution with 6 models can reach 0.566 on private LB.</p>\n<p>この競争の主な課題は、ラベル付けされた画像内のすべてのセルにラベルを付ける方法を見つけることです。これは、セルセグメンテーションモデルが提供されているため、新しいタイプの弱教師ありチャレンジです。これは、弱教師ありのように広く議論されている問題ではないことを意味します。オブジェクトの検出またはセグメンテーション。</p>\n<p>この問題の一般的な方法は、CAMまたは細胞への注意を見つけることですが、CNNネットワークのアクティブ化は、画像の最も識別可能な部分に焦点を当てているため、リコール率が低くなります。この問題を解決するために、FairCellというネットワークを開発しました。パズルCAMに基づくアクティベーションネットワーク（FCAN）。</p>\n<p>FCANから各セルの予測を取得した後、ルールによってセルをラベル[1.0、0.75、0.5、0.25、0]で5レベルに再ラベル付けし、セルラベルを予測するようにSwinTransformerモデルをトレーニングしました。</p>\n<p>この2つのモデルのアンサンブルと、画像境界のセルの信頼性を下げることによる後処理は、プライベートLBで0.555で最初の場所を達成でき、6つのモデルでのより複雑なアンサンブルソリューションは、プライベートLBで0.566に達することができます。</p>\n<ol>\n<li>Methods</li>\n</ol>\n<p>2.1 Fair Cell Activation Network<br>\nThe activations of CNN on feature map of an image is focus on most descriminative instance of a class despite many instances exists. I call this phenomena unfair activation, to address this problem, a network was proposed based on Puzzle-CAM.</p>\n<p>画像の機能マップでのCNNのアクティブ化は、多くのインスタンスが存在するにもかかわらず、クラスの最も識別力のあるインスタンスに焦点を合わせています。 私はこの現象を不公平な活性化と呼んでいますが、この問題に対処するために、Puzzle-CAMに基づいたネットワークが提案されました。</p>\n<p>Training<br>\n<a href=\"https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png\" target=\"_blank\">https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png</a><br>\nInference<br>\n<a href=\"https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png\" target=\"_blank\">https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png</a><br>\nThe main difference to Puzzle-CAM in train part of this model is we can select cells instead of splitting the image to grid.<br>\nThe confidence of a cell is multiplication of image-level prediction and cell level prediction.</p>\n<p>Details</p>\n<p>Images are resized to 512x512 px</p>\n<p>Backbone: EfficientNet-B0.</p>\n<p>Losses<br>\nLcls is FocalLoss + SymmetricLovaszLoss + HardLogLoss<br>\nLml is ArcFaceLoss metric learning supervised by antibody-id.<br>\nLre is MSELoss</p>\n<p>Augmentation<br>\nflip, transpose, scale, rotate, crop<br>\nAdding mitotic spindles with high confidence to other images to generate more positive samples of this type(lead to a boost with 0.02)<br>\nTest Time augmentation: default,flipud,fliplr,transpose.</p>\n<p>Validation<br>\nSelect 433 images in public test set which can be found in public-hpa dataset, and remove them from training set of the model.</p>\n<p>Compare the score of models<br>\nFCAN compare results<br>\nCompare the models real images<br>\nFCAN demo label1<br>\nFCAN demo label2<br>\nFCAN demo label15<br>\nThe left part of these images are the results from traditional CNN model, the middle are the results from the puzzle-cam, the right part is the results from FCAN,.<br>\nThe top part of every figure is image with positive label. the bottom is negative image.<br>\nThe number on each cell is the confidence of this cell.</p>\n<p>2.2 Swin Tranformer based cell classification model</p>\n<p>Data<br>\nCrop the cells in an image by using Cell-Segmentaion model.<br>\nThe cells were labeled to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ], this is a rule based procedure, After getting the outputs of all cells of train set from FCAN introduced above, we can give higher label value if the image probability and cell probability are high, and the cells from an image with label A were given at least 0.25 of this label A. The thresholds of the rule were not sensitive according to my experiments.</p>\n<p>Cell-Segmentaionモデルを使用して、画像内のセルをトリミングします。<br>\nセルはラベル[1.0、0.75、0.5、0.25、0]で5レベルにラベル付けされました。これはルールベースの手順です。上記で紹介したFCANからトレインセットのすべてのセルの出力を取得した後、次の場合に高いラベル値を与えることができます。 画像の確率とセルの確率は高く、ラベルAの画像のセルには、このラベルAの少なくとも0.25が与えられました。私の実験によると、ルールのしきい値は敏感ではありませんでした。</p>\n<p>Model<br>\nSwin Transformer with pretrained weights small_patch4_window7.<br>\nThe cells are resized to 128x128 px to feed into the network</p>\n<p>Loss<br>\nFocalLoss</p>\n<p>Augmentation<br>\nFlip, transpose, scale, rotate, crop<br>\nTest Time augmentation:default,flipud,fliplr,transpose.</p>\n<p>Validation Strategy<br>\nGetting the max cell confidence in an image and use this confidence as the image confidence, calculate the MAP of the image level.<br>\nAlthough this is not a strategy always keep consistency with public LB, but it can reflect the capability of the model to some extend.</p>\n<p>Inference<br>\nThe confidence of a cell is multiplication of FCAN image-level prediction and cell level Swin Transformer prediction.</p>\n<p>2.3 Ensemble<br>\nWeighted average of the prediction of FCAN and Swin transformer.</p>\n<p>2.4 Post-Processing<br>\nAs the host did not label some cells on border, if we give the cell with high confidence, the Fasle Positive cells will increase, so I trained a model to predict the completeness of a cell. If the probability to be a whole cell is very low, the confidence of this cell is multiplied by low value such as 0.3.</p>\n<p><a href=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png\" target=\"_blank\">https://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png</a><br>\nThe data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.</p>\n<p>The backbone of this model is EfficientNet-B0, 3 epochs is enough to get a quite good model.</p>\n<p>The score can improve 0.007 to 0.01 after this step.</p>\n<ol>\n<li>Results</li>\n</ol>\n<p>Results of simple solution<br>\n<a href=\"https://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png\" target=\"_blank\">https://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png</a></p>\n<p>Results of final submission<br>\nimagehttps://bestfitting.github.io/kaggle/hpa2021/figures/Complex.png</p>\n<ol>\n<li>Conclusion</li>\n</ol>\n<p>4.1 The Fair Cell Activation Network(FCAN) can increase cell level recall which is very important to this competition.</p>\n<p>4.2 The vision transformer models have shown promising capability.</p>\n<p>4.3 Larger model not always means better result as most pre-trained models are designed for ImageNet, our models should find relationship of relative position of pixels instead of abstract semantic.</p>\n<p>4.4 I found little differences among JPEG, PNG and 8bit 16bit formats.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1313179,
      "author_name": "Jyot Makadiya",
      "author_url": "",
      "post_date": "2021-05-18T12:36:58.367000",
      "content": "<p>Congratulations on your solo win!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1312517,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-05-18T05:14:51.160000",
      "content": "<p>Congratulations and thank you for a very detailed solution summary, I'm learning so much!</p>\n<p>I have a question related to 'unfair activation'. I had assumed this was a problem for semantic segmentation, but for classification this should be less relevant - given that CNN's operate locally up until GAP layer, they should be able to find discriminative features in every cell in an image. This was the basis of my solution, and also some others like tito's. </p>\n<p>If I understand correctly what FCAN is doing, it does some regularisation, so that a model is pushed to focus on entire cells and activate them consistently across the entire image. Is my thinking correct?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1313975,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-05-18T20:42:54.173000",
          "content": "<p>Hi thedrcat,</p>\n<p>The activation maps/ CAMs can not find all the cells, they will focus on part of cells, this is the key problem with a CNN model.  What puzzle-cam and my network trying to do is force the network not so focus on the most discriminative part.</p>\n<p>You can visualize the activations of an image, you will find out the truth.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1333862,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-06-03T05:24:40.043000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a>, I've done some visualizations as recommended. For a specific image and class, I plot the raw image, activation map, activation map when running the same model on individual cells and applying logits to cell mask, and the same activation map but taking rank order of specific class rather than raw logits. </p>\n<p>The unfair activation is definitely visible in the second column (raw activation map) but it is less salient when running inference on individual cells, especially when considering relative probabilities vs. other classes. For that reason, I thought the main benefit of FCAN model is regularization.</p>\n<p><img src=\"https://pbs.twimg.com/media/E27wfEBXIAEz0s4?format=jpg&amp;name=medium\" alt=\"visualization\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1334043,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-06-03T08:41:36.663000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a>,<br>\nNice to discuss with you!<br>\nI update my post with some visualizations, to get these cell probabilities, I forwarded every cell to the network. The confidences were ranked and then mix-max normalized to [0-1] on each class.</p>\n<p>By the way, there is a problem with grad-cam and other similar methods, the activations or the CAM can not compared between images.  </p>\n<p>And, we need not compare the prob with other class, the order of the confidences to be a class is important.</p>\n<p>As to the regularization, I think it play some role as we force the  network activate correct part of each cell.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1309164,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2021-05-15T18:04:59.117000",
      "content": "<p>Congrats for winning HPA twice with big margin</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1307664,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2021-05-14T15:35:00.640000",
      "content": "<p>Congratulations on defending championship of HPA challenge. This solution is quite different than your previous one. Unlike last one, Densenet was not used here. But the loss function is reused. FCAN is an useful discovery!</p>\n<p>Two questions on the classification model:</p>\n<ol>\n<li>For Swin transformer model, did you assign class labels based on OOF set or the training set? </li>\n<li>Did you filter out the wrong predictions based on the true image level labels or kept everything?</li>\n</ol>\n<p>Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1308581,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-05-15T10:11:30.627000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sgalib\" target=\"_blank\">@sgalib</a> <br>\nThanks!<br>\nAs to you questions:</p>\n<ol>\n<li>Using OOF and training set perform similar but we should set different thresholds, I used training set.</li>\n<li>I trust image level labels.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1309077,
          "author_name": "Shai",
          "author_url": "",
          "post_date": "2021-05-15T16:40:08.433000",
          "content": "<p>Thank you very much!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1307357,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-14T11:59:10.343000",
      "content": "<p>Congrats and thank you for posting solution. FCAM training procedure is clever! It's how we can utilize pre-computed segmentation information with puzzle cam idea!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1307286,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2021-05-14T11:06:28.460000",
      "content": "<p>Congratulations! Very valuable sharing! Sincere thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1313517,
      "author_name": "JMB",
      "author_url": "",
      "post_date": "2021-05-18T15:40:23.903000",
      "content": "<p>Congratulations on the win! I enjoyed your write-up/recap of your implementation. I've actually parsed through several of your write-ups and have found them all very informative. I appreciate you sharing your approaches and techniques, they're definitely helping to expand my knowledge.</p>\n<p>Thanks!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1308355,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2021-05-15T06:26:28.410000",
      "content": "<p>Congratulations for winning HPA again!</p>\n<p>I've got a question: where do you get the antibody-id information?</p>\n<p>I searched around in the <a href=\"https://www.proteinatlas.org/\" target=\"_blank\">https://www.proteinatlas.org/</a> but just cannot find it.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1308869,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-05-15T14:11:37.347000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a>,</p>\n<p>There is a link <a href=\"https://www.proteinatlas.org/about/download\" target=\"_blank\">https://www.proteinatlas.org/about/download</a>, you can download <a href=\"https://www.proteinatlas.org/download/subcellular_location.tsv.zip\" target=\"_blank\">https://www.proteinatlas.org/download/subcellular_location.tsv.zip</a> </p>\n<p>The first column is Gene, we can download related xml for every Gene, for example:<br>\n<a href=\"https://www.proteinatlas.org/ENSG00000134057.xml\" target=\"_blank\">https://www.proteinatlas.org/ENSG00000134057.xml</a></p>\n<p>There are a lot of information, we need  background knowledge to understand them all. But you can search antibody id= you will find one or many antibody for this gene. For exmaple:<br>\n<strong>antibody id=\"CAB000115\"</strong><br>\nFor every antibody, there are images, please search imageUrl, for example, <strong>http://images.proteinatlas.org/115/672_E2_1_blue_red_green.jpg</strong><br>\nyou will find the antibody-id(115) in the link, and you can also find the image_id(672_E2_1), then, we can add antibody-id attribute to this image, you can use other information, such as cell-line, age… </p>\n<p>Perhaps you did not enter the last HPA's competition, there are discussions on how to use information on <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860\" target=\"_blank\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#430860</a><br>\nfor example:<br>\n<a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/430860/10777/Parce%20XML%20and%20Download%20HPAv18%20Image.html\" target=\"_blank\">https://storage.googleapis.com/kaggle-forum-message-attachments/430860/10777/Parce%20XML%20and%20Download%20HPAv18%20Image.html</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1308951,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2021-05-15T14:55:46.553000",
          "content": "<p>Thanks for the detail information, congratulation again for the great achievement! 👍</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1307329,
      "author_name": "Salman",
      "author_url": "",
      "post_date": "2021-05-14T11:32:29.310000",
      "content": "<p>Waited last 2 weeks just for this write-up.<br>\nAmazing work again. <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1307259,
      "author_name": "Correlation",
      "author_url": "",
      "post_date": "2021-05-14T10:53:12.157000",
      "content": "<p>Congratulations. Your work is always so meticulous, sincerely admire.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1307256,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-05-14T10:52:29.690000",
      "content": "<p>Congratulations! My solution looks like a draft of yours (for initial steps) 🙂</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1308510,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-05-15T09:14:53.280000",
          "content": "<p>Thanks, glad to find that our solutions are so similar. :)  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1308519,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-05-15T09:21:10.653000",
          "content": "<p>Yes, but you've designed something really better and clever. <br>\nLessons learned for me: Spend more time on overall solution design at the beginning.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1307177,
      "author_name": "Alexander Riedel",
      "author_url": "",
      "post_date": "2021-05-14T09:52:36.707000",
      "content": "<p>Awesome solution! I also tried and increased the number of grid tiles for Puzzle-CAM but your way is probably the most intuitive to train for nice CAMs</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1534380,
      "author_name": "Salman",
      "author_url": "",
      "post_date": "2021-10-04T20:47:20.880000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a>, Did you use one image at a time in training? or a batch?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1545470,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-10-15T08:42:28.173000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/micheomaano\" target=\"_blank\">@micheomaano</a>,<br>\nSorry for late reply, I did not open kaggle pages often..<br>\nWe can feed the images to the network in batch…As a training image-&gt;5 input images for FCAN, we can feed 2,3,4 training time to the network at same time, and the real batch-size will be 10,15,20.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389078,
      "author_name": "SAMEED",
      "author_url": "",
      "post_date": "2021-07-15T12:48:17.333000",
      "content": "<p><a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> Congrats on winning the competition and thank you for a very detailed solution summary.<br>\nCould you please provide details about the relabelling of cells to 5 levels for Swin transformer. What exactly are the rules and thresholds for label assignments. <br>\nThanks very much.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1545469,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-10-15T08:42:25.343000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sameedhusain\" target=\"_blank\">@sameedhusain</a>,<br>\nSorry for late reply!<br>\nIt's not easy to describe in detail, but the rule is quite easy to understanding, if the cell prediction from FCAN is high and the image level prediction is also high, then, the cell level label will be high…we can set the confidence of image level and cell level prediction to 2-3 levels and then set the cell level label accordingly. What's more, we should assign a label at least 0.25 if the label exists in the image level labels.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1362472,
      "author_name": "wafflebufflo",
      "author_url": "",
      "post_date": "2021-06-23T13:08:19.300000",
      "content": "<p>Congratulations!  Thank you for sharing the solution to this challenge, and to the last one.</p>\n<p>Can you please give a bit more detail about how to identify a border cell?  </p>\n<blockquote>\n  <p>The data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.</p>\n</blockquote>\n<p>Do you mean that the model takes as its input an image with a cell, or part of a cell, inside it, and outputs the area (the number of pixels) occupied by the cell?  If so, how can you tell from this output whether the cell is a border cell or a whole cell?  Or, is the output a number between 0 and 1, where 1 means the cell is complete?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1365984,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2021-06-26T10:32:08.143000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jackchungchiehyu\" target=\"_blank\">@jackchungchiehyu</a>, yes, the output is between 0 and 1, if the value&gt;0.98, it's a border cell, if the value &lt;0.1 or &lt;0.2, we should decrease the confidence of the cell.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1367673,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-28T03:01:49.463000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1495841,
      "author_name": "Fuco",
      "author_url": "",
      "post_date": "2021-08-29T20:38:21.317000",
      "content": "<p>Thanks for sharing!! very helpful</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1307108": "Congrats to all the winners, and thanks to the Human Protein Atlas team and kaggle hosted such an interesting competetion.\n\n**1. Introduction**\n\nThe main challenge of this competition is to find a way to label every cell in a labeled image, it is a new type of weakly supervised challenge as we are provided with a cell segmentation model which means this is not a problem widely discussed like weakly supervised object detection or segmentation.\n\nThe common method for this problem is to find CAM or attention on cells, but the activations of a CNN network is focus on most discriminative parts of an image which lead to a low recall rate, to solve this problem I developed a network called Fair Cell Activation Network(FCAN) based on Puzzle-CAM.\n\nAfter getting the prediction of each cell from FCAN, I relabeled the cells to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ] by rule and trained a Swin Transformer model to predict the cell label. \n\nEnsemble of this two models and post-processing by reducing the confidence of the cells on image border can achieve the first place with 0.555 on private LB, a more complex ensemble solution with 6 models can reach 0.566 on private LB.\n\n**2. Methods**\n\n**2.1 Fair Cell Activation Network**\nThe activations of CNN on feature map of an image is focus on most descriminative instance of a class despite many instances exists. I call this phenomena unfair activation, to address this problem, a network was proposed based on Puzzle-CAM.\n\n**Training**\n![FCAN train](https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png)\n**Inference**\n![FCAN inference](https://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png)\nThe main difference to Puzzle-CAM in train part of this model is we can select cells instead of splitting the image to grid.\nThe confidence of a cell is multiplication of image-level prediction and cell level prediction. \n\n**Details**\n\nImages are resized to 512x512 px\n\n**Backbone**: EfficientNet-B0.\n\n**Losses**\nLcls is FocalLoss + SymmetricLovaszLoss + HardLogLoss\nLml is ArcFaceLoss metric learning supervised by antibody-id.\nLre is MSELoss\n\n**Augmentation**\nflip, transpose, scale, rotate, crop\nAdding mitotic spindles with high confidence to other images to generate more positive samples of this type(lead to a boost with 0.02)\nTest Time augmentation: default,flipud,fliplr,transpose.\n\n**Validation**\nSelect 433 images in public test set which can be found in public-hpa dataset, and remove them from training set of the model.\n\n**Compare the score of models**\n![FCAN compare results](https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/CompareModels.png)\n**Compare the models real images**\n![FCAN demo label1](https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label1_compare.png)\n![FCAN demo label2](https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label2_compare.png)\n![FCAN demo label15](https://raw.githubusercontent.com/bestfitting/kaggle/master/hpa2021/figures/Label15_compare.png)\nThe left part of these images are the results from traditional CNN model,  the middle are the results from the puzzle-cam, the right part is the results from FCAN,.\nThe top part of every figure is image with positive label. the bottom is negative image.\nThe number on each cell is the confidence of this cell.\n\n\n**2.2 Swin Tranformer based cell classification model**\n\n**Data**\nCrop the cells in an image by using Cell-Segmentaion model.\nThe cells were labeled to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ], this is a rule based procedure, After getting the outputs of all cells of train set from FCAN introduced above, we can give higher label value if the image probability and cell probability are high, and the cells from an image with label A were given at least 0.25 of this label A. The thresholds of the rule were not sensitive according to my experiments.\n\n**Model**\nSwin Transformer with pretrained weights small_patch4_window7. \nThe cells are resized to 128x128 px to feed into the network\n\n**Loss**\nFocalLoss\n\n**Augmentation**\nFlip, transpose, scale, rotate, crop\nTest Time augmentation:default,flipud,fliplr,transpose.\n\n**Validation Strategy**\nGetting the max cell confidence in an image and use this confidence as the image confidence, calculate the MAP of the image level. \nAlthough this is not a strategy always keep consistency with public LB, but it can reflect the capability of the model to some extend. \n\n**Inference**\nThe confidence of a cell is multiplication of FCAN image-level prediction and cell level Swin Transformer prediction. \n\n\n**2.3 Ensemble**\nWeighted average of the prediction of FCAN and Swin transformer. \n\n**2.4 Post-Processing**\nAs the host did not label some cells on border, if we give the cell with high confidence, the Fasle Positive cells will increase, so I trained a model to predict the completeness of a cell. If the probability to be a whole cell is very low, the confidence of this cell is multiplied by low value such as 0.3.\n\n![FCAN train](https://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png)\nThe data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.\n\nThe backbone of this model is EfficientNet-B0, 3 epochs is enough to get a quite good model.\n\nThe score can improve 0.007 to 0.01 after this step.  \n\n**3. Results**\n\n**Results of simple solution**\n![Simple-Solution](https://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png)\n\n**Results of final submission**\n![Complex Solution](https://bestfitting.github.io/kaggle/hpa2021/figures/Complex.png)\n\n**4. Conclusion**\n\n4.1 The Fair Cell Activation Network(FCAN) can increase cell level recall which is very important to this competition.\n\n4.2 The vision transformer models have shown promising capability.\n\n4.3 Larger model not always means better result as most pre-trained models are designed for ImageNet, our models should find relationship of relative position of pixels instead of abstract semantic. \n\n4.4 I found little differences among JPEG, PNG  and  8bit 16bit formats.\n  \n\n",
    "1307315": "Awesome, congrats on being back to #1!",
    "1308073": "Congratulations winning HPA two years in a row even after posting your solution from the first win!",
    "1610714": "日本語訳\n\nCongrats to all the winners, and thanks to the Human Protein Atlas team and kaggle hosted such an interesting competetion.\n\n1. Introduction\n\nThe main challenge of this competition is to find a way to label every cell in a labeled image, it is a new type of weakly supervised challenge as we are provided with a cell segmentation model which means this is not a problem widely discussed like weakly supervised object detection or segmentation.\n\nThe common method for this problem is to find CAM or attention on cells, but the activations of a CNN network is focus on most discriminative parts of an image which lead to a low recall rate, to solve this problem I developed a network called Fair Cell Activation Network(FCAN) based on Puzzle-CAM.\n\nAfter getting the prediction of each cell from FCAN, I relabeled the cells to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ] by rule and trained a Swin Transformer model to predict the cell label.\n\nEnsemble of this two models and post-processing by reducing the confidence of the cells on image border can achieve the first place with 0.555 on private LB, a more complex ensemble solution with 6 models can reach 0.566 on private LB.\n\nこの競争の主な課題は、ラベル付けされた画像内のすべてのセルにラベルを付ける方法を見つけることです。これは、セルセグメンテーションモデルが提供されているため、新しいタイプの弱教師ありチャレンジです。これは、弱教師ありのように広く議論されている問題ではないことを意味します。オブジェクトの検出またはセグメンテーション。\n\nこの問題の一般的な方法は、CAMまたは細胞への注意を見つけることですが、CNNネットワークのアクティブ化は、画像の最も識別可能な部分に焦点を当てているため、リコール率が低くなります。この問題を解決するために、FairCellというネットワークを開発しました。パズルCAMに基づくアクティベーションネットワーク（FCAN）。\n\nFCANから各セルの予測を取得した後、ルールによってセルをラベル[1.0、0.75、0.5、0.25、0]で5レベルに再ラベル付けし、セルラベルを予測するようにSwinTransformerモデルをトレーニングしました。\n\nこの2つのモデルのアンサンブルと、画像境界のセルの信頼性を下げることによる後処理は、プライベートLBで0.555で最初の場所を達成でき、6つのモデルでのより複雑なアンサンブルソリューションは、プライベートLBで0.566に達することができます。\n\n2. Methods\n\n2.1 Fair Cell Activation Network\nThe activations of CNN on feature map of an image is focus on most descriminative instance of a class despite many instances exists. I call this phenomena unfair activation, to address this problem, a network was proposed based on Puzzle-CAM.\n\n画像の機能マップでのCNNのアクティブ化は、多くのインスタンスが存在するにもかかわらず、クラスの最も識別力のあるインスタンスに焦点を合わせています。 私はこの現象を不公平な活性化と呼んでいますが、この問題に対処するために、Puzzle-CAMに基づいたネットワークが提案されました。\n\nTraining\nhttps://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_train.png\nInference\nhttps://bestfitting.github.io/kaggle/hpa2021/figures/FCAN_inference.png\nThe main difference to Puzzle-CAM in train part of this model is we can select cells instead of splitting the image to grid.\nThe confidence of a cell is multiplication of image-level prediction and cell level prediction.\n\nDetails\n\nImages are resized to 512x512 px\n\nBackbone: EfficientNet-B0.\n\nLosses\nLcls is FocalLoss + SymmetricLovaszLoss + HardLogLoss\nLml is ArcFaceLoss metric learning supervised by antibody-id.\nLre is MSELoss\n\nAugmentation\nflip, transpose, scale, rotate, crop\nAdding mitotic spindles with high confidence to other images to generate more positive samples of this type(lead to a boost with 0.02)\nTest Time augmentation: default,flipud,fliplr,transpose.\n\nValidation\nSelect 433 images in public test set which can be found in public-hpa dataset, and remove them from training set of the model.\n\nCompare the score of models\nFCAN compare results\nCompare the models real images\nFCAN demo label1\nFCAN demo label2\nFCAN demo label15\nThe left part of these images are the results from traditional CNN model, the middle are the results from the puzzle-cam, the right part is the results from FCAN,.\nThe top part of every figure is image with positive label. the bottom is negative image.\nThe number on each cell is the confidence of this cell.\n\n2.2 Swin Tranformer based cell classification model\n\nData\nCrop the cells in an image by using Cell-Segmentaion model.\nThe cells were labeled to 5 levels with label [1.0, 0.75, 0.5, 0.25, 0 ], this is a rule based procedure, After getting the outputs of all cells of train set from FCAN introduced above, we can give higher label value if the image probability and cell probability are high, and the cells from an image with label A were given at least 0.25 of this label A. The thresholds of the rule were not sensitive according to my experiments.\n\nCell-Segmentaionモデルを使用して、画像内のセルをトリミングします。\nセルはラベル[1.0、0.75、0.5、0.25、0]で5レベルにラベル付けされました。これはルールベースの手順です。上記で紹介したFCANからトレインセットのすべてのセルの出力を取得した後、次の場合に高いラベル値を与えることができます。 画像の確率とセルの確率は高く、ラベルAの画像のセルには、このラベルAの少なくとも0.25が与えられました。私の実験によると、ルールのしきい値は敏感ではありませんでした。\n\nModel\nSwin Transformer with pretrained weights small_patch4_window7.\nThe cells are resized to 128x128 px to feed into the network\n\nLoss\nFocalLoss\n\nAugmentation\nFlip, transpose, scale, rotate, crop\nTest Time augmentation:default,flipud,fliplr,transpose.\n\nValidation Strategy\nGetting the max cell confidence in an image and use this confidence as the image confidence, calculate the MAP of the image level.\nAlthough this is not a strategy always keep consistency with public LB, but it can reflect the capability of the model to some extend.\n\nInference\nThe confidence of a cell is multiplication of FCAN image-level prediction and cell level Swin Transformer prediction.\n\n2.3 Ensemble\nWeighted average of the prediction of FCAN and Swin transformer.\n\n2.4 Post-Processing\nAs the host did not label some cells on border, if we give the cell with high confidence, the Fasle Positive cells will increase, so I trained a model to predict the completeness of a cell. If the probability to be a whole cell is very low, the confidence of this cell is multiplied by low value such as 0.3.\n\nhttps://bestfitting.github.io/kaggle/hpa2021/figures/Border_cell.png\nThe data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.\n\nThe backbone of this model is EfficientNet-B0, 3 epochs is enough to get a quite good model.\n\nThe score can improve 0.007 to 0.01 after this step.\n\n3. Results\n\nResults of simple solution\nhttps://bestfitting.github.io/kaggle/hpa2021/figures/Simple.png\n\nResults of final submission\nimagehttps://bestfitting.github.io/kaggle/hpa2021/figures/Complex.png\n\n4. Conclusion\n\n4.1 The Fair Cell Activation Network(FCAN) can increase cell level recall which is very important to this competition.\n\n4.2 The vision transformer models have shown promising capability.\n\n4.3 Larger model not always means better result as most pre-trained models are designed for ImageNet, our models should find relationship of relative position of pixels instead of abstract semantic.\n\n4.4 I found little differences among JPEG, PNG and 8bit 16bit formats.",
    "1313179": "Congratulations on your solo win!",
    "1312517": "Congratulations and thank you for a very detailed solution summary, I'm learning so much!\n\nI have a question related to 'unfair activation'. I had assumed this was a problem for semantic segmentation, but for classification this should be less relevant - given that CNN's operate locally up until GAP layer, they should be able to find discriminative features in every cell in an image. This was the basis of my solution, and also some others like tito's. \n\nIf I understand correctly what FCAN is doing, it does some regularisation, so that a model is pushed to focus on entire cells and activate them consistently across the entire image. Is my thinking correct?",
    "1309164": "Congrats for winning HPA twice with big margin",
    "1307664": "Congratulations on defending championship of HPA challenge. This solution is quite different than your previous one. Unlike last one, Densenet was not used here. But the loss function is reused. FCAN is an useful discovery!\n\nTwo questions on the classification model:\n1. For Swin transformer model, did you assign class labels based on OOF set or the training set? \n2. Did you filter out the wrong predictions based on the true image level labels or kept everything?\n\nThanks!",
    "1307357": "Congrats and thank you for posting solution. FCAM training procedure is clever! It's how we can utilize pre-computed segmentation information with puzzle cam idea!",
    "1307286": "Congratulations! Very valuable sharing! Sincere thanks!",
    "1313517": "Congratulations on the win! I enjoyed your write-up/recap of your implementation. I've actually parsed through several of your write-ups and have found them all very informative. I appreciate you sharing your approaches and techniques, they're definitely helping to expand my knowledge.\n\nThanks!",
    "1308355": "Congratulations for winning HPA again!\n\nI've got a question: where do you get the antibody-id information?\n\nI searched around in the https://www.proteinatlas.org/ but just cannot find it.",
    "1307329": "Waited last 2 weeks just for this write-up.\nAmazing work again. @bestfitting ",
    "1307259": "Congratulations. Your work is always so meticulous, sincerely admire.",
    "1307256": "Congratulations! My solution looks like a draft of yours (for initial steps) 🙂",
    "1307177": "Awesome solution! I also tried and increased the number of grid tiles for Puzzle-CAM but your way is probably the most intuitive to train for nice CAMs",
    "1534380": "Hi @bestfitting, Did you use one image at a time in training? or a batch?",
    "1389078": "@bestfitting Congrats on winning the competition and thank you for a very detailed solution summary.\nCould you please provide details about the relabelling of cells to 5 levels for Swin transformer. What exactly are the rules and thresholds for label assignments. \nThanks very much.",
    "1362472": "Congratulations!  Thank you for sharing the solution to this challenge, and to the last one.\n\nCan you please give a bit more detail about how to identify a border cell?  \n> The data to train this model is generated by randomly cutting out some area on the border a cell, and target is the area of the remaining part a cell.\n\nDo you mean that the model takes as its input an image with a cell, or part of a cell, inside it, and outputs the area (the number of pixels) occupied by the cell?  If so, how can you tell from this output whether the cell is a border cell or a whole cell?  Or, is the output a number between 0 and 1, where 1 means the cell is complete?",
    "1367673": "",
    "1495841": "Thanks for sharing!! very helpful\n\n"
  }
}