{
  "id": 238507,
  "title": "7th place solution",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238507",
  "author_name": "KaizaburoChubachi",
  "post_date": "2021-05-12T11:41:41.829000",
  "votes": 29,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all, I would like to thank the host for organizing such an exciting competition. I learned a lot from the challenging tasks that I have never dealt with before, and I was helped by the many kind contributions to Discussion by the host. </p>\n<h2>Solution Overview</h2>\n<p>Our team gave up on improving the segmentation mask by HPA-Cell-Segmentation early on, and concentrated on improving the accuracy of multi-label classification for each cell.</p>\n<p>We mainly experimented with following two approaches:</p>\n<ol>\n<li>Predict each cell from a Class Activation Map (CAM) of image level classifiers, as is commonly used in Weakly Supervised Semantic Segmentation.</li>\n<li>Crop the image for each cell and predict one by one</li>\n</ol>\n<p>In approach 1, the image level label is cleaner (compared to using it as a cell level label), so it is easier to train the classifier. On the other hand, it has a disadvantage for images with high SCV because the prediction of individual cells is easily affected by neighboring cells. In approach 2, it is difficult to perform well by simply using the image level label. But it is less affected by SVC because prediction is done for each cell separately. To take advantage of these two complementary approaches, we created models with both approaches and used them as an ensemble.</p>\n<p>The training pipeline is shown in the following image.<br>\n<img src=\"https://user-images.githubusercontent.com/8179588/117969069-eaf42300-b361-11eb-8bdc-719658486cc0.png\" alt=\"\"><br>\nWe repeated the offline pseudo learning process twice, training a new model using pseudo labels from an ensemble of multiple models and TTA. All pseudo labels are soft labels after applying the sigmoid function. All ensembles were done with simple average, and the TTA used all D4 augmentation. The image level classifier was trained using 768 x 768 images except for the 1536 one, and the cell level classifier was trained using 192 x 192 images. Cosine classifier uses cosine similarity between feature map and linear layer instead of linear transformation of feature map (We follow the equation 3 of <a href=\"https://arxiv.org/abs/2103.16370\" target=\"_blank\">https://arxiv.org/abs/2103.16370</a>)</p>\n<h2>Data</h2>\n<p>We used all the training data from this competition and the public HPA data.</p>\n<h2>Validation Strategy</h2>\n<p>To split the data, we used MultilabelStratifiedKFold from <a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">iterative-stratification</a> with 5 folds. We mainly monitored image level mAP, Focal loss, and binary cross entropy, but we could not find any metrics that correlated with public LB, so we relied on feedback from public LB.</p>\n<h2>Image Level Classifier</h2>\n<p><img src=\"https://user-images.githubusercontent.com/8179588/117969094-f6474e80-b361-11eb-89e9-ab9af1e65923.png\" alt=\"\"></p>\n<p>For the image level classifier, in addition to image level Focal loss, we used a consistency loss such that the prediction of the cell level under weak augmentation matches the prediction of the cell level under strong augmentation (CutMix). For cell level prediction, we used the average of the CAM in the region occupied by each cell (since the number of channels in CAMs is small, it worked reasonably fast even using such as scatter_add). The idea of the consistency loss is based on <a href=\"https://arxiv.org/abs/2010.09713\" target=\"_blank\">PseudoSeg</a> and <a href=\"https://arxiv.org/abs/2101.11253\" target=\"_blank\">PuzzleCAM</a> (I think the reconstruction loss in PuzzleCAM can be regarded as a consistency loss using a variant of Cutout).</p>\n<p>We mainly used EfficientNet-B2 as the image level classifier. This is because using other architectures (We tried ResNet and ResNeSt) or the larger EfficientNet would have improved the local image level mAP, but not the public LB. (This choice may have caused the public LB to overfit).</p>\n<p>When using a pseudo label from an ensemble of other models, \"Cell Level Pseudo Label\" in the figure is replaced with the pseudo label from the ensemble.</p>\n<h2>Cell level Classifier</h2>\n<p><img src=\"https://user-images.githubusercontent.com/8179588/117969124-fe9f8980-b361-11eb-9704-859ff2ada179.png\" alt=\"\"></p>\n<p>For the Cell level Classifier, in order to input both the shape of the entire cell and the size of the cell into the CNN at a somewhat small resolution, we concatenated both a fixed-scale, nucleus-centered crop and a variable-scale, whole-cell crop into the CNN. We did not use a model trained with a simple image level label because it did not perform well.</p>\n<h2>Post Processing</h2>\n<p>From the following comment in the <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns#18.-Negative\" target=\"_blank\">single-cell-patterns notebook</a>:</p>\n<blockquote>\n  <p>Please also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it!</p>\n</blockquote>\n<p>We scaled the confidence with a value based on the area of its cell (shown as “edge scale” in the bellow image) so that the confidence of the small cells at the edges of the image would be small. We also scaled the confidence of cells that are not at the edge of the image by a value (shown as non-edge scale), assuming that smaller cells are harder to predict.</p>\n<p><img src=\"https://user-images.githubusercontent.com/8179588/117969457-61912080-b362-11eb-8e3d-40b8d5274a1f.png\" alt=\"image\"></p>\n<h2>Scores</h2>\n<table>\n<thead>\n<tr>\n<th>Training</th>\n<th>Architecture</th>\n<th>Pseudo Label</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2</td>\n<td>-</td>\n<td>0.531</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B5</td>\n<td>-</td>\n<td>0.526</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B7</td>\n<td>-</td>\n<td>0.502</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2 (1536)</td>\n<td>-</td>\n<td>0.526</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2</td>\n<td>1st</td>\n<td>0.554</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2-cos</td>\n<td>2nd</td>\n<td>0.554</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2-cos</td>\n<td>2nd</td>\n<td>0.566</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Cell level classification</td>\n<td>ResNeSt50</td>\n<td>1st</td>\n<td>0.551</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Cell level classification</td>\n<td>ResNeSt50</td>\n<td>2nd</td>\n<td>0.571</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Cell level classification</td>\n<td>ResNeSt50</td>\n<td>2nd</td>\n<td>0.569</td>\n<td>-</td>\n</tr>\n<tr>\n<td>-</td>\n<td>Final Ensemble</td>\n<td>-</td>\n<td>0.580</td>\n<td>-</td>\n</tr>\n<tr>\n<td>-</td>\n<td>Final Ensemble-postprocess</td>\n<td>-</td>\n<td>0.594</td>\n<td>0.540</td>\n</tr>\n</tbody>\n</table>\n<h2>Code</h2>\n<p>(Added on May 28, 2021) We have published the code.<br>\n<a href=\"https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution</a></p>",
  "messages": [
    {
      "id": 1304028,
      "postDate": "2021-05-12T11:41:41.830Z",
      "content": "<p>First of all, I would like to thank the host for organizing such an exciting competition. I learned a lot from the challenging tasks that I have never dealt with before, and I was helped by the many kind contributions to Discussion by the host. </p>\n<h2>Solution Overview</h2>\n<p>Our team gave up on improving the segmentation mask by HPA-Cell-Segmentation early on, and concentrated on improving the accuracy of multi-label classification for each cell.</p>\n<p>We mainly experimented with following two approaches:</p>\n<ol>\n<li>Predict each cell from a Class Activation Map (CAM) of image level classifiers, as is commonly used in Weakly Supervised Semantic Segmentation.</li>\n<li>Crop the image for each cell and predict one by one</li>\n</ol>\n<p>In approach 1, the image level label is cleaner (compared to using it as a cell level label), so it is easier to train the classifier. On the other hand, it has a disadvantage for images with high SCV because the prediction of individual cells is easily affected by neighboring cells. In approach 2, it is difficult to perform well by simply using the image level label. But it is less affected by SVC because prediction is done for each cell separately. To take advantage of these two complementary approaches, we created models with both approaches and used them as an ensemble.</p>\n<p>The training pipeline is shown in the following image.<br>\n<img src=\"https://user-images.githubusercontent.com/8179588/117969069-eaf42300-b361-11eb-8bdc-719658486cc0.png\" alt=\"\"><br>\nWe repeated the offline pseudo learning process twice, training a new model using pseudo labels from an ensemble of multiple models and TTA. All pseudo labels are soft labels after applying the sigmoid function. All ensembles were done with simple average, and the TTA used all D4 augmentation. The image level classifier was trained using 768 x 768 images except for the 1536 one, and the cell level classifier was trained using 192 x 192 images. Cosine classifier uses cosine similarity between feature map and linear layer instead of linear transformation of feature map (We follow the equation 3 of <a href=\"https://arxiv.org/abs/2103.16370\" target=\"_blank\">https://arxiv.org/abs/2103.16370</a>)</p>\n<h2>Data</h2>\n<p>We used all the training data from this competition and the public HPA data.</p>\n<h2>Validation Strategy</h2>\n<p>To split the data, we used MultilabelStratifiedKFold from <a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">iterative-stratification</a> with 5 folds. We mainly monitored image level mAP, Focal loss, and binary cross entropy, but we could not find any metrics that correlated with public LB, so we relied on feedback from public LB.</p>\n<h2>Image Level Classifier</h2>\n<p><img src=\"https://user-images.githubusercontent.com/8179588/117969094-f6474e80-b361-11eb-89e9-ab9af1e65923.png\" alt=\"\"></p>\n<p>For the image level classifier, in addition to image level Focal loss, we used a consistency loss such that the prediction of the cell level under weak augmentation matches the prediction of the cell level under strong augmentation (CutMix). For cell level prediction, we used the average of the CAM in the region occupied by each cell (since the number of channels in CAMs is small, it worked reasonably fast even using such as scatter_add). The idea of the consistency loss is based on <a href=\"https://arxiv.org/abs/2010.09713\" target=\"_blank\">PseudoSeg</a> and <a href=\"https://arxiv.org/abs/2101.11253\" target=\"_blank\">PuzzleCAM</a> (I think the reconstruction loss in PuzzleCAM can be regarded as a consistency loss using a variant of Cutout).</p>\n<p>We mainly used EfficientNet-B2 as the image level classifier. This is because using other architectures (We tried ResNet and ResNeSt) or the larger EfficientNet would have improved the local image level mAP, but not the public LB. (This choice may have caused the public LB to overfit).</p>\n<p>When using a pseudo label from an ensemble of other models, \"Cell Level Pseudo Label\" in the figure is replaced with the pseudo label from the ensemble.</p>\n<h2>Cell level Classifier</h2>\n<p><img src=\"https://user-images.githubusercontent.com/8179588/117969124-fe9f8980-b361-11eb-9704-859ff2ada179.png\" alt=\"\"></p>\n<p>For the Cell level Classifier, in order to input both the shape of the entire cell and the size of the cell into the CNN at a somewhat small resolution, we concatenated both a fixed-scale, nucleus-centered crop and a variable-scale, whole-cell crop into the CNN. We did not use a model trained with a simple image level label because it did not perform well.</p>\n<h2>Post Processing</h2>\n<p>From the following comment in the <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns#18.-Negative\" target=\"_blank\">single-cell-patterns notebook</a>:</p>\n<blockquote>\n  <p>Please also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it!</p>\n</blockquote>\n<p>We scaled the confidence with a value based on the area of its cell (shown as “edge scale” in the bellow image) so that the confidence of the small cells at the edges of the image would be small. We also scaled the confidence of cells that are not at the edge of the image by a value (shown as non-edge scale), assuming that smaller cells are harder to predict.</p>\n<p><img src=\"https://user-images.githubusercontent.com/8179588/117969457-61912080-b362-11eb-8e3d-40b8d5274a1f.png\" alt=\"image\"></p>\n<h2>Scores</h2>\n<table>\n<thead>\n<tr>\n<th>Training</th>\n<th>Architecture</th>\n<th>Pseudo Label</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2</td>\n<td>-</td>\n<td>0.531</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B5</td>\n<td>-</td>\n<td>0.526</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B7</td>\n<td>-</td>\n<td>0.502</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2 (1536)</td>\n<td>-</td>\n<td>0.526</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2</td>\n<td>1st</td>\n<td>0.554</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2-cos</td>\n<td>2nd</td>\n<td>0.554</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Image level classification</td>\n<td>EfficientNet-B2-cos</td>\n<td>2nd</td>\n<td>0.566</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Cell level classification</td>\n<td>ResNeSt50</td>\n<td>1st</td>\n<td>0.551</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Cell level classification</td>\n<td>ResNeSt50</td>\n<td>2nd</td>\n<td>0.571</td>\n<td>-</td>\n</tr>\n<tr>\n<td>Cell level classification</td>\n<td>ResNeSt50</td>\n<td>2nd</td>\n<td>0.569</td>\n<td>-</td>\n</tr>\n<tr>\n<td>-</td>\n<td>Final Ensemble</td>\n<td>-</td>\n<td>0.580</td>\n<td>-</td>\n</tr>\n<tr>\n<td>-</td>\n<td>Final Ensemble-postprocess</td>\n<td>-</td>\n<td>0.594</td>\n<td>0.540</td>\n</tr>\n</tbody>\n</table>\n<h2>Code</h2>\n<p>(Added on May 28, 2021) We have published the code.<br>\n<a href=\"https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution</a></p>",
      "rawMarkdown": "First of all, I would like to thank the host for organizing such an exciting competition. I learned a lot from the challenging tasks that I have never dealt with before, and I was helped by the many kind contributions to Discussion by the host. \n\n## Solution Overview\n\nOur team gave up on improving the segmentation mask by HPA-Cell-Segmentation early on, and concentrated on improving the accuracy of multi-label classification for each cell.\n\nWe mainly experimented with following two approaches:\n1. Predict each cell from a Class Activation Map (CAM) of image level classifiers, as is commonly used in Weakly Supervised Semantic Segmentation.\n2. Crop the image for each cell and predict one by one\n\nIn approach 1, the image level label is cleaner (compared to using it as a cell level label), so it is easier to train the classifier. On the other hand, it has a disadvantage for images with high SCV because the prediction of individual cells is easily affected by neighboring cells. In approach 2, it is difficult to perform well by simply using the image level label. But it is less affected by SVC because prediction is done for each cell separately. To take advantage of these two complementary approaches, we created models with both approaches and used them as an ensemble.\n\nThe training pipeline is shown in the following image.\n![](https://user-images.githubusercontent.com/8179588/117969069-eaf42300-b361-11eb-8bdc-719658486cc0.png)\nWe repeated the offline pseudo learning process twice, training a new model using pseudo labels from an ensemble of multiple models and TTA. All pseudo labels are soft labels after applying the sigmoid function. All ensembles were done with simple average, and the TTA used all D4 augmentation. The image level classifier was trained using 768 x 768 images except for the 1536 one, and the cell level classifier was trained using 192 x 192 images. Cosine classifier uses cosine similarity between feature map and linear layer instead of linear transformation of feature map (We follow the equation 3 of https://arxiv.org/abs/2103.16370)\n\n## Data\nWe used all the training data from this competition and the public HPA data.\n\n## Validation Strategy\nTo split the data, we used MultilabelStratifiedKFold from [iterative-stratification](https://github.com/trent-b/iterative-stratification) with 5 folds. We mainly monitored image level mAP, Focal loss, and binary cross entropy, but we could not find any metrics that correlated with public LB, so we relied on feedback from public LB.\n\n## Image Level Classifier\n![](https://user-images.githubusercontent.com/8179588/117969094-f6474e80-b361-11eb-89e9-ab9af1e65923.png)\n\nFor the image level classifier, in addition to image level Focal loss, we used a consistency loss such that the prediction of the cell level under weak augmentation matches the prediction of the cell level under strong augmentation (CutMix). For cell level prediction, we used the average of the CAM in the region occupied by each cell (since the number of channels in CAMs is small, it worked reasonably fast even using such as scatter_add). The idea of the consistency loss is based on [PseudoSeg](https://arxiv.org/abs/2010.09713) and [PuzzleCAM](https://arxiv.org/abs/2101.11253) (I think the reconstruction loss in PuzzleCAM can be regarded as a consistency loss using a variant of Cutout).\n\nWe mainly used EfficientNet-B2 as the image level classifier. This is because using other architectures (We tried ResNet and ResNeSt) or the larger EfficientNet would have improved the local image level mAP, but not the public LB. (This choice may have caused the public LB to overfit).\n\nWhen using a pseudo label from an ensemble of other models, \"Cell Level Pseudo Label\" in the figure is replaced with the pseudo label from the ensemble.\n\n## Cell level Classifier\n![](https://user-images.githubusercontent.com/8179588/117969124-fe9f8980-b361-11eb-9704-859ff2ada179.png)\n\nFor the Cell level Classifier, in order to input both the shape of the entire cell and the size of the cell into the CNN at a somewhat small resolution, we concatenated both a fixed-scale, nucleus-centered crop and a variable-scale, whole-cell crop into the CNN. We did not use a model trained with a simple image level label because it did not perform well.\n\n## Post Processing\n\nFrom the following comment in the [single-cell-patterns notebook](https://www.kaggle.com/lnhtrang/single-cell-patterns#18.-Negative):\n\n> Please also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it!\n\nWe scaled the confidence with a value based on the area of its cell (shown as “edge scale” in the bellow image) so that the confidence of the small cells at the edges of the image would be small. We also scaled the confidence of cells that are not at the edge of the image by a value (shown as non-edge scale), assuming that smaller cells are harder to predict.\n\n![image](https://user-images.githubusercontent.com/8179588/117969457-61912080-b362-11eb-8e3d-40b8d5274a1f.png)\n\n\n## Scores\n\n|Training|Architecture|Pseudo Label|Public LB|Private LB|\n| --- | --- | --- | --- | --- |\n|Image level classification|EfficientNet-B2|-|0.531|-|\n|Image level classification|EfficientNet-B5|-|0.526|-|\n|Image level classification|EfficientNet-B7|-|0.502|-|\n|Image level classification|EfficientNet-B2 (1536)|-|0.526|-|\n|Image level classification|EfficientNet-B2|1st|0.554|-|\n|Image level classification|EfficientNet-B2-cos|2nd|0.554|-|\n|Image level classification|EfficientNet-B2-cos|2nd|0.566|-|\n|Cell level classification|ResNeSt50|1st|0.551|-|\n|Cell level classification|ResNeSt50|2nd|0.571|-|\n|Cell level classification|ResNeSt50|2nd|0.569|-|\n|-|Final Ensemble|-|0.580|-|\n|-|Final Ensemble-postprocess|-|0.594|0.540|\n\n## Code\n\n(Added on May 28, 2021) We have published the code.\nhttps://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution\n",
      "votes": 29
    },
    {
      "id": 1610761,
      "postDate": "2021-12-07T13:54:02.533Z",
      "content": "<p>日本語訳</p>\n<p>First of all, I would like to thank the host for organizing such an exciting competition. I learned a lot from the challenging tasks that I have never dealt with before, and I was helped by the many kind contributions to Discussion by the host.</p>\n<p>Solution Overview<br>\nOur team gave up on improving the segmentation mask by HPA-Cell-Segmentation early on, and concentrated on improving the accuracy of multi-label classification for each cell.</p>\n<p>We mainly experimented with following two approaches:</p>\n<p>Predict each cell from a Class Activation Map (CAM) of image level classifiers, as is commonly used in Weakly Supervised Semantic Segmentation.<br>\nCrop the image for each cell and predict one by one<br>\nIn approach 1, the image level label is cleaner (compared to using it as a cell level label), so it is easier to train the classifier. On the other hand, it has a disadvantage for images with high SCV because the prediction of individual cells is easily affected by neighboring cells. In approach 2, it is difficult to perform well by simply using the image level label. But it is less affected by SVC because prediction is done for each cell separately. To take advantage of these two complementary approaches, we created models with both approaches and used them as an ensemble.</p>\n<p>The training pipeline is shown in the following image.</p>\n<p>We repeated the offline pseudo learning process twice, training a new model using pseudo labels from an ensemble of multiple models and TTA. All pseudo labels are soft labels after applying the sigmoid function. All ensembles were done with simple average, and the TTA used all D4 augmentation. The image level classifier was trained using 768 x 768 images except for the 1536 one, and the cell level classifier was trained using 192 x 192 images. Cosine classifier uses cosine similarity between feature map and linear layer instead of linear transformation of feature map (We follow the equation 3 of <a href=\"https://arxiv.org/abs/2103.16370\" target=\"_blank\">https://arxiv.org/abs/2103.16370</a>)</p>\n<p>私たちのチームは、HPA-Cell-Segmentationによるセグメンテーションマスクの改善を早い段階で諦め、各セルのマルチラベル分類の精度の向上に集中しました。</p>\n<p>私たちは主に次の2つのアプローチを試しました。</p>\n<p>弱教師ありセマンティックセグメンテーションで一般的に使用されているように、画像レベル分類器のクラスアクティベーションマップ（CAM）から各セルを予測します。<br>\n各セルの画像をトリミングし、1つずつ予測します<br>\nアプローチ1では、画像レベルのラベルが（セルレベルのラベルとして使用する場合と比較して）よりクリーンであるため、分類子のトレーニングが容易になります。一方、個々のセルの予測は隣接するセルの影響を受けやすいため、SCVが高い画像には不利です。アプローチ2では、画像レベルのラベルを使用するだけではうまく機能しません。ただし、予測はセルごとに個別に行われるため、SVCの影響は少なくなります。これらの2つの補完的なアプローチを利用するために、両方のアプローチでモデルを作成し、それらをアンサンブルとして使用しました。</p>\n<p>トレーニングパイプラインを次の画像に示します。</p>\n<p>オフラインの疑似学習プロセスを2回繰り返し、複数のモデルとTTAのアンサンブルからの疑似ラベルを使用して新しいモデルをトレーニングしました。すべての疑似ラベルは、シグモイド関数を適用した後のソフトラベルです。すべてのアンサンブルは単純平均で行われ、TTAはすべてのD4拡張を使用しました。画像レベル分類器は、1536画像を除いて768 x 768画像を使用してトレーニングされ、セルレベル分類器は192 x192画像を使用してトレーニングされました。コサイン分類器は、特徴マップの線形変換の代わりに、特徴マップと線形レイヤーの間のコサイン類似性を使用します（<a href=\"https://arxiv.org/abs/2103.16370の式3に従います）\" target=\"_blank\">https://arxiv.org/abs/2103.16370の式3に従います）</a></p>\n<p>Data<br>\nWe used all the training data from this competition and the public HPA data.</p>\n<p>Validation Strategy<br>\nTo split the data, we used MultilabelStratifiedKFold from iterative-stratification with 5 folds. We mainly monitored image level mAP, Focal loss, and binary cross entropy, but we could not find any metrics that correlated with public LB, so we relied on feedback from public LB.</p>\n<p>Image Level Classifier</p>\n<p>For the image level classifier, in addition to image level Focal loss, we used a consistency loss such that the prediction of the cell level under weak augmentation matches the prediction of the cell level under strong augmentation (CutMix). For cell level prediction, we used the average of the CAM in the region occupied by each cell (since the number of channels in CAMs is small, it worked reasonably fast even using such as scatter_add). The idea of the consistency loss is based on PseudoSeg and PuzzleCAM (I think the reconstruction loss in PuzzleCAM can be regarded as a consistency loss using a variant of Cutout).</p>\n<p>We mainly used EfficientNet-B2 as the image level classifier. This is because using other architectures (We tried ResNet and ResNeSt) or the larger EfficientNet would have improved the local image level mAP, but not the public LB. (This choice may have caused the public LB to overfit).</p>\n<p>When using a pseudo label from an ensemble of other models, \"Cell Level Pseudo Label\" in the figure is replaced with the pseudo label from the ensemble.</p>\n<p>画像レベル分類器では、画像レベルの焦点損失に加えて、弱い増強下の細胞レベルの予測が強い増強下の細胞レベルの予測と一致するように、一貫性損失を使用しました（CutMix）。セルレベルの予測には、各セルが占める領域のCAMの平均を使用しました（CAMのチャネル数が少ないため、scatter_addなどを使用してもかなり高速に動作しました）。一貫性の喪失の考え方は、PseudoSegとPuzzleCAMに基づいています（PuzzleCAMでの再構築の喪失は、Cutoutのバリアントを使用した一貫性の喪失と見なすことができると思います）。</p>\n<p>画像レベル分類器として主にEfficientNet-B2を使用しました。これは、他のアーキテクチャ（ResNetとResNeStを試した）またはより大きなEfficientNetを使用すると、ローカルイメージレベルのmAPは改善されたが、パブリックLBは改善されなかったためです。 （この選択により、パブリックLBが過剰適合した可能性があります）。</p>\n<p>他のモデルのアンサンブルからの疑似ラベルを使用する場合、図の「セルレベルの疑似ラベル」は、アンサンブルからの疑似ラベルに置き換えられます。</p>\n<p>Cell level Classifier</p>\n<p>For the Cell level Classifier, in order to input both the shape of the entire cell and the size of the cell into the CNN at a somewhat small resolution, we concatenated both a fixed-scale, nucleus-centered crop and a variable-scale, whole-cell crop into the CNN. We did not use a model trained with a simple image level label because it did not perform well.</p>\n<p>Post Processing<br>\nFrom the following comment in the single-cell-patterns notebook:</p>\n<p>Please also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it!</p>\n<p>We scaled the confidence with a value based on the area of its cell (shown as “edge scale” in the bellow image) so that the confidence of the small cells at the edges of the image would be small. We also scaled the confidence of cells that are not at the edge of the image by a value (shown as non-edge scale), assuming that smaller cells are harder to predict.</p>\n<p>セルレベル分類器では、セル全体の形状とセルのサイズの両方をやや小さい解像度でCNNに入力するために、固定スケールの核中心の作物と可変スケールの両方を連結しました。 CNNへの全細胞作物。単純な画像レベルのラベルでトレーニングされたモデルは、パフォーマンスが良くなかったため、使用しませんでした。</p>\n<p>後処理<br>\nシングルセルパターンノートブックの次のコメントから：</p>\n<p>また、ほとんどの細胞が視野外にある境界細胞、および損傷した細胞や染色アーチファクトに苦しんでいる細胞にも注意してください。経験則として（グラウンドトゥルースの生成に使用されるアノテーター）、セルの半分以上が存在しない場合は、予測しないでください。</p>\n<p>画像の端にある小さなセルの信頼度が小さくなるように、セルの面積に基づく値（下の画像では「エッジスケール」として表示）で信頼度をスケーリングしました。また、小さいセルは予測が難しいと仮定して、画像のエッジにないセルの信頼度を値（非エッジスケールとして表示）でスケーリングしました。</p>\n<p>image</p>\n<p>Scores<br>\nTraining    Architecture    Pseudo Label    Public LB   Private LB<br>\nImage level classification    EfficientNet-B2 -   0.531   -<br>\nImage level classification    EfficientNet-B5 -   0.526   -<br>\nImage level classification    EfficientNet-B7 -   0.502   -<br>\nImage level classification    EfficientNet-B2 (1536)  -   0.526   -<br>\nImage level classification    EfficientNet-B2 1st 0.554   -<br>\nImage level classification    EfficientNet-B2-cos 2nd 0.554   -<br>\nImage level classification    EfficientNet-B2-cos 2nd 0.566   -<br>\nCell level classification    ResNeSt50   1st 0.551   -<br>\nCell level classification    ResNeSt50   2nd 0.571   -<br>\nCell level classification    ResNeSt50   2nd 0.569   -</p>\n<ul>\n<li>Final Ensemble  -   0.580   -</li>\n<li>Final Ensemble-postprocess  -   0.594   0.540<br>\nCode<br>\n(Added on May 28, 2021) We have published the code.<br>\n<a href=\"https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution</a></li>\n</ul>",
      "rawMarkdown": "日本語訳\n\nFirst of all, I would like to thank the host for organizing such an exciting competition. I learned a lot from the challenging tasks that I have never dealt with before, and I was helped by the many kind contributions to Discussion by the host.\n\nSolution Overview\nOur team gave up on improving the segmentation mask by HPA-Cell-Segmentation early on, and concentrated on improving the accuracy of multi-label classification for each cell.\n\nWe mainly experimented with following two approaches:\n\nPredict each cell from a Class Activation Map (CAM) of image level classifiers, as is commonly used in Weakly Supervised Semantic Segmentation.\nCrop the image for each cell and predict one by one\nIn approach 1, the image level label is cleaner (compared to using it as a cell level label), so it is easier to train the classifier. On the other hand, it has a disadvantage for images with high SCV because the prediction of individual cells is easily affected by neighboring cells. In approach 2, it is difficult to perform well by simply using the image level label. But it is less affected by SVC because prediction is done for each cell separately. To take advantage of these two complementary approaches, we created models with both approaches and used them as an ensemble.\n\nThe training pipeline is shown in the following image.\n\nWe repeated the offline pseudo learning process twice, training a new model using pseudo labels from an ensemble of multiple models and TTA. All pseudo labels are soft labels after applying the sigmoid function. All ensembles were done with simple average, and the TTA used all D4 augmentation. The image level classifier was trained using 768 x 768 images except for the 1536 one, and the cell level classifier was trained using 192 x 192 images. Cosine classifier uses cosine similarity between feature map and linear layer instead of linear transformation of feature map (We follow the equation 3 of https://arxiv.org/abs/2103.16370)\n\n私たちのチームは、HPA-Cell-Segmentationによるセグメンテーションマスクの改善を早い段階で諦め、各セルのマルチラベル分類の精度の向上に集中しました。\n\n私たちは主に次の2つのアプローチを試しました。\n\n弱教師ありセマンティックセグメンテーションで一般的に使用されているように、画像レベル分類器のクラスアクティベーションマップ（CAM）から各セルを予測します。\n各セルの画像をトリミングし、1つずつ予測します\nアプローチ1では、画像レベルのラベルが（セルレベルのラベルとして使用する場合と比較して）よりクリーンであるため、分類子のトレーニングが容易になります。一方、個々のセルの予測は隣接するセルの影響を受けやすいため、SCVが高い画像には不利です。アプローチ2では、画像レベルのラベルを使用するだけではうまく機能しません。ただし、予測はセルごとに個別に行われるため、SVCの影響は少なくなります。これらの2つの補完的なアプローチを利用するために、両方のアプローチでモデルを作成し、それらをアンサンブルとして使用しました。\n\nトレーニングパイプラインを次の画像に示します。\n\nオフラインの疑似学習プロセスを2回繰り返し、複数のモデルとTTAのアンサンブルからの疑似ラベルを使用して新しいモデルをトレーニングしました。すべての疑似ラベルは、シグモイド関数を適用した後のソフトラベルです。すべてのアンサンブルは単純平均で行われ、TTAはすべてのD4拡張を使用しました。画像レベル分類器は、1536画像を除いて768 x 768画像を使用してトレーニングされ、セルレベル分類器は192 x192画像を使用してトレーニングされました。コサイン分類器は、特徴マップの線形変換の代わりに、特徴マップと線形レイヤーの間のコサイン類似性を使用します（https://arxiv.org/abs/2103.16370の式3に従います）\n\nData\nWe used all the training data from this competition and the public HPA data.\n\nValidation Strategy\nTo split the data, we used MultilabelStratifiedKFold from iterative-stratification with 5 folds. We mainly monitored image level mAP, Focal loss, and binary cross entropy, but we could not find any metrics that correlated with public LB, so we relied on feedback from public LB.\n\nImage Level Classifier\n\n\nFor the image level classifier, in addition to image level Focal loss, we used a consistency loss such that the prediction of the cell level under weak augmentation matches the prediction of the cell level under strong augmentation (CutMix). For cell level prediction, we used the average of the CAM in the region occupied by each cell (since the number of channels in CAMs is small, it worked reasonably fast even using such as scatter_add). The idea of the consistency loss is based on PseudoSeg and PuzzleCAM (I think the reconstruction loss in PuzzleCAM can be regarded as a consistency loss using a variant of Cutout).\n\nWe mainly used EfficientNet-B2 as the image level classifier. This is because using other architectures (We tried ResNet and ResNeSt) or the larger EfficientNet would have improved the local image level mAP, but not the public LB. (This choice may have caused the public LB to overfit).\n\nWhen using a pseudo label from an ensemble of other models, \"Cell Level Pseudo Label\" in the figure is replaced with the pseudo label from the ensemble.\n\n画像レベル分類器では、画像レベルの焦点損失に加えて、弱い増強下の細胞レベルの予測が強い増強下の細胞レベルの予測と一致するように、一貫性損失を使用しました（CutMix）。セルレベルの予測には、各セルが占める領域のCAMの平均を使用しました（CAMのチャネル数が少ないため、scatter_addなどを使用してもかなり高速に動作しました）。一貫性の喪失の考え方は、PseudoSegとPuzzleCAMに基づいています（PuzzleCAMでの再構築の喪失は、Cutoutのバリアントを使用した一貫性の喪失と見なすことができると思います）。\n\n画像レベル分類器として主にEfficientNet-B2を使用しました。これは、他のアーキテクチャ（ResNetとResNeStを試した）またはより大きなEfficientNetを使用すると、ローカルイメージレベルのmAPは改善されたが、パブリックLBは改善されなかったためです。 （この選択により、パブリックLBが過剰適合した可能性があります）。\n\n他のモデルのアンサンブルからの疑似ラベルを使用する場合、図の「セルレベルの疑似ラベル」は、アンサンブルからの疑似ラベルに置き換えられます。\n\nCell level Classifier\n\n\nFor the Cell level Classifier, in order to input both the shape of the entire cell and the size of the cell into the CNN at a somewhat small resolution, we concatenated both a fixed-scale, nucleus-centered crop and a variable-scale, whole-cell crop into the CNN. We did not use a model trained with a simple image level label because it did not perform well.\n\nPost Processing\nFrom the following comment in the single-cell-patterns notebook:\n\nPlease also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it!\n\nWe scaled the confidence with a value based on the area of its cell (shown as “edge scale” in the bellow image) so that the confidence of the small cells at the edges of the image would be small. We also scaled the confidence of cells that are not at the edge of the image by a value (shown as non-edge scale), assuming that smaller cells are harder to predict.\n\nセルレベル分類器では、セル全体の形状とセルのサイズの両方をやや小さい解像度でCNNに入力するために、固定スケールの核中心の作物と可変スケールの両方を連結しました。 CNNへの全細胞作物。単純な画像レベルのラベルでトレーニングされたモデルは、パフォーマンスが良くなかったため、使用しませんでした。\n\n後処理\nシングルセルパターンノートブックの次のコメントから：\n\nまた、ほとんどの細胞が視野外にある境界細胞、および損傷した細胞や染色アーチファクトに苦しんでいる細胞にも注意してください。経験則として（グラウンドトゥルースの生成に使用されるアノテーター）、セルの半分以上が存在しない場合は、予測しないでください。\n\n画像の端にある小さなセルの信頼度が小さくなるように、セルの面積に基づく値（下の画像では「エッジスケール」として表示）で信頼度をスケーリングしました。また、小さいセルは予測が難しいと仮定して、画像のエッジにないセルの信頼度を値（非エッジスケールとして表示）でスケーリングしました。\n\nimage\n\nScores\nTraining\tArchitecture\tPseudo Label\tPublic LB\tPrivate LB\nImage level classification\tEfficientNet-B2\t-\t0.531\t-\nImage level classification\tEfficientNet-B5\t-\t0.526\t-\nImage level classification\tEfficientNet-B7\t-\t0.502\t-\nImage level classification\tEfficientNet-B2 (1536)\t-\t0.526\t-\nImage level classification\tEfficientNet-B2\t1st\t0.554\t-\nImage level classification\tEfficientNet-B2-cos\t2nd\t0.554\t-\nImage level classification\tEfficientNet-B2-cos\t2nd\t0.566\t-\nCell level classification\tResNeSt50\t1st\t0.551\t-\nCell level classification\tResNeSt50\t2nd\t0.571\t-\nCell level classification\tResNeSt50\t2nd\t0.569\t-\n-\tFinal Ensemble\t-\t0.580\t-\n-\tFinal Ensemble-postprocess\t-\t0.594\t0.540\nCode\n(Added on May 28, 2021) We have published the code.\nhttps://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution",
      "votes": 1
    },
    {
      "id": 1304032,
      "postDate": "2021-05-12T11:44:09.093Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a> and team. Good job and thanks for sharing the writeup </p>",
      "rawMarkdown": "Congrats @zaburo and team. Good job and thanks for sharing the writeup ",
      "votes": 1,
      "replies": [
        {
          "id": 1304038,
          "postDate": "2021-05-12T11:46:20.907Z",
          "content": "<p>Thanks! <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a></p>",
          "rawMarkdown": "Thanks! @duykhanh99"
        }
      ]
    },
    {
      "id": 1305027,
      "postDate": "2021-05-13T04:26:44.173Z",
      "content": "<p><a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a> Congratulations  and Thanks for sharing the approach</p>",
      "rawMarkdown": "@zaburo Congratulations  and Thanks for sharing the approach",
      "replies": [
        {
          "id": 1305650,
          "postDate": "2021-05-13T12:08:30.180Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> !</p>",
          "rawMarkdown": "Thank you @usharengaraju !"
        }
      ]
    },
    {
      "id": 1304599,
      "postDate": "2021-05-12T18:16:56.690Z",
      "content": "<p>Thanks for your sharings. Really interesting. </p>",
      "rawMarkdown": "Thanks for your sharings. Really interesting. ",
      "replies": [
        {
          "id": 1305646,
          "postDate": "2021-05-13T12:08:02.530Z",
          "content": "<p>Thank you Mathurin!</p>",
          "rawMarkdown": "Thank you Mathurin!\n"
        }
      ]
    },
    {
      "id": 1304152,
      "postDate": "2021-05-12T13:01:35.640Z",
      "content": "<p>Congrats🎉 . thanks for sharing your work through this great write up . well Explained. i have some doubts on the training process of Figure 2. </p>\n<ol>\n<li>is figure 2 the training process. do you train your CNN Feature extractor through this pipeline.</li>\n<li>how is loss backpropagated through fig 2 . can we differentiate CAM</li>\n<li>can you give some explanation on how inverse CutMix is performed on the CAM. is the CutMix region on CAM set to 0?</li>\n</ol>\n<p>and finally , how did you fit this training process on GPU memory…… 😂<br>\nthanks</p>",
      "rawMarkdown": "Congrats🎉 . thanks for sharing your work through this great write up . well Explained. i have some doubts on the training process of Figure 2. \n1. is figure 2 the training process. do you train your CNN Feature extractor through this pipeline.\n2. how is loss backpropagated through fig 2 . can we differentiate CAM\n3. can you give some explanation on how inverse CutMix is performed on the CAM. is the CutMix region on CAM set to 0?\n\nand finally , how did you fit this training process on GPU memory...... 😂\nthanks\n\n",
      "replies": [
        {
          "id": 1305645,
          "postDate": "2021-05-13T12:06:57.783Z",
          "content": "<p>Thanks for the question.</p>\n<ol>\n<li>Yes, Figure 2 is the training process for the image level classifier. We have trained all CNN (imagenet pretrained) weights with this.</li>\n<li>Let me explain about CAM. A normal image classifier obtains the logits of the classes by applying Global Average Pooling -&gt; Linear to the feature map in this order. Here, since GAP and Linear are linear operations, the values of logits does not change even if Linear -&gt; GAP is applied in this order. When Linear is applied first, the tensor immediately after the Linear operation is called CAM. Since both Linear and GAP are differentiable operations, the overall process is also differentiable.</li>\n<li>We implemented it as follows.</li>\n</ol>\n<pre><code># CutMix\nrand_index = torch.randperm(x.size()[0]).cuda()\nbbx1, bby1, bbx2, bby2 = rand_bbox(x.size(), lam)\nx[:, :, bbx1:bbx2, bby1:bby2] = x[rand_index, :, bbx1:bbx2, bby1:bby2]\n\n# Obtain a CAM\ncam = compute_cam(model, x)\n\n# Resize the CAM to the original resolution\ncam = F.interpolate(cam, x.shape[-2:], mode=\"bilinear\", align_corners=False)\n\n# Inverse CutMix\ninv_rand_index = torch.argsort(rand_index)\nmixed_mask = torch.zeros(cam.shape, dtype=torch.bool, device=x.device)\nmixed_mask[:, :, bbx1:bbx2, bby1:bby2] = True\ncam = torch.where(mixed_mask, cam[inv_rand_index], cam)\n</code></pre>\n<p>We are planning to release the code at a later date, so it would be better to check the detailed implementation there. I'll ping you in this thread when I publish it!</p>",
          "rawMarkdown": "Thanks for the question.\n\n1. Yes, Figure 2 is the training process for the image level classifier. We have trained all CNN (imagenet pretrained) weights with this.\n2. Let me explain about CAM. A normal image classifier obtains the logits of the classes by applying Global Average Pooling -> Linear to the feature map in this order. Here, since GAP and Linear are linear operations, the values of logits does not change even if Linear -> GAP is applied in this order. When Linear is applied first, the tensor immediately after the Linear operation is called CAM. Since both Linear and GAP are differentiable operations, the overall process is also differentiable.\n3. We implemented it as follows.\n```\n# CutMix\nrand_index = torch.randperm(x.size()[0]).cuda()\nbbx1, bby1, bbx2, bby2 = rand_bbox(x.size(), lam)\nx[:, :, bbx1:bbx2, bby1:bby2] = x[rand_index, :, bbx1:bbx2, bby1:bby2]\n\n# Obtain a CAM\ncam = compute_cam(model, x)\n\n# Resize the CAM to the original resolution\ncam = F.interpolate(cam, x.shape[-2:], mode=\"bilinear\", align_corners=False)\n\n# Inverse CutMix\ninv_rand_index = torch.argsort(rand_index)\nmixed_mask = torch.zeros(cam.shape, dtype=torch.bool, device=x.device)\nmixed_mask[:, :, bbx1:bbx2, bby1:bby2] = True\ncam = torch.where(mixed_mask, cam[inv_rand_index], cam)\n```\n\nWe are planning to release the code at a later date, so it would be better to check the detailed implementation there. I'll ping you in this thread when I publish it!",
          "votes": 1
        },
        {
          "id": 1307629,
          "postDate": "2021-05-14T15:07:51.797Z",
          "content": "<p>thanks for this detailed reply. waiting for your code … 👍 😄 <a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a> </p>",
          "rawMarkdown": "thanks for this detailed reply. waiting for your code ... 👍 😄 @zaburo "
        },
        {
          "id": 1326001,
          "postDate": "2021-05-28T06:44:20.540Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuvaramsingh\" target=\"_blank\">@yuvaramsingh</a> The code is now available! Sorry for the delay.<br>\n<a href=\"https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution</a></p>",
          "rawMarkdown": "@yuvaramsingh The code is now available! Sorry for the delay.\nhttps://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1610761,
      "author_name": "pixyz0130",
      "author_url": "",
      "post_date": "2021-12-07T13:54:02.533000",
      "content": "<p>日本語訳</p>\n<p>First of all, I would like to thank the host for organizing such an exciting competition. I learned a lot from the challenging tasks that I have never dealt with before, and I was helped by the many kind contributions to Discussion by the host.</p>\n<p>Solution Overview<br>\nOur team gave up on improving the segmentation mask by HPA-Cell-Segmentation early on, and concentrated on improving the accuracy of multi-label classification for each cell.</p>\n<p>We mainly experimented with following two approaches:</p>\n<p>Predict each cell from a Class Activation Map (CAM) of image level classifiers, as is commonly used in Weakly Supervised Semantic Segmentation.<br>\nCrop the image for each cell and predict one by one<br>\nIn approach 1, the image level label is cleaner (compared to using it as a cell level label), so it is easier to train the classifier. On the other hand, it has a disadvantage for images with high SCV because the prediction of individual cells is easily affected by neighboring cells. In approach 2, it is difficult to perform well by simply using the image level label. But it is less affected by SVC because prediction is done for each cell separately. To take advantage of these two complementary approaches, we created models with both approaches and used them as an ensemble.</p>\n<p>The training pipeline is shown in the following image.</p>\n<p>We repeated the offline pseudo learning process twice, training a new model using pseudo labels from an ensemble of multiple models and TTA. All pseudo labels are soft labels after applying the sigmoid function. All ensembles were done with simple average, and the TTA used all D4 augmentation. The image level classifier was trained using 768 x 768 images except for the 1536 one, and the cell level classifier was trained using 192 x 192 images. Cosine classifier uses cosine similarity between feature map and linear layer instead of linear transformation of feature map (We follow the equation 3 of <a href=\"https://arxiv.org/abs/2103.16370\" target=\"_blank\">https://arxiv.org/abs/2103.16370</a>)</p>\n<p>私たちのチームは、HPA-Cell-Segmentationによるセグメンテーションマスクの改善を早い段階で諦め、各セルのマルチラベル分類の精度の向上に集中しました。</p>\n<p>私たちは主に次の2つのアプローチを試しました。</p>\n<p>弱教師ありセマンティックセグメンテーションで一般的に使用されているように、画像レベル分類器のクラスアクティベーションマップ（CAM）から各セルを予測します。<br>\n各セルの画像をトリミングし、1つずつ予測します<br>\nアプローチ1では、画像レベルのラベルが（セルレベルのラベルとして使用する場合と比較して）よりクリーンであるため、分類子のトレーニングが容易になります。一方、個々のセルの予測は隣接するセルの影響を受けやすいため、SCVが高い画像には不利です。アプローチ2では、画像レベルのラベルを使用するだけではうまく機能しません。ただし、予測はセルごとに個別に行われるため、SVCの影響は少なくなります。これらの2つの補完的なアプローチを利用するために、両方のアプローチでモデルを作成し、それらをアンサンブルとして使用しました。</p>\n<p>トレーニングパイプラインを次の画像に示します。</p>\n<p>オフラインの疑似学習プロセスを2回繰り返し、複数のモデルとTTAのアンサンブルからの疑似ラベルを使用して新しいモデルをトレーニングしました。すべての疑似ラベルは、シグモイド関数を適用した後のソフトラベルです。すべてのアンサンブルは単純平均で行われ、TTAはすべてのD4拡張を使用しました。画像レベル分類器は、1536画像を除いて768 x 768画像を使用してトレーニングされ、セルレベル分類器は192 x192画像を使用してトレーニングされました。コサイン分類器は、特徴マップの線形変換の代わりに、特徴マップと線形レイヤーの間のコサイン類似性を使用します（<a href=\"https://arxiv.org/abs/2103.16370の式3に従います）\" target=\"_blank\">https://arxiv.org/abs/2103.16370の式3に従います）</a></p>\n<p>Data<br>\nWe used all the training data from this competition and the public HPA data.</p>\n<p>Validation Strategy<br>\nTo split the data, we used MultilabelStratifiedKFold from iterative-stratification with 5 folds. We mainly monitored image level mAP, Focal loss, and binary cross entropy, but we could not find any metrics that correlated with public LB, so we relied on feedback from public LB.</p>\n<p>Image Level Classifier</p>\n<p>For the image level classifier, in addition to image level Focal loss, we used a consistency loss such that the prediction of the cell level under weak augmentation matches the prediction of the cell level under strong augmentation (CutMix). For cell level prediction, we used the average of the CAM in the region occupied by each cell (since the number of channels in CAMs is small, it worked reasonably fast even using such as scatter_add). The idea of the consistency loss is based on PseudoSeg and PuzzleCAM (I think the reconstruction loss in PuzzleCAM can be regarded as a consistency loss using a variant of Cutout).</p>\n<p>We mainly used EfficientNet-B2 as the image level classifier. This is because using other architectures (We tried ResNet and ResNeSt) or the larger EfficientNet would have improved the local image level mAP, but not the public LB. (This choice may have caused the public LB to overfit).</p>\n<p>When using a pseudo label from an ensemble of other models, \"Cell Level Pseudo Label\" in the figure is replaced with the pseudo label from the ensemble.</p>\n<p>画像レベル分類器では、画像レベルの焦点損失に加えて、弱い増強下の細胞レベルの予測が強い増強下の細胞レベルの予測と一致するように、一貫性損失を使用しました（CutMix）。セルレベルの予測には、各セルが占める領域のCAMの平均を使用しました（CAMのチャネル数が少ないため、scatter_addなどを使用してもかなり高速に動作しました）。一貫性の喪失の考え方は、PseudoSegとPuzzleCAMに基づいています（PuzzleCAMでの再構築の喪失は、Cutoutのバリアントを使用した一貫性の喪失と見なすことができると思います）。</p>\n<p>画像レベル分類器として主にEfficientNet-B2を使用しました。これは、他のアーキテクチャ（ResNetとResNeStを試した）またはより大きなEfficientNetを使用すると、ローカルイメージレベルのmAPは改善されたが、パブリックLBは改善されなかったためです。 （この選択により、パブリックLBが過剰適合した可能性があります）。</p>\n<p>他のモデルのアンサンブルからの疑似ラベルを使用する場合、図の「セルレベルの疑似ラベル」は、アンサンブルからの疑似ラベルに置き換えられます。</p>\n<p>Cell level Classifier</p>\n<p>For the Cell level Classifier, in order to input both the shape of the entire cell and the size of the cell into the CNN at a somewhat small resolution, we concatenated both a fixed-scale, nucleus-centered crop and a variable-scale, whole-cell crop into the CNN. We did not use a model trained with a simple image level label because it did not perform well.</p>\n<p>Post Processing<br>\nFrom the following comment in the single-cell-patterns notebook:</p>\n<p>Please also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it!</p>\n<p>We scaled the confidence with a value based on the area of its cell (shown as “edge scale” in the bellow image) so that the confidence of the small cells at the edges of the image would be small. We also scaled the confidence of cells that are not at the edge of the image by a value (shown as non-edge scale), assuming that smaller cells are harder to predict.</p>\n<p>セルレベル分類器では、セル全体の形状とセルのサイズの両方をやや小さい解像度でCNNに入力するために、固定スケールの核中心の作物と可変スケールの両方を連結しました。 CNNへの全細胞作物。単純な画像レベルのラベルでトレーニングされたモデルは、パフォーマンスが良くなかったため、使用しませんでした。</p>\n<p>後処理<br>\nシングルセルパターンノートブックの次のコメントから：</p>\n<p>また、ほとんどの細胞が視野外にある境界細胞、および損傷した細胞や染色アーチファクトに苦しんでいる細胞にも注意してください。経験則として（グラウンドトゥルースの生成に使用されるアノテーター）、セルの半分以上が存在しない場合は、予測しないでください。</p>\n<p>画像の端にある小さなセルの信頼度が小さくなるように、セルの面積に基づく値（下の画像では「エッジスケール」として表示）で信頼度をスケーリングしました。また、小さいセルは予測が難しいと仮定して、画像のエッジにないセルの信頼度を値（非エッジスケールとして表示）でスケーリングしました。</p>\n<p>image</p>\n<p>Scores<br>\nTraining    Architecture    Pseudo Label    Public LB   Private LB<br>\nImage level classification    EfficientNet-B2 -   0.531   -<br>\nImage level classification    EfficientNet-B5 -   0.526   -<br>\nImage level classification    EfficientNet-B7 -   0.502   -<br>\nImage level classification    EfficientNet-B2 (1536)  -   0.526   -<br>\nImage level classification    EfficientNet-B2 1st 0.554   -<br>\nImage level classification    EfficientNet-B2-cos 2nd 0.554   -<br>\nImage level classification    EfficientNet-B2-cos 2nd 0.566   -<br>\nCell level classification    ResNeSt50   1st 0.551   -<br>\nCell level classification    ResNeSt50   2nd 0.571   -<br>\nCell level classification    ResNeSt50   2nd 0.569   -</p>\n<ul>\n<li>Final Ensemble  -   0.580   -</li>\n<li>Final Ensemble-postprocess  -   0.594   0.540<br>\nCode<br>\n(Added on May 28, 2021) We have published the code.<br>\n<a href=\"https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution</a></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1304032,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-05-12T11:44:09.093000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a> and team. Good job and thanks for sharing the writeup </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304038,
          "author_name": "KaizaburoChubachi",
          "author_url": "",
          "post_date": "2021-05-12T11:46:20.907000",
          "content": "<p>Thanks! <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1305027,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:26:44.173000",
      "content": "<p><a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a> Congratulations  and Thanks for sharing the approach</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1305650,
          "author_name": "KaizaburoChubachi",
          "author_url": "",
          "post_date": "2021-05-13T12:08:30.180000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304599,
      "author_name": "Mathurin Ache",
      "author_url": "",
      "post_date": "2021-05-12T18:16:56.690000",
      "content": "<p>Thanks for your sharings. Really interesting. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1305646,
          "author_name": "KaizaburoChubachi",
          "author_url": "",
          "post_date": "2021-05-13T12:08:02.530000",
          "content": "<p>Thank you Mathurin!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304152,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2021-05-12T13:01:35.640000",
      "content": "<p>Congrats🎉 . thanks for sharing your work through this great write up . well Explained. i have some doubts on the training process of Figure 2. </p>\n<ol>\n<li>is figure 2 the training process. do you train your CNN Feature extractor through this pipeline.</li>\n<li>how is loss backpropagated through fig 2 . can we differentiate CAM</li>\n<li>can you give some explanation on how inverse CutMix is performed on the CAM. is the CutMix region on CAM set to 0?</li>\n</ol>\n<p>and finally , how did you fit this training process on GPU memory…… 😂<br>\nthanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1305645,
          "author_name": "KaizaburoChubachi",
          "author_url": "",
          "post_date": "2021-05-13T12:06:57.783000",
          "content": "<p>Thanks for the question.</p>\n<ol>\n<li>Yes, Figure 2 is the training process for the image level classifier. We have trained all CNN (imagenet pretrained) weights with this.</li>\n<li>Let me explain about CAM. A normal image classifier obtains the logits of the classes by applying Global Average Pooling -&gt; Linear to the feature map in this order. Here, since GAP and Linear are linear operations, the values of logits does not change even if Linear -&gt; GAP is applied in this order. When Linear is applied first, the tensor immediately after the Linear operation is called CAM. Since both Linear and GAP are differentiable operations, the overall process is also differentiable.</li>\n<li>We implemented it as follows.</li>\n</ol>\n<pre><code># CutMix\nrand_index = torch.randperm(x.size()[0]).cuda()\nbbx1, bby1, bbx2, bby2 = rand_bbox(x.size(), lam)\nx[:, :, bbx1:bbx2, bby1:bby2] = x[rand_index, :, bbx1:bbx2, bby1:bby2]\n\n# Obtain a CAM\ncam = compute_cam(model, x)\n\n# Resize the CAM to the original resolution\ncam = F.interpolate(cam, x.shape[-2:], mode=\"bilinear\", align_corners=False)\n\n# Inverse CutMix\ninv_rand_index = torch.argsort(rand_index)\nmixed_mask = torch.zeros(cam.shape, dtype=torch.bool, device=x.device)\nmixed_mask[:, :, bbx1:bbx2, bby1:bby2] = True\ncam = torch.where(mixed_mask, cam[inv_rand_index], cam)\n</code></pre>\n<p>We are planning to release the code at a later date, so it would be better to check the detailed implementation there. I'll ping you in this thread when I publish it!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1307629,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "2021-05-14T15:07:51.797000",
          "content": "<p>thanks for this detailed reply. waiting for your code … 👍 😄 <a href=\"https://www.kaggle.com/zaburo\" target=\"_blank\">@zaburo</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1326001,
          "author_name": "KaizaburoChubachi",
          "author_url": "",
          "post_date": "2021-05-28T06:44:20.540000",
          "content": "<p><a href=\"https://www.kaggle.com/yuvaramsingh\" target=\"_blank\">@yuvaramsingh</a> The code is now available! Sorry for the delay.<br>\n<a href=\"https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution\" target=\"_blank\">https://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1304028": "First of all, I would like to thank the host for organizing such an exciting competition. I learned a lot from the challenging tasks that I have never dealt with before, and I was helped by the many kind contributions to Discussion by the host. \n\n## Solution Overview\n\nOur team gave up on improving the segmentation mask by HPA-Cell-Segmentation early on, and concentrated on improving the accuracy of multi-label classification for each cell.\n\nWe mainly experimented with following two approaches:\n1. Predict each cell from a Class Activation Map (CAM) of image level classifiers, as is commonly used in Weakly Supervised Semantic Segmentation.\n2. Crop the image for each cell and predict one by one\n\nIn approach 1, the image level label is cleaner (compared to using it as a cell level label), so it is easier to train the classifier. On the other hand, it has a disadvantage for images with high SCV because the prediction of individual cells is easily affected by neighboring cells. In approach 2, it is difficult to perform well by simply using the image level label. But it is less affected by SVC because prediction is done for each cell separately. To take advantage of these two complementary approaches, we created models with both approaches and used them as an ensemble.\n\nThe training pipeline is shown in the following image.\n![](https://user-images.githubusercontent.com/8179588/117969069-eaf42300-b361-11eb-8bdc-719658486cc0.png)\nWe repeated the offline pseudo learning process twice, training a new model using pseudo labels from an ensemble of multiple models and TTA. All pseudo labels are soft labels after applying the sigmoid function. All ensembles were done with simple average, and the TTA used all D4 augmentation. The image level classifier was trained using 768 x 768 images except for the 1536 one, and the cell level classifier was trained using 192 x 192 images. Cosine classifier uses cosine similarity between feature map and linear layer instead of linear transformation of feature map (We follow the equation 3 of https://arxiv.org/abs/2103.16370)\n\n## Data\nWe used all the training data from this competition and the public HPA data.\n\n## Validation Strategy\nTo split the data, we used MultilabelStratifiedKFold from [iterative-stratification](https://github.com/trent-b/iterative-stratification) with 5 folds. We mainly monitored image level mAP, Focal loss, and binary cross entropy, but we could not find any metrics that correlated with public LB, so we relied on feedback from public LB.\n\n## Image Level Classifier\n![](https://user-images.githubusercontent.com/8179588/117969094-f6474e80-b361-11eb-89e9-ab9af1e65923.png)\n\nFor the image level classifier, in addition to image level Focal loss, we used a consistency loss such that the prediction of the cell level under weak augmentation matches the prediction of the cell level under strong augmentation (CutMix). For cell level prediction, we used the average of the CAM in the region occupied by each cell (since the number of channels in CAMs is small, it worked reasonably fast even using such as scatter_add). The idea of the consistency loss is based on [PseudoSeg](https://arxiv.org/abs/2010.09713) and [PuzzleCAM](https://arxiv.org/abs/2101.11253) (I think the reconstruction loss in PuzzleCAM can be regarded as a consistency loss using a variant of Cutout).\n\nWe mainly used EfficientNet-B2 as the image level classifier. This is because using other architectures (We tried ResNet and ResNeSt) or the larger EfficientNet would have improved the local image level mAP, but not the public LB. (This choice may have caused the public LB to overfit).\n\nWhen using a pseudo label from an ensemble of other models, \"Cell Level Pseudo Label\" in the figure is replaced with the pseudo label from the ensemble.\n\n## Cell level Classifier\n![](https://user-images.githubusercontent.com/8179588/117969124-fe9f8980-b361-11eb-9704-859ff2ada179.png)\n\nFor the Cell level Classifier, in order to input both the shape of the entire cell and the size of the cell into the CNN at a somewhat small resolution, we concatenated both a fixed-scale, nucleus-centered crop and a variable-scale, whole-cell crop into the CNN. We did not use a model trained with a simple image level label because it did not perform well.\n\n## Post Processing\n\nFrom the following comment in the [single-cell-patterns notebook](https://www.kaggle.com/lnhtrang/single-cell-patterns#18.-Negative):\n\n> Please also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it!\n\nWe scaled the confidence with a value based on the area of its cell (shown as “edge scale” in the bellow image) so that the confidence of the small cells at the edges of the image would be small. We also scaled the confidence of cells that are not at the edge of the image by a value (shown as non-edge scale), assuming that smaller cells are harder to predict.\n\n![image](https://user-images.githubusercontent.com/8179588/117969457-61912080-b362-11eb-8e3d-40b8d5274a1f.png)\n\n\n## Scores\n\n|Training|Architecture|Pseudo Label|Public LB|Private LB|\n| --- | --- | --- | --- | --- |\n|Image level classification|EfficientNet-B2|-|0.531|-|\n|Image level classification|EfficientNet-B5|-|0.526|-|\n|Image level classification|EfficientNet-B7|-|0.502|-|\n|Image level classification|EfficientNet-B2 (1536)|-|0.526|-|\n|Image level classification|EfficientNet-B2|1st|0.554|-|\n|Image level classification|EfficientNet-B2-cos|2nd|0.554|-|\n|Image level classification|EfficientNet-B2-cos|2nd|0.566|-|\n|Cell level classification|ResNeSt50|1st|0.551|-|\n|Cell level classification|ResNeSt50|2nd|0.571|-|\n|Cell level classification|ResNeSt50|2nd|0.569|-|\n|-|Final Ensemble|-|0.580|-|\n|-|Final Ensemble-postprocess|-|0.594|0.540|\n\n## Code\n\n(Added on May 28, 2021) We have published the code.\nhttps://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution\n",
    "1610761": "日本語訳\n\nFirst of all, I would like to thank the host for organizing such an exciting competition. I learned a lot from the challenging tasks that I have never dealt with before, and I was helped by the many kind contributions to Discussion by the host.\n\nSolution Overview\nOur team gave up on improving the segmentation mask by HPA-Cell-Segmentation early on, and concentrated on improving the accuracy of multi-label classification for each cell.\n\nWe mainly experimented with following two approaches:\n\nPredict each cell from a Class Activation Map (CAM) of image level classifiers, as is commonly used in Weakly Supervised Semantic Segmentation.\nCrop the image for each cell and predict one by one\nIn approach 1, the image level label is cleaner (compared to using it as a cell level label), so it is easier to train the classifier. On the other hand, it has a disadvantage for images with high SCV because the prediction of individual cells is easily affected by neighboring cells. In approach 2, it is difficult to perform well by simply using the image level label. But it is less affected by SVC because prediction is done for each cell separately. To take advantage of these two complementary approaches, we created models with both approaches and used them as an ensemble.\n\nThe training pipeline is shown in the following image.\n\nWe repeated the offline pseudo learning process twice, training a new model using pseudo labels from an ensemble of multiple models and TTA. All pseudo labels are soft labels after applying the sigmoid function. All ensembles were done with simple average, and the TTA used all D4 augmentation. The image level classifier was trained using 768 x 768 images except for the 1536 one, and the cell level classifier was trained using 192 x 192 images. Cosine classifier uses cosine similarity between feature map and linear layer instead of linear transformation of feature map (We follow the equation 3 of https://arxiv.org/abs/2103.16370)\n\n私たちのチームは、HPA-Cell-Segmentationによるセグメンテーションマスクの改善を早い段階で諦め、各セルのマルチラベル分類の精度の向上に集中しました。\n\n私たちは主に次の2つのアプローチを試しました。\n\n弱教師ありセマンティックセグメンテーションで一般的に使用されているように、画像レベル分類器のクラスアクティベーションマップ（CAM）から各セルを予測します。\n各セルの画像をトリミングし、1つずつ予測します\nアプローチ1では、画像レベルのラベルが（セルレベルのラベルとして使用する場合と比較して）よりクリーンであるため、分類子のトレーニングが容易になります。一方、個々のセルの予測は隣接するセルの影響を受けやすいため、SCVが高い画像には不利です。アプローチ2では、画像レベルのラベルを使用するだけではうまく機能しません。ただし、予測はセルごとに個別に行われるため、SVCの影響は少なくなります。これらの2つの補完的なアプローチを利用するために、両方のアプローチでモデルを作成し、それらをアンサンブルとして使用しました。\n\nトレーニングパイプラインを次の画像に示します。\n\nオフラインの疑似学習プロセスを2回繰り返し、複数のモデルとTTAのアンサンブルからの疑似ラベルを使用して新しいモデルをトレーニングしました。すべての疑似ラベルは、シグモイド関数を適用した後のソフトラベルです。すべてのアンサンブルは単純平均で行われ、TTAはすべてのD4拡張を使用しました。画像レベル分類器は、1536画像を除いて768 x 768画像を使用してトレーニングされ、セルレベル分類器は192 x192画像を使用してトレーニングされました。コサイン分類器は、特徴マップの線形変換の代わりに、特徴マップと線形レイヤーの間のコサイン類似性を使用します（https://arxiv.org/abs/2103.16370の式3に従います）\n\nData\nWe used all the training data from this competition and the public HPA data.\n\nValidation Strategy\nTo split the data, we used MultilabelStratifiedKFold from iterative-stratification with 5 folds. We mainly monitored image level mAP, Focal loss, and binary cross entropy, but we could not find any metrics that correlated with public LB, so we relied on feedback from public LB.\n\nImage Level Classifier\n\n\nFor the image level classifier, in addition to image level Focal loss, we used a consistency loss such that the prediction of the cell level under weak augmentation matches the prediction of the cell level under strong augmentation (CutMix). For cell level prediction, we used the average of the CAM in the region occupied by each cell (since the number of channels in CAMs is small, it worked reasonably fast even using such as scatter_add). The idea of the consistency loss is based on PseudoSeg and PuzzleCAM (I think the reconstruction loss in PuzzleCAM can be regarded as a consistency loss using a variant of Cutout).\n\nWe mainly used EfficientNet-B2 as the image level classifier. This is because using other architectures (We tried ResNet and ResNeSt) or the larger EfficientNet would have improved the local image level mAP, but not the public LB. (This choice may have caused the public LB to overfit).\n\nWhen using a pseudo label from an ensemble of other models, \"Cell Level Pseudo Label\" in the figure is replaced with the pseudo label from the ensemble.\n\n画像レベル分類器では、画像レベルの焦点損失に加えて、弱い増強下の細胞レベルの予測が強い増強下の細胞レベルの予測と一致するように、一貫性損失を使用しました（CutMix）。セルレベルの予測には、各セルが占める領域のCAMの平均を使用しました（CAMのチャネル数が少ないため、scatter_addなどを使用してもかなり高速に動作しました）。一貫性の喪失の考え方は、PseudoSegとPuzzleCAMに基づいています（PuzzleCAMでの再構築の喪失は、Cutoutのバリアントを使用した一貫性の喪失と見なすことができると思います）。\n\n画像レベル分類器として主にEfficientNet-B2を使用しました。これは、他のアーキテクチャ（ResNetとResNeStを試した）またはより大きなEfficientNetを使用すると、ローカルイメージレベルのmAPは改善されたが、パブリックLBは改善されなかったためです。 （この選択により、パブリックLBが過剰適合した可能性があります）。\n\n他のモデルのアンサンブルからの疑似ラベルを使用する場合、図の「セルレベルの疑似ラベル」は、アンサンブルからの疑似ラベルに置き換えられます。\n\nCell level Classifier\n\n\nFor the Cell level Classifier, in order to input both the shape of the entire cell and the size of the cell into the CNN at a somewhat small resolution, we concatenated both a fixed-scale, nucleus-centered crop and a variable-scale, whole-cell crop into the CNN. We did not use a model trained with a simple image level label because it did not perform well.\n\nPost Processing\nFrom the following comment in the single-cell-patterns notebook:\n\nPlease also note that border cells where most of the cells are out of the field of view and to cells that have been damaged or suffer from staining artifacts. A good rule of thumb (that our annotators used in generating ground truth) is if more than half of the cell is not present, don't predict it!\n\nWe scaled the confidence with a value based on the area of its cell (shown as “edge scale” in the bellow image) so that the confidence of the small cells at the edges of the image would be small. We also scaled the confidence of cells that are not at the edge of the image by a value (shown as non-edge scale), assuming that smaller cells are harder to predict.\n\nセルレベル分類器では、セル全体の形状とセルのサイズの両方をやや小さい解像度でCNNに入力するために、固定スケールの核中心の作物と可変スケールの両方を連結しました。 CNNへの全細胞作物。単純な画像レベルのラベルでトレーニングされたモデルは、パフォーマンスが良くなかったため、使用しませんでした。\n\n後処理\nシングルセルパターンノートブックの次のコメントから：\n\nまた、ほとんどの細胞が視野外にある境界細胞、および損傷した細胞や染色アーチファクトに苦しんでいる細胞にも注意してください。経験則として（グラウンドトゥルースの生成に使用されるアノテーター）、セルの半分以上が存在しない場合は、予測しないでください。\n\n画像の端にある小さなセルの信頼度が小さくなるように、セルの面積に基づく値（下の画像では「エッジスケール」として表示）で信頼度をスケーリングしました。また、小さいセルは予測が難しいと仮定して、画像のエッジにないセルの信頼度を値（非エッジスケールとして表示）でスケーリングしました。\n\nimage\n\nScores\nTraining\tArchitecture\tPseudo Label\tPublic LB\tPrivate LB\nImage level classification\tEfficientNet-B2\t-\t0.531\t-\nImage level classification\tEfficientNet-B5\t-\t0.526\t-\nImage level classification\tEfficientNet-B7\t-\t0.502\t-\nImage level classification\tEfficientNet-B2 (1536)\t-\t0.526\t-\nImage level classification\tEfficientNet-B2\t1st\t0.554\t-\nImage level classification\tEfficientNet-B2-cos\t2nd\t0.554\t-\nImage level classification\tEfficientNet-B2-cos\t2nd\t0.566\t-\nCell level classification\tResNeSt50\t1st\t0.551\t-\nCell level classification\tResNeSt50\t2nd\t0.571\t-\nCell level classification\tResNeSt50\t2nd\t0.569\t-\n-\tFinal Ensemble\t-\t0.580\t-\n-\tFinal Ensemble-postprocess\t-\t0.594\t0.540\nCode\n(Added on May 28, 2021) We have published the code.\nhttps://github.com/pfnet-research/kaggle-hpa-2021-7th-place-solution",
    "1304032": "Congrats @zaburo and team. Good job and thanks for sharing the writeup ",
    "1305027": "@zaburo Congratulations  and Thanks for sharing the approach",
    "1304599": "Thanks for your sharings. Really interesting. ",
    "1304152": "Congrats🎉 . thanks for sharing your work through this great write up . well Explained. i have some doubts on the training process of Figure 2. \n1. is figure 2 the training process. do you train your CNN Feature extractor through this pipeline.\n2. how is loss backpropagated through fig 2 . can we differentiate CAM\n3. can you give some explanation on how inverse CutMix is performed on the CAM. is the CutMix region on CAM set to 0?\n\nand finally , how did you fit this training process on GPU memory...... 😂\nthanks\n\n"
  }
}