{
  "id": 238645,
  "title": "HPA 2nd Place Solution [red.ai]",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238645",
  "author_name": "sheep",
  "post_date": "2021-05-13T00:10:10.014000",
  "votes": 63,
  "comment_count": 29,
  "views": 0,
  "content": "<h2>Preface</h2>\n<p>As promised, we will share our detailed solution within 24 hours. We would like to thank the organizers for this awesome competition since all of us had no experience in dealing with weakly-supervised classification problems and we have learned a lot from the the kind sharings by other kagglers and self-discoveries. The organizers are very active in this competition; huge props to all of you. I am also grateful for my teammates for making my journey to <strong>Kaggle Competition Grandmaster</strong> smooth and gratifying. To say I am excited is a huge under-statement. Without further ado, let's dive into our solution.</p>\n<h2>TLDR</h2>\n<p>Our solution consists of a total of 3 simple pipelines. We did not use any advanced techniques from any paper but we tried to understand the data well and build our model architecture w.r.t the problem statement. Here is a diagram for our final pipeline:</p>\n<p><img src=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&amp;id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" alt=\"\"></p>\n<h2>Pipeline 1: Duo-Branch Cell Model</h2>\n<ul>\n<li>Motivation: A duo-branch(head) cell model was designed in a way that it takes cell tiles as input but has the ability to predict both as cell-level and image-level. Multi-tasking has been shown to be effective in improving model learning. A strong champion in dota2 called Jakiro also has two heads.</li>\n<li>Loss formulation: since the output is cell-level and image-level, we need two losses for both outputs. The final loss is the weighted sum of cell-level loss and image-level loss. We used basic <strong>BCE</strong> loss for both cell-level and image-level. For cell-level, the labels are not certain so it's intuitive to assign a lower weight (=0.1). l = 0.1*loss_cell + loss_image</li>\n<li>Data: we used original data, external data shared by Phil as well as some rare class samples by using the API. The input size for a single cell is 256.</li>\n<li>Training details: we take 4-channel images and crop&amp;resize the cells first; then we random sample N (=16) cells as input of our network. The cells are flattened as a large batch then we feed them into a CNN and backprop. For data augmentations we used dihedral, shift, rotate, scale, distortions, brightness contrast and cutout. The heavy data augmentations allows the model to better generalize as cells can be in any forms in reality. We train 5 folds and 20 epochs each. A single Model (b3, 256) takes about 30 hours to train on a single RTX3090.</li>\n<li>Inference: for each image, the pooled features are concatenated and feed into the last linear layer to predict at a cell-level. We generate image level prediction and cell level prediction and calculate their product as our final prediction.</li>\n<li>Result: with 16xTTA(scale, rotate, flip at random), 256 size cell-tiles, N=16, ensemble of efficientnet B3, B5 resnet200d and se_resnext50 backbone, our model score <code>0.550</code> on the public leaderboard and <code>0.550</code> on the private leaderboard. This single architecture can achieve second place in this competition.</li>\n</ul>\n<p><img src=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&amp;id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" alt=\"\"></p>\n<h2>Pipeline 2: Image-level Model</h2>\n<ul>\n<li><strong>Misc:</strong> The image level model is similar to last HPA's competition. We feed in whole images and perform a multi-label classification. It may surprise you, this pipeline is developed in fastai. We have found fastai's many implementations to be extremely slow (for instance, resize) and had gone through many days of debugging during the final inference phase; we spent a whole last week attempting to figure out how to submit. In the end, we optimized the fast.ai inference code by a lot that helped us cut the inference time almost twice. Luckily, hard work paid off.</li>\n<li><strong>Motivation:</strong> we can train at an image level but predict at a cell-level (with other cells masked) and the result is very promising. We decide to add this our pipeline.</li>\n<li><strong>Data:</strong> we used all data available on HPA official website and resized it to 512 using only RGB 3 channels.</li>\n<li><strong>Training details:</strong> we train 20 epochs with class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and BCE loss for 2 folds only. We used average precision score  for checkpointing. For data augmentations, we used fastai's <code>aug_transforms(flip_vert=True, max_lighting=0.1, max_warp=0.1, p_affine=0.5, p_lighting=0.5)</code></li>\n<li><strong>Result:</strong> we had 10 (5x2folds) models and we took the mean of the final output. And we use them to predict at both cell-level and image-level. We take the mean as our final output.</li>\n</ul>\n<p><img src=\"https://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC\" alt=\"\"></p>\n<h2>Pipeline 3: Cell-level Model</h2>\n<ul>\n<li><strong>Motivation:</strong> we can train at cell-level using the image-level labels but it's a bit counter intuitive. Since his will introduce lots of noise as image-level labels are not ground truth for cells so we think it's beneficial to train less epochs. We ended up only training 2 epochs (1 with backbone freezed and 1 with backbone unfreezed).</li>\n<li><strong>Data:</strong> we used all data available on HPA official website, use the cell segmentor to crop the cells and resized the cells to 168 using only RGB 3 channels. There are a total of 1620178 cropped cells.</li>\n<li><strong>Training details:</strong> we used fastai's built-in <code>finetune</code> and fastai's learning rate finder to train only 2 epochs with the same class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and bce loss. We did not use anything for validation.</li>\n<li><strong>Result:</strong> we had 10 (10x1folds) models. We predicted at cell-level and simply took the mean of the final output.</li>\n</ul>\n<p><img src=\"https://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE\" alt=\"\"></p>\n<h2>Segmentation Model</h2>\n<p>We are inspired by <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> Even Faster HPA Cell Segmentation and <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> Segmentation with a Scaling Factor, we modified the original HPA Segmentator to gain speed but keep the segmentation quality.</p>\n<ul>\n<li>Post-processing: We slightly changed label_cell function from the original implementation. We found that in many cases, border cells are segmented in a wrong way: some of them are combined together with border cells that have no nuclei (or it’s outside of the image). We tweaked the watershed distance threshold in order to separate cells masks a little bit further from each other than they were before, then we ignored the masks on the border that became separated from the main cell. Furthermore, we removed the border cells with nuclei whose area was less than a half of the median area of the non-border nuclei on the image. And we also removed the cells that did not have the corresponding nuclei. Below is an example of a difference between original label_cell implementation (left) and ours (right):</li>\n</ul>\n<p><img src=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&amp;id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" alt=\"\"></p>\n<p><img src=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&amp;id=214638be-245a-4780-81a1-7866fc81081b&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" alt=\"\"></p>\n<h2>Arcface Model</h2>\n<p>We also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.</p>\n<h2>Duplicate samples.</h2>\n<p>We found about ~400 image in public test set duplicated either within the train set or the external data. You can check the csv file at <a href=\"https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample\" target=\"_blank\">https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample</a>. Our public leaderboard score, excluding the boost from duplicates is about 0.58, we have a relative consistent gap w.r.t. the 1st place in both public and private leaderboard.</p>\n<h2>Things that didn't work</h2>\n<ul>\n<li>Segmentation post-processing on scaled-up outputs of the segmentator led to a slight decrease in the score</li>\n<li>Tiling a plot with a single cell and classifying such cells with the image level models.</li>\n</ul>\n<h2>Solution code</h2>\n<ul>\n<li>Pipeline 1's code is now available at <a href=\"https://github.com/iseekwonderful/HPA-singlecell-2nd-dual-head-pipeline\" target=\"_blank\">github</a></li>\n</ul>",
  "messages": [
    {
      "id": 1304849,
      "postDate": "2021-05-13T00:10:10.013Z",
      "content": "<h2>Preface</h2>\n<p>As promised, we will share our detailed solution within 24 hours. We would like to thank the organizers for this awesome competition since all of us had no experience in dealing with weakly-supervised classification problems and we have learned a lot from the the kind sharings by other kagglers and self-discoveries. The organizers are very active in this competition; huge props to all of you. I am also grateful for my teammates for making my journey to <strong>Kaggle Competition Grandmaster</strong> smooth and gratifying. To say I am excited is a huge under-statement. Without further ado, let's dive into our solution.</p>\n<h2>TLDR</h2>\n<p>Our solution consists of a total of 3 simple pipelines. We did not use any advanced techniques from any paper but we tried to understand the data well and build our model architecture w.r.t the problem statement. Here is a diagram for our final pipeline:</p>\n<p><img src=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&amp;id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" alt=\"\"></p>\n<h2>Pipeline 1: Duo-Branch Cell Model</h2>\n<ul>\n<li>Motivation: A duo-branch(head) cell model was designed in a way that it takes cell tiles as input but has the ability to predict both as cell-level and image-level. Multi-tasking has been shown to be effective in improving model learning. A strong champion in dota2 called Jakiro also has two heads.</li>\n<li>Loss formulation: since the output is cell-level and image-level, we need two losses for both outputs. The final loss is the weighted sum of cell-level loss and image-level loss. We used basic <strong>BCE</strong> loss for both cell-level and image-level. For cell-level, the labels are not certain so it's intuitive to assign a lower weight (=0.1). l = 0.1*loss_cell + loss_image</li>\n<li>Data: we used original data, external data shared by Phil as well as some rare class samples by using the API. The input size for a single cell is 256.</li>\n<li>Training details: we take 4-channel images and crop&amp;resize the cells first; then we random sample N (=16) cells as input of our network. The cells are flattened as a large batch then we feed them into a CNN and backprop. For data augmentations we used dihedral, shift, rotate, scale, distortions, brightness contrast and cutout. The heavy data augmentations allows the model to better generalize as cells can be in any forms in reality. We train 5 folds and 20 epochs each. A single Model (b3, 256) takes about 30 hours to train on a single RTX3090.</li>\n<li>Inference: for each image, the pooled features are concatenated and feed into the last linear layer to predict at a cell-level. We generate image level prediction and cell level prediction and calculate their product as our final prediction.</li>\n<li>Result: with 16xTTA(scale, rotate, flip at random), 256 size cell-tiles, N=16, ensemble of efficientnet B3, B5 resnet200d and se_resnext50 backbone, our model score <code>0.550</code> on the public leaderboard and <code>0.550</code> on the private leaderboard. This single architecture can achieve second place in this competition.</li>\n</ul>\n<p><img src=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&amp;id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" alt=\"\"></p>\n<h2>Pipeline 2: Image-level Model</h2>\n<ul>\n<li><strong>Misc:</strong> The image level model is similar to last HPA's competition. We feed in whole images and perform a multi-label classification. It may surprise you, this pipeline is developed in fastai. We have found fastai's many implementations to be extremely slow (for instance, resize) and had gone through many days of debugging during the final inference phase; we spent a whole last week attempting to figure out how to submit. In the end, we optimized the fast.ai inference code by a lot that helped us cut the inference time almost twice. Luckily, hard work paid off.</li>\n<li><strong>Motivation:</strong> we can train at an image level but predict at a cell-level (with other cells masked) and the result is very promising. We decide to add this our pipeline.</li>\n<li><strong>Data:</strong> we used all data available on HPA official website and resized it to 512 using only RGB 3 channels.</li>\n<li><strong>Training details:</strong> we train 20 epochs with class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and BCE loss for 2 folds only. We used average precision score  for checkpointing. For data augmentations, we used fastai's <code>aug_transforms(flip_vert=True, max_lighting=0.1, max_warp=0.1, p_affine=0.5, p_lighting=0.5)</code></li>\n<li><strong>Result:</strong> we had 10 (5x2folds) models and we took the mean of the final output. And we use them to predict at both cell-level and image-level. We take the mean as our final output.</li>\n</ul>\n<p><img src=\"https://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC\" alt=\"\"></p>\n<h2>Pipeline 3: Cell-level Model</h2>\n<ul>\n<li><strong>Motivation:</strong> we can train at cell-level using the image-level labels but it's a bit counter intuitive. Since his will introduce lots of noise as image-level labels are not ground truth for cells so we think it's beneficial to train less epochs. We ended up only training 2 epochs (1 with backbone freezed and 1 with backbone unfreezed).</li>\n<li><strong>Data:</strong> we used all data available on HPA official website, use the cell segmentor to crop the cells and resized the cells to 168 using only RGB 3 channels. There are a total of 1620178 cropped cells.</li>\n<li><strong>Training details:</strong> we used fastai's built-in <code>finetune</code> and fastai's learning rate finder to train only 2 epochs with the same class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and bce loss. We did not use anything for validation.</li>\n<li><strong>Result:</strong> we had 10 (10x1folds) models. We predicted at cell-level and simply took the mean of the final output.</li>\n</ul>\n<p><img src=\"https://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE\" alt=\"\"></p>\n<h2>Segmentation Model</h2>\n<p>We are inspired by <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> Even Faster HPA Cell Segmentation and <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> Segmentation with a Scaling Factor, we modified the original HPA Segmentator to gain speed but keep the segmentation quality.</p>\n<ul>\n<li>Post-processing: We slightly changed label_cell function from the original implementation. We found that in many cases, border cells are segmented in a wrong way: some of them are combined together with border cells that have no nuclei (or it’s outside of the image). We tweaked the watershed distance threshold in order to separate cells masks a little bit further from each other than they were before, then we ignored the masks on the border that became separated from the main cell. Furthermore, we removed the border cells with nuclei whose area was less than a half of the median area of the non-border nuclei on the image. And we also removed the cells that did not have the corresponding nuclei. Below is an example of a difference between original label_cell implementation (left) and ours (right):</li>\n</ul>\n<p><img src=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&amp;id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" alt=\"\"></p>\n<p><img src=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&amp;id=214638be-245a-4780-81a1-7866fc81081b&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" alt=\"\"></p>\n<h2>Arcface Model</h2>\n<p>We also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.</p>\n<h2>Duplicate samples.</h2>\n<p>We found about ~400 image in public test set duplicated either within the train set or the external data. You can check the csv file at <a href=\"https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample\" target=\"_blank\">https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample</a>. Our public leaderboard score, excluding the boost from duplicates is about 0.58, we have a relative consistent gap w.r.t. the 1st place in both public and private leaderboard.</p>\n<h2>Things that didn't work</h2>\n<ul>\n<li>Segmentation post-processing on scaled-up outputs of the segmentator led to a slight decrease in the score</li>\n<li>Tiling a plot with a single cell and classifying such cells with the image level models.</li>\n</ul>\n<h2>Solution code</h2>\n<ul>\n<li>Pipeline 1's code is now available at <a href=\"https://github.com/iseekwonderful/HPA-singlecell-2nd-dual-head-pipeline\" target=\"_blank\">github</a></li>\n</ul>",
      "rawMarkdown": "## Preface\nAs promised, we will share our detailed solution within 24 hours. We would like to thank the organizers for this awesome competition since all of us had no experience in dealing with weakly-supervised classification problems and we have learned a lot from the the kind sharings by other kagglers and self-discoveries. The organizers are very active in this competition; huge props to all of you. I am also grateful for my teammates for making my journey to **Kaggle Competition Grandmaster** smooth and gratifying. To say I am excited is a huge under-statement. Without further ado, let's dive into our solution.\n\n## TLDR\nOur solution consists of a total of 3 simple pipelines. We did not use any advanced techniques from any paper but we tried to understand the data well and build our model architecture w.r.t the problem statement. Here is a diagram for our final pipeline:\n\n![](http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2)\n\n## Pipeline 1: Duo-Branch Cell Model\n- Motivation: A duo-branch(head) cell model was designed in a way that it takes cell tiles as input but has the ability to predict both as cell-level and image-level. Multi-tasking has been shown to be effective in improving model learning. A strong champion in dota2 called Jakiro also has two heads.\n- Loss formulation: since the output is cell-level and image-level, we need two losses for both outputs. The final loss is the weighted sum of cell-level loss and image-level loss. We used basic **BCE** loss for both cell-level and image-level. For cell-level, the labels are not certain so it's intuitive to assign a lower weight (=0.1). l = 0.1*loss_cell + loss_image\n- Data: we used original data, external data shared by Phil as well as some rare class samples by using the API. The input size for a single cell is 256.\n- Training details: we take 4-channel images and crop&resize the cells first; then we random sample N (=16) cells as input of our network. The cells are flattened as a large batch then we feed them into a CNN and backprop. For data augmentations we used dihedral, shift, rotate, scale, distortions, brightness contrast and cutout. The heavy data augmentations allows the model to better generalize as cells can be in any forms in reality. We train 5 folds and 20 epochs each. A single Model (b3, 256) takes about 30 hours to train on a single RTX3090.\n- Inference: for each image, the pooled features are concatenated and feed into the last linear layer to predict at a cell-level. We generate image level prediction and cell level prediction and calculate their product as our final prediction.\n- Result: with 16xTTA(scale, rotate, flip at random), 256 size cell-tiles, N=16, ensemble of efficientnet B3, B5 resnet200d and se_resnext50 backbone, our model score `0.550` on the public leaderboard and `0.550` on the private leaderboard. This single architecture can achieve second place in this competition.\n\n![](http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2)\n\n## Pipeline 2: Image-level Model\n\n- **Misc:** The image level model is similar to last HPA's competition. We feed in whole images and perform a multi-label classification. It may surprise you, this pipeline is developed in fastai. We have found fastai's many implementations to be extremely slow (for instance, resize) and had gone through many days of debugging during the final inference phase; we spent a whole last week attempting to figure out how to submit. In the end, we optimized the fast.ai inference code by a lot that helped us cut the inference time almost twice. Luckily, hard work paid off.\n- **Motivation:** we can train at an image level but predict at a cell-level (with other cells masked) and the result is very promising. We decide to add this our pipeline.\n- **Data:** we used all data available on HPA official website and resized it to 512 using only RGB 3 channels.\n- **Training details:** we train 20 epochs with class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and BCE loss for 2 folds only. We used average precision score  for checkpointing. For data augmentations, we used fastai's `aug_transforms(flip_vert=True, max_lighting=0.1, max_warp=0.1, p_affine=0.5, p_lighting=0.5)`\n- **Result:** we had 10 (5x2folds) models and we took the mean of the final output. And we use them to predict at both cell-level and image-level. We take the mean as our final output.\n\n![](https://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC)\n\n## Pipeline 3: Cell-level Model\n- **Motivation:** we can train at cell-level using the image-level labels but it's a bit counter intuitive. Since his will introduce lots of noise as image-level labels are not ground truth for cells so we think it's beneficial to train less epochs. We ended up only training 2 epochs (1 with backbone freezed and 1 with backbone unfreezed).\n- **Data:** we used all data available on HPA official website, use the cell segmentor to crop the cells and resized the cells to 168 using only RGB 3 channels. There are a total of 1620178 cropped cells.\n- **Training details:** we used fastai's built-in `finetune` and fastai's learning rate finder to train only 2 epochs with the same class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and bce loss. We did not use anything for validation.\n- **Result:** we had 10 (10x1folds) models. We predicted at cell-level and simply took the mean of the final output.\n\n![](https://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE)\n\n## Segmentation Model\n\nWe are inspired by @samusram Even Faster HPA Cell Segmentation and @alexanderriedel Segmentation with a Scaling Factor, we modified the original HPA Segmentator to gain speed but keep the segmentation quality.\n* Post-processing: We slightly changed label_cell function from the original implementation. We found that in many cases, border cells are segmented in a wrong way: some of them are combined together with border cells that have no nuclei (or it’s outside of the image). We tweaked the watershed distance threshold in order to separate cells masks a little bit further from each other than they were before, then we ignored the masks on the border that became separated from the main cell. Furthermore, we removed the border cells with nuclei whose area was less than a half of the median area of the non-border nuclei on the image. And we also removed the cells that did not have the corresponding nuclei. Below is an example of a difference between original label_cell implementation (left) and ours (right):\n\n![](http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2)\n\n![](http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&id=214638be-245a-4780-81a1-7866fc81081b&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2)\n\n## Arcface Model\nWe also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.\n\n## Duplicate samples.\nWe found about ~400 image in public test set duplicated either within the train set or the external data. You can check the csv file at https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample. Our public leaderboard score, excluding the boost from duplicates is about 0.58, we have a relative consistent gap w.r.t. the 1st place in both public and private leaderboard.\n\n## Things that didn't work\n- Segmentation post-processing on scaled-up outputs of the segmentator led to a slight decrease in the score\n- Tiling a plot with a single cell and classifying such cells with the image level models.\n\n## Solution code\n* Pipeline 1's code is now available at [github](https://github.com/iseekwonderful/HPA-singlecell-2nd-dual-head-pipeline)",
      "votes": 63
    },
    {
      "id": 1308072,
      "postDate": "2021-05-15T00:39:40.720Z",
      "content": "<p>Congratulations sheep and team. Great solution. Congratulations sheep on achieving Kaggle Competition Grandmaster!</p>",
      "rawMarkdown": "Congratulations sheep and team. Great solution. Congratulations sheep on achieving Kaggle Competition Grandmaster!",
      "votes": 4
    },
    {
      "id": 1307282,
      "postDate": "2021-05-14T11:05:30.177Z",
      "content": "<p>If anyone interested in the inference code for pipelines 2&amp;3 (including upgraded segmentation postprocessing and some visualizations) I made the kernel public <a href=\"https://www.kaggle.com/vostankovich/crops-full-crops-full-arcface-ensemble-optimize?scriptVersionId=61646005\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "If anyone interested in the inference code for pipelines 2&3 (including upgraded segmentation postprocessing and some visualizations) I made the kernel public [here](https://www.kaggle.com/vostankovich/crops-full-crops-full-arcface-ensemble-optimize?scriptVersionId=61646005)",
      "votes": 4
    },
    {
      "id": 1610733,
      "postDate": "2021-12-07T13:35:19.373Z",
      "content": "<p>日本語訳</p>\n<p>Preface<br>\nAs promised, we will share our detailed solution within 24 hours. We would like to thank the organizers for this awesome competition since all of us had no experience in dealing with weakly-supervised classification problems and we have learned a lot from the the kind sharings by other kagglers and self-discoveries. The organizers are very active in this competition; huge props to all of you. I am also grateful for my teammates for making my journey to Kaggle Competition Grandmaster smooth and gratifying. To say I am excited is a huge under-statement. Without further ado, let's dive into our solution.</p>\n<p>約束どおり、24時間以内に詳細なソリューションを共有します。 教師なし分類の問題に対処した経験がなく、他のカグラーによる親切な共有や自己発見から多くのことを学んだので、この素晴らしいコンテストの主催者に感謝します。 主催者はこの大会に非常に積極的です。 皆さんへの巨大な小道具。 また、Kaggleコンペティションのグランドマスターへの旅をスムーズで満足のいくものにしてくれたチームメートにも感謝しています。 私が興奮していると言うことは、非常に控えめな表現です。 さらに面倒なことはせずに、私たちのソリューションに飛び込みましょう。</p>\n<p>TLDR<br>\nOur solution consists of a total of 3 simple pipelines. We did not use any advanced techniques from any paper but we tried to understand the data well and build our model architecture w.r.t the problem statement. Here is a diagram for our final pipeline:</p>\n<p>私たちのソリューションは、合計3つの単純なパイプラインで構成されています。 どの論文からも高度な手法を使用しませんでしたが、データを十分に理解し、問題ステートメントを使用してモデルアーキテクチャを構築しようとしました。 これが最終パイプラインの図です。</p>\n<p><a href=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&amp;id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" target=\"_blank\">http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&amp;id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2</a></p>\n<p>Pipeline 1: Duo-Branch Cell Model<br>\nMotivation: A duo-branch(head) cell model was designed in a way that it takes cell tiles as input but has the ability to predict both as cell-level and image-level. Multi-tasking has been shown to be effective in improving model learning. A strong champion in dota2 called Jakiro also has two heads.<br>\nデュオブランチ（ヘッド）セルモデルは、セルタイルを入力として受け取るように設計されていますが、セルレベルと画像レベルの両方を予測する機能があります。 マルチタスクは、モデル学習の改善に効果的であることが示されています。 Jakiroと呼ばれるdota2の強力なチャンピオンにも2つの頭があります。</p>\n<p>Loss formulation: since the output is cell-level and image-level, we need two losses for both outputs. The final loss is the weighted sum of cell-level loss and image-level loss. We used basic BCE loss for both cell-level and image-level. For cell-level, the labels are not certain so it's intuitive to assign a lower weight (=0.1). l = 0.1*loss_cell + loss_image<br>\nData: we used original data, external data shared by Phil as well as some rare class samples by using the API. The input size for a single cell is 256.<br>\n出力はセルレベルと画像レベルであるため、両方の出力に2つの損失が必要です。 最終的な損失は、セルレベルの損失と画像レベルの損失の加重和です。 セルレベルと画像レベルの両方で基本的なBCE損失を使用しました。 セルレベルの場合、ラベルは明確ではないため、より低い重み（= 0.1）を割り当てるのは直感的です。 l = 0.1 * loss_cell + loss_image<br>\nデータ：元のデータ、Philが共有する外部データ、およびAPIを使用したいくつかのまれなクラスサンプルを使用しました。 単一セルの入力サイズは256です。</p>\n<p>Training details: we take 4-channel images and crop&amp;resize the cells first; then we random sample N (=16) cells as input of our network. The cells are flattened as a large batch then we feed them into a CNN and backprop. For data augmentations we used dihedral, shift, rotate, scale, distortions, brightness contrast and cutout. The heavy data augmentations allows the model to better generalize as cells can be in any forms in reality. We train 5 folds and 20 epochs each. A single Model (b3, 256) takes about 30 hours to train on a single RTX3090.<br>\n4チャンネルの画像を撮影し、最初にセルをトリミングしてサイズを変更します。 次に、ネットワークの入力としてN（= 16）個のセルをランダムにサンプリングします。 セルは大きなバッチとして平坦化され、CNNとバックプロパゲーションにフィードされます。 データ拡張には、二面角、シフト、回転、スケール、歪み、明るさのコントラスト、カットアウトを使用しました。 大量のデータ拡張により、セルは実際には任意の形式である可能性があるため、モデルをより一般化することができます。 それぞれ5つのフォールドと20のエポックをトレーニングします。 単一のモデル（b3、256）は、単一のRTX3090でトレーニングするのに約30時間かかります。</p>\n<p>Inference: for each image, the pooled features are concatenated and feed into the last linear layer to predict at a cell-level. We generate image level prediction and cell level prediction and calculate their product as our final prediction.<br>\nResult: with 16xTTA(scale, rotate, flip at random), 256 size cell-tiles, N=16, ensemble of efficientnet B3, B5 resnet200d and se_resnext50 backbone, our model score 0.550 on the public leaderboard and 0.550 on the private leaderboard. This single architecture can achieve second place in this competition.<br>\n画像ごとに、プールされた特徴が連結され、最後の線形レイヤーにフィードされて、セルレベルで予測されます。 画像レベル予測と細胞レベル予測を生成し、それらの積を最終予測として計算します。<br>\n結果：16xTTA（スケール、回転、ランダムに反転）、256サイズのセルタイル、N = 16、効率的なネットB3、B5 resnet200d、se_resnext50バックボーンのアンサンブルで、モデルスコアはパブリックリーダーボードで0.550、プライベートリーダーボードで0.550です。 この単一のアーキテクチャは、この競争で2位を獲得することができます。<br>\n<a href=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&amp;id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" target=\"_blank\">http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&amp;id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2</a></p>\n<p>Pipeline 2: Image-level Model<br>\nMisc: The image level model is similar to last HPA's competition. We feed in whole images and perform a multi-label classification. It may surprise you, this pipeline is developed in fastai. We have found fastai's many implementations to be extremely slow (for instance, resize) and had gone through many days of debugging during the final inference phase; we spent a whole last week attempting to figure out how to submit. In the end, we optimized the fast.ai inference code by a lot that helped us cut the inference time almost twice. Luckily, hard work paid off.<br>\nMotivation: we can train at an image level but predict at a cell-level (with other cells masked) and the result is very promising. We decide to add this our pipeline.<br>\n画像レベルのモデルは、前回のHPAの競合製品と似ています。 画像全体をフィードし、マルチラベル分類を実行します。 驚かれるかもしれませんが、このパイプラインはfastaiで開発されています。 fastaiの多くの実装は非常に遅く（たとえば、サイズ変更）、最終的な推論フェーズで何日もデバッグを行っていました。 私たちは先週、提出方法を理解しようと一週間を費やしました。 最終的に、fast.ai推論コードを大幅に最適化して、推論時間をほぼ2倍に短縮しました。 幸いなことに、努力は報われました。<br>\n動機：画像レベルでトレーニングすることはできますが、セルレベルで（他のセルをマスクして）予測することができ、その結果は非常に有望です。 これをパイプラインに追加することにしました。</p>\n<p>Data: we used all data available on HPA official website and resized it to 512 using only RGB 3 channels.<br>\nTraining details: we train 20 epochs with class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and BCE loss for 2 folds only. We used average precision score for checkpointing. For data augmentations, we used fastai's aug_transforms(flip_vert=True, max_lighting=0.1, max_warp=0.1, p_affine=0.5, p_lighting=0.5)<br>\nResult: we had 10 (5x2folds) models and we took the mean of the final output. And we use them to predict at both cell-level and image-level. We take the mean as our final output.<br>\nHPAの公式Webサイトで入手可能なすべてのデータを使用し、RGB3チャネルのみを使用して512にサイズ変更しました。<br>\nトレーニングの詳細：クラスの重み[0.1、1。、0.5、1。、1.、1.、1.、0.5、1。、1.、1.、10.、1.、0.5、0.5で20エポックをトレーニングします 、5、0.2、0.5、1。]および2倍のみのBCE損失。 チェックポイントには平均精度スコアを使用しました。 データ拡張には、fastaiのaug_transforms（flip_vert = True、max_lighting = 0.1、max_warp = 0.1、p_affine = 0.5、p_lighting = 0.5）を使用しました。<br>\n結果：10（5x2folds）モデルがあり、最終出力の平均を取りました。 そして、それらを使用して、セルレベルと画像レベルの両方で予測します。 最終出力として平均を取ります。<br>\n<a href=\"https://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC\" target=\"_blank\">https://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC</a></p>\n<p>Pipeline 3: Cell-level Model<br>\nMotivation: we can train at cell-level using the image-level labels but it's a bit counter intuitive. Since his will introduce lots of noise as image-level labels are not ground truth for cells so we think it's beneficial to train less epochs. We ended up only training 2 epochs (1 with backbone freezed and 1 with backbone unfreezed).<br>\n画像レベルのラベルを使用してセルレベルでトレーニングできますが、少し直感的ではありません。 画像レベルのラベルはセルのグラウンドトゥルースではないため、彼は多くのノイズを導入するため、エポックを少なくすることが有益であると考えています。 最終的には2つのエポックのみをトレーニングしました（1つはバックボーンがフリーズされ、1つはバックボーンがフリーズされていません）。</p>\n<p>Data: we used all data available on HPA official website, use the cell segmentor to crop the cells and resized the cells to 168 using only RGB 3 channels. There are a total of 1620178 cropped cells.<br>\nTraining details: we used fastai's built-in finetune and fastai's learning rate finder to train only 2 epochs with the same class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and bce loss. We did not use anything for validation.<br>\nHPAの公式Webサイトで入手可能なすべてのデータを使用し、セルセグメンターを使用してセルをトリミングし、RGB3チャネルのみを使用してセルのサイズを168に変更しました。 合計1620178個のトリミングされたセルがあります。<br>\nトレーニングの詳細：fastaiの組み込みの微調整とfastaiの学習率ファインダーを使用して、同じクラスの重み[0.1、1。、0.5、1。、1.、1.、1.、0.5、1。、 1.、1.、10.、1.、0.5、0.5、5、0.2、0.5、1。]およびbce損失。 検証には何も使用しませんでした。</p>\n<p>Result: we had 10 (10x1folds) models. We predicted at cell-level and simply took the mean of the final output.<br>\n<a href=\"https://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE\" target=\"_blank\">https://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE</a></p>\n<p>Segmentation Model<br>\nWe are inspired by <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> Even Faster HPA Cell Segmentation and <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> Segmentation with a Scaling Factor, we modified the original HPA Segmentator to gain speed but keep the segmentation quality.</p>\n<p>Post-processing: We slightly changed label_cell function from the original implementation. We found that in many cases, border cells are segmented in a wrong way: some of them are combined together with border cells that have no nuclei (or it’s outside of the image). We tweaked the watershed distance threshold in order to separate cells masks a little bit further from each other than they were before, then we ignored the masks on the border that became separated from the main cell. Furthermore, we removed the border cells with nuclei whose area was less than a half of the median area of the non-border nuclei on the image. And we also removed the cells that did not have the corresponding nuclei. Below is an example of a difference between original label_cell implementation (left) and ours (right):<br>\nlabel_cell関数を元の実装から少し変更しました。 多くの場合、境界セルは間違った方法でセグメント化されていることがわかりました。それらの一部は、核を持たない（または画像の外側にある）境界セルと組み合わされています。 セルマスクを以前よりも少し離すために流域距離のしきい値を微調整し、メインセルから分離された境界のマスクを無視しました。 さらに、画像上の非境界核の中央値の半分未満の面積の核を持つ境界細胞を削除しました。 また、対応する核を持たない細胞も削除しました。 以下は、元のlabel_cell実装（左）と私たちの実装（右）の違いの例です。<br>\n<a href=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&amp;id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" target=\"_blank\">http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&amp;id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2</a></p>\n<p>imagehttp://<a href=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&amp;id=214638be-245a-4780-81a1-7866fc81081b&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" target=\"_blank\">www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&amp;id=214638be-245a-4780-81a1-7866fc81081b&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2</a></p>\n<p>Arcface Model<br>\nWe also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.<br>\nまた、eca_nfnet_l0バックボーンを使用してarcfaceモデルをトレーニングし、抗体IDを分類しました。 Antibody_idは、HPAの公式WebサイトのXMLファイルにあります。 合計11582の抗体IDがあり、トレーニングは非常に困難です。 arc_margin_productとbcelossを使用して15エポックのトレーニングを行い、データセット全体の特徴の埋め込みを抽出しました。 推論中の余弦類似性検索にfaiss_gpuライブラリを使用しました。 パブリックリーダーボードではうまく機能しましたが、プライベートリーダーボードではうまく機能しませんでした。</p>\n<p>Duplicate samples.<br>\nWe found about ~400 image in public test set duplicated either within the train set or the external data. You can check the csv file at <a href=\"https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample\" target=\"_blank\">https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample</a>. Our public leaderboard score, excluding the boost from duplicates is about 0.58, we have a relative consistent gap w.r.t. the 1st place in both public and private leaderboard.</p>\n<p>Things that didn't work<br>\nSegmentation post-processing on scaled-up outputs of the segmentator led to a slight decrease in the score<br>\nTiling a plot with a single cell and classifying such cells with the image level models.<br>\nSolution code<br>\nPipeline 1's code is now available at github</p>",
      "rawMarkdown": "日本語訳\n\nPreface\nAs promised, we will share our detailed solution within 24 hours. We would like to thank the organizers for this awesome competition since all of us had no experience in dealing with weakly-supervised classification problems and we have learned a lot from the the kind sharings by other kagglers and self-discoveries. The organizers are very active in this competition; huge props to all of you. I am also grateful for my teammates for making my journey to Kaggle Competition Grandmaster smooth and gratifying. To say I am excited is a huge under-statement. Without further ado, let's dive into our solution.\n\n約束どおり、24時間以内に詳細なソリューションを共有します。 教師なし分類の問題に対処した経験がなく、他のカグラーによる親切な共有や自己発見から多くのことを学んだので、この素晴らしいコンテストの主催者に感謝します。 主催者はこの大会に非常に積極的です。 皆さんへの巨大な小道具。 また、Kaggleコンペティションのグランドマスターへの旅をスムーズで満足のいくものにしてくれたチームメートにも感謝しています。 私が興奮していると言うことは、非常に控えめな表現です。 さらに面倒なことはせずに、私たちのソリューションに飛び込みましょう。\n\nTLDR\nOur solution consists of a total of 3 simple pipelines. We did not use any advanced techniques from any paper but we tried to understand the data well and build our model architecture w.r.t the problem statement. Here is a diagram for our final pipeline:\n\n私たちのソリューションは、合計3つの単純なパイプラインで構成されています。 どの論文からも高度な手法を使用しませんでしたが、データを十分に理解し、問題ステートメントを使用してモデルアーキテクチャを構築しようとしました。 これが最終パイプラインの図です。\n\nhttp://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2\n\nPipeline 1: Duo-Branch Cell Model\nMotivation: A duo-branch(head) cell model was designed in a way that it takes cell tiles as input but has the ability to predict both as cell-level and image-level. Multi-tasking has been shown to be effective in improving model learning. A strong champion in dota2 called Jakiro also has two heads.\nデュオブランチ（ヘッド）セルモデルは、セルタイルを入力として受け取るように設計されていますが、セルレベルと画像レベルの両方を予測する機能があります。 マルチタスクは、モデル学習の改善に効果的であることが示されています。 Jakiroと呼ばれるdota2の強力なチャンピオンにも2つの頭があります。\n\nLoss formulation: since the output is cell-level and image-level, we need two losses for both outputs. The final loss is the weighted sum of cell-level loss and image-level loss. We used basic BCE loss for both cell-level and image-level. For cell-level, the labels are not certain so it's intuitive to assign a lower weight (=0.1). l = 0.1*loss_cell + loss_image\nData: we used original data, external data shared by Phil as well as some rare class samples by using the API. The input size for a single cell is 256.\n出力はセルレベルと画像レベルであるため、両方の出力に2つの損失が必要です。 最終的な損失は、セルレベルの損失と画像レベルの損失の加重和です。 セルレベルと画像レベルの両方で基本的なBCE損失を使用しました。 セルレベルの場合、ラベルは明確ではないため、より低い重み（= 0.1）を割り当てるのは直感的です。 l = 0.1 * loss_cell + loss_image\nデータ：元のデータ、Philが共有する外部データ、およびAPIを使用したいくつかのまれなクラスサンプルを使用しました。 単一セルの入力サイズは256です。\n\nTraining details: we take 4-channel images and crop&resize the cells first; then we random sample N (=16) cells as input of our network. The cells are flattened as a large batch then we feed them into a CNN and backprop. For data augmentations we used dihedral, shift, rotate, scale, distortions, brightness contrast and cutout. The heavy data augmentations allows the model to better generalize as cells can be in any forms in reality. We train 5 folds and 20 epochs each. A single Model (b3, 256) takes about 30 hours to train on a single RTX3090.\n4チャンネルの画像を撮影し、最初にセルをトリミングしてサイズを変更します。 次に、ネットワークの入力としてN（= 16）個のセルをランダムにサンプリングします。 セルは大きなバッチとして平坦化され、CNNとバックプロパゲーションにフィードされます。 データ拡張には、二面角、シフト、回転、スケール、歪み、明るさのコントラスト、カットアウトを使用しました。 大量のデータ拡張により、セルは実際には任意の形式である可能性があるため、モデルをより一般化することができます。 それぞれ5つのフォールドと20のエポックをトレーニングします。 単一のモデル（b3、256）は、単一のRTX3090でトレーニングするのに約30時間かかります。\n\nInference: for each image, the pooled features are concatenated and feed into the last linear layer to predict at a cell-level. We generate image level prediction and cell level prediction and calculate their product as our final prediction.\nResult: with 16xTTA(scale, rotate, flip at random), 256 size cell-tiles, N=16, ensemble of efficientnet B3, B5 resnet200d and se_resnext50 backbone, our model score 0.550 on the public leaderboard and 0.550 on the private leaderboard. This single architecture can achieve second place in this competition.\n画像ごとに、プールされた特徴が連結され、最後の線形レイヤーにフィードされて、セルレベルで予測されます。 画像レベル予測と細胞レベル予測を生成し、それらの積を最終予測として計算します。\n結果：16xTTA（スケール、回転、ランダムに反転）、256サイズのセルタイル、N = 16、効率的なネットB3、B5 resnet200d、se_resnext50バックボーンのアンサンブルで、モデルスコアはパブリックリーダーボードで0.550、プライベートリーダーボードで0.550です。 この単一のアーキテクチャは、この競争で2位を獲得することができます。\nhttp://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2\n\nPipeline 2: Image-level Model\nMisc: The image level model is similar to last HPA's competition. We feed in whole images and perform a multi-label classification. It may surprise you, this pipeline is developed in fastai. We have found fastai's many implementations to be extremely slow (for instance, resize) and had gone through many days of debugging during the final inference phase; we spent a whole last week attempting to figure out how to submit. In the end, we optimized the fast.ai inference code by a lot that helped us cut the inference time almost twice. Luckily, hard work paid off.\nMotivation: we can train at an image level but predict at a cell-level (with other cells masked) and the result is very promising. We decide to add this our pipeline.\n画像レベルのモデルは、前回のHPAの競合製品と似ています。 画像全体をフィードし、マルチラベル分類を実行します。 驚かれるかもしれませんが、このパイプラインはfastaiで開発されています。 fastaiの多くの実装は非常に遅く（たとえば、サイズ変更）、最終的な推論フェーズで何日もデバッグを行っていました。 私たちは先週、提出方法を理解しようと一週間を費やしました。 最終的に、fast.ai推論コードを大幅に最適化して、推論時間をほぼ2倍に短縮しました。 幸いなことに、努力は報われました。\n動機：画像レベルでトレーニングすることはできますが、セルレベルで（他のセルをマスクして）予測することができ、その結果は非常に有望です。 これをパイプラインに追加することにしました。\n\nData: we used all data available on HPA official website and resized it to 512 using only RGB 3 channels.\nTraining details: we train 20 epochs with class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and BCE loss for 2 folds only. We used average precision score for checkpointing. For data augmentations, we used fastai's aug_transforms(flip_vert=True, max_lighting=0.1, max_warp=0.1, p_affine=0.5, p_lighting=0.5)\nResult: we had 10 (5x2folds) models and we took the mean of the final output. And we use them to predict at both cell-level and image-level. We take the mean as our final output.\nHPAの公式Webサイトで入手可能なすべてのデータを使用し、RGB3チャネルのみを使用して512にサイズ変更しました。\nトレーニングの詳細：クラスの重み[0.1、1。、0.5、1。、1.、1.、1.、0.5、1。、1.、1.、10.、1.、0.5、0.5で20エポックをトレーニングします 、5、0.2、0.5、1。]および2倍のみのBCE損失。 チェックポイントには平均精度スコアを使用しました。 データ拡張には、fastaiのaug_transforms（flip_vert = True、max_lighting = 0.1、max_warp = 0.1、p_affine = 0.5、p_lighting = 0.5）を使用しました。\n結果：10（5x2folds）モデルがあり、最終出力の平均を取りました。 そして、それらを使用して、セルレベルと画像レベルの両方で予測します。 最終出力として平均を取ります。\nhttps://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC\n\nPipeline 3: Cell-level Model\nMotivation: we can train at cell-level using the image-level labels but it's a bit counter intuitive. Since his will introduce lots of noise as image-level labels are not ground truth for cells so we think it's beneficial to train less epochs. We ended up only training 2 epochs (1 with backbone freezed and 1 with backbone unfreezed).\n画像レベルのラベルを使用してセルレベルでトレーニングできますが、少し直感的ではありません。 画像レベルのラベルはセルのグラウンドトゥルースではないため、彼は多くのノイズを導入するため、エポックを少なくすることが有益であると考えています。 最終的には2つのエポックのみをトレーニングしました（1つはバックボーンがフリーズされ、1つはバックボーンがフリーズされていません）。\n\nData: we used all data available on HPA official website, use the cell segmentor to crop the cells and resized the cells to 168 using only RGB 3 channels. There are a total of 1620178 cropped cells.\nTraining details: we used fastai's built-in finetune and fastai's learning rate finder to train only 2 epochs with the same class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and bce loss. We did not use anything for validation.\nHPAの公式Webサイトで入手可能なすべてのデータを使用し、セルセグメンターを使用してセルをトリミングし、RGB3チャネルのみを使用してセルのサイズを168に変更しました。 合計1620178個のトリミングされたセルがあります。\nトレーニングの詳細：fastaiの組み込みの微調整とfastaiの学習率ファインダーを使用して、同じクラスの重み[0.1、1。、0.5、1。、1.、1.、1.、0.5、1。、 1.、1.、10.、1.、0.5、0.5、5、0.2、0.5、1。]およびbce損失。 検証には何も使用しませんでした。\n\nResult: we had 10 (10x1folds) models. We predicted at cell-level and simply took the mean of the final output.\nhttps://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE\n\nSegmentation Model\nWe are inspired by @samusram Even Faster HPA Cell Segmentation and @alexanderriedel Segmentation with a Scaling Factor, we modified the original HPA Segmentator to gain speed but keep the segmentation quality.\n\nPost-processing: We slightly changed label_cell function from the original implementation. We found that in many cases, border cells are segmented in a wrong way: some of them are combined together with border cells that have no nuclei (or it’s outside of the image). We tweaked the watershed distance threshold in order to separate cells masks a little bit further from each other than they were before, then we ignored the masks on the border that became separated from the main cell. Furthermore, we removed the border cells with nuclei whose area was less than a half of the median area of the non-border nuclei on the image. And we also removed the cells that did not have the corresponding nuclei. Below is an example of a difference between original label_cell implementation (left) and ours (right):\nlabel_cell関数を元の実装から少し変更しました。 多くの場合、境界セルは間違った方法でセグメント化されていることがわかりました。それらの一部は、核を持たない（または画像の外側にある）境界セルと組み合わされています。 セルマスクを以前よりも少し離すために流域距離のしきい値を微調整し、メインセルから分離された境界のマスクを無視しました。 さらに、画像上の非境界核の中央値の半分未満の面積の核を持つ境界細胞を削除しました。 また、対応する核を持たない細胞も削除しました。 以下は、元のlabel_cell実装（左）と私たちの実装（右）の違いの例です。\nhttp://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2\n\nimagehttp://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&id=214638be-245a-4780-81a1-7866fc81081b&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2\n\nArcface Model\nWe also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.\nまた、eca_nfnet_l0バックボーンを使用してarcfaceモデルをトレーニングし、抗体IDを分類しました。 Antibody_idは、HPAの公式WebサイトのXMLファイルにあります。 合計11582の抗体IDがあり、トレーニングは非常に困難です。 arc_margin_productとbcelossを使用して15エポックのトレーニングを行い、データセット全体の特徴の埋め込みを抽出しました。 推論中の余弦類似性検索にfaiss_gpuライブラリを使用しました。 パブリックリーダーボードではうまく機能しましたが、プライベートリーダーボードではうまく機能しませんでした。\n\nDuplicate samples.\nWe found about ~400 image in public test set duplicated either within the train set or the external data. You can check the csv file at https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample. Our public leaderboard score, excluding the boost from duplicates is about 0.58, we have a relative consistent gap w.r.t. the 1st place in both public and private leaderboard.\n\nThings that didn't work\nSegmentation post-processing on scaled-up outputs of the segmentator led to a slight decrease in the score\nTiling a plot with a single cell and classifying such cells with the image level models.\nSolution code\nPipeline 1's code is now available at github"
    },
    {
      "id": 1307263,
      "postDate": "2021-05-14T10:55:41.173Z",
      "content": "<p>A really well documented, simple but creative solution!<br>\nI particularly like pipeline 1, the idea is simple but looks promising<br>\nDo u plan to share the source code of your team solution?</p>",
      "rawMarkdown": "A really well documented, simple but creative solution!\nI particularly like pipeline 1, the idea is simple but looks promising\nDo u plan to share the source code of your team solution?",
      "votes": 1
    },
    {
      "id": 1308007,
      "postDate": "2021-05-14T21:34:22.943Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 2,
      "replies": [
        {
          "id": 1308010,
          "postDate": "2021-05-14T21:40:53.603Z",
          "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> wondering what will be the score if we use approach 2 or 3 only?</p>",
          "rawMarkdown": "@steamedsheep wondering what will be the score if we use approach 2 or 3 only?",
          "votes": 1
        },
        {
          "id": 1308019,
          "postDate": "2021-05-14T21:58:03.383Z",
          "content": "<p>We haven't tested these approaches on the private test set separately but from our early experiments I can say that the pipeline 3 alone had the lowest score. I would guess pipeline 3 should score somewhere around 0.52 max.</p>",
          "rawMarkdown": "We haven't tested these approaches on the private test set separately but from our early experiments I can say that the pipeline 3 alone had the lowest score. I would guess pipeline 3 should score somewhere around 0.52 max.",
          "votes": 1
        },
        {
          "id": 1308020,
          "postDate": "2021-05-14T21:58:51.113Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1305343,
      "postDate": "2021-05-13T08:10:14.093Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a>, <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> and <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> for being a great team and congrats <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> on becoming GM! I'm really happy that our team efforts landed us in second place. Great job team [red.ai]!</p>",
      "rawMarkdown": "Thank you @steamedsheep, @underwearfitting and @vostankovich for being a great team and congrats @steamedsheep on becoming GM! I'm really happy that our team efforts landed us in second place. Great job team [red.ai]!",
      "votes": 2,
      "replies": [
        {
          "id": 1305377,
          "postDate": "2021-05-13T08:29:07.837Z",
          "content": "<p>Yay!</p>",
          "rawMarkdown": "                                    Yay!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1305113,
      "postDate": "2021-05-13T05:22:55.317Z",
      "content": "<p>Awesome! I really love your Pipeline 1 approach, for me its the best  solution of all i ready so far, because your model seems to lean so much more and better with the dual head. How did  you determine the loss weight? Die you try to increase or decrease the weight factor on each epoch?</p>",
      "rawMarkdown": "Awesome! I really love your Pipeline 1 approach, for me its the best  solution of all i ready so far, because your model seems to lean so much more and better with the dual head. How did  you determine the loss weight? Die you try to increase or decrease the weight factor on each epoch?",
      "votes": 2,
      "replies": [
        {
          "id": 1305165,
          "postDate": "2021-05-13T06:02:53.880Z",
          "content": "<p>For the dual head model, we initially assigned 0.1 to cells as we use image-level labels and they are not ground truth for cells. Later on, we tried to increase or decrease the weight but saw no improvement. The weight is not changed at each epoch but I personally think it might help if we can change the weight on the fly. Thanks for the suggestion.</p>",
          "rawMarkdown": "For the dual head model, we initially assigned 0.1 to cells as we use image-level labels and they are not ground truth for cells. Later on, we tried to increase or decrease the weight but saw no improvement. The weight is not changed at each epoch but I personally think it might help if we can change the weight on the fly. Thanks for the suggestion.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1305070,
      "postDate": "2021-05-13T04:51:14.610Z",
      "content": "<p>Great Approach. Thanks for sharing and congrats on becoming GM <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> </p>",
      "rawMarkdown": "Great Approach. Thanks for sharing and congrats on becoming GM @steamedsheep ",
      "votes": 2
    },
    {
      "id": 1305021,
      "postDate": "2021-05-13T04:24:52.180Z",
      "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> Congratulations on 2 nd Place and Thanks for sharing the approach</p>",
      "rawMarkdown": "@steamedsheep Congratulations on 2 nd Place and Thanks for sharing the approach",
      "votes": 2
    },
    {
      "id": 1361853,
      "postDate": "2021-06-23T05:13:17.250Z",
      "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> Congratulations on 2nd place, and thank you for sharing. Great solution! I have one question. You wrote the following about metric learning, but what was difficult about it? And what did you do to overcome the difficulties?</p>\n<blockquote>\n  <p>We also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.</p>\n</blockquote>\n<p>Thanks</p>",
      "rawMarkdown": "@steamedsheep Congratulations on 2nd place, and thank you for sharing. Great solution! I have one question. You wrote the following about metric learning, but what was difficult about it? And what did you do to overcome the difficulties?\n\n> We also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.\n\nThanks"
    },
    {
      "id": 1341380,
      "postDate": "2021-06-08T16:05:30.523Z",
      "content": "<p>Thanks for sharing. I wonder what are the dimensions of <code>viewed_pooled</code> and <code>pooled</code>, the dimensions of <code>cell</code> and <code>exp</code>. Thanks </p>",
      "rawMarkdown": "Thanks for sharing. I wonder what are the dimensions of `viewed_pooled` and `pooled`, the dimensions of `cell` and `exp`. Thanks "
    },
    {
      "id": 1306719,
      "postDate": "2021-05-14T03:38:00.377Z",
      "content": "<p>Congrats on the 2nd place and thank you for sharing so interesting solution! I have two question.</p>\n<ol>\n<li>What do you use for ground truth of cell-level in both Pipeline 1 and 2?</li>\n<li>Is my understanding correct that, in pipeline 2, images fed to CNN are masked in any way, not original image? And how to mask them?</li>\n</ol>",
      "rawMarkdown": "Congrats on the 2nd place and thank you for sharing so interesting solution! I have two question.\n1. What do you use for ground truth of cell-level in both Pipeline 1 and 2?\n2. Is my understanding correct that, in pipeline 2, images fed to CNN are masked in any way, not original image? And how to mask them?",
      "replies": [
        {
          "id": 1307217,
          "postDate": "2021-05-14T10:19:41.487Z",
          "content": "<p>For pipeline 2 we trained an image-level model on original resized RGB images without any masks and used the provided labels as ground truth. During inference, we used this model for two kinds of predictions (as you can see on the diagram for pipeline 2):</p>\n<ol>\n<li>We masked small border cells (obtained from the segmentation post-processing) and applied the model to the resulting image. This corresponds to the top branch on the diagram of pipeline 2.</li>\n<li>For every cell we'd like to classify, we masked all but this one cell on a full image and applied the model to all the cells masked this way. Here is an example code to make a list of masked cells: <code>masked_batch = [full_image * (mask == cell_id).astype(np.uint8)[..., None] for cell_id in segmented_cell_ids]</code> where <code>mask</code> is the full-image cell mask we get from the segmentation post-processing and <code>segmented_cell_ids</code> is a list of cells we'd like to classify. Then, we feed this <code>masked_batch</code> to our image-level model.</li>\n</ol>",
          "rawMarkdown": "For pipeline 2 we trained an image-level model on original resized RGB images without any masks and used the provided labels as ground truth. During inference, we used this model for two kinds of predictions (as you can see on the diagram for pipeline 2):\n\n1. We masked small border cells (obtained from the segmentation post-processing) and applied the model to the resulting image. This corresponds to the top branch on the diagram of pipeline 2.\n2. For every cell we'd like to classify, we masked all but this one cell on a full image and applied the model to all the cells masked this way. Here is an example code to make a list of masked cells: `masked_batch = [full_image * (mask == cell_id).astype(np.uint8)[..., None] for cell_id in segmented_cell_ids]` where `mask` is the full-image cell mask we get from the segmentation post-processing and `segmented_cell_ids` is a list of cells we'd like to classify. Then, we feed this `masked_batch` to our image-level model.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1305516,
      "postDate": "2021-05-13T10:35:59.803Z",
      "content": "<p>Very interesting solution , and congrats on the 2nd place and grandmaster too! I had a question - How did you decide on the 16 number of cells to be sampled for the dual headed model? Did you count an average number of cells per image , or just randomly experimented with various number of cells?</p>",
      "rawMarkdown": "Very interesting solution , and congrats on the 2nd place and grandmaster too! I had a question - How did you decide on the 16 number of cells to be sampled for the dual headed model? Did you count an average number of cells per image , or just randomly experimented with various number of cells?",
      "replies": [
        {
          "id": 1306216,
          "postDate": "2021-05-13T17:08:17.010Z",
          "content": "<p>The median of cells in image is 17, due to the limit of the GPU memory, we sample 16 cells.</p>",
          "rawMarkdown": "The median of cells in image is 17, due to the limit of the GPU memory, we sample 16 cells.",
          "votes": 2
        },
        {
          "id": 1306222,
          "postDate": "2021-05-13T17:10:44.807Z",
          "content": "<p>I see. Can you share how the heads of that model were like? I would like to try and implement it :D</p>",
          "rawMarkdown": "I see. Can you share how the heads of that model were like? I would like to try and implement it :D"
        },
        {
          "id": 1306242,
          "postDate": "2021-05-13T17:20:56.730Z",
          "content": "<p>It's pretty simple, please check the code below.</p>\n<pre><code>class EfficinetNet(nn.Module):\n    def __init__(self, name='efficientnet_b0', pretrained='imagenet', out_features=81313, dropout=0.5, feature_dim=512):\n        super().__init__()\n        self.model = torch.hub.load('rwightman/gen-efficientnet-pytorch', name,\n                                    pretrained=(pretrained == 'imagenet'))\n        self.model.conv_stem = Conv2dSame(4, self.model.conv_stem.out_channels, kernel_size=(3, 3), stride=(2, 2), bias=False)\n        self.last_linear = nn.Linear(in_features=self.model.classifier.in_features, out_features=out_features)\n        self.last_linear2 = nn.Linear(in_features=self.model.classifier.in_features, out_features=out_features)\n        self.pool = GeM()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x, cnt=16):\n        x = self.model.features(x)\n        pooled = nn.Flatten()(self.pool(x))\n        viewed_pooled = pooled.view(-1, cnt, pooled.shape[-1])\n        viewed_pooled = viewed_pooled.max(1)[0]\n        return self.last_linear(self.dropout(pooled)), self.last_linear2(self.dropout(viewed_pooled))\n</code></pre>",
          "rawMarkdown": "It's pretty simple, please check the code below.\n```\nclass EfficinetNet(nn.Module):\n    def __init__(self, name='efficientnet_b0', pretrained='imagenet', out_features=81313, dropout=0.5, feature_dim=512):\n        super().__init__()\n        self.model = torch.hub.load('rwightman/gen-efficientnet-pytorch', name,\n                                    pretrained=(pretrained == 'imagenet'))\n        self.model.conv_stem = Conv2dSame(4, self.model.conv_stem.out_channels, kernel_size=(3, 3), stride=(2, 2), bias=False)\n        self.last_linear = nn.Linear(in_features=self.model.classifier.in_features, out_features=out_features)\n        self.last_linear2 = nn.Linear(in_features=self.model.classifier.in_features, out_features=out_features)\n        self.pool = GeM()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x, cnt=16):\n        x = self.model.features(x)\n        pooled = nn.Flatten()(self.pool(x))\n        viewed_pooled = pooled.view(-1, cnt, pooled.shape[-1])\n        viewed_pooled = viewed_pooled.max(1)[0]\n        return self.last_linear(self.dropout(pooled)), self.last_linear2(self.dropout(viewed_pooled))\n```\n",
          "votes": 3
        },
        {
          "id": 1306467,
          "postDate": "2021-05-13T19:24:33.357Z",
          "content": "<p>Thanks for sharing the code. Sorry for so many questions , but I have some more noob questions-</p>\n<ol>\n<li>Why is out_features = 81313?</li>\n<li>the input to the model is a flattened batch of 16 cells ,  so that is stored in pooled , and viewed_pooled is again a tensor with 16 cells separately. What does <code>viewed_pooled = viewed_pooled.max(1)[0]</code> do exactly? It seems counter-intuitive to me.<br>\nThank you again for the prompt replies!</li>\n</ol>",
          "rawMarkdown": "Thanks for sharing the code. Sorry for so many questions , but I have some more noob questions-\n1. Why is out_features = 81313?\n2. the input to the model is a flattened batch of 16 cells ,  so that is stored in pooled , and viewed_pooled is again a tensor with 16 cells separately. What does `viewed_pooled = viewed_pooled.max(1)[0]` do exactly? It seems counter-intuitive to me.\nThank you again for the prompt replies!"
        },
        {
          "id": 1306649,
          "postDate": "2021-05-14T01:06:00.840Z",
          "content": "<p>out_features=81313 is just a default value, 19 is used here. After pooling, every cell in batch have 1280 features. Here I calculate the max of the 16 cells and put input image head.  </p>",
          "rawMarkdown": "out_features=81313 is just a default value, 19 is used here. After pooling, every cell in batch have 1280 features. Here I calculate the max of the 16 cells and put input image head.  ",
          "votes": 2
        },
        {
          "id": 1306864,
          "postDate": "2021-05-14T05:58:08.247Z",
          "content": "<p>I understand. Thank you! But why are we getting the max of the 16 cells? It is possible to pass the entire feature map of the cells as it is in viewed_pooled initially, without getting max right? </p>",
          "rawMarkdown": "I understand. Thank you! But why are we getting the max of the 16 cells? It is possible to pass the entire feature map of the cells as it is in viewed_pooled initially, without getting max right? "
        },
        {
          "id": 1308040,
          "postDate": "2021-05-14T22:49:05.987Z",
          "content": "<p>Wondering if one image do not have 16 cells are you going to resample some duplicate cells?</p>",
          "rawMarkdown": "Wondering if one image do not have 16 cells are you going to resample some duplicate cells?"
        },
        {
          "id": 1308062,
          "postDate": "2021-05-15T00:05:41.180Z",
          "content": "<p>I just input all 0 to model if cell number is less than 16. <br>\nMaximumis the easiest way to get a batch size indepenedent feature map of image and also make sense. Many situation one image contains only one or two postitive cell of certain type.</p>",
          "rawMarkdown": "I just input all 0 to model if cell number is less than 16. \nMaximumis the easiest way to get a batch size indepenedent feature map of image and also make sense. Many situation one image contains only one or two postitive cell of certain type.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1305463,
      "postDate": "2021-05-13T09:42:51.787Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1304907,
      "postDate": "2021-05-13T01:33:14.087Z",
      "content": "<p>great illustrations. Thanks for sharing</p>",
      "rawMarkdown": "great illustrations. Thanks for sharing",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1308072,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-05-15T00:39:40.720000",
      "content": "<p>Congratulations sheep and team. Great solution. Congratulations sheep on achieving Kaggle Competition Grandmaster!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1307282,
      "author_name": "Vladislav Ostankovich",
      "author_url": "",
      "post_date": "2021-05-14T11:05:30.177000",
      "content": "<p>If anyone interested in the inference code for pipelines 2&amp;3 (including upgraded segmentation postprocessing and some visualizations) I made the kernel public <a href=\"https://www.kaggle.com/vostankovich/crops-full-crops-full-arcface-ensemble-optimize?scriptVersionId=61646005\" target=\"_blank\">here</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1610733,
      "author_name": "pixyz0130",
      "author_url": "",
      "post_date": "2021-12-07T13:35:19.373000",
      "content": "<p>日本語訳</p>\n<p>Preface<br>\nAs promised, we will share our detailed solution within 24 hours. We would like to thank the organizers for this awesome competition since all of us had no experience in dealing with weakly-supervised classification problems and we have learned a lot from the the kind sharings by other kagglers and self-discoveries. The organizers are very active in this competition; huge props to all of you. I am also grateful for my teammates for making my journey to Kaggle Competition Grandmaster smooth and gratifying. To say I am excited is a huge under-statement. Without further ado, let's dive into our solution.</p>\n<p>約束どおり、24時間以内に詳細なソリューションを共有します。 教師なし分類の問題に対処した経験がなく、他のカグラーによる親切な共有や自己発見から多くのことを学んだので、この素晴らしいコンテストの主催者に感謝します。 主催者はこの大会に非常に積極的です。 皆さんへの巨大な小道具。 また、Kaggleコンペティションのグランドマスターへの旅をスムーズで満足のいくものにしてくれたチームメートにも感謝しています。 私が興奮していると言うことは、非常に控えめな表現です。 さらに面倒なことはせずに、私たちのソリューションに飛び込みましょう。</p>\n<p>TLDR<br>\nOur solution consists of a total of 3 simple pipelines. We did not use any advanced techniques from any paper but we tried to understand the data well and build our model architecture w.r.t the problem statement. Here is a diagram for our final pipeline:</p>\n<p>私たちのソリューションは、合計3つの単純なパイプラインで構成されています。 どの論文からも高度な手法を使用しませんでしたが、データを十分に理解し、問題ステートメントを使用してモデルアーキテクチャを構築しようとしました。 これが最終パイプラインの図です。</p>\n<p><a href=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&amp;id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" target=\"_blank\">http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&amp;id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2</a></p>\n<p>Pipeline 1: Duo-Branch Cell Model<br>\nMotivation: A duo-branch(head) cell model was designed in a way that it takes cell tiles as input but has the ability to predict both as cell-level and image-level. Multi-tasking has been shown to be effective in improving model learning. A strong champion in dota2 called Jakiro also has two heads.<br>\nデュオブランチ（ヘッド）セルモデルは、セルタイルを入力として受け取るように設計されていますが、セルレベルと画像レベルの両方を予測する機能があります。 マルチタスクは、モデル学習の改善に効果的であることが示されています。 Jakiroと呼ばれるdota2の強力なチャンピオンにも2つの頭があります。</p>\n<p>Loss formulation: since the output is cell-level and image-level, we need two losses for both outputs. The final loss is the weighted sum of cell-level loss and image-level loss. We used basic BCE loss for both cell-level and image-level. For cell-level, the labels are not certain so it's intuitive to assign a lower weight (=0.1). l = 0.1*loss_cell + loss_image<br>\nData: we used original data, external data shared by Phil as well as some rare class samples by using the API. The input size for a single cell is 256.<br>\n出力はセルレベルと画像レベルであるため、両方の出力に2つの損失が必要です。 最終的な損失は、セルレベルの損失と画像レベルの損失の加重和です。 セルレベルと画像レベルの両方で基本的なBCE損失を使用しました。 セルレベルの場合、ラベルは明確ではないため、より低い重み（= 0.1）を割り当てるのは直感的です。 l = 0.1 * loss_cell + loss_image<br>\nデータ：元のデータ、Philが共有する外部データ、およびAPIを使用したいくつかのまれなクラスサンプルを使用しました。 単一セルの入力サイズは256です。</p>\n<p>Training details: we take 4-channel images and crop&amp;resize the cells first; then we random sample N (=16) cells as input of our network. The cells are flattened as a large batch then we feed them into a CNN and backprop. For data augmentations we used dihedral, shift, rotate, scale, distortions, brightness contrast and cutout. The heavy data augmentations allows the model to better generalize as cells can be in any forms in reality. We train 5 folds and 20 epochs each. A single Model (b3, 256) takes about 30 hours to train on a single RTX3090.<br>\n4チャンネルの画像を撮影し、最初にセルをトリミングしてサイズを変更します。 次に、ネットワークの入力としてN（= 16）個のセルをランダムにサンプリングします。 セルは大きなバッチとして平坦化され、CNNとバックプロパゲーションにフィードされます。 データ拡張には、二面角、シフト、回転、スケール、歪み、明るさのコントラスト、カットアウトを使用しました。 大量のデータ拡張により、セルは実際には任意の形式である可能性があるため、モデルをより一般化することができます。 それぞれ5つのフォールドと20のエポックをトレーニングします。 単一のモデル（b3、256）は、単一のRTX3090でトレーニングするのに約30時間かかります。</p>\n<p>Inference: for each image, the pooled features are concatenated and feed into the last linear layer to predict at a cell-level. We generate image level prediction and cell level prediction and calculate their product as our final prediction.<br>\nResult: with 16xTTA(scale, rotate, flip at random), 256 size cell-tiles, N=16, ensemble of efficientnet B3, B5 resnet200d and se_resnext50 backbone, our model score 0.550 on the public leaderboard and 0.550 on the private leaderboard. This single architecture can achieve second place in this competition.<br>\n画像ごとに、プールされた特徴が連結され、最後の線形レイヤーにフィードされて、セルレベルで予測されます。 画像レベル予測と細胞レベル予測を生成し、それらの積を最終予測として計算します。<br>\n結果：16xTTA（スケール、回転、ランダムに反転）、256サイズのセルタイル、N = 16、効率的なネットB3、B5 resnet200d、se_resnext50バックボーンのアンサンブルで、モデルスコアはパブリックリーダーボードで0.550、プライベートリーダーボードで0.550です。 この単一のアーキテクチャは、この競争で2位を獲得することができます。<br>\n<a href=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&amp;id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" target=\"_blank\">http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&amp;id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2</a></p>\n<p>Pipeline 2: Image-level Model<br>\nMisc: The image level model is similar to last HPA's competition. We feed in whole images and perform a multi-label classification. It may surprise you, this pipeline is developed in fastai. We have found fastai's many implementations to be extremely slow (for instance, resize) and had gone through many days of debugging during the final inference phase; we spent a whole last week attempting to figure out how to submit. In the end, we optimized the fast.ai inference code by a lot that helped us cut the inference time almost twice. Luckily, hard work paid off.<br>\nMotivation: we can train at an image level but predict at a cell-level (with other cells masked) and the result is very promising. We decide to add this our pipeline.<br>\n画像レベルのモデルは、前回のHPAの競合製品と似ています。 画像全体をフィードし、マルチラベル分類を実行します。 驚かれるかもしれませんが、このパイプラインはfastaiで開発されています。 fastaiの多くの実装は非常に遅く（たとえば、サイズ変更）、最終的な推論フェーズで何日もデバッグを行っていました。 私たちは先週、提出方法を理解しようと一週間を費やしました。 最終的に、fast.ai推論コードを大幅に最適化して、推論時間をほぼ2倍に短縮しました。 幸いなことに、努力は報われました。<br>\n動機：画像レベルでトレーニングすることはできますが、セルレベルで（他のセルをマスクして）予測することができ、その結果は非常に有望です。 これをパイプラインに追加することにしました。</p>\n<p>Data: we used all data available on HPA official website and resized it to 512 using only RGB 3 channels.<br>\nTraining details: we train 20 epochs with class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and BCE loss for 2 folds only. We used average precision score for checkpointing. For data augmentations, we used fastai's aug_transforms(flip_vert=True, max_lighting=0.1, max_warp=0.1, p_affine=0.5, p_lighting=0.5)<br>\nResult: we had 10 (5x2folds) models and we took the mean of the final output. And we use them to predict at both cell-level and image-level. We take the mean as our final output.<br>\nHPAの公式Webサイトで入手可能なすべてのデータを使用し、RGB3チャネルのみを使用して512にサイズ変更しました。<br>\nトレーニングの詳細：クラスの重み[0.1、1。、0.5、1。、1.、1.、1.、0.5、1。、1.、1.、10.、1.、0.5、0.5で20エポックをトレーニングします 、5、0.2、0.5、1。]および2倍のみのBCE損失。 チェックポイントには平均精度スコアを使用しました。 データ拡張には、fastaiのaug_transforms（flip_vert = True、max_lighting = 0.1、max_warp = 0.1、p_affine = 0.5、p_lighting = 0.5）を使用しました。<br>\n結果：10（5x2folds）モデルがあり、最終出力の平均を取りました。 そして、それらを使用して、セルレベルと画像レベルの両方で予測します。 最終出力として平均を取ります。<br>\n<a href=\"https://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC\" target=\"_blank\">https://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC</a></p>\n<p>Pipeline 3: Cell-level Model<br>\nMotivation: we can train at cell-level using the image-level labels but it's a bit counter intuitive. Since his will introduce lots of noise as image-level labels are not ground truth for cells so we think it's beneficial to train less epochs. We ended up only training 2 epochs (1 with backbone freezed and 1 with backbone unfreezed).<br>\n画像レベルのラベルを使用してセルレベルでトレーニングできますが、少し直感的ではありません。 画像レベルのラベルはセルのグラウンドトゥルースではないため、彼は多くのノイズを導入するため、エポックを少なくすることが有益であると考えています。 最終的には2つのエポックのみをトレーニングしました（1つはバックボーンがフリーズされ、1つはバックボーンがフリーズされていません）。</p>\n<p>Data: we used all data available on HPA official website, use the cell segmentor to crop the cells and resized the cells to 168 using only RGB 3 channels. There are a total of 1620178 cropped cells.<br>\nTraining details: we used fastai's built-in finetune and fastai's learning rate finder to train only 2 epochs with the same class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and bce loss. We did not use anything for validation.<br>\nHPAの公式Webサイトで入手可能なすべてのデータを使用し、セルセグメンターを使用してセルをトリミングし、RGB3チャネルのみを使用してセルのサイズを168に変更しました。 合計1620178個のトリミングされたセルがあります。<br>\nトレーニングの詳細：fastaiの組み込みの微調整とfastaiの学習率ファインダーを使用して、同じクラスの重み[0.1、1。、0.5、1。、1.、1.、1.、0.5、1。、 1.、1.、10.、1.、0.5、0.5、5、0.2、0.5、1。]およびbce損失。 検証には何も使用しませんでした。</p>\n<p>Result: we had 10 (10x1folds) models. We predicted at cell-level and simply took the mean of the final output.<br>\n<a href=\"https://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE\" target=\"_blank\">https://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE</a></p>\n<p>Segmentation Model<br>\nWe are inspired by <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> Even Faster HPA Cell Segmentation and <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> Segmentation with a Scaling Factor, we modified the original HPA Segmentator to gain speed but keep the segmentation quality.</p>\n<p>Post-processing: We slightly changed label_cell function from the original implementation. We found that in many cases, border cells are segmented in a wrong way: some of them are combined together with border cells that have no nuclei (or it’s outside of the image). We tweaked the watershed distance threshold in order to separate cells masks a little bit further from each other than they were before, then we ignored the masks on the border that became separated from the main cell. Furthermore, we removed the border cells with nuclei whose area was less than a half of the median area of the non-border nuclei on the image. And we also removed the cells that did not have the corresponding nuclei. Below is an example of a difference between original label_cell implementation (left) and ours (right):<br>\nlabel_cell関数を元の実装から少し変更しました。 多くの場合、境界セルは間違った方法でセグメント化されていることがわかりました。それらの一部は、核を持たない（または画像の外側にある）境界セルと組み合わされています。 セルマスクを以前よりも少し離すために流域距離のしきい値を微調整し、メインセルから分離された境界のマスクを無視しました。 さらに、画像上の非境界核の中央値の半分未満の面積の核を持つ境界細胞を削除しました。 また、対応する核を持たない細胞も削除しました。 以下は、元のlabel_cell実装（左）と私たちの実装（右）の違いの例です。<br>\n<a href=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&amp;id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" target=\"_blank\">http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&amp;id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2</a></p>\n<p>imagehttp://<a href=\"http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&amp;id=214638be-245a-4780-81a1-7866fc81081b&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2\" target=\"_blank\">www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&amp;id=214638be-245a-4780-81a1-7866fc81081b&amp;width=2610&amp;userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&amp;cache=v2</a></p>\n<p>Arcface Model<br>\nWe also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.<br>\nまた、eca_nfnet_l0バックボーンを使用してarcfaceモデルをトレーニングし、抗体IDを分類しました。 Antibody_idは、HPAの公式WebサイトのXMLファイルにあります。 合計11582の抗体IDがあり、トレーニングは非常に困難です。 arc_margin_productとbcelossを使用して15エポックのトレーニングを行い、データセット全体の特徴の埋め込みを抽出しました。 推論中の余弦類似性検索にfaiss_gpuライブラリを使用しました。 パブリックリーダーボードではうまく機能しましたが、プライベートリーダーボードではうまく機能しませんでした。</p>\n<p>Duplicate samples.<br>\nWe found about ~400 image in public test set duplicated either within the train set or the external data. You can check the csv file at <a href=\"https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample\" target=\"_blank\">https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample</a>. Our public leaderboard score, excluding the boost from duplicates is about 0.58, we have a relative consistent gap w.r.t. the 1st place in both public and private leaderboard.</p>\n<p>Things that didn't work<br>\nSegmentation post-processing on scaled-up outputs of the segmentator led to a slight decrease in the score<br>\nTiling a plot with a single cell and classifying such cells with the image level models.<br>\nSolution code<br>\nPipeline 1's code is now available at github</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1307263,
      "author_name": "Alex Lau",
      "author_url": "",
      "post_date": "2021-05-14T10:55:41.173000",
      "content": "<p>A really well documented, simple but creative solution!<br>\nI particularly like pipeline 1, the idea is simple but looks promising<br>\nDo u plan to share the source code of your team solution?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1308007,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2021-05-14T21:34:22.943000",
      "content": "<p>Congratulations!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1308010,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2021-05-14T21:40:53.603000",
          "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> wondering what will be the score if we use approach 2 or 3 only?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1308019,
          "author_name": "Ilya Makarov",
          "author_url": "",
          "post_date": "2021-05-14T21:58:03.383000",
          "content": "<p>We haven't tested these approaches on the private test set separately but from our early experiments I can say that the pipeline 3 alone had the lowest score. I would guess pipeline 3 should score somewhere around 0.52 max.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1308020,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-14T21:58:51.113000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1305343,
      "author_name": "Ilya Makarov",
      "author_url": "",
      "post_date": "2021-05-13T08:10:14.093000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a>, <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> and <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> for being a great team and congrats <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> on becoming GM! I'm really happy that our team efforts landed us in second place. Great job team [red.ai]!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1305377,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2021-05-13T08:29:07.837000",
          "content": "<p>Yay!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1305113,
      "author_name": "Alexander Riedel",
      "author_url": "",
      "post_date": "2021-05-13T05:22:55.317000",
      "content": "<p>Awesome! I really love your Pipeline 1 approach, for me its the best  solution of all i ready so far, because your model seems to lean so much more and better with the dual head. How did  you determine the loss weight? Die you try to increase or decrease the weight factor on each epoch?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1305165,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2021-05-13T06:02:53.880000",
          "content": "<p>For the dual head model, we initially assigned 0.1 to cells as we use image-level labels and they are not ground truth for cells. Later on, we tried to increase or decrease the weight but saw no improvement. The weight is not changed at each epoch but I personally think it might help if we can change the weight on the fly. Thanks for the suggestion.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1305070,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-05-13T04:51:14.610000",
      "content": "<p>Great Approach. Thanks for sharing and congrats on becoming GM <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1305021,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:24:52.180000",
      "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> Congratulations on 2 nd Place and Thanks for sharing the approach</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1361853,
      "author_name": "KaizaburoChubachi",
      "author_url": "",
      "post_date": "2021-06-23T05:13:17.250000",
      "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> Congratulations on 2nd place, and thank you for sharing. Great solution! I have one question. You wrote the following about metric learning, but what was difficult about it? And what did you do to overcome the difficulties?</p>\n<blockquote>\n  <p>We also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.</p>\n</blockquote>\n<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1341380,
      "author_name": "yizhezx",
      "author_url": "",
      "post_date": "2021-06-08T16:05:30.523000",
      "content": "<p>Thanks for sharing. I wonder what are the dimensions of <code>viewed_pooled</code> and <code>pooled</code>, the dimensions of <code>cell</code> and <code>exp</code>. Thanks </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1306719,
      "author_name": "tsujino",
      "author_url": "",
      "post_date": "2021-05-14T03:38:00.377000",
      "content": "<p>Congrats on the 2nd place and thank you for sharing so interesting solution! I have two question.</p>\n<ol>\n<li>What do you use for ground truth of cell-level in both Pipeline 1 and 2?</li>\n<li>Is my understanding correct that, in pipeline 2, images fed to CNN are masked in any way, not original image? And how to mask them?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 1307217,
          "author_name": "Ilya Makarov",
          "author_url": "",
          "post_date": "2021-05-14T10:19:41.487000",
          "content": "<p>For pipeline 2 we trained an image-level model on original resized RGB images without any masks and used the provided labels as ground truth. During inference, we used this model for two kinds of predictions (as you can see on the diagram for pipeline 2):</p>\n<ol>\n<li>We masked small border cells (obtained from the segmentation post-processing) and applied the model to the resulting image. This corresponds to the top branch on the diagram of pipeline 2.</li>\n<li>For every cell we'd like to classify, we masked all but this one cell on a full image and applied the model to all the cells masked this way. Here is an example code to make a list of masked cells: <code>masked_batch = [full_image * (mask == cell_id).astype(np.uint8)[..., None] for cell_id in segmented_cell_ids]</code> where <code>mask</code> is the full-image cell mask we get from the segmentation post-processing and <code>segmented_cell_ids</code> is a list of cells we'd like to classify. Then, we feed this <code>masked_batch</code> to our image-level model.</li>\n</ol>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1305516,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2021-05-13T10:35:59.803000",
      "content": "<p>Very interesting solution , and congrats on the 2nd place and grandmaster too! I had a question - How did you decide on the 16 number of cells to be sampled for the dual headed model? Did you count an average number of cells per image , or just randomly experimented with various number of cells?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1306216,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-05-13T17:08:17.010000",
          "content": "<p>The median of cells in image is 17, due to the limit of the GPU memory, we sample 16 cells.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1306222,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2021-05-13T17:10:44.807000",
          "content": "<p>I see. Can you share how the heads of that model were like? I would like to try and implement it :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1306242,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-05-13T17:20:56.730000",
          "content": "<p>It's pretty simple, please check the code below.</p>\n<pre><code>class EfficinetNet(nn.Module):\n    def __init__(self, name='efficientnet_b0', pretrained='imagenet', out_features=81313, dropout=0.5, feature_dim=512):\n        super().__init__()\n        self.model = torch.hub.load('rwightman/gen-efficientnet-pytorch', name,\n                                    pretrained=(pretrained == 'imagenet'))\n        self.model.conv_stem = Conv2dSame(4, self.model.conv_stem.out_channels, kernel_size=(3, 3), stride=(2, 2), bias=False)\n        self.last_linear = nn.Linear(in_features=self.model.classifier.in_features, out_features=out_features)\n        self.last_linear2 = nn.Linear(in_features=self.model.classifier.in_features, out_features=out_features)\n        self.pool = GeM()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x, cnt=16):\n        x = self.model.features(x)\n        pooled = nn.Flatten()(self.pool(x))\n        viewed_pooled = pooled.view(-1, cnt, pooled.shape[-1])\n        viewed_pooled = viewed_pooled.max(1)[0]\n        return self.last_linear(self.dropout(pooled)), self.last_linear2(self.dropout(viewed_pooled))\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1306467,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2021-05-13T19:24:33.357000",
          "content": "<p>Thanks for sharing the code. Sorry for so many questions , but I have some more noob questions-</p>\n<ol>\n<li>Why is out_features = 81313?</li>\n<li>the input to the model is a flattened batch of 16 cells ,  so that is stored in pooled , and viewed_pooled is again a tensor with 16 cells separately. What does <code>viewed_pooled = viewed_pooled.max(1)[0]</code> do exactly? It seems counter-intuitive to me.<br>\nThank you again for the prompt replies!</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1306649,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-05-14T01:06:00.840000",
          "content": "<p>out_features=81313 is just a default value, 19 is used here. After pooling, every cell in batch have 1280 features. Here I calculate the max of the 16 cells and put input image head.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1306864,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2021-05-14T05:58:08.247000",
          "content": "<p>I understand. Thank you! But why are we getting the max of the 16 cells? It is possible to pass the entire feature map of the cells as it is in viewed_pooled initially, without getting max right? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1308040,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2021-05-14T22:49:05.987000",
          "content": "<p>Wondering if one image do not have 16 cells are you going to resample some duplicate cells?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1308062,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-05-15T00:05:41.180000",
          "content": "<p>I just input all 0 to model if cell number is less than 16. <br>\nMaximumis the easiest way to get a batch size indepenedent feature map of image and also make sense. Many situation one image contains only one or two postitive cell of certain type.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1305463,
      "author_name": "nbswords",
      "author_url": "",
      "post_date": "2021-05-13T09:42:51.787000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1304907,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2021-05-13T01:33:14.087000",
      "content": "<p>great illustrations. Thanks for sharing</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1304849": "## Preface\nAs promised, we will share our detailed solution within 24 hours. We would like to thank the organizers for this awesome competition since all of us had no experience in dealing with weakly-supervised classification problems and we have learned a lot from the the kind sharings by other kagglers and self-discoveries. The organizers are very active in this competition; huge props to all of you. I am also grateful for my teammates for making my journey to **Kaggle Competition Grandmaster** smooth and gratifying. To say I am excited is a huge under-statement. Without further ado, let's dive into our solution.\n\n## TLDR\nOur solution consists of a total of 3 simple pipelines. We did not use any advanced techniques from any paper but we tried to understand the data well and build our model architecture w.r.t the problem statement. Here is a diagram for our final pipeline:\n\n![](http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2)\n\n## Pipeline 1: Duo-Branch Cell Model\n- Motivation: A duo-branch(head) cell model was designed in a way that it takes cell tiles as input but has the ability to predict both as cell-level and image-level. Multi-tasking has been shown to be effective in improving model learning. A strong champion in dota2 called Jakiro also has two heads.\n- Loss formulation: since the output is cell-level and image-level, we need two losses for both outputs. The final loss is the weighted sum of cell-level loss and image-level loss. We used basic **BCE** loss for both cell-level and image-level. For cell-level, the labels are not certain so it's intuitive to assign a lower weight (=0.1). l = 0.1*loss_cell + loss_image\n- Data: we used original data, external data shared by Phil as well as some rare class samples by using the API. The input size for a single cell is 256.\n- Training details: we take 4-channel images and crop&resize the cells first; then we random sample N (=16) cells as input of our network. The cells are flattened as a large batch then we feed them into a CNN and backprop. For data augmentations we used dihedral, shift, rotate, scale, distortions, brightness contrast and cutout. The heavy data augmentations allows the model to better generalize as cells can be in any forms in reality. We train 5 folds and 20 epochs each. A single Model (b3, 256) takes about 30 hours to train on a single RTX3090.\n- Inference: for each image, the pooled features are concatenated and feed into the last linear layer to predict at a cell-level. We generate image level prediction and cell level prediction and calculate their product as our final prediction.\n- Result: with 16xTTA(scale, rotate, flip at random), 256 size cell-tiles, N=16, ensemble of efficientnet B3, B5 resnet200d and se_resnext50 backbone, our model score `0.550` on the public leaderboard and `0.550` on the private leaderboard. This single architecture can achieve second place in this competition.\n\n![](http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2)\n\n## Pipeline 2: Image-level Model\n\n- **Misc:** The image level model is similar to last HPA's competition. We feed in whole images and perform a multi-label classification. It may surprise you, this pipeline is developed in fastai. We have found fastai's many implementations to be extremely slow (for instance, resize) and had gone through many days of debugging during the final inference phase; we spent a whole last week attempting to figure out how to submit. In the end, we optimized the fast.ai inference code by a lot that helped us cut the inference time almost twice. Luckily, hard work paid off.\n- **Motivation:** we can train at an image level but predict at a cell-level (with other cells masked) and the result is very promising. We decide to add this our pipeline.\n- **Data:** we used all data available on HPA official website and resized it to 512 using only RGB 3 channels.\n- **Training details:** we train 20 epochs with class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and BCE loss for 2 folds only. We used average precision score  for checkpointing. For data augmentations, we used fastai's `aug_transforms(flip_vert=True, max_lighting=0.1, max_warp=0.1, p_affine=0.5, p_lighting=0.5)`\n- **Result:** we had 10 (5x2folds) models and we took the mean of the final output. And we use them to predict at both cell-level and image-level. We take the mean as our final output.\n\n![](https://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC)\n\n## Pipeline 3: Cell-level Model\n- **Motivation:** we can train at cell-level using the image-level labels but it's a bit counter intuitive. Since his will introduce lots of noise as image-level labels are not ground truth for cells so we think it's beneficial to train less epochs. We ended up only training 2 epochs (1 with backbone freezed and 1 with backbone unfreezed).\n- **Data:** we used all data available on HPA official website, use the cell segmentor to crop the cells and resized the cells to 168 using only RGB 3 channels. There are a total of 1620178 cropped cells.\n- **Training details:** we used fastai's built-in `finetune` and fastai's learning rate finder to train only 2 epochs with the same class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and bce loss. We did not use anything for validation.\n- **Result:** we had 10 (10x1folds) models. We predicted at cell-level and simply took the mean of the final output.\n\n![](https://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE)\n\n## Segmentation Model\n\nWe are inspired by @samusram Even Faster HPA Cell Segmentation and @alexanderriedel Segmentation with a Scaling Factor, we modified the original HPA Segmentator to gain speed but keep the segmentation quality.\n* Post-processing: We slightly changed label_cell function from the original implementation. We found that in many cases, border cells are segmented in a wrong way: some of them are combined together with border cells that have no nuclei (or it’s outside of the image). We tweaked the watershed distance threshold in order to separate cells masks a little bit further from each other than they were before, then we ignored the masks on the border that became separated from the main cell. Furthermore, we removed the border cells with nuclei whose area was less than a half of the median area of the non-border nuclei on the image. And we also removed the cells that did not have the corresponding nuclei. Below is an example of a difference between original label_cell implementation (left) and ours (right):\n\n![](http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2)\n\n![](http://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&id=214638be-245a-4780-81a1-7866fc81081b&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2)\n\n## Arcface Model\nWe also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.\n\n## Duplicate samples.\nWe found about ~400 image in public test set duplicated either within the train set or the external data. You can check the csv file at https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample. Our public leaderboard score, excluding the boost from duplicates is about 0.58, we have a relative consistent gap w.r.t. the 1st place in both public and private leaderboard.\n\n## Things that didn't work\n- Segmentation post-processing on scaled-up outputs of the segmentator led to a slight decrease in the score\n- Tiling a plot with a single cell and classifying such cells with the image level models.\n\n## Solution code\n* Pipeline 1's code is now available at [github](https://github.com/iseekwonderful/HPA-singlecell-2nd-dual-head-pipeline)",
    "1308072": "Congratulations sheep and team. Great solution. Congratulations sheep on achieving Kaggle Competition Grandmaster!",
    "1307282": "If anyone interested in the inference code for pipelines 2&3 (including upgraded segmentation postprocessing and some visualizations) I made the kernel public [here](https://www.kaggle.com/vostankovich/crops-full-crops-full-arcface-ensemble-optimize?scriptVersionId=61646005)",
    "1610733": "日本語訳\n\nPreface\nAs promised, we will share our detailed solution within 24 hours. We would like to thank the organizers for this awesome competition since all of us had no experience in dealing with weakly-supervised classification problems and we have learned a lot from the the kind sharings by other kagglers and self-discoveries. The organizers are very active in this competition; huge props to all of you. I am also grateful for my teammates for making my journey to Kaggle Competition Grandmaster smooth and gratifying. To say I am excited is a huge under-statement. Without further ado, let's dive into our solution.\n\n約束どおり、24時間以内に詳細なソリューションを共有します。 教師なし分類の問題に対処した経験がなく、他のカグラーによる親切な共有や自己発見から多くのことを学んだので、この素晴らしいコンテストの主催者に感謝します。 主催者はこの大会に非常に積極的です。 皆さんへの巨大な小道具。 また、Kaggleコンペティションのグランドマスターへの旅をスムーズで満足のいくものにしてくれたチームメートにも感謝しています。 私が興奮していると言うことは、非常に控えめな表現です。 さらに面倒なことはせずに、私たちのソリューションに飛び込みましょう。\n\nTLDR\nOur solution consists of a total of 3 simple pipelines. We did not use any advanced techniques from any paper but we tried to understand the data well and build our model architecture w.r.t the problem statement. Here is a diagram for our final pipeline:\n\n私たちのソリューションは、合計3つの単純なパイプラインで構成されています。 どの論文からも高度な手法を使用しませんでしたが、データを十分に理解し、問題ステートメントを使用してモデルアーキテクチャを構築しようとしました。 これが最終パイプラインの図です。\n\nhttp://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fdf998cf9-f273-4eef-9659-2875d8726a03%2FScreen_Shot_2021-05-12_at_4.53.39_PM.png?table=block&id=d1c0c9a3-db31-4f09-80bf-aa4ffc10eab7&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2\n\nPipeline 1: Duo-Branch Cell Model\nMotivation: A duo-branch(head) cell model was designed in a way that it takes cell tiles as input but has the ability to predict both as cell-level and image-level. Multi-tasking has been shown to be effective in improving model learning. A strong champion in dota2 called Jakiro also has two heads.\nデュオブランチ（ヘッド）セルモデルは、セルタイルを入力として受け取るように設計されていますが、セルレベルと画像レベルの両方を予測する機能があります。 マルチタスクは、モデル学習の改善に効果的であることが示されています。 Jakiroと呼ばれるdota2の強力なチャンピオンにも2つの頭があります。\n\nLoss formulation: since the output is cell-level and image-level, we need two losses for both outputs. The final loss is the weighted sum of cell-level loss and image-level loss. We used basic BCE loss for both cell-level and image-level. For cell-level, the labels are not certain so it's intuitive to assign a lower weight (=0.1). l = 0.1*loss_cell + loss_image\nData: we used original data, external data shared by Phil as well as some rare class samples by using the API. The input size for a single cell is 256.\n出力はセルレベルと画像レベルであるため、両方の出力に2つの損失が必要です。 最終的な損失は、セルレベルの損失と画像レベルの損失の加重和です。 セルレベルと画像レベルの両方で基本的なBCE損失を使用しました。 セルレベルの場合、ラベルは明確ではないため、より低い重み（= 0.1）を割り当てるのは直感的です。 l = 0.1 * loss_cell + loss_image\nデータ：元のデータ、Philが共有する外部データ、およびAPIを使用したいくつかのまれなクラスサンプルを使用しました。 単一セルの入力サイズは256です。\n\nTraining details: we take 4-channel images and crop&resize the cells first; then we random sample N (=16) cells as input of our network. The cells are flattened as a large batch then we feed them into a CNN and backprop. For data augmentations we used dihedral, shift, rotate, scale, distortions, brightness contrast and cutout. The heavy data augmentations allows the model to better generalize as cells can be in any forms in reality. We train 5 folds and 20 epochs each. A single Model (b3, 256) takes about 30 hours to train on a single RTX3090.\n4チャンネルの画像を撮影し、最初にセルをトリミングしてサイズを変更します。 次に、ネットワークの入力としてN（= 16）個のセルをランダムにサンプリングします。 セルは大きなバッチとして平坦化され、CNNとバックプロパゲーションにフィードされます。 データ拡張には、二面角、シフト、回転、スケール、歪み、明るさのコントラスト、カットアウトを使用しました。 大量のデータ拡張により、セルは実際には任意の形式である可能性があるため、モデルをより一般化することができます。 それぞれ5つのフォールドと20のエポックをトレーニングします。 単一のモデル（b3、256）は、単一のRTX3090でトレーニングするのに約30時間かかります。\n\nInference: for each image, the pooled features are concatenated and feed into the last linear layer to predict at a cell-level. We generate image level prediction and cell level prediction and calculate their product as our final prediction.\nResult: with 16xTTA(scale, rotate, flip at random), 256 size cell-tiles, N=16, ensemble of efficientnet B3, B5 resnet200d and se_resnext50 backbone, our model score 0.550 on the public leaderboard and 0.550 on the private leaderboard. This single architecture can achieve second place in this competition.\n画像ごとに、プールされた特徴が連結され、最後の線形レイヤーにフィードされて、セルレベルで予測されます。 画像レベル予測と細胞レベル予測を生成し、それらの積を最終予測として計算します。\n結果：16xTTA（スケール、回転、ランダムに反転）、256サイズのセルタイル、N = 16、効率的なネットB3、B5 resnet200d、se_resnext50バックボーンのアンサンブルで、モデルスコアはパブリックリーダーボードで0.550、プライベートリーダーボードで0.550です。 この単一のアーキテクチャは、この競争で2位を獲得することができます。\nhttp://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2F62cb83d5-d1d5-4c3c-85f5-c6d931998873%2FScreen_Shot_2021-05-12_at_12.51.45_PM.png?table=block&id=be18cfce-9ac3-401e-8c9b-ebd63fc4ff38&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2\n\nPipeline 2: Image-level Model\nMisc: The image level model is similar to last HPA's competition. We feed in whole images and perform a multi-label classification. It may surprise you, this pipeline is developed in fastai. We have found fastai's many implementations to be extremely slow (for instance, resize) and had gone through many days of debugging during the final inference phase; we spent a whole last week attempting to figure out how to submit. In the end, we optimized the fast.ai inference code by a lot that helped us cut the inference time almost twice. Luckily, hard work paid off.\nMotivation: we can train at an image level but predict at a cell-level (with other cells masked) and the result is very promising. We decide to add this our pipeline.\n画像レベルのモデルは、前回のHPAの競合製品と似ています。 画像全体をフィードし、マルチラベル分類を実行します。 驚かれるかもしれませんが、このパイプラインはfastaiで開発されています。 fastaiの多くの実装は非常に遅く（たとえば、サイズ変更）、最終的な推論フェーズで何日もデバッグを行っていました。 私たちは先週、提出方法を理解しようと一週間を費やしました。 最終的に、fast.ai推論コードを大幅に最適化して、推論時間をほぼ2倍に短縮しました。 幸いなことに、努力は報われました。\n動機：画像レベルでトレーニングすることはできますが、セルレベルで（他のセルをマスクして）予測することができ、その結果は非常に有望です。 これをパイプラインに追加することにしました。\n\nData: we used all data available on HPA official website and resized it to 512 using only RGB 3 channels.\nTraining details: we train 20 epochs with class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and BCE loss for 2 folds only. We used average precision score for checkpointing. For data augmentations, we used fastai's aug_transforms(flip_vert=True, max_lighting=0.1, max_warp=0.1, p_affine=0.5, p_lighting=0.5)\nResult: we had 10 (5x2folds) models and we took the mean of the final output. And we use them to predict at both cell-level and image-level. We take the mean as our final output.\nHPAの公式Webサイトで入手可能なすべてのデータを使用し、RGB3チャネルのみを使用して512にサイズ変更しました。\nトレーニングの詳細：クラスの重み[0.1、1。、0.5、1。、1.、1.、1.、0.5、1。、1.、1.、10.、1.、0.5、0.5で20エポックをトレーニングします 、5、0.2、0.5、1。]および2倍のみのBCE損失。 チェックポイントには平均精度スコアを使用しました。 データ拡張には、fastaiのaug_transforms（flip_vert = True、max_lighting = 0.1、max_warp = 0.1、p_affine = 0.5、p_lighting = 0.5）を使用しました。\n結果：10（5x2folds）モデルがあり、最終出力の平均を取りました。 そして、それらを使用して、セルレベルと画像レベルの両方で予測します。 最終出力として平均を取ります。\nhttps://lh4.googleusercontent.com/HB9VR9q024ZJ2meKMjNz1cE_BEi07d4k2hEujYn1rv7-vNkBYsAHj46S1y0kIPaHb6C71x6WZ7DNNw2vpQ2nHi3MXitVIh1Ut20C0NtYog3GdDB0tkM8dTneY94NFq7dtl3VQOjC\n\nPipeline 3: Cell-level Model\nMotivation: we can train at cell-level using the image-level labels but it's a bit counter intuitive. Since his will introduce lots of noise as image-level labels are not ground truth for cells so we think it's beneficial to train less epochs. We ended up only training 2 epochs (1 with backbone freezed and 1 with backbone unfreezed).\n画像レベルのラベルを使用してセルレベルでトレーニングできますが、少し直感的ではありません。 画像レベルのラベルはセルのグラウンドトゥルースではないため、彼は多くのノイズを導入するため、エポックを少なくすることが有益であると考えています。 最終的には2つのエポックのみをトレーニングしました（1つはバックボーンがフリーズされ、1つはバックボーンがフリーズされていません）。\n\nData: we used all data available on HPA official website, use the cell segmentor to crop the cells and resized the cells to 168 using only RGB 3 channels. There are a total of 1620178 cropped cells.\nTraining details: we used fastai's built-in finetune and fastai's learning rate finder to train only 2 epochs with the same class weight [0.1, 1., 0.5, 1., 1., 1., 1., 0.5, 1., 1., 1., 10., 1., 0.5, 0.5, 5, 0.2, 0.5, 1.] and bce loss. We did not use anything for validation.\nHPAの公式Webサイトで入手可能なすべてのデータを使用し、セルセグメンターを使用してセルをトリミングし、RGB3チャネルのみを使用してセルのサイズを168に変更しました。 合計1620178個のトリミングされたセルがあります。\nトレーニングの詳細：fastaiの組み込みの微調整とfastaiの学習率ファインダーを使用して、同じクラスの重み[0.1、1。、0.5、1。、1.、1.、1.、0.5、1。、 1.、1.、10.、1.、0.5、0.5、5、0.2、0.5、1。]およびbce損失。 検証には何も使用しませんでした。\n\nResult: we had 10 (10x1folds) models. We predicted at cell-level and simply took the mean of the final output.\nhttps://lh6.googleusercontent.com/Y2bRKz-YpUF9MDtGrkBai9DRWtRhHfhmOOsXx57GXomcTma8d5J2oChHXk71ljKZaDOxyGs8s72ZrIYki3dyIldBsWx3Q34oKWiYd1ntJdD-Vfakss6aSB82AZ1z2UBPa2VMDCXE\n\nSegmentation Model\nWe are inspired by @samusram Even Faster HPA Cell Segmentation and @alexanderriedel Segmentation with a Scaling Factor, we modified the original HPA Segmentator to gain speed but keep the segmentation quality.\n\nPost-processing: We slightly changed label_cell function from the original implementation. We found that in many cases, border cells are segmented in a wrong way: some of them are combined together with border cells that have no nuclei (or it’s outside of the image). We tweaked the watershed distance threshold in order to separate cells masks a little bit further from each other than they were before, then we ignored the masks on the border that became separated from the main cell. Furthermore, we removed the border cells with nuclei whose area was less than a half of the median area of the non-border nuclei on the image. And we also removed the cells that did not have the corresponding nuclei. Below is an example of a difference between original label_cell implementation (left) and ours (right):\nlabel_cell関数を元の実装から少し変更しました。 多くの場合、境界セルは間違った方法でセグメント化されていることがわかりました。それらの一部は、核を持たない（または画像の外側にある）境界セルと組み合わされています。 セルマスクを以前よりも少し離すために流域距離のしきい値を微調整し、メインセルから分離された境界のマスクを無視しました。 さらに、画像上の非境界核の中央値の半分未満の面積の核を持つ境界細胞を削除しました。 また、対応する核を持たない細胞も削除しました。 以下は、元のlabel_cell実装（左）と私たちの実装（右）の違いの例です。\nhttp://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fa924aedb-fdb4-4cc2-bbda-f7dcf7749e4c%2FUntitled.png?table=block&id=6e5fb18d-9550-4b31-9b4e-09a3f2f28cae&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2\n\nimagehttp://www.notion.so/image/https%3A%2F%2Fs3-us-west-2.amazonaws.com%2Fsecure.notion-static.com%2Fabc60835-6045-45ba-bb9a-923bd4b0d7e9%2FUntitled.png?table=block&id=214638be-245a-4780-81a1-7866fc81081b&width=2610&userId=053c66e4-923c-48af-b90f-a2fe3ed3608c&cache=v2\n\nArcface Model\nWe also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.\nまた、eca_nfnet_l0バックボーンを使用してarcfaceモデルをトレーニングし、抗体IDを分類しました。 Antibody_idは、HPAの公式WebサイトのXMLファイルにあります。 合計11582の抗体IDがあり、トレーニングは非常に困難です。 arc_margin_productとbcelossを使用して15エポックのトレーニングを行い、データセット全体の特徴の埋め込みを抽出しました。 推論中の余弦類似性検索にfaiss_gpuライブラリを使用しました。 パブリックリーダーボードではうまく機能しましたが、プライベートリーダーボードではうまく機能しませんでした。\n\nDuplicate samples.\nWe found about ~400 image in public test set duplicated either within the train set or the external data. You can check the csv file at https://www.kaggle.com/steamedsheep/hpa-2021-duplicated-sample. Our public leaderboard score, excluding the boost from duplicates is about 0.58, we have a relative consistent gap w.r.t. the 1st place in both public and private leaderboard.\n\nThings that didn't work\nSegmentation post-processing on scaled-up outputs of the segmentator led to a slight decrease in the score\nTiling a plot with a single cell and classifying such cells with the image level models.\nSolution code\nPipeline 1's code is now available at github",
    "1307263": "A really well documented, simple but creative solution!\nI particularly like pipeline 1, the idea is simple but looks promising\nDo u plan to share the source code of your team solution?",
    "1308007": "Congratulations!",
    "1305343": "Thank you @steamedsheep, @underwearfitting and @vostankovich for being a great team and congrats @steamedsheep on becoming GM! I'm really happy that our team efforts landed us in second place. Great job team [red.ai]!",
    "1305113": "Awesome! I really love your Pipeline 1 approach, for me its the best  solution of all i ready so far, because your model seems to lean so much more and better with the dual head. How did  you determine the loss weight? Die you try to increase or decrease the weight factor on each epoch?",
    "1305070": "Great Approach. Thanks for sharing and congrats on becoming GM @steamedsheep ",
    "1305021": "@steamedsheep Congratulations on 2 nd Place and Thanks for sharing the approach",
    "1361853": "@steamedsheep Congratulations on 2nd place, and thank you for sharing. Great solution! I have one question. You wrote the following about metric learning, but what was difficult about it? And what did you do to overcome the difficulties?\n\n> We also trained an arcface model with eca_nfnet_l0 backbone to classify antibody_id. Antibody_id can be found on the HPA's official website in a XML file. There are a total of 11582 antibody_id and it's extremely difficult to train. We used arc_margin_product and bce loss to train for 15 epochs then extracted the feature embeddings for the whole dataset. We used faiss_gpu library for cosine similarity search during inference. It worked well on the public leaderboard but it didn't quite work on the private leaderboard.\n\nThanks",
    "1341380": "Thanks for sharing. I wonder what are the dimensions of `viewed_pooled` and `pooled`, the dimensions of `cell` and `exp`. Thanks ",
    "1306719": "Congrats on the 2nd place and thank you for sharing so interesting solution! I have two question.\n1. What do you use for ground truth of cell-level in both Pipeline 1 and 2?\n2. Is my understanding correct that, in pipeline 2, images fed to CNN are masked in any way, not original image? And how to mask them?",
    "1305516": "Very interesting solution , and congrats on the 2nd place and grandmaster too! I had a question - How did you decide on the 16 number of cells to be sampled for the dual headed model? Did you count an average number of cells per image , or just randomly experimented with various number of cells?",
    "1305463": "Thanks for sharing.",
    "1304907": "great illustrations. Thanks for sharing"
  }
}