{
  "id": 239071,
  "title": "4th Place Solution: MILIMED",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/239071",
  "author_name": "CroDoc",
  "post_date": "2021-05-14T14:11:41.404000",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<p>We are a very diverse team of computer scientists and medical doctor/students. It was our great pleasure to participate in this demanding challenge. Hope some of you find this solution useful and/or interesting.</p>\n<h1>Solution overview</h1>\n<ol>\n<li>Segmentation -&gt; HPA-Cell-Segmentation</li>\n<li>Dataset -&gt; 512x512 cell images (20% removed)</li>\n<li>Parallelization -&gt; speed-up -&gt; 3h left for inference</li>\n<li>Manual Labeling -&gt; smaller classes &amp; validation (soft labels)</li>\n<li>Pseudo-Labeling -&gt; negative labeling (&amp; positive for mitotic spindle)</li>\n<li>EfficientNetB0 Ensemble + semi-balanced data sampling</li>\n<li>Fine-tuning -&gt; on manually labeled &amp; non-labeled validation data</li>\n<li>Cell/Image Weighting -&gt; final confidence = 0.7 * cell_confidence + 0.3 * image_confidence</li>\n</ol>\n<h1>1. HPA-Cell-Segmentation</h1>\n<p>The test set was based on this segmentator so it made no sense to spend a lot of time creating a custom segmentator which could make the IoU worse. The authors of the contest said that only 10% modifications were made on the outputs of the segmentator.</p>\n<h1>2. Dataset</h1>\n<p>Our dataset was created from the Train &amp; PublicHPA 16-bit images. Seems most teams used 8-bit images in the end.</p>\n<p>Each image in the final dataset is a 512x512 image of a cell based on the cell masks from the segmentator. Padding (to square) was used to retain original height/width ratio. No surrounding pixels were used (non-cell-mask pixels). My feeling is that it might be better to use a bit larger surrounding, but did not have time to test this (this might be good for some classes such as plasma membrane).</p>\n<p>We decided to go with large images (512x512) since some labels/organelles required higher resolution and their size varied a lot based on the cell size &amp; re-scaling. E.g. sometimes the nucleus was very small and sometimes it as big as the whole image. We even tried to train nuclear organelles on images based on nuclei masks, but since the deadline was too close, we decided not to invest more time on this approach.</p>\n<p>I am eager to find out if diving the problem into nuclear and cytosolic organelles classification would yield better results. I think it would be easier to classify organelles inside the nucleus since it would be approximately the same size for each cell image this way.</p>\n<p>We used a simple heuristic to determine how much of the nuclei was outside of the image and decreased its final predicted confidence accordingly. All images with nuclei that were not present almost completely in the cropped image were removed from the train set.</p>\n<p>Similarly we tried to determine false positive segmentations by finding outliers based on the red channel and a product of the blue and yellow channel. Outliers at inference time got their confidence decreased dramatically. Outliers in the train dataset were removed completely. I assume the accuracy of this heuristic was around 50%. Since false positives were a big score crusher, this seemed acceptable.</p>\n<p>We lost around 20% of the images from the trainset.</p>\n<h1>3. Parallelization</h1>\n<p>HPA-Cell-Segmentation took quite some time so we decided to parallelize most things. Even with only two cores we got a boost in the submission time.</p>\n<p>The first boost was by running the <code>label_cell</code> function in parallell. The second boost was in running all the previously mentioned heuristics and image cropping in parallel as well.</p>\n<p>This left us with more than 3 hours for inference.</p>\n<h1>4. Manual labeling</h1>\n<p>We manually labeled smaller classes or classes with smaller % of occurrence in the initial images (e.g. mitotic spindle, aggresome, intermediate &amp; actin filaments …). We made a simple GUI and relabeled only one label at a time for an image.</p>\n<p>Mostly we would give a score from 1 to 5 on how confident we were that the given cell image contained the image-level label. These scores we transferred to soft labels. Each mapping was different (e.g. 1:0.0, 2:0.2, 3:0.7, 4:0.9, 5:1.0).</p>\n<p>In the end, we tried to create a validation set in the same way with high quality labeling. We managed to do get a few thousand examples for most classes.</p>\n<h1>5. Pseudo-labeling</h1>\n<p>Inspired by the Meta Pseudo Labels paper, we wanted to get rid of some false positives and help our models avoid overfitting. A cut off of 0.3 seemed to remove approx. 15% of image with high accuracy. Here we used an underfitt ResNet18.</p>\n<p>Later we used a better model to find more examples of mitotic spindles in a similar way, but withing the images that did not have mitotic spindle assigned. I think we found around 100 extra mitotic spindles, compared to around 250 that we found in the labeled images.</p>\n<p>In the end, we did not do this for other classes. I think we found a few aggresomes and quickly decided to skip positive labeling.</p>\n<h1>6. EfficientNetB0</h1>\n<p>This network is just awesome :) I am a big fan of solving problems with simple models, so I was quite happy when EfficientNetB0 seemed to be good enough for this challenge. We tried using B4, but it was slower to train and the results did not impress enough to continue playing with it. There was a solution that ensembled some B4-s, but no boost on the private LB.</p>\n<p>We had an 3-part ensemble with weights [0.2, 0.4, 0.4]. All EfficientNetB0s, but trained with different augmentation and loss function combinations.</p>\n<ol>\n<li><p>Single B0 - 0.2 ensemble weight<br>\nAugmentation: Flipping &amp; Rotation<br>\nLoss: FocalLoss<br>\nDid not have time to test if this network actually helped much.</p></li>\n<li><p>2 Checkpoint Ensemble B0s - 0.4 final ensemble weight<br>\nAugmentation: RandomResizer (40% chance), Flipping &amp; Rotation<br>\nLoss: FocalLoss<br>\n*RandomResizer -&gt; Resize to (RSIZE, RSIZE) + Resize back to (512, 512), where RSIZE is a random number between 256 and 384</p></li>\n<li><p>4 Checkpoint Ensemble B0s - 0.4 final ensemble weight<br>\nAugmentation: RandomResizer (30% chance), RandomPad (30% chance), Flipping, Rotation, &amp; Resize(512,512)<br>\nLoss: BCELoss<br>\n*RandomPad -&gt; pad each side (independently) with a random length between 0 and 200</p></li>\n</ol>\n<p>Since it was hard to estimate the \"best\" model on the local validation set, we used the idea of checkpoint ensembling to try to avoid overfitting &amp; boost our score.</p>\n<p>We oversampled classes with less examples or less true positives and tried to avoid overusing images with more assigned labels. </p>\n<h1>7. Fine-tuning</h1>\n<p>We fine-tuned all networks for one \"epoch\" on the validation set. Unlabeled images were used once, while images that we labeled (soft labels) were used multiple times in the \"epoch\". The \"epoch\" was around 200-250k images. We used the same augmentations &amp; loss function for each B0 as it was used during training of that network.</p>\n<h1>8. Cell/Image weighting</h1>\n<p>To extreme outliers we weighted the final confidences with the mean of confidences of all other valid cells (excluding border and outliers). The final confidence was 0.7 * cell_confidence + 0.3 * image_confidence.</p>\n<p>Seems that even extreme values such as 0.6/0.4 tend to work well here. We did not test if smaller weighting gave better score.</p>\n<h1>Conclusion</h1>\n<p>512x512 images, re-labeling, simple network (B0), fine-tuning &amp; final confidence weighting seem to work well enough for this problem.</p>\n<h1>Thank you</h1>\n<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> &amp; <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> -&gt; <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324</a><br>\nWe were not aware of this at that time. This valuable responses made a huge impact on our approach/results.</p>\n<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> -&gt; <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940</a><br>\nThank you for the motivation for the final weighting.</p>\n<p><a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a>, <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> &amp; <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> I hope there will be more opportunities to participate in your journey in the future (Kaggle or non-Kaggle related). Thank you for the ride! :)</p>",
  "messages": [
    {
      "id": 1307555,
      "postDate": "2021-05-14T14:11:41.403Z",
      "content": "<p>We are a very diverse team of computer scientists and medical doctor/students. It was our great pleasure to participate in this demanding challenge. Hope some of you find this solution useful and/or interesting.</p>\n<h1>Solution overview</h1>\n<ol>\n<li>Segmentation -&gt; HPA-Cell-Segmentation</li>\n<li>Dataset -&gt; 512x512 cell images (20% removed)</li>\n<li>Parallelization -&gt; speed-up -&gt; 3h left for inference</li>\n<li>Manual Labeling -&gt; smaller classes &amp; validation (soft labels)</li>\n<li>Pseudo-Labeling -&gt; negative labeling (&amp; positive for mitotic spindle)</li>\n<li>EfficientNetB0 Ensemble + semi-balanced data sampling</li>\n<li>Fine-tuning -&gt; on manually labeled &amp; non-labeled validation data</li>\n<li>Cell/Image Weighting -&gt; final confidence = 0.7 * cell_confidence + 0.3 * image_confidence</li>\n</ol>\n<h1>1. HPA-Cell-Segmentation</h1>\n<p>The test set was based on this segmentator so it made no sense to spend a lot of time creating a custom segmentator which could make the IoU worse. The authors of the contest said that only 10% modifications were made on the outputs of the segmentator.</p>\n<h1>2. Dataset</h1>\n<p>Our dataset was created from the Train &amp; PublicHPA 16-bit images. Seems most teams used 8-bit images in the end.</p>\n<p>Each image in the final dataset is a 512x512 image of a cell based on the cell masks from the segmentator. Padding (to square) was used to retain original height/width ratio. No surrounding pixels were used (non-cell-mask pixels). My feeling is that it might be better to use a bit larger surrounding, but did not have time to test this (this might be good for some classes such as plasma membrane).</p>\n<p>We decided to go with large images (512x512) since some labels/organelles required higher resolution and their size varied a lot based on the cell size &amp; re-scaling. E.g. sometimes the nucleus was very small and sometimes it as big as the whole image. We even tried to train nuclear organelles on images based on nuclei masks, but since the deadline was too close, we decided not to invest more time on this approach.</p>\n<p>I am eager to find out if diving the problem into nuclear and cytosolic organelles classification would yield better results. I think it would be easier to classify organelles inside the nucleus since it would be approximately the same size for each cell image this way.</p>\n<p>We used a simple heuristic to determine how much of the nuclei was outside of the image and decreased its final predicted confidence accordingly. All images with nuclei that were not present almost completely in the cropped image were removed from the train set.</p>\n<p>Similarly we tried to determine false positive segmentations by finding outliers based on the red channel and a product of the blue and yellow channel. Outliers at inference time got their confidence decreased dramatically. Outliers in the train dataset were removed completely. I assume the accuracy of this heuristic was around 50%. Since false positives were a big score crusher, this seemed acceptable.</p>\n<p>We lost around 20% of the images from the trainset.</p>\n<h1>3. Parallelization</h1>\n<p>HPA-Cell-Segmentation took quite some time so we decided to parallelize most things. Even with only two cores we got a boost in the submission time.</p>\n<p>The first boost was by running the <code>label_cell</code> function in parallell. The second boost was in running all the previously mentioned heuristics and image cropping in parallel as well.</p>\n<p>This left us with more than 3 hours for inference.</p>\n<h1>4. Manual labeling</h1>\n<p>We manually labeled smaller classes or classes with smaller % of occurrence in the initial images (e.g. mitotic spindle, aggresome, intermediate &amp; actin filaments …). We made a simple GUI and relabeled only one label at a time for an image.</p>\n<p>Mostly we would give a score from 1 to 5 on how confident we were that the given cell image contained the image-level label. These scores we transferred to soft labels. Each mapping was different (e.g. 1:0.0, 2:0.2, 3:0.7, 4:0.9, 5:1.0).</p>\n<p>In the end, we tried to create a validation set in the same way with high quality labeling. We managed to do get a few thousand examples for most classes.</p>\n<h1>5. Pseudo-labeling</h1>\n<p>Inspired by the Meta Pseudo Labels paper, we wanted to get rid of some false positives and help our models avoid overfitting. A cut off of 0.3 seemed to remove approx. 15% of image with high accuracy. Here we used an underfitt ResNet18.</p>\n<p>Later we used a better model to find more examples of mitotic spindles in a similar way, but withing the images that did not have mitotic spindle assigned. I think we found around 100 extra mitotic spindles, compared to around 250 that we found in the labeled images.</p>\n<p>In the end, we did not do this for other classes. I think we found a few aggresomes and quickly decided to skip positive labeling.</p>\n<h1>6. EfficientNetB0</h1>\n<p>This network is just awesome :) I am a big fan of solving problems with simple models, so I was quite happy when EfficientNetB0 seemed to be good enough for this challenge. We tried using B4, but it was slower to train and the results did not impress enough to continue playing with it. There was a solution that ensembled some B4-s, but no boost on the private LB.</p>\n<p>We had an 3-part ensemble with weights [0.2, 0.4, 0.4]. All EfficientNetB0s, but trained with different augmentation and loss function combinations.</p>\n<ol>\n<li><p>Single B0 - 0.2 ensemble weight<br>\nAugmentation: Flipping &amp; Rotation<br>\nLoss: FocalLoss<br>\nDid not have time to test if this network actually helped much.</p></li>\n<li><p>2 Checkpoint Ensemble B0s - 0.4 final ensemble weight<br>\nAugmentation: RandomResizer (40% chance), Flipping &amp; Rotation<br>\nLoss: FocalLoss<br>\n*RandomResizer -&gt; Resize to (RSIZE, RSIZE) + Resize back to (512, 512), where RSIZE is a random number between 256 and 384</p></li>\n<li><p>4 Checkpoint Ensemble B0s - 0.4 final ensemble weight<br>\nAugmentation: RandomResizer (30% chance), RandomPad (30% chance), Flipping, Rotation, &amp; Resize(512,512)<br>\nLoss: BCELoss<br>\n*RandomPad -&gt; pad each side (independently) with a random length between 0 and 200</p></li>\n</ol>\n<p>Since it was hard to estimate the \"best\" model on the local validation set, we used the idea of checkpoint ensembling to try to avoid overfitting &amp; boost our score.</p>\n<p>We oversampled classes with less examples or less true positives and tried to avoid overusing images with more assigned labels. </p>\n<h1>7. Fine-tuning</h1>\n<p>We fine-tuned all networks for one \"epoch\" on the validation set. Unlabeled images were used once, while images that we labeled (soft labels) were used multiple times in the \"epoch\". The \"epoch\" was around 200-250k images. We used the same augmentations &amp; loss function for each B0 as it was used during training of that network.</p>\n<h1>8. Cell/Image weighting</h1>\n<p>To extreme outliers we weighted the final confidences with the mean of confidences of all other valid cells (excluding border and outliers). The final confidence was 0.7 * cell_confidence + 0.3 * image_confidence.</p>\n<p>Seems that even extreme values such as 0.6/0.4 tend to work well here. We did not test if smaller weighting gave better score.</p>\n<h1>Conclusion</h1>\n<p>512x512 images, re-labeling, simple network (B0), fine-tuning &amp; final confidence weighting seem to work well enough for this problem.</p>\n<h1>Thank you</h1>\n<p><a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> &amp; <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> -&gt; <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324</a><br>\nWe were not aware of this at that time. This valuable responses made a huge impact on our approach/results.</p>\n<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> -&gt; <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940</a><br>\nThank you for the motivation for the final weighting.</p>\n<p><a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a>, <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> &amp; <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> I hope there will be more opportunities to participate in your journey in the future (Kaggle or non-Kaggle related). Thank you for the ride! :)</p>",
      "rawMarkdown": "We are a very diverse team of computer scientists and medical doctor/students. It was our great pleasure to participate in this demanding challenge. Hope some of you find this solution useful and/or interesting.\n\n# Solution overview\n1. Segmentation -> HPA-Cell-Segmentation\n2. Dataset -> 512x512 cell images (20% removed)\n3. Parallelization -> speed-up -> 3h left for inference\n4. Manual Labeling -> smaller classes & validation (soft labels)\n5. Pseudo-Labeling -> negative labeling (& positive for mitotic spindle)\n6. EfficientNetB0 Ensemble + semi-balanced data sampling\n7. Fine-tuning -> on manually labeled & non-labeled validation data\n8. Cell/Image Weighting -> final confidence = 0.7 \\* cell\\_confidence + 0.3 \\* image\\_confidence\n\n# 1. HPA-Cell-Segmentation\n\nThe test set was based on this segmentator so it made no sense to spend a lot of time creating a custom segmentator which could make the IoU worse. The authors of the contest said that only 10% modifications were made on the outputs of the segmentator.\n\n# 2. Dataset\n\nOur dataset was created from the Train & PublicHPA 16-bit images. Seems most teams used 8-bit images in the end.\n\nEach image in the final dataset is a 512x512 image of a cell based on the cell masks from the segmentator. Padding (to square) was used to retain original height/width ratio. No surrounding pixels were used (non-cell-mask pixels). My feeling is that it might be better to use a bit larger surrounding, but did not have time to test this (this might be good for some classes such as plasma membrane).\n\nWe decided to go with large images (512x512) since some labels/organelles required higher resolution and their size varied a lot based on the cell size & re-scaling. E.g. sometimes the nucleus was very small and sometimes it as big as the whole image. We even tried to train nuclear organelles on images based on nuclei masks, but since the deadline was too close, we decided not to invest more time on this approach.\n\nI am eager to find out if diving the problem into nuclear and cytosolic organelles classification would yield better results. I think it would be easier to classify organelles inside the nucleus since it would be approximately the same size for each cell image this way.\n\nWe used a simple heuristic to determine how much of the nuclei was outside of the image and decreased its final predicted confidence accordingly. All images with nuclei that were not present almost completely in the cropped image were removed from the train set.\n\nSimilarly we tried to determine false positive segmentations by finding outliers based on the red channel and a product of the blue and yellow channel. Outliers at inference time got their confidence decreased dramatically. Outliers in the train dataset were removed completely. I assume the accuracy of this heuristic was around 50%. Since false positives were a big score crusher, this seemed acceptable.\n\nWe lost around 20% of the images from the trainset.\n\n# 3. Parallelization\n\nHPA-Cell-Segmentation took quite some time so we decided to parallelize most things. Even with only two cores we got a boost in the submission time.\n\nThe first boost was by running the `label_cell` function in parallell. The second boost was in running all the previously mentioned heuristics and image cropping in parallel as well.\n\nThis left us with more than 3 hours for inference.\n\n# 4. Manual labeling\n\nWe manually labeled smaller classes or classes with smaller % of occurrence in the initial images (e.g. mitotic spindle, aggresome, intermediate & actin filaments ...). We made a simple GUI and relabeled only one label at a time for an image.\n\nMostly we would give a score from 1 to 5 on how confident we were that the given cell image contained the image-level label. These scores we transferred to soft labels. Each mapping was different (e.g. 1:0.0, 2:0.2, 3:0.7, 4:0.9, 5:1.0).\n\nIn the end, we tried to create a validation set in the same way with high quality labeling. We managed to do get a few thousand examples for most classes.\n\n# 5. Pseudo-labeling\nInspired by the Meta Pseudo Labels paper, we wanted to get rid of some false positives and help our models avoid overfitting. A cut off of 0.3 seemed to remove approx. 15% of image with high accuracy. Here we used an underfitt ResNet18.\n\nLater we used a better model to find more examples of mitotic spindles in a similar way, but withing the images that did not have mitotic spindle assigned. I think we found around 100 extra mitotic spindles, compared to around 250 that we found in the labeled images.\n\nIn the end, we did not do this for other classes. I think we found a few aggresomes and quickly decided to skip positive labeling.\n\n# 6. EfficientNetB0\nThis network is just awesome :) I am a big fan of solving problems with simple models, so I was quite happy when EfficientNetB0 seemed to be good enough for this challenge. We tried using B4, but it was slower to train and the results did not impress enough to continue playing with it. There was a solution that ensembled some B4-s, but no boost on the private LB.\n\nWe had an 3-part ensemble with weights [0.2, 0.4, 0.4]. All EfficientNetB0s, but trained with different augmentation and loss function combinations.\n\n1. Single B0 - 0.2 ensemble weight\nAugmentation: Flipping & Rotation\nLoss: FocalLoss\nDid not have time to test if this network actually helped much.\n\n2. 2 Checkpoint Ensemble B0s - 0.4 final ensemble weight\nAugmentation: RandomResizer (40% chance), Flipping & Rotation\nLoss: FocalLoss\n*RandomResizer -> Resize to (RSIZE, RSIZE) + Resize back to (512, 512), where RSIZE is a random number between 256 and 384\n\n3. 4 Checkpoint Ensemble B0s - 0.4 final ensemble weight\nAugmentation: RandomResizer (30% chance), RandomPad (30% chance), Flipping, Rotation, & Resize(512,512)\nLoss: BCELoss\n*RandomPad -> pad each side (independently) with a random length between 0 and 200\n\nSince it was hard to estimate the \"best\" model on the local validation set, we used the idea of checkpoint ensembling to try to avoid overfitting & boost our score.\n\nWe oversampled classes with less examples or less true positives and tried to avoid overusing images with more assigned labels. \n\n# 7. Fine-tuning\n\nWe fine-tuned all networks for one \"epoch\" on the validation set. Unlabeled images were used once, while images that we labeled (soft labels) were used multiple times in the \"epoch\". The \"epoch\" was around 200-250k images. We used the same augmentations & loss function for each B0 as it was used during training of that network.\n\n# 8. Cell/Image weighting\n\nTo extreme outliers we weighted the final confidences with the mean of confidences of all other valid cells (excluding border and outliers). The final confidence was 0.7 \\* cell\\_confidence + 0.3 \\* image\\_confidence.\n\nSeems that even extreme values such as 0.6/0.4 tend to work well here. We did not test if smaller weighting gave better score.\n\n# Conclusion\n\n512x512 images, re-labeling, simple network (B0), fine-tuning & final confidence weighting seem to work well enough for this problem.\n\n# Thank you\n@christofhenkel & @cwinsnes -> [https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324](url)\nWe were not aware of this at that time. This valuable responses made a huge impact on our approach/results.\n\n@h053473666 -> [https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940](url)\nThank you for the motivation for the final weighting.\n\n@emmalumpan, @lnhtrang & @cwinsnes I hope there will be more opportunities to participate in your journey in the future (Kaggle or non-Kaggle related). Thank you for the ride! :)",
      "votes": 15
    },
    {
      "id": 1610751,
      "postDate": "2021-12-07T13:47:39.983Z",
      "content": "<p>日本語訳<br>\nWe are a very diverse team of computer scientists and medical doctor/students. It was our great pleasure to participate in this demanding challenge. Hope some of you find this solution useful and/or interesting.</p>\n<p>Solution overview<br>\nSegmentation -&gt; HPA-Cell-Segmentation<br>\nDataset -&gt; 512x512 cell images (20% removed)<br>\nParallelization -&gt; speed-up -&gt; 3h left for inference<br>\nManual Labeling -&gt; smaller classes &amp; validation (soft labels)<br>\nPseudo-Labeling -&gt; negative labeling (&amp; positive for mitotic spindle)<br>\nEfficientNetB0 Ensemble + semi-balanced data sampling<br>\nFine-tuning -&gt; on manually labeled &amp; non-labeled validation data<br>\nCell/Image Weighting -&gt; final confidence = 0.7 * cell_confidence + 0.3 * image_confidence</p>\n<ol>\n<li><p>HPA-Cell-Segmentation<br>\nThe test set was based on this segmentator so it made no sense to spend a lot of time creating a custom segmentator which could make the IoU worse. The authors of the contest said that only 10% modifications were made on the outputs of the segmentator.</p></li>\n<li><p>Dataset<br>\nOur dataset was created from the Train &amp; PublicHPA 16-bit images. Seems most teams used 8-bit images in the end.</p></li>\n</ol>\n<p>Each image in the final dataset is a 512x512 image of a cell based on the cell masks from the segmentator. Padding (to square) was used to retain original height/width ratio. No surrounding pixels were used (non-cell-mask pixels). My feeling is that it might be better to use a bit larger surrounding, but did not have time to test this (this might be good for some classes such as plasma membrane).</p>\n<p>We decided to go with large images (512x512) since some labels/organelles required higher resolution and their size varied a lot based on the cell size &amp; re-scaling. E.g. sometimes the nucleus was very small and sometimes it as big as the whole image. We even tried to train nuclear organelles on images based on nuclei masks, but since the deadline was too close, we decided not to invest more time on this approach.</p>\n<p>I am eager to find out if diving the problem into nuclear and cytosolic organelles classification would yield better results. I think it would be easier to classify organelles inside the nucleus since it would be approximately the same size for each cell image this way.</p>\n<p>We used a simple heuristic to determine how much of the nuclei was outside of the image and decreased its final predicted confidence accordingly. All images with nuclei that were not present almost completely in the cropped image were removed from the train set.</p>\n<p>Similarly we tried to determine false positive segmentations by finding outliers based on the red channel and a product of the blue and yellow channel. Outliers at inference time got their confidence decreased dramatically. Outliers in the train dataset were removed completely. I assume the accuracy of this heuristic was around 50%. Since false positives were a big score crusher, this seemed acceptable.</p>\n<p>We lost around 20% of the images from the trainset.</p>\n<p>最終的なデータセットの各画像は、セグメンターからのセルマスクに基づくセルの512x512画像です。元の高さ/幅の比率を維持するために、（正方形への）パディングが使用されました。周囲のピクセルは使用されませんでした（セルマスク以外のピクセル）。私の感じでは、少し広い周囲を使用する方が良いかもしれませんが、これをテストする時間がありませんでした（これは原形質膜などの一部のクラスに適している可能性があります）。</p>\n<p>一部のラベル/細胞小器官はより高い解像度を必要とし、それらのサイズはセルサイズと再スケーリングに基づいて大きく変化するため、大きな画像（512x512）を使用することにしました。例えば。核が非常に小さい場合もあれば、画像全体と同じくらい大きい場合もあります。核マスクをベースにした画像で核オルガネラを訓練することも試みましたが、締め切りが近すぎたため、このアプローチにこれ以上時間を費やさないことにしました。</p>\n<p>問題を核および細胞質ゾルの細胞小器官の分類に掘り下げることがより良い結果をもたらすかどうかを知りたいと思っています。このように各細胞画像でほぼ同じサイズになるので、核内の細胞小器官を分類する方が簡単だと思います。</p>\n<p>単純なヒューリスティックを使用して、核のどれだけが画像の外側にあるかを判断し、それに応じて最終的な予測信頼度を下げました。トリミングされた画像にほとんど完全に存在しなかった核を持つすべての画像は、トレインセットから削除されました。</p>\n<p>同様に、赤のチャネルと青と黄色のチャネルの積に基づいて外れ値を見つけることにより、誤検出のセグメンテーションを決定しようとしました。推論時の外れ値は、信頼度が劇的に低下しました。列車データセットの外れ値は完全に削除されました。このヒューリスティックの精度は約50％だったと思います。誤検知は大きなスコアのクラッシャーだったので、これは許容できるようでした。</p>\n<p>トレインセットからの画像の約20％が失われました。</p>\n<ol>\n<li>Parallelization<br>\nHPA-Cell-Segmentation took quite some time so we decided to parallelize most things. Even with only two cores we got a boost in the submission time.</li>\n</ol>\n<p>The first boost was by running the label_cell function in parallell. The second boost was in running all the previously mentioned heuristics and image cropping in parallel as well.</p>\n<p>This left us with more than 3 hours for inference.</p>\n<ol>\n<li>Manual labeling<br>\nWe manually labeled smaller classes or classes with smaller % of occurrence in the initial images (e.g. mitotic spindle, aggresome, intermediate &amp; actin filaments …). We made a simple GUI and relabeled only one label at a time for an image.</li>\n</ol>\n<p>Mostly we would give a score from 1 to 5 on how confident we were that the given cell image contained the image-level label. These scores we transferred to soft labels. Each mapping was different (e.g. 1:0.0, 2:0.2, 3:0.7, 4:0.9, 5:1.0).</p>\n<p>In the end, we tried to create a validation set in the same way with high quality labeling. We managed to do get a few thousand examples for most classes.</p>\n<p>小さいクラスまたは初期画像での発生率が小さいクラス（有糸分裂紡錘体、アグリソーム、中間およびアクチンフィラメントなど）に手動でラベルを付けました。 シンプルなGUIを作成し、画像に対して一度に1つのラベルのみを再ラベル付けしました。</p>\n<p>ほとんどの場合、特定の細胞画像に画像レベルのラベルが含まれていることをどの程度確信しているかについて、1から5のスコアを付けます。 これらのスコアをソフトラベルに転送しました。 各マッピングは異なっていました（例：1：0.0、2：0.2、3：0.7、4：0.9、5：1.0）。</p>\n<p>最終的に、高品質のラベリングを使用して、同じ方法で検証セットを作成しようとしました。 ほとんどのクラスで数千の例を取得することができました。</p>\n<ol>\n<li>Pseudo-labeling<br>\nInspired by the Meta Pseudo Labels paper, we wanted to get rid of some false positives and help our models avoid overfitting. A cut off of 0.3 seemed to remove approx. 15% of image with high accuracy. Here we used an underfitt ResNet18.</li>\n</ol>\n<p>Later we used a better model to find more examples of mitotic spindles in a similar way, but withing the images that did not have mitotic spindle assigned. I think we found around 100 extra mitotic spindles, compared to around 250 that we found in the labeled images.</p>\n<p>In the end, we did not do this for other classes. I think we found a few aggresomes and quickly decided to skip positive labeling.</p>\n<p>Meta Pseudo Labelsの論文に触発されて、いくつかの誤検知を取り除き、モデルが過剰適合を回避できるようにしたいと考えました。 0.3のカットオフは約を削除するようでした。 画像の15％を高精度で。 ここでは、アンダーフィットのResNet18を使用しました。</p>\n<p>その後、より良いモデルを使用して、同様の方法で有糸分裂紡錘体のより多くの例を見つけましたが、有糸分裂紡錘体が割り当てられていない画像を使用しました。 ラベル付けされた画像で見つかった約250と比較して、約100の余分な有糸分裂紡錘体が見つかったと思います。</p>\n<p>結局、他のクラスではこれを行いませんでした。 私たちはいくつかのアグリソームを見つけ、すぐにポジティブラベリングをスキップすることに決めたと思います。</p>\n<ol>\n<li>EfficientNetB0<br>\nThis network is just awesome :) I am a big fan of solving problems with simple models, so I was quite happy when EfficientNetB0 seemed to be good enough for this challenge. We tried using B4, but it was slower to train and the results did not impress enough to continue playing with it. There was a solution that ensembled some B4-s, but no boost on the private LB.</li>\n</ol>\n<p>We had an 3-part ensemble with weights [0.2, 0.4, 0.4]. All EfficientNetB0s, but trained with different augmentation and loss function combinations.</p>\n<p>Single B0 - 0.2 ensemble weight<br>\nAugmentation: Flipping &amp; Rotation<br>\nLoss: FocalLoss<br>\nDid not have time to test if this network actually helped much.</p>\n<p>2 Checkpoint Ensemble B0s - 0.4 final ensemble weight<br>\nAugmentation: RandomResizer (40% chance), Flipping &amp; Rotation<br>\nLoss: FocalLoss<br>\n*RandomResizer -&gt; Resize to (RSIZE, RSIZE) + Resize back to (512, 512), where RSIZE is a random number between 256 and 384</p>\n<p>4 Checkpoint Ensemble B0s - 0.4 final ensemble weight<br>\nAugmentation: RandomResizer (30% chance), RandomPad (30% chance), Flipping, Rotation, &amp; Resize(512,512)<br>\nLoss: BCELoss<br>\n*RandomPad -&gt; pad each side (independently) with a random length between 0 and 200</p>\n<p>Since it was hard to estimate the \"best\" model on the local validation set, we used the idea of checkpoint ensembling to try to avoid overfitting &amp; boost our score.</p>\n<p>We oversampled classes with less examples or less true positives and tried to avoid overusing images with more assigned labels.</p>\n<ol>\n<li><p>Fine-tuning<br>\nWe fine-tuned all networks for one \"epoch\" on the validation set. Unlabeled images were used once, while images that we labeled (soft labels) were used multiple times in the \"epoch\". The \"epoch\" was around 200-250k images. We used the same augmentations &amp; loss function for each B0 as it was used during training of that network.</p></li>\n<li><p>Cell/Image weighting<br>\nTo extreme outliers we weighted the final confidences with the mean of confidences of all other valid cells (excluding border and outliers). The final confidence was 0.7 * cell_confidence + 0.3 * image_confidence.</p></li>\n</ol>\n<p>Seems that even extreme values such as 0.6/0.4 tend to work well here. We did not test if smaller weighting gave better score.</p>\n<p>Conclusion<br>\n512x512 images, re-labeling, simple network (B0), fine-tuning &amp; final confidence weighting seem to work well enough for this problem.</p>\n<p>Thank you<br>\n<a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> &amp; <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> -&gt; <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324</a><br>\nWe were not aware of this at that time. This valuable responses made a huge impact on our approach/results.</p>\n<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> -&gt; <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940</a><br>\nThank you for the motivation for the final weighting.</p>\n<p><a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a>, <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> &amp; <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> I hope there will be more opportunities to participate in your journey in the future (Kaggle or non-Kaggle related). Thank you for the ride! :)</p>",
      "rawMarkdown": "日本語訳\nWe are a very diverse team of computer scientists and medical doctor/students. It was our great pleasure to participate in this demanding challenge. Hope some of you find this solution useful and/or interesting.\n\nSolution overview\nSegmentation -> HPA-Cell-Segmentation\nDataset -> 512x512 cell images (20% removed)\nParallelization -> speed-up -> 3h left for inference\nManual Labeling -> smaller classes & validation (soft labels)\nPseudo-Labeling -> negative labeling (& positive for mitotic spindle)\nEfficientNetB0 Ensemble + semi-balanced data sampling\nFine-tuning -> on manually labeled & non-labeled validation data\nCell/Image Weighting -> final confidence = 0.7 * cell_confidence + 0.3 * image_confidence\n1. HPA-Cell-Segmentation\nThe test set was based on this segmentator so it made no sense to spend a lot of time creating a custom segmentator which could make the IoU worse. The authors of the contest said that only 10% modifications were made on the outputs of the segmentator.\n\n2. Dataset\nOur dataset was created from the Train & PublicHPA 16-bit images. Seems most teams used 8-bit images in the end.\n\nEach image in the final dataset is a 512x512 image of a cell based on the cell masks from the segmentator. Padding (to square) was used to retain original height/width ratio. No surrounding pixels were used (non-cell-mask pixels). My feeling is that it might be better to use a bit larger surrounding, but did not have time to test this (this might be good for some classes such as plasma membrane).\n\nWe decided to go with large images (512x512) since some labels/organelles required higher resolution and their size varied a lot based on the cell size & re-scaling. E.g. sometimes the nucleus was very small and sometimes it as big as the whole image. We even tried to train nuclear organelles on images based on nuclei masks, but since the deadline was too close, we decided not to invest more time on this approach.\n\nI am eager to find out if diving the problem into nuclear and cytosolic organelles classification would yield better results. I think it would be easier to classify organelles inside the nucleus since it would be approximately the same size for each cell image this way.\n\nWe used a simple heuristic to determine how much of the nuclei was outside of the image and decreased its final predicted confidence accordingly. All images with nuclei that were not present almost completely in the cropped image were removed from the train set.\n\nSimilarly we tried to determine false positive segmentations by finding outliers based on the red channel and a product of the blue and yellow channel. Outliers at inference time got their confidence decreased dramatically. Outliers in the train dataset were removed completely. I assume the accuracy of this heuristic was around 50%. Since false positives were a big score crusher, this seemed acceptable.\n\nWe lost around 20% of the images from the trainset.\n\n最終的なデータセットの各画像は、セグメンターからのセルマスクに基づくセルの512x512画像です。元の高さ/幅の比率を維持するために、（正方形への）パディングが使用されました。周囲のピクセルは使用されませんでした（セルマスク以外のピクセル）。私の感じでは、少し広い周囲を使用する方が良いかもしれませんが、これをテストする時間がありませんでした（これは原形質膜などの一部のクラスに適している可能性があります）。\n\n一部のラベル/細胞小器官はより高い解像度を必要とし、それらのサイズはセルサイズと再スケーリングに基づいて大きく変化するため、大きな画像（512x512）を使用することにしました。例えば。核が非常に小さい場合もあれば、画像全体と同じくらい大きい場合もあります。核マスクをベースにした画像で核オルガネラを訓練することも試みましたが、締め切りが近すぎたため、このアプローチにこれ以上時間を費やさないことにしました。\n\n問題を核および細胞質ゾルの細胞小器官の分類に掘り下げることがより良い結果をもたらすかどうかを知りたいと思っています。このように各細胞画像でほぼ同じサイズになるので、核内の細胞小器官を分類する方が簡単だと思います。\n\n単純なヒューリスティックを使用して、核のどれだけが画像の外側にあるかを判断し、それに応じて最終的な予測信頼度を下げました。トリミングされた画像にほとんど完全に存在しなかった核を持つすべての画像は、トレインセットから削除されました。\n\n同様に、赤のチャネルと青と黄色のチャネルの積に基づいて外れ値を見つけることにより、誤検出のセグメンテーションを決定しようとしました。推論時の外れ値は、信頼度が劇的に低下しました。列車データセットの外れ値は完全に削除されました。このヒューリスティックの精度は約50％だったと思います。誤検知は大きなスコアのクラッシャーだったので、これは許容できるようでした。\n\nトレインセットからの画像の約20％が失われました。\n3. Parallelization\nHPA-Cell-Segmentation took quite some time so we decided to parallelize most things. Even with only two cores we got a boost in the submission time.\n\nThe first boost was by running the label_cell function in parallell. The second boost was in running all the previously mentioned heuristics and image cropping in parallel as well.\n\nThis left us with more than 3 hours for inference.\n\n4. Manual labeling\nWe manually labeled smaller classes or classes with smaller % of occurrence in the initial images (e.g. mitotic spindle, aggresome, intermediate & actin filaments …). We made a simple GUI and relabeled only one label at a time for an image.\n\nMostly we would give a score from 1 to 5 on how confident we were that the given cell image contained the image-level label. These scores we transferred to soft labels. Each mapping was different (e.g. 1:0.0, 2:0.2, 3:0.7, 4:0.9, 5:1.0).\n\nIn the end, we tried to create a validation set in the same way with high quality labeling. We managed to do get a few thousand examples for most classes.\n\n小さいクラスまたは初期画像での発生率が小さいクラス（有糸分裂紡錘体、アグリソーム、中間およびアクチンフィラメントなど）に手動でラベルを付けました。 シンプルなGUIを作成し、画像に対して一度に1つのラベルのみを再ラベル付けしました。\n\nほとんどの場合、特定の細胞画像に画像レベルのラベルが含まれていることをどの程度確信しているかについて、1から5のスコアを付けます。 これらのスコアをソフトラベルに転送しました。 各マッピングは異なっていました（例：1：0.0、2：0.2、3：0.7、4：0.9、5：1.0）。\n\n最終的に、高品質のラベリングを使用して、同じ方法で検証セットを作成しようとしました。 ほとんどのクラスで数千の例を取得することができました。\n\n5. Pseudo-labeling\nInspired by the Meta Pseudo Labels paper, we wanted to get rid of some false positives and help our models avoid overfitting. A cut off of 0.3 seemed to remove approx. 15% of image with high accuracy. Here we used an underfitt ResNet18.\n\nLater we used a better model to find more examples of mitotic spindles in a similar way, but withing the images that did not have mitotic spindle assigned. I think we found around 100 extra mitotic spindles, compared to around 250 that we found in the labeled images.\n\nIn the end, we did not do this for other classes. I think we found a few aggresomes and quickly decided to skip positive labeling.\n\nMeta Pseudo Labelsの論文に触発されて、いくつかの誤検知を取り除き、モデルが過剰適合を回避できるようにしたいと考えました。 0.3のカットオフは約を削除するようでした。 画像の15％を高精度で。 ここでは、アンダーフィットのResNet18を使用しました。\n\nその後、より良いモデルを使用して、同様の方法で有糸分裂紡錘体のより多くの例を見つけましたが、有糸分裂紡錘体が割り当てられていない画像を使用しました。 ラベル付けされた画像で見つかった約250と比較して、約100の余分な有糸分裂紡錘体が見つかったと思います。\n\n結局、他のクラスではこれを行いませんでした。 私たちはいくつかのアグリソームを見つけ、すぐにポジティブラベリングをスキップすることに決めたと思います。\n\n6. EfficientNetB0\nThis network is just awesome :) I am a big fan of solving problems with simple models, so I was quite happy when EfficientNetB0 seemed to be good enough for this challenge. We tried using B4, but it was slower to train and the results did not impress enough to continue playing with it. There was a solution that ensembled some B4-s, but no boost on the private LB.\n\nWe had an 3-part ensemble with weights [0.2, 0.4, 0.4]. All EfficientNetB0s, but trained with different augmentation and loss function combinations.\n\nSingle B0 - 0.2 ensemble weight\nAugmentation: Flipping & Rotation\nLoss: FocalLoss\nDid not have time to test if this network actually helped much.\n\n2 Checkpoint Ensemble B0s - 0.4 final ensemble weight\nAugmentation: RandomResizer (40% chance), Flipping & Rotation\nLoss: FocalLoss\n*RandomResizer -> Resize to (RSIZE, RSIZE) + Resize back to (512, 512), where RSIZE is a random number between 256 and 384\n\n4 Checkpoint Ensemble B0s - 0.4 final ensemble weight\nAugmentation: RandomResizer (30% chance), RandomPad (30% chance), Flipping, Rotation, & Resize(512,512)\nLoss: BCELoss\n*RandomPad -> pad each side (independently) with a random length between 0 and 200\n\nSince it was hard to estimate the \"best\" model on the local validation set, we used the idea of checkpoint ensembling to try to avoid overfitting & boost our score.\n\nWe oversampled classes with less examples or less true positives and tried to avoid overusing images with more assigned labels.\n\n7. Fine-tuning\nWe fine-tuned all networks for one \"epoch\" on the validation set. Unlabeled images were used once, while images that we labeled (soft labels) were used multiple times in the \"epoch\". The \"epoch\" was around 200-250k images. We used the same augmentations & loss function for each B0 as it was used during training of that network.\n\n8. Cell/Image weighting\nTo extreme outliers we weighted the final confidences with the mean of confidences of all other valid cells (excluding border and outliers). The final confidence was 0.7 * cell_confidence + 0.3 * image_confidence.\n\nSeems that even extreme values such as 0.6/0.4 tend to work well here. We did not test if smaller weighting gave better score.\n\nConclusion\n512x512 images, re-labeling, simple network (B0), fine-tuning & final confidence weighting seem to work well enough for this problem.\n\nThank you\n@christofhenkel & @cwinsnes -> https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324\nWe were not aware of this at that time. This valuable responses made a huge impact on our approach/results.\n\n@h053473666 -> https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940\nThank you for the motivation for the final weighting.\n\n@emmalumpan, @lnhtrang & @cwinsnes I hope there will be more opportunities to participate in your journey in the future (Kaggle or non-Kaggle related). Thank you for the ride! :)",
      "votes": 1
    },
    {
      "id": 1307655,
      "postDate": "2021-05-14T15:28:54.427Z",
      "content": "<p>Thanks for sharing - some very interesting heuristics!</p>\n<p>I also tried getting label_cells running in parallel using multiprocessing since I wanted to try and do this at full scale on the private set and still have time for inference. This worked fine on my local machine but when I moved to Kaggle's environment I got multiple \"Kaggle Errors\" on submission or other problems related to multiprocessing. Would you be willing to share any code snippets for how you implemented this?</p>",
      "rawMarkdown": "Thanks for sharing - some very interesting heuristics!\n\nI also tried getting label_cells running in parallel using multiprocessing since I wanted to try and do this at full scale on the private set and still have time for inference. This worked fine on my local machine but when I moved to Kaggle's environment I got multiple \"Kaggle Errors\" on submission or other problems related to multiprocessing. Would you be willing to share any code snippets for how you implemented this?",
      "votes": 1,
      "replies": [
        {
          "id": 1307909,
          "postDate": "2021-05-14T19:01:46.237Z",
          "content": "<p>Absolutely :)</p>\n<p>This is the main idea:</p>\n<pre><code>import multiprocessing\nimport multiprocessing as mp\n\npool = mp.Pool(mp.cpu_count())\n\nfor x in mt:\n    pool.apply_async(solve, args=(x,))\n\npool.close()    \npool.join()\n</code></pre>\n<p>And this is the exact code for creating masks:</p>\n<pre><code>def create_masks(cell_segmentation, nuc_segmentation, name):\n    nuclei_mask, cell_mask = label_cell(nuc_segmentation, cell_segmentation)\n\n    ID = name.replace('_red.png','').split('/')[-1]\n    np.save(test_nuclei + ID + '.npy', nuclei_mask)\n    np.save(test_cells + ID + '.npy', cell_mask)\n\n    return\n\nbatch = 32\npos = 0\n\nmt = glob.glob(test_dir + '*_red.png')\ner = [f.replace('red', 'yellow') for f in mt]\nnu = [f.replace('red', 'blue') for f in mt]\n\nwhile pos &lt; len(mt):\n\n    images = [mt[pos:pos+batch], er[pos:pos+batch], nu[pos:pos+batch]]\n    pos += batch\n\n    nuc_segmentations = segmentator.pred_nuclei(images[2])\n    cell_segmentations = segmentator.pred_cells(images)\n\n    pool = mp.Pool(mp.cpu_count())\n\n    for i in range(len(cell_segmentations)):\n        name = images[0][i]\n        pool.apply_async(create_masks, args=(cell_segmentations[i], nuc_segmentations[i], name))\n\n    pool.close()    \n    pool.join()\n</code></pre>",
          "rawMarkdown": "Absolutely :)\n\nThis is the main idea:\n\n\n```\nimport multiprocessing\nimport multiprocessing as mp\n\npool = mp.Pool(mp.cpu_count())\n\nfor x in mt:\n    pool.apply_async(solve, args=(x,))\n    \npool.close()    \npool.join()\n\n```\n\nAnd this is the exact code for creating masks:\n\n```\ndef create_masks(cell_segmentation, nuc_segmentation, name):\n    nuclei_mask, cell_mask = label_cell(nuc_segmentation, cell_segmentation)\n    \n    ID = name.replace('_red.png','').split('/')[-1]\n    np.save(test_nuclei + ID + '.npy', nuclei_mask)\n    np.save(test_cells + ID + '.npy', cell_mask)\n            \n    return\n\nbatch = 32\npos = 0\n\nmt = glob.glob(test_dir + '*_red.png')\ner = [f.replace('red', 'yellow') for f in mt]\nnu = [f.replace('red', 'blue') for f in mt]\n\nwhile pos < len(mt):\n\n    images = [mt[pos:pos+batch], er[pos:pos+batch], nu[pos:pos+batch]]\n    pos += batch\n\n    nuc_segmentations = segmentator.pred_nuclei(images[2])\n    cell_segmentations = segmentator.pred_cells(images)\n\n    pool = mp.Pool(mp.cpu_count())\n\n    for i in range(len(cell_segmentations)):\n        name = images[0][i]\n        pool.apply_async(create_masks, args=(cell_segmentations[i], nuc_segmentations[i], name))\n        \n    pool.close()    \n    pool.join()\n```"
        },
        {
          "id": 1307911,
          "postDate": "2021-05-14T19:04:27.947Z",
          "content": "<p>And if you want to use the results of the functions:</p>\n<pre><code>import multiprocessing\nimport multiprocessing as mp\n\npool = mp.Pool(mp.cpu_count())\nresult = []\n\nfor x in mt:\n    pool.apply_async(solve_cells, args=(x,), callback=result.append)\n\npool.close()    \npool.join()\n</code></pre>",
          "rawMarkdown": "And if you want to use the results of the functions:\n\n```\nimport multiprocessing\nimport multiprocessing as mp\n\npool = mp.Pool(mp.cpu_count())\nresult = []\n\nfor x in mt:\n    pool.apply_async(solve_cells, args=(x,), callback=result.append)\n\npool.close()    \npool.join()\n```"
        },
        {
          "id": 1308088,
          "postDate": "2021-05-15T01:26:36.747Z",
          "content": "<p>Thank you for sharing!</p>",
          "rawMarkdown": "Thank you for sharing!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1610751,
      "author_name": "pixyz0130",
      "author_url": "",
      "post_date": "2021-12-07T13:47:39.983000",
      "content": "<p>日本語訳<br>\nWe are a very diverse team of computer scientists and medical doctor/students. It was our great pleasure to participate in this demanding challenge. Hope some of you find this solution useful and/or interesting.</p>\n<p>Solution overview<br>\nSegmentation -&gt; HPA-Cell-Segmentation<br>\nDataset -&gt; 512x512 cell images (20% removed)<br>\nParallelization -&gt; speed-up -&gt; 3h left for inference<br>\nManual Labeling -&gt; smaller classes &amp; validation (soft labels)<br>\nPseudo-Labeling -&gt; negative labeling (&amp; positive for mitotic spindle)<br>\nEfficientNetB0 Ensemble + semi-balanced data sampling<br>\nFine-tuning -&gt; on manually labeled &amp; non-labeled validation data<br>\nCell/Image Weighting -&gt; final confidence = 0.7 * cell_confidence + 0.3 * image_confidence</p>\n<ol>\n<li><p>HPA-Cell-Segmentation<br>\nThe test set was based on this segmentator so it made no sense to spend a lot of time creating a custom segmentator which could make the IoU worse. The authors of the contest said that only 10% modifications were made on the outputs of the segmentator.</p></li>\n<li><p>Dataset<br>\nOur dataset was created from the Train &amp; PublicHPA 16-bit images. Seems most teams used 8-bit images in the end.</p></li>\n</ol>\n<p>Each image in the final dataset is a 512x512 image of a cell based on the cell masks from the segmentator. Padding (to square) was used to retain original height/width ratio. No surrounding pixels were used (non-cell-mask pixels). My feeling is that it might be better to use a bit larger surrounding, but did not have time to test this (this might be good for some classes such as plasma membrane).</p>\n<p>We decided to go with large images (512x512) since some labels/organelles required higher resolution and their size varied a lot based on the cell size &amp; re-scaling. E.g. sometimes the nucleus was very small and sometimes it as big as the whole image. We even tried to train nuclear organelles on images based on nuclei masks, but since the deadline was too close, we decided not to invest more time on this approach.</p>\n<p>I am eager to find out if diving the problem into nuclear and cytosolic organelles classification would yield better results. I think it would be easier to classify organelles inside the nucleus since it would be approximately the same size for each cell image this way.</p>\n<p>We used a simple heuristic to determine how much of the nuclei was outside of the image and decreased its final predicted confidence accordingly. All images with nuclei that were not present almost completely in the cropped image were removed from the train set.</p>\n<p>Similarly we tried to determine false positive segmentations by finding outliers based on the red channel and a product of the blue and yellow channel. Outliers at inference time got their confidence decreased dramatically. Outliers in the train dataset were removed completely. I assume the accuracy of this heuristic was around 50%. Since false positives were a big score crusher, this seemed acceptable.</p>\n<p>We lost around 20% of the images from the trainset.</p>\n<p>最終的なデータセットの各画像は、セグメンターからのセルマスクに基づくセルの512x512画像です。元の高さ/幅の比率を維持するために、（正方形への）パディングが使用されました。周囲のピクセルは使用されませんでした（セルマスク以外のピクセル）。私の感じでは、少し広い周囲を使用する方が良いかもしれませんが、これをテストする時間がありませんでした（これは原形質膜などの一部のクラスに適している可能性があります）。</p>\n<p>一部のラベル/細胞小器官はより高い解像度を必要とし、それらのサイズはセルサイズと再スケーリングに基づいて大きく変化するため、大きな画像（512x512）を使用することにしました。例えば。核が非常に小さい場合もあれば、画像全体と同じくらい大きい場合もあります。核マスクをベースにした画像で核オルガネラを訓練することも試みましたが、締め切りが近すぎたため、このアプローチにこれ以上時間を費やさないことにしました。</p>\n<p>問題を核および細胞質ゾルの細胞小器官の分類に掘り下げることがより良い結果をもたらすかどうかを知りたいと思っています。このように各細胞画像でほぼ同じサイズになるので、核内の細胞小器官を分類する方が簡単だと思います。</p>\n<p>単純なヒューリスティックを使用して、核のどれだけが画像の外側にあるかを判断し、それに応じて最終的な予測信頼度を下げました。トリミングされた画像にほとんど完全に存在しなかった核を持つすべての画像は、トレインセットから削除されました。</p>\n<p>同様に、赤のチャネルと青と黄色のチャネルの積に基づいて外れ値を見つけることにより、誤検出のセグメンテーションを決定しようとしました。推論時の外れ値は、信頼度が劇的に低下しました。列車データセットの外れ値は完全に削除されました。このヒューリスティックの精度は約50％だったと思います。誤検知は大きなスコアのクラッシャーだったので、これは許容できるようでした。</p>\n<p>トレインセットからの画像の約20％が失われました。</p>\n<ol>\n<li>Parallelization<br>\nHPA-Cell-Segmentation took quite some time so we decided to parallelize most things. Even with only two cores we got a boost in the submission time.</li>\n</ol>\n<p>The first boost was by running the label_cell function in parallell. The second boost was in running all the previously mentioned heuristics and image cropping in parallel as well.</p>\n<p>This left us with more than 3 hours for inference.</p>\n<ol>\n<li>Manual labeling<br>\nWe manually labeled smaller classes or classes with smaller % of occurrence in the initial images (e.g. mitotic spindle, aggresome, intermediate &amp; actin filaments …). We made a simple GUI and relabeled only one label at a time for an image.</li>\n</ol>\n<p>Mostly we would give a score from 1 to 5 on how confident we were that the given cell image contained the image-level label. These scores we transferred to soft labels. Each mapping was different (e.g. 1:0.0, 2:0.2, 3:0.7, 4:0.9, 5:1.0).</p>\n<p>In the end, we tried to create a validation set in the same way with high quality labeling. We managed to do get a few thousand examples for most classes.</p>\n<p>小さいクラスまたは初期画像での発生率が小さいクラス（有糸分裂紡錘体、アグリソーム、中間およびアクチンフィラメントなど）に手動でラベルを付けました。 シンプルなGUIを作成し、画像に対して一度に1つのラベルのみを再ラベル付けしました。</p>\n<p>ほとんどの場合、特定の細胞画像に画像レベルのラベルが含まれていることをどの程度確信しているかについて、1から5のスコアを付けます。 これらのスコアをソフトラベルに転送しました。 各マッピングは異なっていました（例：1：0.0、2：0.2、3：0.7、4：0.9、5：1.0）。</p>\n<p>最終的に、高品質のラベリングを使用して、同じ方法で検証セットを作成しようとしました。 ほとんどのクラスで数千の例を取得することができました。</p>\n<ol>\n<li>Pseudo-labeling<br>\nInspired by the Meta Pseudo Labels paper, we wanted to get rid of some false positives and help our models avoid overfitting. A cut off of 0.3 seemed to remove approx. 15% of image with high accuracy. Here we used an underfitt ResNet18.</li>\n</ol>\n<p>Later we used a better model to find more examples of mitotic spindles in a similar way, but withing the images that did not have mitotic spindle assigned. I think we found around 100 extra mitotic spindles, compared to around 250 that we found in the labeled images.</p>\n<p>In the end, we did not do this for other classes. I think we found a few aggresomes and quickly decided to skip positive labeling.</p>\n<p>Meta Pseudo Labelsの論文に触発されて、いくつかの誤検知を取り除き、モデルが過剰適合を回避できるようにしたいと考えました。 0.3のカットオフは約を削除するようでした。 画像の15％を高精度で。 ここでは、アンダーフィットのResNet18を使用しました。</p>\n<p>その後、より良いモデルを使用して、同様の方法で有糸分裂紡錘体のより多くの例を見つけましたが、有糸分裂紡錘体が割り当てられていない画像を使用しました。 ラベル付けされた画像で見つかった約250と比較して、約100の余分な有糸分裂紡錘体が見つかったと思います。</p>\n<p>結局、他のクラスではこれを行いませんでした。 私たちはいくつかのアグリソームを見つけ、すぐにポジティブラベリングをスキップすることに決めたと思います。</p>\n<ol>\n<li>EfficientNetB0<br>\nThis network is just awesome :) I am a big fan of solving problems with simple models, so I was quite happy when EfficientNetB0 seemed to be good enough for this challenge. We tried using B4, but it was slower to train and the results did not impress enough to continue playing with it. There was a solution that ensembled some B4-s, but no boost on the private LB.</li>\n</ol>\n<p>We had an 3-part ensemble with weights [0.2, 0.4, 0.4]. All EfficientNetB0s, but trained with different augmentation and loss function combinations.</p>\n<p>Single B0 - 0.2 ensemble weight<br>\nAugmentation: Flipping &amp; Rotation<br>\nLoss: FocalLoss<br>\nDid not have time to test if this network actually helped much.</p>\n<p>2 Checkpoint Ensemble B0s - 0.4 final ensemble weight<br>\nAugmentation: RandomResizer (40% chance), Flipping &amp; Rotation<br>\nLoss: FocalLoss<br>\n*RandomResizer -&gt; Resize to (RSIZE, RSIZE) + Resize back to (512, 512), where RSIZE is a random number between 256 and 384</p>\n<p>4 Checkpoint Ensemble B0s - 0.4 final ensemble weight<br>\nAugmentation: RandomResizer (30% chance), RandomPad (30% chance), Flipping, Rotation, &amp; Resize(512,512)<br>\nLoss: BCELoss<br>\n*RandomPad -&gt; pad each side (independently) with a random length between 0 and 200</p>\n<p>Since it was hard to estimate the \"best\" model on the local validation set, we used the idea of checkpoint ensembling to try to avoid overfitting &amp; boost our score.</p>\n<p>We oversampled classes with less examples or less true positives and tried to avoid overusing images with more assigned labels.</p>\n<ol>\n<li><p>Fine-tuning<br>\nWe fine-tuned all networks for one \"epoch\" on the validation set. Unlabeled images were used once, while images that we labeled (soft labels) were used multiple times in the \"epoch\". The \"epoch\" was around 200-250k images. We used the same augmentations &amp; loss function for each B0 as it was used during training of that network.</p></li>\n<li><p>Cell/Image weighting<br>\nTo extreme outliers we weighted the final confidences with the mean of confidences of all other valid cells (excluding border and outliers). The final confidence was 0.7 * cell_confidence + 0.3 * image_confidence.</p></li>\n</ol>\n<p>Seems that even extreme values such as 0.6/0.4 tend to work well here. We did not test if smaller weighting gave better score.</p>\n<p>Conclusion<br>\n512x512 images, re-labeling, simple network (B0), fine-tuning &amp; final confidence weighting seem to work well enough for this problem.</p>\n<p>Thank you<br>\n<a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> &amp; <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> -&gt; <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324</a><br>\nWe were not aware of this at that time. This valuable responses made a huge impact on our approach/results.</p>\n<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> -&gt; <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940</a><br>\nThank you for the motivation for the final weighting.</p>\n<p><a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a>, <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> &amp; <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> I hope there will be more opportunities to participate in your journey in the future (Kaggle or non-Kaggle related). Thank you for the ride! :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1307655,
      "author_name": "Andrew Tratz",
      "author_url": "",
      "post_date": "2021-05-14T15:28:54.427000",
      "content": "<p>Thanks for sharing - some very interesting heuristics!</p>\n<p>I also tried getting label_cells running in parallel using multiprocessing since I wanted to try and do this at full scale on the private set and still have time for inference. This worked fine on my local machine but when I moved to Kaggle's environment I got multiple \"Kaggle Errors\" on submission or other problems related to multiprocessing. Would you be willing to share any code snippets for how you implemented this?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1307909,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-05-14T19:01:46.237000",
          "content": "<p>Absolutely :)</p>\n<p>This is the main idea:</p>\n<pre><code>import multiprocessing\nimport multiprocessing as mp\n\npool = mp.Pool(mp.cpu_count())\n\nfor x in mt:\n    pool.apply_async(solve, args=(x,))\n\npool.close()    \npool.join()\n</code></pre>\n<p>And this is the exact code for creating masks:</p>\n<pre><code>def create_masks(cell_segmentation, nuc_segmentation, name):\n    nuclei_mask, cell_mask = label_cell(nuc_segmentation, cell_segmentation)\n\n    ID = name.replace('_red.png','').split('/')[-1]\n    np.save(test_nuclei + ID + '.npy', nuclei_mask)\n    np.save(test_cells + ID + '.npy', cell_mask)\n\n    return\n\nbatch = 32\npos = 0\n\nmt = glob.glob(test_dir + '*_red.png')\ner = [f.replace('red', 'yellow') for f in mt]\nnu = [f.replace('red', 'blue') for f in mt]\n\nwhile pos &lt; len(mt):\n\n    images = [mt[pos:pos+batch], er[pos:pos+batch], nu[pos:pos+batch]]\n    pos += batch\n\n    nuc_segmentations = segmentator.pred_nuclei(images[2])\n    cell_segmentations = segmentator.pred_cells(images)\n\n    pool = mp.Pool(mp.cpu_count())\n\n    for i in range(len(cell_segmentations)):\n        name = images[0][i]\n        pool.apply_async(create_masks, args=(cell_segmentations[i], nuc_segmentations[i], name))\n\n    pool.close()    \n    pool.join()\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1307911,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-05-14T19:04:27.947000",
          "content": "<p>And if you want to use the results of the functions:</p>\n<pre><code>import multiprocessing\nimport multiprocessing as mp\n\npool = mp.Pool(mp.cpu_count())\nresult = []\n\nfor x in mt:\n    pool.apply_async(solve_cells, args=(x,), callback=result.append)\n\npool.close()    \npool.join()\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1308088,
          "author_name": "Andrew Tratz",
          "author_url": "",
          "post_date": "2021-05-15T01:26:36.747000",
          "content": "<p>Thank you for sharing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1307555": "We are a very diverse team of computer scientists and medical doctor/students. It was our great pleasure to participate in this demanding challenge. Hope some of you find this solution useful and/or interesting.\n\n# Solution overview\n1. Segmentation -> HPA-Cell-Segmentation\n2. Dataset -> 512x512 cell images (20% removed)\n3. Parallelization -> speed-up -> 3h left for inference\n4. Manual Labeling -> smaller classes & validation (soft labels)\n5. Pseudo-Labeling -> negative labeling (& positive for mitotic spindle)\n6. EfficientNetB0 Ensemble + semi-balanced data sampling\n7. Fine-tuning -> on manually labeled & non-labeled validation data\n8. Cell/Image Weighting -> final confidence = 0.7 \\* cell\\_confidence + 0.3 \\* image\\_confidence\n\n# 1. HPA-Cell-Segmentation\n\nThe test set was based on this segmentator so it made no sense to spend a lot of time creating a custom segmentator which could make the IoU worse. The authors of the contest said that only 10% modifications were made on the outputs of the segmentator.\n\n# 2. Dataset\n\nOur dataset was created from the Train & PublicHPA 16-bit images. Seems most teams used 8-bit images in the end.\n\nEach image in the final dataset is a 512x512 image of a cell based on the cell masks from the segmentator. Padding (to square) was used to retain original height/width ratio. No surrounding pixels were used (non-cell-mask pixels). My feeling is that it might be better to use a bit larger surrounding, but did not have time to test this (this might be good for some classes such as plasma membrane).\n\nWe decided to go with large images (512x512) since some labels/organelles required higher resolution and their size varied a lot based on the cell size & re-scaling. E.g. sometimes the nucleus was very small and sometimes it as big as the whole image. We even tried to train nuclear organelles on images based on nuclei masks, but since the deadline was too close, we decided not to invest more time on this approach.\n\nI am eager to find out if diving the problem into nuclear and cytosolic organelles classification would yield better results. I think it would be easier to classify organelles inside the nucleus since it would be approximately the same size for each cell image this way.\n\nWe used a simple heuristic to determine how much of the nuclei was outside of the image and decreased its final predicted confidence accordingly. All images with nuclei that were not present almost completely in the cropped image were removed from the train set.\n\nSimilarly we tried to determine false positive segmentations by finding outliers based on the red channel and a product of the blue and yellow channel. Outliers at inference time got their confidence decreased dramatically. Outliers in the train dataset were removed completely. I assume the accuracy of this heuristic was around 50%. Since false positives were a big score crusher, this seemed acceptable.\n\nWe lost around 20% of the images from the trainset.\n\n# 3. Parallelization\n\nHPA-Cell-Segmentation took quite some time so we decided to parallelize most things. Even with only two cores we got a boost in the submission time.\n\nThe first boost was by running the `label_cell` function in parallell. The second boost was in running all the previously mentioned heuristics and image cropping in parallel as well.\n\nThis left us with more than 3 hours for inference.\n\n# 4. Manual labeling\n\nWe manually labeled smaller classes or classes with smaller % of occurrence in the initial images (e.g. mitotic spindle, aggresome, intermediate & actin filaments ...). We made a simple GUI and relabeled only one label at a time for an image.\n\nMostly we would give a score from 1 to 5 on how confident we were that the given cell image contained the image-level label. These scores we transferred to soft labels. Each mapping was different (e.g. 1:0.0, 2:0.2, 3:0.7, 4:0.9, 5:1.0).\n\nIn the end, we tried to create a validation set in the same way with high quality labeling. We managed to do get a few thousand examples for most classes.\n\n# 5. Pseudo-labeling\nInspired by the Meta Pseudo Labels paper, we wanted to get rid of some false positives and help our models avoid overfitting. A cut off of 0.3 seemed to remove approx. 15% of image with high accuracy. Here we used an underfitt ResNet18.\n\nLater we used a better model to find more examples of mitotic spindles in a similar way, but withing the images that did not have mitotic spindle assigned. I think we found around 100 extra mitotic spindles, compared to around 250 that we found in the labeled images.\n\nIn the end, we did not do this for other classes. I think we found a few aggresomes and quickly decided to skip positive labeling.\n\n# 6. EfficientNetB0\nThis network is just awesome :) I am a big fan of solving problems with simple models, so I was quite happy when EfficientNetB0 seemed to be good enough for this challenge. We tried using B4, but it was slower to train and the results did not impress enough to continue playing with it. There was a solution that ensembled some B4-s, but no boost on the private LB.\n\nWe had an 3-part ensemble with weights [0.2, 0.4, 0.4]. All EfficientNetB0s, but trained with different augmentation and loss function combinations.\n\n1. Single B0 - 0.2 ensemble weight\nAugmentation: Flipping & Rotation\nLoss: FocalLoss\nDid not have time to test if this network actually helped much.\n\n2. 2 Checkpoint Ensemble B0s - 0.4 final ensemble weight\nAugmentation: RandomResizer (40% chance), Flipping & Rotation\nLoss: FocalLoss\n*RandomResizer -> Resize to (RSIZE, RSIZE) + Resize back to (512, 512), where RSIZE is a random number between 256 and 384\n\n3. 4 Checkpoint Ensemble B0s - 0.4 final ensemble weight\nAugmentation: RandomResizer (30% chance), RandomPad (30% chance), Flipping, Rotation, & Resize(512,512)\nLoss: BCELoss\n*RandomPad -> pad each side (independently) with a random length between 0 and 200\n\nSince it was hard to estimate the \"best\" model on the local validation set, we used the idea of checkpoint ensembling to try to avoid overfitting & boost our score.\n\nWe oversampled classes with less examples or less true positives and tried to avoid overusing images with more assigned labels. \n\n# 7. Fine-tuning\n\nWe fine-tuned all networks for one \"epoch\" on the validation set. Unlabeled images were used once, while images that we labeled (soft labels) were used multiple times in the \"epoch\". The \"epoch\" was around 200-250k images. We used the same augmentations & loss function for each B0 as it was used during training of that network.\n\n# 8. Cell/Image weighting\n\nTo extreme outliers we weighted the final confidences with the mean of confidences of all other valid cells (excluding border and outliers). The final confidence was 0.7 \\* cell\\_confidence + 0.3 \\* image\\_confidence.\n\nSeems that even extreme values such as 0.6/0.4 tend to work well here. We did not test if smaller weighting gave better score.\n\n# Conclusion\n\n512x512 images, re-labeling, simple network (B0), fine-tuning & final confidence weighting seem to work well enough for this problem.\n\n# Thank you\n@christofhenkel & @cwinsnes -> [https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324](url)\nWe were not aware of this at that time. This valuable responses made a huge impact on our approach/results.\n\n@h053473666 -> [https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940](url)\nThank you for the motivation for the final weighting.\n\n@emmalumpan, @lnhtrang & @cwinsnes I hope there will be more opportunities to participate in your journey in the future (Kaggle or non-Kaggle related). Thank you for the ride! :)",
    "1610751": "日本語訳\nWe are a very diverse team of computer scientists and medical doctor/students. It was our great pleasure to participate in this demanding challenge. Hope some of you find this solution useful and/or interesting.\n\nSolution overview\nSegmentation -> HPA-Cell-Segmentation\nDataset -> 512x512 cell images (20% removed)\nParallelization -> speed-up -> 3h left for inference\nManual Labeling -> smaller classes & validation (soft labels)\nPseudo-Labeling -> negative labeling (& positive for mitotic spindle)\nEfficientNetB0 Ensemble + semi-balanced data sampling\nFine-tuning -> on manually labeled & non-labeled validation data\nCell/Image Weighting -> final confidence = 0.7 * cell_confidence + 0.3 * image_confidence\n1. HPA-Cell-Segmentation\nThe test set was based on this segmentator so it made no sense to spend a lot of time creating a custom segmentator which could make the IoU worse. The authors of the contest said that only 10% modifications were made on the outputs of the segmentator.\n\n2. Dataset\nOur dataset was created from the Train & PublicHPA 16-bit images. Seems most teams used 8-bit images in the end.\n\nEach image in the final dataset is a 512x512 image of a cell based on the cell masks from the segmentator. Padding (to square) was used to retain original height/width ratio. No surrounding pixels were used (non-cell-mask pixels). My feeling is that it might be better to use a bit larger surrounding, but did not have time to test this (this might be good for some classes such as plasma membrane).\n\nWe decided to go with large images (512x512) since some labels/organelles required higher resolution and their size varied a lot based on the cell size & re-scaling. E.g. sometimes the nucleus was very small and sometimes it as big as the whole image. We even tried to train nuclear organelles on images based on nuclei masks, but since the deadline was too close, we decided not to invest more time on this approach.\n\nI am eager to find out if diving the problem into nuclear and cytosolic organelles classification would yield better results. I think it would be easier to classify organelles inside the nucleus since it would be approximately the same size for each cell image this way.\n\nWe used a simple heuristic to determine how much of the nuclei was outside of the image and decreased its final predicted confidence accordingly. All images with nuclei that were not present almost completely in the cropped image were removed from the train set.\n\nSimilarly we tried to determine false positive segmentations by finding outliers based on the red channel and a product of the blue and yellow channel. Outliers at inference time got their confidence decreased dramatically. Outliers in the train dataset were removed completely. I assume the accuracy of this heuristic was around 50%. Since false positives were a big score crusher, this seemed acceptable.\n\nWe lost around 20% of the images from the trainset.\n\n最終的なデータセットの各画像は、セグメンターからのセルマスクに基づくセルの512x512画像です。元の高さ/幅の比率を維持するために、（正方形への）パディングが使用されました。周囲のピクセルは使用されませんでした（セルマスク以外のピクセル）。私の感じでは、少し広い周囲を使用する方が良いかもしれませんが、これをテストする時間がありませんでした（これは原形質膜などの一部のクラスに適している可能性があります）。\n\n一部のラベル/細胞小器官はより高い解像度を必要とし、それらのサイズはセルサイズと再スケーリングに基づいて大きく変化するため、大きな画像（512x512）を使用することにしました。例えば。核が非常に小さい場合もあれば、画像全体と同じくらい大きい場合もあります。核マスクをベースにした画像で核オルガネラを訓練することも試みましたが、締め切りが近すぎたため、このアプローチにこれ以上時間を費やさないことにしました。\n\n問題を核および細胞質ゾルの細胞小器官の分類に掘り下げることがより良い結果をもたらすかどうかを知りたいと思っています。このように各細胞画像でほぼ同じサイズになるので、核内の細胞小器官を分類する方が簡単だと思います。\n\n単純なヒューリスティックを使用して、核のどれだけが画像の外側にあるかを判断し、それに応じて最終的な予測信頼度を下げました。トリミングされた画像にほとんど完全に存在しなかった核を持つすべての画像は、トレインセットから削除されました。\n\n同様に、赤のチャネルと青と黄色のチャネルの積に基づいて外れ値を見つけることにより、誤検出のセグメンテーションを決定しようとしました。推論時の外れ値は、信頼度が劇的に低下しました。列車データセットの外れ値は完全に削除されました。このヒューリスティックの精度は約50％だったと思います。誤検知は大きなスコアのクラッシャーだったので、これは許容できるようでした。\n\nトレインセットからの画像の約20％が失われました。\n3. Parallelization\nHPA-Cell-Segmentation took quite some time so we decided to parallelize most things. Even with only two cores we got a boost in the submission time.\n\nThe first boost was by running the label_cell function in parallell. The second boost was in running all the previously mentioned heuristics and image cropping in parallel as well.\n\nThis left us with more than 3 hours for inference.\n\n4. Manual labeling\nWe manually labeled smaller classes or classes with smaller % of occurrence in the initial images (e.g. mitotic spindle, aggresome, intermediate & actin filaments …). We made a simple GUI and relabeled only one label at a time for an image.\n\nMostly we would give a score from 1 to 5 on how confident we were that the given cell image contained the image-level label. These scores we transferred to soft labels. Each mapping was different (e.g. 1:0.0, 2:0.2, 3:0.7, 4:0.9, 5:1.0).\n\nIn the end, we tried to create a validation set in the same way with high quality labeling. We managed to do get a few thousand examples for most classes.\n\n小さいクラスまたは初期画像での発生率が小さいクラス（有糸分裂紡錘体、アグリソーム、中間およびアクチンフィラメントなど）に手動でラベルを付けました。 シンプルなGUIを作成し、画像に対して一度に1つのラベルのみを再ラベル付けしました。\n\nほとんどの場合、特定の細胞画像に画像レベルのラベルが含まれていることをどの程度確信しているかについて、1から5のスコアを付けます。 これらのスコアをソフトラベルに転送しました。 各マッピングは異なっていました（例：1：0.0、2：0.2、3：0.7、4：0.9、5：1.0）。\n\n最終的に、高品質のラベリングを使用して、同じ方法で検証セットを作成しようとしました。 ほとんどのクラスで数千の例を取得することができました。\n\n5. Pseudo-labeling\nInspired by the Meta Pseudo Labels paper, we wanted to get rid of some false positives and help our models avoid overfitting. A cut off of 0.3 seemed to remove approx. 15% of image with high accuracy. Here we used an underfitt ResNet18.\n\nLater we used a better model to find more examples of mitotic spindles in a similar way, but withing the images that did not have mitotic spindle assigned. I think we found around 100 extra mitotic spindles, compared to around 250 that we found in the labeled images.\n\nIn the end, we did not do this for other classes. I think we found a few aggresomes and quickly decided to skip positive labeling.\n\nMeta Pseudo Labelsの論文に触発されて、いくつかの誤検知を取り除き、モデルが過剰適合を回避できるようにしたいと考えました。 0.3のカットオフは約を削除するようでした。 画像の15％を高精度で。 ここでは、アンダーフィットのResNet18を使用しました。\n\nその後、より良いモデルを使用して、同様の方法で有糸分裂紡錘体のより多くの例を見つけましたが、有糸分裂紡錘体が割り当てられていない画像を使用しました。 ラベル付けされた画像で見つかった約250と比較して、約100の余分な有糸分裂紡錘体が見つかったと思います。\n\n結局、他のクラスではこれを行いませんでした。 私たちはいくつかのアグリソームを見つけ、すぐにポジティブラベリングをスキップすることに決めたと思います。\n\n6. EfficientNetB0\nThis network is just awesome :) I am a big fan of solving problems with simple models, so I was quite happy when EfficientNetB0 seemed to be good enough for this challenge. We tried using B4, but it was slower to train and the results did not impress enough to continue playing with it. There was a solution that ensembled some B4-s, but no boost on the private LB.\n\nWe had an 3-part ensemble with weights [0.2, 0.4, 0.4]. All EfficientNetB0s, but trained with different augmentation and loss function combinations.\n\nSingle B0 - 0.2 ensemble weight\nAugmentation: Flipping & Rotation\nLoss: FocalLoss\nDid not have time to test if this network actually helped much.\n\n2 Checkpoint Ensemble B0s - 0.4 final ensemble weight\nAugmentation: RandomResizer (40% chance), Flipping & Rotation\nLoss: FocalLoss\n*RandomResizer -> Resize to (RSIZE, RSIZE) + Resize back to (512, 512), where RSIZE is a random number between 256 and 384\n\n4 Checkpoint Ensemble B0s - 0.4 final ensemble weight\nAugmentation: RandomResizer (30% chance), RandomPad (30% chance), Flipping, Rotation, & Resize(512,512)\nLoss: BCELoss\n*RandomPad -> pad each side (independently) with a random length between 0 and 200\n\nSince it was hard to estimate the \"best\" model on the local validation set, we used the idea of checkpoint ensembling to try to avoid overfitting & boost our score.\n\nWe oversampled classes with less examples or less true positives and tried to avoid overusing images with more assigned labels.\n\n7. Fine-tuning\nWe fine-tuned all networks for one \"epoch\" on the validation set. Unlabeled images were used once, while images that we labeled (soft labels) were used multiple times in the \"epoch\". The \"epoch\" was around 200-250k images. We used the same augmentations & loss function for each B0 as it was used during training of that network.\n\n8. Cell/Image weighting\nTo extreme outliers we weighted the final confidences with the mean of confidences of all other valid cells (excluding border and outliers). The final confidence was 0.7 * cell_confidence + 0.3 * image_confidence.\n\nSeems that even extreme values such as 0.6/0.4 tend to work well here. We did not test if smaller weighting gave better score.\n\nConclusion\n512x512 images, re-labeling, simple network (B0), fine-tuning & final confidence weighting seem to work well enough for this problem.\n\nThank you\n@christofhenkel & @cwinsnes -> https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229284#1258324\nWe were not aware of this at that time. This valuable responses made a huge impact on our approach/results.\n\n@h053473666 -> https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/230940\nThank you for the motivation for the final weighting.\n\n@emmalumpan, @lnhtrang & @cwinsnes I hope there will be more opportunities to participate in your journey in the future (Kaggle or non-Kaggle related). Thank you for the ride! :)",
    "1307655": "Thanks for sharing - some very interesting heuristics!\n\nI also tried getting label_cells running in parallel using multiprocessing since I wanted to try and do this at full scale on the private set and still have time for inference. This worked fine on my local machine but when I moved to Kaggle's environment I got multiple \"Kaggle Errors\" on submission or other problems related to multiprocessing. Would you be willing to share any code snippets for how you implemented this?"
  }
}