{
  "id": 239641,
  "title": "20th Place Solution - Two stage model",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/239641",
  "author_name": "Da Yu",
  "post_date": "2021-05-17T06:16:40.051000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thank you to the competition host for such an interesting competition. I had quite some fun and learnt a lot from this competition. Thank you fellow HPA competitors for sharing your ideas and solution. To the winners of this competitions, congratulations!</p>\n<h1>Introduction</h1>\n<p>Like most competitors I used the segmentation algorithms provided by the organizers to segment the cells . The data provides weak positive labels, but an abundance of strong negative labels. The converse is true for the negative class. So I trained a two-stage model, the first on image level labels and the second on cell level labels predicted by the first stage.</p>\n<h1>Stage 1</h1>\n<p>An ensemble of 6 efficientnet-b0 models to predict image level labels, trained with different random seeds but with the full datasets.</p>\n<h2>Augmentation</h2>\n<ul>\n<li>Random flip</li>\n<li>Random rotate 90</li>\n<li>Random brightness (independent in each RGBY channel)</li>\n<li>Random crop</li>\n</ul>\n<h2>Training settings</h2>\n<ul>\n<li>A slightly modified <a href=\"https://openaccess.thecvf.com/content_CVPR_2019/papers/Cui_Class-Balanced_Loss_Based_on_Effective_Number_of_Samples_CVPR_2019_paper.pdf\" target=\"_blank\">balanced focal loss</a>. Instead of assigning a weight to each class I calculated a different alpha for each class assuming a one-vs-all classification.</li>\n<li>cosine annealing with warmup</li>\n<li>1e-2 learning rate</li>\n<li>10 warmup epochs</li>\n<li>90 epochs</li>\n<li>AdamW with 1e-3 weight decay</li>\n</ul>\n<h2>Results on LB</h2>\n<p>The stage 1 model has a LB score of 0.47455 private, 0.48021 public with 8x TTA.</p>\n<h1>Stage 2</h1>\n<p>In stage 2, I made two assumptions on the image level labels provided by the organizers.</p>\n<ol>\n<li>Positive labels indicates at least one cell in the image with the target class.</li>\n<li>Negative labels indicates the all the cells in the image do not contain the target class.</li>\n</ol>\n<p>So I used the stage 1 ensemble to predict individual cell labels but I only keep the cell-wise labels for classes where their ground truth image label is positive. The cell labels for classes which have negative ground truth image label are set to 0.</p>\n<p>For the negative classes, I predict negative_prob = 1 - max(other classes). I downloaded 1123 negative images from the extra HPA dataset and assign a negative class label of 1.</p>\n<p>I then trained this stage 2 model with the following setting.</p>\n<h2>Augmentation</h2>\n<ul>\n<li>Random crop 10% off height and width</li>\n<li>Random flip</li>\n<li>Random rotate 90</li>\n<li>Random brightness (independent in each RGBY channel)</li>\n<li>With 50% probability, I multiply the red, blue and yellow channels with a random number between 0 and 1. This is motivated by the observation that some duplicate images have the same structure but different intensities between different channel.</li>\n<li>Crop or pad images to 224x224</li>\n</ul>\n<h2>Hyperparameters</h2>\n<p>I didn't have time to try different hyperparameters but the settings I used for my solution is</p>\n<ul>\n<li>Cosine annealing with warmup</li>\n<li>learning rate 1e-3</li>\n<li>10 warmup epochs, 30 epochs</li>\n<li>AdamW optimizers with 1-e4 weight decay (this is quite important as the cell-level models overtrain very quickly into NaNs on mixed precision without weight decays).</li>\n</ul>\n<p>This model is very expensive to train as I had more than 500,000 examples. One epoch took 6 minutes on the latest gen GPU.</p>\n<h2>Results of stage 2 training</h2>\n<p>The stage 2 training improve my score by about 0.3 in my best solution, others vary from 0.1-0.8 based on what I tried in stage 1 models. My best submission has a score of 0.50626 private 0.50738 public.</p>\n<h1>Extra Data</h1>\n<p>I downloaded the extra public dataset provided by the organizers but I only selected from classes which had image-level mAP scores of less than 0.8 based on the validation set of my stage 1 model. I also downloaded 1123 negative (no target class) images.</p>\n<h1>Choice of model</h1>\n<p>I noticed efficientnet-b0 performed the best in stage 1. Other models I tried were mobilenet v3 and resnet-50, resnet-101. Larger efficientnet models (I tried b1 and b2) didn't perform as well as b0.</p>\n<h1>Post-processing</h1>\n<p>Unfortunately, I ran out of time towards the end of the competition and I wasn't able to try combining image level predictions with cell level predictions or trying something to deal with edge cells. (One thing I learnt was that time management is quite important in competitions like this.)</p>\n<h1>Things that did not work</h1>\n<ul>\n<li>Concatenate a global max pool and global average pool in the last layer of the stage 1 (image level) models.</li>\n<li><a href=\"https://arxiv.org/abs/2103.07246\" target=\"_blank\">DRS</a> in the global average pooling layer. Maybe there is some hyperparameter setting or something I got wrong.</li>\n</ul>\n<h1>Lack of plots or image</h1>\n<p>It seems I cannot upload images to kaggle at the moment. I hope to update this post with plots or images in the future when that is possible.</p>",
  "messages": [
    {
      "id": 1311060,
      "postDate": "2021-05-17T06:16:40.050Z",
      "content": "<p>Thank you to the competition host for such an interesting competition. I had quite some fun and learnt a lot from this competition. Thank you fellow HPA competitors for sharing your ideas and solution. To the winners of this competitions, congratulations!</p>\n<h1>Introduction</h1>\n<p>Like most competitors I used the segmentation algorithms provided by the organizers to segment the cells . The data provides weak positive labels, but an abundance of strong negative labels. The converse is true for the negative class. So I trained a two-stage model, the first on image level labels and the second on cell level labels predicted by the first stage.</p>\n<h1>Stage 1</h1>\n<p>An ensemble of 6 efficientnet-b0 models to predict image level labels, trained with different random seeds but with the full datasets.</p>\n<h2>Augmentation</h2>\n<ul>\n<li>Random flip</li>\n<li>Random rotate 90</li>\n<li>Random brightness (independent in each RGBY channel)</li>\n<li>Random crop</li>\n</ul>\n<h2>Training settings</h2>\n<ul>\n<li>A slightly modified <a href=\"https://openaccess.thecvf.com/content_CVPR_2019/papers/Cui_Class-Balanced_Loss_Based_on_Effective_Number_of_Samples_CVPR_2019_paper.pdf\" target=\"_blank\">balanced focal loss</a>. Instead of assigning a weight to each class I calculated a different alpha for each class assuming a one-vs-all classification.</li>\n<li>cosine annealing with warmup</li>\n<li>1e-2 learning rate</li>\n<li>10 warmup epochs</li>\n<li>90 epochs</li>\n<li>AdamW with 1e-3 weight decay</li>\n</ul>\n<h2>Results on LB</h2>\n<p>The stage 1 model has a LB score of 0.47455 private, 0.48021 public with 8x TTA.</p>\n<h1>Stage 2</h1>\n<p>In stage 2, I made two assumptions on the image level labels provided by the organizers.</p>\n<ol>\n<li>Positive labels indicates at least one cell in the image with the target class.</li>\n<li>Negative labels indicates the all the cells in the image do not contain the target class.</li>\n</ol>\n<p>So I used the stage 1 ensemble to predict individual cell labels but I only keep the cell-wise labels for classes where their ground truth image label is positive. The cell labels for classes which have negative ground truth image label are set to 0.</p>\n<p>For the negative classes, I predict negative_prob = 1 - max(other classes). I downloaded 1123 negative images from the extra HPA dataset and assign a negative class label of 1.</p>\n<p>I then trained this stage 2 model with the following setting.</p>\n<h2>Augmentation</h2>\n<ul>\n<li>Random crop 10% off height and width</li>\n<li>Random flip</li>\n<li>Random rotate 90</li>\n<li>Random brightness (independent in each RGBY channel)</li>\n<li>With 50% probability, I multiply the red, blue and yellow channels with a random number between 0 and 1. This is motivated by the observation that some duplicate images have the same structure but different intensities between different channel.</li>\n<li>Crop or pad images to 224x224</li>\n</ul>\n<h2>Hyperparameters</h2>\n<p>I didn't have time to try different hyperparameters but the settings I used for my solution is</p>\n<ul>\n<li>Cosine annealing with warmup</li>\n<li>learning rate 1e-3</li>\n<li>10 warmup epochs, 30 epochs</li>\n<li>AdamW optimizers with 1-e4 weight decay (this is quite important as the cell-level models overtrain very quickly into NaNs on mixed precision without weight decays).</li>\n</ul>\n<p>This model is very expensive to train as I had more than 500,000 examples. One epoch took 6 minutes on the latest gen GPU.</p>\n<h2>Results of stage 2 training</h2>\n<p>The stage 2 training improve my score by about 0.3 in my best solution, others vary from 0.1-0.8 based on what I tried in stage 1 models. My best submission has a score of 0.50626 private 0.50738 public.</p>\n<h1>Extra Data</h1>\n<p>I downloaded the extra public dataset provided by the organizers but I only selected from classes which had image-level mAP scores of less than 0.8 based on the validation set of my stage 1 model. I also downloaded 1123 negative (no target class) images.</p>\n<h1>Choice of model</h1>\n<p>I noticed efficientnet-b0 performed the best in stage 1. Other models I tried were mobilenet v3 and resnet-50, resnet-101. Larger efficientnet models (I tried b1 and b2) didn't perform as well as b0.</p>\n<h1>Post-processing</h1>\n<p>Unfortunately, I ran out of time towards the end of the competition and I wasn't able to try combining image level predictions with cell level predictions or trying something to deal with edge cells. (One thing I learnt was that time management is quite important in competitions like this.)</p>\n<h1>Things that did not work</h1>\n<ul>\n<li>Concatenate a global max pool and global average pool in the last layer of the stage 1 (image level) models.</li>\n<li><a href=\"https://arxiv.org/abs/2103.07246\" target=\"_blank\">DRS</a> in the global average pooling layer. Maybe there is some hyperparameter setting or something I got wrong.</li>\n</ul>\n<h1>Lack of plots or image</h1>\n<p>It seems I cannot upload images to kaggle at the moment. I hope to update this post with plots or images in the future when that is possible.</p>",
      "rawMarkdown": "Thank you to the competition host for such an interesting competition. I had quite some fun and learnt a lot from this competition. Thank you fellow HPA competitors for sharing your ideas and solution. To the winners of this competitions, congratulations!\n\n# Introduction\nLike most competitors I used the segmentation algorithms provided by the organizers to segment the cells . The data provides weak positive labels, but an abundance of strong negative labels. The converse is true for the negative class. So I trained a two-stage model, the first on image level labels and the second on cell level labels predicted by the first stage.\n\n# Stage 1\nAn ensemble of 6 efficientnet-b0 models to predict image level labels, trained with different random seeds but with the full datasets.\n\n## Augmentation\n- Random flip\n- Random rotate 90\n- Random brightness (independent in each RGBY channel)\n- Random crop\n\n## Training settings\n- A slightly modified [balanced focal loss](https://openaccess.thecvf.com/content_CVPR_2019/papers/Cui_Class-Balanced_Loss_Based_on_Effective_Number_of_Samples_CVPR_2019_paper.pdf). Instead of assigning a weight to each class I calculated a different alpha for each class assuming a one-vs-all classification.\n- cosine annealing with warmup\n- 1e-2 learning rate\n- 10 warmup epochs\n- 90 epochs\n- AdamW with 1e-3 weight decay\n\n## Results on LB\nThe stage 1 model has a LB score of 0.47455 private, 0.48021 public with 8x TTA.\n\n# Stage 2\nIn stage 2, I made two assumptions on the image level labels provided by the organizers.\n1. Positive labels indicates at least one cell in the image with the target class.\n2. Negative labels indicates the all the cells in the image do not contain the target class.\n\nSo I used the stage 1 ensemble to predict individual cell labels but I only keep the cell-wise labels for classes where their ground truth image label is positive. The cell labels for classes which have negative ground truth image label are set to 0.\n\nFor the negative classes, I predict negative_prob = 1 - max(other classes). I downloaded 1123 negative images from the extra HPA dataset and assign a negative class label of 1.\n\nI then trained this stage 2 model with the following setting.\n\n## Augmentation\n- Random crop 10% off height and width\n- Random flip\n- Random rotate 90\n- Random brightness (independent in each RGBY channel)\n- With 50% probability, I multiply the red, blue and yellow channels with a random number between 0 and 1. This is motivated by the observation that some duplicate images have the same structure but different intensities between different channel.\n- Crop or pad images to 224x224\n\n## Hyperparameters\nI didn't have time to try different hyperparameters but the settings I used for my solution is\n- Cosine annealing with warmup\n- learning rate 1e-3\n- 10 warmup epochs, 30 epochs\n- AdamW optimizers with 1-e4 weight decay (this is quite important as the cell-level models overtrain very quickly into NaNs on mixed precision without weight decays).\n\nThis model is very expensive to train as I had more than 500,000 examples. One epoch took 6 minutes on the latest gen GPU.\n\n## Results of stage 2 training\nThe stage 2 training improve my score by about 0.3 in my best solution, others vary from 0.1-0.8 based on what I tried in stage 1 models. My best submission has a score of 0.50626 private 0.50738 public.\n\n# Extra Data\nI downloaded the extra public dataset provided by the organizers but I only selected from classes which had image-level mAP scores of less than 0.8 based on the validation set of my stage 1 model. I also downloaded 1123 negative (no target class) images.\n\n# Choice of model\nI noticed efficientnet-b0 performed the best in stage 1. Other models I tried were mobilenet v3 and resnet-50, resnet-101. Larger efficientnet models (I tried b1 and b2) didn't perform as well as b0.\n\n# Post-processing\nUnfortunately, I ran out of time towards the end of the competition and I wasn't able to try combining image level predictions with cell level predictions or trying something to deal with edge cells. (One thing I learnt was that time management is quite important in competitions like this.)\n\n# Things that did not work\n- Concatenate a global max pool and global average pool in the last layer of the stage 1 (image level) models.\n- [DRS](https://arxiv.org/abs/2103.07246) in the global average pooling layer. Maybe there is some hyperparameter setting or something I got wrong.\n\n# Lack of plots or image\nIt seems I cannot upload images to kaggle at the moment. I hope to update this post with plots or images in the future when that is possible.\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1311060": "Thank you to the competition host for such an interesting competition. I had quite some fun and learnt a lot from this competition. Thank you fellow HPA competitors for sharing your ideas and solution. To the winners of this competitions, congratulations!\n\n# Introduction\nLike most competitors I used the segmentation algorithms provided by the organizers to segment the cells . The data provides weak positive labels, but an abundance of strong negative labels. The converse is true for the negative class. So I trained a two-stage model, the first on image level labels and the second on cell level labels predicted by the first stage.\n\n# Stage 1\nAn ensemble of 6 efficientnet-b0 models to predict image level labels, trained with different random seeds but with the full datasets.\n\n## Augmentation\n- Random flip\n- Random rotate 90\n- Random brightness (independent in each RGBY channel)\n- Random crop\n\n## Training settings\n- A slightly modified [balanced focal loss](https://openaccess.thecvf.com/content_CVPR_2019/papers/Cui_Class-Balanced_Loss_Based_on_Effective_Number_of_Samples_CVPR_2019_paper.pdf). Instead of assigning a weight to each class I calculated a different alpha for each class assuming a one-vs-all classification.\n- cosine annealing with warmup\n- 1e-2 learning rate\n- 10 warmup epochs\n- 90 epochs\n- AdamW with 1e-3 weight decay\n\n## Results on LB\nThe stage 1 model has a LB score of 0.47455 private, 0.48021 public with 8x TTA.\n\n# Stage 2\nIn stage 2, I made two assumptions on the image level labels provided by the organizers.\n1. Positive labels indicates at least one cell in the image with the target class.\n2. Negative labels indicates the all the cells in the image do not contain the target class.\n\nSo I used the stage 1 ensemble to predict individual cell labels but I only keep the cell-wise labels for classes where their ground truth image label is positive. The cell labels for classes which have negative ground truth image label are set to 0.\n\nFor the negative classes, I predict negative_prob = 1 - max(other classes). I downloaded 1123 negative images from the extra HPA dataset and assign a negative class label of 1.\n\nI then trained this stage 2 model with the following setting.\n\n## Augmentation\n- Random crop 10% off height and width\n- Random flip\n- Random rotate 90\n- Random brightness (independent in each RGBY channel)\n- With 50% probability, I multiply the red, blue and yellow channels with a random number between 0 and 1. This is motivated by the observation that some duplicate images have the same structure but different intensities between different channel.\n- Crop or pad images to 224x224\n\n## Hyperparameters\nI didn't have time to try different hyperparameters but the settings I used for my solution is\n- Cosine annealing with warmup\n- learning rate 1e-3\n- 10 warmup epochs, 30 epochs\n- AdamW optimizers with 1-e4 weight decay (this is quite important as the cell-level models overtrain very quickly into NaNs on mixed precision without weight decays).\n\nThis model is very expensive to train as I had more than 500,000 examples. One epoch took 6 minutes on the latest gen GPU.\n\n## Results of stage 2 training\nThe stage 2 training improve my score by about 0.3 in my best solution, others vary from 0.1-0.8 based on what I tried in stage 1 models. My best submission has a score of 0.50626 private 0.50738 public.\n\n# Extra Data\nI downloaded the extra public dataset provided by the organizers but I only selected from classes which had image-level mAP scores of less than 0.8 based on the validation set of my stage 1 model. I also downloaded 1123 negative (no target class) images.\n\n# Choice of model\nI noticed efficientnet-b0 performed the best in stage 1. Other models I tried were mobilenet v3 and resnet-50, resnet-101. Larger efficientnet models (I tried b1 and b2) didn't perform as well as b0.\n\n# Post-processing\nUnfortunately, I ran out of time towards the end of the competition and I wasn't able to try combining image level predictions with cell level predictions or trying something to deal with edge cells. (One thing I learnt was that time management is quite important in competitions like this.)\n\n# Things that did not work\n- Concatenate a global max pool and global average pool in the last layer of the stage 1 (image level) models.\n- [DRS](https://arxiv.org/abs/2103.07246) in the global average pooling layer. Maybe there is some hyperparameter setting or something I got wrong.\n\n# Lack of plots or image\nIt seems I cannot upload images to kaggle at the moment. I hope to update this post with plots or images in the future when that is possible.\n"
  }
}