{
  "id": 241637,
  "title": "32nd Place Solution",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/241637",
  "author_name": "Hung Quoc To",
  "post_date": "2021-05-25T12:30:09.589000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I had a great experience in the last months attempting to solve an interesting problem. Thank you all the participants for your great ideas and notebooks that helped me achieve my final results. I have learnt a lot from you. Thanks to the host team for always being supportive during the competition.</p>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/quochungto/final-submission-hpa\" target=\"_blank\">https://www.kaggle.com/quochungto/final-submission-hpa</a></p>\n<h3><strong>OVERVIEW</strong></h3>\n<p><strong>Main components of my solution:</strong></p>\n<ul>\n<li>2 stages pipeline - Cell level &amp; Image level predictions</li>\n<li>Eliminating duplicate images - more reliable validation</li>\n<li>Faster version of cell segmentations</li>\n<li>External dataset</li>\n<li>Increasing diversity of models for ensembling: models trained on Green/RGB, 8bit/16bit images</li>\n</ul>\n<p>As the competition stated, we are dealing with a weakly supervised semantic segmentation problem and only have labels for image level, not cell level. The task is to predict the label of each cell for an image with multiple cells.</p>\n<p>Training single stage models to predict each cell’s label is difficult because we do not have the true target and using image labels as target will make it noisy to train. The idea is that each cell’s label inherits part of the image label, so using image prediction as context to correct each cell prediction will be effective compared to just predicting each cell independently.</p>\n<h3><strong>TRAINING DETAILS</strong></h3>\n<h4><strong>Dataset</strong></h4>\n<p>Competition dataset + External dataset</p>\n<h4><strong>Preprocessing</strong></h4>\n<p>As pointed out from other participants, the training dataset contained some duplicated images just like the last competition. Using this dataset directly in training will make validation less reliable. So the first thing I do is to try to remove all the duplicate images in the full dataset.</p>\n<h4><strong>Cell Segmentation</strong></h4>\n<p>All cell masks are extracted using the host's HPA Cell Segmentator as it does quite well in segmenting cells. Then each R, G, B channel is combined  into a single RGB image and saved as a new dataset. This reduces the problem to just weakly supervised multi-label classification.</p>\n<h4><strong>Models</strong></h4>\n<p>In order to maximize the effectiveness of ensembles, different backbones and inputs are used. This increases the diversity of my models.</p>\n<p><strong>Cell level models</strong></p>\n<ul>\n<li>Input: Green/RGB cell images extracted from HPA Cell Segmentator</li>\n<li>Data augmentation: Horizontal &amp; vertical flip, random crop</li>\n<li>Dataset: Competition &amp; external datasets</li>\n<li>Epochs: 2 to 4</li>\n<li>Backbones: Resnet50, Efficientnet-b5</li>\n<li>Loss: BCEWithLogitLoss</li>\n<li>Optimizer: Adam</li>\n<li>TTA: 1 on original data + 4x</li>\n</ul>\n<p><strong>Image level models</strong></p>\n<ul>\n<li>Input: Green/RGB images, 8-bit/16-bit images</li>\n<li>Data augmentation: Horizontal &amp; vertical flip, rotate, shear, shift, zoom, random brightness</li>\n<li>Dataset: Competition &amp; external datasets</li>\n<li>Epochs: 40 to 50</li>\n<li>Backbones: Resnet50, Densenet121, Efficientnet-b0, b1, b2, b3, b5, b7</li>\n<li>Loss: SigmoidFocalCrossEntropy</li>\n<li>Optimizer: Adam</li>\n<li>Learning rate scheduler: One-cycle</li>\n<li>TTA: 1 on original data + 4x</li>\n</ul>\n<h4><strong>Ensemble</strong></h4>\n<ul>\n<li>Weighted average ensemble for each of cell level and image level prediction</li>\n<li>Final prediction as: <code>0.25 * cell_prediction + 0.25 * image_prediction + 0.5 * sqrt(cell_prediction * image_prediction)</code></li>\n</ul>\n<h4><strong>Others</strong></h4>\n<p>Negative class:</p>\n<ul>\n<li>By replacing prediction of negative class with <code>P(neg) = prod(1 - P(cls_i))</code>, my public LB increased ~0.003</li>\n</ul>\n<h4><strong>Things that did not work</strong></h4>\n<ul>\n<li>Hand labeling negative cell by calculating mean pixel intensity (public LB dropped)</li>\n<li>Eliminating border cells (no significant effect on public LB)</li>\n</ul>",
  "messages": [
    {
      "id": 1322389,
      "postDate": "2021-05-25T12:30:09.590Z",
      "content": "<p>I had a great experience in the last months attempting to solve an interesting problem. Thank you all the participants for your great ideas and notebooks that helped me achieve my final results. I have learnt a lot from you. Thanks to the host team for always being supportive during the competition.</p>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/quochungto/final-submission-hpa\" target=\"_blank\">https://www.kaggle.com/quochungto/final-submission-hpa</a></p>\n<h3><strong>OVERVIEW</strong></h3>\n<p><strong>Main components of my solution:</strong></p>\n<ul>\n<li>2 stages pipeline - Cell level &amp; Image level predictions</li>\n<li>Eliminating duplicate images - more reliable validation</li>\n<li>Faster version of cell segmentations</li>\n<li>External dataset</li>\n<li>Increasing diversity of models for ensembling: models trained on Green/RGB, 8bit/16bit images</li>\n</ul>\n<p>As the competition stated, we are dealing with a weakly supervised semantic segmentation problem and only have labels for image level, not cell level. The task is to predict the label of each cell for an image with multiple cells.</p>\n<p>Training single stage models to predict each cell’s label is difficult because we do not have the true target and using image labels as target will make it noisy to train. The idea is that each cell’s label inherits part of the image label, so using image prediction as context to correct each cell prediction will be effective compared to just predicting each cell independently.</p>\n<h3><strong>TRAINING DETAILS</strong></h3>\n<h4><strong>Dataset</strong></h4>\n<p>Competition dataset + External dataset</p>\n<h4><strong>Preprocessing</strong></h4>\n<p>As pointed out from other participants, the training dataset contained some duplicated images just like the last competition. Using this dataset directly in training will make validation less reliable. So the first thing I do is to try to remove all the duplicate images in the full dataset.</p>\n<h4><strong>Cell Segmentation</strong></h4>\n<p>All cell masks are extracted using the host's HPA Cell Segmentator as it does quite well in segmenting cells. Then each R, G, B channel is combined  into a single RGB image and saved as a new dataset. This reduces the problem to just weakly supervised multi-label classification.</p>\n<h4><strong>Models</strong></h4>\n<p>In order to maximize the effectiveness of ensembles, different backbones and inputs are used. This increases the diversity of my models.</p>\n<p><strong>Cell level models</strong></p>\n<ul>\n<li>Input: Green/RGB cell images extracted from HPA Cell Segmentator</li>\n<li>Data augmentation: Horizontal &amp; vertical flip, random crop</li>\n<li>Dataset: Competition &amp; external datasets</li>\n<li>Epochs: 2 to 4</li>\n<li>Backbones: Resnet50, Efficientnet-b5</li>\n<li>Loss: BCEWithLogitLoss</li>\n<li>Optimizer: Adam</li>\n<li>TTA: 1 on original data + 4x</li>\n</ul>\n<p><strong>Image level models</strong></p>\n<ul>\n<li>Input: Green/RGB images, 8-bit/16-bit images</li>\n<li>Data augmentation: Horizontal &amp; vertical flip, rotate, shear, shift, zoom, random brightness</li>\n<li>Dataset: Competition &amp; external datasets</li>\n<li>Epochs: 40 to 50</li>\n<li>Backbones: Resnet50, Densenet121, Efficientnet-b0, b1, b2, b3, b5, b7</li>\n<li>Loss: SigmoidFocalCrossEntropy</li>\n<li>Optimizer: Adam</li>\n<li>Learning rate scheduler: One-cycle</li>\n<li>TTA: 1 on original data + 4x</li>\n</ul>\n<h4><strong>Ensemble</strong></h4>\n<ul>\n<li>Weighted average ensemble for each of cell level and image level prediction</li>\n<li>Final prediction as: <code>0.25 * cell_prediction + 0.25 * image_prediction + 0.5 * sqrt(cell_prediction * image_prediction)</code></li>\n</ul>\n<h4><strong>Others</strong></h4>\n<p>Negative class:</p>\n<ul>\n<li>By replacing prediction of negative class with <code>P(neg) = prod(1 - P(cls_i))</code>, my public LB increased ~0.003</li>\n</ul>\n<h4><strong>Things that did not work</strong></h4>\n<ul>\n<li>Hand labeling negative cell by calculating mean pixel intensity (public LB dropped)</li>\n<li>Eliminating border cells (no significant effect on public LB)</li>\n</ul>",
      "rawMarkdown": "I had a great experience in the last months attempting to solve an interesting problem. Thank you all the participants for your great ideas and notebooks that helped me achieve my final results. I have learnt a lot from you. Thanks to the host team for always being supportive during the competition.\n\nInference Notebook: https://www.kaggle.com/quochungto/final-submission-hpa\n\n### **OVERVIEW**\n\n**Main components of my solution:**\n- 2 stages pipeline - Cell level & Image level predictions\n- Eliminating duplicate images - more reliable validation\n- Faster version of cell segmentations\n- External dataset\n- Increasing diversity of models for ensembling: models trained on Green/RGB, 8bit/16bit images\n\nAs the competition stated, we are dealing with a weakly supervised semantic segmentation problem and only have labels for image level, not cell level. The task is to predict the label of each cell for an image with multiple cells.\n\nTraining single stage models to predict each cell’s label is difficult because we do not have the true target and using image labels as target will make it noisy to train. The idea is that each cell’s label inherits part of the image label, so using image prediction as context to correct each cell prediction will be effective compared to just predicting each cell independently.\n\n### **TRAINING DETAILS**\n\n#### **Dataset**\nCompetition dataset + External dataset\n\n#### **Preprocessing**\nAs pointed out from other participants, the training dataset contained some duplicated images just like the last competition. Using this dataset directly in training will make validation less reliable. So the first thing I do is to try to remove all the duplicate images in the full dataset.\n\n#### **Cell Segmentation**\nAll cell masks are extracted using the host's HPA Cell Segmentator as it does quite well in segmenting cells. Then each R, G, B channel is combined  into a single RGB image and saved as a new dataset. This reduces the problem to just weakly supervised multi-label classification.\n\n#### **Models**\nIn order to maximize the effectiveness of ensembles, different backbones and inputs are used. This increases the diversity of my models.\n\n**Cell level models**\n- Input: Green/RGB cell images extracted from HPA Cell Segmentator\n- Data augmentation: Horizontal & vertical flip, random crop\n- Dataset: Competition & external datasets\n- Epochs: 2 to 4\n- Backbones: Resnet50, Efficientnet-b5\n- Loss: BCEWithLogitLoss\n- Optimizer: Adam\n- TTA: 1 on original data + 4x\n\n**Image level models**\n- Input: Green/RGB images, 8-bit/16-bit images\n- Data augmentation: Horizontal & vertical flip, rotate, shear, shift, zoom, random brightness\n- Dataset: Competition & external datasets\n- Epochs: 40 to 50\n- Backbones: Resnet50, Densenet121, Efficientnet-b0, b1, b2, b3, b5, b7\n- Loss: SigmoidFocalCrossEntropy\n- Optimizer: Adam\n- Learning rate scheduler: One-cycle\n- TTA: 1 on original data + 4x\n\n#### **Ensemble**\n- Weighted average ensemble for each of cell level and image level prediction\n- Final prediction as: `0.25 * cell_prediction + 0.25 * image_prediction + 0.5 * sqrt(cell_prediction * image_prediction)`\n\n#### **Others**\nNegative class:\n- By replacing prediction of negative class with `P(neg) = prod(1 - P(cls_i))`, my public LB increased ~0.003\n\n#### **Things that did not work**\n- Hand labeling negative cell by calculating mean pixel intensity (public LB dropped)\n- Eliminating border cells (no significant effect on public LB)\n    \n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1322389": "I had a great experience in the last months attempting to solve an interesting problem. Thank you all the participants for your great ideas and notebooks that helped me achieve my final results. I have learnt a lot from you. Thanks to the host team for always being supportive during the competition.\n\nInference Notebook: https://www.kaggle.com/quochungto/final-submission-hpa\n\n### **OVERVIEW**\n\n**Main components of my solution:**\n- 2 stages pipeline - Cell level & Image level predictions\n- Eliminating duplicate images - more reliable validation\n- Faster version of cell segmentations\n- External dataset\n- Increasing diversity of models for ensembling: models trained on Green/RGB, 8bit/16bit images\n\nAs the competition stated, we are dealing with a weakly supervised semantic segmentation problem and only have labels for image level, not cell level. The task is to predict the label of each cell for an image with multiple cells.\n\nTraining single stage models to predict each cell’s label is difficult because we do not have the true target and using image labels as target will make it noisy to train. The idea is that each cell’s label inherits part of the image label, so using image prediction as context to correct each cell prediction will be effective compared to just predicting each cell independently.\n\n### **TRAINING DETAILS**\n\n#### **Dataset**\nCompetition dataset + External dataset\n\n#### **Preprocessing**\nAs pointed out from other participants, the training dataset contained some duplicated images just like the last competition. Using this dataset directly in training will make validation less reliable. So the first thing I do is to try to remove all the duplicate images in the full dataset.\n\n#### **Cell Segmentation**\nAll cell masks are extracted using the host's HPA Cell Segmentator as it does quite well in segmenting cells. Then each R, G, B channel is combined  into a single RGB image and saved as a new dataset. This reduces the problem to just weakly supervised multi-label classification.\n\n#### **Models**\nIn order to maximize the effectiveness of ensembles, different backbones and inputs are used. This increases the diversity of my models.\n\n**Cell level models**\n- Input: Green/RGB cell images extracted from HPA Cell Segmentator\n- Data augmentation: Horizontal & vertical flip, random crop\n- Dataset: Competition & external datasets\n- Epochs: 2 to 4\n- Backbones: Resnet50, Efficientnet-b5\n- Loss: BCEWithLogitLoss\n- Optimizer: Adam\n- TTA: 1 on original data + 4x\n\n**Image level models**\n- Input: Green/RGB images, 8-bit/16-bit images\n- Data augmentation: Horizontal & vertical flip, rotate, shear, shift, zoom, random brightness\n- Dataset: Competition & external datasets\n- Epochs: 40 to 50\n- Backbones: Resnet50, Densenet121, Efficientnet-b0, b1, b2, b3, b5, b7\n- Loss: SigmoidFocalCrossEntropy\n- Optimizer: Adam\n- Learning rate scheduler: One-cycle\n- TTA: 1 on original data + 4x\n\n#### **Ensemble**\n- Weighted average ensemble for each of cell level and image level prediction\n- Final prediction as: `0.25 * cell_prediction + 0.25 * image_prediction + 0.5 * sqrt(cell_prediction * image_prediction)`\n\n#### **Others**\nNegative class:\n- By replacing prediction of negative class with `P(neg) = prod(1 - P(cls_i))`, my public LB increased ~0.003\n\n#### **Things that did not work**\n- Hand labeling negative cell by calculating mean pixel intensity (public LB dropped)\n- Eliminating border cells (no significant effect on public LB)\n    \n"
  }
}