{
  "id": 239211,
  "title": "11st Place Solution Summary: Active learning for cell label mining and Ensemble of different models",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/239211",
  "author_name": "rongbo shen",
  "post_date": "2021-05-15T08:57:32.218000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to thank Kaggle and organizers hosting this really interesting competition. Thanks to my teammates <a href=\"https://www.kaggle.com/silversh\" target=\"_blank\">@silversh</a> and <a href=\"https://www.kaggle.com/jianguoguo\" target=\"_blank\">@jianguoguo</a>  for the hard work and collaboration. </p>\n<p>Congratulations to all the winners. This is my first Kaggle competition, I really learn a lot of nice ideas and solutions in the forum.</p>\n<p><strong>1. Overview</strong><br>\nOur final solution includes active learning for cell label mining and ensemble of different models.</p>\n<p><strong>2. Active learning for cell label mining</strong><br>\nBecause I do not participate Kaggle conpetition before, I'm not sure whether data label mining is allowed. Therefore, we try to use active learning to mine cell labels relatively late. Due to time limitation, we only apply it for a few classes, i.e., 9, 11, 12 and 15.</p>\n<p>First, we chose the cells in the single-label image as the initial trainset, and the cell labels directly inherit the image labels. For classes 11 and 15 with few images, we manually selected about 200 cells with ground truth labels into the initial trainset. After the model is trained on the initial trainset, it is used to make predictions for cells in all images. Use the following strategies to prepare the trainset for the next round.</p>\n<p>If an image contains original label X and the prediction confidence of one cell in this image is greater than 0.25, the label X of the cell is retained, otherwise the cell is removed from the next round trainset. If an image does not contain original label X, and the predictive confidence of one cell in this image is greater than 0.5, this cell will be assigned a true label X. </p>\n<p>After several rounds, the noise of most cell labels can be inhibited. For example, we finally retrieve about 2,000 cells for class 11 after 4 rounds. On the final public test set, the AP score for class 11 is improved from 0.001 to 0.035.</p>\n<p>The specific principle of active learning can be referred to my previous papers <a href=\"url\" target=\"_blank\">https://doi.org/10.1016/j.future.2019.07.013</a>.</p>\n<p><strong>3. Ensemble of different models</strong><br>\nWe use ensemble method to fuse Cell-level models and Image-level models. The weights are 0.75 and 0.25, respectively. The Cell-level models employ EfficientNet-B3 network, and the cell score is averaged after 5fold cross-validation. For the Image-level models, we train EfficientNet-B7 networks with RGBY 4-channel image input and Green 1-channel image input, respectively, and the image score is averaged from the two Image-level models. Specifically, when the Cell-level models and the Image-level models are fused, the prediction result of class 11 only use the prediction score of the Cell-level model.</p>\n<p><strong>4. Configuration of training</strong></p>\n<ul>\n<li>Image-level model:</li>\n</ul>\n<ol>\n<li>EfficientNet-B7，Concat-pooling + 2 layers of FC Head, input size 600 × 600</li>\n<li>Data augmentation: flip, rotation</li>\n<li>Focal Loss</li>\n<li>SWA (Nearly useless)</li>\n<li>Fusion of RGBY 4-channel model and Green 1-channel model</li>\n</ol>\n<ul>\n<li>Cell-level model:</li>\n</ul>\n<ol>\n<li>EfficientNet-B3，Concat-pooling + 2 layers of FC Head, input size 300 × 300</li>\n<li>Data augmentation: flip, rotation, mixup</li>\n<li>Focal Loss, Label smooth [0.1, 0.9]</li>\n<li>SWA (Nearly useless)</li>\n<li>5-folds cross validation</li>\n</ol>\n<p><strong>5. Some results</strong><br>\n| One Cell-level model only | active learning for classes | Public LB | Private LB |<br>\n| One Cell-level model only | 11, 15 at 1st round               | 0.446       | 0.465        |<br>\n| One Cell-level model only | 11, 15 at 2nd round             | 0.462        | 0.494       |<br>\n| One Cell-level model only | 11, 15 at 4th round              | 0.502        | -               |</p>\n<p>| Ensemble model        | active learning for classes | Public LB | Private LB |<br>\n| Ensemble model        | 11, 15               | 0.558       | 0.528        |<br>\n| Ensemble model        | 9, 11, 12, 15     | 0.558       | 0.532        | -&gt; Final result in LB<br>\n| Ensemble model        | 9, 11, 12, 15 at 4th round, 4 at 1st round | 0.561        | 0.537       | -&gt;Notebook timeout</p>\n<p><strong>6. Others</strong><br>\nTTA is helpful for prediction, with an increase of 0.001~0.003, but it may tend to time out for our submissions.</p>\n<p>Early in the competition, we tried to solve this problem by using multi-instance learning with transformer encoder-based attention, but the effect was not good.</p>\n<p>As the results above, we additionally use a round of active learning for cell label mining on class 4. The submission was timeout and we resubmited it after the deadline of the competition, and it finally obtained the results of Public LB of 0.561 and Private LB of 0.537, indicating that active learning mining labels can continue to improve the performance if it was extended to other classes. However, we adopted this strategy too late..</p>\n<p>Please forgive my typesetting, I encounter problems inserting images and tables…</p>",
  "messages": [
    {
      "id": 1308490,
      "postDate": "2021-05-15T08:57:32.217Z",
      "content": "<p>First of all, I would like to thank Kaggle and organizers hosting this really interesting competition. Thanks to my teammates <a href=\"https://www.kaggle.com/silversh\" target=\"_blank\">@silversh</a> and <a href=\"https://www.kaggle.com/jianguoguo\" target=\"_blank\">@jianguoguo</a>  for the hard work and collaboration. </p>\n<p>Congratulations to all the winners. This is my first Kaggle competition, I really learn a lot of nice ideas and solutions in the forum.</p>\n<p><strong>1. Overview</strong><br>\nOur final solution includes active learning for cell label mining and ensemble of different models.</p>\n<p><strong>2. Active learning for cell label mining</strong><br>\nBecause I do not participate Kaggle conpetition before, I'm not sure whether data label mining is allowed. Therefore, we try to use active learning to mine cell labels relatively late. Due to time limitation, we only apply it for a few classes, i.e., 9, 11, 12 and 15.</p>\n<p>First, we chose the cells in the single-label image as the initial trainset, and the cell labels directly inherit the image labels. For classes 11 and 15 with few images, we manually selected about 200 cells with ground truth labels into the initial trainset. After the model is trained on the initial trainset, it is used to make predictions for cells in all images. Use the following strategies to prepare the trainset for the next round.</p>\n<p>If an image contains original label X and the prediction confidence of one cell in this image is greater than 0.25, the label X of the cell is retained, otherwise the cell is removed from the next round trainset. If an image does not contain original label X, and the predictive confidence of one cell in this image is greater than 0.5, this cell will be assigned a true label X. </p>\n<p>After several rounds, the noise of most cell labels can be inhibited. For example, we finally retrieve about 2,000 cells for class 11 after 4 rounds. On the final public test set, the AP score for class 11 is improved from 0.001 to 0.035.</p>\n<p>The specific principle of active learning can be referred to my previous papers <a href=\"url\" target=\"_blank\">https://doi.org/10.1016/j.future.2019.07.013</a>.</p>\n<p><strong>3. Ensemble of different models</strong><br>\nWe use ensemble method to fuse Cell-level models and Image-level models. The weights are 0.75 and 0.25, respectively. The Cell-level models employ EfficientNet-B3 network, and the cell score is averaged after 5fold cross-validation. For the Image-level models, we train EfficientNet-B7 networks with RGBY 4-channel image input and Green 1-channel image input, respectively, and the image score is averaged from the two Image-level models. Specifically, when the Cell-level models and the Image-level models are fused, the prediction result of class 11 only use the prediction score of the Cell-level model.</p>\n<p><strong>4. Configuration of training</strong></p>\n<ul>\n<li>Image-level model:</li>\n</ul>\n<ol>\n<li>EfficientNet-B7，Concat-pooling + 2 layers of FC Head, input size 600 × 600</li>\n<li>Data augmentation: flip, rotation</li>\n<li>Focal Loss</li>\n<li>SWA (Nearly useless)</li>\n<li>Fusion of RGBY 4-channel model and Green 1-channel model</li>\n</ol>\n<ul>\n<li>Cell-level model:</li>\n</ul>\n<ol>\n<li>EfficientNet-B3，Concat-pooling + 2 layers of FC Head, input size 300 × 300</li>\n<li>Data augmentation: flip, rotation, mixup</li>\n<li>Focal Loss, Label smooth [0.1, 0.9]</li>\n<li>SWA (Nearly useless)</li>\n<li>5-folds cross validation</li>\n</ol>\n<p><strong>5. Some results</strong><br>\n| One Cell-level model only | active learning for classes | Public LB | Private LB |<br>\n| One Cell-level model only | 11, 15 at 1st round               | 0.446       | 0.465        |<br>\n| One Cell-level model only | 11, 15 at 2nd round             | 0.462        | 0.494       |<br>\n| One Cell-level model only | 11, 15 at 4th round              | 0.502        | -               |</p>\n<p>| Ensemble model        | active learning for classes | Public LB | Private LB |<br>\n| Ensemble model        | 11, 15               | 0.558       | 0.528        |<br>\n| Ensemble model        | 9, 11, 12, 15     | 0.558       | 0.532        | -&gt; Final result in LB<br>\n| Ensemble model        | 9, 11, 12, 15 at 4th round, 4 at 1st round | 0.561        | 0.537       | -&gt;Notebook timeout</p>\n<p><strong>6. Others</strong><br>\nTTA is helpful for prediction, with an increase of 0.001~0.003, but it may tend to time out for our submissions.</p>\n<p>Early in the competition, we tried to solve this problem by using multi-instance learning with transformer encoder-based attention, but the effect was not good.</p>\n<p>As the results above, we additionally use a round of active learning for cell label mining on class 4. The submission was timeout and we resubmited it after the deadline of the competition, and it finally obtained the results of Public LB of 0.561 and Private LB of 0.537, indicating that active learning mining labels can continue to improve the performance if it was extended to other classes. However, we adopted this strategy too late..</p>\n<p>Please forgive my typesetting, I encounter problems inserting images and tables…</p>",
      "rawMarkdown": "First of all, I would like to thank Kaggle and organizers hosting this really interesting competition. Thanks to my teammates @silversh and @jianguoguo  for the hard work and collaboration. \n\nCongratulations to all the winners. This is my first Kaggle competition, I really learn a lot of nice ideas and solutions in the forum.\n\n**1. Overview**\nOur final solution includes active learning for cell label mining and ensemble of different models.\n\n**2. Active learning for cell label mining**\nBecause I do not participate Kaggle conpetition before, I'm not sure whether data label mining is allowed. Therefore, we try to use active learning to mine cell labels relatively late. Due to time limitation, we only apply it for a few classes, i.e., 9, 11, 12 and 15.\n\nFirst, we chose the cells in the single-label image as the initial trainset, and the cell labels directly inherit the image labels. For classes 11 and 15 with few images, we manually selected about 200 cells with ground truth labels into the initial trainset. After the model is trained on the initial trainset, it is used to make predictions for cells in all images. Use the following strategies to prepare the trainset for the next round.\n\nIf an image contains original label X and the prediction confidence of one cell in this image is greater than 0.25, the label X of the cell is retained, otherwise the cell is removed from the next round trainset. If an image does not contain original label X, and the predictive confidence of one cell in this image is greater than 0.5, this cell will be assigned a true label X. \n\nAfter several rounds, the noise of most cell labels can be inhibited. For example, we finally retrieve about 2,000 cells for class 11 after 4 rounds. On the final public test set, the AP score for class 11 is improved from 0.001 to 0.035.\n\nThe specific principle of active learning can be referred to my previous papers [https://doi.org/10.1016/j.future.2019.07.013](url).\n\n**3. Ensemble of different models**\nWe use ensemble method to fuse Cell-level models and Image-level models. The weights are 0.75 and 0.25, respectively. The Cell-level models employ EfficientNet-B3 network, and the cell score is averaged after 5fold cross-validation. For the Image-level models, we train EfficientNet-B7 networks with RGBY 4-channel image input and Green 1-channel image input, respectively, and the image score is averaged from the two Image-level models. Specifically, when the Cell-level models and the Image-level models are fused, the prediction result of class 11 only use the prediction score of the Cell-level model.\n\n**4. Configuration of training**\n- Image-level model:\n1.  EfficientNet-B7，Concat-pooling + 2 layers of FC Head, input size 600 × 600\n2. Data augmentation: flip, rotation\n3. Focal Loss\n4. SWA (Nearly useless)\n5. Fusion of RGBY 4-channel model and Green 1-channel model\n\n- Cell-level model:\n1.  EfficientNet-B3，Concat-pooling + 2 layers of FC Head, input size 300 × 300\n2. Data augmentation: flip, rotation, mixup\n3. Focal Loss, Label smooth [0.1, 0.9]\n4. SWA (Nearly useless)\n5. 5-folds cross validation\n\n**5. Some results**\n| One Cell-level model only | active learning for classes | Public LB | Private LB |\n| One Cell-level model only | 11, 15 at 1st round               | 0.446       | 0.465        |\n| One Cell-level model only | 11, 15 at 2nd round             | 0.462        | 0.494       |\n| One Cell-level model only | 11, 15 at 4th round              | 0.502        | -               |\n\n| Ensemble model        | active learning for classes | Public LB | Private LB |\n| Ensemble model        | 11, 15               | 0.558       | 0.528        |\n| Ensemble model        | 9, 11, 12, 15     | 0.558       | 0.532        | -> Final result in LB\n| Ensemble model        | 9, 11, 12, 15 at 4th round, 4 at 1st round | 0.561        | 0.537       | ->Notebook timeout\n\n\n**6. Others**\nTTA is helpful for prediction, with an increase of 0.001~0.003, but it may tend to time out for our submissions.\n\nEarly in the competition, we tried to solve this problem by using multi-instance learning with transformer encoder-based attention, but the effect was not good.\n\nAs the results above, we additionally use a round of active learning for cell label mining on class 4. The submission was timeout and we resubmited it after the deadline of the competition, and it finally obtained the results of Public LB of 0.561 and Private LB of 0.537, indicating that active learning mining labels can continue to improve the performance if it was extended to other classes. However, we adopted this strategy too late..\n\nPlease forgive my typesetting, I encounter problems inserting images and tables...\n",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1308490": "First of all, I would like to thank Kaggle and organizers hosting this really interesting competition. Thanks to my teammates @silversh and @jianguoguo  for the hard work and collaboration. \n\nCongratulations to all the winners. This is my first Kaggle competition, I really learn a lot of nice ideas and solutions in the forum.\n\n**1. Overview**\nOur final solution includes active learning for cell label mining and ensemble of different models.\n\n**2. Active learning for cell label mining**\nBecause I do not participate Kaggle conpetition before, I'm not sure whether data label mining is allowed. Therefore, we try to use active learning to mine cell labels relatively late. Due to time limitation, we only apply it for a few classes, i.e., 9, 11, 12 and 15.\n\nFirst, we chose the cells in the single-label image as the initial trainset, and the cell labels directly inherit the image labels. For classes 11 and 15 with few images, we manually selected about 200 cells with ground truth labels into the initial trainset. After the model is trained on the initial trainset, it is used to make predictions for cells in all images. Use the following strategies to prepare the trainset for the next round.\n\nIf an image contains original label X and the prediction confidence of one cell in this image is greater than 0.25, the label X of the cell is retained, otherwise the cell is removed from the next round trainset. If an image does not contain original label X, and the predictive confidence of one cell in this image is greater than 0.5, this cell will be assigned a true label X. \n\nAfter several rounds, the noise of most cell labels can be inhibited. For example, we finally retrieve about 2,000 cells for class 11 after 4 rounds. On the final public test set, the AP score for class 11 is improved from 0.001 to 0.035.\n\nThe specific principle of active learning can be referred to my previous papers [https://doi.org/10.1016/j.future.2019.07.013](url).\n\n**3. Ensemble of different models**\nWe use ensemble method to fuse Cell-level models and Image-level models. The weights are 0.75 and 0.25, respectively. The Cell-level models employ EfficientNet-B3 network, and the cell score is averaged after 5fold cross-validation. For the Image-level models, we train EfficientNet-B7 networks with RGBY 4-channel image input and Green 1-channel image input, respectively, and the image score is averaged from the two Image-level models. Specifically, when the Cell-level models and the Image-level models are fused, the prediction result of class 11 only use the prediction score of the Cell-level model.\n\n**4. Configuration of training**\n- Image-level model:\n1.  EfficientNet-B7，Concat-pooling + 2 layers of FC Head, input size 600 × 600\n2. Data augmentation: flip, rotation\n3. Focal Loss\n4. SWA (Nearly useless)\n5. Fusion of RGBY 4-channel model and Green 1-channel model\n\n- Cell-level model:\n1.  EfficientNet-B3，Concat-pooling + 2 layers of FC Head, input size 300 × 300\n2. Data augmentation: flip, rotation, mixup\n3. Focal Loss, Label smooth [0.1, 0.9]\n4. SWA (Nearly useless)\n5. 5-folds cross validation\n\n**5. Some results**\n| One Cell-level model only | active learning for classes | Public LB | Private LB |\n| One Cell-level model only | 11, 15 at 1st round               | 0.446       | 0.465        |\n| One Cell-level model only | 11, 15 at 2nd round             | 0.462        | 0.494       |\n| One Cell-level model only | 11, 15 at 4th round              | 0.502        | -               |\n\n| Ensemble model        | active learning for classes | Public LB | Private LB |\n| Ensemble model        | 11, 15               | 0.558       | 0.528        |\n| Ensemble model        | 9, 11, 12, 15     | 0.558       | 0.532        | -> Final result in LB\n| Ensemble model        | 9, 11, 12, 15 at 4th round, 4 at 1st round | 0.561        | 0.537       | ->Notebook timeout\n\n\n**6. Others**\nTTA is helpful for prediction, with an increase of 0.001~0.003, but it may tend to time out for our submissions.\n\nEarly in the competition, we tried to solve this problem by using multi-instance learning with transformer encoder-based attention, but the effect was not good.\n\nAs the results above, we additionally use a round of active learning for cell label mining on class 4. The submission was timeout and we resubmited it after the deadline of the competition, and it finally obtained the results of Public LB of 0.561 and Private LB of 0.537, indicating that active learning mining labels can continue to improve the performance if it was extended to other classes. However, we adopted this strategy too late..\n\nPlease forgive my typesetting, I encounter problems inserting images and tables...\n"
  }
}