{
  "id": 241355,
  "title": "28th solution/ with self-supervised learning",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/241355",
  "author_name": "Sky walker",
  "post_date": "2021-05-24T08:03:16.374000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to the organizers for this very interesting competition, and we have learned a lot during this competition.</p>\n<p>Code link: <a href=\"https://github.com/tuotuo-1997/HPA\" target=\"_blank\">https://github.com/tuotuo-1997/HPA</a></p>\n<p><strong>Method Overview</strong></p>\n<p>We use the Kaggle provided dataset and the public dataset to train and evaluate using different model architectures. The public tools used include Fastai, Opencv, CellSegmentator, Cleanlab, etc.</p>\n<p>We classify 12 models from the whole image and cells. The 8 different model structures used in the cell classification are: Resnet18, 34, 50, 101, Densenet121, Efficientnet B0, B1, and B2. The cell label assignment is consistent with the whole image. The four different model structures used in the whole image are: Efficientnet B0, B1, B2, and B3.<br>\nWe use Adam as the optimizer. For the cell model, each model trains a total of 4 epochs, and for the whole image model, each model trains 6 epochs. The initial learning rate is set to 3e-2. The predicted probabilities of the cells and the whole image are averaged to obtain the final result.</p>\n<p>The improvement of results in our method is mainly concentrated in the following 5 points:<br>\n1) Self-supervised learning (SSL).<br>\n2) The HPA Public data set.<br>\n3) The combine of cells level and whole image level results.<br>\n4) Asymmetric Loss.<br>\n5) Modify the label of the cell to 18 whose maximum value of the green channel is less than 60.<br>\n6) Ensemble.<br>\nThe training time of the entire 12 models on 8*1080Ti is about 48h. It takes about 1.5 hours with a single 2080Ti to submit  to Kaggle (for public test dataset). After the code is submitted, the prediction time of Kaggle kernel is about 9 hours (for all test datasets).</p>\n<p><strong>Conclusions</strong></p>\n<p>It is important to use Asymmetric Loss can alleviate category imbalance while SSL can be helpful to improve the performance. The combine of the cells and the whole image could also greatly improve experiment result.</p>",
  "messages": [
    {
      "id": 1320640,
      "postDate": "2021-05-24T08:03:16.373Z",
      "content": "<p>Thanks to the organizers for this very interesting competition, and we have learned a lot during this competition.</p>\n<p>Code link: <a href=\"https://github.com/tuotuo-1997/HPA\" target=\"_blank\">https://github.com/tuotuo-1997/HPA</a></p>\n<p><strong>Method Overview</strong></p>\n<p>We use the Kaggle provided dataset and the public dataset to train and evaluate using different model architectures. The public tools used include Fastai, Opencv, CellSegmentator, Cleanlab, etc.</p>\n<p>We classify 12 models from the whole image and cells. The 8 different model structures used in the cell classification are: Resnet18, 34, 50, 101, Densenet121, Efficientnet B0, B1, and B2. The cell label assignment is consistent with the whole image. The four different model structures used in the whole image are: Efficientnet B0, B1, B2, and B3.<br>\nWe use Adam as the optimizer. For the cell model, each model trains a total of 4 epochs, and for the whole image model, each model trains 6 epochs. The initial learning rate is set to 3e-2. The predicted probabilities of the cells and the whole image are averaged to obtain the final result.</p>\n<p>The improvement of results in our method is mainly concentrated in the following 5 points:<br>\n1) Self-supervised learning (SSL).<br>\n2) The HPA Public data set.<br>\n3) The combine of cells level and whole image level results.<br>\n4) Asymmetric Loss.<br>\n5) Modify the label of the cell to 18 whose maximum value of the green channel is less than 60.<br>\n6) Ensemble.<br>\nThe training time of the entire 12 models on 8*1080Ti is about 48h. It takes about 1.5 hours with a single 2080Ti to submit  to Kaggle (for public test dataset). After the code is submitted, the prediction time of Kaggle kernel is about 9 hours (for all test datasets).</p>\n<p><strong>Conclusions</strong></p>\n<p>It is important to use Asymmetric Loss can alleviate category imbalance while SSL can be helpful to improve the performance. The combine of the cells and the whole image could also greatly improve experiment result.</p>",
      "rawMarkdown": "Thanks to the organizers for this very interesting competition, and we have learned a lot during this competition.\n\nCode link: https://github.com/tuotuo-1997/HPA\n\n**Method Overview**\n\nWe use the Kaggle provided dataset and the public dataset to train and evaluate using different model architectures. The public tools used include Fastai, Opencv, CellSegmentator, Cleanlab, etc.\n\nWe classify 12 models from the whole image and cells. The 8 different model structures used in the cell classification are: Resnet18, 34, 50, 101, Densenet121, Efficientnet B0, B1, and B2. The cell label assignment is consistent with the whole image. The four different model structures used in the whole image are: Efficientnet B0, B1, B2, and B3.\nWe use Adam as the optimizer. For the cell model, each model trains a total of 4 epochs, and for the whole image model, each model trains 6 epochs. The initial learning rate is set to 3e-2. The predicted probabilities of the cells and the whole image are averaged to obtain the final result.\n\nThe improvement of results in our method is mainly concentrated in the following 5 points:\n1) Self-supervised learning (SSL).\n2) The HPA Public data set.\n3) The combine of cells level and whole image level results.\n4) Asymmetric Loss.\n5) Modify the label of the cell to 18 whose maximum value of the green channel is less than 60.\n6) Ensemble.\nThe training time of the entire 12 models on 8*1080Ti is about 48h. It takes about 1.5 hours with a single 2080Ti to submit  to Kaggle (for public test dataset). After the code is submitted, the prediction time of Kaggle kernel is about 9 hours (for all test datasets).\n\n**Conclusions**\n\nIt is important to use Asymmetric Loss can alleviate category imbalance while SSL can be helpful to improve the performance. The combine of the cells and the whole image could also greatly improve experiment result.",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1320640": "Thanks to the organizers for this very interesting competition, and we have learned a lot during this competition.\n\nCode link: https://github.com/tuotuo-1997/HPA\n\n**Method Overview**\n\nWe use the Kaggle provided dataset and the public dataset to train and evaluate using different model architectures. The public tools used include Fastai, Opencv, CellSegmentator, Cleanlab, etc.\n\nWe classify 12 models from the whole image and cells. The 8 different model structures used in the cell classification are: Resnet18, 34, 50, 101, Densenet121, Efficientnet B0, B1, and B2. The cell label assignment is consistent with the whole image. The four different model structures used in the whole image are: Efficientnet B0, B1, B2, and B3.\nWe use Adam as the optimizer. For the cell model, each model trains a total of 4 epochs, and for the whole image model, each model trains 6 epochs. The initial learning rate is set to 3e-2. The predicted probabilities of the cells and the whole image are averaged to obtain the final result.\n\nThe improvement of results in our method is mainly concentrated in the following 5 points:\n1) Self-supervised learning (SSL).\n2) The HPA Public data set.\n3) The combine of cells level and whole image level results.\n4) Asymmetric Loss.\n5) Modify the label of the cell to 18 whose maximum value of the green channel is less than 60.\n6) Ensemble.\nThe training time of the entire 12 models on 8*1080Ti is about 48h. It takes about 1.5 hours with a single 2080Ti to submit  to Kaggle (for public test dataset). After the code is submitted, the prediction time of Kaggle kernel is about 9 hours (for all test datasets).\n\n**Conclusions**\n\nIt is important to use Asymmetric Loss can alleviate category imbalance while SSL can be helpful to improve the performance. The combine of the cells and the whole image could also greatly improve experiment result."
  }
}