{
  "id": 238387,
  "title": "Private 22nd Place Solution",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238387",
  "author_name": "cool_rabbit",
  "post_date": "2021-05-12T04:29:34.679000",
  "votes": 31,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Congratulations to all the winners, and thanks so much for hosting such an interesting competition!!<br>\nThis task was really challenging in mostly two points: weak-labels and class imbalance.<br>\nI spent hard time on solving them, and learned a lot in the middle of it.</p>\n<h1>Summary</h1>\n<p>・I tackled this competition as a classification task (I didn't use any segmentation models other than HPA Cell Segmentator).<br>\n・Environment: Kaggle Notebook and Datasets, TPU training, GPU inference, PyTorch<br>\n・Cell tiles: 'nucleus BBox center' chosen as tile center, 'cell BBox short side' chosen as tile one side length→score improved!!<br>\n・2-Stage Training Pipeline (For 2nd stage, pseudo-labels, thresholding and sampling methods were used.)<br>\n・Green Image Level Label prediction further added, shared by <a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> </p>\n<h1>Training</h1>\n<p>My pipeline is the following.<br>\nCV: multilabel stratified group kfold (group by image id)<br>\naugmentation: flip, random rotate, shift scale rotate<br>\nloss: BCEWithLogitsLoss<br>\noptimizer: Adam<br>\nscheduler: cosine annealing<br>\nnumber of cell tiles used as input: about 70000 (1st stage), about 75000 (2nd stage)<br>\nepochs: 5eps w/o early stopping<br>\ntraining time: 1~2 hrs per model</p>\n<p><img alt=\"Screen Shot 2021-05-12 at 12 54 37\" src=\"https://user-images.githubusercontent.com/63890401/117921313-7f8f5e80-b32b-11eb-85cd-aa79dee1314d.png\"></p>\n<h1>Inference</h1>\n<p>I used <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> fast segmentator with a little modified.<br>\nClassification predictions by model above was combined with segmentator instance segmentation result.<br>\n<a href=\"https://www.kaggle.com/drtausamaru/hpa-ct-ill-inference-private\" target=\"_blank\">https://www.kaggle.com/drtausamaru/hpa-ct-ill-inference-private</a><br>\n<img alt=\"Screen Shot 2021-05-12 at 12 48 40\" src=\"https://user-images.githubusercontent.com/63890401/117918045-4b18a400-b325-11eb-84b6-b53930e576b1.png\"></p>\n<h1>What didn't work for me (score dropped)</h1>\n<p>・Cell tiles other than my approach (whole cell tiles, cell BBox long side length, conversion to all values = 0 of the area outside the targeted cell, only green signal used…etc)<br>\n・focal loss<br>\n・BCEWithLogitsLoss with pos-weight argument &gt; 1.0<br>\n・label smoothing<br>\n・pseudo hard labels (0/1)<br>\n・For 1st-stage, using &gt;=4 cell tiles at once<br>\n・MLSKF (not grouped)<br>\n・Models: ResNet200D, SEResNeXt101<br>\n・cellline classification model</p>\n<h1>What I didn't try</h1>\n<p>・RGBY<br>\n・cell tiles: &gt;256x256, uint16 <br>\n・other segmentation models (object detection)<br>\n・other augmentations (mixup, brightness modification, cutout…etc)</p>\n<h1>In the end</h1>\n<p>I really enjoyed this competition because the task itself is interesting and challenging, there were a few public kernels of just ensembling or forking, and I was able to compete with top Kagglers.<br>\nThanks for reading :)</p>",
  "messages": [
    {
      "id": 1303444,
      "postDate": "2021-05-12T04:29:34.680Z",
      "content": "<p>Congratulations to all the winners, and thanks so much for hosting such an interesting competition!!<br>\nThis task was really challenging in mostly two points: weak-labels and class imbalance.<br>\nI spent hard time on solving them, and learned a lot in the middle of it.</p>\n<h1>Summary</h1>\n<p>・I tackled this competition as a classification task (I didn't use any segmentation models other than HPA Cell Segmentator).<br>\n・Environment: Kaggle Notebook and Datasets, TPU training, GPU inference, PyTorch<br>\n・Cell tiles: 'nucleus BBox center' chosen as tile center, 'cell BBox short side' chosen as tile one side length→score improved!!<br>\n・2-Stage Training Pipeline (For 2nd stage, pseudo-labels, thresholding and sampling methods were used.)<br>\n・Green Image Level Label prediction further added, shared by <a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> </p>\n<h1>Training</h1>\n<p>My pipeline is the following.<br>\nCV: multilabel stratified group kfold (group by image id)<br>\naugmentation: flip, random rotate, shift scale rotate<br>\nloss: BCEWithLogitsLoss<br>\noptimizer: Adam<br>\nscheduler: cosine annealing<br>\nnumber of cell tiles used as input: about 70000 (1st stage), about 75000 (2nd stage)<br>\nepochs: 5eps w/o early stopping<br>\ntraining time: 1~2 hrs per model</p>\n<p><img alt=\"Screen Shot 2021-05-12 at 12 54 37\" src=\"https://user-images.githubusercontent.com/63890401/117921313-7f8f5e80-b32b-11eb-85cd-aa79dee1314d.png\"></p>\n<h1>Inference</h1>\n<p>I used <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> fast segmentator with a little modified.<br>\nClassification predictions by model above was combined with segmentator instance segmentation result.<br>\n<a href=\"https://www.kaggle.com/drtausamaru/hpa-ct-ill-inference-private\" target=\"_blank\">https://www.kaggle.com/drtausamaru/hpa-ct-ill-inference-private</a><br>\n<img alt=\"Screen Shot 2021-05-12 at 12 48 40\" src=\"https://user-images.githubusercontent.com/63890401/117918045-4b18a400-b325-11eb-84b6-b53930e576b1.png\"></p>\n<h1>What didn't work for me (score dropped)</h1>\n<p>・Cell tiles other than my approach (whole cell tiles, cell BBox long side length, conversion to all values = 0 of the area outside the targeted cell, only green signal used…etc)<br>\n・focal loss<br>\n・BCEWithLogitsLoss with pos-weight argument &gt; 1.0<br>\n・label smoothing<br>\n・pseudo hard labels (0/1)<br>\n・For 1st-stage, using &gt;=4 cell tiles at once<br>\n・MLSKF (not grouped)<br>\n・Models: ResNet200D, SEResNeXt101<br>\n・cellline classification model</p>\n<h1>What I didn't try</h1>\n<p>・RGBY<br>\n・cell tiles: &gt;256x256, uint16 <br>\n・other segmentation models (object detection)<br>\n・other augmentations (mixup, brightness modification, cutout…etc)</p>\n<h1>In the end</h1>\n<p>I really enjoyed this competition because the task itself is interesting and challenging, there were a few public kernels of just ensembling or forking, and I was able to compete with top Kagglers.<br>\nThanks for reading :)</p>",
      "rawMarkdown": "Congratulations to all the winners, and thanks so much for hosting such an interesting competition!!\nThis task was really challenging in mostly two points: weak-labels and class imbalance.\nI spent hard time on solving them, and learned a lot in the middle of it.\n\n\n\n# Summary\n・I tackled this competition as a classification task (I didn't use any segmentation models other than HPA Cell Segmentator).\n・Environment: Kaggle Notebook and Datasets, TPU training, GPU inference, PyTorch\n・Cell tiles: 'nucleus BBox center' chosen as tile center, 'cell BBox short side' chosen as tile one side length→score improved!!\n・2-Stage Training Pipeline (For 2nd stage, pseudo-labels, thresholding and sampling methods were used.)\n・Green Image Level Label prediction further added, shared by @h053473666 \n\n\n\n# Training\nMy pipeline is the following.\nCV: multilabel stratified group kfold (group by image id)\naugmentation: flip, random rotate, shift scale rotate\nloss: BCEWithLogitsLoss\noptimizer: Adam\nscheduler: cosine annealing\nnumber of cell tiles used as input: about 70000 (1st stage), about 75000 (2nd stage)\nepochs: 5eps w/o early stopping\ntraining time: 1~2 hrs per model\n\n<img width=\"1045\" alt=\"Screen Shot 2021-05-12 at 12 54 37\" src=\"https://user-images.githubusercontent.com/63890401/117921313-7f8f5e80-b32b-11eb-85cd-aa79dee1314d.png\">\n\n\n\n# Inference\nI used @samusram fast segmentator with a little modified.\nClassification predictions by model above was combined with segmentator instance segmentation result.\nhttps://www.kaggle.com/drtausamaru/hpa-ct-ill-inference-private\n<img width=\"987\" alt=\"Screen Shot 2021-05-12 at 12 48 40\" src=\"https://user-images.githubusercontent.com/63890401/117918045-4b18a400-b325-11eb-84b6-b53930e576b1.png\">\n\n\n\n# What didn't work for me (score dropped)\n・Cell tiles other than my approach (whole cell tiles, cell BBox long side length, conversion to all values = 0 of the area outside the targeted cell, only green signal used...etc)\n・focal loss\n・BCEWithLogitsLoss with pos-weight argument > 1.0\n・label smoothing\n・pseudo hard labels (0/1)\n・For 1st-stage, using >=4 cell tiles at once\n・MLSKF (not grouped)\n・Models: ResNet200D, SEResNeXt101\n・cellline classification model\n\n\n\n# What I didn't try\n・RGBY\n・cell tiles: >256x256, uint16 \n・other segmentation models (object detection)\n・other augmentations (mixup, brightness modification, cutout...etc)\n\n\n\n# In the end\nI really enjoyed this competition because the task itself is interesting and challenging, there were a few public kernels of just ensembling or forking, and I was able to compete with top Kagglers.\nThanks for reading :)",
      "votes": 31
    },
    {
      "id": 1305037,
      "postDate": "2021-05-13T04:31:00.293Z",
      "content": "<p><a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> Congratulations and thanks for sharing the approach </p>",
      "rawMarkdown": "@drtausamaru Congratulations and thanks for sharing the approach ",
      "votes": 1
    },
    {
      "id": 1304414,
      "postDate": "2021-05-12T15:53:04.730Z",
      "content": "<p>Congrats! How does the NFNet single model perform?</p>",
      "rawMarkdown": "Congrats! How does the NFNet single model perform?",
      "votes": 1,
      "replies": [
        {
          "id": 1304778,
          "postDate": "2021-05-12T21:41:38.080Z",
          "content": "<p>Thank you!<br>\nI didn't submit only NFNet inference, but LB score increased by 0.003 with it for ensemble.</p>",
          "rawMarkdown": "Thank you!\nI didn't submit only NFNet inference, but LB score increased by 0.003 with it for ensemble.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1304268,
      "postDate": "2021-05-12T14:16:13.743Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> ! Your solo performance has been amazing, a true warrior :) </p>",
      "rawMarkdown": "Congratulations @drtausamaru ! Your solo performance has been amazing, a true warrior :) ",
      "votes": 1,
      "replies": [
        {
          "id": 1304274,
          "postDate": "2021-05-12T14:21:14.910Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a><br>\nHope to work together in the future competition soon.</p>",
          "rawMarkdown": "Thank you @thedrcat\nHope to work together in the future competition soon.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1304043,
      "postDate": "2021-05-12T11:48:05.450Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> on 22th place. Thanks for sharing solution </p>",
      "rawMarkdown": "Congrats @drtausamaru on 22th place. Thanks for sharing solution ",
      "votes": 1
    },
    {
      "id": 1303899,
      "postDate": "2021-05-12T10:27:21.127Z",
      "content": "<p>Great job! Awesome write up! I love reading all of the solutions.</p>",
      "rawMarkdown": "Great job! Awesome write up! I love reading all of the solutions.",
      "votes": 1
    },
    {
      "id": 1303684,
      "postDate": "2021-05-12T07:39:39.797Z",
      "content": "<p>Thanks for sharing, did you tried other backbones ? </p>",
      "rawMarkdown": "Thanks for sharing, did you tried other backbones ? ",
      "votes": 1,
      "replies": [
        {
          "id": 1303715,
          "postDate": "2021-05-12T07:53:22.347Z",
          "content": "<p>You mean, model backbone or pipeline?</p>",
          "rawMarkdown": "You mean, model backbone or pipeline?",
          "votes": 1
        },
        {
          "id": 1305169,
          "postDate": "2021-05-13T06:07:31.983Z",
          "content": "<p>I mean Model backbone </p>",
          "rawMarkdown": "I mean Model backbone "
        },
        {
          "id": 1305174,
          "postDate": "2021-05-13T06:12:09.537Z",
          "content": "<p>I tried the models above plus ResNet200D and SEResNeXt101.<br>\nThose didn’t help the score in my case.</p>",
          "rawMarkdown": "I tried the models above plus ResNet200D and SEResNeXt101.\nThose didn’t help the score in my case."
        }
      ]
    },
    {
      "id": 1303480,
      "postDate": "2021-05-12T04:57:24.293Z",
      "content": "<p>Very interesting approach, and congrats on the silver. I am curious about a few things. While creating pseudo labels, did you take the average of image level label and prediction from your model or only the prediction from your model? Also, what did you use to train TPU models? Do you have any recommended python packages to start off with? </p>",
      "rawMarkdown": "Very interesting approach, and congrats on the silver. I am curious about a few things. While creating pseudo labels, did you take the average of image level label and prediction from your model or only the prediction from your model? Also, what did you use to train TPU models? Do you have any recommended python packages to start off with? ",
      "votes": 1,
      "replies": [
        {
          "id": 1303489,
          "postDate": "2021-05-12T05:03:44.447Z",
          "content": "<p>Thanks for the comment.<br>\nPseudo labels are created by averaging only cell-level models as shown in the solution image above.<br>\nI used XLA for TPU.</p>",
          "rawMarkdown": "Thanks for the comment.\nPseudo labels are created by averaging only cell-level models as shown in the solution image above.\nI used XLA for TPU.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1303450,
      "postDate": "2021-05-12T04:35:14.730Z",
      "content": "<p>Congrats, <a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> ! 😊</p>",
      "rawMarkdown": "Congrats, @drtausamaru ! 😊",
      "votes": 1
    },
    {
      "id": 1304148,
      "postDate": "2021-05-12T12:57:46.857Z",
      "content": "<p>Thanks for sharing with figures!</p>",
      "rawMarkdown": "Thanks for sharing with figures!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1305037,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:31:00.293000",
      "content": "<p><a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> Congratulations and thanks for sharing the approach </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1304414,
      "author_name": "Alien",
      "author_url": "",
      "post_date": "2021-05-12T15:53:04.730000",
      "content": "<p>Congrats! How does the NFNet single model perform?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304778,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-12T21:41:38.080000",
          "content": "<p>Thank you!<br>\nI didn't submit only NFNet inference, but LB score increased by 0.003 with it for ensemble.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1304268,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-05-12T14:16:13.743000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> ! Your solo performance has been amazing, a true warrior :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304274,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-12T14:21:14.910000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a><br>\nHope to work together in the future competition soon.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1304043,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-05-12T11:48:05.450000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> on 22th place. Thanks for sharing solution </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1303899,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-05-12T10:27:21.127000",
      "content": "<p>Great job! Awesome write up! I love reading all of the solutions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1303684,
      "author_name": "Salim Khazem",
      "author_url": "",
      "post_date": "2021-05-12T07:39:39.797000",
      "content": "<p>Thanks for sharing, did you tried other backbones ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1303715,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-12T07:53:22.347000",
          "content": "<p>You mean, model backbone or pipeline?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1305169,
          "author_name": "Salim Khazem",
          "author_url": "",
          "post_date": "2021-05-13T06:07:31.983000",
          "content": "<p>I mean Model backbone </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1305174,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-13T06:12:09.537000",
          "content": "<p>I tried the models above plus ResNet200D and SEResNeXt101.<br>\nThose didn’t help the score in my case.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1303480,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "2021-05-12T04:57:24.293000",
      "content": "<p>Very interesting approach, and congrats on the silver. I am curious about a few things. While creating pseudo labels, did you take the average of image level label and prediction from your model or only the prediction from your model? Also, what did you use to train TPU models? Do you have any recommended python packages to start off with? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1303489,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-12T05:03:44.447000",
          "content": "<p>Thanks for the comment.<br>\nPseudo labels are created by averaging only cell-level models as shown in the solution image above.<br>\nI used XLA for TPU.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1303450,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-12T04:35:14.730000",
      "content": "<p>Congrats, <a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> ! 😊</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1304148,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-12T12:57:46.857000",
      "content": "<p>Thanks for sharing with figures!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1303444": "Congratulations to all the winners, and thanks so much for hosting such an interesting competition!!\nThis task was really challenging in mostly two points: weak-labels and class imbalance.\nI spent hard time on solving them, and learned a lot in the middle of it.\n\n\n\n# Summary\n・I tackled this competition as a classification task (I didn't use any segmentation models other than HPA Cell Segmentator).\n・Environment: Kaggle Notebook and Datasets, TPU training, GPU inference, PyTorch\n・Cell tiles: 'nucleus BBox center' chosen as tile center, 'cell BBox short side' chosen as tile one side length→score improved!!\n・2-Stage Training Pipeline (For 2nd stage, pseudo-labels, thresholding and sampling methods were used.)\n・Green Image Level Label prediction further added, shared by @h053473666 \n\n\n\n# Training\nMy pipeline is the following.\nCV: multilabel stratified group kfold (group by image id)\naugmentation: flip, random rotate, shift scale rotate\nloss: BCEWithLogitsLoss\noptimizer: Adam\nscheduler: cosine annealing\nnumber of cell tiles used as input: about 70000 (1st stage), about 75000 (2nd stage)\nepochs: 5eps w/o early stopping\ntraining time: 1~2 hrs per model\n\n<img width=\"1045\" alt=\"Screen Shot 2021-05-12 at 12 54 37\" src=\"https://user-images.githubusercontent.com/63890401/117921313-7f8f5e80-b32b-11eb-85cd-aa79dee1314d.png\">\n\n\n\n# Inference\nI used @samusram fast segmentator with a little modified.\nClassification predictions by model above was combined with segmentator instance segmentation result.\nhttps://www.kaggle.com/drtausamaru/hpa-ct-ill-inference-private\n<img width=\"987\" alt=\"Screen Shot 2021-05-12 at 12 48 40\" src=\"https://user-images.githubusercontent.com/63890401/117918045-4b18a400-b325-11eb-84b6-b53930e576b1.png\">\n\n\n\n# What didn't work for me (score dropped)\n・Cell tiles other than my approach (whole cell tiles, cell BBox long side length, conversion to all values = 0 of the area outside the targeted cell, only green signal used...etc)\n・focal loss\n・BCEWithLogitsLoss with pos-weight argument > 1.0\n・label smoothing\n・pseudo hard labels (0/1)\n・For 1st-stage, using >=4 cell tiles at once\n・MLSKF (not grouped)\n・Models: ResNet200D, SEResNeXt101\n・cellline classification model\n\n\n\n# What I didn't try\n・RGBY\n・cell tiles: >256x256, uint16 \n・other segmentation models (object detection)\n・other augmentations (mixup, brightness modification, cutout...etc)\n\n\n\n# In the end\nI really enjoyed this competition because the task itself is interesting and challenging, there were a few public kernels of just ensembling or forking, and I was able to compete with top Kagglers.\nThanks for reading :)",
    "1305037": "@drtausamaru Congratulations and thanks for sharing the approach ",
    "1304414": "Congrats! How does the NFNet single model perform?",
    "1304268": "Congratulations @drtausamaru ! Your solo performance has been amazing, a true warrior :) ",
    "1304043": "Congrats @drtausamaru on 22th place. Thanks for sharing solution ",
    "1303899": "Great job! Awesome write up! I love reading all of the solutions.",
    "1303684": "Thanks for sharing, did you tried other backbones ? ",
    "1303480": "Very interesting approach, and congrats on the silver. I am curious about a few things. While creating pseudo labels, did you take the average of image level label and prediction from your model or only the prediction from your model? Also, what did you use to train TPU models? Do you have any recommended python packages to start off with? ",
    "1303450": "Congrats, @drtausamaru ! 😊",
    "1304148": "Thanks for sharing with figures!"
  }
}