{
  "id": 298146,
  "title": "4th place solution",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/298146",
  "author_name": "Yamame🐟",
  "post_date": "2022-01-01T02:48:34.687000",
  "votes": 61,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thanks to Kaggle and Sartorius for this interesting competition. I also thank my teammates, <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> and <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>, <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>. I learned a lot from them.</p>\n<p>Our experiments are based on CBNetV2 <a href=\"https://github.com/VDIGPKU/CBNetV2\" target=\"_blank\">repo</a></p>\n<p>In the following, I want to give a summary of our solution.</p>\n<h1>Overall</h1>\n<p>Our solution consists of three parts: classification part, instance segmentation part, post-processing part. Each cell type has different instance segmentation models and post-processing, so classification is necessary.<br>\n<a href=\"https://postimg.cc/jnNJBNts\" target=\"_blank\"><img src=\"https://i.postimg.cc/RhpQ6QYn/overall.png\" alt=\"overall.png\"></a></p>\n<h1>Classification</h1>\n<p>3class CBNet DBS Cascade-RCNN  (num_classes=3) was used as a classification model. For a given image, this model classifies the image into the cell type with the highest number of detected cells. <br>\nThis model is first pre-trained with LIVE_CELL (num_classes=1) and then fine-tuned with the competition's train data (num_classes=3).<br>\nI tested this on semi-supervised data and it was able to classify them perfectly.</p>\n<h1>Instance Segmentation</h1>\n<p>We prepared at least one model for each of shsy5y, astro, and cort. These models were trained in different settings and with different data. </p>\n<h2>Pseudo Labeling</h2>\n<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> <br>\nFor all cell types, pseudo labeling on semi-supervised data improved both CV and LB.<br>\nBecause of the small number of train data, it was better to use three types of cells when using only train data. However, because of the large number of semi-supervised data, we were able to change the data used for each cell type.<br>\n<a href=\"https://drive.google.com/file/d/1PSIxdNiwtMwTV3Wq6AJbWcp2plnoIVQA/view?usp=sharing\" target=\"_blank\">Pseudo labeling</a></p>\n<h3>Models for pseudo labeling</h3>\n<p>Data: live_cell → train (all cell types)<br>\nModel: 3 types of 1class CBNet DBS Cascade-RCNN<br>\nWe changed the MMdet config file depending on the target cell type.</p>\n<p><a href=\"https://postimg.cc/DW2dsCF1\" target=\"_blank\"><img src=\"https://i.postimg.cc/W4Z9BKHY/train-strategy.png\" alt=\"train-strategy.png\"></a></p>\n<h2>shsy5y models</h2>\n<p><a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> <br>\nData1: live_cell → train (all cell types)<br>\nModel1: CBNet DBS Cascade-RCNN<br>\nData2: live_cell → semi-sup+train (shsy5y &amp; cort) → train (shsy5y)<br>\nModel2: CBNet DBS Cascade-RCNN<br>\nThe number of cells in shsy5y and cort are very different, but the individual cells are similar, so, we used the two as training data when using semi-sup data.</p>\n<h2>astro models</h2>\n<p><a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> <br>\nData: live_cell → semi-sup+train (astro) → train (astro)<br>\nSince astro has a very different shape and size from the other cells, we improved the score by using only astro data fot train data.<br>\nIn astro, I only used one model because the ensemble did not work well.</p>\n<h2>cort models</h2>\n<p><a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <br>\nData: live_cell → semi-sup+train (shsy5y &amp; cort) → train (court)<br>\nModel: 2 x HTC resnext64x4d, 2 x CBNet DBS Cascade-RCNN<br>\nThese four models were combined into one model using the method below.</p>\n<h3>ensemble detection model</h3>\n<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> <br>\nWe only used the two-stage model. So, we used an ensemble method like <a href=\"https://github.com/amirassov/kaggle-imaterialist\" target=\"_blank\">this</a>, where RPNs are connected and treated like a single two-stage model.<br>\nThis improved cort score.<br>\n<a href=\"https://postimg.cc/pmvDb0m3\" target=\"_blank\"><img src=\"https://i.postimg.cc/vZgXnkb8/cort-models.png\" alt=\"cort-models.png\"></a></p>\n<h2>Augmentation</h2>\n<p>Multiscale: astro and cort → (1280, 1280)~(1792, 1792), shsy5y → (1280, 1280)~(1536, 1536)<br>\nHorizontal Flip<br>\nVertical Flip</p>\n<h2>TTA</h2>\n<p>Multiscale + horizontal flip (vertical flip and diagonal flip didn't work)</p>\n<h1>Post-Processing</h1>\n<h2>astro pp</h2>\n<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> <br>\nThe astro annotation was broken, so I reproduced it with cv2.findContours and cv2.fillConvexPoly.</p>\n<h2>fix-overlap</h2>\n<p><a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> <br>\nIn this competition, predictions are not allowed to overlap. So, we have to eliminate overlap part in post-processing. <br>\nIn our fix-overlap process, first we took cells with a higher cofidence score than classwise threshold. Then, we processed from the highest score to the lowest, and deleted instances with a large percentage of already-used area. In addition, only shsy5y score was improved by removing those with pixels lower than threshold.</p>\n<h2>semantic re-lank</h2>\n<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> <br>\nAs shown in the Refinemask <a href=\"https://arxiv.org/abs/2104.08569\" target=\"_blank\">paper</a>, the confidence score of the instance segmentation model does not reflect the correctness of the mask. Therefore, we modified the scores and re-lank the instances with semantic segmentation model (UNet++).</p>\n<h2>fix-overlap ensemble</h2>\n<p><a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> <br>\nIn this ensemble method, we first concat the output of multiple models and sort them by their semantic relank scores. Next, we applied fix-overlap pp to remove the overlap. This improved shsy5y score.<br>\n<a href=\"https://postimg.cc/FfRkNwNC\" target=\"_blank\"><img src=\"https://i.postimg.cc/R017Qz8m/fixoverlap-ensemble.png\" alt=\"fixoverlap-ensemble.png\"></a></p>\n<h2>WBF with mask</h2>\n<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> <br>\nWe tried WBF extended for mask. This improved Public LB score, but 'ensemble detection model' is better for Private LB.</p>\n<h1>Tips</h1>\n<p>The default config file of MMDetection is fitted for COCO. Depending on the shape of the instances and the number of instances in an image, it is necessary to change the settings for training.<br>\ne.g. anchor_generator.ratios, rpn_proposal.nms_pre, and so on…</p>",
  "messages": [
    {
      "id": 1634688,
      "postDate": "2022-01-01T02:48:34.687Z",
      "content": "<p>Thanks to Kaggle and Sartorius for this interesting competition. I also thank my teammates, <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> and <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a>, <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>. I learned a lot from them.</p>\n<p>Our experiments are based on CBNetV2 <a href=\"https://github.com/VDIGPKU/CBNetV2\" target=\"_blank\">repo</a></p>\n<p>In the following, I want to give a summary of our solution.</p>\n<h1>Overall</h1>\n<p>Our solution consists of three parts: classification part, instance segmentation part, post-processing part. Each cell type has different instance segmentation models and post-processing, so classification is necessary.<br>\n<a href=\"https://postimg.cc/jnNJBNts\" target=\"_blank\"><img src=\"https://i.postimg.cc/RhpQ6QYn/overall.png\" alt=\"overall.png\"></a></p>\n<h1>Classification</h1>\n<p>3class CBNet DBS Cascade-RCNN  (num_classes=3) was used as a classification model. For a given image, this model classifies the image into the cell type with the highest number of detected cells. <br>\nThis model is first pre-trained with LIVE_CELL (num_classes=1) and then fine-tuned with the competition's train data (num_classes=3).<br>\nI tested this on semi-supervised data and it was able to classify them perfectly.</p>\n<h1>Instance Segmentation</h1>\n<p>We prepared at least one model for each of shsy5y, astro, and cort. These models were trained in different settings and with different data. </p>\n<h2>Pseudo Labeling</h2>\n<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> <br>\nFor all cell types, pseudo labeling on semi-supervised data improved both CV and LB.<br>\nBecause of the small number of train data, it was better to use three types of cells when using only train data. However, because of the large number of semi-supervised data, we were able to change the data used for each cell type.<br>\n<a href=\"https://drive.google.com/file/d/1PSIxdNiwtMwTV3Wq6AJbWcp2plnoIVQA/view?usp=sharing\" target=\"_blank\">Pseudo labeling</a></p>\n<h3>Models for pseudo labeling</h3>\n<p>Data: live_cell → train (all cell types)<br>\nModel: 3 types of 1class CBNet DBS Cascade-RCNN<br>\nWe changed the MMdet config file depending on the target cell type.</p>\n<p><a href=\"https://postimg.cc/DW2dsCF1\" target=\"_blank\"><img src=\"https://i.postimg.cc/W4Z9BKHY/train-strategy.png\" alt=\"train-strategy.png\"></a></p>\n<h2>shsy5y models</h2>\n<p><a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> <br>\nData1: live_cell → train (all cell types)<br>\nModel1: CBNet DBS Cascade-RCNN<br>\nData2: live_cell → semi-sup+train (shsy5y &amp; cort) → train (shsy5y)<br>\nModel2: CBNet DBS Cascade-RCNN<br>\nThe number of cells in shsy5y and cort are very different, but the individual cells are similar, so, we used the two as training data when using semi-sup data.</p>\n<h2>astro models</h2>\n<p><a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> <br>\nData: live_cell → semi-sup+train (astro) → train (astro)<br>\nSince astro has a very different shape and size from the other cells, we improved the score by using only astro data fot train data.<br>\nIn astro, I only used one model because the ensemble did not work well.</p>\n<h2>cort models</h2>\n<p><a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <br>\nData: live_cell → semi-sup+train (shsy5y &amp; cort) → train (court)<br>\nModel: 2 x HTC resnext64x4d, 2 x CBNet DBS Cascade-RCNN<br>\nThese four models were combined into one model using the method below.</p>\n<h3>ensemble detection model</h3>\n<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> <br>\nWe only used the two-stage model. So, we used an ensemble method like <a href=\"https://github.com/amirassov/kaggle-imaterialist\" target=\"_blank\">this</a>, where RPNs are connected and treated like a single two-stage model.<br>\nThis improved cort score.<br>\n<a href=\"https://postimg.cc/pmvDb0m3\" target=\"_blank\"><img src=\"https://i.postimg.cc/vZgXnkb8/cort-models.png\" alt=\"cort-models.png\"></a></p>\n<h2>Augmentation</h2>\n<p>Multiscale: astro and cort → (1280, 1280)~(1792, 1792), shsy5y → (1280, 1280)~(1536, 1536)<br>\nHorizontal Flip<br>\nVertical Flip</p>\n<h2>TTA</h2>\n<p>Multiscale + horizontal flip (vertical flip and diagonal flip didn't work)</p>\n<h1>Post-Processing</h1>\n<h2>astro pp</h2>\n<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> <br>\nThe astro annotation was broken, so I reproduced it with cv2.findContours and cv2.fillConvexPoly.</p>\n<h2>fix-overlap</h2>\n<p><a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> <br>\nIn this competition, predictions are not allowed to overlap. So, we have to eliminate overlap part in post-processing. <br>\nIn our fix-overlap process, first we took cells with a higher cofidence score than classwise threshold. Then, we processed from the highest score to the lowest, and deleted instances with a large percentage of already-used area. In addition, only shsy5y score was improved by removing those with pixels lower than threshold.</p>\n<h2>semantic re-lank</h2>\n<p><a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> <br>\nAs shown in the Refinemask <a href=\"https://arxiv.org/abs/2104.08569\" target=\"_blank\">paper</a>, the confidence score of the instance segmentation model does not reflect the correctness of the mask. Therefore, we modified the scores and re-lank the instances with semantic segmentation model (UNet++).</p>\n<h2>fix-overlap ensemble</h2>\n<p><a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> <a href=\"https://www.kaggle.com/tanakar\" target=\"_blank\">@tanakar</a> <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> <br>\nIn this ensemble method, we first concat the output of multiple models and sort them by their semantic relank scores. Next, we applied fix-overlap pp to remove the overlap. This improved shsy5y score.<br>\n<a href=\"https://postimg.cc/FfRkNwNC\" target=\"_blank\"><img src=\"https://i.postimg.cc/R017Qz8m/fixoverlap-ensemble.png\" alt=\"fixoverlap-ensemble.png\"></a></p>\n<h2>WBF with mask</h2>\n<p><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> <br>\nWe tried WBF extended for mask. This improved Public LB score, but 'ensemble detection model' is better for Private LB.</p>\n<h1>Tips</h1>\n<p>The default config file of MMDetection is fitted for COCO. Depending on the shape of the instances and the number of instances in an image, it is necessary to change the settings for training.<br>\ne.g. anchor_generator.ratios, rpn_proposal.nms_pre, and so on…</p>",
      "rawMarkdown": "Thanks to Kaggle and Sartorius for this interesting competition. I also thank my teammates, @tanakar and @tereka, @ren4yu. I learned a lot from them.\n\nOur experiments are based on CBNetV2 [repo](https://github.com/VDIGPKU/CBNetV2)\n\nIn the following, I want to give a summary of our solution.\n\n# Overall\nOur solution consists of three parts: classification part, instance segmentation part, post-processing part. Each cell type has different instance segmentation models and post-processing, so classification is necessary.\n[![overall.png](https://i.postimg.cc/RhpQ6QYn/overall.png)](https://postimg.cc/jnNJBNts)\n\n# Classification\n3class CBNet DBS Cascade-RCNN  (num_classes=3) was used as a classification model. For a given image, this model classifies the image into the cell type with the highest number of detected cells. \nThis model is first pre-trained with LIVE_CELL (num_classes=1) and then fine-tuned with the competition's train data (num_classes=3).\nI tested this on semi-supervised data and it was able to classify them perfectly.\n\n# Instance Segmentation\nWe prepared at least one model for each of shsy5y, astro, and cort. These models were trained in different settings and with different data. \n## Pseudo Labeling\n@tereka \nFor all cell types, pseudo labeling on semi-supervised data improved both CV and LB.\nBecause of the small number of train data, it was better to use three types of cells when using only train data. However, because of the large number of semi-supervised data, we were able to change the data used for each cell type.\n[Pseudo labeling](https://drive.google.com/file/d/1PSIxdNiwtMwTV3Wq6AJbWcp2plnoIVQA/view?usp=sharing)\n### Models for pseudo labeling\nData: live_cell → train (all cell types)\nModel: 3 types of 1class CBNet DBS Cascade-RCNN\nWe changed the MMdet config file depending on the target cell type.\n\n[![train-strategy.png](https://i.postimg.cc/W4Z9BKHY/train-strategy.png)](https://postimg.cc/DW2dsCF1)\n## shsy5y models\n@tyaiga \nData1: live_cell → train (all cell types)\nModel1: CBNet DBS Cascade-RCNN\nData2: live_cell → semi-sup+train (shsy5y & cort) → train (shsy5y)\nModel2: CBNet DBS Cascade-RCNN\nThe number of cells in shsy5y and cort are very different, but the individual cells are similar, so, we used the two as training data when using semi-sup data.\n\n## astro models\n@tyaiga \nData: live_cell → semi-sup+train (astro) → train (astro)\nSince astro has a very different shape and size from the other cells, we improved the score by using only astro data fot train data.\nIn astro, I only used one model because the ensemble did not work well.\n\n## cort models\n@tanakar \nData: live_cell → semi-sup+train (shsy5y & cort) → train (court)\nModel: 2 x HTC resnext64x4d, 2 x CBNet DBS Cascade-RCNN\nThese four models were combined into one model using the method below.\n\n### ensemble detection model\n@tereka @ren4yu \nWe only used the two-stage model. So, we used an ensemble method like [this](https://github.com/amirassov/kaggle-imaterialist), where RPNs are connected and treated like a single two-stage model.\nThis improved cort score.\n[![cort-models.png](https://i.postimg.cc/vZgXnkb8/cort-models.png)](https://postimg.cc/pmvDb0m3)\n\n## Augmentation\nMultiscale: astro and cort → (1280, 1280)~(1792, 1792), shsy5y → (1280, 1280)~(1536, 1536)\nHorizontal Flip\nVertical Flip\n\n## TTA\nMultiscale + horizontal flip (vertical flip and diagonal flip didn't work)\n\n# Post-Processing\n\n## astro pp\n@ren4yu \nThe astro annotation was broken, so I reproduced it with cv2.findContours and cv2.fillConvexPoly.\n\n## fix-overlap\n@tyaiga \nIn this competition, predictions are not allowed to overlap. So, we have to eliminate overlap part in post-processing. \nIn our fix-overlap process, first we took cells with a higher cofidence score than classwise threshold. Then, we processed from the highest score to the lowest, and deleted instances with a large percentage of already-used area. In addition, only shsy5y score was improved by removing those with pixels lower than threshold.\n\n## semantic re-lank\n@ren4yu \nAs shown in the Refinemask [paper](https://arxiv.org/abs/2104.08569), the confidence score of the instance segmentation model does not reflect the correctness of the mask. Therefore, we modified the scores and re-lank the instances with semantic segmentation model (UNet++).\n\n## fix-overlap ensemble\n@tyaiga @tanakar @ren4yu \nIn this ensemble method, we first concat the output of multiple models and sort them by their semantic relank scores. Next, we applied fix-overlap pp to remove the overlap. This improved shsy5y score.\n[![fixoverlap-ensemble.png](https://i.postimg.cc/R017Qz8m/fixoverlap-ensemble.png)](https://postimg.cc/FfRkNwNC)\n\n## WBF with mask\n@tereka \nWe tried WBF extended for mask. This improved Public LB score, but 'ensemble detection model' is better for Private LB.\n\n# Tips\nThe default config file of MMDetection is fitted for COCO. Depending on the shape of the instances and the number of instances in an image, it is necessary to change the settings for training.\ne.g. anchor_generator.ratios, rpn_proposal.nms_pre, and so on...",
      "votes": 60
    },
    {
      "id": 1636010,
      "postDate": "2022-01-02T14:02:42.780Z",
      "content": "<p>Great works ! and I have a question.</p>\n<blockquote>\n  <blockquote>\n    <p>The default config file of MMDetection is fitted for COCO<br>\n    e.g. anchor_generator.ratios, rpn_proposal.nms_pre, and so on…</p>\n  </blockquote>\n</blockquote>\n<p>Could I ask you what parameters should be set?<br>\nMy default settings were rpn_proposal. nms_pre=2000<br>\nanchor_generator.ratios=[0.5, 1.0, 2.0]</p>\n<p>how about other settings?</p>",
      "rawMarkdown": "Great works ! and I have a question.\n\n>>The default config file of MMDetection is fitted for COCO\n>>e.g. anchor_generator.ratios, rpn_proposal.nms_pre, and so on…\n\nCould I ask you what parameters should be set?\nMy default settings were rpn_proposal. nms_pre=2000\nanchor_generator.ratios=[0.5, 1.0, 2.0]\n\nhow about other settings?",
      "votes": 1,
      "replies": [
        {
          "id": 1636539,
          "postDate": "2022-01-03T03:12:59.050Z",
          "content": "<pre><code>rpn_proposal.nms_pre=4000,\nrpn_proposal.nms_post=4000,\nrpn_proposal.max_per_img=4000\n\nanchor_generator.ratios=[0.25, 0.5, 1.0, 2.0, 4.0]\n</code></pre>\n<p>This improved astro score.<br>\nThe settings for each cell type are different.</p>",
          "rawMarkdown": "```\nrpn_proposal.nms_pre=4000,\nrpn_proposal.nms_post=4000,\nrpn_proposal.max_per_img=4000\n\nanchor_generator.ratios=[0.25, 0.5, 1.0, 2.0, 4.0]\n```\nThis improved astro score.\nThe settings for each cell type are different.",
          "votes": 2
        },
        {
          "id": 1636717,
          "postDate": "2022-01-03T07:32:39.273Z",
          "content": "<p>I see. Thanks!</p>",
          "rawMarkdown": "I see. Thanks!"
        }
      ]
    },
    {
      "id": 1635694,
      "postDate": "2022-01-02T05:28:25.427Z",
      "content": "<p><img src=\"https://postimg.cc/pmvDb0m3\" alt=\"\"><br>\nHow did you do the ensembling? I mean the code, can you share the ensembling code? or is there some inbuilt function in MMdet</p>",
      "rawMarkdown": "![](https://postimg.cc/pmvDb0m3)\nHow did you do the ensembling? I mean the code, can you share the ensembling code? or is there some inbuilt function in MMdet",
      "votes": 1
    },
    {
      "id": 1635370,
      "postDate": "2022-01-01T17:47:33.077Z",
      "content": "<p>Congratulations !great work</p>",
      "rawMarkdown": "Congratulations !great work"
    },
    {
      "id": 1641225,
      "postDate": "2022-01-07T07:58:24.910Z",
      "content": "<p>which \"double swin-trainsfomer\" do you use ?  ting or small? </p>",
      "rawMarkdown": "which \"double swin-trainsfomer\" do you use ?  ting or small? ",
      "replies": [
        {
          "id": 1641246,
          "postDate": "2022-01-07T08:18:36.970Z",
          "content": "<p>S in 'DBS' means Small.</p>",
          "rawMarkdown": "S in 'DBS' means Small.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1636868,
      "postDate": "2022-01-03T10:45:49.643Z",
      "content": "<p>Thanks for sharing and congrats on 4th place! 🙌</p>",
      "rawMarkdown": "Thanks for sharing and congrats on 4th place! 🙌"
    },
    {
      "id": 1635848,
      "postDate": "2022-01-02T09:20:06.950Z",
      "content": "<p>nice work! </p>",
      "rawMarkdown": "nice work! "
    },
    {
      "id": 1635794,
      "postDate": "2022-01-02T07:42:29.917Z",
      "content": "<p>Great stuff</p>",
      "rawMarkdown": "Great stuff\n"
    },
    {
      "id": 1635559,
      "postDate": "2022-01-01T21:54:18.620Z",
      "content": "<p>Congratulations on becoming Kaggle Master <a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a>! Excellent solution. Thanks for the write-up.</p>",
      "rawMarkdown": "Congratulations on becoming Kaggle Master @tyaiga! Excellent solution. Thanks for the write-up."
    },
    {
      "id": 1635406,
      "postDate": "2022-01-01T18:18:02.603Z",
      "content": "<p>Thanks for information and Congrats for 4th Position!🎉🎉 <a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> </p>",
      "rawMarkdown": "Thanks for information and Congrats for 4th Position!🎉🎉 @tyaiga "
    },
    {
      "id": 1634708,
      "postDate": "2022-01-01T04:03:27.703Z",
      "content": "<p>Congratulations for 2020, to all the winners and the fellow Kaggle community members.</p>",
      "rawMarkdown": "Congratulations for 2020, to all the winners and the fellow Kaggle community members."
    },
    {
      "id": 1636678,
      "postDate": "2022-01-03T06:53:42.570Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1636010,
      "author_name": "opusen",
      "author_url": "",
      "post_date": "2022-01-02T14:02:42.780000",
      "content": "<p>Great works ! and I have a question.</p>\n<blockquote>\n  <blockquote>\n    <p>The default config file of MMDetection is fitted for COCO<br>\n    e.g. anchor_generator.ratios, rpn_proposal.nms_pre, and so on…</p>\n  </blockquote>\n</blockquote>\n<p>Could I ask you what parameters should be set?<br>\nMy default settings were rpn_proposal. nms_pre=2000<br>\nanchor_generator.ratios=[0.5, 1.0, 2.0]</p>\n<p>how about other settings?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1636539,
          "author_name": "Yamame🐟",
          "author_url": "",
          "post_date": "2022-01-03T03:12:59.050000",
          "content": "<pre><code>rpn_proposal.nms_pre=4000,\nrpn_proposal.nms_post=4000,\nrpn_proposal.max_per_img=4000\n\nanchor_generator.ratios=[0.25, 0.5, 1.0, 2.0, 4.0]\n</code></pre>\n<p>This improved astro score.<br>\nThe settings for each cell type are different.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1636717,
          "author_name": "opusen",
          "author_url": "",
          "post_date": "2022-01-03T07:32:39.273000",
          "content": "<p>I see. Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1635694,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2022-01-02T05:28:25.427000",
      "content": "<p><img src=\"https://postimg.cc/pmvDb0m3\" alt=\"\"><br>\nHow did you do the ensembling? I mean the code, can you share the ensembling code? or is there some inbuilt function in MMdet</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1635370,
      "author_name": "Aiswarya Sivakumar",
      "author_url": "",
      "post_date": "2022-01-01T17:47:33.077000",
      "content": "<p>Congratulations !great work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1641225,
      "author_name": "Drzhuzhe",
      "author_url": "",
      "post_date": "2022-01-07T07:58:24.910000",
      "content": "<p>which \"double swin-trainsfomer\" do you use ?  ting or small? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1641246,
          "author_name": "Yamame🐟",
          "author_url": "",
          "post_date": "2022-01-07T08:18:36.970000",
          "content": "<p>S in 'DBS' means Small.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1636868,
      "author_name": "Niek van der Zwaag",
      "author_url": "",
      "post_date": "2022-01-03T10:45:49.643000",
      "content": "<p>Thanks for sharing and congrats on 4th place! 🙌</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1635848,
      "author_name": "Haw Keat",
      "author_url": "",
      "post_date": "2022-01-02T09:20:06.950000",
      "content": "<p>nice work! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1635794,
      "author_name": "YD Glitch",
      "author_url": "",
      "post_date": "2022-01-02T07:42:29.917000",
      "content": "<p>Great stuff</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1635559,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2022-01-01T21:54:18.620000",
      "content": "<p>Congratulations on becoming Kaggle Master <a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a>! Excellent solution. Thanks for the write-up.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1635406,
      "author_name": "Bala Vashan",
      "author_url": "",
      "post_date": "2022-01-01T18:18:02.603000",
      "content": "<p>Thanks for information and Congrats for 4th Position!🎉🎉 <a href=\"https://www.kaggle.com/tyaiga\" target=\"_blank\">@tyaiga</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1634708,
      "author_name": "Ram Jas",
      "author_url": "",
      "post_date": "2022-01-01T04:03:27.703000",
      "content": "<p>Congratulations for 2020, to all the winners and the fellow Kaggle community members.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1636678,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-01-03T06:53:42.570000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1634688": "Thanks to Kaggle and Sartorius for this interesting competition. I also thank my teammates, @tanakar and @tereka, @ren4yu. I learned a lot from them.\n\nOur experiments are based on CBNetV2 [repo](https://github.com/VDIGPKU/CBNetV2)\n\nIn the following, I want to give a summary of our solution.\n\n# Overall\nOur solution consists of three parts: classification part, instance segmentation part, post-processing part. Each cell type has different instance segmentation models and post-processing, so classification is necessary.\n[![overall.png](https://i.postimg.cc/RhpQ6QYn/overall.png)](https://postimg.cc/jnNJBNts)\n\n# Classification\n3class CBNet DBS Cascade-RCNN  (num_classes=3) was used as a classification model. For a given image, this model classifies the image into the cell type with the highest number of detected cells. \nThis model is first pre-trained with LIVE_CELL (num_classes=1) and then fine-tuned with the competition's train data (num_classes=3).\nI tested this on semi-supervised data and it was able to classify them perfectly.\n\n# Instance Segmentation\nWe prepared at least one model for each of shsy5y, astro, and cort. These models were trained in different settings and with different data. \n## Pseudo Labeling\n@tereka \nFor all cell types, pseudo labeling on semi-supervised data improved both CV and LB.\nBecause of the small number of train data, it was better to use three types of cells when using only train data. However, because of the large number of semi-supervised data, we were able to change the data used for each cell type.\n[Pseudo labeling](https://drive.google.com/file/d/1PSIxdNiwtMwTV3Wq6AJbWcp2plnoIVQA/view?usp=sharing)\n### Models for pseudo labeling\nData: live_cell → train (all cell types)\nModel: 3 types of 1class CBNet DBS Cascade-RCNN\nWe changed the MMdet config file depending on the target cell type.\n\n[![train-strategy.png](https://i.postimg.cc/W4Z9BKHY/train-strategy.png)](https://postimg.cc/DW2dsCF1)\n## shsy5y models\n@tyaiga \nData1: live_cell → train (all cell types)\nModel1: CBNet DBS Cascade-RCNN\nData2: live_cell → semi-sup+train (shsy5y & cort) → train (shsy5y)\nModel2: CBNet DBS Cascade-RCNN\nThe number of cells in shsy5y and cort are very different, but the individual cells are similar, so, we used the two as training data when using semi-sup data.\n\n## astro models\n@tyaiga \nData: live_cell → semi-sup+train (astro) → train (astro)\nSince astro has a very different shape and size from the other cells, we improved the score by using only astro data fot train data.\nIn astro, I only used one model because the ensemble did not work well.\n\n## cort models\n@tanakar \nData: live_cell → semi-sup+train (shsy5y & cort) → train (court)\nModel: 2 x HTC resnext64x4d, 2 x CBNet DBS Cascade-RCNN\nThese four models were combined into one model using the method below.\n\n### ensemble detection model\n@tereka @ren4yu \nWe only used the two-stage model. So, we used an ensemble method like [this](https://github.com/amirassov/kaggle-imaterialist), where RPNs are connected and treated like a single two-stage model.\nThis improved cort score.\n[![cort-models.png](https://i.postimg.cc/vZgXnkb8/cort-models.png)](https://postimg.cc/pmvDb0m3)\n\n## Augmentation\nMultiscale: astro and cort → (1280, 1280)~(1792, 1792), shsy5y → (1280, 1280)~(1536, 1536)\nHorizontal Flip\nVertical Flip\n\n## TTA\nMultiscale + horizontal flip (vertical flip and diagonal flip didn't work)\n\n# Post-Processing\n\n## astro pp\n@ren4yu \nThe astro annotation was broken, so I reproduced it with cv2.findContours and cv2.fillConvexPoly.\n\n## fix-overlap\n@tyaiga \nIn this competition, predictions are not allowed to overlap. So, we have to eliminate overlap part in post-processing. \nIn our fix-overlap process, first we took cells with a higher cofidence score than classwise threshold. Then, we processed from the highest score to the lowest, and deleted instances with a large percentage of already-used area. In addition, only shsy5y score was improved by removing those with pixels lower than threshold.\n\n## semantic re-lank\n@ren4yu \nAs shown in the Refinemask [paper](https://arxiv.org/abs/2104.08569), the confidence score of the instance segmentation model does not reflect the correctness of the mask. Therefore, we modified the scores and re-lank the instances with semantic segmentation model (UNet++).\n\n## fix-overlap ensemble\n@tyaiga @tanakar @ren4yu \nIn this ensemble method, we first concat the output of multiple models and sort them by their semantic relank scores. Next, we applied fix-overlap pp to remove the overlap. This improved shsy5y score.\n[![fixoverlap-ensemble.png](https://i.postimg.cc/R017Qz8m/fixoverlap-ensemble.png)](https://postimg.cc/FfRkNwNC)\n\n## WBF with mask\n@tereka \nWe tried WBF extended for mask. This improved Public LB score, but 'ensemble detection model' is better for Private LB.\n\n# Tips\nThe default config file of MMDetection is fitted for COCO. Depending on the shape of the instances and the number of instances in an image, it is necessary to change the settings for training.\ne.g. anchor_generator.ratios, rpn_proposal.nms_pre, and so on...",
    "1636010": "Great works ! and I have a question.\n\n>>The default config file of MMDetection is fitted for COCO\n>>e.g. anchor_generator.ratios, rpn_proposal.nms_pre, and so on…\n\nCould I ask you what parameters should be set?\nMy default settings were rpn_proposal. nms_pre=2000\nanchor_generator.ratios=[0.5, 1.0, 2.0]\n\nhow about other settings?",
    "1635694": "![](https://postimg.cc/pmvDb0m3)\nHow did you do the ensembling? I mean the code, can you share the ensembling code? or is there some inbuilt function in MMdet",
    "1635370": "Congratulations !great work",
    "1641225": "which \"double swin-trainsfomer\" do you use ?  ting or small? ",
    "1636868": "Thanks for sharing and congrats on 4th place! 🙌",
    "1635848": "nice work! ",
    "1635794": "Great stuff\n",
    "1635559": "Congratulations on becoming Kaggle Master @tyaiga! Excellent solution. Thanks for the write-up.",
    "1635406": "Thanks for information and Congrats for 4th Position!🎉🎉 @tyaiga ",
    "1634708": "Congratulations for 2020, to all the winners and the fellow Kaggle community members.",
    "1636678": ""
  }
}