{
  "id": 428447,
  "title": "9th place solution",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/428447",
  "author_name": "D.Imanishi",
  "post_date": "2023-08-01T12:55:24.994000",
  "votes": 43,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thanks to the organizers and congrats to all the winners!</p>\n<h2>Overview</h2>\n<p>From the top 1&amp;2 solutions of the <a href=\"https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation\" target=\"_blank\">Sartorius competition</a>, I decided to go with a two-stage pipeline of object detection and semantic segmentation instead of using a one-stage model (e.g. Mask-RCNN) in the early in the competition. I think that the main advantage of the two-stage pipeline is the ease of the ensemble and TTA.</p>\n<p><img src=\"https://i.postimg.cc/3xDzz1G0/hubmap1.png\" alt=\"\"></p>\n<h2>Detection part</h2>\n<h3>Update of dataset2 bbox-level annotation</h3>\n<p>As been pointed out in the discussion, the dilation significantly improved LB score in models trained from both dataset1 and dataset2. However, models trained only from dataset1 did not show this effect. This led me to believe that the annotations of dataset2 were made smaller than those of dataset1. Considering that there are several types of blood vessels and the possibility that the annotations of them are not uniformly small, I took the approach of updating the bbox-level annotations of dataset2 to dilate them by models trained only from dataset1.</p>\n<p><img src=\"https://i.postimg.cc/MK36b97C/hubmap2.png\" alt=\"\"></p>\n<p>The update was performed by replacing the original annotations with a predicted ones that met the following conditions. Basically, the updated bboxes is larger than the original ones. The total number of bboxes in dataset2 does not change by the updating.</p>\n<ol>\n<li>IoU &gt; 0.4</li>\n<li>FP / (TP + FP) &gt; 0.1</li>\n</ol>\n<p><img src=\"https://i.postimg.cc/jSGRBCkv/hubmap3.png\" alt=\"\"></p>\n<p>In the model trained from the updated dataset2 together with dataset1, the LB improvement by the dilation has almost disappeared and the LB score above 0.5 was achieved without the dilation.</p>\n<h3>CV strategy</h3>\n<p>The CV was carried out by the following special two-fold division.</p>\n<p><img src=\"https://i.postimg.cc/jSWs8yvL/hubmap4.png\" alt=\"\"></p>\n<h3>Model training</h3>\n<p>I trained 2-class (blood_vessel, glomerulus) detection models. \"unsure\" label was ignored. I modified the source code of YOLOv5/v7/v8 and add a 90 degree random rotation augmentation. The following 5 models x 2 folds (total 10 models) were used in the final submission.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Input size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>YOLOv5x6</td>\n<td>512</td>\n</tr>\n<tr>\n<td>YOLOv7x</td>\n<td>512</td>\n</tr>\n<tr>\n<td>YOLOv8l</td>\n<td>512</td>\n</tr>\n<tr>\n<td>YOLOv8x</td>\n<td>512</td>\n</tr>\n<tr>\n<td>YOLOv8l</td>\n<td>768</td>\n</tr>\n</tbody>\n</table>\n<h3>Inference</h3>\n<p>I made a huge ensemble of 10 models x 16 TTAs since I considered the accuracy of detection to be more important than one of the segmentation. The small number of test images made it possible. 10 x 16 = 160 detection results were merged by WBF.</p>\n<ul>\n<li>16 TTAs: 8 for combinations of h-flip, v-flip, 90deg rotation, 2 for 2 scales (base size, base size + 64px), 8 x 2 = 16</li>\n<li>IoU threshold of each model's NMS: 0.6</li>\n<li>IoU threshold of WBF: 0.7</li>\n</ul>\n<h2>Segmentation part</h2>\n<h3>CV strategy</h3>\n<p>Almost the same as detection models, except that dataset2 is not updated.</p>\n<h3>Model training</h3>\n<p>As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations. Unlike the detection models, \"unsure\" label was also used for training. In the final submission, EfficientNetB1-Unet and EfficientNetB2-Unet was used (each 2 folds, total 4 models).</p>\n<h3>Inference</h3>\n<p>TTA was not used because of run-time constraints. A score threshold of 0.5 was used for binarization.</p>\n<h2>Strategy of final submission</h2>\n<p>The effect was smaller by updating dataset2, but the small dilation improved the LB score slightly. I implemented the dilation not by cv2.dilate for the final masks, but by increasing the size of bboxes by a percentage. In my final submission, 3% of bbox dilation increased my LB score about 0.005. I used 3% dilation in one of the two final submissions and not in the other (there are other differences besides the dilation).</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>w/ 3% dilation</td>\n<td>0.580</td>\n<td>0.549</td>\n</tr>\n<tr>\n<td>w/o 3% dilation</td>\n<td>0.572</td>\n<td>0.560</td>\n</tr>\n</tbody>\n</table>\n<h2>Things tried but not worked</h2>\n<ul>\n<li>Semi-supervised learning on dataset3</li>\n</ul>",
  "messages": [
    {
      "id": 2368912,
      "postDate": "2023-08-01T12:55:24.993Z",
      "content": "<p>Thanks to the organizers and congrats to all the winners!</p>\n<h2>Overview</h2>\n<p>From the top 1&amp;2 solutions of the <a href=\"https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation\" target=\"_blank\">Sartorius competition</a>, I decided to go with a two-stage pipeline of object detection and semantic segmentation instead of using a one-stage model (e.g. Mask-RCNN) in the early in the competition. I think that the main advantage of the two-stage pipeline is the ease of the ensemble and TTA.</p>\n<p><img src=\"https://i.postimg.cc/3xDzz1G0/hubmap1.png\" alt=\"\"></p>\n<h2>Detection part</h2>\n<h3>Update of dataset2 bbox-level annotation</h3>\n<p>As been pointed out in the discussion, the dilation significantly improved LB score in models trained from both dataset1 and dataset2. However, models trained only from dataset1 did not show this effect. This led me to believe that the annotations of dataset2 were made smaller than those of dataset1. Considering that there are several types of blood vessels and the possibility that the annotations of them are not uniformly small, I took the approach of updating the bbox-level annotations of dataset2 to dilate them by models trained only from dataset1.</p>\n<p><img src=\"https://i.postimg.cc/MK36b97C/hubmap2.png\" alt=\"\"></p>\n<p>The update was performed by replacing the original annotations with a predicted ones that met the following conditions. Basically, the updated bboxes is larger than the original ones. The total number of bboxes in dataset2 does not change by the updating.</p>\n<ol>\n<li>IoU &gt; 0.4</li>\n<li>FP / (TP + FP) &gt; 0.1</li>\n</ol>\n<p><img src=\"https://i.postimg.cc/jSGRBCkv/hubmap3.png\" alt=\"\"></p>\n<p>In the model trained from the updated dataset2 together with dataset1, the LB improvement by the dilation has almost disappeared and the LB score above 0.5 was achieved without the dilation.</p>\n<h3>CV strategy</h3>\n<p>The CV was carried out by the following special two-fold division.</p>\n<p><img src=\"https://i.postimg.cc/jSWs8yvL/hubmap4.png\" alt=\"\"></p>\n<h3>Model training</h3>\n<p>I trained 2-class (blood_vessel, glomerulus) detection models. \"unsure\" label was ignored. I modified the source code of YOLOv5/v7/v8 and add a 90 degree random rotation augmentation. The following 5 models x 2 folds (total 10 models) were used in the final submission.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Input size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>YOLOv5x6</td>\n<td>512</td>\n</tr>\n<tr>\n<td>YOLOv7x</td>\n<td>512</td>\n</tr>\n<tr>\n<td>YOLOv8l</td>\n<td>512</td>\n</tr>\n<tr>\n<td>YOLOv8x</td>\n<td>512</td>\n</tr>\n<tr>\n<td>YOLOv8l</td>\n<td>768</td>\n</tr>\n</tbody>\n</table>\n<h3>Inference</h3>\n<p>I made a huge ensemble of 10 models x 16 TTAs since I considered the accuracy of detection to be more important than one of the segmentation. The small number of test images made it possible. 10 x 16 = 160 detection results were merged by WBF.</p>\n<ul>\n<li>16 TTAs: 8 for combinations of h-flip, v-flip, 90deg rotation, 2 for 2 scales (base size, base size + 64px), 8 x 2 = 16</li>\n<li>IoU threshold of each model's NMS: 0.6</li>\n<li>IoU threshold of WBF: 0.7</li>\n</ul>\n<h2>Segmentation part</h2>\n<h3>CV strategy</h3>\n<p>Almost the same as detection models, except that dataset2 is not updated.</p>\n<h3>Model training</h3>\n<p>As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations. Unlike the detection models, \"unsure\" label was also used for training. In the final submission, EfficientNetB1-Unet and EfficientNetB2-Unet was used (each 2 folds, total 4 models).</p>\n<h3>Inference</h3>\n<p>TTA was not used because of run-time constraints. A score threshold of 0.5 was used for binarization.</p>\n<h2>Strategy of final submission</h2>\n<p>The effect was smaller by updating dataset2, but the small dilation improved the LB score slightly. I implemented the dilation not by cv2.dilate for the final masks, but by increasing the size of bboxes by a percentage. In my final submission, 3% of bbox dilation increased my LB score about 0.005. I used 3% dilation in one of the two final submissions and not in the other (there are other differences besides the dilation).</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>w/ 3% dilation</td>\n<td>0.580</td>\n<td>0.549</td>\n</tr>\n<tr>\n<td>w/o 3% dilation</td>\n<td>0.572</td>\n<td>0.560</td>\n</tr>\n</tbody>\n</table>\n<h2>Things tried but not worked</h2>\n<ul>\n<li>Semi-supervised learning on dataset3</li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers and congrats to all the winners!\n\n## Overview\nFrom the top 1&2 solutions of the [Sartorius competition](https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation), I decided to go with a two-stage pipeline of object detection and semantic segmentation instead of using a one-stage model (e.g. Mask-RCNN) in the early in the competition. I think that the main advantage of the two-stage pipeline is the ease of the ensemble and TTA.\n\n![](https://i.postimg.cc/3xDzz1G0/hubmap1.png)\n\n## Detection part\n\n### Update of dataset2 bbox-level annotation \nAs been pointed out in the discussion, the dilation significantly improved LB score in models trained from both dataset1 and dataset2. However, models trained only from dataset1 did not show this effect. This led me to believe that the annotations of dataset2 were made smaller than those of dataset1. Considering that there are several types of blood vessels and the possibility that the annotations of them are not uniformly small, I took the approach of updating the bbox-level annotations of dataset2 to dilate them by models trained only from dataset1.\n\n![](https://i.postimg.cc/MK36b97C/hubmap2.png)\n\nThe update was performed by replacing the original annotations with a predicted ones that met the following conditions. Basically, the updated bboxes is larger than the original ones. The total number of bboxes in dataset2 does not change by the updating.\n1. IoU > 0.4\n2. FP / (TP + FP) > 0.1\n\n![](https://i.postimg.cc/jSGRBCkv/hubmap3.png)\n\nIn the model trained from the updated dataset2 together with dataset1, the LB improvement by the dilation has almost disappeared and the LB score above 0.5 was achieved without the dilation.\n\n### CV strategy\nThe CV was carried out by the following special two-fold division.\n\n![](https://i.postimg.cc/jSWs8yvL/hubmap4.png)\n\n### Model training\nI trained 2-class (blood_vessel, glomerulus) detection models. \"unsure\" label was ignored. I modified the source code of YOLOv5/v7/v8 and add a 90 degree random rotation augmentation. The following 5 models x 2 folds (total 10 models) were used in the final submission.\n\n| Model | Input size |\n| ---- | ---- |\n| YOLOv5x6 | 512 |\n| YOLOv7x | 512 |\n| YOLOv8l | 512 |\n| YOLOv8x | 512 |\n| YOLOv8l | 768 |\n\n### Inference\nI made a huge ensemble of 10 models x 16 TTAs since I considered the accuracy of detection to be more important than one of the segmentation. The small number of test images made it possible. 10 x 16 = 160 detection results were merged by WBF.\n\n- 16 TTAs: 8 for combinations of h-flip, v-flip, 90deg rotation, 2 for 2 scales (base size, base size + 64px), 8 x 2 = 16\n- IoU threshold of each model's NMS: 0.6\n- IoU threshold of WBF: 0.7\n\n\n## Segmentation part\n### CV strategy\nAlmost the same as detection models, except that dataset2 is not updated.\n\n### Model training\nAs a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations. Unlike the detection models, \"unsure\" label was also used for training. In the final submission, EfficientNetB1-Unet and EfficientNetB2-Unet was used (each 2 folds, total 4 models).\n\n### Inference\nTTA was not used because of run-time constraints. A score threshold of 0.5 was used for binarization.\n\n## Strategy of final submission\nThe effect was smaller by updating dataset2, but the small dilation improved the LB score slightly. I implemented the dilation not by cv2.dilate for the final masks, but by increasing the size of bboxes by a percentage. In my final submission, 3% of bbox dilation increased my LB score about 0.005. I used 3% dilation in one of the two final submissions and not in the other (there are other differences besides the dilation).\n\n| Submission | Public LB | Private LB |\n| ---- | ---- | ---- |\n| w/ 3% dilation | 0.580 | 0.549 |\n| w/o 3% dilation | 0.572 | 0.560 |\n\n\n## Things tried but not worked\n- Semi-supervised learning on dataset3\n\n\n",
      "votes": 43
    },
    {
      "id": 2375987,
      "postDate": "2023-08-06T04:41:41.127Z",
      "content": "<p>creative one</p>",
      "rawMarkdown": "creative one",
      "votes": 1
    },
    {
      "id": 2371826,
      "postDate": "2023-08-03T09:39:34.153Z",
      "content": "<p>Thanks for sharing! \"Update of dataset2 bbox-level annotation\" has given me great inspiration.</p>",
      "rawMarkdown": "Thanks for sharing! \"Update of dataset2 bbox-level annotation\" has given me great inspiration.",
      "votes": 1
    },
    {
      "id": 2371725,
      "postDate": "2023-08-03T08:54:31.957Z",
      "content": "<p>Congrats on your solo gold. <br>\nIf I understood  correctly, you trained your seg models on both dataset1 &amp; dataset2 with the original annotation. And you didn't dilate the segmentation mask itself during inference, am I correct?</p>",
      "rawMarkdown": "Congrats on your solo gold. \nIf I understood  correctly, you trained your seg models on both dataset1 & dataset2 with the original annotation. And you didn't dilate the segmentation mask itself during inference, am I correct?",
      "votes": 1,
      "replies": [
        {
          "id": 2372771,
          "postDate": "2023-08-03T23:50:45.473Z",
          "content": "<p>Thank you for the question!<br>\nYes, that is the correct understanding.</p>",
          "rawMarkdown": "Thank you for the question!\nYes, that is the correct understanding.",
          "votes": 1,
          "replies": [
            {
              "id": 2373702,
              "postDate": "2023-08-04T12:02:15.780Z",
              "content": "<p>Thank you for your reply.<br>\nI'm still a bit uncertain about how you managed to avoid dilation on masks. Dilating the bounding box would indeed increase the portion of the image being fed to the segmentation model, but it wouldn't enlarge the mask itself. Many participants who trained segmentation models on dataset1 and dataset2 (without fine-tuning on ds1) faced challenges in achieving high scores on the public leaderboard without applying dilation to the masks. </p>",
              "rawMarkdown": "Thank you for your reply.\nI'm still a bit uncertain about how you managed to avoid dilation on masks. Dilating the bounding box would indeed increase the portion of the image being fed to the segmentation model, but it wouldn't enlarge the mask itself. Many participants who trained segmentation models on dataset1 and dataset2 (without fine-tuning on ds1) faced challenges in achieving high scores on the public leaderboard without applying dilation to the masks. "
            },
            {
              "id": 2375149,
              "postDate": "2023-08-05T12:49:11.123Z",
              "content": "<p>Basically, the segmentation model outputs mask predictions of the same size (same width and same height) as the input bbox. Therefore, by dilating the bbox, the mask prediction is also dilated compared to the original bbox input.</p>",
              "rawMarkdown": "Basically, the segmentation model outputs mask predictions of the same size (same width and same height) as the input bbox. Therefore, by dilating the bbox, the mask prediction is also dilated compared to the original bbox input.",
              "votes": 1
            },
            {
              "id": 2379364,
              "postDate": "2023-08-08T06:14:30.347Z",
              "content": "<p>Hi, <a href=\"https://www.kaggle.com/tanjiroll\" target=\"_blank\">@tanjiroll</a>, I see you are looking for a teammate for CommonLit - Evaluate Student Summaries Competition. Not able to contact you through your Profile so contacting you here.</p>",
              "rawMarkdown": "Hi, @tanjiroll, I see you are looking for a teammate for CommonLit - Evaluate Student Summaries Competition. Not able to contact you through your Profile so contacting you here."
            }
          ]
        }
      ]
    },
    {
      "id": 2370474,
      "postDate": "2023-08-02T12:48:55.853Z",
      "content": "<p>Thanks for the clear report and congrats on your achievement with this original solution !</p>",
      "rawMarkdown": "Thanks for the clear report and congrats on your achievement with this original solution !",
      "votes": 1
    },
    {
      "id": 2369003,
      "postDate": "2023-08-01T13:45:14.750Z",
      "content": "<p>Increase the bbox. Very original idea.</p>",
      "rawMarkdown": "Increase the bbox. Very original idea.",
      "votes": 1
    },
    {
      "id": 2370183,
      "postDate": "2023-08-02T08:40:04.987Z",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a> for the report!</p>\n<p>I'm curioius hous did you come up with the following ideas?</p>\n<ul>\n<li>\"As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations.\"  Have you tested both options and settled on this one?</li>\n<li>\"Unlike the detection models, \"unsure\" label was also used for training.\"  Was it a result of testing both options or you had some prior intuition about it? </li>\n<li>\"A score threshold of 0.5 was used for binarization.\".  Another option would be to not have binarization at all. Is binarization always better or you tested both options?</li>\n</ul>",
      "rawMarkdown": "Thank you, @dimanishi for the report!\n\nI'm curioius hous did you come up with the following ideas?\n- \"As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations.\"  Have you tested both options and settled on this one?\n- \"Unlike the detection models, \"unsure\" label was also used for training.\"  Was it a result of testing both options or you had some prior intuition about it? \n- \"A score threshold of 0.5 was used for binarization.\".  Another option would be to not have binarization at all. Is binarization always better or you tested both options?\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 2370470,
          "postDate": "2023-08-02T12:45:05.263Z",
          "content": "<p>Thank you for the question.</p>\n<blockquote>\n  <p>\"As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations.\" Have you tested both options and settled on this one?</p>\n</blockquote>\n<p>No, only this option was tested. The 1st place solution in the satorius competition used this method, so I was confident in my choice.</p>\n<blockquote>\n  <p>\"Unlike the detection models, \"unsure\" label was also used for training.\" Was it a result of testing both options or you had some prior intuition about it?</p>\n</blockquote>\n<p>No, only this option was tested, too. I thought that the benefit of increasing the amount of data would be greater than the loss of data quality due to using \"unsure\" in the segmentation task.</p>\n<blockquote>\n  <p>\"A score threshold of 0.5 was used for binarization.\". Another option would be to not have binarization at all. Is binarization always better or you tested both options?</p>\n</blockquote>\n<p>You may have misunderstood something… I don't think there is an option to not binarize the mask, since the final submission mask must be binarized for RLE in this competition.</p>",
          "rawMarkdown": "Thank you for the question.\n\n> \"As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations.\" Have you tested both options and settled on this one?\n\nNo, only this option was tested. The 1st place solution in the satorius competition used this method, so I was confident in my choice.\n\n> \"Unlike the detection models, \"unsure\" label was also used for training.\" Was it a result of testing both options or you had some prior intuition about it?\n\nNo, only this option was tested, too. I thought that the benefit of increasing the amount of data would be greater than the loss of data quality due to using \"unsure\" in the segmentation task.\n\n> \"A score threshold of 0.5 was used for binarization.\". Another option would be to not have binarization at all. Is binarization always better or you tested both options?\n\nYou may have misunderstood something... I don't think there is an option to not binarize the mask, since the final submission mask must be binarized for RLE in this competition.\n",
          "replies": [
            {
              "id": 2371030,
              "postDate": "2023-08-02T19:24:34.087Z",
              "content": "<p>Thank you, <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a>! Indded, my bad for asking a silly question on masks. I was thinking about confidence levels of course, not binarization per se, but you probably get them from boxes. </p>",
              "rawMarkdown": "Thank you, @dimanishi! Indded, my bad for asking a silly question on masks. I was thinking about confidence levels of course, not binarization per se, but you probably get them from boxes. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2372397,
      "postDate": "2023-08-03T16:35:16.197Z",
      "content": "<p>Congratulations to you to get a gold medal, BBOX is really a very creative idea！</p>",
      "rawMarkdown": "Congratulations to you to get a gold medal, BBOX is really a very creative idea！",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2371811,
      "postDate": "2023-08-03T09:29:31.647Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2375567,
      "postDate": "2023-08-05T18:16:11.483Z",
      "content": "<p>Thanks for your solution!</p>",
      "rawMarkdown": "Thanks for your solution!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2375987,
      "author_name": "MAA",
      "author_url": "",
      "post_date": "2023-08-06T04:41:41.127000",
      "content": "<p>creative one</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2371826,
      "author_name": "FreddyHu",
      "author_url": "",
      "post_date": "2023-08-03T09:39:34.153000",
      "content": "<p>Thanks for sharing! \"Update of dataset2 bbox-level annotation\" has given me great inspiration.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2371725,
      "author_name": "TanjiroLL",
      "author_url": "",
      "post_date": "2023-08-03T08:54:31.957000",
      "content": "<p>Congrats on your solo gold. <br>\nIf I understood  correctly, you trained your seg models on both dataset1 &amp; dataset2 with the original annotation. And you didn't dilate the segmentation mask itself during inference, am I correct?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2372771,
          "author_name": "D.Imanishi",
          "author_url": "",
          "post_date": "2023-08-03T23:50:45.473000",
          "content": "<p>Thank you for the question!<br>\nYes, that is the correct understanding.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2373702,
              "author_name": "TanjiroLL",
              "author_url": "",
              "post_date": "2023-08-04T12:02:15.780000",
              "content": "<p>Thank you for your reply.<br>\nI'm still a bit uncertain about how you managed to avoid dilation on masks. Dilating the bounding box would indeed increase the portion of the image being fed to the segmentation model, but it wouldn't enlarge the mask itself. Many participants who trained segmentation models on dataset1 and dataset2 (without fine-tuning on ds1) faced challenges in achieving high scores on the public leaderboard without applying dilation to the masks. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2375149,
              "author_name": "D.Imanishi",
              "author_url": "",
              "post_date": "2023-08-05T12:49:11.123000",
              "content": "<p>Basically, the segmentation model outputs mask predictions of the same size (same width and same height) as the input bbox. Therefore, by dilating the bbox, the mask prediction is also dilated compared to the original bbox input.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2379364,
              "author_name": "Chirag Chauhan",
              "author_url": "",
              "post_date": "2023-08-08T06:14:30.347000",
              "content": "<p>Hi, <a href=\"https://www.kaggle.com/tanjiroll\" target=\"_blank\">@tanjiroll</a>, I see you are looking for a teammate for CommonLit - Evaluate Student Summaries Competition. Not able to contact you through your Profile so contacting you here.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2370474,
      "author_name": "Hamed JOORATI",
      "author_url": "",
      "post_date": "2023-08-02T12:48:55.853000",
      "content": "<p>Thanks for the clear report and congrats on your achievement with this original solution !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2369003,
      "author_name": "bent1e",
      "author_url": "",
      "post_date": "2023-08-01T13:45:14.750000",
      "content": "<p>Increase the bbox. Very original idea.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2370183,
      "author_name": "Vladimir Slaykovskiy",
      "author_url": "",
      "post_date": "2023-08-02T08:40:04.987000",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a> for the report!</p>\n<p>I'm curioius hous did you come up with the following ideas?</p>\n<ul>\n<li>\"As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations.\"  Have you tested both options and settled on this one?</li>\n<li>\"Unlike the detection models, \"unsure\" label was also used for training.\"  Was it a result of testing both options or you had some prior intuition about it? </li>\n<li>\"A score threshold of 0.5 was used for binarization.\".  Another option would be to not have binarization at all. Is binarization always better or you tested both options?</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 2370470,
          "author_name": "D.Imanishi",
          "author_url": "",
          "post_date": "2023-08-02T12:45:05.263000",
          "content": "<p>Thank you for the question.</p>\n<blockquote>\n  <p>\"As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations.\" Have you tested both options and settled on this one?</p>\n</blockquote>\n<p>No, only this option was tested. The 1st place solution in the satorius competition used this method, so I was confident in my choice.</p>\n<blockquote>\n  <p>\"Unlike the detection models, \"unsure\" label was also used for training.\" Was it a result of testing both options or you had some prior intuition about it?</p>\n</blockquote>\n<p>No, only this option was tested, too. I thought that the benefit of increasing the amount of data would be greater than the loss of data quality due to using \"unsure\" in the segmentation task.</p>\n<blockquote>\n  <p>\"A score threshold of 0.5 was used for binarization.\". Another option would be to not have binarization at all. Is binarization always better or you tested both options?</p>\n</blockquote>\n<p>You may have misunderstood something… I don't think there is an option to not binarize the mask, since the final submission mask must be binarized for RLE in this competition.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2371030,
              "author_name": "Vladimir Slaykovskiy",
              "author_url": "",
              "post_date": "2023-08-02T19:24:34.087000",
              "content": "<p>Thank you, <a href=\"https://www.kaggle.com/dimanishi\" target=\"_blank\">@dimanishi</a>! Indded, my bad for asking a silly question on masks. I was thinking about confidence levels of course, not binarization per se, but you probably get them from boxes. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2372397,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-03T16:35:16.197000",
      "content": "<p>Congratulations to you to get a gold medal, BBOX is really a very creative idea！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2371811,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-03T09:29:31.647000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2375567,
      "author_name": "abhamidi",
      "author_url": "",
      "post_date": "2023-08-05T18:16:11.483000",
      "content": "<p>Thanks for your solution!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2368912": "Thanks to the organizers and congrats to all the winners!\n\n## Overview\nFrom the top 1&2 solutions of the [Sartorius competition](https://www.kaggle.com/competitions/sartorius-cell-instance-segmentation), I decided to go with a two-stage pipeline of object detection and semantic segmentation instead of using a one-stage model (e.g. Mask-RCNN) in the early in the competition. I think that the main advantage of the two-stage pipeline is the ease of the ensemble and TTA.\n\n![](https://i.postimg.cc/3xDzz1G0/hubmap1.png)\n\n## Detection part\n\n### Update of dataset2 bbox-level annotation \nAs been pointed out in the discussion, the dilation significantly improved LB score in models trained from both dataset1 and dataset2. However, models trained only from dataset1 did not show this effect. This led me to believe that the annotations of dataset2 were made smaller than those of dataset1. Considering that there are several types of blood vessels and the possibility that the annotations of them are not uniformly small, I took the approach of updating the bbox-level annotations of dataset2 to dilate them by models trained only from dataset1.\n\n![](https://i.postimg.cc/MK36b97C/hubmap2.png)\n\nThe update was performed by replacing the original annotations with a predicted ones that met the following conditions. Basically, the updated bboxes is larger than the original ones. The total number of bboxes in dataset2 does not change by the updating.\n1. IoU > 0.4\n2. FP / (TP + FP) > 0.1\n\n![](https://i.postimg.cc/jSGRBCkv/hubmap3.png)\n\nIn the model trained from the updated dataset2 together with dataset1, the LB improvement by the dilation has almost disappeared and the LB score above 0.5 was achieved without the dilation.\n\n### CV strategy\nThe CV was carried out by the following special two-fold division.\n\n![](https://i.postimg.cc/jSWs8yvL/hubmap4.png)\n\n### Model training\nI trained 2-class (blood_vessel, glomerulus) detection models. \"unsure\" label was ignored. I modified the source code of YOLOv5/v7/v8 and add a 90 degree random rotation augmentation. The following 5 models x 2 folds (total 10 models) were used in the final submission.\n\n| Model | Input size |\n| ---- | ---- |\n| YOLOv5x6 | 512 |\n| YOLOv7x | 512 |\n| YOLOv8l | 512 |\n| YOLOv8x | 512 |\n| YOLOv8l | 768 |\n\n### Inference\nI made a huge ensemble of 10 models x 16 TTAs since I considered the accuracy of detection to be more important than one of the segmentation. The small number of test images made it possible. 10 x 16 = 160 detection results were merged by WBF.\n\n- 16 TTAs: 8 for combinations of h-flip, v-flip, 90deg rotation, 2 for 2 scales (base size, base size + 64px), 8 x 2 = 16\n- IoU threshold of each model's NMS: 0.6\n- IoU threshold of WBF: 0.7\n\n\n## Segmentation part\n### CV strategy\nAlmost the same as detection models, except that dataset2 is not updated.\n\n### Model training\nAs a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations. Unlike the detection models, \"unsure\" label was also used for training. In the final submission, EfficientNetB1-Unet and EfficientNetB2-Unet was used (each 2 folds, total 4 models).\n\n### Inference\nTTA was not used because of run-time constraints. A score threshold of 0.5 was used for binarization.\n\n## Strategy of final submission\nThe effect was smaller by updating dataset2, but the small dilation improved the LB score slightly. I implemented the dilation not by cv2.dilate for the final masks, but by increasing the size of bboxes by a percentage. In my final submission, 3% of bbox dilation increased my LB score about 0.005. I used 3% dilation in one of the two final submissions and not in the other (there are other differences besides the dilation).\n\n| Submission | Public LB | Private LB |\n| ---- | ---- | ---- |\n| w/ 3% dilation | 0.580 | 0.549 |\n| w/o 3% dilation | 0.572 | 0.560 |\n\n\n## Things tried but not worked\n- Semi-supervised learning on dataset3\n\n\n",
    "2375987": "creative one",
    "2371826": "Thanks for sharing! \"Update of dataset2 bbox-level annotation\" has given me great inspiration.",
    "2371725": "Congrats on your solo gold. \nIf I understood  correctly, you trained your seg models on both dataset1 & dataset2 with the original annotation. And you didn't dilate the segmentation mask itself during inference, am I correct?",
    "2370474": "Thanks for the clear report and congrats on your achievement with this original solution !",
    "2369003": "Increase the bbox. Very original idea.",
    "2370183": "Thank you, @dimanishi for the report!\n\nI'm curioius hous did you come up with the following ideas?\n- \"As a mask of bboxes, I did not use the prediction results of the detection models, but used the bboxes obtained from the original annotations.\"  Have you tested both options and settled on this one?\n- \"Unlike the detection models, \"unsure\" label was also used for training.\"  Was it a result of testing both options or you had some prior intuition about it? \n- \"A score threshold of 0.5 was used for binarization.\".  Another option would be to not have binarization at all. Is binarization always better or you tested both options?\n\n",
    "2372397": "Congratulations to you to get a gold medal, BBOX is really a very creative idea！",
    "2371811": "",
    "2375567": "Thanks for your solution!"
  }
}