{
  "id": 307871,
  "title": "9th place solution: Tiled Training + Seq-NMS",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/307871",
  "author_name": "Bilzard",
  "post_date": "2022-02-16T02:05:48.709000",
  "votes": 24,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Much appreciated for organizing the competition, and thanks to all people who committed to this project.</p>\n<p>I will share my solution with you.</p>\n<h1>Pipeline</h1>\n<h2>Preprocessing</h2>\n<ul>\n<li>Tiled Training<ul>\n<li>Split original image into 320x320 sized tiles, with each tiles have intersections</li>\n<li>Stride is chosen as (sx, sy) = (240, 200), which gives us 5x3 tiles for each image. This setting is according to [6].</li></ul></li>\n<li>Relabeling<ul>\n<li>I noticed a lot of inconsistent labels and missing labels on the train dataset, so I thought relabeling will improve the model's performance.</li>\n<li>I relabeled with convining original lagbels and pseudo labels generated by trained model:<ul>\n<li>NMS(original(with 0.501) + pseudo label)</li>\n<li>threshold(conf&gt;0.5)</li></ul></li></ul></li>\n</ul>\n<h2>Fold Split</h2>\n<ul>\n<li>Split sequences into 400 frames of continuous chunks, then allocate to each folds to balance the sum of COTS is almost the same. Allocating algorithm is greedy algorithm similar to [4]. I know it's leaky split, but I thought to balance the target counts is more important. That is because ensembling models with various performance produces poor prediction when using WBF according to the paper[5].</li>\n</ul>\n<h2>Data Loader</h2>\n<h3>Sampler</h3>\n<ul>\n<li>Class Balanced Sampling<ul>\n<li>I controled foreground/background ratio by implementing class-balanced sampler which is inspired by [1].<ul>\n<li>pos/neg ratio = 1:0.3</li>\n<li>1epoch = 1.3 * pos images</li></ul></li></ul></li>\n</ul>\n<h3>Augumentation</h3>\n<ul>\n<li>Heavy Distortion<ul>\n<li>Scale Shift: 0.3-1.7</li>\n<li>Mosaic, Mixup, perspective transform, color shift etc.</li></ul></li>\n<li>Perspective Transform with Elipse mask<ul>\n<li>In perspective transform, I used elipse mask which is shared by  <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> . It firmly reduced box loss when using with strong distortion.</li></ul></li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>input image size: 640px (x2 of original size)</li>\n<li>model: yolov5l[2]</li>\n<li>epochs: 45-60</li>\n<li>machine: Tesla T4, 16GB (EC2 instance <code>g4dn.xlarge</code>)</li>\n</ul>\n<h2>Inference</h2>\n<ul>\n<li>image size: 2560px (x2 of original size)<ul>\n<li>Note: even though training with tiles, YOLOv5 model accurately detects the object when input with entire image.</li></ul></li>\n<li>TTA<ul>\n<li>(scale, transform) = (0.7, None), (1.0, hflip), (1.3, None) [^1]</li></ul></li>\n<li>PostProcessing<ul>\n<li>Clear out low-confident, near-edge boxes<ul>\n<li>Since I trained the model with tiles, it's predictions is much more likely to predict boxes around edges compared to training with entire images. This causes bad effect when using Seq-NMS, because the model continue to predict FPs around the edges after the target object framed out.</li></ul></li>\n<li>Seq-NMS[3]<ul>\n<li>Original Seq-NMS algorithm is not designed for real-time processing. So I used past 20 frames with exponential decay of factor 0.9. Metrics is <code>max</code>.</li></ul></li></ul></li>\n</ul>\n<p>[^1]: original YOLOv5 TTA algorithm is kind of tricky, since the scale is set <code>(1.0, 0.83, 0.67)</code>. The mean scale is smaller than original size. In the official documentation, it says we should scale input image by x1.3 when using TTA. But this process is slower. So I fixed to <code>(0.7, 1.0, 1.3)</code>, adjusting mean value is equel to 1.0.</p>\n<h1>Experiment Result</h1>\n<p>Figures shown in the table are <strong>private LB</strong> scores.</p>\n<table>\n<thead>\n<tr>\n<th>Id</th>\n<th>Label</th>\n<th>Split</th>\n<th>model</th>\n<th>ensemble</th>\n<th>TTA</th>\n<th>w/o Seq-NMS</th>\n<th>w Seq-NMS</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>original</td>\n<td>all data</td>\n<td>yolov5l</td>\n<td>none</td>\n<td>scale(0.7, 1, 1.3)</td>\n<td>0.713</td>\n<td>0.722</td>\n</tr>\n<tr>\n<td>1</td>\n<td>original + pseudo label</td>\n<td>all data</td>\n<td>yolov5l</td>\n<td>none</td>\n<td>scale(0.7, 1, 1.3)</td>\n<td>0.723</td>\n<td>0.724</td>\n</tr>\n<tr>\n<td>2</td>\n<td>original</td>\n<td>frame chunks 10 fold</td>\n<td>yolov5l</td>\n<td>WBF(fold5-9)</td>\n<td>None</td>\n<td>0.726</td>\n<td><strong>0.737</strong> (late sub.)[7]</td>\n</tr>\n</tbody>\n</table>\n<h1>Discussion</h1>\n<h2>Seq-NMS</h2>\n<p>Seq-NMS gives me gain of <strong>2-5%</strong> when infer with most difficult validation sets (much bubbles and blurly images). So I showed this algorithms is robust to background noise.</p>\n<p>On the contraly, the effect of Seq-NMS is not the same on the LB scores. The Seq-NMS gains <strong>1.1%</strong> when using with the model 2, which is tie score to the 2nd team. However it gives only 0.1% gain when uses with the model 1.</p>\n<h1>Reference</h1>\n<ul>\n<li>[1] <a href=\"https://github.com/ufoym/imbalanced-dataset-sampler\" target=\"_blank\">https://github.com/ufoym/imbalanced-dataset-sampler</a></li>\n<li>[2] <a href=\"https://github.com/ultralytics/yolov5\" target=\"_blank\">https://github.com/ultralytics/yolov5</a></li>\n<li>[3] <a href=\"https://arxiv.org/abs/1602.08465\" target=\"_blank\">https://arxiv.org/abs/1602.08465</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/tatamikenn/balanced-fold-splitting-algorithm\" target=\"_blank\">https://www.kaggle.com/tatamikenn/balanced-fold-splitting-algorithm</a></li>\n<li>[5] <a href=\"https://arxiv.org/abs/1910.13302\" target=\"_blank\">https://arxiv.org/abs/1910.13302</a></li>\n<li>[6] <a href=\"https://openaccess.thecvf.com/content_CVPRW_2019/html/UAVision/Unel_The_Power_of_Tiling_for_Small_Object_Detection_CVPRW_2019_paper.html\" target=\"_blank\">https://openaccess.thecvf.com/content_CVPRW_2019/html/UAVision/Unel_The_Power_of_Tiling_for_Small_Object_Detection_CVPRW_2019_paper.html</a></li>\n<li>[7] <a href=\"https://www.kaggle.com/code/tatamikenn/infer-l-640-5x3-heavyx2-wbf-fold-5-9\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/infer-l-640-5x3-heavyx2-wbf-fold-5-9</a></li>\n</ul>\n<h1>Update Note</h1>\n<ul>\n<li>2022/02/16: updated experiment result: add result of no relabeled data to show the effect of relabeling.</li>\n</ul>",
  "messages": [
    {
      "id": 1692366,
      "postDate": "2022-02-16T02:05:48.710Z",
      "content": "<p>Much appreciated for organizing the competition, and thanks to all people who committed to this project.</p>\n<p>I will share my solution with you.</p>\n<h1>Pipeline</h1>\n<h2>Preprocessing</h2>\n<ul>\n<li>Tiled Training<ul>\n<li>Split original image into 320x320 sized tiles, with each tiles have intersections</li>\n<li>Stride is chosen as (sx, sy) = (240, 200), which gives us 5x3 tiles for each image. This setting is according to [6].</li></ul></li>\n<li>Relabeling<ul>\n<li>I noticed a lot of inconsistent labels and missing labels on the train dataset, so I thought relabeling will improve the model's performance.</li>\n<li>I relabeled with convining original lagbels and pseudo labels generated by trained model:<ul>\n<li>NMS(original(with 0.501) + pseudo label)</li>\n<li>threshold(conf&gt;0.5)</li></ul></li></ul></li>\n</ul>\n<h2>Fold Split</h2>\n<ul>\n<li>Split sequences into 400 frames of continuous chunks, then allocate to each folds to balance the sum of COTS is almost the same. Allocating algorithm is greedy algorithm similar to [4]. I know it's leaky split, but I thought to balance the target counts is more important. That is because ensembling models with various performance produces poor prediction when using WBF according to the paper[5].</li>\n</ul>\n<h2>Data Loader</h2>\n<h3>Sampler</h3>\n<ul>\n<li>Class Balanced Sampling<ul>\n<li>I controled foreground/background ratio by implementing class-balanced sampler which is inspired by [1].<ul>\n<li>pos/neg ratio = 1:0.3</li>\n<li>1epoch = 1.3 * pos images</li></ul></li></ul></li>\n</ul>\n<h3>Augumentation</h3>\n<ul>\n<li>Heavy Distortion<ul>\n<li>Scale Shift: 0.3-1.7</li>\n<li>Mosaic, Mixup, perspective transform, color shift etc.</li></ul></li>\n<li>Perspective Transform with Elipse mask<ul>\n<li>In perspective transform, I used elipse mask which is shared by  <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> . It firmly reduced box loss when using with strong distortion.</li></ul></li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>input image size: 640px (x2 of original size)</li>\n<li>model: yolov5l[2]</li>\n<li>epochs: 45-60</li>\n<li>machine: Tesla T4, 16GB (EC2 instance <code>g4dn.xlarge</code>)</li>\n</ul>\n<h2>Inference</h2>\n<ul>\n<li>image size: 2560px (x2 of original size)<ul>\n<li>Note: even though training with tiles, YOLOv5 model accurately detects the object when input with entire image.</li></ul></li>\n<li>TTA<ul>\n<li>(scale, transform) = (0.7, None), (1.0, hflip), (1.3, None) [^1]</li></ul></li>\n<li>PostProcessing<ul>\n<li>Clear out low-confident, near-edge boxes<ul>\n<li>Since I trained the model with tiles, it's predictions is much more likely to predict boxes around edges compared to training with entire images. This causes bad effect when using Seq-NMS, because the model continue to predict FPs around the edges after the target object framed out.</li></ul></li>\n<li>Seq-NMS[3]<ul>\n<li>Original Seq-NMS algorithm is not designed for real-time processing. So I used past 20 frames with exponential decay of factor 0.9. Metrics is <code>max</code>.</li></ul></li></ul></li>\n</ul>\n<p>[^1]: original YOLOv5 TTA algorithm is kind of tricky, since the scale is set <code>(1.0, 0.83, 0.67)</code>. The mean scale is smaller than original size. In the official documentation, it says we should scale input image by x1.3 when using TTA. But this process is slower. So I fixed to <code>(0.7, 1.0, 1.3)</code>, adjusting mean value is equel to 1.0.</p>\n<h1>Experiment Result</h1>\n<p>Figures shown in the table are <strong>private LB</strong> scores.</p>\n<table>\n<thead>\n<tr>\n<th>Id</th>\n<th>Label</th>\n<th>Split</th>\n<th>model</th>\n<th>ensemble</th>\n<th>TTA</th>\n<th>w/o Seq-NMS</th>\n<th>w Seq-NMS</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>original</td>\n<td>all data</td>\n<td>yolov5l</td>\n<td>none</td>\n<td>scale(0.7, 1, 1.3)</td>\n<td>0.713</td>\n<td>0.722</td>\n</tr>\n<tr>\n<td>1</td>\n<td>original + pseudo label</td>\n<td>all data</td>\n<td>yolov5l</td>\n<td>none</td>\n<td>scale(0.7, 1, 1.3)</td>\n<td>0.723</td>\n<td>0.724</td>\n</tr>\n<tr>\n<td>2</td>\n<td>original</td>\n<td>frame chunks 10 fold</td>\n<td>yolov5l</td>\n<td>WBF(fold5-9)</td>\n<td>None</td>\n<td>0.726</td>\n<td><strong>0.737</strong> (late sub.)[7]</td>\n</tr>\n</tbody>\n</table>\n<h1>Discussion</h1>\n<h2>Seq-NMS</h2>\n<p>Seq-NMS gives me gain of <strong>2-5%</strong> when infer with most difficult validation sets (much bubbles and blurly images). So I showed this algorithms is robust to background noise.</p>\n<p>On the contraly, the effect of Seq-NMS is not the same on the LB scores. The Seq-NMS gains <strong>1.1%</strong> when using with the model 2, which is tie score to the 2nd team. However it gives only 0.1% gain when uses with the model 1.</p>\n<h1>Reference</h1>\n<ul>\n<li>[1] <a href=\"https://github.com/ufoym/imbalanced-dataset-sampler\" target=\"_blank\">https://github.com/ufoym/imbalanced-dataset-sampler</a></li>\n<li>[2] <a href=\"https://github.com/ultralytics/yolov5\" target=\"_blank\">https://github.com/ultralytics/yolov5</a></li>\n<li>[3] <a href=\"https://arxiv.org/abs/1602.08465\" target=\"_blank\">https://arxiv.org/abs/1602.08465</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/tatamikenn/balanced-fold-splitting-algorithm\" target=\"_blank\">https://www.kaggle.com/tatamikenn/balanced-fold-splitting-algorithm</a></li>\n<li>[5] <a href=\"https://arxiv.org/abs/1910.13302\" target=\"_blank\">https://arxiv.org/abs/1910.13302</a></li>\n<li>[6] <a href=\"https://openaccess.thecvf.com/content_CVPRW_2019/html/UAVision/Unel_The_Power_of_Tiling_for_Small_Object_Detection_CVPRW_2019_paper.html\" target=\"_blank\">https://openaccess.thecvf.com/content_CVPRW_2019/html/UAVision/Unel_The_Power_of_Tiling_for_Small_Object_Detection_CVPRW_2019_paper.html</a></li>\n<li>[7] <a href=\"https://www.kaggle.com/code/tatamikenn/infer-l-640-5x3-heavyx2-wbf-fold-5-9\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/infer-l-640-5x3-heavyx2-wbf-fold-5-9</a></li>\n</ul>\n<h1>Update Note</h1>\n<ul>\n<li>2022/02/16: updated experiment result: add result of no relabeled data to show the effect of relabeling.</li>\n</ul>",
      "rawMarkdown": "Much appreciated for organizing the competition, and thanks to all people who committed to this project.\n\nI will share my solution with you.\n\n# Pipeline\n\n## Preprocessing\n\n* Tiled Training\n    * Split original image into 320x320 sized tiles, with each tiles have intersections\n    * Stride is chosen as (sx, sy) = (240, 200), which gives us 5x3 tiles for each image. This setting is according to [6].\n* Relabeling\n    * I noticed a lot of inconsistent labels and missing labels on the train dataset, so I thought relabeling will improve the model's performance.\n    * I relabeled with convining original lagbels and pseudo labels generated by trained model:\n        * NMS(original(with 0.501) + pseudo label)\n        * threshold(conf>0.5)\n\n## Fold Split\n\n* Split sequences into 400 frames of continuous chunks, then allocate to each folds to balance the sum of COTS is almost the same. Allocating algorithm is greedy algorithm similar to [4]. I know it's leaky split, but I thought to balance the target counts is more important. That is because ensembling models with various performance produces poor prediction when using WBF according to the paper[5].\n\n## Data Loader\n\n### Sampler\n\n* Class Balanced Sampling\n    * I controled foreground/background ratio by implementing class-balanced sampler which is inspired by [1].\n        * pos/neg ratio = 1:0.3\n        * 1epoch = 1.3 * pos images\n\n### Augumentation\n\n* Heavy Distortion\n    * Scale Shift: 0.3-1.7\n    * Mosaic, Mixup, perspective transform, color shift etc.\n* Perspective Transform with Elipse mask\n    * In perspective transform, I used elipse mask which is shared by  @hengck23 . It firmly reduced box loss when using with strong distortion.\n\n## Training\n\n* input image size: 640px (x2 of original size)\n* model: yolov5l[2]\n* epochs: 45-60\n* machine: Tesla T4, 16GB (EC2 instance `g4dn.xlarge`)\n\n## Inference\n\n* image size: 2560px (x2 of original size)\n    * Note: even though training with tiles, YOLOv5 model accurately detects the object when input with entire image.\n* TTA\n    * (scale, transform) = (0.7, None), (1.0, hflip), (1.3, None) [^1]\n* PostProcessing\n    * Clear out low-confident, near-edge boxes\n      * Since I trained the model with tiles, it's predictions is much more likely to predict boxes around edges compared to training with entire images. This causes bad effect when using Seq-NMS, because the model continue to predict FPs around the edges after the target object framed out.\n    * Seq-NMS[3]\n      * Original Seq-NMS algorithm is not designed for real-time processing. So I used past 20 frames with exponential decay of factor 0.9. Metrics is `max`.\n\n[^1]: original YOLOv5 TTA algorithm is kind of tricky, since the scale is set `(1.0, 0.83, 0.67)`. The mean scale is smaller than original size. In the official documentation, it says we should scale input image by x1.3 when using TTA. But this process is slower. So I fixed to `(0.7, 1.0, 1.3)`, adjusting mean value is equel to 1.0.\n\n\n# Experiment Result\n\nFigures shown in the table are **private LB** scores.\n\n| Id  | Label                   | Split                | model   | ensemble     | TTA                | w/o Seq-NMS | w Seq-NMS                          |\n| --- | ----------------------- | -------------------- | ------- | ------------ | ------------------ | ---------- | -------------------------------- |\n| 0   |  original                | all data             | yolov5l | none         | scale(0.7, 1, 1.3) | 0.713      | 0.722                           |\n| 1   | original + pseudo label | all data             | yolov5l | none         | scale(0.7, 1, 1.3) | 0.723      | 0.724                            |\n| 2   | original                | frame chunks 10 fold | yolov5l | WBF(fold5-9) | None               | 0.726      | <ins>**0.737**</ins> (late sub.)[7] |\n\n# Discussion\n\n## Seq-NMS\n\nSeq-NMS gives me gain of **2-5%** when infer with most difficult validation sets (much bubbles and blurly images). So I showed this algorithms is robust to background noise.\n\nOn the contraly, the effect of Seq-NMS is not the same on the LB scores. The Seq-NMS gains **1.1%** when using with the model 2, which is tie score to the 2nd team. However it gives only 0.1% gain when uses with the model 1.\n\n# Reference\n\n* [1] https://github.com/ufoym/imbalanced-dataset-sampler\n* [2] https://github.com/ultralytics/yolov5\n* [3] https://arxiv.org/abs/1602.08465\n* [4] https://www.kaggle.com/tatamikenn/balanced-fold-splitting-algorithm\n* [5] https://arxiv.org/abs/1910.13302\n* [6] https://openaccess.thecvf.com/content_CVPRW_2019/html/UAVision/Unel_The_Power_of_Tiling_for_Small_Object_Detection_CVPRW_2019_paper.html\n* [7] https://www.kaggle.com/code/tatamikenn/infer-l-640-5x3-heavyx2-wbf-fold-5-9\n\n# Update Note\n\n* 2022/02/16: updated experiment result: add result of no relabeled data to show the effect of relabeling.",
      "votes": 24
    },
    {
      "id": 1692418,
      "postDate": "2022-02-16T02:56:20.403Z",
      "content": "<p>I also published Seq-NMS code.<br>\n<a href=\"https://www.kaggle.com/tatamikenn/example-code-of-seq-nms\" target=\"_blank\">https://www.kaggle.com/tatamikenn/example-code-of-seq-nms</a></p>",
      "rawMarkdown": "I also published Seq-NMS code.\nhttps://www.kaggle.com/tatamikenn/example-code-of-seq-nms",
      "votes": 3
    },
    {
      "id": 1706357,
      "postDate": "2022-02-27T12:26:41.073Z",
      "content": "<p>congratulate! thanks for sharing your solution. can you further explan the elipse mask  which is new for me? or maybe you can share the link from hengck23. Thanks!</p>",
      "rawMarkdown": "congratulate! thanks for sharing your solution. can you further explan the elipse mask  which is new for me? or maybe you can share the link from hengck23. Thanks!",
      "replies": [
        {
          "id": 1706477,
          "postDate": "2022-02-27T14:11:29.050Z",
          "content": "<p>This is the link for his/her source code:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405</a></p>",
          "rawMarkdown": "This is the link for his/her source code:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405"
        },
        {
          "id": 1706489,
          "postDate": "2022-02-27T14:17:47.390Z",
          "content": "<p>For a quick explanation, most of the target (i.e. COTS) shapes are approximately ellipse which is inscribed to the boxes.</p>\n<p>YOLOv5's default mask is rectangle shape.<br>\nThis will loosen the boxes tightness if applied by heavy geometric distortions (e.g. rotations).<br>\nOn the contrary, using mask of ellipse shape, the tightness of the boxes are almost the same even if applied by heavy distortions.</p>",
          "rawMarkdown": "For a quick explanation, most of the target (i.e. COTS) shapes are approximately ellipse which is inscribed to the boxes.\n\nYOLOv5's default mask is rectangle shape.\nThis will loosen the boxes tightness if applied by heavy geometric distortions (e.g. rotations).\nOn the contrary, using mask of ellipse shape, the tightness of the boxes are almost the same even if applied by heavy distortions."
        },
        {
          "id": 1706518,
          "postDate": "2022-02-27T14:32:04.193Z",
          "content": "<p>Like the picture below:<br>\n<a href=\"https://ibb.co/Z6W2x8x\"><img src=\"https://i.ibb.co/kgQmX8X/Screen-Shot-2022-02-27-at-23-30-11.png\" alt=\"Screen-Shot-2022-02-27-at-23-30-11\"></a></p>",
          "rawMarkdown": "Like the picture below:\n<a href=\"https://ibb.co/Z6W2x8x\"><img src=\"https://i.ibb.co/kgQmX8X/Screen-Shot-2022-02-27-at-23-30-11.png\" alt=\"Screen-Shot-2022-02-27-at-23-30-11\" border=\"0\"></a>"
        },
        {
          "id": 1706535,
          "postDate": "2022-02-27T14:42:24.623Z",
          "content": "<p>Of course this trick is only applicable if the target shape is close to ellipse.<br>\nIf the target shape are close to rectangle (e.g. bus, car etc.), ellipse mask will generate over-tighten boxes after rotation.</p>",
          "rawMarkdown": "Of course this trick is only applicable if the target shape is close to ellipse.\nIf the target shape are close to rectangle (e.g. bus, car etc.), ellipse mask will generate over-tighten boxes after rotation."
        }
      ]
    },
    {
      "id": 1692425,
      "postDate": "2022-02-16T03:11:55.293Z",
      "content": "<p>I like your work, it is very skillful.</p>",
      "rawMarkdown": "I like your work, it is very skillful.",
      "replies": [
        {
          "id": 1692429,
          "postDate": "2022-02-16T03:22:46.610Z",
          "content": "<p>For me, I just customized some open source implementations, so I think it’s not as skillful as you thought. But thanks for saying so.</p>",
          "rawMarkdown": "For me, I just customized some open source implementations, so I think it’s not as skillful as you thought. But thanks for saying so."
        },
        {
          "id": 1692437,
          "postDate": "2022-02-16T03:28:45.460Z",
          "content": "<p>It really impressed me, I am new here and want to learn some skills like this. 😁😁</p>",
          "rawMarkdown": "It really impressed me, I am new here and want to learn some skills like this. 😁😁",
          "votes": 1
        }
      ]
    },
    {
      "id": 1692931,
      "postDate": "2022-02-16T10:53:55.683Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1692418,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-02-16T02:56:20.403000",
      "content": "<p>I also published Seq-NMS code.<br>\n<a href=\"https://www.kaggle.com/tatamikenn/example-code-of-seq-nms\" target=\"_blank\">https://www.kaggle.com/tatamikenn/example-code-of-seq-nms</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1706357,
      "author_name": "Kevin",
      "author_url": "",
      "post_date": "2022-02-27T12:26:41.073000",
      "content": "<p>congratulate! thanks for sharing your solution. can you further explan the elipse mask  which is new for me? or maybe you can share the link from hengck23. Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1706477,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-27T14:11:29.050000",
          "content": "<p>This is the link for his/her source code:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300405</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1706489,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-27T14:17:47.390000",
          "content": "<p>For a quick explanation, most of the target (i.e. COTS) shapes are approximately ellipse which is inscribed to the boxes.</p>\n<p>YOLOv5's default mask is rectangle shape.<br>\nThis will loosen the boxes tightness if applied by heavy geometric distortions (e.g. rotations).<br>\nOn the contrary, using mask of ellipse shape, the tightness of the boxes are almost the same even if applied by heavy distortions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1706518,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-27T14:32:04.193000",
          "content": "<p>Like the picture below:<br>\n<a href=\"https://ibb.co/Z6W2x8x\"><img src=\"https://i.ibb.co/kgQmX8X/Screen-Shot-2022-02-27-at-23-30-11.png\" alt=\"Screen-Shot-2022-02-27-at-23-30-11\"></a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1706535,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-27T14:42:24.623000",
          "content": "<p>Of course this trick is only applicable if the target shape is close to ellipse.<br>\nIf the target shape are close to rectangle (e.g. bus, car etc.), ellipse mask will generate over-tighten boxes after rotation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1692425,
      "author_name": "Good Moon",
      "author_url": "",
      "post_date": "2022-02-16T03:11:55.293000",
      "content": "<p>I like your work, it is very skillful.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1692429,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-16T03:22:46.610000",
          "content": "<p>For me, I just customized some open source implementations, so I think it’s not as skillful as you thought. But thanks for saying so.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1692437,
          "author_name": "Good Moon",
          "author_url": "",
          "post_date": "2022-02-16T03:28:45.460000",
          "content": "<p>It really impressed me, I am new here and want to learn some skills like this. 😁😁</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1692931,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-16T10:53:55.683000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1692366": "Much appreciated for organizing the competition, and thanks to all people who committed to this project.\n\nI will share my solution with you.\n\n# Pipeline\n\n## Preprocessing\n\n* Tiled Training\n    * Split original image into 320x320 sized tiles, with each tiles have intersections\n    * Stride is chosen as (sx, sy) = (240, 200), which gives us 5x3 tiles for each image. This setting is according to [6].\n* Relabeling\n    * I noticed a lot of inconsistent labels and missing labels on the train dataset, so I thought relabeling will improve the model's performance.\n    * I relabeled with convining original lagbels and pseudo labels generated by trained model:\n        * NMS(original(with 0.501) + pseudo label)\n        * threshold(conf>0.5)\n\n## Fold Split\n\n* Split sequences into 400 frames of continuous chunks, then allocate to each folds to balance the sum of COTS is almost the same. Allocating algorithm is greedy algorithm similar to [4]. I know it's leaky split, but I thought to balance the target counts is more important. That is because ensembling models with various performance produces poor prediction when using WBF according to the paper[5].\n\n## Data Loader\n\n### Sampler\n\n* Class Balanced Sampling\n    * I controled foreground/background ratio by implementing class-balanced sampler which is inspired by [1].\n        * pos/neg ratio = 1:0.3\n        * 1epoch = 1.3 * pos images\n\n### Augumentation\n\n* Heavy Distortion\n    * Scale Shift: 0.3-1.7\n    * Mosaic, Mixup, perspective transform, color shift etc.\n* Perspective Transform with Elipse mask\n    * In perspective transform, I used elipse mask which is shared by  @hengck23 . It firmly reduced box loss when using with strong distortion.\n\n## Training\n\n* input image size: 640px (x2 of original size)\n* model: yolov5l[2]\n* epochs: 45-60\n* machine: Tesla T4, 16GB (EC2 instance `g4dn.xlarge`)\n\n## Inference\n\n* image size: 2560px (x2 of original size)\n    * Note: even though training with tiles, YOLOv5 model accurately detects the object when input with entire image.\n* TTA\n    * (scale, transform) = (0.7, None), (1.0, hflip), (1.3, None) [^1]\n* PostProcessing\n    * Clear out low-confident, near-edge boxes\n      * Since I trained the model with tiles, it's predictions is much more likely to predict boxes around edges compared to training with entire images. This causes bad effect when using Seq-NMS, because the model continue to predict FPs around the edges after the target object framed out.\n    * Seq-NMS[3]\n      * Original Seq-NMS algorithm is not designed for real-time processing. So I used past 20 frames with exponential decay of factor 0.9. Metrics is `max`.\n\n[^1]: original YOLOv5 TTA algorithm is kind of tricky, since the scale is set `(1.0, 0.83, 0.67)`. The mean scale is smaller than original size. In the official documentation, it says we should scale input image by x1.3 when using TTA. But this process is slower. So I fixed to `(0.7, 1.0, 1.3)`, adjusting mean value is equel to 1.0.\n\n\n# Experiment Result\n\nFigures shown in the table are **private LB** scores.\n\n| Id  | Label                   | Split                | model   | ensemble     | TTA                | w/o Seq-NMS | w Seq-NMS                          |\n| --- | ----------------------- | -------------------- | ------- | ------------ | ------------------ | ---------- | -------------------------------- |\n| 0   |  original                | all data             | yolov5l | none         | scale(0.7, 1, 1.3) | 0.713      | 0.722                           |\n| 1   | original + pseudo label | all data             | yolov5l | none         | scale(0.7, 1, 1.3) | 0.723      | 0.724                            |\n| 2   | original                | frame chunks 10 fold | yolov5l | WBF(fold5-9) | None               | 0.726      | <ins>**0.737**</ins> (late sub.)[7] |\n\n# Discussion\n\n## Seq-NMS\n\nSeq-NMS gives me gain of **2-5%** when infer with most difficult validation sets (much bubbles and blurly images). So I showed this algorithms is robust to background noise.\n\nOn the contraly, the effect of Seq-NMS is not the same on the LB scores. The Seq-NMS gains **1.1%** when using with the model 2, which is tie score to the 2nd team. However it gives only 0.1% gain when uses with the model 1.\n\n# Reference\n\n* [1] https://github.com/ufoym/imbalanced-dataset-sampler\n* [2] https://github.com/ultralytics/yolov5\n* [3] https://arxiv.org/abs/1602.08465\n* [4] https://www.kaggle.com/tatamikenn/balanced-fold-splitting-algorithm\n* [5] https://arxiv.org/abs/1910.13302\n* [6] https://openaccess.thecvf.com/content_CVPRW_2019/html/UAVision/Unel_The_Power_of_Tiling_for_Small_Object_Detection_CVPRW_2019_paper.html\n* [7] https://www.kaggle.com/code/tatamikenn/infer-l-640-5x3-heavyx2-wbf-fold-5-9\n\n# Update Note\n\n* 2022/02/16: updated experiment result: add result of no relabeled data to show the effect of relabeling.",
    "1692418": "I also published Seq-NMS code.\nhttps://www.kaggle.com/tatamikenn/example-code-of-seq-nms",
    "1706357": "congratulate! thanks for sharing your solution. can you further explan the elipse mask  which is new for me? or maybe you can share the link from hengck23. Thanks!",
    "1692425": "I like your work, it is very skillful.",
    "1692931": ""
  }
}