{
  "id": 307609,
  "title": "69th Place - WBF with Threshold",
  "url": "/competitions/tensorflow-great-barrier-reef/writeups/chris-deotte-69th-place-wbf-with-threshold",
  "author_name": "",
  "post_date": "2022-02-18T00:44:15.687Z",
  "votes": 93,
  "comment_count": 52,
  "views": 0,
  "content": "<h1>Single Model Yolov5L6</h1>\n<p>Thanks Kaggle for a fun competition. Thank you Kagglers for sharing many great discussions and notebooks! My final solution is a single 10-fold Yolo5L6 without TTA without tracking. Train 3072, infer 3072 img size.</p>\n<h1>WBF with Threshold</h1>\n<p>Just this morning, I realized that the awesome WBF GitHub repository <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">here</a> doesn't apply a threshold after fusing boxes. So with 10 hours until competition deadline, I updated my final submission and submitted again! It finished just in time and jumped me into Silver!</p>\n<p>If you have 10 fold models. And each model is inferred box with <code>conf = 0.1</code>. Then afterward when we apply WBF, it is possible that 3 models find a starfish with confidences <code>0.3, 0.2, 0.1</code> and the other 7 fold models do not find that box. Then the average confidence is <code>0.06 = (0.3 + 0.2 + 0.1 + 0 + 0 + 0 + 0 + 0 + 0 + 0)/10</code>. So we need to remove this box (if we don't wish to have boxes under <code>conf = 0.1</code>) as follows:</p>\n<pre><code>WBF_CONF = 0.1\nboxes, scores, labels = weighted_boxes_fusion(ALL_MODELS)\nfiltered boxes=[], filtered_scores=[]\nfor k, box in enumerate(boxes):\n    if scores[k]&lt;WBF_CONF: continue\n    filtered_boxes.append(box)\n    filtered_scores.append(scores[k])\n</code></pre>\n<p>Using this code boosted my final submission Private LB 0.609 to <strong>Private LB 0.684</strong>! Wow. This is important to do if the competition metric is F2 because we are penalized with too many FP (so it's important to establish a conf threshold). However if the competition metric is <code>mAP</code> then it would not matter since with <code>mAP</code> we are not penalized for FP (explained <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">here</a>)</p>\n<h1>Skip_Box_Thr is not WBF_CONF</h1>\n<p>Note that <code>weighted_boxes_fusion</code> has a parameter called <code>skip_box_thr</code>. This filters input boxes not output boxes. In general, this parameter makes no difference because we already use <code>model.conf = 0.1</code> which means that we only input boxes with <code>conf&gt;=0.1</code> into our WBF. There is no parameter for filtering output boxes. So when competition metric is <code>F1</code> or <code>F2</code>, we must write our own code like above.</p>\n<h1>Model Training - 10 Folds</h1>\n<p>For my object detection model, I trained 10 folds of <code>Yolov5L6</code> with <code>image size = 3072</code>, <code>batch_size = 8</code>, and <code>epochs = 10</code> using 4xV100 32GB GPU. Thanks Nvidia! Training took 2.5 hours per fold (i.e. 15 minutes per epoch). I tried a few image sizes and this achieved the best CV score.</p>\n<p>I used <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> training script <a href=\"https://www.kaggle.com/steamedsheep/yolov5-high-resolution-training\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/hjulian3833\" target=\"_blank\">@hjulian3833</a> subsequence CV folds <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">here</a>. I used the best model each epoch as determined by Yolo F2 metric described <a href=\"https://www.kaggle.com/sanchitvj\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638#1658343\" target=\"_blank\">here</a>. I trained using positive frames only.</p>\n<h1>Model Inference - 10 Folds</h1>\n<p>For inference, I inferred 10 fold models with <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> notebook <a href=\"https://www.kaggle.com/awsaf49/great-barrier-reef-yolov5-infer\" target=\"_blank\">here</a>. I inferred at <code>image size = 3072</code> with <code>iou = 0.4</code> and <code>conf = 0.05</code>. I inferred at the same size as training since this achieved the best CV score. I ensembled fold models using <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> awesome WBF GitHub library <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">here</a> and I added my own WBF threshold shown above at <code>WBF_CONF = 0.1</code> (since this was best <code>conf</code> for CV score).</p>\n<h1>No Tracking, No TTA</h1>\n<p>Using the public notebook tracking code did not improve my CV score, so I did not use it. In order to infer 10 folds of Yolov5L6 in 7 hours, I did not use TTA.</p>\n<h1>CV 0.732, Public LB 0.591, Private LB 0.684</h1>\n<p>To compute competition metric, I used <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> notebook <a href=\"https://www.kaggle.com/bamps53/competition-metric-implementation\" target=\"_blank\">here</a>. My final single model achieved 10 fold CV 0.732, Public LB 0.591, and Private LB 0.684</p>\n<h1>Thank You</h1>\n<p>Thank you everyone for sharing helpful discussions and notebooks. My solution was possible because of everyone's generous sharing.</p>\n<h1>UPDATE 1 (after comp ended):</h1>\n<p>Instead of inferring 10 folds. I just submitted with 5 of the 10 folds and added TTA. The public LB boost 0.018 and the private LB boost 0.011. The private rank boost from 75th place to 40th place! </p>\n<h1>UPDATE 2 (after comp ended):</h1>\n<p>Instead of adding TTA, I trained a Yolov5M6 at 3072 with same settings as my Yolov5L6. Then I submit 5 folds out of 10 Yolov5M6 (without TTA) and 5 folds out of 10 Yolo5L6 (without TTA). When using <code>WBF_CONF</code>, the result is private LB 0.709 and 20th place! This confirms that <code>WBF_CONF</code> is a powerful idea. If I add even more models (with different backbones and use different image sizes), I assume the LB will keep climbing!</p>",
  "messages": [
    {
      "id": "1690489",
      "postDate": "02/15/2022 00:55:13",
      "content": "<h1>Single Model Yolov5L6</h1>\n<p>Thanks Kaggle for a fun competition. Thank you Kagglers for sharing many great discussions and notebooks! My final solution is a single 10-fold Yolo5L6 without TTA without tracking. Train 3072, infer 3072 img size.</p>\n<h1>WBF with Threshold</h1>\n<p>Just this morning, I realized that the awesome WBF GitHub repository <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">here</a> doesn't apply a threshold after fusing boxes. So with 10 hours until competition deadline, I updated my final submission and submitted again! It finished just in time and jumped me into Silver!</p>\n<p>If you have 10 fold models. And each model is inferred box with <code>conf = 0.1</code>. Then afterward when we apply WBF, it is possible that 3 models find a starfish with confidences <code>0.3, 0.2, 0.1</code> and the other 7 fold models do not find that box. Then the average confidence is <code>0.06 = (0.3 + 0.2 + 0.1 + 0 + 0 + 0 + 0 + 0 + 0 + 0)/10</code>. So we need to remove this box (if we don't wish to have boxes under <code>conf = 0.1</code>) as follows:</p>\n<pre><code>WBF_CONF = 0.1\nboxes, scores, labels = weighted_boxes_fusion(ALL_MODELS)\nfiltered boxes=[], filtered_scores=[]\nfor k, box in enumerate(boxes):\n    if scores[k]&lt;WBF_CONF: continue\n    filtered_boxes.append(box)\n    filtered_scores.append(scores[k])\n</code></pre>\n<p>Using this code boosted my final submission Private LB 0.609 to <strong>Private LB 0.684</strong>! Wow. This is important to do if the competition metric is F2 because we are penalized with too many FP (so it's important to establish a conf threshold). However if the competition metric is <code>mAP</code> then it would not matter since with <code>mAP</code> we are not penalized for FP (explained <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">here</a>)</p>\n<h1>Skip_Box_Thr is not WBF_CONF</h1>\n<p>Note that <code>weighted_boxes_fusion</code> has a parameter called <code>skip_box_thr</code>. This filters input boxes not output boxes. In general, this parameter makes no difference because we already use <code>model.conf = 0.1</code> which means that we only input boxes with <code>conf&gt;=0.1</code> into our WBF. There is no parameter for filtering output boxes. So when competition metric is <code>F1</code> or <code>F2</code>, we must write our own code like above.</p>\n<h1>Model Training - 10 Folds</h1>\n<p>For my object detection model, I trained 10 folds of <code>Yolov5L6</code> with <code>image size = 3072</code>, <code>batch_size = 8</code>, and <code>epochs = 10</code> using 4xV100 32GB GPU. Thanks Nvidia! Training took 2.5 hours per fold (i.e. 15 minutes per epoch). I tried a few image sizes and this achieved the best CV score.</p>\n<p>I used <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> training script <a href=\"https://www.kaggle.com/steamedsheep/yolov5-high-resolution-training\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/hjulian3833\" target=\"_blank\">@hjulian3833</a> subsequence CV folds <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">here</a>. I used the best model each epoch as determined by Yolo F2 metric described <a href=\"https://www.kaggle.com/sanchitvj\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638#1658343\" target=\"_blank\">here</a>. I trained using positive frames only.</p>\n<h1>Model Inference - 10 Folds</h1>\n<p>For inference, I inferred 10 fold models with <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> notebook <a href=\"https://www.kaggle.com/awsaf49/great-barrier-reef-yolov5-infer\" target=\"_blank\">here</a>. I inferred at <code>image size = 3072</code> with <code>iou = 0.4</code> and <code>conf = 0.05</code>. I inferred at the same size as training since this achieved the best CV score. I ensembled fold models using <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> awesome WBF GitHub library <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">here</a> and I added my own WBF threshold shown above at <code>WBF_CONF = 0.1</code> (since this was best <code>conf</code> for CV score).</p>\n<h1>No Tracking, No TTA</h1>\n<p>Using the public notebook tracking code did not improve my CV score, so I did not use it. In order to infer 10 folds of Yolov5L6 in 7 hours, I did not use TTA.</p>\n<h1>CV 0.732, Public LB 0.591, Private LB 0.684</h1>\n<p>To compute competition metric, I used <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> notebook <a href=\"https://www.kaggle.com/bamps53/competition-metric-implementation\" target=\"_blank\">here</a>. My final single model achieved 10 fold CV 0.732, Public LB 0.591, and Private LB 0.684</p>\n<h1>Thank You</h1>\n<p>Thank you everyone for sharing helpful discussions and notebooks. My solution was possible because of everyone's generous sharing.</p>\n<h1>UPDATE 1 (after comp ended):</h1>\n<p>Instead of inferring 10 folds. I just submitted with 5 of the 10 folds and added TTA. The public LB boost 0.018 and the private LB boost 0.011. The private rank boost from 75th place to 40th place! </p>\n<h1>UPDATE 2 (after comp ended):</h1>\n<p>Instead of adding TTA, I trained a Yolov5M6 at 3072 with same settings as my Yolov5L6. Then I submit 5 folds out of 10 Yolov5M6 (without TTA) and 5 folds out of 10 Yolo5L6 (without TTA). When using <code>WBF_CONF</code>, the result is private LB 0.709 and 20th place! This confirms that <code>WBF_CONF</code> is a powerful idea. If I add even more models (with different backbones and use different image sizes), I assume the LB will keep climbing!</p>",
      "rawMarkdown": "# Single Model Yolov5L6\nThanks Kaggle for a fun competition. Thank you Kagglers for sharing many great discussions and notebooks! My final solution is a single 10-fold Yolo5L6 without TTA without tracking. Train 3072, infer 3072 img size.\n\n# WBF with Threshold\nJust this morning, I realized that the awesome WBF GitHub repository [here][1] doesn't apply a threshold after fusing boxes. So with 10 hours until competition deadline, I updated my final submission and submitted again! It finished just in time and jumped me into Silver!\n\nIf you have 10 fold models. And each model is inferred box with `conf = 0.1`. Then afterward when we apply WBF, it is possible that 3 models find a starfish with confidences `0.3, 0.2, 0.1` and the other 7 fold models do not find that box. Then the average confidence is `0.06 = (0.3 + 0.2 + 0.1 + 0 + 0 + 0 + 0 + 0 + 0 + 0)/10`. So we need to remove this box (if we don't wish to have boxes under `conf = 0.1`) as follows:\n\n    WBF_CONF = 0.1\n    boxes, scores, labels = weighted_boxes_fusion(ALL_MODELS)\n    filtered boxes=[], filtered_scores=[]\n    for k, box in enumerate(boxes):\n        if scores[k]<WBF_CONF: continue\n        filtered_boxes.append(box)\n        filtered_scores.append(scores[k])\n\nUsing this code boosted my final submission Private LB 0.609 to **Private LB 0.684**! Wow. This is important to do if the competition metric is F2 because we are penalized with too many FP (so it's important to establish a conf threshold). However if the competition metric is `mAP` then it would not matter since with `mAP` we are not penalized for FP (explained [here][8])\n\n# Skip_Box_Thr is not WBF_CONF\nNote that `weighted_boxes_fusion` has a parameter called `skip_box_thr`. This filters input boxes not output boxes. In general, this parameter makes no difference because we already use `model.conf = 0.1` which means that we only input boxes with `conf>=0.1` into our WBF. There is no parameter for filtering output boxes. So when competition metric is `F1` or `F2`, we must write our own code like above.\n\n# Model Training - 10 Folds\nFor my object detection model, I trained 10 folds of `Yolov5L6` with `image size = 3072`, `batch_size = 8`, and `epochs = 10` using 4xV100 32GB GPU. Thanks Nvidia! Training took 2.5 hours per fold (i.e. 15 minutes per epoch). I tried a few image sizes and this achieved the best CV score.\n\nI used @steamedsheep training script [here][2] and @hjulian3833 subsequence CV folds [here][3]. I used the best model each epoch as determined by Yolo F2 metric described [here][4] and [here][5]. I trained using positive frames only.\n\n# Model Inference - 10 Folds\nFor inference, I inferred 10 fold models with @awsaf49 notebook [here][6]. I inferred at `image size = 3072` with `iou = 0.4` and `conf = 0.05`. I inferred at the same size as training since this achieved the best CV score. I ensembled fold models using @zfturbo awesome WBF GitHub library [here][1] and I added my own WBF threshold shown above at `WBF_CONF = 0.1` (since this was best `conf` for CV score).\n\n# No Tracking, No TTA\nUsing the public notebook tracking code did not improve my CV score, so I did not use it. In order to infer 10 folds of Yolov5L6 in 7 hours, I did not use TTA.\n\n# CV 0.732, Public LB 0.591, Private LB 0.684\nTo compute competition metric, I used @bamps53 notebook [here][7]. My final single model achieved 10 fold CV 0.732, Public LB 0.591, and Private LB 0.684\n\n# Thank You\nThank you everyone for sharing helpful discussions and notebooks. My solution was possible because of everyone's generous sharing.\n\n# UPDATE 1 (after comp ended): \nInstead of inferring 10 folds. I just submitted with 5 of the 10 folds and added TTA. The public LB boost 0.018 and the private LB boost 0.011. The private rank boost from 75th place to 40th place! \n\n# UPDATE 2 (after comp ended): \nInstead of adding TTA, I trained a Yolov5M6 at 3072 with same settings as my Yolov5L6. Then I submit 5 folds out of 10 Yolov5M6 (without TTA) and 5 folds out of 10 Yolo5L6 (without TTA). When using `WBF_CONF`, the result is private LB 0.709 and 20th place! This confirms that `WBF_CONF` is a powerful idea. If I add even more models (with different backbones and use different image sizes), I assume the LB will keep climbing!\n\n[1]: https://github.com/ZFTurbo/Weighted-Boxes-Fusion\n[2]: https://www.kaggle.com/steamedsheep/yolov5-high-resolution-training\n[3]: https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\n[4]: https://www.kaggle.com/sanchitvj\n[5]: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638#1658343\n[6]: https://www.kaggle.com/awsaf49/great-barrier-reef-yolov5-infer\n[7]: https://www.kaggle.com/bamps53/competition-metric-implementation\n[8]: https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637",
      "votes": null
    },
    {
      "id": "1690507",
      "postDate": "02/15/2022 01:11:39",
      "content": "<p>Our team has the same idea and same code. We set WBF_CONF = 0.5, so it refinds all prediction. We use 9 different sizes for x9 TTA, and I think the model generates different predictions at different sizes, so refining predictions after using WBF is not appropriate. Thank you for disclosing my mistake and unfinished idea. </p>",
      "rawMarkdown": "Our team has the same idea and same code. We set WBF_CONF = 0.5, so it refinds all prediction. We use 9 different sizes for x9 TTA, and I think the model generates different predictions at different sizes, so refining predictions after using WBF is not appropriate. Thank you for disclosing my mistake and unfinished idea.",
      "votes": null
    },
    {
      "id": "1690510",
      "postDate": "02/15/2022 01:12:57",
      "content": "<p>As always, thanks for sharing your approach <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! . I always learn something from reading solutions. </p>",
      "rawMarkdown": "As always, thanks for sharing your approach @cdeotte ! . I always learn something from reading solutions.",
      "votes": null
    },
    {
      "id": "1690514",
      "postDate": "02/15/2022 01:15:23",
      "content": "<p>Did you train on total data or just data having cots?</p>",
      "rawMarkdown": "Did you train on total data or just data having cots?",
      "votes": null
    },
    {
      "id": "1690531",
      "postDate": "02/15/2022 01:46:22",
      "content": "<p>Thank you for the concise summary. I have one question.</p>\n<blockquote>\n  <p>Using this code boosted my final submission Private LB 0.609 to Private LB 0.684! Wow. This is important to do if the competition metric is F2 because we are penalized with too many FN (so it's important to establish a conf threshold). However if the competition metric is mAP then it would not matter since with mAP we are not penalized for FN (explained here)</p>\n</blockquote>\n<p>Can you please explain more about this? Why does setting the WBF threshold lead to an improvement in FN? Since the number of bboxes is reduced, it seems that FN is increased but not decreased. Doesn't it mean that the F2 score improves by decreasing the FP?</p>",
      "rawMarkdown": "Thank you for the concise summary. I have one question.\n\n> Using this code boosted my final submission Private LB 0.609 to Private LB 0.684! Wow. This is important to do if the competition metric is F2 because we are penalized with too many FN (so it's important to establish a conf threshold). However if the competition metric is mAP then it would not matter since with mAP we are not penalized for FN (explained here)\n\nCan you please explain more about this? Why does setting the WBF threshold lead to an improvement in FN? Since the number of bboxes is reduced, it seems that FN is increased but not decreased. Doesn't it mean that the F2 score improves by decreasing the FP?",
      "votes": null
    },
    {
      "id": "1690532",
      "postDate": "02/15/2022 01:46:28",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thanks for sharing! </p>\n<p>My approach is very similar to your single YOLOv5 10-folds model with WBF, but using <code>yolov5m6</code> trained with <code>2560</code> resolution and heavy augmentation (<strong>Private LB: 0.690</strong>):</p>\n<pre><code>[\n    A.RandomSizedBBoxSafeCrop(1280, 1280, erosion_rate=0.2,\n        interpolation=cv2.INTER_LINEAR, p=1.0),\n    A.OneOf([\n        A.MotionBlur(p=.2),\n        A.MedianBlur(blur_limit=3, p=0.3),\n        A.Blur(blur_limit=3, p=0.1),\n    ], p=0.3),\n    A.OneOf([\n        A.CLAHE(clip_limit=2, tile_grid_size=(8, 8), p=0.5),\n        A.RandomBrightnessContrast(brightness_limit=0.2,\n            contrast_limit=0.2, brightness_by_max=False, p=0.5),\n        A.RandomGamma(p = 0.5),\n        A.ShiftScaleRotate(shift_limit=0.0625,\n            scale_limit=0.2, rotate_limit=45, p=0.2),\n    ], p=0.3),\n    A.HorizontalFlip(p=0.5),\n    A.Sharpen(p=0.3),\n    A.RGBShift(p=0.2),\n]\n</code></pre>\n<p>I think the key element would be using <code>RandomSizedBBoxSafeCrop</code> to crop <code>1280x1280</code> areas from <code>2560</code> resolution with <code>erosion_rate=0.2</code>, which helped the model to learn to identify smaller objects. I selected the best F2 version (train about <code>11-20 epochs</code> per fold, with extra <code>20%</code> background images).</p>\n<p>Inference is done at the same size (2560) with <code>TTA</code> in <code>7/10 folds</code> to squeeze time spent (no tracking). As for NMS <code>conf</code> and <code>IOU</code> of the model, they were set to <code>0.5</code> and <code>0.45</code> to keep more bboxes with higher confidence, then for WBF I set <code>iou_thr=0.4</code> and <code>skip_box_thr=0.01</code>.</p>\n<p><a href=\"https://www.kaggle.com/markpeng/yolov5m6-f10-v8-2560-wbf-ensemble/settings?scriptVersionId=87258324\" target=\"_blank\">https://www.kaggle.com/markpeng/yolov5m6-f10-v8-2560-wbf-ensemble/settings?scriptVersionId=87258324</a></p>\n<p>Tried training by tiling images under x2/x4/x8 smaller resolutions resized to <code>1280</code> resolution and merge tile predictions with NMS + WBF, but it turned out to be overfitting the F2 even with CV splits by subsequences (best/F2: <code>0.86987</code>).</p>\n<p>Cheers, I also learnt a lot from the others' great works!</p>",
      "rawMarkdown": "Hey @cdeotte , thanks for sharing! \n\nMy approach is very similar to your single YOLOv5 10-folds model with WBF, but using `yolov5m6` trained with `2560` resolution and heavy augmentation (**Private LB: 0.690**):\n```\n[\n    A.RandomSizedBBoxSafeCrop(1280, 1280, erosion_rate=0.2,\n        interpolation=cv2.INTER_LINEAR, p=1.0),\n    A.OneOf([\n        A.MotionBlur(p=.2),\n        A.MedianBlur(blur_limit=3, p=0.3),\n        A.Blur(blur_limit=3, p=0.1),\n    ], p=0.3),\n    A.OneOf([\n        A.CLAHE(clip_limit=2, tile_grid_size=(8, 8), p=0.5),\n        A.RandomBrightnessContrast(brightness_limit=0.2,\n            contrast_limit=0.2, brightness_by_max=False, p=0.5),\n        A.RandomGamma(p = 0.5),\n        A.ShiftScaleRotate(shift_limit=0.0625,\n            scale_limit=0.2, rotate_limit=45, p=0.2),\n    ], p=0.3),\n    A.HorizontalFlip(p=0.5),\n    A.Sharpen(p=0.3),\n    A.RGBShift(p=0.2),\n]\n```\nI think the key element would be using `RandomSizedBBoxSafeCrop` to crop `1280x1280` areas from `2560` resolution with `erosion_rate=0.2`, which helped the model to learn to identify smaller objects. I selected the best F2 version (train about `11-20 epochs` per fold, with extra `20%` background images).\n\nInference is done at the same size (2560) with `TTA` in `7/10 folds` to squeeze time spent (no tracking). As for NMS `conf` and `IOU` of the model, they were set to `0.5` and `0.45` to keep more bboxes with higher confidence, then for WBF I set `iou_thr=0.4` and `skip_box_thr=0.01`.\n\nhttps://www.kaggle.com/markpeng/yolov5m6-f10-v8-2560-wbf-ensemble/settings?scriptVersionId=87258324\n\nTried training by tiling images under x2/x4/x8 smaller resolutions resized to `1280` resolution and merge tile predictions with NMS + WBF, but it turned out to be overfitting the F2 even with CV splits by subsequences (best/F2: `0.86987`).\n\nCheers, I also learnt a lot from the others' great works!",
      "votes": null
    },
    {
      "id": "1690588",
      "postDate": "02/15/2022 02:42:36",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1690632",
      "postDate": "02/15/2022 03:28:42",
      "content": "<p>Great Work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I can't believe it never occurred to me. WBF does have a <strong>threshold</strong> option <code>skip_box_thr</code>, which only filters only input boxes, not the fused ones.</p>",
      "rawMarkdown": "Great Work @cdeotte. I can't believe it never occurred to me. WBF does have a **threshold** option `skip_box_thr`, which only filters only input boxes, not the fused ones.",
      "votes": null
    },
    {
      "id": "1690667",
      "postDate": "02/15/2022 03:49:21",
      "content": "<p>Yes <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> , I never realized this before either until just this morning. In previous competitions with <code>mAP</code> metric it didn't matter (like our SIIM-FISABIO-RSNA COVID-19 Detection comp) but in this competition with <code>F2</code> (or <code>F1</code>) it does matter.</p>\n<p>I discovered when I increased my inference from using 6 out of 10 folds to 9 out of 10 folds. I couldn't understand why the LB score decreased. Then all of a sudden, it occurred to me.</p>",
      "rawMarkdown": "Yes @awsaf49 , I never realized this before either until just this morning. In previous competitions with `mAP` metric it didn't matter (like our SIIM-FISABIO-RSNA COVID-19 Detection comp) but in this competition with `F2` (or `F1`) it does matter.\n\nI discovered when I increased my inference from using 6 out of 10 folds to 9 out of 10 folds. I couldn't understand why the LB score decreased. Then all of a sudden, it occurred to me.",
      "votes": null
    },
    {
      "id": "1690674",
      "postDate": "02/15/2022 03:54:34",
      "content": "<p>Awesome work <a href=\"https://www.kaggle.com/markpeng\" target=\"_blank\">@markpeng</a> . Congrats on strong Silver solo finish! Using <code>RandomSizedBBoxSafeCrop</code> is a great idea. That let's you train with large images and have large batches. What does <code>erosion_rate</code> do?</p>\n<p>I think since you used a high <code>conf = 0.5</code> for each model, then if only 1 of your 7 models predicted a bbox, it would become <code>0.5 / 7 = 0.07</code> probability which is still acceptable. Alternatively, you could lower each model to <code>conf = 0.1</code> and use <code>WBF_CONF = 0.1</code> like the code in my post above.</p>\n<p>Note that <code>skip_box_thr=0.01</code> filters input boxes which in your case does nothing since each input box is already <code>&gt;=0.5</code>.</p>",
      "rawMarkdown": "Awesome work @markpeng . Congrats on strong Silver solo finish! Using `RandomSizedBBoxSafeCrop` is a great idea. That let's you train with large images and have large batches. What does `erosion_rate` do?\n\nI think since you used a high `conf = 0.5` for each model, then if only 1 of your 7 models predicted a bbox, it would become `0.5 / 7 = 0.07` probability which is still acceptable. Alternatively, you could lower each model to `conf = 0.1` and use `WBF_CONF = 0.1` like the code in my post above.\n\nNote that `skip_box_thr=0.01` filters input boxes which in your case does nothing since each input box is already `>=0.5`.",
      "votes": null
    },
    {
      "id": "1690680",
      "postDate": "02/15/2022 03:58:12",
      "content": "<p>I train on just images with cots. I tried using <code>1:1</code> frames <code>without:with</code>. And <code>2:1</code> and <code>3:1</code>. But using <code>1:0</code> was best. </p>\n<p>Note that even when we train with 100% frames with cots, the model still learns background without cots. Many people don't realize this. For each training image, the model will place thousands of different sized rectangles on the train image. Then some rectangles will have cots and some will not. The model will learning to predict which boxes have cots and which boxes do not. So the model learns what ocean without cots looks like (even when we use all positive images).</p>",
      "rawMarkdown": "I train on just images with cots. I tried using `1:1` frames `without:with`. And `2:1` and `3:1`. But using `1:0` was best. \n\nNote that even when we train with 100% frames with cots, the model still learns background without cots. Many people don't realize this. For each training image, the model will place thousands of different sized rectangles on the train image. Then some rectangles will have cots and some will not. The model will learning to predict which boxes have cots and which boxes do not. So the model learns what ocean without cots looks like (even when we use all positive images).",
      "votes": null
    },
    {
      "id": "1690682",
      "postDate": "02/15/2022 03:59:09",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/trushk\" target=\"_blank\">@trushk</a> . Congrats to you and your team for 6th place Gold. Fantastic!</p>",
      "rawMarkdown": "Thanks @trushk . Congrats to you and your team for 6th place Gold. Fantastic!",
      "votes": null
    },
    {
      "id": "1690747",
      "postDate": "02/15/2022 04:48:47",
      "content": "<p>It is the same idea as the following. If we set <code>model.conf = 0.01</code>, does that help or hurt? This hurts because the model predicts too many boxes.</p>\n<p>When we ensemble 10 models, if the majority of models predict the same box, then it is most likely a TP. But what happens when only 1 model predicts a box, and the other 9 models do not predict that box? (It is most likely FP).</p>\n<p>In this case, WBF takes the average of the first model's box conf with 9 zeros. So the resultant box conf becomes very small like <code>conf = 0.02</code> or something. The library WBF does not remove this box even though the resultant <code>conf</code> is very small. You must write your own code like mine above to remove boxes like this.</p>",
      "rawMarkdown": "It is the same idea as the following. If we set `model.conf = 0.01`, does that help or hurt? This hurts because the model predicts too many boxes.\n\nWhen we ensemble 10 models, if the majority of models predict the same box, then it is most likely a TP. But what happens when only 1 model predicts a box, and the other 9 models do not predict that box? (It is most likely FP).\n\nIn this case, WBF takes the average of the first model's box conf with 9 zeros. So the resultant box conf becomes very small like `conf = 0.02` or something. The library WBF does not remove this box even though the resultant `conf` is very small. You must write your own code like mine above to remove boxes like this.",
      "votes": null
    },
    {
      "id": "1690759",
      "postDate": "02/15/2022 05:00:00",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , Great Work! </p>\n<p>Can I ask what for what augmentations you used or how you tuned your hyperparams? I was planning to use the wandb sweeps that were integrated with YOLOV5, but I didn't have enough time. </p>",
      "rawMarkdown": "Hey @cdeotte , Great Work! \n\nCan I ask what for what augmentations you used or how you tuned your hyperparams? I was planning to use the wandb sweeps that were integrated with YOLOV5, but I didn't have enough time.",
      "votes": null
    },
    {
      "id": "1690773",
      "postDate": "02/15/2022 05:14:46",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> even for late join to competition! :) from our observation WBFT (this is how we called it) was also working</p>",
      "rawMarkdown": "Great work @cdeotte even for late join to competition! :) from our observation WBFT (this is how we called it) was also working",
      "votes": null
    },
    {
      "id": "1690813",
      "postDate": "02/15/2022 05:37:15",
      "content": "<p>Thanks for sharing. Nice work!</p>",
      "rawMarkdown": "Thanks for sharing. Nice work!",
      "votes": null
    },
    {
      "id": "1690881",
      "postDate": "02/15/2022 06:26:32",
      "content": "<p>thank you so much, congrats on your medal</p>",
      "rawMarkdown": "thank you so much, congrats on your medal",
      "votes": null
    },
    {
      "id": "1690896",
      "postDate": "02/15/2022 06:36:17",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> yes we came to the same conclution and develop WBFT (mentioned in my post <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307504\" target=\"_blank\">Experiments list - competition summary from our team perspective</a>. Having 4xGPU was really great in this competition … not only because of image resolution experiments but because of speed up process. I am really wondering how you CV procedure looked like. Could you share your thoughts on that? </p>",
      "rawMarkdown": "cdeotte yes we came to the same conclution and develop WBFT (mentioned in my post [Experiments list - competition summary from our team perspective](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307504). Having 4xGPU was really great in this competition ... not only because of image resolution experiments but because of speed up process. I am really wondering how you CV procedure looked like. Could you share your thoughts on that?",
      "votes": null
    },
    {
      "id": "1690907",
      "postDate": "02/15/2022 06:44:50",
      "content": "<p>thx for sharing.  reading discussion after competition  can learn a lot stuff. </p>",
      "rawMarkdown": "thx for sharing.  reading discussion after competition  can learn a lot stuff.",
      "votes": null
    },
    {
      "id": "1690940",
      "postDate": "02/15/2022 06:59:07",
      "content": "<p>Thanks for your <code>WBF_CONF</code>! Think I will never discover it without your reminder. Now I know why it didn't get better when I use WBF 😂</p>\n<p>We trained with <code>image_size = 3000</code> and infer with <code>image_size = 6400</code>. Instead of yolov5L6, we adopt yolov5s. Tracking didn't improve the score, but TTA yes! Finally, we got a public LB 0.661 and private 0.673. </p>\n<p>By the way, <code>mixup</code> doesn't seem to work, at least for us, does it really work here?</p>\n<p>I'm a new kaggler and really benefit a lot from your Great Work, thank you :)</p>",
      "rawMarkdown": "Thanks for your `WBF_CONF`! Think I will never discover it without your reminder. Now I know why it didn't get better when I use WBF 😂\n\nWe trained with `image_size = 3000` and infer with `image_size = 6400`. Instead of yolov5L6, we adopt yolov5s. Tracking didn't improve the score, but TTA yes! Finally, we got a public LB 0.661 and private 0.673. \n\nBy the way, `mixup` doesn't seem to work, at least for us, does it really work here?\n\nI'm a new kaggler and really benefit a lot from your Great Work, thank you :)",
      "votes": null
    },
    {
      "id": "1691531",
      "postDate": "02/15/2022 13:12:18",
      "content": "<p>Hi, i used sheep's augmentations in his notebook <a href=\"https://www.kaggle.com/steamedsheep/yolov5-high-resolution-training/notebook\" target=\"_blank\">here</a>. Check out code cell 5. I experimented by changing the default anchors in <code>yolov5l6.yaml</code> and changed some augmentations but my final solution uses sheep's script as is and uses yolo default anchors. (I joined late and didn't have enough time to find better).</p>\n<p>The main things i changed were model backbone, image size, batch size, number of epochs, and how many frames without cots to add. I picked one fold and kept training that fold over and over with different settings.</p>",
      "rawMarkdown": "Hi, i used sheep's augmentations in his notebook [here][1]. Check out code cell 5. I experimented by changing the default anchors in `yolov5l6.yaml` and changed some augmentations but my final solution uses sheep's script as is and uses yolo default anchors. (I joined late and didn't have enough time to find better).\n\nThe main things i changed were model backbone, image size, batch size, number of epochs, and how many frames without cots to add. I picked one fold and kept training that fold over and over with different settings.\n\n[1]: https://www.kaggle.com/steamedsheep/yolov5-high-resolution-training/notebook",
      "votes": null
    },
    {
      "id": "1691563",
      "postDate": "02/15/2022 13:31:10",
      "content": "<p>Congratulations Editoxic. You did well.</p>",
      "rawMarkdown": "Congratulations Editoxic. You did well.",
      "votes": null
    },
    {
      "id": "1691577",
      "postDate": "02/15/2022 13:37:28",
      "content": "<p>Hi Remek, congratulations to you and team on achieving 34th place Silver. Thanks for all your sharing.</p>\n<p>For CV, I used public notebook subsequences (not sequences) <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">here</a> and did stratified group K fold. The following code gives each fold the same number of annotated frames and keeps subsequences in their own fold:</p>\n<pre><code>from sklearn.model_selection import StratifiedGroupKFold\nsgkf = StratifiedGroupKFold(n_splits=10)\nfor fold, (t_idx, v_idx) in enumerate( sgkf.split(df, df.has_annotations, df.subsequence_id) ):\n    df.loc[v_idx,'fold'] = fold\n</code></pre>\n<p>In retrospect, since the test data was different videos, it may have been better to use 3 folds with each video in its own fold. Then after finding the best hyparameters, train with all 3 videos.</p>\n<p>For my experiments, i mainly ran the same <code>fold=5</code> over and over with different settings. I mainly tried different Yolo backbone, image size, batch size, number of epochs. I didn't have too much time to adjust augmentations nor try models other than Yolov5.</p>",
      "rawMarkdown": "Hi Remek, congratulations to you and team on achieving 34th place Silver. Thanks for all your sharing.\n\nFor CV, I used public notebook subsequences (not sequences) [here][1] and did stratified group K fold. The following code gives each fold the same number of annotated frames and keeps subsequences in their own fold:\n\n    from sklearn.model_selection import StratifiedGroupKFold\n    sgkf = StratifiedGroupKFold(n_splits=10)\n    for fold, (t_idx, v_idx) in enumerate( sgkf.split(df, df.has_annotations, df.subsequence_id) ):\n        df.loc[v_idx,'fold'] = fold\n\nIn retrospect, since the test data was different videos, it may have been better to use 3 folds with each video in its own fold. Then after finding the best hyparameters, train with all 3 videos.\n\nFor my experiments, i mainly ran the same `fold=5` over and over with different settings. I mainly tried different Yolo backbone, image size, batch size, number of epochs. I didn't have too much time to adjust augmentations nor try models other than Yolov5.\n\n[1]: https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences",
      "votes": null
    },
    {
      "id": "1691580",
      "postDate": "02/15/2022 13:40:28",
      "content": "<p>Thanks Lukasz. Congratulations to you and your team!</p>\n<p>WBFT works great. I guess as an alternative some teams may have just increased <code>model.conf</code>. For example, if you want <code>WBF_CONF = 0.1</code>, you could probably use <code>model.conf = 0.5</code> and use 5 models in ensemble. Then the smallest <code>conf</code> you will get after WBF without WBFT is <code>0.5/5 = 0.1</code>. But using WBFT with <code>model.conf = 0.05</code> and <code>WBF_CONF = 0.1</code> should achieve a better CV LB.</p>",
      "rawMarkdown": "Thanks Lukasz. Congratulations to you and your team!\n\nWBFT works great. I guess as an alternative some teams may have just increased `model.conf`. For example, if you want `WBF_CONF = 0.1`, you could probably use `model.conf = 0.5` and use 5 models in ensemble. Then the smallest `conf` you will get after WBF without WBFT is `0.5/5 = 0.1`. But using WBFT with `model.conf = 0.05` and `WBF_CONF = 0.1` should achieve a better CV LB.",
      "votes": null
    },
    {
      "id": "1691588",
      "postDate": "02/15/2022 13:43:20",
      "content": "<p>Thank you for everything in this competition! 🙏🙏🙏</p>",
      "rawMarkdown": "Thank you for everything in this competition! 🙏🙏🙏",
      "votes": null
    },
    {
      "id": "1691629",
      "postDate": "02/15/2022 14:16:13",
      "content": "<p>Wow, that's it!</p>",
      "rawMarkdown": "Wow, that's it!",
      "votes": null
    },
    {
      "id": "1691646",
      "postDate": "02/15/2022 14:31:13",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> As I understood <code>erosion_rate</code> controls how much area of the original bounding box could be lost after cropping. <code>erosion_rate = 0.2</code> means the augmented bounding box's area could be up to 20% smaller than the area of the original bbox before resizing.<br>\nCheckout this tutorial: <a href=\"https://albumentations.ai/docs/examples/example_bboxes2/\" target=\"_blank\">https://albumentations.ai/docs/examples/example_bboxes2/</a></p>\n<p>Thanks for the extra insight about conf-based post filter for WBF fused bboxes.<br>\nAnd you are right about the <code>skip_box_thr</code> filter, it didn't take any effect in my case.</p>",
      "rawMarkdown": "cdeotte As I understood `erosion_rate` controls how much area of the original bounding box could be lost after cropping. `erosion_rate = 0.2` means the augmented bounding box's area could be up to 20% smaller than the area of the original bbox before resizing.\nCheckout this tutorial: https://albumentations.ai/docs/examples/example_bboxes2/\n\nThanks for the extra insight about conf-based post filter for WBF fused bboxes.\nAnd you are right about the `skip_box_thr` filter, it didn't take any effect in my case.",
      "votes": null
    },
    {
      "id": "1691686",
      "postDate": "02/15/2022 14:53:44",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the reply.<br>\nI understood from your explanation that WBFT is to reduce FP, not FN.<br>\nAs a matter of fact, my team doesn't have such a process like WBFT, so I'll try it later.</p>",
      "rawMarkdown": "cdeotte Thanks for the reply.\nI understood from your explanation that WBFT is to reduce FP, not FN.\nAs a matter of fact, my team doesn't have such a process like WBFT, so I'll try it later.",
      "votes": null
    },
    {
      "id": "1691774",
      "postDate": "02/15/2022 15:47:14",
      "content": "<p><a href=\"https://www.kaggle.com/kmizunoster\" target=\"_blank\">@kmizunoster</a> thanks for pointing this out. I updated my discussion to say \"too many FP\". I incorrectly said \"too many FN\". Using WBFT decreases FP not FN. Using more than 1 model (i.e. ensembling) helps decrease FN but we must balance the addition of boxes with WBFT so that we don't have too many FP.</p>",
      "rawMarkdown": "kmizunoster thanks for pointing this out. I updated my discussion to say \"too many FP\". I incorrectly said \"too many FN\". Using WBFT decreases FP not FN. Using more than 1 model (i.e. ensembling) helps decrease FN but we must balance the addition of boxes with WBFT so that we don't have too many FP.",
      "votes": null
    },
    {
      "id": "1691776",
      "postDate": "02/15/2022 15:48:43",
      "content": "<p>UPDATE: I updated post to say \"WBFT decreases FP\". It does not decrease FN. Using more than 1 model, i.e. ensembling a diversity of models is what decreases FN. But we must use WBFT (instead of WBF) so that the decrease in FN does not add too many FP.</p>",
      "rawMarkdown": "UPDATE: I updated post to say \"WBFT decreases FP\". It does not decrease FN. Using more than 1 model, i.e. ensembling a diversity of models is what decreases FN. But we must use WBFT (instead of WBF) so that the decrease in FN does not add too many FP.",
      "votes": null
    },
    {
      "id": "1691945",
      "postDate": "02/15/2022 17:41:34",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 🌟🎉</p>",
      "rawMarkdown": "congrats @cdeotte 🌟🎉",
      "votes": null
    },
    {
      "id": "1692124",
      "postDate": "02/15/2022 21:00:13",
      "content": "<p>Thanks Aruna</p>",
      "rawMarkdown": "Thanks Aruna",
      "votes": null
    },
    {
      "id": "1692410",
      "postDate": "02/16/2022 02:44:27",
      "content": "<p>UPDATE 1 (after comp ended): Instead of inferring 10 folds (in 7 hours), I just submitted with 5 of the 10 folds and added TTA to each fold (which takes 9 hours). The public LB boost 0.018 and the private LB boost 0.011. The private rank boost from 75th place to 40th place!</p>\n<p>Yolo TTA is very powerful because it infers the image at 3 different image sizes which adds more diversity than adding more folds. TTA also does flips. So using TTA is better than using more of the original folds. From Yolo's page <a href=\"https://github.com/ultralytics/yolov5/issues/303\" target=\"_blank\">here</a></p>\n<blockquote>\n  <p>Note that inference with TTA enabled will typically take about 2-3X the time of normal inference as the images are being left-right flipped and processed at 3 different resolutions, with the outputs merged before NMS. Part of the speed decrease is simply due to larger image sizes (832 vs 640), while part is due to the actual TTA operations.</p>\n</blockquote>",
      "rawMarkdown": "UPDATE 1 (after comp ended): Instead of inferring 10 folds (in 7 hours), I just submitted with 5 of the 10 folds and added TTA to each fold (which takes 9 hours). The public LB boost 0.018 and the private LB boost 0.011. The private rank boost from 75th place to 40th place!\n\nYolo TTA is very powerful because it infers the image at 3 different image sizes which adds more diversity than adding more folds. TTA also does flips. So using TTA is better than using more of the original folds. From Yolo's page [here][1]\n\n>Note that inference with TTA enabled will typically take about 2-3X the time of normal inference as the images are being left-right flipped and processed at 3 different resolutions, with the outputs merged before NMS. Part of the speed decrease is simply due to larger image sizes (832 vs 640), while part is due to the actual TTA operations.\n\n[1]: https://github.com/ultralytics/yolov5/issues/303",
      "votes": null
    },
    {
      "id": "1692970",
      "postDate": "02/16/2022 11:25:15",
      "content": "<p>Thanks vad13irt!</p>",
      "rawMarkdown": "Thanks vad13irt!",
      "votes": null
    },
    {
      "id": "1693138",
      "postDate": "02/16/2022 13:13:55",
      "content": "<p>UPDATE 2 (after comp ended): Instead of adding TTA, I trained a Yolov5M6 at 3072 with same settings as my Yolov5L6. Then I submit 5 folds out of 10 Yolov5M6 (without TTA) and 5 folds out of 10 Yolo5L6 (without TTA). When using <code>WBF_CONF</code>, the result is private LB 0.709 and 20th place! This confirms that <code>WBF_CONF</code> is a powerful idea. If I add even more models (with different backbones and use different image sizes), I assume the LB will keep climbing!</p>",
      "rawMarkdown": "UPDATE 2 (after comp ended): Instead of adding TTA, I trained a Yolov5M6 at 3072 with same settings as my Yolov5L6. Then I submit 5 folds out of 10 Yolov5M6 (without TTA) and 5 folds out of 10 Yolo5L6 (without TTA). When using `WBF_CONF`, the result is private LB 0.709 and 20th place! This confirms that `WBF_CONF` is a powerful idea. If I add even more models (with different backbones and use different image sizes), I assume the LB will keep climbing!",
      "votes": null
    },
    {
      "id": "1693150",
      "postDate": "02/16/2022 13:27:45",
      "content": "<p>Thank you for sharing. I thnink WBFT is better name for this one :) I have similar observation. Having threshold is really great way to controll FP prediction in multi model inference. </p>",
      "rawMarkdown": "Thank you for sharing. I thnink WBFT is better name for this one :) I have similar observation. Having threshold is really great way to controll FP prediction in multi model inference.",
      "votes": null
    },
    {
      "id": "1693157",
      "postDate": "02/16/2022 13:34:26",
      "content": "<p>Could you check different approach? </p>\n<ul>\n<li>No TTA</li>\n<li>Only WBFT - set model IOU on 0.9 … generate many many boxes and let WBF to do a job instead of yolo NMS? We used this trick and it helps us a lot.</li>\n</ul>",
      "rawMarkdown": "Could you check different approach? \n- No TTA\n- Only WBFT - set model IOU on 0.9 ... generate many many boxes and let WBF to do a job instead of yolo NMS? We used this trick and it helps us a lot.",
      "votes": null
    },
    {
      "id": "1693177",
      "postDate": "02/16/2022 13:51:36",
      "content": "<p>Good idea (IOU=0.9), i will try this.</p>\n<blockquote>\n  <p>Having threshold is really great way to controll FP prediction in multi model inference.</p>\n</blockquote>\n<p>Yes, during the comp my best LB was just single model (1 fold) so I didn't focus on ensemble. But with <code>WBFT</code> we can benefit from ensemble. Without <code>WBFT</code> then using multiple model ensemble adds too many low conf FP and hurts competition metric of F2</p>",
      "rawMarkdown": "Good idea (IOU=0.9), i will try this.\n\n>Having threshold is really great way to controll FP prediction in multi model inference.\n\nYes, during the comp my best LB was just single model (1 fold) so I didn't focus on ensemble. But with `WBFT` we can benefit from ensemble. Without `WBFT` then using multiple model ensemble adds too many low conf FP and hurts competition metric of F2",
      "votes": null
    },
    {
      "id": "1693188",
      "postDate": "02/16/2022 13:56:59",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> as you did we just cleared all bboxes below treshold (as you elaborated - if one model claims that it can see bbox we just can using both weight and treshold to filter final bboxes - we used weights as well to create \"more\" bboxes from one model).</p>\n<p>Our configurations was:</p>\n<ul>\n<li>iou even 0.95</li>\n<li>then model.max_det = 100</li>\n<li>wbft to let make the final bbox preds and … coordinations computation - it make bboxes more tight to sf</li>\n</ul>\n<p>See here for implementation (this is one of our work in progress notebook) and debugging code: <a href=\"https://www.kaggle.com/remekkinas/cots-y5-wbf\" target=\"_blank\">https://www.kaggle.com/remekkinas/cots-y5-wbf</a></p>",
      "rawMarkdown": "cdeotte as you did we just cleared all bboxes below treshold (as you elaborated - if one model claims that it can see bbox we just can using both weight and treshold to filter final bboxes - we used weights as well to create \"more\" bboxes from one model).\n\nOur configurations was:\n- iou even 0.95\n- then model.max_det = 100\n- wbft to let make the final bbox preds and ... coordinations computation - it make bboxes more tight to sf\n\nSee here for implementation (this is one of our work in progress notebook) and debugging code: https://www.kaggle.com/remekkinas/cots-y5-wbf",
      "votes": null
    },
    {
      "id": "1693219",
      "postDate": "02/16/2022 14:29:40",
      "content": "<p>I don't understand the following line, can you explain this more?</p>\n<blockquote>\n  <p>we used weights as well to create \"more\" bboxes from one model</p>\n</blockquote>",
      "rawMarkdown": "I don't understand the following line, can you explain this more?\n>we used weights as well to create \"more\" bboxes from one model",
      "votes": null
    },
    {
      "id": "1693236",
      "postDate": "02/16/2022 14:38:27",
      "content": "<p>I think Remek wanted to say we generated more boxes at one sf like this from one model:</p>\n<p><img src=\"https://i.ibb.co/pP0y6br/autowbf1.png\" alt=\"\"></p>\n<p><img src=\"https://i.ibb.co/rHHwBMB/autowbf2.png\" alt=\"\"></p>",
      "rawMarkdown": "I think Remek wanted to say we generated more boxes at one sf like this from one model:\n\n![](https://i.ibb.co/pP0y6br/autowbf1.png)\n\n![](https://i.ibb.co/rHHwBMB/autowbf2.png)",
      "votes": null
    },
    {
      "id": "1693245",
      "postDate": "02/16/2022 14:44:21",
      "content": "<p>Ah, gotcha thanks. I just submitted my models using <code>IOU=0.9</code> and WBFT (with IOU=0.4) afterward, i will report back in 9 hours.</p>",
      "rawMarkdown": "Ah, gotcha thanks. I just submitted my models using `IOU=0.9` and WBFT (with IOU=0.4) afterward, i will report back in 9 hours.",
      "votes": null
    },
    {
      "id": "1693256",
      "postDate": "02/16/2022 14:54:45",
      "content": "<p>We used it to change native nms to wbf. When we mixed f.e. 3 models we applied 3 WBF on each model, and then one WBF summary</p>",
      "rawMarkdown": "We used it to change native nms to wbf. When we mixed f.e. 3 models we applied 3 WBF on each model, and then one WBF summary",
      "votes": null
    },
    {
      "id": "1693334",
      "postDate": "02/16/2022 15:51:53",
      "content": "<p>Congratulations and thank you for sharing knowledge!</p>",
      "rawMarkdown": "Congratulations and thank you for sharing knowledge!",
      "votes": null
    },
    {
      "id": "1693520",
      "postDate": "02/16/2022 17:57:38",
      "content": "<p><a href=\"https://www.kaggle.com/lukaszborecki\" target=\"_blank\">@lukaszborecki</a> We should read the code in GitHub WBF. I'm not sure how it computes the new <code>conf</code> value when one model has multiple boxes (as result of using <code>model.conf=0.9</code>). For example, let's say we ensemble 4 models. And model 1 has 2 boxes (with conf 0.7 and 0.8) that overlap with IOU = 0.6. And model 2 has 1 box (with conf 0.6) that overlaps with those box with IOU = 0.5. And model 3 and 4 has zero boxes.</p>\n<p>How does WBF compute the new <code>conf</code>? For example, I do not think it does <code>(0.7 + 0.8 + 0.6)/3</code> because the denominator should be 4 for 4 models. So does WBF somehow combine the two boxes from model 1 first? And then do <code>(0.75 + 0.6 + 0 + 0)/4</code>? I'm not sure</p>",
      "rawMarkdown": "lukaszborecki We should read the code in GitHub WBF. I'm not sure how it computes the new `conf` value when one model has multiple boxes (as result of using `model.conf=0.9`). For example, let's say we ensemble 4 models. And model 1 has 2 boxes (with conf 0.7 and 0.8) that overlap with IOU = 0.6. And model 2 has 1 box (with conf 0.6) that overlaps with those box with IOU = 0.5. And model 3 and 4 has zero boxes.\n\nHow does WBF compute the new `conf`? For example, I do not think it does `(0.7 + 0.8 + 0.6)/3` because the denominator should be 4 for 4 models. So does WBF somehow combine the two boxes from model 1 first? And then do `(0.75 + 0.6 + 0 + 0)/4`? I'm not sure",
      "votes": null
    },
    {
      "id": "1693528",
      "postDate": "02/16/2022 18:07:26",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I made some calculus when we were submiting.</p>\n<p>1st Scenario IOU = 0.9 and one WBF collecting boxes: in case 3 models</p>\n<p>1st model may give 3 overlapped boxes of conf 0.8 0.9 and 0.7<br>\n2nd model may give 10 overlapped boxes with mean 0.6<br>\n3rd model may give 0 overlapped boxes</p>\n<p>in this scenario wbf got 13 boxes to fuse (10 * 0.6 + 3 * 0.8 +1 * 0) / 14 = 0.6</p>\n<p>2nd scenario: WBF on output each model and then WBF all</p>\n<p>1st model - box 0.8<br>\n2nd model - box 0.6<br>\n3rd model - box 0.0</p>\n<p>Final WBF = 1.4 / 3= 0.46</p>\n<p>Didn't check it with code it was experimental calculation after one ebug inference. So my confidence in this thesis is 0.9 :D</p>\n<p>EDIT</p>\n<p>-- wrong</p>\n<p>in 1st scenario if one model got more than 1 overlap box then models which doesnt predict doesnt lower the score. I got case 9 bounding boxes from 1 model and null from model 2 and 3, and counted mean was from 9 not 11. It seams we need to check source code</p>\n<p>EDIT 2:</p>\n<p>when there were 3 models and one model got 2 overlapped boxes and other models is null it divided by 3. So maybe it check how many models are in box_list , becuase its list of 3 lists. And then divide by 3 or by number of boxes if greater than 3.</p>\n<p>-- Checked it on Excel - it works like EDIT 2</p>",
      "rawMarkdown": "cdeotte I made some calculus when we were submiting.\n\n1st Scenario IOU = 0.9 and one WBF collecting boxes: in case 3 models\n\n1st model may give 3 overlapped boxes of conf 0.8 0.9 and 0.7\n2nd model may give 10 overlapped boxes with mean 0.6\n3rd model may give 0 overlapped boxes\n\nin this scenario wbf got 13 boxes to fuse (10 * 0.6 + 3 * 0.8 +1 * 0) / 14 = 0.6\n\n2nd scenario: WBF on output each model and then WBF all\n\n1st model - box 0.8\n2nd model - box 0.6\n3rd model - box 0.0\n\nFinal WBF = 1.4 / 3= 0.46\n\nDidn't check it with code it was experimental calculation after one ebug inference. So my confidence in this thesis is 0.9 :D\n\nEDIT\n\n-- wrong\n\nin 1st scenario if one model got more than 1 overlap box then models which doesnt predict doesnt lower the score. I got case 9 bounding boxes from 1 model and null from model 2 and 3, and counted mean was from 9 not 11. It seams we need to check source code\n\nEDIT 2:\n\nwhen there were 3 models and one model got 2 overlapped boxes and other models is null it divided by 3. So maybe it check how many models are in box_list , becuase its list of 3 lists. And then divide by 3 or by number of boxes if greater than 3.\n\n-- Checked it on Excel - it works like EDIT 2",
      "votes": null
    },
    {
      "id": "1693827",
      "postDate": "02/17/2022 01:36:06",
      "content": "<p>Congratulations and thank you for sharing knowledge!</p>",
      "rawMarkdown": "Congratulations and thank you for sharing knowledge!",
      "votes": null
    },
    {
      "id": "1695499",
      "postDate": "02/18/2022 07:24:23",
      "content": "<blockquote>\n  <p>Note that even when we train with 100% frames with cots, the model still learns background without cots. <br>\n  I really like your deep thinking and understanding about the competition. Congrats!</p>\n</blockquote>",
      "rawMarkdown": "> Note that even when we train with 100% frames with cots, the model still learns background without cots. \n\nI really like your deep thinking and understanding about the competition. Congrats!",
      "votes": null
    },
    {
      "id": "1695700",
      "postDate": "02/18/2022 10:21:21",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 🤩</p>",
      "rawMarkdown": "congratulations @cdeotte 🤩",
      "votes": null
    },
    {
      "id": "1697184",
      "postDate": "02/19/2022 12:38:37",
      "content": "<p>thanks for this knowledge</p>",
      "rawMarkdown": "thanks for this knowledge",
      "votes": null
    },
    {
      "id": "1698591",
      "postDate": "02/20/2022 14:06:00",
      "content": "<p>May I know how many epochs did you train and which index did you use for splitting? Sequence of Video ID? Stratified or GroupK?</p>",
      "rawMarkdown": "May I know how many epochs did you train and which index did you use for splitting? Sequence of Video ID? Stratified or GroupK?",
      "votes": null
    },
    {
      "id": "1698656",
      "postDate": "02/20/2022 15:02:47",
      "content": "<p>I trained for 10 epochs using cosine schedule with 1 epoch warm up. I used the epoch model weights with the best F2 score which was usually epoch 7. For CV, i used group KFold on subsequences <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">here</a> (which are different than sequences).</p>",
      "rawMarkdown": "I trained for 10 epochs using cosine schedule with 1 epoch warm up. I used the epoch model weights with the best F2 score which was usually epoch 7. For CV, i used group KFold on subsequences [here][1] (which are different than sequences).\n\n[1]: https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences",
      "votes": null
    },
    {
      "id": "1726224",
      "postDate": "03/17/2022 19:18:45",
      "content": "<p>UPDATE: This technique helped win 2nd place in Kaggle NLP competition <a href=\"https://www.kaggle.com/c/feedback-prize-2021/discussion/313389\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "UPDATE: This technique helped win 2nd place in Kaggle NLP competition [here][1]\n\n[1]: https://www.kaggle.com/c/feedback-prize-2021/discussion/313389",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1690507,
      "author_name": "aengusng",
      "author_url": "",
      "post_date": "02/15/2022 01:11:39",
      "content": "<p>Our team has the same idea and same code. We set WBF_CONF = 0.5, so it refinds all prediction. We use 9 different sizes for x9 TTA, and I think the model generates different predictions at different sizes, so refining predictions after using WBF is not appropriate. Thank you for disclosing my mistake and unfinished idea. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1690510,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/15/2022 01:12:57",
      "content": "<p>As always, thanks for sharing your approach <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! . I always learn something from reading solutions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1690682,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 03:59:09",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/trushk\" target=\"_blank\">@trushk</a> . Congrats to you and your team for 6th place Gold. Fantastic!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690514,
      "author_name": "phamthaihoangtung",
      "author_url": "",
      "post_date": "02/15/2022 01:15:23",
      "content": "<p>Did you train on total data or just data having cots?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1690680,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 03:58:12",
          "content": "<p>I train on just images with cots. I tried using <code>1:1</code> frames <code>without:with</code>. And <code>2:1</code> and <code>3:1</code>. But using <code>1:0</code> was best. </p>\n<p>Note that even when we train with 100% frames with cots, the model still learns background without cots. Many people don't realize this. For each training image, the model will place thousands of different sized rectangles on the train image. Then some rectangles will have cots and some will not. The model will learning to predict which boxes have cots and which boxes do not. So the model learns what ocean without cots looks like (even when we use all positive images).</p>",
          "votes": null,
          "replies": [
            {
              "id": 1695499,
              "author_name": "toongzhhang",
              "author_url": "",
              "post_date": "02/18/2022 07:24:23",
              "content": "<blockquote>\n  <p>Note that even when we train with 100% frames with cots, the model still learns background without cots. <br>\n  I really like your deep thinking and understanding about the competition. Congrats!</p>\n</blockquote>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1690881,
          "author_name": "phamthaihoangtung",
          "author_url": "",
          "post_date": "02/15/2022 06:26:32",
          "content": "<p>thank you so much, congrats on your medal</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690531,
      "author_name": "kmizunoster",
      "author_url": "",
      "post_date": "02/15/2022 01:46:22",
      "content": "<p>Thank you for the concise summary. I have one question.</p>\n<blockquote>\n  <p>Using this code boosted my final submission Private LB 0.609 to Private LB 0.684! Wow. This is important to do if the competition metric is F2 because we are penalized with too many FN (so it's important to establish a conf threshold). However if the competition metric is mAP then it would not matter since with mAP we are not penalized for FN (explained here)</p>\n</blockquote>\n<p>Can you please explain more about this? Why does setting the WBF threshold lead to an improvement in FN? Since the number of bboxes is reduced, it seems that FN is increased but not decreased. Doesn't it mean that the F2 score improves by decreasing the FP?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1690747,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 04:48:47",
          "content": "<p>It is the same idea as the following. If we set <code>model.conf = 0.01</code>, does that help or hurt? This hurts because the model predicts too many boxes.</p>\n<p>When we ensemble 10 models, if the majority of models predict the same box, then it is most likely a TP. But what happens when only 1 model predicts a box, and the other 9 models do not predict that box? (It is most likely FP).</p>\n<p>In this case, WBF takes the average of the first model's box conf with 9 zeros. So the resultant box conf becomes very small like <code>conf = 0.02</code> or something. The library WBF does not remove this box even though the resultant <code>conf</code> is very small. You must write your own code like mine above to remove boxes like this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691686,
          "author_name": "kmizunoster",
          "author_url": "",
          "post_date": "02/15/2022 14:53:44",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the reply.<br>\nI understood from your explanation that WBFT is to reduce FP, not FN.<br>\nAs a matter of fact, my team doesn't have such a process like WBFT, so I'll try it later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691774,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 15:47:14",
          "content": "<p><a href=\"https://www.kaggle.com/kmizunoster\" target=\"_blank\">@kmizunoster</a> thanks for pointing this out. I updated my discussion to say \"too many FP\". I incorrectly said \"too many FN\". Using WBFT decreases FP not FN. Using more than 1 model (i.e. ensembling) helps decrease FN but we must balance the addition of boxes with WBFT so that we don't have too many FP.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690532,
      "author_name": "markpeng",
      "author_url": "",
      "post_date": "02/15/2022 01:46:28",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thanks for sharing! </p>\n<p>My approach is very similar to your single YOLOv5 10-folds model with WBF, but using <code>yolov5m6</code> trained with <code>2560</code> resolution and heavy augmentation (<strong>Private LB: 0.690</strong>):</p>\n<pre><code>[\n    A.RandomSizedBBoxSafeCrop(1280, 1280, erosion_rate=0.2,\n        interpolation=cv2.INTER_LINEAR, p=1.0),\n    A.OneOf([\n        A.MotionBlur(p=.2),\n        A.MedianBlur(blur_limit=3, p=0.3),\n        A.Blur(blur_limit=3, p=0.1),\n    ], p=0.3),\n    A.OneOf([\n        A.CLAHE(clip_limit=2, tile_grid_size=(8, 8), p=0.5),\n        A.RandomBrightnessContrast(brightness_limit=0.2,\n            contrast_limit=0.2, brightness_by_max=False, p=0.5),\n        A.RandomGamma(p = 0.5),\n        A.ShiftScaleRotate(shift_limit=0.0625,\n            scale_limit=0.2, rotate_limit=45, p=0.2),\n    ], p=0.3),\n    A.HorizontalFlip(p=0.5),\n    A.Sharpen(p=0.3),\n    A.RGBShift(p=0.2),\n]\n</code></pre>\n<p>I think the key element would be using <code>RandomSizedBBoxSafeCrop</code> to crop <code>1280x1280</code> areas from <code>2560</code> resolution with <code>erosion_rate=0.2</code>, which helped the model to learn to identify smaller objects. I selected the best F2 version (train about <code>11-20 epochs</code> per fold, with extra <code>20%</code> background images).</p>\n<p>Inference is done at the same size (2560) with <code>TTA</code> in <code>7/10 folds</code> to squeeze time spent (no tracking). As for NMS <code>conf</code> and <code>IOU</code> of the model, they were set to <code>0.5</code> and <code>0.45</code> to keep more bboxes with higher confidence, then for WBF I set <code>iou_thr=0.4</code> and <code>skip_box_thr=0.01</code>.</p>\n<p><a href=\"https://www.kaggle.com/markpeng/yolov5m6-f10-v8-2560-wbf-ensemble/settings?scriptVersionId=87258324\" target=\"_blank\">https://www.kaggle.com/markpeng/yolov5m6-f10-v8-2560-wbf-ensemble/settings?scriptVersionId=87258324</a></p>\n<p>Tried training by tiling images under x2/x4/x8 smaller resolutions resized to <code>1280</code> resolution and merge tile predictions with NMS + WBF, but it turned out to be overfitting the F2 even with CV splits by subsequences (best/F2: <code>0.86987</code>).</p>\n<p>Cheers, I also learnt a lot from the others' great works!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1690674,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 03:54:34",
          "content": "<p>Awesome work <a href=\"https://www.kaggle.com/markpeng\" target=\"_blank\">@markpeng</a> . Congrats on strong Silver solo finish! Using <code>RandomSizedBBoxSafeCrop</code> is a great idea. That let's you train with large images and have large batches. What does <code>erosion_rate</code> do?</p>\n<p>I think since you used a high <code>conf = 0.5</code> for each model, then if only 1 of your 7 models predicted a bbox, it would become <code>0.5 / 7 = 0.07</code> probability which is still acceptable. Alternatively, you could lower each model to <code>conf = 0.1</code> and use <code>WBF_CONF = 0.1</code> like the code in my post above.</p>\n<p>Note that <code>skip_box_thr=0.01</code> filters input boxes which in your case does nothing since each input box is already <code>&gt;=0.5</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691646,
          "author_name": "markpeng",
          "author_url": "",
          "post_date": "02/15/2022 14:31:13",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> As I understood <code>erosion_rate</code> controls how much area of the original bounding box could be lost after cropping. <code>erosion_rate = 0.2</code> means the augmented bounding box's area could be up to 20% smaller than the area of the original bbox before resizing.<br>\nCheckout this tutorial: <a href=\"https://albumentations.ai/docs/examples/example_bboxes2/\" target=\"_blank\">https://albumentations.ai/docs/examples/example_bboxes2/</a></p>\n<p>Thanks for the extra insight about conf-based post filter for WBF fused bboxes.<br>\nAnd you are right about the <code>skip_box_thr</code> filter, it didn't take any effect in my case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690588,
      "author_name": "robsonsan",
      "author_url": "",
      "post_date": "02/15/2022 02:42:36",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1690632,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "02/15/2022 03:28:42",
      "content": "<p>Great Work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I can't believe it never occurred to me. WBF does have a <strong>threshold</strong> option <code>skip_box_thr</code>, which only filters only input boxes, not the fused ones.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1690667,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 03:49:21",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> , I never realized this before either until just this morning. In previous competitions with <code>mAP</code> metric it didn't matter (like our SIIM-FISABIO-RSNA COVID-19 Detection comp) but in this competition with <code>F2</code> (or <code>F1</code>) it does matter.</p>\n<p>I discovered when I increased my inference from using 6 out of 10 folds to 9 out of 10 folds. I couldn't understand why the LB score decreased. Then all of a sudden, it occurred to me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690759,
      "author_name": "kennyxie",
      "author_url": "",
      "post_date": "02/15/2022 05:00:00",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , Great Work! </p>\n<p>Can I ask what for what augmentations you used or how you tuned your hyperparams? I was planning to use the wandb sweeps that were integrated with YOLOV5, but I didn't have enough time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1691531,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 13:12:18",
          "content": "<p>Hi, i used sheep's augmentations in his notebook <a href=\"https://www.kaggle.com/steamedsheep/yolov5-high-resolution-training/notebook\" target=\"_blank\">here</a>. Check out code cell 5. I experimented by changing the default anchors in <code>yolov5l6.yaml</code> and changed some augmentations but my final solution uses sheep's script as is and uses yolo default anchors. (I joined late and didn't have enough time to find better).</p>\n<p>The main things i changed were model backbone, image size, batch size, number of epochs, and how many frames without cots to add. I picked one fold and kept training that fold over and over with different settings.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690773,
      "author_name": "lukaszborecki",
      "author_url": "",
      "post_date": "02/15/2022 05:14:46",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> even for late join to competition! :) from our observation WBFT (this is how we called it) was also working</p>",
      "votes": null,
      "replies": [
        {
          "id": 1691580,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 13:40:28",
          "content": "<p>Thanks Lukasz. Congratulations to you and your team!</p>\n<p>WBFT works great. I guess as an alternative some teams may have just increased <code>model.conf</code>. For example, if you want <code>WBF_CONF = 0.1</code>, you could probably use <code>model.conf = 0.5</code> and use 5 models in ensemble. Then the smallest <code>conf</code> you will get after WBF without WBFT is <code>0.5/5 = 0.1</code>. But using WBFT with <code>model.conf = 0.05</code> and <code>WBF_CONF = 0.1</code> should achieve a better CV LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690813,
      "author_name": "kirilproger",
      "author_url": "",
      "post_date": "02/15/2022 05:37:15",
      "content": "<p>Thanks for sharing. Nice work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1690896,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/15/2022 06:36:17",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> yes we came to the same conclution and develop WBFT (mentioned in my post <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307504\" target=\"_blank\">Experiments list - competition summary from our team perspective</a>. Having 4xGPU was really great in this competition … not only because of image resolution experiments but because of speed up process. I am really wondering how you CV procedure looked like. Could you share your thoughts on that? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1691577,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 13:37:28",
          "content": "<p>Hi Remek, congratulations to you and team on achieving 34th place Silver. Thanks for all your sharing.</p>\n<p>For CV, I used public notebook subsequences (not sequences) <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">here</a> and did stratified group K fold. The following code gives each fold the same number of annotated frames and keeps subsequences in their own fold:</p>\n<pre><code>from sklearn.model_selection import StratifiedGroupKFold\nsgkf = StratifiedGroupKFold(n_splits=10)\nfor fold, (t_idx, v_idx) in enumerate( sgkf.split(df, df.has_annotations, df.subsequence_id) ):\n    df.loc[v_idx,'fold'] = fold\n</code></pre>\n<p>In retrospect, since the test data was different videos, it may have been better to use 3 folds with each video in its own fold. Then after finding the best hyparameters, train with all 3 videos.</p>\n<p>For my experiments, i mainly ran the same <code>fold=5</code> over and over with different settings. I mainly tried different Yolo backbone, image size, batch size, number of epochs. I didn't have too much time to adjust augmentations nor try models other than Yolov5.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691588,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/15/2022 13:43:20",
          "content": "<p>Thank you for everything in this competition! 🙏🙏🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690907,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "02/15/2022 06:44:50",
      "content": "<p>thx for sharing.  reading discussion after competition  can learn a lot stuff. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1690940,
      "author_name": "explainaaaaa",
      "author_url": "",
      "post_date": "02/15/2022 06:59:07",
      "content": "<p>Thanks for your <code>WBF_CONF</code>! Think I will never discover it without your reminder. Now I know why it didn't get better when I use WBF 😂</p>\n<p>We trained with <code>image_size = 3000</code> and infer with <code>image_size = 6400</code>. Instead of yolov5L6, we adopt yolov5s. Tracking didn't improve the score, but TTA yes! Finally, we got a public LB 0.661 and private 0.673. </p>\n<p>By the way, <code>mixup</code> doesn't seem to work, at least for us, does it really work here?</p>\n<p>I'm a new kaggler and really benefit a lot from your Great Work, thank you :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1691563,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 13:31:10",
          "content": "<p>Congratulations Editoxic. You did well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691629,
      "author_name": "vad13irt",
      "author_url": "",
      "post_date": "02/15/2022 14:16:13",
      "content": "<p>Wow, that's it!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1692970,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/16/2022 11:25:15",
          "content": "<p>Thanks vad13irt!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691776,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/15/2022 15:48:43",
      "content": "<p>UPDATE: I updated post to say \"WBFT decreases FP\". It does not decrease FN. Using more than 1 model, i.e. ensembling a diversity of models is what decreases FN. But we must use WBFT (instead of WBF) so that the decrease in FN does not add too many FP.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1691945,
      "author_name": "arunasivapragasam",
      "author_url": "",
      "post_date": "02/15/2022 17:41:34",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 🌟🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 1692124,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 21:00:13",
          "content": "<p>Thanks Aruna</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1692410,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/16/2022 02:44:27",
      "content": "<p>UPDATE 1 (after comp ended): Instead of inferring 10 folds (in 7 hours), I just submitted with 5 of the 10 folds and added TTA to each fold (which takes 9 hours). The public LB boost 0.018 and the private LB boost 0.011. The private rank boost from 75th place to 40th place!</p>\n<p>Yolo TTA is very powerful because it infers the image at 3 different image sizes which adds more diversity than adding more folds. TTA also does flips. So using TTA is better than using more of the original folds. From Yolo's page <a href=\"https://github.com/ultralytics/yolov5/issues/303\" target=\"_blank\">here</a></p>\n<blockquote>\n  <p>Note that inference with TTA enabled will typically take about 2-3X the time of normal inference as the images are being left-right flipped and processed at 3 different resolutions, with the outputs merged before NMS. Part of the speed decrease is simply due to larger image sizes (832 vs 640), while part is due to the actual TTA operations.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1693138,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/16/2022 13:13:55",
          "content": "<p>UPDATE 2 (after comp ended): Instead of adding TTA, I trained a Yolov5M6 at 3072 with same settings as my Yolov5L6. Then I submit 5 folds out of 10 Yolov5M6 (without TTA) and 5 folds out of 10 Yolo5L6 (without TTA). When using <code>WBF_CONF</code>, the result is private LB 0.709 and 20th place! This confirms that <code>WBF_CONF</code> is a powerful idea. If I add even more models (with different backbones and use different image sizes), I assume the LB will keep climbing!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693150,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/16/2022 13:27:45",
          "content": "<p>Thank you for sharing. I thnink WBFT is better name for this one :) I have similar observation. Having threshold is really great way to controll FP prediction in multi model inference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693157,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/16/2022 13:34:26",
          "content": "<p>Could you check different approach? </p>\n<ul>\n<li>No TTA</li>\n<li>Only WBFT - set model IOU on 0.9 … generate many many boxes and let WBF to do a job instead of yolo NMS? We used this trick and it helps us a lot.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693177,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/16/2022 13:51:36",
          "content": "<p>Good idea (IOU=0.9), i will try this.</p>\n<blockquote>\n  <p>Having threshold is really great way to controll FP prediction in multi model inference.</p>\n</blockquote>\n<p>Yes, during the comp my best LB was just single model (1 fold) so I didn't focus on ensemble. But with <code>WBFT</code> we can benefit from ensemble. Without <code>WBFT</code> then using multiple model ensemble adds too many low conf FP and hurts competition metric of F2</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693188,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/16/2022 13:56:59",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> as you did we just cleared all bboxes below treshold (as you elaborated - if one model claims that it can see bbox we just can using both weight and treshold to filter final bboxes - we used weights as well to create \"more\" bboxes from one model).</p>\n<p>Our configurations was:</p>\n<ul>\n<li>iou even 0.95</li>\n<li>then model.max_det = 100</li>\n<li>wbft to let make the final bbox preds and … coordinations computation - it make bboxes more tight to sf</li>\n</ul>\n<p>See here for implementation (this is one of our work in progress notebook) and debugging code: <a href=\"https://www.kaggle.com/remekkinas/cots-y5-wbf\" target=\"_blank\">https://www.kaggle.com/remekkinas/cots-y5-wbf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693219,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/16/2022 14:29:40",
          "content": "<p>I don't understand the following line, can you explain this more?</p>\n<blockquote>\n  <p>we used weights as well to create \"more\" bboxes from one model</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693236,
          "author_name": "lukaszborecki",
          "author_url": "",
          "post_date": "02/16/2022 14:38:27",
          "content": "<p>I think Remek wanted to say we generated more boxes at one sf like this from one model:</p>\n<p><img src=\"https://i.ibb.co/pP0y6br/autowbf1.png\" alt=\"\"></p>\n<p><img src=\"https://i.ibb.co/rHHwBMB/autowbf2.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693245,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/16/2022 14:44:21",
          "content": "<p>Ah, gotcha thanks. I just submitted my models using <code>IOU=0.9</code> and WBFT (with IOU=0.4) afterward, i will report back in 9 hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693256,
          "author_name": "lukaszborecki",
          "author_url": "",
          "post_date": "02/16/2022 14:54:45",
          "content": "<p>We used it to change native nms to wbf. When we mixed f.e. 3 models we applied 3 WBF on each model, and then one WBF summary</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693520,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/16/2022 17:57:38",
          "content": "<p><a href=\"https://www.kaggle.com/lukaszborecki\" target=\"_blank\">@lukaszborecki</a> We should read the code in GitHub WBF. I'm not sure how it computes the new <code>conf</code> value when one model has multiple boxes (as result of using <code>model.conf=0.9</code>). For example, let's say we ensemble 4 models. And model 1 has 2 boxes (with conf 0.7 and 0.8) that overlap with IOU = 0.6. And model 2 has 1 box (with conf 0.6) that overlaps with those box with IOU = 0.5. And model 3 and 4 has zero boxes.</p>\n<p>How does WBF compute the new <code>conf</code>? For example, I do not think it does <code>(0.7 + 0.8 + 0.6)/3</code> because the denominator should be 4 for 4 models. So does WBF somehow combine the two boxes from model 1 first? And then do <code>(0.75 + 0.6 + 0 + 0)/4</code>? I'm not sure</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693528,
          "author_name": "lukaszborecki",
          "author_url": "",
          "post_date": "02/16/2022 18:07:26",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I made some calculus when we were submiting.</p>\n<p>1st Scenario IOU = 0.9 and one WBF collecting boxes: in case 3 models</p>\n<p>1st model may give 3 overlapped boxes of conf 0.8 0.9 and 0.7<br>\n2nd model may give 10 overlapped boxes with mean 0.6<br>\n3rd model may give 0 overlapped boxes</p>\n<p>in this scenario wbf got 13 boxes to fuse (10 * 0.6 + 3 * 0.8 +1 * 0) / 14 = 0.6</p>\n<p>2nd scenario: WBF on output each model and then WBF all</p>\n<p>1st model - box 0.8<br>\n2nd model - box 0.6<br>\n3rd model - box 0.0</p>\n<p>Final WBF = 1.4 / 3= 0.46</p>\n<p>Didn't check it with code it was experimental calculation after one ebug inference. So my confidence in this thesis is 0.9 :D</p>\n<p>EDIT</p>\n<p>-- wrong</p>\n<p>in 1st scenario if one model got more than 1 overlap box then models which doesnt predict doesnt lower the score. I got case 9 bounding boxes from 1 model and null from model 2 and 3, and counted mean was from 9 not 11. It seams we need to check source code</p>\n<p>EDIT 2:</p>\n<p>when there were 3 models and one model got 2 overlapped boxes and other models is null it divided by 3. So maybe it check how many models are in box_list , becuase its list of 3 lists. And then divide by 3 or by number of boxes if greater than 3.</p>\n<p>-- Checked it on Excel - it works like EDIT 2</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1693334,
      "author_name": "vivovinco",
      "author_url": "",
      "post_date": "02/16/2022 15:51:53",
      "content": "<p>Congratulations and thank you for sharing knowledge!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1693827,
      "author_name": "kittylina",
      "author_url": "",
      "post_date": "02/17/2022 01:36:06",
      "content": "<p>Congratulations and thank you for sharing knowledge!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1695700,
      "author_name": "arunasivapragasam",
      "author_url": "",
      "post_date": "02/18/2022 10:21:21",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 🤩</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1697184,
      "author_name": "johnandrew321",
      "author_url": "",
      "post_date": "02/19/2022 12:38:37",
      "content": "<p>thanks for this knowledge</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1698591,
      "author_name": "nyanswanaung",
      "author_url": "",
      "post_date": "02/20/2022 14:06:00",
      "content": "<p>May I know how many epochs did you train and which index did you use for splitting? Sequence of Video ID? Stratified or GroupK?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1698656,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/20/2022 15:02:47",
          "content": "<p>I trained for 10 epochs using cosine schedule with 1 epoch warm up. I used the epoch model weights with the best F2 score which was usually epoch 7. For CV, i used group KFold on subsequences <a href=\"https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\" target=\"_blank\">here</a> (which are different than sequences).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1726224,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/17/2022 19:18:45",
      "content": "<p>UPDATE: This technique helped win 2nd place in Kaggle NLP competition <a href=\"https://www.kaggle.com/c/feedback-prize-2021/discussion/313389\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1690489": "# Single Model Yolov5L6\nThanks Kaggle for a fun competition. Thank you Kagglers for sharing many great discussions and notebooks! My final solution is a single 10-fold Yolo5L6 without TTA without tracking. Train 3072, infer 3072 img size.\n\n# WBF with Threshold\nJust this morning, I realized that the awesome WBF GitHub repository [here][1] doesn't apply a threshold after fusing boxes. So with 10 hours until competition deadline, I updated my final submission and submitted again! It finished just in time and jumped me into Silver!\n\nIf you have 10 fold models. And each model is inferred box with `conf = 0.1`. Then afterward when we apply WBF, it is possible that 3 models find a starfish with confidences `0.3, 0.2, 0.1` and the other 7 fold models do not find that box. Then the average confidence is `0.06 = (0.3 + 0.2 + 0.1 + 0 + 0 + 0 + 0 + 0 + 0 + 0)/10`. So we need to remove this box (if we don't wish to have boxes under `conf = 0.1`) as follows:\n\n    WBF_CONF = 0.1\n    boxes, scores, labels = weighted_boxes_fusion(ALL_MODELS)\n    filtered boxes=[], filtered_scores=[]\n    for k, box in enumerate(boxes):\n        if scores[k]<WBF_CONF: continue\n        filtered_boxes.append(box)\n        filtered_scores.append(scores[k])\n\nUsing this code boosted my final submission Private LB 0.609 to **Private LB 0.684**! Wow. This is important to do if the competition metric is F2 because we are penalized with too many FP (so it's important to establish a conf threshold). However if the competition metric is `mAP` then it would not matter since with `mAP` we are not penalized for FP (explained [here][8])\n\n# Skip_Box_Thr is not WBF_CONF\nNote that `weighted_boxes_fusion` has a parameter called `skip_box_thr`. This filters input boxes not output boxes. In general, this parameter makes no difference because we already use `model.conf = 0.1` which means that we only input boxes with `conf>=0.1` into our WBF. There is no parameter for filtering output boxes. So when competition metric is `F1` or `F2`, we must write our own code like above.\n\n# Model Training - 10 Folds\nFor my object detection model, I trained 10 folds of `Yolov5L6` with `image size = 3072`, `batch_size = 8`, and `epochs = 10` using 4xV100 32GB GPU. Thanks Nvidia! Training took 2.5 hours per fold (i.e. 15 minutes per epoch). I tried a few image sizes and this achieved the best CV score.\n\nI used @steamedsheep training script [here][2] and @hjulian3833 subsequence CV folds [here][3]. I used the best model each epoch as determined by Yolo F2 metric described [here][4] and [here][5]. I trained using positive frames only.\n\n# Model Inference - 10 Folds\nFor inference, I inferred 10 fold models with @awsaf49 notebook [here][6]. I inferred at `image size = 3072` with `iou = 0.4` and `conf = 0.05`. I inferred at the same size as training since this achieved the best CV score. I ensembled fold models using @zfturbo awesome WBF GitHub library [here][1] and I added my own WBF threshold shown above at `WBF_CONF = 0.1` (since this was best `conf` for CV score).\n\n# No Tracking, No TTA\nUsing the public notebook tracking code did not improve my CV score, so I did not use it. In order to infer 10 folds of Yolov5L6 in 7 hours, I did not use TTA.\n\n# CV 0.732, Public LB 0.591, Private LB 0.684\nTo compute competition metric, I used @bamps53 notebook [here][7]. My final single model achieved 10 fold CV 0.732, Public LB 0.591, and Private LB 0.684\n\n# Thank You\nThank you everyone for sharing helpful discussions and notebooks. My solution was possible because of everyone's generous sharing.\n\n# UPDATE 1 (after comp ended): \nInstead of inferring 10 folds. I just submitted with 5 of the 10 folds and added TTA. The public LB boost 0.018 and the private LB boost 0.011. The private rank boost from 75th place to 40th place! \n\n# UPDATE 2 (after comp ended): \nInstead of adding TTA, I trained a Yolov5M6 at 3072 with same settings as my Yolov5L6. Then I submit 5 folds out of 10 Yolov5M6 (without TTA) and 5 folds out of 10 Yolo5L6 (without TTA). When using `WBF_CONF`, the result is private LB 0.709 and 20th place! This confirms that `WBF_CONF` is a powerful idea. If I add even more models (with different backbones and use different image sizes), I assume the LB will keep climbing!\n\n[1]: https://github.com/ZFTurbo/Weighted-Boxes-Fusion\n[2]: https://www.kaggle.com/steamedsheep/yolov5-high-resolution-training\n[3]: https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences\n[4]: https://www.kaggle.com/sanchitvj\n[5]: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/300638#1658343\n[6]: https://www.kaggle.com/awsaf49/great-barrier-reef-yolov5-infer\n[7]: https://www.kaggle.com/bamps53/competition-metric-implementation\n[8]: https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637",
    "1690507": "Our team has the same idea and same code. We set WBF_CONF = 0.5, so it refinds all prediction. We use 9 different sizes for x9 TTA, and I think the model generates different predictions at different sizes, so refining predictions after using WBF is not appropriate. Thank you for disclosing my mistake and unfinished idea.",
    "1690510": "As always, thanks for sharing your approach @cdeotte ! . I always learn something from reading solutions.",
    "1690514": "Did you train on total data or just data having cots?",
    "1690531": "Thank you for the concise summary. I have one question.\n\n> Using this code boosted my final submission Private LB 0.609 to Private LB 0.684! Wow. This is important to do if the competition metric is F2 because we are penalized with too many FN (so it's important to establish a conf threshold). However if the competition metric is mAP then it would not matter since with mAP we are not penalized for FN (explained here)\n\nCan you please explain more about this? Why does setting the WBF threshold lead to an improvement in FN? Since the number of bboxes is reduced, it seems that FN is increased but not decreased. Doesn't it mean that the F2 score improves by decreasing the FP?",
    "1690532": "Hey @cdeotte , thanks for sharing! \n\nMy approach is very similar to your single YOLOv5 10-folds model with WBF, but using `yolov5m6` trained with `2560` resolution and heavy augmentation (**Private LB: 0.690**):\n```\n[\n    A.RandomSizedBBoxSafeCrop(1280, 1280, erosion_rate=0.2,\n        interpolation=cv2.INTER_LINEAR, p=1.0),\n    A.OneOf([\n        A.MotionBlur(p=.2),\n        A.MedianBlur(blur_limit=3, p=0.3),\n        A.Blur(blur_limit=3, p=0.1),\n    ], p=0.3),\n    A.OneOf([\n        A.CLAHE(clip_limit=2, tile_grid_size=(8, 8), p=0.5),\n        A.RandomBrightnessContrast(brightness_limit=0.2,\n            contrast_limit=0.2, brightness_by_max=False, p=0.5),\n        A.RandomGamma(p = 0.5),\n        A.ShiftScaleRotate(shift_limit=0.0625,\n            scale_limit=0.2, rotate_limit=45, p=0.2),\n    ], p=0.3),\n    A.HorizontalFlip(p=0.5),\n    A.Sharpen(p=0.3),\n    A.RGBShift(p=0.2),\n]\n```\nI think the key element would be using `RandomSizedBBoxSafeCrop` to crop `1280x1280` areas from `2560` resolution with `erosion_rate=0.2`, which helped the model to learn to identify smaller objects. I selected the best F2 version (train about `11-20 epochs` per fold, with extra `20%` background images).\n\nInference is done at the same size (2560) with `TTA` in `7/10 folds` to squeeze time spent (no tracking). As for NMS `conf` and `IOU` of the model, they were set to `0.5` and `0.45` to keep more bboxes with higher confidence, then for WBF I set `iou_thr=0.4` and `skip_box_thr=0.01`.\n\nhttps://www.kaggle.com/markpeng/yolov5m6-f10-v8-2560-wbf-ensemble/settings?scriptVersionId=87258324\n\nTried training by tiling images under x2/x4/x8 smaller resolutions resized to `1280` resolution and merge tile predictions with NMS + WBF, but it turned out to be overfitting the F2 even with CV splits by subsequences (best/F2: `0.86987`).\n\nCheers, I also learnt a lot from the others' great works!",
    "1690588": "Thanks for sharing",
    "1690632": "Great Work @cdeotte. I can't believe it never occurred to me. WBF does have a **threshold** option `skip_box_thr`, which only filters only input boxes, not the fused ones.",
    "1690667": "Yes @awsaf49 , I never realized this before either until just this morning. In previous competitions with `mAP` metric it didn't matter (like our SIIM-FISABIO-RSNA COVID-19 Detection comp) but in this competition with `F2` (or `F1`) it does matter.\n\nI discovered when I increased my inference from using 6 out of 10 folds to 9 out of 10 folds. I couldn't understand why the LB score decreased. Then all of a sudden, it occurred to me.",
    "1690674": "Awesome work @markpeng . Congrats on strong Silver solo finish! Using `RandomSizedBBoxSafeCrop` is a great idea. That let's you train with large images and have large batches. What does `erosion_rate` do?\n\nI think since you used a high `conf = 0.5` for each model, then if only 1 of your 7 models predicted a bbox, it would become `0.5 / 7 = 0.07` probability which is still acceptable. Alternatively, you could lower each model to `conf = 0.1` and use `WBF_CONF = 0.1` like the code in my post above.\n\nNote that `skip_box_thr=0.01` filters input boxes which in your case does nothing since each input box is already `>=0.5`.",
    "1690680": "I train on just images with cots. I tried using `1:1` frames `without:with`. And `2:1` and `3:1`. But using `1:0` was best. \n\nNote that even when we train with 100% frames with cots, the model still learns background without cots. Many people don't realize this. For each training image, the model will place thousands of different sized rectangles on the train image. Then some rectangles will have cots and some will not. The model will learning to predict which boxes have cots and which boxes do not. So the model learns what ocean without cots looks like (even when we use all positive images).",
    "1690682": "Thanks @trushk . Congrats to you and your team for 6th place Gold. Fantastic!",
    "1690747": "It is the same idea as the following. If we set `model.conf = 0.01`, does that help or hurt? This hurts because the model predicts too many boxes.\n\nWhen we ensemble 10 models, if the majority of models predict the same box, then it is most likely a TP. But what happens when only 1 model predicts a box, and the other 9 models do not predict that box? (It is most likely FP).\n\nIn this case, WBF takes the average of the first model's box conf with 9 zeros. So the resultant box conf becomes very small like `conf = 0.02` or something. The library WBF does not remove this box even though the resultant `conf` is very small. You must write your own code like mine above to remove boxes like this.",
    "1690759": "Hey @cdeotte , Great Work! \n\nCan I ask what for what augmentations you used or how you tuned your hyperparams? I was planning to use the wandb sweeps that were integrated with YOLOV5, but I didn't have enough time.",
    "1690773": "Great work @cdeotte even for late join to competition! :) from our observation WBFT (this is how we called it) was also working",
    "1690813": "Thanks for sharing. Nice work!",
    "1690881": "thank you so much, congrats on your medal",
    "1690896": "cdeotte yes we came to the same conclution and develop WBFT (mentioned in my post [Experiments list - competition summary from our team perspective](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307504). Having 4xGPU was really great in this competition ... not only because of image resolution experiments but because of speed up process. I am really wondering how you CV procedure looked like. Could you share your thoughts on that?",
    "1690907": "thx for sharing.  reading discussion after competition  can learn a lot stuff.",
    "1690940": "Thanks for your `WBF_CONF`! Think I will never discover it without your reminder. Now I know why it didn't get better when I use WBF 😂\n\nWe trained with `image_size = 3000` and infer with `image_size = 6400`. Instead of yolov5L6, we adopt yolov5s. Tracking didn't improve the score, but TTA yes! Finally, we got a public LB 0.661 and private 0.673. \n\nBy the way, `mixup` doesn't seem to work, at least for us, does it really work here?\n\nI'm a new kaggler and really benefit a lot from your Great Work, thank you :)",
    "1691531": "Hi, i used sheep's augmentations in his notebook [here][1]. Check out code cell 5. I experimented by changing the default anchors in `yolov5l6.yaml` and changed some augmentations but my final solution uses sheep's script as is and uses yolo default anchors. (I joined late and didn't have enough time to find better).\n\nThe main things i changed were model backbone, image size, batch size, number of epochs, and how many frames without cots to add. I picked one fold and kept training that fold over and over with different settings.\n\n[1]: https://www.kaggle.com/steamedsheep/yolov5-high-resolution-training/notebook",
    "1691563": "Congratulations Editoxic. You did well.",
    "1691577": "Hi Remek, congratulations to you and team on achieving 34th place Silver. Thanks for all your sharing.\n\nFor CV, I used public notebook subsequences (not sequences) [here][1] and did stratified group K fold. The following code gives each fold the same number of annotated frames and keeps subsequences in their own fold:\n\n    from sklearn.model_selection import StratifiedGroupKFold\n    sgkf = StratifiedGroupKFold(n_splits=10)\n    for fold, (t_idx, v_idx) in enumerate( sgkf.split(df, df.has_annotations, df.subsequence_id) ):\n        df.loc[v_idx,'fold'] = fold\n\nIn retrospect, since the test data was different videos, it may have been better to use 3 folds with each video in its own fold. Then after finding the best hyparameters, train with all 3 videos.\n\nFor my experiments, i mainly ran the same `fold=5` over and over with different settings. I mainly tried different Yolo backbone, image size, batch size, number of epochs. I didn't have too much time to adjust augmentations nor try models other than Yolov5.\n\n[1]: https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences",
    "1691580": "Thanks Lukasz. Congratulations to you and your team!\n\nWBFT works great. I guess as an alternative some teams may have just increased `model.conf`. For example, if you want `WBF_CONF = 0.1`, you could probably use `model.conf = 0.5` and use 5 models in ensemble. Then the smallest `conf` you will get after WBF without WBFT is `0.5/5 = 0.1`. But using WBFT with `model.conf = 0.05` and `WBF_CONF = 0.1` should achieve a better CV LB.",
    "1691588": "Thank you for everything in this competition! 🙏🙏🙏",
    "1691629": "Wow, that's it!",
    "1691646": "cdeotte As I understood `erosion_rate` controls how much area of the original bounding box could be lost after cropping. `erosion_rate = 0.2` means the augmented bounding box's area could be up to 20% smaller than the area of the original bbox before resizing.\nCheckout this tutorial: https://albumentations.ai/docs/examples/example_bboxes2/\n\nThanks for the extra insight about conf-based post filter for WBF fused bboxes.\nAnd you are right about the `skip_box_thr` filter, it didn't take any effect in my case.",
    "1691686": "cdeotte Thanks for the reply.\nI understood from your explanation that WBFT is to reduce FP, not FN.\nAs a matter of fact, my team doesn't have such a process like WBFT, so I'll try it later.",
    "1691774": "kmizunoster thanks for pointing this out. I updated my discussion to say \"too many FP\". I incorrectly said \"too many FN\". Using WBFT decreases FP not FN. Using more than 1 model (i.e. ensembling) helps decrease FN but we must balance the addition of boxes with WBFT so that we don't have too many FP.",
    "1691776": "UPDATE: I updated post to say \"WBFT decreases FP\". It does not decrease FN. Using more than 1 model, i.e. ensembling a diversity of models is what decreases FN. But we must use WBFT (instead of WBF) so that the decrease in FN does not add too many FP.",
    "1691945": "congrats @cdeotte 🌟🎉",
    "1692124": "Thanks Aruna",
    "1692410": "UPDATE 1 (after comp ended): Instead of inferring 10 folds (in 7 hours), I just submitted with 5 of the 10 folds and added TTA to each fold (which takes 9 hours). The public LB boost 0.018 and the private LB boost 0.011. The private rank boost from 75th place to 40th place!\n\nYolo TTA is very powerful because it infers the image at 3 different image sizes which adds more diversity than adding more folds. TTA also does flips. So using TTA is better than using more of the original folds. From Yolo's page [here][1]\n\n>Note that inference with TTA enabled will typically take about 2-3X the time of normal inference as the images are being left-right flipped and processed at 3 different resolutions, with the outputs merged before NMS. Part of the speed decrease is simply due to larger image sizes (832 vs 640), while part is due to the actual TTA operations.\n\n[1]: https://github.com/ultralytics/yolov5/issues/303",
    "1692970": "Thanks vad13irt!",
    "1693138": "UPDATE 2 (after comp ended): Instead of adding TTA, I trained a Yolov5M6 at 3072 with same settings as my Yolov5L6. Then I submit 5 folds out of 10 Yolov5M6 (without TTA) and 5 folds out of 10 Yolo5L6 (without TTA). When using `WBF_CONF`, the result is private LB 0.709 and 20th place! This confirms that `WBF_CONF` is a powerful idea. If I add even more models (with different backbones and use different image sizes), I assume the LB will keep climbing!",
    "1693150": "Thank you for sharing. I thnink WBFT is better name for this one :) I have similar observation. Having threshold is really great way to controll FP prediction in multi model inference.",
    "1693157": "Could you check different approach? \n- No TTA\n- Only WBFT - set model IOU on 0.9 ... generate many many boxes and let WBF to do a job instead of yolo NMS? We used this trick and it helps us a lot.",
    "1693177": "Good idea (IOU=0.9), i will try this.\n\n>Having threshold is really great way to controll FP prediction in multi model inference.\n\nYes, during the comp my best LB was just single model (1 fold) so I didn't focus on ensemble. But with `WBFT` we can benefit from ensemble. Without `WBFT` then using multiple model ensemble adds too many low conf FP and hurts competition metric of F2",
    "1693188": "cdeotte as you did we just cleared all bboxes below treshold (as you elaborated - if one model claims that it can see bbox we just can using both weight and treshold to filter final bboxes - we used weights as well to create \"more\" bboxes from one model).\n\nOur configurations was:\n- iou even 0.95\n- then model.max_det = 100\n- wbft to let make the final bbox preds and ... coordinations computation - it make bboxes more tight to sf\n\nSee here for implementation (this is one of our work in progress notebook) and debugging code: https://www.kaggle.com/remekkinas/cots-y5-wbf",
    "1693219": "I don't understand the following line, can you explain this more?\n>we used weights as well to create \"more\" bboxes from one model",
    "1693236": "I think Remek wanted to say we generated more boxes at one sf like this from one model:\n\n![](https://i.ibb.co/pP0y6br/autowbf1.png)\n\n![](https://i.ibb.co/rHHwBMB/autowbf2.png)",
    "1693245": "Ah, gotcha thanks. I just submitted my models using `IOU=0.9` and WBFT (with IOU=0.4) afterward, i will report back in 9 hours.",
    "1693256": "We used it to change native nms to wbf. When we mixed f.e. 3 models we applied 3 WBF on each model, and then one WBF summary",
    "1693334": "Congratulations and thank you for sharing knowledge!",
    "1693520": "lukaszborecki We should read the code in GitHub WBF. I'm not sure how it computes the new `conf` value when one model has multiple boxes (as result of using `model.conf=0.9`). For example, let's say we ensemble 4 models. And model 1 has 2 boxes (with conf 0.7 and 0.8) that overlap with IOU = 0.6. And model 2 has 1 box (with conf 0.6) that overlaps with those box with IOU = 0.5. And model 3 and 4 has zero boxes.\n\nHow does WBF compute the new `conf`? For example, I do not think it does `(0.7 + 0.8 + 0.6)/3` because the denominator should be 4 for 4 models. So does WBF somehow combine the two boxes from model 1 first? And then do `(0.75 + 0.6 + 0 + 0)/4`? I'm not sure",
    "1693528": "cdeotte I made some calculus when we were submiting.\n\n1st Scenario IOU = 0.9 and one WBF collecting boxes: in case 3 models\n\n1st model may give 3 overlapped boxes of conf 0.8 0.9 and 0.7\n2nd model may give 10 overlapped boxes with mean 0.6\n3rd model may give 0 overlapped boxes\n\nin this scenario wbf got 13 boxes to fuse (10 * 0.6 + 3 * 0.8 +1 * 0) / 14 = 0.6\n\n2nd scenario: WBF on output each model and then WBF all\n\n1st model - box 0.8\n2nd model - box 0.6\n3rd model - box 0.0\n\nFinal WBF = 1.4 / 3= 0.46\n\nDidn't check it with code it was experimental calculation after one ebug inference. So my confidence in this thesis is 0.9 :D\n\nEDIT\n\n-- wrong\n\nin 1st scenario if one model got more than 1 overlap box then models which doesnt predict doesnt lower the score. I got case 9 bounding boxes from 1 model and null from model 2 and 3, and counted mean was from 9 not 11. It seams we need to check source code\n\nEDIT 2:\n\nwhen there were 3 models and one model got 2 overlapped boxes and other models is null it divided by 3. So maybe it check how many models are in box_list , becuase its list of 3 lists. And then divide by 3 or by number of boxes if greater than 3.\n\n-- Checked it on Excel - it works like EDIT 2",
    "1693827": "Congratulations and thank you for sharing knowledge!",
    "1695499": "> Note that even when we train with 100% frames with cots, the model still learns background without cots. \n\nI really like your deep thinking and understanding about the competition. Congrats!",
    "1695700": "congratulations @cdeotte 🤩",
    "1697184": "thanks for this knowledge",
    "1698591": "May I know how many epochs did you train and which index did you use for splitting? Sequence of Video ID? Stratified or GroupK?",
    "1698656": "I trained for 10 epochs using cosine schedule with 1 epoch warm up. I used the epoch model weights with the best F2 score which was usually epoch 7. For CV, i used group KFold on subsequences [here][1] (which are different than sequences).\n\n[1]: https://www.kaggle.com/julian3833/reef-a-cv-strategy-subsequences",
    "1726224": "UPDATE: This technique helped win 2nd place in Kaggle NLP competition [here][1]\n\n[1]: https://www.kaggle.com/c/feedback-prize-2021/discussion/313389"
  },
  "source": "meta"
}