{
  "id": 307718,
  "title": "11th Place solution - Team COTS ",
  "url": "/competitions/tensorflow-great-barrier-reef/writeups/cots-11th-place-solution-team-cots",
  "author_name": "",
  "post_date": "2022-02-15T11:14:40.790Z",
  "votes": 34,
  "comment_count": 9,
  "views": 0,
  "content": "<p>We are happy to survive in the shakeup and it turns out trusting CV is indeed the key here! It is a great team effort and it has been a great journey to work with <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>,  <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> ,  <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> and <a href=\"https://www.kaggle.com/markunys\" target=\"_blank\">@markunys</a> . With this gold medal, 3 of us ( <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> , <a href=\"https://www.kaggle.com/markunys\" target=\"_blank\">@markunys</a>, and me) will become Competition Master, it is just too good to be true 😆</p>\n<p>In general, it is very important to observe the video clip with those model predictions, it helps us a lot for observing some improvement areas, like tracking, ensembling, and adding new data. Here are some key points about our solutions:</p>\n<p>Best private LB model pipeline:<br>\n<img src=\"https://i.postimg.cc/QdFy25t6/COTS-ens-multitracker-0775-drawio-1.png\" alt=\"pipeline\"></p>\n<ol>\n<li><strong>Final Submission Model (2 best CV,  1 best CV &amp; LB,  1 best LB)</strong><ol>\n<li>5 YOLO WBF + 1 Cascade RCNN WBF: only used those predictions from RCNN that has &gt;= 0.3 IOU with YOLO prediction in order to correct YOLO's bounding box).  <strong>CV 775 Public LB 626  Private LB 720</strong></li>\n<li>7 YOLO WBF: <strong>CV 781 Public LB 628  Private 719</strong></li>\n<li>4 YOLO WBF + 1 best LB model WBF:  4 YOLO WBF CV 771,  <strong>Public LB 653, Private LB 706</strong></li>\n<li>2 best LB model WBF: Not sure CV,  <strong>Public LB 718, Private LB 666</strong>.</li></ol></li>\n<li><strong>Ensemble, final submission <a href=\"https://www.kaggle.com/vincentwang25/11th-place-solution-100-iterations-journey\" target=\"_blank\">notebook</a></strong><ol>\n<li>Model inference with 1.3 x training size (multiscale inference in yolov5)</li>\n<li>low confidence threshold for every single model and high confidence threshold after WBF  (<strong>~0.005 CV improvement</strong> compared with using high conf threshold before and low conf threshold after WBF)</li>\n<li>Model picking criteria:  diversity from data (CLAHE preprocessing or not), training method (1 stage, 2 stages), and model (yolov5 or rcnn)</li></ol></li>\n<li><strong>CV scheme</strong><ol>\n<li>manually picked sequence that represents 20% of data (4707 images)  --  balancing between representative and data size. <strong>This results in a final CV and PVT LB correlation of 90%.</strong> <ol>\n<li>Pick criteria: # of COTS per frame + how hard it is to predict by f2 score (we don't wanna pick a subset of sequence that is too easy or too hard to predict) + sequence across different video.</li></ol></li>\n<li>Not using 5 folds because of computational constraints.  Not using video_id because it takes too much data away.</li></ol></li>\n<li><strong>Training</strong><ol>\n<li>training with all annotated images +  ~5% background image  (those background images are picked from those high FP images or randomly, it <strong>improves CV round 0.005</strong>).</li>\n<li>Using CV to choose hyperparameter and best epoch number --&gt; training with all data for submission.</li>\n<li>YOLOv5: default hyperparameter (mixup, mosaic, flip, HSV, translate, scale) but with learning rate 0.001 and Adam optimizer.<ol>\n<li>1 stage model: training with GT data for 20 ~ 40 epochs, batch size 4, training image size 2400~3600 (depending on model size).</li>\n<li>2 stage model: training with original data like 1 stage model, then with GT data for 10 epochs with 0.0001 LR rate.</li>\n<li>Some models also used <code>--label-smoothing 0.2</code> for faster convergence and higher CV score.</li></ol></li>\n<li>Cascade-RCNN:  mmdetection framework with augmentation flip, rotate, CLAHE, HueSaturation, RandomBrightnessContrast, RandomSizedBBoxSafeCrop.</li></ol></li>\n<li><strong>Postprocessing</strong><ol>\n<li>Tracker for each model before WBF (<strong>~ 0.01 CV improvement</strong> compared with using tracker after WBF)</li>\n<li>For the tracker, we filtered out those predictions from trackers that are right on the edge because they are usually FP (this <strong>increases CV by ~0.002</strong>).  Also, we use <code>initialization_delay=2</code> instead of 1 to depress too many FP, this <strong>improves CV by ~0.005</strong>.</li>\n<li>Using Optical Flow (<a href=\"https://www.kaggle.com/yamsam/optical-flow-by-raft\" target=\"_blank\">notebook</a>) to reset tracker to reduce FP (~ 0.01 private LB improvement)</li></ol></li>\n<li><strong>Data</strong><ol>\n<li>GT data: Add COTS through WBF prediction: Some real COTS is not marked in original data. We added it manually through observing the predicted video clip. In total we added 1213 COTS, around 10% of the original COTS number. This <strong>improves CV ~0.01</strong>.</li></ol></li>\n<li><strong>Things didn't work</strong><ol>\n<li>SAHI</li>\n<li>modify tracker to reduce obvious FP during each sequence.</li>\n<li>Using other datasets, like <a href=\"https://github.com/chongweiliu/DUO\" target=\"_blank\">DUO</a></li>\n<li>Trying to use best CV model and best LB model together to improve LB score but never get better than original best LB score (due to the tightness theory…..)</li>\n<li>YOLOX, YOLOR, YOLOv5-p7, Swin Transformer with Retina, fasterRCNN and TPH YOLOV5.</li>\n<li>Tried to add mosaic and mixup in mmdetection framework but didn't work.</li></ol></li>\n<li><strong>Hardware / # of gpu hours</strong><ol>\n<li>3 x RTX 3090 +  2 x 2080Ti + 1~2 GPU with memory 34GB + colab pro</li>\n<li>GPU hours 850+ hours</li></ol></li>\n</ol>\n<p>List of final ensemble candidates and their CV score in case someone is interested.<br>\n<img src=\"https://i.postimg.cc/VNcWKbb0/image.png\" alt=\"pic\"></p>\n<p>In the end, thank you to the great Kaggle community for many amazing sharing throughout the process, it is an exciting learning experience!🔥🔥</p>",
  "messages": [
    {
      "id": "1691344",
      "postDate": "02/15/2022 11:07:07",
      "content": "<p>We are happy to survive in the shakeup and it turns out trusting CV is indeed the key here! It is a great team effort and it has been a great journey to work with <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a>,  <a href=\"https://www.kaggle.com/yamsam\" target=\"_blank\">@yamsam</a> ,  <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> and <a href=\"https://www.kaggle.com/markunys\" target=\"_blank\">@markunys</a> . With this gold medal, 3 of us ( <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> , <a href=\"https://www.kaggle.com/markunys\" target=\"_blank\">@markunys</a>, and me) will become Competition Master, it is just too good to be true 😆</p>\n<p>In general, it is very important to observe the video clip with those model predictions, it helps us a lot for observing some improvement areas, like tracking, ensembling, and adding new data. Here are some key points about our solutions:</p>\n<p>Best private LB model pipeline:<br>\n<img src=\"https://i.postimg.cc/QdFy25t6/COTS-ens-multitracker-0775-drawio-1.png\" alt=\"pipeline\"></p>\n<ol>\n<li><strong>Final Submission Model (2 best CV,  1 best CV &amp; LB,  1 best LB)</strong><ol>\n<li>5 YOLO WBF + 1 Cascade RCNN WBF: only used those predictions from RCNN that has &gt;= 0.3 IOU with YOLO prediction in order to correct YOLO's bounding box).  <strong>CV 775 Public LB 626  Private LB 720</strong></li>\n<li>7 YOLO WBF: <strong>CV 781 Public LB 628  Private 719</strong></li>\n<li>4 YOLO WBF + 1 best LB model WBF:  4 YOLO WBF CV 771,  <strong>Public LB 653, Private LB 706</strong></li>\n<li>2 best LB model WBF: Not sure CV,  <strong>Public LB 718, Private LB 666</strong>.</li></ol></li>\n<li><strong>Ensemble, final submission <a href=\"https://www.kaggle.com/vincentwang25/11th-place-solution-100-iterations-journey\" target=\"_blank\">notebook</a></strong><ol>\n<li>Model inference with 1.3 x training size (multiscale inference in yolov5)</li>\n<li>low confidence threshold for every single model and high confidence threshold after WBF  (<strong>~0.005 CV improvement</strong> compared with using high conf threshold before and low conf threshold after WBF)</li>\n<li>Model picking criteria:  diversity from data (CLAHE preprocessing or not), training method (1 stage, 2 stages), and model (yolov5 or rcnn)</li></ol></li>\n<li><strong>CV scheme</strong><ol>\n<li>manually picked sequence that represents 20% of data (4707 images)  --  balancing between representative and data size. <strong>This results in a final CV and PVT LB correlation of 90%.</strong> <ol>\n<li>Pick criteria: # of COTS per frame + how hard it is to predict by f2 score (we don't wanna pick a subset of sequence that is too easy or too hard to predict) + sequence across different video.</li></ol></li>\n<li>Not using 5 folds because of computational constraints.  Not using video_id because it takes too much data away.</li></ol></li>\n<li><strong>Training</strong><ol>\n<li>training with all annotated images +  ~5% background image  (those background images are picked from those high FP images or randomly, it <strong>improves CV round 0.005</strong>).</li>\n<li>Using CV to choose hyperparameter and best epoch number --&gt; training with all data for submission.</li>\n<li>YOLOv5: default hyperparameter (mixup, mosaic, flip, HSV, translate, scale) but with learning rate 0.001 and Adam optimizer.<ol>\n<li>1 stage model: training with GT data for 20 ~ 40 epochs, batch size 4, training image size 2400~3600 (depending on model size).</li>\n<li>2 stage model: training with original data like 1 stage model, then with GT data for 10 epochs with 0.0001 LR rate.</li>\n<li>Some models also used <code>--label-smoothing 0.2</code> for faster convergence and higher CV score.</li></ol></li>\n<li>Cascade-RCNN:  mmdetection framework with augmentation flip, rotate, CLAHE, HueSaturation, RandomBrightnessContrast, RandomSizedBBoxSafeCrop.</li></ol></li>\n<li><strong>Postprocessing</strong><ol>\n<li>Tracker for each model before WBF (<strong>~ 0.01 CV improvement</strong> compared with using tracker after WBF)</li>\n<li>For the tracker, we filtered out those predictions from trackers that are right on the edge because they are usually FP (this <strong>increases CV by ~0.002</strong>).  Also, we use <code>initialization_delay=2</code> instead of 1 to depress too many FP, this <strong>improves CV by ~0.005</strong>.</li>\n<li>Using Optical Flow (<a href=\"https://www.kaggle.com/yamsam/optical-flow-by-raft\" target=\"_blank\">notebook</a>) to reset tracker to reduce FP (~ 0.01 private LB improvement)</li></ol></li>\n<li><strong>Data</strong><ol>\n<li>GT data: Add COTS through WBF prediction: Some real COTS is not marked in original data. We added it manually through observing the predicted video clip. In total we added 1213 COTS, around 10% of the original COTS number. This <strong>improves CV ~0.01</strong>.</li></ol></li>\n<li><strong>Things didn't work</strong><ol>\n<li>SAHI</li>\n<li>modify tracker to reduce obvious FP during each sequence.</li>\n<li>Using other datasets, like <a href=\"https://github.com/chongweiliu/DUO\" target=\"_blank\">DUO</a></li>\n<li>Trying to use best CV model and best LB model together to improve LB score but never get better than original best LB score (due to the tightness theory…..)</li>\n<li>YOLOX, YOLOR, YOLOv5-p7, Swin Transformer with Retina, fasterRCNN and TPH YOLOV5.</li>\n<li>Tried to add mosaic and mixup in mmdetection framework but didn't work.</li></ol></li>\n<li><strong>Hardware / # of gpu hours</strong><ol>\n<li>3 x RTX 3090 +  2 x 2080Ti + 1~2 GPU with memory 34GB + colab pro</li>\n<li>GPU hours 850+ hours</li></ol></li>\n</ol>\n<p>List of final ensemble candidates and their CV score in case someone is interested.<br>\n<img src=\"https://i.postimg.cc/VNcWKbb0/image.png\" alt=\"pic\"></p>\n<p>In the end, thank you to the great Kaggle community for many amazing sharing throughout the process, it is an exciting learning experience!🔥🔥</p>",
      "rawMarkdown": "We are happy to survive in the shakeup and it turns out trusting CV is indeed the key here! It is a great team effort and it has been a great journey to work with @anjum48,  @yamsam ,  @imeintanis and @markunys . With this gold medal, 3 of us ( @imeintanis , @markunys, and me) will become Competition Master, it is just too good to be true 😆\n\nIn general, it is very important to observe the video clip with those model predictions, it helps us a lot for observing some improvement areas, like tracking, ensembling, and adding new data. Here are some key points about our solutions:\n\nBest private LB model pipeline:\n![pipeline](https://i.postimg.cc/QdFy25t6/COTS-ens-multitracker-0775-drawio-1.png)\n\n1. **Final Submission Model (2 best CV,  1 best CV & LB,  1 best LB)**\n    1. 5 YOLO WBF + 1 Cascade RCNN WBF: only used those predictions from RCNN that has >= 0.3 IOU with YOLO prediction in order to correct YOLO's bounding box).  **CV 775 Public LB 626  Private LB 720**\n    2. 7 YOLO WBF: **CV 781 Public LB 628  Private 719**\n    3. 4 YOLO WBF + 1 best LB model WBF:  4 YOLO WBF CV 771,  **Public LB 653, Private LB 706**\n    4. 2 best LB model WBF: Not sure CV,  **Public LB 718, Private LB 666**.\n2. **Ensemble, final submission [notebook](https://www.kaggle.com/vincentwang25/11th-place-solution-100-iterations-journey)**\n    1. Model inference with 1.3 x training size (multiscale inference in yolov5)\n    2. low confidence threshold for every single model and high confidence threshold after WBF  (**~0.005 CV improvement** compared with using high conf threshold before and low conf threshold after WBF)\n    3. Model picking criteria:  diversity from data (CLAHE preprocessing or not), training method (1 stage, 2 stages), and model (yolov5 or rcnn)\n3. **CV scheme**\n    1. manually picked sequence that represents 20% of data (4707 images)  --  balancing between representative and data size. **This results in a final CV and PVT LB correlation of 90%.** \n         1. Pick criteria: # of COTS per frame + how hard it is to predict by f2 score (we don't wanna pick a subset of sequence that is too easy or too hard to predict) + sequence across different video.\n    2. Not using 5 folds because of computational constraints.  Not using video_id because it takes too much data away.\n4. **Training**\n    1. training with all annotated images +  ~5% background image  (those background images are picked from those high FP images or randomly, it **improves CV round 0.005**).\n    2. Using CV to choose hyperparameter and best epoch number --> training with all data for submission.\n    3. YOLOv5: default hyperparameter (mixup, mosaic, flip, HSV, translate, scale) but with learning rate 0.001 and Adam optimizer.\n        1. 1 stage model: training with GT data for 20 ~ 40 epochs, batch size 4, training image size 2400~3600 (depending on model size).\n        2. 2 stage model: training with original data like 1 stage model, then with GT data for 10 epochs with 0.0001 LR rate.\n        3. Some models also used `--label-smoothing 0.2` for faster convergence and higher CV score.\n    4. Cascade-RCNN:  mmdetection framework with augmentation flip, rotate, CLAHE, HueSaturation, RandomBrightnessContrast, RandomSizedBBoxSafeCrop.\n5. **Postprocessing**\n    1. Tracker for each model before WBF (**~ 0.01 CV improvement** compared with using tracker after WBF)\n    2. For the tracker, we filtered out those predictions from trackers that are right on the edge because they are usually FP (this **increases CV by ~0.002**).  Also, we use `initialization_delay=2` instead of 1 to depress too many FP, this **improves CV by ~0.005**.\n    3. Using Optical Flow ([notebook](https://www.kaggle.com/yamsam/optical-flow-by-raft)) to reset tracker to reduce FP (~ 0.01 private LB improvement)\n4. **Data**\n    1. GT data: Add COTS through WBF prediction: Some real COTS is not marked in original data. We added it manually through observing the predicted video clip. In total we added 1213 COTS, around 10% of the original COTS number. This **improves CV ~0.01**.\n5. **Things didn't work**\n    1. SAHI\n    2. modify tracker to reduce obvious FP during each sequence.\n    3. Using other datasets, like [DUO](https://github.com/chongweiliu/DUO)\n    4. Trying to use best CV model and best LB model together to improve LB score but never get better than original best LB score (due to the tightness theory.....)\n    5. YOLOX, YOLOR, YOLOv5-p7, Swin Transformer with Retina, fasterRCNN and TPH YOLOV5.\n    6. Tried to add mosaic and mixup in mmdetection framework but didn't work.\n6. **Hardware / # of gpu hours**\n    1. 3 x RTX 3090 +  2 x 2080Ti + 1~2 GPU with memory 34GB + colab pro\n    2. GPU hours 850+ hours\n\nList of final ensemble candidates and their CV score in case someone is interested.\n![pic](https://i.postimg.cc/VNcWKbb0/image.png)\n\nIn the end, thank you to the great Kaggle community for many amazing sharing throughout the process, it is an exciting learning experience!🔥🔥",
      "votes": null
    },
    {
      "id": "1691391",
      "postDate": "02/15/2022 11:40:17",
      "content": "<p>WoW! Great solution description. Thank you for sharing and congratulations! As I understad:<br>\nf2_t stands for avg f2 score, P_t - precision, R_t recall?  </p>",
      "rawMarkdown": "WoW! Great solution description. Thank you for sharing and congratulations! As I understad:\nf2_t stands for avg f2 score, P_t - precision, R_t recall?",
      "votes": null
    },
    {
      "id": "1691413",
      "postDate": "02/15/2022 11:56:37",
      "content": "<p>Thank you Remek, I learnt a lot from your sharing, thank you!<br>\nYes! <code>_t</code> stands for tracking, so those score are with tracking added :D </p>",
      "rawMarkdown": "Thank you Remek, I learnt a lot from your sharing, thank you!\nYes! `_t` stands for tracking, so those score are with tracking added :D",
      "votes": null
    },
    {
      "id": "1691457",
      "postDate": "02/15/2022 12:20:01",
      "content": "<p>Ohhh … now I understand. I am analyzing your solution to achieve better result in future :) Now it is time to learn from top guys :)</p>",
      "rawMarkdown": "Ohhh ... now I understand. I am analyzing your solution to achieve better result in future :) Now it is time to learn from top guys :)",
      "votes": null
    },
    {
      "id": "1691623",
      "postDate": "02/15/2022 14:08:41",
      "content": "<p>My congratulations, <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a>! I learn a lot from your solution! </p>",
      "rawMarkdown": "My congratulations, @vincentwang25! I learn a lot from your solution!",
      "votes": null
    },
    {
      "id": "1691624",
      "postDate": "02/15/2022 14:10:36",
      "content": "<blockquote>\n  <p>1 stage model: training with GT data for 20 ~ 40 epochs, batch size 4, training image size 2400~3600 (depending on model size).</p>\n</blockquote>\n<p>The batch size is 4, it is quite small, did you try Gradient Accumulation? </p>",
      "rawMarkdown": "> 1 stage model: training with GT data for 20 ~ 40 epochs, batch size 4, training image size 2400~3600 (depending on model size).\n\nThe batch size is 4, it is quite small, did you try Gradient Accumulation?",
      "votes": null
    },
    {
      "id": "1691647",
      "postDate": "02/15/2022 14:33:38",
      "content": "<p>YoloV5 itself uses nominal batch size 64 and grad accumulation accordingly </p>",
      "rawMarkdown": "YoloV5 itself uses nominal batch size 64 and grad accumulation accordingly",
      "votes": null
    },
    {
      "id": "1692576",
      "postDate": "02/16/2022 05:53:49",
      "content": "<p>Congrates on the gold for your team! <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a> <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> Congrats on becoming Competition Master!</p>",
      "rawMarkdown": "Congrates on the gold for your team! @vincentwang25 @imeintanis Congrats on becoming Competition Master!",
      "votes": null
    },
    {
      "id": "1693098",
      "postDate": "02/16/2022 12:35:07",
      "content": "<p>Thank you Richard 😆</p>",
      "rawMarkdown": "Thank you Richard 😆",
      "votes": null
    },
    {
      "id": "1720212",
      "postDate": "03/12/2022 15:06:42",
      "content": "<p>Late congrats! Could you please share a clearer img of the ensemble candidates? I am quite interested in evaluating how much ensemble can help. Thanks a lot!</p>",
      "rawMarkdown": "Late congrats! Could you please share a clearer img of the ensemble candidates? I am quite interested in evaluating how much ensemble can help. Thanks a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1691391,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/15/2022 11:40:17",
      "content": "<p>WoW! Great solution description. Thank you for sharing and congratulations! As I understad:<br>\nf2_t stands for avg f2 score, P_t - precision, R_t recall?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1691413,
          "author_name": "vincentwang25",
          "author_url": "",
          "post_date": "02/15/2022 11:56:37",
          "content": "<p>Thank you Remek, I learnt a lot from your sharing, thank you!<br>\nYes! <code>_t</code> stands for tracking, so those score are with tracking added :D </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691457,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/15/2022 12:20:01",
          "content": "<p>Ohhh … now I understand. I am analyzing your solution to achieve better result in future :) Now it is time to learn from top guys :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691623,
      "author_name": "vad13irt",
      "author_url": "",
      "post_date": "02/15/2022 14:08:41",
      "content": "<p>My congratulations, <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a>! I learn a lot from your solution! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1691624,
          "author_name": "vad13irt",
          "author_url": "",
          "post_date": "02/15/2022 14:10:36",
          "content": "<blockquote>\n  <p>1 stage model: training with GT data for 20 ~ 40 epochs, batch size 4, training image size 2400~3600 (depending on model size).</p>\n</blockquote>\n<p>The batch size is 4, it is quite small, did you try Gradient Accumulation? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691647,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "02/15/2022 14:33:38",
          "content": "<p>YoloV5 itself uses nominal batch size 64 and grad accumulation accordingly </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1692576,
      "author_name": "richx86",
      "author_url": "",
      "post_date": "02/16/2022 05:53:49",
      "content": "<p>Congrates on the gold for your team! <a href=\"https://www.kaggle.com/vincentwang25\" target=\"_blank\">@vincentwang25</a> <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> Congrats on becoming Competition Master!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1693098,
          "author_name": "vincentwang25",
          "author_url": "",
          "post_date": "02/16/2022 12:35:07",
          "content": "<p>Thank you Richard 😆</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1720212,
      "author_name": "toongzhhang",
      "author_url": "",
      "post_date": "03/12/2022 15:06:42",
      "content": "<p>Late congrats! Could you please share a clearer img of the ensemble candidates? I am quite interested in evaluating how much ensemble can help. Thanks a lot!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1691344": "We are happy to survive in the shakeup and it turns out trusting CV is indeed the key here! It is a great team effort and it has been a great journey to work with @anjum48,  @yamsam ,  @imeintanis and @markunys . With this gold medal, 3 of us ( @imeintanis , @markunys, and me) will become Competition Master, it is just too good to be true 😆\n\nIn general, it is very important to observe the video clip with those model predictions, it helps us a lot for observing some improvement areas, like tracking, ensembling, and adding new data. Here are some key points about our solutions:\n\nBest private LB model pipeline:\n![pipeline](https://i.postimg.cc/QdFy25t6/COTS-ens-multitracker-0775-drawio-1.png)\n\n1. **Final Submission Model (2 best CV,  1 best CV & LB,  1 best LB)**\n    1. 5 YOLO WBF + 1 Cascade RCNN WBF: only used those predictions from RCNN that has >= 0.3 IOU with YOLO prediction in order to correct YOLO's bounding box).  **CV 775 Public LB 626  Private LB 720**\n    2. 7 YOLO WBF: **CV 781 Public LB 628  Private 719**\n    3. 4 YOLO WBF + 1 best LB model WBF:  4 YOLO WBF CV 771,  **Public LB 653, Private LB 706**\n    4. 2 best LB model WBF: Not sure CV,  **Public LB 718, Private LB 666**.\n2. **Ensemble, final submission [notebook](https://www.kaggle.com/vincentwang25/11th-place-solution-100-iterations-journey)**\n    1. Model inference with 1.3 x training size (multiscale inference in yolov5)\n    2. low confidence threshold for every single model and high confidence threshold after WBF  (**~0.005 CV improvement** compared with using high conf threshold before and low conf threshold after WBF)\n    3. Model picking criteria:  diversity from data (CLAHE preprocessing or not), training method (1 stage, 2 stages), and model (yolov5 or rcnn)\n3. **CV scheme**\n    1. manually picked sequence that represents 20% of data (4707 images)  --  balancing between representative and data size. **This results in a final CV and PVT LB correlation of 90%.** \n         1. Pick criteria: # of COTS per frame + how hard it is to predict by f2 score (we don't wanna pick a subset of sequence that is too easy or too hard to predict) + sequence across different video.\n    2. Not using 5 folds because of computational constraints.  Not using video_id because it takes too much data away.\n4. **Training**\n    1. training with all annotated images +  ~5% background image  (those background images are picked from those high FP images or randomly, it **improves CV round 0.005**).\n    2. Using CV to choose hyperparameter and best epoch number --> training with all data for submission.\n    3. YOLOv5: default hyperparameter (mixup, mosaic, flip, HSV, translate, scale) but with learning rate 0.001 and Adam optimizer.\n        1. 1 stage model: training with GT data for 20 ~ 40 epochs, batch size 4, training image size 2400~3600 (depending on model size).\n        2. 2 stage model: training with original data like 1 stage model, then with GT data for 10 epochs with 0.0001 LR rate.\n        3. Some models also used `--label-smoothing 0.2` for faster convergence and higher CV score.\n    4. Cascade-RCNN:  mmdetection framework with augmentation flip, rotate, CLAHE, HueSaturation, RandomBrightnessContrast, RandomSizedBBoxSafeCrop.\n5. **Postprocessing**\n    1. Tracker for each model before WBF (**~ 0.01 CV improvement** compared with using tracker after WBF)\n    2. For the tracker, we filtered out those predictions from trackers that are right on the edge because they are usually FP (this **increases CV by ~0.002**).  Also, we use `initialization_delay=2` instead of 1 to depress too many FP, this **improves CV by ~0.005**.\n    3. Using Optical Flow ([notebook](https://www.kaggle.com/yamsam/optical-flow-by-raft)) to reset tracker to reduce FP (~ 0.01 private LB improvement)\n4. **Data**\n    1. GT data: Add COTS through WBF prediction: Some real COTS is not marked in original data. We added it manually through observing the predicted video clip. In total we added 1213 COTS, around 10% of the original COTS number. This **improves CV ~0.01**.\n5. **Things didn't work**\n    1. SAHI\n    2. modify tracker to reduce obvious FP during each sequence.\n    3. Using other datasets, like [DUO](https://github.com/chongweiliu/DUO)\n    4. Trying to use best CV model and best LB model together to improve LB score but never get better than original best LB score (due to the tightness theory.....)\n    5. YOLOX, YOLOR, YOLOv5-p7, Swin Transformer with Retina, fasterRCNN and TPH YOLOV5.\n    6. Tried to add mosaic and mixup in mmdetection framework but didn't work.\n6. **Hardware / # of gpu hours**\n    1. 3 x RTX 3090 +  2 x 2080Ti + 1~2 GPU with memory 34GB + colab pro\n    2. GPU hours 850+ hours\n\nList of final ensemble candidates and their CV score in case someone is interested.\n![pic](https://i.postimg.cc/VNcWKbb0/image.png)\n\nIn the end, thank you to the great Kaggle community for many amazing sharing throughout the process, it is an exciting learning experience!🔥🔥",
    "1691391": "WoW! Great solution description. Thank you for sharing and congratulations! As I understad:\nf2_t stands for avg f2 score, P_t - precision, R_t recall?",
    "1691413": "Thank you Remek, I learnt a lot from your sharing, thank you!\nYes! `_t` stands for tracking, so those score are with tracking added :D",
    "1691457": "Ohhh ... now I understand. I am analyzing your solution to achieve better result in future :) Now it is time to learn from top guys :)",
    "1691623": "My congratulations, @vincentwang25! I learn a lot from your solution!",
    "1691624": "> 1 stage model: training with GT data for 20 ~ 40 epochs, batch size 4, training image size 2400~3600 (depending on model size).\n\nThe batch size is 4, it is quite small, did you try Gradient Accumulation?",
    "1691647": "YoloV5 itself uses nominal batch size 64 and grad accumulation accordingly",
    "1692576": "Congrates on the gold for your team! @vincentwang25 @imeintanis Congrats on becoming Competition Master!",
    "1693098": "Thank you Richard 😆",
    "1720212": "Late congrats! Could you please share a clearer img of the ensemble candidates? I am quite interested in evaluating how much ensemble can help. Thanks a lot!"
  },
  "source": "meta"
}