{
  "id": 307825,
  "title": "17th position on public LB to 45th in private [Beginners approach and mistakes we did]",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/307825",
  "author_name": "",
  "post_date": "2022-02-15T19:34:42.501188600Z",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I would like to thank organizers for such an intriguing competition.<br>\nBeing a beginner with object detection we got to understand lot of new concepts and lessons learned from our mistakes. I will share those in this post.<br>\nThanks to my teammates <a href=\"https://www.kaggle.com/tiquasar\" target=\"_blank\">@tiquasar</a> and <a href=\"https://www.kaggle.com/anshulkhadse\" target=\"_blank\">@anshulkhadse</a> <br>\n<strong>GPU: Tesla T4 16 GB and Tesla V100 64GB</strong></p>\n<h3>Training</h3>\n<h5>For training we followed 2 strategies</h5>\n<ol>\n<li>Using 0.1 and 0.2 subsequence split for cross-validation </li>\n<li>Using video-wise split strategy to train baseline and then using split 0.1 or 0.2 for finetuning it. </li>\n</ol>\n<h5>Parameters</h5>\n<ol>\n<li><code>Adam</code> for first strategy and for second one <code>Adam</code> for baseline and <code>SGD</code> for finetuning</li>\n<li>Learning rate of 0.01 for <code>SGD</code> and 0.0025 for <code>Adam</code>.</li>\n<li><code>scale: 0.5</code>, <code>shear:0.2-0.4</code>, <code>degree: 0-0.3</code>, <code>translate: 0.1-0.2</code>, <code>mosaic:1</code>, <code>mixup:0.2-0.5</code>, <code>fliplr:0.5</code>, <code>flipud: 0-0.005</code>, <code>perspective: 0-0.0002</code>, (<code>copy_paste</code>: we made a mistake and didn't use it for most of the competition)</li>\n</ol>\n<h5>Data and Augmentations</h5>\n<ol>\n<li>Augmentations during training on original data with bboxes .</li>\n<li>Augmentations on customised data along with original, generated by randomly selecting images with bbox.</li>\n</ol>\n<h3>Models</h3>\n<h5>YOLOv5</h5>\n<ol>\n<li>m and m6- <code>imgsz: 2560-4000</code>; <code>batch-size: 4</code></li>\n<li>l6- <code>imgsz: 3200</code>; <code>batch-size: 4</code></li>\n<li>s6- <code>imgsz: 2560-5200</code>; <code>​batch-size: 2, 4, 8</code></li>\n</ol>\n<h5>YOLOR</h5>\n<ol>\n<li>p6 and w6- <code>imgsz: 3200</code>; <code>batch-size: 4</code></li>\n</ol>\n<h3>Analyses</h3>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>imgsz(train)</th>\n<th>imgsz(infer)</th>\n<th>strategy</th>\n<th>Data</th>\n<th>TTA</th>\n<th>PrivateLB</th>\n<th>PublicLB</th>\n<th>Final</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>m</td>\n<td>2560</td>\n<td>3200</td>\n<td>1</td>\n<td>original</td>\n<td>default</td>\n<td>0.651</td>\n<td>0.571</td>\n<td></td>\n</tr>\n<tr>\n<td>m6</td>\n<td>2560</td>\n<td>3200</td>\n<td>1</td>\n<td>original</td>\n<td>default</td>\n<td>0.677</td>\n<td>0.552</td>\n<td></td>\n</tr>\n<tr>\n<td>m6</td>\n<td>4000</td>\n<td>6400</td>\n<td>2</td>\n<td>custom</td>\n<td>default</td>\n<td>0.672</td>\n<td>0.613</td>\n<td></td>\n</tr>\n<tr>\n<td>m6</td>\n<td>4000</td>\n<td>8000</td>\n<td>2</td>\n<td>custom</td>\n<td>s=[1,0.9,0.8]</td>\n<td>0.663</td>\n<td>0.650</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>l6</td>\n<td>3200</td>\n<td>6400</td>\n<td>1</td>\n<td>custom</td>\n<td>s=[1,0.85,0.72]</td>\n<td>0.651</td>\n<td>0.647</td>\n<td></td>\n</tr>\n<tr>\n<td>s6</td>\n<td>3584</td>\n<td>6000</td>\n<td>2</td>\n<td>original</td>\n<td>s=[1,0.7,0.6]</td>\n<td>0.626</td>\n<td>0.746</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>s6</td>\n<td>3584</td>\n<td>3600</td>\n<td>2</td>\n<td>original</td>\n<td>s=[1,0.8,0.6]</td>\n<td>0.608</td>\n<td>0.704</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>w6</td>\n<td>3200</td>\n<td>6400</td>\n<td>1</td>\n<td>custom</td>\n<td>s=[1,0.85,0.72]</td>\n<td>0.634</td>\n<td>0.626</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>For evaluation we used <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302241#1659077\" target=\"_blank\">F2-score</a> as metric the training data and flipped it horizontally and used some augmentations like CLAHE, equlize, gaussian noise. Code for this <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo\" target=\"_blank\">here</a> with name <code>evaluation_data_gen.py</code>. Bigger model were giving less FN and FP and but were not good on public lb with high resolution. Smaller model (s6) was performing good on public lb after finetuning but was giving more FN and FP then m6 &amp; l6.</p>\n<h3>Ensembles</h3>\n<p>Our best ensemble scored 0.701 on private but didn't select it :(<br>\n<a href=\"https://www.kaggle.com/tiquasar\" target=\"_blank\">@tiquasar</a> will elaborate on this</p>\n<h4>Generating Custom Data</h4>\n<p>We selected around 500 images from all three videos with bboxes count ranging from 1-18 such that each bbox count is there in selection with fair proportion as in the original dataset. Then used albumentations to flip them horizontally with CLAHE, equilize and gaussian noise. Code for this available <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo\" target=\"_blank\">here</a> with the name <code>augment_data_generate.py</code>. This extra data helped bigger models in generalizing and finetuning specially.</p>\n<h3>Mistakes we did</h3>\n<ul>\n<li>Didn't try training and inference with lower resolution(Bigger GPU size distracted us)</li>\n<li>Didn't trust CV and focused too much on public lb(Got emotionally attached to public lb rank)</li>\n<li>Lack of experiments with ensembles and different models like centernet, RCNN, efficientdet</li>\n</ul>\n<hr>\n<p>Becuase <a href=\"https://github.com/WongKinYiu/yolor/tree/paper\" target=\"_blank\">yolor</a> codebase was similar to <a href=\"https://github.com/ultralytics/yolov5\" target=\"_blank\">yolov5</a> we merged the extra features of yolor into yolov5 codebase and used it train both yolov5 and yolor.<br>\nWeights and code used in this competition <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": "1692068",
      "postDate": "02/15/2022 19:34:42",
      "content": "<p>I would like to thank organizers for such an intriguing competition.<br>\nBeing a beginner with object detection we got to understand lot of new concepts and lessons learned from our mistakes. I will share those in this post.<br>\nThanks to my teammates <a href=\"https://www.kaggle.com/tiquasar\" target=\"_blank\">@tiquasar</a> and <a href=\"https://www.kaggle.com/anshulkhadse\" target=\"_blank\">@anshulkhadse</a> <br>\n<strong>GPU: Tesla T4 16 GB and Tesla V100 64GB</strong></p>\n<h3>Training</h3>\n<h5>For training we followed 2 strategies</h5>\n<ol>\n<li>Using 0.1 and 0.2 subsequence split for cross-validation </li>\n<li>Using video-wise split strategy to train baseline and then using split 0.1 or 0.2 for finetuning it. </li>\n</ol>\n<h5>Parameters</h5>\n<ol>\n<li><code>Adam</code> for first strategy and for second one <code>Adam</code> for baseline and <code>SGD</code> for finetuning</li>\n<li>Learning rate of 0.01 for <code>SGD</code> and 0.0025 for <code>Adam</code>.</li>\n<li><code>scale: 0.5</code>, <code>shear:0.2-0.4</code>, <code>degree: 0-0.3</code>, <code>translate: 0.1-0.2</code>, <code>mosaic:1</code>, <code>mixup:0.2-0.5</code>, <code>fliplr:0.5</code>, <code>flipud: 0-0.005</code>, <code>perspective: 0-0.0002</code>, (<code>copy_paste</code>: we made a mistake and didn't use it for most of the competition)</li>\n</ol>\n<h5>Data and Augmentations</h5>\n<ol>\n<li>Augmentations during training on original data with bboxes .</li>\n<li>Augmentations on customised data along with original, generated by randomly selecting images with bbox.</li>\n</ol>\n<h3>Models</h3>\n<h5>YOLOv5</h5>\n<ol>\n<li>m and m6- <code>imgsz: 2560-4000</code>; <code>batch-size: 4</code></li>\n<li>l6- <code>imgsz: 3200</code>; <code>batch-size: 4</code></li>\n<li>s6- <code>imgsz: 2560-5200</code>; <code>​batch-size: 2, 4, 8</code></li>\n</ol>\n<h5>YOLOR</h5>\n<ol>\n<li>p6 and w6- <code>imgsz: 3200</code>; <code>batch-size: 4</code></li>\n</ol>\n<h3>Analyses</h3>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>imgsz(train)</th>\n<th>imgsz(infer)</th>\n<th>strategy</th>\n<th>Data</th>\n<th>TTA</th>\n<th>PrivateLB</th>\n<th>PublicLB</th>\n<th>Final</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>m</td>\n<td>2560</td>\n<td>3200</td>\n<td>1</td>\n<td>original</td>\n<td>default</td>\n<td>0.651</td>\n<td>0.571</td>\n<td></td>\n</tr>\n<tr>\n<td>m6</td>\n<td>2560</td>\n<td>3200</td>\n<td>1</td>\n<td>original</td>\n<td>default</td>\n<td>0.677</td>\n<td>0.552</td>\n<td></td>\n</tr>\n<tr>\n<td>m6</td>\n<td>4000</td>\n<td>6400</td>\n<td>2</td>\n<td>custom</td>\n<td>default</td>\n<td>0.672</td>\n<td>0.613</td>\n<td></td>\n</tr>\n<tr>\n<td>m6</td>\n<td>4000</td>\n<td>8000</td>\n<td>2</td>\n<td>custom</td>\n<td>s=[1,0.9,0.8]</td>\n<td>0.663</td>\n<td>0.650</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>l6</td>\n<td>3200</td>\n<td>6400</td>\n<td>1</td>\n<td>custom</td>\n<td>s=[1,0.85,0.72]</td>\n<td>0.651</td>\n<td>0.647</td>\n<td></td>\n</tr>\n<tr>\n<td>s6</td>\n<td>3584</td>\n<td>6000</td>\n<td>2</td>\n<td>original</td>\n<td>s=[1,0.7,0.6]</td>\n<td>0.626</td>\n<td>0.746</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>s6</td>\n<td>3584</td>\n<td>3600</td>\n<td>2</td>\n<td>original</td>\n<td>s=[1,0.8,0.6]</td>\n<td>0.608</td>\n<td>0.704</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>w6</td>\n<td>3200</td>\n<td>6400</td>\n<td>1</td>\n<td>custom</td>\n<td>s=[1,0.85,0.72]</td>\n<td>0.634</td>\n<td>0.626</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>For evaluation we used <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302241#1659077\" target=\"_blank\">F2-score</a> as metric the training data and flipped it horizontally and used some augmentations like CLAHE, equlize, gaussian noise. Code for this <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo\" target=\"_blank\">here</a> with name <code>evaluation_data_gen.py</code>. Bigger model were giving less FN and FP and but were not good on public lb with high resolution. Smaller model (s6) was performing good on public lb after finetuning but was giving more FN and FP then m6 &amp; l6.</p>\n<h3>Ensembles</h3>\n<p>Our best ensemble scored 0.701 on private but didn't select it :(<br>\n<a href=\"https://www.kaggle.com/tiquasar\" target=\"_blank\">@tiquasar</a> will elaborate on this</p>\n<h4>Generating Custom Data</h4>\n<p>We selected around 500 images from all three videos with bboxes count ranging from 1-18 such that each bbox count is there in selection with fair proportion as in the original dataset. Then used albumentations to flip them horizontally with CLAHE, equilize and gaussian noise. Code for this available <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo\" target=\"_blank\">here</a> with the name <code>augment_data_generate.py</code>. This extra data helped bigger models in generalizing and finetuning specially.</p>\n<h3>Mistakes we did</h3>\n<ul>\n<li>Didn't try training and inference with lower resolution(Bigger GPU size distracted us)</li>\n<li>Didn't trust CV and focused too much on public lb(Got emotionally attached to public lb rank)</li>\n<li>Lack of experiments with ensembles and different models like centernet, RCNN, efficientdet</li>\n</ul>\n<hr>\n<p>Becuase <a href=\"https://github.com/WongKinYiu/yolor/tree/paper\" target=\"_blank\">yolor</a> codebase was similar to <a href=\"https://github.com/ultralytics/yolov5\" target=\"_blank\">yolov5</a> we merged the extra features of yolor into yolov5 codebase and used it train both yolov5 and yolor.<br>\nWeights and code used in this competition <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I would like to thank organizers for such an intriguing competition.\nBeing a beginner with object detection we got to understand lot of new concepts and lessons learned from our mistakes. I will share those in this post.\nThanks to my teammates @tiquasar and @anshulkhadse \n**GPU: Tesla T4 16 GB and Tesla V100 64GB**\n\n### Training\n##### For training we followed 2 strategies\n1. Using 0.1 and 0.2 subsequence split for cross-validation \n2. Using video-wise split strategy to train baseline and then using split 0.1 or 0.2 for finetuning it. \n\n##### Parameters\n1. `Adam` for first strategy and for second one `Adam` for baseline and `SGD` for finetuning\n2. Learning rate of 0.01 for `SGD` and 0.0025 for `Adam`.\n3. `scale: 0.5`, `shear:0.2-0.4`, `degree: 0-0.3`, `translate: 0.1-0.2`, `mosaic:1`, `mixup:0.2-0.5`, `fliplr:0.5`, `flipud: 0-0.005`, `perspective: 0-0.0002`, (`copy_paste`: we made a mistake and didn't use it for most of the competition)\n\n##### Data and Augmentations\n1. Augmentations during training on original data with bboxes .\n2. Augmentations on customised data along with original, generated by randomly selecting images with bbox.\n\n### Models\n##### YOLOv5\n1. m and m6- `imgsz: 2560-4000`; `batch-size: 4`\n2. l6- `imgsz: 3200`; `batch-size: 4`\n3. s6- `imgsz: 2560-5200`; `​batch-size: 2, 4, 8`\n\n##### YOLOR\n1. p6 and w6- `imgsz: 3200`; `batch-size: 4`\n\n### Analyses\n| Model | imgsz(train) | imgsz(infer) | strategy | Data | TTA | PrivateLB | PublicLB | Final |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n| m | 2560 | 3200 | 1 | original | default | 0.651 | 0.571 | |\n| m6 | 2560 | 3200 | 1 | original |default | 0.677 | 0.552 | |\n| m6 | 4000 | 6400 | 2 | custom |default | 0.672 | 0.613 | |\n| m6 | 4000 | 8000 | 2 | custom |s=[1,0.9,0.8] | 0.663 | 0.650 | ✔️ |\n| l6 | 3200 | 6400 | 1 | custom |s=[1,0.85,0.72] | 0.651 | 0.647 | |\n| s6 | 3584 | 6000 | 2 | original |s=[1,0.7,0.6] | 0.626 | 0.746 |✔️ |\n| s6 | 3584 | 3600 | 2 | original |s=[1,0.8,0.6] | 0.608 | 0.704 | ✔️ |\n| w6 | 3200 | 6400 | 1 | custom |s=[1,0.85,0.72] | 0.634 | 0.626 | |\n\nFor evaluation we used [F2-score](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302241#1659077) as metric the training data and flipped it horizontally and used some augmentations like CLAHE, equlize, gaussian noise. Code for this [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo) with name `evaluation_data_gen.py`. Bigger model were giving less FN and FP and but were not good on public lb with high resolution. Smaller model (s6) was performing good on public lb after finetuning but was giving more FN and FP then m6 & l6.\n\n### Ensembles\nOur best ensemble scored 0.701 on private but didn't select it :(\n@tiquasar will elaborate on this\n\n#### Generating Custom Data\nWe selected around 500 images from all three videos with bboxes count ranging from 1-18 such that each bbox count is there in selection with fair proportion as in the original dataset. Then used albumentations to flip them horizontally with CLAHE, equilize and gaussian noise. Code for this available [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo) with the name `augment_data_generate.py`. This extra data helped bigger models in generalizing and finetuning specially.\n\n### Mistakes we did\n- Didn't try training and inference with lower resolution(Bigger GPU size distracted us)\n- Didn't trust CV and focused too much on public lb(Got emotionally attached to public lb rank)\n- Lack of experiments with ensembles and different models like centernet, RCNN, efficientdet\n\n------------------------------------------------------------------------------------\nBecuase [yolor](https://github.com/WongKinYiu/yolor/tree/paper) codebase was similar to [yolov5](https://github.com/ultralytics/yolov5) we merged the extra features of yolor into yolov5 codebase and used it train both yolov5 and yolor.\nWeights and code used in this competition [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition).",
      "votes": null
    },
    {
      "id": "1692470",
      "postDate": "02/16/2022 04:03:04",
      "content": "<p>Congratulations! And thank you for sharing!</p>\n<p>Especially, \"Mistakes we did\" is interesting.<br>\nIt will be helpful for beginners (like me).</p>",
      "rawMarkdown": "Congratulations! And thank you for sharing!\n\nEspecially, \"Mistakes we did\" is interesting.\nIt will be helpful for beginners (like me).",
      "votes": null
    },
    {
      "id": "1692512",
      "postDate": "02/16/2022 04:46:27",
      "content": "<p>Thanks. Mistakes must be identified :)</p>",
      "rawMarkdown": "Thanks. Mistakes must be identified :)",
      "votes": null
    },
    {
      "id": "1694653",
      "postDate": "02/17/2022 16:01:49",
      "content": "<h2>Thread Update,</h2>\n<blockquote>\n  <p>I would like to thank organizers for such an intriguing competition.<br>\n  Being a beginner with object detection we got to understand lot of new concepts and lessons learned from our mistakes. I will share those in this post.<br>\n  Thanks to my teammates <a href=\"https://www.kaggle.com/tiquasar\" target=\"_blank\">@tiquasar</a> and <a href=\"https://www.kaggle.com/anshulkhadse\" target=\"_blank\">@anshulkhadse</a> <br>\n  <strong>GPU: Tesla T4 16 GB and Tesla V100 64GB</strong></p>\n  <h3>Training</h3>\n  <h5>For training we followed 2 strategies</h5>\n  <ol>\n  <li>Using 0.1 and 0.2 subsequence split for cross-validation </li>\n  <li>Using video-wise split strategy to train baseline and then using split 0.1 or 0.2 for finetuning it. </li>\n  </ol>\n  <h5>Parameters</h5>\n  <ol>\n  <li><code>Adam</code> for first strategy and for second one <code>Adam</code> for baseline and <code>SGD</code> for finetuning</li>\n  <li>Learning rate of 0.01 for <code>SGD</code> and 0.0025 for <code>Adam</code>.</li>\n  <li><code>scale: 0.5</code>, <code>shear:0.2-0.4</code>, <code>degree: 0-0.3</code>, <code>translate: 0.1-0.2</code>, <code>mosaic:1</code>, <code>mixup:0.2-0.5</code>, <code>fliplr:0.5</code>, <code>flipud: 0-0.005</code>, <code>perspective: 0-0.0002</code>, (<code>copy_paste</code>: we made a mistake and didn't use it for most of the competition)</li>\n  </ol>\n  <h5>Data and Augmentations</h5>\n  <ol>\n  <li>Augmentations during training on original data with bboxes .</li>\n  <li>Augmentations on customised data along with original, generated by randomly selecting images with bbox.</li>\n  </ol>\n  <h3>Models</h3>\n  <h5>YOLOv5</h5>\n  <ol>\n  <li>m and m6- <code>imgsz: 2560-4000</code>; <code>batch-size: 4</code></li>\n  <li>l6- <code>imgsz: 3200</code>; <code>batch-size: 4</code></li>\n  <li>s6- <code>imgsz: 2560-5200</code>; <code>​batch-size: 2, 4, 8</code></li>\n  </ol>\n  <h5>YOLOR</h5>\n  <ol>\n  <li>p6 and w6- <code>imgsz: 3200</code>; <code>batch-size: 4</code></li>\n  </ol>\n  <h3>Analyses</h3>\n  <table>\n  <thead>\n  <tr>\n  <th>Model</th>\n  <th>imgsz(train)</th>\n  <th>imgsz(infer)</th>\n  <th>strategy</th>\n  <th>Data</th>\n  <th>TTA</th>\n  <th>PrivateLB</th>\n  <th>PublicLB</th>\n  <th>Final</th>\n  </tr>\n  </thead>\n  <tbody>\n  <tr>\n  <td>m</td>\n  <td>2560</td>\n  <td>3200</td>\n  <td>1</td>\n  <td>original</td>\n  <td>default</td>\n  <td>0.651</td>\n  <td>0.571</td>\n  <td></td>\n  </tr>\n  <tr>\n  <td>m6</td>\n  <td>2560</td>\n  <td>3200</td>\n  <td>1</td>\n  <td>original</td>\n  <td>default</td>\n  <td>0.677</td>\n  <td>0.552</td>\n  <td></td>\n  </tr>\n  <tr>\n  <td>m6</td>\n  <td>4000</td>\n  <td>6400</td>\n  <td>2</td>\n  <td>custom</td>\n  <td>default</td>\n  <td>0.672</td>\n  <td>0.613</td>\n  <td></td>\n  </tr>\n  <tr>\n  <td>m6</td>\n  <td>4000</td>\n  <td>8000</td>\n  <td>2</td>\n  <td>custom</td>\n  <td>s=[1,0.9,0.8]</td>\n  <td>0.663</td>\n  <td>0.650</td>\n  <td>✔️</td>\n  </tr>\n  <tr>\n  <td>l6</td>\n  <td>3200</td>\n  <td>6400</td>\n  <td>1</td>\n  <td>custom</td>\n  <td>s=[1,0.85,0.72]</td>\n  <td>0.651</td>\n  <td>0.647</td>\n  <td></td>\n  </tr>\n  <tr>\n  <td>s6</td>\n  <td>3584</td>\n  <td>6000</td>\n  <td>2</td>\n  <td>original</td>\n  <td>s=[1,0.7,0.6]</td>\n  <td>0.626</td>\n  <td>0.746</td>\n  <td>✔️</td>\n  </tr>\n  <tr>\n  <td>s6</td>\n  <td>3584</td>\n  <td>3600</td>\n  <td>2</td>\n  <td>original</td>\n  <td>s=[1,0.8,0.6]</td>\n  <td>0.608</td>\n  <td>0.704</td>\n  <td>✔️</td>\n  </tr>\n  <tr>\n  <td>w6</td>\n  <td>3200</td>\n  <td>6400</td>\n  <td>1</td>\n  <td>custom</td>\n  <td>s=[1,0.85,0.72]</td>\n  <td>0.634</td>\n  <td>0.626</td>\n  <td></td>\n  </tr>\n  </tbody>\n  </table>\n  <p>For evaluation we used <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302241#1659077\" target=\"_blank\">F2-score</a> as metric the training data and flipped it horizontally and used some augmentations like CLAHE, equlize, gaussian noise. Code for this <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo\" target=\"_blank\">here</a> with name <code>evaluation_data_gen.py</code>. Bigger model were giving less FN and FP and but were not good on public lb with high resolution. Smaller model (s6) was performing good on public lb after finetuning but was giving more FN and FP then m6 &amp; l6.</p>\n  <h3>Ensembles</h3>\n  <p>In this competition we tried ensemble with three yolov5 series models with individual TTA for each models as well as default TTA for all models and inference image size scaled between 1xtrain_image_size to 1.8xtrain_image_size.<br>\n  The Notebook that achieved 0.701 on private leaderboard was inferred on 3600 image size with default TTA.</p>\n  <h4>Generating Custom Data</h4>\n  <p>We selected around 500 images from all three videos with bboxes count ranging from 1-18 such that each bbox count is there in selection with fair proportion as in the original dataset. Then used albumentations to flip them horizontally with CLAHE, equilize and gaussian noise. Code for this available <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo\" target=\"_blank\">here</a> with the name <code>augment_data_generate.py</code>. This extra data helped bigger models in generalizing and finetuning specially.</p>\n  <h3>Mistakes we did</h3>\n  <ul>\n  <li>Didn't try training and inference with lower resolution(Bigger GPU size distracted us)</li>\n  <li>Didn't trust CV and focused too much on public lb(Got emotionally attached to public lb rank)</li>\n  <li>Lack of experiments with ensembles and different models like centernet, RCNN, efficientdet</li>\n  </ul>\n  <hr>\n  <p>Becuase <a href=\"https://github.com/WongKinYiu/yolor/tree/paper\" target=\"_blank\">yolor</a> codebase was similar to <a href=\"https://github.com/ultralytics/yolov5\" target=\"_blank\">yolov5</a> we merged the extra features of yolor into yolov5 codebase and used it train both yolov5 and yolor.<br>\n  Weights and code used in this competition <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition\" target=\"_blank\">here</a>.</p>\n</blockquote>\n<h4>And some additional points observed from this competition :</h4>\n<blockquote>\n  <ul>\n  <li>Adding score threshold before the boxes are averaged during ensemble helped in score increment as the non detected boxes from individual models were bringing down the average detection probability of the final boxes. </li>\n  <li>Experimented with adding ASFF(Adaptively spatial feature fusion) with yolov5, although I didn't get enough time on experimenting more with this however in initial runs it was giving a more than decent precision but poor recall.</li>\n  <li>Adding tracking with optimal parameters gave a stable boost of 0.01-0.05 on Public/Private Leaderboards.</li>\n  <li>In my observation Slicing Aided Hyper Inference (SAHI) was one algorithm that gave a huge difference between private and public leaderboards( 0.431 public LB to 0.591 private LB in one of the notebooks ), definitely will be researching and working on this algorithm later.</li>\n  </ul>\n</blockquote>",
      "rawMarkdown": "## Thread Update,\n> I would like to thank organizers for such an intriguing competition.\n> Being a beginner with object detection we got to understand lot of new concepts and lessons learned from our mistakes. I will share those in this post.\n> Thanks to my teammates @tiquasar and @anshulkhadse \n> **GPU: Tesla T4 16 GB and Tesla V100 64GB**\n> \n> ### Training\n> ##### For training we followed 2 strategies\n> 1. Using 0.1 and 0.2 subsequence split for cross-validation \n> 2. Using video-wise split strategy to train baseline and then using split 0.1 or 0.2 for finetuning it. \n> \n> ##### Parameters\n> 1. `Adam` for first strategy and for second one `Adam` for baseline and `SGD` for finetuning\n> 2. Learning rate of 0.01 for `SGD` and 0.0025 for `Adam`.\n> 3. `scale: 0.5`, `shear:0.2-0.4`, `degree: 0-0.3`, `translate: 0.1-0.2`, `mosaic:1`, `mixup:0.2-0.5`, `fliplr:0.5`, `flipud: 0-0.005`, `perspective: 0-0.0002`, (`copy_paste`: we made a mistake and didn't use it for most of the competition)\n> \n> ##### Data and Augmentations\n> 1. Augmentations during training on original data with bboxes .\n> 2. Augmentations on customised data along with original, generated by randomly selecting images with bbox.\n> \n> ### Models\n> ##### YOLOv5\n> 1. m and m6- `imgsz: 2560-4000`; `batch-size: 4`\n> 2. l6- `imgsz: 3200`; `batch-size: 4`\n> 3. s6- `imgsz: 2560-5200`; `​batch-size: 2, 4, 8`\n> \n> ##### YOLOR\n> 1. p6 and w6- `imgsz: 3200`; `batch-size: 4`\n> \n> ### Analyses\n> | Model | imgsz(train) | imgsz(infer) | strategy | Data | TTA | PrivateLB | PublicLB | Final |\n> | --- | --- | --- | --- | --- | --- | --- | --- |\n> | m | 2560 | 3200 | 1 | original | default | 0.651 | 0.571 | |\n> | m6 | 2560 | 3200 | 1 | original |default | 0.677 | 0.552 | |\n> | m6 | 4000 | 6400 | 2 | custom |default | 0.672 | 0.613 | |\n> | m6 | 4000 | 8000 | 2 | custom |s=[1,0.9,0.8] | 0.663 | 0.650 | ✔️ |\n> | l6 | 3200 | 6400 | 1 | custom |s=[1,0.85,0.72] | 0.651 | 0.647 | |\n> | s6 | 3584 | 6000 | 2 | original |s=[1,0.7,0.6] | 0.626 | 0.746 |✔️ |\n> | s6 | 3584 | 3600 | 2 | original |s=[1,0.8,0.6] | 0.608 | 0.704 | ✔️ |\n> | w6 | 3200 | 6400 | 1 | custom |s=[1,0.85,0.72] | 0.634 | 0.626 | |\n> \n> For evaluation we used [F2-score](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302241#1659077) as metric the training data and flipped it horizontally and used some augmentations like CLAHE, equlize, gaussian noise. Code for this [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo) with name `evaluation_data_gen.py`. Bigger model were giving less FN and FP and but were not good on public lb with high resolution. Smaller model (s6) was performing good on public lb after finetuning but was giving more FN and FP then m6 & l6.\n> \n> ### Ensembles\n> In this competition we tried ensemble with three yolov5 series models with individual TTA for each models as well as default TTA for all models and inference image size scaled between 1xtrain_image_size to 1.8xtrain_image_size.\n> The Notebook that achieved 0.701 on private leaderboard was inferred on 3600 image size with default TTA.\n>\n> #### Generating Custom Data\n> We selected around 500 images from all three videos with bboxes count ranging from 1-18 such that each bbox count is there in selection with fair proportion as in the original dataset. Then used albumentations to flip them horizontally with CLAHE, equilize and gaussian noise. Code for this available [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo) with the name `augment_data_generate.py`. This extra data helped bigger models in generalizing and finetuning specially.\n> \n> ### Mistakes we did\n> - Didn't try training and inference with lower resolution(Bigger GPU size distracted us)\n> - Didn't trust CV and focused too much on public lb(Got emotionally attached to public lb rank)\n> - Lack of experiments with ensembles and different models like centernet, RCNN, efficientdet\n> \n> ------------------------------------------------------------------------------------\n> Becuase [yolor](https://github.com/WongKinYiu/yolor/tree/paper) codebase was similar to [yolov5](https://github.com/ultralytics/yolov5) we merged the extra features of yolor into yolov5 codebase and used it train both yolov5 and yolor.\n> Weights and code used in this competition [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition).\n\n#### And some additional points observed from this competition :\n> - Adding score threshold before the boxes are averaged during ensemble helped in score increment as the non detected boxes from individual models were bringing down the average detection probability of the final boxes. \n> - Experimented with adding ASFF(Adaptively spatial feature fusion) with yolov5, although I didn't get enough time on experimenting more with this however in initial runs it was giving a more than decent precision but poor recall.\n> - Adding tracking with optimal parameters gave a stable boost of 0.01-0.05 on Public/Private Leaderboards.\n> - In my observation Slicing Aided Hyper Inference (SAHI) was one algorithm that gave a huge difference between private and public leaderboards( 0.431 public LB to 0.591 private LB in one of the notebooks ), definitely will be researching and working on this algorithm later.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1692470,
      "author_name": "takasukohei",
      "author_url": "",
      "post_date": "02/16/2022 04:03:04",
      "content": "<p>Congratulations! And thank you for sharing!</p>\n<p>Especially, \"Mistakes we did\" is interesting.<br>\nIt will be helpful for beginners (like me).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1692512,
          "author_name": "sanchitvj",
          "author_url": "",
          "post_date": "02/16/2022 04:46:27",
          "content": "<p>Thanks. Mistakes must be identified :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1694653,
      "author_name": "tiquasar",
      "author_url": "",
      "post_date": "02/17/2022 16:01:49",
      "content": "<h2>Thread Update,</h2>\n<blockquote>\n  <p>I would like to thank organizers for such an intriguing competition.<br>\n  Being a beginner with object detection we got to understand lot of new concepts and lessons learned from our mistakes. I will share those in this post.<br>\n  Thanks to my teammates <a href=\"https://www.kaggle.com/tiquasar\" target=\"_blank\">@tiquasar</a> and <a href=\"https://www.kaggle.com/anshulkhadse\" target=\"_blank\">@anshulkhadse</a> <br>\n  <strong>GPU: Tesla T4 16 GB and Tesla V100 64GB</strong></p>\n  <h3>Training</h3>\n  <h5>For training we followed 2 strategies</h5>\n  <ol>\n  <li>Using 0.1 and 0.2 subsequence split for cross-validation </li>\n  <li>Using video-wise split strategy to train baseline and then using split 0.1 or 0.2 for finetuning it. </li>\n  </ol>\n  <h5>Parameters</h5>\n  <ol>\n  <li><code>Adam</code> for first strategy and for second one <code>Adam</code> for baseline and <code>SGD</code> for finetuning</li>\n  <li>Learning rate of 0.01 for <code>SGD</code> and 0.0025 for <code>Adam</code>.</li>\n  <li><code>scale: 0.5</code>, <code>shear:0.2-0.4</code>, <code>degree: 0-0.3</code>, <code>translate: 0.1-0.2</code>, <code>mosaic:1</code>, <code>mixup:0.2-0.5</code>, <code>fliplr:0.5</code>, <code>flipud: 0-0.005</code>, <code>perspective: 0-0.0002</code>, (<code>copy_paste</code>: we made a mistake and didn't use it for most of the competition)</li>\n  </ol>\n  <h5>Data and Augmentations</h5>\n  <ol>\n  <li>Augmentations during training on original data with bboxes .</li>\n  <li>Augmentations on customised data along with original, generated by randomly selecting images with bbox.</li>\n  </ol>\n  <h3>Models</h3>\n  <h5>YOLOv5</h5>\n  <ol>\n  <li>m and m6- <code>imgsz: 2560-4000</code>; <code>batch-size: 4</code></li>\n  <li>l6- <code>imgsz: 3200</code>; <code>batch-size: 4</code></li>\n  <li>s6- <code>imgsz: 2560-5200</code>; <code>​batch-size: 2, 4, 8</code></li>\n  </ol>\n  <h5>YOLOR</h5>\n  <ol>\n  <li>p6 and w6- <code>imgsz: 3200</code>; <code>batch-size: 4</code></li>\n  </ol>\n  <h3>Analyses</h3>\n  <table>\n  <thead>\n  <tr>\n  <th>Model</th>\n  <th>imgsz(train)</th>\n  <th>imgsz(infer)</th>\n  <th>strategy</th>\n  <th>Data</th>\n  <th>TTA</th>\n  <th>PrivateLB</th>\n  <th>PublicLB</th>\n  <th>Final</th>\n  </tr>\n  </thead>\n  <tbody>\n  <tr>\n  <td>m</td>\n  <td>2560</td>\n  <td>3200</td>\n  <td>1</td>\n  <td>original</td>\n  <td>default</td>\n  <td>0.651</td>\n  <td>0.571</td>\n  <td></td>\n  </tr>\n  <tr>\n  <td>m6</td>\n  <td>2560</td>\n  <td>3200</td>\n  <td>1</td>\n  <td>original</td>\n  <td>default</td>\n  <td>0.677</td>\n  <td>0.552</td>\n  <td></td>\n  </tr>\n  <tr>\n  <td>m6</td>\n  <td>4000</td>\n  <td>6400</td>\n  <td>2</td>\n  <td>custom</td>\n  <td>default</td>\n  <td>0.672</td>\n  <td>0.613</td>\n  <td></td>\n  </tr>\n  <tr>\n  <td>m6</td>\n  <td>4000</td>\n  <td>8000</td>\n  <td>2</td>\n  <td>custom</td>\n  <td>s=[1,0.9,0.8]</td>\n  <td>0.663</td>\n  <td>0.650</td>\n  <td>✔️</td>\n  </tr>\n  <tr>\n  <td>l6</td>\n  <td>3200</td>\n  <td>6400</td>\n  <td>1</td>\n  <td>custom</td>\n  <td>s=[1,0.85,0.72]</td>\n  <td>0.651</td>\n  <td>0.647</td>\n  <td></td>\n  </tr>\n  <tr>\n  <td>s6</td>\n  <td>3584</td>\n  <td>6000</td>\n  <td>2</td>\n  <td>original</td>\n  <td>s=[1,0.7,0.6]</td>\n  <td>0.626</td>\n  <td>0.746</td>\n  <td>✔️</td>\n  </tr>\n  <tr>\n  <td>s6</td>\n  <td>3584</td>\n  <td>3600</td>\n  <td>2</td>\n  <td>original</td>\n  <td>s=[1,0.8,0.6]</td>\n  <td>0.608</td>\n  <td>0.704</td>\n  <td>✔️</td>\n  </tr>\n  <tr>\n  <td>w6</td>\n  <td>3200</td>\n  <td>6400</td>\n  <td>1</td>\n  <td>custom</td>\n  <td>s=[1,0.85,0.72]</td>\n  <td>0.634</td>\n  <td>0.626</td>\n  <td></td>\n  </tr>\n  </tbody>\n  </table>\n  <p>For evaluation we used <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302241#1659077\" target=\"_blank\">F2-score</a> as metric the training data and flipped it horizontally and used some augmentations like CLAHE, equlize, gaussian noise. Code for this <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo\" target=\"_blank\">here</a> with name <code>evaluation_data_gen.py</code>. Bigger model were giving less FN and FP and but were not good on public lb with high resolution. Smaller model (s6) was performing good on public lb after finetuning but was giving more FN and FP then m6 &amp; l6.</p>\n  <h3>Ensembles</h3>\n  <p>In this competition we tried ensemble with three yolov5 series models with individual TTA for each models as well as default TTA for all models and inference image size scaled between 1xtrain_image_size to 1.8xtrain_image_size.<br>\n  The Notebook that achieved 0.701 on private leaderboard was inferred on 3600 image size with default TTA.</p>\n  <h4>Generating Custom Data</h4>\n  <p>We selected around 500 images from all three videos with bboxes count ranging from 1-18 such that each bbox count is there in selection with fair proportion as in the original dataset. Then used albumentations to flip them horizontally with CLAHE, equilize and gaussian noise. Code for this available <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo\" target=\"_blank\">here</a> with the name <code>augment_data_generate.py</code>. This extra data helped bigger models in generalizing and finetuning specially.</p>\n  <h3>Mistakes we did</h3>\n  <ul>\n  <li>Didn't try training and inference with lower resolution(Bigger GPU size distracted us)</li>\n  <li>Didn't trust CV and focused too much on public lb(Got emotionally attached to public lb rank)</li>\n  <li>Lack of experiments with ensembles and different models like centernet, RCNN, efficientdet</li>\n  </ul>\n  <hr>\n  <p>Becuase <a href=\"https://github.com/WongKinYiu/yolor/tree/paper\" target=\"_blank\">yolor</a> codebase was similar to <a href=\"https://github.com/ultralytics/yolov5\" target=\"_blank\">yolov5</a> we merged the extra features of yolor into yolov5 codebase and used it train both yolov5 and yolor.<br>\n  Weights and code used in this competition <a href=\"https://www.kaggle.com/sanchitvj/barrier-reef-competition\" target=\"_blank\">here</a>.</p>\n</blockquote>\n<h4>And some additional points observed from this competition :</h4>\n<blockquote>\n  <ul>\n  <li>Adding score threshold before the boxes are averaged during ensemble helped in score increment as the non detected boxes from individual models were bringing down the average detection probability of the final boxes. </li>\n  <li>Experimented with adding ASFF(Adaptively spatial feature fusion) with yolov5, although I didn't get enough time on experimenting more with this however in initial runs it was giving a more than decent precision but poor recall.</li>\n  <li>Adding tracking with optimal parameters gave a stable boost of 0.01-0.05 on Public/Private Leaderboards.</li>\n  <li>In my observation Slicing Aided Hyper Inference (SAHI) was one algorithm that gave a huge difference between private and public leaderboards( 0.431 public LB to 0.591 private LB in one of the notebooks ), definitely will be researching and working on this algorithm later.</li>\n  </ul>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1692068": "I would like to thank organizers for such an intriguing competition.\nBeing a beginner with object detection we got to understand lot of new concepts and lessons learned from our mistakes. I will share those in this post.\nThanks to my teammates @tiquasar and @anshulkhadse \n**GPU: Tesla T4 16 GB and Tesla V100 64GB**\n\n### Training\n##### For training we followed 2 strategies\n1. Using 0.1 and 0.2 subsequence split for cross-validation \n2. Using video-wise split strategy to train baseline and then using split 0.1 or 0.2 for finetuning it. \n\n##### Parameters\n1. `Adam` for first strategy and for second one `Adam` for baseline and `SGD` for finetuning\n2. Learning rate of 0.01 for `SGD` and 0.0025 for `Adam`.\n3. `scale: 0.5`, `shear:0.2-0.4`, `degree: 0-0.3`, `translate: 0.1-0.2`, `mosaic:1`, `mixup:0.2-0.5`, `fliplr:0.5`, `flipud: 0-0.005`, `perspective: 0-0.0002`, (`copy_paste`: we made a mistake and didn't use it for most of the competition)\n\n##### Data and Augmentations\n1. Augmentations during training on original data with bboxes .\n2. Augmentations on customised data along with original, generated by randomly selecting images with bbox.\n\n### Models\n##### YOLOv5\n1. m and m6- `imgsz: 2560-4000`; `batch-size: 4`\n2. l6- `imgsz: 3200`; `batch-size: 4`\n3. s6- `imgsz: 2560-5200`; `​batch-size: 2, 4, 8`\n\n##### YOLOR\n1. p6 and w6- `imgsz: 3200`; `batch-size: 4`\n\n### Analyses\n| Model | imgsz(train) | imgsz(infer) | strategy | Data | TTA | PrivateLB | PublicLB | Final |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n| m | 2560 | 3200 | 1 | original | default | 0.651 | 0.571 | |\n| m6 | 2560 | 3200 | 1 | original |default | 0.677 | 0.552 | |\n| m6 | 4000 | 6400 | 2 | custom |default | 0.672 | 0.613 | |\n| m6 | 4000 | 8000 | 2 | custom |s=[1,0.9,0.8] | 0.663 | 0.650 | ✔️ |\n| l6 | 3200 | 6400 | 1 | custom |s=[1,0.85,0.72] | 0.651 | 0.647 | |\n| s6 | 3584 | 6000 | 2 | original |s=[1,0.7,0.6] | 0.626 | 0.746 |✔️ |\n| s6 | 3584 | 3600 | 2 | original |s=[1,0.8,0.6] | 0.608 | 0.704 | ✔️ |\n| w6 | 3200 | 6400 | 1 | custom |s=[1,0.85,0.72] | 0.634 | 0.626 | |\n\nFor evaluation we used [F2-score](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302241#1659077) as metric the training data and flipped it horizontally and used some augmentations like CLAHE, equlize, gaussian noise. Code for this [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo) with name `evaluation_data_gen.py`. Bigger model were giving less FN and FP and but were not good on public lb with high resolution. Smaller model (s6) was performing good on public lb after finetuning but was giving more FN and FP then m6 & l6.\n\n### Ensembles\nOur best ensemble scored 0.701 on private but didn't select it :(\n@tiquasar will elaborate on this\n\n#### Generating Custom Data\nWe selected around 500 images from all three videos with bboxes count ranging from 1-18 such that each bbox count is there in selection with fair proportion as in the original dataset. Then used albumentations to flip them horizontally with CLAHE, equilize and gaussian noise. Code for this available [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo) with the name `augment_data_generate.py`. This extra data helped bigger models in generalizing and finetuning specially.\n\n### Mistakes we did\n- Didn't try training and inference with lower resolution(Bigger GPU size distracted us)\n- Didn't trust CV and focused too much on public lb(Got emotionally attached to public lb rank)\n- Lack of experiments with ensembles and different models like centernet, RCNN, efficientdet\n\n------------------------------------------------------------------------------------\nBecuase [yolor](https://github.com/WongKinYiu/yolor/tree/paper) codebase was similar to [yolov5](https://github.com/ultralytics/yolov5) we merged the extra features of yolor into yolov5 codebase and used it train both yolov5 and yolor.\nWeights and code used in this competition [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition).",
    "1692470": "Congratulations! And thank you for sharing!\n\nEspecially, \"Mistakes we did\" is interesting.\nIt will be helpful for beginners (like me).",
    "1692512": "Thanks. Mistakes must be identified :)",
    "1694653": "## Thread Update,\n> I would like to thank organizers for such an intriguing competition.\n> Being a beginner with object detection we got to understand lot of new concepts and lessons learned from our mistakes. I will share those in this post.\n> Thanks to my teammates @tiquasar and @anshulkhadse \n> **GPU: Tesla T4 16 GB and Tesla V100 64GB**\n> \n> ### Training\n> ##### For training we followed 2 strategies\n> 1. Using 0.1 and 0.2 subsequence split for cross-validation \n> 2. Using video-wise split strategy to train baseline and then using split 0.1 or 0.2 for finetuning it. \n> \n> ##### Parameters\n> 1. `Adam` for first strategy and for second one `Adam` for baseline and `SGD` for finetuning\n> 2. Learning rate of 0.01 for `SGD` and 0.0025 for `Adam`.\n> 3. `scale: 0.5`, `shear:0.2-0.4`, `degree: 0-0.3`, `translate: 0.1-0.2`, `mosaic:1`, `mixup:0.2-0.5`, `fliplr:0.5`, `flipud: 0-0.005`, `perspective: 0-0.0002`, (`copy_paste`: we made a mistake and didn't use it for most of the competition)\n> \n> ##### Data and Augmentations\n> 1. Augmentations during training on original data with bboxes .\n> 2. Augmentations on customised data along with original, generated by randomly selecting images with bbox.\n> \n> ### Models\n> ##### YOLOv5\n> 1. m and m6- `imgsz: 2560-4000`; `batch-size: 4`\n> 2. l6- `imgsz: 3200`; `batch-size: 4`\n> 3. s6- `imgsz: 2560-5200`; `​batch-size: 2, 4, 8`\n> \n> ##### YOLOR\n> 1. p6 and w6- `imgsz: 3200`; `batch-size: 4`\n> \n> ### Analyses\n> | Model | imgsz(train) | imgsz(infer) | strategy | Data | TTA | PrivateLB | PublicLB | Final |\n> | --- | --- | --- | --- | --- | --- | --- | --- |\n> | m | 2560 | 3200 | 1 | original | default | 0.651 | 0.571 | |\n> | m6 | 2560 | 3200 | 1 | original |default | 0.677 | 0.552 | |\n> | m6 | 4000 | 6400 | 2 | custom |default | 0.672 | 0.613 | |\n> | m6 | 4000 | 8000 | 2 | custom |s=[1,0.9,0.8] | 0.663 | 0.650 | ✔️ |\n> | l6 | 3200 | 6400 | 1 | custom |s=[1,0.85,0.72] | 0.651 | 0.647 | |\n> | s6 | 3584 | 6000 | 2 | original |s=[1,0.7,0.6] | 0.626 | 0.746 |✔️ |\n> | s6 | 3584 | 3600 | 2 | original |s=[1,0.8,0.6] | 0.608 | 0.704 | ✔️ |\n> | w6 | 3200 | 6400 | 1 | custom |s=[1,0.85,0.72] | 0.634 | 0.626 | |\n> \n> For evaluation we used [F2-score](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302241#1659077) as metric the training data and flipped it horizontally and used some augmentations like CLAHE, equlize, gaussian noise. Code for this [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo) with name `evaluation_data_gen.py`. Bigger model were giving less FN and FP and but were not good on public lb with high resolution. Smaller model (s6) was performing good on public lb after finetuning but was giving more FN and FP then m6 & l6.\n> \n> ### Ensembles\n> In this competition we tried ensemble with three yolov5 series models with individual TTA for each models as well as default TTA for all models and inference image size scaled between 1xtrain_image_size to 1.8xtrain_image_size.\n> The Notebook that achieved 0.701 on private leaderboard was inferred on 3600 image size with default TTA.\n>\n> #### Generating Custom Data\n> We selected around 500 images from all three videos with bboxes count ranging from 1-18 such that each bbox count is there in selection with fair proportion as in the original dataset. Then used albumentations to flip them horizontally with CLAHE, equilize and gaussian noise. Code for this available [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition?select=yolo_repo) with the name `augment_data_generate.py`. This extra data helped bigger models in generalizing and finetuning specially.\n> \n> ### Mistakes we did\n> - Didn't try training and inference with lower resolution(Bigger GPU size distracted us)\n> - Didn't trust CV and focused too much on public lb(Got emotionally attached to public lb rank)\n> - Lack of experiments with ensembles and different models like centernet, RCNN, efficientdet\n> \n> ------------------------------------------------------------------------------------\n> Becuase [yolor](https://github.com/WongKinYiu/yolor/tree/paper) codebase was similar to [yolov5](https://github.com/ultralytics/yolov5) we merged the extra features of yolor into yolov5 codebase and used it train both yolov5 and yolor.\n> Weights and code used in this competition [here](https://www.kaggle.com/sanchitvj/barrier-reef-competition).\n\n#### And some additional points observed from this competition :\n> - Adding score threshold before the boxes are averaged during ensemble helped in score increment as the non detected boxes from individual models were bringing down the average detection probability of the final boxes. \n> - Experimented with adding ASFF(Adaptively spatial feature fusion) with yolov5, although I didn't get enough time on experimenting more with this however in initial runs it was giving a more than decent precision but poor recall.\n> - Adding tracking with optimal parameters gave a stable boost of 0.01-0.05 on Public/Private Leaderboards.\n> - In my observation Slicing Aided Hyper Inference (SAHI) was one algorithm that gave a huge difference between private and public leaderboards( 0.431 public LB to 0.591 private LB in one of the notebooks ), definitely will be researching and working on this algorithm later."
  },
  "source": "meta"
}