{
  "id": 307756,
  "title": "10th place solution — 2xYOLOv5 + tracking",
  "url": "/competitions/tensorflow-great-barrier-reef/writeups/where-s-crown-of-thorns-starfish-10th-place-soluti",
  "author_name": "",
  "post_date": "2022-02-15T14:18:24.800Z",
  "votes": 55,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Congratulates all the participants and many thanks to organizers. It was very challenging and interesting competition.</p>\n<p>This post was written jointly with&nbsp;<a href=\"https://www.kaggle.com/danjafish\" target=\"_blank\">@danjafish</a> and we want to say a big thank you to our team.</p>\n<h1>Summary</h1>\n<p>Our solution is an ensemble of two differently trained yolov5 models + tracking.<br>\n<img src=\"https://i.imgur.com/hrGOAv1.png\" alt=\"summary\"></p>\n<p>Our final submit code is available here<br>\n<a href=\"https://www.kaggle.com/parapapapam/2xyolov5l6-tracking-lb-0-678-private-0-722\" target=\"_blank\">https://www.kaggle.com/parapapapam/2xyolov5l6-tracking-lb-0-678-private-0-722</a></p>\n<h1>Models</h1>\n<p>In our solution we used two yolov5 models trained in different ways.</p>\n<h2>YOLOv5l6 1280</h2>\n<p>(by <a href=\"https://www.kaggle.com/parapapapam\" target=\"_blank\">@parapapapam</a>)</p>\n<h3>Idea</h3>\n<p>The main trick of this competition is the use of high resolution. But when you just resize the image, you increase the noise without adding any additional information, so the reason why it works better not in the image.</p>\n<p>My hypothesis was that the reason of that effect is in achors assignment for small object. When you increase image resolution you make your objects bigger, but anchors grid step is the same and you get better allignment so model works better.</p>\n<p>Therefore I decided do not increase input resolution, but use higher resolution inside the model. For that I changed stride of second convolutional layer from 2 to 1:</p>\n<pre><code>...\n[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n...\n</code></pre>\n<p>So this model will have 2 times bigger resolution starting from 3rd convolutional layer and should behaves like 2560 model but without resizing the original image.</p>\n<h3>Validation Strategy</h3>\n<p>Split by video. Video 1 was used as a validation. For submit model was retrained on videos 0 1 with validation on video 2.</p>\n<h3>Data</h3>\n<p>I’ve used kaggle data with 6 iterations of fixing the labels. I’ve checked FP errors of models with high confidence and fix them manually using labelImg. It improves CV without any effect on LB. <br>\nI’ve trained on labeled images + 10% unlabeled data picked at random.</p>\n<h3>Augmentations</h3>\n<p>I used albumentations with the following pipeline:</p>\n<pre><code>aug_strong = [\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=1.0),\n    A.OneOf([\n        HueSaturationValue(p=0.5, hue_limit=0.05, sat_limit=0.7, val_limit=0.4),\n        A.RandomBrightnessContrast(p=0.5, brightness_limit=0.15, contrast_limit=0.15),\n        A.ToGray(p=0.1),\n    ], p=0.8),\n    A.OneOf([\n        ShiftScaleRotate(p=0.4, shift_limit=0.1, rotate_limit=3, scale_limit=(0.5, 2.0), min_bbox=27, max_bbox=None, border_mode=0, value=(114, 114, 114)),\n        A.Perspective(p=0.1, scale=0.1, pad_mode=0, pad_val=(114, 114, 114)),\n    ], p=0.8),\n    A.OneOf([\n        A.GaussianBlur(p=0.5, blur_limit=(3, 5)),\n        A.MotionBlur(p=0.5, blur_limit=(3, 5)),\n        A.MultiplicativeNoise(p=0.5, multiplier=(0.9, 1.1), per_channel=False, elementwise=True),\n    ], p=0.2),\n]\nbbox_params = A.BboxParams(format=\"yolo\", label_fields=[\"class_labels\"], min_area=16, min_visibility=0.2)\ntransform = A.Compose(aug_strong, bbox_params=bbox_params)\n</code></pre>\n<p>And from YOLOv5 mosaic=1.0 and mixup=0.5.</p>\n<p>Also I’ve used enchancement based on clahe and channel stitching for train and inference. It gave around +0.005 CV/LB.</p>\n<h3>Train params</h3>\n<p>10 epochs, 8 batch size, SGD 0.01, OneCycle, input size 1280x720.</p>\n<p>Also I’ve changed obj: 8.0 because it depends on image size and I have to multiply it by factor of 2 ^ 2 = 4 because I’ve used two times bigger resolution inside the model and also experiment shows that multiplying it by 2 also a bit increased CV.</p>\n<h3>Results</h3>\n<p>CV ~0.65 (+ tracking 0.66) / Public 0.67 (+ tracking 0.685) / Private 0.699 (+ tracking 0.718)</p>\n<h3>What didn’t work</h3>\n<p>Copy paste COTS bboxes augmentation. Style transfer. External data.</p>\n<h2>YOLOv5l6 3100</h2>\n<p>(by <a href=\"https://www.kaggle.com/danjafish\" target=\"_blank\">@danjafish</a>)</p>\n<p>My approach was quite straightforward. I got most of my ideas from <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/172569\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/172569</a> and <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/172418\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/172418</a>. However, I decided to use yolov5 for a start since I found it the most promising.</p>\n<ul>\n<li>Validation strategy: split by video. Video 1 was used as a validation. After selecting thresholds and other parameters, the model was re-trained on all data.</li>\n<li>Data: I used only kaggle data. Trained on not empty images only. Adding empty ones slightly improved scoring on validation, but significantly reduced LB. That's why I gave it up.</li>\n<li>Various augmentations from&nbsp;<a href=\"https://albumentations.ai/\" target=\"_blank\">albumentations</a> (I borrowed pipeline from <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/172569\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/172569</a>):<ul>\n<li>HorizontalFlip, ShiftScaleRotate, RandomRotate90</li>\n<li>RandomBrightnessContrast, HueSaturationValue, RGBShift</li>\n<li>RandomGamma</li>\n<li>CLAHE</li>\n<li>Blur, MotionBlur</li>\n<li>GaussNoise</li>\n<li>ImageCompression</li>\n<li>CoarseDropout</li></ul></li>\n<li>I also changed the default mosaic implementation to get images of about the same size and made up parts of images with CenterCrop.</li>\n<li>Inference size: 3100. I tried different sizes and different architectures. The best results on validation were shown by yolov5l on size 3000X3000. However it was much worse on LB than l6 3100X3100. That's why I used the later one.</li>\n<li>Train params: 20 epochs, Adam optimizer, lr 0.001 ony cycle sheduler. bs=4</li>\n<li>Hardware: most of the time I used 4 V100 GPUs with 16GB RAM. Model takes about 2 hours to train.</li>\n<li>What did’t work:<ul>\n<li>Manual reannotations of data - some COTS were not labeled because they appeared in frames later than updated annotation. (my assumption). Adding these COTS slightly improved CV, but significantly worsened LB.</li>\n<li>TTA. Slightly impoved LB and CV, but takes much longer to run.</li>\n<li>Blend of different yolo architectures</li>\n<li>Blend one stage and two stage detectors. I spent the last week completely devoted to the models on mmdetection, hoping to improve the ensemble score. However, my cascade rcnn showed much worse results on both CV and LB.</li></ul></li>\n<li>Final scores CV/Public+tracking/Private+tracking: ~0.643(video 1)/0.693/0.722</li>\n</ul>\n<h1>Ensemble</h1>\n<p>We took predicts of two above models and use WBF with iou_threshold = 0.6 to blend them with weights 0.3 and 0.7. After that we filter results with confidence threshold ≥ 0.1.</p>\n<p>But we didn’t spend enough time to tune the weights and the final result seems not much better that each model, but we have strong models, so each of them + tracking already is in gold :).</p>\n<h1>Tracking</h1>\n<p>I’ve already posted tracking approach based on norfair library<br>\n<a href=\"https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539\" target=\"_blank\">https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539</a><br>\nSo I will not go into details of norfair tracking implementation based on SORT algorithm.</p>\n<p>Here I want to mention that this problem is not a MOT problem because objects are almost static and we have to deal with only camera movement.<br>\nFor that we’ve developed algorithm of homography calculation between every two consecutive frames using ORB keypoint descriptors. Next we transform our bboxes using homography matrix to obtain predictions on the next frame.It gives much accurate predictions on two consecutive frames comparing to norfair:<br>\n<img src=\"https://i.imgur.com/vwMG9AP.png\" alt=\"frame1\"><br>\n<img src=\"https://i.imgur.com/zwZjqRu.png\" alt=\"frame2\"></p>\n<p>Green labels is a ground truth, the blue ones are predictions using just norfair tracking, red ones — homography transform from previous frame + norfair. As we can see, homography + norfair match much better than just norfair.<br>\nWhen we have no detection for object that was detected at least two frames before we predict it 2 more frames using tracking bboxes and if the object doesn’t appear remove it from tracker.<br>\nAlso we filtered objects that are moving out of the image.</p>\n<p>Our tracking approach gives stable +0.01-0.02 CV/Public/Private.</p>\n<h1>Other ideas</h1>\n<h3>Rounding</h3>\n<p>We found that using of bboxes = bboxes.astype(int) instead of bboxes = bboxes.round().astype(int) improves CV and Public ~ 0.005 and have no effect on Private. It seems the same like was already discussed<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307605\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307605</a>.</p>\n<p>Thank you for reading and happy Kaggling! :)</p>",
  "messages": [
    {
      "id": "1691631",
      "postDate": "02/15/2022 14:16:50",
      "content": "<p>Congratulates all the participants and many thanks to organizers. It was very challenging and interesting competition.</p>\n<p>This post was written jointly with&nbsp;<a href=\"https://www.kaggle.com/danjafish\" target=\"_blank\">@danjafish</a> and we want to say a big thank you to our team.</p>\n<h1>Summary</h1>\n<p>Our solution is an ensemble of two differently trained yolov5 models + tracking.<br>\n<img src=\"https://i.imgur.com/hrGOAv1.png\" alt=\"summary\"></p>\n<p>Our final submit code is available here<br>\n<a href=\"https://www.kaggle.com/parapapapam/2xyolov5l6-tracking-lb-0-678-private-0-722\" target=\"_blank\">https://www.kaggle.com/parapapapam/2xyolov5l6-tracking-lb-0-678-private-0-722</a></p>\n<h1>Models</h1>\n<p>In our solution we used two yolov5 models trained in different ways.</p>\n<h2>YOLOv5l6 1280</h2>\n<p>(by <a href=\"https://www.kaggle.com/parapapapam\" target=\"_blank\">@parapapapam</a>)</p>\n<h3>Idea</h3>\n<p>The main trick of this competition is the use of high resolution. But when you just resize the image, you increase the noise without adding any additional information, so the reason why it works better not in the image.</p>\n<p>My hypothesis was that the reason of that effect is in achors assignment for small object. When you increase image resolution you make your objects bigger, but anchors grid step is the same and you get better allignment so model works better.</p>\n<p>Therefore I decided do not increase input resolution, but use higher resolution inside the model. For that I changed stride of second convolutional layer from 2 to 1:</p>\n<pre><code>...\n[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n...\n</code></pre>\n<p>So this model will have 2 times bigger resolution starting from 3rd convolutional layer and should behaves like 2560 model but without resizing the original image.</p>\n<h3>Validation Strategy</h3>\n<p>Split by video. Video 1 was used as a validation. For submit model was retrained on videos 0 1 with validation on video 2.</p>\n<h3>Data</h3>\n<p>I’ve used kaggle data with 6 iterations of fixing the labels. I’ve checked FP errors of models with high confidence and fix them manually using labelImg. It improves CV without any effect on LB. <br>\nI’ve trained on labeled images + 10% unlabeled data picked at random.</p>\n<h3>Augmentations</h3>\n<p>I used albumentations with the following pipeline:</p>\n<pre><code>aug_strong = [\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=1.0),\n    A.OneOf([\n        HueSaturationValue(p=0.5, hue_limit=0.05, sat_limit=0.7, val_limit=0.4),\n        A.RandomBrightnessContrast(p=0.5, brightness_limit=0.15, contrast_limit=0.15),\n        A.ToGray(p=0.1),\n    ], p=0.8),\n    A.OneOf([\n        ShiftScaleRotate(p=0.4, shift_limit=0.1, rotate_limit=3, scale_limit=(0.5, 2.0), min_bbox=27, max_bbox=None, border_mode=0, value=(114, 114, 114)),\n        A.Perspective(p=0.1, scale=0.1, pad_mode=0, pad_val=(114, 114, 114)),\n    ], p=0.8),\n    A.OneOf([\n        A.GaussianBlur(p=0.5, blur_limit=(3, 5)),\n        A.MotionBlur(p=0.5, blur_limit=(3, 5)),\n        A.MultiplicativeNoise(p=0.5, multiplier=(0.9, 1.1), per_channel=False, elementwise=True),\n    ], p=0.2),\n]\nbbox_params = A.BboxParams(format=\"yolo\", label_fields=[\"class_labels\"], min_area=16, min_visibility=0.2)\ntransform = A.Compose(aug_strong, bbox_params=bbox_params)\n</code></pre>\n<p>And from YOLOv5 mosaic=1.0 and mixup=0.5.</p>\n<p>Also I’ve used enchancement based on clahe and channel stitching for train and inference. It gave around +0.005 CV/LB.</p>\n<h3>Train params</h3>\n<p>10 epochs, 8 batch size, SGD 0.01, OneCycle, input size 1280x720.</p>\n<p>Also I’ve changed obj: 8.0 because it depends on image size and I have to multiply it by factor of 2 ^ 2 = 4 because I’ve used two times bigger resolution inside the model and also experiment shows that multiplying it by 2 also a bit increased CV.</p>\n<h3>Results</h3>\n<p>CV ~0.65 (+ tracking 0.66) / Public 0.67 (+ tracking 0.685) / Private 0.699 (+ tracking 0.718)</p>\n<h3>What didn’t work</h3>\n<p>Copy paste COTS bboxes augmentation. Style transfer. External data.</p>\n<h2>YOLOv5l6 3100</h2>\n<p>(by <a href=\"https://www.kaggle.com/danjafish\" target=\"_blank\">@danjafish</a>)</p>\n<p>My approach was quite straightforward. I got most of my ideas from <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/172569\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/172569</a> and <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/172418\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/172418</a>. However, I decided to use yolov5 for a start since I found it the most promising.</p>\n<ul>\n<li>Validation strategy: split by video. Video 1 was used as a validation. After selecting thresholds and other parameters, the model was re-trained on all data.</li>\n<li>Data: I used only kaggle data. Trained on not empty images only. Adding empty ones slightly improved scoring on validation, but significantly reduced LB. That's why I gave it up.</li>\n<li>Various augmentations from&nbsp;<a href=\"https://albumentations.ai/\" target=\"_blank\">albumentations</a> (I borrowed pipeline from <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/172569\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/172569</a>):<ul>\n<li>HorizontalFlip, ShiftScaleRotate, RandomRotate90</li>\n<li>RandomBrightnessContrast, HueSaturationValue, RGBShift</li>\n<li>RandomGamma</li>\n<li>CLAHE</li>\n<li>Blur, MotionBlur</li>\n<li>GaussNoise</li>\n<li>ImageCompression</li>\n<li>CoarseDropout</li></ul></li>\n<li>I also changed the default mosaic implementation to get images of about the same size and made up parts of images with CenterCrop.</li>\n<li>Inference size: 3100. I tried different sizes and different architectures. The best results on validation were shown by yolov5l on size 3000X3000. However it was much worse on LB than l6 3100X3100. That's why I used the later one.</li>\n<li>Train params: 20 epochs, Adam optimizer, lr 0.001 ony cycle sheduler. bs=4</li>\n<li>Hardware: most of the time I used 4 V100 GPUs with 16GB RAM. Model takes about 2 hours to train.</li>\n<li>What did’t work:<ul>\n<li>Manual reannotations of data - some COTS were not labeled because they appeared in frames later than updated annotation. (my assumption). Adding these COTS slightly improved CV, but significantly worsened LB.</li>\n<li>TTA. Slightly impoved LB and CV, but takes much longer to run.</li>\n<li>Blend of different yolo architectures</li>\n<li>Blend one stage and two stage detectors. I spent the last week completely devoted to the models on mmdetection, hoping to improve the ensemble score. However, my cascade rcnn showed much worse results on both CV and LB.</li></ul></li>\n<li>Final scores CV/Public+tracking/Private+tracking: ~0.643(video 1)/0.693/0.722</li>\n</ul>\n<h1>Ensemble</h1>\n<p>We took predicts of two above models and use WBF with iou_threshold = 0.6 to blend them with weights 0.3 and 0.7. After that we filter results with confidence threshold ≥ 0.1.</p>\n<p>But we didn’t spend enough time to tune the weights and the final result seems not much better that each model, but we have strong models, so each of them + tracking already is in gold :).</p>\n<h1>Tracking</h1>\n<p>I’ve already posted tracking approach based on norfair library<br>\n<a href=\"https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539\" target=\"_blank\">https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539</a><br>\nSo I will not go into details of norfair tracking implementation based on SORT algorithm.</p>\n<p>Here I want to mention that this problem is not a MOT problem because objects are almost static and we have to deal with only camera movement.<br>\nFor that we’ve developed algorithm of homography calculation between every two consecutive frames using ORB keypoint descriptors. Next we transform our bboxes using homography matrix to obtain predictions on the next frame.It gives much accurate predictions on two consecutive frames comparing to norfair:<br>\n<img src=\"https://i.imgur.com/vwMG9AP.png\" alt=\"frame1\"><br>\n<img src=\"https://i.imgur.com/zwZjqRu.png\" alt=\"frame2\"></p>\n<p>Green labels is a ground truth, the blue ones are predictions using just norfair tracking, red ones — homography transform from previous frame + norfair. As we can see, homography + norfair match much better than just norfair.<br>\nWhen we have no detection for object that was detected at least two frames before we predict it 2 more frames using tracking bboxes and if the object doesn’t appear remove it from tracker.<br>\nAlso we filtered objects that are moving out of the image.</p>\n<p>Our tracking approach gives stable +0.01-0.02 CV/Public/Private.</p>\n<h1>Other ideas</h1>\n<h3>Rounding</h3>\n<p>We found that using of bboxes = bboxes.astype(int) instead of bboxes = bboxes.round().astype(int) improves CV and Public ~ 0.005 and have no effect on Private. It seems the same like was already discussed<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307605\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307605</a>.</p>\n<p>Thank you for reading and happy Kaggling! :)</p>",
      "rawMarkdown": "Congratulates all the participants and many thanks to organizers. It was very challenging and interesting competition.\n\nThis post was written jointly with [@danjafish](https://www.kaggle.com/danjafish) and we want to say a big thank you to our team.\n\n# Summary\nOur solution is an ensemble of two differently trained yolov5 models + tracking.\n![summary](https://i.imgur.com/hrGOAv1.png)\n\nOur final submit code is available here\n[https://www.kaggle.com/parapapapam/2xyolov5l6-tracking-lb-0-678-private-0-722](https://www.kaggle.com/parapapapam/2xyolov5l6-tracking-lb-0-678-private-0-722)\n\n# Models\nIn our solution we used two yolov5 models trained in different ways.\n\n## YOLOv5l6 1280 \n(by [@parapapapam](https://www.kaggle.com/parapapapam))\n\n### Idea\nThe main trick of this competition is the use of high resolution. But when you just resize the image, you increase the noise without adding any additional information, so the reason why it works better not in the image.\n\nMy hypothesis was that the reason of that effect is in achors assignment for small object. When you increase image resolution you make your objects bigger, but anchors grid step is the same and you get better allignment so model works better.\n\nTherefore I decided do not increase input resolution, but use higher resolution inside the model. For that I changed stride of second convolutional layer from 2 to 1:\n\n```yaml\n...\n[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n...\n```\n\nSo this model will have 2 times bigger resolution starting from 3rd convolutional layer and should behaves like 2560 model but without resizing the original image.\n\n### Validation Strategy\n\nSplit by video. Video 1 was used as a validation. For submit model was retrained on videos 0 1 with validation on video 2.\n\n### Data\n\nI’ve used kaggle data with 6 iterations of fixing the labels. I’ve checked FP errors of models with high confidence and fix them manually using labelImg. It improves CV without any effect on LB. \nI’ve trained on labeled images + 10% unlabeled data picked at random.\n\n### Augmentations\n\nI used albumentations with the following pipeline:\n\n```python\naug_strong = [\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=1.0),\n    A.OneOf([\n        HueSaturationValue(p=0.5, hue_limit=0.05, sat_limit=0.7, val_limit=0.4),\n        A.RandomBrightnessContrast(p=0.5, brightness_limit=0.15, contrast_limit=0.15),\n        A.ToGray(p=0.1),\n    ], p=0.8),\n    A.OneOf([\n        ShiftScaleRotate(p=0.4, shift_limit=0.1, rotate_limit=3, scale_limit=(0.5, 2.0), min_bbox=27, max_bbox=None, border_mode=0, value=(114, 114, 114)),\n        A.Perspective(p=0.1, scale=0.1, pad_mode=0, pad_val=(114, 114, 114)),\n    ], p=0.8),\n    A.OneOf([\n        A.GaussianBlur(p=0.5, blur_limit=(3, 5)),\n        A.MotionBlur(p=0.5, blur_limit=(3, 5)),\n        A.MultiplicativeNoise(p=0.5, multiplier=(0.9, 1.1), per_channel=False, elementwise=True),\n    ], p=0.2),\n]\nbbox_params = A.BboxParams(format=\"yolo\", label_fields=[\"class_labels\"], min_area=16, min_visibility=0.2)\ntransform = A.Compose(aug_strong, bbox_params=bbox_params)\n```\n\nAnd from YOLOv5 mosaic=1.0 and mixup=0.5.\n\nAlso I’ve used enchancement based on clahe and channel stitching for train and inference. It gave around +0.005 CV/LB.\n\n### Train params\n\n10 epochs, 8 batch size, SGD 0.01, OneCycle, input size 1280x720.\n\nAlso I’ve changed obj: 8.0 because it depends on image size and I have to multiply it by factor of 2 ^ 2 = 4 because I’ve used two times bigger resolution inside the model and also experiment shows that multiplying it by 2 also a bit increased CV.\n\n### Results\n\nCV ~0.65 (+ tracking 0.66) / Public 0.67 (+ tracking 0.685) / Private 0.699 (+ tracking 0.718)\n\n### What didn’t work\n\nCopy paste COTS bboxes augmentation. Style transfer. External data.\n\n## YOLOv5l6 3100 \n(by [@danjafish](https://www.kaggle.com/danjafish))\n\nMy approach was quite straightforward. I got most of my ideas from [https://www.kaggle.com/c/global-wheat-detection/discussion/172569](https://www.kaggle.com/c/global-wheat-detection/discussion/172569) and [https://www.kaggle.com/c/global-wheat-detection/discussion/172418](https://www.kaggle.com/c/global-wheat-detection/discussion/172418). However, I decided to use yolov5 for a start since I found it the most promising.\n\n- Validation strategy: split by video. Video 1 was used as a validation. After selecting thresholds and other parameters, the model was re-trained on all data.\n- Data: I used only kaggle data. Trained on not empty images only. Adding empty ones slightly improved scoring on validation, but significantly reduced LB. That's why I gave it up.\n- Various augmentations from [albumentations](https://albumentations.ai/) (I borrowed pipeline from [https://www.kaggle.com/c/global-wheat-detection/discussion/172569](https://www.kaggle.com/c/global-wheat-detection/discussion/172569)):\n    - HorizontalFlip, ShiftScaleRotate, RandomRotate90\n    - RandomBrightnessContrast, HueSaturationValue, RGBShift\n    - RandomGamma\n    - CLAHE\n    - Blur, MotionBlur\n    - GaussNoise\n    - ImageCompression\n    - CoarseDropout\n- I also changed the default mosaic implementation to get images of about the same size and made up parts of images with CenterCrop.\n- Inference size: 3100. I tried different sizes and different architectures. The best results on validation were shown by yolov5l on size 3000X3000. However it was much worse on LB than l6 3100X3100. That's why I used the later one.\n- Train params: 20 epochs, Adam optimizer, lr 0.001 ony cycle sheduler. bs=4\n- Hardware: most of the time I used 4 V100 GPUs with 16GB RAM. Model takes about 2 hours to train.\n- What did’t work:\n    - Manual reannotations of data - some COTS were not labeled because they appeared in frames later than updated annotation. (my assumption). Adding these COTS slightly improved CV, but significantly worsened LB.\n    - TTA. Slightly impoved LB and CV, but takes much longer to run.\n    - Blend of different yolo architectures\n    - Blend one stage and two stage detectors. I spent the last week completely devoted to the models on mmdetection, hoping to improve the ensemble score. However, my cascade rcnn showed much worse results on both CV and LB.\n- Final scores CV/Public+tracking/Private+tracking: ~0.643(video 1)/0.693/0.722\n\n# Ensemble\n\nWe took predicts of two above models and use WBF with iou_threshold = 0.6 to blend them with weights 0.3 and 0.7. After that we filter results with confidence threshold ≥ 0.1.\n\nBut we didn’t spend enough time to tune the weights and the final result seems not much better that each model, but we have strong models, so each of them + tracking already is in gold :).\n\n# Tracking\n\nI’ve already posted tracking approach based on norfair library\n[https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539](https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539)\nSo I will not go into details of norfair tracking implementation based on SORT algorithm.\n\nHere I want to mention that this problem is not a MOT problem because objects are almost static and we have to deal with only camera movement.\nFor that we’ve developed algorithm of homography calculation between every two consecutive frames using ORB keypoint descriptors. Next we transform our bboxes using homography matrix to obtain predictions on the next frame.It gives much accurate predictions on two consecutive frames comparing to norfair:\n![frame1](https://i.imgur.com/vwMG9AP.png)\n![frame2](https://i.imgur.com/zwZjqRu.png)\n\nGreen labels is a ground truth, the blue ones are predictions using just norfair tracking, red ones — homography transform from previous frame + norfair. As we can see, homography + norfair match much better than just norfair.\nWhen we have no detection for object that was detected at least two frames before we predict it 2 more frames using tracking bboxes and if the object doesn’t appear remove it from tracker.\nAlso we filtered objects that are moving out of the image.\n\nOur tracking approach gives stable +0.01-0.02 CV/Public/Private.\n\n# Other ideas\n\n### Rounding\nWe found that using of bboxes = bboxes.astype(int) instead of bboxes = bboxes.round().astype(int) improves CV and Public ~ 0.005 and have no effect on Private. It seems the same like was already discussed\n[https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307605](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307605).\n\nThank you for reading and happy Kaggling! :)",
      "votes": null
    },
    {
      "id": "1691640",
      "postDate": "02/15/2022 14:25:44",
      "content": "<blockquote>\n  <p>[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4</p>\n</blockquote>\n<p>That's awesome. Thanks for sharing this trick! I will use this in the future!</p>\n<p>Congratulations to you and your team for finishing 10th place Gold. Fantastic job! </p>\n<p>I like how simple and elegant your solution is. Using homography transform is a great idea.</p>",
      "rawMarkdown": ">[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n\nThat's awesome. Thanks for sharing this trick! I will use this in the future!\n\nCongratulations to you and your team for finishing 10th place Gold. Fantastic job! \n\nI like how simple and elegant your solution is. Using homography transform is a great idea.",
      "votes": null
    },
    {
      "id": "1691726",
      "postDate": "02/15/2022 15:15:09",
      "content": "<p>A doubt, how does changing stride, increases the resolution? and doesn't it affect the pretrained weights, since its been trained on different stride and the following layers getting different input shape</p>",
      "rawMarkdown": "A doubt, how does changing stride, increases the resolution? and doesn't it affect the pretrained weights, since its been trained on different stride and the following layers getting different input shape",
      "votes": null
    },
    {
      "id": "1691734",
      "postDate": "02/15/2022 15:19:18",
      "content": "<p><code>...\n[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n...</code></p>\n<p>coool! Outstanding tip. Congratulations!</p>",
      "rawMarkdown": "`...\n[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n...`\n\ncoool! Outstanding tip. Congratulations!",
      "votes": null
    },
    {
      "id": "1691737",
      "postDate": "02/15/2022 15:22:16",
      "content": "<p>These models (and layers) are fully convolutional. Therefore they don't care what input shape you give them.</p>\n<p>Each layer is trained to recognize a certain pattern in the image. If the previous layer doesn't do stride 2 after convolution, then the current layer just receives a 2x larger image. The current layer can still use its pretraining. It will just see an image that is 2x larger.</p>",
      "rawMarkdown": "These models (and layers) are fully convolutional. Therefore they don't care what input shape you give them.\n\nEach layer is trained to recognize a certain pattern in the image. If the previous layer doesn't do stride 2 after convolution, then the current layer just receives a 2x larger image. The current layer can still use its pretraining. It will just see an image that is 2x larger.",
      "votes": null
    },
    {
      "id": "1692202",
      "postDate": "02/15/2022 22:24:16",
      "content": "<p>Great ideas. When changing the stride, were you able to directly use pretrained models from the yolov5 repo ?</p>",
      "rawMarkdown": "Great ideas. When changing the stride, were you able to directly use pretrained models from the yolov5 repo ?",
      "votes": null
    },
    {
      "id": "1692376",
      "postDate": "02/16/2022 02:14:16",
      "content": "<p>Cong bro! How did you make sure non-overfitting of your model when retrain on the whole set?</p>",
      "rawMarkdown": "Cong bro! How did you make sure non-overfitting of your model when retrain on the whole set?",
      "votes": null
    },
    {
      "id": "1692379",
      "postDate": "02/16/2022 02:18:45",
      "content": "<p>I also tried cascade rcnn for a week and got 0.15 in LB 😂</p>",
      "rawMarkdown": "I also tried cascade rcnn for a week and got 0.15 in LB 😂",
      "votes": null
    },
    {
      "id": "1692521",
      "postDate": "02/16/2022 05:03:02",
      "content": "<p>\"When you increase image resolution you make your objects bigger, but anchors grid step is the same and you get better allignment so model works better.\" Should this be resolved by recomputing the anchor boxes with yolo k-means clustering and using custom anchors? I had tried using custom anchors with no improvement in CV. Curious if you had tried using custom anchors. Congrats on the medal.</p>",
      "rawMarkdown": "\"When you increase image resolution you make your objects bigger, but anchors grid step is the same and you get better allignment so model works better.\" Should this be resolved by recomputing the anchor boxes with yolo k-means clustering and using custom anchors? I had tried using custom anchors with no improvement in CV. Curious if you had tried using custom anchors. Congrats on the medal.",
      "votes": null
    },
    {
      "id": "1692588",
      "postDate": "02/16/2022 06:10:01",
      "content": "<p>Thanks Chris.<br>\nCongratulations to you!<br>\nAnd thank you for all knowledge you shared on kaggle. :)</p>",
      "rawMarkdown": "Thanks Chris.\nCongratulations to you!\nAnd thank you for all knowledge you shared on kaggle. :)",
      "votes": null
    },
    {
      "id": "1692594",
      "postDate": "02/16/2022 06:13:11",
      "content": "<p>Thanks Remek.<br>\nCongratulations to you and your team!<br>\nYou and your team did a big contribution in this competition.</p>",
      "rawMarkdown": "Thanks Remek.\nCongratulations to you and your team!\nYou and your team did a big contribution in this competition.",
      "votes": null
    },
    {
      "id": "1692597",
      "postDate": "02/16/2022 06:15:43",
      "content": "<p>Thanks Alexandre.<br>\nYes because the model is fully convolutional and it doesn't depend on the shape of input tensor.</p>",
      "rawMarkdown": "Thanks Alexandre.\nYes because the model is fully convolutional and it doesn't depend on the shape of input tensor.",
      "votes": null
    },
    {
      "id": "1692600",
      "postDate": "02/16/2022 06:17:46",
      "content": "<p>Thanks Sean.<br>\nYou have to keep exactly the same config. And also we observed that the validation metrics don't overfit when you train longer.</p>",
      "rawMarkdown": "Thanks Sean.\nYou have to keep exactly the same config. And also we observed that the validation metrics don't overfit when you train longer.",
      "votes": null
    },
    {
      "id": "1692604",
      "postDate": "02/16/2022 06:20:42",
      "content": "<p>Thanks Rajaram.<br>\nThe difference not only in anchor sizes, but also in a grid size.<br>\nIf you make resolution 2 times bigger inside the model you make your anchor grid step 2 times smaller and it makes assignment between objects and anchors better so this is the key.<br>\nRegarding custom anchors — yes we used custom anchors for 1280 model, but it has not so much effect comparing to the default anchors.</p>",
      "rawMarkdown": "Thanks Rajaram.\nThe difference not only in anchor sizes, but also in a grid size.\nIf you make resolution 2 times bigger inside the model you make your anchor grid step 2 times smaller and it makes assignment between objects and anchors better so this is the key.\nRegarding custom anchors — yes we used custom anchors for 1280 model, but it has not so much effect comparing to the default anchors.",
      "votes": null
    },
    {
      "id": "1692619",
      "postDate": "02/16/2022 06:34:37",
      "content": "<p>got it - thanks for clarifying</p>",
      "rawMarkdown": "got it - thanks for clarifying",
      "votes": null
    },
    {
      "id": "1692620",
      "postDate": "02/16/2022 06:37:33",
      "content": "<p>thanks for sharing.  </p>",
      "rawMarkdown": "thanks for sharing.",
      "votes": null
    },
    {
      "id": "1711215",
      "postDate": "03/03/2022 18:23:17",
      "content": "<p>👍👍👍👍👍👍👍👍👍👍</p>",
      "rawMarkdown": "👍👍👍👍👍👍👍👍👍👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1691640,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/15/2022 14:25:44",
      "content": "<blockquote>\n  <p>[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4</p>\n</blockquote>\n<p>That's awesome. Thanks for sharing this trick! I will use this in the future!</p>\n<p>Congratulations to you and your team for finishing 10th place Gold. Fantastic job! </p>\n<p>I like how simple and elegant your solution is. Using homography transform is a great idea.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1691726,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/15/2022 15:15:09",
          "content": "<p>A doubt, how does changing stride, increases the resolution? and doesn't it affect the pretrained weights, since its been trained on different stride and the following layers getting different input shape</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691737,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/15/2022 15:22:16",
          "content": "<p>These models (and layers) are fully convolutional. Therefore they don't care what input shape you give them.</p>\n<p>Each layer is trained to recognize a certain pattern in the image. If the previous layer doesn't do stride 2 after convolution, then the current layer just receives a 2x larger image. The current layer can still use its pretraining. It will just see an image that is 2x larger.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692588,
          "author_name": "parapapapam",
          "author_url": "",
          "post_date": "02/16/2022 06:10:01",
          "content": "<p>Thanks Chris.<br>\nCongratulations to you!<br>\nAnd thank you for all knowledge you shared on kaggle. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691734,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/15/2022 15:19:18",
      "content": "<p><code>...\n[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n...</code></p>\n<p>coool! Outstanding tip. Congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1692594,
          "author_name": "parapapapam",
          "author_url": "",
          "post_date": "02/16/2022 06:13:11",
          "content": "<p>Thanks Remek.<br>\nCongratulations to you and your team!<br>\nYou and your team did a big contribution in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1692202,
      "author_name": "alexandrecc",
      "author_url": "",
      "post_date": "02/15/2022 22:24:16",
      "content": "<p>Great ideas. When changing the stride, were you able to directly use pretrained models from the yolov5 repo ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1692597,
          "author_name": "parapapapam",
          "author_url": "",
          "post_date": "02/16/2022 06:15:43",
          "content": "<p>Thanks Alexandre.<br>\nYes because the model is fully convolutional and it doesn't depend on the shape of input tensor.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1692376,
      "author_name": "lixxxxx",
      "author_url": "",
      "post_date": "02/16/2022 02:14:16",
      "content": "<p>Cong bro! How did you make sure non-overfitting of your model when retrain on the whole set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1692600,
          "author_name": "parapapapam",
          "author_url": "",
          "post_date": "02/16/2022 06:17:46",
          "content": "<p>Thanks Sean.<br>\nYou have to keep exactly the same config. And also we observed that the validation metrics don't overfit when you train longer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1692379,
      "author_name": "lixxxxx",
      "author_url": "",
      "post_date": "02/16/2022 02:18:45",
      "content": "<p>I also tried cascade rcnn for a week and got 0.15 in LB 😂</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1692521,
      "author_name": "mpdroid",
      "author_url": "",
      "post_date": "02/16/2022 05:03:02",
      "content": "<p>\"When you increase image resolution you make your objects bigger, but anchors grid step is the same and you get better allignment so model works better.\" Should this be resolved by recomputing the anchor boxes with yolo k-means clustering and using custom anchors? I had tried using custom anchors with no improvement in CV. Curious if you had tried using custom anchors. Congrats on the medal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1692604,
          "author_name": "parapapapam",
          "author_url": "",
          "post_date": "02/16/2022 06:20:42",
          "content": "<p>Thanks Rajaram.<br>\nThe difference not only in anchor sizes, but also in a grid size.<br>\nIf you make resolution 2 times bigger inside the model you make your anchor grid step 2 times smaller and it makes assignment between objects and anchors better so this is the key.<br>\nRegarding custom anchors — yes we used custom anchors for 1280 model, but it has not so much effect comparing to the default anchors.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692619,
          "author_name": "mpdroid",
          "author_url": "",
          "post_date": "02/16/2022 06:34:37",
          "content": "<p>got it - thanks for clarifying</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1692620,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "02/16/2022 06:37:33",
      "content": "<p>thanks for sharing.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1711215,
      "author_name": "rustembek158",
      "author_url": "",
      "post_date": "03/03/2022 18:23:17",
      "content": "<p>👍👍👍👍👍👍👍👍👍👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1691631": "Congratulates all the participants and many thanks to organizers. It was very challenging and interesting competition.\n\nThis post was written jointly with [@danjafish](https://www.kaggle.com/danjafish) and we want to say a big thank you to our team.\n\n# Summary\nOur solution is an ensemble of two differently trained yolov5 models + tracking.\n![summary](https://i.imgur.com/hrGOAv1.png)\n\nOur final submit code is available here\n[https://www.kaggle.com/parapapapam/2xyolov5l6-tracking-lb-0-678-private-0-722](https://www.kaggle.com/parapapapam/2xyolov5l6-tracking-lb-0-678-private-0-722)\n\n# Models\nIn our solution we used two yolov5 models trained in different ways.\n\n## YOLOv5l6 1280 \n(by [@parapapapam](https://www.kaggle.com/parapapapam))\n\n### Idea\nThe main trick of this competition is the use of high resolution. But when you just resize the image, you increase the noise without adding any additional information, so the reason why it works better not in the image.\n\nMy hypothesis was that the reason of that effect is in achors assignment for small object. When you increase image resolution you make your objects bigger, but anchors grid step is the same and you get better allignment so model works better.\n\nTherefore I decided do not increase input resolution, but use higher resolution inside the model. For that I changed stride of second convolutional layer from 2 to 1:\n\n```yaml\n...\n[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n...\n```\n\nSo this model will have 2 times bigger resolution starting from 3rd convolutional layer and should behaves like 2560 model but without resizing the original image.\n\n### Validation Strategy\n\nSplit by video. Video 1 was used as a validation. For submit model was retrained on videos 0 1 with validation on video 2.\n\n### Data\n\nI’ve used kaggle data with 6 iterations of fixing the labels. I’ve checked FP errors of models with high confidence and fix them manually using labelImg. It improves CV without any effect on LB. \nI’ve trained on labeled images + 10% unlabeled data picked at random.\n\n### Augmentations\n\nI used albumentations with the following pipeline:\n\n```python\naug_strong = [\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=1.0),\n    A.OneOf([\n        HueSaturationValue(p=0.5, hue_limit=0.05, sat_limit=0.7, val_limit=0.4),\n        A.RandomBrightnessContrast(p=0.5, brightness_limit=0.15, contrast_limit=0.15),\n        A.ToGray(p=0.1),\n    ], p=0.8),\n    A.OneOf([\n        ShiftScaleRotate(p=0.4, shift_limit=0.1, rotate_limit=3, scale_limit=(0.5, 2.0), min_bbox=27, max_bbox=None, border_mode=0, value=(114, 114, 114)),\n        A.Perspective(p=0.1, scale=0.1, pad_mode=0, pad_val=(114, 114, 114)),\n    ], p=0.8),\n    A.OneOf([\n        A.GaussianBlur(p=0.5, blur_limit=(3, 5)),\n        A.MotionBlur(p=0.5, blur_limit=(3, 5)),\n        A.MultiplicativeNoise(p=0.5, multiplier=(0.9, 1.1), per_channel=False, elementwise=True),\n    ], p=0.2),\n]\nbbox_params = A.BboxParams(format=\"yolo\", label_fields=[\"class_labels\"], min_area=16, min_visibility=0.2)\ntransform = A.Compose(aug_strong, bbox_params=bbox_params)\n```\n\nAnd from YOLOv5 mosaic=1.0 and mixup=0.5.\n\nAlso I’ve used enchancement based on clahe and channel stitching for train and inference. It gave around +0.005 CV/LB.\n\n### Train params\n\n10 epochs, 8 batch size, SGD 0.01, OneCycle, input size 1280x720.\n\nAlso I’ve changed obj: 8.0 because it depends on image size and I have to multiply it by factor of 2 ^ 2 = 4 because I’ve used two times bigger resolution inside the model and also experiment shows that multiplying it by 2 also a bit increased CV.\n\n### Results\n\nCV ~0.65 (+ tracking 0.66) / Public 0.67 (+ tracking 0.685) / Private 0.699 (+ tracking 0.718)\n\n### What didn’t work\n\nCopy paste COTS bboxes augmentation. Style transfer. External data.\n\n## YOLOv5l6 3100 \n(by [@danjafish](https://www.kaggle.com/danjafish))\n\nMy approach was quite straightforward. I got most of my ideas from [https://www.kaggle.com/c/global-wheat-detection/discussion/172569](https://www.kaggle.com/c/global-wheat-detection/discussion/172569) and [https://www.kaggle.com/c/global-wheat-detection/discussion/172418](https://www.kaggle.com/c/global-wheat-detection/discussion/172418). However, I decided to use yolov5 for a start since I found it the most promising.\n\n- Validation strategy: split by video. Video 1 was used as a validation. After selecting thresholds and other parameters, the model was re-trained on all data.\n- Data: I used only kaggle data. Trained on not empty images only. Adding empty ones slightly improved scoring on validation, but significantly reduced LB. That's why I gave it up.\n- Various augmentations from [albumentations](https://albumentations.ai/) (I borrowed pipeline from [https://www.kaggle.com/c/global-wheat-detection/discussion/172569](https://www.kaggle.com/c/global-wheat-detection/discussion/172569)):\n    - HorizontalFlip, ShiftScaleRotate, RandomRotate90\n    - RandomBrightnessContrast, HueSaturationValue, RGBShift\n    - RandomGamma\n    - CLAHE\n    - Blur, MotionBlur\n    - GaussNoise\n    - ImageCompression\n    - CoarseDropout\n- I also changed the default mosaic implementation to get images of about the same size and made up parts of images with CenterCrop.\n- Inference size: 3100. I tried different sizes and different architectures. The best results on validation were shown by yolov5l on size 3000X3000. However it was much worse on LB than l6 3100X3100. That's why I used the later one.\n- Train params: 20 epochs, Adam optimizer, lr 0.001 ony cycle sheduler. bs=4\n- Hardware: most of the time I used 4 V100 GPUs with 16GB RAM. Model takes about 2 hours to train.\n- What did’t work:\n    - Manual reannotations of data - some COTS were not labeled because they appeared in frames later than updated annotation. (my assumption). Adding these COTS slightly improved CV, but significantly worsened LB.\n    - TTA. Slightly impoved LB and CV, but takes much longer to run.\n    - Blend of different yolo architectures\n    - Blend one stage and two stage detectors. I spent the last week completely devoted to the models on mmdetection, hoping to improve the ensemble score. However, my cascade rcnn showed much worse results on both CV and LB.\n- Final scores CV/Public+tracking/Private+tracking: ~0.643(video 1)/0.693/0.722\n\n# Ensemble\n\nWe took predicts of two above models and use WBF with iou_threshold = 0.6 to blend them with weights 0.3 and 0.7. After that we filter results with confidence threshold ≥ 0.1.\n\nBut we didn’t spend enough time to tune the weights and the final result seems not much better that each model, but we have strong models, so each of them + tracking already is in gold :).\n\n# Tracking\n\nI’ve already posted tracking approach based on norfair library\n[https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539](https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539)\nSo I will not go into details of norfair tracking implementation based on SORT algorithm.\n\nHere I want to mention that this problem is not a MOT problem because objects are almost static and we have to deal with only camera movement.\nFor that we’ve developed algorithm of homography calculation between every two consecutive frames using ORB keypoint descriptors. Next we transform our bboxes using homography matrix to obtain predictions on the next frame.It gives much accurate predictions on two consecutive frames comparing to norfair:\n![frame1](https://i.imgur.com/vwMG9AP.png)\n![frame2](https://i.imgur.com/zwZjqRu.png)\n\nGreen labels is a ground truth, the blue ones are predictions using just norfair tracking, red ones — homography transform from previous frame + norfair. As we can see, homography + norfair match much better than just norfair.\nWhen we have no detection for object that was detected at least two frames before we predict it 2 more frames using tracking bboxes and if the object doesn’t appear remove it from tracker.\nAlso we filtered objects that are moving out of the image.\n\nOur tracking approach gives stable +0.01-0.02 CV/Public/Private.\n\n# Other ideas\n\n### Rounding\nWe found that using of bboxes = bboxes.astype(int) instead of bboxes = bboxes.round().astype(int) improves CV and Public ~ 0.005 and have no effect on Private. It seems the same like was already discussed\n[https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307605](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307605).\n\nThank you for reading and happy Kaggling! :)",
    "1691640": ">[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n\nThat's awesome. Thanks for sharing this trick! I will use this in the future!\n\nCongratulations to you and your team for finishing 10th place Gold. Fantastic job! \n\nI like how simple and elegant your solution is. Using homography transform is a great idea.",
    "1691726": "A doubt, how does changing stride, increases the resolution? and doesn't it affect the pretrained weights, since its been trained on different stride and the following layers getting different input shape",
    "1691734": "`...\n[-1, 1, Conv, [128, 3, 1]],  # 1-P2/4\n...`\n\ncoool! Outstanding tip. Congratulations!",
    "1691737": "These models (and layers) are fully convolutional. Therefore they don't care what input shape you give them.\n\nEach layer is trained to recognize a certain pattern in the image. If the previous layer doesn't do stride 2 after convolution, then the current layer just receives a 2x larger image. The current layer can still use its pretraining. It will just see an image that is 2x larger.",
    "1692202": "Great ideas. When changing the stride, were you able to directly use pretrained models from the yolov5 repo ?",
    "1692376": "Cong bro! How did you make sure non-overfitting of your model when retrain on the whole set?",
    "1692379": "I also tried cascade rcnn for a week and got 0.15 in LB 😂",
    "1692521": "\"When you increase image resolution you make your objects bigger, but anchors grid step is the same and you get better allignment so model works better.\" Should this be resolved by recomputing the anchor boxes with yolo k-means clustering and using custom anchors? I had tried using custom anchors with no improvement in CV. Curious if you had tried using custom anchors. Congrats on the medal.",
    "1692588": "Thanks Chris.\nCongratulations to you!\nAnd thank you for all knowledge you shared on kaggle. :)",
    "1692594": "Thanks Remek.\nCongratulations to you and your team!\nYou and your team did a big contribution in this competition.",
    "1692597": "Thanks Alexandre.\nYes because the model is fully convolutional and it doesn't depend on the shape of input tensor.",
    "1692600": "Thanks Sean.\nYou have to keep exactly the same config. And also we observed that the validation metrics don't overfit when you train longer.",
    "1692604": "Thanks Rajaram.\nThe difference not only in anchor sizes, but also in a grid size.\nIf you make resolution 2 times bigger inside the model you make your anchor grid step 2 times smaller and it makes assignment between objects and anchors better so this is the key.\nRegarding custom anchors — yes we used custom anchors for 1280 model, but it has not so much effect comparing to the default anchors.",
    "1692619": "got it - thanks for clarifying",
    "1692620": "thanks for sharing.",
    "1711215": "👍👍👍👍👍👍👍👍👍👍"
  },
  "source": "meta"
}