{
  "id": 307753,
  "title": "12th place solution - YOLOv5 + Optical flow tracker",
  "url": "/competitions/tensorflow-great-barrier-reef/writeups/d-imanishi-12th-place-solution-yolov5-optical-flow",
  "author_name": "",
  "post_date": "2022-02-18T12:39:06.027Z",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, I would like to thank the host for organizing such an exciting competition. I am glad to survive this big shake!</p>\n<h2>Detection model training</h2>\n<p>Because I utilized only the kaggle-kernel's GPU, I wanted to reduce GPU memory usage and training time. So I upscaled the training images by two times (2560x1440) and split them into four 1280x720 images. After that, I discarded the split image not containing bounding boxes. So that means the training size was 1280. In the inference phase, 2560 size was used. By doing this, I thought that the results are almost same to those obtained when training with 2560 size with reduced memory usage.</p>\n<ul>\n<li>YOLOv5m6</li>\n<li>4 folds Group-Kfold grouped by sequence</li>\n<li>1280 size with x2 upsacaled &amp; 4 split images</li>\n<li>Augmentation: Mainly, flipud p=0.5, mixup p=0.5, rot90 p=0.5 was changed from \"hyp.scratch.yaml\". To implement rot90, I modified \"augmentations.py\".</li>\n</ul>\n<h2>Tracking</h2>\n<p>I reused my optical flow based tracking code which I made in the <a href=\"https://www.kaggle.com/c/nfl-health-and-safety-helmet-assignment/discussion/285156#1569583\" target=\"_blank\">NFL competition</a>. Utilizing optical flow, next frame's bouding boxes can be estimated from previous frame's detection result. So it can continue tracking even if the detection model lost detection. I varied the number of frames continuing tracking in the lost detection condition according to the number of tracking counts up to that frame, and its maximum number is five.</p>\n<h2>Choice of final submission</h2>\n<p>In my CV, I noticed that upscaling the image size equally both of training size and inference size from original 1280x720 size contributed to both CV score and LB score. However, using different  size between training and inference (for example training 2560 and inference 3840) was only contributed to LB score and got worse the CV score. So I suspected the possibility of overfitting and I chose the following four submissions.</p>\n<h5>1. Best public LB</h5>\n<ul>\n<li>Ensemble of following 3 models<ul>\n<li>Best 2 public LB score models out of 4 folds with inference size of 3840 (x1.5 upscaled from training size)</li>\n<li>YOLOv5s model shared by <a href=\"https://www.kaggle.com/freshair1996\" target=\"_blank\">@freshair1996</a> (score 0.665) with inference size of 6400</li></ul></li>\n<li>TTA: original, LR flip, UD flip (*)</li>\n<li>Ensemble: 3 models x 3 TTAs were ensembled by WBF<br>\n(*) The TTA pattern was limited by 9 hours limitation. If more TTA pattern is used, the more score will be got.</li>\n</ul>\n<h5>2. Best CV</h5>\n<ul>\n<li>All of 4 folds with inference size of 2560 (= training size)</li>\n<li>TTA: original, LR flip, UD flip, Rot90</li>\n<li>Ensemble: 4 models x 4 TTAs were ensembled by WBF</li>\n</ul>\n<h5>3. Middle of CV and public LB</h5>\n<ul>\n<li>Best 2 public LB score models out of 4 folds</li>\n<li>TTA: 2 inference size (2560, 3840) x original, LR flip, UD flip</li>\n<li>Ensemble: 2 models x 2 sizes x 3 TTAs were ensembled by WBF</li>\n</ul>\n<h5>4. TensorFlow EfficientDet</h5>\n<p>To get the TensorFlow prize, I trained the EfficientDet-D2 model, but the private LB score of it was 0.543…</p>\n<p>The scores are following.</p>\n<table>\n<thead>\n<tr>\n<th>submission</th>\n<th>Private LB w/ Tracking</th>\n<th>Private LB w/o Tracking</th>\n<th>Public LB w/ Tracking</th>\n<th>Public LB w/o Tracking</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.712</td>\n<td>0.697</td>\n<td><strong>0.718</strong></td>\n<td>0.703</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.708</td>\n<td>0.700</td>\n<td>0.621</td>\n<td>0.609</td>\n</tr>\n<tr>\n<td>3</td>\n<td><strong>0.718</strong></td>\n<td>0.705</td>\n<td>0.667</td>\n<td>0.658</td>\n</tr>\n</tbody>\n</table>\n<p># Updated 13th place to 12th place</p>",
  "messages": [
    {
      "id": "1691612",
      "postDate": "02/15/2022 13:57:32",
      "content": "<p>First of all, I would like to thank the host for organizing such an exciting competition. I am glad to survive this big shake!</p>\n<h2>Detection model training</h2>\n<p>Because I utilized only the kaggle-kernel's GPU, I wanted to reduce GPU memory usage and training time. So I upscaled the training images by two times (2560x1440) and split them into four 1280x720 images. After that, I discarded the split image not containing bounding boxes. So that means the training size was 1280. In the inference phase, 2560 size was used. By doing this, I thought that the results are almost same to those obtained when training with 2560 size with reduced memory usage.</p>\n<ul>\n<li>YOLOv5m6</li>\n<li>4 folds Group-Kfold grouped by sequence</li>\n<li>1280 size with x2 upsacaled &amp; 4 split images</li>\n<li>Augmentation: Mainly, flipud p=0.5, mixup p=0.5, rot90 p=0.5 was changed from \"hyp.scratch.yaml\". To implement rot90, I modified \"augmentations.py\".</li>\n</ul>\n<h2>Tracking</h2>\n<p>I reused my optical flow based tracking code which I made in the <a href=\"https://www.kaggle.com/c/nfl-health-and-safety-helmet-assignment/discussion/285156#1569583\" target=\"_blank\">NFL competition</a>. Utilizing optical flow, next frame's bouding boxes can be estimated from previous frame's detection result. So it can continue tracking even if the detection model lost detection. I varied the number of frames continuing tracking in the lost detection condition according to the number of tracking counts up to that frame, and its maximum number is five.</p>\n<h2>Choice of final submission</h2>\n<p>In my CV, I noticed that upscaling the image size equally both of training size and inference size from original 1280x720 size contributed to both CV score and LB score. However, using different  size between training and inference (for example training 2560 and inference 3840) was only contributed to LB score and got worse the CV score. So I suspected the possibility of overfitting and I chose the following four submissions.</p>\n<h5>1. Best public LB</h5>\n<ul>\n<li>Ensemble of following 3 models<ul>\n<li>Best 2 public LB score models out of 4 folds with inference size of 3840 (x1.5 upscaled from training size)</li>\n<li>YOLOv5s model shared by <a href=\"https://www.kaggle.com/freshair1996\" target=\"_blank\">@freshair1996</a> (score 0.665) with inference size of 6400</li></ul></li>\n<li>TTA: original, LR flip, UD flip (*)</li>\n<li>Ensemble: 3 models x 3 TTAs were ensembled by WBF<br>\n(*) The TTA pattern was limited by 9 hours limitation. If more TTA pattern is used, the more score will be got.</li>\n</ul>\n<h5>2. Best CV</h5>\n<ul>\n<li>All of 4 folds with inference size of 2560 (= training size)</li>\n<li>TTA: original, LR flip, UD flip, Rot90</li>\n<li>Ensemble: 4 models x 4 TTAs were ensembled by WBF</li>\n</ul>\n<h5>3. Middle of CV and public LB</h5>\n<ul>\n<li>Best 2 public LB score models out of 4 folds</li>\n<li>TTA: 2 inference size (2560, 3840) x original, LR flip, UD flip</li>\n<li>Ensemble: 2 models x 2 sizes x 3 TTAs were ensembled by WBF</li>\n</ul>\n<h5>4. TensorFlow EfficientDet</h5>\n<p>To get the TensorFlow prize, I trained the EfficientDet-D2 model, but the private LB score of it was 0.543…</p>\n<p>The scores are following.</p>\n<table>\n<thead>\n<tr>\n<th>submission</th>\n<th>Private LB w/ Tracking</th>\n<th>Private LB w/o Tracking</th>\n<th>Public LB w/ Tracking</th>\n<th>Public LB w/o Tracking</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.712</td>\n<td>0.697</td>\n<td><strong>0.718</strong></td>\n<td>0.703</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.708</td>\n<td>0.700</td>\n<td>0.621</td>\n<td>0.609</td>\n</tr>\n<tr>\n<td>3</td>\n<td><strong>0.718</strong></td>\n<td>0.705</td>\n<td>0.667</td>\n<td>0.658</td>\n</tr>\n</tbody>\n</table>\n<p># Updated 13th place to 12th place</p>",
      "rawMarkdown": "First of all, I would like to thank the host for organizing such an exciting competition. I am glad to survive this big shake!\n\n## Detection model training\nBecause I utilized only the kaggle-kernel's GPU, I wanted to reduce GPU memory usage and training time. So I upscaled the training images by two times (2560x1440) and split them into four 1280x720 images. After that, I discarded the split image not containing bounding boxes. So that means the training size was 1280. In the inference phase, 2560 size was used. By doing this, I thought that the results are almost same to those obtained when training with 2560 size with reduced memory usage.\n \n- YOLOv5m6\n- 4 folds Group-Kfold grouped by sequence\n- 1280 size with x2 upsacaled & 4 split images\n- Augmentation: Mainly, flipud p=0.5, mixup p=0.5, rot90 p=0.5 was changed from \"hyp.scratch.yaml\". To implement rot90, I modified \"augmentations.py\".\n\n## Tracking\nI reused my optical flow based tracking code which I made in the [NFL competition](https://www.kaggle.com/c/nfl-health-and-safety-helmet-assignment/discussion/285156#1569583). Utilizing optical flow, next frame's bouding boxes can be estimated from previous frame's detection result. So it can continue tracking even if the detection model lost detection. I varied the number of frames continuing tracking in the lost detection condition according to the number of tracking counts up to that frame, and its maximum number is five.\n\n\n## Choice of final submission\nIn my CV, I noticed that upscaling the image size equally both of training size and inference size from original 1280x720 size contributed to both CV score and LB score. However, using different  size between training and inference (for example training 2560 and inference 3840) was only contributed to LB score and got worse the CV score. So I suspected the possibility of overfitting and I chose the following four submissions.\n\n##### 1. Best public LB\n- Ensemble of following 3 models\n    - Best 2 public LB score models out of 4 folds with inference size of 3840 (x1.5 upscaled from training size)\n    - YOLOv5s model shared by @freshair1996 (score 0.665) with inference size of 6400\n- TTA: original, LR flip, UD flip (*)\n- Ensemble: 3 models x 3 TTAs were ensembled by WBF\n(*) The TTA pattern was limited by 9 hours limitation. If more TTA pattern is used, the more score will be got.\n\n##### 2. Best CV\n- All of 4 folds with inference size of 2560 (= training size)\n- TTA: original, LR flip, UD flip, Rot90\n- Ensemble: 4 models x 4 TTAs were ensembled by WBF\n\n##### 3. Middle of CV and public LB\n- Best 2 public LB score models out of 4 folds\n- TTA: 2 inference size (2560, 3840) x original, LR flip, UD flip\n- Ensemble: 2 models x 2 sizes x 3 TTAs were ensembled by WBF\n\n##### 4. TensorFlow EfficientDet\nTo get the TensorFlow prize, I trained the EfficientDet-D2 model, but the private LB score of it was 0.543...\n\nThe scores are following.\n\n| submission | Private LB w/ Tracking | Private LB w/o Tracking | Public LB w/ Tracking | Public LB w/o Tracking |\n| --- | --- | --- | --- | --- |\n| 1 | 0.712 | 0.697 | **0.718** | 0.703 |\n| 2 | 0.708 | 0.700 | 0.621 | 0.609 |\n| 3 | **0.718** | 0.705 | 0.667 | 0.658 |\n\n\\# Updated 13th place to 12th place",
      "votes": null
    },
    {
      "id": "1691628",
      "postDate": "02/15/2022 14:16:01",
      "content": "<p>Thanks for sharing and congrats on gold! I was also in NFL competition (61st…). Many of public notebook have used norfair as a tracking method. Did you compare optical flow with that? It would be appreciated if you share the experiment results.</p>",
      "rawMarkdown": "Thanks for sharing and congrats on gold! I was also in NFL competition (61st...). Many of public notebook have used norfair as a tracking method. Did you compare optical flow with that? It would be appreciated if you share the experiment results.",
      "votes": null
    },
    {
      "id": "1691708",
      "postDate": "02/15/2022 15:06:25",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "1691747",
      "postDate": "02/15/2022 15:31:53",
      "content": "<p>Thank you for sharing your model. It improved my public score significantly!</p>",
      "rawMarkdown": "Thank you for sharing your model. It improved my public score significantly!",
      "votes": null
    },
    {
      "id": "1691753",
      "postDate": "02/15/2022 15:33:33",
      "content": "<p>Thanks for comments. In my experiment, norfair did not worked for my CV score.</p>",
      "rawMarkdown": "Thanks for comments. In my experiment, norfair did not worked for my CV score.",
      "votes": null
    },
    {
      "id": "1692726",
      "postDate": "02/16/2022 07:52:56",
      "content": "<p>Thanks for your kindness!</p>",
      "rawMarkdown": "Thanks for your kindness!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1691628,
      "author_name": "ttkagglett",
      "author_url": "",
      "post_date": "02/15/2022 14:16:01",
      "content": "<p>Thanks for sharing and congrats on gold! I was also in NFL competition (61st…). Many of public notebook have used norfair as a tracking method. Did you compare optical flow with that? It would be appreciated if you share the experiment results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1691753,
          "author_name": "dimanishi",
          "author_url": "",
          "post_date": "02/15/2022 15:33:33",
          "content": "<p>Thanks for comments. In my experiment, norfair did not worked for my CV score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692726,
          "author_name": "ttkagglett",
          "author_url": "",
          "post_date": "02/16/2022 07:52:56",
          "content": "<p>Thanks for your kindness!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691708,
      "author_name": "freshair1996",
      "author_url": "",
      "post_date": "02/15/2022 15:06:25",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1691747,
          "author_name": "dimanishi",
          "author_url": "",
          "post_date": "02/15/2022 15:31:53",
          "content": "<p>Thank you for sharing your model. It improved my public score significantly!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1691612": "First of all, I would like to thank the host for organizing such an exciting competition. I am glad to survive this big shake!\n\n## Detection model training\nBecause I utilized only the kaggle-kernel's GPU, I wanted to reduce GPU memory usage and training time. So I upscaled the training images by two times (2560x1440) and split them into four 1280x720 images. After that, I discarded the split image not containing bounding boxes. So that means the training size was 1280. In the inference phase, 2560 size was used. By doing this, I thought that the results are almost same to those obtained when training with 2560 size with reduced memory usage.\n \n- YOLOv5m6\n- 4 folds Group-Kfold grouped by sequence\n- 1280 size with x2 upsacaled & 4 split images\n- Augmentation: Mainly, flipud p=0.5, mixup p=0.5, rot90 p=0.5 was changed from \"hyp.scratch.yaml\". To implement rot90, I modified \"augmentations.py\".\n\n## Tracking\nI reused my optical flow based tracking code which I made in the [NFL competition](https://www.kaggle.com/c/nfl-health-and-safety-helmet-assignment/discussion/285156#1569583). Utilizing optical flow, next frame's bouding boxes can be estimated from previous frame's detection result. So it can continue tracking even if the detection model lost detection. I varied the number of frames continuing tracking in the lost detection condition according to the number of tracking counts up to that frame, and its maximum number is five.\n\n\n## Choice of final submission\nIn my CV, I noticed that upscaling the image size equally both of training size and inference size from original 1280x720 size contributed to both CV score and LB score. However, using different  size between training and inference (for example training 2560 and inference 3840) was only contributed to LB score and got worse the CV score. So I suspected the possibility of overfitting and I chose the following four submissions.\n\n##### 1. Best public LB\n- Ensemble of following 3 models\n    - Best 2 public LB score models out of 4 folds with inference size of 3840 (x1.5 upscaled from training size)\n    - YOLOv5s model shared by @freshair1996 (score 0.665) with inference size of 6400\n- TTA: original, LR flip, UD flip (*)\n- Ensemble: 3 models x 3 TTAs were ensembled by WBF\n(*) The TTA pattern was limited by 9 hours limitation. If more TTA pattern is used, the more score will be got.\n\n##### 2. Best CV\n- All of 4 folds with inference size of 2560 (= training size)\n- TTA: original, LR flip, UD flip, Rot90\n- Ensemble: 4 models x 4 TTAs were ensembled by WBF\n\n##### 3. Middle of CV and public LB\n- Best 2 public LB score models out of 4 folds\n- TTA: 2 inference size (2560, 3840) x original, LR flip, UD flip\n- Ensemble: 2 models x 2 sizes x 3 TTAs were ensembled by WBF\n\n##### 4. TensorFlow EfficientDet\nTo get the TensorFlow prize, I trained the EfficientDet-D2 model, but the private LB score of it was 0.543...\n\nThe scores are following.\n\n| submission | Private LB w/ Tracking | Private LB w/o Tracking | Public LB w/ Tracking | Public LB w/o Tracking |\n| --- | --- | --- | --- | --- |\n| 1 | 0.712 | 0.697 | **0.718** | 0.703 |\n| 2 | 0.708 | 0.700 | 0.621 | 0.609 |\n| 3 | **0.718** | 0.705 | 0.667 | 0.658 |\n\n\\# Updated 13th place to 12th place",
    "1691628": "Thanks for sharing and congrats on gold! I was also in NFL competition (61st...). Many of public notebook have used norfair as a tracking method. Did you compare optical flow with that? It would be appreciated if you share the experiment results.",
    "1691708": "Congratulations!",
    "1691747": "Thank you for sharing your model. It improved my public score significantly!",
    "1691753": "Thanks for comments. In my experiment, norfair did not worked for my CV score.",
    "1692726": "Thanks for your kindness!"
  },
  "source": "meta"
}