{
  "id": 307807,
  "title": "41st place Solution (7th Public LB) YOLO-Z Inspired Model with Tracking",
  "url": "/competitions/tensorflow-great-barrier-reef/writeups/outwrest-41st-place-solution-7th-public-lb-yolo-z-",
  "author_name": "",
  "post_date": "2022-02-15T17:37:41.705010200Z",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I want to start by saying that I am not experienced in data science. I don't exactly know all the bells and whistles in data science but I love to learn and experiment. In terms of this competition, I learned a ton about object detection and it was really fun!</p>\n<h1>Data</h1>\n<p>I used a groupkfold split on sequences since the data in private LB will be new sequences, the best way to validate a model (at least in my noob eyes) is with new sequences it has not seen before. I only trained models on the same fold and did not ensemble or use validate models via oof predictions. </p>\n<h1>Model</h1>\n<p>My solution revolved around the <a href=\"https://arxiv.org/pdf/2112.11798.pdf\" target=\"_blank\">YOLO-Z</a> paper that looked into improving the original YOLO5 architecture in small object detection. It looks like their proposed model is somewhat included in YOLO5v6. They mention that in their experiments modifying the \"width\" multiplier to that of the next tier (using L or M \"width\" for M or S models) they achieved better results in small object detection. Since YOLO5v6 already includes smaller layers (see <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/303766\" target=\"_blank\">this</a> discussion thread), I went into the direction of modifying YOLO5v6 models to increase the width multiplier (to see if this does better than the original). This increases the number of parameters but not the number of layers. This would keep inference time lower than using a bigger model and give the best of both tiers.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Layers</th>\n<th>params (M)</th>\n<th>FLOPs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>YOLOv5s6</td>\n<td>280</td>\n<td>12.6</td>\n<td>16.8</td>\n</tr>\n<tr>\n<td>YOLOv5s6-wide</td>\n<td>280</td>\n<td>27.6</td>\n<td>35.8</td>\n</tr>\n<tr>\n<td>YOLOv5m6</td>\n<td>378</td>\n<td>35.7</td>\n<td>50.0</td>\n</tr>\n<tr>\n<td>YOLOv5m6-wide</td>\n<td>378</td>\n<td>62.6</td>\n<td>86.5</td>\n</tr>\n<tr>\n<td>YOLOv5l6</td>\n<td>476</td>\n<td>76.7</td>\n<td>111.4</td>\n</tr>\n</tbody>\n</table>\n<p>To not have to re-train the models, I had to use to transfer weights from the next model (if I want to use YOLOv5s6-wide, use YOLOv5m6 weights, etc.). I still wanted to pre-train weights since I deleted a lot of layers in the process.</p>\n<p>I only selected YOLOv5s6-wide and YOLOv5m6-wide to train and use.</p>\n<h1>Training</h1>\n<p>Because the \"wider\" models deleted many layers I thought re-training would help better fit the models. After searching a bit I found <a href=\"https://github.com/chongweiliu/DUO\" target=\"_blank\">Detecting Underwater Objects</a> dataset based on competitions listed <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/304367\" target=\"_blank\">here</a>. </p>\n<p>I trained for 100 epochs with heavy augmentations (see <a href=\"https://wandb.ai/outwrest/gbr/reports/m-wide-duo-1280--VmlldzoxNTYzMDk0?accessToken=q1nh3kbk1u39ydfs53h1ffnuchpig7sntec09i7rhkszjd2dcjw6wenilritoyno\" target=\"_blank\">wandb output</a>). </p>\n<p>I used the COT mask dataset to make 10,000 \"fake\" images of cots pasted in places with different opacity levels. I found that training with this dataset increases recall in the final model (thanks <a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a>). But I don't 100% know if it helps in public LB.<br>\nSee this <a href=\"https://www.kaggle.com/outwrest/augmentation-using-image-blending-cot-masks\" target=\"_blank\">notebook</a> and <a href=\"https://wandb.ai/outwrest/gbr/reports/m-wide-fake-1280--VmlldzoxNTYzMTQz?accessToken=341vkrc709f82dx5xs0vxlv0j17wqj5asq8wptfc20x7jyun27ud4r2208pu9448\" target=\"_blank\">results</a>.</p>\n<p>Trained again on sequences with 3000 IMG size -&gt; and fined tuned again by training with a much lower LR and less augmentations. I used simple norfair tracking notebook as a postprocessing technique.</p>\n<h1>Overfitting Public LB</h1>\n<p>I noticed that by submitting 2x img size (6000) without TTA I was able to overfit LB pretty easily and achieved 7th place in LB with some CONF tunning. </p>\n<p>The best overfit was YOLOv5s6-wide with 7200 img size and no TTA trained on 3600 img size, 0.769 Public-0.641 Private.<br>\nThe best notebook I had is YOLOv5m6-wide with 6000 img size and TTA trained on 3000 img size, 0.673 Public-0.706 Private.</p>\n<h1>TLDR</h1>\n<p>Modified YOLOv5m6 model to include more parameters -&gt; pretrained on DUO dataset -&gt; pretrainedx2 on a \"fake\" dataset (I am not 100% sure this helped) -&gt; trained on sequences -&gt; finedtuned with lower LR again.</p>\n<p>Hopefully, this helps. I feel like the score can be improved even with an ensemble and more conf/iou tuning since I did not explore that much. Thanks to other competitors and I hope to learn from your writeups. </p>",
  "messages": [
    {
      "id": "1691938",
      "postDate": "02/15/2022 17:37:41",
      "content": "<p>I want to start by saying that I am not experienced in data science. I don't exactly know all the bells and whistles in data science but I love to learn and experiment. In terms of this competition, I learned a ton about object detection and it was really fun!</p>\n<h1>Data</h1>\n<p>I used a groupkfold split on sequences since the data in private LB will be new sequences, the best way to validate a model (at least in my noob eyes) is with new sequences it has not seen before. I only trained models on the same fold and did not ensemble or use validate models via oof predictions. </p>\n<h1>Model</h1>\n<p>My solution revolved around the <a href=\"https://arxiv.org/pdf/2112.11798.pdf\" target=\"_blank\">YOLO-Z</a> paper that looked into improving the original YOLO5 architecture in small object detection. It looks like their proposed model is somewhat included in YOLO5v6. They mention that in their experiments modifying the \"width\" multiplier to that of the next tier (using L or M \"width\" for M or S models) they achieved better results in small object detection. Since YOLO5v6 already includes smaller layers (see <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/303766\" target=\"_blank\">this</a> discussion thread), I went into the direction of modifying YOLO5v6 models to increase the width multiplier (to see if this does better than the original). This increases the number of parameters but not the number of layers. This would keep inference time lower than using a bigger model and give the best of both tiers.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Layers</th>\n<th>params (M)</th>\n<th>FLOPs</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>YOLOv5s6</td>\n<td>280</td>\n<td>12.6</td>\n<td>16.8</td>\n</tr>\n<tr>\n<td>YOLOv5s6-wide</td>\n<td>280</td>\n<td>27.6</td>\n<td>35.8</td>\n</tr>\n<tr>\n<td>YOLOv5m6</td>\n<td>378</td>\n<td>35.7</td>\n<td>50.0</td>\n</tr>\n<tr>\n<td>YOLOv5m6-wide</td>\n<td>378</td>\n<td>62.6</td>\n<td>86.5</td>\n</tr>\n<tr>\n<td>YOLOv5l6</td>\n<td>476</td>\n<td>76.7</td>\n<td>111.4</td>\n</tr>\n</tbody>\n</table>\n<p>To not have to re-train the models, I had to use to transfer weights from the next model (if I want to use YOLOv5s6-wide, use YOLOv5m6 weights, etc.). I still wanted to pre-train weights since I deleted a lot of layers in the process.</p>\n<p>I only selected YOLOv5s6-wide and YOLOv5m6-wide to train and use.</p>\n<h1>Training</h1>\n<p>Because the \"wider\" models deleted many layers I thought re-training would help better fit the models. After searching a bit I found <a href=\"https://github.com/chongweiliu/DUO\" target=\"_blank\">Detecting Underwater Objects</a> dataset based on competitions listed <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/304367\" target=\"_blank\">here</a>. </p>\n<p>I trained for 100 epochs with heavy augmentations (see <a href=\"https://wandb.ai/outwrest/gbr/reports/m-wide-duo-1280--VmlldzoxNTYzMDk0?accessToken=q1nh3kbk1u39ydfs53h1ffnuchpig7sntec09i7rhkszjd2dcjw6wenilritoyno\" target=\"_blank\">wandb output</a>). </p>\n<p>I used the COT mask dataset to make 10,000 \"fake\" images of cots pasted in places with different opacity levels. I found that training with this dataset increases recall in the final model (thanks <a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a>). But I don't 100% know if it helps in public LB.<br>\nSee this <a href=\"https://www.kaggle.com/outwrest/augmentation-using-image-blending-cot-masks\" target=\"_blank\">notebook</a> and <a href=\"https://wandb.ai/outwrest/gbr/reports/m-wide-fake-1280--VmlldzoxNTYzMTQz?accessToken=341vkrc709f82dx5xs0vxlv0j17wqj5asq8wptfc20x7jyun27ud4r2208pu9448\" target=\"_blank\">results</a>.</p>\n<p>Trained again on sequences with 3000 IMG size -&gt; and fined tuned again by training with a much lower LR and less augmentations. I used simple norfair tracking notebook as a postprocessing technique.</p>\n<h1>Overfitting Public LB</h1>\n<p>I noticed that by submitting 2x img size (6000) without TTA I was able to overfit LB pretty easily and achieved 7th place in LB with some CONF tunning. </p>\n<p>The best overfit was YOLOv5s6-wide with 7200 img size and no TTA trained on 3600 img size, 0.769 Public-0.641 Private.<br>\nThe best notebook I had is YOLOv5m6-wide with 6000 img size and TTA trained on 3000 img size, 0.673 Public-0.706 Private.</p>\n<h1>TLDR</h1>\n<p>Modified YOLOv5m6 model to include more parameters -&gt; pretrained on DUO dataset -&gt; pretrainedx2 on a \"fake\" dataset (I am not 100% sure this helped) -&gt; trained on sequences -&gt; finedtuned with lower LR again.</p>\n<p>Hopefully, this helps. I feel like the score can be improved even with an ensemble and more conf/iou tuning since I did not explore that much. Thanks to other competitors and I hope to learn from your writeups. </p>",
      "rawMarkdown": "I want to start by saying that I am not experienced in data science. I don't exactly know all the bells and whistles in data science but I love to learn and experiment. In terms of this competition, I learned a ton about object detection and it was really fun!\n\n# Data\nI used a groupkfold split on sequences since the data in private LB will be new sequences, the best way to validate a model (at least in my noob eyes) is with new sequences it has not seen before. I only trained models on the same fold and did not ensemble or use validate models via oof predictions. \n\n# Model\nMy solution revolved around the [YOLO-Z](https://arxiv.org/pdf/2112.11798.pdf) paper that looked into improving the original YOLO5 architecture in small object detection. It looks like their proposed model is somewhat included in YOLO5v6. They mention that in their experiments modifying the \"width\" multiplier to that of the next tier (using L or M \"width\" for M or S models) they achieved better results in small object detection. Since YOLO5v6 already includes smaller layers (see [this](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/303766) discussion thread), I went into the direction of modifying YOLO5v6 models to increase the width multiplier (to see if this does better than the original). This increases the number of parameters but not the number of layers. This would keep inference time lower than using a bigger model and give the best of both tiers.\n\n|     Model     | Layers | params (M) | FLOPs |\n|:-------------:|:------:|:----------:|:-----:|\n| YOLOv5s6      | 280    | 12.6       | 16.8  |\n| YOLOv5s6-wide | 280    | 27.6       | 35.8  |\n| YOLOv5m6      | 378    | 35.7       | 50.0  |\n| YOLOv5m6-wide | 378    | 62.6       | 86.5  |\n| YOLOv5l6      | 476    | 76.7       | 111.4 |\n\nTo not have to re-train the models, I had to use to transfer weights from the next model (if I want to use YOLOv5s6-wide, use YOLOv5m6 weights, etc.). I still wanted to pre-train weights since I deleted a lot of layers in the process.\n\nI only selected YOLOv5s6-wide and YOLOv5m6-wide to train and use.\n\n# Training\nBecause the \"wider\" models deleted many layers I thought re-training would help better fit the models. After searching a bit I found [Detecting Underwater Objects](https://github.com/chongweiliu/DUO) dataset based on competitions listed [here](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/304367). \n\nI trained for 100 epochs with heavy augmentations (see [wandb output](https://wandb.ai/outwrest/gbr/reports/m-wide-duo-1280--VmlldzoxNTYzMDk0?accessToken=q1nh3kbk1u39ydfs53h1ffnuchpig7sntec09i7rhkszjd2dcjw6wenilritoyno)). \n\nI used the COT mask dataset to make 10,000 \"fake\" images of cots pasted in places with different opacity levels. I found that training with this dataset increases recall in the final model (thanks @alexandrecc). But I don't 100% know if it helps in public LB.\nSee this [notebook](https://www.kaggle.com/outwrest/augmentation-using-image-blending-cot-masks) and [results](https://wandb.ai/outwrest/gbr/reports/m-wide-fake-1280--VmlldzoxNTYzMTQz?accessToken=341vkrc709f82dx5xs0vxlv0j17wqj5asq8wptfc20x7jyun27ud4r2208pu9448).\n\nTrained again on sequences with 3000 IMG size -> and fined tuned again by training with a much lower LR and less augmentations. I used simple norfair tracking notebook as a postprocessing technique.\n\n# Overfitting Public LB\nI noticed that by submitting 2x img size (6000) without TTA I was able to overfit LB pretty easily and achieved 7th place in LB with some CONF tunning. \n\nThe best overfit was YOLOv5s6-wide with 7200 img size and no TTA trained on 3600 img size, 0.769 Public-0.641 Private.\nThe best notebook I had is YOLOv5m6-wide with 6000 img size and TTA trained on 3000 img size, 0.673 Public-0.706 Private.\n\n# TLDR\nModified YOLOv5m6 model to include more parameters -> pretrained on DUO dataset -> pretrainedx2 on a \"fake\" dataset (I am not 100% sure this helped) -> trained on sequences -> finedtuned with lower LR again.\n\n\nHopefully, this helps. I feel like the score can be improved even with an ensemble and more conf/iou tuning since I did not explore that much. Thanks to other competitors and I hope to learn from your writeups.",
      "votes": null
    },
    {
      "id": "1692177",
      "postDate": "02/15/2022 21:59:20",
      "content": "<p>I was looking for Yolo-Z sources. For me it seems to be more changes to the Yolo-Z than you describe - such as a BiFPN head (we tried to create model using this head but we have not achieve good result) and a different backbone (ResNet/DenseNet). Definitely great attitude you presented and thank you very much for sharing your attitude. </p>",
      "rawMarkdown": "I was looking for Yolo-Z sources. For me it seems to be more changes to the Yolo-Z than you describe - such as a BiFPN head (we tried to create model using this head but we have not achieve good result) and a different backbone (ResNet/DenseNet). Definitely great attitude you presented and thank you very much for sharing your attitude.",
      "votes": null
    },
    {
      "id": "1692265",
      "postDate": "02/15/2022 23:49:07",
      "content": "<p>Thank you! I learned a lot from your posts</p>",
      "rawMarkdown": "Thank you! I learned a lot from your posts",
      "votes": null
    },
    {
      "id": "1695546",
      "postDate": "02/18/2022 07:53:50",
      "content": "<p>thanks for sharing.  </p>",
      "rawMarkdown": "thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1692177,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/15/2022 21:59:20",
      "content": "<p>I was looking for Yolo-Z sources. For me it seems to be more changes to the Yolo-Z than you describe - such as a BiFPN head (we tried to create model using this head but we have not achieve good result) and a different backbone (ResNet/DenseNet). Definitely great attitude you presented and thank you very much for sharing your attitude. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1692265,
          "author_name": "outwrest",
          "author_url": "",
          "post_date": "02/15/2022 23:49:07",
          "content": "<p>Thank you! I learned a lot from your posts</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1695546,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "02/18/2022 07:53:50",
      "content": "<p>thanks for sharing.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1691938": "I want to start by saying that I am not experienced in data science. I don't exactly know all the bells and whistles in data science but I love to learn and experiment. In terms of this competition, I learned a ton about object detection and it was really fun!\n\n# Data\nI used a groupkfold split on sequences since the data in private LB will be new sequences, the best way to validate a model (at least in my noob eyes) is with new sequences it has not seen before. I only trained models on the same fold and did not ensemble or use validate models via oof predictions. \n\n# Model\nMy solution revolved around the [YOLO-Z](https://arxiv.org/pdf/2112.11798.pdf) paper that looked into improving the original YOLO5 architecture in small object detection. It looks like their proposed model is somewhat included in YOLO5v6. They mention that in their experiments modifying the \"width\" multiplier to that of the next tier (using L or M \"width\" for M or S models) they achieved better results in small object detection. Since YOLO5v6 already includes smaller layers (see [this](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/303766) discussion thread), I went into the direction of modifying YOLO5v6 models to increase the width multiplier (to see if this does better than the original). This increases the number of parameters but not the number of layers. This would keep inference time lower than using a bigger model and give the best of both tiers.\n\n|     Model     | Layers | params (M) | FLOPs |\n|:-------------:|:------:|:----------:|:-----:|\n| YOLOv5s6      | 280    | 12.6       | 16.8  |\n| YOLOv5s6-wide | 280    | 27.6       | 35.8  |\n| YOLOv5m6      | 378    | 35.7       | 50.0  |\n| YOLOv5m6-wide | 378    | 62.6       | 86.5  |\n| YOLOv5l6      | 476    | 76.7       | 111.4 |\n\nTo not have to re-train the models, I had to use to transfer weights from the next model (if I want to use YOLOv5s6-wide, use YOLOv5m6 weights, etc.). I still wanted to pre-train weights since I deleted a lot of layers in the process.\n\nI only selected YOLOv5s6-wide and YOLOv5m6-wide to train and use.\n\n# Training\nBecause the \"wider\" models deleted many layers I thought re-training would help better fit the models. After searching a bit I found [Detecting Underwater Objects](https://github.com/chongweiliu/DUO) dataset based on competitions listed [here](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/304367). \n\nI trained for 100 epochs with heavy augmentations (see [wandb output](https://wandb.ai/outwrest/gbr/reports/m-wide-duo-1280--VmlldzoxNTYzMDk0?accessToken=q1nh3kbk1u39ydfs53h1ffnuchpig7sntec09i7rhkszjd2dcjw6wenilritoyno)). \n\nI used the COT mask dataset to make 10,000 \"fake\" images of cots pasted in places with different opacity levels. I found that training with this dataset increases recall in the final model (thanks @alexandrecc). But I don't 100% know if it helps in public LB.\nSee this [notebook](https://www.kaggle.com/outwrest/augmentation-using-image-blending-cot-masks) and [results](https://wandb.ai/outwrest/gbr/reports/m-wide-fake-1280--VmlldzoxNTYzMTQz?accessToken=341vkrc709f82dx5xs0vxlv0j17wqj5asq8wptfc20x7jyun27ud4r2208pu9448).\n\nTrained again on sequences with 3000 IMG size -> and fined tuned again by training with a much lower LR and less augmentations. I used simple norfair tracking notebook as a postprocessing technique.\n\n# Overfitting Public LB\nI noticed that by submitting 2x img size (6000) without TTA I was able to overfit LB pretty easily and achieved 7th place in LB with some CONF tunning. \n\nThe best overfit was YOLOv5s6-wide with 7200 img size and no TTA trained on 3600 img size, 0.769 Public-0.641 Private.\nThe best notebook I had is YOLOv5m6-wide with 6000 img size and TTA trained on 3000 img size, 0.673 Public-0.706 Private.\n\n# TLDR\nModified YOLOv5m6 model to include more parameters -> pretrained on DUO dataset -> pretrainedx2 on a \"fake\" dataset (I am not 100% sure this helped) -> trained on sequences -> finedtuned with lower LR again.\n\n\nHopefully, this helps. I feel like the score can be improved even with an ensemble and more conf/iou tuning since I did not explore that much. Thanks to other competitors and I hope to learn from your writeups.",
    "1692177": "I was looking for Yolo-Z sources. For me it seems to be more changes to the Yolo-Z than you describe - such as a BiFPN head (we tried to create model using this head but we have not achieve good result) and a different backbone (ResNet/DenseNet). Definitely great attitude you presented and thank you very much for sharing your attitude.",
    "1692265": "Thank you! I learned a lot from your posts",
    "1695546": "thanks for sharing."
  },
  "source": "meta"
}