{
  "id": 308336,
  "title": "16th Simple Solution - Only Yolov5",
  "url": "/competitions/tensorflow-great-barrier-reef/writeups/ian-hwigeon-statking-16th-simple-solution-only-yol",
  "author_name": "",
  "post_date": "2022-02-18T10:43:55.123Z",
  "votes": 26,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I would like to thank the organizers for hosting the great competition.<br>\nalso, I was very honored to become a teammate with <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> and <a href=\"https://www.kaggle.com/hwigeon\" target=\"_blank\">@hwigeon</a>. <br>\nI really learned a lot from our teammates. I think it will be the unforgettable good experience.</p>\n<h1>Summary</h1>\n<p>We reached LB 709 so quickly using a large-size inference trick and naive iou-tracker.<br>\nbut at the same time, we started to doubt the public LB since we couldn't reproduce the result clearly. we found out that Increasing image size increased recall but also yielded unnecessary bounding boxes.<br>\nSo, we started to build robust models by checking CV F2 scores.<br>\nIn CV, only inference with originally trained size worked.<br>\nWe tried size threshold-based prediction, IOU tracker(i might have made a mistake), larger size than originally trained size didn't work in the local.</p>\n<h1>Models</h1>\n<p>We found out that a large model necessarily isn't needed.<br>\nAlso what we found out is multi-scale ensemble boosts both cv and public LB scores.<br>\nInstead of investing our time in searching complex models, we prepared yolov5m trained with the sizes of [2400,2560,2688,2880,3000]<br>\nand yolov5s trained with the sizes of [3000,3600,3720,3840,3960,4032,4080,4244]</p>\n<h1>Traning Method</h1>\n<p>We changed yolo's hyperparameters slightly<br>\nApplying mixup 0.5, mosaic 1.0, strong hsv change helped a lot<br>\n10 epoch training and adam optimizer were chosen.<br>\nModel was chosen by weighting Recall vs Precision  4:1</p>\n<h1>Validation</h1>\n<p>We chose video-split folds since these are more realistic.<br>\nEnsembled CV scores for video-split were each [0.71, 0.64, 0.77]<br>\nAfter that, we had searched optimized WBF coefficients and confidence scores using a method of grid search.</p>\n<p>For video fold 0-1, WBF coefficient 0.5, skip bbox 0.01, confidence score 0.1 was best<br>\nFor video fold 2, WBF coefficient 0.5, skip_bbox 0.1, confidence score 0.2 was best</p>\n<p>With these values, we trained models with all the datasets and regarded the public LB as a holdout<br>\nthe public LB scores of each model were ranged in [0.55~0.60]</p>\n<p>The result is private LB 0.713 and public LB 0.648.<br>\nWe missed the gold but we couldn't choose the best private LB since this competition is a kind of shake-up competition</p>\n<p>Thank you for reading this!</p>",
  "messages": [
    {
      "id": "1695523",
      "postDate": "02/18/2022 07:41:45",
      "content": "<p>I would like to thank the organizers for hosting the great competition.<br>\nalso, I was very honored to become a teammate with <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> and <a href=\"https://www.kaggle.com/hwigeon\" target=\"_blank\">@hwigeon</a>. <br>\nI really learned a lot from our teammates. I think it will be the unforgettable good experience.</p>\n<h1>Summary</h1>\n<p>We reached LB 709 so quickly using a large-size inference trick and naive iou-tracker.<br>\nbut at the same time, we started to doubt the public LB since we couldn't reproduce the result clearly. we found out that Increasing image size increased recall but also yielded unnecessary bounding boxes.<br>\nSo, we started to build robust models by checking CV F2 scores.<br>\nIn CV, only inference with originally trained size worked.<br>\nWe tried size threshold-based prediction, IOU tracker(i might have made a mistake), larger size than originally trained size didn't work in the local.</p>\n<h1>Models</h1>\n<p>We found out that a large model necessarily isn't needed.<br>\nAlso what we found out is multi-scale ensemble boosts both cv and public LB scores.<br>\nInstead of investing our time in searching complex models, we prepared yolov5m trained with the sizes of [2400,2560,2688,2880,3000]<br>\nand yolov5s trained with the sizes of [3000,3600,3720,3840,3960,4032,4080,4244]</p>\n<h1>Traning Method</h1>\n<p>We changed yolo's hyperparameters slightly<br>\nApplying mixup 0.5, mosaic 1.0, strong hsv change helped a lot<br>\n10 epoch training and adam optimizer were chosen.<br>\nModel was chosen by weighting Recall vs Precision  4:1</p>\n<h1>Validation</h1>\n<p>We chose video-split folds since these are more realistic.<br>\nEnsembled CV scores for video-split were each [0.71, 0.64, 0.77]<br>\nAfter that, we had searched optimized WBF coefficients and confidence scores using a method of grid search.</p>\n<p>For video fold 0-1, WBF coefficient 0.5, skip bbox 0.01, confidence score 0.1 was best<br>\nFor video fold 2, WBF coefficient 0.5, skip_bbox 0.1, confidence score 0.2 was best</p>\n<p>With these values, we trained models with all the datasets and regarded the public LB as a holdout<br>\nthe public LB scores of each model were ranged in [0.55~0.60]</p>\n<p>The result is private LB 0.713 and public LB 0.648.<br>\nWe missed the gold but we couldn't choose the best private LB since this competition is a kind of shake-up competition</p>\n<p>Thank you for reading this!</p>",
      "rawMarkdown": "I would like to thank the organizers for hosting the great competition.\nalso, I was very honored to become a teammate with @vaillant and @hwigeon. \nI really learned a lot from our teammates. I think it will be the unforgettable good experience.\n\n# Summary\nWe reached LB 709 so quickly using a large-size inference trick and naive iou-tracker.\nbut at the same time, we started to doubt the public LB since we couldn't reproduce the result clearly. we found out that Increasing image size increased recall but also yielded unnecessary bounding boxes.\nSo, we started to build robust models by checking CV F2 scores.\nIn CV, only inference with originally trained size worked.\nWe tried size threshold-based prediction, IOU tracker(i might have made a mistake), larger size than originally trained size didn't work in the local.\n\n# Models\nWe found out that a large model necessarily isn't needed.\nAlso what we found out is multi-scale ensemble boosts both cv and public LB scores.\nInstead of investing our time in searching complex models, we prepared yolov5m trained with the sizes of [2400,2560,2688,2880,3000]\nand yolov5s trained with the sizes of [3000,3600,3720,3840,3960,4032,4080,4244]\n\n# Traning Method\nWe changed yolo's hyperparameters slightly\nApplying mixup 0.5, mosaic 1.0, strong hsv change helped a lot\n10 epoch training and adam optimizer were chosen.\nModel was chosen by weighting Recall vs Precision  4:1\n\n# Validation\nWe chose video-split folds since these are more realistic.\nEnsembled CV scores for video-split were each [0.71, 0.64, 0.77]\nAfter that, we had searched optimized WBF coefficients and confidence scores using a method of grid search.\n\nFor video fold 0-1, WBF coefficient 0.5, skip bbox 0.01, confidence score 0.1 was best\nFor video fold 2, WBF coefficient 0.5, skip_bbox 0.1, confidence score 0.2 was best\n\nWith these values, we trained models with all the datasets and regarded the public LB as a holdout\nthe public LB scores of each model were ranged in [0.55~0.60]\n\nThe result is private LB 0.713 and public LB 0.648.\nWe missed the gold but we couldn't choose the best private LB since this competition is a kind of shake-up competition\n\nThank you for reading this!",
      "votes": null
    },
    {
      "id": "1695880",
      "postDate": "02/18/2022 12:28:38",
      "content": "<blockquote>\n  <p>In CV, only inference with originally trained size worked.</p>\n</blockquote>\n<p>This line explains why most of public inference codes didn't work well.<br>\nAfter competition, I found that our small models (which have training image size = inference image size) work as good as high score public codes in private dataset.</p>\n<p>I wish I checked CV-LB correlation with image size.<br>\nCongrats for taking 16th!</p>",
      "rawMarkdown": "> In CV, only inference with originally trained size worked.\n\nThis line explains why most of public inference codes didn't work well.\nAfter competition, I found that our small models (which have training image size = inference image size) work as good as high score public codes in private dataset.\n\nI wish I checked CV-LB correlation with image size.\nCongrats for taking 16th!",
      "votes": null
    },
    {
      "id": "1695882",
      "postDate": "02/18/2022 12:32:31",
      "content": "<p>I also checked the same situation.<br>\nOur yoloX (1280,720 size) 0.399 in public LB but 0.597 in private LB<br>\nI learned again that we should verify cv-lb very carefully!<br>\nThank you for reading this post!</p>",
      "rawMarkdown": "I also checked the same situation.\nOur yoloX (1280,720 size) 0.399 in public LB but 0.597 in private LB\nI learned again that we should verify cv-lb very carefully!\nThank you for reading this post!",
      "votes": null
    },
    {
      "id": "1696015",
      "postDate": "02/18/2022 14:41:39",
      "content": "<p>Great to know about ensemble with multi-scale images. Thanks and congrats on silver medal!</p>",
      "rawMarkdown": "Great to know about ensemble with multi-scale images. Thanks and congrats on silver medal!",
      "votes": null
    },
    {
      "id": "1696931",
      "postDate": "02/19/2022 08:17:20",
      "content": "<p>Thanks! Good kaggle!</p>",
      "rawMarkdown": "Thanks! Good kaggle!",
      "votes": null
    },
    {
      "id": "1698768",
      "postDate": "02/20/2022 16:34:02",
      "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Thanks for sharing your solution and congratulations to your team on the silver!</p>\n<p>I didn't understand the point shared by you here, could you kindly clarify what you mean by:</p>\n<blockquote>\n  <p>we found out that Increasing image size increased recall but also yielded unnecessary bounding boxes.</p>\n</blockquote>\n<p>TIA! :) </p>",
      "rawMarkdown": "deepkim Thanks for sharing your solution and congratulations to your team on the silver!\n\nI didn't understand the point shared by you here, could you kindly clarify what you mean by:\n\n> we found out that Increasing image size increased recall but also yielded unnecessary bounding boxes.\n\nTIA! :)",
      "votes": null
    },
    {
      "id": "1698806",
      "postDate": "02/20/2022 17:19:38",
      "content": "<p>I mean increasing image size increases <strong>TRUE POSITIVES</strong> but it also increases many many <strong>FALSE POSITIVES</strong>.<br>\nif <strong>FALSE POSITIVES</strong> increases a lot, then It swallows the merit of increasing <strong>TRUE POSITIVES</strong> and results in a drop of F2 SCORE.</p>",
      "rawMarkdown": "I mean increasing image size increases **TRUE POSITIVES** but it also increases many many **FALSE POSITIVES**.\nif **FALSE POSITIVES** increases a lot, then It swallows the merit of increasing **TRUE POSITIVES** and results in a drop of F2 SCORE.",
      "votes": null
    },
    {
      "id": "1699184",
      "postDate": "02/21/2022 01:54:24",
      "content": "<p>Many Thanks for clarifying! 🙏</p>",
      "rawMarkdown": "Many Thanks for clarifying! 🙏",
      "votes": null
    },
    {
      "id": "1703217",
      "postDate": "02/24/2022 09:56:08",
      "content": "<p>Congratulate for taking the 16th and thanks for sharing your solution. I wonder why you trained the yolov5s with large sizes when the yolov5m trained with smaller one.</p>",
      "rawMarkdown": "Congratulate for taking the 16th and thanks for sharing your solution. I wonder why you trained the yolov5s with large sizes when the yolov5m trained with smaller one.",
      "votes": null
    },
    {
      "id": "1706639",
      "postDate": "02/27/2022 17:12:50",
      "content": "<p><a href=\"https://www.kaggle.com/jarviskevin\" target=\"_blank\">@jarviskevin</a> It's from experiments. or to maximize variabilities. Relatively yolov5m was good at a smaller size but yolov5s was good at a larger size. also, performance was almost the same between yolov5m and yolov5s</p>",
      "rawMarkdown": "jarviskevin It's from experiments. or to maximize variabilities. Relatively yolov5m was good at a smaller size but yolov5s was good at a larger size. also, performance was almost the same between yolov5m and yolov5s",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1695880,
      "author_name": "mutsumichieda",
      "author_url": "",
      "post_date": "02/18/2022 12:28:38",
      "content": "<blockquote>\n  <p>In CV, only inference with originally trained size worked.</p>\n</blockquote>\n<p>This line explains why most of public inference codes didn't work well.<br>\nAfter competition, I found that our small models (which have training image size = inference image size) work as good as high score public codes in private dataset.</p>\n<p>I wish I checked CV-LB correlation with image size.<br>\nCongrats for taking 16th!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1695882,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "02/18/2022 12:32:31",
          "content": "<p>I also checked the same situation.<br>\nOur yoloX (1280,720 size) 0.399 in public LB but 0.597 in private LB<br>\nI learned again that we should verify cv-lb very carefully!<br>\nThank you for reading this post!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1696015,
      "author_name": "ttkagglett",
      "author_url": "",
      "post_date": "02/18/2022 14:41:39",
      "content": "<p>Great to know about ensemble with multi-scale images. Thanks and congrats on silver medal!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1696931,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "02/19/2022 08:17:20",
          "content": "<p>Thanks! Good kaggle!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1698768,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/20/2022 16:34:02",
      "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Thanks for sharing your solution and congratulations to your team on the silver!</p>\n<p>I didn't understand the point shared by you here, could you kindly clarify what you mean by:</p>\n<blockquote>\n  <p>we found out that Increasing image size increased recall but also yielded unnecessary bounding boxes.</p>\n</blockquote>\n<p>TIA! :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1698806,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "02/20/2022 17:19:38",
          "content": "<p>I mean increasing image size increases <strong>TRUE POSITIVES</strong> but it also increases many many <strong>FALSE POSITIVES</strong>.<br>\nif <strong>FALSE POSITIVES</strong> increases a lot, then It swallows the merit of increasing <strong>TRUE POSITIVES</strong> and results in a drop of F2 SCORE.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699184,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/21/2022 01:54:24",
          "content": "<p>Many Thanks for clarifying! 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1703217,
      "author_name": "jarviskevin",
      "author_url": "",
      "post_date": "02/24/2022 09:56:08",
      "content": "<p>Congratulate for taking the 16th and thanks for sharing your solution. I wonder why you trained the yolov5s with large sizes when the yolov5m trained with smaller one.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1706639,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "02/27/2022 17:12:50",
          "content": "<p><a href=\"https://www.kaggle.com/jarviskevin\" target=\"_blank\">@jarviskevin</a> It's from experiments. or to maximize variabilities. Relatively yolov5m was good at a smaller size but yolov5s was good at a larger size. also, performance was almost the same between yolov5m and yolov5s</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1695523": "I would like to thank the organizers for hosting the great competition.\nalso, I was very honored to become a teammate with @vaillant and @hwigeon. \nI really learned a lot from our teammates. I think it will be the unforgettable good experience.\n\n# Summary\nWe reached LB 709 so quickly using a large-size inference trick and naive iou-tracker.\nbut at the same time, we started to doubt the public LB since we couldn't reproduce the result clearly. we found out that Increasing image size increased recall but also yielded unnecessary bounding boxes.\nSo, we started to build robust models by checking CV F2 scores.\nIn CV, only inference with originally trained size worked.\nWe tried size threshold-based prediction, IOU tracker(i might have made a mistake), larger size than originally trained size didn't work in the local.\n\n# Models\nWe found out that a large model necessarily isn't needed.\nAlso what we found out is multi-scale ensemble boosts both cv and public LB scores.\nInstead of investing our time in searching complex models, we prepared yolov5m trained with the sizes of [2400,2560,2688,2880,3000]\nand yolov5s trained with the sizes of [3000,3600,3720,3840,3960,4032,4080,4244]\n\n# Traning Method\nWe changed yolo's hyperparameters slightly\nApplying mixup 0.5, mosaic 1.0, strong hsv change helped a lot\n10 epoch training and adam optimizer were chosen.\nModel was chosen by weighting Recall vs Precision  4:1\n\n# Validation\nWe chose video-split folds since these are more realistic.\nEnsembled CV scores for video-split were each [0.71, 0.64, 0.77]\nAfter that, we had searched optimized WBF coefficients and confidence scores using a method of grid search.\n\nFor video fold 0-1, WBF coefficient 0.5, skip bbox 0.01, confidence score 0.1 was best\nFor video fold 2, WBF coefficient 0.5, skip_bbox 0.1, confidence score 0.2 was best\n\nWith these values, we trained models with all the datasets and regarded the public LB as a holdout\nthe public LB scores of each model were ranged in [0.55~0.60]\n\nThe result is private LB 0.713 and public LB 0.648.\nWe missed the gold but we couldn't choose the best private LB since this competition is a kind of shake-up competition\n\nThank you for reading this!",
    "1695880": "> In CV, only inference with originally trained size worked.\n\nThis line explains why most of public inference codes didn't work well.\nAfter competition, I found that our small models (which have training image size = inference image size) work as good as high score public codes in private dataset.\n\nI wish I checked CV-LB correlation with image size.\nCongrats for taking 16th!",
    "1695882": "I also checked the same situation.\nOur yoloX (1280,720 size) 0.399 in public LB but 0.597 in private LB\nI learned again that we should verify cv-lb very carefully!\nThank you for reading this post!",
    "1696015": "Great to know about ensemble with multi-scale images. Thanks and congrats on silver medal!",
    "1696931": "Thanks! Good kaggle!",
    "1698768": "deepkim Thanks for sharing your solution and congratulations to your team on the silver!\n\nI didn't understand the point shared by you here, could you kindly clarify what you mean by:\n\n> we found out that Increasing image size increased recall but also yielded unnecessary bounding boxes.\n\nTIA! :)",
    "1698806": "I mean increasing image size increases **TRUE POSITIVES** but it also increases many many **FALSE POSITIVES**.\nif **FALSE POSITIVES** increases a lot, then It swallows the merit of increasing **TRUE POSITIVES** and results in a drop of F2 SCORE.",
    "1699184": "Many Thanks for clarifying! 🙏",
    "1703217": "Congratulate for taking the 16th and thanks for sharing your solution. I wonder why you trained the yolov5s with large sizes when the yolov5m trained with smaller one.",
    "1706639": "jarviskevin It's from experiments. or to maximize variabilities. Relatively yolov5m was good at a smaller size but yolov5s was good at a larger size. also, performance was almost the same between yolov5m and yolov5s"
  },
  "source": "meta"
}