{
  "id": 307691,
  "title": "15th place solution: YOLO-X only, Seq-NMS",
  "url": "/competitions/tensorflow-great-barrier-reef/writeups/maxwell-15th-place-solution-yolo-x-only-seq-nms",
  "author_name": "",
  "post_date": "2022-02-18T05:29:31.393Z",
  "votes": 33,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congratulations to the winners and all participants who finished this competition 😊<br>\nIn this competition, many participants were focused on increasing the score of Public LB, and I was not able to keep up with Public LB at least halfway through.<br>\nHowever, in my local experiments, I have found that methods that increase the image size only during inference, even when the threshold and ensemble methods are properly adjusted, result in extremely low scores for the entire training data.<br>\nSo, I thought there would be shake. Although I was not able to get into the gold zone, I am glad that I participated because I was able to recognize the necessity of believing in my own experiments.</p>\n<p>Ok, my solution is a pretty simple one, but I would like to share.</p>\n<p><strong>1. Preprocessing</strong><br>\nSince the video was basically switching almost every sequence, I built a validation with 4 folds based on the sequence (However, I think there were some sequences that were a little continuous).<br>\nOne coin I did was to make the number of CoTS almost the same in each fold. In training, I did not use images without CoTS, but in validation, I used all images without CoTS in order to check the False Positive properly.<br>\n  The image was simply divided by 255 without any noise removal, and the size was 1952 x 3520.<br>\nDue to GPU limitations, I couldn't try a larger size, but at least up to this size, I was able to get a small gain with Local CV. However I got the most gain up to about twice the size of the original image.</p>\n<p><strong>2. Model: All I want to use is YOLO-X</strong><br>\nThis time I really wanted to use <code>anchor-free</code> YOLO-X, which I had never used before, so I stuck to it, even though many people have had success with YOLO-v5 in Public LB.<br>\n  It was hard to modify YOLO-X because many parts were hard-coded, but I tried to tweak the <a href=\"https://arxiv.org/pdf/2107.08430.pdf\" target=\"_blank\">top-K selection algorithm</a>, add augmentation, and some other dubious modifications. But in the end, only augmentation seemed to have an effect. In that sense, YOLO-X's perfection as a model may be high 😏</p>\n<p><strong>3. Augmentation and learning strategy</strong><br>\nIn addition to the augmentation (mosaic, mixup, etc.) that YOLO-X has as a default function, I added RandomGamma, RGBshift, Sharpen, GaussNoise, etc.<br>\nInstead, the probability of applying mixup and mosaic has been slightly reduced from the default value and assigned the rest probability to the newly added augmentation path. This may have increased the diversity of the input images somewhat, and I was able to get a gain in Local CV(I don't know the specific ablation values in detail, as the experiment took some twists and turns).<br>\n  In addition, I used the <a href=\"https://arxiv.org/abs/2104.00298\" target=\"_blank\">progressive learning</a> method used in EfficientNetV2: gradually increasing the size of the image (e.g. 1280 =&gt; … =&gt; 3520) as the learning progressed. At the same time, I remember that regularization (increasing the probability of application of augmentation) was also strengthened somewhat.</p>\n<p><strong>4. Inference</strong><br>\nThe inference was very simple: I did TTA with Flip and ensembled with WBF for 4-fold (so 8 models). I think this is almost same as many of the participants.<br>\nNote that I set thresholds to improve the local CV. The F2 metrics were sensitive to the threshold settings to some extent because our evaluation metric was based on the confusion matrix.</p>\n<p><strong>5. An original point that I have worked out: Seq-NMS</strong><br>\nOne point that I devised a little is the post-processing.<br>\n  Since the task is object recognition in video, I made a post in <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812\" target=\"_blank\">this thread</a> early stage in this competition, and I was reading almost all papers. Many of them reported that feature aggregation, which is a method of enriching input features by using past images in the neighborhood, is more performing than post-processing methods such as tracking, and I thought that was probably true. However, since the majority of feature aggregation methods is based on RPNs, it was too much of a hurdle for me to apply, since I was sticking to YOLO-X, which is anchor-free. So I decided to use <a href=\"https://arxiv.org/abs/1602.08465\" target=\"_blank\">Seq-NMS</a>: a typical post-processing method. Specifically, the confidence is adjusted by the degree of overlap between the predicted bbox of the previous images and the predicted bbox of the current image (Ref: <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812#1611428)\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812#1611428)</a>.<br>\n(The implementation can be found <a href=\"https://github.com/tmoopenn/seq-nms\" target=\"_blank\">here</a>. The core processing is written in cython and is fast enough, but as far as I understand, there is a serious bug in the python script that makes it not work properly as is, so I fixed the bug to use it.)<br>\n  As long as I experimented, there is no need to consider many frames for Seq-NMS, just the prediction of the previous image. I also tried to use <a href=\"https://docs.opencv.org/3.4/d4/dee/tutorial_optical_flow.html#:~:text=Dense%20Optical%20Flow%20in%20OpenCV\" target=\"_blank\">Dense Optical Flow</a> to compensate for the shift of predicted bboxes of the previous frame, but in the end I did not use it because Dense Optical Flow is expensive to compute and does not have much gain.</p>\n<hr>\n<p>That's all I can remember right now. If I remember anything else, I'll add it.<br>\nThank you all for your hard work, and see you at another competition :)</p>\n<p>Happy Kaggling😊</p>\n<p><img src=\"https://i.imgur.com/fuUtT1i.jpg\" alt=\"\"></p>\n<p>Feb. 18. 2022, Model pipeline figure updated</p>",
  "messages": [
    {
      "id": "1691165",
      "postDate": "02/15/2022 09:08:03",
      "content": "<p>Congratulations to the winners and all participants who finished this competition 😊<br>\nIn this competition, many participants were focused on increasing the score of Public LB, and I was not able to keep up with Public LB at least halfway through.<br>\nHowever, in my local experiments, I have found that methods that increase the image size only during inference, even when the threshold and ensemble methods are properly adjusted, result in extremely low scores for the entire training data.<br>\nSo, I thought there would be shake. Although I was not able to get into the gold zone, I am glad that I participated because I was able to recognize the necessity of believing in my own experiments.</p>\n<p>Ok, my solution is a pretty simple one, but I would like to share.</p>\n<p><strong>1. Preprocessing</strong><br>\nSince the video was basically switching almost every sequence, I built a validation with 4 folds based on the sequence (However, I think there were some sequences that were a little continuous).<br>\nOne coin I did was to make the number of CoTS almost the same in each fold. In training, I did not use images without CoTS, but in validation, I used all images without CoTS in order to check the False Positive properly.<br>\n  The image was simply divided by 255 without any noise removal, and the size was 1952 x 3520.<br>\nDue to GPU limitations, I couldn't try a larger size, but at least up to this size, I was able to get a small gain with Local CV. However I got the most gain up to about twice the size of the original image.</p>\n<p><strong>2. Model: All I want to use is YOLO-X</strong><br>\nThis time I really wanted to use <code>anchor-free</code> YOLO-X, which I had never used before, so I stuck to it, even though many people have had success with YOLO-v5 in Public LB.<br>\n  It was hard to modify YOLO-X because many parts were hard-coded, but I tried to tweak the <a href=\"https://arxiv.org/pdf/2107.08430.pdf\" target=\"_blank\">top-K selection algorithm</a>, add augmentation, and some other dubious modifications. But in the end, only augmentation seemed to have an effect. In that sense, YOLO-X's perfection as a model may be high 😏</p>\n<p><strong>3. Augmentation and learning strategy</strong><br>\nIn addition to the augmentation (mosaic, mixup, etc.) that YOLO-X has as a default function, I added RandomGamma, RGBshift, Sharpen, GaussNoise, etc.<br>\nInstead, the probability of applying mixup and mosaic has been slightly reduced from the default value and assigned the rest probability to the newly added augmentation path. This may have increased the diversity of the input images somewhat, and I was able to get a gain in Local CV(I don't know the specific ablation values in detail, as the experiment took some twists and turns).<br>\n  In addition, I used the <a href=\"https://arxiv.org/abs/2104.00298\" target=\"_blank\">progressive learning</a> method used in EfficientNetV2: gradually increasing the size of the image (e.g. 1280 =&gt; … =&gt; 3520) as the learning progressed. At the same time, I remember that regularization (increasing the probability of application of augmentation) was also strengthened somewhat.</p>\n<p><strong>4. Inference</strong><br>\nThe inference was very simple: I did TTA with Flip and ensembled with WBF for 4-fold (so 8 models). I think this is almost same as many of the participants.<br>\nNote that I set thresholds to improve the local CV. The F2 metrics were sensitive to the threshold settings to some extent because our evaluation metric was based on the confusion matrix.</p>\n<p><strong>5. An original point that I have worked out: Seq-NMS</strong><br>\nOne point that I devised a little is the post-processing.<br>\n  Since the task is object recognition in video, I made a post in <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812\" target=\"_blank\">this thread</a> early stage in this competition, and I was reading almost all papers. Many of them reported that feature aggregation, which is a method of enriching input features by using past images in the neighborhood, is more performing than post-processing methods such as tracking, and I thought that was probably true. However, since the majority of feature aggregation methods is based on RPNs, it was too much of a hurdle for me to apply, since I was sticking to YOLO-X, which is anchor-free. So I decided to use <a href=\"https://arxiv.org/abs/1602.08465\" target=\"_blank\">Seq-NMS</a>: a typical post-processing method. Specifically, the confidence is adjusted by the degree of overlap between the predicted bbox of the previous images and the predicted bbox of the current image (Ref: <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812#1611428)\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812#1611428)</a>.<br>\n(The implementation can be found <a href=\"https://github.com/tmoopenn/seq-nms\" target=\"_blank\">here</a>. The core processing is written in cython and is fast enough, but as far as I understand, there is a serious bug in the python script that makes it not work properly as is, so I fixed the bug to use it.)<br>\n  As long as I experimented, there is no need to consider many frames for Seq-NMS, just the prediction of the previous image. I also tried to use <a href=\"https://docs.opencv.org/3.4/d4/dee/tutorial_optical_flow.html#:~:text=Dense%20Optical%20Flow%20in%20OpenCV\" target=\"_blank\">Dense Optical Flow</a> to compensate for the shift of predicted bboxes of the previous frame, but in the end I did not use it because Dense Optical Flow is expensive to compute and does not have much gain.</p>\n<hr>\n<p>That's all I can remember right now. If I remember anything else, I'll add it.<br>\nThank you all for your hard work, and see you at another competition :)</p>\n<p>Happy Kaggling😊</p>\n<p><img src=\"https://i.imgur.com/fuUtT1i.jpg\" alt=\"\"></p>\n<p>Feb. 18. 2022, Model pipeline figure updated</p>",
      "rawMarkdown": "Congratulations to the winners and all participants who finished this competition 😊\nIn this competition, many participants were focused on increasing the score of Public LB, and I was not able to keep up with Public LB at least halfway through.\nHowever, in my local experiments, I have found that methods that increase the image size only during inference, even when the threshold and ensemble methods are properly adjusted, result in extremely low scores for the entire training data.\nSo, I thought there would be shake. Although I was not able to get into the gold zone, I am glad that I participated because I was able to recognize the necessity of believing in my own experiments.\n\nOk, my solution is a pretty simple one, but I would like to share.\n\n\n**1. Preprocessing**\nSince the video was basically switching almost every sequence, I built a validation with 4 folds based on the sequence (However, I think there were some sequences that were a little continuous).\nOne coin I did was to make the number of CoTS almost the same in each fold. In training, I did not use images without CoTS, but in validation, I used all images without CoTS in order to check the False Positive properly.\n  The image was simply divided by 255 without any noise removal, and the size was 1952 x 3520.\nDue to GPU limitations, I couldn't try a larger size, but at least up to this size, I was able to get a small gain with Local CV. However I got the most gain up to about twice the size of the original image.\n\n\n**2. Model: All I want to use is YOLO-X**\nThis time I really wanted to use `anchor-free` YOLO-X, which I had never used before, so I stuck to it, even though many people have had success with YOLO-v5 in Public LB.\n  It was hard to modify YOLO-X because many parts were hard-coded, but I tried to tweak the [top-K selection algorithm](https://arxiv.org/pdf/2107.08430.pdf), add augmentation, and some other dubious modifications. But in the end, only augmentation seemed to have an effect. In that sense, YOLO-X's perfection as a model may be high 😏\n\n\n**3. Augmentation and learning strategy**\nIn addition to the augmentation (mosaic, mixup, etc.) that YOLO-X has as a default function, I added RandomGamma, RGBshift, Sharpen, GaussNoise, etc.\nInstead, the probability of applying mixup and mosaic has been slightly reduced from the default value and assigned the rest probability to the newly added augmentation path. This may have increased the diversity of the input images somewhat, and I was able to get a gain in Local CV(I don't know the specific ablation values in detail, as the experiment took some twists and turns).\n  In addition, I used the [progressive learning](https://arxiv.org/abs/2104.00298) method used in EfficientNetV2: gradually increasing the size of the image (e.g. 1280 => ... => 3520) as the learning progressed. At the same time, I remember that regularization (increasing the probability of application of augmentation) was also strengthened somewhat.\n\n\n**4. Inference**\nThe inference was very simple: I did TTA with Flip and ensembled with WBF for 4-fold (so 8 models). I think this is almost same as many of the participants.\nNote that I set thresholds to improve the local CV. The F2 metrics were sensitive to the threshold settings to some extent because our evaluation metric was based on the confusion matrix.\n\n\n**5. An original point that I have worked out: Seq-NMS**\nOne point that I devised a little is the post-processing.\n  Since the task is object recognition in video, I made a post in [this thread](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812) early stage in this competition, and I was reading almost all papers. Many of them reported that feature aggregation, which is a method of enriching input features by using past images in the neighborhood, is more performing than post-processing methods such as tracking, and I thought that was probably true. However, since the majority of feature aggregation methods is based on RPNs, it was too much of a hurdle for me to apply, since I was sticking to YOLO-X, which is anchor-free. So I decided to use [Seq-NMS](https://arxiv.org/abs/1602.08465): a typical post-processing method. Specifically, the confidence is adjusted by the degree of overlap between the predicted bbox of the previous images and the predicted bbox of the current image (Ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812#1611428).\n(The implementation can be found [here](https://github.com/tmoopenn/seq-nms). The core processing is written in cython and is fast enough, but as far as I understand, there is a serious bug in the python script that makes it not work properly as is, so I fixed the bug to use it.)\n  As long as I experimented, there is no need to consider many frames for Seq-NMS, just the prediction of the previous image. I also tried to use [Dense Optical Flow](https://docs.opencv.org/3.4/d4/dee/tutorial_optical_flow.html#:~:text=Dense%20Optical%20Flow%20in%20OpenCV) to compensate for the shift of predicted bboxes of the previous frame, but in the end I did not use it because Dense Optical Flow is expensive to compute and does not have much gain.\n\n---\n\nThat's all I can remember right now. If I remember anything else, I'll add it.\nThank you all for your hard work, and see you at another competition :)\n\nHappy Kaggling😊\n\n\n![](https://i.imgur.com/fuUtT1i.jpg)\n\nFeb. 18. 2022, Model pipeline figure updated",
      "votes": null
    },
    {
      "id": "1691177",
      "postDate": "02/15/2022 09:16:35",
      "content": "<p>Congratulations! Yolo-X 👍👍💪😄😍😍😍 Great!</p>",
      "rawMarkdown": "Congratulations! Yolo-X 👍👍💪😄😍😍😍 Great!",
      "votes": null
    },
    {
      "id": "1691187",
      "postDate": "02/15/2022 09:27:45",
      "content": "<p>Congratulations! How much improvement has Seq-NMS brought you?</p>",
      "rawMarkdown": "Congratulations! How much improvement has Seq-NMS brought you?",
      "votes": null
    },
    {
      "id": "1691210",
      "postDate": "02/15/2022 09:48:53",
      "content": "<p>Thank you. Congratulations, too!</p>\n<blockquote>\n  <p>Q. How much improvement has Seq-NMS brought you?</p>\n</blockquote>\n<p>A. about 0.015. I only have my experiments to base this on, but I think it's a little better than the tracking method that was shared on kernel.</p>\n<p>By the way, since you seem to like memes, I'll give you this image.<br>\n<img src=\"https://i.imgur.com/Tg7MjJS.jpg\" alt=\"\"></p>",
      "rawMarkdown": "Thank you. Congratulations, too!\n\n> Q. How much improvement has Seq-NMS brought you?\n\nA. about 0.015. I only have my experiments to base this on, but I think it's a little better than the tracking method that was shared on kernel.\n\nBy the way, since you seem to like memes, I'll give you this image.\n![](https://i.imgur.com/Tg7MjJS.jpg)",
      "votes": null
    },
    {
      "id": "1691241",
      "postDate": "02/15/2022 10:05:54",
      "content": "<p>lol, thanks for your reply</p>",
      "rawMarkdown": "lol, thanks for your reply",
      "votes": null
    },
    {
      "id": "1691587",
      "postDate": "02/15/2022 13:43:08",
      "content": "<p>I think the simple solution is the better way in real-world problem. Thanks for sharing and congrats on 16th!</p>",
      "rawMarkdown": "I think the simple solution is the better way in real-world problem. Thanks for sharing and congrats on 16th!",
      "votes": null
    },
    {
      "id": "1693022",
      "postDate": "02/16/2022 11:54:48",
      "content": "<p>Thanks for this solution mate, I'm glad that YoloX is able to make it up there. I had used YoloX myself but got hard stucked at 0.4X, I'm pretty new to kaggle so no suprise there lol. Seems high training resolution gave the largest boost? at least from the 0.4X perspective. Due to GPU limitation was only able to train on 800ish resolution.</p>\n<p>Anyway, congrats on the 16th, especially with YoloX 🎉🎉.</p>",
      "rawMarkdown": "Thanks for this solution mate, I'm glad that YoloX is able to make it up there. I had used YoloX myself but got hard stucked at 0.4X, I'm pretty new to kaggle so no suprise there lol. Seems high training resolution gave the largest boost? at least from the 0.4X perspective. Due to GPU limitation was only able to train on 800ish resolution.\n\nAnyway, congrats on the 16th, especially with YoloX 🎉🎉.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1691177,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/15/2022 09:16:35",
      "content": "<p>Congratulations! Yolo-X 👍👍💪😄😍😍😍 Great!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1691187,
      "author_name": "zzy990106",
      "author_url": "",
      "post_date": "02/15/2022 09:27:45",
      "content": "<p>Congratulations! How much improvement has Seq-NMS brought you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1691210,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "02/15/2022 09:48:53",
          "content": "<p>Thank you. Congratulations, too!</p>\n<blockquote>\n  <p>Q. How much improvement has Seq-NMS brought you?</p>\n</blockquote>\n<p>A. about 0.015. I only have my experiments to base this on, but I think it's a little better than the tracking method that was shared on kernel.</p>\n<p>By the way, since you seem to like memes, I'll give you this image.<br>\n<img src=\"https://i.imgur.com/Tg7MjJS.jpg\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691241,
          "author_name": "zzy990106",
          "author_url": "",
          "post_date": "02/15/2022 10:05:54",
          "content": "<p>lol, thanks for your reply</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691587,
      "author_name": "ttkagglett",
      "author_url": "",
      "post_date": "02/15/2022 13:43:08",
      "content": "<p>I think the simple solution is the better way in real-world problem. Thanks for sharing and congrats on 16th!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1693022,
      "author_name": "zhenhuazhang",
      "author_url": "",
      "post_date": "02/16/2022 11:54:48",
      "content": "<p>Thanks for this solution mate, I'm glad that YoloX is able to make it up there. I had used YoloX myself but got hard stucked at 0.4X, I'm pretty new to kaggle so no suprise there lol. Seems high training resolution gave the largest boost? at least from the 0.4X perspective. Due to GPU limitation was only able to train on 800ish resolution.</p>\n<p>Anyway, congrats on the 16th, especially with YoloX 🎉🎉.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1691165": "Congratulations to the winners and all participants who finished this competition 😊\nIn this competition, many participants were focused on increasing the score of Public LB, and I was not able to keep up with Public LB at least halfway through.\nHowever, in my local experiments, I have found that methods that increase the image size only during inference, even when the threshold and ensemble methods are properly adjusted, result in extremely low scores for the entire training data.\nSo, I thought there would be shake. Although I was not able to get into the gold zone, I am glad that I participated because I was able to recognize the necessity of believing in my own experiments.\n\nOk, my solution is a pretty simple one, but I would like to share.\n\n\n**1. Preprocessing**\nSince the video was basically switching almost every sequence, I built a validation with 4 folds based on the sequence (However, I think there were some sequences that were a little continuous).\nOne coin I did was to make the number of CoTS almost the same in each fold. In training, I did not use images without CoTS, but in validation, I used all images without CoTS in order to check the False Positive properly.\n  The image was simply divided by 255 without any noise removal, and the size was 1952 x 3520.\nDue to GPU limitations, I couldn't try a larger size, but at least up to this size, I was able to get a small gain with Local CV. However I got the most gain up to about twice the size of the original image.\n\n\n**2. Model: All I want to use is YOLO-X**\nThis time I really wanted to use `anchor-free` YOLO-X, which I had never used before, so I stuck to it, even though many people have had success with YOLO-v5 in Public LB.\n  It was hard to modify YOLO-X because many parts were hard-coded, but I tried to tweak the [top-K selection algorithm](https://arxiv.org/pdf/2107.08430.pdf), add augmentation, and some other dubious modifications. But in the end, only augmentation seemed to have an effect. In that sense, YOLO-X's perfection as a model may be high 😏\n\n\n**3. Augmentation and learning strategy**\nIn addition to the augmentation (mosaic, mixup, etc.) that YOLO-X has as a default function, I added RandomGamma, RGBshift, Sharpen, GaussNoise, etc.\nInstead, the probability of applying mixup and mosaic has been slightly reduced from the default value and assigned the rest probability to the newly added augmentation path. This may have increased the diversity of the input images somewhat, and I was able to get a gain in Local CV(I don't know the specific ablation values in detail, as the experiment took some twists and turns).\n  In addition, I used the [progressive learning](https://arxiv.org/abs/2104.00298) method used in EfficientNetV2: gradually increasing the size of the image (e.g. 1280 => ... => 3520) as the learning progressed. At the same time, I remember that regularization (increasing the probability of application of augmentation) was also strengthened somewhat.\n\n\n**4. Inference**\nThe inference was very simple: I did TTA with Flip and ensembled with WBF for 4-fold (so 8 models). I think this is almost same as many of the participants.\nNote that I set thresholds to improve the local CV. The F2 metrics were sensitive to the threshold settings to some extent because our evaluation metric was based on the confusion matrix.\n\n\n**5. An original point that I have worked out: Seq-NMS**\nOne point that I devised a little is the post-processing.\n  Since the task is object recognition in video, I made a post in [this thread](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812) early stage in this competition, and I was reading almost all papers. Many of them reported that feature aggregation, which is a method of enriching input features by using past images in the neighborhood, is more performing than post-processing methods such as tracking, and I thought that was probably true. However, since the majority of feature aggregation methods is based on RPNs, it was too much of a hurdle for me to apply, since I was sticking to YOLO-X, which is anchor-free. So I decided to use [Seq-NMS](https://arxiv.org/abs/1602.08465): a typical post-processing method. Specifically, the confidence is adjusted by the degree of overlap between the predicted bbox of the previous images and the predicted bbox of the current image (Ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293812#1611428).\n(The implementation can be found [here](https://github.com/tmoopenn/seq-nms). The core processing is written in cython and is fast enough, but as far as I understand, there is a serious bug in the python script that makes it not work properly as is, so I fixed the bug to use it.)\n  As long as I experimented, there is no need to consider many frames for Seq-NMS, just the prediction of the previous image. I also tried to use [Dense Optical Flow](https://docs.opencv.org/3.4/d4/dee/tutorial_optical_flow.html#:~:text=Dense%20Optical%20Flow%20in%20OpenCV) to compensate for the shift of predicted bboxes of the previous frame, but in the end I did not use it because Dense Optical Flow is expensive to compute and does not have much gain.\n\n---\n\nThat's all I can remember right now. If I remember anything else, I'll add it.\nThank you all for your hard work, and see you at another competition :)\n\nHappy Kaggling😊\n\n\n![](https://i.imgur.com/fuUtT1i.jpg)\n\nFeb. 18. 2022, Model pipeline figure updated",
    "1691177": "Congratulations! Yolo-X 👍👍💪😄😍😍😍 Great!",
    "1691187": "Congratulations! How much improvement has Seq-NMS brought you?",
    "1691210": "Thank you. Congratulations, too!\n\n> Q. How much improvement has Seq-NMS brought you?\n\nA. about 0.015. I only have my experiments to base this on, but I think it's a little better than the tracking method that was shared on kernel.\n\nBy the way, since you seem to like memes, I'll give you this image.\n![](https://i.imgur.com/Tg7MjJS.jpg)",
    "1691241": "lol, thanks for your reply",
    "1691587": "I think the simple solution is the better way in real-world problem. Thanks for sharing and congrats on 16th!",
    "1693022": "Thanks for this solution mate, I'm glad that YoloX is able to make it up there. I had used YoloX myself but got hard stucked at 0.4X, I'm pretty new to kaggle so no suprise there lol. Seems high training resolution gave the largest boost? at least from the 0.4X perspective. Due to GPU limitation was only able to train on 800ish resolution.\n\nAnyway, congrats on the 16th, especially with YoloX 🎉🎉."
  },
  "source": "meta"
}