{
  "id": 293812,
  "title": "Papers on Video Object Detection",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/293812",
  "author_name": "",
  "post_date": "2021-12-07T06:01:19.146242700Z",
  "votes": 75,
  "comment_count": 17,
  "views": 0,
  "content": "<p>At the moment (Dec.07.2021), many discussions seem to be focusing on object detection in still images, as represented by augmentation and detection models.<br>\nHowever, as anyone who has done or seen EDA will be aware, I think the major key to this competition is that the images are sequence (video). In the case of video detection, many methods have been proposed that make good use of continuous images, i.e., images that are close in time, because they contain information that can reduce detection errors (FP and FN).</p>\n<p>So, I would like to list some of the papers I found by exploring <a href=\"https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid\" target=\"_blank\">the ImageNet VID benchmarks in Paper with Code</a>. Although there are some old papers, I think that understanding the origin of the idea of the proposed methods will help us to come up with an original method for this competition. In addition, the post-processing method can be applied directly to models trained on still images, so it will be of great help to those who are currently considering using models based on still images.</p>\n<p>If you are interested in reading them and would like to post a summary, please feel free to do so in this thread (or in other threads).</p>\n<hr>\n<p>Paper with Code ImageNet VID: <br>\n<a href=\"https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid\" target=\"_blank\">https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid</a></p>\n<ul>\n<li><p>Seq-NMS for Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/1602.08465\" target=\"_blank\">https://arxiv.org/abs/1602.08465</a></p></li>\n<li><p>Deep Feature Flow for Video Recognition<br>\n<a href=\"https://arxiv.org/abs/1611.07715\" target=\"_blank\">https://arxiv.org/abs/1611.07715</a></p></li>\n<li><p>T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos<br>\n<a href=\"https://arxiv.org/abs/1604.02532\" target=\"_blank\">https://arxiv.org/abs/1604.02532</a></p></li>\n<li><p>Object Detection in Videos with Tubelet Proposal Networks<br>\n<a href=\"https://arxiv.org/abs/1702.06355\" target=\"_blank\">https://arxiv.org/abs/1702.06355</a></p></li>\n<li><p>Flow-Guided Feature Aggregation for Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/1703.10025\" target=\"_blank\">https://arxiv.org/abs/1703.10025</a></p></li>\n<li><p>Detect to Track and Track to Detect<br>\n<a href=\"https://arxiv.org/abs/1710.03958\" target=\"_blank\">https://arxiv.org/abs/1710.03958</a></p></li>\n<li><p>Towards High Performance Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/1711.11577\" target=\"_blank\">https://arxiv.org/abs/1711.11577</a></p></li>\n<li><p>Mobile Video Object Detection with Temporally-Aware Feature Maps<br>\n<a href=\"https://arxiv.org/abs/1711.06368\" target=\"_blank\">https://arxiv.org/abs/1711.06368</a></p></li>\n<li><p>Video Object Detection with an Aligned Spatial-Temporal Memory<br>\n<a href=\"https://arxiv.org/abs/1712.06317\" target=\"_blank\">https://arxiv.org/abs/1712.06317</a></p></li>\n<li><p>Object Detection in Video with Spatiotemporal Sampling Networks<br>\n<a href=\"https://arxiv.org/abs/1803.05549\" target=\"_blank\">https://arxiv.org/abs/1803.05549</a></p></li>\n<li><p>Sequence Level Semantics Aggregation for Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/1907.06390v2\" target=\"_blank\">https://arxiv.org/abs/1907.06390v2</a></p></li>\n<li><p>Memory Enhanced Global-Local Aggregation for Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/2003.12063v1\" target=\"_blank\">https://arxiv.org/abs/2003.12063v1</a></p></li>\n<li><p>Robust and efficient post-processing for video object detection<br>\n<a href=\"https://arxiv.org/abs/2009.11050\" target=\"_blank\">https://arxiv.org/abs/2009.11050</a></p></li>\n<li><p>Mining Inter-Video Proposal Relations for Video Object Detection<br>\n<a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/html/3764_ECCV_2020_paper.php\" target=\"_blank\">https://www.ecva.net/papers/eccv_2020/papers_ECCV/html/3764_ECCV_2020_paper.php</a></p></li>\n</ul>\n<hr>\n<p>Happy Kaggling ✨</p>",
  "messages": [
    {
      "id": "1610305",
      "postDate": "12/07/2021 06:01:19",
      "content": "<p>At the moment (Dec.07.2021), many discussions seem to be focusing on object detection in still images, as represented by augmentation and detection models.<br>\nHowever, as anyone who has done or seen EDA will be aware, I think the major key to this competition is that the images are sequence (video). In the case of video detection, many methods have been proposed that make good use of continuous images, i.e., images that are close in time, because they contain information that can reduce detection errors (FP and FN).</p>\n<p>So, I would like to list some of the papers I found by exploring <a href=\"https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid\" target=\"_blank\">the ImageNet VID benchmarks in Paper with Code</a>. Although there are some old papers, I think that understanding the origin of the idea of the proposed methods will help us to come up with an original method for this competition. In addition, the post-processing method can be applied directly to models trained on still images, so it will be of great help to those who are currently considering using models based on still images.</p>\n<p>If you are interested in reading them and would like to post a summary, please feel free to do so in this thread (or in other threads).</p>\n<hr>\n<p>Paper with Code ImageNet VID: <br>\n<a href=\"https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid\" target=\"_blank\">https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid</a></p>\n<ul>\n<li><p>Seq-NMS for Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/1602.08465\" target=\"_blank\">https://arxiv.org/abs/1602.08465</a></p></li>\n<li><p>Deep Feature Flow for Video Recognition<br>\n<a href=\"https://arxiv.org/abs/1611.07715\" target=\"_blank\">https://arxiv.org/abs/1611.07715</a></p></li>\n<li><p>T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos<br>\n<a href=\"https://arxiv.org/abs/1604.02532\" target=\"_blank\">https://arxiv.org/abs/1604.02532</a></p></li>\n<li><p>Object Detection in Videos with Tubelet Proposal Networks<br>\n<a href=\"https://arxiv.org/abs/1702.06355\" target=\"_blank\">https://arxiv.org/abs/1702.06355</a></p></li>\n<li><p>Flow-Guided Feature Aggregation for Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/1703.10025\" target=\"_blank\">https://arxiv.org/abs/1703.10025</a></p></li>\n<li><p>Detect to Track and Track to Detect<br>\n<a href=\"https://arxiv.org/abs/1710.03958\" target=\"_blank\">https://arxiv.org/abs/1710.03958</a></p></li>\n<li><p>Towards High Performance Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/1711.11577\" target=\"_blank\">https://arxiv.org/abs/1711.11577</a></p></li>\n<li><p>Mobile Video Object Detection with Temporally-Aware Feature Maps<br>\n<a href=\"https://arxiv.org/abs/1711.06368\" target=\"_blank\">https://arxiv.org/abs/1711.06368</a></p></li>\n<li><p>Video Object Detection with an Aligned Spatial-Temporal Memory<br>\n<a href=\"https://arxiv.org/abs/1712.06317\" target=\"_blank\">https://arxiv.org/abs/1712.06317</a></p></li>\n<li><p>Object Detection in Video with Spatiotemporal Sampling Networks<br>\n<a href=\"https://arxiv.org/abs/1803.05549\" target=\"_blank\">https://arxiv.org/abs/1803.05549</a></p></li>\n<li><p>Sequence Level Semantics Aggregation for Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/1907.06390v2\" target=\"_blank\">https://arxiv.org/abs/1907.06390v2</a></p></li>\n<li><p>Memory Enhanced Global-Local Aggregation for Video Object Detection<br>\n<a href=\"https://arxiv.org/abs/2003.12063v1\" target=\"_blank\">https://arxiv.org/abs/2003.12063v1</a></p></li>\n<li><p>Robust and efficient post-processing for video object detection<br>\n<a href=\"https://arxiv.org/abs/2009.11050\" target=\"_blank\">https://arxiv.org/abs/2009.11050</a></p></li>\n<li><p>Mining Inter-Video Proposal Relations for Video Object Detection<br>\n<a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/html/3764_ECCV_2020_paper.php\" target=\"_blank\">https://www.ecva.net/papers/eccv_2020/papers_ECCV/html/3764_ECCV_2020_paper.php</a></p></li>\n</ul>\n<hr>\n<p>Happy Kaggling ✨</p>",
      "rawMarkdown": "At the moment (Dec.07.2021), many discussions seem to be focusing on object detection in still images, as represented by augmentation and detection models.\nHowever, as anyone who has done or seen EDA will be aware, I think the major key to this competition is that the images are sequence (video). In the case of video detection, many methods have been proposed that make good use of continuous images, i.e., images that are close in time, because they contain information that can reduce detection errors (FP and FN).\n\nSo, I would like to list some of the papers I found by exploring [the ImageNet VID benchmarks in Paper with Code](https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid). Although there are some old papers, I think that understanding the origin of the idea of the proposed methods will help us to come up with an original method for this competition. In addition, the post-processing method can be applied directly to models trained on still images, so it will be of great help to those who are currently considering using models based on still images.\n\nIf you are interested in reading them and would like to post a summary, please feel free to do so in this thread (or in other threads).\n\n---\n\nPaper with Code ImageNet VID: \nhttps://paperswithcode.com/sota/video-object-detection-on-imagenet-vid\n\n\n- Seq-NMS for Video Object Detection\nhttps://arxiv.org/abs/1602.08465\n\n- Deep Feature Flow for Video Recognition\nhttps://arxiv.org/abs/1611.07715\n\n- T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos\nhttps://arxiv.org/abs/1604.02532\n\n- Object Detection in Videos with Tubelet Proposal Networks\nhttps://arxiv.org/abs/1702.06355\n\n- Flow-Guided Feature Aggregation for Video Object Detection\nhttps://arxiv.org/abs/1703.10025\n\n- Detect to Track and Track to Detect\nhttps://arxiv.org/abs/1710.03958\n\n- Towards High Performance Video Object Detection\nhttps://arxiv.org/abs/1711.11577\n\n- Mobile Video Object Detection with Temporally-Aware Feature Maps\nhttps://arxiv.org/abs/1711.06368\n\n- Video Object Detection with an Aligned Spatial-Temporal Memory\nhttps://arxiv.org/abs/1712.06317\n\n- Object Detection in Video with Spatiotemporal Sampling Networks\nhttps://arxiv.org/abs/1803.05549\n\n- Sequence Level Semantics Aggregation for Video Object Detection\nhttps://arxiv.org/abs/1907.06390v2\n\n- Memory Enhanced Global-Local Aggregation for Video Object Detection\nhttps://arxiv.org/abs/2003.12063v1\n\n- Robust and efficient post-processing for video object detection\nhttps://arxiv.org/abs/2009.11050\n\n- Mining Inter-Video Proposal Relations for Video Object Detection\nhttps://www.ecva.net/papers/eccv_2020/papers_ECCV/html/3764_ECCV_2020_paper.php\n\n---\n\nHappy Kaggling ✨",
      "votes": null
    },
    {
      "id": "1611401",
      "postDate": "12/07/2021 23:59:45",
      "content": "<p>Like What algorithm are you mainly hinting at here?</p>",
      "rawMarkdown": "Like What algorithm are you mainly hinting at here?",
      "votes": null
    },
    {
      "id": "1611428",
      "postDate": "12/08/2021 00:49:51",
      "content": "<p>I am still investigating, but for example, even the oldest paper listed above, \"Seq-NMS for Video Object Detection,\" may be helpful.</p>\n<p><img src=\"https://i.imgur.com/CTVFUW5.png\" alt=\"Seq-NMS\"></p>\n<p>Seq-NMS is a technique for post-processing the results of object detection using a video frame as a still image. In this method, bboxes with large IoU are considered as a sequence (the same object moving with time), and the scores of the sequence bboxes are adjusted to raise the ones with low probability (confidence).<br>\nAfter that, NMS is performed on all bboxes, and bboxes that are likely to be FNs in a given frame can be detected thanks to the adjustment made earlier.<br>\nThis technique uses the probability of bboxes to find sequences in post-processing, but I think there are several papers that detect sequences in the model without post-processing.<br>\n(e.g. <a href=\"https://arxiv.org/abs/1702.06355\" target=\"_blank\">Object Detection in Videos with Tubelet Proposal Networks</a>)</p>\n<p>We have to experiment and find out which solution is more suitable for this competition, but as I mentioned in my post above, post-processing techniques may be easier to apply.</p>",
      "rawMarkdown": "I am still investigating, but for example, even the oldest paper listed above, \"Seq-NMS for Video Object Detection,\" may be helpful.\n\n![Seq-NMS](https://i.imgur.com/CTVFUW5.png)\n\nSeq-NMS is a technique for post-processing the results of object detection using a video frame as a still image. In this method, bboxes with large IoU are considered as a sequence (the same object moving with time), and the scores of the sequence bboxes are adjusted to raise the ones with low probability (confidence).\nAfter that, NMS is performed on all bboxes, and bboxes that are likely to be FNs in a given frame can be detected thanks to the adjustment made earlier.\nThis technique uses the probability of bboxes to find sequences in post-processing, but I think there are several papers that detect sequences in the model without post-processing.\n(e.g. [Object Detection in Videos with Tubelet Proposal Networks](https://arxiv.org/abs/1702.06355))\n\nWe have to experiment and find out which solution is more suitable for this competition, but as I mentioned in my post above, post-processing techniques may be easier to apply.",
      "votes": null
    },
    {
      "id": "1611871",
      "postDate": "12/08/2021 10:49:32",
      "content": "<p>I was looking for papers on video object detection. Thanks for sharing a good thesis!</p>",
      "rawMarkdown": "I was looking for papers on video object detection. Thanks for sharing a good thesis!",
      "votes": null
    },
    {
      "id": "1612946",
      "postDate": "12/09/2021 13:44:54",
      "content": "<p>Very useful, thanks for sharing these papers <a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a> </p>",
      "rawMarkdown": "Very useful, thanks for sharing these papers @maxwell110",
      "votes": null
    },
    {
      "id": "1614997",
      "postDate": "12/11/2021 17:38:25",
      "content": "<p>Wouldn't the submission requirements prevent you from doing any post-processing?</p>\n<p><code>for (pixel_array, sample_prediction_df) in iter_test:\n    sample_prediction_df['annotations'] = '0.5 0 0 100 100'  # make your predictions here\n    env.predict(sample_prediction_df)   # register your predictions</code></p>\n<p>Sure, you can store the last n frames and use them to enhance the predictions on the current frame, but you can't go back and update your previous predictions on previous frames like the Seq-NMS paper does, right?</p>",
      "rawMarkdown": "Wouldn't the submission requirements prevent you from doing any post-processing?\n\n`for (pixel_array, sample_prediction_df) in iter_test:\n    sample_prediction_df['annotations'] = '0.5 0 0 100 100'  # make your predictions here\n    env.predict(sample_prediction_df)   # register your predictions`\n\nSure, you can store the last n frames and use them to enhance the predictions on the current frame, but you can't go back and update your previous predictions on previous frames like the Seq-NMS paper does, right?",
      "votes": null
    },
    {
      "id": "1615004",
      "postDate": "12/11/2021 17:46:17",
      "content": "<h1><a href=\"https://arxiv.org/pdf/2105.10920.pdf\" target=\"_blank\">Video Object Detection with Spatial-Temporal Transformers</a></h1>\n<p>Code (not yet released). <a href=\"https://github.com/SJTU-LuHe/TransVOD\" target=\"_blank\">https://github.com/SJTU-LuHe/TransVOD</a><br>\n<img src=\"https://user-images.githubusercontent.com/17668390/145686428-9b894e7f-8db6-41dd-902d-f820a3f314f9.png\" alt=\"image\"></p>",
      "rawMarkdown": "# [Video Object Detection with Spatial-Temporal Transformers](https://arxiv.org/pdf/2105.10920.pdf)\nCode (not yet released). https://github.com/SJTU-LuHe/TransVOD\n![image](https://user-images.githubusercontent.com/17668390/145686428-9b894e7f-8db6-41dd-902d-f820a3f314f9.png)",
      "votes": null
    },
    {
      "id": "1615010",
      "postDate": "12/11/2021 17:51:31",
      "content": "<h1><a href=\"https://reader.elsevier.com/reader/sd/pii/S092427162100099X?token=A7D35964A21C14BC5604C0829C9E5A5BD64B902093205E47FC9F0B22BE50A6C85290A970E9D29E25F03FCE1138FE5515&amp;originRegion=eu-west-1&amp;originCreation=20211211173701\" target=\"_blank\">Video object detection with a convolutional regression tracker</a></h1>\n<p>Code: None.</p>\n<blockquote>\n  <p><strong>ABSTRAC</strong><br>\n  … we propose to augment a well-trained image object detector with an efficient and effective class-agnostic convolutional regression tracker for the video object detection task. The tracker learns to track objects by reusing the features from the image object detector, which is a light-weighted increment to the detector, with only a slight speed drop for the video object detection task. …</p>\n</blockquote>\n<p><img src=\"https://user-images.githubusercontent.com/17668390/145686533-772124d1-e779-4cad-bbeb-aebd1645e145.png\" alt=\"image\"></p>",
      "rawMarkdown": "# [Video object detection with a convolutional regression tracker](https://reader.elsevier.com/reader/sd/pii/S092427162100099X?token=A7D35964A21C14BC5604C0829C9E5A5BD64B902093205E47FC9F0B22BE50A6C85290A970E9D29E25F03FCE1138FE5515&originRegion=eu-west-1&originCreation=20211211173701)\nCode: None.\n\n> **ABSTRAC**\n... we propose to augment a well-trained image object detector with an efficient and effective class-agnostic convolutional regression tracker for the video object detection task. The tracker learns to track objects by reusing the features from the image object detector, which is a light-weighted increment to the detector, with only a slight speed drop for the video object detection task. ...\n\n![image](https://user-images.githubusercontent.com/17668390/145686533-772124d1-e779-4cad-bbeb-aebd1645e145.png)",
      "votes": null
    },
    {
      "id": "1615014",
      "postDate": "12/11/2021 18:02:21",
      "content": "<p><a href=\"https://github.com/OlafenwaMoses/ImageAI\" target=\"_blank\">ImageAI</a></p>\n<ul>\n<li><a href=\"https://github.com/OlafenwaMoses/ImageAI#detection\" target=\"_blank\">Object Detection</a></li>\n<li><a href=\"https://github.com/OlafenwaMoses/ImageAI#videodetection\" target=\"_blank\">Video Object Detection and Tracking</a></li>\n<li><a href=\"https://github.com/OlafenwaMoses/ImageAI#customvideodetection\" target=\"_blank\">Custom Video Object Detection &amp; Analysis</a></li>\n<li>Interesting Post: <a href=\"https://towardsdatascience.com/ug-vod-the-ultimate-guide-to-video-object-detection-816a76073aef\" target=\"_blank\">Guide to Video Object Detection</a></li>\n<li><a href=\"https://github.com/msracver/Flow-Guided-Feature-Aggregation?utm_source=catalyzex.com\" target=\"_blank\">Flow-Guided Feature Aggregation for Video Object Detection</a></li>\n</ul>",
      "rawMarkdown": "[ImageAI](https://github.com/OlafenwaMoses/ImageAI)\n- [Object Detection](https://github.com/OlafenwaMoses/ImageAI#detection)\n- [Video Object Detection and Tracking](https://github.com/OlafenwaMoses/ImageAI#videodetection)\n- [Custom Video Object Detection & Analysis](https://github.com/OlafenwaMoses/ImageAI#customvideodetection)\n- Interesting Post: [Guide to Video Object Detection](https://towardsdatascience.com/ug-vod-the-ultimate-guide-to-video-object-detection-816a76073aef)\n- [Flow-Guided Feature Aggregation for Video Object Detection](https://github.com/msracver/Flow-Guided-Feature-Aggregation?utm_source=catalyzex.com)",
      "votes": null
    },
    {
      "id": "1615017",
      "postDate": "12/11/2021 18:08:24",
      "content": "<p>Was really considering using TansVOD until I realized the model isn't released. I could follow the paper and make my own and train it on the ImageNet VID dataset, but that would be a lot, not sure if it's worth my time at that point :/</p>",
      "rawMarkdown": "Was really considering using TansVOD until I realized the model isn't released. I could follow the paper and make my own and train it on the ImageNet VID dataset, but that would be a lot, not sure if it's worth my time at that point :/",
      "votes": null
    },
    {
      "id": "1615184",
      "postDate": "12/12/2021 00:46:04",
      "content": "<p>Thank you for your comment 😊<br>\nThat's a very good point.</p>\n<p>Yes. As you pointed out, as far as I can figure out from the API specification, it is not possible to use the prediction of future frames<br>\n(Ref: <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/291520#1599795\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/291520#1599795</a>)</p>\n<p>I think there are two things to keep in mind when applying the various methods of Video Object Detection to this competition.</p>\n<ol>\n<li><p>It is not possible to use the information of future frames, only the information from past frames.</p></li>\n<li><p>There is no information about the start and end frames of the video clip, so there must be some process to determine whether the video is continuous or not.</p></li>\n</ol>\n<p>I think it's inevitable that the information on future frames cannot be used because this competition is intended for real-time object detection, but I would have liked the organizer to provide some information on the continuity of video clips 😩</p>",
      "rawMarkdown": "Thank you for your comment 😊\nThat's a very good point.\n\nYes. As you pointed out, as far as I can figure out from the API specification, it is not possible to use the prediction of future frames\n(Ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/291520#1599795)\n\nI think there are two things to keep in mind when applying the various methods of Video Object Detection to this competition.\n\n1. It is not possible to use the information of future frames, only the information from past frames.\n\n2. There is no information about the start and end frames of the video clip, so there must be some process to determine whether the video is continuous or not.\n\n\nI think it's inevitable that the information on future frames cannot be used because this competition is intended for real-time object detection, but I would have liked the organizer to provide some information on the continuity of video clips 😩",
      "votes": null
    },
    {
      "id": "1615188",
      "postDate": "12/12/2021 00:57:02",
      "content": "<p>I completely agree, they could have made clear indication of the separate videos, maybe you loop through an n-tuple of iterables provided by the environment API, each of which is it's own continuous video. It's not like the model will have to deal with non-temporally coherent data when deployed in the real world anyways.</p>\n<p>Not a big deal though, I really think you can do a simple subtraction or something like it of the previous and current frame to see if they are different. We also have three separate videos so you could emperically test what a good threshold value would be.</p>",
      "rawMarkdown": "I completely agree, they could have made clear indication of the separate videos, maybe you loop through an n-tuple of iterables provided by the environment API, each of which is it's own continuous video. It's not like the model will have to deal with non-temporally coherent data when deployed in the real world anyways.\n\nNot a big deal though, I really think you can do a simple subtraction or something like it of the previous and current frame to see if they are different. We also have three separate videos so you could emperically test what a good threshold value would be.",
      "votes": null
    },
    {
      "id": "1615281",
      "postDate": "12/12/2021 05:17:46",
      "content": "<p>Very Useful Information</p>",
      "rawMarkdown": "Very Useful Information",
      "votes": null
    },
    {
      "id": "1615299",
      "postDate": "12/12/2021 05:57:41",
      "content": "<p>One possibility is to use the optical flow of <code>opencv2</code>.<br>\n(Ref: <a href=\"https://learnopencv.com/optical-flow-in-opencv/\" target=\"_blank\">https://learnopencv.com/optical-flow-in-opencv/</a>)</p>\n<p>I experimented a bit, and it seemed to be able to get at least the movement vector of the camera in continuous frames. If there is a discontinuous frame, the vector will be very disordered or the optical flow may not be able to be obtained.</p>\n<p>The figure below shows an example of one of my experiments.</p>\n<p><img src=\"https://i.imgur.com/lMs0xMb.jpeg\" alt=\"\"></p>",
      "rawMarkdown": "One possibility is to use the optical flow of `opencv2`.\n(Ref: https://learnopencv.com/optical-flow-in-opencv/)\n\nI experimented a bit, and it seemed to be able to get at least the movement vector of the camera in continuous frames. If there is a discontinuous frame, the vector will be very disordered or the optical flow may not be able to be obtained.\n\nThe figure below shows an example of one of my experiments.\n\n![](https://i.imgur.com/lMs0xMb.jpeg)",
      "votes": null
    },
    {
      "id": "1615549",
      "postDate": "12/12/2021 11:33:09",
      "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a> </p>",
      "rawMarkdown": "thanks for sharing @maxwell110",
      "votes": null
    },
    {
      "id": "1615551",
      "postDate": "12/12/2021 11:36:40",
      "content": "<p><a href=\"https://www.kaggle.com/altruisticemphasis\" target=\"_blank\">@altruisticemphasis</a>  DO HAVE A LOOK THIS WILL HELP YOU!!!</p>",
      "rawMarkdown": "altruisticemphasis  DO HAVE A LOOK THIS WILL HELP YOU!!!",
      "votes": null
    },
    {
      "id": "1616299",
      "postDate": "12/13/2021 10:14:51",
      "content": "<p>GOOD WORK!</p>",
      "rawMarkdown": "GOOD WORK!",
      "votes": null
    },
    {
      "id": "1661244",
      "postDate": "01/23/2022 10:25:32",
      "content": "<p>Thank you for your meaningful discussion, <a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a> .<br>\nI also think an optical flow-based method is sufficient for this task, because COTSs does not move by itself.</p>",
      "rawMarkdown": "Thank you for your meaningful discussion, @maxwell110 .\nI also think an optical flow-based method is sufficient for this task, because COTSs does not move by itself.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1611401,
      "author_name": "muhammadammarjamshed",
      "author_url": "",
      "post_date": "12/07/2021 23:59:45",
      "content": "<p>Like What algorithm are you mainly hinting at here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1611428,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "12/08/2021 00:49:51",
          "content": "<p>I am still investigating, but for example, even the oldest paper listed above, \"Seq-NMS for Video Object Detection,\" may be helpful.</p>\n<p><img src=\"https://i.imgur.com/CTVFUW5.png\" alt=\"Seq-NMS\"></p>\n<p>Seq-NMS is a technique for post-processing the results of object detection using a video frame as a still image. In this method, bboxes with large IoU are considered as a sequence (the same object moving with time), and the scores of the sequence bboxes are adjusted to raise the ones with low probability (confidence).<br>\nAfter that, NMS is performed on all bboxes, and bboxes that are likely to be FNs in a given frame can be detected thanks to the adjustment made earlier.<br>\nThis technique uses the probability of bboxes to find sequences in post-processing, but I think there are several papers that detect sequences in the model without post-processing.<br>\n(e.g. <a href=\"https://arxiv.org/abs/1702.06355\" target=\"_blank\">Object Detection in Videos with Tubelet Proposal Networks</a>)</p>\n<p>We have to experiment and find out which solution is more suitable for this competition, but as I mentioned in my post above, post-processing techniques may be easier to apply.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1614997,
          "author_name": "togden1",
          "author_url": "",
          "post_date": "12/11/2021 17:38:25",
          "content": "<p>Wouldn't the submission requirements prevent you from doing any post-processing?</p>\n<p><code>for (pixel_array, sample_prediction_df) in iter_test:\n    sample_prediction_df['annotations'] = '0.5 0 0 100 100'  # make your predictions here\n    env.predict(sample_prediction_df)   # register your predictions</code></p>\n<p>Sure, you can store the last n frames and use them to enhance the predictions on the current frame, but you can't go back and update your previous predictions on previous frames like the Seq-NMS paper does, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1615184,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "12/12/2021 00:46:04",
          "content": "<p>Thank you for your comment 😊<br>\nThat's a very good point.</p>\n<p>Yes. As you pointed out, as far as I can figure out from the API specification, it is not possible to use the prediction of future frames<br>\n(Ref: <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/291520#1599795\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/291520#1599795</a>)</p>\n<p>I think there are two things to keep in mind when applying the various methods of Video Object Detection to this competition.</p>\n<ol>\n<li><p>It is not possible to use the information of future frames, only the information from past frames.</p></li>\n<li><p>There is no information about the start and end frames of the video clip, so there must be some process to determine whether the video is continuous or not.</p></li>\n</ol>\n<p>I think it's inevitable that the information on future frames cannot be used because this competition is intended for real-time object detection, but I would have liked the organizer to provide some information on the continuity of video clips 😩</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1615188,
          "author_name": "togden1",
          "author_url": "",
          "post_date": "12/12/2021 00:57:02",
          "content": "<p>I completely agree, they could have made clear indication of the separate videos, maybe you loop through an n-tuple of iterables provided by the environment API, each of which is it's own continuous video. It's not like the model will have to deal with non-temporally coherent data when deployed in the real world anyways.</p>\n<p>Not a big deal though, I really think you can do a simple subtraction or something like it of the previous and current frame to see if they are different. We also have three separate videos so you could emperically test what a good threshold value would be.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1615299,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "12/12/2021 05:57:41",
          "content": "<p>One possibility is to use the optical flow of <code>opencv2</code>.<br>\n(Ref: <a href=\"https://learnopencv.com/optical-flow-in-opencv/\" target=\"_blank\">https://learnopencv.com/optical-flow-in-opencv/</a>)</p>\n<p>I experimented a bit, and it seemed to be able to get at least the movement vector of the camera in continuous frames. If there is a discontinuous frame, the vector will be very disordered or the optical flow may not be able to be obtained.</p>\n<p>The figure below shows an example of one of my experiments.</p>\n<p><img src=\"https://i.imgur.com/lMs0xMb.jpeg\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1661244,
          "author_name": "nobmatsushima",
          "author_url": "",
          "post_date": "01/23/2022 10:25:32",
          "content": "<p>Thank you for your meaningful discussion, <a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a> .<br>\nI also think an optical flow-based method is sufficient for this task, because COTSs does not move by itself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1611871,
      "author_name": "kimbosoek",
      "author_url": "",
      "post_date": "12/08/2021 10:49:32",
      "content": "<p>I was looking for papers on video object detection. Thanks for sharing a good thesis!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1612946,
      "author_name": "shivamb",
      "author_url": "",
      "post_date": "12/09/2021 13:44:54",
      "content": "<p>Very useful, thanks for sharing these papers <a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1615004,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "12/11/2021 17:46:17",
      "content": "<h1><a href=\"https://arxiv.org/pdf/2105.10920.pdf\" target=\"_blank\">Video Object Detection with Spatial-Temporal Transformers</a></h1>\n<p>Code (not yet released). <a href=\"https://github.com/SJTU-LuHe/TransVOD\" target=\"_blank\">https://github.com/SJTU-LuHe/TransVOD</a><br>\n<img src=\"https://user-images.githubusercontent.com/17668390/145686428-9b894e7f-8db6-41dd-902d-f820a3f314f9.png\" alt=\"image\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1615017,
          "author_name": "togden1",
          "author_url": "",
          "post_date": "12/11/2021 18:08:24",
          "content": "<p>Was really considering using TansVOD until I realized the model isn't released. I could follow the paper and make my own and train it on the ImageNet VID dataset, but that would be a lot, not sure if it's worth my time at that point :/</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1615010,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "12/11/2021 17:51:31",
      "content": "<h1><a href=\"https://reader.elsevier.com/reader/sd/pii/S092427162100099X?token=A7D35964A21C14BC5604C0829C9E5A5BD64B902093205E47FC9F0B22BE50A6C85290A970E9D29E25F03FCE1138FE5515&amp;originRegion=eu-west-1&amp;originCreation=20211211173701\" target=\"_blank\">Video object detection with a convolutional regression tracker</a></h1>\n<p>Code: None.</p>\n<blockquote>\n  <p><strong>ABSTRAC</strong><br>\n  … we propose to augment a well-trained image object detector with an efficient and effective class-agnostic convolutional regression tracker for the video object detection task. The tracker learns to track objects by reusing the features from the image object detector, which is a light-weighted increment to the detector, with only a slight speed drop for the video object detection task. …</p>\n</blockquote>\n<p><img src=\"https://user-images.githubusercontent.com/17668390/145686533-772124d1-e779-4cad-bbeb-aebd1645e145.png\" alt=\"image\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1615014,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "12/11/2021 18:02:21",
      "content": "<p><a href=\"https://github.com/OlafenwaMoses/ImageAI\" target=\"_blank\">ImageAI</a></p>\n<ul>\n<li><a href=\"https://github.com/OlafenwaMoses/ImageAI#detection\" target=\"_blank\">Object Detection</a></li>\n<li><a href=\"https://github.com/OlafenwaMoses/ImageAI#videodetection\" target=\"_blank\">Video Object Detection and Tracking</a></li>\n<li><a href=\"https://github.com/OlafenwaMoses/ImageAI#customvideodetection\" target=\"_blank\">Custom Video Object Detection &amp; Analysis</a></li>\n<li>Interesting Post: <a href=\"https://towardsdatascience.com/ug-vod-the-ultimate-guide-to-video-object-detection-816a76073aef\" target=\"_blank\">Guide to Video Object Detection</a></li>\n<li><a href=\"https://github.com/msracver/Flow-Guided-Feature-Aggregation?utm_source=catalyzex.com\" target=\"_blank\">Flow-Guided Feature Aggregation for Video Object Detection</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1615281,
      "author_name": "ingalevarad",
      "author_url": "",
      "post_date": "12/12/2021 05:17:46",
      "content": "<p>Very Useful Information</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1615549,
      "author_name": "vinayaktiwari28",
      "author_url": "",
      "post_date": "12/12/2021 11:33:09",
      "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1615551,
      "author_name": "vinayaktiwari28",
      "author_url": "",
      "post_date": "12/12/2021 11:36:40",
      "content": "<p><a href=\"https://www.kaggle.com/altruisticemphasis\" target=\"_blank\">@altruisticemphasis</a>  DO HAVE A LOOK THIS WILL HELP YOU!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1616299,
      "author_name": "gollakeerthana",
      "author_url": "",
      "post_date": "12/13/2021 10:14:51",
      "content": "<p>GOOD WORK!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1610305": "At the moment (Dec.07.2021), many discussions seem to be focusing on object detection in still images, as represented by augmentation and detection models.\nHowever, as anyone who has done or seen EDA will be aware, I think the major key to this competition is that the images are sequence (video). In the case of video detection, many methods have been proposed that make good use of continuous images, i.e., images that are close in time, because they contain information that can reduce detection errors (FP and FN).\n\nSo, I would like to list some of the papers I found by exploring [the ImageNet VID benchmarks in Paper with Code](https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid). Although there are some old papers, I think that understanding the origin of the idea of the proposed methods will help us to come up with an original method for this competition. In addition, the post-processing method can be applied directly to models trained on still images, so it will be of great help to those who are currently considering using models based on still images.\n\nIf you are interested in reading them and would like to post a summary, please feel free to do so in this thread (or in other threads).\n\n---\n\nPaper with Code ImageNet VID: \nhttps://paperswithcode.com/sota/video-object-detection-on-imagenet-vid\n\n\n- Seq-NMS for Video Object Detection\nhttps://arxiv.org/abs/1602.08465\n\n- Deep Feature Flow for Video Recognition\nhttps://arxiv.org/abs/1611.07715\n\n- T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos\nhttps://arxiv.org/abs/1604.02532\n\n- Object Detection in Videos with Tubelet Proposal Networks\nhttps://arxiv.org/abs/1702.06355\n\n- Flow-Guided Feature Aggregation for Video Object Detection\nhttps://arxiv.org/abs/1703.10025\n\n- Detect to Track and Track to Detect\nhttps://arxiv.org/abs/1710.03958\n\n- Towards High Performance Video Object Detection\nhttps://arxiv.org/abs/1711.11577\n\n- Mobile Video Object Detection with Temporally-Aware Feature Maps\nhttps://arxiv.org/abs/1711.06368\n\n- Video Object Detection with an Aligned Spatial-Temporal Memory\nhttps://arxiv.org/abs/1712.06317\n\n- Object Detection in Video with Spatiotemporal Sampling Networks\nhttps://arxiv.org/abs/1803.05549\n\n- Sequence Level Semantics Aggregation for Video Object Detection\nhttps://arxiv.org/abs/1907.06390v2\n\n- Memory Enhanced Global-Local Aggregation for Video Object Detection\nhttps://arxiv.org/abs/2003.12063v1\n\n- Robust and efficient post-processing for video object detection\nhttps://arxiv.org/abs/2009.11050\n\n- Mining Inter-Video Proposal Relations for Video Object Detection\nhttps://www.ecva.net/papers/eccv_2020/papers_ECCV/html/3764_ECCV_2020_paper.php\n\n---\n\nHappy Kaggling ✨",
    "1611401": "Like What algorithm are you mainly hinting at here?",
    "1611428": "I am still investigating, but for example, even the oldest paper listed above, \"Seq-NMS for Video Object Detection,\" may be helpful.\n\n![Seq-NMS](https://i.imgur.com/CTVFUW5.png)\n\nSeq-NMS is a technique for post-processing the results of object detection using a video frame as a still image. In this method, bboxes with large IoU are considered as a sequence (the same object moving with time), and the scores of the sequence bboxes are adjusted to raise the ones with low probability (confidence).\nAfter that, NMS is performed on all bboxes, and bboxes that are likely to be FNs in a given frame can be detected thanks to the adjustment made earlier.\nThis technique uses the probability of bboxes to find sequences in post-processing, but I think there are several papers that detect sequences in the model without post-processing.\n(e.g. [Object Detection in Videos with Tubelet Proposal Networks](https://arxiv.org/abs/1702.06355))\n\nWe have to experiment and find out which solution is more suitable for this competition, but as I mentioned in my post above, post-processing techniques may be easier to apply.",
    "1611871": "I was looking for papers on video object detection. Thanks for sharing a good thesis!",
    "1612946": "Very useful, thanks for sharing these papers @maxwell110",
    "1614997": "Wouldn't the submission requirements prevent you from doing any post-processing?\n\n`for (pixel_array, sample_prediction_df) in iter_test:\n    sample_prediction_df['annotations'] = '0.5 0 0 100 100'  # make your predictions here\n    env.predict(sample_prediction_df)   # register your predictions`\n\nSure, you can store the last n frames and use them to enhance the predictions on the current frame, but you can't go back and update your previous predictions on previous frames like the Seq-NMS paper does, right?",
    "1615004": "# [Video Object Detection with Spatial-Temporal Transformers](https://arxiv.org/pdf/2105.10920.pdf)\nCode (not yet released). https://github.com/SJTU-LuHe/TransVOD\n![image](https://user-images.githubusercontent.com/17668390/145686428-9b894e7f-8db6-41dd-902d-f820a3f314f9.png)",
    "1615010": "# [Video object detection with a convolutional regression tracker](https://reader.elsevier.com/reader/sd/pii/S092427162100099X?token=A7D35964A21C14BC5604C0829C9E5A5BD64B902093205E47FC9F0B22BE50A6C85290A970E9D29E25F03FCE1138FE5515&originRegion=eu-west-1&originCreation=20211211173701)\nCode: None.\n\n> **ABSTRAC**\n... we propose to augment a well-trained image object detector with an efficient and effective class-agnostic convolutional regression tracker for the video object detection task. The tracker learns to track objects by reusing the features from the image object detector, which is a light-weighted increment to the detector, with only a slight speed drop for the video object detection task. ...\n\n![image](https://user-images.githubusercontent.com/17668390/145686533-772124d1-e779-4cad-bbeb-aebd1645e145.png)",
    "1615014": "[ImageAI](https://github.com/OlafenwaMoses/ImageAI)\n- [Object Detection](https://github.com/OlafenwaMoses/ImageAI#detection)\n- [Video Object Detection and Tracking](https://github.com/OlafenwaMoses/ImageAI#videodetection)\n- [Custom Video Object Detection & Analysis](https://github.com/OlafenwaMoses/ImageAI#customvideodetection)\n- Interesting Post: [Guide to Video Object Detection](https://towardsdatascience.com/ug-vod-the-ultimate-guide-to-video-object-detection-816a76073aef)\n- [Flow-Guided Feature Aggregation for Video Object Detection](https://github.com/msracver/Flow-Guided-Feature-Aggregation?utm_source=catalyzex.com)",
    "1615017": "Was really considering using TansVOD until I realized the model isn't released. I could follow the paper and make my own and train it on the ImageNet VID dataset, but that would be a lot, not sure if it's worth my time at that point :/",
    "1615184": "Thank you for your comment 😊\nThat's a very good point.\n\nYes. As you pointed out, as far as I can figure out from the API specification, it is not possible to use the prediction of future frames\n(Ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/291520#1599795)\n\nI think there are two things to keep in mind when applying the various methods of Video Object Detection to this competition.\n\n1. It is not possible to use the information of future frames, only the information from past frames.\n\n2. There is no information about the start and end frames of the video clip, so there must be some process to determine whether the video is continuous or not.\n\n\nI think it's inevitable that the information on future frames cannot be used because this competition is intended for real-time object detection, but I would have liked the organizer to provide some information on the continuity of video clips 😩",
    "1615188": "I completely agree, they could have made clear indication of the separate videos, maybe you loop through an n-tuple of iterables provided by the environment API, each of which is it's own continuous video. It's not like the model will have to deal with non-temporally coherent data when deployed in the real world anyways.\n\nNot a big deal though, I really think you can do a simple subtraction or something like it of the previous and current frame to see if they are different. We also have three separate videos so you could emperically test what a good threshold value would be.",
    "1615281": "Very Useful Information",
    "1615299": "One possibility is to use the optical flow of `opencv2`.\n(Ref: https://learnopencv.com/optical-flow-in-opencv/)\n\nI experimented a bit, and it seemed to be able to get at least the movement vector of the camera in continuous frames. If there is a discontinuous frame, the vector will be very disordered or the optical flow may not be able to be obtained.\n\nThe figure below shows an example of one of my experiments.\n\n![](https://i.imgur.com/lMs0xMb.jpeg)",
    "1615549": "thanks for sharing @maxwell110",
    "1615551": "altruisticemphasis  DO HAVE A LOOK THIS WILL HELP YOU!!!",
    "1616299": "GOOD WORK!",
    "1661244": "Thank you for your meaningful discussion, @maxwell110 .\nI also think an optical flow-based method is sufficient for this task, because COTSs does not move by itself."
  },
  "source": "meta"
}