{
  "id": 290350,
  "title": "Problem statement is not only Object detection it's \"Video\" Object Detection, given data is in a video sequence ",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/290350",
  "author_name": "",
  "post_date": "2021-11-24T08:19:24.695337200Z",
  "votes": 33,
  "comment_count": 14,
  "views": 0,
  "content": "<p>\n<img src=\"https://i.imgur.com/tTOe7cV.png\">\n</p>\n<h1>Using Video Object Detection [Sequential Model + Object detection], because the images are in a video sequence</h1>\n<p><strong><em>Howdy!!!</em></strong> Seems like <strong>Patrick</strong> is lost, and  <strong>SpongeBob</strong> and others need our help to find him. I was discussing with my friends about different ideas we can use to make a better model(and find Patrick). We came up with an idea that is <strong><code>using a kind of Sequential model like LSTM or RNN to capture the sequential information and then do Object detection.</code></strong>  <strong><code>The Idea is similar to video object detection</code></strong>. As most people are doing only object detection, it will be great to throw a Sequential model with object detection and see if that can help or not. <em><code>It is like taking the context of the previous frame and using it to predict the object's position for the next frame.</code></em></p>\n<p>If you see the video from the <strong>Overview</strong> tab, it also mentions that model related to time series. And one picture from the video, </p>\n<p>\n<img src=\"https://i.imgur.com/XesZdt4.png\">\n</p>\n<p>also shows[not sure that info is right or wrong] that starfishes borns in groups like you might not find a single(or only 2 or 3) starfish in the sea, it means if you see a starfish that means you will find starfish in some nearby areas too. <strong>So, taking this into context and using the sequential image data[kind of like a video] it would be great to create a video object detection model that can detect starfishes(and help find Patrick).</strong></p>\n<p><strong><code>It also opens a lot of opportunities for attention and transformer-based models</code></strong>[DART might do really well]. I don't have experience in video object detection but here are some Benchmark papers on video object detection that i could find,</p>\n<ul>\n<li><p><a href=\"https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid\" target=\"_blank\">Mining Inter-Video Proposal Relations for Video Object Detection</a></p></li>\n<li><p><a href=\"https://paperswithcode.com/sota/video-object-detection-on-davis-2017\" target=\"_blank\">Emerging Properties in Self-Supervised Vision Transformers</a></p></li>\n<li><p><a href=\"https://paperswithcode.com/sota/video-object-detection-on-epic-kitchens-seen\" target=\"_blank\">Temporal RoI Align for Video Object Recognition</a></p></li>\n<li><p><a href=\"https://paperswithcode.com/sota/video-object-detection-on-epic-kitchens-1\" target=\"_blank\">Temporal RoI Align for Video Object Recognition</a></p></li>\n</ul>\n<p>Check more here -&gt; <a href=\"https://paperswithcode.com/task/video-object-detection\" target=\"_blank\">https://paperswithcode.com/task/video-object-detection</a></p>\n\n<p>I will try to implement a model based on this idea. Would like to hear your thoughts on this, and any point or advice you want to give😇. Happy kaggling, and <strong>Thanksgiving</strong>!</p>",
  "messages": [
    {
      "id": "1593743",
      "postDate": "11/24/2021 08:19:24",
      "content": "<p>\n<img src=\"https://i.imgur.com/tTOe7cV.png\">\n</p>\n<h1>Using Video Object Detection [Sequential Model + Object detection], because the images are in a video sequence</h1>\n<p><strong><em>Howdy!!!</em></strong> Seems like <strong>Patrick</strong> is lost, and  <strong>SpongeBob</strong> and others need our help to find him. I was discussing with my friends about different ideas we can use to make a better model(and find Patrick). We came up with an idea that is <strong><code>using a kind of Sequential model like LSTM or RNN to capture the sequential information and then do Object detection.</code></strong>  <strong><code>The Idea is similar to video object detection</code></strong>. As most people are doing only object detection, it will be great to throw a Sequential model with object detection and see if that can help or not. <em><code>It is like taking the context of the previous frame and using it to predict the object's position for the next frame.</code></em></p>\n<p>If you see the video from the <strong>Overview</strong> tab, it also mentions that model related to time series. And one picture from the video, </p>\n<p>\n<img src=\"https://i.imgur.com/XesZdt4.png\">\n</p>\n<p>also shows[not sure that info is right or wrong] that starfishes borns in groups like you might not find a single(or only 2 or 3) starfish in the sea, it means if you see a starfish that means you will find starfish in some nearby areas too. <strong>So, taking this into context and using the sequential image data[kind of like a video] it would be great to create a video object detection model that can detect starfishes(and help find Patrick).</strong></p>\n<p><strong><code>It also opens a lot of opportunities for attention and transformer-based models</code></strong>[DART might do really well]. I don't have experience in video object detection but here are some Benchmark papers on video object detection that i could find,</p>\n<ul>\n<li><p><a href=\"https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid\" target=\"_blank\">Mining Inter-Video Proposal Relations for Video Object Detection</a></p></li>\n<li><p><a href=\"https://paperswithcode.com/sota/video-object-detection-on-davis-2017\" target=\"_blank\">Emerging Properties in Self-Supervised Vision Transformers</a></p></li>\n<li><p><a href=\"https://paperswithcode.com/sota/video-object-detection-on-epic-kitchens-seen\" target=\"_blank\">Temporal RoI Align for Video Object Recognition</a></p></li>\n<li><p><a href=\"https://paperswithcode.com/sota/video-object-detection-on-epic-kitchens-1\" target=\"_blank\">Temporal RoI Align for Video Object Recognition</a></p></li>\n</ul>\n<p>Check more here -&gt; <a href=\"https://paperswithcode.com/task/video-object-detection\" target=\"_blank\">https://paperswithcode.com/task/video-object-detection</a></p>\n\n<p>I will try to implement a model based on this idea. Would like to hear your thoughts on this, and any point or advice you want to give😇. Happy kaggling, and <strong>Thanksgiving</strong>!</p>",
      "rawMarkdown": "<p align=\"center\">\n<img width = \"700\" src=\"https://i.imgur.com/tTOe7cV.png\">\n</p>\n<h1 style=\"text-align: center; font-family: Verdana; font-size: 32px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; font-variant: small-caps; letter-spacing: 3px; color: #468282; background-color: #ffffff;\">Using Video Object Detection [Sequential Model + Object detection], because the images are in a video sequence</h1>\n***Howdy!!!*** Seems like **Patrick** is lost, and  **SpongeBob** and others need our help to find him. I was discussing with my friends about different ideas we can use to make a better model(and find Patrick). We came up with an idea that is **`using a kind of Sequential model like LSTM or RNN to capture the sequential information and then do Object detection.`**  **`The Idea is similar to video object detection`**. As most people are doing only object detection, it will be great to throw a Sequential model with object detection and see if that can help or not. *`It is like taking the context of the previous frame and using it to predict the object's position for the next frame.`*\n\nIf you see the video from the **Overview** tab, it also mentions that model related to time series. And one picture from the video, \n<p align=\"center\">\n<img width = \"900\" src=\"https://i.imgur.com/XesZdt4.png\">\n</p>\n\nalso shows[not sure that info is right or wrong] that starfishes borns in groups like you might not find a single(or only 2 or 3) starfish in the sea, it means if you see a starfish that means you will find starfish in some nearby areas too. **So, taking this into context and using the sequential image data[kind of like a video] it would be great to create a video object detection model that can detect starfishes(and help find Patrick).**\n\n**`It also opens a lot of opportunities for attention and transformer-based models`**[DART might do really well]. I don't have experience in video object detection but here are some Benchmark papers on video object detection that i could find,\n\n\t\n- [Mining Inter-Video Proposal Relations for Video Object Detection](https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid)\n- [Emerging Properties in Self-Supervised Vision Transformers](https://paperswithcode.com/sota/video-object-detection-on-davis-2017)\n\n- [Temporal RoI Align for Video Object Recognition](https://paperswithcode.com/sota/video-object-detection-on-epic-kitchens-seen)\n\t\n- [Temporal RoI Align for Video Object Recognition](https://paperswithcode.com/sota/video-object-detection-on-epic-kitchens-1)\n\n\nCheck more here -> https://paperswithcode.com/task/video-object-detection\n<!--\nI am trying to make FasterRCNN work using Detectron2 and do some EDA [its not done yet though],  check out my NB here,\n- [Finding Patrick w/ Detectron2 fasterRCNN + EDA](https://www.kaggle.com/soumya9977/finding-patrick-w-detectron2-fasterrcnn-eda)\n-->\nI will try to implement a model based on this idea. Would like to hear your thoughts on this, and any point or advice you want to give😇. Happy kaggling, and **Thanksgiving**!",
      "votes": null
    },
    {
      "id": "1593810",
      "postDate": "11/24/2021 09:43:00",
      "content": "<p>How long sequence you are going to use in one sample?<br>\nHere we have a limited amount of video. Will be there a high chance of overfitting?</p>\n<p>In addition to your brilliant idea - how about Inflated 3D convolution (I am not familiar with transformed yet, sorry). May it be helpful to perform convolution by width/height and time as well?</p>",
      "rawMarkdown": "How long sequence you are going to use in one sample?\nHere we have a limited amount of video. Will be there a high chance of overfitting?\n\nIn addition to your brilliant idea - how about Inflated 3D convolution (I am not familiar with transformed yet, sorry). May it be helpful to perform convolution by width/height and time as well?",
      "votes": null
    },
    {
      "id": "1593816",
      "postDate": "11/24/2021 09:48:07",
      "content": "<p>yeah, I guess it is a very bad idea to go with, so many problems, sorry for posting it. i guess its not much helpful taking sequence information into consideration after all. </p>\n<p>it's just my opinion though.</p>",
      "rawMarkdown": "yeah, I guess it is a very bad idea to go with, so many problems, sorry for posting it. i guess its not much helpful taking sequence information into consideration after all. \n\nit's just my opinion though.",
      "votes": null
    },
    {
      "id": "1593825",
      "postDate": "11/24/2021 09:53:49",
      "content": "<p>Hey, mate. I believe your approach is good. Anyway, thank you for posting your thoughts!</p>",
      "rawMarkdown": "Hey, mate. I believe your approach is good. Anyway, thank you for posting your thoughts!",
      "votes": null
    },
    {
      "id": "1594332",
      "postDate": "11/24/2021 17:52:23",
      "content": "<p>I strongly agree that using multiple frames is likely to be fruitful. Another point I’d make is that, unlike general challenges with object tracking, the COTS are (essentially) stationery and the tow is (relatively) constant and unidirectional. So as the tow passes over a target, we can assume that the target will shift, over the course of <code>n</code> frames, not much in <code>x</code> but increasing in <code>y</code>. </p>\n<p>That might be projecting too much heuristic thinking — “just let the model decide” — but given the desire for efficiency, there <em>is</em> a part of me that is looking for a solution that takes advantage of that specific aspect of the domain.</p>",
      "rawMarkdown": "I strongly agree that using multiple frames is likely to be fruitful. Another point I’d make is that, unlike general challenges with object tracking, the COTS are (essentially) stationery and the tow is (relatively) constant and unidirectional. So as the tow passes over a target, we can assume that the target will shift, over the course of `n` frames, not much in `x` but increasing in `y`. \n\nThat might be projecting too much heuristic thinking — “just let the model decide” — but given the desire for efficiency, there _is_ a part of me that is looking for a solution that takes advantage of that specific aspect of the domain.",
      "votes": null
    },
    {
      "id": "1594489",
      "postDate": "11/24/2021 21:51:30",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/lobrien\" target=\"_blank\">@lobrien</a> , I really appreciate your thoughtfulness. <br>\n<code>So as the tow passes over a target, we can assume that the target will shift, over the course of n frames, not much in x but increasing in y.</code> <br>\nby that, you mean that over the course of n frames right after passing a target, the next target would be a bit higher[increase in y] compared to the previous target but the difference in x won't be much compared to y? </p>\n<ul>\n<li>This might be true cause I have observed in the data that there are many times when I saw that the camera angle was from the top[not straight top though], so we are getting a tilted top view, like looking at something at the bottom from a tilted top. And when you see something from the top, then the next element looks a bit further than the previous element. but at the same time, the next element might be at the same x, this will happen if the targets happen to be in a row.</li>\n</ul>\n<p>\n<img src=\"https://i.imgur.com/3vPdBlK.png\">\n</p>\n\n<p>My interpretation might be wrong though.</p>",
      "rawMarkdown": "Thank you @lobrien , I really appreciate your thoughtfulness. \n`So as the tow passes over a target, we can assume that the target will shift, over the course of n frames, not much in x but increasing in y.` \nby that, you mean that over the course of n frames right after passing a target, the next target would be a bit higher[increase in y] compared to the previous target but the difference in x won't be much compared to y? \n\n- This might be true cause I have observed in the data that there are many times when I saw that the camera angle was from the top[not straight top though], so we are getting a tilted top view, like looking at something at the bottom from a tilted top. And when you see something from the top, then the next element looks a bit further than the previous element. but at the same time, the next element might be at the same x, this will happen if the targets happen to be in a row.\n\n\n<p align=\"center\">\n<img src=\"https://i.imgur.com/3vPdBlK.png\">\n</p>\n\n<!--\n|                                                                                 |\n|                                                                                 |\n|                     0   -> t2                                               |              0  -> t2\n|                                                                                 |\n|       o  -> t1                                                               |              o  -> t1\n|___________________________                                      |____________________________\n[Top view of two targets]\nnot in the same row, but y is increased                     y increased but x is the same\n-->\n\nMy interpretation might be wrong though.",
      "votes": null
    },
    {
      "id": "1594499",
      "postDate": "11/24/2021 21:59:13",
      "content": "<p>hey <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a> , your point about the less data is true, and I also saw that there are many empty annotations in the dataset, so using a pre-trained model might yield some good results. and if someone wants to try it then, they can check out the total length of the video, and divide them into equal parts accordingly.  I don't know the specifics though, cause I have not tried video object detection before, I'm sorry about that.</p>",
      "rawMarkdown": "hey @meowmeowmeowmeowmeow , your point about the less data is true, and I also saw that there are many empty annotations in the dataset, so using a pre-trained model might yield some good results. and if someone wants to try it then, they can check out the total length of the video, and divide them into equal parts accordingly.  I don't know the specifics though, cause I have not tried video object detection before, I'm sorry about that.",
      "votes": null
    },
    {
      "id": "1596486",
      "postDate": "11/26/2021 14:41:05",
      "content": "<p>Great catch! I got the same idea when I checked the training data carefully. The main question is: can we expect that the test data will contains frames from one video in order? If yes, then in this case modified tracking algos (e.g.: IoU tracking) would help a lot.</p>",
      "rawMarkdown": "Great catch! I got the same idea when I checked the training data carefully. The main question is: can we expect that the test data will contains frames from one video in order? If yes, then in this case modified tracking algos (e.g.: IoU tracking) would help a lot.",
      "votes": null
    },
    {
      "id": "1596555",
      "postDate": "11/26/2021 16:11:29",
      "content": "<p>In that competition, test results must be submitted via a special API that feeds data frame by frame in the correct order. </p>",
      "rawMarkdown": "In that competition, test results must be submitted via a special API that feeds data frame by frame in the correct order.",
      "votes": null
    },
    {
      "id": "1596699",
      "postDate": "11/26/2021 18:31:38",
      "content": "<p>Yeah as <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a> said,  this is from the data tab,<br>\n<code>This competition uses a hidden test set that will be served by an API to ensure you evaluate the images in the same order they were recorded within each video.</code></p>",
      "rawMarkdown": "Yeah as @meowmeowmeowmeowmeow said,  this is from the data tab,\n`This competition uses a hidden test set that will be served by an API to ensure you evaluate the images in the same order they were recorded within each video. `",
      "votes": null
    },
    {
      "id": "1596730",
      "postDate": "11/26/2021 19:17:55",
      "content": "<p>Well, turns out I was wrong that \"…as the tow passes over a target, we can assume that the target will shift, over the course of n frames, not much in x but increasing in y.\"</p>\n<p>I plotted the bounding boxes from video 0, frames 16-32 (lines 18-34 in <code>train.csv</code>):</p>\n<p><img src=\"https://imgur.com/a/MqU2PjI\" alt=\"2 plots, showing shift in X and Y\"></p>\n<p>There was actually a significant shift in X and the shift in Y does not increase every time. </p>\n<p>Why I was wrong: as the FOV moves, anything not in the very center is going to shift in X and Y due to both the camera's motion in Z, but because of the angle in the viewing frustrum. Objects to either side of the vertical centerline will shift further to the side as the camera passes by, objects well above the horizontal centerline may initially shift up in Y as distance decreases. </p>",
      "rawMarkdown": "Well, turns out I was wrong that \"...as the tow passes over a target, we can assume that the target will shift, over the course of n frames, not much in x but increasing in y.\"\n\nI plotted the bounding boxes from video 0, frames 16-32 (lines 18-34 in `train.csv`):\n\n![2 plots, showing shift in X and Y](https://imgur.com/a/MqU2PjI)\n\nThere was actually a significant shift in X and the shift in Y does not increase every time. \n\nWhy I was wrong: as the FOV moves, anything not in the very center is going to shift in X and Y due to both the camera's motion in Z, but because of the angle in the viewing frustrum. Objects to either side of the vertical centerline will shift further to the side as the camera passes by, objects well above the horizontal centerline may initially shift up in Y as distance decreases.",
      "votes": null
    },
    {
      "id": "1598699",
      "postDate": "11/28/2021 17:49:29",
      "content": "<p>Oh, I see, thank you for explaining <a href=\"https://www.kaggle.com/lobrien\" target=\"_blank\">@lobrien</a>, I guess my interpretation was wrong.<br>\nI'm really curious what is this <code>tow</code> that you are referring to?</p>",
      "rawMarkdown": "Oh, I see, thank you for explaining @lobrien, I guess my interpretation was wrong.\nI'm really curious what is this `tow` that you are referring to?",
      "votes": null
    },
    {
      "id": "1599051",
      "postDate": "11/29/2021 04:58:44",
      "content": "<p>I mean the camera that is being towed. While it is likely true that every frame is basically the reef \"passing under\" the towed camera, you can see from my graph that the region of interest is not just \"pretty much constant in X, uniformly increasing in Y.\"</p>",
      "rawMarkdown": "I mean the camera that is being towed. While it is likely true that every frame is basically the reef \"passing under\" the towed camera, you can see from my graph that the region of interest is not just \"pretty much constant in X, uniformly increasing in Y.\"",
      "votes": null
    },
    {
      "id": "1599347",
      "postDate": "11/29/2021 11:10:32",
      "content": "<p>This is natural that BB coordinates change linearly. Moreover, size of BB changes as well. This is due to camera optical distortion. But anyway, we should expect COTS in the same region (with some padding of course) on the next frame. </p>\n<p>Can we try to feed the expected region as an additional ROI to the FasterRCNN, for example?</p>",
      "rawMarkdown": "This is natural that BB coordinates change linearly. Moreover, size of BB changes as well. This is due to camera optical distortion. But anyway, we should expect COTS in the same region (with some padding of course) on the next frame. \n\nCan we try to feed the expected region as an additional ROI to the FasterRCNN, for example?",
      "votes": null
    },
    {
      "id": "1599645",
      "postDate": "11/29/2021 16:41:00",
      "content": "<p>Yes, I'd like to try something like that, but whether to do it manually or with something like a TimeDistributed layer, I don't know. I'm concerned about the memory cost of using TimeDistributed or Conv3D since performance is a goal of the contest.</p>",
      "rawMarkdown": "Yes, I'd like to try something like that, but whether to do it manually or with something like a TimeDistributed layer, I don't know. I'm concerned about the memory cost of using TimeDistributed or Conv3D since performance is a goal of the contest.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1593810,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "11/24/2021 09:43:00",
      "content": "<p>How long sequence you are going to use in one sample?<br>\nHere we have a limited amount of video. Will be there a high chance of overfitting?</p>\n<p>In addition to your brilliant idea - how about Inflated 3D convolution (I am not familiar with transformed yet, sorry). May it be helpful to perform convolution by width/height and time as well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1593816,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "11/24/2021 09:48:07",
          "content": "<p>yeah, I guess it is a very bad idea to go with, so many problems, sorry for posting it. i guess its not much helpful taking sequence information into consideration after all. </p>\n<p>it's just my opinion though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1593825,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "11/24/2021 09:53:49",
          "content": "<p>Hey, mate. I believe your approach is good. Anyway, thank you for posting your thoughts!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1594499,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "11/24/2021 21:59:13",
          "content": "<p>hey <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a> , your point about the less data is true, and I also saw that there are many empty annotations in the dataset, so using a pre-trained model might yield some good results. and if someone wants to try it then, they can check out the total length of the video, and divide them into equal parts accordingly.  I don't know the specifics though, cause I have not tried video object detection before, I'm sorry about that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1594332,
      "author_name": "lobrien",
      "author_url": "",
      "post_date": "11/24/2021 17:52:23",
      "content": "<p>I strongly agree that using multiple frames is likely to be fruitful. Another point I’d make is that, unlike general challenges with object tracking, the COTS are (essentially) stationery and the tow is (relatively) constant and unidirectional. So as the tow passes over a target, we can assume that the target will shift, over the course of <code>n</code> frames, not much in <code>x</code> but increasing in <code>y</code>. </p>\n<p>That might be projecting too much heuristic thinking — “just let the model decide” — but given the desire for efficiency, there <em>is</em> a part of me that is looking for a solution that takes advantage of that specific aspect of the domain.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1594489,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "11/24/2021 21:51:30",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/lobrien\" target=\"_blank\">@lobrien</a> , I really appreciate your thoughtfulness. <br>\n<code>So as the tow passes over a target, we can assume that the target will shift, over the course of n frames, not much in x but increasing in y.</code> <br>\nby that, you mean that over the course of n frames right after passing a target, the next target would be a bit higher[increase in y] compared to the previous target but the difference in x won't be much compared to y? </p>\n<ul>\n<li>This might be true cause I have observed in the data that there are many times when I saw that the camera angle was from the top[not straight top though], so we are getting a tilted top view, like looking at something at the bottom from a tilted top. And when you see something from the top, then the next element looks a bit further than the previous element. but at the same time, the next element might be at the same x, this will happen if the targets happen to be in a row.</li>\n</ul>\n<p>\n<img src=\"https://i.imgur.com/3vPdBlK.png\">\n</p>\n\n<p>My interpretation might be wrong though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1596730,
          "author_name": "lobrien",
          "author_url": "",
          "post_date": "11/26/2021 19:17:55",
          "content": "<p>Well, turns out I was wrong that \"…as the tow passes over a target, we can assume that the target will shift, over the course of n frames, not much in x but increasing in y.\"</p>\n<p>I plotted the bounding boxes from video 0, frames 16-32 (lines 18-34 in <code>train.csv</code>):</p>\n<p><img src=\"https://imgur.com/a/MqU2PjI\" alt=\"2 plots, showing shift in X and Y\"></p>\n<p>There was actually a significant shift in X and the shift in Y does not increase every time. </p>\n<p>Why I was wrong: as the FOV moves, anything not in the very center is going to shift in X and Y due to both the camera's motion in Z, but because of the angle in the viewing frustrum. Objects to either side of the vertical centerline will shift further to the side as the camera passes by, objects well above the horizontal centerline may initially shift up in Y as distance decreases. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1598699,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "11/28/2021 17:49:29",
          "content": "<p>Oh, I see, thank you for explaining <a href=\"https://www.kaggle.com/lobrien\" target=\"_blank\">@lobrien</a>, I guess my interpretation was wrong.<br>\nI'm really curious what is this <code>tow</code> that you are referring to?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1599051,
          "author_name": "lobrien",
          "author_url": "",
          "post_date": "11/29/2021 04:58:44",
          "content": "<p>I mean the camera that is being towed. While it is likely true that every frame is basically the reef \"passing under\" the towed camera, you can see from my graph that the region of interest is not just \"pretty much constant in X, uniformly increasing in Y.\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1599347,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "11/29/2021 11:10:32",
          "content": "<p>This is natural that BB coordinates change linearly. Moreover, size of BB changes as well. This is due to camera optical distortion. But anyway, we should expect COTS in the same region (with some padding of course) on the next frame. </p>\n<p>Can we try to feed the expected region as an additional ROI to the FasterRCNN, for example?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1599645,
          "author_name": "lobrien",
          "author_url": "",
          "post_date": "11/29/2021 16:41:00",
          "content": "<p>Yes, I'd like to try something like that, but whether to do it manually or with something like a TimeDistributed layer, I don't know. I'm concerned about the memory cost of using TimeDistributed or Conv3D since performance is a goal of the contest.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1596486,
      "author_name": "benkelaci",
      "author_url": "",
      "post_date": "11/26/2021 14:41:05",
      "content": "<p>Great catch! I got the same idea when I checked the training data carefully. The main question is: can we expect that the test data will contains frames from one video in order? If yes, then in this case modified tracking algos (e.g.: IoU tracking) would help a lot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1596555,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "11/26/2021 16:11:29",
          "content": "<p>In that competition, test results must be submitted via a special API that feeds data frame by frame in the correct order. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1596699,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "11/26/2021 18:31:38",
          "content": "<p>Yeah as <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a> said,  this is from the data tab,<br>\n<code>This competition uses a hidden test set that will be served by an API to ensure you evaluate the images in the same order they were recorded within each video.</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1593743": "<p align=\"center\">\n<img width = \"700\" src=\"https://i.imgur.com/tTOe7cV.png\">\n</p>\n<h1 style=\"text-align: center; font-family: Verdana; font-size: 32px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; font-variant: small-caps; letter-spacing: 3px; color: #468282; background-color: #ffffff;\">Using Video Object Detection [Sequential Model + Object detection], because the images are in a video sequence</h1>\n***Howdy!!!*** Seems like **Patrick** is lost, and  **SpongeBob** and others need our help to find him. I was discussing with my friends about different ideas we can use to make a better model(and find Patrick). We came up with an idea that is **`using a kind of Sequential model like LSTM or RNN to capture the sequential information and then do Object detection.`**  **`The Idea is similar to video object detection`**. As most people are doing only object detection, it will be great to throw a Sequential model with object detection and see if that can help or not. *`It is like taking the context of the previous frame and using it to predict the object's position for the next frame.`*\n\nIf you see the video from the **Overview** tab, it also mentions that model related to time series. And one picture from the video, \n<p align=\"center\">\n<img width = \"900\" src=\"https://i.imgur.com/XesZdt4.png\">\n</p>\n\nalso shows[not sure that info is right or wrong] that starfishes borns in groups like you might not find a single(or only 2 or 3) starfish in the sea, it means if you see a starfish that means you will find starfish in some nearby areas too. **So, taking this into context and using the sequential image data[kind of like a video] it would be great to create a video object detection model that can detect starfishes(and help find Patrick).**\n\n**`It also opens a lot of opportunities for attention and transformer-based models`**[DART might do really well]. I don't have experience in video object detection but here are some Benchmark papers on video object detection that i could find,\n\n\t\n- [Mining Inter-Video Proposal Relations for Video Object Detection](https://paperswithcode.com/sota/video-object-detection-on-imagenet-vid)\n- [Emerging Properties in Self-Supervised Vision Transformers](https://paperswithcode.com/sota/video-object-detection-on-davis-2017)\n\n- [Temporal RoI Align for Video Object Recognition](https://paperswithcode.com/sota/video-object-detection-on-epic-kitchens-seen)\n\t\n- [Temporal RoI Align for Video Object Recognition](https://paperswithcode.com/sota/video-object-detection-on-epic-kitchens-1)\n\n\nCheck more here -> https://paperswithcode.com/task/video-object-detection\n<!--\nI am trying to make FasterRCNN work using Detectron2 and do some EDA [its not done yet though],  check out my NB here,\n- [Finding Patrick w/ Detectron2 fasterRCNN + EDA](https://www.kaggle.com/soumya9977/finding-patrick-w-detectron2-fasterrcnn-eda)\n-->\nI will try to implement a model based on this idea. Would like to hear your thoughts on this, and any point or advice you want to give😇. Happy kaggling, and **Thanksgiving**!",
    "1593810": "How long sequence you are going to use in one sample?\nHere we have a limited amount of video. Will be there a high chance of overfitting?\n\nIn addition to your brilliant idea - how about Inflated 3D convolution (I am not familiar with transformed yet, sorry). May it be helpful to perform convolution by width/height and time as well?",
    "1593816": "yeah, I guess it is a very bad idea to go with, so many problems, sorry for posting it. i guess its not much helpful taking sequence information into consideration after all. \n\nit's just my opinion though.",
    "1593825": "Hey, mate. I believe your approach is good. Anyway, thank you for posting your thoughts!",
    "1594332": "I strongly agree that using multiple frames is likely to be fruitful. Another point I’d make is that, unlike general challenges with object tracking, the COTS are (essentially) stationery and the tow is (relatively) constant and unidirectional. So as the tow passes over a target, we can assume that the target will shift, over the course of `n` frames, not much in `x` but increasing in `y`. \n\nThat might be projecting too much heuristic thinking — “just let the model decide” — but given the desire for efficiency, there _is_ a part of me that is looking for a solution that takes advantage of that specific aspect of the domain.",
    "1594489": "Thank you @lobrien , I really appreciate your thoughtfulness. \n`So as the tow passes over a target, we can assume that the target will shift, over the course of n frames, not much in x but increasing in y.` \nby that, you mean that over the course of n frames right after passing a target, the next target would be a bit higher[increase in y] compared to the previous target but the difference in x won't be much compared to y? \n\n- This might be true cause I have observed in the data that there are many times when I saw that the camera angle was from the top[not straight top though], so we are getting a tilted top view, like looking at something at the bottom from a tilted top. And when you see something from the top, then the next element looks a bit further than the previous element. but at the same time, the next element might be at the same x, this will happen if the targets happen to be in a row.\n\n\n<p align=\"center\">\n<img src=\"https://i.imgur.com/3vPdBlK.png\">\n</p>\n\n<!--\n|                                                                                 |\n|                                                                                 |\n|                     0   -> t2                                               |              0  -> t2\n|                                                                                 |\n|       o  -> t1                                                               |              o  -> t1\n|___________________________                                      |____________________________\n[Top view of two targets]\nnot in the same row, but y is increased                     y increased but x is the same\n-->\n\nMy interpretation might be wrong though.",
    "1594499": "hey @meowmeowmeowmeowmeow , your point about the less data is true, and I also saw that there are many empty annotations in the dataset, so using a pre-trained model might yield some good results. and if someone wants to try it then, they can check out the total length of the video, and divide them into equal parts accordingly.  I don't know the specifics though, cause I have not tried video object detection before, I'm sorry about that.",
    "1596486": "Great catch! I got the same idea when I checked the training data carefully. The main question is: can we expect that the test data will contains frames from one video in order? If yes, then in this case modified tracking algos (e.g.: IoU tracking) would help a lot.",
    "1596555": "In that competition, test results must be submitted via a special API that feeds data frame by frame in the correct order.",
    "1596699": "Yeah as @meowmeowmeowmeowmeow said,  this is from the data tab,\n`This competition uses a hidden test set that will be served by an API to ensure you evaluate the images in the same order they were recorded within each video. `",
    "1596730": "Well, turns out I was wrong that \"...as the tow passes over a target, we can assume that the target will shift, over the course of n frames, not much in x but increasing in y.\"\n\nI plotted the bounding boxes from video 0, frames 16-32 (lines 18-34 in `train.csv`):\n\n![2 plots, showing shift in X and Y](https://imgur.com/a/MqU2PjI)\n\nThere was actually a significant shift in X and the shift in Y does not increase every time. \n\nWhy I was wrong: as the FOV moves, anything not in the very center is going to shift in X and Y due to both the camera's motion in Z, but because of the angle in the viewing frustrum. Objects to either side of the vertical centerline will shift further to the side as the camera passes by, objects well above the horizontal centerline may initially shift up in Y as distance decreases.",
    "1598699": "Oh, I see, thank you for explaining @lobrien, I guess my interpretation was wrong.\nI'm really curious what is this `tow` that you are referring to?",
    "1599051": "I mean the camera that is being towed. While it is likely true that every frame is basically the reef \"passing under\" the towed camera, you can see from my graph that the region of interest is not just \"pretty much constant in X, uniformly increasing in Y.\"",
    "1599347": "This is natural that BB coordinates change linearly. Moreover, size of BB changes as well. This is due to camera optical distortion. But anyway, we should expect COTS in the same region (with some padding of course) on the next frame. \n\nCan we try to feed the expected region as an additional ROI to the FasterRCNN, for example?",
    "1599645": "Yes, I'd like to try something like that, but whether to do it manually or with something like a TimeDistributed layer, I don't know. I'm concerned about the memory cost of using TimeDistributed or Conv3D since performance is a goal of the contest."
  },
  "source": "meta"
}