{
  "id": 290656,
  "title": "Thoughts on Conv3D vs TimeDistributed(Conv2D)? ",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/290656",
  "author_name": "",
  "post_date": "2021-11-25T16:37:01.596129900Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I’m guessing that success is going to require multiframe inference rather than just per-image work. It looks to me like Keras has 2 pre-written layers that are natural fits for working with video sequences: <code>Conv3D</code> is the obvious match but there’s also the <code>TimeDistributed</code> layer, which could be used to wrap a <code>Conv2D</code>.  The temptation with <code>TimeDistributed</code> is that it seems like it might be easier to integrate with known object detection architectures. The temptation with a <code>Conv3D</code>-based architecture is that it feels like there’s better affinity between the features. </p>\n<p>Any experiences or thoughts with these options or others?</p>",
  "messages": [
    {
      "id": "1595384",
      "postDate": "11/25/2021 16:37:01",
      "content": "<p>I’m guessing that success is going to require multiframe inference rather than just per-image work. It looks to me like Keras has 2 pre-written layers that are natural fits for working with video sequences: <code>Conv3D</code> is the obvious match but there’s also the <code>TimeDistributed</code> layer, which could be used to wrap a <code>Conv2D</code>.  The temptation with <code>TimeDistributed</code> is that it seems like it might be easier to integrate with known object detection architectures. The temptation with a <code>Conv3D</code>-based architecture is that it feels like there’s better affinity between the features. </p>\n<p>Any experiences or thoughts with these options or others?</p>",
      "rawMarkdown": "I’m guessing that success is going to require multiframe inference rather than just per-image work. It looks to me like Keras has 2 pre-written layers that are natural fits for working with video sequences: `Conv3D` is the obvious match but there’s also the `TimeDistributed` layer, which could be used to wrap a `Conv2D`.  The temptation with `TimeDistributed` is that it seems like it might be easier to integrate with known object detection architectures. The temptation with a `Conv3D`-based architecture is that it feels like there’s better affinity between the features. \n\nAny experiences or thoughts with these options or others?",
      "votes": null
    },
    {
      "id": "1595422",
      "postDate": "11/25/2021 17:17:16",
      "content": "<p>What about <code>Conv3D</code> - I think that <code>Conv3D</code> will information about multiple frames. My intuition says me that it would be useful for predicting one thing in multiple frames (e.g. some video action). Would it be useful in our case when we need to predict boxes per frame?<br>\nOr may we pass only several subsequence frames to the <code>Conv3D</code>? In such a case ground truth bounding boxes are not changed dramatically. But frames are very similar as well.<br>\nAny thought about it?</p>",
      "rawMarkdown": "What about `Conv3D` - I think that `Conv3D` will information about multiple frames. My intuition says me that it would be useful for predicting one thing in multiple frames (e.g. some video action). Would it be useful in our case when we need to predict boxes per frame?\nOr may we pass only several subsequence frames to the `Conv3D`? In such a case ground truth bounding boxes are not changed dramatically. But frames are very similar as well.\nAny thought about it?",
      "votes": null
    },
    {
      "id": "1595705",
      "postDate": "11/25/2021 23:20:52",
      "content": "<p>Well, the thing is that ROIs <em>do</em> shift over frames as the tow passes over the reef. I was thinking of doing some of that visualization today, at least for those sequences that do not have many ROIs in the same frame (there’s one sequence that has 18 COTS in the image). Once you have multiple or overlapping ROIs, the difficult part of object tracking kicks in. </p>",
      "rawMarkdown": "Well, the thing is that ROIs _do_ shift over frames as the tow passes over the reef. I was thinking of doing some of that visualization today, at least for those sequences that do not have many ROIs in the same frame (there’s one sequence that has 18 COTS in the image). Once you have multiple or overlapping ROIs, the difficult part of object tracking kicks in.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1595422,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "11/25/2021 17:17:16",
      "content": "<p>What about <code>Conv3D</code> - I think that <code>Conv3D</code> will information about multiple frames. My intuition says me that it would be useful for predicting one thing in multiple frames (e.g. some video action). Would it be useful in our case when we need to predict boxes per frame?<br>\nOr may we pass only several subsequence frames to the <code>Conv3D</code>? In such a case ground truth bounding boxes are not changed dramatically. But frames are very similar as well.<br>\nAny thought about it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1595705,
          "author_name": "lobrien",
          "author_url": "",
          "post_date": "11/25/2021 23:20:52",
          "content": "<p>Well, the thing is that ROIs <em>do</em> shift over frames as the tow passes over the reef. I was thinking of doing some of that visualization today, at least for those sequences that do not have many ROIs in the same frame (there’s one sequence that has 18 COTS in the image). Once you have multiple or overlapping ROIs, the difficult part of object tracking kicks in. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1595384": "I’m guessing that success is going to require multiframe inference rather than just per-image work. It looks to me like Keras has 2 pre-written layers that are natural fits for working with video sequences: `Conv3D` is the obvious match but there’s also the `TimeDistributed` layer, which could be used to wrap a `Conv2D`.  The temptation with `TimeDistributed` is that it seems like it might be easier to integrate with known object detection architectures. The temptation with a `Conv3D`-based architecture is that it feels like there’s better affinity between the features. \n\nAny experiences or thoughts with these options or others?",
    "1595422": "What about `Conv3D` - I think that `Conv3D` will information about multiple frames. My intuition says me that it would be useful for predicting one thing in multiple frames (e.g. some video action). Would it be useful in our case when we need to predict boxes per frame?\nOr may we pass only several subsequence frames to the `Conv3D`? In such a case ground truth bounding boxes are not changed dramatically. But frames are very similar as well.\nAny thought about it?",
    "1595705": "Well, the thing is that ROIs _do_ shift over frames as the tow passes over the reef. I was thinking of doing some of that visualization today, at least for those sequences that do not have many ROIs in the same frame (there’s one sequence that has 18 COTS in the image). Once you have multiple or overlapping ROIs, the difficult part of object tracking kicks in."
  },
  "source": "meta"
}