{
  "id": 290456,
  "title": "Scale Invariance and \"Perspective Scaling\"",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/290456",
  "author_name": "",
  "post_date": "2021-11-24T18:22:19.546837800Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm not very familiar with Object Detection, so forgive the newbie questions.   The first question is:   What about scale?   Even if we can assume that COTS have a fixed size, they will appear to have different sizes in the video image due to the effects of perspective projection.   It seems to me that using CNNs in the feature extraction imposes a kind of scaling assumption that is implied by the size and stride of the convolution kernels, and thus may be able to detect COTS at a certain scale but not at others.   I understand that this can be addressed by tapping features at different points in a chain of convolutions and down-samples, but that would seem to require much additional processing.</p>\n<p>Second question:   What can we assume about the video camera that takes the video and its pose relative to the imaged objects?   Can we assume that the camera model does not change (i.e., the video is not taken through a zoom lens)?   Can we further assume that the 2D x-axis is more or less horizontal (i.e., subject to yaw and pitch but not roll)?   Finally, can we assume that in 3D, the reef surface is more or less horizontal?   </p>\n<p>What I'm thinking is perhaps we can determine a \"perspective scale factor\" at each point in the 2D image.   If we can answer the second questions affirmatively, then I think it's possible to do so, and that, in turn, could be used as another image-like input to the COTS detection problem, looking for \"smaller\" COTS near the top of the image and \"larger\" COTS near the bottom.</p>\n<p>Make any sense?   Or are the assumptions too restrictive?</p>\n<ul>\n<li>- - - - - - - - - - - - - - - - - - - - - - - Edit 27Nov2021  1:13pmEST - - - - - - - - - - - - - - - - - - - - -</li>\n</ul>\n<p>Well, the good news is that it turned out to be pretty easy to test the hypothesis.   I took all the bounding boxes in the training set, pulled out several hundred more or less at random, and for each, I plotted bounding box size as a function of image Y coordinate.   The results need a little<br>\nexplanation, since it appears that size <em>increases</em> rather than decreases with Y coordinate.   I believe this is because Y coordinates are inverted, with the largest at the bottom of the images and the smallest at the top.   So, the slope has the right sign but, as the plot shows, the correlation between Y coordinate and size is pretty weak:   the correlation coefficient is 0.25420354.   Sorry I couldn't post the plot, seems the Discussions preclude that.</p>",
  "messages": [
    {
      "id": "1594366",
      "postDate": "11/24/2021 18:22:19",
      "content": "<p>I'm not very familiar with Object Detection, so forgive the newbie questions.   The first question is:   What about scale?   Even if we can assume that COTS have a fixed size, they will appear to have different sizes in the video image due to the effects of perspective projection.   It seems to me that using CNNs in the feature extraction imposes a kind of scaling assumption that is implied by the size and stride of the convolution kernels, and thus may be able to detect COTS at a certain scale but not at others.   I understand that this can be addressed by tapping features at different points in a chain of convolutions and down-samples, but that would seem to require much additional processing.</p>\n<p>Second question:   What can we assume about the video camera that takes the video and its pose relative to the imaged objects?   Can we assume that the camera model does not change (i.e., the video is not taken through a zoom lens)?   Can we further assume that the 2D x-axis is more or less horizontal (i.e., subject to yaw and pitch but not roll)?   Finally, can we assume that in 3D, the reef surface is more or less horizontal?   </p>\n<p>What I'm thinking is perhaps we can determine a \"perspective scale factor\" at each point in the 2D image.   If we can answer the second questions affirmatively, then I think it's possible to do so, and that, in turn, could be used as another image-like input to the COTS detection problem, looking for \"smaller\" COTS near the top of the image and \"larger\" COTS near the bottom.</p>\n<p>Make any sense?   Or are the assumptions too restrictive?</p>\n<ul>\n<li>- - - - - - - - - - - - - - - - - - - - - - - Edit 27Nov2021  1:13pmEST - - - - - - - - - - - - - - - - - - - - -</li>\n</ul>\n<p>Well, the good news is that it turned out to be pretty easy to test the hypothesis.   I took all the bounding boxes in the training set, pulled out several hundred more or less at random, and for each, I plotted bounding box size as a function of image Y coordinate.   The results need a little<br>\nexplanation, since it appears that size <em>increases</em> rather than decreases with Y coordinate.   I believe this is because Y coordinates are inverted, with the largest at the bottom of the images and the smallest at the top.   So, the slope has the right sign but, as the plot shows, the correlation between Y coordinate and size is pretty weak:   the correlation coefficient is 0.25420354.   Sorry I couldn't post the plot, seems the Discussions preclude that.</p>",
      "rawMarkdown": "I'm not very familiar with Object Detection, so forgive the newbie questions.   The first question is:   What about scale?   Even if we can assume that COTS have a fixed size, they will appear to have different sizes in the video image due to the effects of perspective projection.   It seems to me that using CNNs in the feature extraction imposes a kind of scaling assumption that is implied by the size and stride of the convolution kernels, and thus may be able to detect COTS at a certain scale but not at others.   I understand that this can be addressed by tapping features at different points in a chain of convolutions and down-samples, but that would seem to require much additional processing.\n\nSecond question:   What can we assume about the video camera that takes the video and its pose relative to the imaged objects?   Can we assume that the camera model does not change (i.e., the video is not taken through a zoom lens)?   Can we further assume that the 2D x-axis is more or less horizontal (i.e., subject to yaw and pitch but not roll)?   Finally, can we assume that in 3D, the reef surface is more or less horizontal?   \n\nWhat I'm thinking is perhaps we can determine a \"perspective scale factor\" at each point in the 2D image.   If we can answer the second questions affirmatively, then I think it's possible to do so, and that, in turn, could be used as another image-like input to the COTS detection problem, looking for \"smaller\" COTS near the top of the image and \"larger\" COTS near the bottom.\n\nMake any sense?   Or are the assumptions too restrictive?\n\n- - - - - - - - - - - - - - - - - - - - - - - - Edit 27Nov2021  1:13pmEST - - - - - - - - - - - - - - - - - - - - -\n\nWell, the good news is that it turned out to be pretty easy to test the hypothesis.   I took all the bounding boxes in the training set, pulled out several hundred more or less at random, and for each, I plotted bounding box size as a function of image Y coordinate.   The results need a little\nexplanation, since it appears that size *increases* rather than decreases with Y coordinate.   I believe this is because Y coordinates are inverted, with the largest at the bottom of the images and the smallest at the top.   So, the slope has the right sign but, as the plot shows, the correlation between Y coordinate and size is pretty weak:   the correlation coefficient is 0.25420354.   Sorry I couldn't post the plot, seems the Discussions preclude that.",
      "votes": null
    },
    {
      "id": "1594425",
      "postDate": "11/24/2021 20:01:22",
      "content": "<p>Usually scaling problem in CNN is solved by FPN - exactly the same as what you mentioned - using features from different CNN layers. Different Conv layers have different perceptive fields.</p>\n<p>What about perspective transform - there was a nice paper in which a trainable perspective transform layer was introduced. I don't remember the name of that paper, but remember that it was about the OCR task (probably license plate recognition). That layer performs perspective transform and has trainable parameters of that transformation. But in that case, it would be really difficult to find perspective transform for specific are of the image. Take into account such facts, that we have to deal with uneven terrain. I suspect the it would be much easier to find BB on the raw images.</p>\n<p>It's my opinion. Would love to hear any other thoughts.</p>",
      "rawMarkdown": "Usually scaling problem in CNN is solved by FPN - exactly the same as what you mentioned - using features from different CNN layers. Different Conv layers have different perceptive fields.\n\nWhat about perspective transform - there was a nice paper in which a trainable perspective transform layer was introduced. I don't remember the name of that paper, but remember that it was about the OCR task (probably license plate recognition). That layer performs perspective transform and has trainable parameters of that transformation. But in that case, it would be really difficult to find perspective transform for specific are of the image. Take into account such facts, that we have to deal with uneven terrain. I suspect the it would be much easier to find BB on the raw images.\n\nIt's my opinion. Would love to hear any other thoughts.",
      "votes": null
    },
    {
      "id": "1595366",
      "postDate": "11/25/2021 16:17:33",
      "content": "<p>\"can we assume that in 3D, the reef surface is more or less horizontal?\" No. The reefs in the testing data have a very significant amount of verticality. There are some slopes, but mostly the verticality comes from coral “bommies” that rise from the seafloor. The reef also exhibits a huge amount of complexity, with many vertical surfaces, cracks, caves, overhangs, etc. The COTS roughly, but not entirely, conform to the surface on which they are crawling.</p>",
      "rawMarkdown": "\"can we assume that in 3D, the reef surface is more or less horizontal?\" No. The reefs in the testing data have a very significant amount of verticality. There are some slopes, but mostly the verticality comes from coral “bommies” that rise from the seafloor. The reef also exhibits a huge amount of complexity, with many vertical surfaces, cracks, caves, overhangs, etc. The COTS roughly, but not entirely, conform to the surface on which they are crawling.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1594425,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "11/24/2021 20:01:22",
      "content": "<p>Usually scaling problem in CNN is solved by FPN - exactly the same as what you mentioned - using features from different CNN layers. Different Conv layers have different perceptive fields.</p>\n<p>What about perspective transform - there was a nice paper in which a trainable perspective transform layer was introduced. I don't remember the name of that paper, but remember that it was about the OCR task (probably license plate recognition). That layer performs perspective transform and has trainable parameters of that transformation. But in that case, it would be really difficult to find perspective transform for specific are of the image. Take into account such facts, that we have to deal with uneven terrain. I suspect the it would be much easier to find BB on the raw images.</p>\n<p>It's my opinion. Would love to hear any other thoughts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1595366,
      "author_name": "lobrien",
      "author_url": "",
      "post_date": "11/25/2021 16:17:33",
      "content": "<p>\"can we assume that in 3D, the reef surface is more or less horizontal?\" No. The reefs in the testing data have a very significant amount of verticality. There are some slopes, but mostly the verticality comes from coral “bommies” that rise from the seafloor. The reef also exhibits a huge amount of complexity, with many vertical surfaces, cracks, caves, overhangs, etc. The COTS roughly, but not entirely, conform to the surface on which they are crawling.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1594366": "I'm not very familiar with Object Detection, so forgive the newbie questions.   The first question is:   What about scale?   Even if we can assume that COTS have a fixed size, they will appear to have different sizes in the video image due to the effects of perspective projection.   It seems to me that using CNNs in the feature extraction imposes a kind of scaling assumption that is implied by the size and stride of the convolution kernels, and thus may be able to detect COTS at a certain scale but not at others.   I understand that this can be addressed by tapping features at different points in a chain of convolutions and down-samples, but that would seem to require much additional processing.\n\nSecond question:   What can we assume about the video camera that takes the video and its pose relative to the imaged objects?   Can we assume that the camera model does not change (i.e., the video is not taken through a zoom lens)?   Can we further assume that the 2D x-axis is more or less horizontal (i.e., subject to yaw and pitch but not roll)?   Finally, can we assume that in 3D, the reef surface is more or less horizontal?   \n\nWhat I'm thinking is perhaps we can determine a \"perspective scale factor\" at each point in the 2D image.   If we can answer the second questions affirmatively, then I think it's possible to do so, and that, in turn, could be used as another image-like input to the COTS detection problem, looking for \"smaller\" COTS near the top of the image and \"larger\" COTS near the bottom.\n\nMake any sense?   Or are the assumptions too restrictive?\n\n- - - - - - - - - - - - - - - - - - - - - - - - Edit 27Nov2021  1:13pmEST - - - - - - - - - - - - - - - - - - - - -\n\nWell, the good news is that it turned out to be pretty easy to test the hypothesis.   I took all the bounding boxes in the training set, pulled out several hundred more or less at random, and for each, I plotted bounding box size as a function of image Y coordinate.   The results need a little\nexplanation, since it appears that size *increases* rather than decreases with Y coordinate.   I believe this is because Y coordinates are inverted, with the largest at the bottom of the images and the smallest at the top.   So, the slope has the right sign but, as the plot shows, the correlation between Y coordinate and size is pretty weak:   the correlation coefficient is 0.25420354.   Sorry I couldn't post the plot, seems the Discussions preclude that.",
    "1594425": "Usually scaling problem in CNN is solved by FPN - exactly the same as what you mentioned - using features from different CNN layers. Different Conv layers have different perceptive fields.\n\nWhat about perspective transform - there was a nice paper in which a trainable perspective transform layer was introduced. I don't remember the name of that paper, but remember that it was about the OCR task (probably license plate recognition). That layer performs perspective transform and has trainable parameters of that transformation. But in that case, it would be really difficult to find perspective transform for specific are of the image. Take into account such facts, that we have to deal with uneven terrain. I suspect the it would be much easier to find BB on the raw images.\n\nIt's my opinion. Would love to hear any other thoughts.",
    "1595366": "\"can we assume that in 3D, the reef surface is more or less horizontal?\" No. The reefs in the testing data have a very significant amount of verticality. There are some slopes, but mostly the verticality comes from coral “bommies” that rise from the seafloor. The reef also exhibits a huge amount of complexity, with many vertical surfaces, cracks, caves, overhangs, etc. The COTS roughly, but not entirely, conform to the surface on which they are crawling."
  },
  "source": "meta"
}