{
  "id": 307121,
  "title": "Image size greater than native resolution: theory?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/307121",
  "author_name": "",
  "post_date": "2022-02-12T17:35:21.145970Z",
  "votes": 28,
  "comment_count": 9,
  "views": 0,
  "content": "<p>It seems like one aspect of higher-scoring kernels was an image size larger than native resolution. The native resolution was 1280x720, but apparently lots of people trained with 3K, 6K, or even 10K and had good success. Why would this work?</p>\n<p>If you upscale an image, you don't gain any information. I would think upscaling would actually work against a CNN because you're just making the first several sets of layers work with a blurry average of the original pixels. What am I missing?</p>",
  "messages": [
    {
      "id": "1687254",
      "postDate": "02/12/2022 17:35:21",
      "content": "<p>It seems like one aspect of higher-scoring kernels was an image size larger than native resolution. The native resolution was 1280x720, but apparently lots of people trained with 3K, 6K, or even 10K and had good success. Why would this work?</p>\n<p>If you upscale an image, you don't gain any information. I would think upscaling would actually work against a CNN because you're just making the first several sets of layers work with a blurry average of the original pixels. What am I missing?</p>",
      "rawMarkdown": "It seems like one aspect of higher-scoring kernels was an image size larger than native resolution. The native resolution was 1280x720, but apparently lots of people trained with 3K, 6K, or even 10K and had good success. Why would this work?\n\nIf you upscale an image, you don't gain any information. I would think upscaling would actually work against a CNN because you're just making the first several sets of layers work with a blurry average of the original pixels. What am I missing?",
      "votes": null
    },
    {
      "id": "1687440",
      "postDate": "02/12/2022 21:28:21",
      "content": "<p>In Yolov5, minimum anchor stride is fixed in transformed coordinates depending on the level of extracted features of the network [8, 16, 32, 64]. So if you have more features with higher resolution, the stride in native space coordinates is actually smaller which can improve recall. If mininum stride is too large according to the size of objects, you can miss objects in between the stride interval.</p>\n<p>Also, depending on the training hyperparameters, the default anchors in Yolov5 (which are fixed even if you change resolution) can be more appropriate with higher resolution for the size of the objects in the dataset (and/or in the public LB subset). The default metric parameter (in autoanchors.py) to evaluate the best possible recall associated with the default anchors is not strict enough for this dataset. Consequently, new anchors are never automatically generated with the current code.</p>\n<p>P.S. I have no idea why you got downvotes for your question because that is seriously a very interesting question.</p>",
      "rawMarkdown": "In Yolov5, minimum anchor stride is fixed in transformed coordinates depending on the level of extracted features of the network [8, 16, 32, 64]. So if you have more features with higher resolution, the stride in native space coordinates is actually smaller which can improve recall. If mininum stride is too large according to the size of objects, you can miss objects in between the stride interval.\n\nAlso, depending on the training hyperparameters, the default anchors in Yolov5 (which are fixed even if you change resolution) can be more appropriate with higher resolution for the size of the objects in the dataset (and/or in the public LB subset). The default metric parameter (in autoanchors.py) to evaluate the best possible recall associated with the default anchors is not strict enough for this dataset. Consequently, new anchors are never automatically generated with the current code.\n\nP.S. I have no idea why you got downvotes for your question because that is seriously a very interesting question.",
      "votes": null
    },
    {
      "id": "1687636",
      "postDate": "02/13/2022 02:58:52",
      "content": "<p><a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a> Response is really good here. </p>\n<p>I would add that I wouldn't necessarily count image interpolation as working against a CNN in all instances. In a sense the interpolation of an image at a larger size is still new information to the network even if it isn't to us because we know contextually the process. </p>",
      "rawMarkdown": "alexandrecc Response is really good here. \n\nI would add that I wouldn't necessarily count image interpolation as working against a CNN in all instances. In a sense the interpolation of an image at a larger size is still new information to the network even if it isn't to us because we know contextually the process.",
      "votes": null
    },
    {
      "id": "1687665",
      "postDate": "02/13/2022 03:32:58",
      "content": "<p>Have wondered this as well.  From the organisers' paper <a href=\"https://arxiv.org/abs/2111.14311\" target=\"_blank\">https://arxiv.org/abs/2111.14311</a></p>\n<p>\"We set the GoPro cameras to record videos continuously at 24 frames per second at 3840x2160 resolution and manually removed the periods of no activity between transects.\" </p>\n<p>So original images were 3x the competition jpegs. The private test set is delivered as .npy files so not really sure what the res is on them originally but imagine is the same.  Thought if you compress an image to a smaller size then upscale it, that would be a better image than if you are just upscaling an image from its original smaller resolution.  But might depend on its jpeq quality factor.  </p>",
      "rawMarkdown": "Have wondered this as well.  From the organisers' paper https://arxiv.org/abs/2111.14311\n\n\"We set the GoPro cameras to record videos continuously at 24 frames per second at 3840x2160 resolution and manually removed the periods of no activity between transects.\" \n\nSo original images were 3x the competition jpegs. The private test set is delivered as .npy files so not really sure what the res is on them originally but imagine is the same.  Thought if you compress an image to a smaller size then upscale it, that would be a better image than if you are just upscaling an image from its original smaller resolution.  But might depend on its jpeq quality factor.",
      "votes": null
    },
    {
      "id": "1687679",
      "postDate": "02/13/2022 04:01:04",
      "content": "<p>why microscope exists?  It depends on targets size and distribution.</p>",
      "rawMarkdown": "why microscope exists?  It depends on targets size and distribution.",
      "votes": null
    },
    {
      "id": "1687960",
      "postDate": "02/13/2022 09:09:42",
      "content": "<p>Based on upvotes, a lot of people agreeing with this? Original question was about \"trained with 3K, 6K, or even 10K and had good success\" and train and test data should really hit some criteria for your explanation to work well (because improving recall comes with some drawbacks). I dont think it is true in general sense, you cant say go as high-res as you can. While of course there is still techniques like ms-training and ms-testing, but i never saw x10 ms-anything </p>",
      "rawMarkdown": "Based on upvotes, a lot of people agreeing with this? Original question was about \"trained with 3K, 6K, or even 10K and had good success\" and train and test data should really hit some criteria for your explanation to work well (because improving recall comes with some drawbacks). I dont think it is true in general sense, you cant say go as high-res as you can. While of course there is still techniques like ms-training and ms-testing, but i never saw x10 ms-anything",
      "votes": null
    },
    {
      "id": "1688974",
      "postDate": "02/13/2022 23:52:49",
      "content": "<p>We have to separate this into two question:</p>\n<ol>\n<li>Why higher resolution training result in higher score in LB? (e.g. train x3, infer x3)</li>\n<li>Why magnifying infer scale greater than training scale result in higher LB score? (e.g. train x3, infer x6 time)</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a> have explained about the first question, but the reasoning of second question is still unknown.</p>\n<p>One intuitive explanation is that the distribution of target size of the test set is smaller than that of train set. However, we can't know the truth unless the host publishes the hidden test set.</p>",
      "rawMarkdown": "We have to separate this into two question:\n\n1. Why higher resolution training result in higher score in LB? (e.g. train x3, infer x3)\n2. Why magnifying infer scale greater than training scale result in higher LB score? (e.g. train x3, infer x6 time)\n\n@alexandrecc have explained about the first question, but the reasoning of second question is still unknown.\n\nOne intuitive explanation is that the distribution of target size of the test set is smaller than that of train set. However, we can't know the truth unless the host publishes the hidden test set.",
      "votes": null
    },
    {
      "id": "1689004",
      "postDate": "02/14/2022 00:46:39",
      "content": "<p>Maybe I was not clear enough but both the minimum anchor stride and the anchor sizes can also explain the upscaling benefit during inference if there are a lot more smaller cots in public LB than the training set. That was of course an implicit assumption of the entire discussion.</p>",
      "rawMarkdown": "Maybe I was not clear enough but both the minimum anchor stride and the anchor sizes can also explain the upscaling benefit during inference if there are a lot more smaller cots in public LB than the training set. That was of course an implicit assumption of the entire discussion.",
      "votes": null
    },
    {
      "id": "1689021",
      "postDate": "02/14/2022 01:18:35",
      "content": "<blockquote>\n  <p>That was of course an implicit assumption of the entire discussion.</p>\n</blockquote>\n<p>I see. Yes, anchor size and stride explains the second question if the assumption is true.</p>\n<p>BTW, if the phenomenon is partly based on the fitting problem of anchors, I wonder if this phenomenon also reproduces when using the non anchor-based model like YOLOX.</p>",
      "rawMarkdown": "> That was of course an implicit assumption of the entire discussion.\n\nI see. Yes, anchor size and stride explains the second question if the assumption is true.\n\nBTW, if the phenomenon is partly based on the fitting problem of anchors, I wonder if this phenomenon also reproduces when using the non anchor-based model like YOLOX.",
      "votes": null
    },
    {
      "id": "1689095",
      "postDate": "02/14/2022 03:14:41",
      "content": "<p>Bigger image size not always works. For my experiments, I use DBSCAN to get my anchors , and train model in size 2016x3584 and infer in the same size(0.577), the fact is that x2(3600*6400) infer size will get a nice score(0.632), if I increase the size to x2.2 or more times, the LB score will decrease. So LB score will not always increase with resolution, but is related with your anchor and size in training.<br>\nAs why the scores of sheep's model will increase with the resolution increasing, I think that may the sheep's model take the yolov5 default anchor size, which can adapt in a bigger range.</p>",
      "rawMarkdown": "Bigger image size not always works. For my experiments, I use DBSCAN to get my anchors , and train model in size 2016x3584 and infer in the same size(0.577), the fact is that x2(3600*6400) infer size will get a nice score(0.632), if I increase the size to x2.2 or more times, the LB score will decrease. So LB score will not always increase with resolution, but is related with your anchor and size in training.\nAs why the scores of sheep's model will increase with the resolution increasing, I think that may the sheep's model take the yolov5 default anchor size, which can adapt in a bigger range.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1687440,
      "author_name": "alexandrecc",
      "author_url": "",
      "post_date": "02/12/2022 21:28:21",
      "content": "<p>In Yolov5, minimum anchor stride is fixed in transformed coordinates depending on the level of extracted features of the network [8, 16, 32, 64]. So if you have more features with higher resolution, the stride in native space coordinates is actually smaller which can improve recall. If mininum stride is too large according to the size of objects, you can miss objects in between the stride interval.</p>\n<p>Also, depending on the training hyperparameters, the default anchors in Yolov5 (which are fixed even if you change resolution) can be more appropriate with higher resolution for the size of the objects in the dataset (and/or in the public LB subset). The default metric parameter (in autoanchors.py) to evaluate the best possible recall associated with the default anchors is not strict enough for this dataset. Consequently, new anchors are never automatically generated with the current code.</p>\n<p>P.S. I have no idea why you got downvotes for your question because that is seriously a very interesting question.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1687960,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "02/13/2022 09:09:42",
          "content": "<p>Based on upvotes, a lot of people agreeing with this? Original question was about \"trained with 3K, 6K, or even 10K and had good success\" and train and test data should really hit some criteria for your explanation to work well (because improving recall comes with some drawbacks). I dont think it is true in general sense, you cant say go as high-res as you can. While of course there is still techniques like ms-training and ms-testing, but i never saw x10 ms-anything </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1688974,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "02/13/2022 23:52:49",
          "content": "<p>We have to separate this into two question:</p>\n<ol>\n<li>Why higher resolution training result in higher score in LB? (e.g. train x3, infer x3)</li>\n<li>Why magnifying infer scale greater than training scale result in higher LB score? (e.g. train x3, infer x6 time)</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a> have explained about the first question, but the reasoning of second question is still unknown.</p>\n<p>One intuitive explanation is that the distribution of target size of the test set is smaller than that of train set. However, we can't know the truth unless the host publishes the hidden test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689004,
          "author_name": "alexandrecc",
          "author_url": "",
          "post_date": "02/14/2022 00:46:39",
          "content": "<p>Maybe I was not clear enough but both the minimum anchor stride and the anchor sizes can also explain the upscaling benefit during inference if there are a lot more smaller cots in public LB than the training set. That was of course an implicit assumption of the entire discussion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689021,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "02/14/2022 01:18:35",
          "content": "<blockquote>\n  <p>That was of course an implicit assumption of the entire discussion.</p>\n</blockquote>\n<p>I see. Yes, anchor size and stride explains the second question if the assumption is true.</p>\n<p>BTW, if the phenomenon is partly based on the fitting problem of anchors, I wonder if this phenomenon also reproduces when using the non anchor-based model like YOLOX.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1687636,
      "author_name": "taranmarley",
      "author_url": "",
      "post_date": "02/13/2022 02:58:52",
      "content": "<p><a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a> Response is really good here. </p>\n<p>I would add that I wouldn't necessarily count image interpolation as working against a CNN in all instances. In a sense the interpolation of an image at a larger size is still new information to the network even if it isn't to us because we know contextually the process. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1687665,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "02/13/2022 03:32:58",
      "content": "<p>Have wondered this as well.  From the organisers' paper <a href=\"https://arxiv.org/abs/2111.14311\" target=\"_blank\">https://arxiv.org/abs/2111.14311</a></p>\n<p>\"We set the GoPro cameras to record videos continuously at 24 frames per second at 3840x2160 resolution and manually removed the periods of no activity between transects.\" </p>\n<p>So original images were 3x the competition jpegs. The private test set is delivered as .npy files so not really sure what the res is on them originally but imagine is the same.  Thought if you compress an image to a smaller size then upscale it, that would be a better image than if you are just upscaling an image from its original smaller resolution.  But might depend on its jpeq quality factor.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1687679,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "02/13/2022 04:01:04",
      "content": "<p>why microscope exists?  It depends on targets size and distribution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1689095,
      "author_name": "freshair1996",
      "author_url": "",
      "post_date": "02/14/2022 03:14:41",
      "content": "<p>Bigger image size not always works. For my experiments, I use DBSCAN to get my anchors , and train model in size 2016x3584 and infer in the same size(0.577), the fact is that x2(3600*6400) infer size will get a nice score(0.632), if I increase the size to x2.2 or more times, the LB score will decrease. So LB score will not always increase with resolution, but is related with your anchor and size in training.<br>\nAs why the scores of sheep's model will increase with the resolution increasing, I think that may the sheep's model take the yolov5 default anchor size, which can adapt in a bigger range.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1687254": "It seems like one aspect of higher-scoring kernels was an image size larger than native resolution. The native resolution was 1280x720, but apparently lots of people trained with 3K, 6K, or even 10K and had good success. Why would this work?\n\nIf you upscale an image, you don't gain any information. I would think upscaling would actually work against a CNN because you're just making the first several sets of layers work with a blurry average of the original pixels. What am I missing?",
    "1687440": "In Yolov5, minimum anchor stride is fixed in transformed coordinates depending on the level of extracted features of the network [8, 16, 32, 64]. So if you have more features with higher resolution, the stride in native space coordinates is actually smaller which can improve recall. If mininum stride is too large according to the size of objects, you can miss objects in between the stride interval.\n\nAlso, depending on the training hyperparameters, the default anchors in Yolov5 (which are fixed even if you change resolution) can be more appropriate with higher resolution for the size of the objects in the dataset (and/or in the public LB subset). The default metric parameter (in autoanchors.py) to evaluate the best possible recall associated with the default anchors is not strict enough for this dataset. Consequently, new anchors are never automatically generated with the current code.\n\nP.S. I have no idea why you got downvotes for your question because that is seriously a very interesting question.",
    "1687636": "alexandrecc Response is really good here. \n\nI would add that I wouldn't necessarily count image interpolation as working against a CNN in all instances. In a sense the interpolation of an image at a larger size is still new information to the network even if it isn't to us because we know contextually the process.",
    "1687665": "Have wondered this as well.  From the organisers' paper https://arxiv.org/abs/2111.14311\n\n\"We set the GoPro cameras to record videos continuously at 24 frames per second at 3840x2160 resolution and manually removed the periods of no activity between transects.\" \n\nSo original images were 3x the competition jpegs. The private test set is delivered as .npy files so not really sure what the res is on them originally but imagine is the same.  Thought if you compress an image to a smaller size then upscale it, that would be a better image than if you are just upscaling an image from its original smaller resolution.  But might depend on its jpeq quality factor.",
    "1687679": "why microscope exists?  It depends on targets size and distribution.",
    "1687960": "Based on upvotes, a lot of people agreeing with this? Original question was about \"trained with 3K, 6K, or even 10K and had good success\" and train and test data should really hit some criteria for your explanation to work well (because improving recall comes with some drawbacks). I dont think it is true in general sense, you cant say go as high-res as you can. While of course there is still techniques like ms-training and ms-testing, but i never saw x10 ms-anything",
    "1688974": "We have to separate this into two question:\n\n1. Why higher resolution training result in higher score in LB? (e.g. train x3, infer x3)\n2. Why magnifying infer scale greater than training scale result in higher LB score? (e.g. train x3, infer x6 time)\n\n@alexandrecc have explained about the first question, but the reasoning of second question is still unknown.\n\nOne intuitive explanation is that the distribution of target size of the test set is smaller than that of train set. However, we can't know the truth unless the host publishes the hidden test set.",
    "1689004": "Maybe I was not clear enough but both the minimum anchor stride and the anchor sizes can also explain the upscaling benefit during inference if there are a lot more smaller cots in public LB than the training set. That was of course an implicit assumption of the entire discussion.",
    "1689021": "> That was of course an implicit assumption of the entire discussion.\n\nI see. Yes, anchor size and stride explains the second question if the assumption is true.\n\nBTW, if the phenomenon is partly based on the fitting problem of anchors, I wonder if this phenomenon also reproduces when using the non anchor-based model like YOLOX.",
    "1689095": "Bigger image size not always works. For my experiments, I use DBSCAN to get my anchors , and train model in size 2016x3584 and infer in the same size(0.577), the fact is that x2(3600*6400) infer size will get a nice score(0.632), if I increase the size to x2.2 or more times, the LB score will decrease. So LB score will not always increase with resolution, but is related with your anchor and size in training.\nAs why the scores of sheep's model will increase with the resolution increasing, I think that may the sheep's model take the yolov5 default anchor size, which can adapt in a bigger range."
  },
  "source": "meta"
}