{
  "id": 262321,
  "title": "Increasing image size during inference gives a little boost. Why?",
  "url": "/competitions/seti-breakthrough-listen/discussion/262321",
  "author_name": "Parth Dhameliya",
  "post_date": "2021-08-06T10:42:57.203000",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>During inference I sometime increase the image size (eg. 512 used in training but 620 in inference). I don't know the reason why it gives a boost in LB. In many competitions I had seen many solution in which image size during inference is different than the image size used in train. I want to know the reason….why most of the time it increases the lb score.  Even sometimes scaling down works too ?</p>",
  "messages": [
    {
      "id": 1454830,
      "postDate": "2021-08-06T10:42:57.203Z",
      "content": "<p>During inference I sometime increase the image size (eg. 512 used in training but 620 in inference). I don't know the reason why it gives a boost in LB. In many competitions I had seen many solution in which image size during inference is different than the image size used in train. I want to know the reason….why most of the time it increases the lb score.  Even sometimes scaling down works too ?</p>",
      "rawMarkdown": "During inference I sometime increase the image size (eg. 512 used in training but 620 in inference). I don't know the reason why it gives a boost in LB. In many competitions I had seen many solution in which image size during inference is different than the image size used in train. I want to know the reason....why most of the time it increases the lb score.  Even sometimes scaling down works too ?",
      "votes": 7
    },
    {
      "id": 1455073,
      "postDate": "2021-08-06T12:46:41.117Z",
      "content": "<p>During training, do you use random resize augmentation? If so, then you are training on multiple sizes including 620. (Note that if you resize a 512 image to 620 and then crop a 512 square, that is still \"training on 620\" because the final product has the same pixel resolution, i.e. same pixel configuration as 620)</p>",
      "rawMarkdown": "During training, do you use random resize augmentation? If so, then you are training on multiple sizes including 620. (Note that if you resize a 512 image to 620 and then crop a 512 square, that is still \"training on 620\" because the final product has the same pixel resolution, i.e. same pixel configuration as 620)",
      "votes": 6,
      "replies": [
        {
          "id": 1455152,
          "postDate": "2021-08-06T13:16:58.357Z",
          "content": "<p>No my img size is fixed i.e 512, 512 during training and during inference I resize images to 600, 600 . No random resize augmentation during training. </p>",
          "rawMarkdown": "No my img size is fixed i.e 512, 512 during training and during inference I resize images to 600, 600 . No random resize augmentation during training. "
        },
        {
          "id": 1455230,
          "postDate": "2021-08-06T13:43:58.827Z",
          "content": "<p>Ok, so nothing like <code>albu.ShiftScaleRotate(rotate_limit=0, scale_limit=0.15, shift_limit=0, p=0.5)</code> ? Did you try random scale augmentation?</p>",
          "rawMarkdown": "Ok, so nothing like `albu.ShiftScaleRotate(rotate_limit=0, scale_limit=0.15, shift_limit=0, p=0.5)` ? Did you try random scale augmentation?"
        },
        {
          "id": 1455246,
          "postDate": "2021-08-06T13:49:52.103Z",
          "content": "<p>Yeah i am using shift scale rotate. This could be the reason ?</p>",
          "rawMarkdown": "Yeah i am using shift scale rotate. This could be the reason ?",
          "votes": 1
        },
        {
          "id": 1455260,
          "postDate": "2021-08-06T13:55:45.960Z",
          "content": "<p>Scale augmentation is \"random resize augmentation then crop to 512\". So you are effectively training your model on 384 images, 512 images, 620 images. </p>\n<p>So the answer to your question, is that you are training on multiple image sizes. So one of these image sizes will be the best during infernce.</p>\n<p>Note that you could even try TTA where you infer 384, 512, 620 and ensemble the 3 predictions.</p>",
          "rawMarkdown": "Scale augmentation is \"random resize augmentation then crop to 512\". So you are effectively training your model on 384 images, 512 images, 620 images. \n\nSo the answer to your question, is that you are training on multiple image sizes. So one of these image sizes will be the best during infernce.\n\nNote that you could even try TTA where you infer 384, 512, 620 and ensemble the 3 predictions.",
          "votes": 20
        },
        {
          "id": 1455319,
          "postDate": "2021-08-06T14:27:50.073Z",
          "content": "<p>This is enlightening. Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "rawMarkdown": "This is enlightening. Thanks @cdeotte ",
          "votes": 1
        },
        {
          "id": 1455378,
          "postDate": "2021-08-06T14:48:37.083Z",
          "content": "<p>Thank you so much👍😄  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "rawMarkdown": "Thank you so much👍😄  @cdeotte "
        }
      ]
    },
    {
      "id": 1456169,
      "postDate": "2021-08-06T20:02:06.427Z",
      "content": "<p>This is well-known fact, you could read more about it in the Facebook research papers:<br>\n<a href=\"https://arxiv.org/pdf/1906.06423.pdf\" target=\"_blank\">https://arxiv.org/pdf/1906.06423.pdf</a><br>\n<a href=\"https://arxiv.org/pdf/2003.08237.pdf\" target=\"_blank\">https://arxiv.org/pdf/2003.08237.pdf</a></p>",
      "rawMarkdown": "This is well-known fact, you could read more about it in the Facebook research papers:\nhttps://arxiv.org/pdf/1906.06423.pdf\nhttps://arxiv.org/pdf/2003.08237.pdf",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1455073,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-06T12:46:41.117000",
      "content": "<p>During training, do you use random resize augmentation? If so, then you are training on multiple sizes including 620. (Note that if you resize a 512 image to 620 and then crop a 512 square, that is still \"training on 620\" because the final product has the same pixel resolution, i.e. same pixel configuration as 620)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1455152,
          "author_name": "Parth Dhameliya",
          "author_url": "",
          "post_date": "2021-08-06T13:16:58.357000",
          "content": "<p>No my img size is fixed i.e 512, 512 during training and during inference I resize images to 600, 600 . No random resize augmentation during training. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1455230,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-06T13:43:58.827000",
          "content": "<p>Ok, so nothing like <code>albu.ShiftScaleRotate(rotate_limit=0, scale_limit=0.15, shift_limit=0, p=0.5)</code> ? Did you try random scale augmentation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1455246,
          "author_name": "Parth Dhameliya",
          "author_url": "",
          "post_date": "2021-08-06T13:49:52.103000",
          "content": "<p>Yeah i am using shift scale rotate. This could be the reason ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1455260,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-06T13:55:45.960000",
          "content": "<p>Scale augmentation is \"random resize augmentation then crop to 512\". So you are effectively training your model on 384 images, 512 images, 620 images. </p>\n<p>So the answer to your question, is that you are training on multiple image sizes. So one of these image sizes will be the best during infernce.</p>\n<p>Note that you could even try TTA where you infer 384, 512, 620 and ensemble the 3 predictions.</p>",
          "votes": 20,
          "replies": []
        },
        {
          "id": 1455319,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-08-06T14:27:50.073000",
          "content": "<p>This is enlightening. Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1455378,
          "author_name": "Parth Dhameliya",
          "author_url": "",
          "post_date": "2021-08-06T14:48:37.083000",
          "content": "<p>Thank you so much👍😄  <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1456169,
      "author_name": "Mykola Lavreniuk",
      "author_url": "",
      "post_date": "2021-08-06T20:02:06.427000",
      "content": "<p>This is well-known fact, you could read more about it in the Facebook research papers:<br>\n<a href=\"https://arxiv.org/pdf/1906.06423.pdf\" target=\"_blank\">https://arxiv.org/pdf/1906.06423.pdf</a><br>\n<a href=\"https://arxiv.org/pdf/2003.08237.pdf\" target=\"_blank\">https://arxiv.org/pdf/2003.08237.pdf</a></p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1454830": "During inference I sometime increase the image size (eg. 512 used in training but 620 in inference). I don't know the reason why it gives a boost in LB. In many competitions I had seen many solution in which image size during inference is different than the image size used in train. I want to know the reason....why most of the time it increases the lb score.  Even sometimes scaling down works too ?",
    "1455073": "During training, do you use random resize augmentation? If so, then you are training on multiple sizes including 620. (Note that if you resize a 512 image to 620 and then crop a 512 square, that is still \"training on 620\" because the final product has the same pixel resolution, i.e. same pixel configuration as 620)",
    "1456169": "This is well-known fact, you could read more about it in the Facebook research papers:\nhttps://arxiv.org/pdf/1906.06423.pdf\nhttps://arxiv.org/pdf/2003.08237.pdf"
  }
}