{
  "id": 98736,
  "title": "Big size vs Small size",
  "url": "/competitions/aptos2019-blindness-detection/discussion/98736",
  "author_name": "",
  "post_date": "2019-07-05T23:09:16.037500600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am using 224 as input size and I get 0.91 CV but when I try to increase the size to 500 I have lower CV 0.87.\nCan someone explain this? I was thinking that the more information the model has the better result we can get but this is not working with the competition's data.</p>",
  "messages": [
    {
      "id": "569039",
      "postDate": "07/05/2019 23:09:16",
      "content": "<p>I am using 224 as input size and I get 0.91 CV but when I try to increase the size to 500 I have lower CV 0.87.\nCan someone explain this? I was thinking that the more information the model has the better result we can get but this is not working with the competition's data.</p>",
      "rawMarkdown": "I am using 224 as input size and I get 0.91 CV but when I try to increase the size to 500 I have lower CV 0.87.\nCan someone explain this? I was thinking that the more information the model has the better result we can get but this is not working with the competition's data.",
      "votes": null
    },
    {
      "id": "584581",
      "postDate": "07/26/2019 07:36:58",
      "content": "<p>Are you using a DenseNet conv network? DenseNet-121 pre-trained has 224 as input size (was the input size which was used when it was trained in the first place). See PDF on <a href=\"https://arxiv.org/abs/1608.06993\">https://arxiv.org/abs/1608.06993</a>, chapter 3, section \"Implementation details\".</p>",
      "rawMarkdown": "Are you using a DenseNet conv network? DenseNet-121 pre-trained has 224 as input size (was the input size which was used when it was trained in the first place). See PDF on [https://arxiv.org/abs/1608.06993](https://arxiv.org/abs/1608.06993), chapter 3, section \"Implementation details\".",
      "votes": null
    },
    {
      "id": "585043",
      "postDate": "07/26/2019 20:46:35",
      "content": "<p>As the pre-trained networks are trained on completely different images (ordinary pictures) I don't think it matters much what the original input size was.</p>\n\n<p>One of the reasons I can think of is the change in receptive field. Depending on the depth of your network, the value at some position in the output volume is only based on N pixels in its surrounding.</p>\n\n<p>An example situation:\nYou want to predict whether a hamburger contains both tomato and a patty, but there's cheese in the middle.\nIf the receptive field is 12 pixels and the cheese is 8 pixels in size. there are indices in the activation volume at sufficient depth that were based on pixels that both \"saw\" the tomato and the patty.\nNow if you make your stretched the image, and the cheese would now be 16 pixels thick, this would no longer be possible, which may make it more difficult to classify. The \"real world\" receptive field is smaller if you increase the size of your input.</p>",
      "rawMarkdown": "As the pre-trained networks are trained on completely different images (ordinary pictures) I don't think it matters much what the original input size was.\n\nOne of the reasons I can think of is the change in receptive field. Depending on the depth of your network, the value at some position in the output volume is only based on N pixels in its surrounding.\n\nAn example situation:\nYou want to predict whether a hamburger contains both tomato and a patty, but there's cheese in the middle.\nIf the receptive field is 12 pixels and the cheese is 8 pixels in size. there are indices in the activation volume at sufficient depth that were based on pixels that both \"saw\" the tomato and the patty.\nNow if you make your stretched the image, and the cheese would now be 16 pixels thick, this would no longer be possible, which may make it more difficult to classify. The \"real world\" receptive field is smaller if you increase the size of your input.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 584581,
      "author_name": "tairosonloa",
      "author_url": "",
      "post_date": "07/26/2019 07:36:58",
      "content": "<p>Are you using a DenseNet conv network? DenseNet-121 pre-trained has 224 as input size (was the input size which was used when it was trained in the first place). See PDF on <a href=\"https://arxiv.org/abs/1608.06993\">https://arxiv.org/abs/1608.06993</a>, chapter 3, section \"Implementation details\".</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 585043,
      "author_name": "gzuidhof",
      "author_url": "",
      "post_date": "07/26/2019 20:46:35",
      "content": "<p>As the pre-trained networks are trained on completely different images (ordinary pictures) I don't think it matters much what the original input size was.</p>\n\n<p>One of the reasons I can think of is the change in receptive field. Depending on the depth of your network, the value at some position in the output volume is only based on N pixels in its surrounding.</p>\n\n<p>An example situation:\nYou want to predict whether a hamburger contains both tomato and a patty, but there's cheese in the middle.\nIf the receptive field is 12 pixels and the cheese is 8 pixels in size. there are indices in the activation volume at sufficient depth that were based on pixels that both \"saw\" the tomato and the patty.\nNow if you make your stretched the image, and the cheese would now be 16 pixels thick, this would no longer be possible, which may make it more difficult to classify. The \"real world\" receptive field is smaller if you increase the size of your input.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "569039": "I am using 224 as input size and I get 0.91 CV but when I try to increase the size to 500 I have lower CV 0.87.\nCan someone explain this? I was thinking that the more information the model has the better result we can get but this is not working with the competition's data.",
    "584581": "Are you using a DenseNet conv network? DenseNet-121 pre-trained has 224 as input size (was the input size which was used when it was trained in the first place). See PDF on [https://arxiv.org/abs/1608.06993](https://arxiv.org/abs/1608.06993), chapter 3, section \"Implementation details\".",
    "585043": "As the pre-trained networks are trained on completely different images (ordinary pictures) I don't think it matters much what the original input size was.\n\nOne of the reasons I can think of is the change in receptive field. Depending on the depth of your network, the value at some position in the output volume is only based on N pixels in its surrounding.\n\nAn example situation:\nYou want to predict whether a hamburger contains both tomato and a patty, but there's cheese in the middle.\nIf the receptive field is 12 pixels and the cheese is 8 pixels in size. there are indices in the activation volume at sufficient depth that were based on pixels that both \"saw\" the tomato and the patty.\nNow if you make your stretched the image, and the cheese would now be 16 pixels thick, this would no longer be possible, which may make it more difficult to classify. The \"real world\" receptive field is smaller if you increase the size of your input."
  },
  "source": "meta"
}