{
  "id": 39278,
  "title": "Will big input size such as 1280 * 1280 overfit?",
  "url": "/competitions/carvana-image-masking-challenge/discussion/39278",
  "author_name": "",
  "post_date": "2017-09-11T08:03:28.095266Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>rt</p>",
  "messages": [
    {
      "id": "220109",
      "postDate": "09/11/2017 08:03:28",
      "content": "<p>rt</p>",
      "rawMarkdown": "rt",
      "votes": null
    },
    {
      "id": "220120",
      "postDate": "09/11/2017 08:49:26",
      "content": "<p>Overfit is generally because your network is larger than needed, what is a large network is of course hard to determine. I would rather assume if you train network A on small images, say 256x256, that network is more prone to overfit than if you'd input 512x512 into the same network.</p>\n\n<p>I'm able to overfit using the full resolution images, you should look at the loss / accuracy curves. If you see your validation accuracy increase, then after a while start decreasing again, you have most likely achieved overfitting (Assuming your validation  /  train split is done right). Similar to this, if you see that your training loss keeps decreasing, but your validation loss is steadily increasing, you have an overfit on your hands.</p>",
      "rawMarkdown": "Overfit is generally because your network is larger than needed, what is a large network is of course hard to determine. I would rather assume if you train network A on small images, say 256x256, that network is more prone to overfit than if you'd input 512x512 into the same network.\n\nI'm able to overfit using the full resolution images, you should look at the loss / accuracy curves. If you see your validation accuracy increase, then after a while start decreasing again, you have most likely achieved overfitting (Assuming your validation  /  train split is done right). Similar to this, if you see that your training loss keeps decreasing, but your validation loss is steadily increasing, you have an overfit on your hands.",
      "votes": null
    },
    {
      "id": "220470",
      "postDate": "09/12/2017 07:45:22",
      "content": "<p>Generally it depends on the model but surely it can. \nYou get overfitting when the ratio of your network size to training examples grows big enough. As a measure of network size you can use number of weights and as a measure of training examples you can use pixel count (for fully convolutional architectures as U-net  ). \nLets say you have a U-net which doubles the filter number after each pooling operations. So on the second level you will have 4 times as many weights and 4 times less pixels to train on. At level 5 the ratio of weights_no/pixels_no is a million times bigger than on level 1. \nSo the more filters you get on the pooled layers - the bigger overfit you will see. On the other hand if you don't have enough filters your network will stop learning when it reaches some accuracy level, and adding filters to unpooled layers is very computationally expensive. </p>",
      "rawMarkdown": "Generally it depends on the model but surely it can. \nYou get overfitting when the ratio of your network size to training examples grows big enough. As a measure of network size you can use number of weights and as a measure of training examples you can use pixel count (for fully convolutional architectures as U-net  ). \nLets say you have a U-net which doubles the filter number after each pooling operations. So on the second level you will have 4 times as many weights and 4 times less pixels to train on. At level 5 the ratio of weights_no/pixels_no is a million times bigger than on level 1. \nSo the more filters you get on the pooled layers - the bigger overfit you will see. On the other hand if you don't have enough filters your network will stop learning when it reaches some accuracy level, and adding filters to unpooled layers is very computationally expensive.",
      "votes": null
    },
    {
      "id": "221168",
      "postDate": "09/14/2017 11:31:18",
      "content": "<p>I have tested recently with 1280*1280 input size, and I can say my model over fitted and I got the worst result in comparison to my 1024*1024 model.  Due to limited GPU memory, I could fit the batch size of one for 1280 input image only. I think ensembling of previous predicated outptus would give better result as most of top LBs have done smiliar approach. </p>",
      "rawMarkdown": "I have tested recently with 1280*1280 input size, and I can say my model over fitted and I got the worst result in comparison to my 1024*1024 model.  Due to limited GPU memory, I could fit the batch size of one for 1280 input image only. I think ensembling of previous predicated outptus would give better result as most of top LBs have done smiliar approach.",
      "votes": null
    },
    {
      "id": "221335",
      "postDate": "09/14/2017 20:41:16",
      "content": "<p>It might give a marginal advantage once you start ensembling, but I trained it + 1536^2 with batch size 2 and it worked just fine.</p>",
      "rawMarkdown": "It might give a marginal advantage once you start ensembling, but I trained it + 1536^2 with batch size 2 and it worked just fine.",
      "votes": null
    },
    {
      "id": "221553",
      "postDate": "09/15/2017 16:16:02",
      "content": "<p>@Steven what do you mean by 1536^2? do you mean about image size? I have not tried with larger scale images, but if it worked for you then I could give a try also.</p>",
      "rawMarkdown": "Steven what do you mean by 1536^2? do you mean about image size? I have not tried with larger scale images, but if it worked for you then I could give a try also.",
      "votes": null
    },
    {
      "id": "221555",
      "postDate": "09/15/2017 16:19:58",
      "content": "<p>I made the input 1536 x 1536 (taller than the actual input). I only made one for the sake of ensembling.</p>",
      "rawMarkdown": "I made the input 1536 x 1536 (taller than the actual input). I only made one for the sake of ensembling.",
      "votes": null
    },
    {
      "id": "221620",
      "postDate": "09/15/2017 21:18:44",
      "content": "<p>If you want to keep the same aspect ratio 1.5:1 as the original image but smaller in size try 1536x1024. </p>",
      "rawMarkdown": "If you want to keep the same aspect ratio 1.5:1 as the original image but smaller in size try 1536x1024.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 220120,
      "author_name": "adamhart",
      "author_url": "",
      "post_date": "09/11/2017 08:49:26",
      "content": "<p>Overfit is generally because your network is larger than needed, what is a large network is of course hard to determine. I would rather assume if you train network A on small images, say 256x256, that network is more prone to overfit than if you'd input 512x512 into the same network.</p>\n\n<p>I'm able to overfit using the full resolution images, you should look at the loss / accuracy curves. If you see your validation accuracy increase, then after a while start decreasing again, you have most likely achieved overfitting (Assuming your validation  /  train split is done right). Similar to this, if you see that your training loss keeps decreasing, but your validation loss is steadily increasing, you have an overfit on your hands.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 220470,
      "author_name": "mitiau",
      "author_url": "",
      "post_date": "09/12/2017 07:45:22",
      "content": "<p>Generally it depends on the model but surely it can. \nYou get overfitting when the ratio of your network size to training examples grows big enough. As a measure of network size you can use number of weights and as a measure of training examples you can use pixel count (for fully convolutional architectures as U-net  ). \nLets say you have a U-net which doubles the filter number after each pooling operations. So on the second level you will have 4 times as many weights and 4 times less pixels to train on. At level 5 the ratio of weights_no/pixels_no is a million times bigger than on level 1. \nSo the more filters you get on the pooled layers - the bigger overfit you will see. On the other hand if you don't have enough filters your network will stop learning when it reaches some accuracy level, and adding filters to unpooled layers is very computationally expensive. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 221168,
      "author_name": "svesal",
      "author_url": "",
      "post_date": "09/14/2017 11:31:18",
      "content": "<p>I have tested recently with 1280*1280 input size, and I can say my model over fitted and I got the worst result in comparison to my 1024*1024 model.  Due to limited GPU memory, I could fit the batch size of one for 1280 input image only. I think ensembling of previous predicated outptus would give better result as most of top LBs have done smiliar approach. </p>",
      "votes": null,
      "replies": [
        {
          "id": 221335,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "09/14/2017 20:41:16",
          "content": "<p>It might give a marginal advantage once you start ensembling, but I trained it + 1536^2 with batch size 2 and it worked just fine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 221553,
          "author_name": "svesal",
          "author_url": "",
          "post_date": "09/15/2017 16:16:02",
          "content": "<p>@Steven what do you mean by 1536^2? do you mean about image size? I have not tried with larger scale images, but if it worked for you then I could give a try also.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 221555,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "09/15/2017 16:19:58",
          "content": "<p>I made the input 1536 x 1536 (taller than the actual input). I only made one for the sake of ensembling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 221620,
      "author_name": "jamesrequa",
      "author_url": "",
      "post_date": "09/15/2017 21:18:44",
      "content": "<p>If you want to keep the same aspect ratio 1.5:1 as the original image but smaller in size try 1536x1024. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "220109": "rt",
    "220120": "Overfit is generally because your network is larger than needed, what is a large network is of course hard to determine. I would rather assume if you train network A on small images, say 256x256, that network is more prone to overfit than if you'd input 512x512 into the same network.\n\nI'm able to overfit using the full resolution images, you should look at the loss / accuracy curves. If you see your validation accuracy increase, then after a while start decreasing again, you have most likely achieved overfitting (Assuming your validation  /  train split is done right). Similar to this, if you see that your training loss keeps decreasing, but your validation loss is steadily increasing, you have an overfit on your hands.",
    "220470": "Generally it depends on the model but surely it can. \nYou get overfitting when the ratio of your network size to training examples grows big enough. As a measure of network size you can use number of weights and as a measure of training examples you can use pixel count (for fully convolutional architectures as U-net  ). \nLets say you have a U-net which doubles the filter number after each pooling operations. So on the second level you will have 4 times as many weights and 4 times less pixels to train on. At level 5 the ratio of weights_no/pixels_no is a million times bigger than on level 1. \nSo the more filters you get on the pooled layers - the bigger overfit you will see. On the other hand if you don't have enough filters your network will stop learning when it reaches some accuracy level, and adding filters to unpooled layers is very computationally expensive.",
    "221168": "I have tested recently with 1280*1280 input size, and I can say my model over fitted and I got the worst result in comparison to my 1024*1024 model.  Due to limited GPU memory, I could fit the batch size of one for 1280 input image only. I think ensembling of previous predicated outptus would give better result as most of top LBs have done smiliar approach.",
    "221335": "It might give a marginal advantage once you start ensembling, but I trained it + 1536^2 with batch size 2 and it worked just fine.",
    "221553": "Steven what do you mean by 1536^2? do you mean about image size? I have not tried with larger scale images, but if it worked for you then I could give a try also.",
    "221555": "I made the input 1536 x 1536 (taller than the actual input). I only made one for the sake of ensembling.",
    "221620": "If you want to keep the same aspect ratio 1.5:1 as the original image but smaller in size try 1536x1024."
  },
  "source": "meta"
}