{
  "id": 133160,
  "title": "How Image size effects pretrained model(Ex ResNet) performance?",
  "url": "/competitions/bengaliai-cv19/discussion/133160",
  "author_name": "",
  "post_date": "2020-03-01T02:36:42.956662Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, I'm curious to know-how image size affects the performance(feature extraction) of pre-trained models such as ResNet or DenseNet?. \nI'll try to explain my doubt clearly, I'm using ResNet with pre-trained weights from imagenet data(image size 224x224). I'm resizing my Bengali dataset images to 64x64. When i'm retraining, i observed my training accuracy improving (7-60+), but my validation accuracy was stuck there at 30%.</p>",
  "messages": [
    {
      "id": "760237",
      "postDate": "03/01/2020 02:36:42",
      "content": "<p>Hi, I'm curious to know-how image size affects the performance(feature extraction) of pre-trained models such as ResNet or DenseNet?. \nI'll try to explain my doubt clearly, I'm using ResNet with pre-trained weights from imagenet data(image size 224x224). I'm resizing my Bengali dataset images to 64x64. When i'm retraining, i observed my training accuracy improving (7-60+), but my validation accuracy was stuck there at 30%.</p>",
      "rawMarkdown": "Hi, I'm curious to know-how image size affects the performance(feature extraction) of pre-trained models such as ResNet or DenseNet?. \nI'll try to explain my doubt clearly, I'm using ResNet with pre-trained weights from imagenet data(image size 224x224). I'm resizing my Bengali dataset images to 64x64. When i'm retraining, i observed my training accuracy improving (7-60+), but my validation accuracy was stuck there at 30%.",
      "votes": null
    },
    {
      "id": "761112",
      "postDate": "03/02/2020 06:45:09",
      "content": "<p>How much images ur taken for training </p>",
      "rawMarkdown": "How much images ur taken for training",
      "votes": null
    },
    {
      "id": "761398",
      "postDate": "03/02/2020 13:20:40",
      "content": "<p>This may be a case of overfitting. Essentially with the smaller dataset, the network has a much easier time memorizing the exact images in the training dataset, but is doing so without generalising. When it is given an image in the test dataset, it says \"teacher didn't tell us that question would be on the exam, but its a good thing its a multiple choice question\".</p>\n\n<p>With the larger pixel dataset, the images are all a bit of a blur. There is a squiggly thing in the corner, I guess that must be a diacritic.</p>",
      "rawMarkdown": "This may be a case of overfitting. Essentially with the smaller dataset, the network has a much easier time memorizing the exact images in the training dataset, but is doing so without generalising. When it is given an image in the test dataset, it says \"teacher didn't tell us that question would be on the exam, but its a good thing its a multiple choice question\".\n\nWith the larger pixel dataset, the images are all a bit of a blur. There is a squiggly thing in the corner, I guess that must be a diacritic.",
      "votes": null
    },
    {
      "id": "762497",
      "postDate": "03/03/2020 15:04:29",
      "content": "<p>I think it would be over fitting too soon on the 64x64 images as the pre-trained model was trained on 224x224 so this would seem too easy for the model to learn on train data, thus over fitting. One thing you can do is to k-fold cross validation to let it generalize better over the entire dataset. and probably lowering learning rate might help. </p>\n\n<p>If you are using freezed model for feature extraction, then I think it wouldn't have much difference. as the kernels of 3x3, 5x5 are locally operating.</p>",
      "rawMarkdown": "I think it would be over fitting too soon on the 64x64 images as the pre-trained model was trained on 224x224 so this would seem too easy for the model to learn on train data, thus over fitting. One thing you can do is to k-fold cross validation to let it generalize better over the entire dataset. and probably lowering learning rate might help. \n\nIf you are using freezed model for feature extraction, then I think it wouldn't have much difference. as the kernels of 3x3, 5x5 are locally operating.",
      "votes": null
    },
    {
      "id": "765829",
      "postDate": "03/07/2020 07:02:42",
      "content": "<p>I split the whole dataset into 5 stratified parts. 4*<em>(160672 images)</em>* parts for training 1*<em>(40168 images)</em>* part is for testing(CV).</p>",
      "rawMarkdown": "I split the whole dataset into 5 stratified parts. 4**(160672 images)** parts for training 1**(40168 images)** part is for testing(CV).",
      "votes": null
    },
    {
      "id": "927163",
      "postDate": "07/13/2020 08:30:58",
      "content": "<p>Hello <a href=\"/ramireddym\">@ramireddym</a> <a href=\"/parmarsuraj99\">@parmarsuraj99</a> <a href=\"/jamesmcguigan\">@jamesmcguigan</a>  i have a doubt.. When you are using ResNet for finetuning and we know that ResNet accepts 224x224 image size then how you are using 64x64 image size to retrain your network?</p>",
      "rawMarkdown": "Hello @ramireddym @parmarsuraj99 @jamesmcguigan  i have a doubt.. When you are using ResNet for finetuning and we know that ResNet accepts 224x224 image size then how you are using 64x64 image size to retrain your network?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 761112,
      "author_name": "jatinhabibkar",
      "author_url": "",
      "post_date": "03/02/2020 06:45:09",
      "content": "<p>How much images ur taken for training </p>",
      "votes": null,
      "replies": [
        {
          "id": 765829,
          "author_name": "ramireddym",
          "author_url": "",
          "post_date": "03/07/2020 07:02:42",
          "content": "<p>I split the whole dataset into 5 stratified parts. 4*<em>(160672 images)</em>* parts for training 1*<em>(40168 images)</em>* part is for testing(CV).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761398,
      "author_name": "jamesmcguigan",
      "author_url": "",
      "post_date": "03/02/2020 13:20:40",
      "content": "<p>This may be a case of overfitting. Essentially with the smaller dataset, the network has a much easier time memorizing the exact images in the training dataset, but is doing so without generalising. When it is given an image in the test dataset, it says \"teacher didn't tell us that question would be on the exam, but its a good thing its a multiple choice question\".</p>\n\n<p>With the larger pixel dataset, the images are all a bit of a blur. There is a squiggly thing in the corner, I guess that must be a diacritic.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 762497,
      "author_name": "parmarsuraj99",
      "author_url": "",
      "post_date": "03/03/2020 15:04:29",
      "content": "<p>I think it would be over fitting too soon on the 64x64 images as the pre-trained model was trained on 224x224 so this would seem too easy for the model to learn on train data, thus over fitting. One thing you can do is to k-fold cross validation to let it generalize better over the entire dataset. and probably lowering learning rate might help. </p>\n\n<p>If you are using freezed model for feature extraction, then I think it wouldn't have much difference. as the kernels of 3x3, 5x5 are locally operating.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 927163,
      "author_name": "anamika06",
      "author_url": "",
      "post_date": "07/13/2020 08:30:58",
      "content": "<p>Hello <a href=\"/ramireddym\">@ramireddym</a> <a href=\"/parmarsuraj99\">@parmarsuraj99</a> <a href=\"/jamesmcguigan\">@jamesmcguigan</a>  i have a doubt.. When you are using ResNet for finetuning and we know that ResNet accepts 224x224 image size then how you are using 64x64 image size to retrain your network?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "760237": "Hi, I'm curious to know-how image size affects the performance(feature extraction) of pre-trained models such as ResNet or DenseNet?. \nI'll try to explain my doubt clearly, I'm using ResNet with pre-trained weights from imagenet data(image size 224x224). I'm resizing my Bengali dataset images to 64x64. When i'm retraining, i observed my training accuracy improving (7-60+), but my validation accuracy was stuck there at 30%.",
    "761112": "How much images ur taken for training",
    "761398": "This may be a case of overfitting. Essentially with the smaller dataset, the network has a much easier time memorizing the exact images in the training dataset, but is doing so without generalising. When it is given an image in the test dataset, it says \"teacher didn't tell us that question would be on the exam, but its a good thing its a multiple choice question\".\n\nWith the larger pixel dataset, the images are all a bit of a blur. There is a squiggly thing in the corner, I guess that must be a diacritic.",
    "762497": "I think it would be over fitting too soon on the 64x64 images as the pre-trained model was trained on 224x224 so this would seem too easy for the model to learn on train data, thus over fitting. One thing you can do is to k-fold cross validation to let it generalize better over the entire dataset. and probably lowering learning rate might help. \n\nIf you are using freezed model for feature extraction, then I think it wouldn't have much difference. as the kernels of 3x3, 5x5 are locally operating.",
    "765829": "I split the whole dataset into 5 stratified parts. 4**(160672 images)** parts for training 1**(40168 images)** part is for testing(CV).",
    "927163": "Hello @ramireddym @parmarsuraj99 @jamesmcguigan  i have a doubt.. When you are using ResNet for finetuning and we know that ResNet accepts 224x224 image size then how you are using 64x64 image size to retrain your network?"
  },
  "source": "meta"
}