{
  "id": 156339,
  "title": "Is 128x128 better than 32x32",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156339",
  "author_name": "",
  "post_date": "2020-06-05T14:41:50.033917500Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I trained model using 32x32 images for 50 epochs and got score of 0.84 so just wondering if I use 128 or more image size will get better results ? I am using efficientnet.</p>",
  "messages": [
    {
      "id": "875126",
      "postDate": "06/05/2020 14:41:50",
      "content": "<p>I trained model using 32x32 images for 50 epochs and got score of 0.84 so just wondering if I use 128 or more image size will get better results ? I am using efficientnet.</p>",
      "rawMarkdown": "I trained model using 32x32 images for 50 epochs and got score of 0.84 so just wondering if I use 128 or more image size will get better results ? I am using efficientnet.",
      "votes": null
    },
    {
      "id": "875146",
      "postDate": "06/05/2020 14:50:28",
      "content": "<p>Yes, if you use 128x128 you will get a better score. Both the image size (i.e. 32, 64, 128, 256, 512, 1024) and model size (EfficientNet B0, B1, B2, B3, B4, B5, B6, B7) will affect your models accuracy. Try different combinations and see which does better.</p>",
      "rawMarkdown": "Yes, if you use 128x128 you will get a better score. Both the image size (i.e. 32, 64, 128, 256, 512, 1024) and model size (EfficientNet B0, B1, B2, B3, B4, B5, B6, B7) will affect your models accuracy. Try different combinations and see which does better.",
      "votes": null
    },
    {
      "id": "875210",
      "postDate": "06/05/2020 15:23:50",
      "content": "<p>Ok </p>",
      "rawMarkdown": "Ok",
      "votes": null
    },
    {
      "id": "875427",
      "postDate": "06/05/2020 19:13:07",
      "content": "<p>bigger image sizes mean more time to compute. Advanced experiments often use K-Folds (that means splitting the training data into a train and a val part in a special way - for each iteration of a fold one uses different splits of the training data to get a statistical benefit). This means for 5 folds, that you train 5 different versions of the model/training data = 5 times more computing time -&gt; one has to find a good mix between model and image size, to get a better result, if you use this folding trick.</p>",
      "rawMarkdown": "bigger image sizes mean more time to compute. Advanced experiments often use K-Folds (that means splitting the training data into a train and a val part in a special way - for each iteration of a fold one uses different splits of the training data to get a statistical benefit). This means for 5 folds, that you train 5 different versions of the model/training data = 5 times more computing time -&gt; one has to find a good mix between model and image size, to get a better result, if you use this folding trick.",
      "votes": null
    },
    {
      "id": "875795",
      "postDate": "06/06/2020 06:45:18",
      "content": "<p>I am not using kfolds right now, for starting I decided to choose between different image sizes and models once that is fixed I will proceed with kfolds and other techniques to improve my score. I hope this is the right approach. </p>",
      "rawMarkdown": "I am not using kfolds right now, for starting I decided to choose between different image sizes and models once that is fixed I will proceed with kfolds and other techniques to improve my score. I hope this is the right approach.",
      "votes": null
    },
    {
      "id": "875876",
      "postDate": "06/06/2020 08:24:28",
      "content": "<p>you can try different EfficientNets as Chris mentioned. Try to stick to the image size each of them is made for. Then ensemble them. </p>",
      "rawMarkdown": "you can try different EfficientNets as Chris mentioned. Try to stick to the image size each of them is made for. Then ensemble them.",
      "votes": null
    },
    {
      "id": "876079",
      "postDate": "06/06/2020 12:26:38",
      "content": "<p>There is often confusion about which image size to use.The bigger the size,more will be the computation cost.But models are trained better on bigger images.So choose images sizes as per your convenience.I had better results when I used larger image sizes such as 224 or 512 on efficientnet.</p>\n\n<p>It's always safe to train models on those sizes on which they are pretrained.I would recommend to follow that.Surely, you will get better results. Here is the list for efficientnet:</p>\n\n<pre><code>    Coefficients     :   width,depth,res,dropout\n    'efficientnet-b0':  (1.0, 1.0, 224, 0.2),\n    'efficientnet-b1': (1.0, 1.1, 240, 0.2),\n    'efficientnet-b2': (1.1, 1.2, 260, 0.3),\n    'efficientnet-b3': (1.2, 1.4, 300, 0.3),\n    'efficientnet-b4': (1.4, 1.8, 380, 0.4),\n    'efficientnet-b5': (1.6, 2.2, 456, 0.4),\n    'efficientnet-b6': (1.8, 2.6, 528, 0.5),\n    'efficientnet-b7': (2.0, 3.1, 600, 0.5)\n</code></pre>\n\n<p>Hope this helps you.</p>",
      "rawMarkdown": "There is often confusion about which image size to use.The bigger the size,more will be the computation cost.But models are trained better on bigger images.So choose images sizes as per your convenience.I had better results when I used larger image sizes such as 224 or 512 on efficientnet.\n\nIt's always safe to train models on those sizes on which they are pretrained.I would recommend to follow that.Surely, you will get better results. Here is the list for efficientnet:\n\n        Coefficients     :   width,depth,res,dropout\n        'efficientnet-b0':  (1.0, 1.0, 224, 0.2),\n        'efficientnet-b1': (1.0, 1.1, 240, 0.2),\n        'efficientnet-b2': (1.1, 1.2, 260, 0.3),\n        'efficientnet-b3': (1.2, 1.4, 300, 0.3),\n        'efficientnet-b4': (1.4, 1.8, 380, 0.4),\n        'efficientnet-b5': (1.6, 2.2, 456, 0.4),\n        'efficientnet-b6': (1.8, 2.6, 528, 0.5),\n        'efficientnet-b7': (2.0, 3.1, 600, 0.5)\n\nHope this helps you.",
      "votes": null
    },
    {
      "id": "876309",
      "postDate": "06/06/2020 16:04:30",
      "content": "<p>I don't think it is important to stick with the image size that they were made for. They are all fully convolutional nets. That means they looks for features in the images. For example, maybe they look for little circles or maybe they look for triangles, etc. If you give them 1024x1024 or 256x256, they will still look for circles and triangles.</p>\n\n<p>Then when you use <code>GlobalAveragePooling2D()</code> you get back a vector (let's call is <code>x</code>) of some length, let's say <code>len(x) = 2048</code>. Then the each element in that vector refers to the presence of a certain pattern. For example <code>x[0]</code> indicates the presence of circles in the image and <code>x[1]</code> indicates the presence of triangles in the image. You see, it doesn't matter what size image you started with.</p>",
      "rawMarkdown": "I don't think it is important to stick with the image size that they were made for. They are all fully convolutional nets. That means they looks for features in the images. For example, maybe they look for little circles or maybe they look for triangles, etc. If you give them 1024x1024 or 256x256, they will still look for circles and triangles.\n\nThen when you use `GlobalAveragePooling2D()` you get back a vector (let's call is `x`) of some length, let's say `len(x) = 2048`. Then the each element in that vector refers to the presence of a certain pattern. For example `x[0]` indicates the presence of circles in the image and `x[1]` indicates the presence of triangles in the image. You see, it doesn't matter what size image you started with.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 875146,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/05/2020 14:50:28",
      "content": "<p>Yes, if you use 128x128 you will get a better score. Both the image size (i.e. 32, 64, 128, 256, 512, 1024) and model size (EfficientNet B0, B1, B2, B3, B4, B5, B6, B7) will affect your models accuracy. Try different combinations and see which does better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 875210,
          "author_name": "priyt00",
          "author_url": "",
          "post_date": "06/05/2020 15:23:50",
          "content": "<p>Ok </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 875427,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "06/05/2020 19:13:07",
      "content": "<p>bigger image sizes mean more time to compute. Advanced experiments often use K-Folds (that means splitting the training data into a train and a val part in a special way - for each iteration of a fold one uses different splits of the training data to get a statistical benefit). This means for 5 folds, that you train 5 different versions of the model/training data = 5 times more computing time -&gt; one has to find a good mix between model and image size, to get a better result, if you use this folding trick.</p>",
      "votes": null,
      "replies": [
        {
          "id": 875795,
          "author_name": "priyt00",
          "author_url": "",
          "post_date": "06/06/2020 06:45:18",
          "content": "<p>I am not using kfolds right now, for starting I decided to choose between different image sizes and models once that is fixed I will proceed with kfolds and other techniques to improve my score. I hope this is the right approach. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 875876,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "06/06/2020 08:24:28",
          "content": "<p>you can try different EfficientNets as Chris mentioned. Try to stick to the image size each of them is made for. Then ensemble them. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 876309,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/06/2020 16:04:30",
          "content": "<p>I don't think it is important to stick with the image size that they were made for. They are all fully convolutional nets. That means they looks for features in the images. For example, maybe they look for little circles or maybe they look for triangles, etc. If you give them 1024x1024 or 256x256, they will still look for circles and triangles.</p>\n\n<p>Then when you use <code>GlobalAveragePooling2D()</code> you get back a vector (let's call is <code>x</code>) of some length, let's say <code>len(x) = 2048</code>. Then the each element in that vector refers to the presence of a certain pattern. For example <code>x[0]</code> indicates the presence of circles in the image and <code>x[1]</code> indicates the presence of triangles in the image. You see, it doesn't matter what size image you started with.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 876079,
      "author_name": "",
      "author_url": "",
      "post_date": "06/06/2020 12:26:38",
      "content": "<p>There is often confusion about which image size to use.The bigger the size,more will be the computation cost.But models are trained better on bigger images.So choose images sizes as per your convenience.I had better results when I used larger image sizes such as 224 or 512 on efficientnet.</p>\n\n<p>It's always safe to train models on those sizes on which they are pretrained.I would recommend to follow that.Surely, you will get better results. Here is the list for efficientnet:</p>\n\n<pre><code>    Coefficients     :   width,depth,res,dropout\n    'efficientnet-b0':  (1.0, 1.0, 224, 0.2),\n    'efficientnet-b1': (1.0, 1.1, 240, 0.2),\n    'efficientnet-b2': (1.1, 1.2, 260, 0.3),\n    'efficientnet-b3': (1.2, 1.4, 300, 0.3),\n    'efficientnet-b4': (1.4, 1.8, 380, 0.4),\n    'efficientnet-b5': (1.6, 2.2, 456, 0.4),\n    'efficientnet-b6': (1.8, 2.6, 528, 0.5),\n    'efficientnet-b7': (2.0, 3.1, 600, 0.5)\n</code></pre>\n\n<p>Hope this helps you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "875126": "I trained model using 32x32 images for 50 epochs and got score of 0.84 so just wondering if I use 128 or more image size will get better results ? I am using efficientnet.",
    "875146": "Yes, if you use 128x128 you will get a better score. Both the image size (i.e. 32, 64, 128, 256, 512, 1024) and model size (EfficientNet B0, B1, B2, B3, B4, B5, B6, B7) will affect your models accuracy. Try different combinations and see which does better.",
    "875210": "Ok",
    "875427": "bigger image sizes mean more time to compute. Advanced experiments often use K-Folds (that means splitting the training data into a train and a val part in a special way - for each iteration of a fold one uses different splits of the training data to get a statistical benefit). This means for 5 folds, that you train 5 different versions of the model/training data = 5 times more computing time -&gt; one has to find a good mix between model and image size, to get a better result, if you use this folding trick.",
    "875795": "I am not using kfolds right now, for starting I decided to choose between different image sizes and models once that is fixed I will proceed with kfolds and other techniques to improve my score. I hope this is the right approach.",
    "875876": "you can try different EfficientNets as Chris mentioned. Try to stick to the image size each of them is made for. Then ensemble them.",
    "876079": "There is often confusion about which image size to use.The bigger the size,more will be the computation cost.But models are trained better on bigger images.So choose images sizes as per your convenience.I had better results when I used larger image sizes such as 224 or 512 on efficientnet.\n\nIt's always safe to train models on those sizes on which they are pretrained.I would recommend to follow that.Surely, you will get better results. Here is the list for efficientnet:\n\n        Coefficients     :   width,depth,res,dropout\n        'efficientnet-b0':  (1.0, 1.0, 224, 0.2),\n        'efficientnet-b1': (1.0, 1.1, 240, 0.2),\n        'efficientnet-b2': (1.1, 1.2, 260, 0.3),\n        'efficientnet-b3': (1.2, 1.4, 300, 0.3),\n        'efficientnet-b4': (1.4, 1.8, 380, 0.4),\n        'efficientnet-b5': (1.6, 2.2, 456, 0.4),\n        'efficientnet-b6': (1.8, 2.6, 528, 0.5),\n        'efficientnet-b7': (2.0, 3.1, 600, 0.5)\n\nHope this helps you.",
    "876309": "I don't think it is important to stick with the image size that they were made for. They are all fully convolutional nets. That means they looks for features in the images. For example, maybe they look for little circles or maybe they look for triangles, etc. If you give them 1024x1024 or 256x256, they will still look for circles and triangles.\n\nThen when you use `GlobalAveragePooling2D()` you get back a vector (let's call is `x`) of some length, let's say `len(x) = 2048`. Then the each element in that vector refers to the presence of a certain pattern. For example `x[0]` indicates the presence of circles in the image and `x[1]` indicates the presence of triangles in the image. You see, it doesn't matter what size image you started with."
  },
  "source": "meta"
}