{
  "id": 227032,
  "title": "Massive Performance boost - resizing images",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/227032",
  "author_name": "AnkurSingh",
  "post_date": "2021-03-18T16:49:19.297000",
  "votes": 36,
  "comment_count": 17,
  "views": 0,
  "content": "<p>In this competition, the original images are quite large (~(2672, 4000)). It took me 42 mins to train a ResNet50 (one epoch only). This is extremely slow to conduct any sort of experiment. </p>\n<p>So, after some digging, I look at the image size and decide to create a new resized dataset. With resized dataset, it took only 6 mins (originally 42 mins) per epoch. </p>\n<p>This is huge time saver (7x faster). Now, I can do more experiments, try different model, and do so much more. </p>\n<p>So, I decided to create multiple resized versions (particularly 256, 384, 512, 640) of the dataset. You can find the <a href=\"https://www.kaggle.com/ankursingh12/resized-plant2021\" target=\"_blank\">resized dataset here</a></p>\n<p>Also the code used for resizing the dataset can be found in the description.<br>\nHope, all the participants will save at least a couple of hours by using it. Cheers!</p>",
  "messages": [
    {
      "id": 1244020,
      "postDate": "2021-03-18T16:49:19.297Z",
      "content": "<p>In this competition, the original images are quite large (~(2672, 4000)). It took me 42 mins to train a ResNet50 (one epoch only). This is extremely slow to conduct any sort of experiment. </p>\n<p>So, after some digging, I look at the image size and decide to create a new resized dataset. With resized dataset, it took only 6 mins (originally 42 mins) per epoch. </p>\n<p>This is huge time saver (7x faster). Now, I can do more experiments, try different model, and do so much more. </p>\n<p>So, I decided to create multiple resized versions (particularly 256, 384, 512, 640) of the dataset. You can find the <a href=\"https://www.kaggle.com/ankursingh12/resized-plant2021\" target=\"_blank\">resized dataset here</a></p>\n<p>Also the code used for resizing the dataset can be found in the description.<br>\nHope, all the participants will save at least a couple of hours by using it. Cheers!</p>",
      "rawMarkdown": "In this competition, the original images are quite large (~(2672, 4000)). It took me 42 mins to train a ResNet50 (one epoch only). This is extremely slow to conduct any sort of experiment. \n\nSo, after some digging, I look at the image size and decide to create a new resized dataset. With resized dataset, it took only 6 mins (originally 42 mins) per epoch. \n\nThis is huge time saver (7x faster). Now, I can do more experiments, try different model, and do so much more. \n\nSo, I decided to create multiple resized versions (particularly 256, 384, 512, 640) of the dataset. You can find the [resized dataset here](https://www.kaggle.com/ankursingh12/resized-plant2021)\n\nAlso the code used for resizing the dataset can be found in the description.\nHope, all the participants will save at least a couple of hours by using it. Cheers!",
      "votes": 35
    },
    {
      "id": 1248544,
      "postDate": "2021-03-22T17:07:14.100Z",
      "content": "<p>Thank you so much. </p>\n<p>I am getting large difference between CV and LB. (CV : 0.603, LB : 0.164). [with the original dataset]. Are you facing any such problem. This is my <a href=\"https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-tpu-training\" target=\"_blank\">training notebook</a>.</p>",
      "rawMarkdown": "Thank you so much. \n\nI am getting large difference between CV and LB. (CV : 0.603, LB : 0.164). [with the original dataset]. Are you facing any such problem. This is my [training notebook](https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-tpu-training).",
      "votes": 1,
      "replies": [
        {
          "id": 1250299,
          "postDate": "2021-03-23T23:30:27.210Z",
          "content": "<p>Actually, everyone is facing the same issue. Here is a <strong><a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227237\" target=\"_blank\">discussion thread</a></strong> about the same</p>",
          "rawMarkdown": "Actually, everyone is facing the same issue. Here is a **[discussion thread](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227237)** about the same"
        }
      ]
    },
    {
      "id": 1349132,
      "postDate": "2021-06-14T14:27:28.740Z",
      "content": "<p>i am using efficientnet _b2 pytorch ,input_size=260*260, batch_size=16, on kaggle notebook one epoch taking a lot of time it has been more than 15 min and it's still not done,any help?</p>",
      "rawMarkdown": "i am using efficientnet _b2 pytorch ,input_size=260*260, batch_size=16, on kaggle notebook one epoch taking a lot of time it has been more than 15 min and it's still not done,any help?"
    },
    {
      "id": 1308635,
      "postDate": "2021-05-15T10:53:07.367Z",
      "content": "<p>hey, I'm trying different size ratio's using your code but getting outputs as visualization instead of folders. can you help me? <a href=\"https://www.kaggle.com/ankursingh12\" target=\"_blank\">@ankursingh12</a> </p>",
      "rawMarkdown": "hey, I'm trying different size ratio's using your code but getting outputs as visualization instead of folders. can you help me? @ankursingh12 "
    },
    {
      "id": 1252990,
      "postDate": "2021-03-26T09:15:19.013Z",
      "content": "<p>Why not resize the images in the processing of training?</p>",
      "rawMarkdown": "Why not resize the images in the processing of training?",
      "replies": [
        {
          "id": 1253241,
          "postDate": "2021-03-26T14:06:59.440Z",
          "content": "<p>Takes a lot of time! You will have to resize the image for every epoch (if you do it within dataloader). And if you do it before the dataloader, then you will have to resize them every time you run the notebook. </p>\n<p>These are some big images!</p>",
          "rawMarkdown": "Takes a lot of time! You will have to resize the image for every epoch (if you do it within dataloader). And if you do it before the dataloader, then you will have to resize them every time you run the notebook. \n\nThese are some big images!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1246708,
      "postDate": "2021-03-21T02:51:05.600Z",
      "content": "<p>This is great for train images. But then we need the same resizing for test images too, and we want to make sure that the transformation is the same.</p>\n<p>So, I tried to repeat the same transformation for some image, but all my approaches failed. For example, the following code were always returning slightly different output for resized_from_datased and resized_here, even though I tried different interpolations.</p>\n<pre><code>from fastai.vision.all import Image\nimport numpy as np\nimport os\nf = '9fad869f21b5b240.jpg'\nresized_from_dataset = Image.open(os.path.join('../input/resized-plant2021/img_sz_512/',f))\nresized_here = Image.open(os.path.join('../input/plant-pathology-2021-fgvc8/train_images/',f))\nresized_here = resized_here.resize((512,341),resample=Image.BILINEAR)\nnp.array(resized_from_dataset).astype(np.int32) - np.array(resized_here).astype(np.int32)\n</code></pre>\n<p>I am starting to suspect that the resized images undergo some changes at the moment when they are saved to disc. Hopefully it is not the case. What is the correct way to get resized_from_dataset==resized_here?</p>",
      "rawMarkdown": "This is great for train images. But then we need the same resizing for test images too, and we want to make sure that the transformation is the same.\n\nSo, I tried to repeat the same transformation for some image, but all my approaches failed. For example, the following code were always returning slightly different output for resized_from_datased and resized_here, even though I tried different interpolations.\n```\nfrom fastai.vision.all import Image\nimport numpy as np\nimport os\nf = '9fad869f21b5b240.jpg'\nresized_from_dataset = Image.open(os.path.join('../input/resized-plant2021/img_sz_512/',f))\nresized_here = Image.open(os.path.join('../input/plant-pathology-2021-fgvc8/train_images/',f))\nresized_here = resized_here.resize((512,341),resample=Image.BILINEAR)\nnp.array(resized_from_dataset).astype(np.int32) - np.array(resized_here).astype(np.int32)\n```\n\nI am starting to suspect that the resized images undergo some changes at the moment when they are saved to disc. Hopefully it is not the case. What is the correct way to get resized_from_dataset==resized_here?\n\n",
      "replies": [
        {
          "id": 1246723,
          "postDate": "2021-03-21T03:08:46.170Z",
          "content": "<p>Looks like saving to disc does distort the images… The default quality is 75.</p>\n<pre><code>from fastai.vision.all import Image\nimport numpy as np\nimport os\nf = '9fad869f21b5b240.jpg'\nresized_from_dataset = Image.open(os.path.join('../input/resized-plant2021/img_sz_512/',f))\nresized_here = Image.open(os.path.join('../input/plant-pathology-2021-fgvc8/train_images/',f))\nresized_here = resized_here.resize((512,341),resample=Image.BILINEAR)\n\nprint('Images are different:\\n', np.array(resized_from_dataset).astype(np.int32) - np.array(resized_here).astype(np.int32))\nresized_here.save('new_file_'+f, img_format='jpg')\nresized_here_recovered = Image.open('new_file_'+f)\nprint('\\nImages are the same:\\n', np.array(resized_from_dataset).astype(np.int32) - np.array(resized_here_recovered).astype(np.int32))\n</code></pre>",
          "rawMarkdown": "Looks like saving to disc does distort the images... The default quality is 75.\n```\nfrom fastai.vision.all import Image\nimport numpy as np\nimport os\nf = '9fad869f21b5b240.jpg'\nresized_from_dataset = Image.open(os.path.join('../input/resized-plant2021/img_sz_512/',f))\nresized_here = Image.open(os.path.join('../input/plant-pathology-2021-fgvc8/train_images/',f))\nresized_here = resized_here.resize((512,341),resample=Image.BILINEAR)\n\nprint('Images are different:\\n', np.array(resized_from_dataset).astype(np.int32) - np.array(resized_here).astype(np.int32))\nresized_here.save('new_file_'+f, img_format='jpg')\nresized_here_recovered = Image.open('new_file_'+f)\nprint('\\nImages are the same:\\n', np.array(resized_from_dataset).astype(np.int32) - np.array(resized_here_recovered).astype(np.int32))\n```"
        },
        {
          "id": 1250302,
          "postDate": "2021-03-23T23:33:35.870Z",
          "content": "<blockquote>\n  <p>What is the correct way to get resized_from_dataset==resized_here?</p>\n</blockquote>\n<p>Interesting! I will have to look into it 😅</p>",
          "rawMarkdown": "> What is the correct way to get resized_from_dataset==resized_here?\n\nInteresting! I will have to look into it 😅"
        }
      ]
    },
    {
      "id": 1245672,
      "postDate": "2021-03-20T02:48:19.480Z",
      "content": "<p>Thank you for sharing. It helped speed up learning.</p>",
      "rawMarkdown": "Thank you for sharing. It helped speed up learning.",
      "replies": [
        {
          "id": 1245845,
          "postDate": "2021-03-20T08:49:28.587Z",
          "content": "<p>Happy to help 😁</p>",
          "rawMarkdown": "Happy to help 😁"
        }
      ]
    },
    {
      "id": 1244399,
      "postDate": "2021-03-19T00:45:55.097Z",
      "content": "<p>Edit: Actually the aspect ratio is preserved in the dataset. Size 512 actually means 342*512. Thanks for sharing.</p>\n<p>Hi! Thanks. <br>\nDo you think if there is a benefit in preserving the aspect ratio? E.g. creating a version with 512x766 size?</p>",
      "rawMarkdown": "Edit: Actually the aspect ratio is preserved in the dataset. Size 512 actually means 342*512. Thanks for sharing.\n\nHi! Thanks. \nDo you think if there is a benefit in preserving the aspect ratio? E.g. creating a version with 512x766 size?",
      "replies": [
        {
          "id": 1244552,
          "postDate": "2021-03-19T04:10:45.027Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/buinyi\" target=\"_blank\">@buinyi</a>, yes the actual aspect ratio is preserved. There are many different image sizes in the dataset </p>\n<pre><code>{(2672, 4000): 16485,\n(1728, 2592): 1027,\n(3000, 4000): 665,\n(3456, 4608): 123,\n(3456, 5184): 193,\n(3024, 4032): 132,\n(3024, 3024): 3,\n(4032, 3024): 3,\n(2248, 4000): 1}\n</code></pre>\n<p>One obvious benefit of preserving the aspect ratio is, not disturbing the image. If there is considerable difference in image's height and width, then making them equal will result in a distorted image. If you want square images, then it can be easily done later. </p>\n<p>For example: When loading the image, I am using <code>RandomizedCrop</code> to crop square areas from the image. This makes sure that you are not disturbing the image throughout the pipeline. </p>\n<p>Hope it helps, Cheers 😁</p>",
          "rawMarkdown": "Hi @buinyi, yes the actual aspect ratio is preserved. There are many different image sizes in the dataset \n\n```\n{(2672, 4000): 16485,\n(1728, 2592): 1027,\n(3000, 4000): 665,\n(3456, 4608): 123,\n(3456, 5184): 193,\n(3024, 4032): 132,\n(3024, 3024): 3,\n(4032, 3024): 3,\n(2248, 4000): 1}\n```\n\nOne obvious benefit of preserving the aspect ratio is, not disturbing the image. If there is considerable difference in image's height and width, then making them equal will result in a distorted image. If you want square images, then it can be easily done later. \n\nFor example: When loading the image, I am using `RandomizedCrop` to crop square areas from the image. This makes sure that you are not disturbing the image throughout the pipeline. \n\nHope it helps, Cheers 😁"
        }
      ]
    },
    {
      "id": 1276686,
      "postDate": "2021-04-17T20:56:51.460Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1248385,
      "postDate": "2021-03-22T15:08:05.510Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1246236,
      "postDate": "2021-03-20T15:08:37.423Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1246262,
          "postDate": "2021-03-20T15:40:57.177Z",
          "content": "<p><a href=\"https://www.kaggle.com/mervynyang\" target=\"_blank\">@mervynyang</a> there can be many possible reasons. 😅</p>",
          "rawMarkdown": "@mervynyang there can be many possible reasons. 😅"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1248544,
      "author_name": "Shanmukh",
      "author_url": "",
      "post_date": "2021-03-22T17:07:14.100000",
      "content": "<p>Thank you so much. </p>\n<p>I am getting large difference between CV and LB. (CV : 0.603, LB : 0.164). [with the original dataset]. Are you facing any such problem. This is my <a href=\"https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-tpu-training\" target=\"_blank\">training notebook</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1250299,
          "author_name": "AnkurSingh",
          "author_url": "",
          "post_date": "2021-03-23T23:30:27.210000",
          "content": "<p>Actually, everyone is facing the same issue. Here is a <strong><a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227237\" target=\"_blank\">discussion thread</a></strong> about the same</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1349132,
      "author_name": "Mohd Kashif Akhtar",
      "author_url": "",
      "post_date": "2021-06-14T14:27:28.740000",
      "content": "<p>i am using efficientnet _b2 pytorch ,input_size=260*260, batch_size=16, on kaggle notebook one epoch taking a lot of time it has been more than 15 min and it's still not done,any help?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1308635,
      "author_name": "Ranjeet shrivastav",
      "author_url": "",
      "post_date": "2021-05-15T10:53:07.367000",
      "content": "<p>hey, I'm trying different size ratio's using your code but getting outputs as visualization instead of folders. can you help me? <a href=\"https://www.kaggle.com/ankursingh12\" target=\"_blank\">@ankursingh12</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1252990,
      "author_name": "Jianye He",
      "author_url": "",
      "post_date": "2021-03-26T09:15:19.013000",
      "content": "<p>Why not resize the images in the processing of training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1253241,
          "author_name": "AnkurSingh",
          "author_url": "",
          "post_date": "2021-03-26T14:06:59.440000",
          "content": "<p>Takes a lot of time! You will have to resize the image for every epoch (if you do it within dataloader). And if you do it before the dataloader, then you will have to resize them every time you run the notebook. </p>\n<p>These are some big images!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1246708,
      "author_name": "Igor Buinyi",
      "author_url": "",
      "post_date": "2021-03-21T02:51:05.600000",
      "content": "<p>This is great for train images. But then we need the same resizing for test images too, and we want to make sure that the transformation is the same.</p>\n<p>So, I tried to repeat the same transformation for some image, but all my approaches failed. For example, the following code were always returning slightly different output for resized_from_datased and resized_here, even though I tried different interpolations.</p>\n<pre><code>from fastai.vision.all import Image\nimport numpy as np\nimport os\nf = '9fad869f21b5b240.jpg'\nresized_from_dataset = Image.open(os.path.join('../input/resized-plant2021/img_sz_512/',f))\nresized_here = Image.open(os.path.join('../input/plant-pathology-2021-fgvc8/train_images/',f))\nresized_here = resized_here.resize((512,341),resample=Image.BILINEAR)\nnp.array(resized_from_dataset).astype(np.int32) - np.array(resized_here).astype(np.int32)\n</code></pre>\n<p>I am starting to suspect that the resized images undergo some changes at the moment when they are saved to disc. Hopefully it is not the case. What is the correct way to get resized_from_dataset==resized_here?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1246723,
          "author_name": "Igor Buinyi",
          "author_url": "",
          "post_date": "2021-03-21T03:08:46.170000",
          "content": "<p>Looks like saving to disc does distort the images… The default quality is 75.</p>\n<pre><code>from fastai.vision.all import Image\nimport numpy as np\nimport os\nf = '9fad869f21b5b240.jpg'\nresized_from_dataset = Image.open(os.path.join('../input/resized-plant2021/img_sz_512/',f))\nresized_here = Image.open(os.path.join('../input/plant-pathology-2021-fgvc8/train_images/',f))\nresized_here = resized_here.resize((512,341),resample=Image.BILINEAR)\n\nprint('Images are different:\\n', np.array(resized_from_dataset).astype(np.int32) - np.array(resized_here).astype(np.int32))\nresized_here.save('new_file_'+f, img_format='jpg')\nresized_here_recovered = Image.open('new_file_'+f)\nprint('\\nImages are the same:\\n', np.array(resized_from_dataset).astype(np.int32) - np.array(resized_here_recovered).astype(np.int32))\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1250302,
          "author_name": "AnkurSingh",
          "author_url": "",
          "post_date": "2021-03-23T23:33:35.870000",
          "content": "<blockquote>\n  <p>What is the correct way to get resized_from_dataset==resized_here?</p>\n</blockquote>\n<p>Interesting! I will have to look into it 😅</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1245672,
      "author_name": "Sh1r0",
      "author_url": "",
      "post_date": "2021-03-20T02:48:19.480000",
      "content": "<p>Thank you for sharing. It helped speed up learning.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1245845,
          "author_name": "AnkurSingh",
          "author_url": "",
          "post_date": "2021-03-20T08:49:28.587000",
          "content": "<p>Happy to help 😁</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1244399,
      "author_name": "Igor Buinyi",
      "author_url": "",
      "post_date": "2021-03-19T00:45:55.097000",
      "content": "<p>Edit: Actually the aspect ratio is preserved in the dataset. Size 512 actually means 342*512. Thanks for sharing.</p>\n<p>Hi! Thanks. <br>\nDo you think if there is a benefit in preserving the aspect ratio? E.g. creating a version with 512x766 size?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1244552,
          "author_name": "AnkurSingh",
          "author_url": "",
          "post_date": "2021-03-19T04:10:45.027000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/buinyi\" target=\"_blank\">@buinyi</a>, yes the actual aspect ratio is preserved. There are many different image sizes in the dataset </p>\n<pre><code>{(2672, 4000): 16485,\n(1728, 2592): 1027,\n(3000, 4000): 665,\n(3456, 4608): 123,\n(3456, 5184): 193,\n(3024, 4032): 132,\n(3024, 3024): 3,\n(4032, 3024): 3,\n(2248, 4000): 1}\n</code></pre>\n<p>One obvious benefit of preserving the aspect ratio is, not disturbing the image. If there is considerable difference in image's height and width, then making them equal will result in a distorted image. If you want square images, then it can be easily done later. </p>\n<p>For example: When loading the image, I am using <code>RandomizedCrop</code> to crop square areas from the image. This makes sure that you are not disturbing the image throughout the pipeline. </p>\n<p>Hope it helps, Cheers 😁</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1276686,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-17T20:56:51.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1248385,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-22T15:08:05.510000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1246236,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-20T15:08:37.423000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1246262,
          "author_name": "AnkurSingh",
          "author_url": "",
          "post_date": "2021-03-20T15:40:57.177000",
          "content": "<p><a href=\"https://www.kaggle.com/mervynyang\" target=\"_blank\">@mervynyang</a> there can be many possible reasons. 😅</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1244020": "In this competition, the original images are quite large (~(2672, 4000)). It took me 42 mins to train a ResNet50 (one epoch only). This is extremely slow to conduct any sort of experiment. \n\nSo, after some digging, I look at the image size and decide to create a new resized dataset. With resized dataset, it took only 6 mins (originally 42 mins) per epoch. \n\nThis is huge time saver (7x faster). Now, I can do more experiments, try different model, and do so much more. \n\nSo, I decided to create multiple resized versions (particularly 256, 384, 512, 640) of the dataset. You can find the [resized dataset here](https://www.kaggle.com/ankursingh12/resized-plant2021)\n\nAlso the code used for resizing the dataset can be found in the description.\nHope, all the participants will save at least a couple of hours by using it. Cheers!",
    "1248544": "Thank you so much. \n\nI am getting large difference between CV and LB. (CV : 0.603, LB : 0.164). [with the original dataset]. Are you facing any such problem. This is my [training notebook](https://www.kaggle.com/shanmukh05/plant-pathology-2k21-baseline-tpu-training).",
    "1349132": "i am using efficientnet _b2 pytorch ,input_size=260*260, batch_size=16, on kaggle notebook one epoch taking a lot of time it has been more than 15 min and it's still not done,any help?",
    "1308635": "hey, I'm trying different size ratio's using your code but getting outputs as visualization instead of folders. can you help me? @ankursingh12 ",
    "1252990": "Why not resize the images in the processing of training?",
    "1246708": "This is great for train images. But then we need the same resizing for test images too, and we want to make sure that the transformation is the same.\n\nSo, I tried to repeat the same transformation for some image, but all my approaches failed. For example, the following code were always returning slightly different output for resized_from_datased and resized_here, even though I tried different interpolations.\n```\nfrom fastai.vision.all import Image\nimport numpy as np\nimport os\nf = '9fad869f21b5b240.jpg'\nresized_from_dataset = Image.open(os.path.join('../input/resized-plant2021/img_sz_512/',f))\nresized_here = Image.open(os.path.join('../input/plant-pathology-2021-fgvc8/train_images/',f))\nresized_here = resized_here.resize((512,341),resample=Image.BILINEAR)\nnp.array(resized_from_dataset).astype(np.int32) - np.array(resized_here).astype(np.int32)\n```\n\nI am starting to suspect that the resized images undergo some changes at the moment when they are saved to disc. Hopefully it is not the case. What is the correct way to get resized_from_dataset==resized_here?\n\n",
    "1245672": "Thank you for sharing. It helped speed up learning.",
    "1244399": "Edit: Actually the aspect ratio is preserved in the dataset. Size 512 actually means 342*512. Thanks for sharing.\n\nHi! Thanks. \nDo you think if there is a benefit in preserving the aspect ratio? E.g. creating a version with 512x766 size?",
    "1276686": "",
    "1248385": "",
    "1246236": ""
  }
}