{
  "id": 171562,
  "title": "Fast way of reading images",
  "url": "/competitions/landmark-recognition-2020/discussion/171562",
  "author_name": "",
  "post_date": "2020-08-01T12:30:05.258262300Z",
  "votes": 11,
  "comment_count": 13,
  "views": 0,
  "content": "<p>The train set contains ~1.5 million images. Imageio is able to read ~20 images in a second on average (the average is based on first 38k images). Reading 1.5 million images with this average rate would give a total reading time of 20.8 hr. Any ideas how to reduce this?</p>\n\n<p>Update: I solved this problem using python multiprocessing. The total reading time is 1hr 46min now. The kernel is available <a href=\"https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size\">here</a>. This kernel also presents the image shape and channel information for all train images.</p>",
  "messages": [
    {
      "id": "954146",
      "postDate": "08/01/2020 12:30:05",
      "content": "<p>The train set contains ~1.5 million images. Imageio is able to read ~20 images in a second on average (the average is based on first 38k images). Reading 1.5 million images with this average rate would give a total reading time of 20.8 hr. Any ideas how to reduce this?</p>\n\n<p>Update: I solved this problem using python multiprocessing. The total reading time is 1hr 46min now. The kernel is available <a href=\"https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size\">here</a>. This kernel also presents the image shape and channel information for all train images.</p>",
      "rawMarkdown": "The train set contains ~1.5 million images. Imageio is able to read ~20 images in a second on average (the average is based on first 38k images). Reading 1.5 million images with this average rate would give a total reading time of 20.8 hr. Any ideas how to reduce this?\n\nUpdate: I solved this problem using python multiprocessing. The total reading time is 1hr 46min now. The kernel is available [here](https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size). This kernel also presents the image shape and channel information for all train images.",
      "votes": null
    },
    {
      "id": "954159",
      "postDate": "08/01/2020 12:58:49",
      "content": "<p>Have you done this with cv2?\nIf cv2 isn't faster you have to decode it into tfrecord or an other faster image type like bitmap and use it as dataset.</p>",
      "rawMarkdown": "Have you done this with cv2?\nIf cv2 isn't faster you have to decode it into tfrecord or an other faster image type like bitmap and use it as dataset.",
      "votes": null
    },
    {
      "id": "954180",
      "postDate": "08/01/2020 13:21:31",
      "content": "<p>Thank you! I haven't tried anything other than using imageio yet. I'll try your suggestions.</p>",
      "rawMarkdown": "Thank you! I haven't tried anything other than using imageio yet. I'll try your suggestions.",
      "votes": null
    },
    {
      "id": "954192",
      "postDate": "08/01/2020 13:32:33",
      "content": "<ul>\n<li>Resize and store smaller images</li>\n<li>Convert (already smaller resized images) as TFRecords, along with any metadata, for faster loading</li>\n<li>Use <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171513\">caching and other techniques</a></li>\n</ul>",
      "rawMarkdown": "Resize and store smaller images\n- Convert (already smaller resized images) as TFRecords, along with any metadata, for faster loading\n- Use [caching and other techniques](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171513)",
      "votes": null
    },
    {
      "id": "954215",
      "postDate": "08/01/2020 13:58:02",
      "content": "<p>Resizing the images requires loading the images and loading the images takes long. That's the problem.</p>",
      "rawMarkdown": "Resizing the images requires loading the images and loading the images takes long. That's the problem.",
      "votes": null
    },
    {
      "id": "954226",
      "postDate": "08/01/2020 14:14:18",
      "content": "<p>Yes, my tips were for loading images quickly during training.</p>\n\n<p>Regarding pre-processing of input, it may take time, but you will only be loading images one-time in offline mode to prepare TFRecords. You can also use multi-threading/multi-vm-processing to complete this step faster.</p>",
      "rawMarkdown": "Yes, my tips were for loading images quickly during training.\n\nRegarding pre-processing of input, it may take time, but you will only be loading images one-time in offline mode to prepare TFRecords. You can also use multi-threading/multi-vm-processing to complete this step faster.",
      "votes": null
    },
    {
      "id": "955276",
      "postDate": "08/02/2020 14:08:52",
      "content": "<p>Now I have a solution! I used multiprocessing to reduce the reading time. The total reading time is 1hr 46min now. The kernel is available <a href=\"https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size\">here</a>. This kernel also presents the image size and channel information for all train images.</p>",
      "rawMarkdown": "Now I have a solution! I used multiprocessing to reduce the reading time. The total reading time is 1hr 46min now. The kernel is available [here](https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size). This kernel also presents the image size and channel information for all train images.",
      "votes": null
    },
    {
      "id": "960840",
      "postDate": "08/06/2020 18:03:51",
      "content": "<p>Well done, thanks for sharing your work! Is imageio the fastest CV library for reading images? Have you tried opencv or PIL? </p>",
      "rawMarkdown": "Well done, thanks for sharing your work! Is imageio the fastest CV library for reading images? Have you tried opencv or PIL?",
      "votes": null
    },
    {
      "id": "960846",
      "postDate": "08/06/2020 18:10:09",
      "content": "<p>I have tried only cv2 as an alternative. Its reading rate is similar to that of imageio.</p>",
      "rawMarkdown": "I have tried only cv2 as an alternative. Its reading rate is similar to that of imageio.",
      "votes": null
    },
    {
      "id": "961771",
      "postDate": "08/07/2020 13:39:03",
      "content": "<p>How do handle such huge dataset ? Do you use images of low resolution or maybe something else?</p>",
      "rawMarkdown": "How do handle such huge dataset ? Do you use images of low resolution or maybe something else?",
      "votes": null
    },
    {
      "id": "961778",
      "postDate": "08/07/2020 13:48:04",
      "content": "<p>Could you be more specific about handling?</p>",
      "rawMarkdown": "Could you be more specific about handling?",
      "votes": null
    },
    {
      "id": "962728",
      "postDate": "08/08/2020 11:40:23",
      "content": "<p>As in using the data locally on our machine is not possible due large size.  So is their any image data of resized small size</p>",
      "rawMarkdown": "As in using the data locally on our machine is not possible due large size.  So is their any image data of resized small size",
      "votes": null
    },
    {
      "id": "965159",
      "postDate": "08/10/2020 12:25:12",
      "content": "<p>The general approach is to use gpu, tpu, multiprocessing on cpu, or a combination of all. Otherwise you can't process 1.5 million images in a reasonably short time.</p>",
      "rawMarkdown": "The general approach is to use gpu, tpu, multiprocessing on cpu, or a combination of all. Otherwise you can't process 1.5 million images in a reasonably short time.",
      "votes": null
    },
    {
      "id": "968539",
      "postDate": "08/13/2020 05:08:02",
      "content": "<p>Can the same not be achieved using ImageDataGenerator???<br>\nI am not sure. </p>",
      "rawMarkdown": "Can the same not be achieved using ImageDataGenerator???\nI am not sure.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961771,
      "author_name": "swarajshinde",
      "author_url": "",
      "post_date": "08/07/2020 13:39:03",
      "content": "<p>How do handle such huge dataset ? Do you use images of low resolution or maybe something else?</p>",
      "votes": null,
      "replies": [
        {
          "id": 961778,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "08/07/2020 13:48:04",
          "content": "<p>Could you be more specific about handling?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962728,
          "author_name": "swarajshinde",
          "author_url": "",
          "post_date": "08/08/2020 11:40:23",
          "content": "<p>As in using the data locally on our machine is not possible due large size.  So is their any image data of resized small size</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965159,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "08/10/2020 12:25:12",
          "content": "<p>The general approach is to use gpu, tpu, multiprocessing on cpu, or a combination of all. Otherwise you can't process 1.5 million images in a reasonably short time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 968539,
      "author_name": "palaksood97",
      "author_url": "",
      "post_date": "08/13/2020 05:08:02",
      "content": "<p>Can the same not be achieved using ImageDataGenerator???<br>\nI am not sure. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 954159,
      "author_name": "derinformatiker",
      "author_url": "",
      "post_date": "08/01/2020 12:58:49",
      "content": "<p>Have you done this with cv2?\nIf cv2 isn't faster you have to decode it into tfrecord or an other faster image type like bitmap and use it as dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 954180,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "08/01/2020 13:21:31",
          "content": "<p>Thank you! I haven't tried anything other than using imageio yet. I'll try your suggestions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 954192,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "08/01/2020 13:32:33",
      "content": "<ul>\n<li>Resize and store smaller images</li>\n<li>Convert (already smaller resized images) as TFRecords, along with any metadata, for faster loading</li>\n<li>Use <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171513\">caching and other techniques</a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 954215,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "08/01/2020 13:58:02",
          "content": "<p>Resizing the images requires loading the images and loading the images takes long. That's the problem.</p>",
          "votes": null,
          "replies": [
            {
              "id": 954226,
              "author_name": "sirishks",
              "author_url": "",
              "post_date": "08/01/2020 14:14:18",
              "content": "<p>Yes, my tips were for loading images quickly during training.</p>\n\n<p>Regarding pre-processing of input, it may take time, but you will only be loading images one-time in offline mode to prepare TFRecords. You can also use multi-threading/multi-vm-processing to complete this step faster.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 955276,
      "author_name": "tolgadincer",
      "author_url": "",
      "post_date": "08/02/2020 14:08:52",
      "content": "<p>Now I have a solution! I used multiprocessing to reduce the reading time. The total reading time is 1hr 46min now. The kernel is available <a href=\"https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size\">here</a>. This kernel also presents the image size and channel information for all train images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 960840,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/06/2020 18:03:51",
          "content": "<p>Well done, thanks for sharing your work! Is imageio the fastest CV library for reading images? Have you tried opencv or PIL? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960846,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "08/06/2020 18:10:09",
          "content": "<p>I have tried only cv2 as an alternative. Its reading rate is similar to that of imageio.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "954146": "The train set contains ~1.5 million images. Imageio is able to read ~20 images in a second on average (the average is based on first 38k images). Reading 1.5 million images with this average rate would give a total reading time of 20.8 hr. Any ideas how to reduce this?\n\nUpdate: I solved this problem using python multiprocessing. The total reading time is 1hr 46min now. The kernel is available [here](https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size). This kernel also presents the image shape and channel information for all train images.",
    "954159": "Have you done this with cv2?\nIf cv2 isn't faster you have to decode it into tfrecord or an other faster image type like bitmap and use it as dataset.",
    "954180": "Thank you! I haven't tried anything other than using imageio yet. I'll try your suggestions.",
    "954192": "Resize and store smaller images\n- Convert (already smaller resized images) as TFRecords, along with any metadata, for faster loading\n- Use [caching and other techniques](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171513)",
    "954215": "Resizing the images requires loading the images and loading the images takes long. That's the problem.",
    "954226": "Yes, my tips were for loading images quickly during training.\n\nRegarding pre-processing of input, it may take time, but you will only be loading images one-time in offline mode to prepare TFRecords. You can also use multi-threading/multi-vm-processing to complete this step faster.",
    "955276": "Now I have a solution! I used multiprocessing to reduce the reading time. The total reading time is 1hr 46min now. The kernel is available [here](https://www.kaggle.com/tolgadincer/landmark-recognition-multiprocessing-image-size). This kernel also presents the image size and channel information for all train images.",
    "960840": "Well done, thanks for sharing your work! Is imageio the fastest CV library for reading images? Have you tried opencv or PIL?",
    "960846": "I have tried only cv2 as an alternative. Its reading rate is similar to that of imageio.",
    "961771": "How do handle such huge dataset ? Do you use images of low resolution or maybe something else?",
    "961778": "Could you be more specific about handling?",
    "962728": "As in using the data locally on our machine is not possible due large size.  So is their any image data of resized small size",
    "965159": "The general approach is to use gpu, tpu, multiprocessing on cpu, or a combination of all. Otherwise you can't process 1.5 million images in a reasonably short time.",
    "968539": "Can the same not be achieved using ImageDataGenerator???\nI am not sure."
  },
  "source": "meta"
}