{
  "id": 228426,
  "title": "How do you load the images without it taking very long time?",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/228426",
  "author_name": "",
  "post_date": "2021-03-24T16:38:58.287390200Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello,<br>\nI am pretty new to Kaggle, I only completed the Titanic competition.<br>\nI have one very basic question, how usually people handled the loading of the img such that it doesn't take too much time or memory?</p>",
  "messages": [
    {
      "id": "1251284",
      "postDate": "03/24/2021 16:38:58",
      "content": "<p>Hello,<br>\nI am pretty new to Kaggle, I only completed the Titanic competition.<br>\nI have one very basic question, how usually people handled the loading of the img such that it doesn't take too much time or memory?</p>",
      "rawMarkdown": "Hello,\nI am pretty new to Kaggle, I only completed the Titanic competition.\nI have one very basic question, how usually people handled the loading of the img such that it doesn't take too much time or memory?",
      "votes": null
    },
    {
      "id": "1251956",
      "postDate": "03/25/2021 09:35:40",
      "content": "<p>Hello!</p>\n<p>I'm afraid, your question is too general to give a proper answer. I suggest picking a good-looking notebook written in a framework you're familiar with (there's a bunch of them listed in <strong><a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226316\" target=\"_blank\">this discussion</a></strong> thanks to <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a>) as a point of quick start. The more you tweak it, the more you learn from it.<br>\nConsidering you are a Keras user, some points to consider for faster speed and less memory consumption:</p>\n<ul>\n<li>do not use original image size (it's huge even for TPU), consider starting from 384x384 or 512x512 size</li>\n<li>consider switching from DataFrame or NumPy arrays feeding to efficient and highly optimized <code>tf.data.Dataset</code> or <code>tf.data.TFRecordDataset</code> (with .tfrec files) classes. These pipelines are way harder to understand from the first glance, but that's a reasonable cost for a tremendous performance boost (for example, see <strong><a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training\" target=\"_blank\">this notebook</a></strong> )</li>\n<li>when running on GPU turn on mixed-precision (TPUs use mixed-precision by default): this means all your tensors are stored in <code>float32</code> format (as usual) but all the computations in between are carried in <code>float16</code> format. Add this line to the top of your notebook:</li>\n</ul>\n<pre><code>tf.keras.mixed_precision.set_global_policy('mixed_float16')\n</code></pre>\n<p>and make sure the last layer of your model ends on activation with specified <code>dtype</code> argument explicitly:</p>\n<pre><code>tf.keras.layers.Activation('sigmoid', dtype='float32')\n</code></pre>\n<p>Easy as that. With mixed-precision, you will be able to run nearly 2 times faster and double your batch size.</p>",
      "rawMarkdown": "Hello!\n\nI'm afraid, your question is too general to give a proper answer. I suggest picking a good-looking notebook written in a framework you're familiar with (there's a bunch of them listed in **[this discussion](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226316)** thanks to @usharengaraju) as a point of quick start. The more you tweak it, the more you learn from it.\nConsidering you are a Keras user, some points to consider for faster speed and less memory consumption:\n* do not use original image size (it's huge even for TPU), consider starting from 384x384 or 512x512 size\n* consider switching from DataFrame or NumPy arrays feeding to efficient and highly optimized `tf.data.Dataset` or `tf.data.TFRecordDataset` (with .tfrec files) classes. These pipelines are way harder to understand from the first glance, but that's a reasonable cost for a tremendous performance boost (for example, see **[this notebook](https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training)** )\n* when running on GPU turn on mixed-precision (TPUs use mixed-precision by default): this means all your tensors are stored in `float32` format (as usual) but all the computations in between are carried in `float16` format. Add this line to the top of your notebook:\n```\ntf.keras.mixed_precision.set_global_policy('mixed_float16')\n```\nand make sure the last layer of your model ends on activation with specified `dtype` argument explicitly:\n```\ntf.keras.layers.Activation('sigmoid', dtype='float32')\n```\nEasy as that. With mixed-precision, you will be able to run nearly 2 times faster and double your batch size.",
      "votes": null
    },
    {
      "id": "1251970",
      "postDate": "03/25/2021 09:55:43",
      "content": "<p>Thanks for your answer but I think my question is even more basic than that!</p>\n<p>Most basic models are asking np.array instead of .jpg, so it is required to converted all of the .jpg files into ndarray. To do so I use this piece of code:</p>\n<pre><code>def load_image(path):\n    img_path = '/kaggle/input/plant-pathology-2021-fgvc8/train_images/' + path\n    img = PIL.Image.open(img_path)\n    img = np.asarray(img.resize((512, 512)))\n    return np.ravel(img)\n\ndata = pd.read_csv(\"/kaggle/input/plant-pathology-2021-fgvc8/train.csv\")\ndata['loaded_image'] = np.nan\ndata['loaded_image'] = data['image'].apply(load_image)\n</code></pre>\n<p>(Quite similar to what is done here: <a href=\"https://www.kaggle.com/tarunpaparaju/plant-pathology-2020-eda-models\" target=\"_blank\">https://www.kaggle.com/tarunpaparaju/plant-pathology-2020-eda-models</a> )</p>\n<p>But this seems to take forever to run (which kind of make sense)</p>\n<p>Is it the correct way to proceed?</p>",
      "rawMarkdown": "Thanks for your answer but I think my question is even more basic than that!\n\nMost basic models are asking np.array instead of .jpg, so it is required to converted all of the .jpg files into ndarray. To do so I use this piece of code:\n\n```\ndef load_image(path):\n    img_path = '/kaggle/input/plant-pathology-2021-fgvc8/train_images/' + path\n    img = PIL.Image.open(img_path)\n    img = np.asarray(img.resize((512, 512)))\n    return np.ravel(img)\n\ndata = pd.read_csv(\"/kaggle/input/plant-pathology-2021-fgvc8/train.csv\")\ndata['loaded_image'] = np.nan\ndata['loaded_image'] = data['image'].apply(load_image)\n```\n\n(Quite similar to what is done here: https://www.kaggle.com/tarunpaparaju/plant-pathology-2020-eda-models )\n\nBut this seems to take forever to run (which kind of make sense)\n\nIs it the correct way to proceed?",
      "votes": null
    },
    {
      "id": "1251988",
      "postDate": "03/25/2021 10:17:37",
      "content": "<p>By doing so you're trying to load the entire dataset of 3500x3500 images into memory which would definitely end up with OOM even if it is able to finish. As I previously mentioned, downscaling is required first, which can be done by various methods. I also suggest tracking your progress which is easy with <code>tqdm</code>. <br>\nE.g. you can use tensorflow for downscaling and track your progress like this:</p>\n<pre><code>import tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport os\n\nroot = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\n\ndef load_image(path):\n    img = tf.io.read_file(path)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.resize(img, [512, 512])\n    img = tf.cast(img, tf.uint8)\n    return img.numpy()\n\ndf = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv')\n\ndf['image'] = [os.path.join(root, x) for x in df['image']]\n\ndata = np.fromiter([load_image(x) for x in tqdm(df['image'], total=len(df))])\nprint(data.shape)\n</code></pre>\n<p>Set lower image size if you get an OOM.</p>",
      "rawMarkdown": "By doing so you're trying to load the entire dataset of 3500x3500 images into memory which would definitely end up with OOM even if it is able to finish. As I previously mentioned, downscaling is required first, which can be done by various methods. I also suggest tracking your progress which is easy with `tqdm`. \nE.g. you can use tensorflow for downscaling and track your progress like this:\n```\nimport tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport os\n\nroot = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\n\ndef load_image(path):\n    img = tf.io.read_file(path)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.resize(img, [512, 512])\n    img = tf.cast(img, tf.uint8)\n    return img.numpy()\n\ndf = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv')\n\ndf['image'] = [os.path.join(root, x) for x in df['image']]\n\ndata = np.fromiter([load_image(x) for x in tqdm(df['image'], total=len(df))])\nprint(data.shape)\n```\nSet lower image size if you get an OOM.",
      "votes": null
    },
    {
      "id": "1251995",
      "postDate": "03/25/2021 10:24:32",
      "content": "<p>Thanks for your help! I think I'll try to tackle some of the \"Getting Started\" first before coming back to this one. :D</p>",
      "rawMarkdown": "Thanks for your help! I think I'll try to tackle some of the \"Getting Started\" first before coming back to this one. :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1251956,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "03/25/2021 09:35:40",
      "content": "<p>Hello!</p>\n<p>I'm afraid, your question is too general to give a proper answer. I suggest picking a good-looking notebook written in a framework you're familiar with (there's a bunch of them listed in <strong><a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226316\" target=\"_blank\">this discussion</a></strong> thanks to <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a>) as a point of quick start. The more you tweak it, the more you learn from it.<br>\nConsidering you are a Keras user, some points to consider for faster speed and less memory consumption:</p>\n<ul>\n<li>do not use original image size (it's huge even for TPU), consider starting from 384x384 or 512x512 size</li>\n<li>consider switching from DataFrame or NumPy arrays feeding to efficient and highly optimized <code>tf.data.Dataset</code> or <code>tf.data.TFRecordDataset</code> (with .tfrec files) classes. These pipelines are way harder to understand from the first glance, but that's a reasonable cost for a tremendous performance boost (for example, see <strong><a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training\" target=\"_blank\">this notebook</a></strong> )</li>\n<li>when running on GPU turn on mixed-precision (TPUs use mixed-precision by default): this means all your tensors are stored in <code>float32</code> format (as usual) but all the computations in between are carried in <code>float16</code> format. Add this line to the top of your notebook:</li>\n</ul>\n<pre><code>tf.keras.mixed_precision.set_global_policy('mixed_float16')\n</code></pre>\n<p>and make sure the last layer of your model ends on activation with specified <code>dtype</code> argument explicitly:</p>\n<pre><code>tf.keras.layers.Activation('sigmoid', dtype='float32')\n</code></pre>\n<p>Easy as that. With mixed-precision, you will be able to run nearly 2 times faster and double your batch size.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1251970,
          "author_name": "coilysum",
          "author_url": "",
          "post_date": "03/25/2021 09:55:43",
          "content": "<p>Thanks for your answer but I think my question is even more basic than that!</p>\n<p>Most basic models are asking np.array instead of .jpg, so it is required to converted all of the .jpg files into ndarray. To do so I use this piece of code:</p>\n<pre><code>def load_image(path):\n    img_path = '/kaggle/input/plant-pathology-2021-fgvc8/train_images/' + path\n    img = PIL.Image.open(img_path)\n    img = np.asarray(img.resize((512, 512)))\n    return np.ravel(img)\n\ndata = pd.read_csv(\"/kaggle/input/plant-pathology-2021-fgvc8/train.csv\")\ndata['loaded_image'] = np.nan\ndata['loaded_image'] = data['image'].apply(load_image)\n</code></pre>\n<p>(Quite similar to what is done here: <a href=\"https://www.kaggle.com/tarunpaparaju/plant-pathology-2020-eda-models\" target=\"_blank\">https://www.kaggle.com/tarunpaparaju/plant-pathology-2020-eda-models</a> )</p>\n<p>But this seems to take forever to run (which kind of make sense)</p>\n<p>Is it the correct way to proceed?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251988,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "03/25/2021 10:17:37",
          "content": "<p>By doing so you're trying to load the entire dataset of 3500x3500 images into memory which would definitely end up with OOM even if it is able to finish. As I previously mentioned, downscaling is required first, which can be done by various methods. I also suggest tracking your progress which is easy with <code>tqdm</code>. <br>\nE.g. you can use tensorflow for downscaling and track your progress like this:</p>\n<pre><code>import tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport os\n\nroot = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\n\ndef load_image(path):\n    img = tf.io.read_file(path)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.resize(img, [512, 512])\n    img = tf.cast(img, tf.uint8)\n    return img.numpy()\n\ndf = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv')\n\ndf['image'] = [os.path.join(root, x) for x in df['image']]\n\ndata = np.fromiter([load_image(x) for x in tqdm(df['image'], total=len(df))])\nprint(data.shape)\n</code></pre>\n<p>Set lower image size if you get an OOM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251995,
          "author_name": "coilysum",
          "author_url": "",
          "post_date": "03/25/2021 10:24:32",
          "content": "<p>Thanks for your help! I think I'll try to tackle some of the \"Getting Started\" first before coming back to this one. :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1251284": "Hello,\nI am pretty new to Kaggle, I only completed the Titanic competition.\nI have one very basic question, how usually people handled the loading of the img such that it doesn't take too much time or memory?",
    "1251956": "Hello!\n\nI'm afraid, your question is too general to give a proper answer. I suggest picking a good-looking notebook written in a framework you're familiar with (there's a bunch of them listed in **[this discussion](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226316)** thanks to @usharengaraju) as a point of quick start. The more you tweak it, the more you learn from it.\nConsidering you are a Keras user, some points to consider for faster speed and less memory consumption:\n* do not use original image size (it's huge even for TPU), consider starting from 384x384 or 512x512 size\n* consider switching from DataFrame or NumPy arrays feeding to efficient and highly optimized `tf.data.Dataset` or `tf.data.TFRecordDataset` (with .tfrec files) classes. These pipelines are way harder to understand from the first glance, but that's a reasonable cost for a tremendous performance boost (for example, see **[this notebook](https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training)** )\n* when running on GPU turn on mixed-precision (TPUs use mixed-precision by default): this means all your tensors are stored in `float32` format (as usual) but all the computations in between are carried in `float16` format. Add this line to the top of your notebook:\n```\ntf.keras.mixed_precision.set_global_policy('mixed_float16')\n```\nand make sure the last layer of your model ends on activation with specified `dtype` argument explicitly:\n```\ntf.keras.layers.Activation('sigmoid', dtype='float32')\n```\nEasy as that. With mixed-precision, you will be able to run nearly 2 times faster and double your batch size.",
    "1251970": "Thanks for your answer but I think my question is even more basic than that!\n\nMost basic models are asking np.array instead of .jpg, so it is required to converted all of the .jpg files into ndarray. To do so I use this piece of code:\n\n```\ndef load_image(path):\n    img_path = '/kaggle/input/plant-pathology-2021-fgvc8/train_images/' + path\n    img = PIL.Image.open(img_path)\n    img = np.asarray(img.resize((512, 512)))\n    return np.ravel(img)\n\ndata = pd.read_csv(\"/kaggle/input/plant-pathology-2021-fgvc8/train.csv\")\ndata['loaded_image'] = np.nan\ndata['loaded_image'] = data['image'].apply(load_image)\n```\n\n(Quite similar to what is done here: https://www.kaggle.com/tarunpaparaju/plant-pathology-2020-eda-models )\n\nBut this seems to take forever to run (which kind of make sense)\n\nIs it the correct way to proceed?",
    "1251988": "By doing so you're trying to load the entire dataset of 3500x3500 images into memory which would definitely end up with OOM even if it is able to finish. As I previously mentioned, downscaling is required first, which can be done by various methods. I also suggest tracking your progress which is easy with `tqdm`. \nE.g. you can use tensorflow for downscaling and track your progress like this:\n```\nimport tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport os\n\nroot = '/kaggle/input/plant-pathology-2021-fgvc8/train_images'\n\ndef load_image(path):\n    img = tf.io.read_file(path)\n    img = tf.image.decode_jpeg(img, channels=3)\n    img = tf.image.resize(img, [512, 512])\n    img = tf.cast(img, tf.uint8)\n    return img.numpy()\n\ndf = pd.read_csv('/kaggle/input/plant-pathology-2021-fgvc8/train.csv')\n\ndf['image'] = [os.path.join(root, x) for x in df['image']]\n\ndata = np.fromiter([load_image(x) for x in tqdm(df['image'], total=len(df))])\nprint(data.shape)\n```\nSet lower image size if you get an OOM.",
    "1251995": "Thanks for your help! I think I'll try to tackle some of the \"Getting Started\" first before coming back to this one. :D"
  },
  "source": "meta"
}