{
  "id": 210299,
  "title": "How to do predictions faster with Tensorflow?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210299",
  "author_name": "",
  "post_date": "2021-01-10T10:33:53.892168400Z",
  "votes": 2,
  "comment_count": 17,
  "views": 0,
  "content": "<p>list_of_preds_per_class = model.predict(array_of_images)<br>\nWill update this for better clarity.<br>\n<a href=\"https://youtu.be/RqLD1INA_cQ\" target=\"_blank\">How to do predictions with tensorflow</a><br>\n<code>### libraries</code><br>\n<code>import tensorflow as tf #load model #load image</code><br>\n<code>from PIL import Image</code></p>\n<p><code>test_dir = '../input/cassava-leaf-disease-classification/test_images/'</code></p>\n<h3>takes 0.08 seconds per image.</h3>\n<p>(This was tested on 3500 train images took 281 seconds)</p>\n<h4>The code</h4>\n<p><code>import time</code><br>\n<code>start = time.process_time()</code><br>\n<code></code><br>\n<code>preds = []</code><br>\n<code>for img in test_images:</code><br>\n<code>img = Image.open(test_dir + img)</code><br>\n<code>img = np.expand_dims(img, axis=0) #add extra dimension don't know why yet.</code><br>\n<code>preds.extend(model.predict(img).argmax(axis = 1)) #predict list of probabilities per class then takes max prediction class as prediction.</code></p>\n<p><code>print('Time taken in seconds : ',time.process_time() - start)</code></p>\n<h3>takes 0.073 seconds per image</h3>\n<p>(This was tested on 3500 train images took 256 seconds)</p>\n<h4>The code</h4>\n<p><code>import time</code><br>\n<code>start = time.process_time()</code></p>\n<p><code>preds = []</code><br>\n<code>for img in test_images:</code><br>\n<code>img = tf.keras.preprocessing.image.load_img(test_dir + img) #loads image</code><br>\n<code>img = np.expand_dims(img, axis=0) #add extra dimension dont know why yet.</code><br>\n<code>preds.extend(model.predict(img).argmax(axis = 1)) #predict list of probabilities per class then takes max prediction class as prediction.</code></p>\n<p><code>print('Time taken in seconds : ',time.process_time() - start)</code></p>\n<h3>Any faster Ideas?</h3>\n<p>Some very useful comments below on using Generators.💯 💯<br>\nAlso someone emailed me this link:<br>\n<a href=\"https://stanford.edu/~shervine/blog/keras-how-to-generate-data-on-the-fly\" target=\"_blank\">keras-how-to-generate-data-on-the-fly-Stanford</a></p>\n<h5>Faster Idea 1 (from comments)</h5>\n<h3>1. Using <code>tf.keras.utils.Sequence</code>,</h3>\n<h4>code from code cell 33 <a href=\"https://www.kaggle.com/awsaf49/efficientnetb6-512-cutmixupdropout-tpu-infer/notebook?scriptVersionId=47469988\" target=\"_blank\">here as mentioned by Chris Deotte in the comments and coded by awsaf49</a>  and using <code>model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)</code></h4>\n<h2>0.0356 seconds per image.</h2>\n<p>1.8 seconds on a single image and 761 seconds on 21397 images</p>\n<h3>The code.</h3>\n<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/211269\" target=\"_blank\">Link to Code , it a bit long that why its on another discussion</a></p>\n<h4>2. [Up Next] Using tf.data.Dataset</h4>",
  "messages": [
    {
      "id": "1147171",
      "postDate": "01/10/2021 10:33:53",
      "content": "<p>list_of_preds_per_class = model.predict(array_of_images)<br>\nWill update this for better clarity.<br>\n<a href=\"https://youtu.be/RqLD1INA_cQ\" target=\"_blank\">How to do predictions with tensorflow</a><br>\n<code>### libraries</code><br>\n<code>import tensorflow as tf #load model #load image</code><br>\n<code>from PIL import Image</code></p>\n<p><code>test_dir = '../input/cassava-leaf-disease-classification/test_images/'</code></p>\n<h3>takes 0.08 seconds per image.</h3>\n<p>(This was tested on 3500 train images took 281 seconds)</p>\n<h4>The code</h4>\n<p><code>import time</code><br>\n<code>start = time.process_time()</code><br>\n<code></code><br>\n<code>preds = []</code><br>\n<code>for img in test_images:</code><br>\n<code>img = Image.open(test_dir + img)</code><br>\n<code>img = np.expand_dims(img, axis=0) #add extra dimension don't know why yet.</code><br>\n<code>preds.extend(model.predict(img).argmax(axis = 1)) #predict list of probabilities per class then takes max prediction class as prediction.</code></p>\n<p><code>print('Time taken in seconds : ',time.process_time() - start)</code></p>\n<h3>takes 0.073 seconds per image</h3>\n<p>(This was tested on 3500 train images took 256 seconds)</p>\n<h4>The code</h4>\n<p><code>import time</code><br>\n<code>start = time.process_time()</code></p>\n<p><code>preds = []</code><br>\n<code>for img in test_images:</code><br>\n<code>img = tf.keras.preprocessing.image.load_img(test_dir + img) #loads image</code><br>\n<code>img = np.expand_dims(img, axis=0) #add extra dimension dont know why yet.</code><br>\n<code>preds.extend(model.predict(img).argmax(axis = 1)) #predict list of probabilities per class then takes max prediction class as prediction.</code></p>\n<p><code>print('Time taken in seconds : ',time.process_time() - start)</code></p>\n<h3>Any faster Ideas?</h3>\n<p>Some very useful comments below on using Generators.💯 💯<br>\nAlso someone emailed me this link:<br>\n<a href=\"https://stanford.edu/~shervine/blog/keras-how-to-generate-data-on-the-fly\" target=\"_blank\">keras-how-to-generate-data-on-the-fly-Stanford</a></p>\n<h5>Faster Idea 1 (from comments)</h5>\n<h3>1. Using <code>tf.keras.utils.Sequence</code>,</h3>\n<h4>code from code cell 33 <a href=\"https://www.kaggle.com/awsaf49/efficientnetb6-512-cutmixupdropout-tpu-infer/notebook?scriptVersionId=47469988\" target=\"_blank\">here as mentioned by Chris Deotte in the comments and coded by awsaf49</a>  and using <code>model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)</code></h4>\n<h2>0.0356 seconds per image.</h2>\n<p>1.8 seconds on a single image and 761 seconds on 21397 images</p>\n<h3>The code.</h3>\n<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/211269\" target=\"_blank\">Link to Code , it a bit long that why its on another discussion</a></p>\n<h4>2. [Up Next] Using tf.data.Dataset</h4>",
      "rawMarkdown": "list_of_preds_per_class = model.predict(array_of_images)\nWill update this for better clarity.\n[How to do predictions with tensorflow](https://youtu.be/RqLD1INA_cQ)\n`### libraries`\n`import tensorflow as tf #load model #load image `\n`from PIL import Image `\n\n`test_dir = '../input/cassava-leaf-disease-classification/test_images/'`\n\n\n###  takes 0.08 seconds per image.\n(This was tested on 3500 train images took 281 seconds)\n#### The code\n`import time`\n`start = time.process_time()`\n`  `\n`preds = []`\n`for img in test_images:`\n`    img = Image.open(test_dir + img)`\n`    img = np.expand_dims(img, axis=0) #add extra dimension don't know why yet.`\n`    preds.extend(model.predict(img).argmax(axis = 1)) #predict list of probabilities per class then takes max prediction class as prediction. `\n\n`print('Time taken in seconds : ',time.process_time() - start)`\n\n### takes 0.073 seconds per image \n(This was tested on 3500 train images took 256 seconds)\n#### The code\n`import time`\n`start = time.process_time()`\n\n`preds = []`\n`for img in test_images:`\n`    img = tf.keras.preprocessing.image.load_img(test_dir + img) #loads image `\n`    img = np.expand_dims(img, axis=0) #add extra dimension dont know why yet.`\n`    preds.extend(model.predict(img).argmax(axis = 1)) #predict list of probabilities per class then takes max prediction class as prediction. `\n\n`print('Time taken in seconds : ',time.process_time() - start)`\n\n\n\n\n### Any faster Ideas?\n\nSome very useful comments below on using Generators.💯 💯\nAlso someone emailed me this link:\n[keras-how-to-generate-data-on-the-fly-Stanford](https://stanford.edu/~shervine/blog/keras-how-to-generate-data-on-the-fly)\n\n\n##### Faster Idea 1 (from comments)\n### 1. Using `tf.keras.utils.Sequence `,\n#### code from code cell 33 [here as mentioned by Chris Deotte in the comments and coded by awsaf49](https://www.kaggle.com/awsaf49/efficientnetb6-512-cutmixupdropout-tpu-infer/notebook?scriptVersionId=47469988)  and using `model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)`\n## 0.0356 seconds per image.\n 1.8 seconds on a single image and 761 seconds on 21397 images\n\n### The code.\n[Link to Code , it a bit long that why its on another discussion](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/211269)\n\n\n#### 2. [Up Next] Using tf.data.Dataset",
      "votes": null
    },
    {
      "id": "1147827",
      "postDate": "01/10/2021 18:22:21",
      "content": "<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207280\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207280</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207280",
      "votes": null
    },
    {
      "id": "1147906",
      "postDate": "01/10/2021 19:49:05",
      "content": "<p>If you're interested, for the line : <br>\nimg = np.expand_dims(img, axis=0) #add extra dimension don't know why yet.<br>\nan extra dimension is added because in TensorFlow, batches of images are 4 dimensional arrays of size [batch_size, height, width, channels]. Even if your batch contains only one image, it has to be a 4d array, with batch_size = 1.</p>\n<p>Maybe using ImageDataGenerator and flow_from_directory or flow_from_dataframe would be faster, as it loads batches of images, which might be faster than loading each image one by one (if the source code is more efficient than pure Python code, but I haven't checked). At least, if you use ImageDataGenerator for training, it is very convenient (in my opinion) to use it for prediction too.</p>",
      "rawMarkdown": "If you're interested, for the line : \nimg = np.expand_dims(img, axis=0) #add extra dimension don't know why yet.\nan extra dimension is added because in TensorFlow, batches of images are 4 dimensional arrays of size [batch_size, height, width, channels]. Even if your batch contains only one image, it has to be a 4d array, with batch_size = 1.\n\nMaybe using ImageDataGenerator and flow_from_directory or flow_from_dataframe would be faster, as it loads batches of images, which might be faster than loading each image one by one (if the source code is more efficient than pure Python code, but I haven't checked). At least, if you use ImageDataGenerator for training, it is very convenient (in my opinion) to use it for prediction too.",
      "votes": null
    },
    {
      "id": "1147992",
      "postDate": "01/10/2021 20:52:28",
      "content": "<p>Hey and  Thank you for your reply.<br>\nI have tried using ImageDataGenerator and flow_from_directory, I have not gotten the hang of it yet and Errors! 😪 Also I have not found a way to manipulate the data in that state. Still looking at code snippets to understand this part better.</p>",
      "rawMarkdown": "Hey and  Thank you for your reply.\nI have tried using ImageDataGenerator and flow_from_directory, I have not gotten the hang of it yet and Errors! 😪 Also I have not found a way to manipulate the data in that state. Still looking at code snippets to understand this part better.",
      "votes": null
    },
    {
      "id": "1148002",
      "postDate": "01/10/2021 20:56:08",
      "content": "<p>Thank you for this. Might you take a minute to explain this code please. <br>\n<code>@tf.function(input_signature=[</code><br>\n<code>tf.TensorSpec(shape=None, dtype=tf.int64),</code><br>\n<code>tf.TensorSpec(shape=None, dtype=tf.int32),</code><br>\n<code>tf.TensorSpec(shape=[], dtype=tf.bool),</code><br>\n<code>]</code><br>\n<code>)</code><br>\n <code>def call(self, user_id, content_id, training=False):</code><br>\n<code>N = len(user_ix)</code><br>\n<code>user_id = tf.reshape(user_id, (-1,))</code><br>\n<code>content_id = tf.reshape(content_id, (-1,))</code></p>",
      "rawMarkdown": "Thank you for this. Might you take a minute to explain this code please. \n`@tf.function(input_signature=[`\n`        tf.TensorSpec(shape=None, dtype=tf.int64),`\n`        tf.TensorSpec(shape=None, dtype=tf.int32),`\n`        tf.TensorSpec(shape=[], dtype=tf.bool),`\n`    ]`\n`    )`\n `   def call(self, user_id, content_id, training=False):`\n`        N = len(user_ix)`\n`        user_id = tf.reshape(user_id, (-1,))`\n`        content_id = tf.reshape(content_id, (-1,))`",
      "votes": null
    },
    {
      "id": "1148214",
      "postDate": "01/11/2021 02:24:46",
      "content": "<p>Hi, ignore my comment below. Using <code>@tf.function()</code> will only help if the bottleneck is your model. </p>\n<p>Your bottleneck is reading images from disk, and starving your GPU by inferring only 1 image at a time. To increase speed, you will want to read multiple images from the disk in parallel and feed your GPU larger batches. (GPUs work faster when they have lots of work to do)(and disk I/O is faster with multiprocessing).</p>\n<p>You will need to create a dataloader to feed your model batches of images during inference, and enable I/O multiprocessing to read from disk. Below are 3 choices with code examples</p>\n<ul>\n<li><code>tf.keras.utils.Sequence</code> this is my favorite dataloader. There is an example in code cell 33 <a href=\"https://www.kaggle.com/awsaf49/efficientnetb6-512-cutmixupdropout-tpu-infer/notebook?scriptVersionId=47469988\" target=\"_blank\">here</a>. It is the easiest to add augmentation to. When you call predict <code>model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)</code> enable multiprocessing and workers </li>\n<li><code>tf.data.Dataset</code>. This can be used with images (does not need TFRecords). There is an example in code cell 6 <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference?scriptVersionId=49613680\" target=\"_blank\">here</a>. (You will need to unhide this code cell 6). When reading the files, use <code>dataset.map(process_path, num_parallel_calls=AUTO)</code> with <code>num_parallel_calls=AUTO</code> where <code>AUTO = tf.data.experimental.AUTOTUNE</code>. This enables I/O multiprocessing.</li>\n<li><code>tensorflow.keras.preprocessing.image.ImageDataGenerator</code>. There is an example in code cell 2 <a href=\"https://www.kaggle.com/harveenchadha/efficientnetb3-baseline-inference-keras-tf2-tta?scriptVersionId=50249323\" target=\"_blank\">here</a>. Once again when you predict, use <code>my_model.predict(test_gen,  verbose = True, use_multiprocessing=True, workers=4)</code></li>\n</ul>\n<p>To find more examples, view public notebooks that are inference notebooks for tensorflow models.</p>",
      "rawMarkdown": "Hi, ignore my comment below. Using `@tf.function()` will only help if the bottleneck is your model. \n\nYour bottleneck is reading images from disk, and starving your GPU by inferring only 1 image at a time. To increase speed, you will want to read multiple images from the disk in parallel and feed your GPU larger batches. (GPUs work faster when they have lots of work to do)(and disk I/O is faster with multiprocessing).\n\nYou will need to create a dataloader to feed your model batches of images during inference, and enable I/O multiprocessing to read from disk. Below are 3 choices with code examples\n* `tf.keras.utils.Sequence` this is my favorite dataloader. There is an example in code cell 33 [here][1]. It is the easiest to add augmentation to. When you call predict `model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)` enable multiprocessing and workers \n* `tf.data.Dataset`. This can be used with images (does not need TFRecords). There is an example in code cell 6 [here][2]. (You will need to unhide this code cell 6). When reading the files, use `dataset.map(process_path, num_parallel_calls=AUTO)` with `num_parallel_calls=AUTO` where `AUTO = tf.data.experimental.AUTOTUNE`. This enables I/O multiprocessing.\n* `tensorflow.keras.preprocessing.image.ImageDataGenerator`. There is an example in code cell 2 [here][3]. Once again when you predict, use `my_model.predict(test_gen,  verbose = True, use_multiprocessing=True, workers=4)`\n\nTo find more examples, view public notebooks that are inference notebooks for tensorflow models.\n\n[1]: https://www.kaggle.com/awsaf49/efficientnetb6-512-cutmixupdropout-tpu-infer/notebook?scriptVersionId=47469988\n[2]: https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference?scriptVersionId=49613680\n[3]: https://www.kaggle.com/harveenchadha/efficientnetb3-baseline-inference-keras-tf2-tta?scriptVersionId=50249323",
      "votes": null
    },
    {
      "id": "1148492",
      "postDate": "01/11/2021 07:12:45",
      "content": "<p>Welcome back. 😃</p>",
      "rawMarkdown": "Welcome back. 😃",
      "votes": null
    },
    {
      "id": "1149382",
      "postDate": "01/11/2021 19:55:15",
      "content": "<p>To use flow_from_directory, your images have to be stored in subdirectories. Each subdirectory contains all the images of a given class, and the name of the subdirectory is the class label. Thus it is not very convenient for the test dataset ; it is easier with flow_from_dataframe. You need to have a dataframe such as train.csv, with one column containing the images ids, and one column giving the class labels. For the test dataset, you can create the dataframe yourself, with a dummy column for the class labels (setting every value to 0 for example).</p>",
      "rawMarkdown": "To use flow_from_directory, your images have to be stored in subdirectories. Each subdirectory contains all the images of a given class, and the name of the subdirectory is the class label. Thus it is not very convenient for the test dataset ; it is easier with flow_from_dataframe. You need to have a dataframe such as train.csv, with one column containing the images ids, and one column giving the class labels. For the test dataset, you can create the dataframe yourself, with a dummy column for the class labels (setting every value to 0 for example).",
      "votes": null
    },
    {
      "id": "1149475",
      "postDate": "01/11/2021 21:25:56",
      "content": "<p>Thanks Bruce. Good to see you here.</p>",
      "rawMarkdown": "Thanks Bruce. Good to see you here.",
      "votes": null
    },
    {
      "id": "1149687",
      "postDate": "01/12/2021 04:20:18",
      "content": "<p>You want to use <code>tf.data.Dataset</code> like<br>\n<code>ds = tf.data.Dataset.from_tensor_slices(df['path']).map(tf.io.read_file).map(tf.io.decode_image)</code>.<br>\nAlso, in machine learning and big data world, thing should always be vectorized. Doing a python <code>for</code> loop to iterate through each data is not good.</p>",
      "rawMarkdown": "You want to use `tf.data.Dataset` like\n`ds = tf.data.Dataset.from_tensor_slices(df['path']).map(tf.io.read_file).map(tf.io.decode_image)`.\nAlso, in machine learning and big data world, thing should always be vectorized. Doing a python `for` loop to iterate through each data is not good.",
      "votes": null
    },
    {
      "id": "1150185",
      "postDate": "01/12/2021 12:49:13",
      "content": "<p>Nice explanation 🙂 , I will definitely try them out and get back to you.</p>",
      "rawMarkdown": "Nice explanation 🙂 , I will definitely try them out and get back to you.",
      "votes": null
    },
    {
      "id": "1150190",
      "postDate": "01/12/2021 12:55:03",
      "content": "<p>Hey, thanks, I was hoping for a tf.Dataset explanation.  I'll try this and get back to you. Also Check out Chris Deotte's explanation.</p>",
      "rawMarkdown": "Hey, thanks, I was hoping for a tf.Dataset explanation.  I'll try this and get back to you. Also Check out Chris Deotte's explanation.",
      "votes": null
    },
    {
      "id": "1150199",
      "postDate": "01/12/2021 12:59:31",
      "content": "<p>🙂 Nice Hack with the dataframe. I will definitely try this and get back to you.</p>",
      "rawMarkdown": "🙂 Nice Hack with the dataframe. I will definitely try this and get back to you.",
      "votes": null
    },
    {
      "id": "1152822",
      "postDate": "01/14/2021 13:24:00",
      "content": "<p>Hey 🙂, I tried out <code>tf.keras.utils.Sequence</code> code. I saw some reasons it's your favourite, like this bar during prediction <code>21397/21397 [==============================] - 762s 36ms/step</code> and how the code is very abstracted.  A few follow up questions : </p>\n<ol>\n<li><code>workers = 4</code> have you chosen 4 because of the number of cores and is there some interesting theory behind this or is it a random number? Also does it get affected when I use a CPU instead of a GPU.</li>\n<li>How do you learn about these more complex Tensorflow methods and techniques, do you just read the documentation or a book or online courses? Any information you share would be very useful for learning purposes.</li>\n</ol>",
      "rawMarkdown": "Hey 🙂, I tried out `tf.keras.utils.Sequence` code. I saw some reasons it's your favourite, like this bar during prediction `21397/21397 [==============================] - 762s 36ms/step` and how the code is very abstracted.  A few follow up questions : \n\n1. `workers = 4` have you chosen 4 because of the number of cores and is there some interesting theory behind this or is it a random number? Also does it get affected when I use a CPU instead of a GPU.\n2. How do you learn about these more complex Tensorflow methods and techniques, do you just read the documentation or a book or online courses? Any information you share would be very useful for learning purposes.",
      "votes": null
    },
    {
      "id": "1153227",
      "postDate": "01/14/2021 17:41:40",
      "content": "<p>Question ? 🤔<br>\nWhen I predict using <code>tf.keras.utils.Sequence</code> on train images  (just for model testing purposes) I get all predictions as 4 ie <code>21397/21397 [==============================] - 764s 36ms/step\n[4 4 4 ... 4 4 4]\nPrediction time :  762.911825714</code></p>\n<p>When I run a for loop like this <br>\n<code>import time</code><br>\n<code>images_test = []</code><br>\n<code>preds_try = []</code></p>\n<h6>#### load all images</h6>\n<p><code>start_load = time.process_time()</code><br>\n<code>train_images = os.listdir('../input/cassava-leaf-disease-classification/'+'train_images/')</code><br>\n<code>for img in train_images:</code><br>\n<code>img = Image.open('../input/cassava-leaf-disease-classification/'+'train_images/' + img)</code><br>\n<code>img = img.resize(size)</code><br>\n<code>img = np.expand_dims(img, axis=0)</code><br>\n<code>preds_try.extend(model.predict(img).argmax(axis = 1))</code><br>\n<code>print('Prediction time : ',time.process_time() - start_load)</code></p>\n<p>I get the desired output. <code>[4 3 1... 1 2 3]</code> but takes<code>Prediction time :  1675.7722212400004 sec</code> .  What could be the cause of this error . Based on the second method the issue is not the model, it is the generator. De😷 (debugging) challenge anyone? 🤓</p>",
      "rawMarkdown": "Question ? 🤔\nWhen I predict using `tf.keras.utils.Sequence` on train images  (just for model testing purposes) I get all predictions as 4 ie `21397/21397 [==============================] - 764s 36ms/step\n[4 4 4 ... 4 4 4]\nPrediction time :  762.911825714`\n\nWhen I run a for loop like this \n`import time`\n`images_test = []`\n`preds_try = []`\n########## load all images\n`start_load = time.process_time()`\n`train_images = os.listdir('../input/cassava-leaf-disease-classification/'+'train_images/')`\n`for img in train_images:`\n`    img = Image.open('../input/cassava-leaf-disease-classification/'+'train_images/' + img)`\n`    img = img.resize(size)`\n`    img = np.expand_dims(img, axis=0)`\n`    preds_try.extend(model.predict(img).argmax(axis = 1))`\n`print('Prediction time : ',time.process_time() - start_load)`\n \nI get the desired output. `[4 3 1... 1 2 3]` but takes`Prediction time :  1675.7722212400004 sec` .  What could be the cause of this error . Based on the second method the issue is not the model, it is the generator. De😷 (debugging) challenge anyone? 🤓",
      "votes": null
    },
    {
      "id": "1153268",
      "postDate": "01/14/2021 18:19:15",
      "content": "<blockquote>\n  <p>workers = 4 have you chosen 4 because of the number of cores and is there some interesting theory behind this or is it a random number?</p>\n</blockquote>\n<p>Using <code>workers = 4</code> may not be the best. You should try 1, 2, 4, 8 and see what is fastest. When you enable GPU at Kaggle, the associated CPU only has 2 cores, so for parallel computation perhaps using <code>workers = 2</code> is best but for disk I/O (which is what we need here) another number may be best. Try different numbers and see. (Number of workers can be anything because 1 core and have an unlimited number of threads, but different numbers have different performance).</p>\n<blockquote>\n  <p>How do you learn about these more complex Tensorflow methods and techniques</p>\n</blockquote>\n<p>Everything i've learned about TF is from reading other Kaggler's code and Google searching for how to do things. And then writing code and watching how it performs.</p>\n<p>When i use TF, I usually use Keras with is a high level framework for making models. You can also write TF code directly using tensors, graphs, gradient tape, and <code>@tf.function</code>. This allows you to make everything faster but it is harder to code. (I've only done this a few times when i need very custom stuff or maximum speed).</p>",
      "rawMarkdown": "> workers = 4 have you chosen 4 because of the number of cores and is there some interesting theory behind this or is it a random number?\n\nUsing `workers = 4` may not be the best. You should try 1, 2, 4, 8 and see what is fastest. When you enable GPU at Kaggle, the associated CPU only has 2 cores, so for parallel computation perhaps using `workers = 2` is best but for disk I/O (which is what we need here) another number may be best. Try different numbers and see. (Number of workers can be anything because 1 core and have an unlimited number of threads, but different numbers have different performance).\n\n> How do you learn about these more complex Tensorflow methods and techniques\n\nEverything i've learned about TF is from reading other Kaggler's code and Google searching for how to do things. And then writing code and watching how it performs.\n\nWhen i use TF, I usually use Keras with is a high level framework for making models. You can also write TF code directly using tensors, graphs, gradient tape, and `@tf.function`. This allows you to make everything faster but it is harder to code. (I've only done this a few times when i need very custom stuff or maximum speed).",
      "votes": null
    },
    {
      "id": "1153275",
      "postDate": "01/14/2021 18:25:05",
      "content": "<p>Great experiments. I'm looking forward to seeing the speed differences comparing <code>tf.keras.utils.Sequence</code>, <code>tf.data.Dataset</code>, and <code>tensorflow.keras.preprocessing.image.ImageDataGenerator</code>. Personally, i don't which is faster but i would like to know.</p>\n<p>Also try varying <code>workers=4</code>. Try <code>1, 2, 4, 8, 16</code> and find which is fastest for <code>tf.keras.utils.Sequence</code> and <code>tensorflow.keras.preprocessing.image.ImageDataGenerator</code></p>",
      "rawMarkdown": "Great experiments. I'm looking forward to seeing the speed differences comparing `tf.keras.utils.Sequence`, `tf.data.Dataset`, and `tensorflow.keras.preprocessing.image.ImageDataGenerator`. Personally, i don't which is faster but i would like to know.\n\nAlso try varying `workers=4`. Try `1, 2, 4, 8, 16` and find which is fastest for `tf.keras.utils.Sequence` and `tensorflow.keras.preprocessing.image.ImageDataGenerator`",
      "votes": null
    },
    {
      "id": "1157448",
      "postDate": "01/17/2021 22:53:30",
      "content": "<p>I can't keep calm, my mentor just mentioned my kernel. One of the happiest days of my life!</p>",
      "rawMarkdown": "I can't keep calm, my mentor just mentioned my kernel. One of the happiest days of my life!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1147827,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/10/2021 18:22:21",
      "content": "<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207280\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207280</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1148002,
          "author_name": "hypnotu",
          "author_url": "",
          "post_date": "01/10/2021 20:56:08",
          "content": "<p>Thank you for this. Might you take a minute to explain this code please. <br>\n<code>@tf.function(input_signature=[</code><br>\n<code>tf.TensorSpec(shape=None, dtype=tf.int64),</code><br>\n<code>tf.TensorSpec(shape=None, dtype=tf.int32),</code><br>\n<code>tf.TensorSpec(shape=[], dtype=tf.bool),</code><br>\n<code>]</code><br>\n<code>)</code><br>\n <code>def call(self, user_id, content_id, training=False):</code><br>\n<code>N = len(user_ix)</code><br>\n<code>user_id = tf.reshape(user_id, (-1,))</code><br>\n<code>content_id = tf.reshape(content_id, (-1,))</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1147906,
      "author_name": "florian12",
      "author_url": "",
      "post_date": "01/10/2021 19:49:05",
      "content": "<p>If you're interested, for the line : <br>\nimg = np.expand_dims(img, axis=0) #add extra dimension don't know why yet.<br>\nan extra dimension is added because in TensorFlow, batches of images are 4 dimensional arrays of size [batch_size, height, width, channels]. Even if your batch contains only one image, it has to be a 4d array, with batch_size = 1.</p>\n<p>Maybe using ImageDataGenerator and flow_from_directory or flow_from_dataframe would be faster, as it loads batches of images, which might be faster than loading each image one by one (if the source code is more efficient than pure Python code, but I haven't checked). At least, if you use ImageDataGenerator for training, it is very convenient (in my opinion) to use it for prediction too.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147992,
          "author_name": "hypnotu",
          "author_url": "",
          "post_date": "01/10/2021 20:52:28",
          "content": "<p>Hey and  Thank you for your reply.<br>\nI have tried using ImageDataGenerator and flow_from_directory, I have not gotten the hang of it yet and Errors! 😪 Also I have not found a way to manipulate the data in that state. Still looking at code snippets to understand this part better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149382,
          "author_name": "florian12",
          "author_url": "",
          "post_date": "01/11/2021 19:55:15",
          "content": "<p>To use flow_from_directory, your images have to be stored in subdirectories. Each subdirectory contains all the images of a given class, and the name of the subdirectory is the class label. Thus it is not very convenient for the test dataset ; it is easier with flow_from_dataframe. You need to have a dataframe such as train.csv, with one column containing the images ids, and one column giving the class labels. For the test dataset, you can create the dataframe yourself, with a dummy column for the class labels (setting every value to 0 for example).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150199,
          "author_name": "hypnotu",
          "author_url": "",
          "post_date": "01/12/2021 12:59:31",
          "content": "<p>🙂 Nice Hack with the dataframe. I will definitely try this and get back to you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148214,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/11/2021 02:24:46",
      "content": "<p>Hi, ignore my comment below. Using <code>@tf.function()</code> will only help if the bottleneck is your model. </p>\n<p>Your bottleneck is reading images from disk, and starving your GPU by inferring only 1 image at a time. To increase speed, you will want to read multiple images from the disk in parallel and feed your GPU larger batches. (GPUs work faster when they have lots of work to do)(and disk I/O is faster with multiprocessing).</p>\n<p>You will need to create a dataloader to feed your model batches of images during inference, and enable I/O multiprocessing to read from disk. Below are 3 choices with code examples</p>\n<ul>\n<li><code>tf.keras.utils.Sequence</code> this is my favorite dataloader. There is an example in code cell 33 <a href=\"https://www.kaggle.com/awsaf49/efficientnetb6-512-cutmixupdropout-tpu-infer/notebook?scriptVersionId=47469988\" target=\"_blank\">here</a>. It is the easiest to add augmentation to. When you call predict <code>model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)</code> enable multiprocessing and workers </li>\n<li><code>tf.data.Dataset</code>. This can be used with images (does not need TFRecords). There is an example in code cell 6 <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference?scriptVersionId=49613680\" target=\"_blank\">here</a>. (You will need to unhide this code cell 6). When reading the files, use <code>dataset.map(process_path, num_parallel_calls=AUTO)</code> with <code>num_parallel_calls=AUTO</code> where <code>AUTO = tf.data.experimental.AUTOTUNE</code>. This enables I/O multiprocessing.</li>\n<li><code>tensorflow.keras.preprocessing.image.ImageDataGenerator</code>. There is an example in code cell 2 <a href=\"https://www.kaggle.com/harveenchadha/efficientnetb3-baseline-inference-keras-tf2-tta?scriptVersionId=50249323\" target=\"_blank\">here</a>. Once again when you predict, use <code>my_model.predict(test_gen,  verbose = True, use_multiprocessing=True, workers=4)</code></li>\n</ul>\n<p>To find more examples, view public notebooks that are inference notebooks for tensorflow models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1148492,
          "author_name": "mutantspore",
          "author_url": "",
          "post_date": "01/11/2021 07:12:45",
          "content": "<p>Welcome back. 😃</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149475,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/11/2021 21:25:56",
          "content": "<p>Thanks Bruce. Good to see you here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150185,
          "author_name": "hypnotu",
          "author_url": "",
          "post_date": "01/12/2021 12:49:13",
          "content": "<p>Nice explanation 🙂 , I will definitely try them out and get back to you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1152822,
          "author_name": "hypnotu",
          "author_url": "",
          "post_date": "01/14/2021 13:24:00",
          "content": "<p>Hey 🙂, I tried out <code>tf.keras.utils.Sequence</code> code. I saw some reasons it's your favourite, like this bar during prediction <code>21397/21397 [==============================] - 762s 36ms/step</code> and how the code is very abstracted.  A few follow up questions : </p>\n<ol>\n<li><code>workers = 4</code> have you chosen 4 because of the number of cores and is there some interesting theory behind this or is it a random number? Also does it get affected when I use a CPU instead of a GPU.</li>\n<li>How do you learn about these more complex Tensorflow methods and techniques, do you just read the documentation or a book or online courses? Any information you share would be very useful for learning purposes.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1153268,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "01/14/2021 18:19:15",
          "content": "<blockquote>\n  <p>workers = 4 have you chosen 4 because of the number of cores and is there some interesting theory behind this or is it a random number?</p>\n</blockquote>\n<p>Using <code>workers = 4</code> may not be the best. You should try 1, 2, 4, 8 and see what is fastest. When you enable GPU at Kaggle, the associated CPU only has 2 cores, so for parallel computation perhaps using <code>workers = 2</code> is best but for disk I/O (which is what we need here) another number may be best. Try different numbers and see. (Number of workers can be anything because 1 core and have an unlimited number of threads, but different numbers have different performance).</p>\n<blockquote>\n  <p>How do you learn about these more complex Tensorflow methods and techniques</p>\n</blockquote>\n<p>Everything i've learned about TF is from reading other Kaggler's code and Google searching for how to do things. And then writing code and watching how it performs.</p>\n<p>When i use TF, I usually use Keras with is a high level framework for making models. You can also write TF code directly using tensors, graphs, gradient tape, and <code>@tf.function</code>. This allows you to make everything faster but it is harder to code. (I've only done this a few times when i need very custom stuff or maximum speed).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1157448,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "01/17/2021 22:53:30",
          "content": "<p>I can't keep calm, my mentor just mentioned my kernel. One of the happiest days of my life!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1149687,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "01/12/2021 04:20:18",
      "content": "<p>You want to use <code>tf.data.Dataset</code> like<br>\n<code>ds = tf.data.Dataset.from_tensor_slices(df['path']).map(tf.io.read_file).map(tf.io.decode_image)</code>.<br>\nAlso, in machine learning and big data world, thing should always be vectorized. Doing a python <code>for</code> loop to iterate through each data is not good.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1150190,
          "author_name": "hypnotu",
          "author_url": "",
          "post_date": "01/12/2021 12:55:03",
          "content": "<p>Hey, thanks, I was hoping for a tf.Dataset explanation.  I'll try this and get back to you. Also Check out Chris Deotte's explanation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1153227,
      "author_name": "hypnotu",
      "author_url": "",
      "post_date": "01/14/2021 17:41:40",
      "content": "<p>Question ? 🤔<br>\nWhen I predict using <code>tf.keras.utils.Sequence</code> on train images  (just for model testing purposes) I get all predictions as 4 ie <code>21397/21397 [==============================] - 764s 36ms/step\n[4 4 4 ... 4 4 4]\nPrediction time :  762.911825714</code></p>\n<p>When I run a for loop like this <br>\n<code>import time</code><br>\n<code>images_test = []</code><br>\n<code>preds_try = []</code></p>\n<h6>#### load all images</h6>\n<p><code>start_load = time.process_time()</code><br>\n<code>train_images = os.listdir('../input/cassava-leaf-disease-classification/'+'train_images/')</code><br>\n<code>for img in train_images:</code><br>\n<code>img = Image.open('../input/cassava-leaf-disease-classification/'+'train_images/' + img)</code><br>\n<code>img = img.resize(size)</code><br>\n<code>img = np.expand_dims(img, axis=0)</code><br>\n<code>preds_try.extend(model.predict(img).argmax(axis = 1))</code><br>\n<code>print('Prediction time : ',time.process_time() - start_load)</code></p>\n<p>I get the desired output. <code>[4 3 1... 1 2 3]</code> but takes<code>Prediction time :  1675.7722212400004 sec</code> .  What could be the cause of this error . Based on the second method the issue is not the model, it is the generator. De😷 (debugging) challenge anyone? 🤓</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1153275,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/14/2021 18:25:05",
      "content": "<p>Great experiments. I'm looking forward to seeing the speed differences comparing <code>tf.keras.utils.Sequence</code>, <code>tf.data.Dataset</code>, and <code>tensorflow.keras.preprocessing.image.ImageDataGenerator</code>. Personally, i don't which is faster but i would like to know.</p>\n<p>Also try varying <code>workers=4</code>. Try <code>1, 2, 4, 8, 16</code> and find which is fastest for <code>tf.keras.utils.Sequence</code> and <code>tensorflow.keras.preprocessing.image.ImageDataGenerator</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1147171": "list_of_preds_per_class = model.predict(array_of_images)\nWill update this for better clarity.\n[How to do predictions with tensorflow](https://youtu.be/RqLD1INA_cQ)\n`### libraries`\n`import tensorflow as tf #load model #load image `\n`from PIL import Image `\n\n`test_dir = '../input/cassava-leaf-disease-classification/test_images/'`\n\n\n###  takes 0.08 seconds per image.\n(This was tested on 3500 train images took 281 seconds)\n#### The code\n`import time`\n`start = time.process_time()`\n`  `\n`preds = []`\n`for img in test_images:`\n`    img = Image.open(test_dir + img)`\n`    img = np.expand_dims(img, axis=0) #add extra dimension don't know why yet.`\n`    preds.extend(model.predict(img).argmax(axis = 1)) #predict list of probabilities per class then takes max prediction class as prediction. `\n\n`print('Time taken in seconds : ',time.process_time() - start)`\n\n### takes 0.073 seconds per image \n(This was tested on 3500 train images took 256 seconds)\n#### The code\n`import time`\n`start = time.process_time()`\n\n`preds = []`\n`for img in test_images:`\n`    img = tf.keras.preprocessing.image.load_img(test_dir + img) #loads image `\n`    img = np.expand_dims(img, axis=0) #add extra dimension dont know why yet.`\n`    preds.extend(model.predict(img).argmax(axis = 1)) #predict list of probabilities per class then takes max prediction class as prediction. `\n\n`print('Time taken in seconds : ',time.process_time() - start)`\n\n\n\n\n### Any faster Ideas?\n\nSome very useful comments below on using Generators.💯 💯\nAlso someone emailed me this link:\n[keras-how-to-generate-data-on-the-fly-Stanford](https://stanford.edu/~shervine/blog/keras-how-to-generate-data-on-the-fly)\n\n\n##### Faster Idea 1 (from comments)\n### 1. Using `tf.keras.utils.Sequence `,\n#### code from code cell 33 [here as mentioned by Chris Deotte in the comments and coded by awsaf49](https://www.kaggle.com/awsaf49/efficientnetb6-512-cutmixupdropout-tpu-infer/notebook?scriptVersionId=47469988)  and using `model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)`\n## 0.0356 seconds per image.\n 1.8 seconds on a single image and 761 seconds on 21397 images\n\n### The code.\n[Link to Code , it a bit long that why its on another discussion](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/211269)\n\n\n#### 2. [Up Next] Using tf.data.Dataset",
    "1147827": "https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207280",
    "1147906": "If you're interested, for the line : \nimg = np.expand_dims(img, axis=0) #add extra dimension don't know why yet.\nan extra dimension is added because in TensorFlow, batches of images are 4 dimensional arrays of size [batch_size, height, width, channels]. Even if your batch contains only one image, it has to be a 4d array, with batch_size = 1.\n\nMaybe using ImageDataGenerator and flow_from_directory or flow_from_dataframe would be faster, as it loads batches of images, which might be faster than loading each image one by one (if the source code is more efficient than pure Python code, but I haven't checked). At least, if you use ImageDataGenerator for training, it is very convenient (in my opinion) to use it for prediction too.",
    "1147992": "Hey and  Thank you for your reply.\nI have tried using ImageDataGenerator and flow_from_directory, I have not gotten the hang of it yet and Errors! 😪 Also I have not found a way to manipulate the data in that state. Still looking at code snippets to understand this part better.",
    "1148002": "Thank you for this. Might you take a minute to explain this code please. \n`@tf.function(input_signature=[`\n`        tf.TensorSpec(shape=None, dtype=tf.int64),`\n`        tf.TensorSpec(shape=None, dtype=tf.int32),`\n`        tf.TensorSpec(shape=[], dtype=tf.bool),`\n`    ]`\n`    )`\n `   def call(self, user_id, content_id, training=False):`\n`        N = len(user_ix)`\n`        user_id = tf.reshape(user_id, (-1,))`\n`        content_id = tf.reshape(content_id, (-1,))`",
    "1148214": "Hi, ignore my comment below. Using `@tf.function()` will only help if the bottleneck is your model. \n\nYour bottleneck is reading images from disk, and starving your GPU by inferring only 1 image at a time. To increase speed, you will want to read multiple images from the disk in parallel and feed your GPU larger batches. (GPUs work faster when they have lots of work to do)(and disk I/O is faster with multiprocessing).\n\nYou will need to create a dataloader to feed your model batches of images during inference, and enable I/O multiprocessing to read from disk. Below are 3 choices with code examples\n* `tf.keras.utils.Sequence` this is my favorite dataloader. There is an example in code cell 33 [here][1]. It is the easiest to add augmentation to. When you call predict `model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)` enable multiprocessing and workers \n* `tf.data.Dataset`. This can be used with images (does not need TFRecords). There is an example in code cell 6 [here][2]. (You will need to unhide this code cell 6). When reading the files, use `dataset.map(process_path, num_parallel_calls=AUTO)` with `num_parallel_calls=AUTO` where `AUTO = tf.data.experimental.AUTOTUNE`. This enables I/O multiprocessing.\n* `tensorflow.keras.preprocessing.image.ImageDataGenerator`. There is an example in code cell 2 [here][3]. Once again when you predict, use `my_model.predict(test_gen,  verbose = True, use_multiprocessing=True, workers=4)`\n\nTo find more examples, view public notebooks that are inference notebooks for tensorflow models.\n\n[1]: https://www.kaggle.com/awsaf49/efficientnetb6-512-cutmixupdropout-tpu-infer/notebook?scriptVersionId=47469988\n[2]: https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tpu-v2-pods-inference?scriptVersionId=49613680\n[3]: https://www.kaggle.com/harveenchadha/efficientnetb3-baseline-inference-keras-tf2-tta?scriptVersionId=50249323",
    "1148492": "Welcome back. 😃",
    "1149382": "To use flow_from_directory, your images have to be stored in subdirectories. Each subdirectory contains all the images of a given class, and the name of the subdirectory is the class label. Thus it is not very convenient for the test dataset ; it is easier with flow_from_dataframe. You need to have a dataframe such as train.csv, with one column containing the images ids, and one column giving the class labels. For the test dataset, you can create the dataframe yourself, with a dummy column for the class labels (setting every value to 0 for example).",
    "1149475": "Thanks Bruce. Good to see you here.",
    "1149687": "You want to use `tf.data.Dataset` like\n`ds = tf.data.Dataset.from_tensor_slices(df['path']).map(tf.io.read_file).map(tf.io.decode_image)`.\nAlso, in machine learning and big data world, thing should always be vectorized. Doing a python `for` loop to iterate through each data is not good.",
    "1150185": "Nice explanation 🙂 , I will definitely try them out and get back to you.",
    "1150190": "Hey, thanks, I was hoping for a tf.Dataset explanation.  I'll try this and get back to you. Also Check out Chris Deotte's explanation.",
    "1150199": "🙂 Nice Hack with the dataframe. I will definitely try this and get back to you.",
    "1152822": "Hey 🙂, I tried out `tf.keras.utils.Sequence` code. I saw some reasons it's your favourite, like this bar during prediction `21397/21397 [==============================] - 762s 36ms/step` and how the code is very abstracted.  A few follow up questions : \n\n1. `workers = 4` have you chosen 4 because of the number of cores and is there some interesting theory behind this or is it a random number? Also does it get affected when I use a CPU instead of a GPU.\n2. How do you learn about these more complex Tensorflow methods and techniques, do you just read the documentation or a book or online courses? Any information you share would be very useful for learning purposes.",
    "1153227": "Question ? 🤔\nWhen I predict using `tf.keras.utils.Sequence` on train images  (just for model testing purposes) I get all predictions as 4 ie `21397/21397 [==============================] - 764s 36ms/step\n[4 4 4 ... 4 4 4]\nPrediction time :  762.911825714`\n\nWhen I run a for loop like this \n`import time`\n`images_test = []`\n`preds_try = []`\n########## load all images\n`start_load = time.process_time()`\n`train_images = os.listdir('../input/cassava-leaf-disease-classification/'+'train_images/')`\n`for img in train_images:`\n`    img = Image.open('../input/cassava-leaf-disease-classification/'+'train_images/' + img)`\n`    img = img.resize(size)`\n`    img = np.expand_dims(img, axis=0)`\n`    preds_try.extend(model.predict(img).argmax(axis = 1))`\n`print('Prediction time : ',time.process_time() - start_load)`\n \nI get the desired output. `[4 3 1... 1 2 3]` but takes`Prediction time :  1675.7722212400004 sec` .  What could be the cause of this error . Based on the second method the issue is not the model, it is the generator. De😷 (debugging) challenge anyone? 🤓",
    "1153268": "> workers = 4 have you chosen 4 because of the number of cores and is there some interesting theory behind this or is it a random number?\n\nUsing `workers = 4` may not be the best. You should try 1, 2, 4, 8 and see what is fastest. When you enable GPU at Kaggle, the associated CPU only has 2 cores, so for parallel computation perhaps using `workers = 2` is best but for disk I/O (which is what we need here) another number may be best. Try different numbers and see. (Number of workers can be anything because 1 core and have an unlimited number of threads, but different numbers have different performance).\n\n> How do you learn about these more complex Tensorflow methods and techniques\n\nEverything i've learned about TF is from reading other Kaggler's code and Google searching for how to do things. And then writing code and watching how it performs.\n\nWhen i use TF, I usually use Keras with is a high level framework for making models. You can also write TF code directly using tensors, graphs, gradient tape, and `@tf.function`. This allows you to make everything faster but it is harder to code. (I've only done this a few times when i need very custom stuff or maximum speed).",
    "1153275": "Great experiments. I'm looking forward to seeing the speed differences comparing `tf.keras.utils.Sequence`, `tf.data.Dataset`, and `tensorflow.keras.preprocessing.image.ImageDataGenerator`. Personally, i don't which is faster but i would like to know.\n\nAlso try varying `workers=4`. Try `1, 2, 4, 8, 16` and find which is fastest for `tf.keras.utils.Sequence` and `tensorflow.keras.preprocessing.image.ImageDataGenerator`",
    "1157448": "I can't keep calm, my mentor just mentioned my kernel. One of the happiest days of my life!"
  },
  "source": "meta"
}