{
  "id": 225017,
  "title": "Speeding up keras predictions via tfrecords (or otherwise)",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/225017",
  "author_name": "",
  "post_date": "2021-03-10T14:55:41.505200900Z",
  "votes": 8,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I've got a pretty large model I trained in <code>keras</code> and inference time is an issue. I wondered about using the tfrecords to speed up predictions and used these options (as described <a href=\"https://keras.io/examples/keras_recipes/tfrecord/\" target=\"_blank\">here in the keras documentation</a> and I also looked at this <a href=\"https://www.kaggle.com/tpothjuan/efficientnetb7-tfrecords\" target=\"_blank\">notebook</a>) for inspiration:</p>\n<pre><code>def load_dataset(filenames):\n    ignore_order = tf.data.Options()\n    ignore_order.experimental_deterministic = False  # disable order, increase speed, keep order ensure ordering\n    dataset = tf.data.TFRecordDataset(filenames)  # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order)  # uses data as soon as it streams in, rather than in its original order\n    dataset = dataset.map(partial(read_tfrecord), num_parallel_calls=AUTOTUNE)\n    # returns a dataset of just images (test data are not labelled)\n    return dataset\n</code></pre>\n<p>and then I did <code>preds = model.predict(test_dataset, verbose = 1, workers=2, use_multiprocessing=True)</code>. However, while my <code>read_tfrecord</code> function does return a tuple of image and image_id, and iterating on the dataset I create via <code>test_dataset = get_dataset(TEST_FILENAMES)</code> does indeed return both, the <code>model.predict</code> function does not return these. I.e. there seems to be no obvious way to figure out what prediction corresponds to which record. Clearly, I'm overlooking something here, since this seems to follow a simple example from the documentation and nobody would want predictions without knowing what records they are for.</p>\n<p>When I instead used</p>\n<pre><code>preds = []\nfor batch in test_dataset:\n        imgs, StudyInstanceUIDs = batch \n        yhat = model.predict(imgs)\n        preds.append( (yhat, StudyInstanceUIDs) )\n</code></pre>\n<p>that achieves what I want, but the version with <code>preds = model.predict(test_dataset, verbose = 1, workers=2, use_multiprocessing=True)</code> is about twice faster (or is that just all from the multi-processing??). It's 8 min 32 s vs. 3 min 29 s on the public LB set - thus, probably 34 vs. 14 min for the full LB dataset. With a model for 5 folds (never mind using TTA), that turns into 2h 50 min vs. 1h 10 min. </p>\n<p>So, once I want to use multiple models, this is a real obstacle. Does anyone have any ideas or insights on what to do about this and to speed up the inference?</p>",
  "messages": [
    {
      "id": "1233607",
      "postDate": "03/10/2021 14:55:41",
      "content": "<p>I've got a pretty large model I trained in <code>keras</code> and inference time is an issue. I wondered about using the tfrecords to speed up predictions and used these options (as described <a href=\"https://keras.io/examples/keras_recipes/tfrecord/\" target=\"_blank\">here in the keras documentation</a> and I also looked at this <a href=\"https://www.kaggle.com/tpothjuan/efficientnetb7-tfrecords\" target=\"_blank\">notebook</a>) for inspiration:</p>\n<pre><code>def load_dataset(filenames):\n    ignore_order = tf.data.Options()\n    ignore_order.experimental_deterministic = False  # disable order, increase speed, keep order ensure ordering\n    dataset = tf.data.TFRecordDataset(filenames)  # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order)  # uses data as soon as it streams in, rather than in its original order\n    dataset = dataset.map(partial(read_tfrecord), num_parallel_calls=AUTOTUNE)\n    # returns a dataset of just images (test data are not labelled)\n    return dataset\n</code></pre>\n<p>and then I did <code>preds = model.predict(test_dataset, verbose = 1, workers=2, use_multiprocessing=True)</code>. However, while my <code>read_tfrecord</code> function does return a tuple of image and image_id, and iterating on the dataset I create via <code>test_dataset = get_dataset(TEST_FILENAMES)</code> does indeed return both, the <code>model.predict</code> function does not return these. I.e. there seems to be no obvious way to figure out what prediction corresponds to which record. Clearly, I'm overlooking something here, since this seems to follow a simple example from the documentation and nobody would want predictions without knowing what records they are for.</p>\n<p>When I instead used</p>\n<pre><code>preds = []\nfor batch in test_dataset:\n        imgs, StudyInstanceUIDs = batch \n        yhat = model.predict(imgs)\n        preds.append( (yhat, StudyInstanceUIDs) )\n</code></pre>\n<p>that achieves what I want, but the version with <code>preds = model.predict(test_dataset, verbose = 1, workers=2, use_multiprocessing=True)</code> is about twice faster (or is that just all from the multi-processing??). It's 8 min 32 s vs. 3 min 29 s on the public LB set - thus, probably 34 vs. 14 min for the full LB dataset. With a model for 5 folds (never mind using TTA), that turns into 2h 50 min vs. 1h 10 min. </p>\n<p>So, once I want to use multiple models, this is a real obstacle. Does anyone have any ideas or insights on what to do about this and to speed up the inference?</p>",
      "rawMarkdown": "I've got a pretty large model I trained in `keras` and inference time is an issue. I wondered about using the tfrecords to speed up predictions and used these options (as described [here in the keras documentation](https://keras.io/examples/keras_recipes/tfrecord/) and I also looked at this [notebook]( https://www.kaggle.com/tpothjuan/efficientnetb7-tfrecords)) for inspiration:\n```\ndef load_dataset(filenames):\n    ignore_order = tf.data.Options()\n    ignore_order.experimental_deterministic = False  # disable order, increase speed, keep order ensure ordering\n    dataset = tf.data.TFRecordDataset(filenames)  # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order)  # uses data as soon as it streams in, rather than in its original order\n    dataset = dataset.map(partial(read_tfrecord), num_parallel_calls=AUTOTUNE)\n    # returns a dataset of just images (test data are not labelled)\n    return dataset\n```\nand then I did `preds = model.predict(test_dataset, verbose = 1, workers=2, use_multiprocessing=True)`. However, while my `read_tfrecord` function does return a tuple of image and image_id, and iterating on the dataset I create via `test_dataset = get_dataset(TEST_FILENAMES)` does indeed return both, the `model.predict` function does not return these. I.e. there seems to be no obvious way to figure out what prediction corresponds to which record. Clearly, I'm overlooking something here, since this seems to follow a simple example from the documentation and nobody would want predictions without knowing what records they are for.\n\nWhen I instead used\n```\npreds = []\nfor batch in test_dataset:\n        imgs, StudyInstanceUIDs = batch \n        yhat = model.predict(imgs)\n        preds.append( (yhat, StudyInstanceUIDs) )\n```\nthat achieves what I want, but the version with `preds = model.predict(test_dataset, verbose = 1, workers=2, use_multiprocessing=True)` is about twice faster (or is that just all from the multi-processing??). It's 8 min 32 s vs. 3 min 29 s on the public LB set - thus, probably 34 vs. 14 min for the full LB dataset. With a model for 5 folds (never mind using TTA), that turns into 2h 50 min vs. 1h 10 min. \n\nSo, once I want to use multiple models, this is a real obstacle. Does anyone have any ideas or insights on what to do about this and to speed up the inference?",
      "votes": null
    },
    {
      "id": "1233981",
      "postDate": "03/10/2021 20:30:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> </p>\n<p>Have look here.</p>\n<p><a href=\"https://www.kaggle.com/c/jane-street-market-prediction/discussion/218752\" target=\"_blank\">https://www.kaggle.com/c/jane-street-market-prediction/discussion/218752</a></p>\n<p>all the best </p>",
      "rawMarkdown": "Hi @bjoernholzhauer \n\nHave look here.\n\nhttps://www.kaggle.com/c/jane-street-market-prediction/discussion/218752\n\nall the best",
      "votes": null
    },
    {
      "id": "1234011",
      "postDate": "03/10/2021 21:23:38",
      "content": "<p>For this competition, I am using this code</p>\n<pre><code>model1 = models.load_model('../input/.....h5')\nprint('model1.predict&gt;&gt;&gt;&gt;')\np1 = model1.predict(test_df, batch_size=20000)\ndel model1\n\nmodel2 = models.load_model('../input/....h5')\nprint('model2.predict&gt;&gt;&gt;&gt;')\np2 = model2.predict(test_df, batch_size=20000)\ndel model2\n\nmodel3 = models.load_model('../input/....h5')\nprint('model3.predict&gt;&gt;&gt;&gt;')\np3 = model3.predict(test_df, batch_size=20000)\ndel model3\n\nmodel4 = models.load_model('../inps.h5')\nprint('model4.predict&gt;&gt;&gt;&gt;')\np4 = model4.predict(test_df, batch_size=20000)\ndel model4\n\nmodel5 = models.load_model('../inps.h5')\n print('model5.predict&gt;&gt;&gt;&gt;')\n p5 = model5.predict(test_df, batch_size=20000)\n\nmodel6 = models.load_model('../inps.h5')\nprint('model6.predict&gt;&gt;&gt;&gt;')\np6 = model6.predict(test_df, batch_size=20000)\n\n\npredictions = (p1 + p2 + p3 + p4 + p5 +p6) / 6.0\n\n\nfrom numba import cuda\ncuda.select_device(0)\ncuda.close()\ncuda.select_device(0)\n</code></pre>",
      "rawMarkdown": "For this competition, I am using this code\n\n```\nmodel1 = models.load_model('../input/.....h5')\nprint('model1.predict>>>>')\np1 = model1.predict(test_df, batch_size=20000)\ndel model1\n\nmodel2 = models.load_model('../input/....h5')\nprint('model2.predict>>>>')\np2 = model2.predict(test_df, batch_size=20000)\ndel model2\n\nmodel3 = models.load_model('../input/....h5')\nprint('model3.predict>>>>')\np3 = model3.predict(test_df, batch_size=20000)\ndel model3\n\nmodel4 = models.load_model('../inps.h5')\nprint('model4.predict>>>>')\np4 = model4.predict(test_df, batch_size=20000)\ndel model4\n\nmodel5 = models.load_model('../inps.h5')\n print('model5.predict>>>>')\n p5 = model5.predict(test_df, batch_size=20000)\n\nmodel6 = models.load_model('../inps.h5')\nprint('model6.predict>>>>')\np6 = model6.predict(test_df, batch_size=20000)\n\n\npredictions = (p1 + p2 + p3 + p4 + p5 +p6) / 6.0\n\n\nfrom numba import cuda\ncuda.select_device(0)\ncuda.close()\ncuda.select_device(0)\n```",
      "votes": null
    },
    {
      "id": "1234026",
      "postDate": "03/10/2021 21:45:13",
      "content": "<p>I guess your dataset test_df is loaded in some known order? My concern was that the supposedly fastest method of reading tfrecords does not ensure any particular order…</p>",
      "rawMarkdown": "I guess your dataset test_df is loaded in some known order? My concern was that the supposedly fastest method of reading tfrecords does not ensure any particular order...",
      "votes": null
    },
    {
      "id": "1234322",
      "postDate": "03/11/2021 06:18:36",
      "content": "<p>for the **pipeline **   I am using this great code </p>\n<p><strong>Prediction</strong><br>\n<a href=\"https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-prediction\" target=\"_blank\">https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-prediction</a></p>\n<p><strong>Training</strong><br>\n<a href=\"https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline\" target=\"_blank\">https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline</a></p>\n<p>Thanks to <a href=\"https://www.kaggle.com/maksymshkliarevskyi\" target=\"_blank\">@maksymshkliarevskyi</a> </p>",
      "rawMarkdown": "for the **pipeline **   I am using this great code \n\n**Prediction**\nhttps://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-prediction\n\n**Training**\nhttps://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline\n\nThanks to @maksymshkliarevskyi",
      "votes": null
    },
    {
      "id": "1234326",
      "postDate": "03/11/2021 06:27:09",
      "content": "<p>for stratified Kfold data splitting I am using This great code</p>\n<p><a href=\"https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3\" target=\"_blank\">https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3</a></p>\n<p>Thanks to <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> </p>",
      "rawMarkdown": "for stratified Kfold data splitting I am using This great code\n\nhttps://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3\n\nThanks to @virilo",
      "votes": null
    },
    {
      "id": "1234343",
      "postDate": "03/11/2021 06:38:50",
      "content": "<p>Here he is using </p>\n<p><code>cache_dir='/kaggle/tf_cache</code></p>\n<p><a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit\" target=\"_blank\">https://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit</a></p>\n<p>I need to try this </p>\n<p>Thanks to <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> </p>",
      "rawMarkdown": "Here he is using \n\n`cache_dir='/kaggle/tf_cache`\n\nhttps://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit\n\nI need to try this \n\nThanks to @xhlulu",
      "votes": null
    },
    {
      "id": "1234383",
      "postDate": "03/11/2021 07:41:30",
      "content": "<p>For The order of Predictions using Test Time Augmentation (TTA):</p>\n<p>See This code: (2nd of Flower Classification on TPU) competition<br>\n<a href=\"https://www.kaggle.com/atamazian/fc-ensemble-external-data-effnet-densenet\" target=\"_blank\">https://www.kaggle.com/atamazian/fc-ensemble-external-data-effnet-densenet</a></p>\n<pre><code>def predict_tta(model, n_iter):\n    probs  = []\n    for i in range(n_iter):\n        test_ds = get_test_dataset(ordered=True) # since we are splitting the dataset and iterating separately on images and ids, order matters.\n        test_images_ds = test_ds.map(lambda image, idnum: image)\n        probs.append(model.predict(test_images_ds,verbose=0))\n\n    return probs\n</code></pre>\n<p>Thanks to <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a></p>",
      "rawMarkdown": "For The order of Predictions using Test Time Augmentation (TTA):\n\nSee This code: (2nd of Flower Classification on TPU) competition\nhttps://www.kaggle.com/atamazian/fc-ensemble-external-data-effnet-densenet\n\n\n```\n\ndef predict_tta(model, n_iter):\n    probs  = []\n    for i in range(n_iter):\n        test_ds = get_test_dataset(ordered=True) # since we are splitting the dataset and iterating separately on images and ids, order matters.\n        test_images_ds = test_ds.map(lambda image, idnum: image)\n        probs.append(model.predict(test_images_ds,verbose=0))\n        \n    return probs\n```\n\nThanks to @atamazian",
      "votes": null
    },
    {
      "id": "1236395",
      "postDate": "03/13/2021 05:21:54",
      "content": "<p>Hello!</p>\n<p>As for me, I use <code>image_ids = [x[1].decode('utf-8') for x in dataset.unbatch().as_numpy_iterator()]</code> to get back the ids.</p>\n<p>Another option I found useful is setting mixed precision for inference like this: <code>tf.keras.mixed_precision.set_global_policy('mixed_float16')</code> (your model should than end with <code>Activation('sigmoid', dtype='float32')</code> layer explicitly). This additionally saves up to 2.5x time for me, ceteris paribus (203s vs 537s for making predictions for 5K 768x768 images with EfficientNetB6)</p>",
      "rawMarkdown": "Hello!\n\nAs for me, I use `image_ids = [x[1].decode('utf-8') for x in dataset.unbatch().as_numpy_iterator()]` to get back the ids.\n\nAnother option I found useful is setting mixed precision for inference like this: `tf.keras.mixed_precision.set_global_policy('mixed_float16')` (your model should than end with `Activation('sigmoid', dtype='float32')` layer explicitly). This additionally saves up to 2.5x time for me, ceteris paribus (203s vs 537s for making predictions for 5K 768x768 images with EfficientNetB6)",
      "votes": null
    },
    {
      "id": "1236956",
      "postDate": "03/13/2021 15:23:45",
      "content": "<p>That sounds incredibly useful, Thank you, I will try out later today.</p>",
      "rawMarkdown": "That sounds incredibly useful, Thank you, I will try out later today.",
      "votes": null
    },
    {
      "id": "1237046",
      "postDate": "03/13/2021 17:17:44",
      "content": "<p>My notebook is in TF 2.3.1, in which case it's <code>tf.keras.mixed_precision.experimental.set_policy('mixed_float16')</code>, as it turns out.</p>",
      "rawMarkdown": "My notebook is in TF 2.3.1, in which case it's `tf.keras.mixed_precision.experimental.set_policy('mixed_float16')`, as it turns out.",
      "votes": null
    },
    {
      "id": "1237105",
      "postDate": "03/13/2021 18:48:30",
      "content": "<p>This was only 8% faster for me on the public LB (of course, there's a bunch of set-up overhead, so presumably somewhat more speed gains on the full LB) with the notebook I tried. That is still great, if it costs me next to nothing. I've now submitted the notebook to see whether using mixed precision affects the public LB score.</p>\n<p><strong>Update</strong>: In terms of public LB score, this did no worse than the non-mixed precision version (in fact, the score looks the same to the decimal places I can see, but is ranked better when I sort \"My submissions\" by \"Public Score\").</p>",
      "rawMarkdown": "This was only 8% faster for me on the public LB (of course, there's a bunch of set-up overhead, so presumably somewhat more speed gains on the full LB) with the notebook I tried. That is still great, if it costs me next to nothing. I've now submitted the notebook to see whether using mixed precision affects the public LB score.\n\n**Update**: In terms of public LB score, this did no worse than the non-mixed precision version (in fact, the score looks the same to the decimal places I can see, but is ranked better when I sort \"My submissions\" by \"Public Score\").",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1233981,
      "author_name": "faisalalsrheed",
      "author_url": "",
      "post_date": "03/10/2021 20:30:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> </p>\n<p>Have look here.</p>\n<p><a href=\"https://www.kaggle.com/c/jane-street-market-prediction/discussion/218752\" target=\"_blank\">https://www.kaggle.com/c/jane-street-market-prediction/discussion/218752</a></p>\n<p>all the best </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1234011,
      "author_name": "faisalalsrheed",
      "author_url": "",
      "post_date": "03/10/2021 21:23:38",
      "content": "<p>For this competition, I am using this code</p>\n<pre><code>model1 = models.load_model('../input/.....h5')\nprint('model1.predict&gt;&gt;&gt;&gt;')\np1 = model1.predict(test_df, batch_size=20000)\ndel model1\n\nmodel2 = models.load_model('../input/....h5')\nprint('model2.predict&gt;&gt;&gt;&gt;')\np2 = model2.predict(test_df, batch_size=20000)\ndel model2\n\nmodel3 = models.load_model('../input/....h5')\nprint('model3.predict&gt;&gt;&gt;&gt;')\np3 = model3.predict(test_df, batch_size=20000)\ndel model3\n\nmodel4 = models.load_model('../inps.h5')\nprint('model4.predict&gt;&gt;&gt;&gt;')\np4 = model4.predict(test_df, batch_size=20000)\ndel model4\n\nmodel5 = models.load_model('../inps.h5')\n print('model5.predict&gt;&gt;&gt;&gt;')\n p5 = model5.predict(test_df, batch_size=20000)\n\nmodel6 = models.load_model('../inps.h5')\nprint('model6.predict&gt;&gt;&gt;&gt;')\np6 = model6.predict(test_df, batch_size=20000)\n\n\npredictions = (p1 + p2 + p3 + p4 + p5 +p6) / 6.0\n\n\nfrom numba import cuda\ncuda.select_device(0)\ncuda.close()\ncuda.select_device(0)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1234026,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/10/2021 21:45:13",
          "content": "<p>I guess your dataset test_df is loaded in some known order? My concern was that the supposedly fastest method of reading tfrecords does not ensure any particular order…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1234322,
          "author_name": "faisalalsrheed",
          "author_url": "",
          "post_date": "03/11/2021 06:18:36",
          "content": "<p>for the **pipeline **   I am using this great code </p>\n<p><strong>Prediction</strong><br>\n<a href=\"https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-prediction\" target=\"_blank\">https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-prediction</a></p>\n<p><strong>Training</strong><br>\n<a href=\"https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline\" target=\"_blank\">https://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline</a></p>\n<p>Thanks to <a href=\"https://www.kaggle.com/maksymshkliarevskyi\" target=\"_blank\">@maksymshkliarevskyi</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1234326,
          "author_name": "faisalalsrheed",
          "author_url": "",
          "post_date": "03/11/2021 06:27:09",
          "content": "<p>for stratified Kfold data splitting I am using This great code</p>\n<p><a href=\"https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3\" target=\"_blank\">https://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3</a></p>\n<p>Thanks to <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1234343,
          "author_name": "faisalalsrheed",
          "author_url": "",
          "post_date": "03/11/2021 06:38:50",
          "content": "<p>Here he is using </p>\n<p><code>cache_dir='/kaggle/tf_cache</code></p>\n<p><a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit\" target=\"_blank\">https://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit</a></p>\n<p>I need to try this </p>\n<p>Thanks to <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1234383,
          "author_name": "faisalalsrheed",
          "author_url": "",
          "post_date": "03/11/2021 07:41:30",
          "content": "<p>For The order of Predictions using Test Time Augmentation (TTA):</p>\n<p>See This code: (2nd of Flower Classification on TPU) competition<br>\n<a href=\"https://www.kaggle.com/atamazian/fc-ensemble-external-data-effnet-densenet\" target=\"_blank\">https://www.kaggle.com/atamazian/fc-ensemble-external-data-effnet-densenet</a></p>\n<pre><code>def predict_tta(model, n_iter):\n    probs  = []\n    for i in range(n_iter):\n        test_ds = get_test_dataset(ordered=True) # since we are splitting the dataset and iterating separately on images and ids, order matters.\n        test_images_ds = test_ds.map(lambda image, idnum: image)\n        probs.append(model.predict(test_images_ds,verbose=0))\n\n    return probs\n</code></pre>\n<p>Thanks to <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1236395,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "03/13/2021 05:21:54",
      "content": "<p>Hello!</p>\n<p>As for me, I use <code>image_ids = [x[1].decode('utf-8') for x in dataset.unbatch().as_numpy_iterator()]</code> to get back the ids.</p>\n<p>Another option I found useful is setting mixed precision for inference like this: <code>tf.keras.mixed_precision.set_global_policy('mixed_float16')</code> (your model should than end with <code>Activation('sigmoid', dtype='float32')</code> layer explicitly). This additionally saves up to 2.5x time for me, ceteris paribus (203s vs 537s for making predictions for 5K 768x768 images with EfficientNetB6)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1236956,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/13/2021 15:23:45",
          "content": "<p>That sounds incredibly useful, Thank you, I will try out later today.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1237046,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/13/2021 17:17:44",
          "content": "<p>My notebook is in TF 2.3.1, in which case it's <code>tf.keras.mixed_precision.experimental.set_policy('mixed_float16')</code>, as it turns out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1237105,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/13/2021 18:48:30",
          "content": "<p>This was only 8% faster for me on the public LB (of course, there's a bunch of set-up overhead, so presumably somewhat more speed gains on the full LB) with the notebook I tried. That is still great, if it costs me next to nothing. I've now submitted the notebook to see whether using mixed precision affects the public LB score.</p>\n<p><strong>Update</strong>: In terms of public LB score, this did no worse than the non-mixed precision version (in fact, the score looks the same to the decimal places I can see, but is ranked better when I sort \"My submissions\" by \"Public Score\").</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1233607": "I've got a pretty large model I trained in `keras` and inference time is an issue. I wondered about using the tfrecords to speed up predictions and used these options (as described [here in the keras documentation](https://keras.io/examples/keras_recipes/tfrecord/) and I also looked at this [notebook]( https://www.kaggle.com/tpothjuan/efficientnetb7-tfrecords)) for inspiration:\n```\ndef load_dataset(filenames):\n    ignore_order = tf.data.Options()\n    ignore_order.experimental_deterministic = False  # disable order, increase speed, keep order ensure ordering\n    dataset = tf.data.TFRecordDataset(filenames)  # automatically interleaves reads from multiple files\n    dataset = dataset.with_options(ignore_order)  # uses data as soon as it streams in, rather than in its original order\n    dataset = dataset.map(partial(read_tfrecord), num_parallel_calls=AUTOTUNE)\n    # returns a dataset of just images (test data are not labelled)\n    return dataset\n```\nand then I did `preds = model.predict(test_dataset, verbose = 1, workers=2, use_multiprocessing=True)`. However, while my `read_tfrecord` function does return a tuple of image and image_id, and iterating on the dataset I create via `test_dataset = get_dataset(TEST_FILENAMES)` does indeed return both, the `model.predict` function does not return these. I.e. there seems to be no obvious way to figure out what prediction corresponds to which record. Clearly, I'm overlooking something here, since this seems to follow a simple example from the documentation and nobody would want predictions without knowing what records they are for.\n\nWhen I instead used\n```\npreds = []\nfor batch in test_dataset:\n        imgs, StudyInstanceUIDs = batch \n        yhat = model.predict(imgs)\n        preds.append( (yhat, StudyInstanceUIDs) )\n```\nthat achieves what I want, but the version with `preds = model.predict(test_dataset, verbose = 1, workers=2, use_multiprocessing=True)` is about twice faster (or is that just all from the multi-processing??). It's 8 min 32 s vs. 3 min 29 s on the public LB set - thus, probably 34 vs. 14 min for the full LB dataset. With a model for 5 folds (never mind using TTA), that turns into 2h 50 min vs. 1h 10 min. \n\nSo, once I want to use multiple models, this is a real obstacle. Does anyone have any ideas or insights on what to do about this and to speed up the inference?",
    "1233981": "Hi @bjoernholzhauer \n\nHave look here.\n\nhttps://www.kaggle.com/c/jane-street-market-prediction/discussion/218752\n\nall the best",
    "1234011": "For this competition, I am using this code\n\n```\nmodel1 = models.load_model('../input/.....h5')\nprint('model1.predict>>>>')\np1 = model1.predict(test_df, batch_size=20000)\ndel model1\n\nmodel2 = models.load_model('../input/....h5')\nprint('model2.predict>>>>')\np2 = model2.predict(test_df, batch_size=20000)\ndel model2\n\nmodel3 = models.load_model('../input/....h5')\nprint('model3.predict>>>>')\np3 = model3.predict(test_df, batch_size=20000)\ndel model3\n\nmodel4 = models.load_model('../inps.h5')\nprint('model4.predict>>>>')\np4 = model4.predict(test_df, batch_size=20000)\ndel model4\n\nmodel5 = models.load_model('../inps.h5')\n print('model5.predict>>>>')\n p5 = model5.predict(test_df, batch_size=20000)\n\nmodel6 = models.load_model('../inps.h5')\nprint('model6.predict>>>>')\np6 = model6.predict(test_df, batch_size=20000)\n\n\npredictions = (p1 + p2 + p3 + p4 + p5 +p6) / 6.0\n\n\nfrom numba import cuda\ncuda.select_device(0)\ncuda.close()\ncuda.select_device(0)\n```",
    "1234026": "I guess your dataset test_df is loaded in some known order? My concern was that the supposedly fastest method of reading tfrecords does not ensure any particular order...",
    "1234322": "for the **pipeline **   I am using this great code \n\n**Prediction**\nhttps://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-prediction\n\n**Training**\nhttps://www.kaggle.com/maksymshkliarevskyi/ranzcr-xception-tpu-baseline\n\nThanks to @maksymshkliarevskyi",
    "1234326": "for stratified Kfold data splitting I am using This great code\n\nhttps://www.kaggle.com/virilo/ranzcr-clip-stratified-kfold-to-team-up-v3\n\nThanks to @virilo",
    "1234343": "Here he is using \n\n`cache_dir='/kaggle/tf_cache`\n\nhttps://www.kaggle.com/xhlulu/ranzcr-efficientnet-gpu-starter-train-submit\n\nI need to try this \n\nThanks to @xhlulu",
    "1234383": "For The order of Predictions using Test Time Augmentation (TTA):\n\nSee This code: (2nd of Flower Classification on TPU) competition\nhttps://www.kaggle.com/atamazian/fc-ensemble-external-data-effnet-densenet\n\n\n```\n\ndef predict_tta(model, n_iter):\n    probs  = []\n    for i in range(n_iter):\n        test_ds = get_test_dataset(ordered=True) # since we are splitting the dataset and iterating separately on images and ids, order matters.\n        test_images_ds = test_ds.map(lambda image, idnum: image)\n        probs.append(model.predict(test_images_ds,verbose=0))\n        \n    return probs\n```\n\nThanks to @atamazian",
    "1236395": "Hello!\n\nAs for me, I use `image_ids = [x[1].decode('utf-8') for x in dataset.unbatch().as_numpy_iterator()]` to get back the ids.\n\nAnother option I found useful is setting mixed precision for inference like this: `tf.keras.mixed_precision.set_global_policy('mixed_float16')` (your model should than end with `Activation('sigmoid', dtype='float32')` layer explicitly). This additionally saves up to 2.5x time for me, ceteris paribus (203s vs 537s for making predictions for 5K 768x768 images with EfficientNetB6)",
    "1236956": "That sounds incredibly useful, Thank you, I will try out later today.",
    "1237046": "My notebook is in TF 2.3.1, in which case it's `tf.keras.mixed_precision.experimental.set_policy('mixed_float16')`, as it turns out.",
    "1237105": "This was only 8% faster for me on the public LB (of course, there's a bunch of set-up overhead, so presumably somewhat more speed gains on the full LB) with the notebook I tried. That is still great, if it costs me next to nothing. I've now submitted the notebook to see whether using mixed precision affects the public LB score.\n\n**Update**: In terms of public LB score, this did no worse than the non-mixed precision version (in fact, the score looks the same to the decimal places I can see, but is ranked better when I sort \"My submissions\" by \"Public Score\")."
  },
  "source": "meta"
}