{
  "id": 266008,
  "title": "Get dataset ID faster during TPU inference",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/266008",
  "author_name": "",
  "post_date": "2021-08-17T15:38:09.189405200Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Many great notebooks using this code in tpu inference:</p>\n<pre><code>ds_test = get_dataset(files_test_all, batch_size=BATCH_SIZE * 2, repeat=False, shuffle=False, aug=False, labeled=False, return_image_ids=True)\nstart_time = perf_counter()\nfile_ids = np.array([target.numpy().decode(\"utf-8\") for img, target in tqdm(ds_test.unbatch())])\nprint('iter by for loop', f'finished; duration = {perf_counter() - start_time} s')\n</code></pre>\n<p><em>Get test tfrecordes dataset id about 7 mins</em></p>\n<p>Instead, We can do it the same way using<br>\n<strong>Putting all the ids in one batch and retrieve it back</strong></p>\n<pre><code>ds = get_dataset(files_test_all, batch_size=BATCH_SIZE * 2, repeat=False, shuffle=False, aug=False, labeled=False, return_image_ids=True)\nid_ds = ds.map(lambda image, label: label).unbatch()\nstart_time = perf_counter()\nids = next(iter(id_ds.batch(NUM_TEST_IMAGES))).numpy().astype('U')\nprint('one batch', f'finished; duration = {perf_counter() - start_time} s')\n</code></pre>\n<p><em>Get test tfrecordes dataset id about 2 mins</em></p>\n<p>Saving time, Maybe useful when you are doing oof inference while the data is huge…</p>\n<p>I make a detail notebook FYI <a href=\"https://www.kaggle.com/mozhiwenmzw/get-data-id-faster-during-tpu-inference?scriptVersionId=72131222\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "1477702",
      "postDate": "08/17/2021 15:38:09",
      "content": "<p>Many great notebooks using this code in tpu inference:</p>\n<pre><code>ds_test = get_dataset(files_test_all, batch_size=BATCH_SIZE * 2, repeat=False, shuffle=False, aug=False, labeled=False, return_image_ids=True)\nstart_time = perf_counter()\nfile_ids = np.array([target.numpy().decode(\"utf-8\") for img, target in tqdm(ds_test.unbatch())])\nprint('iter by for loop', f'finished; duration = {perf_counter() - start_time} s')\n</code></pre>\n<p><em>Get test tfrecordes dataset id about 7 mins</em></p>\n<p>Instead, We can do it the same way using<br>\n<strong>Putting all the ids in one batch and retrieve it back</strong></p>\n<pre><code>ds = get_dataset(files_test_all, batch_size=BATCH_SIZE * 2, repeat=False, shuffle=False, aug=False, labeled=False, return_image_ids=True)\nid_ds = ds.map(lambda image, label: label).unbatch()\nstart_time = perf_counter()\nids = next(iter(id_ds.batch(NUM_TEST_IMAGES))).numpy().astype('U')\nprint('one batch', f'finished; duration = {perf_counter() - start_time} s')\n</code></pre>\n<p><em>Get test tfrecordes dataset id about 2 mins</em></p>\n<p>Saving time, Maybe useful when you are doing oof inference while the data is huge…</p>\n<p>I make a detail notebook FYI <a href=\"https://www.kaggle.com/mozhiwenmzw/get-data-id-faster-during-tpu-inference?scriptVersionId=72131222\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Many great notebooks using this code in tpu inference:\n\n```python3\nds_test = get_dataset(files_test_all, batch_size=BATCH_SIZE * 2, repeat=False, shuffle=False, aug=False, labeled=False, return_image_ids=True)\nstart_time = perf_counter()\nfile_ids = np.array([target.numpy().decode(\"utf-8\") for img, target in tqdm(ds_test.unbatch())])\nprint('iter by for loop', f'finished; duration = {perf_counter() - start_time} s')\n```\n*Get test tfrecordes dataset id about 7 mins*\n\nInstead, We can do it the same way using\n**Putting all the ids in one batch and retrieve it back**\n```python3\nds = get_dataset(files_test_all, batch_size=BATCH_SIZE * 2, repeat=False, shuffle=False, aug=False, labeled=False, return_image_ids=True)\nid_ds = ds.map(lambda image, label: label).unbatch()\nstart_time = perf_counter()\nids = next(iter(id_ds.batch(NUM_TEST_IMAGES))).numpy().astype('U')\nprint('one batch', f'finished; duration = {perf_counter() - start_time} s')\n```\n*Get test tfrecordes dataset id about 2 mins*\n\nSaving time, Maybe useful when you are doing oof inference while the data is huge...\n\nI make a detail notebook FYI [here](https://www.kaggle.com/mozhiwenmzw/get-data-id-faster-during-tpu-inference?scriptVersionId=72131222)",
      "votes": null
    },
    {
      "id": "1478247",
      "postDate": "08/17/2021 21:51:03",
      "content": "<p>It is very useful indeed! Thanks for sharing.</p>",
      "rawMarkdown": "It is very useful indeed! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1478422",
      "postDate": "08/18/2021 01:42:44",
      "content": "<p>Glad to hear that!</p>",
      "rawMarkdown": "Glad to hear that!",
      "votes": null
    },
    {
      "id": "1478956",
      "postDate": "08/18/2021 08:12:07",
      "content": "<p>Thanks ! Why is is faster though ?</p>",
      "rawMarkdown": "Thanks ! Why is is faster though ?",
      "votes": null
    },
    {
      "id": "1480384",
      "postDate": "08/19/2021 02:16:16",
      "content": "<p>I didn't dive into it, but I guess tf dataset good at batch processing. If we fetch data by for loop, then we can only make little use of the advantage of tf dataset.</p>",
      "rawMarkdown": "I didn't dive into it, but I guess tf dataset good at batch processing. If we fetch data by for loop, then we can only make little use of the advantage of tf dataset.",
      "votes": null
    },
    {
      "id": "1561218",
      "postDate": "10/27/2021 12:20:51",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1478247,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "08/17/2021 21:51:03",
      "content": "<p>It is very useful indeed! Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1478422,
          "author_name": "mozhiwenmzw",
          "author_url": "",
          "post_date": "08/18/2021 01:42:44",
          "content": "<p>Glad to hear that!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1478956,
      "author_name": "zarif98sjs",
      "author_url": "",
      "post_date": "08/18/2021 08:12:07",
      "content": "<p>Thanks ! Why is is faster though ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1480384,
          "author_name": "mozhiwenmzw",
          "author_url": "",
          "post_date": "08/19/2021 02:16:16",
          "content": "<p>I didn't dive into it, but I guess tf dataset good at batch processing. If we fetch data by for loop, then we can only make little use of the advantage of tf dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1561218,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 12:20:51",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1477702": "Many great notebooks using this code in tpu inference:\n\n```python3\nds_test = get_dataset(files_test_all, batch_size=BATCH_SIZE * 2, repeat=False, shuffle=False, aug=False, labeled=False, return_image_ids=True)\nstart_time = perf_counter()\nfile_ids = np.array([target.numpy().decode(\"utf-8\") for img, target in tqdm(ds_test.unbatch())])\nprint('iter by for loop', f'finished; duration = {perf_counter() - start_time} s')\n```\n*Get test tfrecordes dataset id about 7 mins*\n\nInstead, We can do it the same way using\n**Putting all the ids in one batch and retrieve it back**\n```python3\nds = get_dataset(files_test_all, batch_size=BATCH_SIZE * 2, repeat=False, shuffle=False, aug=False, labeled=False, return_image_ids=True)\nid_ds = ds.map(lambda image, label: label).unbatch()\nstart_time = perf_counter()\nids = next(iter(id_ds.batch(NUM_TEST_IMAGES))).numpy().astype('U')\nprint('one batch', f'finished; duration = {perf_counter() - start_time} s')\n```\n*Get test tfrecordes dataset id about 2 mins*\n\nSaving time, Maybe useful when you are doing oof inference while the data is huge...\n\nI make a detail notebook FYI [here](https://www.kaggle.com/mozhiwenmzw/get-data-id-faster-during-tpu-inference?scriptVersionId=72131222)",
    "1478247": "It is very useful indeed! Thanks for sharing.",
    "1478422": "Glad to hear that!",
    "1478956": "Thanks ! Why is is faster though ?",
    "1480384": "I didn't dive into it, but I guess tf dataset good at batch processing. If we fetch data by for loop, then we can only make little use of the advantage of tf dataset.",
    "1561218": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}