{
  "id": 267718,
  "title": "TFRecords for all train data res 384",
  "url": "/competitions/landmark-recognition-2021/discussion/267718",
  "author_name": "Mark Wijkhuizen",
  "post_date": "2021-08-24T11:40:23.234000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>Since the training dataset is huge with over 1.5M images I created TFRecords to read the images quickly in a dataset pipeline. On a TPU the TFRecords can be read with ~6000 images per second, which is far quicker than reading them one by one. There are 3 separate datasets, since the dataset limit is 20GB. A notebook demonstrating the creation of the TFRecords can be found <a href=\"https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-2021-tfrecords-res-384\" target=\"_blank\">here</a>. The TFRecords datasets can be found here:</p>\n<p><a href=\"https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-1\" target=\"_blank\">Part 1</a></p>\n<p><a href=\"https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-2\" target=\"_blank\">Part 2</a></p>\n<p><a href=\"https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-3\" target=\"_blank\">Part 3</a></p>\n<p>To create a TFRecordDataset the following code can be used:</p>\n<pre><code>GCS_DS_PATH_1 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-1')\nGCS_DS_PATH_2 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-2')\nGCS_DS_PATH_3 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-3')\n\ndef get_train_dataset():\n\n    FNAMES_TRAIN_TFRECORDS = (\n        tf.io.gfile.glob(f'{GCS_DS_PATH_1}/*.tfrecords') +\n        tf.io.gfile.glob(f'{GCS_DS_PATH_2}/*.tfrecords') +\n        tf.io.gfile.glob(f'{GCS_DS_PATH_3}/*.tfrecords')\n    )\n\n    train_dataset = tf.data.TFRecordDataset(FNAMES_TRAIN_TFRECORDS)\n    train_dataset = train_dataset.prefetch(AUTO)\n    train_dataset = train_dataset.repeat()\n    train_dataset = train_dataset.map(decode_tfrecord)\n    train_dataset = train_dataset.batch(32)\n\n    return train_dataset\n</code></pre>\n<p>Hope these datasets help with creating training pipelines!</p>",
  "messages": [
    {
      "id": 1488584,
      "postDate": "2021-08-24T11:40:23.233Z",
      "content": "<p>Hi all,</p>\n<p>Since the training dataset is huge with over 1.5M images I created TFRecords to read the images quickly in a dataset pipeline. On a TPU the TFRecords can be read with ~6000 images per second, which is far quicker than reading them one by one. There are 3 separate datasets, since the dataset limit is 20GB. A notebook demonstrating the creation of the TFRecords can be found <a href=\"https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-2021-tfrecords-res-384\" target=\"_blank\">here</a>. The TFRecords datasets can be found here:</p>\n<p><a href=\"https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-1\" target=\"_blank\">Part 1</a></p>\n<p><a href=\"https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-2\" target=\"_blank\">Part 2</a></p>\n<p><a href=\"https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-3\" target=\"_blank\">Part 3</a></p>\n<p>To create a TFRecordDataset the following code can be used:</p>\n<pre><code>GCS_DS_PATH_1 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-1')\nGCS_DS_PATH_2 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-2')\nGCS_DS_PATH_3 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-3')\n\ndef get_train_dataset():\n\n    FNAMES_TRAIN_TFRECORDS = (\n        tf.io.gfile.glob(f'{GCS_DS_PATH_1}/*.tfrecords') +\n        tf.io.gfile.glob(f'{GCS_DS_PATH_2}/*.tfrecords') +\n        tf.io.gfile.glob(f'{GCS_DS_PATH_3}/*.tfrecords')\n    )\n\n    train_dataset = tf.data.TFRecordDataset(FNAMES_TRAIN_TFRECORDS)\n    train_dataset = train_dataset.prefetch(AUTO)\n    train_dataset = train_dataset.repeat()\n    train_dataset = train_dataset.map(decode_tfrecord)\n    train_dataset = train_dataset.batch(32)\n\n    return train_dataset\n</code></pre>\n<p>Hope these datasets help with creating training pipelines!</p>",
      "rawMarkdown": "Hi all,\n\nSince the training dataset is huge with over 1.5M images I created TFRecords to read the images quickly in a dataset pipeline. On a TPU the TFRecords can be read with ~6000 images per second, which is far quicker than reading them one by one. There are 3 separate datasets, since the dataset limit is 20GB. A notebook demonstrating the creation of the TFRecords can be found [here](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-2021-tfrecords-res-384). The TFRecords datasets can be found here:\n\n[Part 1](https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-1)\n\n[Part 2](https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-2)\n\n[Part 3](https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-3)\n\nTo create a TFRecordDataset the following code can be used:\n\n```\nGCS_DS_PATH_1 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-1')\nGCS_DS_PATH_2 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-2')\nGCS_DS_PATH_3 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-3')\n\ndef get_train_dataset():\n    \n    FNAMES_TRAIN_TFRECORDS = (\n        tf.io.gfile.glob(f'{GCS_DS_PATH_1}/*.tfrecords') +\n        tf.io.gfile.glob(f'{GCS_DS_PATH_2}/*.tfrecords') +\n        tf.io.gfile.glob(f'{GCS_DS_PATH_3}/*.tfrecords')\n    )\n    \n    train_dataset = tf.data.TFRecordDataset(FNAMES_TRAIN_TFRECORDS)\n    train_dataset = train_dataset.prefetch(AUTO)\n    train_dataset = train_dataset.repeat()\n    train_dataset = train_dataset.map(decode_tfrecord)\n    train_dataset = train_dataset.batch(32)\n    \n    return train_dataset\n```\n\nHope these datasets help with creating training pipelines!",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1488584": "Hi all,\n\nSince the training dataset is huge with over 1.5M images I created TFRecords to read the images quickly in a dataset pipeline. On a TPU the TFRecords can be read with ~6000 images per second, which is far quicker than reading them one by one. There are 3 separate datasets, since the dataset limit is 20GB. A notebook demonstrating the creation of the TFRecords can be found [here](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-2021-tfrecords-res-384). The TFRecords datasets can be found here:\n\n[Part 1](https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-1)\n\n[Part 2](https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-2)\n\n[Part 3](https://www.kaggle.com/markwijkhuizen/landmark-recognition-2021-tfrecords-384-part-3)\n\nTo create a TFRecordDataset the following code can be used:\n\n```\nGCS_DS_PATH_1 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-1')\nGCS_DS_PATH_2 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-2')\nGCS_DS_PATH_3 = KaggleDatasets().get_gcs_path('landmark-recognition-2021-tfrecords-384-part-3')\n\ndef get_train_dataset():\n    \n    FNAMES_TRAIN_TFRECORDS = (\n        tf.io.gfile.glob(f'{GCS_DS_PATH_1}/*.tfrecords') +\n        tf.io.gfile.glob(f'{GCS_DS_PATH_2}/*.tfrecords') +\n        tf.io.gfile.glob(f'{GCS_DS_PATH_3}/*.tfrecords')\n    )\n    \n    train_dataset = tf.data.TFRecordDataset(FNAMES_TRAIN_TFRECORDS)\n    train_dataset = train_dataset.prefetch(AUTO)\n    train_dataset = train_dataset.repeat()\n    train_dataset = train_dataset.map(decode_tfrecord)\n    train_dataset = train_dataset.batch(32)\n    \n    return train_dataset\n```\n\nHope these datasets help with creating training pipelines!"
  }
}