{
  "id": 267241,
  "title": "Read data from TFRecords, train with PyTorch - Mixture of two frameworks",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/267241",
  "author_name": "",
  "post_date": "2021-08-22T11:56:22.330585Z",
  "votes": 74,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi kagglers,</p>\n<p>In this post, I'll share a way to read data from tfrecords while using PyTorch framework to train your model. There are two advantage of doing this:</p>\n<ol>\n<li>You can speed up your training.</li>\n<li>You don't need to download dataset. This is quite good for those who use Colab (Pro/Pro+) to train their models.</li>\n</ol>\n<h2>First step - get the TFRecord paths</h2>\n<p>To read TFRecord data, you only need to know GCS path of the TFRecord. If you use Kaggle Datasets, you can get GCS path of the Kaggle Dataset as follows:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\n\n\ngcs_path = KaggleDatasets().get_gcs_path(\"&lt;name of your dataset&gt;\")\n</code></pre>\n<p>This GCS path changes in around 1 week. If it changes you need to rerun this code. Note that it would take long time to get GCS path for the first time, and you may get Timeout error. In that case you only need to rerun the code until you get the path.</p>\n<h2>Second step - define custom Data Loader</h2>\n<p>Next, you need to define tensorflow dataset. Here's an example of a tensorflow dataset that reads waveform from my dataset.</p>\n<pre><code>import tensorflow as tf\nimport tensorflow_datasets as tfds\n\n\ndef count_data_items(fileids):\n    \"\"\"\n    Count the number of samples.\n    Each of the TFRecord datasets is designed to contain 28000 samples.\n    \"\"\"\n    return len(fileids) * 28000\n\n\nAUTO = tf.data.experimental.AUTOTUNE\n\n\ndef prepare_wave(wave):\n    wave = tf.reshape(tf.io.decode_raw(wave, tf.float64), (3, 4096))\n    normalized_waves = []\n    for i in range(3):\n        normalized_wave = wave[i] / tf.math.reduce_max(wave[i])\n        normalized_waves.append(normalized_wave)\n    wave = tf.stack(normalized_waves, axis=0)\n    wave = tf.cast(wave, tf.float32)\n    return wave\n\n\ndef read_labeled_tfrecord(example):\n    tfrec_format = {\n        \"wave\": tf.io.FixedLenFeature([], tf.string),\n        \"wave_id\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    return prepare_wave(example[\"wave\"]), tf.reshape(tf.cast(example[\"target\"], tf.float32), [1]), example[\"wave_id\"]\n\n\ndef read_unlabeled_tfrecord(example, return_image_id):\n    tfrec_format = {\n        \"wave\": tf.io.FixedLenFeature([], tf.string),\n        \"wave_id\": tf.io.FixedLenFeature([], tf.string)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    return prepare_wave(example[\"wave\"]), example[\"wave_id\"] if return_image_id else 0\n\n\ndef get_dataset(files, batch_size=16, repeat=False, cache=False, shuffle=False, labeled=True, return_image_ids=True):\n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=AUTO, compression_type=\"GZIP\")\n    if cache:\n        # You'll need around 15GB RAM if you'd like to cache val dataset, and 50~60GB RAM for train dataset.\n        ds = ds.cache()\n\n    if repeat:\n        ds = ds.repeat()\n\n    if shuffle:\n        ds = ds.shuffle(1024 * 2)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n\n    if labeled:\n        ds = ds.map(read_labeled_tfrecord, num_parallel_calls=AUTO)\n    else:\n        ds = ds.map(lambda example: read_unlabeled_tfrecord(example, return_image_ids), num_parallel_calls=AUTO)\n\n    ds = ds.batch(batch_size)\n    ds = ds.prefetch(AUTO)\n    return tfds.as_numpy(ds)\n</code></pre>\n<p>and then define a DataLoader class which takes the role of <code>torch.utils.data.DataLoader</code></p>\n<pre><code>class TFRecordDataLoader:\n    def __init__(self, files, batch_size=32, cache=False, train=True, repeat=False, shuffle=False, labeled=True, return_image_ids=True):\n        self.ds = get_dataset(\n            files, \n            batch_size=batch_size,\n            cache=cache,\n            repeat=repeat,\n            shuffle=shuffle,\n            labeled=labeled,\n            return_image_ids=return_image_ids)\n\n        if train:\n            self.num_examples = count_data_items(files)\n        else:\n            self.num_examples = count_data_items_test(files)\n\n        self.batch_size = batch_size\n        self.labeled = labeled\n        self.return_image_ids = return_image_ids\n        self._iterator = None\n\n    def __iter__(self):\n        if self._iterator is None:\n            self._iterator = iter(self.ds)\n        else:\n            self._reset()\n        return self._iterator\n\n    def _reset(self):\n        self._iterator = iter(self.ds)\n\n    def __next__(self):\n        batch = next(self._iterator)\n        return batch\n\n    def __len__(self):\n        n_batches = self.num_examples // self.batch_size\n        if self.num_examples % self.batch_size == 0:\n            return n_batches\n        else:\n            return n_batches + 1\n</code></pre>\n<p>Finally, you need to use custom DataLoader instead of using pytorch's DataLoader.</p>\n<p>Here's an example Notebook.</p>\n<p><a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch</a></p>\n<h2>Who will benefit from this?</h2>\n<p>This technique is for PyTorch users. Especially who:</p>\n<ul>\n<li>use Colab or other cloud platform to train models.<ul>\n<li>With this technique, you don't need to worry about the storage of the data.</li></ul></li>\n<li>want to speed up training<ul>\n<li>If you have RAM more than 15GB you can cache val dataset which speed up the training even more. If you have RAM more than 50 ~ 60GB you can cache the whole dataset.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1485758",
      "postDate": "08/22/2021 11:56:22",
      "content": "<p>Hi kagglers,</p>\n<p>In this post, I'll share a way to read data from tfrecords while using PyTorch framework to train your model. There are two advantage of doing this:</p>\n<ol>\n<li>You can speed up your training.</li>\n<li>You don't need to download dataset. This is quite good for those who use Colab (Pro/Pro+) to train their models.</li>\n</ol>\n<h2>First step - get the TFRecord paths</h2>\n<p>To read TFRecord data, you only need to know GCS path of the TFRecord. If you use Kaggle Datasets, you can get GCS path of the Kaggle Dataset as follows:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\n\n\ngcs_path = KaggleDatasets().get_gcs_path(\"&lt;name of your dataset&gt;\")\n</code></pre>\n<p>This GCS path changes in around 1 week. If it changes you need to rerun this code. Note that it would take long time to get GCS path for the first time, and you may get Timeout error. In that case you only need to rerun the code until you get the path.</p>\n<h2>Second step - define custom Data Loader</h2>\n<p>Next, you need to define tensorflow dataset. Here's an example of a tensorflow dataset that reads waveform from my dataset.</p>\n<pre><code>import tensorflow as tf\nimport tensorflow_datasets as tfds\n\n\ndef count_data_items(fileids):\n    \"\"\"\n    Count the number of samples.\n    Each of the TFRecord datasets is designed to contain 28000 samples.\n    \"\"\"\n    return len(fileids) * 28000\n\n\nAUTO = tf.data.experimental.AUTOTUNE\n\n\ndef prepare_wave(wave):\n    wave = tf.reshape(tf.io.decode_raw(wave, tf.float64), (3, 4096))\n    normalized_waves = []\n    for i in range(3):\n        normalized_wave = wave[i] / tf.math.reduce_max(wave[i])\n        normalized_waves.append(normalized_wave)\n    wave = tf.stack(normalized_waves, axis=0)\n    wave = tf.cast(wave, tf.float32)\n    return wave\n\n\ndef read_labeled_tfrecord(example):\n    tfrec_format = {\n        \"wave\": tf.io.FixedLenFeature([], tf.string),\n        \"wave_id\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    return prepare_wave(example[\"wave\"]), tf.reshape(tf.cast(example[\"target\"], tf.float32), [1]), example[\"wave_id\"]\n\n\ndef read_unlabeled_tfrecord(example, return_image_id):\n    tfrec_format = {\n        \"wave\": tf.io.FixedLenFeature([], tf.string),\n        \"wave_id\": tf.io.FixedLenFeature([], tf.string)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    return prepare_wave(example[\"wave\"]), example[\"wave_id\"] if return_image_id else 0\n\n\ndef get_dataset(files, batch_size=16, repeat=False, cache=False, shuffle=False, labeled=True, return_image_ids=True):\n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=AUTO, compression_type=\"GZIP\")\n    if cache:\n        # You'll need around 15GB RAM if you'd like to cache val dataset, and 50~60GB RAM for train dataset.\n        ds = ds.cache()\n\n    if repeat:\n        ds = ds.repeat()\n\n    if shuffle:\n        ds = ds.shuffle(1024 * 2)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n\n    if labeled:\n        ds = ds.map(read_labeled_tfrecord, num_parallel_calls=AUTO)\n    else:\n        ds = ds.map(lambda example: read_unlabeled_tfrecord(example, return_image_ids), num_parallel_calls=AUTO)\n\n    ds = ds.batch(batch_size)\n    ds = ds.prefetch(AUTO)\n    return tfds.as_numpy(ds)\n</code></pre>\n<p>and then define a DataLoader class which takes the role of <code>torch.utils.data.DataLoader</code></p>\n<pre><code>class TFRecordDataLoader:\n    def __init__(self, files, batch_size=32, cache=False, train=True, repeat=False, shuffle=False, labeled=True, return_image_ids=True):\n        self.ds = get_dataset(\n            files, \n            batch_size=batch_size,\n            cache=cache,\n            repeat=repeat,\n            shuffle=shuffle,\n            labeled=labeled,\n            return_image_ids=return_image_ids)\n\n        if train:\n            self.num_examples = count_data_items(files)\n        else:\n            self.num_examples = count_data_items_test(files)\n\n        self.batch_size = batch_size\n        self.labeled = labeled\n        self.return_image_ids = return_image_ids\n        self._iterator = None\n\n    def __iter__(self):\n        if self._iterator is None:\n            self._iterator = iter(self.ds)\n        else:\n            self._reset()\n        return self._iterator\n\n    def _reset(self):\n        self._iterator = iter(self.ds)\n\n    def __next__(self):\n        batch = next(self._iterator)\n        return batch\n\n    def __len__(self):\n        n_batches = self.num_examples // self.batch_size\n        if self.num_examples % self.batch_size == 0:\n            return n_batches\n        else:\n            return n_batches + 1\n</code></pre>\n<p>Finally, you need to use custom DataLoader instead of using pytorch's DataLoader.</p>\n<p>Here's an example Notebook.</p>\n<p><a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch</a></p>\n<h2>Who will benefit from this?</h2>\n<p>This technique is for PyTorch users. Especially who:</p>\n<ul>\n<li>use Colab or other cloud platform to train models.<ul>\n<li>With this technique, you don't need to worry about the storage of the data.</li></ul></li>\n<li>want to speed up training<ul>\n<li>If you have RAM more than 15GB you can cache val dataset which speed up the training even more. If you have RAM more than 50 ~ 60GB you can cache the whole dataset.</li></ul></li>\n</ul>",
      "rawMarkdown": "Hi kagglers,\n\nIn this post, I'll share a way to read data from tfrecords while using PyTorch framework to train your model. There are two advantage of doing this:\n\n1. You can speed up your training.\n2. You don't need to download dataset. This is quite good for those who use Colab (Pro/Pro+) to train their models.\n\n## First step - get the TFRecord paths\n\nTo read TFRecord data, you only need to know GCS path of the TFRecord. If you use Kaggle Datasets, you can get GCS path of the Kaggle Dataset as follows:\n\n```python\nfrom kaggle_datasets import KaggleDatasets\n\n\ngcs_path = KaggleDatasets().get_gcs_path(\"<name of your dataset>\")\n```\n\nThis GCS path changes in around 1 week. If it changes you need to rerun this code. Note that it would take long time to get GCS path for the first time, and you may get Timeout error. In that case you only need to rerun the code until you get the path.\n\n## Second step - define custom Data Loader\n\nNext, you need to define tensorflow dataset. Here's an example of a tensorflow dataset that reads waveform from my dataset.\n\n```python\nimport tensorflow as tf\nimport tensorflow_datasets as tfds\n\n\ndef count_data_items(fileids):\n    \"\"\"\n    Count the number of samples.\n    Each of the TFRecord datasets is designed to contain 28000 samples.\n    \"\"\"\n    return len(fileids) * 28000\n\n\nAUTO = tf.data.experimental.AUTOTUNE\n\n\ndef prepare_wave(wave):\n    wave = tf.reshape(tf.io.decode_raw(wave, tf.float64), (3, 4096))\n    normalized_waves = []\n    for i in range(3):\n        normalized_wave = wave[i] / tf.math.reduce_max(wave[i])\n        normalized_waves.append(normalized_wave)\n    wave = tf.stack(normalized_waves, axis=0)\n    wave = tf.cast(wave, tf.float32)\n    return wave\n\n\ndef read_labeled_tfrecord(example):\n    tfrec_format = {\n        \"wave\": tf.io.FixedLenFeature([], tf.string),\n        \"wave_id\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    return prepare_wave(example[\"wave\"]), tf.reshape(tf.cast(example[\"target\"], tf.float32), [1]), example[\"wave_id\"]\n\n\ndef read_unlabeled_tfrecord(example, return_image_id):\n    tfrec_format = {\n        \"wave\": tf.io.FixedLenFeature([], tf.string),\n        \"wave_id\": tf.io.FixedLenFeature([], tf.string)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    return prepare_wave(example[\"wave\"]), example[\"wave_id\"] if return_image_id else 0\n\n\ndef get_dataset(files, batch_size=16, repeat=False, cache=False, shuffle=False, labeled=True, return_image_ids=True):\n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=AUTO, compression_type=\"GZIP\")\n    if cache:\n        # You'll need around 15GB RAM if you'd like to cache val dataset, and 50~60GB RAM for train dataset.\n        ds = ds.cache()\n\n    if repeat:\n        ds = ds.repeat()\n\n    if shuffle:\n        ds = ds.shuffle(1024 * 2)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n\n    if labeled:\n        ds = ds.map(read_labeled_tfrecord, num_parallel_calls=AUTO)\n    else:\n        ds = ds.map(lambda example: read_unlabeled_tfrecord(example, return_image_ids), num_parallel_calls=AUTO)\n\n    ds = ds.batch(batch_size)\n    ds = ds.prefetch(AUTO)\n    return tfds.as_numpy(ds)\n```\n\nand then define a DataLoader class which takes the role of `torch.utils.data.DataLoader`\n\n```python\nclass TFRecordDataLoader:\n    def __init__(self, files, batch_size=32, cache=False, train=True, repeat=False, shuffle=False, labeled=True, return_image_ids=True):\n        self.ds = get_dataset(\n            files, \n            batch_size=batch_size,\n            cache=cache,\n            repeat=repeat,\n            shuffle=shuffle,\n            labeled=labeled,\n            return_image_ids=return_image_ids)\n        \n        if train:\n            self.num_examples = count_data_items(files)\n        else:\n            self.num_examples = count_data_items_test(files)\n\n        self.batch_size = batch_size\n        self.labeled = labeled\n        self.return_image_ids = return_image_ids\n        self._iterator = None\n    \n    def __iter__(self):\n        if self._iterator is None:\n            self._iterator = iter(self.ds)\n        else:\n            self._reset()\n        return self._iterator\n\n    def _reset(self):\n        self._iterator = iter(self.ds)\n\n    def __next__(self):\n        batch = next(self._iterator)\n        return batch\n\n    def __len__(self):\n        n_batches = self.num_examples // self.batch_size\n        if self.num_examples % self.batch_size == 0:\n            return n_batches\n        else:\n            return n_batches + 1\n```\n\nFinally, you need to use custom DataLoader instead of using pytorch's DataLoader.\n\nHere's an example Notebook.\n\nhttps://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\n\n## Who will benefit from this?\n\nThis technique is for PyTorch users. Especially who:\n\n* use Colab or other cloud platform to train models.\n  - With this technique, you don't need to worry about the storage of the data.\n* want to speed up training\n  - If you have RAM more than 15GB you can cache val dataset which speed up the training even more. If you have RAM more than 50 ~ 60GB you can cache the whole dataset.",
      "votes": null
    },
    {
      "id": "1485829",
      "postDate": "08/22/2021 13:03:00",
      "content": "<p>…or you can use this <a href=\"https://github.com/vahidk/tfrecord\" target=\"_blank\">https://github.com/vahidk/tfrecord</a></p>",
      "rawMarkdown": "...or you can use this https://github.com/vahidk/tfrecord",
      "votes": null
    },
    {
      "id": "1485847",
      "postDate": "08/22/2021 13:18:48",
      "content": "<p>I recommend you to read my Notebook. I've written about that. That library is not optimized, and slower than PyTorch's usual Dataset and DataLoader.</p>",
      "rawMarkdown": "I recommend you to read my Notebook. I've written about that. That library is not optimized, and slower than PyTorch's usual Dataset and DataLoader.",
      "votes": null
    },
    {
      "id": "1485864",
      "postDate": "08/22/2021 13:26:40",
      "content": "<p>I waited for this functionality for many days :) It's Awesome, will test and give you my feedback <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a></p>\n<p>gcs_path = KaggleDatasets().get_gcs_path(\"\"), I used this functionality as you explained and working fine</p>\n<p><strong>cache</strong> It's smooth and working so well. Need to find a workaround to use split the data based on available ram :). Thanks Again, u save a lot of copy stuff and stay with Pytorch :D</p>",
      "rawMarkdown": "I waited for this functionality for many days :) It's Awesome, will test and give you my feedback @hidehisaarai1213\n\ngcs_path = KaggleDatasets().get_gcs_path(\"<name of your dataset>\"), I used this functionality as you explained and working fine\n\n**cache** It's smooth and working so well. Need to find a workaround to use split the data based on available ram :). Thanks Again, u save a lot of copy stuff and stay with Pytorch :D",
      "votes": null
    },
    {
      "id": "1486621",
      "postDate": "08/23/2021 05:34:57",
      "content": "<p>Hi , thanks for sharing… Could you share how you created the TFRecords. Is bandpass filter already applied ?</p>",
      "rawMarkdown": "Hi , thanks for sharing... Could you share how you created the TFRecords. Is bandpass filter already applied ?",
      "votes": null
    },
    {
      "id": "1488139",
      "postDate": "08/24/2021 06:02:05",
      "content": "<p>I think its the raw np arrays as TF Recs. Any processing you can apply on top of these.</p>",
      "rawMarkdown": "I think its the raw np arrays as TF Recs. Any processing you can apply on top of these.",
      "votes": null
    },
    {
      "id": "1488167",
      "postDate": "08/24/2021 06:19:45",
      "content": "<p>Yes.. I think so too… Just saw the <a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords\" target=\"_blank\">notebooks</a> that were used to generate these … Thanks</p>",
      "rawMarkdown": "Yes.. I think so too... Just saw the [notebooks](https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords) that were used to generate these ... Thanks",
      "votes": null
    },
    {
      "id": "1488205",
      "postDate": "08/24/2021 06:53:51",
      "content": "<p>Thank you for sharing.  Can I ask about this:</p>\n<blockquote>\n  <p>If you have RAM more than 15GB you can cache val dataset which speed up the training even more. If you have RAM more than 50 ~ 60GB you can cache the whole dataset.</p>\n</blockquote>\n<p>What does it mean to cache val dataset ? Is there anything extra that we need to do before calling the train_loop? </p>",
      "rawMarkdown": "Thank you for sharing.  Can I ask about this:\n\n> If you have RAM more than 15GB you can cache val dataset which speed up the training even more. If you have RAM more than 50 ~ 60GB you can cache the whole dataset.\n\nWhat does it mean to cache val dataset ? Is there anything extra that we need to do before calling the train_loop?",
      "votes": null
    },
    {
      "id": "1488618",
      "postDate": "08/24/2021 12:04:08",
      "content": "<p>Yes, tf.data.Dataset can cache the data<br>\nPlease check out the detail on the docs.<br>\n<a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#cache\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/data/Dataset#cache</a></p>",
      "rawMarkdown": "Yes, tf.data.Dataset can cache the data\nPlease check out the detail on the docs.\nhttps://www.tensorflow.org/api_docs/python/tf/data/Dataset#cache",
      "votes": null
    },
    {
      "id": "1488822",
      "postDate": "08/24/2021 14:14:39",
      "content": "<p>Very nice work! Thanks for sharing! Your idea seems to read the tfrecs through <code>tf.data.TFRecordDataset</code> , and the class TFRecordDataLoader is just like a Iterator outside the TFRecordDataset, I hope I do't misunderstand your idea~<br>\nOne question is that in this task I know the lenght or the total number of the data items, what if the number is unkown? this seems important to define the dataloader.<br>\nThanks again for such a great job!</p>",
      "rawMarkdown": "Very nice work! Thanks for sharing! Your idea seems to read the tfrecs through `tf.data.TFRecordDataset` , and the class TFRecordDataLoader is just like a Iterator outside the TFRecordDataset, I hope I do't misunderstand your idea~\nOne question is that in this task I know the lenght or the total number of the data items, what if the number is unkown? this seems important to define the dataloader.\nThanks again for such a great job!",
      "votes": null
    },
    {
      "id": "1489143",
      "postDate": "08/24/2021 18:28:50",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> . This speeds up my pipeline by quite a bit, although I can only cache the val dataset for now. I do have 1 question, are the tf recs stratified on labels? If not, how would I go about doing it in your pipeline.</p>",
      "rawMarkdown": "Great work @hidehisaarai1213 . This speeds up my pipeline by quite a bit, although I can only cache the val dataset for now. I do have 1 question, are the tf recs stratified on labels? If not, how would I go about doing it in your pipeline.",
      "votes": null
    },
    {
      "id": "1489507",
      "postDate": "08/25/2021 03:56:47",
      "content": "<p>Thank you for the reply. I will check it.</p>",
      "rawMarkdown": "Thank you for the reply. I will check it.",
      "votes": null
    },
    {
      "id": "1489571",
      "postDate": "08/25/2021 05:26:17",
      "content": "<p>I think tf records are not stratified. If you see the kernel <a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords</a> which was used then it is simply a generation of TF records. So you can generate tf records gain with stratified data frame and then use Kfold to train the models. </p>",
      "rawMarkdown": "I think tf records are not stratified. If you see the kernel https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords which was used then it is simply a generation of TF records. So you can generate tf records gain with stratified data frame and then use Kfold to train the models.",
      "votes": null
    },
    {
      "id": "1489828",
      "postDate": "08/25/2021 10:05:10",
      "content": "<p>I don't think we need to create new tf recs, perhaps some modifications to the dataloader should do it.</p>",
      "rawMarkdown": "I don't think we need to create new tf recs, perhaps some modifications to the dataloader should do it.",
      "votes": null
    },
    {
      "id": "1489847",
      "postDate": "08/25/2021 10:23:43",
      "content": "<p>yeah, that might be possible. not sure as I have not tried this set-up yet. Would it be possible for you to tell me how fast your pipeline is now after using this data loader? I am curious to know the improvement coming from a combination of two frameworks. </p>",
      "rawMarkdown": "yeah, that might be possible. not sure as I have not tried this set-up yet. Would it be possible for you to tell me how fast your pipeline is now after using this data loader? I am curious to know the improvement coming from a combination of two frameworks.",
      "votes": null
    },
    {
      "id": "1503331",
      "postDate": "09/05/2021 09:11:59",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> for this idea.<br>\n<a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a> here are the speed-ups I got with this new pipeline</p>\n<p>Initial pipeline -&gt; Current Tfrecords pipeline</p>\n<ul>\n<li>Train (per epoch) =&gt; (2hrs → 30min)</li>\n<li>Eval (per epoch) =&gt; (30mins → 5min)</li>\n<li>Inference =&gt; (30min → 10min)</li>\n<li>Per experiment time =&gt; (8hrs → 2hrs 15min)</li>\n</ul>\n<p>All the above results are with <code>efficientnet_b0</code> and Tesla P100 GPU</p>\n<p>I want to take this pipeline to TPU, but not sure how to implement CQT transforms on TPU (with PyTorch)</p>",
      "rawMarkdown": "Thanks @hidehisaarai1213 for this idea.\n@urvishp80 here are the speed-ups I got with this new pipeline\n\nInitial pipeline -> Current Tfrecords pipeline\n- Train (per epoch) => (2hrs → 30min)\n- Eval (per epoch) => (30mins → 5min)\n- Inference => (30min → 10min)\n- Per experiment time => (8hrs → 2hrs 15min)\n\nAll the above results are with `efficientnet_b0` and Tesla P100 GPU\n\nI want to take this pipeline to TPU, but not sure how to implement CQT transforms on TPU (with PyTorch)",
      "votes": null
    },
    {
      "id": "1504097",
      "postDate": "09/06/2021 04:53:47",
      "content": "<p>Thanks for letting me know how fast your pipeline is now due to this setup <a href=\"https://www.kaggle.com/atharvaingle\" target=\"_blank\">@atharvaingle</a> I am definitely going to try this now.</p>",
      "rawMarkdown": "Thanks for letting me know how fast your pipeline is now due to this setup @atharvaingle I am definitely going to try this now.",
      "votes": null
    },
    {
      "id": "1559959",
      "postDate": "10/27/2021 08:38:39",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1485829,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "08/22/2021 13:03:00",
      "content": "<p>…or you can use this <a href=\"https://github.com/vahidk/tfrecord\" target=\"_blank\">https://github.com/vahidk/tfrecord</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1485847,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/22/2021 13:18:48",
          "content": "<p>I recommend you to read my Notebook. I've written about that. That library is not optimized, and slower than PyTorch's usual Dataset and DataLoader.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1485864,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "08/22/2021 13:26:40",
          "content": "<p>I waited for this functionality for many days :) It's Awesome, will test and give you my feedback <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a></p>\n<p>gcs_path = KaggleDatasets().get_gcs_path(\"\"), I used this functionality as you explained and working fine</p>\n<p><strong>cache</strong> It's smooth and working so well. Need to find a workaround to use split the data based on available ram :). Thanks Again, u save a lot of copy stuff and stay with Pytorch :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1486621,
      "author_name": "benihime91",
      "author_url": "",
      "post_date": "08/23/2021 05:34:57",
      "content": "<p>Hi , thanks for sharing… Could you share how you created the TFRecords. Is bandpass filter already applied ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1488139,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "08/24/2021 06:02:05",
          "content": "<p>I think its the raw np arrays as TF Recs. Any processing you can apply on top of these.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1488167,
          "author_name": "benihime91",
          "author_url": "",
          "post_date": "08/24/2021 06:19:45",
          "content": "<p>Yes.. I think so too… Just saw the <a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords\" target=\"_blank\">notebooks</a> that were used to generate these … Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1488205,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "08/24/2021 06:53:51",
      "content": "<p>Thank you for sharing.  Can I ask about this:</p>\n<blockquote>\n  <p>If you have RAM more than 15GB you can cache val dataset which speed up the training even more. If you have RAM more than 50 ~ 60GB you can cache the whole dataset.</p>\n</blockquote>\n<p>What does it mean to cache val dataset ? Is there anything extra that we need to do before calling the train_loop? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1488618,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/24/2021 12:04:08",
          "content": "<p>Yes, tf.data.Dataset can cache the data<br>\nPlease check out the detail on the docs.<br>\n<a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#cache\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/data/Dataset#cache</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1489507,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "08/25/2021 03:56:47",
          "content": "<p>Thank you for the reply. I will check it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1488822,
      "author_name": "cittasj",
      "author_url": "",
      "post_date": "08/24/2021 14:14:39",
      "content": "<p>Very nice work! Thanks for sharing! Your idea seems to read the tfrecs through <code>tf.data.TFRecordDataset</code> , and the class TFRecordDataLoader is just like a Iterator outside the TFRecordDataset, I hope I do't misunderstand your idea~<br>\nOne question is that in this task I know the lenght or the total number of the data items, what if the number is unkown? this seems important to define the dataloader.<br>\nThanks again for such a great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1489143,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "08/24/2021 18:28:50",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> . This speeds up my pipeline by quite a bit, although I can only cache the val dataset for now. I do have 1 question, are the tf recs stratified on labels? If not, how would I go about doing it in your pipeline.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1489571,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "08/25/2021 05:26:17",
          "content": "<p>I think tf records are not stratified. If you see the kernel <a href=\"https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords</a> which was used then it is simply a generation of TF records. So you can generate tf records gain with stratified data frame and then use Kfold to train the models. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1489828,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "08/25/2021 10:05:10",
          "content": "<p>I don't think we need to create new tf recs, perhaps some modifications to the dataloader should do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1489847,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "08/25/2021 10:23:43",
          "content": "<p>yeah, that might be possible. not sure as I have not tried this set-up yet. Would it be possible for you to tell me how fast your pipeline is now after using this data loader? I am curious to know the improvement coming from a combination of two frameworks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1503331,
          "author_name": "atharvaingle",
          "author_url": "",
          "post_date": "09/05/2021 09:11:59",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> for this idea.<br>\n<a href=\"https://www.kaggle.com/urvishp80\" target=\"_blank\">@urvishp80</a> here are the speed-ups I got with this new pipeline</p>\n<p>Initial pipeline -&gt; Current Tfrecords pipeline</p>\n<ul>\n<li>Train (per epoch) =&gt; (2hrs → 30min)</li>\n<li>Eval (per epoch) =&gt; (30mins → 5min)</li>\n<li>Inference =&gt; (30min → 10min)</li>\n<li>Per experiment time =&gt; (8hrs → 2hrs 15min)</li>\n</ul>\n<p>All the above results are with <code>efficientnet_b0</code> and Tesla P100 GPU</p>\n<p>I want to take this pipeline to TPU, but not sure how to implement CQT transforms on TPU (with PyTorch)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1504097,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "09/06/2021 04:53:47",
          "content": "<p>Thanks for letting me know how fast your pipeline is now due to this setup <a href=\"https://www.kaggle.com/atharvaingle\" target=\"_blank\">@atharvaingle</a> I am definitely going to try this now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559959,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:38:39",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1485758": "Hi kagglers,\n\nIn this post, I'll share a way to read data from tfrecords while using PyTorch framework to train your model. There are two advantage of doing this:\n\n1. You can speed up your training.\n2. You don't need to download dataset. This is quite good for those who use Colab (Pro/Pro+) to train their models.\n\n## First step - get the TFRecord paths\n\nTo read TFRecord data, you only need to know GCS path of the TFRecord. If you use Kaggle Datasets, you can get GCS path of the Kaggle Dataset as follows:\n\n```python\nfrom kaggle_datasets import KaggleDatasets\n\n\ngcs_path = KaggleDatasets().get_gcs_path(\"<name of your dataset>\")\n```\n\nThis GCS path changes in around 1 week. If it changes you need to rerun this code. Note that it would take long time to get GCS path for the first time, and you may get Timeout error. In that case you only need to rerun the code until you get the path.\n\n## Second step - define custom Data Loader\n\nNext, you need to define tensorflow dataset. Here's an example of a tensorflow dataset that reads waveform from my dataset.\n\n```python\nimport tensorflow as tf\nimport tensorflow_datasets as tfds\n\n\ndef count_data_items(fileids):\n    \"\"\"\n    Count the number of samples.\n    Each of the TFRecord datasets is designed to contain 28000 samples.\n    \"\"\"\n    return len(fileids) * 28000\n\n\nAUTO = tf.data.experimental.AUTOTUNE\n\n\ndef prepare_wave(wave):\n    wave = tf.reshape(tf.io.decode_raw(wave, tf.float64), (3, 4096))\n    normalized_waves = []\n    for i in range(3):\n        normalized_wave = wave[i] / tf.math.reduce_max(wave[i])\n        normalized_waves.append(normalized_wave)\n    wave = tf.stack(normalized_waves, axis=0)\n    wave = tf.cast(wave, tf.float32)\n    return wave\n\n\ndef read_labeled_tfrecord(example):\n    tfrec_format = {\n        \"wave\": tf.io.FixedLenFeature([], tf.string),\n        \"wave_id\": tf.io.FixedLenFeature([], tf.string),\n        \"target\": tf.io.FixedLenFeature([], tf.int64)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    return prepare_wave(example[\"wave\"]), tf.reshape(tf.cast(example[\"target\"], tf.float32), [1]), example[\"wave_id\"]\n\n\ndef read_unlabeled_tfrecord(example, return_image_id):\n    tfrec_format = {\n        \"wave\": tf.io.FixedLenFeature([], tf.string),\n        \"wave_id\": tf.io.FixedLenFeature([], tf.string)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    return prepare_wave(example[\"wave\"]), example[\"wave_id\"] if return_image_id else 0\n\n\ndef get_dataset(files, batch_size=16, repeat=False, cache=False, shuffle=False, labeled=True, return_image_ids=True):\n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=AUTO, compression_type=\"GZIP\")\n    if cache:\n        # You'll need around 15GB RAM if you'd like to cache val dataset, and 50~60GB RAM for train dataset.\n        ds = ds.cache()\n\n    if repeat:\n        ds = ds.repeat()\n\n    if shuffle:\n        ds = ds.shuffle(1024 * 2)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n\n    if labeled:\n        ds = ds.map(read_labeled_tfrecord, num_parallel_calls=AUTO)\n    else:\n        ds = ds.map(lambda example: read_unlabeled_tfrecord(example, return_image_ids), num_parallel_calls=AUTO)\n\n    ds = ds.batch(batch_size)\n    ds = ds.prefetch(AUTO)\n    return tfds.as_numpy(ds)\n```\n\nand then define a DataLoader class which takes the role of `torch.utils.data.DataLoader`\n\n```python\nclass TFRecordDataLoader:\n    def __init__(self, files, batch_size=32, cache=False, train=True, repeat=False, shuffle=False, labeled=True, return_image_ids=True):\n        self.ds = get_dataset(\n            files, \n            batch_size=batch_size,\n            cache=cache,\n            repeat=repeat,\n            shuffle=shuffle,\n            labeled=labeled,\n            return_image_ids=return_image_ids)\n        \n        if train:\n            self.num_examples = count_data_items(files)\n        else:\n            self.num_examples = count_data_items_test(files)\n\n        self.batch_size = batch_size\n        self.labeled = labeled\n        self.return_image_ids = return_image_ids\n        self._iterator = None\n    \n    def __iter__(self):\n        if self._iterator is None:\n            self._iterator = iter(self.ds)\n        else:\n            self._reset()\n        return self._iterator\n\n    def _reset(self):\n        self._iterator = iter(self.ds)\n\n    def __next__(self):\n        batch = next(self._iterator)\n        return batch\n\n    def __len__(self):\n        n_batches = self.num_examples // self.batch_size\n        if self.num_examples % self.batch_size == 0:\n            return n_batches\n        else:\n            return n_batches + 1\n```\n\nFinally, you need to use custom DataLoader instead of using pytorch's DataLoader.\n\nHere's an example Notebook.\n\nhttps://www.kaggle.com/hidehisaarai1213/g2net-read-from-tfrecord-train-with-pytorch\n\n## Who will benefit from this?\n\nThis technique is for PyTorch users. Especially who:\n\n* use Colab or other cloud platform to train models.\n  - With this technique, you don't need to worry about the storage of the data.\n* want to speed up training\n  - If you have RAM more than 15GB you can cache val dataset which speed up the training even more. If you have RAM more than 50 ~ 60GB you can cache the whole dataset.",
    "1485829": "...or you can use this https://github.com/vahidk/tfrecord",
    "1485847": "I recommend you to read my Notebook. I've written about that. That library is not optimized, and slower than PyTorch's usual Dataset and DataLoader.",
    "1485864": "I waited for this functionality for many days :) It's Awesome, will test and give you my feedback @hidehisaarai1213\n\ngcs_path = KaggleDatasets().get_gcs_path(\"<name of your dataset>\"), I used this functionality as you explained and working fine\n\n**cache** It's smooth and working so well. Need to find a workaround to use split the data based on available ram :). Thanks Again, u save a lot of copy stuff and stay with Pytorch :D",
    "1486621": "Hi , thanks for sharing... Could you share how you created the TFRecords. Is bandpass filter already applied ?",
    "1488139": "I think its the raw np arrays as TF Recs. Any processing you can apply on top of these.",
    "1488167": "Yes.. I think so too... Just saw the [notebooks](https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords) that were used to generate these ... Thanks",
    "1488205": "Thank you for sharing.  Can I ask about this:\n\n> If you have RAM more than 15GB you can cache val dataset which speed up the training even more. If you have RAM more than 50 ~ 60GB you can cache the whole dataset.\n\nWhat does it mean to cache val dataset ? Is there anything extra that we need to do before calling the train_loop?",
    "1488618": "Yes, tf.data.Dataset can cache the data\nPlease check out the detail on the docs.\nhttps://www.tensorflow.org/api_docs/python/tf/data/Dataset#cache",
    "1488822": "Very nice work! Thanks for sharing! Your idea seems to read the tfrecs through `tf.data.TFRecordDataset` , and the class TFRecordDataLoader is just like a Iterator outside the TFRecordDataset, I hope I do't misunderstand your idea~\nOne question is that in this task I know the lenght or the total number of the data items, what if the number is unkown? this seems important to define the dataloader.\nThanks again for such a great job!",
    "1489143": "Great work @hidehisaarai1213 . This speeds up my pipeline by quite a bit, although I can only cache the val dataset for now. I do have 1 question, are the tf recs stratified on labels? If not, how would I go about doing it in your pipeline.",
    "1489507": "Thank you for the reply. I will check it.",
    "1489571": "I think tf records are not stratified. If you see the kernel https://www.kaggle.com/hidehisaarai1213/g2net-waveform-tfrecords which was used then it is simply a generation of TF records. So you can generate tf records gain with stratified data frame and then use Kfold to train the models.",
    "1489828": "I don't think we need to create new tf recs, perhaps some modifications to the dataloader should do it.",
    "1489847": "yeah, that might be possible. not sure as I have not tried this set-up yet. Would it be possible for you to tell me how fast your pipeline is now after using this data loader? I am curious to know the improvement coming from a combination of two frameworks.",
    "1503331": "Thanks @hidehisaarai1213 for this idea.\n@urvishp80 here are the speed-ups I got with this new pipeline\n\nInitial pipeline -> Current Tfrecords pipeline\n- Train (per epoch) => (2hrs → 30min)\n- Eval (per epoch) => (30mins → 5min)\n- Inference => (30min → 10min)\n- Per experiment time => (8hrs → 2hrs 15min)\n\nAll the above results are with `efficientnet_b0` and Tesla P100 GPU\n\nI want to take this pipeline to TPU, but not sure how to implement CQT transforms on TPU (with PyTorch)",
    "1504097": "Thanks for letting me know how fast your pipeline is now due to this setup @atharvaingle I am definitely going to try this now.",
    "1559959": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}