{
  "id": 390503,
  "title": "Efficient methods for loading training data",
  "url": "/competitions/asl-signs/discussion/390503",
  "author_name": "aapo kossi",
  "post_date": "2023-02-25T22:35:50.625000",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I'm wondering what efficient method people are using for loading the training data, since I'm having trouble with my first option, which is parallelized loading by using tfio.IODataset.from_parquet for each sample. I am now working with tf.py_function and reading with pandas, but that is slow. I know I could simply save the dataset after waiting for it to load once, but a faster method would be nice.</p>\n<p>Trying to load a parquet file into a dataset with tfio silently crashes the kernel. The error seems to be the following:<br>\n<code>terminate called after throwing an instance of 'parquet::ParquetException'</code><br>\n<code>what(): Unexpected end of stream</code></p>",
  "messages": [
    {
      "id": 2159642,
      "postDate": "2023-02-25T22:35:50.627Z",
      "content": "<p>I'm wondering what efficient method people are using for loading the training data, since I'm having trouble with my first option, which is parallelized loading by using tfio.IODataset.from_parquet for each sample. I am now working with tf.py_function and reading with pandas, but that is slow. I know I could simply save the dataset after waiting for it to load once, but a faster method would be nice.</p>\n<p>Trying to load a parquet file into a dataset with tfio silently crashes the kernel. The error seems to be the following:<br>\n<code>terminate called after throwing an instance of 'parquet::ParquetException'</code><br>\n<code>what(): Unexpected end of stream</code></p>",
      "rawMarkdown": "I'm wondering what efficient method people are using for loading the training data, since I'm having trouble with my first option, which is parallelized loading by using tfio.IODataset.from_parquet for each sample. I am now working with tf.py_function and reading with pandas, but that is slow. I know I could simply save the dataset after waiting for it to load once, but a faster method would be nice.\n\nTrying to load a parquet file into a dataset with tfio silently crashes the kernel. The error seems to be the following:\n`terminate called after throwing an instance of 'parquet::ParquetException'`\n`what(): Unexpected end of stream`",
      "votes": 4
    },
    {
      "id": 2163022,
      "postDate": "2023-02-28T14:39:24.657Z",
      "content": "<p>I have now published a <a href=\"https://www.kaggle.com/datasets/aapokossi/saved-tfdataset-of-google-isl-recognition-data\" target=\"_blank\">tf.Dataset</a> that quite significantly speeds up loading for me (700 it/sec vs. 60 it/sec originally, with no batching), along with <a href=\"https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-tf-dataset\" target=\"_blank\">a notebook</a> on how to easily save the data yourself, which should make it easier for anyone to apply their own desired level of preprocessing before saving, if you want even more performance. I myself will probably be using a variation that is padded  to a specific sequence length and pre-batched. </p>",
      "rawMarkdown": "I have now published a [tf.Dataset](https://www.kaggle.com/datasets/aapokossi/saved-tfdataset-of-google-isl-recognition-data) that quite significantly speeds up loading for me (700 it/sec vs. 60 it/sec originally, with no batching), along with [a notebook](https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-tf-dataset) on how to easily save the data yourself, which should make it easier for anyone to apply their own desired level of preprocessing before saving, if you want even more performance. I myself will probably be using a variation that is padded  to a specific sequence length and pre-batched. ",
      "votes": 2
    },
    {
      "id": 2162219,
      "postDate": "2023-02-28T04:10:16.230Z",
      "content": "<p>I'm quite new to TF and Kaggle competitions, but taking reference some of the codes shared (shoutout to <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a>), I believe the competition has specified a method they will use to load the data, available in the Evaluation page. I'm guessing we can adopt the same method to load the data, then incorporate any further processing into your model pipeline.</p>\n<pre><code> ():\n    data_columns = [, , ]\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = ((data) / ROWS_PER_FRAME)\n    data = data.values.reshape(n_frames, ROWS_PER_FRAME, (data_columns))\n     data.astype(np.float32)\n</code></pre>",
      "rawMarkdown": "I'm quite new to TF and Kaggle competitions, but taking reference some of the codes shared (shoutout to @lonnieqin), I believe the competition has specified a method they will use to load the data, available in the Evaluation page. I'm guessing we can adopt the same method to load the data, then incorporate any further processing into your model pipeline.\n\n```python\ndef load_relevant_data_subset(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = int(len(data) / ROWS_PER_FRAME)\n    data = data.values.reshape(n_frames, ROWS_PER_FRAME, len(data_columns))\n    return data.astype(np.float32)\n```",
      "replies": [
        {
          "id": 2162401,
          "postDate": "2023-02-28T07:21:43.663Z",
          "content": "<p>This only loads a single sample from disk. What I'm asking for is how to incorporate this into a Tf Dataset object to automatically load batches from disk during training. There is standard code for standard data formats (e.g. images), but writing this from scratch requires in-depth understanding of Tensorflow.</p>\n<p>Based on the leaderboard, I's assuming that at least some people must have managed to solve this …</p>",
          "rawMarkdown": "This only loads a single sample from disk. What I'm asking for is how to incorporate this into a Tf Dataset object to automatically load batches from disk during training. There is standard code for standard data formats (e.g. images), but writing this from scratch requires in-depth understanding of Tensorflow.\n\nBased on the leaderboard, I's assuming that at least some people must have managed to solve this ..."
        }
      ]
    },
    {
      "id": 2161922,
      "postDate": "2023-02-27T20:36:11.827Z",
      "content": "<p>I have the same question. I've been browsing through multiple explanations about creating a Tf Dataset from parquet files, but this is way above my level of understanding. When I try pre-padding the data for training as a  time-series model, the data becomes WAY too large, e.g. window size of 30 -&gt;  94477x30x543x3 numbers</p>\n<p>Actually, since this is really essential for proper training here, and since documentation for this is really lacking, I feel the Tensorflow development team should provide example code for this. Otherwise, anyone who is not a really advanced Tf software engineer is excluded from properly participating because of getting stuck on something I'm sure the organisers have already solved.</p>\n<p>So please, can someone provide example code for a (train and val) Dataset with batch prefetching that can be used for efficient training, so we can focus on the model design.</p>",
      "rawMarkdown": "I have the same question. I've been browsing through multiple explanations about creating a Tf Dataset from parquet files, but this is way above my level of understanding. When I try pre-padding the data for training as a  time-series model, the data becomes WAY too large, e.g. window size of 30 ->  94477x30x543x3 numbers\n\nActually, since this is really essential for proper training here, and since documentation for this is really lacking, I feel the Tensorflow development team should provide example code for this. Otherwise, anyone who is not a really advanced Tf software engineer is excluded from properly participating because of getting stuck on something I'm sure the organisers have already solved.\n\nSo please, can someone provide example code for a (train and val) Dataset with batch prefetching that can be used for efficient training, so we can focus on the model design.",
      "replies": [
        {
          "id": 2162762,
          "postDate": "2023-02-28T11:58:55.640Z",
          "content": "<p>I'm not sure about your data dimension here. </p>\n<ul>\n<li>a frame is represented by a vector of size (543,3)</li>\n<li>the maximum number of frames in a sequence is 537</li>\n<li>Therefore if you padded to the maximum length, each example should be of shape: <strong>(537,543,3)</strong></li>\n</ul>\n<p>This is indeed pretty big though.</p>",
          "rawMarkdown": "I'm not sure about your data dimension here. \n\n* a frame is represented by a vector of size (543,3)\n* the maximum number of frames in a sequence is 537\n* Therefore if you padded to the maximum length, each example should be of shape: **(537,543,3)**\n\nThis is indeed pretty big though."
        }
      ]
    },
    {
      "id": 2160746,
      "postDate": "2023-02-27T00:12:00.680Z",
      "content": "<p>Haven’t tried it but maybe this:<br>\n<a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/experimental/IODataset#from_parquet\" target=\"_blank\">https://www.tensorflow.org/io/api_docs/python/tfio/experimental/IODataset#from_parquet</a></p>\n<p>It’s in experimental so maybe the issue your having has been fixed?</p>\n<p>I’ll dig more into this tomorrow.</p>",
      "rawMarkdown": "Haven’t tried it but maybe this:\nhttps://www.tensorflow.org/io/api_docs/python/tfio/experimental/IODataset#from_parquet\n\nIt’s in experimental so maybe the issue your having has been fixed?\n\nI’ll dig more into this tomorrow.\n",
      "replies": [
        {
          "id": 2162411,
          "postDate": "2023-02-28T07:24:05.443Z",
          "content": "<p>Eager to hear whether you've made any progress ??</p>",
          "rawMarkdown": "Eager to hear whether you've made any progress ??\n",
          "replies": [
            {
              "id": 2162720,
              "postDate": "2023-02-28T11:32:42.310Z",
              "content": "<p>No, I have the same issue as you where it crashes when I try to load it.</p>\n<p>I am in the process of making TFRecords for efficient data loading. I'll share the notebook shortly and you can use that to create the dataset that reflects what you want.</p>\n<p><strong>Some Random Notes:</strong></p>\n<ul>\n<li>More than likely you need to do dimensionality reduction PRIOR TO (or inline with) dataset creation.</li>\n<li>This is because as you posted, the raw arrays would be too large.</li>\n<li>An alternative would be to scale the x,y,z (or discard the z) to be from 0-255 and the byte encode it as a string to save on space (just like you would an image)</li>\n</ul>\n<p>I'll include everything I can in my notebook.</p>",
              "rawMarkdown": "No, I have the same issue as you where it crashes when I try to load it.\n\nI am in the process of making TFRecords for efficient data loading. I'll share the notebook shortly and you can use that to create the dataset that reflects what you want.\n\n**Some Random Notes:**\n* More than likely you need to do dimensionality reduction PRIOR TO (or inline with) dataset creation.\n* This is because as you posted, the raw arrays would be too large.\n* An alternative would be to scale the x,y,z (or discard the z) to be from 0-255 and the byte encode it as a string to save on space (just like you would an image)\n\nI'll include everything I can in my notebook.",
              "votes": 1
            },
            {
              "id": 2163027,
              "postDate": "2023-02-28T14:43:06.747Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2163029,
              "postDate": "2023-02-28T14:43:43.217Z",
              "content": "<p>I prefer the tf.data aligned Dataset.save compared to TFRecords, but I like that we will now have multiple alternatives! You can see my top level comment for a saved tf.Dataset and notebook on how to save your own if you want to try with tf.Datasets without the intermediate step of TFRecords.</p>",
              "rawMarkdown": "I prefer the tf.data aligned Dataset.save compared to TFRecords, but I like that we will now have multiple alternatives! You can see my top level comment for a saved tf.Dataset and notebook on how to save your own if you want to try with tf.Datasets without the intermediate step of TFRecords."
            },
            {
              "id": 2163421,
              "postDate": "2023-02-28T19:47:38.763Z",
              "content": "<p>Dimensionality reduction: The face landmarks might be total overkill, if removing them or grouping/averaging them, then that gets the dataset down to a more reasonable size. At the risk of removing even a little bit of useful data, though.</p>\n<p>If taking the raw data without padding, it would be around:</p>\n<p>94477<em>40</em>75<em>3</em>[float32 size aka 4 bytes] ~= 3.5 GB.</p>",
              "rawMarkdown": "Dimensionality reduction: The face landmarks might be total overkill, if removing them or grouping/averaging them, then that gets the dataset down to a more reasonable size. At the risk of removing even a little bit of useful data, though.\n\nIf taking the raw data without padding, it would be around:\n\n94477*40*75*3*[float32 size aka 4 bytes] ~= 3.5 GB.",
              "votes": 1
            },
            {
              "id": 2175334,
              "postDate": "2023-03-09T19:31:21.503Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/aapokossi\" target=\"_blank\">@aapokossi</a> </p>\n<p>Thanks a lot for your notebook. I've been using it for a while, but now it seems that some of the operations I want to do do not work on a ragged batch. I'm afraid the Dataset stuff is beyond me (and very poorly documented): do you happen to know how to adapt it to generate padded batches (i.e., each batch padded to the length of the longes sample it it)?</p>\n<p>Also, do I understand correctly that your code pre-forms batches, i.e. the batches would always be the same (but in a random order) during training?</p>",
              "rawMarkdown": "Hi @aapokossi \n\nThanks a lot for your notebook. I've been using it for a while, but now it seems that some of the operations I want to do do not work on a ragged batch. I'm afraid the Dataset stuff is beyond me (and very poorly documented): do you happen to know how to adapt it to generate padded batches (i.e., each batch padded to the length of the longes sample it it)?\n\nAlso, do I understand correctly that your code pre-forms batches, i.e. the batches would always be the same (but in a random order) during training?\n\n\n"
            },
            {
              "id": 2175491,
              "postDate": "2023-03-09T21:57:09.763Z",
              "content": "<p>Yes, you are correct that the batches will always be the same if you don't specifically unbatch and then shuffle them. You can generate padded batches for example with the following transformation:<br>\n<code>ds = ds.map(lambda x, y: (x.to_tensor(), y)</code><br>\nRagged tensors have the to_tensor method which pads ragged dimensions to the longest dimension in the tensor (or optionally some static shape which may truncate the dim size).<br>\nAdditionally, <code>tf.keras.layers.Masking()</code> may or may not be useful when you are using padding, depending on your model architecture and if your layers support masking.</p>",
              "rawMarkdown": "Yes, you are correct that the batches will always be the same if you don't specifically unbatch and then shuffle them. You can generate padded batches for example with the following transformation:\n`ds = ds.map(lambda x, y: (x.to_tensor(), y)`\nRagged tensors have the to_tensor method which pads ragged dimensions to the longest dimension in the tensor (or optionally some static shape which may truncate the dim size).\nAdditionally, `tf.keras.layers.Masking()` may or may not be useful when you are using padding, depending on your model architecture and if your layers support masking.",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2163022,
      "author_name": "aapo kossi",
      "author_url": "",
      "post_date": "2023-02-28T14:39:24.657000",
      "content": "<p>I have now published a <a href=\"https://www.kaggle.com/datasets/aapokossi/saved-tfdataset-of-google-isl-recognition-data\" target=\"_blank\">tf.Dataset</a> that quite significantly speeds up loading for me (700 it/sec vs. 60 it/sec originally, with no batching), along with <a href=\"https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-tf-dataset\" target=\"_blank\">a notebook</a> on how to easily save the data yourself, which should make it easier for anyone to apply their own desired level of preprocessing before saving, if you want even more performance. I myself will probably be using a variation that is padded  to a specific sequence length and pre-batched. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2162219,
      "author_name": "Shawn",
      "author_url": "",
      "post_date": "2023-02-28T04:10:16.230000",
      "content": "<p>I'm quite new to TF and Kaggle competitions, but taking reference some of the codes shared (shoutout to <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a>), I believe the competition has specified a method they will use to load the data, available in the Evaluation page. I'm guessing we can adopt the same method to load the data, then incorporate any further processing into your model pipeline.</p>\n<pre><code> ():\n    data_columns = [, , ]\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = ((data) / ROWS_PER_FRAME)\n    data = data.values.reshape(n_frames, ROWS_PER_FRAME, (data_columns))\n     data.astype(np.float32)\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2162401,
          "author_name": "Wondering Alice",
          "author_url": "",
          "post_date": "2023-02-28T07:21:43.663000",
          "content": "<p>This only loads a single sample from disk. What I'm asking for is how to incorporate this into a Tf Dataset object to automatically load batches from disk during training. There is standard code for standard data formats (e.g. images), but writing this from scratch requires in-depth understanding of Tensorflow.</p>\n<p>Based on the leaderboard, I's assuming that at least some people must have managed to solve this …</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2161922,
      "author_name": "Wondering Alice",
      "author_url": "",
      "post_date": "2023-02-27T20:36:11.827000",
      "content": "<p>I have the same question. I've been browsing through multiple explanations about creating a Tf Dataset from parquet files, but this is way above my level of understanding. When I try pre-padding the data for training as a  time-series model, the data becomes WAY too large, e.g. window size of 30 -&gt;  94477x30x543x3 numbers</p>\n<p>Actually, since this is really essential for proper training here, and since documentation for this is really lacking, I feel the Tensorflow development team should provide example code for this. Otherwise, anyone who is not a really advanced Tf software engineer is excluded from properly participating because of getting stuck on something I'm sure the organisers have already solved.</p>\n<p>So please, can someone provide example code for a (train and val) Dataset with batch prefetching that can be used for efficient training, so we can focus on the model design.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2162762,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2023-02-28T11:58:55.640000",
          "content": "<p>I'm not sure about your data dimension here. </p>\n<ul>\n<li>a frame is represented by a vector of size (543,3)</li>\n<li>the maximum number of frames in a sequence is 537</li>\n<li>Therefore if you padded to the maximum length, each example should be of shape: <strong>(537,543,3)</strong></li>\n</ul>\n<p>This is indeed pretty big though.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2160746,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2023-02-27T00:12:00.680000",
      "content": "<p>Haven’t tried it but maybe this:<br>\n<a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/experimental/IODataset#from_parquet\" target=\"_blank\">https://www.tensorflow.org/io/api_docs/python/tfio/experimental/IODataset#from_parquet</a></p>\n<p>It’s in experimental so maybe the issue your having has been fixed?</p>\n<p>I’ll dig more into this tomorrow.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2162411,
          "author_name": "Wondering Alice",
          "author_url": "",
          "post_date": "2023-02-28T07:24:05.443000",
          "content": "<p>Eager to hear whether you've made any progress ??</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2162720,
              "author_name": "Darien Schettler",
              "author_url": "",
              "post_date": "2023-02-28T11:32:42.310000",
              "content": "<p>No, I have the same issue as you where it crashes when I try to load it.</p>\n<p>I am in the process of making TFRecords for efficient data loading. I'll share the notebook shortly and you can use that to create the dataset that reflects what you want.</p>\n<p><strong>Some Random Notes:</strong></p>\n<ul>\n<li>More than likely you need to do dimensionality reduction PRIOR TO (or inline with) dataset creation.</li>\n<li>This is because as you posted, the raw arrays would be too large.</li>\n<li>An alternative would be to scale the x,y,z (or discard the z) to be from 0-255 and the byte encode it as a string to save on space (just like you would an image)</li>\n</ul>\n<p>I'll include everything I can in my notebook.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2163027,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-02-28T14:43:06.747000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2163029,
              "author_name": "aapo kossi",
              "author_url": "",
              "post_date": "2023-02-28T14:43:43.217000",
              "content": "<p>I prefer the tf.data aligned Dataset.save compared to TFRecords, but I like that we will now have multiple alternatives! You can see my top level comment for a saved tf.Dataset and notebook on how to save your own if you want to try with tf.Datasets without the intermediate step of TFRecords.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2163421,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-28T19:47:38.763000",
              "content": "<p>Dimensionality reduction: The face landmarks might be total overkill, if removing them or grouping/averaging them, then that gets the dataset down to a more reasonable size. At the risk of removing even a little bit of useful data, though.</p>\n<p>If taking the raw data without padding, it would be around:</p>\n<p>94477<em>40</em>75<em>3</em>[float32 size aka 4 bytes] ~= 3.5 GB.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2175334,
              "author_name": "Wondering Alice",
              "author_url": "",
              "post_date": "2023-03-09T19:31:21.503000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/aapokossi\" target=\"_blank\">@aapokossi</a> </p>\n<p>Thanks a lot for your notebook. I've been using it for a while, but now it seems that some of the operations I want to do do not work on a ragged batch. I'm afraid the Dataset stuff is beyond me (and very poorly documented): do you happen to know how to adapt it to generate padded batches (i.e., each batch padded to the length of the longes sample it it)?</p>\n<p>Also, do I understand correctly that your code pre-forms batches, i.e. the batches would always be the same (but in a random order) during training?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2175491,
              "author_name": "aapo kossi",
              "author_url": "",
              "post_date": "2023-03-09T21:57:09.763000",
              "content": "<p>Yes, you are correct that the batches will always be the same if you don't specifically unbatch and then shuffle them. You can generate padded batches for example with the following transformation:<br>\n<code>ds = ds.map(lambda x, y: (x.to_tensor(), y)</code><br>\nRagged tensors have the to_tensor method which pads ragged dimensions to the longest dimension in the tensor (or optionally some static shape which may truncate the dim size).<br>\nAdditionally, <code>tf.keras.layers.Masking()</code> may or may not be useful when you are using padding, depending on your model architecture and if your layers support masking.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2159642": "I'm wondering what efficient method people are using for loading the training data, since I'm having trouble with my first option, which is parallelized loading by using tfio.IODataset.from_parquet for each sample. I am now working with tf.py_function and reading with pandas, but that is slow. I know I could simply save the dataset after waiting for it to load once, but a faster method would be nice.\n\nTrying to load a parquet file into a dataset with tfio silently crashes the kernel. The error seems to be the following:\n`terminate called after throwing an instance of 'parquet::ParquetException'`\n`what(): Unexpected end of stream`",
    "2163022": "I have now published a [tf.Dataset](https://www.kaggle.com/datasets/aapokossi/saved-tfdataset-of-google-isl-recognition-data) that quite significantly speeds up loading for me (700 it/sec vs. 60 it/sec originally, with no batching), along with [a notebook](https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-tf-dataset) on how to easily save the data yourself, which should make it easier for anyone to apply their own desired level of preprocessing before saving, if you want even more performance. I myself will probably be using a variation that is padded  to a specific sequence length and pre-batched. ",
    "2162219": "I'm quite new to TF and Kaggle competitions, but taking reference some of the codes shared (shoutout to @lonnieqin), I believe the competition has specified a method they will use to load the data, available in the Evaluation page. I'm guessing we can adopt the same method to load the data, then incorporate any further processing into your model pipeline.\n\n```python\ndef load_relevant_data_subset(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = int(len(data) / ROWS_PER_FRAME)\n    data = data.values.reshape(n_frames, ROWS_PER_FRAME, len(data_columns))\n    return data.astype(np.float32)\n```",
    "2161922": "I have the same question. I've been browsing through multiple explanations about creating a Tf Dataset from parquet files, but this is way above my level of understanding. When I try pre-padding the data for training as a  time-series model, the data becomes WAY too large, e.g. window size of 30 ->  94477x30x543x3 numbers\n\nActually, since this is really essential for proper training here, and since documentation for this is really lacking, I feel the Tensorflow development team should provide example code for this. Otherwise, anyone who is not a really advanced Tf software engineer is excluded from properly participating because of getting stuck on something I'm sure the organisers have already solved.\n\nSo please, can someone provide example code for a (train and val) Dataset with batch prefetching that can be used for efficient training, so we can focus on the model design.",
    "2160746": "Haven’t tried it but maybe this:\nhttps://www.tensorflow.org/io/api_docs/python/tfio/experimental/IODataset#from_parquet\n\nIt’s in experimental so maybe the issue your having has been fixed?\n\nI’ll dig more into this tomorrow.\n"
  }
}