{
  "id": 398825,
  "title": "What is the best way to read the parquet files and preprocess them",
  "url": "/competitions/asl-signs/discussion/398825",
  "author_name": "",
  "post_date": "2023-04-01T03:27:35.030684300Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>I was trying to create tensorflow data set pipeline with the given parquet file locations. But I get errors. Here is my code</h2>\n<pre><code>import tensorflow as tf\nimport glob\n\ndef parquet_dataset(file_pattern):\n    files = glob.glob(file_pattern)\n    ds = tf.data.Dataset.from_tensor_slices(files)\n    ds = ds.map(lambda x: tf.py_function(read_parquet_file, [tf.strings.strip(x)], (tf.float32, tf.int64)))\n    return ds\n\ndef read_parquet_file(file_path):\n    df = pd.read_parquet(file_path)\n   data = df[['x', 'y', 'z']].to_numpy()\n    num_rows = data.shape[0]\n    return data, num_rows\n\ndef parquet_dataset(file_pattern):\n    files = glob.glob(file_pattern)\n    ds = tf.data.Dataset.from_tensor_slices(files)\n    ds = ds.map(lambda x: tf.py_function(read_parquet_file, [tf.strings.strip(x)], (tf.float32, tf.int64)))\n    return ds\n\nfile_pattern = '/kaggle/input/asl-signs/train_landmark_files/16069/*'\ntrain_dataset = parquet_dataset(file_pattern)\n\nfor item in train_dataset:\n    print(item)\n    break\n</code></pre>\n<h2>And I get the following error</h2>\n<p>TypeError                                 Traceback (most recent call last)<br>\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2285328250.parquet', shape=(), dtype=string)<br>\nException ignored in: 'pyarrow._dataset._make_file_source'<br>\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2285328250.parquet', shape=(), dtype=string)<br>\nException ignored in: 'pyarrow._dataset._make_file_source'<br>\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2596699720.parquet', shape=(), dtype=string)</p>\n<h2>Please can someone help me to solve the issue? Thanks</h2>",
  "messages": [
    {
      "id": "2204856",
      "postDate": "04/01/2023 03:27:35",
      "content": "<h2>I was trying to create tensorflow data set pipeline with the given parquet file locations. But I get errors. Here is my code</h2>\n<pre><code>import tensorflow as tf\nimport glob\n\ndef parquet_dataset(file_pattern):\n    files = glob.glob(file_pattern)\n    ds = tf.data.Dataset.from_tensor_slices(files)\n    ds = ds.map(lambda x: tf.py_function(read_parquet_file, [tf.strings.strip(x)], (tf.float32, tf.int64)))\n    return ds\n\ndef read_parquet_file(file_path):\n    df = pd.read_parquet(file_path)\n   data = df[['x', 'y', 'z']].to_numpy()\n    num_rows = data.shape[0]\n    return data, num_rows\n\ndef parquet_dataset(file_pattern):\n    files = glob.glob(file_pattern)\n    ds = tf.data.Dataset.from_tensor_slices(files)\n    ds = ds.map(lambda x: tf.py_function(read_parquet_file, [tf.strings.strip(x)], (tf.float32, tf.int64)))\n    return ds\n\nfile_pattern = '/kaggle/input/asl-signs/train_landmark_files/16069/*'\ntrain_dataset = parquet_dataset(file_pattern)\n\nfor item in train_dataset:\n    print(item)\n    break\n</code></pre>\n<h2>And I get the following error</h2>\n<p>TypeError                                 Traceback (most recent call last)<br>\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2285328250.parquet', shape=(), dtype=string)<br>\nException ignored in: 'pyarrow._dataset._make_file_source'<br>\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2285328250.parquet', shape=(), dtype=string)<br>\nException ignored in: 'pyarrow._dataset._make_file_source'<br>\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2596699720.parquet', shape=(), dtype=string)</p>\n<h2>Please can someone help me to solve the issue? Thanks</h2>",
      "rawMarkdown": "I was trying to create tensorflow data set pipeline with the given parquet file locations. But I get errors. Here is my code\n---------------------------------------------------------------------------\n```\nimport tensorflow as tf\nimport glob\n\ndef parquet_dataset(file_pattern):\n    files = glob.glob(file_pattern)\n    ds = tf.data.Dataset.from_tensor_slices(files)\n    ds = ds.map(lambda x: tf.py_function(read_parquet_file, [tf.strings.strip(x)], (tf.float32, tf.int64)))\n    return ds\n\ndef read_parquet_file(file_path):\n    df = pd.read_parquet(file_path)\n   data = df[['x', 'y', 'z']].to_numpy()\n    num_rows = data.shape[0]\n    return data, num_rows\n\ndef parquet_dataset(file_pattern):\n    files = glob.glob(file_pattern)\n    ds = tf.data.Dataset.from_tensor_slices(files)\n    ds = ds.map(lambda x: tf.py_function(read_parquet_file, [tf.strings.strip(x)], (tf.float32, tf.int64)))\n    return ds\n\nfile_pattern = '/kaggle/input/asl-signs/train_landmark_files/16069/*'\ntrain_dataset = parquet_dataset(file_pattern)\n\nfor item in train_dataset:\n    print(item)\n    break\n```\nAnd I get the following error\n---------------------------------------------------------------------------\n\nTypeError                                 Traceback (most recent call last)\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2285328250.parquet', shape=(), dtype=string)\nException ignored in: 'pyarrow._dataset._make_file_source'\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2285328250.parquet', shape=(), dtype=string)\nException ignored in: 'pyarrow._dataset._make_file_source'\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2596699720.parquet', shape=(), dtype=string)\n\n\nPlease can someone help me to solve the issue? Thanks\n---------------------------------------------------------------------------",
      "votes": null
    },
    {
      "id": "2204891",
      "postDate": "04/01/2023 04:22:29",
      "content": "<p>I would suggest you look at this super helpful notebook related to building a tf dataset from the input data. <a href=\"https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-ragged-tf-dataset\" target=\"_blank\">https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-ragged-tf-dataset</a> <br>\nHappy Kaggle-ing,<br>\nPeter</p>",
      "rawMarkdown": "I would suggest you look at this super helpful notebook related to building a tf dataset from the input data. https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-ragged-tf-dataset \nHappy Kaggle-ing,\nPeter",
      "votes": null
    },
    {
      "id": "2207548",
      "postDate": "04/03/2023 13:50:21",
      "content": "<p><code>df = pd.read_parquet(file_path.numpy().decode('utf-8'))</code><br>\ntry the above code. the reason for the error is that the filepath you passed to pd.read_parquet is not a python string. It is tf.Tensor, so you need to first convert to numpy array, and you also need to decode it because tf.string is in bytes format. </p>",
      "rawMarkdown": "`df = pd.read_parquet(file_path.numpy().decode('utf-8'))`\n\ntry the above code. the reason for the error is that the filepath you passed to pd.read_parquet is not a python string. It is tf.Tensor, so you need to first convert to numpy array, and you also need to decode it because tf.string is in bytes format.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2204891,
      "author_name": "koehlepe",
      "author_url": "",
      "post_date": "04/01/2023 04:22:29",
      "content": "<p>I would suggest you look at this super helpful notebook related to building a tf dataset from the input data. <a href=\"https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-ragged-tf-dataset\" target=\"_blank\">https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-ragged-tf-dataset</a> <br>\nHappy Kaggle-ing,<br>\nPeter</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2207548,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "04/03/2023 13:50:21",
      "content": "<p><code>df = pd.read_parquet(file_path.numpy().decode('utf-8'))</code><br>\ntry the above code. the reason for the error is that the filepath you passed to pd.read_parquet is not a python string. It is tf.Tensor, so you need to first convert to numpy array, and you also need to decode it because tf.string is in bytes format. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2204856": "I was trying to create tensorflow data set pipeline with the given parquet file locations. But I get errors. Here is my code\n---------------------------------------------------------------------------\n```\nimport tensorflow as tf\nimport glob\n\ndef parquet_dataset(file_pattern):\n    files = glob.glob(file_pattern)\n    ds = tf.data.Dataset.from_tensor_slices(files)\n    ds = ds.map(lambda x: tf.py_function(read_parquet_file, [tf.strings.strip(x)], (tf.float32, tf.int64)))\n    return ds\n\ndef read_parquet_file(file_path):\n    df = pd.read_parquet(file_path)\n   data = df[['x', 'y', 'z']].to_numpy()\n    num_rows = data.shape[0]\n    return data, num_rows\n\ndef parquet_dataset(file_pattern):\n    files = glob.glob(file_pattern)\n    ds = tf.data.Dataset.from_tensor_slices(files)\n    ds = ds.map(lambda x: tf.py_function(read_parquet_file, [tf.strings.strip(x)], (tf.float32, tf.int64)))\n    return ds\n\nfile_pattern = '/kaggle/input/asl-signs/train_landmark_files/16069/*'\ntrain_dataset = parquet_dataset(file_pattern)\n\nfor item in train_dataset:\n    print(item)\n    break\n```\nAnd I get the following error\n---------------------------------------------------------------------------\n\nTypeError                                 Traceback (most recent call last)\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2285328250.parquet', shape=(), dtype=string)\nException ignored in: 'pyarrow._dataset._make_file_source'\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2285328250.parquet', shape=(), dtype=string)\nException ignored in: 'pyarrow._dataset._make_file_source'\nTypeError: cannot construct a FileSource from tf.Tensor(b'/kaggle/input/asl-signs/train_landmark_files/16069/2596699720.parquet', shape=(), dtype=string)\n\n\nPlease can someone help me to solve the issue? Thanks\n---------------------------------------------------------------------------",
    "2204891": "I would suggest you look at this super helpful notebook related to building a tf dataset from the input data. https://www.kaggle.com/code/aapokossi/how-to-save-parquet-data-as-ragged-tf-dataset \nHappy Kaggle-ing,\nPeter",
    "2207548": "`df = pd.read_parquet(file_path.numpy().decode('utf-8'))`\n\ntry the above code. the reason for the error is that the filepath you passed to pd.read_parquet is not a python string. It is tf.Tensor, so you need to first convert to numpy array, and you also need to decode it because tf.string is in bytes format."
  },
  "source": "meta"
}