{
  "id": 543637,
  "title": "Is anyone else having issues reading test.parquet?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543637",
  "author_name": "Schwartzman",
  "post_date": "2024-10-31T18:43:31.347000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I know test.parquet has just 1 partition, but I am seeing these errors when reading it either via pandas or dask_dataframe. </p>\n<p>I was able to read train.parquet just fine using dask_dataframe.</p>\n<p>===============</p>\n<pre><code> pandas  pd\n\npd.read_parquet()\nArrowTypeError: Unable to merge: Field date_id has incompatible types: int16 vs dictionary&lt;values=int32, indices=int32, ordered=&gt;\n</code></pre>\n<p>================</p>\n<pre><code>dd.read_parquet()\n\nArrowTypeError: Unable to merge: Field date_id has incompatible types: int16 vs int32\n\nThe above exception was the direct cause of the following exception:\n\nArrowTypeError                            Traceback (most recent call last)\n/usr/local/lib/python3/dist-packages/dask/backends.py  wrapper(*args, **kwargs)\n                              e\n                         :\n--&gt;                           exc  e\n     \n                 wrapper.__name__ = dispatch_name\n\nArrowTypeError: An error occurred  calling the read_parquet method registered to the pandas backend.\nOriginal Message: Unable to merge: Field date_id has incompatible types: int16 vs int32\n</code></pre>\n<p>In response, I have tried things like specify a dtype, but still get the same error</p>\n<pre><code>test_data = dd.read_parquet(, dtype={: })\n()\ntest_data.head()\n</code></pre>\n<p>Thank you. </p>",
  "messages": [
    {
      "id": 3033584,
      "postDate": "2024-11-01T10:26:46.413Z",
      "content": "<p>try this. <br>\nIt works for me<br>\ndata = pl.scan_parquet(f\"{data_dir}/train.parquet\")</p>",
      "rawMarkdown": "try this. \nIt works for me\ndata = pl.scan_parquet(f\"{data_dir}/train.parquet\")",
      "votes": 1,
      "replies": [
        {
          "id": 3033620,
          "postDate": "2024-11-01T11:11:43.420Z",
          "content": "<p>the error suggest there is a mismatch in data types for the date_id<br>\nyou can either read it lazily with polars then converted it back to pandas  :<br>\ndata= data.collect().to_pandas <br>\nor since it has only one partition you can read it directly with :<br>\ndata = pd.read_parquet('test.parquet/date_id=0/part-0.parquet') </p>",
          "rawMarkdown": "the error suggest there is a mismatch in data types for the date_id\nyou can either read it lazily with polars then converted it back to pandas  :\ndata= data.collect().to_pandas \nor since it has only one partition you can read it directly with :\ndata = pd.read_parquet('test.parquet/date_id=0/part-0.parquet') \n"
        }
      ]
    },
    {
      "id": 3033094,
      "postDate": "2024-10-31T18:43:31.347Z",
      "content": "<p>I know test.parquet has just 1 partition, but I am seeing these errors when reading it either via pandas or dask_dataframe. </p>\n<p>I was able to read train.parquet just fine using dask_dataframe.</p>\n<p>===============</p>\n<pre><code> pandas  pd\n\npd.read_parquet()\nArrowTypeError: Unable to merge: Field date_id has incompatible types: int16 vs dictionary&lt;values=int32, indices=int32, ordered=&gt;\n</code></pre>\n<p>================</p>\n<pre><code>dd.read_parquet()\n\nArrowTypeError: Unable to merge: Field date_id has incompatible types: int16 vs int32\n\nThe above exception was the direct cause of the following exception:\n\nArrowTypeError                            Traceback (most recent call last)\n/usr/local/lib/python3/dist-packages/dask/backends.py  wrapper(*args, **kwargs)\n                              e\n                         :\n--&gt;                           exc  e\n     \n                 wrapper.__name__ = dispatch_name\n\nArrowTypeError: An error occurred  calling the read_parquet method registered to the pandas backend.\nOriginal Message: Unable to merge: Field date_id has incompatible types: int16 vs int32\n</code></pre>\n<p>In response, I have tried things like specify a dtype, but still get the same error</p>\n<pre><code>test_data = dd.read_parquet(, dtype={: })\n()\ntest_data.head()\n</code></pre>\n<p>Thank you. </p>",
      "rawMarkdown": "I know test.parquet has just 1 partition, but I am seeing these errors when reading it either via pandas or dask_dataframe. \n\nI was able to read train.parquet just fine using dask_dataframe.\n\n===============\n```python\nimport pandas as pd\n\npd.read_parquet('test.parquet')\nArrowTypeError: Unable to merge: Field date_id has incompatible types: int16 vs dictionary<values=int32, indices=int32, ordered=0>\n```\n================\n```python\ndd.read_parquet('test.parquet')\n\nArrowTypeError: Unable to merge: Field date_id has incompatible types: int16 vs int32\n\nThe above exception was the direct cause of the following exception:\n\nArrowTypeError                            Traceback (most recent call last)\n/usr/local/lib/python3.10/dist-packages/dask/backends.py in wrapper(*args, **kwargs)\n    149                         raise e\n    150                     else:\n--> 151                         raise exc from e\n    152 \n    153             wrapper.__name__ = dispatch_name\n\nArrowTypeError: An error occurred while calling the read_parquet method registered to the pandas backend.\nOriginal Message: Unable to merge: Field date_id has incompatible types: int16 vs int32\n```\n\nIn response, I have tried things like specify a dtype, but still get the same error\n\n```python\ntest_data = dd.read_parquet('test.parquet', dtype={'date_id': 'int16'})\nprint(f\"Num rows: {len(test_data)}\")\ntest_data.head(2)\n```\n\nThank you. ",
      "votes": 2
    },
    {
      "id": 3033337,
      "postDate": "2024-11-01T02:13:47.337Z",
      "content": "<p>pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0/')</p>",
      "rawMarkdown": "pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0/')"
    },
    {
      "id": 3033306,
      "postDate": "2024-11-01T01:27:20.607Z",
      "content": "<p>Ok. This does work when reading with <code>engine = 'fastparquet'</code>, but not with the default engine or <code>engine=pyarrow</code></p>",
      "rawMarkdown": "Ok. This does work when reading with `engine = 'fastparquet'`, but not with the default engine or `engine=pyarrow`"
    },
    {
      "id": 3033207,
      "postDate": "2024-10-31T22:04:11.877Z",
      "content": "<p>Thank you, Younes.</p>\n<p>I even downloaded the zip again and extracted the files again. </p>\n<p>Please see the attached small notebook where I reproduce the error.</p>",
      "rawMarkdown": "Thank you, Younes.\n\nI even downloaded the zip again and extracted the files again. \n \nPlease see the attached small notebook where I reproduce the error."
    },
    {
      "id": 3033163,
      "postDate": "2024-10-31T20:23:59.990Z",
      "content": "<p>can you provide an example notebook that produces this error? </p>",
      "rawMarkdown": "can you provide an example notebook that produces this error? "
    },
    {
      "id": 3037963,
      "postDate": "2024-11-06T12:17:48.413Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3033584,
      "author_name": "Farhan Kardan",
      "author_url": "",
      "post_date": "2024-11-01T10:26:46.413000",
      "content": "<p>try this. <br>\nIt works for me<br>\ndata = pl.scan_parquet(f\"{data_dir}/train.parquet\")</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3033620,
          "author_name": "Younes Benalia",
          "author_url": "",
          "post_date": "2024-11-01T11:11:43.420000",
          "content": "<p>the error suggest there is a mismatch in data types for the date_id<br>\nyou can either read it lazily with polars then converted it back to pandas  :<br>\ndata= data.collect().to_pandas <br>\nor since it has only one partition you can read it directly with :<br>\ndata = pd.read_parquet('test.parquet/date_id=0/part-0.parquet') </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3033337,
      "author_name": "Jack",
      "author_url": "",
      "post_date": "2024-11-01T02:13:47.337000",
      "content": "<p>pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0/')</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033306,
      "author_name": "Schwartzman",
      "author_url": "",
      "post_date": "2024-11-01T01:27:20.607000",
      "content": "<p>Ok. This does work when reading with <code>engine = 'fastparquet'</code>, but not with the default engine or <code>engine=pyarrow</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033207,
      "author_name": "Schwartzman",
      "author_url": "",
      "post_date": "2024-10-31T22:04:11.877000",
      "content": "<p>Thank you, Younes.</p>\n<p>I even downloaded the zip again and extracted the files again. </p>\n<p>Please see the attached small notebook where I reproduce the error.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033163,
      "author_name": "Younes Benalia",
      "author_url": "",
      "post_date": "2024-10-31T20:23:59.990000",
      "content": "<p>can you provide an example notebook that produces this error? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3037963,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-06T12:17:48.413000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3033584": "try this. \nIt works for me\ndata = pl.scan_parquet(f\"{data_dir}/train.parquet\")",
    "3033094": "I know test.parquet has just 1 partition, but I am seeing these errors when reading it either via pandas or dask_dataframe. \n\nI was able to read train.parquet just fine using dask_dataframe.\n\n===============\n```python\nimport pandas as pd\n\npd.read_parquet('test.parquet')\nArrowTypeError: Unable to merge: Field date_id has incompatible types: int16 vs dictionary<values=int32, indices=int32, ordered=0>\n```\n================\n```python\ndd.read_parquet('test.parquet')\n\nArrowTypeError: Unable to merge: Field date_id has incompatible types: int16 vs int32\n\nThe above exception was the direct cause of the following exception:\n\nArrowTypeError                            Traceback (most recent call last)\n/usr/local/lib/python3.10/dist-packages/dask/backends.py in wrapper(*args, **kwargs)\n    149                         raise e\n    150                     else:\n--> 151                         raise exc from e\n    152 \n    153             wrapper.__name__ = dispatch_name\n\nArrowTypeError: An error occurred while calling the read_parquet method registered to the pandas backend.\nOriginal Message: Unable to merge: Field date_id has incompatible types: int16 vs int32\n```\n\nIn response, I have tried things like specify a dtype, but still get the same error\n\n```python\ntest_data = dd.read_parquet('test.parquet', dtype={'date_id': 'int16'})\nprint(f\"Num rows: {len(test_data)}\")\ntest_data.head(2)\n```\n\nThank you. ",
    "3033337": "pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0/')",
    "3033306": "Ok. This does work when reading with `engine = 'fastparquet'`, but not with the default engine or `engine=pyarrow`",
    "3033207": "Thank you, Younes.\n\nI even downloaded the zip again and extracted the files again. \n \nPlease see the attached small notebook where I reproduce the error.",
    "3033163": "can you provide an example notebook that produces this error? ",
    "3037963": ""
  }
}