{
  "id": 543252,
  "title": "Fast read for cross validation",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543252",
  "author_name": "",
  "post_date": "2024-10-29T14:02:32.743186800Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>What is the best way to save different pre-processed version of the data, so you can load different dataset/columns based on different validation folds. I find it is very time consuming to read data in each fold!</p>",
  "messages": [
    {
      "id": "3031247",
      "postDate": "10/29/2024 14:02:32",
      "content": "<p>What is the best way to save different pre-processed version of the data, so you can load different dataset/columns based on different validation folds. I find it is very time consuming to read data in each fold!</p>",
      "rawMarkdown": "What is the best way to save different pre-processed version of the data, so you can load different dataset/columns based on different validation folds. I find it is very time consuming to read data in each fold!",
      "votes": null
    },
    {
      "id": "3031891",
      "postDate": "10/30/2024 08:55:57",
      "content": "<p><a href=\"https://www.kaggle.com/chaotic\" target=\"_blank\">@chaotic</a> Is it not possible to save the parquet file with Polars and then read it using a LazyFrame?</p>\n<p>example</p>\n<pre><code>#  pre-\ntrain.write_parquet()\n</code></pre>\n<pre><code>df = pl.scan_parquet()\n\n# example ( kfold )\n  in ():\n    train = df.(pl.() != ).collect()\n    valid = df.(pl.() == ).collect()\n\n…\n</code></pre>",
      "rawMarkdown": "chaotic Is it not possible to save the parquet file with Polars and then read it using a LazyFrame?\n\nexample\n~~~\n# after pre-process\ntrain.write_parquet(\"train.parquet\")\n~~~\n\n~~~\ndf = pl.scan_parquet(\"train.parquet\")\n\n# example ( 5kfold )\nfor fold in range(5):\n    train = df.filter(pl.col(\"fold\") != fold).collect()\n    valid = df.filter(pl.col(\"fold\") == fold).collect()\n\n…\n~~~",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3031891,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "10/30/2024 08:55:57",
      "content": "<p><a href=\"https://www.kaggle.com/chaotic\" target=\"_blank\">@chaotic</a> Is it not possible to save the parquet file with Polars and then read it using a LazyFrame?</p>\n<p>example</p>\n<pre><code>#  pre-\ntrain.write_parquet()\n</code></pre>\n<pre><code>df = pl.scan_parquet()\n\n# example ( kfold )\n  in ():\n    train = df.(pl.() != ).collect()\n    valid = df.(pl.() == ).collect()\n\n…\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3031247": "What is the best way to save different pre-processed version of the data, so you can load different dataset/columns based on different validation folds. I find it is very time consuming to read data in each fold!",
    "3031891": "chaotic Is it not possible to save the parquet file with Polars and then read it using a LazyFrame?\n\nexample\n~~~\n# after pre-process\ntrain.write_parquet(\"train.parquet\")\n~~~\n\n~~~\ndf = pl.scan_parquet(\"train.parquet\")\n\n# example ( 5kfold )\nfor fold in range(5):\n    train = df.filter(pl.col(\"fold\") != fold).collect()\n    valid = df.filter(pl.col(\"fold\") == fold).collect()\n\n…\n~~~"
  },
  "source": "meta"
}