{
  "id": 540508,
  "title": "Save time and resources with train data imports",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/540508",
  "author_name": "Ravi Ramakrishnan",
  "post_date": "2024-10-15T00:24:54.163000",
  "votes": 62,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>The train dataset here is partitioned into 10 sections, with 1 parquet file per section. Most of us using pandas would be inclined to manually providing the requisite folder paths in a loop and importing files sequentially. <strong>This is however not needed here</strong> as polars automatically does this under the hood. <strong>One may simply use the below code and read-in all relevant files at once</strong> - </p>\n<pre><code> polars  pl\ntrain = \\\npl.scan_parquet(\n    \n).\\\nselect(\n    pl.int_range(pl.(), dtype=pl.UInt64).alias(),\n    pl.(),\n)\n</code></pre>\n<p>Polars automatically scans through the directory path and reads-in the 10-files at once. One could then filter the lazyframe as per one's requirement using the <strong>partition_id</strong> column. </p>",
  "messages": [
    {
      "id": 3017456,
      "postDate": "2024-10-15T00:24:54.163Z",
      "content": "<p>Hello all,</p>\n<p>The train dataset here is partitioned into 10 sections, with 1 parquet file per section. Most of us using pandas would be inclined to manually providing the requisite folder paths in a loop and importing files sequentially. <strong>This is however not needed here</strong> as polars automatically does this under the hood. <strong>One may simply use the below code and read-in all relevant files at once</strong> - </p>\n<pre><code> polars  pl\ntrain = \\\npl.scan_parquet(\n    \n).\\\nselect(\n    pl.int_range(pl.(), dtype=pl.UInt64).alias(),\n    pl.(),\n)\n</code></pre>\n<p>Polars automatically scans through the directory path and reads-in the 10-files at once. One could then filter the lazyframe as per one's requirement using the <strong>partition_id</strong> column. </p>",
      "rawMarkdown": "Hello all,\n\nThe train dataset here is partitioned into 10 sections, with 1 parquet file per section. Most of us using pandas would be inclined to manually providing the requisite folder paths in a loop and importing files sequentially. **This is however not needed here** as polars automatically does this under the hood. **One may simply use the below code and read-in all relevant files at once** - \n\n```python\nimport polars as pl\ntrain = \\\npl.scan_parquet(\n    f\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet\"\n).\\\nselect(\n    pl.int_range(pl.len(), dtype=pl.UInt64).alias(\"id\"),\n    pl.all(),\n)\n```\n\nPolars automatically scans through the directory path and reads-in the 10-files at once. One could then filter the lazyframe as per one's requirement using the **partition_id** column. \n",
      "votes": 62
    },
    {
      "id": 3018897,
      "postDate": "2024-10-16T05:44:45.163Z",
      "content": "<ol>\n<li>This fails on OSX once the over-eager OS taints the dirs with <code>.DS_Store</code> files</li>\n<li>It will fail silently, since it is a lazyframe - failure only percolates when the frame is collected.</li>\n<li>Use PEP8 … it works so well with polars </li>\n</ol>\n<pre><code>import polars as pl\n\ntrain = (\n    .scan_parquet(  )\n    .select(\n        .int_range(.len(), dtype=pl.UInt64).alias(),\n        pl.all(),\n    )\n)\n</code></pre>",
      "rawMarkdown": "1. This fails on OSX once the over-eager OS taints the dirs with `.DS_Store` files\n2. It will fail silently, since it is a lazyframe - failure only percolates when the frame is collected.\n3. Use PEP8 ... it works so well with polars \n\n```\nimport polars as pl\n\ntrain = (\n    pl.scan_parquet( f\"train.parquet/**/*.parquet\" )\n    .select(\n        pl.int_range(pl.len(), dtype=pl.UInt64).alias(\"id\"),\n        pl.all(),\n    )\n)\n```",
      "votes": 5
    },
    {
      "id": 3030761,
      "postDate": "2024-10-28T21:11:02.887Z",
      "content": "<p>since last three days. i am not able to make submissions. it fails and says \"notebook inference server error\" i tried with more than 10 notebooks. last week all submission were getting accepted. I am using same predict function as them but still not working with new models. i tried asking on discord as well. someone please guide. this is my first competition.</p>",
      "rawMarkdown": "since last three days. i am not able to make submissions. it fails and says \"notebook inference server error\" i tried with more than 10 notebooks. last week all submission were getting accepted. I am using same predict function as them but still not working with new models. i tried asking on discord as well. someone please guide. this is my first competition.\n",
      "votes": 3
    },
    {
      "id": 3023670,
      "postDate": "2024-10-20T20:16:33.867Z",
      "content": "<p>Hey!<br>\nWhenever I proceed with the training data loading, the kernel dies. I have tried using the Lazy Dataframe and then collecting to an Eager Dataframe. After that, any processing leads to the kernel being dead due to high RAM and CPU utilization. What is the best optimal way to load the entire dataset?</p>",
      "rawMarkdown": "Hey!\nWhenever I proceed with the training data loading, the kernel dies. I have tried using the Lazy Dataframe and then collecting to an Eager Dataframe. After that, any processing leads to the kernel being dead due to high RAM and CPU utilization. What is the best optimal way to load the entire dataset?",
      "votes": 3,
      "replies": [
        {
          "id": 3024101,
          "postDate": "2024-10-21T10:28:25.347Z",
          "content": "<p>Try and load it outside of Kaggle - this is one of the better ways to load the entire data. I am using an A6000 chip and 128GB RAM and this works well <a href=\"https://www.kaggle.com/vaibhavsomani04\" target=\"_blank\">@vaibhavsomani04</a> </p>",
          "rawMarkdown": "Try and load it outside of Kaggle - this is one of the better ways to load the entire data. I am using an A6000 chip and 128GB RAM and this works well @vaibhavsomani04 ",
          "votes": 1,
          "replies": [
            {
              "id": 3029044,
              "postDate": "2024-10-26T19:31:34.700Z",
              "content": "<p>Thanks for the post and the comment. <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> I have a question for you - how can you do it outside of Kaggle? I thought this competition is only available via Kaggle's notebooks and there is no way to process data outside and upload it to a notebook for this competition. Thanks in advance!</p>",
              "rawMarkdown": "Thanks for the post and the comment. @ravi20076 I have a question for you - how can you do it outside of Kaggle? I thought this competition is only available via Kaggle's notebooks and there is no way to process data outside and upload it to a notebook for this competition. Thanks in advance!",
              "votes": 3
            },
            {
              "id": 3029050,
              "postDate": "2024-10-26T19:36:32.313Z",
              "content": "<p>Not at all <a href=\"https://www.kaggle.com/dinezra11\" target=\"_blank\">@dinezra11</a> <br>\nOne is recommended to train models outside of Kaggle and submit <strong>using Kaggle kernels</strong> </p>",
              "rawMarkdown": "Not at all @dinezra11 \nOne is recommended to train models outside of Kaggle and submit **using Kaggle kernels** ",
              "votes": 2
            },
            {
              "id": 3029067,
              "postDate": "2024-10-26T20:04:27.113Z",
              "content": "<p>I'm also using A6000 😀</p>",
              "rawMarkdown": "I'm also using A6000 😀",
              "votes": 1
            },
            {
              "id": 3029449,
              "postDate": "2024-10-27T10:40:52.833Z",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Oh wow, so I guess I have a lot of knowledge missing regrading Kaggle competitions. (This is the first \"serious\" competition I'm participating in).<br>\nI'll search now to learn about the Kaggle kernels and how to submis there. Thank you!</p>",
              "rawMarkdown": "@ravi20076 Oh wow, so I guess I have a lot of knowledge missing regrading Kaggle competitions. (This is the first \"serious\" competition I'm participating in).\nI'll search now to learn about the Kaggle kernels and how to submis there. Thank you!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3021017,
      "postDate": "2024-10-18T06:26:11.727Z",
      "content": "<p>Thanks for this information <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> . Might be my first competition here on kaggle. wish me luck :)</p>",
      "rawMarkdown": "Thanks for this information @ravi20076 . Might be my first competition here on kaggle. wish me luck :)",
      "votes": 3,
      "replies": [
        {
          "id": 3021101,
          "postDate": "2024-10-18T07:36:03.773Z",
          "content": "<p>All the best <a href=\"https://www.kaggle.com/cptverma\" target=\"_blank\">@cptverma</a> </p>",
          "rawMarkdown": "All the best @cptverma "
        }
      ]
    },
    {
      "id": 3017469,
      "postDate": "2024-10-15T00:57:19.043Z",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>  I just learned this for the first time! It automatically concatenates, doesn't it? Thank you! This is really helpful.</p>",
      "rawMarkdown": "@ravi20076  I just learned this for the first time! It automatically concatenates, doesn't it? Thank you! This is really helpful.",
      "votes": 4,
      "replies": [
        {
          "id": 3017627,
          "postDate": "2024-10-15T04:24:17.673Z",
          "content": "<p>Yes sir you are right <br>\nPolars reads in and concats all component files under the hood <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
          "rawMarkdown": "Yes sir you are right \nPolars reads in and concats all component files under the hood @chumajin ",
          "votes": 1
        },
        {
          "id": 3030255,
          "postDate": "2024-10-28T10:43:09.867Z",
          "content": "<p>Details here - <a href=\"https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory\" target=\"_blank\">https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory</a></p>",
          "rawMarkdown": "Details here - https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory",
          "votes": 1
        }
      ]
    },
    {
      "id": 3031018,
      "postDate": "2024-10-29T07:56:24.670Z",
      "content": "<p>Thanks for sharing this tip, Polars seems like a great time saver for handling multiple parquet files. Automating the import process is definitely a win. Excited to try this out in my next project.👍</p>",
      "rawMarkdown": "Thanks for sharing this tip, Polars seems like a great time saver for handling multiple parquet files. Automating the import process is definitely a win. Excited to try this out in my next project.👍",
      "votes": 1
    },
    {
      "id": 3030928,
      "postDate": "2024-10-29T04:11:58.193Z",
      "content": "<p>thanks for the information.Keep up the good work</p>",
      "rawMarkdown": "thanks for the information.Keep up the good work",
      "votes": 1
    },
    {
      "id": 3020753,
      "postDate": "2024-10-17T21:03:24.487Z",
      "content": "<p>Very interesting and informative! </p>",
      "rawMarkdown": "Very interesting and informative! ",
      "votes": 1
    },
    {
      "id": 3019659,
      "postDate": "2024-10-16T18:25:23.737Z",
      "content": "<p>It is worthy, <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "It is worthy, @ravi20076 ",
      "votes": 1
    },
    {
      "id": 3019076,
      "postDate": "2024-10-16T08:40:35.773Z",
      "content": "<p>Very useful information , kudos !!</p>",
      "rawMarkdown": "Very useful information , kudos !!",
      "votes": 1
    },
    {
      "id": 3018932,
      "postDate": "2024-10-16T06:40:23.363Z",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> so what about pandas or glob?? Aren't they helpful</p>",
      "rawMarkdown": "@ravi20076 so what about pandas or glob?? Aren't they helpful",
      "votes": 2,
      "replies": [
        {
          "id": 3019004,
          "postDate": "2024-10-16T07:41:57.097Z",
          "content": "<p><a href=\"https://www.kaggle.com/evanhislupus\" target=\"_blank\">@evanhislupus</a>, they are also useful, but you will need to code it well and ensure that the data in the test set API aligns to your code. Just check for memory consumption and you are good with glob as well!</p>",
          "rawMarkdown": "@evanhislupus, they are also useful, but you will need to code it well and ensure that the data in the test set API aligns to your code. Just check for memory consumption and you are good with glob as well!",
          "votes": 2,
          "replies": [
            {
              "id": 3019141,
              "postDate": "2024-10-16T10:00:16.490Z",
              "content": "<p>ohky! thanks a lot Sir!! <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> This was something new to me. Your discussions are always insightful. :-)</p>",
              "rawMarkdown": "ohky! thanks a lot Sir!! @ravi20076 This was something new to me. Your discussions are always insightful. :-)",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 3017839,
      "postDate": "2024-10-15T08:48:50.117Z",
      "content": "<p>thanks very much. Polars seems better than pandas for large datasets.</p>",
      "rawMarkdown": "thanks very much. Polars seems better than pandas for large datasets.",
      "votes": 2
    },
    {
      "id": 3017795,
      "postDate": "2024-10-15T07:55:54.967Z",
      "content": "<p>You could also just… <code>pl.read_parquet(\n    '../input/jane-street-realtime-marketdata-forecasting/train.parquet')</code>, no? :) </p>",
      "rawMarkdown": "You could also just... `pl.read_parquet(\n    '../input/jane-street-realtime-marketdata-forecasting/train.parquet')`, no? :) ",
      "votes": 2,
      "replies": [
        {
          "id": 3017808,
          "postDate": "2024-10-15T08:18:56.040Z",
          "content": "<p>You could do this if you want to work with eager frames. I usually avoid this and work in lazy-frame mode and collect with cpu/ gpu at the end of the data pipeline <a href=\"https://www.kaggle.com/missgranger\" target=\"_blank\">@missgranger</a> </p>",
          "rawMarkdown": "You could do this if you want to work with eager frames. I usually avoid this and work in lazy-frame mode and collect with cpu/ gpu at the end of the data pipeline @missgranger ",
          "votes": 2
        }
      ]
    },
    {
      "id": 3017706,
      "postDate": "2024-10-15T06:17:53.520Z",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Hello sir, I want to ask that how I can start participating in Jane street competition and what are the requirements for this like what things are required basically to start it with? </p>",
      "rawMarkdown": "@ravi20076 Hello sir, I want to ask that how I can start participating in Jane street competition and what are the requirements for this like what things are required basically to start it with? ",
      "votes": 2,
      "replies": [
        {
          "id": 3017739,
          "postDate": "2024-10-15T06:56:15.300Z",
          "content": "<p><a href=\"https://www.kaggle.com/manishtiwari07\" target=\"_blank\">@manishtiwari07</a> anyone can participate - there is no entry barrier here<br>\nPlease be mindful of the resources needed as this dataset is very huge</p>",
          "rawMarkdown": "@manishtiwari07 anyone can participate - there is no entry barrier here\nPlease be mindful of the resources needed as this dataset is very huge",
          "votes": 1
        }
      ]
    },
    {
      "id": 3030757,
      "postDate": "2024-10-28T21:07:08.937Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3029575,
      "postDate": "2024-10-27T13:20:01.107Z",
      "content": "<p>very interesting</p>",
      "rawMarkdown": "very interesting",
      "votes": 1
    },
    {
      "id": 3019847,
      "postDate": "2024-10-16T23:51:03.910Z",
      "content": "<p>Thank your the information.</p>",
      "rawMarkdown": "Thank your the information.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3018897,
      "author_name": "Andrew Pullin",
      "author_url": "",
      "post_date": "2024-10-16T05:44:45.163000",
      "content": "<ol>\n<li>This fails on OSX once the over-eager OS taints the dirs with <code>.DS_Store</code> files</li>\n<li>It will fail silently, since it is a lazyframe - failure only percolates when the frame is collected.</li>\n<li>Use PEP8 … it works so well with polars </li>\n</ol>\n<pre><code>import polars as pl\n\ntrain = (\n    .scan_parquet(  )\n    .select(\n        .int_range(.len(), dtype=pl.UInt64).alias(),\n        pl.all(),\n    )\n)\n</code></pre>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3030761,
      "author_name": "Meet Babariya",
      "author_url": "",
      "post_date": "2024-10-28T21:11:02.887000",
      "content": "<p>since last three days. i am not able to make submissions. it fails and says \"notebook inference server error\" i tried with more than 10 notebooks. last week all submission were getting accepted. I am using same predict function as them but still not working with new models. i tried asking on discord as well. someone please guide. this is my first competition.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3023670,
      "author_name": "Vaibhav Somani",
      "author_url": "",
      "post_date": "2024-10-20T20:16:33.867000",
      "content": "<p>Hey!<br>\nWhenever I proceed with the training data loading, the kernel dies. I have tried using the Lazy Dataframe and then collecting to an Eager Dataframe. After that, any processing leads to the kernel being dead due to high RAM and CPU utilization. What is the best optimal way to load the entire dataset?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3024101,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-10-21T10:28:25.347000",
          "content": "<p>Try and load it outside of Kaggle - this is one of the better ways to load the entire data. I am using an A6000 chip and 128GB RAM and this works well <a href=\"https://www.kaggle.com/vaibhavsomani04\" target=\"_blank\">@vaibhavsomani04</a> </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3029044,
              "author_name": "Din Ezra",
              "author_url": "",
              "post_date": "2024-10-26T19:31:34.700000",
              "content": "<p>Thanks for the post and the comment. <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> I have a question for you - how can you do it outside of Kaggle? I thought this competition is only available via Kaggle's notebooks and there is no way to process data outside and upload it to a notebook for this competition. Thanks in advance!</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3029050,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-10-26T19:36:32.313000",
              "content": "<p>Not at all <a href=\"https://www.kaggle.com/dinezra11\" target=\"_blank\">@dinezra11</a> <br>\nOne is recommended to train models outside of Kaggle and submit <strong>using Kaggle kernels</strong> </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3029067,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-10-26T20:04:27.113000",
              "content": "<p>I'm also using A6000 😀</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3029449,
              "author_name": "Din Ezra",
              "author_url": "",
              "post_date": "2024-10-27T10:40:52.833000",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Oh wow, so I guess I have a lot of knowledge missing regrading Kaggle competitions. (This is the first \"serious\" competition I'm participating in).<br>\nI'll search now to learn about the Kaggle kernels and how to submis there. Thank you!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3021017,
      "author_name": "Pranav Verma",
      "author_url": "",
      "post_date": "2024-10-18T06:26:11.727000",
      "content": "<p>Thanks for this information <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> . Might be my first competition here on kaggle. wish me luck :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3021101,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-10-18T07:36:03.773000",
          "content": "<p>All the best <a href=\"https://www.kaggle.com/cptverma\" target=\"_blank\">@cptverma</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3017469,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "2024-10-15T00:57:19.043000",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>  I just learned this for the first time! It automatically concatenates, doesn't it? Thank you! This is really helpful.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3017627,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-10-15T04:24:17.673000",
          "content": "<p>Yes sir you are right <br>\nPolars reads in and concats all component files under the hood <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3030255,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-10-28T10:43:09.867000",
          "content": "<p>Details here - <a href=\"https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory\" target=\"_blank\">https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3031018,
      "author_name": "pranshu meshram",
      "author_url": "",
      "post_date": "2024-10-29T07:56:24.670000",
      "content": "<p>Thanks for sharing this tip, Polars seems like a great time saver for handling multiple parquet files. Automating the import process is definitely a win. Excited to try this out in my next project.👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3030928,
      "author_name": "Michael Fernandes",
      "author_url": "",
      "post_date": "2024-10-29T04:11:58.193000",
      "content": "<p>thanks for the information.Keep up the good work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3020753,
      "author_name": "Asma Hachaichi",
      "author_url": "",
      "post_date": "2024-10-17T21:03:24.487000",
      "content": "<p>Very interesting and informative! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3019659,
      "author_name": "Muhammed Tausif",
      "author_url": "",
      "post_date": "2024-10-16T18:25:23.737000",
      "content": "<p>It is worthy, <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3019076,
      "author_name": "Akshay M Bhat",
      "author_url": "",
      "post_date": "2024-10-16T08:40:35.773000",
      "content": "<p>Very useful information , kudos !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3018932,
      "author_name": "Evanhis",
      "author_url": "",
      "post_date": "2024-10-16T06:40:23.363000",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> so what about pandas or glob?? Aren't they helpful</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3019004,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-10-16T07:41:57.097000",
          "content": "<p><a href=\"https://www.kaggle.com/evanhislupus\" target=\"_blank\">@evanhislupus</a>, they are also useful, but you will need to code it well and ensure that the data in the test set API aligns to your code. Just check for memory consumption and you are good with glob as well!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3019141,
              "author_name": "Evanhis",
              "author_url": "",
              "post_date": "2024-10-16T10:00:16.490000",
              "content": "<p>ohky! thanks a lot Sir!! <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> This was something new to me. Your discussions are always insightful. :-)</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3017839,
      "author_name": "KennyT",
      "author_url": "",
      "post_date": "2024-10-15T08:48:50.117000",
      "content": "<p>thanks very much. Polars seems better than pandas for large datasets.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3017795,
      "author_name": "Faridka M",
      "author_url": "",
      "post_date": "2024-10-15T07:55:54.967000",
      "content": "<p>You could also just… <code>pl.read_parquet(\n    '../input/jane-street-realtime-marketdata-forecasting/train.parquet')</code>, no? :) </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3017808,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-10-15T08:18:56.040000",
          "content": "<p>You could do this if you want to work with eager frames. I usually avoid this and work in lazy-frame mode and collect with cpu/ gpu at the end of the data pipeline <a href=\"https://www.kaggle.com/missgranger\" target=\"_blank\">@missgranger</a> </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3017706,
      "author_name": "Manish Tiwari07",
      "author_url": "",
      "post_date": "2024-10-15T06:17:53.520000",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Hello sir, I want to ask that how I can start participating in Jane street competition and what are the requirements for this like what things are required basically to start it with? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3017739,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-10-15T06:56:15.300000",
          "content": "<p><a href=\"https://www.kaggle.com/manishtiwari07\" target=\"_blank\">@manishtiwari07</a> anyone can participate - there is no entry barrier here<br>\nPlease be mindful of the resources needed as this dataset is very huge</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3030757,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-28T21:07:08.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3029575,
      "author_name": "Azermed",
      "author_url": "",
      "post_date": "2024-10-27T13:20:01.107000",
      "content": "<p>very interesting</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3019847,
      "author_name": "Learner",
      "author_url": "",
      "post_date": "2024-10-16T23:51:03.910000",
      "content": "<p>Thank your the information.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3017456": "Hello all,\n\nThe train dataset here is partitioned into 10 sections, with 1 parquet file per section. Most of us using pandas would be inclined to manually providing the requisite folder paths in a loop and importing files sequentially. **This is however not needed here** as polars automatically does this under the hood. **One may simply use the below code and read-in all relevant files at once** - \n\n```python\nimport polars as pl\ntrain = \\\npl.scan_parquet(\n    f\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet\"\n).\\\nselect(\n    pl.int_range(pl.len(), dtype=pl.UInt64).alias(\"id\"),\n    pl.all(),\n)\n```\n\nPolars automatically scans through the directory path and reads-in the 10-files at once. One could then filter the lazyframe as per one's requirement using the **partition_id** column. \n",
    "3018897": "1. This fails on OSX once the over-eager OS taints the dirs with `.DS_Store` files\n2. It will fail silently, since it is a lazyframe - failure only percolates when the frame is collected.\n3. Use PEP8 ... it works so well with polars \n\n```\nimport polars as pl\n\ntrain = (\n    pl.scan_parquet( f\"train.parquet/**/*.parquet\" )\n    .select(\n        pl.int_range(pl.len(), dtype=pl.UInt64).alias(\"id\"),\n        pl.all(),\n    )\n)\n```",
    "3030761": "since last three days. i am not able to make submissions. it fails and says \"notebook inference server error\" i tried with more than 10 notebooks. last week all submission were getting accepted. I am using same predict function as them but still not working with new models. i tried asking on discord as well. someone please guide. this is my first competition.\n",
    "3023670": "Hey!\nWhenever I proceed with the training data loading, the kernel dies. I have tried using the Lazy Dataframe and then collecting to an Eager Dataframe. After that, any processing leads to the kernel being dead due to high RAM and CPU utilization. What is the best optimal way to load the entire dataset?",
    "3021017": "Thanks for this information @ravi20076 . Might be my first competition here on kaggle. wish me luck :)",
    "3017469": "@ravi20076  I just learned this for the first time! It automatically concatenates, doesn't it? Thank you! This is really helpful.",
    "3031018": "Thanks for sharing this tip, Polars seems like a great time saver for handling multiple parquet files. Automating the import process is definitely a win. Excited to try this out in my next project.👍",
    "3030928": "thanks for the information.Keep up the good work",
    "3020753": "Very interesting and informative! ",
    "3019659": "It is worthy, @ravi20076 ",
    "3019076": "Very useful information , kudos !!",
    "3018932": "@ravi20076 so what about pandas or glob?? Aren't they helpful",
    "3017839": "thanks very much. Polars seems better than pandas for large datasets.",
    "3017795": "You could also just... `pl.read_parquet(\n    '../input/jane-street-realtime-marketdata-forecasting/train.parquet')`, no? :) ",
    "3017706": "@ravi20076 Hello sir, I want to ask that how I can start participating in Jane street competition and what are the requirements for this like what things are required basically to start it with? ",
    "3030757": "",
    "3029575": "very interesting",
    "3019847": "Thank your the information."
  }
}