{
  "id": 539703,
  "title": "Most efficient way to read and aggregate all the series files at once",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/539703",
  "author_name": "",
  "post_date": "2024-10-10T11:41:47.780682800Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>All the public notebooks that I saw so far are using the same reading / processing functions like load_time_series(dirname) and then a series of manipulations to extract the id from folder name, add it as a column, doing the aggregation stuff and then concatenating all the dataframes. (By the way cheers to the person who made it first, I'd like a link to the real author's notebook).<br>\nBut there is a simpler way, because it's a hive data structure (more <a href=\"https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory\" target=\"_blank\">here</a>) and polars knows by default to extract the id as a new column, so try this snippet to get all the files at once and to create the id column directly, it takes about: 1.5 minutes and 25gb of RAM at the peak:</p>\n<pre><code>root_path = \ntrain_series_path = \npl_df = pl.scan_parquet(train_series_path).group_by().agg().collect()\n</code></pre>",
  "messages": [
    {
      "id": "3013653",
      "postDate": "10/10/2024 11:41:47",
      "content": "<p>All the public notebooks that I saw so far are using the same reading / processing functions like load_time_series(dirname) and then a series of manipulations to extract the id from folder name, add it as a column, doing the aggregation stuff and then concatenating all the dataframes. (By the way cheers to the person who made it first, I'd like a link to the real author's notebook).<br>\nBut there is a simpler way, because it's a hive data structure (more <a href=\"https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory\" target=\"_blank\">here</a>) and polars knows by default to extract the id as a new column, so try this snippet to get all the files at once and to create the id column directly, it takes about: 1.5 minutes and 25gb of RAM at the peak:</p>\n<pre><code>root_path = \ntrain_series_path = \npl_df = pl.scan_parquet(train_series_path).group_by().agg().collect()\n</code></pre>",
      "rawMarkdown": "All the public notebooks that I saw so far are using the same reading / processing functions like load_time_series(dirname) and then a series of manipulations to extract the id from folder name, add it as a column, doing the aggregation stuff and then concatenating all the dataframes. (By the way cheers to the person who made it first, I'd like a link to the real author's notebook).\nBut there is a simpler way, because it's a hive data structure (more [here](https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory)) and polars knows by default to extract the id as a new column, so try this snippet to get all the files at once and to create the id column directly, it takes about: 1.5 minutes and 25gb of RAM at the peak:\n```python\nroot_path = '/kaggle/input/child-mind-institute-problematic-internet-use'\ntrain_series_path = f'{root_path}/series_train.parquet'\npl_df = pl.scan_parquet(train_series_path).group_by(\"id\").agg('your_aggregations_here').collect()\n```",
      "votes": null
    },
    {
      "id": "3013807",
      "postDate": "10/10/2024 14:49:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> I think that load_time_series() first appeared in the <a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-modell\" target=\"_blank\">CMI | Best Single Model</a> notebook by <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> .</p>\n<p>You can see the chronology of Notebooks in Code and Posts in Discussions by selecting Recently Created and Recently Posted, respectively, and working from the back forward 🙂</p>",
      "rawMarkdown": "Hi @eu1234 I think that load_time_series() first appeared in the [CMI | Best Single Model](https://www.kaggle.com/code/abdmental01/cmi-best-single-modell) notebook by @abdmental01 .\n\nYou can see the chronology of Notebooks in Code and Posts in Discussions by selecting Recently Created and Recently Posted, respectively, and working from the back forward 🙂",
      "votes": null
    },
    {
      "id": "3014749",
      "postDate": "10/11/2024 14:53:40",
      "content": "<p>That's a good approach, I'm using it but sometimes depends on the aggregation the ram explodes. It's exploding also on my local machine that has 32gb of ram. So I'll suggest to compute ids in batch.</p>",
      "rawMarkdown": "That's a good approach, I'm using it but sometimes depends on the aggregation the ram explodes. It's exploding also on my local machine that has 32gb of ram. So I'll suggest to compute ids in batch.",
      "votes": null
    },
    {
      "id": "3014765",
      "postDate": "10/11/2024 15:06:35",
      "content": "<p>Yes, it depends on how many columns and aggregations you are doing, but even so you can just split the columns in 2 parts and aggregate them separately instead of going in file batches. At least for 8 columns it has enough RAM to process multiple aggregations</p>",
      "rawMarkdown": "Yes, it depends on how many columns and aggregations you are doing, but even so you can just split the columns in 2 parts and aggregate them separately instead of going in file batches. At least for 8 columns it has enough RAM to process multiple aggregations",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3013807,
      "author_name": "dan3dewey",
      "author_url": "",
      "post_date": "10/10/2024 14:49:54",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> I think that load_time_series() first appeared in the <a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-modell\" target=\"_blank\">CMI | Best Single Model</a> notebook by <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> .</p>\n<p>You can see the chronology of Notebooks in Code and Posts in Discussions by selecting Recently Created and Recently Posted, respectively, and working from the back forward 🙂</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3014749,
      "author_name": "faibioss",
      "author_url": "",
      "post_date": "10/11/2024 14:53:40",
      "content": "<p>That's a good approach, I'm using it but sometimes depends on the aggregation the ram explodes. It's exploding also on my local machine that has 32gb of ram. So I'll suggest to compute ids in batch.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3014765,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "10/11/2024 15:06:35",
          "content": "<p>Yes, it depends on how many columns and aggregations you are doing, but even so you can just split the columns in 2 parts and aggregate them separately instead of going in file batches. At least for 8 columns it has enough RAM to process multiple aggregations</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3013653": "All the public notebooks that I saw so far are using the same reading / processing functions like load_time_series(dirname) and then a series of manipulations to extract the id from folder name, add it as a column, doing the aggregation stuff and then concatenating all the dataframes. (By the way cheers to the person who made it first, I'd like a link to the real author's notebook).\nBut there is a simpler way, because it's a hive data structure (more [here](https://docs.pola.rs/user-guide/io/hive/#scanning-a-hive-directory)) and polars knows by default to extract the id as a new column, so try this snippet to get all the files at once and to create the id column directly, it takes about: 1.5 minutes and 25gb of RAM at the peak:\n```python\nroot_path = '/kaggle/input/child-mind-institute-problematic-internet-use'\ntrain_series_path = f'{root_path}/series_train.parquet'\npl_df = pl.scan_parquet(train_series_path).group_by(\"id\").agg('your_aggregations_here').collect()\n```",
    "3013807": "Hi @eu1234 I think that load_time_series() first appeared in the [CMI | Best Single Model](https://www.kaggle.com/code/abdmental01/cmi-best-single-modell) notebook by @abdmental01 .\n\nYou can see the chronology of Notebooks in Code and Posts in Discussions by selecting Recently Created and Recently Posted, respectively, and working from the back forward 🙂",
    "3014749": "That's a good approach, I'm using it but sometimes depends on the aggregation the ram explodes. It's exploding also on my local machine that has 32gb of ram. So I'll suggest to compute ids in batch.",
    "3014765": "Yes, it depends on how many columns and aggregations you are doing, but even so you can just split the columns in 2 parts and aggregate them separately instead of going in file batches. At least for 8 columns it has enough RAM to process multiple aggregations"
  },
  "source": "meta"
}