{
  "id": 496954,
  "title": "Using dask to handle this larger than life dataset. ",
  "url": "/competitions/leash-BELKA/discussion/496954",
  "author_name": "",
  "post_date": "2024-04-23T04:37:34.970085700Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>One of the interesting things of this competition is <em>it's larger than life</em> dataset and dealing with it (going to come with a lot of learnings along the way). <br>\nIf you're struggling with the very first step of loading this dataset using pandas (ikr), switch to dask to load the parquet files. Here's the code-snippet attached on how to do it. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2875376%2F147594485803a0e2bab182f65e2992f4%2FScreen%20Shot%202024-04-23%20at%209.54.35%20AM.png?generation=1713846717396710&amp;alt=media\">  </p>",
  "messages": [
    {
      "id": "2768877",
      "postDate": "04/23/2024 04:37:34",
      "content": "<p>One of the interesting things of this competition is <em>it's larger than life</em> dataset and dealing with it (going to come with a lot of learnings along the way). <br>\nIf you're struggling with the very first step of loading this dataset using pandas (ikr), switch to dask to load the parquet files. Here's the code-snippet attached on how to do it. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2875376%2F147594485803a0e2bab182f65e2992f4%2FScreen%20Shot%202024-04-23%20at%209.54.35%20AM.png?generation=1713846717396710&amp;alt=media\">  </p>",
      "rawMarkdown": "One of the interesting things of this competition is *it's larger than life* dataset and dealing with it (going to come with a lot of learnings along the way). \nIf you're struggling with the very first step of loading this dataset using pandas (ikr), switch to dask to load the parquet files. Here's the code-snippet attached on how to do it. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2875376%2F147594485803a0e2bab182f65e2992f4%2FScreen%20Shot%202024-04-23%20at%209.54.35%20AM.png?generation=1713846717396710&alt=media)",
      "votes": null
    },
    {
      "id": "2771666",
      "postDate": "04/24/2024 10:55:24",
      "content": "<p>try using polars for kaggle notebooks  <a href=\"url\" target=\"_blank\">https://docs.pola.rs/py-polars/html/reference/api/polars.read_parquet.html</a></p>\n<ul>\n<li><p>For in-memory data processing, especially for large datasets that fit in memory, and you value speed and ease of use: <strong>Choose Polars.</strong><br>\nPolars: Designed for multi-threaded parallelism, utilizing multiple CPU cores efficiently.</p></li>\n<li><p>For working with truly massive datasets that exceed available memory and require distributed computing: <strong>Choose Dask.</strong><br>\nDask: Supports both multi-threaded and distributed parallelism, allowing scaling out to multiple machines in a cluster.</p></li>\n</ul>\n<p>Note : <strong>Dask</strong>: Offers both multi-threaded and distributed parallelism. It can run computations on a single machine using threads or distribute them across multiple machines in a cluster.</p>",
      "rawMarkdown": "try using polars for kaggle notebooks  [https://docs.pola.rs/py-polars/html/reference/api/polars.read_parquet.html](url)\n\n- For in-memory data processing, especially for large datasets that fit in memory, and you value speed and ease of use: **Choose Polars.**\nPolars: Designed for multi-threaded parallelism, utilizing multiple CPU cores efficiently.\n\n- For working with truly massive datasets that exceed available memory and require distributed computing: **Choose Dask.**\nDask: Supports both multi-threaded and distributed parallelism, allowing scaling out to multiple machines in a cluster.\n\nNote : **Dask**: Offers both multi-threaded and distributed parallelism. It can run computations on a single machine using threads or distribute them across multiple machines in a cluster.",
      "votes": null
    },
    {
      "id": "2771893",
      "postDate": "04/24/2024 12:50:44",
      "content": "<p>this is great, thanks! </p>",
      "rawMarkdown": "this is great, thanks!",
      "votes": null
    },
    {
      "id": "2773514",
      "postDate": "04/24/2024 19:09:42",
      "content": "<p>anywhere know of a R variant of dask?</p>",
      "rawMarkdown": "anywhere know of a R variant of dask?",
      "votes": null
    },
    {
      "id": "2774748",
      "postDate": "04/25/2024 10:08:13",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/tuttlen\" target=\"_blank\">@tuttlen</a> <br>\ndisk.frame is what people use in R. You can look into it <a href=\"https://diskframe.com/\" target=\"_blank\">here.</a> </p>",
      "rawMarkdown": "Hey @tuttlen \ndisk.frame is what people use in R. You can look into it [here.](https://diskframe.com/)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2771666,
      "author_name": "kunduruanil",
      "author_url": "",
      "post_date": "04/24/2024 10:55:24",
      "content": "<p>try using polars for kaggle notebooks  <a href=\"url\" target=\"_blank\">https://docs.pola.rs/py-polars/html/reference/api/polars.read_parquet.html</a></p>\n<ul>\n<li><p>For in-memory data processing, especially for large datasets that fit in memory, and you value speed and ease of use: <strong>Choose Polars.</strong><br>\nPolars: Designed for multi-threaded parallelism, utilizing multiple CPU cores efficiently.</p></li>\n<li><p>For working with truly massive datasets that exceed available memory and require distributed computing: <strong>Choose Dask.</strong><br>\nDask: Supports both multi-threaded and distributed parallelism, allowing scaling out to multiple machines in a cluster.</p></li>\n</ul>\n<p>Note : <strong>Dask</strong>: Offers both multi-threaded and distributed parallelism. It can run computations on a single machine using threads or distribute them across multiple machines in a cluster.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2771893,
      "author_name": "samansalmasi",
      "author_url": "",
      "post_date": "04/24/2024 12:50:44",
      "content": "<p>this is great, thanks! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2773514,
      "author_name": "tuttlen",
      "author_url": "",
      "post_date": "04/24/2024 19:09:42",
      "content": "<p>anywhere know of a R variant of dask?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2774748,
          "author_name": "ahsuna123",
          "author_url": "",
          "post_date": "04/25/2024 10:08:13",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/tuttlen\" target=\"_blank\">@tuttlen</a> <br>\ndisk.frame is what people use in R. You can look into it <a href=\"https://diskframe.com/\" target=\"_blank\">here.</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2768877": "One of the interesting things of this competition is *it's larger than life* dataset and dealing with it (going to come with a lot of learnings along the way). \nIf you're struggling with the very first step of loading this dataset using pandas (ikr), switch to dask to load the parquet files. Here's the code-snippet attached on how to do it. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2875376%2F147594485803a0e2bab182f65e2992f4%2FScreen%20Shot%202024-04-23%20at%209.54.35%20AM.png?generation=1713846717396710&alt=media)",
    "2771666": "try using polars for kaggle notebooks  [https://docs.pola.rs/py-polars/html/reference/api/polars.read_parquet.html](url)\n\n- For in-memory data processing, especially for large datasets that fit in memory, and you value speed and ease of use: **Choose Polars.**\nPolars: Designed for multi-threaded parallelism, utilizing multiple CPU cores efficiently.\n\n- For working with truly massive datasets that exceed available memory and require distributed computing: **Choose Dask.**\nDask: Supports both multi-threaded and distributed parallelism, allowing scaling out to multiple machines in a cluster.\n\nNote : **Dask**: Offers both multi-threaded and distributed parallelism. It can run computations on a single machine using threads or distribute them across multiple machines in a cluster.",
    "2771893": "this is great, thanks!",
    "2773514": "anywhere know of a R variant of dask?",
    "2774748": "Hey @tuttlen \ndisk.frame is what people use in R. You can look into it [here.](https://diskframe.com/)"
  },
  "source": "meta"
}