{
  "id": 540551,
  "title": "Using dask to handle this large dataset.",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/540551",
  "author_name": "",
  "post_date": "2024-10-15T03:35:32.533789100Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Given, how large the dataset is, it's obvious to encounter OOM error when trying to load the entire dataset in one go. If you're familiar with LeashBio competition that was hosted few months back, I opened a discussion thread on using dask to deal with larger than life datasets. And you can see the code implementation in the same discussion thread: <br>\n<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496954\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/496954</a></p>\n<p>For this competition, it took me only 5.31 sec to load the entire dataset.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2875376%2Fc000f8d6bb1243da7eb235de17cbdbd9%2FScreen%20Shot%202024-10-15%20at%209.03.16%20AM.png?generation=1728963235861626&amp;alt=media\" alt=\"\"> </p>",
  "messages": [
    {
      "id": "3017608",
      "postDate": "10/15/2024 03:35:32",
      "content": "<p>Given, how large the dataset is, it's obvious to encounter OOM error when trying to load the entire dataset in one go. If you're familiar with LeashBio competition that was hosted few months back, I opened a discussion thread on using dask to deal with larger than life datasets. And you can see the code implementation in the same discussion thread: <br>\n<a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/496954\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/496954</a></p>\n<p>For this competition, it took me only 5.31 sec to load the entire dataset.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2875376%2Fc000f8d6bb1243da7eb235de17cbdbd9%2FScreen%20Shot%202024-10-15%20at%209.03.16%20AM.png?generation=1728963235861626&amp;alt=media\" alt=\"\"> </p>",
      "rawMarkdown": "Given, how large the dataset is, it's obvious to encounter OOM error when trying to load the entire dataset in one go. If you're familiar with LeashBio competition that was hosted few months back, I opened a discussion thread on using dask to deal with larger than life datasets. And you can see the code implementation in the same discussion thread: \nhttps://www.kaggle.com/competitions/leash-BELKA/discussion/496954\n\nFor this competition, it took me only 5.31 sec to load the entire dataset.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2875376%2Fc000f8d6bb1243da7eb235de17cbdbd9%2FScreen%20Shot%202024-10-15%20at%209.03.16%20AM.png?generation=1728963235861626&alt=media)",
      "votes": null
    },
    {
      "id": "3018328",
      "postDate": "10/15/2024 17:14:21",
      "content": "<p>It took 5.31s to load the first few rows. Dask is similar to spark in the sense that it does lazy compute - it doesn't load anything into memory until it's needed.</p>",
      "rawMarkdown": "It took 5.31s to load the first few rows. Dask is similar to spark in the sense that it does lazy compute - it doesn't load anything into memory until it's needed.",
      "votes": null
    },
    {
      "id": "3018727",
      "postDate": "10/16/2024 02:45:19",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": null
    },
    {
      "id": "3020354",
      "postDate": "10/17/2024 12:53:45",
      "content": "<p>It sounds better to use polars its written based on Rust.</p>",
      "rawMarkdown": "It sounds better to use polars its written based on Rust.",
      "votes": null
    },
    {
      "id": "3073012",
      "postDate": "12/15/2024 23:11:20",
      "content": "<p>however, when you load it into the memory, the memory explodes again.</p>",
      "rawMarkdown": "however, when you load it into the memory, the memory explodes again.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3018328,
      "author_name": "madcarrot",
      "author_url": "",
      "post_date": "10/15/2024 17:14:21",
      "content": "<p>It took 5.31s to load the first few rows. Dask is similar to spark in the sense that it does lazy compute - it doesn't load anything into memory until it's needed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3018727,
      "author_name": "zeminglin",
      "author_url": "",
      "post_date": "10/16/2024 02:45:19",
      "content": "<p>Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3020354,
      "author_name": "farhankardan",
      "author_url": "",
      "post_date": "10/17/2024 12:53:45",
      "content": "<p>It sounds better to use polars its written based on Rust.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3073012,
      "author_name": "billyrobert",
      "author_url": "",
      "post_date": "12/15/2024 23:11:20",
      "content": "<p>however, when you load it into the memory, the memory explodes again.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3017608": "Given, how large the dataset is, it's obvious to encounter OOM error when trying to load the entire dataset in one go. If you're familiar with LeashBio competition that was hosted few months back, I opened a discussion thread on using dask to deal with larger than life datasets. And you can see the code implementation in the same discussion thread: \nhttps://www.kaggle.com/competitions/leash-BELKA/discussion/496954\n\nFor this competition, it took me only 5.31 sec to load the entire dataset.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2875376%2Fc000f8d6bb1243da7eb235de17cbdbd9%2FScreen%20Shot%202024-10-15%20at%209.03.16%20AM.png?generation=1728963235861626&alt=media)",
    "3018328": "It took 5.31s to load the first few rows. Dask is similar to spark in the sense that it does lazy compute - it doesn't load anything into memory until it's needed.",
    "3018727": "Great work!",
    "3020354": "It sounds better to use polars its written based on Rust.",
    "3073012": "however, when you load it into the memory, the memory explodes again."
  },
  "source": "meta"
}