{
  "id": 327110,
  "title": "Handling large datasets with Dask",
  "url": "/competitions/amex-default-prediction/discussion/327110",
  "author_name": "",
  "post_date": "2022-05-25T16:30:10.712556400Z",
  "votes": 33,
  "comment_count": 3,
  "views": 0,
  "content": "<blockquote>\n  <p>Dask provides advanced parallelism for analytics, enabling performance at scale for the tools you love.</p>\n  <p>Dask is open source and freely available. It is developed in coordination with other community projects like NumPy, pandas, and scikit-learn.</p>\n</blockquote>\n<p><a href=\"https://dask.org/\" target=\"_blank\">Dask official website</a></p>\n<h2>Installing Dask</h2>\n<p><strong>Pip</strong></p>\n<pre><code>pip install \"dask[complete]\"\n</code></pre>\n<p><strong>Conda</strong></p>\n<pre><code>conda install dask\n</code></pre>\n<h2>DataFrame</h2>\n<blockquote>\n  <p>A Dask DataFrame is a large parallel DataFrame composed of many smaller Pandas DataFrames, split along the index. These Pandas DataFrames may live on disk for larger-than-memory computing on a single machine or many different machines in a cluster. One Dask DataFrame operation triggers many operations on the constituent Pandas DataFrames.</p>\n</blockquote>\n<p><a href=\"https://docs.dask.org/en/latest/dataframe.html\" target=\"_blank\">Reference</a></p>\n<h2>Using dask</h2>\n<pre><code>import dask.dataframe as dd\n\n# Dask dataframe\nddf = dd.read_csv(\"./input/train_data.csv\")\nprint(df.head())\n\n# customer_ID and S_2 are columns in our dataset\ndcmd = ddf.groupby(ddf.customer_ID).S_2.count()\nresult = dcmd.compute()\n\n# Number of unique customer\nprint(result.shape)\n\n# More details (In this case this is a pandas.Series object)\nprint(result.head())\n</code></pre>",
  "messages": [
    {
      "id": "1801349",
      "postDate": "05/25/2022 16:30:10",
      "content": "<blockquote>\n  <p>Dask provides advanced parallelism for analytics, enabling performance at scale for the tools you love.</p>\n  <p>Dask is open source and freely available. It is developed in coordination with other community projects like NumPy, pandas, and scikit-learn.</p>\n</blockquote>\n<p><a href=\"https://dask.org/\" target=\"_blank\">Dask official website</a></p>\n<h2>Installing Dask</h2>\n<p><strong>Pip</strong></p>\n<pre><code>pip install \"dask[complete]\"\n</code></pre>\n<p><strong>Conda</strong></p>\n<pre><code>conda install dask\n</code></pre>\n<h2>DataFrame</h2>\n<blockquote>\n  <p>A Dask DataFrame is a large parallel DataFrame composed of many smaller Pandas DataFrames, split along the index. These Pandas DataFrames may live on disk for larger-than-memory computing on a single machine or many different machines in a cluster. One Dask DataFrame operation triggers many operations on the constituent Pandas DataFrames.</p>\n</blockquote>\n<p><a href=\"https://docs.dask.org/en/latest/dataframe.html\" target=\"_blank\">Reference</a></p>\n<h2>Using dask</h2>\n<pre><code>import dask.dataframe as dd\n\n# Dask dataframe\nddf = dd.read_csv(\"./input/train_data.csv\")\nprint(df.head())\n\n# customer_ID and S_2 are columns in our dataset\ndcmd = ddf.groupby(ddf.customer_ID).S_2.count()\nresult = dcmd.compute()\n\n# Number of unique customer\nprint(result.shape)\n\n# More details (In this case this is a pandas.Series object)\nprint(result.head())\n</code></pre>",
      "rawMarkdown": "> Dask provides advanced parallelism for analytics, enabling performance at scale for the tools you love.\n>\n> Dask is open source and freely available. It is developed in coordination with other community projects like NumPy, pandas, and scikit-learn.\n\n[Dask official website](https://dask.org/)\n\n## Installing Dask\n**Pip**\n```\npip install \"dask[complete]\"\n```\n\n**Conda**\n```\nconda install dask\n```\n\n\n## DataFrame\n\n> A Dask DataFrame is a large parallel DataFrame composed of many smaller Pandas DataFrames, split along the index. These Pandas DataFrames may live on disk for larger-than-memory computing on a single machine or many different machines in a cluster. One Dask DataFrame operation triggers many operations on the constituent Pandas DataFrames.\n\n[Reference](https://docs.dask.org/en/latest/dataframe.html)\n\n## Using dask\n```python\nimport dask.dataframe as dd\n\n# Dask dataframe\nddf = dd.read_csv(\"./input/train_data.csv\")\nprint(df.head())\n\n# customer_ID and S_2 are columns in our dataset\ndcmd = ddf.groupby(ddf.customer_ID).S_2.count()\nresult = dcmd.compute()\n\n# Number of unique customer\nprint(result.shape)\n\n# More details (In this case this is a pandas.Series object)\nprint(result.head())\n```",
      "votes": null
    },
    {
      "id": "1802195",
      "postDate": "05/26/2022 14:40:58",
      "content": "<p>Thanks for sharing!👍</p>",
      "rawMarkdown": "Thanks for sharing!👍",
      "votes": null
    },
    {
      "id": "1850740",
      "postDate": "07/10/2022 17:09:24",
      "content": "<p>It's print(ddf.head()) ?</p>",
      "rawMarkdown": "It's print(ddf.head()) ?",
      "votes": null
    },
    {
      "id": "2068113",
      "postDate": "12/17/2022 14:22:06",
      "content": "<p>Here is a notebook on how to deal with <a href=\"https://www.kaggle.com/code/ezzzio/large-datasets-on-limited-hardware\" target=\"_blank\">large datasets of images on limited hardware</a></p>",
      "rawMarkdown": "Here is a notebook on how to deal with [large datasets of images on limited hardware](https://www.kaggle.com/code/ezzzio/large-datasets-on-limited-hardware)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1802195,
      "author_name": "gomohit",
      "author_url": "",
      "post_date": "05/26/2022 14:40:58",
      "content": "<p>Thanks for sharing!👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1850740,
      "author_name": "amellalibrahim",
      "author_url": "",
      "post_date": "07/10/2022 17:09:24",
      "content": "<p>It's print(ddf.head()) ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2068113,
      "author_name": "ezzzio",
      "author_url": "",
      "post_date": "12/17/2022 14:22:06",
      "content": "<p>Here is a notebook on how to deal with <a href=\"https://www.kaggle.com/code/ezzzio/large-datasets-on-limited-hardware\" target=\"_blank\">large datasets of images on limited hardware</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801349": "> Dask provides advanced parallelism for analytics, enabling performance at scale for the tools you love.\n>\n> Dask is open source and freely available. It is developed in coordination with other community projects like NumPy, pandas, and scikit-learn.\n\n[Dask official website](https://dask.org/)\n\n## Installing Dask\n**Pip**\n```\npip install \"dask[complete]\"\n```\n\n**Conda**\n```\nconda install dask\n```\n\n\n## DataFrame\n\n> A Dask DataFrame is a large parallel DataFrame composed of many smaller Pandas DataFrames, split along the index. These Pandas DataFrames may live on disk for larger-than-memory computing on a single machine or many different machines in a cluster. One Dask DataFrame operation triggers many operations on the constituent Pandas DataFrames.\n\n[Reference](https://docs.dask.org/en/latest/dataframe.html)\n\n## Using dask\n```python\nimport dask.dataframe as dd\n\n# Dask dataframe\nddf = dd.read_csv(\"./input/train_data.csv\")\nprint(df.head())\n\n# customer_ID and S_2 are columns in our dataset\ndcmd = ddf.groupby(ddf.customer_ID).S_2.count()\nresult = dcmd.compute()\n\n# Number of unique customer\nprint(result.shape)\n\n# More details (In this case this is a pandas.Series object)\nprint(result.head())\n```",
    "1802195": "Thanks for sharing!👍",
    "1850740": "It's print(ddf.head()) ?",
    "2068113": "Here is a notebook on how to deal with [large datasets of images on limited hardware](https://www.kaggle.com/code/ezzzio/large-datasets-on-limited-hardware)"
  },
  "source": "meta"
}