{
  "id": 333917,
  "title": "Taking too long to turn test.csv into a dataframe",
  "url": "/competitions/amex-default-prediction/discussion/333917",
  "author_name": "",
  "post_date": "2022-06-29T00:21:39.727965500Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Pandas is taking too long to turn a CSV containing 924621 rows into a dataframe. How are you dealing with this issue?</p>",
  "messages": [
    {
      "id": "1836665",
      "postDate": "06/29/2022 00:21:39",
      "content": "<p>Pandas is taking too long to turn a CSV containing 924621 rows into a dataframe. How are you dealing with this issue?</p>",
      "rawMarkdown": "Pandas is taking too long to turn a CSV containing 924621 rows into a dataframe. How are you dealing with this issue?",
      "votes": null
    },
    {
      "id": "1836703",
      "postDate": "06/29/2022 01:54:05",
      "content": "<p>You can load it in chunks by using the parameter \"chunksize\".</p>",
      "rawMarkdown": "You can load it in chunks by using the parameter \"chunksize\".",
      "votes": null
    },
    {
      "id": "1836728",
      "postDate": "06/29/2022 02:26:27",
      "content": "<h2><strong>Using Dask</strong></h2>\n<p>Dask is an open-source python library that includes features of parallelism and scalability in Python by using the existing libraries like pandas, NumPy, or sklearn.</p>\n<h3><strong>To Install</strong></h3>\n<p><code>!pip install dask</code></p>\n<h3><strong>Usage of Dask</strong></h3>\n<pre><code># import required modules\nimport pandas as pd\nimport numpy as np\nimport time\nfrom dask import dataframe as df1\n\ndask_df = df1.read_csv('test.csv')\n\n# data\ndask_df.head(10)\n</code></pre>\n<ul>\n<li>Dask can enable efficient parallel computations on single machines by leveraging their multi-core CPUs and streaming data efficiently from disk.</li>\n</ul>",
      "rawMarkdown": "## **Using Dask**\n\nDask is an open-source python library that includes features of parallelism and scalability in Python by using the existing libraries like pandas, NumPy, or sklearn.\n\n### **To Install**\n`!pip install dask`\n\n### **Usage of Dask**\n```\n# import required modules\nimport pandas as pd\nimport numpy as np\nimport time\nfrom dask import dataframe as df1\n\ndask_df = df1.read_csv('test.csv')\n\n# data\ndask_df.head(10)\n```\n\n- Dask can enable efficient parallel computations on single machines by leveraging their multi-core CPUs and streaming data efficiently from disk.",
      "votes": null
    },
    {
      "id": "1837077",
      "postDate": "06/29/2022 10:12:25",
      "content": "<p>there are a bunch of good discussions in the competition forum, and several already-compressed datasets available. In particular <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> did some cool work in denoising the data in this dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></p>\n<p>Discussions:<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327205\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327205</a><br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143</a><br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514</a><br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054</a></p>",
      "rawMarkdown": "there are a bunch of good discussions in the competition forum, and several already-compressed datasets available. In particular @raddar did some cool work in denoising the data in this dataset:\nhttps://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n\nDiscussions:\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327205\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328054",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1836703,
      "author_name": "hskhawaja",
      "author_url": "",
      "post_date": "06/29/2022 01:54:05",
      "content": "<p>You can load it in chunks by using the parameter \"chunksize\".</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1836728,
      "author_name": "muki2003",
      "author_url": "",
      "post_date": "06/29/2022 02:26:27",
      "content": "<h2><strong>Using Dask</strong></h2>\n<p>Dask is an open-source python library that includes features of parallelism and scalability in Python by using the existing libraries like pandas, NumPy, or sklearn.</p>\n<h3><strong>To Install</strong></h3>\n<p><code>!pip install dask</code></p>\n<h3><strong>Usage of Dask</strong></h3>\n<pre><code># import required modules\nimport pandas as pd\nimport numpy as np\nimport time\nfrom dask import dataframe as df1\n\ndask_df = df1.read_csv('test.csv')\n\n# data\ndask_df.head(10)\n</code></pre>\n<ul>\n<li>Dask can enable efficient parallel computations on single machines by leveraging their multi-core CPUs and streaming data efficiently from disk.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1837077,
      "author_name": "eduus710",
      "author_url": "",
      "post_date": "06/29/2022 10:12:25",
      "content": "<p>there are a bunch of good discussions in the competition forum, and several already-compressed datasets available. In particular <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> did some cool work in denoising the data in this dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></p>\n<p>Discussions:<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327205\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327205</a><br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327143</a><br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514</a><br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1836665": "Pandas is taking too long to turn a CSV containing 924621 rows into a dataframe. How are you dealing with this issue?",
    "1836703": "You can load it in chunks by using the parameter \"chunksize\".",
    "1836728": "## **Using Dask**\n\nDask is an open-source python library that includes features of parallelism and scalability in Python by using the existing libraries like pandas, NumPy, or sklearn.\n\n### **To Install**\n`!pip install dask`\n\n### **Usage of Dask**\n```\n# import required modules\nimport pandas as pd\nimport numpy as np\nimport time\nfrom dask import dataframe as df1\n\ndask_df = df1.read_csv('test.csv')\n\n# data\ndask_df.head(10)\n```\n\n- Dask can enable efficient parallel computations on single machines by leveraging their multi-core CPUs and streaming data efficiently from disk.",
    "1837077": "there are a bunch of good discussions in the competition forum, and several already-compressed datasets available. In particular @raddar did some cool work in denoising the data in this dataset:\nhttps://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n\nDiscussions:\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327205\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/327143\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328054"
  },
  "source": "meta"
}