{
  "id": 190120,
  "title": "Options to read a large csv file",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190120",
  "author_name": "Subbu Vidyasekar",
  "post_date": "2020-10-10T08:48:34.225000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello Kagglers, </p>\n<p>Reading a huge file in pandas is not adviseable as per the Expert.</p>\n<p>There are multiple options to read a huge csv file</p>\n<ol>\n<li>csv.DictReader</li>\n<li>pandas.read_csv</li>\n<li>pandas.read_csv using Chunks</li>\n<li>Dask</li>\n<li>Datatable </li>\n</ol>\n<p>Here is a comparison of the above options<br>\n<a href=\"https://medium.com/casual-inference/the-most-time-efficient-ways-to-import-csv-data-in-python-cc159b44063d\" target=\"_blank\">https://medium.com/casual-inference/the-most-time-efficient-ways-to-import-csv-data-in-python-cc159b44063d</a></p>\n<p>The final recommendation would be either <strong>Dask</strong> or <strong>Datatable</strong>.</p>\n<p>Those who familiar with Pandas:<br>\nDask will return pandas dataframe or series when it called with <code>.compute()</code></p>",
  "messages": [
    {
      "id": 1044959,
      "postDate": "2020-10-10T08:48:34.227Z",
      "content": "<p>Hello Kagglers, </p>\n<p>Reading a huge file in pandas is not adviseable as per the Expert.</p>\n<p>There are multiple options to read a huge csv file</p>\n<ol>\n<li>csv.DictReader</li>\n<li>pandas.read_csv</li>\n<li>pandas.read_csv using Chunks</li>\n<li>Dask</li>\n<li>Datatable </li>\n</ol>\n<p>Here is a comparison of the above options<br>\n<a href=\"https://medium.com/casual-inference/the-most-time-efficient-ways-to-import-csv-data-in-python-cc159b44063d\" target=\"_blank\">https://medium.com/casual-inference/the-most-time-efficient-ways-to-import-csv-data-in-python-cc159b44063d</a></p>\n<p>The final recommendation would be either <strong>Dask</strong> or <strong>Datatable</strong>.</p>\n<p>Those who familiar with Pandas:<br>\nDask will return pandas dataframe or series when it called with <code>.compute()</code></p>",
      "rawMarkdown": "Hello Kagglers, \n\nReading a huge file in pandas is not adviseable as per the Expert.\n\nThere are multiple options to read a huge csv file\n1. csv.DictReader\n2. pandas.read_csv\n3. pandas.read_csv using Chunks\n4. Dask\n5. Datatable \n\nHere is a comparison of the above options\nhttps://medium.com/casual-inference/the-most-time-efficient-ways-to-import-csv-data-in-python-cc159b44063d\n\nThe final recommendation would be either **Dask** or **Datatable**.\n\nThose who familiar with Pandas:\nDask will return pandas dataframe or series when it called with ```.compute()```",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1044959": "Hello Kagglers, \n\nReading a huge file in pandas is not adviseable as per the Expert.\n\nThere are multiple options to read a huge csv file\n1. csv.DictReader\n2. pandas.read_csv\n3. pandas.read_csv using Chunks\n4. Dask\n5. Datatable \n\nHere is a comparison of the above options\nhttps://medium.com/casual-inference/the-most-time-efficient-ways-to-import-csv-data-in-python-cc159b44063d\n\nThe final recommendation would be either **Dask** or **Datatable**.\n\nThose who familiar with Pandas:\nDask will return pandas dataframe or series when it called with ```.compute()```"
  }
}