{
  "id": 337055,
  "title": "Loading the data ",
  "url": "/competitions/amex-default-prediction/discussion/337055",
  "author_name": "",
  "post_date": "2022-07-14T08:41:53.668113600Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi , I am unable to load the data with pandas due to the huge size </p>\n<p>Could any help me in this regard…</p>\n<p>is there any alternate library to load the data…</p>",
  "messages": [
    {
      "id": "1855040",
      "postDate": "07/14/2022 08:41:53",
      "content": "<p>Hi , I am unable to load the data with pandas due to the huge size </p>\n<p>Could any help me in this regard…</p>\n<p>is there any alternate library to load the data…</p>",
      "rawMarkdown": "Hi , I am unable to load the data with pandas due to the huge size \n\nCould any help me in this regard...\n\nis there any alternate library to load the data...",
      "votes": null
    },
    {
      "id": "1855052",
      "postDate": "07/14/2022 08:50:29",
      "content": "<p><a href=\"https://www.kaggle.com/pnsaimanitejaswaroop\" target=\"_blank\">@pnsaimanitejaswaroop</a> Did you try chunksize parameter in Pandas read_csv? Here is where you can find the link to <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_csv.html\" target=\"_blank\">chunksize</a>.</p>\n<pre><code>chunksize = 10 ** 6\nfor chunk in pd.read_csv(filename, chunksize=chunksize):\n    process(chunk)\n</code></pre>",
      "rawMarkdown": "pnsaimanitejaswaroop Did you try chunksize parameter in Pandas read_csv? Here is where you can find the link to [chunksize](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_csv.html).\n```python\nchunksize = 10 ** 6\nfor chunk in pd.read_csv(filename, chunksize=chunksize):\n    process(chunk)\n```",
      "votes": null
    },
    {
      "id": "1855071",
      "postDate": "07/14/2022 09:03:07",
      "content": "<p>I recommend checking the discussion forum. Many threads related to this. You can also check my notebook to load data faster <a href=\"https://www.kaggle.com/code/naiborhujosua/loading-your-data-faster\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I recommend checking the discussion forum. Many threads related to this. You can also check my notebook to load data faster [here](https://www.kaggle.com/code/naiborhujosua/loading-your-data-faster)",
      "votes": null
    },
    {
      "id": "1855091",
      "postDate": "07/14/2022 09:22:04",
      "content": "<p>To be able to load the data with pandas, you first need to reduce its size. Chris Deotte described in his <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">discussion topic</a> how to do this. Alternatively, you can use a dataset that has already been preprocessed, such as <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">this</a> dataset from raddar.</p>",
      "rawMarkdown": "To be able to load the data with pandas, you first need to reduce its size. Chris Deotte described in his [discussion topic](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054) how to do this. Alternatively, you can use a dataset that has already been preprocessed, such as [this](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) dataset from raddar.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1855052,
      "author_name": "thanatoz",
      "author_url": "",
      "post_date": "07/14/2022 08:50:29",
      "content": "<p><a href=\"https://www.kaggle.com/pnsaimanitejaswaroop\" target=\"_blank\">@pnsaimanitejaswaroop</a> Did you try chunksize parameter in Pandas read_csv? Here is where you can find the link to <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_csv.html\" target=\"_blank\">chunksize</a>.</p>\n<pre><code>chunksize = 10 ** 6\nfor chunk in pd.read_csv(filename, chunksize=chunksize):\n    process(chunk)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1855071,
      "author_name": "naiborhujosua",
      "author_url": "",
      "post_date": "07/14/2022 09:03:07",
      "content": "<p>I recommend checking the discussion forum. Many threads related to this. You can also check my notebook to load data faster <a href=\"https://www.kaggle.com/code/naiborhujosua/loading-your-data-faster\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1855091,
      "author_name": "viktorfairuschin",
      "author_url": "",
      "post_date": "07/14/2022 09:22:04",
      "content": "<p>To be able to load the data with pandas, you first need to reduce its size. Chris Deotte described in his <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">discussion topic</a> how to do this. Alternatively, you can use a dataset that has already been preprocessed, such as <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">this</a> dataset from raddar.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1855040": "Hi , I am unable to load the data with pandas due to the huge size \n\nCould any help me in this regard...\n\nis there any alternate library to load the data...",
    "1855052": "pnsaimanitejaswaroop Did you try chunksize parameter in Pandas read_csv? Here is where you can find the link to [chunksize](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_csv.html).\n```python\nchunksize = 10 ** 6\nfor chunk in pd.read_csv(filename, chunksize=chunksize):\n    process(chunk)\n```",
    "1855071": "I recommend checking the discussion forum. Many threads related to this. You can also check my notebook to load data faster [here](https://www.kaggle.com/code/naiborhujosua/loading-your-data-faster)",
    "1855091": "To be able to load the data with pandas, you first need to reduce its size. Chris Deotte described in his [discussion topic](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054) how to do this. Alternatively, you can use a dataset that has already been preprocessed, such as [this](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format) dataset from raddar."
  },
  "source": "meta"
}