{
  "id": 334886,
  "title": "Reading the CSV file",
  "url": "/competitions/amex-default-prediction/discussion/334886",
  "author_name": "",
  "post_date": "2022-07-03T15:06:53.393255Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello Everyone,</p>\n<p>I have a rather basic question. My notebook crashes when trying to read the train CSV file itself due to RAM overload possibly because of size of data.</p>\n<p>I did happen to read on forums that it is best to convert it to parquet files. Now I wanted to understand how can I integrate the parquet files in my notebook? I can see some folks have already converted and hosted on datasets (e.g. - <a href=\"https://www.kaggle.com/datasets/helendace/american-express-parquet\" target=\"_blank\">https://www.kaggle.com/datasets/helendace/american-express-parquet</a>)</p>\n<p>I just wanted to understand how can I leverage that into my notebook.</p>\n<p>Thanks</p>\n<p>Edit: Thanks everyone..I was able to do this thanks to help from everybody here.</p>",
  "messages": [
    {
      "id": "1841943",
      "postDate": "07/03/2022 15:06:53",
      "content": "<p>Hello Everyone,</p>\n<p>I have a rather basic question. My notebook crashes when trying to read the train CSV file itself due to RAM overload possibly because of size of data.</p>\n<p>I did happen to read on forums that it is best to convert it to parquet files. Now I wanted to understand how can I integrate the parquet files in my notebook? I can see some folks have already converted and hosted on datasets (e.g. - <a href=\"https://www.kaggle.com/datasets/helendace/american-express-parquet\" target=\"_blank\">https://www.kaggle.com/datasets/helendace/american-express-parquet</a>)</p>\n<p>I just wanted to understand how can I leverage that into my notebook.</p>\n<p>Thanks</p>\n<p>Edit: Thanks everyone..I was able to do this thanks to help from everybody here.</p>",
      "rawMarkdown": "Hello Everyone,\n\nI have a rather basic question. My notebook crashes when trying to read the train CSV file itself due to RAM overload possibly because of size of data.\n\nI did happen to read on forums that it is best to convert it to parquet files. Now I wanted to understand how can I integrate the parquet files in my notebook? I can see some folks have already converted and hosted on datasets (e.g. - https://www.kaggle.com/datasets/helendace/american-express-parquet)\n\nI just wanted to understand how can I leverage that into my notebook.\n\nThanks\n\nEdit: Thanks everyone..I was able to do this thanks to help from everybody here.",
      "votes": null
    },
    {
      "id": "1841967",
      "postDate": "07/03/2022 15:28:15",
      "content": "<p>The easiest thing is to use Raddar's dataset <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>. Attach his dataset to your notebook. Then instead of <code>pd.read_csv(KAGGLE_CSV)</code> just use <code>pd.read_parquet(RADDAR_PARQUET)</code> instead. The resultant dataframe will have the same rows and same columns. And Raddar's parquet will not crash your memory.</p>",
      "rawMarkdown": "The easiest thing is to use Raddar's dataset [here][1]. Attach his dataset to your notebook. Then instead of `pd.read_csv(KAGGLE_CSV)` just use `pd.read_parquet(RADDAR_PARQUET)` instead. The resultant dataframe will have the same rows and same columns. And Raddar's parquet will not crash your memory.\n\n[1]: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
      "votes": null
    },
    {
      "id": "1841968",
      "postDate": "07/03/2022 15:30:28",
      "content": "<p>Add one of the hosted data sets into your notebook ( + Add Data).   You can copy the file name and folder address and just use a pd.read_parquet(xxx) to load the smaller file into your notebook.</p>",
      "rawMarkdown": "Add one of the hosted data sets into your notebook ( + Add Data).   You can copy the file name and folder address and just use a pd.read_parquet(xxx) to load the smaller file into your notebook.",
      "votes": null
    },
    {
      "id": "1842088",
      "postDate": "07/03/2022 17:37:52",
      "content": "<p>Read the parquet file using pandas read_parquet command and you will be fine</p>",
      "rawMarkdown": "Read the parquet file using pandas read_parquet command and you will be fine",
      "votes": null
    },
    {
      "id": "1842711",
      "postDate": "07/04/2022 08:24:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jayeshgokhale\" target=\"_blank\">@jayeshgokhale</a> It was same with me and below mentioned users have given me instructions but to add converted parquet data to your notebook was something which i had concerns so here is the solution. When you go to parquet notebook (through your mentioned link ) click on new notebook and in that you can see it would have been added in dataset section right top corner. Let me know if you have any questions. Thanks</p>",
      "rawMarkdown": "Hi @jayeshgokhale It was same with me and below mentioned users have given me instructions but to add converted parquet data to your notebook was something which i had concerns so here is the solution. When you go to parquet notebook (through your mentioned link ) click on new notebook and in that you can see it would have been added in dataset section right top corner. Let me know if you have any questions. Thanks",
      "votes": null
    },
    {
      "id": "1843621",
      "postDate": "07/05/2022 03:44:33",
      "content": "<p>Thanks everyone..I was able to add parquet dataset from <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> Thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for that too. Actually what I missed earlier from the \"Add Data\" option that there was a section to add URL as well instead of files. I guess I just need to keep my eyes open :)</p>",
      "rawMarkdown": "Thanks everyone..I was able to add parquet dataset from @raddar Thanks to @raddar for that too. Actually what I missed earlier from the \"Add Data\" option that there was a section to add URL as well instead of files. I guess I just need to keep my eyes open :)",
      "votes": null
    },
    {
      "id": "1848776",
      "postDate": "07/08/2022 22:11:59",
      "content": "<p>This worked perfectly. Thank you for asking, <a href=\"https://www.kaggle.com/jayeshgokhale\" target=\"_blank\">@jayeshgokhale</a> - it saved me a lot of time and aggravation. </p>",
      "rawMarkdown": "This worked perfectly. Thank you for asking, @jayeshgokhale - it saved me a lot of time and aggravation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1841967,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/03/2022 15:28:15",
      "content": "<p>The easiest thing is to use Raddar's dataset <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>. Attach his dataset to your notebook. Then instead of <code>pd.read_csv(KAGGLE_CSV)</code> just use <code>pd.read_parquet(RADDAR_PARQUET)</code> instead. The resultant dataframe will have the same rows and same columns. And Raddar's parquet will not crash your memory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1841968,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/03/2022 15:30:28",
      "content": "<p>Add one of the hosted data sets into your notebook ( + Add Data).   You can copy the file name and folder address and just use a pd.read_parquet(xxx) to load the smaller file into your notebook.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1842088,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/03/2022 17:37:52",
      "content": "<p>Read the parquet file using pandas read_parquet command and you will be fine</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1842711,
      "author_name": "nknarendra7",
      "author_url": "",
      "post_date": "07/04/2022 08:24:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jayeshgokhale\" target=\"_blank\">@jayeshgokhale</a> It was same with me and below mentioned users have given me instructions but to add converted parquet data to your notebook was something which i had concerns so here is the solution. When you go to parquet notebook (through your mentioned link ) click on new notebook and in that you can see it would have been added in dataset section right top corner. Let me know if you have any questions. Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1843621,
      "author_name": "jayeshgokhale",
      "author_url": "",
      "post_date": "07/05/2022 03:44:33",
      "content": "<p>Thanks everyone..I was able to add parquet dataset from <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> Thanks to <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for that too. Actually what I missed earlier from the \"Add Data\" option that there was a section to add URL as well instead of files. I guess I just need to keep my eyes open :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1848776,
      "author_name": "megan3",
      "author_url": "",
      "post_date": "07/08/2022 22:11:59",
      "content": "<p>This worked perfectly. Thank you for asking, <a href=\"https://www.kaggle.com/jayeshgokhale\" target=\"_blank\">@jayeshgokhale</a> - it saved me a lot of time and aggravation. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1841943": "Hello Everyone,\n\nI have a rather basic question. My notebook crashes when trying to read the train CSV file itself due to RAM overload possibly because of size of data.\n\nI did happen to read on forums that it is best to convert it to parquet files. Now I wanted to understand how can I integrate the parquet files in my notebook? I can see some folks have already converted and hosted on datasets (e.g. - https://www.kaggle.com/datasets/helendace/american-express-parquet)\n\nI just wanted to understand how can I leverage that into my notebook.\n\nThanks\n\nEdit: Thanks everyone..I was able to do this thanks to help from everybody here.",
    "1841967": "The easiest thing is to use Raddar's dataset [here][1]. Attach his dataset to your notebook. Then instead of `pd.read_csv(KAGGLE_CSV)` just use `pd.read_parquet(RADDAR_PARQUET)` instead. The resultant dataframe will have the same rows and same columns. And Raddar's parquet will not crash your memory.\n\n[1]: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
    "1841968": "Add one of the hosted data sets into your notebook ( + Add Data).   You can copy the file name and folder address and just use a pd.read_parquet(xxx) to load the smaller file into your notebook.",
    "1842088": "Read the parquet file using pandas read_parquet command and you will be fine",
    "1842711": "Hi @jayeshgokhale It was same with me and below mentioned users have given me instructions but to add converted parquet data to your notebook was something which i had concerns so here is the solution. When you go to parquet notebook (through your mentioned link ) click on new notebook and in that you can see it would have been added in dataset section right top corner. Let me know if you have any questions. Thanks",
    "1843621": "Thanks everyone..I was able to add parquet dataset from @raddar Thanks to @raddar for that too. Actually what I missed earlier from the \"Add Data\" option that there was a section to add URL as well instead of files. I guess I just need to keep my eyes open :)",
    "1848776": "This worked perfectly. Thank you for asking, @jayeshgokhale - it saved me a lot of time and aggravation."
  },
  "source": "meta"
}