{
  "id": 335629,
  "title": "Reading data through parquet is also memory constraints?",
  "url": "/competitions/amex-default-prediction/discussion/335629",
  "author_name": "",
  "post_date": "2022-07-07T05:23:15.167772200Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi Everyone,<br>\nI get converted data into parquet but still my jupiter notebook memory is getting exhausted frequently while doing simple preprocessing and notebook keeps on restarting. Do we have any options here to tackle such kind of issues? </p>\n<p>For converting csv to parquet info is here.<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/334956#1842426\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/334956#1842426</a></p>",
  "messages": [
    {
      "id": "1846477",
      "postDate": "07/07/2022 05:23:15",
      "content": "<p>Hi Everyone,<br>\nI get converted data into parquet but still my jupiter notebook memory is getting exhausted frequently while doing simple preprocessing and notebook keeps on restarting. Do we have any options here to tackle such kind of issues? </p>\n<p>For converting csv to parquet info is here.<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/334956#1842426\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/334956#1842426</a></p>",
      "rawMarkdown": "Hi Everyone,\nI get converted data into parquet but still my jupiter notebook memory is getting exhausted frequently while doing simple preprocessing and notebook keeps on restarting. Do we have any options here to tackle such kind of issues? \n\nFor converting csv to parquet info is here.\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/334956#1842426",
      "votes": null
    },
    {
      "id": "1846612",
      "postDate": "07/07/2022 08:00:39",
      "content": "<p><a href=\"https://www.kaggle.com/nknarendra7\" target=\"_blank\">@nknarendra7</a> Here you need two things. </p>\n<p>First, size reduced version of datasets. Second, use libraries like DASK that can efficiently manage large datasets.</p>\n<p>You can have a look at my notebook \"<a href=\"https://www.kaggle.com/code/mirfanazam/large-dataset-csv-dask-parquet\" target=\"_blank\">Large Dataset - CSV - DASK - Parquet</a>\"</p>",
      "rawMarkdown": "nknarendra7 Here you need two things. \n\nFirst, size reduced version of datasets. Second, use libraries like DASK that can efficiently manage large datasets.\n\nYou can have a look at my notebook \"[Large Dataset - CSV - DASK - Parquet](https://www.kaggle.com/code/mirfanazam/large-dataset-csv-dask-parquet)\"",
      "votes": null
    },
    {
      "id": "1846672",
      "postDate": "07/07/2022 09:16:16",
      "content": "<p>Hi, I saw your code. It was great and thank you for sharing.<br>\nI just wonder if I have converted data and store it on the notebook, do I have to convert and save every time I reconnect to kernel or is there some way that I save them permanently?<br>\nDownload the saved data and add to the notebook whenever kernel start newly is not a considered option. Because if it needs to be uploaded every time then it cannot be used for submit though (unless I submit the file solely)</p>",
      "rawMarkdown": "Hi, I saw your code. It was great and thank you for sharing.\nI just wonder if I have converted data and store it on the notebook, do I have to convert and save every time I reconnect to kernel or is there some way that I save them permanently?\nDownload the saved data and add to the notebook whenever kernel start newly is not a considered option. Because if it needs to be uploaded every time then it cannot be used for submit though (unless I submit the file solely)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1846612,
      "author_name": "mirfanazam",
      "author_url": "",
      "post_date": "07/07/2022 08:00:39",
      "content": "<p><a href=\"https://www.kaggle.com/nknarendra7\" target=\"_blank\">@nknarendra7</a> Here you need two things. </p>\n<p>First, size reduced version of datasets. Second, use libraries like DASK that can efficiently manage large datasets.</p>\n<p>You can have a look at my notebook \"<a href=\"https://www.kaggle.com/code/mirfanazam/large-dataset-csv-dask-parquet\" target=\"_blank\">Large Dataset - CSV - DASK - Parquet</a>\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 1846672,
          "author_name": "gilgarad",
          "author_url": "",
          "post_date": "07/07/2022 09:16:16",
          "content": "<p>Hi, I saw your code. It was great and thank you for sharing.<br>\nI just wonder if I have converted data and store it on the notebook, do I have to convert and save every time I reconnect to kernel or is there some way that I save them permanently?<br>\nDownload the saved data and add to the notebook whenever kernel start newly is not a considered option. Because if it needs to be uploaded every time then it cannot be used for submit though (unless I submit the file solely)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1846477": "Hi Everyone,\nI get converted data into parquet but still my jupiter notebook memory is getting exhausted frequently while doing simple preprocessing and notebook keeps on restarting. Do we have any options here to tackle such kind of issues? \n\nFor converting csv to parquet info is here.\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/334956#1842426",
    "1846612": "nknarendra7 Here you need two things. \n\nFirst, size reduced version of datasets. Second, use libraries like DASK that can efficiently manage large datasets.\n\nYou can have a look at my notebook \"[Large Dataset - CSV - DASK - Parquet](https://www.kaggle.com/code/mirfanazam/large-dataset-csv-dask-parquet)\"",
    "1846672": "Hi, I saw your code. It was great and thank you for sharing.\nI just wonder if I have converted data and store it on the notebook, do I have to convert and save every time I reconnect to kernel or is there some way that I save them permanently?\nDownload the saved data and add to the notebook whenever kernel start newly is not a considered option. Because if it needs to be uploaded every time then it cannot be used for submit though (unless I submit the file solely)"
  },
  "source": "meta"
}