{
  "id": 337846,
  "title": "Out of memory reading test file (parquet, 3.3GiB). Any suggestions?",
  "url": "/competitions/amex-default-prediction/discussion/337846",
  "author_name": "",
  "post_date": "2022-07-18T02:23:36.428958900Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I am trying to run with edits this wonderful notebook from <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart</a>), but everytime I tried to read the test file (3.3 GiB, parquet) I run out of memory.</p>\n<p>I have delete everything possible related to train (except for the model) before loading up test for preprocessing, but to no avail.</p>\n<p>Here is my NB: <a href=\"https://www.kaggle.com/code/jsmithperera/amex-lightgbm-f-gen\" target=\"_blank\">https://www.kaggle.com/code/jsmithperera/amex-lightgbm-f-gen</a></p>\n<p>Any suggestions will be appreciated. </p>\n<p>Gracias…</p>",
  "messages": [
    {
      "id": "1859847",
      "postDate": "07/18/2022 02:23:36",
      "content": "<p>I am trying to run with edits this wonderful notebook from <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart</a>), but everytime I tried to read the test file (3.3 GiB, parquet) I run out of memory.</p>\n<p>I have delete everything possible related to train (except for the model) before loading up test for preprocessing, but to no avail.</p>\n<p>Here is my NB: <a href=\"https://www.kaggle.com/code/jsmithperera/amex-lightgbm-f-gen\" target=\"_blank\">https://www.kaggle.com/code/jsmithperera/amex-lightgbm-f-gen</a></p>\n<p>Any suggestions will be appreciated. </p>\n<p>Gracias…</p>",
      "rawMarkdown": "I am trying to run with edits this wonderful notebook from @ambrosm ([https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart](url)), but everytime I tried to read the test file (3.3 GiB, parquet) I run out of memory.\n\nI have delete everything possible related to train (except for the model) before loading up test for preprocessing, but to no avail.\n\nHere is my NB: https://www.kaggle.com/code/jsmithperera/amex-lightgbm-f-gen\n\nAny suggestions will be appreciated. \n\nGracias...",
      "votes": null
    },
    {
      "id": "1859922",
      "postDate": "07/18/2022 04:58:15",
      "content": "<p>Converting the data type to int8 and recording index is a good choice, but there is a loss of precision.<br>\nOf course there are other ways to do this, but in practice, not practical enough.</p>",
      "rawMarkdown": "Converting the data type to int8 and recording index is a good choice, but there is a loss of precision.\nOf course there are other ways to do this, but in practice, not practical enough.",
      "votes": null
    },
    {
      "id": "1859930",
      "postDate": "07/18/2022 05:03:11",
      "content": "<p><a href=\"https://www.kaggle.com/jsmithperera\" target=\"_blank\">@jsmithperera</a> The notebook needs 16 GB RAM. GPU notebooks have only 13 GB. Run it without GPU.</p>",
      "rawMarkdown": "jsmithperera The notebook needs 16 GB RAM. GPU notebooks have only 13 GB. Run it without GPU.",
      "votes": null
    },
    {
      "id": "1860033",
      "postDate": "07/18/2022 06:10:54",
      "content": "<p><a href=\"https://www.kaggle.com/jsmithperera\" target=\"_blank\">@jsmithperera</a> explore the parquet format and try to perform the task sequentially and remove the unnecessary Data frame and variable after training the model; Gpu can boost the performance by manifolds</p>",
      "rawMarkdown": "jsmithperera explore the parquet format and try to perform the task sequentially and remove the unnecessary Data frame and variable after training the model; Gpu can boost the performance by manifolds",
      "votes": null
    },
    {
      "id": "1860157",
      "postDate": "07/18/2022 06:58:13",
      "content": "<p>try the generator </p>",
      "rawMarkdown": "try the generator",
      "votes": null
    },
    {
      "id": "1860646",
      "postDate": "07/18/2022 12:49:44",
      "content": "<p>I did not know about generator functions. I will look at it. </p>\n<p>I am currently trying to read the test file in chunks.</p>\n<p>Thanks for your comment!</p>",
      "rawMarkdown": "I did not know about generator functions. I will look at it. \n\nI am currently trying to read the test file in chunks.\n\nThanks for your comment!",
      "votes": null
    },
    {
      "id": "1860652",
      "postDate": "07/18/2022 12:53:17",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dhruv8\" target=\"_blank\">@dhruv8</a>. </p>\n<p>I delete everything after training and I am currently trying to read the test file in chunks.</p>\n<p>Thanks again!</p>",
      "rawMarkdown": "Thanks @dhruv8. \n\nI delete everything after training and I am currently trying to read the test file in chunks.\n\nThanks again!",
      "votes": null
    },
    {
      "id": "1860726",
      "postDate": "07/18/2022 13:53:10",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>. Thanks for responding. Your notebook is great.</p>\n<p>I ran the notebook without GPU but got the same message: out of memory.  I had deleted everything about training except for the model. I will try it again. </p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @ambrosm. Thanks for responding. Your notebook is great.\n\nI ran the notebook without GPU but got the same message: out of memory.  I had deleted everything about training except for the model. I will try it again. \n\nThanks!",
      "votes": null
    },
    {
      "id": "1860779",
      "postDate": "07/18/2022 14:58:28",
      "content": "<p>You can try virtual memory</p>",
      "rawMarkdown": "You can try virtual memory",
      "votes": null
    },
    {
      "id": "1863190",
      "postDate": "07/20/2022 07:54:00",
      "content": "<p>I'm also facing the same issue. So I use two notebooks - one for train and the other for prediction. I add the output of the train (models and any additional file you use) to the prediction notebook. You could try the same.</p>",
      "rawMarkdown": "I'm also facing the same issue. So I use two notebooks - one for train and the other for prediction. I add the output of the train (models and any additional file you use) to the prediction notebook. You could try the same.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1859922,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "07/18/2022 04:58:15",
      "content": "<p>Converting the data type to int8 and recording index is a good choice, but there is a loss of precision.<br>\nOf course there are other ways to do this, but in practice, not practical enough.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1859930,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "07/18/2022 05:03:11",
      "content": "<p><a href=\"https://www.kaggle.com/jsmithperera\" target=\"_blank\">@jsmithperera</a> The notebook needs 16 GB RAM. GPU notebooks have only 13 GB. Run it without GPU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1860726,
          "author_name": "jsmithperera",
          "author_url": "",
          "post_date": "07/18/2022 13:53:10",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>. Thanks for responding. Your notebook is great.</p>\n<p>I ran the notebook without GPU but got the same message: out of memory.  I had deleted everything about training except for the model. I will try it again. </p>\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1860033,
      "author_name": "dhruv8",
      "author_url": "",
      "post_date": "07/18/2022 06:10:54",
      "content": "<p><a href=\"https://www.kaggle.com/jsmithperera\" target=\"_blank\">@jsmithperera</a> explore the parquet format and try to perform the task sequentially and remove the unnecessary Data frame and variable after training the model; Gpu can boost the performance by manifolds</p>",
      "votes": null,
      "replies": [
        {
          "id": 1860652,
          "author_name": "jsmithperera",
          "author_url": "",
          "post_date": "07/18/2022 12:53:17",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/dhruv8\" target=\"_blank\">@dhruv8</a>. </p>\n<p>I delete everything after training and I am currently trying to read the test file in chunks.</p>\n<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1860157,
      "author_name": "bingdaoqwq",
      "author_url": "",
      "post_date": "07/18/2022 06:58:13",
      "content": "<p>try the generator </p>",
      "votes": null,
      "replies": [
        {
          "id": 1860646,
          "author_name": "jsmithperera",
          "author_url": "",
          "post_date": "07/18/2022 12:49:44",
          "content": "<p>I did not know about generator functions. I will look at it. </p>\n<p>I am currently trying to read the test file in chunks.</p>\n<p>Thanks for your comment!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1860779,
      "author_name": "liuzengyu",
      "author_url": "",
      "post_date": "07/18/2022 14:58:28",
      "content": "<p>You can try virtual memory</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1863190,
      "author_name": "bhawinbengani",
      "author_url": "",
      "post_date": "07/20/2022 07:54:00",
      "content": "<p>I'm also facing the same issue. So I use two notebooks - one for train and the other for prediction. I add the output of the train (models and any additional file you use) to the prediction notebook. You could try the same.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1859847": "I am trying to run with edits this wonderful notebook from @ambrosm ([https://www.kaggle.com/code/ambrosm/amex-lightgbm-quickstart](url)), but everytime I tried to read the test file (3.3 GiB, parquet) I run out of memory.\n\nI have delete everything possible related to train (except for the model) before loading up test for preprocessing, but to no avail.\n\nHere is my NB: https://www.kaggle.com/code/jsmithperera/amex-lightgbm-f-gen\n\nAny suggestions will be appreciated. \n\nGracias...",
    "1859922": "Converting the data type to int8 and recording index is a good choice, but there is a loss of precision.\nOf course there are other ways to do this, but in practice, not practical enough.",
    "1859930": "jsmithperera The notebook needs 16 GB RAM. GPU notebooks have only 13 GB. Run it without GPU.",
    "1860033": "jsmithperera explore the parquet format and try to perform the task sequentially and remove the unnecessary Data frame and variable after training the model; Gpu can boost the performance by manifolds",
    "1860157": "try the generator",
    "1860646": "I did not know about generator functions. I will look at it. \n\nI am currently trying to read the test file in chunks.\n\nThanks for your comment!",
    "1860652": "Thanks @dhruv8. \n\nI delete everything after training and I am currently trying to read the test file in chunks.\n\nThanks again!",
    "1860726": "Hi @ambrosm. Thanks for responding. Your notebook is great.\n\nI ran the notebook without GPU but got the same message: out of memory.  I had deleted everything about training except for the model. I will try it again. \n\nThanks!",
    "1860779": "You can try virtual memory",
    "1863190": "I'm also facing the same issue. So I use two notebooks - one for train and the other for prediction. I add the output of the train (models and any additional file you use) to the prediction notebook. You could try the same."
  },
  "source": "meta"
}