{
  "id": 438373,
  "title": "Beginner issue: Running out of memory on kaggle notebook",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/438373",
  "author_name": "",
  "post_date": "2023-09-10T20:19:31.339590600Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As the title suggest I run out of memory when trying to load in the train data:</p>\n<p>`import pandas as pd</p>\n<p>train_data = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv')</p>\n<p>print(train_data.head())<br>\nprint(train_data['sequence'].dtype)`</p>\n<p>--&gt; results in \"Your notebook tried to allocate more memory than is available. It has restarted.\"</p>\n<p>Should the dataset be split up and be loaded in piece by piece? </p>\n<p>Thanks for any suggestions.</p>",
  "messages": [
    {
      "id": "2432411",
      "postDate": "09/10/2023 20:19:31",
      "content": "<p>As the title suggest I run out of memory when trying to load in the train data:</p>\n<p>`import pandas as pd</p>\n<p>train_data = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv')</p>\n<p>print(train_data.head())<br>\nprint(train_data['sequence'].dtype)`</p>\n<p>--&gt; results in \"Your notebook tried to allocate more memory than is available. It has restarted.\"</p>\n<p>Should the dataset be split up and be loaded in piece by piece? </p>\n<p>Thanks for any suggestions.</p>",
      "rawMarkdown": "As the title suggest I run out of memory when trying to load in the train data:\n\n`import pandas as pd\n\ntrain_data = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv')\n\nprint(train_data.head())\nprint(train_data['sequence'].dtype)`\n\n--> results in \"Your notebook tried to allocate more memory than is available. It has restarted.\"\n\nShould the dataset be split up and be loaded in piece by piece? \n\nThanks for any suggestions.",
      "votes": null
    },
    {
      "id": "2432493",
      "postDate": "09/10/2023 22:18:45",
      "content": "<p>Hi. If you're doing exploratory data analysis and you only need a couple of rows just add <code>nrows =1000</code>in the <code>pd.read_csv</code> function.</p>\n<p>Otherwise use chunking (<code>chunksize=100000</code> for example)</p>",
      "rawMarkdown": "Hi. If you're doing exploratory data analysis and you only need a couple of rows just add `nrows =1000`in the `pd.read_csv` function.\n\nOtherwise use chunking (`chunksize=100000` for example)",
      "votes": null
    },
    {
      "id": "2432499",
      "postDate": "09/10/2023 22:34:41",
      "content": "<p>Hi,</p>\n<p>You can download them in chunks using the following code:<br>\nchunksize = 1000 #read 1000 rows at a time<br>\nwith pd.read_csv(data, chunksize=chunksize) as file:<br>\n        for c in file:<br>\n              process(c)</p>\n<p>you can try using dask which is optimized for large data<br>\nimport dask.dataframe as ddf<br>\ndf = ddf.read_csv(data)</p>",
      "rawMarkdown": "Hi,\n \nYou can download them in chunks using the following code:\nchunksize = 1000 #read 1000 rows at a time\nwith pd.read_csv(data, chunksize=chunksize) as file:\n        for c in file:\n              process(c)\n\nyou can try using dask which is optimized for large data\nimport dask.dataframe as ddf\ndf = ddf.read_csv(data)",
      "votes": null
    },
    {
      "id": "2432569",
      "postDate": "09/11/2023 01:55:06",
      "content": "<p><a href=\"https://www.kaggle.com/datasets/horikitasaku/rna-low-m-data\" target=\"_blank\">https://www.kaggle.com/datasets/horikitasaku/rna-low-m-data</a><br>\n<a href=\"https://www.kaggle.com/code/horikitasaku/rna-data-memory-reducing\" target=\"_blank\">https://www.kaggle.com/code/horikitasaku/rna-data-memory-reducing</a><br>\nI made a dataset which may use less memory, maybe you can have a try?</p>\n<p>However the submission file is still huge</p>",
      "rawMarkdown": "https://www.kaggle.com/datasets/horikitasaku/rna-low-m-data\nhttps://www.kaggle.com/code/horikitasaku/rna-data-memory-reducing\nI made a dataset which may use less memory, maybe you can have a try?\n\nHowever the submission file is still huge",
      "votes": null
    },
    {
      "id": "2432596",
      "postDate": "09/11/2023 02:59:13",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/anthonyshes0teacup\" target=\"_blank\">@anthonyshes0teacup</a> I having too much problem  with allocate memory and now I can solve it easily</p>",
      "rawMarkdown": "Thanks @anthonyshes0teacup I having too much problem  with allocate memory and now I can solve it easily",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2432493,
      "author_name": "attilaimre",
      "author_url": "",
      "post_date": "09/10/2023 22:18:45",
      "content": "<p>Hi. If you're doing exploratory data analysis and you only need a couple of rows just add <code>nrows =1000</code>in the <code>pd.read_csv</code> function.</p>\n<p>Otherwise use chunking (<code>chunksize=100000</code> for example)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2432499,
      "author_name": "stitch",
      "author_url": "",
      "post_date": "09/10/2023 22:34:41",
      "content": "<p>Hi,</p>\n<p>You can download them in chunks using the following code:<br>\nchunksize = 1000 #read 1000 rows at a time<br>\nwith pd.read_csv(data, chunksize=chunksize) as file:<br>\n        for c in file:<br>\n              process(c)</p>\n<p>you can try using dask which is optimized for large data<br>\nimport dask.dataframe as ddf<br>\ndf = ddf.read_csv(data)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2432569,
      "author_name": "horikitasaku",
      "author_url": "",
      "post_date": "09/11/2023 01:55:06",
      "content": "<p><a href=\"https://www.kaggle.com/datasets/horikitasaku/rna-low-m-data\" target=\"_blank\">https://www.kaggle.com/datasets/horikitasaku/rna-low-m-data</a><br>\n<a href=\"https://www.kaggle.com/code/horikitasaku/rna-data-memory-reducing\" target=\"_blank\">https://www.kaggle.com/code/horikitasaku/rna-data-memory-reducing</a><br>\nI made a dataset which may use less memory, maybe you can have a try?</p>\n<p>However the submission file is still huge</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2432596,
      "author_name": "sourabhsingh03993493",
      "author_url": "",
      "post_date": "09/11/2023 02:59:13",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/anthonyshes0teacup\" target=\"_blank\">@anthonyshes0teacup</a> I having too much problem  with allocate memory and now I can solve it easily</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2432411": "As the title suggest I run out of memory when trying to load in the train data:\n\n`import pandas as pd\n\ntrain_data = pd.read_csv('/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv')\n\nprint(train_data.head())\nprint(train_data['sequence'].dtype)`\n\n--> results in \"Your notebook tried to allocate more memory than is available. It has restarted.\"\n\nShould the dataset be split up and be loaded in piece by piece? \n\nThanks for any suggestions.",
    "2432493": "Hi. If you're doing exploratory data analysis and you only need a couple of rows just add `nrows =1000`in the `pd.read_csv` function.\n\nOtherwise use chunking (`chunksize=100000` for example)",
    "2432499": "Hi,\n \nYou can download them in chunks using the following code:\nchunksize = 1000 #read 1000 rows at a time\nwith pd.read_csv(data, chunksize=chunksize) as file:\n        for c in file:\n              process(c)\n\nyou can try using dask which is optimized for large data\nimport dask.dataframe as ddf\ndf = ddf.read_csv(data)",
    "2432569": "https://www.kaggle.com/datasets/horikitasaku/rna-low-m-data\nhttps://www.kaggle.com/code/horikitasaku/rna-data-memory-reducing\nI made a dataset which may use less memory, maybe you can have a try?\n\nHowever the submission file is still huge",
    "2432596": "Thanks @anthonyshes0teacup I having too much problem  with allocate memory and now I can solve it easily"
  },
  "source": "meta"
}