{
  "id": 190595,
  "title": "MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332) and data type int64",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190595",
  "author_name": "",
  "post_date": "2020-10-12T13:52:35.423511900Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello Kagglers,</p>\n<p>This is my first time competing and this may be a silly question. I was trying to run stratified k-fold cross-validation on train.csv and it gives me a MemoryError. Specifically, MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332) and data type int64. </p>\n<p>My specs:<br>\n16 GB RAM<br>\nGTX 1060<br>\nCore i7 </p>\n<p>I do not think it is much of a hardware issue. I ran the fold code on Spyder. Any help is appreciated. Thank you very much.</p>",
  "messages": [
    {
      "id": "1047346",
      "postDate": "10/12/2020 13:52:35",
      "content": "<p>Hello Kagglers,</p>\n<p>This is my first time competing and this may be a silly question. I was trying to run stratified k-fold cross-validation on train.csv and it gives me a MemoryError. Specifically, MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332) and data type int64. </p>\n<p>My specs:<br>\n16 GB RAM<br>\nGTX 1060<br>\nCore i7 </p>\n<p>I do not think it is much of a hardware issue. I ran the fold code on Spyder. Any help is appreciated. Thank you very much.</p>",
      "rawMarkdown": "Hello Kagglers,\n\nThis is my first time competing and this may be a silly question. I was trying to run stratified k-fold cross-validation on train.csv and it gives me a MemoryError. Specifically, MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332) and data type int64. \n\nMy specs:\n16 GB RAM\nGTX 1060\nCore i7 \n\nI do not think it is much of a hardware issue. I ran the fold code on Spyder. Any help is appreciated. Thank you very much.",
      "votes": null
    },
    {
      "id": "1047379",
      "postDate": "10/12/2020 14:34:58",
      "content": "<p>It is a hardware issue.<br>\nEven kaggle kernels also have 16GB of RAM and by loading this much size of any arrays will have memory errors</p>",
      "rawMarkdown": "It is a hardware issue.\nEven kaggle kernels also have 16GB of RAM and by loading this much size of any arrays will have memory errors",
      "votes": null
    },
    {
      "id": "1047924",
      "postDate": "10/13/2020 04:12:03",
      "content": "<p>\"MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332)\" - This error usually arises when you initialize a numpy array (by doing something like np.zeros((8, 101230332)) ), however such an array cant be fit on the memory, so this step fails. More details - <a href=\"https://stackoverflow.com/questions/57507832/unable-to-allocate-array-with-shape-and-data-type\" target=\"_blank\">https://stackoverflow.com/questions/57507832/unable-to-allocate-array-with-shape-and-data-type</a></p>",
      "rawMarkdown": "\"MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332)\" - This error usually arises when you initialize a numpy array (by doing something like np.zeros((8, 101230332)) ), however such an array cant be fit on the memory, so this step fails. More details - https://stackoverflow.com/questions/57507832/unable-to-allocate-array-with-shape-and-data-type",
      "votes": null
    },
    {
      "id": "1047943",
      "postDate": "10/13/2020 04:36:23",
      "content": "<p>I have also noticed that the program first tries to allocate a much larger memory, may be while making copies of the original dataset i.e. memory goes from 6GB to 20GB and then back down to 14GB. So, even though actual requirement is less than 16GB, training code throws an error at the intermediate stage where a larger memory is allocated.</p>\n<p>Even though train data size is about 7.5GB, we were not able to load it in kaggle kernel, which is the primary reason for several previous discussions on using Dask, cuDF, etc. Also see <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188908\" target=\"_blank\">pandas solution</a> and <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a> using chunksize.</p>\n<ol>\n<li>Change data type from int64 to smaller sizes:<br>\ndtype={'timestamp': 'int64',<br>\n      'user_id': 'int32',<br>\n      'content_id': 'int16',<br>\n      'content_type_id': 'int8',<br>\n      'answered_correctly':'uint8',<br>\n      'prior_question_elapsed_time': 'float32',<br>\n      'prior_question_had_explanation': 'boolean'}<br>\n)</li>\n<li>Run training on a personal desktop / cloud VM with 32GB RAM</li>\n<li>Try a subset of the data for training</li>\n<li>Another option is to try OS settings:<br>\necho 1 &gt; /proc/sys/vm/overcommit_memory</li>\n</ol>",
      "rawMarkdown": "I have also noticed that the program first tries to allocate a much larger memory, may be while making copies of the original dataset i.e. memory goes from 6GB to 20GB and then back down to 14GB. So, even though actual requirement is less than 16GB, training code throws an error at the intermediate stage where a larger memory is allocated.\n\nEven though train data size is about 7.5GB, we were not able to load it in kaggle kernel, which is the primary reason for several previous discussions on using Dask, cuDF, etc. Also see [pandas solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188908) and [kernel](https://www.kaggle.com/sirishks/cant-load-train-data) using chunksize.\n\n1. Change data type from int64 to smaller sizes:\n   dtype={'timestamp': 'int64',\n          'user_id': 'int32',\n          'content_id': 'int16',\n          'content_type_id': 'int8',\n          'answered_correctly':'uint8',\n          'prior_question_elapsed_time': 'float32',\n          'prior_question_had_explanation': 'boolean'}\n   )\n2. Run training on a personal desktop / cloud VM with 32GB RAM\n3. Try a subset of the data for training\n4. Another option is to try OS settings:\n   echo 1 > /proc/sys/vm/overcommit_memory",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1047379,
      "author_name": "msharuk589",
      "author_url": "",
      "post_date": "10/12/2020 14:34:58",
      "content": "<p>It is a hardware issue.<br>\nEven kaggle kernels also have 16GB of RAM and by loading this much size of any arrays will have memory errors</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1047924,
      "author_name": "abhimanyud",
      "author_url": "",
      "post_date": "10/13/2020 04:12:03",
      "content": "<p>\"MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332)\" - This error usually arises when you initialize a numpy array (by doing something like np.zeros((8, 101230332)) ), however such an array cant be fit on the memory, so this step fails. More details - <a href=\"https://stackoverflow.com/questions/57507832/unable-to-allocate-array-with-shape-and-data-type\" target=\"_blank\">https://stackoverflow.com/questions/57507832/unable-to-allocate-array-with-shape-and-data-type</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1047943,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "10/13/2020 04:36:23",
      "content": "<p>I have also noticed that the program first tries to allocate a much larger memory, may be while making copies of the original dataset i.e. memory goes from 6GB to 20GB and then back down to 14GB. So, even though actual requirement is less than 16GB, training code throws an error at the intermediate stage where a larger memory is allocated.</p>\n<p>Even though train data size is about 7.5GB, we were not able to load it in kaggle kernel, which is the primary reason for several previous discussions on using Dask, cuDF, etc. Also see <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188908\" target=\"_blank\">pandas solution</a> and <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a> using chunksize.</p>\n<ol>\n<li>Change data type from int64 to smaller sizes:<br>\ndtype={'timestamp': 'int64',<br>\n      'user_id': 'int32',<br>\n      'content_id': 'int16',<br>\n      'content_type_id': 'int8',<br>\n      'answered_correctly':'uint8',<br>\n      'prior_question_elapsed_time': 'float32',<br>\n      'prior_question_had_explanation': 'boolean'}<br>\n)</li>\n<li>Run training on a personal desktop / cloud VM with 32GB RAM</li>\n<li>Try a subset of the data for training</li>\n<li>Another option is to try OS settings:<br>\necho 1 &gt; /proc/sys/vm/overcommit_memory</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1047346": "Hello Kagglers,\n\nThis is my first time competing and this may be a silly question. I was trying to run stratified k-fold cross-validation on train.csv and it gives me a MemoryError. Specifically, MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332) and data type int64. \n\nMy specs:\n16 GB RAM\nGTX 1060\nCore i7 \n\nI do not think it is much of a hardware issue. I ran the fold code on Spyder. Any help is appreciated. Thank you very much.",
    "1047379": "It is a hardware issue.\nEven kaggle kernels also have 16GB of RAM and by loading this much size of any arrays will have memory errors",
    "1047924": "\"MemoryError: Unable to allocate 6.03 GiB for an array with shape (8, 101230332)\" - This error usually arises when you initialize a numpy array (by doing something like np.zeros((8, 101230332)) ), however such an array cant be fit on the memory, so this step fails. More details - https://stackoverflow.com/questions/57507832/unable-to-allocate-array-with-shape-and-data-type",
    "1047943": "I have also noticed that the program first tries to allocate a much larger memory, may be while making copies of the original dataset i.e. memory goes from 6GB to 20GB and then back down to 14GB. So, even though actual requirement is less than 16GB, training code throws an error at the intermediate stage where a larger memory is allocated.\n\nEven though train data size is about 7.5GB, we were not able to load it in kaggle kernel, which is the primary reason for several previous discussions on using Dask, cuDF, etc. Also see [pandas solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188908) and [kernel](https://www.kaggle.com/sirishks/cant-load-train-data) using chunksize.\n\n1. Change data type from int64 to smaller sizes:\n   dtype={'timestamp': 'int64',\n          'user_id': 'int32',\n          'content_id': 'int16',\n          'content_type_id': 'int8',\n          'answered_correctly':'uint8',\n          'prior_question_elapsed_time': 'float32',\n          'prior_question_had_explanation': 'boolean'}\n   )\n2. Run training on a personal desktop / cloud VM with 32GB RAM\n3. Try a subset of the data for training\n4. Another option is to try OS settings:\n   echo 1 > /proc/sys/vm/overcommit_memory"
  },
  "source": "meta"
}