{
  "id": 188908,
  "title": "[Solution found] Load entire Train data in kernel without error",
  "url": "/competitions/riiid-test-answer-prediction/discussion/188908",
  "author_name": "Sirish Somanchi",
  "post_date": "2020-10-05T19:57:44.780000",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<p>[Edit] <strong>Solution:</strong> (using chunksize in a for loop)</p>\n<pre><code>train = pd.DataFrame()\nfor chunk in pd.read_csv('../input/riiid-test-answer-prediction/train.csv', chunksize=1000000, low_memory=False):\n    train = pd.concat([train, chunk], ignore_index=True)\n</code></pre>\n<p>PS: Updated the same example <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a> below to demonstrate that we can load the <strong>entire original</strong> train dataset without error!</p>\n<hr>\n<p>Train data has  101,230,332 rows (7.5+ GB).<br>\nBut Kaggle kernel with 16GB RAM is unable to load this file using pd.read_csv!</p>\n<p>See my <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a> which crashes with below error message when I try to load the train data.<br>\n<strong>Your notebook tried to allocate more memory than is available. It has restarted.</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F656212%2F27799bd9d61f3bf405ff9662fdb7e0f0%2Fcant_load_train_data.png?generation=1601928130593303&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1038472,
      "postDate": "2020-10-05T19:57:44.780Z",
      "content": "<p>[Edit] <strong>Solution:</strong> (using chunksize in a for loop)</p>\n<pre><code>train = pd.DataFrame()\nfor chunk in pd.read_csv('../input/riiid-test-answer-prediction/train.csv', chunksize=1000000, low_memory=False):\n    train = pd.concat([train, chunk], ignore_index=True)\n</code></pre>\n<p>PS: Updated the same example <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a> below to demonstrate that we can load the <strong>entire original</strong> train dataset without error!</p>\n<hr>\n<p>Train data has  101,230,332 rows (7.5+ GB).<br>\nBut Kaggle kernel with 16GB RAM is unable to load this file using pd.read_csv!</p>\n<p>See my <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a> which crashes with below error message when I try to load the train data.<br>\n<strong>Your notebook tried to allocate more memory than is available. It has restarted.</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F656212%2F27799bd9d61f3bf405ff9662fdb7e0f0%2Fcant_load_train_data.png?generation=1601928130593303&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "[Edit] **Solution:** (using chunksize in a for loop)\n```\ntrain = pd.DataFrame()\nfor chunk in pd.read_csv('../input/riiid-test-answer-prediction/train.csv', chunksize=1000000, low_memory=False):\n    train = pd.concat([train, chunk], ignore_index=True)\n\n```\nPS: Updated the same example [kernel](https://www.kaggle.com/sirishks/cant-load-train-data) below to demonstrate that we can load the **entire original** train dataset without error!\n____________________________________________________________________________\nTrain data has  101,230,332 rows (7.5+ GB).\nBut Kaggle kernel with 16GB RAM is unable to load this file using pd.read_csv!\n\nSee my [kernel](https://www.kaggle.com/sirishks/cant-load-train-data) which crashes with below error message when I try to load the train data.\n**Your notebook tried to allocate more memory than is available. It has restarted.**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F656212%2F27799bd9d61f3bf405ff9662fdb7e0f0%2Fcant_load_train_data.png?generation=1601928130593303&alt=media)",
      "votes": 14
    },
    {
      "id": 1042147,
      "postDate": "2020-10-08T05:05:38.383Z",
      "content": "<p>Good solution</p>",
      "rawMarkdown": "Good solution",
      "votes": 1
    },
    {
      "id": 1038476,
      "postDate": "2020-10-05T20:04:17.860Z",
      "content": "<p>You should use the datatypes shown in the <a href=\"https://www.kaggle.com/sohier/competition-api-detailed-introduction\" target=\"_blank\">starter notebook</a>. Depending on what else you end up doing you may also need to reduce the number of rows you read at any  one time.</p>",
      "rawMarkdown": "You should use the datatypes shown in the [starter notebook](https://www.kaggle.com/sohier/competition-api-detailed-introduction). Depending on what else you end up doing you may also need to reduce the number of rows you read at any  one time.",
      "votes": 1,
      "replies": [
        {
          "id": 1038481,
          "postDate": "2020-10-05T20:08:51.173Z",
          "content": "<p>I see that you have used \"nrows=10**5\" in read_csv call so that you are only reading 100,000 rows (less than 0.1% of the data) at a time.</p>",
          "rawMarkdown": "I see that you have used \"nrows=10**5\" in read_csv call so that you are only reading 100,000 rows (less than 0.1% of the data) at a time.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1039726,
      "postDate": "2020-10-06T18:38:21.317Z",
      "content": "<p>Finally figured out how to use use chunksize in a for loop, and loaded train data successfully!<br>\nSee my <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a></p>",
      "rawMarkdown": "Finally figured out how to use use chunksize in a for loop, and loaded train data successfully!\nSee my [kernel](https://www.kaggle.com/sirishks/cant-load-train-data)"
    },
    {
      "id": 1042658,
      "postDate": "2020-10-08T11:38:07.447Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1042727,
      "postDate": "2020-10-08T12:28:52.693Z",
      "content": "<p>Thanks its working now !</p>",
      "rawMarkdown": "Thanks its working now !"
    }
  ],
  "comments": [
    {
      "id": 1042147,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-08T05:05:38.383000",
      "content": "<p>Good solution</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1038476,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2020-10-05T20:04:17.860000",
      "content": "<p>You should use the datatypes shown in the <a href=\"https://www.kaggle.com/sohier/competition-api-detailed-introduction\" target=\"_blank\">starter notebook</a>. Depending on what else you end up doing you may also need to reduce the number of rows you read at any  one time.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1038481,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-10-05T20:08:51.173000",
          "content": "<p>I see that you have used \"nrows=10**5\" in read_csv call so that you are only reading 100,000 rows (less than 0.1% of the data) at a time.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1039726,
      "author_name": "Sirish Somanchi",
      "author_url": "",
      "post_date": "2020-10-06T18:38:21.317000",
      "content": "<p>Finally figured out how to use use chunksize in a for loop, and loaded train data successfully!<br>\nSee my <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1042658,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-08T11:38:07.447000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1042727,
      "author_name": "Mayank_bhandari0385",
      "author_url": "",
      "post_date": "2020-10-08T12:28:52.693000",
      "content": "<p>Thanks its working now !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1038472": "[Edit] **Solution:** (using chunksize in a for loop)\n```\ntrain = pd.DataFrame()\nfor chunk in pd.read_csv('../input/riiid-test-answer-prediction/train.csv', chunksize=1000000, low_memory=False):\n    train = pd.concat([train, chunk], ignore_index=True)\n\n```\nPS: Updated the same example [kernel](https://www.kaggle.com/sirishks/cant-load-train-data) below to demonstrate that we can load the **entire original** train dataset without error!\n____________________________________________________________________________\nTrain data has  101,230,332 rows (7.5+ GB).\nBut Kaggle kernel with 16GB RAM is unable to load this file using pd.read_csv!\n\nSee my [kernel](https://www.kaggle.com/sirishks/cant-load-train-data) which crashes with below error message when I try to load the train data.\n**Your notebook tried to allocate more memory than is available. It has restarted.**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F656212%2F27799bd9d61f3bf405ff9662fdb7e0f0%2Fcant_load_train_data.png?generation=1601928130593303&alt=media)",
    "1042147": "Good solution",
    "1038476": "You should use the datatypes shown in the [starter notebook](https://www.kaggle.com/sohier/competition-api-detailed-introduction). Depending on what else you end up doing you may also need to reduce the number of rows you read at any  one time.",
    "1039726": "Finally figured out how to use use chunksize in a for loop, and loaded train data successfully!\nSee my [kernel](https://www.kaggle.com/sirishks/cant-load-train-data)",
    "1042658": "",
    "1042727": "Thanks its working now !"
  }
}