{
  "id": 73434,
  "title": "Too Small RAM. Cant't Read All Test Data",
  "url": "/competitions/PLAsTiCC-2018/discussion/73434",
  "author_name": "",
  "post_date": "2018-12-03T06:00:44.234596200Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi All,</p>\n\n<p>I am using Kaggle Kernel to read training data, which RAM is 17.2G. However, every time I use pd.read_csv('..//input//test_set.csv') to read the data, the Kernel will crush. Is there any way to address this issue? </p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "431939",
      "postDate": "12/03/2018 06:00:44",
      "content": "<p>Hi All,</p>\n\n<p>I am using Kaggle Kernel to read training data, which RAM is 17.2G. However, every time I use pd.read_csv('..//input//test_set.csv') to read the data, the Kernel will crush. Is there any way to address this issue? </p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi All,\n\nI am using Kaggle Kernel to read training data, which RAM is 17.2G. However, every time I use pd.read_csv('..//input//test_set.csv') to read the data, the Kernel will crush. Is there any way to address this issue? \n\nThanks!",
      "votes": null
    },
    {
      "id": "431985",
      "postDate": "12/03/2018 08:12:12",
      "content": "<p>you could try using python generators so memory size is reduced\n-<a href=\"https://realpython.com/introduction-to-python-generators/\">https://realpython.com/introduction-to-python-generators/</a></p>",
      "rawMarkdown": "you could try using python generators so memory size is reduced\n-https://realpython.com/introduction-to-python-generators/",
      "votes": null
    },
    {
      "id": "431992",
      "postDate": "12/03/2018 08:22:29",
      "content": "<p>See how to process test data by chunks in public kernels.  There is no need to load all test data at once.</p>",
      "rawMarkdown": "See how to process test data by chunks in public kernels.  There is no need to load all test data at once.",
      "votes": null
    },
    {
      "id": "432352",
      "postDate": "12/03/2018 18:49:44",
      "content": "<p>There is a public kernel called 'fast-test-set-reading' that is really good.  I integrated it with my custom features (which aren't very good) in 'something different - test set edition.'  You could probably fork that and then replace my custom features with those of your choosing.</p>",
      "rawMarkdown": "There is a public kernel called 'fast-test-set-reading' that is really good.  I integrated it with my custom features (which aren't very good) in 'something different - test set edition.'  You could probably fork that and then replace my custom features with those of your choosing.",
      "votes": null
    },
    {
      "id": "432925",
      "postDate": "12/04/2018 13:31:12",
      "content": "<p>you could use a dask dataframe composed of many smaller pandas dataframes.\n- <a href=\"https://dask.org\">https://dask.org</a></p>",
      "rawMarkdown": "you could use a dask dataframe composed of many smaller pandas dataframes.\n- https://dask.org",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 431985,
      "author_name": "richarde",
      "author_url": "",
      "post_date": "12/03/2018 08:12:12",
      "content": "<p>you could try using python generators so memory size is reduced\n-<a href=\"https://realpython.com/introduction-to-python-generators/\">https://realpython.com/introduction-to-python-generators/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 431992,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/03/2018 08:22:29",
      "content": "<p>See how to process test data by chunks in public kernels.  There is no need to load all test data at once.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 432352,
      "author_name": "jimpsull",
      "author_url": "",
      "post_date": "12/03/2018 18:49:44",
      "content": "<p>There is a public kernel called 'fast-test-set-reading' that is really good.  I integrated it with my custom features (which aren't very good) in 'something different - test set edition.'  You could probably fork that and then replace my custom features with those of your choosing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 432925,
      "author_name": "hattan0523",
      "author_url": "",
      "post_date": "12/04/2018 13:31:12",
      "content": "<p>you could use a dask dataframe composed of many smaller pandas dataframes.\n- <a href=\"https://dask.org\">https://dask.org</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "431939": "Hi All,\n\nI am using Kaggle Kernel to read training data, which RAM is 17.2G. However, every time I use pd.read_csv('..//input//test_set.csv') to read the data, the Kernel will crush. Is there any way to address this issue? \n\nThanks!",
    "431985": "you could try using python generators so memory size is reduced\n-https://realpython.com/introduction-to-python-generators/",
    "431992": "See how to process test data by chunks in public kernels.  There is no need to load all test data at once.",
    "432352": "There is a public kernel called 'fast-test-set-reading' that is really good.  I integrated it with my custom features (which aren't very good) in 'something different - test set edition.'  You could probably fork that and then replace my custom features with those of your choosing.",
    "432925": "you could use a dask dataframe composed of many smaller pandas dataframes.\n- https://dask.org"
  },
  "source": "meta"
}