{
  "id": 370091,
  "title": "problem, how to solve it?   the data is too big",
  "url": "/competitions/otto-recommender-system/discussion/370091",
  "author_name": "",
  "post_date": "2022-12-03T01:17:01.714886500Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Your notebook tried to allocate more memory than is available. It has restarted.</p>",
  "messages": [
    {
      "id": "2053203",
      "postDate": "12/03/2022 01:17:01",
      "content": "<p>Your notebook tried to allocate more memory than is available. It has restarted.</p>",
      "rawMarkdown": "Your notebook tried to allocate more memory than is available. It has restarted.",
      "votes": null
    },
    {
      "id": "2053568",
      "postDate": "12/03/2022 12:08:22",
      "content": "<p>Yep, it's a lot of GBs of data here, so you can't just load it all with pandas like you normally would. You can use the <code>chunksize</code> argument in <code>pandas.read_json</code> to load it in chunks (fragments of data), using much less memory at a time: <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.read_json.html\" target=\"_blank\">documentation</a>, <a href=\"https://stackoverflow.com/questions/10238340/whats-the-best-way-to-load-large-json-lists-in-python\" target=\"_blank\">example</a>.</p>\n<p>Keep in mind that there really is A LOT of data, so reading it all raw would take hours and, likely, the full loaded dataset will still take up more memory than you have. For exploration and some light experimentation, you could limit the number of chunks you're loading.</p>",
      "rawMarkdown": "Yep, it's a lot of GBs of data here, so you can't just load it all with pandas like you normally would. You can use the `chunksize` argument in `pandas.read_json` to load it in chunks (fragments of data), using much less memory at a time: [documentation] (https://pandas.pydata.org/docs/reference/api/pandas.read_json.html), [example](https://stackoverflow.com/questions/10238340/whats-the-best-way-to-load-large-json-lists-in-python).\n\nKeep in mind that there really is A LOT of data, so reading it all raw would take hours and, likely, the full loaded dataset will still take up more memory than you have. For exploration and some light experimentation, you could limit the number of chunks you're loading.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2053568,
      "author_name": "andriikliachkin",
      "author_url": "",
      "post_date": "12/03/2022 12:08:22",
      "content": "<p>Yep, it's a lot of GBs of data here, so you can't just load it all with pandas like you normally would. You can use the <code>chunksize</code> argument in <code>pandas.read_json</code> to load it in chunks (fragments of data), using much less memory at a time: <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.read_json.html\" target=\"_blank\">documentation</a>, <a href=\"https://stackoverflow.com/questions/10238340/whats-the-best-way-to-load-large-json-lists-in-python\" target=\"_blank\">example</a>.</p>\n<p>Keep in mind that there really is A LOT of data, so reading it all raw would take hours and, likely, the full loaded dataset will still take up more memory than you have. For exploration and some light experimentation, you could limit the number of chunks you're loading.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2053203": "Your notebook tried to allocate more memory than is available. It has restarted.",
    "2053568": "Yep, it's a lot of GBs of data here, so you can't just load it all with pandas like you normally would. You can use the `chunksize` argument in `pandas.read_json` to load it in chunks (fragments of data), using much less memory at a time: [documentation] (https://pandas.pydata.org/docs/reference/api/pandas.read_json.html), [example](https://stackoverflow.com/questions/10238340/whats-the-best-way-to-load-large-json-lists-in-python).\n\nKeep in mind that there really is A LOT of data, so reading it all raw would take hours and, likely, the full loaded dataset will still take up more memory than you have. For exploration and some light experimentation, you could limit the number of chunks you're loading."
  },
  "source": "meta"
}