{
  "id": 341248,
  "title": "Strategy for dealing with prediction files that are too big to hold in computer RAM",
  "url": "/competitions/amex-default-prediction/discussion/341248",
  "author_name": "",
  "post_date": "2022-08-02T03:12:37.560384500Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The test data is too big to read in and hold inside a dataframe all at once. I'd like to ask how what the solution to this problem is. Is it best to read the csv in as a batch, make predictions on those, and then write to a submission file?</p>",
  "messages": [
    {
      "id": "1880750",
      "postDate": "08/02/2022 03:12:37",
      "content": "<p>The test data is too big to read in and hold inside a dataframe all at once. I'd like to ask how what the solution to this problem is. Is it best to read the csv in as a batch, make predictions on those, and then write to a submission file?</p>",
      "rawMarkdown": "The test data is too big to read in and hold inside a dataframe all at once. I'd like to ask how what the solution to this problem is. Is it best to read the csv in as a batch, make predictions on those, and then write to a submission file?",
      "votes": null
    },
    {
      "id": "1880772",
      "postDate": "08/02/2022 03:42:38",
      "content": "<p>Yes doing the inference in batches is probably your best best</p>\n<p>This notebook has a great example of how to do this<br>\n<a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></p>",
      "rawMarkdown": "Yes doing the inference in batches is probably your best best\n\nThis notebook has a great example of how to do this\nhttps://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1880772,
      "author_name": "illidan7",
      "author_url": "",
      "post_date": "08/02/2022 03:42:38",
      "content": "<p>Yes doing the inference in batches is probably your best best</p>\n<p>This notebook has a great example of how to do this<br>\n<a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1880750": "The test data is too big to read in and hold inside a dataframe all at once. I'd like to ask how what the solution to this problem is. Is it best to read the csv in as a batch, make predictions on those, and then write to a submission file?",
    "1880772": "Yes doing the inference in batches is probably your best best\n\nThis notebook has a great example of how to do this\nhttps://www.kaggle.com/code/cdeotte/xgboost-starter-0-793"
  },
  "source": "meta"
}