{
  "id": 478165,
  "title": "How the test set is processed?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/478165",
  "author_name": "Danu A.",
  "post_date": "2024-02-19T12:26:04.428000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Processing the train data is already a challenge for the RAM size (I know that Polars can help), but I'd like to know how the real test set is processed? As I'm not quite sure what kaggle do when switches the test data with the real one - the notebook is processed in the same environment that we have or in background on some other resources?</p>\n<p>In conclusion:<br>\n1) Should we do parallel processing (train set and our fake test together)<br>\nor<br>\n2) Should we process the train set, create the model and fit it, then delete train set from RAM and just after - start the test set process?</p>",
  "messages": [
    {
      "id": 2658814,
      "postDate": "2024-02-19T12:26:04.427Z",
      "content": "<p>Processing the train data is already a challenge for the RAM size (I know that Polars can help), but I'd like to know how the real test set is processed? As I'm not quite sure what kaggle do when switches the test data with the real one - the notebook is processed in the same environment that we have or in background on some other resources?</p>\n<p>In conclusion:<br>\n1) Should we do parallel processing (train set and our fake test together)<br>\nor<br>\n2) Should we process the train set, create the model and fit it, then delete train set from RAM and just after - start the test set process?</p>",
      "rawMarkdown": "Processing the train data is already a challenge for the RAM size (I know that Polars can help), but I'd like to know how the real test set is processed? As I'm not quite sure what kaggle do when switches the test data with the real one - the notebook is processed in the same environment that we have or in background on some other resources?\n\nIn conclusion:\n1) Should we do parallel processing (train set and our fake test together)\nor\n2) Should we process the train set, create the model and fit it, then delete train set from RAM and just after - start the test set process?\n\n",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2658814": "Processing the train data is already a challenge for the RAM size (I know that Polars can help), but I'd like to know how the real test set is processed? As I'm not quite sure what kaggle do when switches the test data with the real one - the notebook is processed in the same environment that we have or in background on some other resources?\n\nIn conclusion:\n1) Should we do parallel processing (train set and our fake test together)\nor\n2) Should we process the train set, create the model and fit it, then delete train set from RAM and just after - start the test set process?\n\n"
  }
}