{
  "id": 53560,
  "title": "XGBoost Python - trouble loading test data ?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53560",
  "author_name": "",
  "post_date": "2018-04-01T14:46:03.463356900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I having trouble in using entire test set on XGB model. If I understand right, when we use predict on test data, we got to load the entire data ? If so, please advise best way to prepare the final submission file ? </p>",
  "messages": [
    {
      "id": "307407",
      "postDate": "04/01/2018 14:46:03",
      "content": "<p>I having trouble in using entire test set on XGB model. If I understand right, when we use predict on test data, we got to load the entire data ? If so, please advise best way to prepare the final submission file ? </p>",
      "rawMarkdown": "I having trouble in using entire test set on XGB model. If I understand right, when we use predict on test data, we got to load the entire data ? If so, please advise best way to prepare the final submission file ?",
      "votes": null
    },
    {
      "id": "307521",
      "postDate": "04/01/2018 23:04:55",
      "content": "<p>You could predict by batch if limited in memory, loading from file and predicting by batch, clearing the batches from memory at the end of each step. Without looking at your code I think that if you run out of memory when loading the whole test sample is because you did not clear the train samples after training the classifier.</p>\n\n<p>Look at this kernel for instance <a href=\"https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code\">https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code</a> . The lines </p>\n\n<pre><code>del dtrain\ngc.collect()\n</code></pre>\n\n<p>are doing this job.</p>",
      "rawMarkdown": "You could predict by batch if limited in memory, loading from file and predicting by batch, clearing the batches from memory at the end of each step. Without looking at your code I think that if you run out of memory when loading the whole test sample is because you did not clear the train samples after training the classifier.\n\nLook at this kernel for instance https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code . The lines \n\n    del dtrain\n    gc.collect()\n\nare doing this job.",
      "votes": null
    },
    {
      "id": "307525",
      "postDate": "04/01/2018 23:34:11",
      "content": "<p>Great, Thank you giim.  Yes I am clearing with </p>\n\n<p>del train\n del test\n gc.collect()</p>\n\n<p>Still not sure on batch and storing the results. I will look into the link you have sent.</p>\n\n<p>Yep, I understood my issue is with OHE.  Thank you.</p>",
      "rawMarkdown": "Great, Thank you giim.  Yes I am clearing with \n\n del train\n del test\n gc.collect()\n\nStill not sure on batch and storing the results. I will look into the link you have sent.\n\nYep, I understood my issue is with OHE.  Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 307521,
      "author_name": "gimunu",
      "author_url": "",
      "post_date": "04/01/2018 23:04:55",
      "content": "<p>You could predict by batch if limited in memory, loading from file and predicting by batch, clearing the batches from memory at the end of each step. Without looking at your code I think that if you run out of memory when loading the whole test sample is because you did not clear the train samples after training the classifier.</p>\n\n<p>Look at this kernel for instance <a href=\"https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code\">https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code</a> . The lines </p>\n\n<pre><code>del dtrain\ngc.collect()\n</code></pre>\n\n<p>are doing this job.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 307525,
      "author_name": "bgopalakrishnan",
      "author_url": "",
      "post_date": "04/01/2018 23:34:11",
      "content": "<p>Great, Thank you giim.  Yes I am clearing with </p>\n\n<p>del train\n del test\n gc.collect()</p>\n\n<p>Still not sure on batch and storing the results. I will look into the link you have sent.</p>\n\n<p>Yep, I understood my issue is with OHE.  Thank you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "307407": "I having trouble in using entire test set on XGB model. If I understand right, when we use predict on test data, we got to load the entire data ? If so, please advise best way to prepare the final submission file ?",
    "307521": "You could predict by batch if limited in memory, loading from file and predicting by batch, clearing the batches from memory at the end of each step. Without looking at your code I think that if you run out of memory when loading the whole test sample is because you did not clear the train samples after training the classifier.\n\nLook at this kernel for instance https://www.kaggle.com/joaopmpeinado/talkingdata-xgboost-lb-0-966/code . The lines \n\n    del dtrain\n    gc.collect()\n\nare doing this job.",
    "307525": "Great, Thank you giim.  Yes I am clearing with \n\n del train\n del test\n gc.collect()\n\nStill not sure on batch and storing the results. I will look into the link you have sent.\n\nYep, I understood my issue is with OHE.  Thank you."
  },
  "source": "meta"
}