{
  "id": 21343,
  "title": "How to train xgb by chunk?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21343",
  "author_name": "",
  "post_date": "2016-05-31T15:16:02.427Z",
  "votes": null,
  "comment_count": 1,
  "views": 426,
  "content": "<p>How could we train xgb by chunk? If we split our train data into 10 chunks and train each chunk and predict the test data and then average these 10 predictions, how could guarantee each chunk has enough information of all users? Thank you.</p>",
  "messages": [
    {
      "id": "121998",
      "postDate": "05/31/2016 15:16:02",
      "content": "<p>How could we train xgb by chunk? If we split our train data into 10 chunks and train each chunk and predict the test data and then average these 10 predictions, how could guarantee each chunk has enough information of all users? Thank you.</p>",
      "rawMarkdown": "How could we train xgb by chunk? If we split our train data into 10 chunks and train each chunk and predict the test data and then average these 10 predictions, how could guarantee each chunk has enough information of all users? Thank you.",
      "votes": null
    },
    {
      "id": "122039",
      "postDate": "05/31/2016 22:55:37",
      "content": "<p>You can try to break the training set by country or other parameters that are clearly present in the test data. That will allow similarly split test data. This can be considered when preparing the solutions. I tried, but I get the results of 0.1987 (split on user_location_country). </p>\n\n<p>Maybe I have not tried it on a complete set of data (only is_booking==1 about 3M records in training set), so it gave this result. You can spend a lot of time, but do not get normal results.</p>\n\n<p>So I stopped trying. In addition, this calculation takes a very long time.</p>",
      "rawMarkdown": "You can try to break the training set by country or other parameters that are clearly present in the test data. That will allow similarly split test data. This can be considered when preparing the solutions. I tried, but I get the results of 0.1987 (split on user_location_country). \r\n\r\nMaybe I have not tried it on a complete set of data (only is_booking==1 about 3M records in training set), so it gave this result. You can spend a lot of time, but do not get normal results.\r\n\r\nSo I stopped trying. In addition, this calculation takes a very long time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122039,
      "author_name": "sorokinv",
      "author_url": "",
      "post_date": "05/31/2016 22:55:37",
      "content": "<p>You can try to break the training set by country or other parameters that are clearly present in the test data. That will allow similarly split test data. This can be considered when preparing the solutions. I tried, but I get the results of 0.1987 (split on user_location_country). </p>\n\n<p>Maybe I have not tried it on a complete set of data (only is_booking==1 about 3M records in training set), so it gave this result. You can spend a lot of time, but do not get normal results.</p>\n\n<p>So I stopped trying. In addition, this calculation takes a very long time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "121998": "How could we train xgb by chunk? If we split our train data into 10 chunks and train each chunk and predict the test data and then average these 10 predictions, how could guarantee each chunk has enough information of all users? Thank you.",
    "122039": "You can try to break the training set by country or other parameters that are clearly present in the test data. That will allow similarly split test data. This can be considered when preparing the solutions. I tried, but I get the results of 0.1987 (split on user_location_country). \r\n\r\nMaybe I have not tried it on a complete set of data (only is_booking==1 about 3M records in training set), so it gave this result. You can spend a lot of time, but do not get normal results.\r\n\r\nSo I stopped trying. In addition, this calculation takes a very long time."
  },
  "source": "meta"
}