{
  "id": 346443,
  "title": "train test split vs CV",
  "url": "/competitions/amex-default-prediction/discussion/346443",
  "author_name": "Andy Atkinson",
  "post_date": "2022-08-19T13:35:13.731000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Models take a long time to train on this large dataset.  I'm doing a simple train test split, like 85/15, for 5X faster training speed.  Of course, this puts more reliance on the lb score to gauge true model performance, but this risk might be lower than normal because CV has correlated so well with lb.  For me the train test split seems good enough to fit the number of boosting rounds, which in most cases is \"a lot\" and hitting the nail on the head exactly is probably not too important.</p>\n<p>Anyone else succumbing to impatience and saying farewell to CV in the final stages of this particular competition?  It's been noted that even CV is volatile, so a train test split eval is even more-so. But we have a pretty reliable lb given the huge test set.</p>",
  "messages": [
    {
      "id": 1905960,
      "postDate": "2022-08-19T13:35:13.730Z",
      "content": "<p>Models take a long time to train on this large dataset.  I'm doing a simple train test split, like 85/15, for 5X faster training speed.  Of course, this puts more reliance on the lb score to gauge true model performance, but this risk might be lower than normal because CV has correlated so well with lb.  For me the train test split seems good enough to fit the number of boosting rounds, which in most cases is \"a lot\" and hitting the nail on the head exactly is probably not too important.</p>\n<p>Anyone else succumbing to impatience and saying farewell to CV in the final stages of this particular competition?  It's been noted that even CV is volatile, so a train test split eval is even more-so. But we have a pretty reliable lb given the huge test set.</p>",
      "rawMarkdown": "Models take a long time to train on this large dataset.  I'm doing a simple train test split, like 85/15, for 5X faster training speed.  Of course, this puts more reliance on the lb score to gauge true model performance, but this risk might be lower than normal because CV has correlated so well with lb.  For me the train test split seems good enough to fit the number of boosting rounds, which in most cases is \"a lot\" and hitting the nail on the head exactly is probably not too important.\n\nAnyone else succumbing to impatience and saying farewell to CV in the final stages of this particular competition?  It's been noted that even CV is volatile, so a train test split eval is even more-so. But we have a pretty reliable lb given the huge test set.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1905960": "Models take a long time to train on this large dataset.  I'm doing a simple train test split, like 85/15, for 5X faster training speed.  Of course, this puts more reliance on the lb score to gauge true model performance, but this risk might be lower than normal because CV has correlated so well with lb.  For me the train test split seems good enough to fit the number of boosting rounds, which in most cases is \"a lot\" and hitting the nail on the head exactly is probably not too important.\n\nAnyone else succumbing to impatience and saying farewell to CV in the final stages of this particular competition?  It's been noted that even CV is volatile, so a train test split eval is even more-so. But we have a pretty reliable lb given the huge test set."
  }
}