{
  "id": 47832,
  "title": "Never mix test and training data",
  "url": "/competitions/sp-society-camera-model-identification/discussion/47832",
  "author_name": "",
  "post_date": "2018-01-19T13:44:29.077495200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I've heard that statement on multiple occuasions. And I definitely agree. </p>\n\n<p>But being fairly new to ML, I have a question about this; </p>\n\n<p>When I use the training-set to train my model, and also create a lot columns based on a single text-column (pd.get_dummies), I end up with 15 columns let's say.</p>\n\n<p>However, when I go through the same process on my test-set, I cannot garantee that it will also end up being 15 columns, and this is a must if I want to predict on the test-set when my model was trained on the training-set.</p>\n\n<p>What is best practices with regards to this sort of problem?</p>",
  "messages": [
    {
      "id": "271019",
      "postDate": "01/19/2018 13:44:29",
      "content": "<p>I've heard that statement on multiple occuasions. And I definitely agree. </p>\n\n<p>But being fairly new to ML, I have a question about this; </p>\n\n<p>When I use the training-set to train my model, and also create a lot columns based on a single text-column (pd.get_dummies), I end up with 15 columns let's say.</p>\n\n<p>However, when I go through the same process on my test-set, I cannot garantee that it will also end up being 15 columns, and this is a must if I want to predict on the test-set when my model was trained on the training-set.</p>\n\n<p>What is best practices with regards to this sort of problem?</p>",
      "rawMarkdown": "I've heard that statement on multiple occuasions. And I definitely agree. \n\nBut being fairly new to ML, I have a question about this; \n\nWhen I use the training-set to train my model, and also create a lot columns based on a single text-column (pd.get_dummies), I end up with 15 columns let's say.\n\nHowever, when I go through the same process on my test-set, I cannot garantee that it will also end up being 15 columns, and this is a must if I want to predict on the test-set when my model was trained on the training-set.\n\nWhat is best practices with regards to this sort of problem?",
      "votes": null
    },
    {
      "id": "271022",
      "postDate": "01/19/2018 13:46:03",
      "content": "<p>Apologies - I posted this in the wrong forum.</p>",
      "rawMarkdown": "Apologies - I posted this in the wrong forum.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 271022,
      "author_name": "nilzone",
      "author_url": "",
      "post_date": "01/19/2018 13:46:03",
      "content": "<p>Apologies - I posted this in the wrong forum.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "271019": "I've heard that statement on multiple occuasions. And I definitely agree. \n\nBut being fairly new to ML, I have a question about this; \n\nWhen I use the training-set to train my model, and also create a lot columns based on a single text-column (pd.get_dummies), I end up with 15 columns let's say.\n\nHowever, when I go through the same process on my test-set, I cannot garantee that it will also end up being 15 columns, and this is a must if I want to predict on the test-set when my model was trained on the training-set.\n\nWhat is best practices with regards to this sort of problem?",
    "271022": "Apologies - I posted this in the wrong forum."
  },
  "source": "meta"
}