{
  "id": 2610,
  "title": "Can we use additional fields available in 6GB file?",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2610",
  "author_name": "",
  "post_date": "2012-09-07T12:44:01.453Z",
  "votes": null,
  "comment_count": 2,
  "views": 1326,
  "content": "<p>There are many new fields available in the 2012-07StackOverflow.7z file. Can we use these additional fields in the model ?</p>",
  "messages": [
    {
      "id": "14056",
      "postDate": "09/07/2012 12:44:01",
      "content": "<p>There are many new fields available in the 2012-07StackOverflow.7z file. Can we use these additional fields in the model ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14062",
      "postDate": "09/07/2012 16:17:00",
      "content": "<p>This was <a href=\"http://www.kaggle.com/c/predict-closed-questions-on-stack-overflow/forums/t/2532/this-data-will-not-be-available-as-inputs\">\r\npreviously answered here</a>.</p>\r\n<p>The short of it is, you can use what you learn from the full data dump to choose and refine your approach; but you can't use it as actual training data. &nbsp;It won't be available for the final leaderboard.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14071",
      "postDate": "09/07/2012 16:48:05",
      "content": "<p>This is still a bit unclear. The additional data cannot be used in the final solution, what if I deduce rules from the data set and add them as code? What if the rules are like 'if user_id == 301 then ...'?</p>\r\n<p>Also, what if I clean up the obvious mistakes in classification?</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 14062,
      "author_name": "kevinmontrose",
      "author_url": "",
      "post_date": "09/07/2012 16:17:00",
      "content": "<p>This was <a href=\"http://www.kaggle.com/c/predict-closed-questions-on-stack-overflow/forums/t/2532/this-data-will-not-be-available-as-inputs\">\r\npreviously answered here</a>.</p>\r\n<p>The short of it is, you can use what you learn from the full data dump to choose and refine your approach; but you can't use it as actual training data. &nbsp;It won't be available for the final leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14071,
      "author_name": "melisgl",
      "author_url": "",
      "post_date": "09/07/2012 16:48:05",
      "content": "<p>This is still a bit unclear. The additional data cannot be used in the final solution, what if I deduce rules from the data set and add them as code? What if the rules are like 'if user_id == 301 then ...'?</p>\r\n<p>Also, what if I clean up the obvious mistakes in classification?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "14056": "",
    "14062": "",
    "14071": ""
  },
  "source": "meta"
}