{
  "id": 2532,
  "title": "This data will not be available as inputs",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2532",
  "author_name": "",
  "post_date": "2012-09-01T13:05:33.533Z",
  "votes": null,
  "comment_count": 3,
  "views": 1928,
  "content": "<p>Hi,</p>\r\n<p>I'm a little bit confused with this (related to the 6GB dataset) :</p>\r\n<blockquote>\r\n<p><span>This data will not be available as inputs, but may be useful in building your solution.</span></p>\r\n</blockquote>\r\n<p>Does it means that we can't use theses data to improove the quality of our models?&nbsp;For exemple, i think that some variables can add a lots of informations like the User location, or the user AboutMe page.</p>",
  "messages": [
    {
      "id": "13751",
      "postDate": "09/01/2012 13:05:33",
      "content": "<p>Hi,</p>\r\n<p>I'm a little bit confused with this (related to the 6GB dataset) :</p>\r\n<blockquote>\r\n<p><span>This data will not be available as inputs, but may be useful in building your solution.</span></p>\r\n</blockquote>\r\n<p>Does it means that we can't use theses data to improove the quality of our models?&nbsp;For exemple, i think that some variables can add a lots of informations like the User location, or the user AboutMe page.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13837",
      "postDate": "09/03/2012 16:17:09",
      "content": "<p>The data is made available to provide further insight into how Stack Overflow functions, you may be able to glean something useful when it comes to choosing algorithms, weighting features, or what have you.</p>\r\n<p>What it shouldn't be used as is a training set, as it won't be available for the final submission.</p>\r\n<p>While there is some data in there that's both public and probably useful, when we structured this contest we started from a known &quot;safe to publish&quot; data set and added new things; we naturally didn't get absolutely everything in there. We expect any solution\r\n to benefit from additional data (and before hitting production we'd definitely be incorporating some private data), so we're not particularly concerned about a few omissions.</p>\r\n<p>As an aside, the reason those two columns in particular aren't included in the training set is that we can't reconstruct them historically; we have their current state, not their state at an arbitrary point in time.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14452",
      "postDate": "09/18/2012 07:18:29",
      "content": "<p>&quot;This data will not be available as inputs&quot;</p>\r\n<p>&nbsp;</p>\r\n<p>Which interpretation is intended?</p>\r\n<p>&quot;You are free to train on this and submit the 6G file as external data, just remember that no more data of the same kind is going to be provided when the final training set is published so for the final evaluation you'll only have inputs like public_leaderboard.csv.&quot;</p>\r\n<p>or</p>\r\n<p>&quot;You cannot even train on this data, hence you cannot submit it as external data either.&quot;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14491",
      "postDate": "09/18/2012 19:01:42",
      "content": "<blockquote>\r\n<p><span>&quot;What it shouldn't be used as is a training set, as it won't be available for the final submission.&quot;</span></p>\r\n</blockquote>\r\n<p><span>It's the second one.</span></p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13837,
      "author_name": "kevinmontrose",
      "author_url": "",
      "post_date": "09/03/2012 16:17:09",
      "content": "<p>The data is made available to provide further insight into how Stack Overflow functions, you may be able to glean something useful when it comes to choosing algorithms, weighting features, or what have you.</p>\r\n<p>What it shouldn't be used as is a training set, as it won't be available for the final submission.</p>\r\n<p>While there is some data in there that's both public and probably useful, when we structured this contest we started from a known &quot;safe to publish&quot; data set and added new things; we naturally didn't get absolutely everything in there. We expect any solution\r\n to benefit from additional data (and before hitting production we'd definitely be incorporating some private data), so we're not particularly concerned about a few omissions.</p>\r\n<p>As an aside, the reason those two columns in particular aren't included in the training set is that we can't reconstruct them historically; we have their current state, not their state at an arbitrary point in time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14452,
      "author_name": "melisgl",
      "author_url": "",
      "post_date": "09/18/2012 07:18:29",
      "content": "<p>&quot;This data will not be available as inputs&quot;</p>\r\n<p>&nbsp;</p>\r\n<p>Which interpretation is intended?</p>\r\n<p>&quot;You are free to train on this and submit the 6G file as external data, just remember that no more data of the same kind is going to be provided when the final training set is published so for the final evaluation you'll only have inputs like public_leaderboard.csv.&quot;</p>\r\n<p>or</p>\r\n<p>&quot;You cannot even train on this data, hence you cannot submit it as external data either.&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 14491,
      "author_name": "kevinmontrose",
      "author_url": "",
      "post_date": "09/18/2012 19:01:42",
      "content": "<blockquote>\r\n<p><span>&quot;What it shouldn't be used as is a training set, as it won't be available for the final submission.&quot;</span></p>\r\n</blockquote>\r\n<p><span>It's the second one.</span></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13751": "",
    "13837": "",
    "14452": "",
    "14491": ""
  },
  "source": "meta"
}