{
  "id": 2491,
  "title": "Some straightforward descriptions on using \"data\" files",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2491",
  "author_name": "",
  "post_date": "2012-08-29T09:06:12.670Z",
  "votes": null,
  "comment_count": 2,
  "views": 1480,
  "content": "<p>Hey everyone,</p>\r\n<p>well, I'm pretty new here, and I'm not a professional in this area also, however I've been spending the recent few years on implimanting AI in some little and simple projects of my own. I found Kaggle and this contest accidentally, however I'm wondering\r\n to join the competition and having fun while learning something with you guys. :)</p>\r\n<p>My misunderstood, which I'm sure sounds so beginners like questions - which I am!, is which 'data' files should I use on my algorithm and when? For example we have 'train' file, plus 'train-sample' and 'public_leaderboard'. In my mind, I have to build the\r\n way to predict whatever wanted, then I should provide some amount of real data for the input, and see what's gonna happenes then. By real data I mean all the open and closed questions plus some other useful data like users' account information and so on, for\r\n a particular time for example first week of August. The algorithm will provide some output which you can 'compare' with the real data in the next week to see how was the result accurate, right? ... Here we have three different sources, all similar, and without\r\n too much description. Just a few lines that this is the file and it contains these fields and finished.</p>\r\n<p>So, I'm just wondering if you guys just give me a little bit more clear description, it would be really good, and I could join the competition in a correct way! :)</p>\r\n<p>&nbsp;</p>\r\n<p>Thanks,</p>\r\n<p>Mahdi</p>",
  "messages": [
    {
      "id": "13597",
      "postDate": "08/29/2012 09:06:12",
      "content": "<p>Hey everyone,</p>\r\n<p>well, I'm pretty new here, and I'm not a professional in this area also, however I've been spending the recent few years on implimanting AI in some little and simple projects of my own. I found Kaggle and this contest accidentally, however I'm wondering\r\n to join the competition and having fun while learning something with you guys. :)</p>\r\n<p>My misunderstood, which I'm sure sounds so beginners like questions - which I am!, is which 'data' files should I use on my algorithm and when? For example we have 'train' file, plus 'train-sample' and 'public_leaderboard'. In my mind, I have to build the\r\n way to predict whatever wanted, then I should provide some amount of real data for the input, and see what's gonna happenes then. By real data I mean all the open and closed questions plus some other useful data like users' account information and so on, for\r\n a particular time for example first week of August. The algorithm will provide some output which you can 'compare' with the real data in the next week to see how was the result accurate, right? ... Here we have three different sources, all similar, and without\r\n too much description. Just a few lines that this is the file and it contains these fields and finished.</p>\r\n<p>So, I'm just wondering if you guys just give me a little bit more clear description, it would be really good, and I could join the competition in a correct way! :)</p>\r\n<p>&nbsp;</p>\r\n<p>Thanks,</p>\r\n<p>Mahdi</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13653",
      "postDate": "08/30/2012 17:12:56",
      "content": "<p>Use the <code>train</code> file to construct your solution, it contains all the input fields and the results.\r\n<code>train-sample</code> is just a smaller sampling of <code>train</code> data, which contains all closed questions but only a small set of the open ones.</p>\r\n<p>The <code>public_leaderboard</code> file contains just the inputs for some questions (that aren't found in\r\n<code>train</code>), your submissions consist of running your algorithm against <code>\r\npublic_leaderboard</code> and submitting it's results.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13664",
      "postDate": "08/30/2012 18:09:53",
      "content": "<p>Hey Kevin,</p>\r\n<p>Thank you! I was totally disappointed that anyone would give me an answer! Thanks! :)</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13653,
      "author_name": "kevinmontrose",
      "author_url": "",
      "post_date": "08/30/2012 17:12:56",
      "content": "<p>Use the <code>train</code> file to construct your solution, it contains all the input fields and the results.\r\n<code>train-sample</code> is just a smaller sampling of <code>train</code> data, which contains all closed questions but only a small set of the open ones.</p>\r\n<p>The <code>public_leaderboard</code> file contains just the inputs for some questions (that aren't found in\r\n<code>train</code>), your submissions consist of running your algorithm against <code>\r\npublic_leaderboard</code> and submitting it's results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13664,
      "author_name": "mahdi6",
      "author_url": "",
      "post_date": "08/30/2012 18:09:53",
      "content": "<p>Hey Kevin,</p>\r\n<p>Thank you! I was totally disappointed that anyone would give me an answer! Thanks! :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13597": "",
    "13653": "",
    "13664": ""
  },
  "source": "meta"
}