{
  "id": 2475,
  "title": "Class distribution",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2475",
  "author_name": "",
  "post_date": "2012-08-28T05:40:08.300Z",
  "votes": null,
  "comment_count": 4,
  "views": 1778,
  "content": "<p>Can anyone confirm this?</p>\r\n<p>Does train.csv have the following class distributions.</p>\r\n<p>open questions:&nbsp;3300392<br>\r\nclosed questions: 70136</p>\r\n<p>The reason for my confusion is from the fact that the sample-train.csv has the exact same number (70136) of data points.</p>",
  "messages": [
    {
      "id": "13547",
      "postDate": "08/28/2012 05:40:08",
      "content": "<p>Can anyone confirm this?</p>\r\n<p>Does train.csv have the following class distributions.</p>\r\n<p>open questions:&nbsp;3300392<br>\r\nclosed questions: 70136</p>\r\n<p>The reason for my confusion is from the fact that the sample-train.csv has the exact same number (70136) of data points.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13549",
      "postDate": "08/28/2012 06:42:20",
      "content": "<p>edit: nevermind.</p>\r\n<p>I don't have numbers handy, heh. &nbsp;I was just going to say that there should be an equal number of open and closed questions in train-sample.csv, so there should be twice that number. &nbsp;Hm.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13553",
      "postDate": "08/28/2012 09:29:49",
      "content": "<p>From the 'Data' page:</p>\r\n<p>&quot;The file train-sample.csv is a stratified sample of the training data: it contains every closed question and an equally-sized random sample of the open questions in the training data.&quot;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13556",
      "postDate": "08/28/2012 12:16:58",
      "content": "<p>I can confirm, I got the exact same results after parsing 'train.csv':</p>\r\n<pre>Open questions -  3300392<br>Closed questions -   70136</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13560",
      "postDate": "08/28/2012 14:18:34",
      "content": "<p>Great. Thank you all for confirming this.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13549,
      "author_name": "andysloane",
      "author_url": "",
      "post_date": "08/28/2012 06:42:20",
      "content": "<p>edit: nevermind.</p>\r\n<p>I don't have numbers handy, heh. &nbsp;I was just going to say that there should be an equal number of open and closed questions in train-sample.csv, so there should be twice that number. &nbsp;Hm.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13553,
      "author_name": "ashelly",
      "author_url": "",
      "post_date": "08/28/2012 09:29:49",
      "content": "<p>From the 'Data' page:</p>\r\n<p>&quot;The file train-sample.csv is a stratified sample of the training data: it contains every closed question and an equally-sized random sample of the open questions in the training data.&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13556,
      "author_name": "adamroth",
      "author_url": "",
      "post_date": "08/28/2012 12:16:58",
      "content": "<p>I can confirm, I got the exact same results after parsing 'train.csv':</p>\r\n<pre>Open questions -  3300392<br>Closed questions -   70136</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13560,
      "author_name": "hackinghabits",
      "author_url": "",
      "post_date": "08/28/2012 14:18:34",
      "content": "<p>Great. Thank you all for confirming this.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13547": "",
    "13549": "",
    "13553": "",
    "13556": "",
    "13560": ""
  },
  "source": "meta"
}