{
  "id": 2609,
  "title": "How to Get User Reputation at Post Creation Time",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2609",
  "author_name": "",
  "post_date": "2012-09-07T12:23:18.120Z",
  "votes": null,
  "comment_count": 1,
  "views": 1155,
  "content": "<p>Hi,</p>\r\n<p>In the train-sample.csv files there are few fields which and we are supposed to use only those to train our models.</p>\r\n<pre>However I am unable to find few fields like :</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>1) ReputationAtPostCreation --&gt; users.xml ( But this is not same as User_Reputation at creation time)<br><br></pre>\r\n<pre>2) OwnerUndeletedAnswerCountAtPostTime&nbsp;</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>Please let us know how to derive these fields.</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>Thanks<br><br></pre>",
  "messages": [
    {
      "id": "14055",
      "postDate": "09/07/2012 12:23:18",
      "content": "<p>Hi,</p>\r\n<p>In the train-sample.csv files there are few fields which and we are supposed to use only those to train our models.</p>\r\n<pre>However I am unable to find few fields like :</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>1) ReputationAtPostCreation --&gt; users.xml ( But this is not same as User_Reputation at creation time)<br><br></pre>\r\n<pre>2) OwnerUndeletedAnswerCountAtPostTime&nbsp;</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>Please let us know how to derive these fields.</pre>\r\n<pre>&nbsp;</pre>\r\n<pre>Thanks<br><br></pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "14065",
      "postDate": "09/07/2012 16:25:45",
      "content": "<p>ReputationAtPostCreation isn't really available in the full data dumps, we went through some pain to reconstruct it for the purposes of this contest. The trick is that not all voting is public, so we only make some of it available in the data dumps (in particular\r\n you can't determine most votes that have <em>negative</em> effects on reputation).</p>\r\n<p>Since we're pretty confident Reputation is an important feature we wanted to make it available for the contest, but we can't release a full history without NDA'ing everyone (which we don't want to do).</p>\r\n<p>OwnerUndeletedAnswerCountAtPostTime is likewise a bit tricky to reconstruct, you can approximate it by counting answers from the same user (Question.OwnerUserId = Answer.OwnerUserId_ with CreationDates less than the target Question's. Deleted content isn't\r\n available in the data dumps, so there will be some differences where an Answer was deleted\r\n<em>after</em> a Question was posted.</p>\r\n<p>You should be treating the training CSVs as canonical, and the XML data as supplementary, in part because the XML data dumps weren't specially constructed for this contest (but mostly because the CSVs are all that will be available for the final solution\r\n judgement).</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 14065,
      "author_name": "kevinmontrose",
      "author_url": "",
      "post_date": "09/07/2012 16:25:45",
      "content": "<p>ReputationAtPostCreation isn't really available in the full data dumps, we went through some pain to reconstruct it for the purposes of this contest. The trick is that not all voting is public, so we only make some of it available in the data dumps (in particular\r\n you can't determine most votes that have <em>negative</em> effects on reputation).</p>\r\n<p>Since we're pretty confident Reputation is an important feature we wanted to make it available for the contest, but we can't release a full history without NDA'ing everyone (which we don't want to do).</p>\r\n<p>OwnerUndeletedAnswerCountAtPostTime is likewise a bit tricky to reconstruct, you can approximate it by counting answers from the same user (Question.OwnerUserId = Answer.OwnerUserId_ with CreationDates less than the target Question's. Deleted content isn't\r\n available in the data dumps, so there will be some differences where an Answer was deleted\r\n<em>after</em> a Question was posted.</p>\r\n<p>You should be treating the training CSVs as canonical, and the XML data as supplementary, in part because the XML data dumps weren't specially constructed for this contest (but mostly because the CSVs are all that will be available for the final solution\r\n judgement).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "14055": "",
    "14065": ""
  },
  "source": "meta"
}