{
  "id": 2437,
  "title": "Training/testing label accuracy?",
  "url": "/competitions/predict-closed-questions-on-stack-overflow/discussion/2437",
  "author_name": "",
  "post_date": "2012-08-24T22:44:47.057Z",
  "votes": 5,
  "comment_count": 3,
  "views": 2710,
  "content": "<p>Are the datasets used for evaluation hand-labelled, or are they drawn from the same population as the training dataset?</p>\r\n<p>It seems clear from the stackoverflow blog post that this is a raw data dump, and there are plenty of examples which should have been closed but were not, and some which shouldn't have been closed but were. &nbsp;Also, intuitively it's pretty hard to tell the\r\n difference between off-topic and non-constructive comments but there apparently are some differences in the data. &nbsp;My classifier finds things which are labeled open but probably fit the definition of non-constructive, and is penalized -8 or so for its trouble.</p>\r\n<p>The reason I ask is because the evaluation function we're using strongly penalizes confident decisions that don't agree with the label. &nbsp;If the label is wrong, you're better off hedging your bets slightly and leaving a little bit of weight on 'open'. &nbsp;I\r\n guess in some sense that's already built-in to the problem: P(closed) = P(should be closed) * P(was actually moderated), but we probably want to learn the &quot;should be closed&quot; probability instead in order to have a useful system.</p>\r\n<p>If you either hand-curate the training data by fixing up labels which are obviously wrong, or find an automatic method to do that (e.g. a hierarchical model with a strong but finite prior on the labels being true), you might have a more accurate classifier\r\n at the expense (for this contest) of a bigger penalty for badly labeled target data.</p>",
  "messages": [
    {
      "id": "13431",
      "postDate": "08/24/2012 22:44:47",
      "content": "<p>Are the datasets used for evaluation hand-labelled, or are they drawn from the same population as the training dataset?</p>\r\n<p>It seems clear from the stackoverflow blog post that this is a raw data dump, and there are plenty of examples which should have been closed but were not, and some which shouldn't have been closed but were. &nbsp;Also, intuitively it's pretty hard to tell the\r\n difference between off-topic and non-constructive comments but there apparently are some differences in the data. &nbsp;My classifier finds things which are labeled open but probably fit the definition of non-constructive, and is penalized -8 or so for its trouble.</p>\r\n<p>The reason I ask is because the evaluation function we're using strongly penalizes confident decisions that don't agree with the label. &nbsp;If the label is wrong, you're better off hedging your bets slightly and leaving a little bit of weight on 'open'. &nbsp;I\r\n guess in some sense that's already built-in to the problem: P(closed) = P(should be closed) * P(was actually moderated), but we probably want to learn the &quot;should be closed&quot; probability instead in order to have a useful system.</p>\r\n<p>If you either hand-curate the training data by fixing up labels which are obviously wrong, or find an automatic method to do that (e.g. a hierarchical model with a strong but finite prior on the labels being true), you might have a more accurate classifier\r\n at the expense (for this contest) of a bigger penalty for badly labeled target data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13450",
      "postDate": "08/25/2012 16:53:14",
      "content": "<p>This was essentially my question. I'd like to know the error rate in classification of the training set. Even if we had say 100 closed samples reviewed by two experts, that might be enough to construct a transition matrix where each entry is the probability\r\n that one label was classified as another label by the second expert.<br>\r\nWe have 5 labels (4 closed and 1 open) so it would be a 5x5 matrix<br>\r\nP(x,y) = probability that label x was re-classified as label y by another evaluation</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13535",
      "postDate": "08/27/2012 23:21:37",
      "content": "<p>Andy, thanks for the detailed comments / questions.</p>\r\n<p>[quote=Andy Sloane;13431]</p>\r\n<p>Are the datasets used for evaluation hand-labelled, or are they drawn from the same population as the training dataset?</p>\r\n<p>[/quote]None of the datasets are hand-labelled. The evaluation sets are an out-of-time sample, so these are only drawn from the same distribution as the training data to the extent that there is no temporal variation in the data.</p>\r\n<p>[quote=Andy Sloane;13431]It seems clear from the stackoverflow blog post that this is a raw data dump, and there are plenty of examples which should have been closed but were not, and some which shouldn't have been closed but were. &nbsp;Also, intuitively it's\r\n pretty hard to tell the difference between off-topic and non-constructive comments but there apparently are some differences in the data. &nbsp;My classifier finds things which are labeled open but probably fit the definition of non-constructive, and is penalized\r\n -8 or so for its trouble.[/quote]In making the evaluation dependent on whether the question is closed in practice as opposed to creating a &quot;gold standard&quot; set of questions that have been evaluated in a controlled setting, we are asking you to predict whether\r\n a question <strong>will be closed</strong>, instead of whether it <strong>should be closed</strong>. There are pros and cons to each approach, but we decided to go with the former, at least for this initial study.&nbsp;</p>\r\n<p>[quote=Andy Sloane;13431]The reason I ask is because the evaluation function we're using strongly penalizes confident decisions that don't agree with the label. &nbsp;If the label is wrong, you're better off hedging your bets slightly and leaving a little bit\r\n of weight on 'open'. &nbsp;I guess in some sense that's already built-in to the problem: P(closed) = P(should be closed) * P(was actually moderated), but we probably want to learn the &quot;should be closed&quot; probability instead in order to have a useful system.[/quote]Though\r\n the target of this problem is to predict whether a question will be closed instead of whether a question should be closed (which is less well-defined), do you have any ideas on how we could have directly targeted P(should be closed) without generating a &quot;gold\r\n standard&quot; set of manually labelled questions?</p>\r\n<p>[quote=Andy Sloane;13431]If you either hand-curate the training data by fixing up labels which are obviously wrong, or find an automatic method to do that (e.g. a hierarchical model with a strong but finite prior on the labels being true), you might have\r\n a more accurate classifier at the expense (for this contest) of a bigger penalty for badly labeled target data.[/quote]I'd considered doing this, but introducing a human in the loop makes it more difficult to retrain the algorithm on new data points and keep\r\n the model up to date. Additionally, it would introduce additional distortions in the source labels (/ reduce the amount of labeled training data we have to work with) while not fully correcting for the issue of label noise.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "13548",
      "postDate": "08/28/2012 06:41:01",
      "content": "<p>[quote=Ben Hamner;13535]</p>\r\n<p>None of the datasets are hand-labelled. The evaluation sets are an out-of-time sample, so these are only drawn from the same distribution as the training data to the extent that there is no temporal variation in the data.[/quote]</p>\r\n<p>It just dawned on me that the public leaderboard set is August 2012 only and the training set goes back to 2008. &nbsp;That explains a lot of phenomena I saw while training and should have been obvious in hindsight. &nbsp;(And I think that's why Amro is kicking everyone\r\n in the pants right now but I could be wrong.)</p>\r\n<p>[quote]Though the target of this problem is to predict whether a question will be closed instead of whether a question should be closed (which is less well-defined), do you have any ideas on how we could have directly targeted P(should be closed) without\r\n generating a &quot;gold standard&quot; set of manually labelled questions?[/quote]</p>\r\n<p>No, you pretty much have to hand-label. &nbsp;Yeah, it's fine that it's not a gold standard data set we're testing against, it's just good to confirm that this is what we're dealing with.</p>\r\n<p>Thanks for the response.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 13450,
      "author_name": "darkoram",
      "author_url": "",
      "post_date": "08/25/2012 16:53:14",
      "content": "<p>This was essentially my question. I'd like to know the error rate in classification of the training set. Even if we had say 100 closed samples reviewed by two experts, that might be enough to construct a transition matrix where each entry is the probability\r\n that one label was classified as another label by the second expert.<br>\r\nWe have 5 labels (4 closed and 1 open) so it would be a 5x5 matrix<br>\r\nP(x,y) = probability that label x was re-classified as label y by another evaluation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13535,
      "author_name": "benhamner",
      "author_url": "",
      "post_date": "08/27/2012 23:21:37",
      "content": "<p>Andy, thanks for the detailed comments / questions.</p>\r\n<p>[quote=Andy Sloane;13431]</p>\r\n<p>Are the datasets used for evaluation hand-labelled, or are they drawn from the same population as the training dataset?</p>\r\n<p>[/quote]None of the datasets are hand-labelled. The evaluation sets are an out-of-time sample, so these are only drawn from the same distribution as the training data to the extent that there is no temporal variation in the data.</p>\r\n<p>[quote=Andy Sloane;13431]It seems clear from the stackoverflow blog post that this is a raw data dump, and there are plenty of examples which should have been closed but were not, and some which shouldn't have been closed but were. &nbsp;Also, intuitively it's\r\n pretty hard to tell the difference between off-topic and non-constructive comments but there apparently are some differences in the data. &nbsp;My classifier finds things which are labeled open but probably fit the definition of non-constructive, and is penalized\r\n -8 or so for its trouble.[/quote]In making the evaluation dependent on whether the question is closed in practice as opposed to creating a &quot;gold standard&quot; set of questions that have been evaluated in a controlled setting, we are asking you to predict whether\r\n a question <strong>will be closed</strong>, instead of whether it <strong>should be closed</strong>. There are pros and cons to each approach, but we decided to go with the former, at least for this initial study.&nbsp;</p>\r\n<p>[quote=Andy Sloane;13431]The reason I ask is because the evaluation function we're using strongly penalizes confident decisions that don't agree with the label. &nbsp;If the label is wrong, you're better off hedging your bets slightly and leaving a little bit\r\n of weight on 'open'. &nbsp;I guess in some sense that's already built-in to the problem: P(closed) = P(should be closed) * P(was actually moderated), but we probably want to learn the &quot;should be closed&quot; probability instead in order to have a useful system.[/quote]Though\r\n the target of this problem is to predict whether a question will be closed instead of whether a question should be closed (which is less well-defined), do you have any ideas on how we could have directly targeted P(should be closed) without generating a &quot;gold\r\n standard&quot; set of manually labelled questions?</p>\r\n<p>[quote=Andy Sloane;13431]If you either hand-curate the training data by fixing up labels which are obviously wrong, or find an automatic method to do that (e.g. a hierarchical model with a strong but finite prior on the labels being true), you might have\r\n a more accurate classifier at the expense (for this contest) of a bigger penalty for badly labeled target data.[/quote]I'd considered doing this, but introducing a human in the loop makes it more difficult to retrain the algorithm on new data points and keep\r\n the model up to date. Additionally, it would introduce additional distortions in the source labels (/ reduce the amount of labeled training data we have to work with) while not fully correcting for the issue of label noise.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 13548,
      "author_name": "andysloane",
      "author_url": "",
      "post_date": "08/28/2012 06:41:01",
      "content": "<p>[quote=Ben Hamner;13535]</p>\r\n<p>None of the datasets are hand-labelled. The evaluation sets are an out-of-time sample, so these are only drawn from the same distribution as the training data to the extent that there is no temporal variation in the data.[/quote]</p>\r\n<p>It just dawned on me that the public leaderboard set is August 2012 only and the training set goes back to 2008. &nbsp;That explains a lot of phenomena I saw while training and should have been obvious in hindsight. &nbsp;(And I think that's why Amro is kicking everyone\r\n in the pants right now but I could be wrong.)</p>\r\n<p>[quote]Though the target of this problem is to predict whether a question will be closed instead of whether a question should be closed (which is less well-defined), do you have any ideas on how we could have directly targeted P(should be closed) without\r\n generating a &quot;gold standard&quot; set of manually labelled questions?[/quote]</p>\r\n<p>No, you pretty much have to hand-label. &nbsp;Yeah, it's fine that it's not a gold standard data set we're testing against, it's just good to confirm that this is what we're dealing with.</p>\r\n<p>Thanks for the response.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "13431": "",
    "13450": "",
    "13535": "",
    "13548": ""
  },
  "source": "meta"
}