{
  "id": 10070,
  "title": "Probabilistic score normalization",
  "url": "/competitions/seizure-detection/discussion/10070",
  "author_name": "",
  "post_date": "2014-08-18T19:43:27.587Z",
  "votes": null,
  "comment_count": 2,
  "views": 1469,
  "content": "<p>Probably for calculating AUC, different thresholds will be applied on the whole results of all subjects entirely. Am I right?</p>\n<p>As the range of probabilistic outputs may differ from subject to subject, and from seizure to early detection (surely all of them are in the range of [0 1]), I was wondering if it is&nbsp;allowed to normalize the probabilistic outputs of each patient separately to [0 1], so that the minimum be 0 and maximum be one? And then make the final submission file?</p>\n<p>Appreciate your response.</p>",
  "messages": [
    {
      "id": "52153",
      "postDate": "08/18/2014 19:43:27",
      "content": "<p>Probably for calculating AUC, different thresholds will be applied on the whole results of all subjects entirely. Am I right?</p>\n<p>As the range of probabilistic outputs may differ from subject to subject, and from seizure to early detection (surely all of them are in the range of [0 1]), I was wondering if it is&nbsp;allowed to normalize the probabilistic outputs of each patient separately to [0 1], so that the minimum be 0 and maximum be one? And then make the final submission file?</p>\n<p>Appreciate your response.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52156",
      "postDate": "08/18/2014 20:46:40",
      "content": "<p>My understanding based upon previous posts is that this is okay if the normalization constant is calculated using the training data. &nbsp;But it seems that dividing by a constant that is the sum of the test prediction probabilities is making use of a &quot;time machine&quot; to calculate the constant.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52199",
      "postDate": "08/19/2014 21:06:59",
      "content": "<p>Yes- George has it right. Use any normalization you like on the training data.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 52156,
      "author_name": "george04",
      "author_url": "",
      "post_date": "08/18/2014 20:46:40",
      "content": "<p>My understanding based upon previous posts is that this is okay if the normalization constant is calculated using the training data. &nbsp;But it seems that dividing by a constant that is the sum of the test prediction probabilities is making use of a &quot;time machine&quot; to calculate the constant.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52199,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "08/19/2014 21:06:59",
      "content": "<p>Yes- George has it right. Use any normalization you like on the training data.&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "52153": "",
    "52156": "",
    "52199": ""
  },
  "source": "meta"
}