{
  "id": 10792,
  "title": "Seizure prevalence",
  "url": "/competitions/seizure-prediction/discussion/10792",
  "author_name": "",
  "post_date": "2014-10-29T19:24:52.460Z",
  "votes": null,
  "comment_count": 3,
  "views": 1125,
  "content": "<p>Hi,&nbsp;</p>\n<p>It's sometimes nice to know the actual prevalence. I want to modify my model offset to reflect the subjects actual seizure rate. Can admin provide seizure prevalence for each subject?&nbsp;</p>",
  "messages": [
    {
      "id": "57012",
      "postDate": "10/29/2014 19:24:52",
      "content": "<p>Hi,&nbsp;</p>\n<p>It's sometimes nice to know the actual prevalence. I want to modify my model offset to reflect the subjects actual seizure rate. Can admin provide seizure prevalence for each subject?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57016",
      "postDate": "10/29/2014 19:58:19",
      "content": "<p>No. In order to mirror the real world problem we cannot disclose how many seizures are in the test set. You would need to estimate this from the training set, or some other combination of prior information.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57223",
      "postDate": "11/03/2014 01:20:53",
      "content": "<p>Thanks for reply. I don't necessarily want to know rate in test set. I'm more interested to know overall rate of patient 1 having a seizure or dog-1, etc. The logistic regression intercept is an estimate of &quot;success&quot; prevalence. If the case (preictal) to control (interictal) ratio is not representative of that subject (i.e. just by convenience) then the model will be biased.This isn't a prior. It's&nbsp;the&nbsp;overall seizure rate. For example, if we fit a model to predict heart failure with 160 case and 360 control then logistic intercept will be 160/(360+160) on logit scale. However, lets say the actual rate within our sampling area is 0.05, then estimated intercept is not correct (if we know true prevalence from population then we can simply adjust). If we combine trained models from multiple sampling areas with varying prevalence, then an overall AUC might get thrown off, I think</p>\n<p>I'm guessing you still won't provide, though</p>\n\n<p>-Brian</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57437",
      "postDate": "11/06/2014 21:06:03",
      "content": "<p>My apologies... I understand the dilemma but don't think I can provide that info.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 57016,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "10/29/2014 19:58:19",
      "content": "<p>No. In order to mirror the real world problem we cannot disclose how many seizures are in the test set. You would need to estimate this from the training set, or some other combination of prior information.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57223,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "11/03/2014 01:20:53",
      "content": "<p>Thanks for reply. I don't necessarily want to know rate in test set. I'm more interested to know overall rate of patient 1 having a seizure or dog-1, etc. The logistic regression intercept is an estimate of &quot;success&quot; prevalence. If the case (preictal) to control (interictal) ratio is not representative of that subject (i.e. just by convenience) then the model will be biased.This isn't a prior. It's&nbsp;the&nbsp;overall seizure rate. For example, if we fit a model to predict heart failure with 160 case and 360 control then logistic intercept will be 160/(360+160) on logit scale. However, lets say the actual rate within our sampling area is 0.05, then estimated intercept is not correct (if we know true prevalence from population then we can simply adjust). If we combine trained models from multiple sampling areas with varying prevalence, then an overall AUC might get thrown off, I think</p>\n<p>I'm guessing you still won't provide, though</p>\n\n<p>-Brian</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57437,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "11/06/2014 21:06:03",
      "content": "<p>My apologies... I understand the dilemma but don't think I can provide that info.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57012": "",
    "57016": "",
    "57223": "",
    "57437": ""
  },
  "source": "meta"
}