{
  "id": 10801,
  "title": "Using Patient-Specific Features",
  "url": "/competitions/seizure-prediction/discussion/10801",
  "author_name": "",
  "post_date": "2014-10-30T20:02:44.053Z",
  "votes": null,
  "comment_count": 4,
  "views": 1322,
  "content": "<p>Hey all,</p>\n<p>This is my first time working on a Kaggle problem, and I was wondering whether it is allowed to choose quite different features for each of the subjects. For example, after doing some exploratory analysis, I think I can &quot;see&quot; the difference between the two types of segments in both Patients 1 and 2. However, the way I would quantify these differences (and thus the features I would collect)&nbsp;would be quite different for the two different patients.</p>\n\n<p>Is it allowed to extract different kinds of features per patient and then do a per patient classifier? On the one hand, this approach is rather ad hoc and so for real world applications may not be useful. On the other hand, since the electrode placements within the (human) patients are not standardized, it would be quite hard to expect the same features to be relevant for each, although one might hope for a general feature extraction procedure which can be applied to new patients.</p>\n\n<p>One way to reconcile this would just be to find features that appear to work on a patient by patient basis, then collect ALL those features for ALL patients and find some sort of automated model selection which ends up choosing the useful ones for each. This might be allowed, but sounds like it would end up being highly computationally inefficient.&nbsp;</p>\n\n<p>Thanks,</p>\n<p>Mike</p>",
  "messages": [
    {
      "id": "57067",
      "postDate": "10/30/2014 20:02:44",
      "content": "<p>Hey all,</p>\n<p>This is my first time working on a Kaggle problem, and I was wondering whether it is allowed to choose quite different features for each of the subjects. For example, after doing some exploratory analysis, I think I can &quot;see&quot; the difference between the two types of segments in both Patients 1 and 2. However, the way I would quantify these differences (and thus the features I would collect)&nbsp;would be quite different for the two different patients.</p>\n\n<p>Is it allowed to extract different kinds of features per patient and then do a per patient classifier? On the one hand, this approach is rather ad hoc and so for real world applications may not be useful. On the other hand, since the electrode placements within the (human) patients are not standardized, it would be quite hard to expect the same features to be relevant for each, although one might hope for a general feature extraction procedure which can be applied to new patients.</p>\n\n<p>One way to reconcile this would just be to find features that appear to work on a patient by patient basis, then collect ALL those features for ALL patients and find some sort of automated model selection which ends up choosing the useful ones for each. This might be allowed, but sounds like it would end up being highly computationally inefficient.&nbsp;</p>\n\n<p>Thanks,</p>\n<p>Mike</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57076",
      "postDate": "10/30/2014 21:37:21",
      "content": "<p>The latter strategy is what you'll need to do here. If you read the&nbsp;<a href=\"http://www.kaggle.com/c/seizure-prediction/rules\">rules</a> of the competition, it pretty clear that you can't have a subject-specific solution (e.g., manually choosing specific features as you suggest).&nbsp;</p>\n<p>There are plenty of algorithms out there that have variable selection built into them (e.g., random forests). So, chances are, if you use these, the features that are predictive for a given dog/patient will emerge. Your best bet is to start with a lot of features.</p>\n<p>Alternatively, as long as feature selection is automated in some way, then you can &quot;choose&quot; separate features for each subject. For instance, if you first included some sort of feature selection algorithm (e.g., a simple t-test or F-test), and then only use those features showing good discrimination in the subsequent learning algorithm, that would be ok, too. The key is that the feature selection is algorithmic, and thus generalizable to new instances.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57175",
      "postDate": "11/02/2014 02:27:46",
      "content": "<p>if given a test file, <strong>can we use the filename to decide what fn to call (knowing the filename allows us to know this signal belongs to which subject)</strong>? seems answer will be no, but then what is the difference between:</p>\n<p># Not allowed:<br>if 'Dog_1' then foo()<br>if 'Patient_1' then bar()<br>...</p>\n<p>#Also Allowed</p>\n<p>for subject=1:N<br> train(subject)<br> test(subject)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57176",
      "postDate": "11/02/2014 02:34:36",
      "content": "<p>think I understand what above means but would like to confirm. Are the rules saying, we can train and test individual subjects separately but the same algorithm should be used to build the classifier for each subject? please let me know. If above is correct, then we do not need to build a classifier to classify if input signal belongs to dog 1, 2, 3, 4, 5 or patient 1, 2.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57276",
      "postDate": "11/04/2014 07:39:16",
      "content": "<p>Yep!</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 57076,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "10/30/2014 21:37:21",
      "content": "<p>The latter strategy is what you'll need to do here. If you read the&nbsp;<a href=\"http://www.kaggle.com/c/seizure-prediction/rules\">rules</a> of the competition, it pretty clear that you can't have a subject-specific solution (e.g., manually choosing specific features as you suggest).&nbsp;</p>\n<p>There are plenty of algorithms out there that have variable selection built into them (e.g., random forests). So, chances are, if you use these, the features that are predictive for a given dog/patient will emerge. Your best bet is to start with a lot of features.</p>\n<p>Alternatively, as long as feature selection is automated in some way, then you can &quot;choose&quot; separate features for each subject. For instance, if you first included some sort of feature selection algorithm (e.g., a simple t-test or F-test), and then only use those features showing good discrimination in the subsequent learning algorithm, that would be ok, too. The key is that the feature selection is algorithmic, and thus generalizable to new instances.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57175,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "11/02/2014 02:27:46",
      "content": "<p>if given a test file, <strong>can we use the filename to decide what fn to call (knowing the filename allows us to know this signal belongs to which subject)</strong>? seems answer will be no, but then what is the difference between:</p>\n<p># Not allowed:<br>if 'Dog_1' then foo()<br>if 'Patient_1' then bar()<br>...</p>\n<p>#Also Allowed</p>\n<p>for subject=1:N<br> train(subject)<br> test(subject)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57176,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "11/02/2014 02:34:36",
      "content": "<p>think I understand what above means but would like to confirm. Are the rules saying, we can train and test individual subjects separately but the same algorithm should be used to build the classifier for each subject? please let me know. If above is correct, then we do not need to build a classifier to classify if input signal belongs to dog 1, 2, 3, 4, 5 or patient 1, 2.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57276,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "11/04/2014 07:39:16",
      "content": "<p>Yep!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57067": "",
    "57076": "",
    "57175": "",
    "57176": "",
    "57276": ""
  },
  "source": "meta"
}