{
  "id": 9412,
  "title": "Per Subject Classifiers or Universal Classifier?",
  "url": "/competitions/seizure-detection/discussion/9412",
  "author_name": "",
  "post_date": "2014-06-07T22:03:18.803Z",
  "votes": null,
  "comment_count": 14,
  "views": 3326,
  "content": "<p>Hi,</p>\n\n<p>I would like to know whether we are allowed to train our classifiers separately for each subject (i.e. dog 1, dog 2, human 1, etc.), or whether we should only train one universal classifier once and perform the predictions for all subjects using that one.</p>\n\n<p>Thanks,</p>\n<p>Anthony</p>",
  "messages": [
    {
      "id": "48847",
      "postDate": "06/07/2014 22:03:18",
      "content": "<p>Hi,</p>\n\n<p>I would like to know whether we are allowed to train our classifiers separately for each subject (i.e. dog 1, dog 2, human 1, etc.), or whether we should only train one universal classifier once and perform the predictions for all subjects using that one.</p>\n\n<p>Thanks,</p>\n<p>Anthony</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48854",
      "postDate": "06/08/2014 03:08:19",
      "content": "<p>I have the same question. I suppose either way is valid, insofar as this is not an &quot;asking for too much information&quot; question; how will the predictions be made in real life? Will lab techs be training models for each new patient, or classifying universally with the same models?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48867",
      "postDate": "06/08/2014 17:27:42",
      "content": "<p>As per the rules there is no restriction on subject specific models. Im sure its perfectly fine to train subject-specific models based on reapective subjects train data or pooled trainig data. As long as the model can be applied to a new subject given its trainig data.</p>\n<p>Cheers</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48897",
      "postDate": "06/09/2014 12:42:46",
      "content": "<p>The rules page was fairly clear on this, IMHO:</p>\n\n<p style=\"padding-left: 30px\">Any changes to the methodology across species/subjects must be done in an automated way, so that your approach will generalize to new subjects.</p>\n<pre style=\"padding-left: 60px\"># Not allowed:<br>if 'Dog_1' then foo()<br>if 'Patient_1' then bar()<br>...<br><br># Allowed<br>if f(signal) &lt; 2 then foo()<br>else bar()</pre>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48899",
      "postDate": "06/09/2014 12:49:44",
      "content": "<p>Thanks Will! Hoping you can clarify the evaluation criteria for me as well in <a href=\"http://www.kaggle.com/c/seizure-detection/forums/t/8311/evaluation-criteria-written-by-google-translate\">this</a>&nbsp;thread.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48916",
      "postDate": "06/09/2014 17:06:45",
      "content": "<p>Hi William!</p>\n\n<p>Does the automated way include a case where we consider each folder separately?&nbsp;For example, for&nbsp;each folder (i.e. subject folder) train using the training data within that folder and test using the test data within that folder - with no subject specific code (e.g. for dog 1 do that, etc.). That is what was not clear to me from the rules page.</p>\n\n<p>Thanks,</p>\n<p>Anthony</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49021",
      "postDate": "06/11/2014 18:40:37",
      "content": "<p>This guideline is not a clear. The following function will generalize to new subjects, yet allows discrimination between Dogs and Patients in the current contest.</p>\n\n<p>if 'freq' &lt;= 400 then foo()&nbsp;&nbsp;&nbsp;&nbsp; # Dog case</p>\n<p>if 400 &lt; 'freq' &lt; 600 then bar()&nbsp;&nbsp; # Patient 1 case</p>\n<p>if 'freq' &gt; 600 then foobar()&nbsp;&nbsp;&nbsp; # Other patients</p>\n\n<p>Not allowed?</p>\n\n<p>[quote=William Cukierski;48897]</p>\n<p>The rules page was fairly clear on this, IMHO:</p>\n<p style=\"padding-left: 30px\">Any changes to the methodology across species/subjects must be done in an automated way, so that your approach will generalize to new subjects.</p>\n<pre style=\"padding-left: 60px\"># Not allowed:<br>if 'Dog_1' then foo()<br>if 'Patient_1' then bar()<br>...<br><br># Allowed<br>if f(signal) &lt; 2 then foo()<br>else bar()</pre>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49065",
      "postDate": "06/12/2014 17:24:58",
      "content": "<p>The guideline requires using your judgement. In your example, frequency is acting as a proxy for the individual subjects. It would not generalize to future subjects or new data,&nbsp;so you should not use such a rule.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49097",
      "postDate": "06/13/2014 11:29:22",
      "content": "<p>[quote=Anthony Platanios;48916]</p>\n<p>Hi William!</p>\n<p>Does the automated way include a case where we consider each folder separately?&nbsp;For example, for&nbsp;each folder (i.e. subject folder) train using the training data within that folder and test using the test data within that folder - with no subject specific code (e.g. for dog 1 do that, etc.). That is what was not clear to me from the rules page.</p>\n<p>Thanks,</p>\n<p>Anthony</p>\n<p>[/quote]</p>\n<p>This is a good questions. It is also not clear to me. From an application point-of-view, I don't think it makes too much sense to enforce&nbsp;solutions&nbsp;that generalise over dogs and human patients. However, now that I read this thread, I am not sure anymore.</p>\n\n\n<p>Also: &quot;Any changes to the methodology across species/subjects must be done in an automated way, so that your approach will generalize to new subjects.&quot; (from the rules) is a very broad statement, given that the data of each subject is so different (frequency, #channels, etc).&nbsp;For example, should our approaches generalise to subjects&nbsp;that are monitored by only one electrode? Or to snippets with a length of 2 seconds?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49129",
      "postDate": "06/13/2014 21:50:35",
      "content": "<p>Ah I am really confused now. I thought the reason the test and train for each subject are provided in their separate respective folder is to train subject specific models in a way that given training data for a subject the generalized model can predict the test trial. If this was not the case, then why all the test trials not provided in a general/common pooled directory!</p>\n\n<p>I think the way the data has been provided the following should be allowed:</p>\n\n<p><strong>Predict label for subject <em>n</em> given the training data for that subject <em>(n)</em> using a generalized common model.</strong></p>\n\n<p>i.e.</p>\n\n<p><strong>if f(testSignal<sub>n</sub> | trainSignal<sub>n</sub>) &lt; 2 then foo()<br>else bar()</strong></p>\n<p><strong>where <em>n</em> is the subject number</strong></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49247",
      "postDate": "06/17/2014 09:54:17",
      "content": "<p>This is important.</p>\n\n<p>William, is a per-subject classification scheme valid or not?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49252",
      "postDate": "06/17/2014 13:17:44",
      "content": "<p>Let me try to clarify-</p>\n\n<p>It is fine to train subject specific models. Chaos is right- this is why the folders were provided as they were. In the responsive neurostimulation application, we would expect to have to train an algorithm/approach for each patient. Also, as Ulo Gulo accuratly points out, electrode placement and configurations in this data set differ, so there's not a very good way to generalize an algorithm trained on pooled data. </p>\n\n<p>In real applications, data will be acquired in a continuous stream and processed in near real time, rather than as 1 second clips. The clips are a necessity for the contest. The requirement is to generalize over the range of data presented in this contest. You're right, there's an infinite range of possible electrode and data configurations, but we don't expect contestants to account for this.</p>\n\n<p>The sampling frequency case is interesting because you can separate dogs and humans by this metric - this was unintended, and please bear in mind that much clinical iEEG data currently is acquired sampled at 256-500 Hz. The fact that the contest data can be separated by sampling frequency is a peculiarity of us trying to provide the best and most novel iEEG data possible for the contest. So it would really not be appropriate to use the sampling frequency as a way to artificially use different approaches for dogs and humans, as described by Alexander VR.&nbsp; </p>\n\n<p><br>I hope this helps. Keep the questions coming...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "51105",
      "postDate": "07/23/2014 21:20:10",
      "content": "<p>Hi</p>\n<p>Following up on this thread..Did anyone try the&nbsp; classification (subject specific and pooled subjects)..Any thoughts which seems to be better? :)</p>\n<p>Thanks</p>\n<p>Navin</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "51107",
      "postDate": "07/23/2014 23:08:06",
      "content": "<p>My bet is on the subject specific approach. That's why I made so many (unsuccessful) submissions! :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "51110",
      "postDate": "07/23/2014 23:13:39",
      "content": "<p>Thanks Jose for your thoughts !!</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 48854,
      "author_name": "cavaunpeu",
      "author_url": "",
      "post_date": "06/08/2014 03:08:19",
      "content": "<p>I have the same question. I suppose either way is valid, insofar as this is not an &quot;asking for too much information&quot; question; how will the predictions be made in real life? Will lab techs be training models for each new patient, or classifying universally with the same models?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48867,
      "author_name": "chaotic",
      "author_url": "",
      "post_date": "06/08/2014 17:27:42",
      "content": "<p>As per the rules there is no restriction on subject specific models. Im sure its perfectly fine to train subject-specific models based on reapective subjects train data or pooled trainig data. As long as the model can be applied to a new subject given its trainig data.</p>\n<p>Cheers</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48897,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "06/09/2014 12:42:46",
      "content": "<p>The rules page was fairly clear on this, IMHO:</p>\n\n<p style=\"padding-left: 30px\">Any changes to the methodology across species/subjects must be done in an automated way, so that your approach will generalize to new subjects.</p>\n<pre style=\"padding-left: 60px\"># Not allowed:<br>if 'Dog_1' then foo()<br>if 'Patient_1' then bar()<br>...<br><br># Allowed<br>if f(signal) &lt; 2 then foo()<br>else bar()</pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48899,
      "author_name": "cavaunpeu",
      "author_url": "",
      "post_date": "06/09/2014 12:49:44",
      "content": "<p>Thanks Will! Hoping you can clarify the evaluation criteria for me as well in <a href=\"http://www.kaggle.com/c/seizure-detection/forums/t/8311/evaluation-criteria-written-by-google-translate\">this</a>&nbsp;thread.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48916,
      "author_name": "eaplatanios",
      "author_url": "",
      "post_date": "06/09/2014 17:06:45",
      "content": "<p>Hi William!</p>\n\n<p>Does the automated way include a case where we consider each folder separately?&nbsp;For example, for&nbsp;each folder (i.e. subject folder) train using the training data within that folder and test using the test data within that folder - with no subject specific code (e.g. for dog 1 do that, etc.). That is what was not clear to me from the rules page.</p>\n\n<p>Thanks,</p>\n<p>Anthony</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49021,
      "author_name": "alexvr",
      "author_url": "",
      "post_date": "06/11/2014 18:40:37",
      "content": "<p>This guideline is not a clear. The following function will generalize to new subjects, yet allows discrimination between Dogs and Patients in the current contest.</p>\n\n<p>if 'freq' &lt;= 400 then foo()&nbsp;&nbsp;&nbsp;&nbsp; # Dog case</p>\n<p>if 400 &lt; 'freq' &lt; 600 then bar()&nbsp;&nbsp; # Patient 1 case</p>\n<p>if 'freq' &gt; 600 then foobar()&nbsp;&nbsp;&nbsp; # Other patients</p>\n\n<p>Not allowed?</p>\n\n<p>[quote=William Cukierski;48897]</p>\n<p>The rules page was fairly clear on this, IMHO:</p>\n<p style=\"padding-left: 30px\">Any changes to the methodology across species/subjects must be done in an automated way, so that your approach will generalize to new subjects.</p>\n<pre style=\"padding-left: 60px\"># Not allowed:<br>if 'Dog_1' then foo()<br>if 'Patient_1' then bar()<br>...<br><br># Allowed<br>if f(signal) &lt; 2 then foo()<br>else bar()</pre>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49065,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "06/12/2014 17:24:58",
      "content": "<p>The guideline requires using your judgement. In your example, frequency is acting as a proxy for the individual subjects. It would not generalize to future subjects or new data,&nbsp;so you should not use such a rule.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49097,
      "author_name": "ugrossek",
      "author_url": "",
      "post_date": "06/13/2014 11:29:22",
      "content": "<p>[quote=Anthony Platanios;48916]</p>\n<p>Hi William!</p>\n<p>Does the automated way include a case where we consider each folder separately?&nbsp;For example, for&nbsp;each folder (i.e. subject folder) train using the training data within that folder and test using the test data within that folder - with no subject specific code (e.g. for dog 1 do that, etc.). That is what was not clear to me from the rules page.</p>\n<p>Thanks,</p>\n<p>Anthony</p>\n<p>[/quote]</p>\n<p>This is a good questions. It is also not clear to me. From an application point-of-view, I don't think it makes too much sense to enforce&nbsp;solutions&nbsp;that generalise over dogs and human patients. However, now that I read this thread, I am not sure anymore.</p>\n\n\n<p>Also: &quot;Any changes to the methodology across species/subjects must be done in an automated way, so that your approach will generalize to new subjects.&quot; (from the rules) is a very broad statement, given that the data of each subject is so different (frequency, #channels, etc).&nbsp;For example, should our approaches generalise to subjects&nbsp;that are monitored by only one electrode? Or to snippets with a length of 2 seconds?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49129,
      "author_name": "chaotic",
      "author_url": "",
      "post_date": "06/13/2014 21:50:35",
      "content": "<p>Ah I am really confused now. I thought the reason the test and train for each subject are provided in their separate respective folder is to train subject specific models in a way that given training data for a subject the generalized model can predict the test trial. If this was not the case, then why all the test trials not provided in a general/common pooled directory!</p>\n\n<p>I think the way the data has been provided the following should be allowed:</p>\n\n<p><strong>Predict label for subject <em>n</em> given the training data for that subject <em>(n)</em> using a generalized common model.</strong></p>\n\n<p>i.e.</p>\n\n<p><strong>if f(testSignal<sub>n</sub> | trainSignal<sub>n</sub>) &lt; 2 then foo()<br>else bar()</strong></p>\n<p><strong>where <em>n</em> is the subject number</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49247,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "06/17/2014 09:54:17",
      "content": "<p>This is important.</p>\n\n<p>William, is a per-subject classification scheme valid or not?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49252,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "06/17/2014 13:17:44",
      "content": "<p>Let me try to clarify-</p>\n\n<p>It is fine to train subject specific models. Chaos is right- this is why the folders were provided as they were. In the responsive neurostimulation application, we would expect to have to train an algorithm/approach for each patient. Also, as Ulo Gulo accuratly points out, electrode placement and configurations in this data set differ, so there's not a very good way to generalize an algorithm trained on pooled data. </p>\n\n<p>In real applications, data will be acquired in a continuous stream and processed in near real time, rather than as 1 second clips. The clips are a necessity for the contest. The requirement is to generalize over the range of data presented in this contest. You're right, there's an infinite range of possible electrode and data configurations, but we don't expect contestants to account for this.</p>\n\n<p>The sampling frequency case is interesting because you can separate dogs and humans by this metric - this was unintended, and please bear in mind that much clinical iEEG data currently is acquired sampled at 256-500 Hz. The fact that the contest data can be separated by sampling frequency is a peculiarity of us trying to provide the best and most novel iEEG data possible for the contest. So it would really not be appropriate to use the sampling frequency as a way to artificially use different approaches for dogs and humans, as described by Alexander VR.&nbsp; </p>\n\n<p><br>I hope this helps. Keep the questions coming...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 51105,
      "author_name": "navincota",
      "author_url": "",
      "post_date": "07/23/2014 21:20:10",
      "content": "<p>Hi</p>\n<p>Following up on this thread..Did anyone try the&nbsp; classification (subject specific and pooled subjects)..Any thoughts which seems to be better? :)</p>\n<p>Thanks</p>\n<p>Navin</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 51107,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "07/23/2014 23:08:06",
      "content": "<p>My bet is on the subject specific approach. That's why I made so many (unsuccessful) submissions! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 51110,
      "author_name": "navincota",
      "author_url": "",
      "post_date": "07/23/2014 23:13:39",
      "content": "<p>Thanks Jose for your thoughts !!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "48847": "",
    "48854": "",
    "48867": "",
    "48897": "",
    "48899": "",
    "48916": "",
    "49021": "",
    "49065": "",
    "49097": "",
    "49129": "",
    "49247": "",
    "49252": "",
    "51105": "",
    "51107": "",
    "51110": ""
  },
  "source": "meta"
}