{
  "id": 8347,
  "title": "Using test set distribution",
  "url": "/competitions/seizure-detection/discussion/8347",
  "author_name": "",
  "post_date": "2014-05-31T21:26:00.317Z",
  "votes": 1,
  "comment_count": 11,
  "views": 2878,
  "content": "<p>Hello everyone,</p>\n\n<p>I was wondering that is it ok to use the test set distribution information in training the prediction model. I am just worried that the test set distribution will not be available in real time application hence the model will not be practical. Such as implementing the covariate shift method by using the test set distribution to tweak the prediction model.</p>\n<p>In short, can the information from test set be used to readjust the prediction model or only the training samples are allowed to be used in the prediction model?</p>\n<p>Cheers</p>",
  "messages": [
    {
      "id": "47413",
      "postDate": "05/31/2014 21:26:00",
      "content": "<p>Hello everyone,</p>\n\n<p>I was wondering that is it ok to use the test set distribution information in training the prediction model. I am just worried that the test set distribution will not be available in real time application hence the model will not be practical. Such as implementing the covariate shift method by using the test set distribution to tweak the prediction model.</p>\n<p>In short, can the information from test set be used to readjust the prediction model or only the training samples are allowed to be used in the prediction model?</p>\n<p>Cheers</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48481",
      "postDate": "06/02/2014 14:28:35",
      "content": "<p>I just spoke&nbsp;with the competition hosts&nbsp;and they ruled&nbsp;that semi-supervised learning is&nbsp;off limits for this competition. One&nbsp;would not have access to the test set in the clinical setting (because it would be in the future), and we want the competition to mirror the clinical setting as closely as possible.</p>\n<p>Thanks for the great question.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48509",
      "postDate": "06/02/2014 18:41:39",
      "content": "<p>Thanks for the precision.</p>\n<p>Does this limitation also apply to preprocessing or other kind of unsupervised processing of the data (for example, a re-scaling of the classification scores or a normalization of the features). I understand that the hosts want a method which can be used online, but there is a lot of ways to implement this kind of thing in an adaptive (and efficient) way and therefore respect the causality of the signals.</p>\n<p>There is a clear shift between training and test for some subjects, a simple normalization of the data can make a huge&nbsp;difference&nbsp;on&nbsp;the results.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48510",
      "postDate": "06/02/2014 18:51:13",
      "content": "<p>I informally call it the time machine rule: if you need a time machine for any part of the training process, it is not allowed. Your entire training process (training, parameters, cross-validation, scaling, etc.) can't use the test set. &nbsp;The real-world test set hasn't been created yet.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48513",
      "postDate": "06/02/2014 19:42:01",
      "content": "<p>Hmm, Interesting,</p>\n<p>I think this is a flaw in the way the competition has been setup, because in the real-time application, a baseline could be adjusted to know the ground truth (using upto t-1 samples), whereas, here we are given the test trials where the sequence of the test samples is also randomized. In real-time application the algorithms can adapt to the shift in distribution by using causal previous samples. This can be implemented in the current data if the sequence of the trials is also provided so that we can use test samples til trial <strong>t-1</strong> to adapt to the shift for trail <strong>t</strong>. Currently, the methods developed here would be suboptimal because algorithms would not be able to readjust to the baseline. A better test set would have been a continuous stream&nbsp; of data, lets say an hour test data&nbsp; and algorithms would have to detect where the seizure started and how long did it last (which is more closer to how a real-time application would be).&nbsp;</p>\n<p>Furthermore, I think the leader-board do not represent the valid positions as many teams would have used test-set distribution to re shift their prediction models.</p>\n<p>Its just a thought but i think these are some valid points.</p>\n<p>Cheers</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48517",
      "postDate": "06/02/2014 20:54:48",
      "content": "<p>[quote=Chaos;48513]</p>\n<p>Hmm, Interesting,</p>\n<p>I think this is a flaw in the way the competition has been setup, because in the real-time application, a baseline could be adjusted to know the ground truth (using upto t-1 samples), whereas, here we are given the test trials where the sequence of the test samples is also randomized. In real-time application the algorithms can adapt to the shift in distribution by using causal previous samples. This can be implemented in the current data if the sequence of the trials is also provided so that we can use test samples til trial <strong>t-1</strong> to adapt to the shift for trail <strong>t</strong>. Currently, the methods developed here would be suboptimal because algorithms would not be able to readjust to the baseline. A better test set would have been a continuous stream&nbsp; of data, lets say an hour test data&nbsp; and algorithms would have to detect where the seizure started and how long did it last (which is more closer to how a real-time application would be).&nbsp;</p>\n<p>Furthermore, I think the leader-board do not represent the valid positions as many teams would have used test-set distribution to re shift their prediction models.</p>\n<p>Its just a thought but i think these are some valid points.</p>\n<p>Cheers</p>\n<p>[/quote]</p>\n<p>I can relate.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48518",
      "postDate": "06/02/2014 21:11:06",
      "content": "<p>[quote=Chaos;48513]</p>\n<p>This can be implemented in the current data if the sequence of the trials is also provided so that we can use test samples til trial <strong>t-1</strong> to adapt to the shift for trail <strong>t</strong>. Currently, the methods developed here would be suboptimal because algorithms would not be able to readjust to the baseline. A better test set would have been a continuous stream&nbsp; of data, lets say an hour test data&nbsp; and algorithms would have to detect where the seizure started and how long did it last (which is more closer to how a real-time application would be).&nbsp;</p>\n<p>[/quote]</p>\n<p>Exactly, this is what i was thinking when i said adaptation. It is well know that the EEG (or iEEG) signals are not stable over time. For a viable long term application, adaptation must be done to take into account this shift in the baseline. I've done this in several of my online experiments, with a great success.</p>\n<p>On the other hand, i understand that the hosts didn't want to provide continuous signals, as it will be too easy to crack the test signals by a visual inspection and fine-tune all the parameters to achieves a perfect score.</p>\n<p>The best way to set-up this kind of competition is to give continuous signal for both training and test, and evaluates the final performances on a unknown dataset. I've seen this in the past,&nbsp;but i'm not sure it is possible to do this with kaggle.</p>\n<p>my fear is that the current competition and dataset may be too restricted to see any significant improvement over the state-of-the-art.</p>\n<p>Anyway, i don't&nbsp;make any use of the test data in my method, so my score remains&nbsp;valid ;-)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48543",
      "postDate": "06/03/2014 02:07:22",
      "content": "<p>Thanks everyone for your excellent comments and suggestions. We are planning a competition to follow this one on a related topic, and we may use these suggestions in its design.&nbsp;</p>\n<p>It's true that by dividing the data up into discrete segments we've prevented&nbsp;algorithms from doing adaptive baseline correction, which many prior authors in seizure detection have used&nbsp;with some success. However, the practical limitations of providing&nbsp;any kind of continuous data for a contest like this made this approach necessary. As Alexandre points out - providing continuous data allows any contestant to begin by identifying the late portion of a seizure, which often has a high amplitude and can be&nbsp;fairly easy to detect, and then look&nbsp;backwards in time to find where the EEG data first shows a significant change. This would have vastly reduced the difficulty of the challenge, and would have made the results useless for responsive neurostimulation.&nbsp;The only way we could think of to do this was by streaming test data live to contestants, which we don't have the technical ability to do, and would require all contestants to provide classifications online simultaneously. An interesting possibility, but not practical currently.</p>\n<p>In addition, by making the challenge more difficult, it's possible&nbsp;a great algorithm can be easily made to perform better by adding adaptive normalization as a pre-processing step. The EEG data will be made available on ieeg.org after the contest, and interested contestants could do exactly this and post their results.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48771",
      "postDate": "06/06/2014 11:14:39",
      "content": "<p>[quote=bbrinkm;48543]</p>\n\n<p>It's true that by dividing the data up into discrete segments we've prevented&nbsp;algorithms from doing adaptive baseline correction, which many prior authors in seizure detection have used&nbsp;with some success.</p>\n<p>[/quote]</p>\n<p>But, still, we could be able to build the continuous signals from the discrete segments. For example, for each segment we could look for another segment whose beginning is the most correlated with it's ending.</p>\n<p>Is acceptable to do this with the training data?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48774",
      "postDate": "06/06/2014 13:03:26",
      "content": "<p>[quote=Jose M.;48771]</p>\n<p>Is acceptable to do this with the training data?</p>\n<p>[/quote]</p>\n\n<p>The training data are provided in sequential order, so there is no need to rearrange them.</p>\n\n<p>Regarding this rule, it could be good after all. It will force the contestants to find features which are more &quot;baseline invariant&quot;.</p>\n<p>Anyway, you should&nbsp;add&nbsp;this limitation in the rules. I'm not sure everyone is (or will be) aware of this. Things need to be clear, people here are smart enough to exploit every leaks in the data.</p>\n<p>I think something like &quot;<strong>The output probabilities of a test segment should not be function of another test segment in any way</strong>&quot; would do the job.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48777",
      "postDate": "06/06/2014 13:09:29",
      "content": "<p>[quote=Alexandre;48774]</p>\n<p>[quote=Jose M.;48771]</p>\n<p>Is acceptable to do this with the training data?</p>\n<p>[/quote]</p>\n<p>The training data are provided in sequential order, so there is no need to rearrange them.</p>\n<p>[/quote]</p>\n\n<p>The ictal examples are provided with an order reference, but the interictal ones are not. What I mean in my post is to reconstruct the interictal-ictal-interictal sequences in order to work on them as a whole.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48778",
      "postDate": "06/06/2014 13:15:11",
      "content": "<p>[quote=Jose M.;48777]</p>\n<p>The ictal examples are provided with an order reference, but the interictal ones are not. What I mean in my post is to reconstruct the interictal-ictal-interictal sequences in order to work on them as a whole.</p>\n<p>[/quote]</p>\n\n<p>the data description page says : &quot;Training data are arranged sequentially while testing data are in random order.&quot;&nbsp;</p>\n<p>This is true for both ictal and interictal trainin data. The only thing we don't know about interictal data is when a new sequence end or begin (we know this for ictal through the latency variable).</p>\n<p>But i agree, it would be nice to have this information.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 48481,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "06/02/2014 14:28:35",
      "content": "<p>I just spoke&nbsp;with the competition hosts&nbsp;and they ruled&nbsp;that semi-supervised learning is&nbsp;off limits for this competition. One&nbsp;would not have access to the test set in the clinical setting (because it would be in the future), and we want the competition to mirror the clinical setting as closely as possible.</p>\n<p>Thanks for the great question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48509,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/02/2014 18:41:39",
      "content": "<p>Thanks for the precision.</p>\n<p>Does this limitation also apply to preprocessing or other kind of unsupervised processing of the data (for example, a re-scaling of the classification scores or a normalization of the features). I understand that the hosts want a method which can be used online, but there is a lot of ways to implement this kind of thing in an adaptive (and efficient) way and therefore respect the causality of the signals.</p>\n<p>There is a clear shift between training and test for some subjects, a simple normalization of the data can make a huge&nbsp;difference&nbsp;on&nbsp;the results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48510,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "06/02/2014 18:51:13",
      "content": "<p>I informally call it the time machine rule: if you need a time machine for any part of the training process, it is not allowed. Your entire training process (training, parameters, cross-validation, scaling, etc.) can't use the test set. &nbsp;The real-world test set hasn't been created yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48513,
      "author_name": "chaotic",
      "author_url": "",
      "post_date": "06/02/2014 19:42:01",
      "content": "<p>Hmm, Interesting,</p>\n<p>I think this is a flaw in the way the competition has been setup, because in the real-time application, a baseline could be adjusted to know the ground truth (using upto t-1 samples), whereas, here we are given the test trials where the sequence of the test samples is also randomized. In real-time application the algorithms can adapt to the shift in distribution by using causal previous samples. This can be implemented in the current data if the sequence of the trials is also provided so that we can use test samples til trial <strong>t-1</strong> to adapt to the shift for trail <strong>t</strong>. Currently, the methods developed here would be suboptimal because algorithms would not be able to readjust to the baseline. A better test set would have been a continuous stream&nbsp; of data, lets say an hour test data&nbsp; and algorithms would have to detect where the seizure started and how long did it last (which is more closer to how a real-time application would be).&nbsp;</p>\n<p>Furthermore, I think the leader-board do not represent the valid positions as many teams would have used test-set distribution to re shift their prediction models.</p>\n<p>Its just a thought but i think these are some valid points.</p>\n<p>Cheers</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48517,
      "author_name": "rybnik",
      "author_url": "",
      "post_date": "06/02/2014 20:54:48",
      "content": "<p>[quote=Chaos;48513]</p>\n<p>Hmm, Interesting,</p>\n<p>I think this is a flaw in the way the competition has been setup, because in the real-time application, a baseline could be adjusted to know the ground truth (using upto t-1 samples), whereas, here we are given the test trials where the sequence of the test samples is also randomized. In real-time application the algorithms can adapt to the shift in distribution by using causal previous samples. This can be implemented in the current data if the sequence of the trials is also provided so that we can use test samples til trial <strong>t-1</strong> to adapt to the shift for trail <strong>t</strong>. Currently, the methods developed here would be suboptimal because algorithms would not be able to readjust to the baseline. A better test set would have been a continuous stream&nbsp; of data, lets say an hour test data&nbsp; and algorithms would have to detect where the seizure started and how long did it last (which is more closer to how a real-time application would be).&nbsp;</p>\n<p>Furthermore, I think the leader-board do not represent the valid positions as many teams would have used test-set distribution to re shift their prediction models.</p>\n<p>Its just a thought but i think these are some valid points.</p>\n<p>Cheers</p>\n<p>[/quote]</p>\n<p>I can relate.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48518,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/02/2014 21:11:06",
      "content": "<p>[quote=Chaos;48513]</p>\n<p>This can be implemented in the current data if the sequence of the trials is also provided so that we can use test samples til trial <strong>t-1</strong> to adapt to the shift for trail <strong>t</strong>. Currently, the methods developed here would be suboptimal because algorithms would not be able to readjust to the baseline. A better test set would have been a continuous stream&nbsp; of data, lets say an hour test data&nbsp; and algorithms would have to detect where the seizure started and how long did it last (which is more closer to how a real-time application would be).&nbsp;</p>\n<p>[/quote]</p>\n<p>Exactly, this is what i was thinking when i said adaptation. It is well know that the EEG (or iEEG) signals are not stable over time. For a viable long term application, adaptation must be done to take into account this shift in the baseline. I've done this in several of my online experiments, with a great success.</p>\n<p>On the other hand, i understand that the hosts didn't want to provide continuous signals, as it will be too easy to crack the test signals by a visual inspection and fine-tune all the parameters to achieves a perfect score.</p>\n<p>The best way to set-up this kind of competition is to give continuous signal for both training and test, and evaluates the final performances on a unknown dataset. I've seen this in the past,&nbsp;but i'm not sure it is possible to do this with kaggle.</p>\n<p>my fear is that the current competition and dataset may be too restricted to see any significant improvement over the state-of-the-art.</p>\n<p>Anyway, i don't&nbsp;make any use of the test data in my method, so my score remains&nbsp;valid ;-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48543,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "06/03/2014 02:07:22",
      "content": "<p>Thanks everyone for your excellent comments and suggestions. We are planning a competition to follow this one on a related topic, and we may use these suggestions in its design.&nbsp;</p>\n<p>It's true that by dividing the data up into discrete segments we've prevented&nbsp;algorithms from doing adaptive baseline correction, which many prior authors in seizure detection have used&nbsp;with some success. However, the practical limitations of providing&nbsp;any kind of continuous data for a contest like this made this approach necessary. As Alexandre points out - providing continuous data allows any contestant to begin by identifying the late portion of a seizure, which often has a high amplitude and can be&nbsp;fairly easy to detect, and then look&nbsp;backwards in time to find where the EEG data first shows a significant change. This would have vastly reduced the difficulty of the challenge, and would have made the results useless for responsive neurostimulation.&nbsp;The only way we could think of to do this was by streaming test data live to contestants, which we don't have the technical ability to do, and would require all contestants to provide classifications online simultaneously. An interesting possibility, but not practical currently.</p>\n<p>In addition, by making the challenge more difficult, it's possible&nbsp;a great algorithm can be easily made to perform better by adding adaptive normalization as a pre-processing step. The EEG data will be made available on ieeg.org after the contest, and interested contestants could do exactly this and post their results.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48771,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "06/06/2014 11:14:39",
      "content": "<p>[quote=bbrinkm;48543]</p>\n\n<p>It's true that by dividing the data up into discrete segments we've prevented&nbsp;algorithms from doing adaptive baseline correction, which many prior authors in seizure detection have used&nbsp;with some success.</p>\n<p>[/quote]</p>\n<p>But, still, we could be able to build the continuous signals from the discrete segments. For example, for each segment we could look for another segment whose beginning is the most correlated with it's ending.</p>\n<p>Is acceptable to do this with the training data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48774,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/06/2014 13:03:26",
      "content": "<p>[quote=Jose M.;48771]</p>\n<p>Is acceptable to do this with the training data?</p>\n<p>[/quote]</p>\n\n<p>The training data are provided in sequential order, so there is no need to rearrange them.</p>\n\n<p>Regarding this rule, it could be good after all. It will force the contestants to find features which are more &quot;baseline invariant&quot;.</p>\n<p>Anyway, you should&nbsp;add&nbsp;this limitation in the rules. I'm not sure everyone is (or will be) aware of this. Things need to be clear, people here are smart enough to exploit every leaks in the data.</p>\n<p>I think something like &quot;<strong>The output probabilities of a test segment should not be function of another test segment in any way</strong>&quot; would do the job.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48777,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "06/06/2014 13:09:29",
      "content": "<p>[quote=Alexandre;48774]</p>\n<p>[quote=Jose M.;48771]</p>\n<p>Is acceptable to do this with the training data?</p>\n<p>[/quote]</p>\n<p>The training data are provided in sequential order, so there is no need to rearrange them.</p>\n<p>[/quote]</p>\n\n<p>The ictal examples are provided with an order reference, but the interictal ones are not. What I mean in my post is to reconstruct the interictal-ictal-interictal sequences in order to work on them as a whole.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48778,
      "author_name": "alexandrebarachant",
      "author_url": "",
      "post_date": "06/06/2014 13:15:11",
      "content": "<p>[quote=Jose M.;48777]</p>\n<p>The ictal examples are provided with an order reference, but the interictal ones are not. What I mean in my post is to reconstruct the interictal-ictal-interictal sequences in order to work on them as a whole.</p>\n<p>[/quote]</p>\n\n<p>the data description page says : &quot;Training data are arranged sequentially while testing data are in random order.&quot;&nbsp;</p>\n<p>This is true for both ictal and interictal trainin data. The only thing we don't know about interictal data is when a new sequence end or begin (we know this for ictal through the latency variable).</p>\n<p>But i agree, it would be nice to have this information.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "47413": "",
    "48481": "",
    "48509": "",
    "48510": "",
    "48513": "",
    "48517": "",
    "48518": "",
    "48543": "",
    "48771": "",
    "48774": "",
    "48777": "",
    "48778": ""
  },
  "source": "meta"
}