{
  "id": 10790,
  "title": "Use of test data",
  "url": "/competitions/seizure-prediction/discussion/10790",
  "author_name": "",
  "post_date": "2014-10-29T18:18:57.617Z",
  "votes": 2,
  "comment_count": 29,
  "views": 6761,
  "content": "<p>Hi all,</p>\n<p>We've received a number of questions offline about using the test data. It took some time to reach a decision, but <strong>we have decided to allow use of the test data to calibrate your predictions</strong>. Apologies for any interim confusion and lack of clarity on whether this was allowed. It was a difficult choice given the tradeoffs between using the algorithm in the real world vs. enforcing a fair competition.</p>\n<p>We are not able to extend the competition deadline due to timing of the AES meeting in December. Instead, we will be increasing the daily submission limit (starting today) to allow you some extra leeway. Thanks for your continued hard work on this problem.</p>",
  "messages": [
    {
      "id": "57006",
      "postDate": "10/29/2014 18:18:57",
      "content": "<p>Hi all,</p>\n<p>We've received a number of questions offline about using the test data. It took some time to reach a decision, but <strong>we have decided to allow use of the test data to calibrate your predictions</strong>. Apologies for any interim confusion and lack of clarity on whether this was allowed. It was a difficult choice given the tradeoffs between using the algorithm in the real world vs. enforcing a fair competition.</p>\n<p>We are not able to extend the competition deadline due to timing of the AES meeting in December. Instead, we will be increasing the daily submission limit (starting today) to allow you some extra leeway. Thanks for your continued hard work on this problem.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57011",
      "postDate": "10/29/2014 19:21:35",
      "content": "<p>will there still be a blind test set for final results?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57013",
      "postDate": "10/29/2014 19:36:59",
      "content": "<p>The test set you have is the test set used to score the competition.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57014",
      "postDate": "10/29/2014 19:38:37",
      "content": "<p>Right, but is this statement still true &quot;The final results will be based on the other 60%, so the final standings may be different.&quot;</p>\n<p>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57015",
      "postDate": "10/29/2014 19:42:43",
      "content": "<p>Yes</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57022",
      "postDate": "10/29/2014 21:51:12",
      "content": "<p>Since the procedure of providing the full test set (keeping unknown the public/private identification) to the competitors is the standard one in Kaggle competitions, I would guess the same issue of using test data for improving the predictions on the test set itself should have risen before. What was the rule then?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57025",
      "postDate": "10/29/2014 22:06:50",
      "content": "<p>We have historically allowed semi-supervised learning on the grounds that it is very hard to enforce a rule against it. However, it's a decision that is still made on a case-by-case basis and can be&nbsp;more detracting for some problems than others. e.g. a machine learning&nbsp;problem&nbsp;on a text corpus is a very different beast&nbsp;than a time-series forecast.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57107",
      "postDate": "10/31/2014 13:44:29",
      "content": "<p>Sorry to ask a stupid question here, but how does one use unlabeled test data to calibrate a model? Are there any good resources we might read for this?</p>\n<p>Both isotonic regression and platt scaling (discussed <a href=\"http://www.kaggle.com/c/seizure-prediction/forums/t/10542/good-cross-validation-poor-test-set-result\">here</a>) seem to require mapping model output (e.g., probabilities) to a known class label. Using the same data for training and calibration clearly introduces bias. Thus it seems we're stuck with withholding part of the training set. In this case, though, there is such a small number of positive (i.e., 'preictal') instances, this seem like it would be unwise.</p>\n<p>One obvious answer, then, seems to be to use out-of-bag data to calibrate in each bagging iteration. Still, this begs the question, where does (or could) the unlabeled test data come into play?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57851",
      "postDate": "11/11/2014 21:58:29",
      "content": "<p>William,</p>\n<p>Is the 40/60 test split completely at random? So that seizure rates within an individual remain the same between 40% LB and 60% blind sets</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57858",
      "postDate": "11/11/2014 22:31:16",
      "content": "<p>[quote=Maineiac;57107]</p>\n<p>Sorry to ask a stupid question here, but how does one use unlabeled test data to calibrate a model? Are there any good resources we might read for this?</p>\n<p>One obvious answer, then, seems to be to use out-of-bag data to calibrate in each bagging iteration. Still, this begs the question, where does (or could) the unlabeled test data come into play?</p>\n<p>[/quote]</p>\n<p>In brief, in some computer vision related tasks the pipeline might be the following:</p>\n<p>1)&nbsp;Extract some patches or descriptors from images</p>\n<p>2) Cluster them to create vocabulary</p>\n<p>3) Then&nbsp;represent images as vectors of vocabulary entries</p>\n<p>4) Train smth on this vectors</p>\n<p>This is called&nbsp;<a href=\"http://en.wikipedia.org/wiki/Bag-of-words_model_in_computer_vision\" target=\"_blank\">Bag of&nbsp;Visual Words</a>. Since clustering is unsupervised, you can create vocabulary using unlabeled data (for example test data), which might improve overall performance.</p>\n\n<p>P.S. That's not an example of model calibration but test data might help sometimes.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58166",
      "postDate": "11/17/2014 00:39:24",
      "content": "<p>Since today is the last day, can we increase the number of submission by 10 more? I have a lot of models I want to check.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58237",
      "postDate": "11/17/2014 23:45:33",
      "content": "<p>Hi if you are able to extend the deadline at least until sometime later today that would be great.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58238",
      "postDate": "11/17/2014 23:54:20",
      "content": "<p>I had to find some control data (i.e.- no seizures)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58256",
      "postDate": "11/18/2014 02:15:18",
      "content": "<p>Kaggle &#8211; in the future, in similar competitions, can you cut out a small random gap between segments? Even taking out few samples will help the competition a lot.</p>\n<p>I this competition I noticed that you can undo the random shuffling of test segments and restore the original sequences. Knowing the sequence of segments in the test data can be used to improve the score but it will not help much the real problem we are trying to solve.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58276",
      "postDate": "11/18/2014 04:44:25",
      "content": "<p>[quote=zzspar;58256]</p>\n<p>Kaggle &#8211; in the future, in similar competitions, can you cut out a small random gap between segments? Even taking out few samples will help the competition a lot.</p>\n<p>I this competition I noticed that you can undo the random shuffling of test segments and restore the original sequences. Knowing the sequence of segments in the test data can be used to improve the score but it will not help much the real problem we are trying to solve.</p>\n<p>[/quote]</p>\n<p>The host did trim the ends of the clips and mean centered them. How do you know you were able to reverse engineer the order?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58295",
      "postDate": "11/18/2014 10:17:21",
      "content": "<p>I tried it out and it helpd the public LB...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58304",
      "postDate": "11/18/2014 11:42:27",
      "content": "<p>Now that the competition is over, can someone (Will perhaps?) explain <strong>how we can use unlabeled test data to calibrate our model predictions?</strong></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58308",
      "postDate": "11/18/2014 12:06:00",
      "content": "<p>I used min-max scaling of &nbsp;test probabilities for each subject separately. This gave an improvement</p>\n<p>of ~0.015 for my best model on public LB. &nbsp;I think this trick worked because my predicted probabilities were in very different ranges for each of the subjects, but&nbsp;I am not sure if it's helpful for other models.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58329",
      "postDate": "11/18/2014 18:32:22",
      "content": "<p>[quote=Maineiac;58304]</p>\n<p>Now that the competition is over, can someone (Will perhaps?) explain <strong>how we can use unlabeled test data to calibrate our model predictions?</strong></p>\n<p>[/quote]</p>\n<p>Imagine you do CV to select a model. You can do a random CV (lets name it CV1), you can do random preserving chunk integrity if you have more than a single frame from a 10min chunk (CV2) or you can do random preserving an event integrity (CV3) that is all chunks from an event are either used for test or for train. Your CV performance will increase from CV3 to CV1, say 75-85-95, first by having examples of a particular preictal event, second by having examples of a particular event chunk. The real-life testing stage follows&nbsp;the CV3 condition. However, if you create your own labels from the test data and then incorporate this data to your classifier then you can approximate CV2 condition with a corresponding increase in the performance. This is given that the organisers used a CV1 40-60 split between public and private. If they used CV3 split, then using testing data could be a trap. Personally, I disagree with allowing to use test data in any form or sense. It makes the solutions practically useless. Even assuming the instantaneous training of a classifier, there should be a clinician that would say, hey here is a 10min chunk I believe it is preictal, quickly retrain your classifier and tell me that the remaining 5 chunks are also preictal as if I don't know it yet. You can use previous preictal events of the same patients to predict future one but you can't in real-life have examples of would-be preictal events. Whether using the test data to retrain the classifier helps or not I still have to figure out as my account is blocked for suspected cheating :) I have to see which of the two models submitted (using testing data and without) brought me to the 3rd public and 5th private.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58397",
      "postDate": "11/19/2014 12:54:04",
      "content": "<p>I hope this can help&nbsp;https://www.kaggle.com/c/seizure-prediction/forums/t/10945/congratulations-to-the-winners/58396#post58396</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58513",
      "postDate": "11/20/2014 11:49:16",
      "content": "<p>[quote=William Cukierski;58276]</p>\n<p>[quote=zzspar;58256]</p>\n<p>Kaggle &#8211; in the future, in similar competitions, can you cut out a small random gap between segments? Even taking out few samples will help the competition a lot.</p>\n<p>I this competition I noticed that you can undo the random shuffling of test segments and restore the original sequences. Knowing the sequence of segments in the test data can be used to improve the score but it will not help much the real problem we are trying to solve.</p>\n<p>[/quote]</p>\n<p>The host did trim the ends of the clips and mean centered them. How do you know you were able to reverse engineer the order?</p>\n<p>[/quote]</p>\n<p>Hi William, I add a link to a viewer to my notebook&nbsp;</p>\n<p>http://nbviewer.ipython.org/github/udibr/seizure-detection-boost/blob/master/postprocessing.ipynb</p>\n<p>so you can just browse the code, without running it, and see that it is possible to get &quot;trimming&quot; information from the test data.</p>\n<p>Hope this help</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "216340",
      "postDate": "08/25/2017 09:53:02",
      "content": "<p>Hello everyone, I work some project on my faculty based on this competition. I need more data for my algorithm. Can I get annotated test data for Patient_1 and Patient_2 ? That would help me with my further work Thanks in advance, Nikola</p>",
      "rawMarkdown": "Hello everyone, I work some project on my faculty based on this competition. I need more data for my algorithm. Can I get annotated test data for Patient_1 and Patient_2 ? That would help me with my further work Thanks in advance, Nikola",
      "votes": null
    },
    {
      "id": "617460",
      "postDate": "09/04/2019 06:43:52",
      "content": "<p>hello,\nI am also working on my project on seizure prediction can you please help me to get the labeled test data? I have no time so please if you can help me. my mail xuwang_hbu@163.com</p>",
      "rawMarkdown": "hello,\nI am also working on my project on seizure prediction can you please help me to get the labeled test data? I have no time so please if you can help me. my mail xuwang_hbu@163.com",
      "votes": null
    },
    {
      "id": "730455",
      "postDate": "01/27/2020 14:03:47",
      "content": "<p>Hi , I would like to know if you succeeded to get more data? I would like to do the same. I am using the dataset for a current study but Patient1 and 2 only have 3 labeled seizures and it seems like my algorithm needs more training data. </p>",
      "rawMarkdown": "Hi , I would like to know if you succeeded to get more data? I would like to do the same. I am using the dataset for a current study but Patient1 and 2 only have 3 labeled seizures and it seems like my algorithm needs more training data.",
      "votes": null
    },
    {
      "id": "786710",
      "postDate": "03/26/2020 06:10:21",
      "content": "<p>I was wondering if anyone could please help me by answering my questions below:</p>\n\n<p>So for patients 1 and 2, I know that in every 6th data segment there is a seizure onset (because each segment is 10min and we have an hour prior to each seizure). This means we only have 3 seizures since we only have 18 pre-ictal segments( seizures occur in the middle of 6th 12th and 18th pre-ictal segments). \nAccording to the paragraph above, did I understand the data properly?</p>\n\n<p>My goal is to make a continuous signal with interictal and then preictal and then ictal ...again interictal then preictal and ictal and so on....in order to train my data. My questions are the following:</p>\n\n<p>1- does Patient_1_interictal_segment_0001.mat occur in time(in the real world) right before Patient_1_prerictal_segment_0001.mat</p>\n\n<p>2-When do the seizures occur in each of the test-segments? (Do you have a labelled test data?, if not I can label </p>\n\n<p>the test data myself if you tell me for example the seizures occur in the middle of test segments 5 13 or whatever...)</p>\n\n<p>Please help me with this I would be grateful,</p>\n\n<p>Thank you.</p>",
      "rawMarkdown": "I was wondering if anyone could please help me by answering my questions below:\n\nSo for patients 1 and 2, I know that in every 6th data segment there is a seizure onset (because each segment is 10min and we have an hour prior to each seizure). This means we only have 3 seizures since we only have 18 pre-ictal segments( seizures occur in the middle of 6th 12th and 18th pre-ictal segments). \nAccording to the paragraph above, did I understand the data properly?\n\nMy goal is to make a continuous signal with interictal and then preictal and then ictal ...again interictal then preictal and ictal and so on....in order to train my data. My questions are the following:\n\n1- does Patient_1_interictal_segment_0001.mat occur in time(in the real world) right before Patient_1_prerictal_segment_0001.mat\n\n2-When do the seizures occur in each of the test-segments? (Do you have a labelled test data?, if not I can label \n\nthe test data myself if you tell me for example the seizures occur in the middle of test segments 5 13 or whatever...)\n\nPlease help me with this I would be grateful,\n\nThank you.",
      "votes": null
    },
    {
      "id": "786711",
      "postDate": "03/26/2020 06:12:19",
      "content": "<p>I was wondering if anyone could please help me by answering my questions below:</p>\n\n<p>So for patients 1 and 2, I know that in every 6th data segment there is a seizure onset (because each segment is 10min and we have an hour prior to each seizure). This means we only have 3 seizures since we only have 18 pre-ictal segments( seizures occur in the middle of 6th 12th and 18th pre-ictal segments). \nAccording to the paragraph above, did I understand the data properly?</p>\n\n<p>My goal is to make a continuous signal with interictal and then preictal and then ictal ...again interictal then preictal and ictal and so on....in order to train my data. My questions are the following:</p>\n\n<p>1- does Patient_1_interictal_segment_0001.mat occur in time(in the real world) right before Patient_1_prerictal_segment_0001.mat</p>\n\n<p>2-When do the seizures occur in each of the test-segments? (Do you have a labelled test data?, if not I can label </p>\n\n<p>the test data myself if you tell me for example the seizures occur in the middle of test segments 5 13 or whatever...)</p>\n\n<p>Please help me with this I would be grateful,</p>\n\n<p>Thank you,</p>",
      "rawMarkdown": "I was wondering if anyone could please help me by answering my questions below:\n\nSo for patients 1 and 2, I know that in every 6th data segment there is a seizure onset (because each segment is 10min and we have an hour prior to each seizure). This means we only have 3 seizures since we only have 18 pre-ictal segments( seizures occur in the middle of 6th 12th and 18th pre-ictal segments). \nAccording to the paragraph above, did I understand the data properly?\n\nMy goal is to make a continuous signal with interictal and then preictal and then ictal ...again interictal then preictal and ictal and so on....in order to train my data. My questions are the following:\n\n1- does Patient_1_interictal_segment_0001.mat occur in time(in the real world) right before Patient_1_prerictal_segment_0001.mat\n\n2-When do the seizures occur in each of the test-segments? (Do you have a labelled test data?, if not I can label \n\nthe test data myself if you tell me for example the seizures occur in the middle of test segments 5 13 or whatever...)\n\nPlease help me with this I would be grateful,\n\nThank you,",
      "votes": null
    },
    {
      "id": "802588",
      "postDate": "04/09/2020 16:32:03",
      "content": "<p>Thank you for your reply. I just have two more questions:</p>\n\n<p>1- In the test data, the segments where the seizure occurs, say for example Patient1_testsegment_0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)</p>\n\n<p>2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?</p>\n\n<p>Again Thank you for your reply.</p>",
      "rawMarkdown": "Thank you for your reply. I just have two more questions:\n\n1- In the test data, the segments where the seizure occurs, say for example Patient1_testsegment_0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)\n\n2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?\n\nAgain Thank you for your reply.",
      "votes": null
    },
    {
      "id": "802589",
      "postDate": "04/09/2020 16:33:25",
      "content": "<p>Thank you for your reply. I just have two more questions:</p>\n\n<p>1- In the test data, the segments where the seizure occurs; say for example Patient1testsegment0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)</p>\n\n<p>2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?</p>\n\n<p>Again Thank you for your reply.</p>",
      "rawMarkdown": "Thank you for your reply. I just have two more questions:\n\n1- In the test data, the segments where the seizure occurs; say for example Patient1testsegment0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)\n\n2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?\n\nAgain Thank you for your reply.",
      "votes": null
    },
    {
      "id": "818656",
      "postDate": "04/24/2020 03:11:41",
      "content": "<p>Thank you for your reply. I just have two more questions:</p>\n\n<p>1- In the test data, the segments where the seizure occurs, say for example Patient1testsegment0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)</p>\n\n<p>2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?</p>\n\n<p>Again Thank you for your reply.</p>",
      "rawMarkdown": "Thank you for your reply. I just have two more questions:\n\n1- In the test data, the segments where the seizure occurs, say for example Patient1testsegment0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)\n\n2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?\n\nAgain Thank you for your reply.",
      "votes": null
    },
    {
      "id": "1774358",
      "postDate": "05/02/2022 05:54:33",
      "content": "<p>Hello, <br>\nI am an ML teacher, I would like to use this challenge in my class, would it be possible to get the test labels ? <br>\nWhat about Patient_2, Dog_1 and Dog_5 (they is no data right now) ? <br>\nThank you very much ! <br>\nEmmanuel</p>",
      "rawMarkdown": "Hello, \nI am an ML teacher, I would like to use this challenge in my class, would it be possible to get the test labels ? \nWhat about Patient_2, Dog_1 and Dog_5 (they is no data right now) ? \nThank you very much ! \nEmmanuel",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1774358,
      "author_name": "emmanuelbacry",
      "author_url": "",
      "post_date": "05/02/2022 05:54:33",
      "content": "<p>Hello, <br>\nI am an ML teacher, I would like to use this challenge in my class, would it be possible to get the test labels ? <br>\nWhat about Patient_2, Dog_1 and Dog_5 (they is no data right now) ? <br>\nThank you very much ! <br>\nEmmanuel</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57011,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "10/29/2014 19:21:35",
      "content": "<p>will there still be a blind test set for final results?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57013,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "10/29/2014 19:36:59",
      "content": "<p>The test set you have is the test set used to score the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57014,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "10/29/2014 19:38:37",
      "content": "<p>Right, but is this statement still true &quot;The final results will be based on the other 60%, so the final standings may be different.&quot;</p>\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57015,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "10/29/2014 19:42:43",
      "content": "<p>Yes</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57022,
      "author_name": "mdemenech",
      "author_url": "",
      "post_date": "10/29/2014 21:51:12",
      "content": "<p>Since the procedure of providing the full test set (keeping unknown the public/private identification) to the competitors is the standard one in Kaggle competitions, I would guess the same issue of using test data for improving the predictions on the test set itself should have risen before. What was the rule then?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57025,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "10/29/2014 22:06:50",
      "content": "<p>We have historically allowed semi-supervised learning on the grounds that it is very hard to enforce a rule against it. However, it's a decision that is still made on a case-by-case basis and can be&nbsp;more detracting for some problems than others. e.g. a machine learning&nbsp;problem&nbsp;on a text corpus is a very different beast&nbsp;than a time-series forecast.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57107,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "10/31/2014 13:44:29",
      "content": "<p>Sorry to ask a stupid question here, but how does one use unlabeled test data to calibrate a model? Are there any good resources we might read for this?</p>\n<p>Both isotonic regression and platt scaling (discussed <a href=\"http://www.kaggle.com/c/seizure-prediction/forums/t/10542/good-cross-validation-poor-test-set-result\">here</a>) seem to require mapping model output (e.g., probabilities) to a known class label. Using the same data for training and calibration clearly introduces bias. Thus it seems we're stuck with withholding part of the training set. In this case, though, there is such a small number of positive (i.e., 'preictal') instances, this seem like it would be unwise.</p>\n<p>One obvious answer, then, seems to be to use out-of-bag data to calibrate in each bagging iteration. Still, this begs the question, where does (or could) the unlabeled test data come into play?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57851,
      "author_name": "bgeier",
      "author_url": "",
      "post_date": "11/11/2014 21:58:29",
      "content": "<p>William,</p>\n<p>Is the 40/60 test split completely at random? So that seizure rates within an individual remain the same between 40% LB and 60% blind sets</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57858,
      "author_name": "thenx00",
      "author_url": "",
      "post_date": "11/11/2014 22:31:16",
      "content": "<p>[quote=Maineiac;57107]</p>\n<p>Sorry to ask a stupid question here, but how does one use unlabeled test data to calibrate a model? Are there any good resources we might read for this?</p>\n<p>One obvious answer, then, seems to be to use out-of-bag data to calibrate in each bagging iteration. Still, this begs the question, where does (or could) the unlabeled test data come into play?</p>\n<p>[/quote]</p>\n<p>In brief, in some computer vision related tasks the pipeline might be the following:</p>\n<p>1)&nbsp;Extract some patches or descriptors from images</p>\n<p>2) Cluster them to create vocabulary</p>\n<p>3) Then&nbsp;represent images as vectors of vocabulary entries</p>\n<p>4) Train smth on this vectors</p>\n<p>This is called&nbsp;<a href=\"http://en.wikipedia.org/wiki/Bag-of-words_model_in_computer_vision\" target=\"_blank\">Bag of&nbsp;Visual Words</a>. Since clustering is unsupervised, you can create vocabulary using unlabeled data (for example test data), which might improve overall performance.</p>\n\n<p>P.S. That's not an example of model calibration but test data might help sometimes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58166,
      "author_name": "mahi83",
      "author_url": "",
      "post_date": "11/17/2014 00:39:24",
      "content": "<p>Since today is the last day, can we increase the number of submission by 10 more? I have a lot of models I want to check.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58237,
      "author_name": "tjames",
      "author_url": "",
      "post_date": "11/17/2014 23:45:33",
      "content": "<p>Hi if you are able to extend the deadline at least until sometime later today that would be great.</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58238,
      "author_name": "tjames",
      "author_url": "",
      "post_date": "11/17/2014 23:54:20",
      "content": "<p>I had to find some control data (i.e.- no seizures)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58256,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "11/18/2014 02:15:18",
      "content": "<p>Kaggle &#8211; in the future, in similar competitions, can you cut out a small random gap between segments? Even taking out few samples will help the competition a lot.</p>\n<p>I this competition I noticed that you can undo the random shuffling of test segments and restore the original sequences. Knowing the sequence of segments in the test data can be used to improve the score but it will not help much the real problem we are trying to solve.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58276,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "11/18/2014 04:44:25",
      "content": "<p>[quote=zzspar;58256]</p>\n<p>Kaggle &#8211; in the future, in similar competitions, can you cut out a small random gap between segments? Even taking out few samples will help the competition a lot.</p>\n<p>I this competition I noticed that you can undo the random shuffling of test segments and restore the original sequences. Knowing the sequence of segments in the test data can be used to improve the score but it will not help much the real problem we are trying to solve.</p>\n<p>[/quote]</p>\n<p>The host did trim the ends of the clips and mean centered them. How do you know you were able to reverse engineer the order?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58295,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "11/18/2014 10:17:21",
      "content": "<p>I tried it out and it helpd the public LB...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58304,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "11/18/2014 11:42:27",
      "content": "<p>Now that the competition is over, can someone (Will perhaps?) explain <strong>how we can use unlabeled test data to calibrate our model predictions?</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58308,
      "author_name": "golondrina",
      "author_url": "",
      "post_date": "11/18/2014 12:06:00",
      "content": "<p>I used min-max scaling of &nbsp;test probabilities for each subject separately. This gave an improvement</p>\n<p>of ~0.015 for my best model on public LB. &nbsp;I think this trick worked because my predicted probabilities were in very different ranges for each of the subjects, but&nbsp;I am not sure if it's helpful for other models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58329,
      "author_name": "andtem2000",
      "author_url": "",
      "post_date": "11/18/2014 18:32:22",
      "content": "<p>[quote=Maineiac;58304]</p>\n<p>Now that the competition is over, can someone (Will perhaps?) explain <strong>how we can use unlabeled test data to calibrate our model predictions?</strong></p>\n<p>[/quote]</p>\n<p>Imagine you do CV to select a model. You can do a random CV (lets name it CV1), you can do random preserving chunk integrity if you have more than a single frame from a 10min chunk (CV2) or you can do random preserving an event integrity (CV3) that is all chunks from an event are either used for test or for train. Your CV performance will increase from CV3 to CV1, say 75-85-95, first by having examples of a particular preictal event, second by having examples of a particular event chunk. The real-life testing stage follows&nbsp;the CV3 condition. However, if you create your own labels from the test data and then incorporate this data to your classifier then you can approximate CV2 condition with a corresponding increase in the performance. This is given that the organisers used a CV1 40-60 split between public and private. If they used CV3 split, then using testing data could be a trap. Personally, I disagree with allowing to use test data in any form or sense. It makes the solutions practically useless. Even assuming the instantaneous training of a classifier, there should be a clinician that would say, hey here is a 10min chunk I believe it is preictal, quickly retrain your classifier and tell me that the remaining 5 chunks are also preictal as if I don't know it yet. You can use previous preictal events of the same patients to predict future one but you can't in real-life have examples of would-be preictal events. Whether using the test data to retrain the classifier helps or not I still have to figure out as my account is blocked for suspected cheating :) I have to see which of the two models submitted (using testing data and without) brought me to the 3rd public and 5th private.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58397,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "11/19/2014 12:54:04",
      "content": "<p>I hope this can help&nbsp;https://www.kaggle.com/c/seizure-prediction/forums/t/10945/congratulations-to-the-winners/58396#post58396</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58513,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "11/20/2014 11:49:16",
      "content": "<p>[quote=William Cukierski;58276]</p>\n<p>[quote=zzspar;58256]</p>\n<p>Kaggle &#8211; in the future, in similar competitions, can you cut out a small random gap between segments? Even taking out few samples will help the competition a lot.</p>\n<p>I this competition I noticed that you can undo the random shuffling of test segments and restore the original sequences. Knowing the sequence of segments in the test data can be used to improve the score but it will not help much the real problem we are trying to solve.</p>\n<p>[/quote]</p>\n<p>The host did trim the ends of the clips and mean centered them. How do you know you were able to reverse engineer the order?</p>\n<p>[/quote]</p>\n<p>Hi William, I add a link to a viewer to my notebook&nbsp;</p>\n<p>http://nbviewer.ipython.org/github/udibr/seizure-detection-boost/blob/master/postprocessing.ipynb</p>\n<p>so you can just browse the code, without running it, and see that it is possible to get &quot;trimming&quot; information from the test data.</p>\n<p>Hope this help</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 216340,
      "author_name": "nikola993",
      "author_url": "",
      "post_date": "08/25/2017 09:53:02",
      "content": "<p>Hello everyone, I work some project on my faculty based on this competition. I need more data for my algorithm. Can I get annotated test data for Patient_1 and Patient_2 ? That would help me with my further work Thanks in advance, Nikola</p>",
      "votes": null,
      "replies": [
        {
          "id": 730455,
          "author_name": "rmybenmessaoud",
          "author_url": "",
          "post_date": "01/27/2020 14:03:47",
          "content": "<p>Hi , I would like to know if you succeeded to get more data? I would like to do the same. I am using the dataset for a current study but Patient1 and 2 only have 3 labeled seizures and it seems like my algorithm needs more training data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 617460,
      "author_name": "mengyun1993",
      "author_url": "",
      "post_date": "09/04/2019 06:43:52",
      "content": "<p>hello,\nI am also working on my project on seizure prediction can you please help me to get the labeled test data? I have no time so please if you can help me. my mail xuwang_hbu@163.com</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 786710,
      "author_name": "sepehrrashidi",
      "author_url": "",
      "post_date": "03/26/2020 06:10:21",
      "content": "<p>I was wondering if anyone could please help me by answering my questions below:</p>\n\n<p>So for patients 1 and 2, I know that in every 6th data segment there is a seizure onset (because each segment is 10min and we have an hour prior to each seizure). This means we only have 3 seizures since we only have 18 pre-ictal segments( seizures occur in the middle of 6th 12th and 18th pre-ictal segments). \nAccording to the paragraph above, did I understand the data properly?</p>\n\n<p>My goal is to make a continuous signal with interictal and then preictal and then ictal ...again interictal then preictal and ictal and so on....in order to train my data. My questions are the following:</p>\n\n<p>1- does Patient_1_interictal_segment_0001.mat occur in time(in the real world) right before Patient_1_prerictal_segment_0001.mat</p>\n\n<p>2-When do the seizures occur in each of the test-segments? (Do you have a labelled test data?, if not I can label </p>\n\n<p>the test data myself if you tell me for example the seizures occur in the middle of test segments 5 13 or whatever...)</p>\n\n<p>Please help me with this I would be grateful,</p>\n\n<p>Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 786711,
      "author_name": "sepehrrashidi",
      "author_url": "",
      "post_date": "03/26/2020 06:12:19",
      "content": "<p>I was wondering if anyone could please help me by answering my questions below:</p>\n\n<p>So for patients 1 and 2, I know that in every 6th data segment there is a seizure onset (because each segment is 10min and we have an hour prior to each seizure). This means we only have 3 seizures since we only have 18 pre-ictal segments( seizures occur in the middle of 6th 12th and 18th pre-ictal segments). \nAccording to the paragraph above, did I understand the data properly?</p>\n\n<p>My goal is to make a continuous signal with interictal and then preictal and then ictal ...again interictal then preictal and ictal and so on....in order to train my data. My questions are the following:</p>\n\n<p>1- does Patient_1_interictal_segment_0001.mat occur in time(in the real world) right before Patient_1_prerictal_segment_0001.mat</p>\n\n<p>2-When do the seizures occur in each of the test-segments? (Do you have a labelled test data?, if not I can label </p>\n\n<p>the test data myself if you tell me for example the seizures occur in the middle of test segments 5 13 or whatever...)</p>\n\n<p>Please help me with this I would be grateful,</p>\n\n<p>Thank you,</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 802588,
      "author_name": "sepehrrashidi",
      "author_url": "",
      "post_date": "04/09/2020 16:32:03",
      "content": "<p>Thank you for your reply. I just have two more questions:</p>\n\n<p>1- In the test data, the segments where the seizure occurs, say for example Patient1_testsegment_0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)</p>\n\n<p>2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?</p>\n\n<p>Again Thank you for your reply.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 802589,
      "author_name": "sepehrrashidi",
      "author_url": "",
      "post_date": "04/09/2020 16:33:25",
      "content": "<p>Thank you for your reply. I just have two more questions:</p>\n\n<p>1- In the test data, the segments where the seizure occurs; say for example Patient1testsegment0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)</p>\n\n<p>2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?</p>\n\n<p>Again Thank you for your reply.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 818656,
      "author_name": "sepehrrashidi",
      "author_url": "",
      "post_date": "04/24/2020 03:11:41",
      "content": "<p>Thank you for your reply. I just have two more questions:</p>\n\n<p>1- In the test data, the segments where the seizure occurs, say for example Patient1testsegment0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)</p>\n\n<p>2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?</p>\n\n<p>Again Thank you for your reply.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57006": "",
    "57011": "",
    "57013": "",
    "57014": "",
    "57015": "",
    "57022": "",
    "57025": "",
    "57107": "",
    "57851": "",
    "57858": "",
    "58166": "",
    "58237": "",
    "58238": "",
    "58256": "",
    "58276": "",
    "58295": "",
    "58304": "",
    "58308": "",
    "58329": "",
    "58397": "",
    "58513": "",
    "216340": "Hello everyone, I work some project on my faculty based on this competition. I need more data for my algorithm. Can I get annotated test data for Patient_1 and Patient_2 ? That would help me with my further work Thanks in advance, Nikola",
    "617460": "hello,\nI am also working on my project on seizure prediction can you please help me to get the labeled test data? I have no time so please if you can help me. my mail xuwang_hbu@163.com",
    "730455": "Hi , I would like to know if you succeeded to get more data? I would like to do the same. I am using the dataset for a current study but Patient1 and 2 only have 3 labeled seizures and it seems like my algorithm needs more training data.",
    "786710": "I was wondering if anyone could please help me by answering my questions below:\n\nSo for patients 1 and 2, I know that in every 6th data segment there is a seizure onset (because each segment is 10min and we have an hour prior to each seizure). This means we only have 3 seizures since we only have 18 pre-ictal segments( seizures occur in the middle of 6th 12th and 18th pre-ictal segments). \nAccording to the paragraph above, did I understand the data properly?\n\nMy goal is to make a continuous signal with interictal and then preictal and then ictal ...again interictal then preictal and ictal and so on....in order to train my data. My questions are the following:\n\n1- does Patient_1_interictal_segment_0001.mat occur in time(in the real world) right before Patient_1_prerictal_segment_0001.mat\n\n2-When do the seizures occur in each of the test-segments? (Do you have a labelled test data?, if not I can label \n\nthe test data myself if you tell me for example the seizures occur in the middle of test segments 5 13 or whatever...)\n\nPlease help me with this I would be grateful,\n\nThank you.",
    "786711": "I was wondering if anyone could please help me by answering my questions below:\n\nSo for patients 1 and 2, I know that in every 6th data segment there is a seizure onset (because each segment is 10min and we have an hour prior to each seizure). This means we only have 3 seizures since we only have 18 pre-ictal segments( seizures occur in the middle of 6th 12th and 18th pre-ictal segments). \nAccording to the paragraph above, did I understand the data properly?\n\nMy goal is to make a continuous signal with interictal and then preictal and then ictal ...again interictal then preictal and ictal and so on....in order to train my data. My questions are the following:\n\n1- does Patient_1_interictal_segment_0001.mat occur in time(in the real world) right before Patient_1_prerictal_segment_0001.mat\n\n2-When do the seizures occur in each of the test-segments? (Do you have a labelled test data?, if not I can label \n\nthe test data myself if you tell me for example the seizures occur in the middle of test segments 5 13 or whatever...)\n\nPlease help me with this I would be grateful,\n\nThank you,",
    "802588": "Thank you for your reply. I just have two more questions:\n\n1- In the test data, the segments where the seizure occurs, say for example Patient1_testsegment_0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)\n\n2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?\n\nAgain Thank you for your reply.",
    "802589": "Thank you for your reply. I just have two more questions:\n\n1- In the test data, the segments where the seizure occurs; say for example Patient1testsegment0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)\n\n2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?\n\nAgain Thank you for your reply.",
    "818656": "Thank you for your reply. I just have two more questions:\n\n1- In the test data, the segments where the seizure occurs, say for example Patient1testsegment0007.mat, that segment is also 10mins. does the seizure occur in the middle of that segment? (If so how long does it last; what is the duration of the seizure?)\n\n2- Is this statement correct?:\nThe competition does not care about when the seizure occurs within each test segment. The test segments are provided and we have to predict which ones have a seizure in them( just like the answer key with 0 and 1 labels) Am I correct?\n\nAgain Thank you for your reply.",
    "1774358": "Hello, \nI am an ML teacher, I would like to use this challenge in my class, would it be possible to get the test labels ? \nWhat about Patient_2, Dog_1 and Dog_5 (they is no data right now) ? \nThank you very much ! \nEmmanuel"
  },
  "source": "meta"
}