{
  "id": 10909,
  "title": "Increase number of samples by using short time segments",
  "url": "/competitions/seizure-prediction/discussion/10909",
  "author_name": "",
  "post_date": "2014-11-12T09:22:31.677Z",
  "votes": 2,
  "comment_count": 9,
  "views": 2109,
  "content": "<p>My general approach to the problem is based on extracting various features on short data clips, i.e. I split the 10 min segments into 200 separate 3s long samples. Each of these is then fed into a classifier (tried only random forest so far). So now I have 200 times more samples (both for training and predicting) and this should clearly help. However, I also get 200 different probability predictions for each test segment and I am not sure how to best proceed from there on. The simplest approach would be to take the average&nbsp;or median or some fixed percentile of these 200 different probabilities. A more sophisticated approach would be to feed these probabilities (or some characteristics of the probability distribution such as max, min, std, mean etc)&nbsp; into a new classifier.</p>\n<p>To my surprise, I get the best results from simply taking a fixed percentile of the probabilities and the re-training on probabilities does not really work too well.</p>\n<p>Yet a different strategy is to increase not the number of samples but the numbers of features. So I keep only one sample from each 10 min segment but take e.g. from each feature the max, min, std and mean from the 200 different 3s clips. So, in this example, I would get 4 features from every feature I calculate on the short clips. But I dont get good results from that either. The best strategy I found so far is to increase the number of samples and use a simple percentile of the probability distribution of each segment to get a probability for each segment.</p>\n<p>I wonder if anybody else here is trying something similar? I am sure a lot of people are looking at short time segments, does anybody have a better suggestion how to handle the short-segment information?</p>",
  "messages": [
    {
      "id": "57885",
      "postDate": "11/12/2014 09:22:31",
      "content": "<p>My general approach to the problem is based on extracting various features on short data clips, i.e. I split the 10 min segments into 200 separate 3s long samples. Each of these is then fed into a classifier (tried only random forest so far). So now I have 200 times more samples (both for training and predicting) and this should clearly help. However, I also get 200 different probability predictions for each test segment and I am not sure how to best proceed from there on. The simplest approach would be to take the average&nbsp;or median or some fixed percentile of these 200 different probabilities. A more sophisticated approach would be to feed these probabilities (or some characteristics of the probability distribution such as max, min, std, mean etc)&nbsp; into a new classifier.</p>\n<p>To my surprise, I get the best results from simply taking a fixed percentile of the probabilities and the re-training on probabilities does not really work too well.</p>\n<p>Yet a different strategy is to increase not the number of samples but the numbers of features. So I keep only one sample from each 10 min segment but take e.g. from each feature the max, min, std and mean from the 200 different 3s clips. So, in this example, I would get 4 features from every feature I calculate on the short clips. But I dont get good results from that either. The best strategy I found so far is to increase the number of samples and use a simple percentile of the probability distribution of each segment to get a probability for each segment.</p>\n<p>I wonder if anybody else here is trying something similar? I am sure a lot of people are looking at short time segments, does anybody have a better suggestion how to handle the short-segment information?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57888",
      "postDate": "11/12/2014 09:42:21",
      "content": "<p>I tried dividing the 10 min clips into 15 sec chunks, deriving time correlations for each chunk and concatinating all of them as features vector but it didn't yield any benefit over using the entire 10min for feature extraction as a single chunk. I hope there must be people who have tried something similar.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57890",
      "postDate": "11/12/2014 09:43:31",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57896",
      "postDate": "11/12/2014 12:05:39",
      "content": "<p>Thanks, gaurav. But actually I would not expect simply concatinating the features from different time slices to work well. This would force the classifier to compare the features from the nth slice of a given segment to only those of the nth slice of another segment. But there is no reason why it should not also be compared against all the other time slices from other segments.</p>\n<p>So I still hope there is some value in partitioning the segments, did anybody out there get good results from partitioning?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57917",
      "postDate": "11/12/2014 15:29:09",
      "content": "<p>We are traversing data using windows of 60s length, and using one window to classify as precital/interictal. The length of the windows is inspired by the reference which is in the problem description:</p>\n<p><em>Howbert JJ, Patterson EE, Stead SM, Brinkmann B, Vasoli V, Crepeau D, Vite CH, Sturges B, Ruedebusch V, Mavoori J, Leyde K, Sheffield WD, Litt B, Worrell GA (2014) Forecasting seizures in dogs with naturally occurring epilepsy. PLoS One 9(1):e81920.</em></p>\n<p>In this way, you only need to combine a few probabilities, and using the mean, maximum, percentil, etc. you can obtain similar results.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57961",
      "postDate": "11/13/2014 01:28:27",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57972",
      "postDate": "11/13/2014 08:09:40",
      "content": "<p>Morpheus, I&nbsp;will sugest you to try with a different dog because dog 1 is with difference the worst.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58009",
      "postDate": "11/13/2014 16:49:17",
      "content": "<p>Thanks Francisco. I tried with dog 2 and dog 5 as well. For dog 2 also I did not find any difference in power spectra. I did detect a difference for dog 5 between interictal vs preictal. Btw, I deleted my old post because for a moment I was wary if it violated any rules - I hadn't seen people posting any code.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58027",
      "postDate": "11/13/2014 20:31:29",
      "content": "<p>It is legal to write code in the forum. Unfortunatelly I don't used to write fft part in R, but in my case, fft filter captures different information in every bin.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58239",
      "postDate": "11/18/2014 00:09:22",
      "content": "<p>i was going by the &quot;dprime&quot; number and kept thinking my results are no good but when i actually computed the AUC using a varying threshold today I got a score of 0.6559 for dog 1 using logistic regression and a feature vector consisting of power in the 6 bands averaged over the channels. This was a lesson to me since the Dprime numbers are extremely low for the 6 bands and I kept thinking any feature with D less than 1 will be useless.</p>\n<p>0.129016855 0.079619722 0.063978125 0.074278229 0.043733561 0.068145816</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 57888,
      "author_name": "scatterbrain333",
      "author_url": "",
      "post_date": "11/12/2014 09:42:21",
      "content": "<p>I tried dividing the 10 min clips into 15 sec chunks, deriving time correlations for each chunk and concatinating all of them as features vector but it didn't yield any benefit over using the entire 10min for feature extraction as a single chunk. I hope there must be people who have tried something similar.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57890,
      "author_name": "scatterbrain333",
      "author_url": "",
      "post_date": "11/12/2014 09:43:31",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 57896,
      "author_name": "hasselmann",
      "author_url": "",
      "post_date": "11/12/2014 12:05:39",
      "content": "<p>Thanks, gaurav. But actually I would not expect simply concatinating the features from different time slices to work well. This would force the classifier to compare the features from the nth slice of a given segment to only those of the nth slice of another segment. But there is no reason why it should not also be compared against all the other time slices from other segments.</p>\n<p>So I still hope there is some value in partitioning the segments, did anybody out there get good results from partitioning?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57917,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/12/2014 15:29:09",
      "content": "<p>We are traversing data using windows of 60s length, and using one window to classify as precital/interictal. The length of the windows is inspired by the reference which is in the problem description:</p>\n<p><em>Howbert JJ, Patterson EE, Stead SM, Brinkmann B, Vasoli V, Crepeau D, Vite CH, Sturges B, Ruedebusch V, Mavoori J, Leyde K, Sheffield WD, Litt B, Worrell GA (2014) Forecasting seizures in dogs with naturally occurring epilepsy. PLoS One 9(1):e81920.</em></p>\n<p>In this way, you only need to combine a few probabilities, and using the mean, maximum, percentil, etc. you can obtain similar results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57961,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "11/13/2014 01:28:27",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 57972,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/13/2014 08:09:40",
      "content": "<p>Morpheus, I&nbsp;will sugest you to try with a different dog because dog 1 is with difference the worst.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58009,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "11/13/2014 16:49:17",
      "content": "<p>Thanks Francisco. I tried with dog 2 and dog 5 as well. For dog 2 also I did not find any difference in power spectra. I did detect a difference for dog 5 between interictal vs preictal. Btw, I deleted my old post because for a moment I was wary if it violated any rules - I hadn't seen people posting any code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58027,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "11/13/2014 20:31:29",
      "content": "<p>It is legal to write code in the forum. Unfortunatelly I don't used to write fft part in R, but in my case, fft filter captures different information in every bin.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58239,
      "author_name": "siddjain",
      "author_url": "",
      "post_date": "11/18/2014 00:09:22",
      "content": "<p>i was going by the &quot;dprime&quot; number and kept thinking my results are no good but when i actually computed the AUC using a varying threshold today I got a score of 0.6559 for dog 1 using logistic regression and a feature vector consisting of power in the 6 bands averaged over the channels. This was a lesson to me since the Dprime numbers are extremely low for the 6 bands and I kept thinking any feature with D less than 1 will be useless.</p>\n<p>0.129016855 0.079619722 0.063978125 0.074278229 0.043733561 0.068145816</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57885": "",
    "57888": "",
    "57890": "",
    "57896": "",
    "57917": "",
    "57961": "",
    "57972": "",
    "58009": "",
    "58027": "",
    "58239": ""
  },
  "source": "meta"
}