{
  "id": 10904,
  "title": "Feature vector size",
  "url": "/competitions/seizure-prediction/discussion/10904",
  "author_name": "",
  "post_date": "2014-11-11T22:44:19.017Z",
  "votes": null,
  "comment_count": 3,
  "views": 1087,
  "content": "<p>Hello Everyone,</p>\n<p>I got into the competition late. But nonetheless, the task as I understand it is to predict/classify if each 10 minute segment is interictal or preictal. I extracted some features based on the fft and correlation of fft coefficients between channels. Even doing an analysis on 10 seconds chunks yields a really large dimensionality on the input data feature per 10 min segment. If one throws this into a logistic regression classifier, I find that the model (tested only on a single subject e.g. Dog4) horribly overfits to the training data given the super-large number of parameters. Is anyone else seeing similar behaviour?&nbsp;</p>\n<p>Also neural networks would suffer from the sheer size of the input layer. Is L2/dropout helping anyone who tried that?</p>\n<p>Also has anyone looked at any dimensionality reduction algorithms?</p>",
  "messages": [
    {
      "id": "57859",
      "postDate": "11/11/2014 22:44:19",
      "content": "<p>Hello Everyone,</p>\n<p>I got into the competition late. But nonetheless, the task as I understand it is to predict/classify if each 10 minute segment is interictal or preictal. I extracted some features based on the fft and correlation of fft coefficients between channels. Even doing an analysis on 10 seconds chunks yields a really large dimensionality on the input data feature per 10 min segment. If one throws this into a logistic regression classifier, I find that the model (tested only on a single subject e.g. Dog4) horribly overfits to the training data given the super-large number of parameters. Is anyone else seeing similar behaviour?&nbsp;</p>\n<p>Also neural networks would suffer from the sheer size of the input layer. Is L2/dropout helping anyone who tried that?</p>\n<p>Also has anyone looked at any dimensionality reduction algorithms?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57966",
      "postDate": "11/13/2014 04:23:12",
      "content": "<p>I am having trouble understanding how to convert a time series into features, something perhaps more basic than your post. I am grappling with the following question.&nbsp;Does each feature (ie correlation, FFT, etc) have a single column for&nbsp;the entire 10 minute clip, or is there a group&nbsp;of features for a smaller time interval, say 10 seconds, that gets repeated as a group?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57967",
      "postDate": "11/13/2014 04:41:27",
      "content": "<p>Hi Hugh. From what I understood from the literature I have read -- it is the latter, i.e. there is a group of features for a smaller time interval that gets repeated over. The length of the smaller time interval depends on the paper you read. Some chose a smaller time window of 5s, some 10s and some 60s. In the <a href=\"http://www.plosone.org/article/fetchObject.action?uri=info:doi%2F10.1371%2Fjournal.pone.0081920&representation=PDF\">paper </a>that accompanies this competition does it in chunks of 60s.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "58023",
      "postDate": "11/13/2014 19:35:31",
      "content": "<p>I've also seen some papers that use&nbsp;a moving, overlapping window. In other words, if you have 100 points, you break that into 0-20, 10-30, 20-40, etc. and calculate features on each of the epochs.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 57966,
      "author_name": "hughkroghfreeman",
      "author_url": "",
      "post_date": "11/13/2014 04:23:12",
      "content": "<p>I am having trouble understanding how to convert a time series into features, something perhaps more basic than your post. I am grappling with the following question.&nbsp;Does each feature (ie correlation, FFT, etc) have a single column for&nbsp;the entire 10 minute clip, or is there a group&nbsp;of features for a smaller time interval, say 10 seconds, that gets repeated as a group?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57967,
      "author_name": "speechmachine",
      "author_url": "",
      "post_date": "11/13/2014 04:41:27",
      "content": "<p>Hi Hugh. From what I understood from the literature I have read -- it is the latter, i.e. there is a group of features for a smaller time interval that gets repeated over. The length of the smaller time interval depends on the paper you read. Some chose a smaller time window of 5s, some 10s and some 60s. In the <a href=\"http://www.plosone.org/article/fetchObject.action?uri=info:doi%2F10.1371%2Fjournal.pone.0081920&representation=PDF\">paper </a>that accompanies this competition does it in chunks of 60s.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 58023,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "11/13/2014 19:35:31",
      "content": "<p>I've also seen some papers that use&nbsp;a moving, overlapping window. In other words, if you have 100 points, you break that into 0-20, 10-30, 20-40, etc. and calculate features on each of the epochs.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "57859": "",
    "57966": "",
    "57967": "",
    "58023": ""
  },
  "source": "meta"
}