{
  "id": 10495,
  "title": "Construct the data set",
  "url": "/competitions/seizure-prediction/discussion/10495",
  "author_name": "",
  "post_date": "2014-10-01T02:15:15.063Z",
  "votes": null,
  "comment_count": 3,
  "views": 1098,
  "content": "<p>Hi everyone,</p>\n<p>I just downloaded part of the data and am trying to understand the data first. And I am a bit confused right now.</p>\n<p>So the competition can be viewed as a two-class classification problem, and it is almost all about feature extraction. My question is, for instance, can we take files</p>\n<p>Dog_1_interictal_segment_0001.mat</p>\n<p>Dog_1_interictal_segment_0002.mat</p>\n<p>.....</p>\n<p>Dog_1_interictal_segment_0006.mat</p>\n<p>and do whatever extraction to them, then classify it as &quot;interictal&quot;, use it as one example in our training data.</p>\n<p>and then take interictal_segment_0007 to 00012, classify as another &quot;interictal&quot;</p>\n<p>etc.</p>\n<p>And similarly, take</p>\n<p>Dog_1_preictal_segment_0001.mat ----&nbsp;Dog_1_preictal_segment_0006.mat</p>\n<p>do feature engineering, and classified as &quot;preictal&quot;.</p>\n<p>Since 6 mat files combined represents one hour prior activity, I think every 6 files should be considered as one example in our models.</p>\n<p>But the problem is, the test mat files are not organized this way, so it seems not appropriate to do it this way.&nbsp;</p>\n<p>Or, maybe, we can consider each mat file as one row in constructing the training data space?</p>\n\n<p>Please anyone gives anyone suggestion would be very appreciated!</p>\n<p>Best,</p>",
  "messages": [
    {
      "id": "55448",
      "postDate": "10/01/2014 02:15:15",
      "content": "<p>Hi everyone,</p>\n<p>I just downloaded part of the data and am trying to understand the data first. And I am a bit confused right now.</p>\n<p>So the competition can be viewed as a two-class classification problem, and it is almost all about feature extraction. My question is, for instance, can we take files</p>\n<p>Dog_1_interictal_segment_0001.mat</p>\n<p>Dog_1_interictal_segment_0002.mat</p>\n<p>.....</p>\n<p>Dog_1_interictal_segment_0006.mat</p>\n<p>and do whatever extraction to them, then classify it as &quot;interictal&quot;, use it as one example in our training data.</p>\n<p>and then take interictal_segment_0007 to 00012, classify as another &quot;interictal&quot;</p>\n<p>etc.</p>\n<p>And similarly, take</p>\n<p>Dog_1_preictal_segment_0001.mat ----&nbsp;Dog_1_preictal_segment_0006.mat</p>\n<p>do feature engineering, and classified as &quot;preictal&quot;.</p>\n<p>Since 6 mat files combined represents one hour prior activity, I think every 6 files should be considered as one example in our models.</p>\n<p>But the problem is, the test mat files are not organized this way, so it seems not appropriate to do it this way.&nbsp;</p>\n<p>Or, maybe, we can consider each mat file as one row in constructing the training data space?</p>\n\n<p>Please anyone gives anyone suggestion would be very appreciated!</p>\n<p>Best,</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55468",
      "postDate": "10/01/2014 19:35:15",
      "content": "<p>Hi</p>\n<p>As each test file is one &quot;example&quot; to be classified, also each &quot;interictal&quot; or &quot;preictal&quot; must be an example to be classified. If in your training approach you wish to aggregate mat files or (as I did) divide them in pieces to create new training examples is another question...&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55472",
      "postDate": "10/01/2014 21:08:48",
      "content": "<p>Thanks&nbsp;RuiR,</p>\n<p>That is what I figured. Then it seems difficult to use the 1-6 sequence values....</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55552",
      "postDate": "10/03/2014 14:44:45",
      "content": "<p>Or maybe it's just useless. I think they try to dynamically estimates (i.e. for every 10 minuates, or with time window of 10 min) the risk of seizure occurring in the next hour, and you need to segregate those pre-seizure and&nbsp;interictal clips clearly. Although seq_id=6 of preictal clip seems more &quot;preictal&quot;, it's weighted as important as those of seq_id=1, at least in the context of this competition.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 55468,
      "author_name": "ruiprodrigues",
      "author_url": "",
      "post_date": "10/01/2014 19:35:15",
      "content": "<p>Hi</p>\n<p>As each test file is one &quot;example&quot; to be classified, also each &quot;interictal&quot; or &quot;preictal&quot; must be an example to be classified. If in your training approach you wish to aggregate mat files or (as I did) divide them in pieces to create new training examples is another question...&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55472,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "10/01/2014 21:08:48",
      "content": "<p>Thanks&nbsp;RuiR,</p>\n<p>That is what I figured. Then it seems difficult to use the 1-6 sequence values....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55552,
      "author_name": "",
      "author_url": "",
      "post_date": "10/03/2014 14:44:45",
      "content": "<p>Or maybe it's just useless. I think they try to dynamically estimates (i.e. for every 10 minuates, or with time window of 10 min) the risk of seizure occurring in the next hour, and you need to segregate those pre-seizure and&nbsp;interictal clips clearly. Although seq_id=6 of preictal clip seems more &quot;preictal&quot;, it's weighted as important as those of seq_id=1, at least in the context of this competition.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "55448": "",
    "55468": "",
    "55472": "",
    "55552": ""
  },
  "source": "meta"
}