{
  "id": 10324,
  "title": "are adjacent segments from the same hour?",
  "url": "/competitions/seizure-prediction/discussion/10324",
  "author_name": "",
  "post_date": "2014-09-13T17:30:45.453Z",
  "votes": null,
  "comment_count": 6,
  "views": 1796,
  "content": "<p>Pre- and inter-ictal 10m segments are not labelled with the hour they came from.&nbsp; Are adjacent segments in the default numbering from the same hour?&nbsp; That is, if preictal_0032 is segment 4 and preictal_0033 is segment 5, are they actually adjacent in time?</p>",
  "messages": [
    {
      "id": "53652",
      "postDate": "09/13/2014 17:30:45",
      "content": "<p>Pre- and inter-ictal 10m segments are not labelled with the hour they came from.&nbsp; Are adjacent segments in the default numbering from the same hour?&nbsp; That is, if preictal_0032 is segment 4 and preictal_0033 is segment 5, are they actually adjacent in time?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53957",
      "postDate": "09/16/2014 09:06:05",
      "content": "<p>In general, they are. For example, Dog 1 has 24 preictal examples where 1 to 6 are adjacent, as well as 7 to 12, 13 to 18 and 19 to 24.</p>\n<p>For the fourth dog, however, there are 97 examples, so it is not clear where the 6-fold pattern is broken.</p>\n<p>I guess it is possible to find out which preictal and interictal segments are adjacent, but I have not tried to do it.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53977",
      "postDate": "09/16/2014 10:51:23",
      "content": "<p>Each segment in preictal and interictal have a sequence number stored in the metadata alongside channels/data_length_sec/sampling_frequeuncy. I assume they are part of the same sequence whenever the sequence increases by 1 and a new sequence starts when the number wraps around again.</p>\n<p>e.g. Dog_1 preictal sequences are 1,2,3,4,5,6,1,2,3,4,5,6,1,2,3,4,5,6,1,2,3,4,5,6 which I assume to mean 4 sequences of 1-6. For Dog_4 I have the sequences down as&nbsp;[1-6, 1-6, 1, 1-6, 2-6, 4-6, 1-6, 1-6, 2-6, 1-6, 1-6, 1-6, 1-6, 1-6, 1-6, 1-6, 1-6, 1-5]</p>\n<p>The separation is something I speculated on though.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53996",
      "postDate": "09/16/2014 13:03:37",
      "content": "<p>Michael is correct - the clips are usually in sequences of 6 as described on the data page (https://www.kaggle.com/c/seizure-prediction/data)- see the graphic in particular. However, there are places where some clips were omitted. Some may have been omitted if there were corrupt or unreadable data segments, or recording gaps in the original file. It looks like Dog 4's data had a few of these issues. Also, data segments may have been omitted from the end of the last data sequence as part of my effort to limit the size of the downloads, and in this instance clips may have been deleted from the end of a block of 6.</p>\n<p>I'm not seeing the single &quot;1&quot; clip in the original data bundle (#13 in Michael's post). Can you post which training clip this was so I can verify?</p>\n<p>We realize this irregularity in the data adds an additional wrinkle to your modelling efforts, but this, in addition to artifacts and noise as observed elsewhere in the forum, is a reality with real-world data. There are gaps where the equipment is being maintained or batteries are replaced, and electromagnetic induction (60Hz as these recordings are all from the US) does show up sometimes despite our best efforts at shielding. The data you're working with in this contest are rather unique, research long-duration ambulatory recordings (from the dogs) and wide-bandwidth recordings (from the human patients). However, even in routine clinical recordings some of these data issues are unavoidable, and this should be viewed as part of the challenge of the contest.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53999",
      "postDate": "09/16/2014 13:11:08",
      "content": "<p>It was for Dog_4 preictal, but I can't find the single &quot;1&quot; clip anymore.&nbsp;Might have been a programming error or made a mistake when&nbsp;I was manually massaging&nbsp;the metadata to take up less screen space so I could see it more easily.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55556",
      "postDate": "10/03/2014 15:08:25",
      "content": "<p>Do you think the seq_id is even relevant? After all, the information is not even provided for test data, and we generally don't know if some of the seq_id&nbsp;is missing in test data, as appears in training samples.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55773",
      "postDate": "10/08/2014 11:12:21",
      "content": "<p>As Byronyi pointed out, is it true for test data.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 53957,
      "author_name": "joseleiva",
      "author_url": "",
      "post_date": "09/16/2014 09:06:05",
      "content": "<p>In general, they are. For example, Dog 1 has 24 preictal examples where 1 to 6 are adjacent, as well as 7 to 12, 13 to 18 and 19 to 24.</p>\n<p>For the fourth dog, however, there are 97 examples, so it is not clear where the 6-fold pattern is broken.</p>\n<p>I guess it is possible to find out which preictal and interictal segments are adjacent, but I have not tried to do it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53977,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "09/16/2014 10:51:23",
      "content": "<p>Each segment in preictal and interictal have a sequence number stored in the metadata alongside channels/data_length_sec/sampling_frequeuncy. I assume they are part of the same sequence whenever the sequence increases by 1 and a new sequence starts when the number wraps around again.</p>\n<p>e.g. Dog_1 preictal sequences are 1,2,3,4,5,6,1,2,3,4,5,6,1,2,3,4,5,6,1,2,3,4,5,6 which I assume to mean 4 sequences of 1-6. For Dog_4 I have the sequences down as&nbsp;[1-6, 1-6, 1, 1-6, 2-6, 4-6, 1-6, 1-6, 2-6, 1-6, 1-6, 1-6, 1-6, 1-6, 1-6, 1-6, 1-6, 1-5]</p>\n<p>The separation is something I speculated on though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53996,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "09/16/2014 13:03:37",
      "content": "<p>Michael is correct - the clips are usually in sequences of 6 as described on the data page (https://www.kaggle.com/c/seizure-prediction/data)- see the graphic in particular. However, there are places where some clips were omitted. Some may have been omitted if there were corrupt or unreadable data segments, or recording gaps in the original file. It looks like Dog 4's data had a few of these issues. Also, data segments may have been omitted from the end of the last data sequence as part of my effort to limit the size of the downloads, and in this instance clips may have been deleted from the end of a block of 6.</p>\n<p>I'm not seeing the single &quot;1&quot; clip in the original data bundle (#13 in Michael's post). Can you post which training clip this was so I can verify?</p>\n<p>We realize this irregularity in the data adds an additional wrinkle to your modelling efforts, but this, in addition to artifacts and noise as observed elsewhere in the forum, is a reality with real-world data. There are gaps where the equipment is being maintained or batteries are replaced, and electromagnetic induction (60Hz as these recordings are all from the US) does show up sometimes despite our best efforts at shielding. The data you're working with in this contest are rather unique, research long-duration ambulatory recordings (from the dogs) and wide-bandwidth recordings (from the human patients). However, even in routine clinical recordings some of these data issues are unavoidable, and this should be viewed as part of the challenge of the contest.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53999,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "09/16/2014 13:11:08",
      "content": "<p>It was for Dog_4 preictal, but I can't find the single &quot;1&quot; clip anymore.&nbsp;Might have been a programming error or made a mistake when&nbsp;I was manually massaging&nbsp;the metadata to take up less screen space so I could see it more easily.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55556,
      "author_name": "",
      "author_url": "",
      "post_date": "10/03/2014 15:08:25",
      "content": "<p>Do you think the seq_id is even relevant? After all, the information is not even provided for test data, and we generally don't know if some of the seq_id&nbsp;is missing in test data, as appears in training samples.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55773,
      "author_name": "saikumarallaka",
      "author_url": "",
      "post_date": "10/08/2014 11:12:21",
      "content": "<p>As Byronyi pointed out, is it true for test data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "53652": "",
    "53957": "",
    "53977": "",
    "53996": "",
    "53999": "",
    "55556": "",
    "55773": ""
  },
  "source": "meta"
}