{
  "id": 10417,
  "title": "Problem with sequences",
  "url": "/competitions/seizure-prediction/discussion/10417",
  "author_name": "",
  "post_date": "2014-09-22T20:36:13.467Z",
  "votes": null,
  "comment_count": 5,
  "views": 1463,
  "content": "<p>In the competition description it is stated that one hour records have been performed, and every sample file has been taken segmenting hours into 10 minute files. So, it is expected to have 6 files for every hour. However, I found the following sequences which are&nbsp;discontinuous&nbsp;in different ways. Could someone else check if I'm correct? The sequence data can be important to make a proper block&nbsp;partition for cross-validation.</p>\n<p><strong>Dog_2_interictal_segment_0500 2&nbsp;&lt;-- lost sequences 3,4,5,6</strong><br>Dog_3_preictal_segment_0001 1<br><br>Dog_4_preictal_segment_0018 6<br><strong>Dog_4_preictal_segment_0019 2 &lt;-- lost sequence 1</strong><br>Dog_4_preictal_segment_0020 3<br><br>Dog_4_preictal_segment_0023 6<br><strong>Dog_4_preictal_segment_0024 4 &lt;-- lost sequences 1,2,3</strong><br>Dog_4_preictal_segment_0025 5<br><br>Dog_4_preictal_segment_0038 6<br><strong>Dog_4_preictal_segment_0039 2 &lt;-- lost sequence 1</strong><br>Dog_4_preictal_segment_0040 3<br><br><strong>Patient_1_interictal_segment_0050 2 &lt;-- lost&nbsp;sequences 3,4,5,6</strong><br>Patient_2_preictal_segment_0001 1<br><br></p>",
  "messages": [
    {
      "id": "54480",
      "postDate": "09/22/2014 20:36:13",
      "content": "<p>In the competition description it is stated that one hour records have been performed, and every sample file has been taken segmenting hours into 10 minute files. So, it is expected to have 6 files for every hour. However, I found the following sequences which are&nbsp;discontinuous&nbsp;in different ways. Could someone else check if I'm correct? The sequence data can be important to make a proper block&nbsp;partition for cross-validation.</p>\n<p><strong>Dog_2_interictal_segment_0500 2&nbsp;&lt;-- lost sequences 3,4,5,6</strong><br>Dog_3_preictal_segment_0001 1<br><br>Dog_4_preictal_segment_0018 6<br><strong>Dog_4_preictal_segment_0019 2 &lt;-- lost sequence 1</strong><br>Dog_4_preictal_segment_0020 3<br><br>Dog_4_preictal_segment_0023 6<br><strong>Dog_4_preictal_segment_0024 4 &lt;-- lost sequences 1,2,3</strong><br>Dog_4_preictal_segment_0025 5<br><br>Dog_4_preictal_segment_0038 6<br><strong>Dog_4_preictal_segment_0039 2 &lt;-- lost sequence 1</strong><br>Dog_4_preictal_segment_0040 3<br><br><strong>Patient_1_interictal_segment_0050 2 &lt;-- lost&nbsp;sequences 3,4,5,6</strong><br>Patient_2_preictal_segment_0001 1<br><br></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54482",
      "postDate": "09/22/2014 20:51:26",
      "content": "<p>I posted a link to my partitioned file list in <a href=\"http://www.kaggle.com/c/seizure-prediction/forums/t/10405/python-code-for-cv-splitting\">this thread</a>.</p>\n<p>Seems like we agree, although I didn't check all&nbsp;of your examples.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54493",
      "postDate": "09/23/2014 01:31:51",
      "content": "<p>The sequence gaps are correct. In addition, you would expect that you can simply concatenate samples when the segments have consecutive sequence number, however, in some cases (e.g. Patient_2) there is a large value jump which can indicate that the sequences do not exactly follow each other in time.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55774",
      "postDate": "10/08/2014 11:21:47",
      "content": "<p>@zzspar: In test data, sequence number is not provided in data structure.</p>\n<p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Still&nbsp;can we consider 6 segments in sequence belong to 1 hour?</p>\n<p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; If that is the case, in patient2 data segments from&nbsp;<strong>Patient_2_test_segment_0001</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; to&nbsp;<strong>Patient_2_test_segment_0006&nbsp;</strong>must belong to either&nbsp;<strong>interictal&nbsp;</strong>or&nbsp;<strong>preictal.</strong></p>\n<p><strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</strong>Please correct me if i am wrong.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55786",
      "postDate": "10/08/2014 13:39:45",
      "content": "<p>As far as I know, the competition states that test samples are taken randomly and shuffled, so you can't make any assumption about sequences in test set.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55797",
      "postDate": "10/08/2014 16:32:33",
      "content": "<p>This is correct. Shuffling was done on a data clip basis - no effort was made to keep sequential test segments together.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 54482,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "09/22/2014 20:51:26",
      "content": "<p>I posted a link to my partitioned file list in <a href=\"http://www.kaggle.com/c/seizure-prediction/forums/t/10405/python-code-for-cv-splitting\">this thread</a>.</p>\n<p>Seems like we agree, although I didn't check all&nbsp;of your examples.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54493,
      "author_name": "udibr1",
      "author_url": "",
      "post_date": "09/23/2014 01:31:51",
      "content": "<p>The sequence gaps are correct. In addition, you would expect that you can simply concatenate samples when the segments have consecutive sequence number, however, in some cases (e.g. Patient_2) there is a large value jump which can indicate that the sequences do not exactly follow each other in time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55774,
      "author_name": "saikumarallaka",
      "author_url": "",
      "post_date": "10/08/2014 11:21:47",
      "content": "<p>@zzspar: In test data, sequence number is not provided in data structure.</p>\n<p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Still&nbsp;can we consider 6 segments in sequence belong to 1 hour?</p>\n<p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; If that is the case, in patient2 data segments from&nbsp;<strong>Patient_2_test_segment_0001</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; to&nbsp;<strong>Patient_2_test_segment_0006&nbsp;</strong>must belong to either&nbsp;<strong>interictal&nbsp;</strong>or&nbsp;<strong>preictal.</strong></p>\n<p><strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</strong>Please correct me if i am wrong.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55786,
      "author_name": "pakozm",
      "author_url": "",
      "post_date": "10/08/2014 13:39:45",
      "content": "<p>As far as I know, the competition states that test samples are taken randomly and shuffled, so you can't make any assumption about sequences in test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55797,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "10/08/2014 16:32:33",
      "content": "<p>This is correct. Shuffling was done on a data clip basis - no effort was made to keep sequential test segments together.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "54480": "",
    "54482": "",
    "54493": "",
    "55774": "",
    "55786": "",
    "55797": ""
  },
  "source": "meta"
}