{
  "id": 10128,
  "title": "Fields in Matlab file - can't find 'data' etc",
  "url": "/competitions/seizure-prediction/discussion/10128",
  "author_name": "",
  "post_date": "2014-08-27T19:23:33.173Z",
  "votes": 1,
  "comment_count": 13,
  "views": 4713,
  "content": "<p>After I import the files into Python, I can't find the field names 'Data' , 'data_length_sec', 'sampling_frequency' etc etc. There's an array with</p>\n<p>'preictal_segment_1': array</p>\n<p>etc</p>\n<p>but I can't access by the field names given in data description. Any idea? I can see fields in R but not Python</p>",
  "messages": [
    {
      "id": "52537",
      "postDate": "08/27/2014 19:23:33",
      "content": "<p>After I import the files into Python, I can't find the field names 'Data' , 'data_length_sec', 'sampling_frequency' etc etc. There's an array with</p>\n<p>'preictal_segment_1': array</p>\n<p>etc</p>\n<p>but I can't access by the field names given in data description. Any idea? I can see fields in R but not Python</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52556",
      "postDate": "08/28/2014 07:02:16",
      "content": "<p>I hope it's help.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52560",
      "postDate": "08/28/2014 08:29:02",
      "content": "<p>BTW, it's seems we spend many GB for zeroes ) All data traces saved in DBL (double), but they appears in SHORT (i16)... At least the examples I tried..</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52570",
      "postDate": "08/28/2014 13:04:40",
      "content": "<p>&nbsp;If you use scipy.io.loadmat() then you will have all data you mentioned in a python dictionary , but nested in other data structures. If you print the result you can figure out how to access each field.</p>\n<p>For example ,if you have</p>\n<p>data_struct = scipy.io.loadmat('Dog_1_interictal_segment_0008.mat')</p>\n<p>with code:</p>\n<p>data_struct['interictal_segment_8'][0][0][0][ch][samp]</p>\n<p>you can extract the sample with index 'samp' from chanel&nbsp; 'ch'&nbsp;&nbsp; from file 'Dog_1_interictal_segment_0008.mat'</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52598",
      "postDate": "08/28/2014 19:39:25",
      "content": "<p>[quote=claudiu1989;52570]</p>\n<p>&nbsp;If you use scipy.io.loadmat() then you will have all data you mentioned in a python dictionary , but nested in other data structures. If you print the result you can figure out how to access each field.</p>\n<p>For example ,if you have</p>\n<p>data_struct = scipy.io.loadmat('Dog_1_interictal_segment_0008.mat')</p>\n<p>with code:</p>\n<p>data_struct['interictal_segment_8'][0][0][0][ch][samp]</p>\n<p>you can extract the sample with index 'samp' from chanel&nbsp; 'ch'&nbsp;&nbsp; from file 'Dog_1_interictal_segment_0008.mat'</p>\n<p>[/quote]</p>\n<p>Expanding on Claudiu1989's great response:</p>\n<p>data_struct[sample][0][0] is your 'base' level to access the clip's information. From there:</p>\n<ul>\n<li>data_struct[sample][0][0][0][x] - The series of electrode measurements corresponding to electrode x. For the example I was looking at from Dog_5, this = 239766 measurements (399 Hz * 600 seconds).</li>\n<li>data_struct[sample][0][0][1][0] - The length (in seconds) of the clip. If they are all 10 minutes as described, this should be 600.</li>\n<li>data_struct[sample][0][0][2][0] - The sampling rate in Hz (e.g. ~399 for Dog_5)</li>\n<li>data_struct[sample][0][0][3][0][x] - The name of the xth electrode.</li>\n<li>data_struct[sample][0][0][4][0] - The index of the clip's location within the hour (e.g. 4 = from minutes 40-50)<br><br></li>\n</ul>\n<p>EDIT: For the last bullet, you should only expect to have a '4' value for the third index for preictal and interictal files. Since we are not given the test files in the context of an hour long series, they are not provided with this field. Thanks to Lawrence for pointing out the lack of clarity! :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "52600",
      "postDate": "08/28/2014 20:36:37",
      "content": "<p>In case if you prefer R:</p>\n<p><code>library(R.matlab)<br>pat &lt;- readMat('Patient_1\\\\Patient_1_interictal_segment_0018.mat')<br>#pat[[1]][[2]] == 600<br>#pat[[1]][[3]] == 5000<br>df &lt;- data.frame(t(pat[[1]][[1]]))<br>names(df) &lt;- unlist(pat[[1]][[4]])<br>head(df)<br></code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53073",
      "postDate": "09/03/2014 21:09:21",
      "content": "<p>Thanks all. Can I confirm that for each segment (eg Patient 1), I would end up with a 300000 * 15 data frame (in R)? ta</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53203",
      "postDate": "09/05/2014 20:04:45",
      "content": "<p>I do not see any values for the third index set to &quot;4&quot; as LAD wrote, and as a sanity check: take&nbsp;</p>\n<p>data_length_sec =&nbsp;&nbsp;data_struct[sample][0][0][<strong>1</strong>][0][0] &nbsp; &nbsp; &nbsp;(e.g. 600 seconds)&nbsp;</p>\n<p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;multiplied by&nbsp;</p>\n<p>sampling_frequency = data_struct[sample][0][0][<strong>2</strong>][0][0] &nbsp; &nbsp;(e.g. 399.61 Hz)</p>\n<p>to see that it matches the amount of data in the clip for that electrode:</p>\n<p>len(&nbsp;data_struct[sample][0][0][0][0] ) &nbsp; &nbsp;(e.g.&nbsp;239766)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53206",
      "postDate": "09/05/2014 20:50:06",
      "content": "<p>Are you missing the '4' in third index for any interictal or priectal&nbsp;files? I&nbsp;should have been more explicit in the original post I made, but I don't believe you should expect one for the test files, since we are not given them in the context of an hour long segment (as we are for preictal and interictal).</p>\n<p>I only have Dog_5 locally on this computer to verify, but some quick scanning seems to confirm&nbsp;this to be true for substituent data files (i.e. preictal and interictal files have a '4' vale for third index, test files do not).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53207",
      "postDate": "09/05/2014 21:07:26",
      "content": "<p>Sorry, my mistake. It was missing in a test file.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53209",
      "postDate": "09/05/2014 21:09:17",
      "content": "<p>No worries! It was good that you pointed it out, I should have been more explicit in the original post (I'll go back and edit it now for clarity) :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53327",
      "postDate": "09/08/2014 06:43:43",
      "content": "<p>[quote=Vadym Gnatkovsky;52560]</p>\n<p>BTW, it's seems we spend many GB for zeroes ) All data traces saved in DBL (double), but they appears in SHORT (i16)... At least the examples I tried..</p>\n<p>[/quote]</p>\n\n<p>I have only looked at Dog_5 (the short one). Can anyone confirm that these are all i16's taken straight from the ADC, so I can safely shrink them?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56108",
      "postDate": "10/16/2014 00:04:02",
      "content": "<p>Is anybody having issues loading some of the files using scipy.io.loadmat. So far 2 files are giving me: &quot;IOError:could not read bytes&quot; (currently checking all files)</p>\n<p>1) Dog_1_test_segment_0088.mat</p>\n<p>2) Dog_2_interictal_segment_0111.mat</p>\n\n<p>Anyone else experiencing this and are matlab/octave people seeing this?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "216263",
      "postDate": "08/25/2017 03:03:29",
      "content": "<p>could anyone share a scrip for reading the data using scipy.io.loadmat? it's would be great</p>",
      "rawMarkdown": "could anyone share a scrip for reading the data using scipy.io.loadmat? it's would be great",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 52556,
      "author_name": "vadymgnatkovsky",
      "author_url": "",
      "post_date": "08/28/2014 07:02:16",
      "content": "<p>I hope it's help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52560,
      "author_name": "vadymgnatkovsky",
      "author_url": "",
      "post_date": "08/28/2014 08:29:02",
      "content": "<p>BTW, it's seems we spend many GB for zeroes ) All data traces saved in DBL (double), but they appears in SHORT (i16)... At least the examples I tried..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52570,
      "author_name": "claudiu1989",
      "author_url": "",
      "post_date": "08/28/2014 13:04:40",
      "content": "<p>&nbsp;If you use scipy.io.loadmat() then you will have all data you mentioned in a python dictionary , but nested in other data structures. If you print the result you can figure out how to access each field.</p>\n<p>For example ,if you have</p>\n<p>data_struct = scipy.io.loadmat('Dog_1_interictal_segment_0008.mat')</p>\n<p>with code:</p>\n<p>data_struct['interictal_segment_8'][0][0][0][ch][samp]</p>\n<p>you can extract the sample with index 'samp' from chanel&nbsp; 'ch'&nbsp;&nbsp; from file 'Dog_1_interictal_segment_0008.mat'</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52598,
      "author_name": "lad215562",
      "author_url": "",
      "post_date": "08/28/2014 19:39:25",
      "content": "<p>[quote=claudiu1989;52570]</p>\n<p>&nbsp;If you use scipy.io.loadmat() then you will have all data you mentioned in a python dictionary , but nested in other data structures. If you print the result you can figure out how to access each field.</p>\n<p>For example ,if you have</p>\n<p>data_struct = scipy.io.loadmat('Dog_1_interictal_segment_0008.mat')</p>\n<p>with code:</p>\n<p>data_struct['interictal_segment_8'][0][0][0][ch][samp]</p>\n<p>you can extract the sample with index 'samp' from chanel&nbsp; 'ch'&nbsp;&nbsp; from file 'Dog_1_interictal_segment_0008.mat'</p>\n<p>[/quote]</p>\n<p>Expanding on Claudiu1989's great response:</p>\n<p>data_struct[sample][0][0] is your 'base' level to access the clip's information. From there:</p>\n<ul>\n<li>data_struct[sample][0][0][0][x] - The series of electrode measurements corresponding to electrode x. For the example I was looking at from Dog_5, this = 239766 measurements (399 Hz * 600 seconds).</li>\n<li>data_struct[sample][0][0][1][0] - The length (in seconds) of the clip. If they are all 10 minutes as described, this should be 600.</li>\n<li>data_struct[sample][0][0][2][0] - The sampling rate in Hz (e.g. ~399 for Dog_5)</li>\n<li>data_struct[sample][0][0][3][0][x] - The name of the xth electrode.</li>\n<li>data_struct[sample][0][0][4][0] - The index of the clip's location within the hour (e.g. 4 = from minutes 40-50)<br><br></li>\n</ul>\n<p>EDIT: For the last bullet, you should only expect to have a '4' value for the third index for preictal and interictal files. Since we are not given the test files in the context of an hour long series, they are not provided with this field. Thanks to Lawrence for pointing out the lack of clarity! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 52600,
      "author_name": "spimenov",
      "author_url": "",
      "post_date": "08/28/2014 20:36:37",
      "content": "<p>In case if you prefer R:</p>\n<p><code>library(R.matlab)<br>pat &lt;- readMat('Patient_1\\\\Patient_1_interictal_segment_0018.mat')<br>#pat[[1]][[2]] == 600<br>#pat[[1]][[3]] == 5000<br>df &lt;- data.frame(t(pat[[1]][[1]]))<br>names(df) &lt;- unlist(pat[[1]][[4]])<br>head(df)<br></code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53073,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "09/03/2014 21:09:21",
      "content": "<p>Thanks all. Can I confirm that for each segment (eg Patient 1), I would end up with a 300000 * 15 data frame (in R)? ta</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53203,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "09/05/2014 20:04:45",
      "content": "<p>I do not see any values for the third index set to &quot;4&quot; as LAD wrote, and as a sanity check: take&nbsp;</p>\n<p>data_length_sec =&nbsp;&nbsp;data_struct[sample][0][0][<strong>1</strong>][0][0] &nbsp; &nbsp; &nbsp;(e.g. 600 seconds)&nbsp;</p>\n<p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;multiplied by&nbsp;</p>\n<p>sampling_frequency = data_struct[sample][0][0][<strong>2</strong>][0][0] &nbsp; &nbsp;(e.g. 399.61 Hz)</p>\n<p>to see that it matches the amount of data in the clip for that electrode:</p>\n<p>len(&nbsp;data_struct[sample][0][0][0][0] ) &nbsp; &nbsp;(e.g.&nbsp;239766)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53206,
      "author_name": "lad215562",
      "author_url": "",
      "post_date": "09/05/2014 20:50:06",
      "content": "<p>Are you missing the '4' in third index for any interictal or priectal&nbsp;files? I&nbsp;should have been more explicit in the original post I made, but I don't believe you should expect one for the test files, since we are not given them in the context of an hour long segment (as we are for preictal and interictal).</p>\n<p>I only have Dog_5 locally on this computer to verify, but some quick scanning seems to confirm&nbsp;this to be true for substituent data files (i.e. preictal and interictal files have a '4' vale for third index, test files do not).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53207,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "09/05/2014 21:07:26",
      "content": "<p>Sorry, my mistake. It was missing in a test file.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53209,
      "author_name": "lad215562",
      "author_url": "",
      "post_date": "09/05/2014 21:09:17",
      "content": "<p>No worries! It was good that you pointed it out, I should have been more explicit in the original post (I'll go back and edit it now for clarity) :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53327,
      "author_name": "eumeswil0",
      "author_url": "",
      "post_date": "09/08/2014 06:43:43",
      "content": "<p>[quote=Vadym Gnatkovsky;52560]</p>\n<p>BTW, it's seems we spend many GB for zeroes ) All data traces saved in DBL (double), but they appears in SHORT (i16)... At least the examples I tried..</p>\n<p>[/quote]</p>\n\n<p>I have only looked at Dog_5 (the short one). Can anyone confirm that these are all i16's taken straight from the ADC, so I can safely shrink them?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56108,
      "author_name": "corticalharmony",
      "author_url": "",
      "post_date": "10/16/2014 00:04:02",
      "content": "<p>Is anybody having issues loading some of the files using scipy.io.loadmat. So far 2 files are giving me: &quot;IOError:could not read bytes&quot; (currently checking all files)</p>\n<p>1) Dog_1_test_segment_0088.mat</p>\n<p>2) Dog_2_interictal_segment_0111.mat</p>\n\n<p>Anyone else experiencing this and are matlab/octave people seeing this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 216263,
      "author_name": "ngduyanhece",
      "author_url": "",
      "post_date": "08/25/2017 03:03:29",
      "content": "<p>could anyone share a scrip for reading the data using scipy.io.loadmat? it's would be great</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "52537": "",
    "52556": "",
    "52560": "",
    "52570": "",
    "52598": "",
    "52600": "",
    "53073": "",
    "53203": "",
    "53206": "",
    "53207": "",
    "53209": "",
    "53327": "",
    "56108": "",
    "216263": "could anyone share a scrip for reading the data using scipy.io.loadmat? it's would be great"
  },
  "source": "meta"
}