{
  "id": 7882,
  "title": "Python preprocess of MAT Files",
  "url": "/competitions/decoding-the-human-brain/discussion/7882",
  "author_name": "",
  "post_date": "2014-04-26T14:40:06.143Z",
  "votes": null,
  "comment_count": 8,
  "views": 2837,
  "content": "<p>Hey,</p>\n\n<p>I'm C/C++ programmer and need want to have a look in to the data.</p>\n<p>With youre Example-Code on Kaggle ive tried to convert the train MAT files to one large ascii file with the format</p>\n\n<p>ID&nbsp;&nbsp;&nbsp; CHANNEL&nbsp; TARGET&nbsp;&nbsp; MEASURE#1&nbsp;&nbsp;&nbsp;&nbsp; MEASURE#2&nbsp;&nbsp; ... MEASURE#375</p>\n<p>this is the code :</p>\n<p>#!/usr/bin/python</p>\n<p>import numpy as np<br>from scipy.io import loadmat</p>\n<p>subjects_train = range(1, 17) <br>fp = open('train.dat', 'w')</p>\n<p>for subject in subjects_train:<br> <br> filename = 'train_subject%02d.mat' % subject<br> data = loadmat(filename, squeeze_me=True)</p>\n<p>X = data['X' ]<br> Y = data['y' ]<br> <br> #X -= X.mean(0)<br> #X = np.nan_to_num(X / X.std(0))</p>\n<p>trails=len( X )<br> for trail in range( trails ):<br> id='%02d%03d' % (subject, trail)<br> print id<br> channels=len(X[0])<br> for channel in range( channels ):<br> times=len( X[0][0] )<br> fp.write( id + &quot;\\t&quot; + str(channel) + &quot;\\t&quot;+str(Y[ trail]) )<br> for time in range( times ):<br> fp.write( &quot;\\t&quot; + &quot;%20.17f &quot; % ( X[ trail ][ channel ][ time ]) )<br> fp.write(&quot;\\n&quot;)<br>fp.close()</p>\n\n<p>is there anything wrong with it ? Couze the values lookes really small ? :-/</p>",
  "messages": [
    {
      "id": "43080",
      "postDate": "04/26/2014 14:40:06",
      "content": "<p>Hey,</p>\n\n<p>I'm C/C++ programmer and need want to have a look in to the data.</p>\n<p>With youre Example-Code on Kaggle ive tried to convert the train MAT files to one large ascii file with the format</p>\n\n<p>ID&nbsp;&nbsp;&nbsp; CHANNEL&nbsp; TARGET&nbsp;&nbsp; MEASURE#1&nbsp;&nbsp;&nbsp;&nbsp; MEASURE#2&nbsp;&nbsp; ... MEASURE#375</p>\n<p>this is the code :</p>\n<p>#!/usr/bin/python</p>\n<p>import numpy as np<br>from scipy.io import loadmat</p>\n<p>subjects_train = range(1, 17) <br>fp = open('train.dat', 'w')</p>\n<p>for subject in subjects_train:<br> <br> filename = 'train_subject%02d.mat' % subject<br> data = loadmat(filename, squeeze_me=True)</p>\n<p>X = data['X' ]<br> Y = data['y' ]<br> <br> #X -= X.mean(0)<br> #X = np.nan_to_num(X / X.std(0))</p>\n<p>trails=len( X )<br> for trail in range( trails ):<br> id='%02d%03d' % (subject, trail)<br> print id<br> channels=len(X[0])<br> for channel in range( channels ):<br> times=len( X[0][0] )<br> fp.write( id + &quot;\\t&quot; + str(channel) + &quot;\\t&quot;+str(Y[ trail]) )<br> for time in range( times ):<br> fp.write( &quot;\\t&quot; + &quot;%20.17f &quot; % ( X[ trail ][ channel ][ time ]) )<br> fp.write(&quot;\\n&quot;)<br>fp.close()</p>\n\n<p>is there anything wrong with it ? Couze the values lookes really small ? :-/</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43088",
      "postDate": "04/26/2014 17:01:20",
      "content": "<p>Hi,</p>\n<p>The magnetic field induced by the currents in the brain is very very low. Consider that the unit of measure is Tesla (T) and the actual values are in the order of 10^-15 T, called femtoTesla (fT). More precisely, each&nbsp;magnetometer - i.e. the sensor measuring the z component of the magnetic field in one location - returns Tesla. Each gradiometer - i.e. the sensor measuring the x or y component of the <em>gradient</em> of the magnetic field in one location - returns Tesla/cm. In each of the 102 locations there is one magnetometer and two gradiometers.</p>\n<p>Having said that, it may be convenient for you to save the data in floating point exponential format, with &quot;%e&quot; (or &quot;%g&quot;), instead of &nbsp;the decimal format &quot;%20.17f&quot; that you use. I guess &quot;%1.20e&quot; should be enough. See here for additional details:</p>\n<p>&nbsp;&nbsp;<a href=\"https://docs.python.org/2/library/stdtypes.html#string-formatting\">https://docs.python.org/2/library/stdtypes.html#string-formatting</a></p>\n<p>&nbsp;&nbsp;<a href=\"https://docs.python.org/3/library/stdtypes.html#printf-style-string-formatting\">https://docs.python.org/3/library/stdtypes.html#printf-style-string-formatting</a></p>\n\n<p>Note: the setup of gradiometers/magnetometers and number of locations changes across different MEG devices. The one used for acquiring the data of this competition, i.e. Elekta Neuromag VectorView, is as described above.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43131",
      "postDate": "04/27/2014 08:51:27",
      "content": "<p>Hey,</p>\n<p>thank you really mutch for youre time to answer the question. You make mutch things clear for me.</p>\n<p>But in one lines, the challenge is to predict 2 classes ( Face and Scramble ) from the measure. There is no importans of analyset this https://github.com/mne-tools/mne-python/blob/master/mne/layouts/Vectorview-all.lout file ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43134",
      "postDate": "04/27/2014 10:16:59",
      "content": "<p>Hi,</p>\n<p>The&nbsp;<a href=\"https://github.com/mne-tools/mne-python/blob/master/mne/layouts/Vectorview-all.lout\">Vectorview-all.lout</a> is a text file that describes the approximate 2D position of the sensors with respect to the head of the subject. I used that file and the data from subject 1 to create the&nbsp;<a href=\"http://nilab.cimec.unitn.it/people/olivetti/decmeg2014/MEG_data2.png\">figure</a> in the description page of the competition. It might be useful to know these approximate positions in order to account for spatial correlations between the timeseries. But I have no strong opinion if this information can affect the score of the competition.</p>\n<p>I plan to release a bit more code in the next days, in the Github repository. For example the one I used to create the figure.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43157",
      "postDate": "04/27/2014 15:28:41",
      "content": "<p>I'm new to this sort of data, so my apologies if these questions can be inferred from the existing data description.&nbsp;</p>\n<ul>\n<li>Are the 306 channels identically ordered in all of the test subjects?&nbsp; E.g., does channel 107 in trial subject 01 correspond to the same sensor-type &amp; location in test subject 17?</li>\n<li>Also, can the layouts in your github repsitory (Vectorview-mag &amp; Vectorview-grad) be used to reliably associate a channel with a specific type of sensor (since there are two types of sensor in use)?</li>\n</ul>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43162",
      "postDate": "04/27/2014 16:09:08",
      "content": "<p>Hi,</p>\n<p>Thank you for your questions.</p>\n<p>The sensors are in the same order over all 23 subjects. The order of the channels in the dataset is the same as in <a href=\"https://github.com/FBK-NILab/DecMeg2014/blob/master/additional_files/Vectorview-all.lout\">Vectorview-all.lout</a> that is in the Github repository of the competition. The other two files that you mention, which are available from a different repository (specifically <a href=\"https://github.com/mne-tools/mne-python/tree/master/mne/layouts\">here</a>), provide the same information of Vectorview-all.lout, but split in two parts: Vectorview-mag for magnetometers and Vectorview-grad for gradiometers. So, in my opinion, the right way to proceed, is to use Vectorview-all.lout in order to retrieve the name and position of each sensor. For example, sensor &quot;MEG 0113&quot; has x=-73.416206 and y=33.416687 (see the 2nd line of the file). Notice that if the number in the sensor name ends with &quot;1&quot; then it is a magnetometer. If it ends with &quot;2&quot; or &quot;3&quot;, then it is a gradiometer. For example, &quot;MEG 0113&quot; is a gradiometer and &quot;MEG 0111&quot; is a magnetometer.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43601",
      "postDate": "05/03/2014 12:07:55",
      "content": "<p>Hi Emanuele,</p>\n<p>I would suggest that you add this:</p>\n<p>[quote=Emanuele;43162]</p>\n<p>Notice that if the number in the sensor name ends with &quot;1&quot; then it is a magnetometer. If it ends with &quot;2&quot; or &quot;3&quot;, then it is a gradiometer. For example, &quot;MEG 0113&quot; is a gradiometer and &quot;MEG 0111&quot; is a magnetometer.</p>\n<p>[/quote]</p>\n<p>to the description of the data. It may be useful for everyone (not just people reading &quot;Python preprocess of MAT files&quot;).</p>\n<p>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43617",
      "postDate": "05/03/2014 15:40:25",
      "content": "<p>Good point. I'll ask to the Kaggle people to change the data description page. We don't have permission to do that anymore.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "47345",
      "postDate": "05/30/2014 16:09:10",
      "content": "<p>Hello,</p>\n<p>What does the first column in Vectorview-all.lout specify?</p>\n<p>Thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 43088,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "04/26/2014 17:01:20",
      "content": "<p>Hi,</p>\n<p>The magnetic field induced by the currents in the brain is very very low. Consider that the unit of measure is Tesla (T) and the actual values are in the order of 10^-15 T, called femtoTesla (fT). More precisely, each&nbsp;magnetometer - i.e. the sensor measuring the z component of the magnetic field in one location - returns Tesla. Each gradiometer - i.e. the sensor measuring the x or y component of the <em>gradient</em> of the magnetic field in one location - returns Tesla/cm. In each of the 102 locations there is one magnetometer and two gradiometers.</p>\n<p>Having said that, it may be convenient for you to save the data in floating point exponential format, with &quot;%e&quot; (or &quot;%g&quot;), instead of &nbsp;the decimal format &quot;%20.17f&quot; that you use. I guess &quot;%1.20e&quot; should be enough. See here for additional details:</p>\n<p>&nbsp;&nbsp;<a href=\"https://docs.python.org/2/library/stdtypes.html#string-formatting\">https://docs.python.org/2/library/stdtypes.html#string-formatting</a></p>\n<p>&nbsp;&nbsp;<a href=\"https://docs.python.org/3/library/stdtypes.html#printf-style-string-formatting\">https://docs.python.org/3/library/stdtypes.html#printf-style-string-formatting</a></p>\n\n<p>Note: the setup of gradiometers/magnetometers and number of locations changes across different MEG devices. The one used for acquiring the data of this competition, i.e. Elekta Neuromag VectorView, is as described above.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43131,
      "author_name": "robstr",
      "author_url": "",
      "post_date": "04/27/2014 08:51:27",
      "content": "<p>Hey,</p>\n<p>thank you really mutch for youre time to answer the question. You make mutch things clear for me.</p>\n<p>But in one lines, the challenge is to predict 2 classes ( Face and Scramble ) from the measure. There is no importans of analyset this https://github.com/mne-tools/mne-python/blob/master/mne/layouts/Vectorview-all.lout file ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43134,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "04/27/2014 10:16:59",
      "content": "<p>Hi,</p>\n<p>The&nbsp;<a href=\"https://github.com/mne-tools/mne-python/blob/master/mne/layouts/Vectorview-all.lout\">Vectorview-all.lout</a> is a text file that describes the approximate 2D position of the sensors with respect to the head of the subject. I used that file and the data from subject 1 to create the&nbsp;<a href=\"http://nilab.cimec.unitn.it/people/olivetti/decmeg2014/MEG_data2.png\">figure</a> in the description page of the competition. It might be useful to know these approximate positions in order to account for spatial correlations between the timeseries. But I have no strong opinion if this information can affect the score of the competition.</p>\n<p>I plan to release a bit more code in the next days, in the Github repository. For example the one I used to create the figure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43157,
      "author_name": "alexstephens",
      "author_url": "",
      "post_date": "04/27/2014 15:28:41",
      "content": "<p>I'm new to this sort of data, so my apologies if these questions can be inferred from the existing data description.&nbsp;</p>\n<ul>\n<li>Are the 306 channels identically ordered in all of the test subjects?&nbsp; E.g., does channel 107 in trial subject 01 correspond to the same sensor-type &amp; location in test subject 17?</li>\n<li>Also, can the layouts in your github repsitory (Vectorview-mag &amp; Vectorview-grad) be used to reliably associate a channel with a specific type of sensor (since there are two types of sensor in use)?</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43162,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "04/27/2014 16:09:08",
      "content": "<p>Hi,</p>\n<p>Thank you for your questions.</p>\n<p>The sensors are in the same order over all 23 subjects. The order of the channels in the dataset is the same as in <a href=\"https://github.com/FBK-NILab/DecMeg2014/blob/master/additional_files/Vectorview-all.lout\">Vectorview-all.lout</a> that is in the Github repository of the competition. The other two files that you mention, which are available from a different repository (specifically <a href=\"https://github.com/mne-tools/mne-python/tree/master/mne/layouts\">here</a>), provide the same information of Vectorview-all.lout, but split in two parts: Vectorview-mag for magnetometers and Vectorview-grad for gradiometers. So, in my opinion, the right way to proceed, is to use Vectorview-all.lout in order to retrieve the name and position of each sensor. For example, sensor &quot;MEG 0113&quot; has x=-73.416206 and y=33.416687 (see the 2nd line of the file). Notice that if the number in the sensor name ends with &quot;1&quot; then it is a magnetometer. If it ends with &quot;2&quot; or &quot;3&quot;, then it is a gradiometer. For example, &quot;MEG 0113&quot; is a gradiometer and &quot;MEG 0111&quot; is a magnetometer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43601,
      "author_name": "gbrdata",
      "author_url": "",
      "post_date": "05/03/2014 12:07:55",
      "content": "<p>Hi Emanuele,</p>\n<p>I would suggest that you add this:</p>\n<p>[quote=Emanuele;43162]</p>\n<p>Notice that if the number in the sensor name ends with &quot;1&quot; then it is a magnetometer. If it ends with &quot;2&quot; or &quot;3&quot;, then it is a gradiometer. For example, &quot;MEG 0113&quot; is a gradiometer and &quot;MEG 0111&quot; is a magnetometer.</p>\n<p>[/quote]</p>\n<p>to the description of the data. It may be useful for everyone (not just people reading &quot;Python preprocess of MAT files&quot;).</p>\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43617,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "05/03/2014 15:40:25",
      "content": "<p>Good point. I'll ask to the Kaggle people to change the data description page. We don't have permission to do that anymore.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 47345,
      "author_name": "jaidevd",
      "author_url": "",
      "post_date": "05/30/2014 16:09:10",
      "content": "<p>Hello,</p>\n<p>What does the first column in Vectorview-all.lout specify?</p>\n<p>Thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "43080": "",
    "43088": "",
    "43131": "",
    "43134": "",
    "43157": "",
    "43162": "",
    "43601": "",
    "43617": "",
    "47345": ""
  },
  "source": "meta"
}