{
  "id": 7899,
  "title": "MAT to CSV converter",
  "url": "/competitions/decoding-the-human-brain/discussion/7899",
  "author_name": "",
  "post_date": "2014-04-27T17:28:44.650Z",
  "votes": 9,
  "comment_count": 4,
  "views": 6297,
  "content": "<p>Hi,</p>\n<p>I prepared a converter to create a comma-separated value (CSV) file from each of the original Matlab v5 (MAT) files of the competition.The converter is in Python (v2.7.x) and requires only SciPy to be installed in order to work:</p>\n<p><a href=\"https://github.com/FBK-NILab/DecMeg2014/blob/master/python/mat2csv.py\">https://github.com/FBK-NILab/DecMeg2014/blob/master/python/mat2csv.py</a></p>\n<p>Just issuing &quot;python mat2csv.py&quot;, or executing mat2csv.py, is enough to convert all files, provided that the variables mat_dir and csv_dir are correct in your setting. Please have a look to the code - it should be pretty easy to understand and to adapt.</p>\n<p>The format of the CSV file is straightforward: each line of the CSV file is a trial. For train data, the first value is the class-label of the trial (y), i.e. the category of the stimulus presented to the subject (1=Face, 0=Scramble). For test data, the first value is the Id of the trial. In both cases, the remaining values in the line are the MEG values (X) of the multiple timeseries of that trial stored sequentially. Each timeseries consists of 375 timepoints, i.e. 1.5sec recorded at 250Hz. The groups of 375 values, one for each of the 306 channel/sensor, are stored sequentially, with the channels/sensors in the same order as in the MAT file. For this reason, each line consists of (1 + 375 x 306) = 114751 values.</p>\n<p>I think that it is not possible to upload the CSV (or compressed CSV) files in Kaggle, because they are too large. Consider that each CSV file is approximately 2Gb. In my preliminary tests, the gzip compressor shrinks each CSV file to approximately 900Mb. For comparison, the corresponding MAT file is 260Mb only. Part of the reason for this notable difference is the txt vs. binary format. Another reason is that, in MAT files, values are stored as float32, while in CSV files I used put 20-digits IEEE exponential format - which may be too conservative.</p>\n<p>So, in my opinion, if you need CSV data, it is better to download the MAT files distributed with the competition, and then to transform them in CSV format on your local disk with this converter. The only difficulty is to run the converter, i.e. to have Python and SciPy installed.</p>\n<p>Note: the converter is not extensively tested, so if you find a bug or want to propose an enhancement, you are warmly invited to contact me or to send pull requests through Gihub.</p>",
  "messages": [
    {
      "id": "43169",
      "postDate": "04/27/2014 17:28:44",
      "content": "<p>Hi,</p>\n<p>I prepared a converter to create a comma-separated value (CSV) file from each of the original Matlab v5 (MAT) files of the competition.The converter is in Python (v2.7.x) and requires only SciPy to be installed in order to work:</p>\n<p><a href=\"https://github.com/FBK-NILab/DecMeg2014/blob/master/python/mat2csv.py\">https://github.com/FBK-NILab/DecMeg2014/blob/master/python/mat2csv.py</a></p>\n<p>Just issuing &quot;python mat2csv.py&quot;, or executing mat2csv.py, is enough to convert all files, provided that the variables mat_dir and csv_dir are correct in your setting. Please have a look to the code - it should be pretty easy to understand and to adapt.</p>\n<p>The format of the CSV file is straightforward: each line of the CSV file is a trial. For train data, the first value is the class-label of the trial (y), i.e. the category of the stimulus presented to the subject (1=Face, 0=Scramble). For test data, the first value is the Id of the trial. In both cases, the remaining values in the line are the MEG values (X) of the multiple timeseries of that trial stored sequentially. Each timeseries consists of 375 timepoints, i.e. 1.5sec recorded at 250Hz. The groups of 375 values, one for each of the 306 channel/sensor, are stored sequentially, with the channels/sensors in the same order as in the MAT file. For this reason, each line consists of (1 + 375 x 306) = 114751 values.</p>\n<p>I think that it is not possible to upload the CSV (or compressed CSV) files in Kaggle, because they are too large. Consider that each CSV file is approximately 2Gb. In my preliminary tests, the gzip compressor shrinks each CSV file to approximately 900Mb. For comparison, the corresponding MAT file is 260Mb only. Part of the reason for this notable difference is the txt vs. binary format. Another reason is that, in MAT files, values are stored as float32, while in CSV files I used put 20-digits IEEE exponential format - which may be too conservative.</p>\n<p>So, in my opinion, if you need CSV data, it is better to download the MAT files distributed with the competition, and then to transform them in CSV format on your local disk with this converter. The only difficulty is to run the converter, i.e. to have Python and SciPy installed.</p>\n<p>Note: the converter is not extensively tested, so if you find a bug or want to propose an enhancement, you are warmly invited to contact me or to send pull requests through Gihub.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44343",
      "postDate": "05/10/2014 14:17:35",
      "content": "<p>Since I spent some essential time to process mat-file in pure C++, I put here potentially useful information.</p>\n<p>The main hint is that Matlab uses unusual (for C users) order for walking through indices: first index is incremented first.</p>\n<p>First two attached plots are wrong because of mistake on indices walking order, but the last two plots look right.</p>\n<p>For subject=01, trial=100, sensor=MEG1731 (195-th sensor)&nbsp; I got the values: 5.82E-13, 5.39E-13, 1.66E-13, 8.88E-14, 3.22E-13, etc...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44359",
      "postDate": "05/11/2014 05:29:14",
      "content": "<p>Since I have only 2 Gb RAM on my PC I unable to reproduce the benchmark. Emanuele, could you upload the array of coefficients obtained by logistic regression?</p>\n<p>The last plot (with Z-transformation) in previous post looks to be essentially different from the plot given in description. Has anybody reproduced such plots?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44399",
      "postDate": "05/12/2014 10:24:17",
      "content": "<p>Hi,</p>\n<p>If you have only 2Gb of RAM you may want to have a look to&nbsp;<a href=\"https://www.kaggle.com/c/decoding-the-human-brain/forums/t/7805/beating-the-benchmark-with-hinge-loss-0-66100\">this thread</a>&nbsp;about using very efficient online learning.</p>\n<p>Notice that the timeseries shown in the <a href=\"http://nilab.cimec.unitn.it/people/olivetti/decmeg2014/MEG_data2.png\">figure of the description page</a>, are <em>averages</em> over all trials of each&nbsp;class (blue = Face, red = Scrambled Face). This is a common&nbsp;way to show the recordings in MEG data analysis. Moreover I picked the sensors for which this average difference is clearly visible, for didactical purpose. If you look at single trials, as you did, the noise covers almost everything of the signal of the mental process of interest. Moreover, in many locations the signal is too low to be measured (= there is no signal at all).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "875904",
      "postDate": "06/06/2020 08:52:44",
      "content": "<p><a href=\"https://www.kaggle.com/vinayak123tyagi/mat-to-csv-code\">https://www.kaggle.com/vinayak123tyagi/mat-to-csv-code</a></p>",
      "rawMarkdown": "https://www.kaggle.com/vinayak123tyagi/mat-to-csv-code",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 44343,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "05/10/2014 14:17:35",
      "content": "<p>Since I spent some essential time to process mat-file in pure C++, I put here potentially useful information.</p>\n<p>The main hint is that Matlab uses unusual (for C users) order for walking through indices: first index is incremented first.</p>\n<p>First two attached plots are wrong because of mistake on indices walking order, but the last two plots look right.</p>\n<p>For subject=01, trial=100, sensor=MEG1731 (195-th sensor)&nbsp; I got the values: 5.82E-13, 5.39E-13, 1.66E-13, 8.88E-14, 3.22E-13, etc...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44359,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "05/11/2014 05:29:14",
      "content": "<p>Since I have only 2 Gb RAM on my PC I unable to reproduce the benchmark. Emanuele, could you upload the array of coefficients obtained by logistic regression?</p>\n<p>The last plot (with Z-transformation) in previous post looks to be essentially different from the plot given in description. Has anybody reproduced such plots?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44399,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "05/12/2014 10:24:17",
      "content": "<p>Hi,</p>\n<p>If you have only 2Gb of RAM you may want to have a look to&nbsp;<a href=\"https://www.kaggle.com/c/decoding-the-human-brain/forums/t/7805/beating-the-benchmark-with-hinge-loss-0-66100\">this thread</a>&nbsp;about using very efficient online learning.</p>\n<p>Notice that the timeseries shown in the <a href=\"http://nilab.cimec.unitn.it/people/olivetti/decmeg2014/MEG_data2.png\">figure of the description page</a>, are <em>averages</em> over all trials of each&nbsp;class (blue = Face, red = Scrambled Face). This is a common&nbsp;way to show the recordings in MEG data analysis. Moreover I picked the sensors for which this average difference is clearly visible, for didactical purpose. If you look at single trials, as you did, the noise covers almost everything of the signal of the mental process of interest. Moreover, in many locations the signal is too low to be measured (= there is no signal at all).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 875904,
      "author_name": "vinayak123tyagi",
      "author_url": "",
      "post_date": "06/06/2020 08:52:44",
      "content": "<p><a href=\"https://www.kaggle.com/vinayak123tyagi/mat-to-csv-code\">https://www.kaggle.com/vinayak123tyagi/mat-to-csv-code</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "43169": "",
    "44343": "",
    "44359": "",
    "44399": "",
    "875904": "https://www.kaggle.com/vinayak123tyagi/mat-to-csv-code"
  },
  "source": "meta"
}