{
  "id": 10319,
  "title": ".mat to .csv",
  "url": "/competitions/seizure-prediction/discussion/10319",
  "author_name": "",
  "post_date": "2014-09-13T13:07:48.227Z",
  "votes": 1,
  "comment_count": 20,
  "views": 7221,
  "content": "<p>Is there a way to convert the .mat files to .csv files? I am not using matlab so I can't read the data.</p>",
  "messages": [
    {
      "id": "53638",
      "postDate": "09/13/2014 13:07:48",
      "content": "<p>Is there a way to convert the .mat files to .csv files? I am not using matlab so I can't read the data.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53658",
      "postDate": "09/13/2014 18:50:39",
      "content": "<p>What are you using?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53659",
      "postDate": "09/13/2014 19:03:09",
      "content": "<p>scipy has a loadmat function that will allow you to load the file, from there you could export to csv if you're not comfortable with python.</p>\n<p>But to be honest you're better off using .mat or .hkl (hickle) because they are binary file types, i think the CSV will be even bigger.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53669",
      "postDate": "09/14/2014 00:37:17",
      "content": "<p>@elyase I am using c++.</p>\n\n<p>@franklyn I know a little bit of python and I'll look into it. Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53670",
      "postDate": "09/14/2014 01:37:45",
      "content": "<p>You might find this library helpful</p>\n\n<p>http://sourceforge.net/projects/matio/</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53676",
      "postDate": "09/14/2014 03:41:37",
      "content": "<p>Thanks franklyn. It made my life a lot easier. I've done a ton of programming in c++ but I barely know the basics of any other language.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53678",
      "postDate": "09/14/2014 03:43:12",
      "content": "<p>what machine learning library do you use in c++ ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53682",
      "postDate": "09/14/2014 08:05:30",
      "content": "<p>I am not using any machine learning library. I am planning on coding everything myself.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53688",
      "postDate": "09/14/2014 12:30:25",
      "content": "<p>You can use a C-library called 'matio' to read the mat file and then, if you want, save it it in usual binary format.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "53745",
      "postDate": "09/15/2014 02:04:53",
      "content": "<p>In the previous competition I was using hickle (wraps h5py), this time around with a much bigger dataset I'm using h5py directly as hickle seemed to behave strangely slow for me when dealing with 1000s of files. FYI h5py uses the hdf5 file format.</p>\n<p>During my testing the original mat format seems to perform the same as hdf5 with regards to speed of loading data. So there seems to be no harm in sticking with mat.</p>\n<p>Based on my experience with csv in the MNIST dataset, you really don't want to use csv. It's going to be bigger and much much slower to load.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54037",
      "postDate": "09/16/2014 16:56:10",
      "content": "<p>Hi everyone,</p>\n<p>I downloaded the matio library, but the documentation is nearly inexistant.</p>\n<p>If someone using this library could give an example of code for the transformation of the .mat into a .csv it would be nice.</p>\n<p>I think I should use the mat_varRead function, but it doesn't work with the &quot;data&quot; variable.</p>\n<p>Which variable name should I give to mat_varRead to get the real data?</p>\n<p>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54058",
      "postDate": "09/16/2014 18:46:03",
      "content": "<p>Hi</p>\n<p>In the example I include as atacchment, the data matrix &nbsp;from&nbsp;&quot;Patient_2_interictal_segment_0001.mat&quot; is read into an array of double.</p>\n<p>Then you can put it into &nbsp;a file in any format (csv or binary) or just use it in your C/C++ functions.</p>\n<p>you can compile it with :</p>\n<p>gcc readdatamatrixfrom_matfile.c -lmatio</p>\n\n<p>I'm using matio version 1.3 . If you use the new version (1.5) the instructions:</p>\n<p>field = Mat_VarGetStructField(structname, &amp;index,BY_INDEX,0);</p>\n<p>have to be replaced by&nbsp;</p>\n<p>field = Mat_VarGetStructFieldByIndex(structname, index,0);</p>\n<p>There an example &nbsp;that helped me here&nbsp;</p>\n<p>http://sourceforge.net/p/matio/discussion/609377/thread/b703ce7a/</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54150",
      "postDate": "09/17/2014 14:12:37",
      "content": "<p>Thank you very much, it's very helpful.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54261",
      "postDate": "09/18/2014 20:35:58",
      "content": "<p>can someone give me an example of data format from the mat file ? I am trying to improve on the example given ?&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54304",
      "postDate": "09/19/2014 15:30:19",
      "content": "<p>Why is everything stored as double if they are just short ints?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54306",
      "postDate": "09/19/2014 16:03:30",
      "content": "<p>'double' is the default type for matlab (*.mat are matlabfiles). So if you want to read the binary mat file you get the original data as 'double'. In 'Dog* ' data the sampling frequency &nbsp;is not an int, it's value is is 399.66 . So I would no risk to convert all the data to 'int' after reading it to 'double'.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54573",
      "postDate": "09/24/2014 04:35:54",
      "content": "<p>[quote=Apple314;53682]</p>\n<p>I am not using any machine learning library. I am planning on coding everything myself.</p>\n<p>[/quote]</p>\n\n<p>Thank you for this.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54574",
      "postDate": "09/24/2014 04:38:56",
      "content": "<p>be nice.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56251",
      "postDate": "10/19/2014 07:36:12",
      "content": "<p>Has anyone used matio with visual C+ 2010 express? I cant even get the thing to compile.</p>\n<p>I was able to use one of the project files that came with it. but then I needed a library called &quot;hdf5&quot; which i downloaded from somewhere, but it didn't come with the header file. This is costing me way too much time, and life would be a lot simpler if kaggle just gave the files in zipped csv format which would be the same size anyway.</p>\n<p>if anyone has a ready to use recipe for me to convert the files to csv, either with matio or another way, I would be really greatful.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "106593",
      "postDate": "02/02/2016 09:30:10",
      "content": "<pre><code>import scipy.io\nimport numpy as np\ndata = scipy.io.loadmat(&quot;file.mat&quot;)\n\nfor i in data:\n     if '__' not in i and 'readme' not in i:\n           np.savetxt((&quot;file.csv&quot;),data[i],delimiter=',')\n</code></pre>",
      "rawMarkdown": "import scipy.io\r\n    import numpy as np\r\n    data = scipy.io.loadmat(\"file.mat\")\r\n\r\n    for i in data:\r\n         if '__' not in i and 'readme' not in i:\r\n               np.savetxt((\"file.csv\"),data[i],delimiter=',')",
      "votes": null
    },
    {
      "id": "748710",
      "postDate": "02/17/2020 21:59:09",
      "content": "<p>This gives the TypeError: only size-1 arrays can be converted to Python scalars</p>",
      "rawMarkdown": "This gives the TypeError: only size-1 arrays can be converted to Python scalars",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 53658,
      "author_name": "elyase",
      "author_url": "",
      "post_date": "09/13/2014 18:50:39",
      "content": "<p>What are you using?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53659,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "09/13/2014 19:03:09",
      "content": "<p>scipy has a loadmat function that will allow you to load the file, from there you could export to csv if you're not comfortable with python.</p>\n<p>But to be honest you're better off using .mat or .hkl (hickle) because they are binary file types, i think the CSV will be even bigger.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53669,
      "author_name": "",
      "author_url": "",
      "post_date": "09/14/2014 00:37:17",
      "content": "<p>@elyase I am using c++.</p>\n\n<p>@franklyn I know a little bit of python and I'll look into it. Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53670,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "09/14/2014 01:37:45",
      "content": "<p>You might find this library helpful</p>\n\n<p>http://sourceforge.net/projects/matio/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53676,
      "author_name": "",
      "author_url": "",
      "post_date": "09/14/2014 03:41:37",
      "content": "<p>Thanks franklyn. It made my life a lot easier. I've done a ton of programming in c++ but I barely know the basics of any other language.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53678,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "09/14/2014 03:43:12",
      "content": "<p>what machine learning library do you use in c++ ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53682,
      "author_name": "",
      "author_url": "",
      "post_date": "09/14/2014 08:05:30",
      "content": "<p>I am not using any machine learning library. I am planning on coding everything myself.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53688,
      "author_name": "ruiprodrigues",
      "author_url": "",
      "post_date": "09/14/2014 12:30:25",
      "content": "<p>You can use a C-library called 'matio' to read the mat file and then, if you want, save it it in usual binary format.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 53745,
      "author_name": "michaelhills",
      "author_url": "",
      "post_date": "09/15/2014 02:04:53",
      "content": "<p>In the previous competition I was using hickle (wraps h5py), this time around with a much bigger dataset I'm using h5py directly as hickle seemed to behave strangely slow for me when dealing with 1000s of files. FYI h5py uses the hdf5 file format.</p>\n<p>During my testing the original mat format seems to perform the same as hdf5 with regards to speed of loading data. So there seems to be no harm in sticking with mat.</p>\n<p>Based on my experience with csv in the MNIST dataset, you really don't want to use csv. It's going to be bigger and much much slower to load.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54037,
      "author_name": "ws45100529",
      "author_url": "",
      "post_date": "09/16/2014 16:56:10",
      "content": "<p>Hi everyone,</p>\n<p>I downloaded the matio library, but the documentation is nearly inexistant.</p>\n<p>If someone using this library could give an example of code for the transformation of the .mat into a .csv it would be nice.</p>\n<p>I think I should use the mat_varRead function, but it doesn't work with the &quot;data&quot; variable.</p>\n<p>Which variable name should I give to mat_varRead to get the real data?</p>\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54058,
      "author_name": "ruiprodrigues",
      "author_url": "",
      "post_date": "09/16/2014 18:46:03",
      "content": "<p>Hi</p>\n<p>In the example I include as atacchment, the data matrix &nbsp;from&nbsp;&quot;Patient_2_interictal_segment_0001.mat&quot; is read into an array of double.</p>\n<p>Then you can put it into &nbsp;a file in any format (csv or binary) or just use it in your C/C++ functions.</p>\n<p>you can compile it with :</p>\n<p>gcc readdatamatrixfrom_matfile.c -lmatio</p>\n\n<p>I'm using matio version 1.3 . If you use the new version (1.5) the instructions:</p>\n<p>field = Mat_VarGetStructField(structname, &amp;index,BY_INDEX,0);</p>\n<p>have to be replaced by&nbsp;</p>\n<p>field = Mat_VarGetStructFieldByIndex(structname, index,0);</p>\n<p>There an example &nbsp;that helped me here&nbsp;</p>\n<p>http://sourceforge.net/p/matio/discussion/609377/thread/b703ce7a/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54150,
      "author_name": "ws45100529",
      "author_url": "",
      "post_date": "09/17/2014 14:12:37",
      "content": "<p>Thank you very much, it's very helpful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54261,
      "author_name": "chamarts",
      "author_url": "",
      "post_date": "09/18/2014 20:35:58",
      "content": "<p>can someone give me an example of data format from the mat file ? I am trying to improve on the example given ?&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54304,
      "author_name": "",
      "author_url": "",
      "post_date": "09/19/2014 15:30:19",
      "content": "<p>Why is everything stored as double if they are just short ints?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54306,
      "author_name": "ruiprodrigues",
      "author_url": "",
      "post_date": "09/19/2014 16:03:30",
      "content": "<p>'double' is the default type for matlab (*.mat are matlabfiles). So if you want to read the binary mat file you get the original data as 'double'. In 'Dog* ' data the sampling frequency &nbsp;is not an int, it's value is is 399.66 . So I would no risk to convert all the data to 'int' after reading it to 'double'.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54573,
      "author_name": "frederickhadley",
      "author_url": "",
      "post_date": "09/24/2014 04:35:54",
      "content": "<p>[quote=Apple314;53682]</p>\n<p>I am not using any machine learning library. I am planning on coding everything myself.</p>\n<p>[/quote]</p>\n\n<p>Thank you for this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54574,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "09/24/2014 04:38:56",
      "content": "<p>be nice.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56251,
      "author_name": "jossmit",
      "author_url": "",
      "post_date": "10/19/2014 07:36:12",
      "content": "<p>Has anyone used matio with visual C+ 2010 express? I cant even get the thing to compile.</p>\n<p>I was able to use one of the project files that came with it. but then I needed a library called &quot;hdf5&quot; which i downloaded from somewhere, but it didn't come with the header file. This is costing me way too much time, and life would be a lot simpler if kaggle just gave the files in zipped csv format which would be the same size anyway.</p>\n<p>if anyone has a ready to use recipe for me to convert the files to csv, either with matio or another way, I would be really greatful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 106593,
      "author_name": "kaushalaman",
      "author_url": "",
      "post_date": "02/02/2016 09:30:10",
      "content": "<pre><code>import scipy.io\nimport numpy as np\ndata = scipy.io.loadmat(&quot;file.mat&quot;)\n\nfor i in data:\n     if '__' not in i and 'readme' not in i:\n           np.savetxt((&quot;file.csv&quot;),data[i],delimiter=',')\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 748710,
          "author_name": "shrutikaggle",
          "author_url": "",
          "post_date": "02/17/2020 21:59:09",
          "content": "<p>This gives the TypeError: only size-1 arrays can be converted to Python scalars</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "53638": "",
    "53658": "",
    "53659": "",
    "53669": "",
    "53670": "",
    "53676": "",
    "53678": "",
    "53682": "",
    "53688": "",
    "53745": "",
    "54037": "",
    "54058": "",
    "54150": "",
    "54261": "",
    "54304": "",
    "54306": "",
    "54573": "",
    "54574": "",
    "56251": "",
    "106593": "import scipy.io\r\n    import numpy as np\r\n    data = scipy.io.loadmat(\"file.mat\")\r\n\r\n    for i in data:\r\n         if '__' not in i and 'readme' not in i:\r\n               np.savetxt((\"file.csv\"),data[i],delimiter=',')",
    "748710": "This gives the TypeError: only size-1 arrays can be converted to Python scalars"
  },
  "source": "meta"
}