{
  "id": 4992,
  "title": "Starter code for mat file conversion in Python",
  "url": "/competitions/belkin-energy-disaggregation-competition/discussion/4992",
  "author_name": "",
  "post_date": "2013-07-02T02:15:56.690Z",
  "votes": 2,
  "comment_count": 15,
  "views": 10176,
  "content": "<p>One of the admins was kind enough to provide some starter code for Python users. &nbsp;See here:&nbsp;http://pastebin.com/TPKWwWey</p>\r\n<p>(Comes with the usual warnings and caveats that it may not be correct and that you should verify it's doing what you want it to do...)</p>",
  "messages": [
    {
      "id": "26569",
      "postDate": "07/02/2013 02:15:56",
      "content": "<p>One of the admins was kind enough to provide some starter code for Python users. &nbsp;See here:&nbsp;http://pastebin.com/TPKWwWey</p>\r\n<p>(Comes with the usual warnings and caveats that it may not be correct and that you should verify it's doing what you want it to do...)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "26575",
      "postDate": "07/02/2013 04:22:57",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27178",
      "postDate": "07/14/2013 12:18:11",
      "content": "<p>Hello and thanks for the Python code.</p>\r\n<p>I hope my question is not too basic, but which files do I have to feed into the script?</p>\r\n<p>I tried it with the .mat files from the H1 dataset. For the testData I use e.g. Testing_07_09_1341817201.mat what seems to work. But what file do I have to load for the taggingData? I tried all of the 3 data &quot;types&quot; in the folders, but I keep getting some\r\n errors like &quot;KeyError: 'TaggingInfo'&quot; when I use AllTaggingInfo.mat. </p>\r\n<p>When I use &quot;Tagged_Training_04_13_1334300401.mat&quot; there seems to be some crash/error deep within Python, and I just get the output &quot;Memory Error&quot; at the end.</p>\r\n<p>Maybe I miss something really obvious, if you could drop me a line, I would be very happy.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27211",
      "postDate": "07/15/2013 05:23:33",
      "content": "<p>It is possible that you are using the 32-bit version of Python. Please try with 64-bit to resolve the Memory Error. For training datasets you should be loading a file from &quot;Tagged_Training_*.mat&quot;.</p>\r\n<p>You can also look at the sample Matlab code provided to better inform yourself on what the code looks like. The Python code sample is really just a quick starting point and has not been tested the same way as the Matlab code.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27224",
      "postDate": "07/15/2013 11:03:15",
      "content": "<p>Thanks for clarification on which file to use. I encountered, that with the pre-installed Python environments, there seem to be some packages missing when you use them out of the box. Just in case anybody encounters similar problems here are some issues\r\n I found out:</p>\r\n<p>When I use Anaconda (64 bit), I have problems to create/show a plot with matplotlib. The demo code opens an empty window but doesn't plot anything. Moreover when I execute the Python code from pastbin, my memory usage grows to 10GB, then falls down to 1GB,\r\n goes up to 10GB again, ...</p>\r\n<p>I then tried Canopy from Enthougt (32 bit). Plotting worked immediately. The Python code from Pastebin gives me sometimes memory errors (the first few launches work, then I get memory errors) or an error when I want to read in the tagginf file (&quot;KeyError:\r\n 'TaggingInfo&quot;). I dug a little bit deeper and it seems, that one needs to have H5py installed, to properly read the .mat files. Canopy offers this as an easy download only if you have a paid subscription.</p>\r\n<p>I guess to get a well defined environment, I will set up Python with all libraries from scratch.</p>\r\n<p>Maybe this experience helps others :-)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27415",
      "postDate": "07/18/2013 15:04:47",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27744",
      "postDate": "07/29/2013 13:36:06",
      "content": "<p>Hi Konkordan,</p>\n<p>Did you manage to load the files through python? And if so, can you please let us know what version of python and what libraries did you use?</p>\n<p><span style=\"line-height: 1.4\">I tried with python(x,y) 32 bit and WinPython 64 bit (both on Pyhton 2.7) and I got two different errors. For first option was some longer MemoryError, which started with line 9 in the provided script. The second option ended up with KeyError: &quot;Tagging info&quot; on line 42 of the same script.</span></p>\n<p>Also the files that I'm using are:&nbsp;</p>\n<p>testData = io.loadmat('Testing_09_12_1347433201.mat')<br>taggingData = io.loadmat('Tagged_Training_07_26_1343286001.mat')</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27747",
      "postDate": "07/29/2013 13:48:21",
      "content": "<p>I got one of the files loaded. It just took about 35 minutes to complete loading it into a numpy array.</p>\n<p>I had to upgrade online server at linode to a 4Gb machine. It costs 80USD already so i cannot afford more specs. The reason I am using this is because the downloads will take forever in South-Africa. It takes about 4 minutes per file from Kaggle to my Linode.</p>\n<p>I also implemented a 64bit ubuntu version. With that using basic code below, it takes 35 minutes per file. At that rate it could take days just to load data, not even processing.&nbsp;<br><br>Any ideas anyone on how to tackle this problem?</p>\n<p>&nbsp;</p>\n<p><code>from scipy.io import&nbsp;</code><code style=\"font-size: 16px; line-height: 1.4;\">matlab</code></p>\n<p><code style=\"font-size: 16px; line-height: 1.4;\">dummy = matlab.loadmat('taggedxxx.mat')</code></p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27772",
      "postDate": "07/29/2013 16:40:15",
      "content": "<p>Alexandra - regarding the KeyError, the tagged training sets (when loaded with io.loadmat, at least) have their data contained within a top level called &quot;Buffer&quot;. &nbsp;So the assignment in the sample code:</p>\n<p>taggingInfo&nbsp;=&nbsp;taggingData['TaggingInfo']</p>\n<p>won't work for those data sets. &nbsp;If you want to pull tagging data from that data set, use</p>\n<p>taggingInfo&nbsp;=&nbsp;taggingData['Buffer']['TaggingInfo']</p>\n<p>Regarding your memory error - yep, stick with the 64-bit version.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27774",
      "postDate": "07/29/2013 17:22:56",
      "content": "<p>Tobie - sounds like a CPU problem. &nbsp;On my Macbook Pro loading a tagged training set takes 10-15 seconds at the outside. &nbsp;The Linode CPUs (decently-clocked Xeons, IIRC) shouldn't be doing that badly comparatively. &nbsp;May be a priority issue - folks with more expensive Linode accounts get higher-priority access to CPU time, as I recall.</p>\n<p>Is there anything you can do to improve access to CPU time? &nbsp;Try logging on / doing your processing at different times? &nbsp;Also, once you've gotten in and massaged / chopped up your data it should be much smaller, and perhaps you could transfer it to your local machine then.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27778",
      "postDate": "07/29/2013 19:10:24",
      "content": "<p>Phil - I started the whole Linode&nbsp;set-up&nbsp;from scratch. I installed Ubuntu 64bit and then Anaconda 64 bit...now it loads in about 3 to 5 seconds...Whoop!</p>\n<p>&nbsp;</p>\n<p>I think it is save to&nbsp;recommend&nbsp;that if you wanna use Python for analysis it is easier to just get anaconda or EPD! Saves a lot of hasstles.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27779",
      "postDate": "07/29/2013 19:36:39",
      "content": "<p>[quote=Tobie Nortje;27778]</p>\n<p>Phil - I started the whole Linode&nbsp;set-up&nbsp;from scratch. I installed Ubuntu 64bit and then Anaconda 64 bit...now it loads in about 3 to 5 seconds...Whoop!</p>\n<p>[/quote]</p>\n<p>Awesome! &nbsp;Very glad to hear that!</p>\n<p>&nbsp;</p>\n<p>[quote=Tobie Nortje;27778]</p>\n<p>I think it is save to&nbsp;recommend&nbsp;that if you wanna use Python for analysis it is easier to just get anaconda or EPD! Saves a lot of hasstles.</p>\n<p>[/quote]</p>\n<p>Yep! &nbsp;Much easier. &nbsp;There are a heap of options, too.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27780",
      "postDate": "07/29/2013 20:01:17",
      "content": "<p>OK so here is a link to what I have managed to do. I am happy that I can now realy start playing around -</p>\n<p>&nbsp;</p>\n<p><code>&nbsp;http://nbviewer.ipython.org/6107240</code></p>\n<p>&nbsp;</p>\n<p>can anybody give advise on the frequency plot - that was not implemented in the started code</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "27781",
      "postDate": "07/29/2013 20:02:22",
      "content": "<p>http://nbviewer.ipython.org/6107240</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "28824",
      "postDate": "08/16/2013 15:55:50",
      "content": "<p>Hi Alexandra (and others),</p>\n<p>here a short update on the Python environment: I now use the Canopy (=EPD) free 64bit environment.</p>\n<p>When I pasted the demo code and fix the line with the tagging info to</p>\n<p>taggingInfo&nbsp;=&nbsp;taggingData['Buffer']['TaggingInfo']</p>\n<p>I can load the files. I just have to comment out the for loop in the plotting routine, this still doesn't work.</p>\n<p>I tried for several hours to get Python and libs running on a Mac from scratch. First you have to install the Python version from the Python website (the version shipped with MacOS won't work). Then I installed scipy and numpy (pay attention to the install sequence, I forgot it). Long story short: I didn't manage to get it to run, uninstalled everything and took the Canopy stuff.</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "30480",
      "postDate": "09/08/2013 13:45:20",
      "content": "<p>A new version of the&nbsp; pastebin from the frist post:</p>\n<p>http://pastebin.com/BpTNvJBn</p>\n<p>The difference is that I added&nbsp; these parameters to the loadmat function</p>\n<p><code style=\"padding-left: 30px\">struct_as_record=False, squeeze_me=True</code></p>\n<p>&nbsp;This creates numpy arrays which are much easier to use. The initial pastebin created arrays where every element was a single element array, that's why&nbsp; [0][0] was needed all over the code. The above parameters make&nbsp;&nbsp; taggingData['Buffer'] behave like an objet with fields, and the fields are 1D or 2D arrays.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 26575,
      "author_name": "phillipadkins",
      "author_url": "",
      "post_date": "07/02/2013 04:22:57",
      "content": "<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27178,
      "author_name": "konkordan",
      "author_url": "",
      "post_date": "07/14/2013 12:18:11",
      "content": "<p>Hello and thanks for the Python code.</p>\r\n<p>I hope my question is not too basic, but which files do I have to feed into the script?</p>\r\n<p>I tried it with the .mat files from the H1 dataset. For the testData I use e.g. Testing_07_09_1341817201.mat what seems to work. But what file do I have to load for the taggingData? I tried all of the 3 data &quot;types&quot; in the folders, but I keep getting some\r\n errors like &quot;KeyError: 'TaggingInfo'&quot; when I use AllTaggingInfo.mat. </p>\r\n<p>When I use &quot;Tagged_Training_04_13_1334300401.mat&quot; there seems to be some crash/error deep within Python, and I just get the output &quot;Memory Error&quot; at the end.</p>\r\n<p>Maybe I miss something really obvious, if you could drop me a line, I would be very happy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27211,
      "author_name": "sidhantgupta",
      "author_url": "",
      "post_date": "07/15/2013 05:23:33",
      "content": "<p>It is possible that you are using the 32-bit version of Python. Please try with 64-bit to resolve the Memory Error. For training datasets you should be loading a file from &quot;Tagged_Training_*.mat&quot;.</p>\r\n<p>You can also look at the sample Matlab code provided to better inform yourself on what the code looks like. The Python code sample is really just a quick starting point and has not been tested the same way as the Matlab code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27224,
      "author_name": "konkordan",
      "author_url": "",
      "post_date": "07/15/2013 11:03:15",
      "content": "<p>Thanks for clarification on which file to use. I encountered, that with the pre-installed Python environments, there seem to be some packages missing when you use them out of the box. Just in case anybody encounters similar problems here are some issues\r\n I found out:</p>\r\n<p>When I use Anaconda (64 bit), I have problems to create/show a plot with matplotlib. The demo code opens an empty window but doesn't plot anything. Moreover when I execute the Python code from pastbin, my memory usage grows to 10GB, then falls down to 1GB,\r\n goes up to 10GB again, ...</p>\r\n<p>I then tried Canopy from Enthougt (32 bit). Plotting worked immediately. The Python code from Pastebin gives me sometimes memory errors (the first few launches work, then I get memory errors) or an error when I want to read in the tagginf file (&quot;KeyError:\r\n 'TaggingInfo&quot;). I dug a little bit deeper and it seems, that one needs to have H5py installed, to properly read the .mat files. Canopy offers this as an easy download only if you have a paid subscription.</p>\r\n<p>I guess to get a well defined environment, I will set up Python with all libraries from scratch.</p>\r\n<p>Maybe this experience helps others :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27415,
      "author_name": "yzhao0527",
      "author_url": "",
      "post_date": "07/18/2013 15:04:47",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 27744,
      "author_name": "alexandra6",
      "author_url": "",
      "post_date": "07/29/2013 13:36:06",
      "content": "<p>Hi Konkordan,</p>\n<p>Did you manage to load the files through python? And if so, can you please let us know what version of python and what libraries did you use?</p>\n<p><span style=\"line-height: 1.4\">I tried with python(x,y) 32 bit and WinPython 64 bit (both on Pyhton 2.7) and I got two different errors. For first option was some longer MemoryError, which started with line 9 in the provided script. The second option ended up with KeyError: &quot;Tagging info&quot; on line 42 of the same script.</span></p>\n<p>Also the files that I'm using are:&nbsp;</p>\n<p>testData = io.loadmat('Testing_09_12_1347433201.mat')<br>taggingData = io.loadmat('Tagged_Training_07_26_1343286001.mat')</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27747,
      "author_name": "tobienortje",
      "author_url": "",
      "post_date": "07/29/2013 13:48:21",
      "content": "<p>I got one of the files loaded. It just took about 35 minutes to complete loading it into a numpy array.</p>\n<p>I had to upgrade online server at linode to a 4Gb machine. It costs 80USD already so i cannot afford more specs. The reason I am using this is because the downloads will take forever in South-Africa. It takes about 4 minutes per file from Kaggle to my Linode.</p>\n<p>I also implemented a 64bit ubuntu version. With that using basic code below, it takes 35 minutes per file. At that rate it could take days just to load data, not even processing.&nbsp;<br><br>Any ideas anyone on how to tackle this problem?</p>\n<p>&nbsp;</p>\n<p><code>from scipy.io import&nbsp;</code><code style=\"font-size: 16px; line-height: 1.4;\">matlab</code></p>\n<p><code style=\"font-size: 16px; line-height: 1.4;\">dummy = matlab.loadmat('taggedxxx.mat')</code></p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27772,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "07/29/2013 16:40:15",
      "content": "<p>Alexandra - regarding the KeyError, the tagged training sets (when loaded with io.loadmat, at least) have their data contained within a top level called &quot;Buffer&quot;. &nbsp;So the assignment in the sample code:</p>\n<p>taggingInfo&nbsp;=&nbsp;taggingData['TaggingInfo']</p>\n<p>won't work for those data sets. &nbsp;If you want to pull tagging data from that data set, use</p>\n<p>taggingInfo&nbsp;=&nbsp;taggingData['Buffer']['TaggingInfo']</p>\n<p>Regarding your memory error - yep, stick with the 64-bit version.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27774,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "07/29/2013 17:22:56",
      "content": "<p>Tobie - sounds like a CPU problem. &nbsp;On my Macbook Pro loading a tagged training set takes 10-15 seconds at the outside. &nbsp;The Linode CPUs (decently-clocked Xeons, IIRC) shouldn't be doing that badly comparatively. &nbsp;May be a priority issue - folks with more expensive Linode accounts get higher-priority access to CPU time, as I recall.</p>\n<p>Is there anything you can do to improve access to CPU time? &nbsp;Try logging on / doing your processing at different times? &nbsp;Also, once you've gotten in and massaged / chopped up your data it should be much smaller, and perhaps you could transfer it to your local machine then.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27778,
      "author_name": "tobienortje",
      "author_url": "",
      "post_date": "07/29/2013 19:10:24",
      "content": "<p>Phil - I started the whole Linode&nbsp;set-up&nbsp;from scratch. I installed Ubuntu 64bit and then Anaconda 64 bit...now it loads in about 3 to 5 seconds...Whoop!</p>\n<p>&nbsp;</p>\n<p>I think it is save to&nbsp;recommend&nbsp;that if you wanna use Python for analysis it is easier to just get anaconda or EPD! Saves a lot of hasstles.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27779,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "07/29/2013 19:36:39",
      "content": "<p>[quote=Tobie Nortje;27778]</p>\n<p>Phil - I started the whole Linode&nbsp;set-up&nbsp;from scratch. I installed Ubuntu 64bit and then Anaconda 64 bit...now it loads in about 3 to 5 seconds...Whoop!</p>\n<p>[/quote]</p>\n<p>Awesome! &nbsp;Very glad to hear that!</p>\n<p>&nbsp;</p>\n<p>[quote=Tobie Nortje;27778]</p>\n<p>I think it is save to&nbsp;recommend&nbsp;that if you wanna use Python for analysis it is easier to just get anaconda or EPD! Saves a lot of hasstles.</p>\n<p>[/quote]</p>\n<p>Yep! &nbsp;Much easier. &nbsp;There are a heap of options, too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27780,
      "author_name": "tobienortje",
      "author_url": "",
      "post_date": "07/29/2013 20:01:17",
      "content": "<p>OK so here is a link to what I have managed to do. I am happy that I can now realy start playing around -</p>\n<p>&nbsp;</p>\n<p><code>&nbsp;http://nbviewer.ipython.org/6107240</code></p>\n<p>&nbsp;</p>\n<p>can anybody give advise on the frequency plot - that was not implemented in the started code</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 27781,
      "author_name": "tobienortje",
      "author_url": "",
      "post_date": "07/29/2013 20:02:22",
      "content": "<p>http://nbviewer.ipython.org/6107240</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 28824,
      "author_name": "konkordan",
      "author_url": "",
      "post_date": "08/16/2013 15:55:50",
      "content": "<p>Hi Alexandra (and others),</p>\n<p>here a short update on the Python environment: I now use the Canopy (=EPD) free 64bit environment.</p>\n<p>When I pasted the demo code and fix the line with the tagging info to</p>\n<p>taggingInfo&nbsp;=&nbsp;taggingData['Buffer']['TaggingInfo']</p>\n<p>I can load the files. I just have to comment out the for loop in the plotting routine, this still doesn't work.</p>\n<p>I tried for several hours to get Python and libs running on a Mac from scratch. First you have to install the Python version from the Python website (the version shipped with MacOS won't work). Then I installed scipy and numpy (pay attention to the install sequence, I forgot it). Long story short: I didn't manage to get it to run, uninstalled everything and took the Canopy stuff.</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 30480,
      "author_name": "vojtekb",
      "author_url": "",
      "post_date": "09/08/2013 13:45:20",
      "content": "<p>A new version of the&nbsp; pastebin from the frist post:</p>\n<p>http://pastebin.com/BpTNvJBn</p>\n<p>The difference is that I added&nbsp; these parameters to the loadmat function</p>\n<p><code style=\"padding-left: 30px\">struct_as_record=False, squeeze_me=True</code></p>\n<p>&nbsp;This creates numpy arrays which are much easier to use. The initial pastebin created arrays where every element was a single element array, that's why&nbsp; [0][0] was needed all over the code. The above parameters make&nbsp;&nbsp; taggingData['Buffer'] behave like an objet with fields, and the fields are 1D or 2D arrays.</p>\n<p>&nbsp;</p>\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "26569": "",
    "26575": "",
    "27178": "",
    "27211": "",
    "27224": "",
    "27415": "",
    "27744": "",
    "27747": "",
    "27772": "",
    "27774": "",
    "27778": "",
    "27779": "",
    "27780": "",
    "27781": "",
    "28824": "",
    "30480": ""
  },
  "source": "meta"
}