{
  "id": 8227,
  "title": "Contents clips.tar unclear",
  "url": "/competitions/seizure-detection/discussion/8227",
  "author_name": "",
  "post_date": "2014-05-20T08:44:15.610Z",
  "votes": null,
  "comment_count": 8,
  "views": 2083,
  "content": "<p>Hi all! First off, great competition! Thank you all for organizing.</p>\n<p>My question concerns the contents of clips.gz</p>\n<p>I downloaded the 10GB file, but unzipping it gives a single huge file, with no extension, named &quot;clips&quot;. The first line of this file is:</p>\n<p>&quot;Volumes/Seagate/seizure_detection/competition_data/clips/&quot;</p>\n<p>Followed by some Null-pointers and binary. Next line is:</p>\n<p>O&#175;T%&#192; 4&#8364;&#183;@&#162;[&#192; \b&#206;...</p>\n<p>I am in the slow process of unzipping the entire file (I ran out of HD-space, so need to unzip to network drive). Maybe if you rename the file to .tar it is another volume, but that seems less likely and a bit counter-intuitive. I'll have to try renaming to .tar when it is done unzipping.</p>\n<p>Perhaps something went wrong with this dataset? As it is 10GB+ of a download, this could be somewhat costly. I do not wish to cry wolf, and create a scare for the competition organizers, but if this could be checked by someone else and confirmed to be a problem, then we know for sure that some action is required.</p>\n<p>I am not quite sure how to get to the dataset contents in an other way.&nbsp;</p>\n<p>Edit: renaming a partial unzip of the file &quot;clips&quot; to &quot;clips.tar&quot; did give me another volume to open. In it were folders and .mat files. Folder structure is:</p>\n<p>&quot;\\Volumes\\Seagate\\seizure_detection\\competition_data\\clips\\Patient_8&quot;</p>\n<p>will need to wait another 12 or so hours before I have it all uncompressed :D, but it seems to be working with a few minor changes.</p>",
  "messages": [
    {
      "id": "44917",
      "postDate": "05/20/2014 08:44:15",
      "content": "<p>Hi all! First off, great competition! Thank you all for organizing.</p>\n<p>My question concerns the contents of clips.gz</p>\n<p>I downloaded the 10GB file, but unzipping it gives a single huge file, with no extension, named &quot;clips&quot;. The first line of this file is:</p>\n<p>&quot;Volumes/Seagate/seizure_detection/competition_data/clips/&quot;</p>\n<p>Followed by some Null-pointers and binary. Next line is:</p>\n<p>O&#175;T%&#192; 4&#8364;&#183;@&#162;[&#192; \b&#206;...</p>\n<p>I am in the slow process of unzipping the entire file (I ran out of HD-space, so need to unzip to network drive). Maybe if you rename the file to .tar it is another volume, but that seems less likely and a bit counter-intuitive. I'll have to try renaming to .tar when it is done unzipping.</p>\n<p>Perhaps something went wrong with this dataset? As it is 10GB+ of a download, this could be somewhat costly. I do not wish to cry wolf, and create a scare for the competition organizers, but if this could be checked by someone else and confirmed to be a problem, then we know for sure that some action is required.</p>\n<p>I am not quite sure how to get to the dataset contents in an other way.&nbsp;</p>\n<p>Edit: renaming a partial unzip of the file &quot;clips&quot; to &quot;clips.tar&quot; did give me another volume to open. In it were folders and .mat files. Folder structure is:</p>\n<p>&quot;\\Volumes\\Seagate\\seizure_detection\\competition_data\\clips\\Patient_8&quot;</p>\n<p>will need to wait another 12 or so hours before I have it all uncompressed :D, but it seems to be working with a few minor changes.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44926",
      "postDate": "05/20/2014 09:26:56",
      "content": "<p>That's weird, it worked fine for me. Shouldn't the file suffix be .tar.gz, as it originally is?</p>\n<p>The folder structure is: /clips/Volumes/Seagate/seizure_detection/competition_data/clips/ and then 12 subfolders, Dog_i (i==1...4) and Patient_i(i==i...8).&nbsp; In total: 58837 files, 48.4 GB</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44937",
      "postDate": "05/20/2014 12:40:30",
      "content": "<p>Thanks Triskelion and Herra.</p>\n<p>Herra's post is correct - the data bundle should have the tar.gz suffix when downloaded. It's possible your browser or ftp client could strip off the suffix, in which case manually replacing it should fix the problem. Inside the tarball there should be 12 folders - 4 dogs and 8 patients. The data was compressed using gzip 1.3.12 on a BSD platform.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44962",
      "postDate": "05/20/2014 16:54:59",
      "content": "<p>Same happens to me. I am trying to unzip in Windows 7 using 7-zip. I obtain a big file of ~ 40GB, but not the directory structure.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44964",
      "postDate": "05/20/2014 17:00:21",
      "content": "<p><span style=\"line-height: 1.4\">Have you tried renaming the big file to a .tar?</span></p>\n<p><span style=\"line-height: 1.4\">Using the command line (gunzip, tar) should work as well.</span></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44991",
      "postDate": "05/21/2014 01:17:54",
      "content": "<p>gunzip and tar works inside Cygwin on Windows.</p>\n<p><a href=\"http://tartool.codeplex.com/releases/view/31505\">This program</a>&nbsp;can unzip on Windows without Cygwin (no huge file without extension is created, no need to unzip twice).</p>\n<p><code>TarTool clips.gz data<br></code></p>\n<p>tar.gz and Windows don't seem to play so nicely together, or maybe I am missing something, anyway I can start with this competition soon!</p>\n<p>Edit: Rename the downloaded file from clips.gz to clips.tar.gz and it will unpack fine.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "46113",
      "postDate": "05/23/2014 08:15:06",
      "content": "<p>from scipy import io</p>\n<p>data = io.loadmat(&quot;test/Patient_1_ictal_segment_1.mat&quot;)</p>\n<p>print data</p>\n<p>{'latency': array([ 0.]), '__header__': 'MATLAB 5.0 MAT-file, Platform: unix, Software: R v3.0.2, Created on: Mon Apr 7 22:16:14 2014', '__globals__': [], 'channels': array([[ ([u'LFG1'], [u'LFG10'], [u'LFG11'], [u'LFG12'], [u'LFG13'], [u'LFG14'], [u'LFG15'], [u'LFG16'], [u'LFG17'], [u'LFG18'], [u'LFG19'], [u'LFG2'], [u'LFG20'], [u'LFG21'], [u'LFG22'], [u'LFG23'], [u'LFG24'], [u'LFG25'], [u'LFG26'], [u'LFG27'], [u'LFG28'], [u'LFG29'], [u'LFG3'], [u'LFG30'], [u'LFG31'], [u'LFG32'], [u'LFG33'], [u'LFG34'], [u'LFG35'], [u'LFG36'], [u'LFG37'], [u'LFG38'], [u'LFG39'], [u'LFG4'], [u'LFG40'], [u'LFG41'], [u'LFG42'], [u'LFG43'], [u'LFG44'], [u'LFG45'], [u'LFG46'], [u'LFG47'], [u'LFG48'], [u'LFG49'], [u'LFG5'], [u'LFG50'], [u'LFG51'], [u'LFG52'], [u'LFG53'], [u'LFG54'], [u'LFG55'], [u'LFG56'], [u'LFG57'], [u'LFG58'], [u'LFG59'], [u'LFG6'], [u'LFG60'], [u'LFG61'], [u'LFG62'], [u'LFG63'], [u'LFG64'], [u'LFG7'], [u'LFG8'], [u'LFG9'], [u'LFS1'], [u'LFS2'], [u'LFS3'], [u'LFS4'])]], <br> dtype=[('X.LFG1.', 'O'), ('X.LFG10.', 'O'), ('X.LFG11.', 'O'), ('X.LFG12.', 'O'), ('X.LFG13.', 'O'), ('X.LFG14.', 'O'), ('X.LFG15.', 'O'), ('X.LFG16.', 'O'), ('X.LFG17.', 'O'), ('X.LFG18.', 'O'), ('X.LFG19.', 'O'), ('X.LFG2.', 'O'), ('X.LFG20.', 'O'), ('X.LFG21.', 'O'), ('X.LFG22.', 'O'), ('X.LFG23.', 'O'), ('X.LFG24.', 'O'), ('X.LFG25.', 'O'), ('X.LFG26.', 'O'), ('X.LFG27.', 'O'), ('X.LFG28.', 'O'), ('X.LFG29.', 'O'), ('X.LFG3.', 'O'), ('X.LFG30.', 'O'), ('X.LFG31.', 'O'), ('X.LFG32.', 'O'), ('X.LFG33.', 'O'), ('X.LFG34.', 'O'), ('X.LFG35.', 'O'), ('X.LFG36.', 'O'), ('X.LFG37.', 'O'), ('X.LFG38.', 'O'), ('X.LFG39.', 'O'), ('X.LFG4.', 'O'), ('X.LFG40.', 'O'), ('X.LFG41.', 'O'), ('X.LFG42.', 'O'), ('X.LFG43.', 'O'), ('X.LFG44.', 'O'), ('X.LFG45.', 'O'), ('X.LFG46.', 'O'), ('X.LFG47.', 'O'), ('X.LFG48.', 'O'), ('X.LFG49.', 'O'), ('X.LFG5.', 'O'), ('X.LFG50.', 'O'), ('X.LFG51.', 'O'), ('X.LFG52.', 'O'), ('X.LFG53.', 'O'), ('X.LFG54.', 'O'), ('X.LFG55.', 'O'), ('X.LFG56.', 'O'), ('X.LFG57.', 'O'), ('X.LFG58.', 'O'), ('X.LFG59.', 'O'), ('X.LFG6.', 'O'), ('X.LFG60.', 'O'), ('X.LFG61.', 'O'), ('X.LFG62.', 'O'), ('X.LFG63.', 'O'), ('X.LFG64.', 'O'), ('X.LFG7.', 'O'), ('X.LFG8.', 'O'), ('X.LFG9.', 'O'), ('X.LFS1.', 'O'), ('X.LFS2.', 'O'), ('X.LFS3.', 'O'), ('X.LFS4.', 'O')]), 'freq': array([ 499.906994]), '__version__': '1.0', 'data': array([[ 204.984, 213.984, 217.984, ..., -195.016, -216.016, -214.016],<br> [-138.314, -127.314, -111.314, ..., 17.686, -17.314, -40.314],<br> [ -64.03 , -57.03 , -28.03 , ..., -55.03 , -83.03 , -66.03 ],<br> ..., <br> [-281.344, -261.344, -232.344, ..., 90.656, 97.656, 112.656],<br> [ 70.398, 99.398, 117.398, ..., -58.602, -44.602, -1.602],<br> [-112.468, -116.468, -123.468, ..., -142.468, -147.468, -139.468]])}</p>\n<p>As a newbie &quot;what is this&quot;? Don&#8217;t the ellipsis mean the continuation of a pattern? Or is this just the readings from 36 sensors at 1 time interval? And how about the stuff at the top? Looks like you don&#8217;t have much to worry about from me, if I can&#8217;t even find the pattern of what the data represents. LOL</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "46124",
      "postDate": "05/23/2014 13:44:57",
      "content": "<p>[quote=Kevin Barbian;46113]</p>\n<p>Don&#8217;t the ellipsis mean the continuation of a pattern? Or is this just the readings from 36 sensors at 1 time interval? And how about the stuff at the top?</p>\n<p>[/quote]</p>\n<p>Those are Numpy arrays. Numpy arrays usually don't get printed fully (as they can be quite large). The ellipsis is to denote columns or rows that are not shown (truncated). To print the entire array see:&nbsp;http://docs.scipy.org/doc/numpy/reference/generated/numpy.set_printoptions.html</p>\n<p><code>import numpy<br>numpy.set_printoptions(threshold=numpy.nan)</code></p>\n<p>What you get in Python after loadmat is a dictionary. This dictionary has information on the data:</p>\n<p><code>print data['data']</code></p>\n<p>But also on metadata, like the headers:</p>\n<p><code>print data['__headers__']</code></p>\n<p>I believe the data is [channel x bins], so this could be 16 x 5000 for human subjects. Try it with:</p>\n<p><code>print data['data'].shape</code></p>\n<p>To familiarize yourself with the .mat files loaded into a dictionary:</p>\n<p><code>for key, value in data.items():</code><code>&nbsp; </code></p>\n<p><code>&nbsp; print key, &quot;\\n&quot;, value, &quot;\\n\\n&quot;</code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "51549",
      "postDate": "08/03/2014 13:57:41",
      "content": "<p>Internet Explorer was [slightly] slower than Google Chrome in downloading the clips.gz file on a Windows 7 laptop.</p>\n<p>Extraction of the clips.gz file using WinRAR program gave a huge no-extension file. After renaming the clips.gz file as clips.tar.gz, the extraction gave several .mat files.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 44926,
      "author_name": "herrahuu",
      "author_url": "",
      "post_date": "05/20/2014 09:26:56",
      "content": "<p>That's weird, it worked fine for me. Shouldn't the file suffix be .tar.gz, as it originally is?</p>\n<p>The folder structure is: /clips/Volumes/Seagate/seizure_detection/competition_data/clips/ and then 12 subfolders, Dog_i (i==1...4) and Patient_i(i==i...8).&nbsp; In total: 58837 files, 48.4 GB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44937,
      "author_name": "bbrinkm",
      "author_url": "",
      "post_date": "05/20/2014 12:40:30",
      "content": "<p>Thanks Triskelion and Herra.</p>\n<p>Herra's post is correct - the data bundle should have the tar.gz suffix when downloaded. It's possible your browser or ftp client could strip off the suffix, in which case manually replacing it should fix the problem. Inside the tarball there should be 12 folders - 4 dogs and 8 patients. The data was compressed using gzip 1.3.12 on a BSD platform.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44962,
      "author_name": "mlopezm",
      "author_url": "",
      "post_date": "05/20/2014 16:54:59",
      "content": "<p>Same happens to me. I am trying to unzip in Windows 7 using 7-zip. I obtain a big file of ~ 40GB, but not the directory structure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44964,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "05/20/2014 17:00:21",
      "content": "<p><span style=\"line-height: 1.4\">Have you tried renaming the big file to a .tar?</span></p>\n<p><span style=\"line-height: 1.4\">Using the command line (gunzip, tar) should work as well.</span></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44991,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "05/21/2014 01:17:54",
      "content": "<p>gunzip and tar works inside Cygwin on Windows.</p>\n<p><a href=\"http://tartool.codeplex.com/releases/view/31505\">This program</a>&nbsp;can unzip on Windows without Cygwin (no huge file without extension is created, no need to unzip twice).</p>\n<p><code>TarTool clips.gz data<br></code></p>\n<p>tar.gz and Windows don't seem to play so nicely together, or maybe I am missing something, anyway I can start with this competition soon!</p>\n<p>Edit: Rename the downloaded file from clips.gz to clips.tar.gz and it will unpack fine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 46113,
      "author_name": "kevinbarbian1",
      "author_url": "",
      "post_date": "05/23/2014 08:15:06",
      "content": "<p>from scipy import io</p>\n<p>data = io.loadmat(&quot;test/Patient_1_ictal_segment_1.mat&quot;)</p>\n<p>print data</p>\n<p>{'latency': array([ 0.]), '__header__': 'MATLAB 5.0 MAT-file, Platform: unix, Software: R v3.0.2, Created on: Mon Apr 7 22:16:14 2014', '__globals__': [], 'channels': array([[ ([u'LFG1'], [u'LFG10'], [u'LFG11'], [u'LFG12'], [u'LFG13'], [u'LFG14'], [u'LFG15'], [u'LFG16'], [u'LFG17'], [u'LFG18'], [u'LFG19'], [u'LFG2'], [u'LFG20'], [u'LFG21'], [u'LFG22'], [u'LFG23'], [u'LFG24'], [u'LFG25'], [u'LFG26'], [u'LFG27'], [u'LFG28'], [u'LFG29'], [u'LFG3'], [u'LFG30'], [u'LFG31'], [u'LFG32'], [u'LFG33'], [u'LFG34'], [u'LFG35'], [u'LFG36'], [u'LFG37'], [u'LFG38'], [u'LFG39'], [u'LFG4'], [u'LFG40'], [u'LFG41'], [u'LFG42'], [u'LFG43'], [u'LFG44'], [u'LFG45'], [u'LFG46'], [u'LFG47'], [u'LFG48'], [u'LFG49'], [u'LFG5'], [u'LFG50'], [u'LFG51'], [u'LFG52'], [u'LFG53'], [u'LFG54'], [u'LFG55'], [u'LFG56'], [u'LFG57'], [u'LFG58'], [u'LFG59'], [u'LFG6'], [u'LFG60'], [u'LFG61'], [u'LFG62'], [u'LFG63'], [u'LFG64'], [u'LFG7'], [u'LFG8'], [u'LFG9'], [u'LFS1'], [u'LFS2'], [u'LFS3'], [u'LFS4'])]], <br> dtype=[('X.LFG1.', 'O'), ('X.LFG10.', 'O'), ('X.LFG11.', 'O'), ('X.LFG12.', 'O'), ('X.LFG13.', 'O'), ('X.LFG14.', 'O'), ('X.LFG15.', 'O'), ('X.LFG16.', 'O'), ('X.LFG17.', 'O'), ('X.LFG18.', 'O'), ('X.LFG19.', 'O'), ('X.LFG2.', 'O'), ('X.LFG20.', 'O'), ('X.LFG21.', 'O'), ('X.LFG22.', 'O'), ('X.LFG23.', 'O'), ('X.LFG24.', 'O'), ('X.LFG25.', 'O'), ('X.LFG26.', 'O'), ('X.LFG27.', 'O'), ('X.LFG28.', 'O'), ('X.LFG29.', 'O'), ('X.LFG3.', 'O'), ('X.LFG30.', 'O'), ('X.LFG31.', 'O'), ('X.LFG32.', 'O'), ('X.LFG33.', 'O'), ('X.LFG34.', 'O'), ('X.LFG35.', 'O'), ('X.LFG36.', 'O'), ('X.LFG37.', 'O'), ('X.LFG38.', 'O'), ('X.LFG39.', 'O'), ('X.LFG4.', 'O'), ('X.LFG40.', 'O'), ('X.LFG41.', 'O'), ('X.LFG42.', 'O'), ('X.LFG43.', 'O'), ('X.LFG44.', 'O'), ('X.LFG45.', 'O'), ('X.LFG46.', 'O'), ('X.LFG47.', 'O'), ('X.LFG48.', 'O'), ('X.LFG49.', 'O'), ('X.LFG5.', 'O'), ('X.LFG50.', 'O'), ('X.LFG51.', 'O'), ('X.LFG52.', 'O'), ('X.LFG53.', 'O'), ('X.LFG54.', 'O'), ('X.LFG55.', 'O'), ('X.LFG56.', 'O'), ('X.LFG57.', 'O'), ('X.LFG58.', 'O'), ('X.LFG59.', 'O'), ('X.LFG6.', 'O'), ('X.LFG60.', 'O'), ('X.LFG61.', 'O'), ('X.LFG62.', 'O'), ('X.LFG63.', 'O'), ('X.LFG64.', 'O'), ('X.LFG7.', 'O'), ('X.LFG8.', 'O'), ('X.LFG9.', 'O'), ('X.LFS1.', 'O'), ('X.LFS2.', 'O'), ('X.LFS3.', 'O'), ('X.LFS4.', 'O')]), 'freq': array([ 499.906994]), '__version__': '1.0', 'data': array([[ 204.984, 213.984, 217.984, ..., -195.016, -216.016, -214.016],<br> [-138.314, -127.314, -111.314, ..., 17.686, -17.314, -40.314],<br> [ -64.03 , -57.03 , -28.03 , ..., -55.03 , -83.03 , -66.03 ],<br> ..., <br> [-281.344, -261.344, -232.344, ..., 90.656, 97.656, 112.656],<br> [ 70.398, 99.398, 117.398, ..., -58.602, -44.602, -1.602],<br> [-112.468, -116.468, -123.468, ..., -142.468, -147.468, -139.468]])}</p>\n<p>As a newbie &quot;what is this&quot;? Don&#8217;t the ellipsis mean the continuation of a pattern? Or is this just the readings from 36 sensors at 1 time interval? And how about the stuff at the top? Looks like you don&#8217;t have much to worry about from me, if I can&#8217;t even find the pattern of what the data represents. LOL</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 46124,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "05/23/2014 13:44:57",
      "content": "<p>[quote=Kevin Barbian;46113]</p>\n<p>Don&#8217;t the ellipsis mean the continuation of a pattern? Or is this just the readings from 36 sensors at 1 time interval? And how about the stuff at the top?</p>\n<p>[/quote]</p>\n<p>Those are Numpy arrays. Numpy arrays usually don't get printed fully (as they can be quite large). The ellipsis is to denote columns or rows that are not shown (truncated). To print the entire array see:&nbsp;http://docs.scipy.org/doc/numpy/reference/generated/numpy.set_printoptions.html</p>\n<p><code>import numpy<br>numpy.set_printoptions(threshold=numpy.nan)</code></p>\n<p>What you get in Python after loadmat is a dictionary. This dictionary has information on the data:</p>\n<p><code>print data['data']</code></p>\n<p>But also on metadata, like the headers:</p>\n<p><code>print data['__headers__']</code></p>\n<p>I believe the data is [channel x bins], so this could be 16 x 5000 for human subjects. Try it with:</p>\n<p><code>print data['data'].shape</code></p>\n<p>To familiarize yourself with the .mat files loaded into a dictionary:</p>\n<p><code>for key, value in data.items():</code><code>&nbsp; </code></p>\n<p><code>&nbsp; print key, &quot;\\n&quot;, value, &quot;\\n\\n&quot;</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 51549,
      "author_name": "lalitapatel",
      "author_url": "",
      "post_date": "08/03/2014 13:57:41",
      "content": "<p>Internet Explorer was [slightly] slower than Google Chrome in downloading the clips.gz file on a Windows 7 laptop.</p>\n<p>Extraction of the clips.gz file using WinRAR program gave a huge no-extension file. After renaming the clips.gz file as clips.tar.gz, the extraction gave several .mat files.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "44917": "",
    "44926": "",
    "44937": "",
    "44962": "",
    "44964": "",
    "44991": "",
    "46113": "",
    "46124": "",
    "51549": ""
  },
  "source": "meta"
}