{
  "id": 12545,
  "title": "Can't unzip data",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/12545",
  "author_name": "",
  "post_date": "2015-02-17T21:05:10.173Z",
  "votes": 8,
  "comment_count": 31,
  "views": 12300,
  "content": "<p>I've downloaded train.zip.005 twice and have not been able to unzip it with either unzip or 7z.&nbsp;</p>\n<p>Suggestions?</p>",
  "messages": [
    {
      "id": "64426",
      "postDate": "02/17/2015 21:05:10",
      "content": "<p>I've downloaded train.zip.005 twice and have not been able to unzip it with either unzip or 7z.&nbsp;</p>\n<p>Suggestions?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64429",
      "postDate": "02/17/2015 21:19:06",
      "content": "<p>You need to download all train parts, put them into one folder and only then try to unzip.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64436",
      "postDate": "02/17/2015 22:52:58",
      "content": "<p>what's the size of the total data after unzipping?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64449",
      "postDate": "02/18/2015 01:09:53",
      "content": "<p>@rcarson</p>\n\n<p>total size of <strong>training only</strong></p>\n\n<p>unzipped is 37.9 gigs</p>\n<p>~35,000 images</p>\n<p>image sizes range from 7kB (300 * 400 pixels) on the small size photos up to (6000 * 5000 pixels) 2.2MB for the larger ones.</p>\n\n<p>as Vlad noted. gotta have all the files in same folder to unarchive. &nbsp;&nbsp;&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64488",
      "postDate": "02/18/2015 13:14:57",
      "content": "<p>[quote=Timothy Scharf;64449]</p>\n<p>@rcarson</p>\n<p>total size of <strong>training only</strong></p>\n<p>unzipped is 37.9 gigs</p>\n<p>~35,000 images</p>\n<p>image sizes range from 7kB (300 * 400 pixels) on the small size photos up to (6000 * 5000 pixels) 2.2MB for the larger ones.</p>\n<p>as Vlad noted. gotta have all the files in same folder to unarchive. &nbsp;&nbsp;&nbsp;</p>\n<p>[/quote]</p>\n<p>Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64536",
      "postDate": "02/19/2015 02:53:51",
      "content": "<p>I'm unzipping all the files following this example on Mac OSX :&nbsp;https://gist.github.com/4np/2913012&nbsp;</p>\n<p><code>cat *.zip &gt; combined.zip;zip -FF combined.zip --out combined-fixed.zip;rm combined.zip;yes A|unzip -qq combined-fixed.zip;rm combined-fixed.zip</code></p>\n<p>but I'm still failing to extract the files. &nbsp;Should I be doing something else? Is there a hash for the zip files that we can use to confirm proper download?<br><br></p>\n<p>Thx</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64540",
      "postDate": "02/19/2015 03:09:33",
      "content": "<p>The MD5 checksums are here:</p>\n<p>http://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12543/could-you-please-generate-md5-hashes-for-files</p>\n\n\n<p>MD5 (./test.zip.001) = ebc9de8927479400920c100060910fe9<br>MD5 (./test.zip.002) = 06a36b7bb17cd5d3b81b0edbaae6098b<br>MD5 (./test.zip.003) = 2fc232c0eba9fc4436fbd86134c4a22f<br>MD5 (./test.zip.004) = e4a807cf1e8b975b23b67344515034a9<br>MD5 (./test.zip.005) = 9db444a8efd018b0a4f08626f7f273c5<br>MD5 (./test.zip.006) = 9da2d6838f625a871c1adb5dcad3b369<br>MD5 (./test.zip.007) = 79af54ef8aa81239570ad411b63e1184</p>\n<p>MD5 (./train.zip.001) = 39aa61ba09604d79b63ea6c4df44db3d<br>MD5 (./train.zip.002) = 78fe5bbd4835cc3036f911ee77146652<br>MD5 (./train.zip.003) = 6bdf19c38477267576f4bcfbc3de669d<br>MD5 (./train.zip.004) = 303365bdc425c0033c2e42a75274917a<br>MD5 (./train.zip.005) = 8b1745259c2963f3c3d4b953881c23b9</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64550",
      "postDate": "02/19/2015 04:43:21",
      "content": "<p><code>for the Mac use http://brew.sh/ to install http://p7zip.sourceforge.net/ :<br>&nbsp; brew install p7zip<br>With the 5 files in the same directory type:<br>&nbsp; 7zr x train.zip.001</code></p>\n<p>The extraction program knows to go from one archive file to the next, provided that they are in order (i.e. 001, 002, &#8230;)</p>\n<p>This will create a file train.zip after several minutes that's 33G in size (34988445506).</p>\n<p>&gt; 7zr x train.zip.001</p>\n<p>7-Zip (A) [64] 9.20 Copyright (c) 1999-2010 Igor Pavlov 2010-11-18<br>p7zip Version 9.20 (locale=utf8,Utf16=on,HugeFiles=on,8 CPUs)</p>\n<p>Processing archive: train.zip.001</p>\n<p>Extracting train.zip</p>\n<p>Everything is Ok</p>\n<p>Size: 34988445506<br>Compressed: 8388608000</p>\n\n<p>However I'm also getting an error with it:</p>\n<p>unzip -t train.zip</p>\n<p><br>Archive: train.zip<br>warning [train.zip]: 30690745840 extra bytes at beginning or within zipfile<br> (attempting to process anyway)<br>error [train.zip]: start of central directory not found;<br> zipfile corrupt.<br> (please check that you have transferred or created the zipfile in the<br> appropriate BINARY mode and that you have compiled UnZip properly)</p>\n\n<p>my md5's match up with Zero's post above.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64552",
      "postDate": "02/19/2015 04:55:23",
      "content": "<p>I tried to do this:<br>&nbsp; zip -FFv train.zip --out train_fixed.zip<br>But ran into errors like<br>&nbsp; zip warning: Illegal PK version mapping in local header: 785<br>Googling I found<br>&nbsp; http://sourceforge.net/p/p7zip/bugs/106/<br>Which says to use -Fv and implies that there's some weird Windows to unix problem.</p>\n<p>zip -Fv train.zip --out train_fixed.zip</p>\n<p>Anyway doing that I get a ton of errors like</p>\n<p>copying: train/99_left.jpeg<br> zip warning: Local Version Needed To Extract does not match CD: train/99_left.jpeg<br> copying: train/99_right.jpeg<br> zip warning: Local Version Needed To Extract does not match CD: train/99_right.jpeg</p>\n<p>The new train_fixed.zip file is the same size as train.zip still gives me the same error as in my previous post.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64577",
      "postDate": "02/19/2015 14:10:14",
      "content": "<p>@cwilkes, after the 7zr command I did:</p>\n<p><code>&gt; 7z x train.zip</code><code></code></p>\n<p>and that worked for me. Thanks for sending me down that path.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64581",
      "postDate": "02/19/2015 16:41:02",
      "content": "<p>It looks like the files are not real multi-part zips, but just chunks of a single zip file.</p>\n<p>This worked for me:</p>\n<p>&gt; cat&nbsp;train.zip.* &gt; train.zip</p>\n<p>&gt; unzip train.zip</p>\n<p>(make sure you have enough disk space)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64609",
      "postDate": "02/19/2015 21:57:34",
      "content": "<p>What worked for me on MacOS X with the Homebrew package manager:</p>\n<p style=\"padding-left: 30px\">brew install p7zip</p>\n<p style=\"padding-left: 30px\">7z x train.zip.000</p>\n<p>This results in:</p>\n<p style=\"padding-left: 30px\">Everything is Ok</p>\n<p style=\"padding-left: 30px\">Folders: 1<br>Files: 35126<br>Size: 37944223074<br>Compressed: 8388608000</p>\n<p>Re-creating the big zip file with 7zr is not necessary.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64610",
      "postDate": "02/19/2015 22:11:00",
      "content": "<p>I tried the simple approach (as suggested by Seyn) and got the same error as cwilkes (see below) using a Max OSX 10.8.5. I looked at the Keka software but that seems to just be a compressor.&nbsp;</p>\n<p>Can someone describe the most straightforward way to unzip these files?</p>\n<p>This is infuriating. I'm ready to give up on the project and I haven't started yet! Is downloading the data part of the challenge? ;-)</p>\n<p>**********************</p>\n<p>unzip train.zip</p>\n<p>Archive: train.zip<br>warning [train.zip]: 30690745840 extra bytes at beginning or within zipfile<br> (attempting to process anyway)<br>error [train.zip]: start of central directory not found;<br> zipfile corrupt.<br> (please check that you have transferred or created the zipfile in the<br> appropriate BINARY mode and that you have compiled UnZip properly)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64613",
      "postDate": "02/19/2015 23:10:27",
      "content": "<p>Keka works on Yosemite. &nbsp;Just right click the first t*zip.001 file and open with keka and it'll decompress all related files in a single go. Not sure how to do this directly from the application, however.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64619",
      "postDate": "02/20/2015 03:22:30",
      "content": "<p>1) Download all the train files into one&nbsp;directory.</p>\n<p>2) Download all the test files into&nbsp;another directory.</p>\n<p>3) Install 7-Zip (http://www.7-zip.org)</p>\n<p>4) Go to each of these directories from 7-Zip, select all the files within the directory and simply click Extract (the one with the blue '-' sign)</p>\n<p>Your files will be extracted...DONE</p>\n<p>Note: This worked on a Windows system.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64629",
      "postDate": "02/20/2015 04:46:56",
      "content": "<p>[quote=paulperry;64577]</p>\n<p>@cwilkes, after the 7zr command I did:</p>\n<p><code>&gt; 7z x train.zip</code><code></code></p>\n<p>and that worked for me. Thanks for sending me down that path.</p>\n<p>[/quote]</p>\n\n<p>Thanks, that worked for me as well. &nbsp;Extracted all&nbsp;35126 jpegs.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64650",
      "postDate": "02/20/2015 15:14:42",
      "content": "<p>In the end, what worked for me (on a Mac OSX 10.8.5) was:</p>\n<p>1) cat train.zip.* &gt; train.zip</p>\n<p>2) Using OSX's archive utility (which I think is the default associated program for zip files) from the GUI. (Running &quot;unzip train.zip&quot; failed)</p>\n<p>3) Going out to dinner while the Archive Utility ran. (It took at least an hour or so)</p>\n<p>Thanks&nbsp;to all for the suggestions &amp; trying to help.</p>\n\n<p>Dan</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64714",
      "postDate": "02/21/2015 18:15:00",
      "content": "<p>Painfully learned:&nbsp;if the hashes correspond but you&nbsp;still can't extract the damn thing, then the problem may come from the name or type of the files.</p>\n<p>Using windows+7zip it worked for me when the name [type] were: &nbsp;</p>\n<p>train.zip [zip file (.001)]</p>\n<p>train.zip.002 [002 File (.002)]</p>\n<p>train.zip.003 [003 File (.003)]</p>\n<p>train.zip.004 [004 File (.004)]</p>\n<p>train.zip.005 [005 File (.005)]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64917",
      "postDate": "02/26/2015 07:41:06",
      "content": "<p>For beginners like me,</p>\n<p>I couldn't unzip by 7zip in any ways. But the site below (Mr. Retrofire's suggestion) was very helpful for me.</p>\n<p>http://forums.macrumors.com/showthread.php?t=1271752</p>\n<p>It might be same as ArtherDent's suggestion. This site was helpful for folks like me who don't know how to use command line or file path.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65062",
      "postDate": "02/27/2015 19:20:46",
      "content": "<p>[quote=Timothy Scharf;64449]</p>\n<p>@rcarson</p>\n<p>total size of <strong>training only</strong></p>\n<p>unzipped is 37.9 gigs</p>\n<p>[/quote]</p>\n<p>I don't know if that's a relief or frustration; I was worried that it would be closer to the 100 GB range when unzipped, but 37.9 GB vs 29 GB of zip files is a poor compression ratio!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65159",
      "postDate": "03/01/2015 00:50:57",
      "content": "<p>If anyone else is using AWS, first you need to install p7zip</p>\n<p><code>sudo yum install p7zip --enablerepo=epel</code></p>\n<p>then just switch to the folder containing all the zip files</p>\n<p><code>7za x train.zip.001</code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65230",
      "postDate": "03/02/2015 12:01:39",
      "content": "<p>sorry for stupid question, but may be someone has already uploaded those huge files to AWS and can share it?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65485",
      "postDate": "03/05/2015 10:26:22",
      "content": "<p>If anyone uses Linux, this might be helpful. First, concatenate&nbsp;all partial zip files into single .zip file:</p>\n<p>cat train.zip.* &gt; train.zip</p>\n<p>cat test.zip.* &gt; test.zip</p>\n<p>Then just unzip these two files.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66379",
      "postDate": "03/16/2015 15:09:56",
      "content": "<p>I don't think '7zr x train.zip.001' is zipping all of the train.zip.* files based on the resulting size.&nbsp;</p>\n<p>Any suggestions?</p>\n\n<p>=======================================================================</p>\n<p>$&nbsp;7zr x train.zip.001</p>\n<p>7-Zip (A) [64] 9.20 Copyright (c) 1999-2010 Igor Pavlov 2010-11-18<br>p7zip Version 9.20 (locale=utf8,Utf16=on,HugeFiles=on,4 CPUs)</p>\n<p>Processing archive: train.zip.001</p>\n<p>Extracting train.zip</p>\n<p>Everything is Ok</p>\n<p>Size: 72465<br>Compressed: 14493</p>\n<p><br>$ unzip -t train.zip</p>\n<p><br>Archive: train.zip<br> End-of-central-directory signature not found. Either this file is not<br> a zipfile, or it constitutes one disk of a multi-part archive. In the<br> latter case the central directory and zipfile comment will be found on<br> the last disk(s) of this archive.<br>unzip: cannot find zipfile directory in one of train.zip or<br> train.zip.zip, and cannot find train.zip.ZIP, period.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66465",
      "postDate": "03/16/2015 19:48:37",
      "content": "<p>Thank you @seyn</p>\n<p>&gt; cat train.zip.* &gt; train.zip</p>\n<p>&gt; unzip train.zip</p>\n<p>I tried 7z first but it took forever, &nbsp;and didn't work in the end. &nbsp;This method was 10x faster and actually worked.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "67598",
      "postDate": "03/22/2015 07:06:25",
      "content": "<p>Check the MD5 hashes. OSX downloaded the files out of order, so I had to manually rename them to match the correct order. After that it unzipped without problems</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "68996",
      "postDate": "03/30/2015 12:13:34",
      "content": "<p>[quote=Torgos;65062]</p>\n<p>[quote=Timothy Scharf;64449]</p>\n<p>@rcarson</p>\n<p>total size of <strong>training only</strong></p>\n<p>unzipped is 37.9 gigs</p>\n<p>[/quote]</p>\n<p>I don't know if that's a relief or frustration; I was worried that it would be closer to the 100 GB range when unzipped, but 37.9 GB vs 29 GB of zip files is a poor compression ratio!</p>\n<p>[/quote]</p>\n<p>The JPEG images are already compressed, so the 7zip is essentially a way to bundle them conveniently.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69001",
      "postDate": "03/30/2015 12:38:16",
      "content": "<p>I'm getting a train set composed of 4277 images taking 4.66 GB&nbsp;and a&nbsp;test set is composed of 4311 images taking the same space.&nbsp;Why did I have to download 10 files of 7.81 GB each?</p>\n<p>Edit 1: this thread <a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12871/how-many-images-are-there\">How many images are there?</a> mentions many more images than what I got. So I did something wrong.</p>\n<p>Edit 2: OK, I got it. The command that worked for me (on MacOS) were:</p>\n<p><code>7zr x train.zip.001</code></p>\n<p>followed by</p>\n<p><code>7zr x train.zip</code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "75364",
      "postDate": "04/29/2015 06:04:13",
      "content": "<p>question to the organizers.</p>\n<p>&nbsp;48GB is a lot of disk space for me and potentially other people. Is any reason, we need to be aware of, why all files were put into a single archive? &nbsp; &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "77691",
      "postDate": "05/09/2015 14:02:29",
      "content": "<p>Thanks for that Thomas! Works perfectly on my AWS Ubuntu setup.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "993938",
      "postDate": "09/01/2020 08:43:44",
      "content": "<p>I've downloaded my dataset in google drive through google colab and i want to extract training and testing dataset how can i do that? all things are in same folder</p>",
      "rawMarkdown": "I've downloaded my dataset in google drive through google colab and i want to extract training and testing dataset how can i do that? all things are in same folder",
      "votes": null
    },
    {
      "id": "1266109",
      "postDate": "04/07/2021 13:44:26",
      "content": "<p>If you are using in google colab and are constrined by the drive space,then try using 7za x train.zip.001 in your case,this helps to download data directly to session storage whereas using cat to make it into a single file and extract is extremely tiresome.</p>",
      "rawMarkdown": "If you are using in google colab and are constrined by the drive space,then try using 7za x train.zip.001 in your case,this helps to download data directly to session storage whereas using cat to make it into a single file and extract is extremely tiresome.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 993938,
      "author_name": "paurav2912",
      "author_url": "",
      "post_date": "09/01/2020 08:43:44",
      "content": "<p>I've downloaded my dataset in google drive through google colab and i want to extract training and testing dataset how can i do that? all things are in same folder</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1266109,
      "author_name": "trivik261",
      "author_url": "",
      "post_date": "04/07/2021 13:44:26",
      "content": "<p>If you are using in google colab and are constrined by the drive space,then try using 7za x train.zip.001 in your case,this helps to download data directly to session storage whereas using cat to make it into a single file and extract is extremely tiresome.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64429,
      "author_name": "vladislavbelyaev",
      "author_url": "",
      "post_date": "02/17/2015 21:19:06",
      "content": "<p>You need to download all train parts, put them into one folder and only then try to unzip.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64436,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "02/17/2015 22:52:58",
      "content": "<p>what's the size of the total data after unzipping?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64449,
      "author_name": "scharf",
      "author_url": "",
      "post_date": "02/18/2015 01:09:53",
      "content": "<p>@rcarson</p>\n\n<p>total size of <strong>training only</strong></p>\n\n<p>unzipped is 37.9 gigs</p>\n<p>~35,000 images</p>\n<p>image sizes range from 7kB (300 * 400 pixels) on the small size photos up to (6000 * 5000 pixels) 2.2MB for the larger ones.</p>\n\n<p>as Vlad noted. gotta have all the files in same folder to unarchive. &nbsp;&nbsp;&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64488,
      "author_name": "asaini",
      "author_url": "",
      "post_date": "02/18/2015 13:14:57",
      "content": "<p>[quote=Timothy Scharf;64449]</p>\n<p>@rcarson</p>\n<p>total size of <strong>training only</strong></p>\n<p>unzipped is 37.9 gigs</p>\n<p>~35,000 images</p>\n<p>image sizes range from 7kB (300 * 400 pixels) on the small size photos up to (6000 * 5000 pixels) 2.2MB for the larger ones.</p>\n<p>as Vlad noted. gotta have all the files in same folder to unarchive. &nbsp;&nbsp;&nbsp;</p>\n<p>[/quote]</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64536,
      "author_name": "paulperry",
      "author_url": "",
      "post_date": "02/19/2015 02:53:51",
      "content": "<p>I'm unzipping all the files following this example on Mac OSX :&nbsp;https://gist.github.com/4np/2913012&nbsp;</p>\n<p><code>cat *.zip &gt; combined.zip;zip -FF combined.zip --out combined-fixed.zip;rm combined.zip;yes A|unzip -qq combined-fixed.zip;rm combined-fixed.zip</code></p>\n<p>but I'm still failing to extract the files. &nbsp;Should I be doing something else? Is there a hash for the zip files that we can use to confirm proper download?<br><br></p>\n<p>Thx</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64540,
      "author_name": "asaini",
      "author_url": "",
      "post_date": "02/19/2015 03:09:33",
      "content": "<p>The MD5 checksums are here:</p>\n<p>http://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12543/could-you-please-generate-md5-hashes-for-files</p>\n\n\n<p>MD5 (./test.zip.001) = ebc9de8927479400920c100060910fe9<br>MD5 (./test.zip.002) = 06a36b7bb17cd5d3b81b0edbaae6098b<br>MD5 (./test.zip.003) = 2fc232c0eba9fc4436fbd86134c4a22f<br>MD5 (./test.zip.004) = e4a807cf1e8b975b23b67344515034a9<br>MD5 (./test.zip.005) = 9db444a8efd018b0a4f08626f7f273c5<br>MD5 (./test.zip.006) = 9da2d6838f625a871c1adb5dcad3b369<br>MD5 (./test.zip.007) = 79af54ef8aa81239570ad411b63e1184</p>\n<p>MD5 (./train.zip.001) = 39aa61ba09604d79b63ea6c4df44db3d<br>MD5 (./train.zip.002) = 78fe5bbd4835cc3036f911ee77146652<br>MD5 (./train.zip.003) = 6bdf19c38477267576f4bcfbc3de669d<br>MD5 (./train.zip.004) = 303365bdc425c0033c2e42a75274917a<br>MD5 (./train.zip.005) = 8b1745259c2963f3c3d4b953881c23b9</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64550,
      "author_name": "cwilkes",
      "author_url": "",
      "post_date": "02/19/2015 04:43:21",
      "content": "<p><code>for the Mac use http://brew.sh/ to install http://p7zip.sourceforge.net/ :<br>&nbsp; brew install p7zip<br>With the 5 files in the same directory type:<br>&nbsp; 7zr x train.zip.001</code></p>\n<p>The extraction program knows to go from one archive file to the next, provided that they are in order (i.e. 001, 002, &#8230;)</p>\n<p>This will create a file train.zip after several minutes that's 33G in size (34988445506).</p>\n<p>&gt; 7zr x train.zip.001</p>\n<p>7-Zip (A) [64] 9.20 Copyright (c) 1999-2010 Igor Pavlov 2010-11-18<br>p7zip Version 9.20 (locale=utf8,Utf16=on,HugeFiles=on,8 CPUs)</p>\n<p>Processing archive: train.zip.001</p>\n<p>Extracting train.zip</p>\n<p>Everything is Ok</p>\n<p>Size: 34988445506<br>Compressed: 8388608000</p>\n\n<p>However I'm also getting an error with it:</p>\n<p>unzip -t train.zip</p>\n<p><br>Archive: train.zip<br>warning [train.zip]: 30690745840 extra bytes at beginning or within zipfile<br> (attempting to process anyway)<br>error [train.zip]: start of central directory not found;<br> zipfile corrupt.<br> (please check that you have transferred or created the zipfile in the<br> appropriate BINARY mode and that you have compiled UnZip properly)</p>\n\n<p>my md5's match up with Zero's post above.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64552,
      "author_name": "cwilkes",
      "author_url": "",
      "post_date": "02/19/2015 04:55:23",
      "content": "<p>I tried to do this:<br>&nbsp; zip -FFv train.zip --out train_fixed.zip<br>But ran into errors like<br>&nbsp; zip warning: Illegal PK version mapping in local header: 785<br>Googling I found<br>&nbsp; http://sourceforge.net/p/p7zip/bugs/106/<br>Which says to use -Fv and implies that there's some weird Windows to unix problem.</p>\n<p>zip -Fv train.zip --out train_fixed.zip</p>\n<p>Anyway doing that I get a ton of errors like</p>\n<p>copying: train/99_left.jpeg<br> zip warning: Local Version Needed To Extract does not match CD: train/99_left.jpeg<br> copying: train/99_right.jpeg<br> zip warning: Local Version Needed To Extract does not match CD: train/99_right.jpeg</p>\n<p>The new train_fixed.zip file is the same size as train.zip still gives me the same error as in my previous post.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64577,
      "author_name": "paulperry",
      "author_url": "",
      "post_date": "02/19/2015 14:10:14",
      "content": "<p>@cwilkes, after the 7zr command I did:</p>\n<p><code>&gt; 7z x train.zip</code><code></code></p>\n<p>and that worked for me. Thanks for sending me down that path.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64581,
      "author_name": "mikaelrousson",
      "author_url": "",
      "post_date": "02/19/2015 16:41:02",
      "content": "<p>It looks like the files are not real multi-part zips, but just chunks of a single zip file.</p>\n<p>This worked for me:</p>\n<p>&gt; cat&nbsp;train.zip.* &gt; train.zip</p>\n<p>&gt; unzip train.zip</p>\n<p>(make sure you have enough disk space)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64609,
      "author_name": "mlnoga",
      "author_url": "",
      "post_date": "02/19/2015 21:57:34",
      "content": "<p>What worked for me on MacOS X with the Homebrew package manager:</p>\n<p style=\"padding-left: 30px\">brew install p7zip</p>\n<p style=\"padding-left: 30px\">7z x train.zip.000</p>\n<p>This results in:</p>\n<p style=\"padding-left: 30px\">Everything is Ok</p>\n<p style=\"padding-left: 30px\">Folders: 1<br>Files: 35126<br>Size: 37944223074<br>Compressed: 8388608000</p>\n<p>Re-creating the big zip file with 7zr is not necessary.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64610,
      "author_name": "arthurdent",
      "author_url": "",
      "post_date": "02/19/2015 22:11:00",
      "content": "<p>I tried the simple approach (as suggested by Seyn) and got the same error as cwilkes (see below) using a Max OSX 10.8.5. I looked at the Keka software but that seems to just be a compressor.&nbsp;</p>\n<p>Can someone describe the most straightforward way to unzip these files?</p>\n<p>This is infuriating. I'm ready to give up on the project and I haven't started yet! Is downloading the data part of the challenge? ;-)</p>\n<p>**********************</p>\n<p>unzip train.zip</p>\n<p>Archive: train.zip<br>warning [train.zip]: 30690745840 extra bytes at beginning or within zipfile<br> (attempting to process anyway)<br>error [train.zip]: start of central directory not found;<br> zipfile corrupt.<br> (please check that you have transferred or created the zipfile in the<br> appropriate BINARY mode and that you have compiled UnZip properly)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64613,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "02/19/2015 23:10:27",
      "content": "<p>Keka works on Yosemite. &nbsp;Just right click the first t*zip.001 file and open with keka and it'll decompress all related files in a single go. Not sure how to do this directly from the application, however.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64619,
      "author_name": "",
      "author_url": "",
      "post_date": "02/20/2015 03:22:30",
      "content": "<p>1) Download all the train files into one&nbsp;directory.</p>\n<p>2) Download all the test files into&nbsp;another directory.</p>\n<p>3) Install 7-Zip (http://www.7-zip.org)</p>\n<p>4) Go to each of these directories from 7-Zip, select all the files within the directory and simply click Extract (the one with the blue '-' sign)</p>\n<p>Your files will be extracted...DONE</p>\n<p>Note: This worked on a Windows system.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64629,
      "author_name": "cwilkes",
      "author_url": "",
      "post_date": "02/20/2015 04:46:56",
      "content": "<p>[quote=paulperry;64577]</p>\n<p>@cwilkes, after the 7zr command I did:</p>\n<p><code>&gt; 7z x train.zip</code><code></code></p>\n<p>and that worked for me. Thanks for sending me down that path.</p>\n<p>[/quote]</p>\n\n<p>Thanks, that worked for me as well. &nbsp;Extracted all&nbsp;35126 jpegs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64650,
      "author_name": "arthurdent",
      "author_url": "",
      "post_date": "02/20/2015 15:14:42",
      "content": "<p>In the end, what worked for me (on a Mac OSX 10.8.5) was:</p>\n<p>1) cat train.zip.* &gt; train.zip</p>\n<p>2) Using OSX's archive utility (which I think is the default associated program for zip files) from the GUI. (Running &quot;unzip train.zip&quot; failed)</p>\n<p>3) Going out to dinner while the Archive Utility ran. (It took at least an hour or so)</p>\n<p>Thanks&nbsp;to all for the suggestions &amp; trying to help.</p>\n\n<p>Dan</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64714,
      "author_name": "woolsey",
      "author_url": "",
      "post_date": "02/21/2015 18:15:00",
      "content": "<p>Painfully learned:&nbsp;if the hashes correspond but you&nbsp;still can't extract the damn thing, then the problem may come from the name or type of the files.</p>\n<p>Using windows+7zip it worked for me when the name [type] were: &nbsp;</p>\n<p>train.zip [zip file (.001)]</p>\n<p>train.zip.002 [002 File (.002)]</p>\n<p>train.zip.003 [003 File (.003)]</p>\n<p>train.zip.004 [004 File (.004)]</p>\n<p>train.zip.005 [005 File (.005)]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64917,
      "author_name": "miwachu",
      "author_url": "",
      "post_date": "02/26/2015 07:41:06",
      "content": "<p>For beginners like me,</p>\n<p>I couldn't unzip by 7zip in any ways. But the site below (Mr. Retrofire's suggestion) was very helpful for me.</p>\n<p>http://forums.macrumors.com/showthread.php?t=1271752</p>\n<p>It might be same as ArtherDent's suggestion. This site was helpful for folks like me who don't know how to use command line or file path.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65062,
      "author_name": "telser",
      "author_url": "",
      "post_date": "02/27/2015 19:20:46",
      "content": "<p>[quote=Timothy Scharf;64449]</p>\n<p>@rcarson</p>\n<p>total size of <strong>training only</strong></p>\n<p>unzipped is 37.9 gigs</p>\n<p>[/quote]</p>\n<p>I don't know if that's a relief or frustration; I was worried that it would be closer to the 100 GB range when unzipped, but 37.9 GB vs 29 GB of zip files is a poor compression ratio!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65159,
      "author_name": "thomaswatson",
      "author_url": "",
      "post_date": "03/01/2015 00:50:57",
      "content": "<p>If anyone else is using AWS, first you need to install p7zip</p>\n<p><code>sudo yum install p7zip --enablerepo=epel</code></p>\n<p>then just switch to the folder containing all the zip files</p>\n<p><code>7za x train.zip.001</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65230,
      "author_name": "alex180",
      "author_url": "",
      "post_date": "03/02/2015 12:01:39",
      "content": "<p>sorry for stupid question, but may be someone has already uploaded those huge files to AWS and can share it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65485,
      "author_name": "nikogamulin",
      "author_url": "",
      "post_date": "03/05/2015 10:26:22",
      "content": "<p>If anyone uses Linux, this might be helpful. First, concatenate&nbsp;all partial zip files into single .zip file:</p>\n<p>cat train.zip.* &gt; train.zip</p>\n<p>cat test.zip.* &gt; test.zip</p>\n<p>Then just unzip these two files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66379,
      "author_name": "mesmes",
      "author_url": "",
      "post_date": "03/16/2015 15:09:56",
      "content": "<p>I don't think '7zr x train.zip.001' is zipping all of the train.zip.* files based on the resulting size.&nbsp;</p>\n<p>Any suggestions?</p>\n\n<p>=======================================================================</p>\n<p>$&nbsp;7zr x train.zip.001</p>\n<p>7-Zip (A) [64] 9.20 Copyright (c) 1999-2010 Igor Pavlov 2010-11-18<br>p7zip Version 9.20 (locale=utf8,Utf16=on,HugeFiles=on,4 CPUs)</p>\n<p>Processing archive: train.zip.001</p>\n<p>Extracting train.zip</p>\n<p>Everything is Ok</p>\n<p>Size: 72465<br>Compressed: 14493</p>\n<p><br>$ unzip -t train.zip</p>\n<p><br>Archive: train.zip<br> End-of-central-directory signature not found. Either this file is not<br> a zipfile, or it constitutes one disk of a multi-part archive. In the<br> latter case the central directory and zipfile comment will be found on<br> the last disk(s) of this archive.<br>unzip: cannot find zipfile directory in one of train.zip or<br> train.zip.zip, and cannot find train.zip.ZIP, period.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66465,
      "author_name": "kevinsoucy",
      "author_url": "",
      "post_date": "03/16/2015 19:48:37",
      "content": "<p>Thank you @seyn</p>\n<p>&gt; cat train.zip.* &gt; train.zip</p>\n<p>&gt; unzip train.zip</p>\n<p>I tried 7z first but it took forever, &nbsp;and didn't work in the end. &nbsp;This method was 10x faster and actually worked.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 67598,
      "author_name": "mistakenot",
      "author_url": "",
      "post_date": "03/22/2015 07:06:25",
      "content": "<p>Check the MD5 hashes. OSX downloaded the files out of order, so I had to manually rename them to match the correct order. After that it unzipped without problems</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 68996,
      "author_name": "hugues",
      "author_url": "",
      "post_date": "03/30/2015 12:13:34",
      "content": "<p>[quote=Torgos;65062]</p>\n<p>[quote=Timothy Scharf;64449]</p>\n<p>@rcarson</p>\n<p>total size of <strong>training only</strong></p>\n<p>unzipped is 37.9 gigs</p>\n<p>[/quote]</p>\n<p>I don't know if that's a relief or frustration; I was worried that it would be closer to the 100 GB range when unzipped, but 37.9 GB vs 29 GB of zip files is a poor compression ratio!</p>\n<p>[/quote]</p>\n<p>The JPEG images are already compressed, so the 7zip is essentially a way to bundle them conveniently.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69001,
      "author_name": "hugues",
      "author_url": "",
      "post_date": "03/30/2015 12:38:16",
      "content": "<p>I'm getting a train set composed of 4277 images taking 4.66 GB&nbsp;and a&nbsp;test set is composed of 4311 images taking the same space.&nbsp;Why did I have to download 10 files of 7.81 GB each?</p>\n<p>Edit 1: this thread <a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12871/how-many-images-are-there\">How many images are there?</a> mentions many more images than what I got. So I did something wrong.</p>\n<p>Edit 2: OK, I got it. The command that worked for me (on MacOS) were:</p>\n<p><code>7zr x train.zip.001</code></p>\n<p>followed by</p>\n<p><code>7zr x train.zip</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 75364,
      "author_name": "agloolga",
      "author_url": "",
      "post_date": "04/29/2015 06:04:13",
      "content": "<p>question to the organizers.</p>\n<p>&nbsp;48GB is a lot of disk space for me and potentially other people. Is any reason, we need to be aware of, why all files were put into a single archive? &nbsp; &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 77691,
      "author_name": "bhm20038",
      "author_url": "",
      "post_date": "05/09/2015 14:02:29",
      "content": "<p>Thanks for that Thomas! Works perfectly on my AWS Ubuntu setup.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "64426": "",
    "64429": "",
    "64436": "",
    "64449": "",
    "64488": "",
    "64536": "",
    "64540": "",
    "64550": "",
    "64552": "",
    "64577": "",
    "64581": "",
    "64609": "",
    "64610": "",
    "64613": "",
    "64619": "",
    "64629": "",
    "64650": "",
    "64714": "",
    "64917": "",
    "65062": "",
    "65159": "",
    "65230": "",
    "65485": "",
    "66379": "",
    "66465": "",
    "67598": "",
    "68996": "",
    "69001": "",
    "75364": "",
    "77691": "",
    "993938": "I've downloaded my dataset in google drive through google colab and i want to extract training and testing dataset how can i do that? all things are in same folder",
    "1266109": "If you are using in google colab and are constrined by the drive space,then try using 7za x train.zip.001 in your case,this helps to download data directly to session storage whereas using cat to make it into a single file and extract is extremely tiresome."
  },
  "source": "meta"
}