{
  "id": 12543,
  "title": "Could you please generate MD5 hashes for files",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/12543",
  "author_name": "",
  "post_date": "2015-02-17T20:28:53.973Z",
  "votes": 7,
  "comment_count": 10,
  "views": 4210,
  "content": "<p>We found these handy in other Kaggles when checking large downloads</p>\n<p>Thanks, Alastair</p>",
  "messages": [
    {
      "id": "64422",
      "postDate": "02/17/2015 20:28:53",
      "content": "<p>We found these handy in other Kaggles when checking large downloads</p>\n<p>Thanks, Alastair</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64424",
      "postDate": "02/17/2015 20:57:49",
      "content": "<p>MD5 (./test.zip.001) = ebc9de8927479400920c100060910fe9<br>MD5 (./test.zip.002) = 06a36b7bb17cd5d3b81b0edbaae6098b<br>MD5 (./test.zip.003) = 2fc232c0eba9fc4436fbd86134c4a22f<br>MD5 (./test.zip.004) = e4a807cf1e8b975b23b67344515034a9<br>MD5 (./test.zip.005) = 9db444a8efd018b0a4f08626f7f273c5<br>MD5 (./test.zip.006) = 9da2d6838f625a871c1adb5dcad3b369<br>MD5 (./test.zip.007) = 79af54ef8aa81239570ad411b63e1184</p>\n<p>MD5 (./train.zip.001) = 39aa61ba09604d79b63ea6c4df44db3d<br>MD5 (./train.zip.002) = 78fe5bbd4835cc3036f911ee77146652<br>MD5 (./train.zip.003) = 6bdf19c38477267576f4bcfbc3de669d<br>MD5 (./train.zip.004) = 303365bdc425c0033c2e42a75274917a<br>MD5 (./train.zip.005) = 8b1745259c2963f3c3d4b953881c23b9</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64432",
      "postDate": "02/17/2015 21:55:50",
      "content": "<p>@William, thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64433",
      "postDate": "02/17/2015 22:17:53",
      "content": "<p>[quote=William Cukierski;64424]</p>\n<p>MD5 (./test.zip.001) = ebc9de8927479400920c100060910fe9<br>MD5 (./test.zip.002) = 06a36b7bb17cd5d3b81b0edbaae6098b<br>MD5 (./test.zip.003) = 2fc232c0eba9fc4436fbd86134c4a22f<br>MD5 (./test.zip.004) = e4a807cf1e8b975b23b67344515034a9<br>MD5 (./test.zip.005) = 9db444a8efd018b0a4f08626f7f273c5<br>MD5 (./test.zip.006) = 9da2d6838f625a871c1adb5dcad3b369<br>MD5 (./test.zip.007) = 79af54ef8aa81239570ad411b63e1184</p>\n<p>MD5 (./train.zip.001) = 39aa61ba09604d79b63ea6c4df44db3d<br>MD5 (./train.zip.002) = 78fe5bbd4835cc3036f911ee77146652<br>MD5 (./train.zip.003) = 6bdf19c38477267576f4bcfbc3de669d<br>MD5 (./train.zip.004) = 303365bdc425c0033c2e42a75274917a<br>MD5 (./train.zip.005) = 8b1745259c2963f3c3d4b953881c23b9</p>\n<p>[/quote]</p>\n<p>@William: Could you please consider adding this information to the data page?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64668",
      "postDate": "02/20/2015 20:54:32",
      "content": "<p>My train.zip.001 has MD5 starting with 61f1d, and is&nbsp;8388608249 bytes long. &nbsp;The rest are&nbsp;8388608000 bytes, 249 bytes shorter. &nbsp;Anyone else run into this? &nbsp;I seem to have wound up with a sane set of jpegs nonetheless.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64675",
      "postDate": "02/20/2015 22:58:31",
      "content": "<p>@zeros_benchmark</p>\n<p>My train.zip.001 is 8,388,608,000 bytes with hash of 39aa61ba09604d79b63ea6c4df44db3d, just as expected. I have 35.126 total training items.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64683",
      "postDate": "02/21/2015 00:59:43",
      "content": "<p>Thanks, Alistair. &nbsp;I have the &nbsp;same number of training images. &nbsp;Downloading it again, just to be sure.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64693",
      "postDate": "02/21/2015 04:24:51",
      "content": "<p>[quote=throwaway;64668]</p>\n<p>My train.zip.001 has MD5 starting with 61f1d, and is&nbsp;8388608249 bytes long. &nbsp;The rest are&nbsp;8388608000 bytes, 249 bytes shorter. &nbsp;Anyone else run into this? &nbsp;I seem to have wound up with a sane set of jpegs nonetheless.</p>\n<p>[/quote]</p>\n<p>My train.zip.001 is also as expected:39aa61ba09604d79b63ea6c4df44db3d</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64713",
      "postDate": "02/21/2015 16:53:14",
      "content": "<p>Thanks for the confirmation, I have a matching file now. &nbsp;One image was corrupted in the previous download. &nbsp;I think it was probably a problem with continuing an interrupted download using curl.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "64715",
      "postDate": "02/21/2015 18:17:30",
      "content": "<p>@William, thanks</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65698",
      "postDate": "03/08/2015 04:27:20",
      "content": "<p>Seeing as my version of unzip does not read stdin, or operate on more than one input file at a time, I constructed the full zip files in order to unzip them.&nbsp; After concatenating train and test zip files I get:</p>\n<p>md5sum train.zip f96babfd62e9b8286e52246f5d7cd553</p>\n<p>md5sum test.zip&nbsp; 528c93bdc390988f291502dba80d36b1</p>\n<p>Can anyone confirm this?&nbsp;</p>\n<p>It's what I expected in this thread.</p>\n<p>ref:</p>\n<p>https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12545/can-t-unzip-data</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 64424,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "02/17/2015 20:57:49",
      "content": "<p>MD5 (./test.zip.001) = ebc9de8927479400920c100060910fe9<br>MD5 (./test.zip.002) = 06a36b7bb17cd5d3b81b0edbaae6098b<br>MD5 (./test.zip.003) = 2fc232c0eba9fc4436fbd86134c4a22f<br>MD5 (./test.zip.004) = e4a807cf1e8b975b23b67344515034a9<br>MD5 (./test.zip.005) = 9db444a8efd018b0a4f08626f7f273c5<br>MD5 (./test.zip.006) = 9da2d6838f625a871c1adb5dcad3b369<br>MD5 (./test.zip.007) = 79af54ef8aa81239570ad411b63e1184</p>\n<p>MD5 (./train.zip.001) = 39aa61ba09604d79b63ea6c4df44db3d<br>MD5 (./train.zip.002) = 78fe5bbd4835cc3036f911ee77146652<br>MD5 (./train.zip.003) = 6bdf19c38477267576f4bcfbc3de669d<br>MD5 (./train.zip.004) = 303365bdc425c0033c2e42a75274917a<br>MD5 (./train.zip.005) = 8b1745259c2963f3c3d4b953881c23b9</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64432,
      "author_name": "alastairk",
      "author_url": "",
      "post_date": "02/17/2015 21:55:50",
      "content": "<p>@William, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64433,
      "author_name": "asaini",
      "author_url": "",
      "post_date": "02/17/2015 22:17:53",
      "content": "<p>[quote=William Cukierski;64424]</p>\n<p>MD5 (./test.zip.001) = ebc9de8927479400920c100060910fe9<br>MD5 (./test.zip.002) = 06a36b7bb17cd5d3b81b0edbaae6098b<br>MD5 (./test.zip.003) = 2fc232c0eba9fc4436fbd86134c4a22f<br>MD5 (./test.zip.004) = e4a807cf1e8b975b23b67344515034a9<br>MD5 (./test.zip.005) = 9db444a8efd018b0a4f08626f7f273c5<br>MD5 (./test.zip.006) = 9da2d6838f625a871c1adb5dcad3b369<br>MD5 (./test.zip.007) = 79af54ef8aa81239570ad411b63e1184</p>\n<p>MD5 (./train.zip.001) = 39aa61ba09604d79b63ea6c4df44db3d<br>MD5 (./train.zip.002) = 78fe5bbd4835cc3036f911ee77146652<br>MD5 (./train.zip.003) = 6bdf19c38477267576f4bcfbc3de669d<br>MD5 (./train.zip.004) = 303365bdc425c0033c2e42a75274917a<br>MD5 (./train.zip.005) = 8b1745259c2963f3c3d4b953881c23b9</p>\n<p>[/quote]</p>\n<p>@William: Could you please consider adding this information to the data page?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64668,
      "author_name": "alexcoventry",
      "author_url": "",
      "post_date": "02/20/2015 20:54:32",
      "content": "<p>My train.zip.001 has MD5 starting with 61f1d, and is&nbsp;8388608249 bytes long. &nbsp;The rest are&nbsp;8388608000 bytes, 249 bytes shorter. &nbsp;Anyone else run into this? &nbsp;I seem to have wound up with a sane set of jpegs nonetheless.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64675,
      "author_name": "alastairk",
      "author_url": "",
      "post_date": "02/20/2015 22:58:31",
      "content": "<p>@zeros_benchmark</p>\n<p>My train.zip.001 is 8,388,608,000 bytes with hash of 39aa61ba09604d79b63ea6c4df44db3d, just as expected. I have 35.126 total training items.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64683,
      "author_name": "alexcoventry",
      "author_url": "",
      "post_date": "02/21/2015 00:59:43",
      "content": "<p>Thanks, Alistair. &nbsp;I have the &nbsp;same number of training images. &nbsp;Downloading it again, just to be sure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64693,
      "author_name": "asaini",
      "author_url": "",
      "post_date": "02/21/2015 04:24:51",
      "content": "<p>[quote=throwaway;64668]</p>\n<p>My train.zip.001 has MD5 starting with 61f1d, and is&nbsp;8388608249 bytes long. &nbsp;The rest are&nbsp;8388608000 bytes, 249 bytes shorter. &nbsp;Anyone else run into this? &nbsp;I seem to have wound up with a sane set of jpegs nonetheless.</p>\n<p>[/quote]</p>\n<p>My train.zip.001 is also as expected:39aa61ba09604d79b63ea6c4df44db3d</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64713,
      "author_name": "alexcoventry",
      "author_url": "",
      "post_date": "02/21/2015 16:53:14",
      "content": "<p>Thanks for the confirmation, I have a matching file now. &nbsp;One image was corrupted in the previous download. &nbsp;I think it was probably a problem with continuing an interrupted download using curl.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 64715,
      "author_name": "woolsey",
      "author_url": "",
      "post_date": "02/21/2015 18:17:30",
      "content": "<p>@William, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65698,
      "author_name": "lance10",
      "author_url": "",
      "post_date": "03/08/2015 04:27:20",
      "content": "<p>Seeing as my version of unzip does not read stdin, or operate on more than one input file at a time, I constructed the full zip files in order to unzip them.&nbsp; After concatenating train and test zip files I get:</p>\n<p>md5sum train.zip f96babfd62e9b8286e52246f5d7cd553</p>\n<p>md5sum test.zip&nbsp; 528c93bdc390988f291502dba80d36b1</p>\n<p>Can anyone confirm this?&nbsp;</p>\n<p>It's what I expected in this thread.</p>\n<p>ref:</p>\n<p>https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12545/can-t-unzip-data</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "64422": "",
    "64424": "",
    "64432": "",
    "64433": "",
    "64668": "",
    "64675": "",
    "64683": "",
    "64693": "",
    "64713": "",
    "64715": "",
    "65698": ""
  },
  "source": "meta"
}