{
  "id": 16176,
  "title": "Please re-download competition files",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16176",
  "author_name": "",
  "post_date": "2015-08-28T18:59:09.073Z",
  "votes": 3,
  "comment_count": 9,
  "views": 2045,
  "content": "<p>Hi all,</p>\n\n<p>There were some image labels lost in the processing pipeline code, and also <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/forums/t/16165/bad-zipfile-offsets/90607#post90607\">reported</a> that for some people (usually on Windows) the un-zipping of images files give errors. </p>\n\n<p>We have now fixed both, and we are uploading new files (train.csv, sample_submission.csv, imgs.zip). The upload should be done in the next hour. Please download the files again. Previously, the zipping was done with 7z. This time it's done in normal zip (I used Keka on a Mac). </p>\n\n<p>Some details about the data files for those that want to double check:</p>\n\n<ul>\n<li>There are 11469 images in the imgs folder, named w_0.jpg to w_11468.jpg. </li>\n<li>There are 4544 images for training</li>\n<li>There are 6925 images for test</li>\n<li>4544+6925=11469</li>\n</ul>\n\n<p>Sorry for the inconvenience.</p>",
  "messages": [
    {
      "id": "90683",
      "postDate": "08/28/2015 18:59:09",
      "content": "<p>Hi all,</p>\n\n<p>There were some image labels lost in the processing pipeline code, and also <a href=\"https://www.kaggle.com/c/noaa-right-whale-recognition/forums/t/16165/bad-zipfile-offsets/90607#post90607\">reported</a> that for some people (usually on Windows) the un-zipping of images files give errors. </p>\n\n<p>We have now fixed both, and we are uploading new files (train.csv, sample_submission.csv, imgs.zip). The upload should be done in the next hour. Please download the files again. Previously, the zipping was done with 7z. This time it's done in normal zip (I used Keka on a Mac). </p>\n\n<p>Some details about the data files for those that want to double check:</p>\n\n<ul>\n<li>There are 11469 images in the imgs folder, named w_0.jpg to w_11468.jpg. </li>\n<li>There are 4544 images for training</li>\n<li>There are 6925 images for test</li>\n<li>4544+6925=11469</li>\n</ul>\n\n<p>Sorry for the inconvenience.</p>",
      "rawMarkdown": "Hi all,\r\n\r\nThere were some image labels lost in the processing pipeline code, and also [reported][1] that for some people (usually on Windows) the un-zipping of images files give errors. \r\n\r\nWe have now fixed both, and we are uploading new files (train.csv, sample_submission.csv, imgs.zip). The upload should be done in the next hour. Please download the files again. Previously, the zipping was done with 7z. This time it's done in normal zip (I used Keka on a Mac). \r\n\r\nSome details about the data files for those that want to double check:\r\n\r\n- There are 11469 images in the imgs folder, named w_0.jpg to w_11468.jpg. \r\n- There are 4544 images for training\r\n- There are 6925 images for test\r\n- 4544+6925=11469\r\n\r\nSorry for the inconvenience.\r\n\r\n  [1]: https://www.kaggle.com/c/noaa-right-whale-recognition/forums/t/16165/bad-zipfile-offsets/90607#post90607",
      "votes": null
    },
    {
      "id": "90690",
      "postDate": "08/28/2015 19:48:45",
      "content": "<p>That explains it! Thanks!!</p>",
      "rawMarkdown": "That explains it! Thanks!!",
      "votes": null
    },
    {
      "id": "90701",
      "postDate": "08/28/2015 20:23:10",
      "content": "<p>I didn't see the imgs.zip file on the data page. Only the two csv files. Could you check this? Thank you.</p>",
      "rawMarkdown": "I didn't see the imgs.zip file on the data page. Only the two csv files. Could you check this? Thank you.",
      "votes": null
    },
    {
      "id": "90716",
      "postDate": "08/28/2015 20:56:02",
      "content": "<p>Yes, you are correct. Due to the size of imgs.zip, it's still uploading. </p>\n\n<p>I'm at 5.75 GB / 8.73 GB now. </p>",
      "rawMarkdown": "Yes, you are correct. Due to the size of imgs.zip, it's still uploading. \r\n\r\nI'm at 5.75 GB / 8.73 GB now.",
      "votes": null
    },
    {
      "id": "90724",
      "postDate": "08/28/2015 21:30:59",
      "content": "<p>Upload is finished now. You may now download imgs.zip. </p>",
      "rawMarkdown": "Upload is finished now. You may now download imgs.zip.",
      "votes": null
    },
    {
      "id": "90739",
      "postDate": "08/28/2015 23:14:46",
      "content": "<p>Anyone else missing w_7489.jpg?</p>",
      "rawMarkdown": "Anyone else missing w_7489.jpg?",
      "votes": null
    },
    {
      "id": "90753",
      "postDate": "08/29/2015 00:07:07",
      "content": "<p>Yep, I don't see that one. </p>",
      "rawMarkdown": "Yep, I don't see that one.",
      "votes": null
    },
    {
      "id": "90757",
      "postDate": "08/29/2015 00:09:31",
      "content": "<p>Good catch, @James King. </p>\n\n<p>One image had fallen through the image processing pipeline. I have uploaded the single image to the data page. You may download it now. </p>",
      "rawMarkdown": "Good catch, @James King. \r\n\r\nOne image had fallen through the image processing pipeline. I have uploaded the single image to the data page. You may download it now.",
      "votes": null
    },
    {
      "id": "91292",
      "postDate": "09/02/2015 00:56:55",
      "content": "<p>I'm having issues unzipping the file on my macbook (although it unzips just fine when transferred to a EC2 ubuntu instance);  </p>\n\n<p><em>warning [imgs.zip]:  5073236012 extra bytes at beginning or within zipfile\n  (attempting to process anyway)\nerror [imgs.zip]:  start of central directory not found;\n  zipfile corrupt.\n  (please check that you have transferred or created the zipfile in the\n  appropriate BINARY mode and that you have compiled UnZip properly)</em></p>\n\n<p>Any idea why this might be?</p>",
      "rawMarkdown": "I'm having issues unzipping the file on my macbook (although it unzips just fine when transferred to a EC2 ubuntu instance);  \r\n\r\n*warning [imgs.zip]:  5073236012 extra bytes at beginning or within zipfile\r\n  (attempting to process anyway)\r\nerror [imgs.zip]:  start of central directory not found;\r\n  zipfile corrupt.\r\n  (please check that you have transferred or created the zipfile in the\r\n  appropriate BINARY mode and that you have compiled UnZip properly)*\r\n\r\nAny idea why this might be?",
      "votes": null
    },
    {
      "id": "91681",
      "postDate": "09/06/2015 04:11:59",
      "content": "<p>Try to open with Archive utility on mac instead of unzip.</p>",
      "rawMarkdown": "Try to open with Archive utility on mac instead of unzip.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 90690,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "08/28/2015 19:48:45",
      "content": "<p>That explains it! Thanks!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90701,
      "author_name": "zhao0046",
      "author_url": "",
      "post_date": "08/28/2015 20:23:10",
      "content": "<p>I didn't see the imgs.zip file on the data page. Only the two csv files. Could you check this? Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90716,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "08/28/2015 20:56:02",
      "content": "<p>Yes, you are correct. Due to the size of imgs.zip, it's still uploading. </p>\n\n<p>I'm at 5.75 GB / 8.73 GB now. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90724,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "08/28/2015 21:30:59",
      "content": "<p>Upload is finished now. You may now download imgs.zip. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90739,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "08/28/2015 23:14:46",
      "content": "<p>Anyone else missing w_7489.jpg?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90753,
      "author_name": "ftlftw",
      "author_url": "",
      "post_date": "08/29/2015 00:07:07",
      "content": "<p>Yep, I don't see that one. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90757,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "08/29/2015 00:09:31",
      "content": "<p>Good catch, @James King. </p>\n\n<p>One image had fallen through the image processing pipeline. I have uploaded the single image to the data page. You may download it now. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91292,
      "author_name": "telser",
      "author_url": "",
      "post_date": "09/02/2015 00:56:55",
      "content": "<p>I'm having issues unzipping the file on my macbook (although it unzips just fine when transferred to a EC2 ubuntu instance);  </p>\n\n<p><em>warning [imgs.zip]:  5073236012 extra bytes at beginning or within zipfile\n  (attempting to process anyway)\nerror [imgs.zip]:  start of central directory not found;\n  zipfile corrupt.\n  (please check that you have transferred or created the zipfile in the\n  appropriate BINARY mode and that you have compiled UnZip properly)</em></p>\n\n<p>Any idea why this might be?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91681,
      "author_name": "sunilv",
      "author_url": "",
      "post_date": "09/06/2015 04:11:59",
      "content": "<p>Try to open with Archive utility on mac instead of unzip.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "90683": "Hi all,\r\n\r\nThere were some image labels lost in the processing pipeline code, and also [reported][1] that for some people (usually on Windows) the un-zipping of images files give errors. \r\n\r\nWe have now fixed both, and we are uploading new files (train.csv, sample_submission.csv, imgs.zip). The upload should be done in the next hour. Please download the files again. Previously, the zipping was done with 7z. This time it's done in normal zip (I used Keka on a Mac). \r\n\r\nSome details about the data files for those that want to double check:\r\n\r\n- There are 11469 images in the imgs folder, named w_0.jpg to w_11468.jpg. \r\n- There are 4544 images for training\r\n- There are 6925 images for test\r\n- 4544+6925=11469\r\n\r\nSorry for the inconvenience.\r\n\r\n  [1]: https://www.kaggle.com/c/noaa-right-whale-recognition/forums/t/16165/bad-zipfile-offsets/90607#post90607",
    "90690": "That explains it! Thanks!!",
    "90701": "I didn't see the imgs.zip file on the data page. Only the two csv files. Could you check this? Thank you.",
    "90716": "Yes, you are correct. Due to the size of imgs.zip, it's still uploading. \r\n\r\nI'm at 5.75 GB / 8.73 GB now.",
    "90724": "Upload is finished now. You may now download imgs.zip.",
    "90739": "Anyone else missing w_7489.jpg?",
    "90753": "Yep, I don't see that one.",
    "90757": "Good catch, @James King. \r\n\r\nOne image had fallen through the image processing pipeline. I have uploaded the single image to the data page. You may download it now.",
    "91292": "I'm having issues unzipping the file on my macbook (although it unzips just fine when transferred to a EC2 ubuntu instance);  \r\n\r\n*warning [imgs.zip]:  5073236012 extra bytes at beginning or within zipfile\r\n  (attempting to process anyway)\r\nerror [imgs.zip]:  start of central directory not found;\r\n  zipfile corrupt.\r\n  (please check that you have transferred or created the zipfile in the\r\n  appropriate BINARY mode and that you have compiled UnZip properly)*\r\n\r\nAny idea why this might be?",
    "91681": "Try to open with Archive utility on mac instead of unzip."
  },
  "source": "meta"
}