{
  "id": 16165,
  "title": "Bad zipfile offsets",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16165",
  "author_name": "James King",
  "post_date": "2015-08-28T04:09:42.983000",
  "votes": 1,
  "comment_count": 8,
  "views": 1318,
  "content": "<p>Is anyone else getting bad zipfile offsets?</p>\n\n<pre><code>file #9566:  bad zipfile offset (lseek):  12112568320\nfile #9567:  bad zipfile offset (lseek):  12112986112\n</code></pre>\n\n<p>etc.</p>\n\n<p>My md5sum for imgs.zip is 645772cb6592bbf904d1a2ed054046f9 </p>",
  "messages": [
    {
      "id": 90607,
      "postDate": "2015-08-28T04:09:42.983Z",
      "content": "<p>Is anyone else getting bad zipfile offsets?</p>\n\n<pre><code>file #9566:  bad zipfile offset (lseek):  12112568320\nfile #9567:  bad zipfile offset (lseek):  12112986112\n</code></pre>\n\n<p>etc.</p>\n\n<p>My md5sum for imgs.zip is 645772cb6592bbf904d1a2ed054046f9 </p>",
      "rawMarkdown": "Is anyone else getting bad zipfile offsets?\r\n\r\n    file #9566:  bad zipfile offset (lseek):  12112568320\r\n    file #9567:  bad zipfile offset (lseek):  12112986112\r\n\r\n \r\netc.\r\n\r\nMy md5sum for imgs.zip is 645772cb6592bbf904d1a2ed054046f9 ",
      "votes": 1
    },
    {
      "id": 90856,
      "postDate": "2015-08-29T16:00:02.870Z",
      "content": "<p>Got it jar the zip and you put up the missing image.</p>\n\n<p>Thank you Wendy!</p>",
      "rawMarkdown": "Got it jar the zip and you put up the missing image.\r\n\r\nThank you Wendy!"
    },
    {
      "id": 90673,
      "postDate": "2015-08-28T17:07:08.210Z",
      "content": "<p>Hi all,</p>\n\n<p>There are 11468 images in total. The zipped file fails for some when unzipping, I'll upload a correct version soon (today). Sorry for the inconvenience. </p>\n\n<p>EDIT: correction, there are 11469 images named from w_0.jpg to w_11468.jpg</p>",
      "rawMarkdown": "Hi all,\r\n\r\nThere are 11468 images in total. The zipped file fails for some when unzipping, I'll upload a correct version soon (today). Sorry for the inconvenience. \r\n\r\nEDIT: correction, there are 11469 images named from w_0.jpg to w_11468.jpg"
    },
    {
      "id": 90668,
      "postDate": "2015-08-28T16:37:00.103Z",
      "content": "<p>I got same md5sum as James which is 645772cb6592bbf904d1a2ed054046f9. There are 6276 images if I use unzip imgs.zip. Then I open the archive with jar -xf imgs.zip, I got 11468 images. Maybe Admin should clear this confusion for us.</p>",
      "rawMarkdown": "I got same md5sum as James which is 645772cb6592bbf904d1a2ed054046f9. There are 6276 images if I use unzip imgs.zip. Then I open the archive with jar -xf imgs.zip, I got 11468 images. Maybe Admin should clear this confusion for us.\r\n\r\n "
    },
    {
      "id": 90654,
      "postDate": "2015-08-28T15:19:49.543Z",
      "content": "<p>It looks like there should be 4156 training images and 6915 testing images based on the number of rows in train.csv and sample_submission.csv, so a total of 11071. I was able to open the archive with</p>\n\n<pre><code>jar -xf imgs.zip\n</code></pre>\n\n<p>but that gives me 11468 images. I think we need a corrected dataset.</p>",
      "rawMarkdown": "It looks like there should be 4156 training images and 6915 testing images based on the number of rows in train.csv and sample_submission.csv, so a total of 11071. I was able to open the archive with\r\n\r\n    jar -xf imgs.zip\r\n\r\nbut that gives me 11468 images. I think we need a corrected dataset."
    },
    {
      "id": 90625,
      "postDate": "2015-08-28T10:23:55.073Z",
      "content": "<p>Same here.</p>",
      "rawMarkdown": "Same here."
    },
    {
      "id": 90620,
      "postDate": "2015-08-28T07:58:54.383Z",
      "content": "<p>Yes, I've tried downloading the zip twice now with the same results and the same md5sum.</p>\n\n<p>I also get the following results when I unzip the file:</p>\n\n<p>$ ls imgs/ | wc -l</p>\n\n<p>6276</p>\n\n<p>6K images seem a little light given the numbering scheme</p>",
      "rawMarkdown": "Yes, I've tried downloading the zip twice now with the same results and the same md5sum.\r\n\r\nI also get the following results when I unzip the file:\r\n\r\n$ ls imgs/ | wc -l\r\n\r\n   6276\r\n\r\n6K images seem a little light given the numbering scheme"
    },
    {
      "id": 90702,
      "postDate": "2015-08-28T20:23:12.527Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 90674,
      "postDate": "2015-08-28T17:11:38.717Z",
      "content": "<p>Thank Wendy for the clarification. </p>",
      "rawMarkdown": "Thank Wendy for the clarification. "
    }
  ],
  "comments": [
    {
      "id": 90856,
      "author_name": "Telesphore",
      "author_url": "",
      "post_date": "2015-08-29T16:00:02.870000",
      "content": "<p>Got it jar the zip and you put up the missing image.</p>\n\n<p>Thank you Wendy!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 90673,
      "author_name": "Wendy Kan",
      "author_url": "",
      "post_date": "2015-08-28T17:07:08.210000",
      "content": "<p>Hi all,</p>\n\n<p>There are 11468 images in total. The zipped file fails for some when unzipping, I'll upload a correct version soon (today). Sorry for the inconvenience. </p>\n\n<p>EDIT: correction, there are 11469 images named from w_0.jpg to w_11468.jpg</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 90668,
      "author_name": "Jianmin Sun",
      "author_url": "",
      "post_date": "2015-08-28T16:37:00.103000",
      "content": "<p>I got same md5sum as James which is 645772cb6592bbf904d1a2ed054046f9. There are 6276 images if I use unzip imgs.zip. Then I open the archive with jar -xf imgs.zip, I got 11468 images. Maybe Admin should clear this confusion for us.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 90654,
      "author_name": "James King",
      "author_url": "",
      "post_date": "2015-08-28T15:19:49.543000",
      "content": "<p>It looks like there should be 4156 training images and 6915 testing images based on the number of rows in train.csv and sample_submission.csv, so a total of 11071. I was able to open the archive with</p>\n\n<pre><code>jar -xf imgs.zip\n</code></pre>\n\n<p>but that gives me 11468 images. I think we need a corrected dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 90625,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-08-28T10:23:55.073000",
      "content": "<p>Same here.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 90620,
      "author_name": "Telesphore",
      "author_url": "",
      "post_date": "2015-08-28T07:58:54.383000",
      "content": "<p>Yes, I've tried downloading the zip twice now with the same results and the same md5sum.</p>\n\n<p>I also get the following results when I unzip the file:</p>\n\n<p>$ ls imgs/ | wc -l</p>\n\n<p>6276</p>\n\n<p>6K images seem a little light given the numbering scheme</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 90702,
      "author_name": "",
      "author_url": "",
      "post_date": "2015-08-28T20:23:12.527000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 90674,
      "author_name": "Jianmin Sun",
      "author_url": "",
      "post_date": "2015-08-28T17:11:38.717000",
      "content": "<p>Thank Wendy for the clarification. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "90607": "Is anyone else getting bad zipfile offsets?\r\n\r\n    file #9566:  bad zipfile offset (lseek):  12112568320\r\n    file #9567:  bad zipfile offset (lseek):  12112986112\r\n\r\n \r\netc.\r\n\r\nMy md5sum for imgs.zip is 645772cb6592bbf904d1a2ed054046f9 ",
    "90856": "Got it jar the zip and you put up the missing image.\r\n\r\nThank you Wendy!",
    "90673": "Hi all,\r\n\r\nThere are 11468 images in total. The zipped file fails for some when unzipping, I'll upload a correct version soon (today). Sorry for the inconvenience. \r\n\r\nEDIT: correction, there are 11469 images named from w_0.jpg to w_11468.jpg",
    "90668": "I got same md5sum as James which is 645772cb6592bbf904d1a2ed054046f9. There are 6276 images if I use unzip imgs.zip. Then I open the archive with jar -xf imgs.zip, I got 11468 images. Maybe Admin should clear this confusion for us.\r\n\r\n ",
    "90654": "It looks like there should be 4156 training images and 6915 testing images based on the number of rows in train.csv and sample_submission.csv, so a total of 11071. I was able to open the archive with\r\n\r\n    jar -xf imgs.zip\r\n\r\nbut that gives me 11468 images. I think we need a corrected dataset.",
    "90625": "Same here.",
    "90620": "Yes, I've tried downloading the zip twice now with the same results and the same md5sum.\r\n\r\nI also get the following results when I unzip the file:\r\n\r\n$ ls imgs/ | wc -l\r\n\r\n   6276\r\n\r\n6K images seem a little light given the numbering scheme",
    "90702": "",
    "90674": "Thank Wendy for the clarification. "
  }
}