{
  "id": 35351,
  "title": "Dataset incomplete",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/35351",
  "author_name": "",
  "post_date": "2017-06-27T08:24:01.602487Z",
  "votes": 2,
  "comment_count": 11,
  "views": 2,
  "content": "<p>I tried unzipping the data recently after the download but it cited some header error for some of the files and I was only able to unzip 11,371 images. But I read in the forum that there are some 35k images in total. I have already re-downloading the data and unzipping again but there was no change. Is the dataset corrupted or am I doing something wrong? </p>",
  "messages": [
    {
      "id": "196389",
      "postDate": "06/27/2017 08:24:01",
      "content": "<p>I tried unzipping the data recently after the download but it cited some header error for some of the files and I was only able to unzip 11,371 images. But I read in the forum that there are some 35k images in total. I have already re-downloading the data and unzipping again but there was no change. Is the dataset corrupted or am I doing something wrong? </p>",
      "rawMarkdown": "I tried unzipping the data recently after the download but it cited some header error for some of the files and I was only able to unzip 11,371 images. But I read in the forum that there are some 35k images in total. I have already re-downloading the data and unzipping again but there was no change. Is the dataset corrupted or am I doing something wrong?",
      "votes": null
    },
    {
      "id": "196438",
      "postDate": "06/27/2017 11:02:12",
      "content": "<p>It is a split zipped file\nkeep all the 5 files in a single folder, and then used 7z for uncompressing, just the first file, it will automatically uncompress all the 5 part files</p>",
      "rawMarkdown": "It is a split zipped file\nkeep all the 5 files in a single folder, and then used 7z for uncompressing, just the first file, it will automatically uncompress all the 5 part files",
      "votes": null
    },
    {
      "id": "196501",
      "postDate": "06/27/2017 14:02:16",
      "content": "<p>Thanks for the reply. But when I do that, It shows \"Headers error: train\\xxx_left.jpeg\". And I don't think it's a dataset problem because I downloaded it many times just to be sure.</p>",
      "rawMarkdown": "Thanks for the reply. But when I do that, It shows \"Headers error: train\\xxx_left.jpeg\". And I don't think it's a dataset problem because I downloaded it many times just to be sure.",
      "votes": null
    },
    {
      "id": "196757",
      "postDate": "06/28/2017 03:11:15",
      "content": "<p>I have recently tried downloading the dataset multiple time from 3 different machines across 3 different internet connections. I have also tried to download directly to a AWS instance using wget as well as kaggle-cli. I received the same error as the author of this post along with only 11,371 images. </p>\n\n<p>Each time I tried downloading the dataset the md5 hashes are identical across each download but different to the hashes provided by the organiser.</p>\n\n<p>I truely believe that the problem is with the file hosted on the server. Was anyone else able to successfully download the entire dataset recently?</p>",
      "rawMarkdown": "I have recently tried downloading the dataset multiple time from 3 different machines across 3 different internet connections. I have also tried to download directly to a AWS instance using wget as well as kaggle-cli. I received the same error as the author of this post along with only 11,371 images. \n\nEach time I tried downloading the dataset the md5 hashes are identical across each download but different to the hashes provided by the organiser.\n\nI truely believe that the problem is with the file hosted on the server. Was anyone else able to successfully download the entire dataset recently?",
      "votes": null
    },
    {
      "id": "196931",
      "postDate": "06/28/2017 12:26:36",
      "content": "<p>Yes, you are correct\nThe dataset actually seems to be corrupt</p>\n\n<p>We checked the md5sum, only the part5 was correct rest all were corrupted </p>",
      "rawMarkdown": "Yes, you are correct\nThe dataset actually seems to be corrupt\n\nWe checked the md5sum, only the part5 was correct rest all were corrupted",
      "votes": null
    },
    {
      "id": "196933",
      "postDate": "06/28/2017 12:27:20",
      "content": "<p>yes, we also did the same thing, verified the checksum, I think the data from the source is corrupted</p>",
      "rawMarkdown": "yes, we also did the same thing, verified the checksum, I think the data from the source is corrupted",
      "votes": null
    },
    {
      "id": "196953",
      "postDate": "06/28/2017 13:18:59",
      "content": "<p>Thanks. Good to know I'm not the only one. I'm new here so I don't know, is there anything we can do about it?</p>",
      "rawMarkdown": "Thanks. Good to know I'm not the only one. I'm new here so I don't know, is there anything we can do about it?",
      "votes": null
    },
    {
      "id": "196973",
      "postDate": "06/28/2017 14:04:33",
      "content": "<p>Wait for the person who uploaded the dataset to respond to this unusual activity, on this 2 year competition</p>",
      "rawMarkdown": "Wait for the person who uploaded the dataset to respond to this unusual activity, on this 2 year competition",
      "votes": null
    },
    {
      "id": "196985",
      "postDate": "06/28/2017 14:30:44",
      "content": "<p>Thanks for pointing this out! It looks like it was caused by a migration error when copying to Google Cloud Storage. For now, we've temporarily reverted the files back to their original versions. Try downloading again and you should get complete files.</p>",
      "rawMarkdown": "Thanks for pointing this out! It looks like it was caused by a migration error when copying to Google Cloud Storage. For now, we've temporarily reverted the files back to their original versions. Try downloading again and you should get complete files.",
      "votes": null
    },
    {
      "id": "197894",
      "postDate": "06/30/2017 12:45:48",
      "content": "<p>Thank you. I was able to download all the data successfully!</p>",
      "rawMarkdown": "Thank you. I was able to download all the data successfully!",
      "votes": null
    },
    {
      "id": "197939",
      "postDate": "06/30/2017 15:37:00",
      "content": "<p>Thank You.</p>",
      "rawMarkdown": "Thank You.",
      "votes": null
    },
    {
      "id": "202812",
      "postDate": "07/13/2017 10:23:52",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 196438,
      "author_name": "saran95",
      "author_url": "",
      "post_date": "06/27/2017 11:02:12",
      "content": "<p>It is a split zipped file\nkeep all the 5 files in a single folder, and then used 7z for uncompressing, just the first file, it will automatically uncompress all the 5 part files</p>",
      "votes": null,
      "replies": [
        {
          "id": 196501,
          "author_name": "ricksanchez",
          "author_url": "",
          "post_date": "06/27/2017 14:02:16",
          "content": "<p>Thanks for the reply. But when I do that, It shows \"Headers error: train\\xxx_left.jpeg\". And I don't think it's a dataset problem because I downloaded it many times just to be sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 196931,
          "author_name": "saran95",
          "author_url": "",
          "post_date": "06/28/2017 12:26:36",
          "content": "<p>Yes, you are correct\nThe dataset actually seems to be corrupt</p>\n\n<p>We checked the md5sum, only the part5 was correct rest all were corrupted </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 196953,
          "author_name": "ricksanchez",
          "author_url": "",
          "post_date": "06/28/2017 13:18:59",
          "content": "<p>Thanks. Good to know I'm not the only one. I'm new here so I don't know, is there anything we can do about it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 196973,
          "author_name": "saran95",
          "author_url": "",
          "post_date": "06/28/2017 14:04:33",
          "content": "<p>Wait for the person who uploaded the dataset to respond to this unusual activity, on this 2 year competition</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 196757,
      "author_name": "aengustran",
      "author_url": "",
      "post_date": "06/28/2017 03:11:15",
      "content": "<p>I have recently tried downloading the dataset multiple time from 3 different machines across 3 different internet connections. I have also tried to download directly to a AWS instance using wget as well as kaggle-cli. I received the same error as the author of this post along with only 11,371 images. </p>\n\n<p>Each time I tried downloading the dataset the md5 hashes are identical across each download but different to the hashes provided by the organiser.</p>\n\n<p>I truely believe that the problem is with the file hosted on the server. Was anyone else able to successfully download the entire dataset recently?</p>",
      "votes": null,
      "replies": [
        {
          "id": 196933,
          "author_name": "saran95",
          "author_url": "",
          "post_date": "06/28/2017 12:27:20",
          "content": "<p>yes, we also did the same thing, verified the checksum, I think the data from the source is corrupted</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 196985,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "06/28/2017 14:30:44",
      "content": "<p>Thanks for pointing this out! It looks like it was caused by a migration error when copying to Google Cloud Storage. For now, we've temporarily reverted the files back to their original versions. Try downloading again and you should get complete files.</p>",
      "votes": null,
      "replies": [
        {
          "id": 197894,
          "author_name": "aengustran",
          "author_url": "",
          "post_date": "06/30/2017 12:45:48",
          "content": "<p>Thank you. I was able to download all the data successfully!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 197939,
          "author_name": "ricksanchez",
          "author_url": "",
          "post_date": "06/30/2017 15:37:00",
          "content": "<p>Thank You.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 202812,
      "author_name": "anikam04",
      "author_url": "",
      "post_date": "07/13/2017 10:23:52",
      "content": "",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "196389": "I tried unzipping the data recently after the download but it cited some header error for some of the files and I was only able to unzip 11,371 images. But I read in the forum that there are some 35k images in total. I have already re-downloading the data and unzipping again but there was no change. Is the dataset corrupted or am I doing something wrong?",
    "196438": "It is a split zipped file\nkeep all the 5 files in a single folder, and then used 7z for uncompressing, just the first file, it will automatically uncompress all the 5 part files",
    "196501": "Thanks for the reply. But when I do that, It shows \"Headers error: train\\xxx_left.jpeg\". And I don't think it's a dataset problem because I downloaded it many times just to be sure.",
    "196757": "I have recently tried downloading the dataset multiple time from 3 different machines across 3 different internet connections. I have also tried to download directly to a AWS instance using wget as well as kaggle-cli. I received the same error as the author of this post along with only 11,371 images. \n\nEach time I tried downloading the dataset the md5 hashes are identical across each download but different to the hashes provided by the organiser.\n\nI truely believe that the problem is with the file hosted on the server. Was anyone else able to successfully download the entire dataset recently?",
    "196931": "Yes, you are correct\nThe dataset actually seems to be corrupt\n\nWe checked the md5sum, only the part5 was correct rest all were corrupted",
    "196933": "yes, we also did the same thing, verified the checksum, I think the data from the source is corrupted",
    "196953": "Thanks. Good to know I'm not the only one. I'm new here so I don't know, is there anything we can do about it?",
    "196973": "Wait for the person who uploaded the dataset to respond to this unusual activity, on this 2 year competition",
    "196985": "Thanks for pointing this out! It looks like it was caused by a migration error when copying to Google Cloud Storage. For now, we've temporarily reverted the files back to their original versions. Try downloading again and you should get complete files.",
    "197894": "Thank you. I was able to download all the data successfully!",
    "197939": "Thank You.",
    "202812": ""
  },
  "source": "meta"
}