{
  "id": 164941,
  "title": "Corrupt files ?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/164941",
  "author_name": "",
  "post_date": "2020-07-08T01:08:52.768849600Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am unable to read 32 files - I have posted the code here <a href=\"https://www.kaggle.com/rashmibanthia/corrupt-files\">https://www.kaggle.com/rashmibanthia/corrupt-files</a> </p>\n\n<p>Are these files corrupt or I am missing something ?  (Rest of the files are working) </p>",
  "messages": [
    {
      "id": "919559",
      "postDate": "07/08/2020 01:08:52",
      "content": "<p>I am unable to read 32 files - I have posted the code here <a href=\"https://www.kaggle.com/rashmibanthia/corrupt-files\">https://www.kaggle.com/rashmibanthia/corrupt-files</a> </p>\n\n<p>Are these files corrupt or I am missing something ?  (Rest of the files are working) </p>",
      "rawMarkdown": "I am unable to read 32 files - I have posted the code here https://www.kaggle.com/rashmibanthia/corrupt-files \n\nAre these files corrupt or I am missing something ?  (Rest of the files are working)",
      "votes": null
    },
    {
      "id": "919583",
      "postDate": "07/08/2020 01:50:23",
      "content": "<p>Thanks for the list of corrupted files, I did get a similar error when looking at the data.  The error states that it expects an array of 524288 bytes but got 322744 bytes. We can calculate the number of bytes by looking at the header info.</p>\n\n<p>For example for this image:'../input/osic-pulmonary-fibrosis-progression/train/ID00011637202177653955184/31.dcm'</p>\n\n<p>The number of bytes is calculated by this formula: Rows * Columns * Number of frames * Samples per pixel * (bits_allocated/8)</p>\n\n<p>512 * 512 * 1 * 1 * (16/8) = 524,288 bytes</p>\n\n<p>So the header is expecting 524,288 bytes but is only getting 322, 744 bytes. It seems the Pixel Data is corrupted in these images.</p>",
      "rawMarkdown": "Thanks for the list of corrupted files, I did get a similar error when looking at the data.  The error states that it expects an array of 524288 bytes but got 322744 bytes. We can calculate the number of bytes by looking at the header info.\n\nFor example for this image:'../input/osic-pulmonary-fibrosis-progression/train/ID00011637202177653955184/31.dcm'\n\nThe number of bytes is calculated by this formula: Rows * Columns * Number of frames * Samples per pixel * (bits_allocated/8)\n\n512 * 512 * 1 * 1 * (16/8) = 524,288 bytes\n\nSo the header is expecting 524,288 bytes but is only getting 322, 744 bytes. It seems the Pixel Data is corrupted in these images.",
      "votes": null
    },
    {
      "id": "919641",
      "postDate": "07/08/2020 03:03:23",
      "content": "<p>Thanks for this explanation - all files for patient ID00011637202177653955184 seems corrupt, and this file - ID00052637202186188008618/4.dcm</p>",
      "rawMarkdown": "Thanks for this explanation - all files for patient ID00011637202177653955184 seems corrupt, and this file - ID00052637202186188008618/4.dcm",
      "votes": null
    },
    {
      "id": "920291",
      "postDate": "07/08/2020 13:33:48",
      "content": "<p>Thanks for reporting. We're investigating this.</p>",
      "rawMarkdown": "Thanks for reporting. We're investigating this.",
      "votes": null
    },
    {
      "id": "920335",
      "postDate": "07/08/2020 14:10:27",
      "content": "<p>This file: ID00052637202186188008618/4.dcm displays in a Dicom Viewer (K-Pacs). Based on the DICOM tags, it looks like a compressed JPEG. It might be a \"screen capture\". The Window/Level is a little off.  I suspect gdcm will be able to decompress it, but I don't have that package loaded (if anybody has a trusted source for an easy install in Python on Windows without Conda, please share).</p>",
      "rawMarkdown": "This file: ID00052637202186188008618/4.dcm displays in a Dicom Viewer (K-Pacs). Based on the DICOM tags, it looks like a compressed JPEG. It might be a \"screen capture\". The Window/Level is a little off.  I suspect gdcm will be able to decompress it, but I don't have that package loaded (if anybody has a trusted source for an easy install in Python on Windows without Conda, please share).",
      "votes": null
    },
    {
      "id": "924633",
      "postDate": "07/11/2020 14:52:01",
      "content": "<p><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 919583,
      "author_name": "avirdee",
      "author_url": "",
      "post_date": "07/08/2020 01:50:23",
      "content": "<p>Thanks for the list of corrupted files, I did get a similar error when looking at the data.  The error states that it expects an array of 524288 bytes but got 322744 bytes. We can calculate the number of bytes by looking at the header info.</p>\n\n<p>For example for this image:'../input/osic-pulmonary-fibrosis-progression/train/ID00011637202177653955184/31.dcm'</p>\n\n<p>The number of bytes is calculated by this formula: Rows * Columns * Number of frames * Samples per pixel * (bits_allocated/8)</p>\n\n<p>512 * 512 * 1 * 1 * (16/8) = 524,288 bytes</p>\n\n<p>So the header is expecting 524,288 bytes but is only getting 322, 744 bytes. It seems the Pixel Data is corrupted in these images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 919641,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "07/08/2020 03:03:23",
          "content": "<p>Thanks for this explanation - all files for patient ID00011637202177653955184 seems corrupt, and this file - ID00052637202186188008618/4.dcm</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920291,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "07/08/2020 13:33:48",
          "content": "<p>Thanks for reporting. We're investigating this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920335,
          "author_name": "richardepstein",
          "author_url": "",
          "post_date": "07/08/2020 14:10:27",
          "content": "<p>This file: ID00052637202186188008618/4.dcm displays in a Dicom Viewer (K-Pacs). Based on the DICOM tags, it looks like a compressed JPEG. It might be a \"screen capture\". The Window/Level is a little off.  I suspect gdcm will be able to decompress it, but I don't have that package loaded (if anybody has a trusted source for an easy install in Python on Windows without Conda, please share).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 924633,
      "author_name": "ahmedhshahin",
      "author_url": "",
      "post_date": "07/11/2020 14:52:01",
      "content": "<p><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "919559": "I am unable to read 32 files - I have posted the code here https://www.kaggle.com/rashmibanthia/corrupt-files \n\nAre these files corrupt or I am missing something ?  (Rest of the files are working)",
    "919583": "Thanks for the list of corrupted files, I did get a similar error when looking at the data.  The error states that it expects an array of 524288 bytes but got 322744 bytes. We can calculate the number of bytes by looking at the header info.\n\nFor example for this image:'../input/osic-pulmonary-fibrosis-progression/train/ID00011637202177653955184/31.dcm'\n\nThe number of bytes is calculated by this formula: Rows * Columns * Number of frames * Samples per pixel * (bits_allocated/8)\n\n512 * 512 * 1 * 1 * (16/8) = 524,288 bytes\n\nSo the header is expecting 524,288 bytes but is only getting 322, 744 bytes. It seems the Pixel Data is corrupted in these images.",
    "919641": "Thanks for this explanation - all files for patient ID00011637202177653955184 seems corrupt, and this file - ID00052637202186188008618/4.dcm",
    "920291": "Thanks for reporting. We're investigating this.",
    "920335": "This file: ID00052637202186188008618/4.dcm displays in a Dicom Viewer (K-Pacs). Based on the DICOM tags, it looks like a compressed JPEG. It might be a \"screen capture\". The Window/Level is a little off.  I suspect gdcm will be able to decompress it, but I don't have that package loaded (if anybody has a trusted source for an easy install in Python on Windows without Conda, please share).",
    "924633": "https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723"
  },
  "source": "meta"
}