{
  "id": 64346,
  "title": "Reading images in R",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/64346",
  "author_name": "",
  "post_date": "2018-08-28T12:49:55.056192800Z",
  "votes": 3,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I tried to read the images in R using\nreadDICOMFile(\"../input/stage_1_train_images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\")\nbut all I get is an error:\nError in parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, : Number of bytes in PixelData not specified; guess = 1</p>\n\n<p>How do I read the data correctly, i.e. how do I get the images and all the metadata included in that file?</p>\n\n<p>Regards\nDaniel</p>",
  "messages": [
    {
      "id": "376996",
      "postDate": "08/28/2018 12:49:55",
      "content": "<p>Hi,</p>\n\n<p>I tried to read the images in R using\nreadDICOMFile(\"../input/stage_1_train_images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\")\nbut all I get is an error:\nError in parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, : Number of bytes in PixelData not specified; guess = 1</p>\n\n<p>How do I read the data correctly, i.e. how do I get the images and all the metadata included in that file?</p>\n\n<p>Regards\nDaniel</p>",
      "rawMarkdown": "Hi,\n\nI tried to read the images in R using\nreadDICOMFile(\"../input/stage_1_train_images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\")\nbut all I get is an error:\nError in parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, : Number of bytes in PixelData not specified; guess = 1\n\nHow do I read the data correctly, i.e. how do I get the images and all the metadata included in that file?\n\nRegards\nDaniel",
      "votes": null
    },
    {
      "id": "378491",
      "postDate": "08/30/2018 09:12:06",
      "content": "<p>push...</p>",
      "rawMarkdown": "push...",
      "votes": null
    },
    {
      "id": "380337",
      "postDate": "09/02/2018 11:26:27",
      "content": "<p>Why there is not much people to use R for Kaggle competitions? I am a R user, looking for more R discussion </p>",
      "rawMarkdown": "Why there is not much people to use R for Kaggle competitions? I am a R user, looking for more R discussion",
      "votes": null
    },
    {
      "id": "380340",
      "postDate": "09/02/2018 11:32:21",
      "content": "<p>Guys, it's still a good time to learn python ;)</p>",
      "rawMarkdown": "Guys, it's still a good time to learn python ;)",
      "votes": null
    },
    {
      "id": "380477",
      "postDate": "09/02/2018 18:59:17",
      "content": "<blockquote>\n  <p><strong>Godfrey Cheung wrote</strong></p>\n  \n  <blockquote>\n    <p>Why there is not much people to use R for Kaggle competitions? I am a R user, looking for more R discussion </p>\n  </blockquote>\n</blockquote>\n\n<p>You're not likely to find many R kernels and discussions on deep learning and particularly computer vision competitions. </p>\n\n<p>As said by Henrique,  it's not too late to learn python ;)</p>",
      "rawMarkdown": "&gt; **Godfrey Cheung wrote**\n&gt; \n&gt; &gt; Why there is not much people to use R for Kaggle competitions? I am a R user, looking for more R discussion \n\nYou're not likely to find many R kernels and discussions on deep learning and particularly computer vision competitions. \n\nAs said by Henrique,  it's not too late to learn python ;)",
      "votes": null
    },
    {
      "id": "380812",
      "postDate": "09/03/2018 13:19:36",
      "content": "<p>;)</p>",
      "rawMarkdown": ";)",
      "votes": null
    },
    {
      "id": "381077",
      "postDate": "09/04/2018 03:00:41",
      "content": "<p>Hi Daniel,\nI'm in the same boat as you (trying to complete this very interesting exercise with R). The first 128 bits of each DCM file are zero's. This can be determined by using the debug=TRUE argument. </p>\n\n<p>test1 &lt;- readDICOMFile(\"/Users/davem/OneDrive/Desktop/R Prog Files/MSDS692/train images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", debug=TRUE)</p>\n\n<p>*# First 128 bytes of DICOM header =\n[1] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[42] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[83] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[124] 00 00 00 00 00</p>\n\n<h1>DICM = TRUE*</h1>\n\n<p>You can use the 'boffset' argument to skip the first 128 bits but this only returns the header information.</p>\n\n<p>test2 &lt;- readDICOMFile(\"/Users/davem/OneDrive/Desktop/R Prog Files/MSDS692/train images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", boffset = 128)\nnames(test2)</p>\n\n<p>[1] \"hdr\" \"img\"</p>\n\n<p>Unfortunately, the image data is NULL. I'm working through that. If you have any suggestions, I would appreciate them.</p>\n\n<p>Dave</p>",
      "rawMarkdown": "Hi Daniel,\nI'm in the same boat as you (trying to complete this very interesting exercise with R). The first 128 bits of each DCM file are zero's. This can be determined by using the debug=TRUE argument. \n\ntest1 &lt;- readDICOMFile(\"/Users/davem/OneDrive/Desktop/R Prog Files/MSDS692/train images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", debug=TRUE)\n\n*# First 128 bytes of DICOM header =\n[1] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[42] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[83] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[124] 00 00 00 00 00\n# DICM = TRUE*\n\nYou can use the 'boffset' argument to skip the first 128 bits but this only returns the header information.\n\ntest2 &lt;- readDICOMFile(\"/Users/davem/OneDrive/Desktop/R Prog Files/MSDS692/train images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", boffset = 128)\nnames(test2)\n\n[1] \"hdr\" \"img\"\n\nUnfortunately, the image data is NULL. I'm working through that. If you have any suggestions, I would appreciate them.\n\nDave",
      "votes": null
    },
    {
      "id": "381872",
      "postDate": "09/05/2018 10:34:55",
      "content": "<p>Hi Dave,</p>\n\n<p>interesting approach. Unfortunately I still get the same error. First I get those 128 \"00\"s and then some content or information, but in the end I get:</p>\n\n<p>Error in parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, : Number of bytes in PixelData not specified; guess = 1\nTraceback:</p>\n\n<ol>\n<li>readDICOMFile(\"../input/stage_1_train_images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", \n.     debug = TRUE)</li>\n<li>parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, \n.     flipud)</li>\n<li>stop(paste(\"Number of bytes in PixelData not specified; guess =\", \n.     guess))</li>\n</ol>\n\n<p>I think, this task is not meant to be solved in R....but I have to admit that I find Python somehow confusing... ;-)</p>\n\n<p>Daniel</p>",
      "rawMarkdown": "Hi Dave,\n\ninteresting approach. Unfortunately I still get the same error. First I get those 128 \"00\"s and then some content or information, but in the end I get:\n\nError in parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, : Number of bytes in PixelData not specified; guess = 1\nTraceback:\n\n1. readDICOMFile(\"../input/stage_1_train_images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", \n .     debug = TRUE)\n2. parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, \n .     flipud)\n3. stop(paste(\"Number of bytes in PixelData not specified; guess =\", \n .     guess))\n\nI think, this task is not meant to be solved in R....but I have to admit that I find Python somehow confusing... ;-)\n\nDaniel",
      "votes": null
    },
    {
      "id": "383051",
      "postDate": "09/07/2018 15:48:22",
      "content": "<p>I think I will try to convert the images to JPG or PNG using Python and extract the metadata with R. This is not very nice but it should work.\nThat of course leads to a new structure of the sets as the image itself and the metadata are stored in two different places...</p>",
      "rawMarkdown": "I think I will try to convert the images to JPG or PNG using Python and extract the metadata with R. This is not very nice but it should work.\nThat of course leads to a new structure of the sets as the image itself and the metadata are stored in two different places...",
      "votes": null
    },
    {
      "id": "383133",
      "postDate": "09/07/2018 18:46:22",
      "content": "<p>David, would it be possible to expose in this thread the DICOM header fields from the test2$hdr data.frame?  There is a lot of information in the header, something there may help diagnose the issue.  The <code>pixelData = FALSE</code> setting will ignore the image.  </p>\n\n<p>Brandon</p>",
      "rawMarkdown": "David, would it be possible to expose in this thread the DICOM header fields from the test2$hdr data.frame?  There is a lot of information in the header, something there may help diagnose the issue.  The `pixelData = FALSE` setting will ignore the image.  \n\nBrandon",
      "votes": null
    },
    {
      "id": "385183",
      "postDate": "09/10/2018 13:59:52",
      "content": "<p>I'm converting the files now with irfanview and a plugin. For now every picture has the size of about 1MB...and I cannot upload more than 50 files at one time to add as external data. So I think I will run some analysis on my computer...unless someone finds a way to view the images in R</p>",
      "rawMarkdown": "I'm converting the files now with irfanview and a plugin. For now every picture has the size of about 1MB...and I cannot upload more than 50 files at one time to add as external data. So I think I will run some analysis on my computer...unless someone finds a way to view the images in R",
      "votes": null
    },
    {
      "id": "385835",
      "postDate": "09/11/2018 17:06:11",
      "content": "<p>Hi DanielAC:\n you can easily read files with pydicom. I talked to Brandon Whitcher, and told me he is getting back into the medical image analysis field again,  but it will not happen for at least a few months. Hopefully he will upgrade <strong>oro.dicom</strong> which is a very good tool. He suggested using <strong>pydicom</strong> or <strong>DCMTK</strong>.  Using pydicom within R is very simple if you use library <strong>reticulate</strong>. You have to download pyton or anaconda and then you can install with pip in python or install in python within R. Here it goes an example; I downloaded anaconda:</p>\n\n<p>library(data.table)</p>\n\n<p>library(imager) #very good package</p>\n\n<p>library(oro.dicom)</p>\n\n<p>library(reticulate)</p>\n\n<h1>install py packages, including pydicom. Only neded one time to install in python</h1>\n\n<p>if(F){\n   conda_version()</p>\n\n<p>py_available(initialize = T)</p>\n\n<p>py_numpy_available(initialize = FALSE)</p>\n\n<p>py_install(\"pandas\")</p>\n\n<p>py_install(\"gdcm\")</p>\n\n<p>py_install(\"pydicom\")</p>\n\n<h1>some others you might need</h1>\n\n<p>py_install(\"keras\")</p>\n\n<p>py_install(\"tensorflow\")</p>\n\n<p>py_install(\"matplotlib\")\n} </p>\n\n<h3>#</h3>\n\n<p>conda_version()</p>\n\n<h1>import pydicom and use it</h1>\n\n<p>pydcm=import(\"pydicom\")</p>\n\n<p>trainfolder=paste0(getwd(),\"/input/stage_1_train_images/\")</p>\n\n<p>trainpatno=list.files(trainfolder)</p>\n\n<p>idx=1</p>\n\n<p>fname=paste0(trainfolder,trainpatno[idx])</p>\n\n<p>hdr=readDICOMFile(fname,pixelData=F)$hdr</p>\n\n<p>img=pydcm$read_file(fname)</p>\n\n<p>atribs=py_list_attributes(img)</p>\n\n<p>img=t(img$pixel_array)</p>\n\n<p>as.cimg(img)%&gt;%plot</p>\n\n<p>I hope it helps.</p>",
      "rawMarkdown": "Hi DanielAC:\n you can easily read files with pydicom. I talked to Brandon Whitcher, and told me he is getting back into the medical image analysis field again,  but it will not happen for at least a few months. Hopefully he will upgrade **oro.dicom** which is a very good tool. He suggested using **pydicom** or **DCMTK**.  Using pydicom within R is very simple if you use library **reticulate**. You have to download pyton or anaconda and then you can install with pip in python or install in python within R. Here it goes an example; I downloaded anaconda:\n\nlibrary(data.table)\n\nlibrary(imager) #very good package\n\nlibrary(oro.dicom)\n\nlibrary(reticulate)\n\n#install py packages, including pydicom. Only neded one time to install in python\n\nif(F){\n   conda_version()\n\n   py_available(initialize = T)\n\n   py_numpy_available(initialize = FALSE)\n\n   py_install(\"pandas\")\n\n   py_install(\"gdcm\")\n\n   py_install(\"pydicom\")\n\n#some others you might need\n\n   py_install(\"keras\")\n\n   py_install(\"tensorflow\")\n\n   py_install(\"matplotlib\")\n} \n####\nconda_version()\n#import pydicom and use it\npydcm=import(\"pydicom\")\n\ntrainfolder=paste0(getwd(),\"/input/stage_1_train_images/\")\n\ntrainpatno=list.files(trainfolder)\n\nidx=1\n\nfname=paste0(trainfolder,trainpatno[idx])\n\nhdr=readDICOMFile(fname,pixelData=F)$hdr\n\nimg=pydcm$read_file(fname)\n\natribs=py_list_attributes(img)\n\n\nimg=t(img$pixel_array)\n\nas.cimg(img)%&gt;%plot\n\n\n\nI hope it helps.",
      "votes": null
    },
    {
      "id": "385842",
      "postDate": "09/11/2018 17:22:08",
      "content": "<p>Hi Brandon! I posted some ideas you gave me the other day below. I didn´t know you were in Kaggle. Thanks again.</p>",
      "rawMarkdown": "Hi Brandon! I posted some ideas you gave me the other day below. I didn´t know you were in Kaggle. Thanks again.",
      "votes": null
    },
    {
      "id": "385869",
      "postDate": "09/11/2018 18:05:09",
      "content": "<p>Thanks, Guillermo.  I haven't visited the Kaggle site in years but your DICOM question piqued my interest :)</p>",
      "rawMarkdown": "Thanks, Guillermo.  I haven't visited the Kaggle site in years but your DICOM question piqued my interest :)",
      "votes": null
    },
    {
      "id": "385884",
      "postDate": "09/11/2018 18:19:46",
      "content": "<p>I have also played around a little bit with the data and here is an example of using DCMTK to uncompress all of the DICOM files provided in this challenge.  Then they can be read into R via <strong>oro.dicom</strong> no problem.  My suggestion would be to stack the images into a 3D array and convert the array into a NIfTI object using <strong>oro.nifti</strong>.  This would allow for easier manipulation both in R and via the filesystem.  Some example code:</p>\n\n<pre><code>library(oro.dicom)\n\nrsna &lt;- \"~/Kaggle/rsna_pneumonia_detection_challenge\"\ntestDir &lt;- \"stage_1_test_images\"\ndir.create(paste(file.path(rsna, testDir), \"uncompressed\", sep = \"_\"))\ntrainDir &lt;- sub(\"test\", \"train\", testDir)\ndir.create(paste(file.path(rsna, trainDir), \"uncompressed\", sep = \"_\"))\ndcmFile &lt;- file.path(rsna, trainDir, \"00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\")\ndcm &lt;- readDICOMFile(dcmFile, debug = TRUE) # fails\n# system(\"brew install dcmtk\")\ncmd &lt;- paste(\"dcmdjpeg\",\n             dcmFile,\n             sub(\"images\", \"images_uncompressed\", dcmFile))\nsystem(cmd)\ndcmFile &lt;- file.path(sub(\"images\", \"images_uncompressed\", dcmFile))\ndcm &lt;- readDICOMFile(dcmFile) # succeeds\nimage(t(dcm$img), col = grey(0:64/64), axes = FALSE)\n# now do this for all files in the test and training directories\ndcmList &lt;- c(list.files(file.path(rsna, testDir), full.names = TRUE),\n             list.files(file.path(rsna, trainDir), full.names = TRUE))\nfor (f in dcmList) {\n  cmd &lt;- paste(\"dcmdjpeg\", f, sub(\"images\", \"images_uncompressed\", f))\n  system(cmd)\n}\n</code></pre>",
      "rawMarkdown": "I have also played around a little bit with the data and here is an example of using DCMTK to uncompress all of the DICOM files provided in this challenge.  Then they can be read into R via **oro.dicom** no problem.  My suggestion would be to stack the images into a 3D array and convert the array into a NIfTI object using **oro.nifti**.  This would allow for easier manipulation both in R and via the filesystem.  Some example code:\n\n    library(oro.dicom)\n    \n    rsna &lt;- \"~/Kaggle/rsna_pneumonia_detection_challenge\"\n    testDir &lt;- \"stage_1_test_images\"\n    dir.create(paste(file.path(rsna, testDir), \"uncompressed\", sep = \"_\"))\n    trainDir &lt;- sub(\"test\", \"train\", testDir)\n    dir.create(paste(file.path(rsna, trainDir), \"uncompressed\", sep = \"_\"))\n    dcmFile &lt;- file.path(rsna, trainDir, \"00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\")\n    dcm &lt;- readDICOMFile(dcmFile, debug = TRUE) # fails\n    # system(\"brew install dcmtk\")\n    cmd &lt;- paste(\"dcmdjpeg\",\n                 dcmFile,\n                 sub(\"images\", \"images_uncompressed\", dcmFile))\n    system(cmd)\n    dcmFile &lt;- file.path(sub(\"images\", \"images_uncompressed\", dcmFile))\n    dcm &lt;- readDICOMFile(dcmFile) # succeeds\n    image(t(dcm$img), col = grey(0:64/64), axes = FALSE)\n    # now do this for all files in the test and training directories\n    dcmList &lt;- c(list.files(file.path(rsna, testDir), full.names = TRUE),\n                 list.files(file.path(rsna, trainDir), full.names = TRUE))\n    for (f in dcmList) {\n      cmd &lt;- paste(\"dcmdjpeg\", f, sub(\"images\", \"images_uncompressed\", f))\n      system(cmd)\n    }",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 378491,
      "author_name": "danielac",
      "author_url": "",
      "post_date": "08/30/2018 09:12:06",
      "content": "<p>push...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 380337,
      "author_name": "godfreycheungowl",
      "author_url": "",
      "post_date": "09/02/2018 11:26:27",
      "content": "<p>Why there is not much people to use R for Kaggle competitions? I am a R user, looking for more R discussion </p>",
      "votes": null,
      "replies": [
        {
          "id": 380340,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "09/02/2018 11:32:21",
          "content": "<p>Guys, it's still a good time to learn python ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 380477,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "09/02/2018 18:59:17",
          "content": "<blockquote>\n  <p><strong>Godfrey Cheung wrote</strong></p>\n  \n  <blockquote>\n    <p>Why there is not much people to use R for Kaggle competitions? I am a R user, looking for more R discussion </p>\n  </blockquote>\n</blockquote>\n\n<p>You're not likely to find many R kernels and discussions on deep learning and particularly computer vision competitions. </p>\n\n<p>As said by Henrique,  it's not too late to learn python ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 380812,
          "author_name": "godfreycheungowl",
          "author_url": "",
          "post_date": "09/03/2018 13:19:36",
          "content": "<p>;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 381077,
      "author_name": "davelovesdata",
      "author_url": "",
      "post_date": "09/04/2018 03:00:41",
      "content": "<p>Hi Daniel,\nI'm in the same boat as you (trying to complete this very interesting exercise with R). The first 128 bits of each DCM file are zero's. This can be determined by using the debug=TRUE argument. </p>\n\n<p>test1 &lt;- readDICOMFile(\"/Users/davem/OneDrive/Desktop/R Prog Files/MSDS692/train images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", debug=TRUE)</p>\n\n<p>*# First 128 bytes of DICOM header =\n[1] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[42] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[83] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[124] 00 00 00 00 00</p>\n\n<h1>DICM = TRUE*</h1>\n\n<p>You can use the 'boffset' argument to skip the first 128 bits but this only returns the header information.</p>\n\n<p>test2 &lt;- readDICOMFile(\"/Users/davem/OneDrive/Desktop/R Prog Files/MSDS692/train images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", boffset = 128)\nnames(test2)</p>\n\n<p>[1] \"hdr\" \"img\"</p>\n\n<p>Unfortunately, the image data is NULL. I'm working through that. If you have any suggestions, I would appreciate them.</p>\n\n<p>Dave</p>",
      "votes": null,
      "replies": [
        {
          "id": 381872,
          "author_name": "danielac",
          "author_url": "",
          "post_date": "09/05/2018 10:34:55",
          "content": "<p>Hi Dave,</p>\n\n<p>interesting approach. Unfortunately I still get the same error. First I get those 128 \"00\"s and then some content or information, but in the end I get:</p>\n\n<p>Error in parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, : Number of bytes in PixelData not specified; guess = 1\nTraceback:</p>\n\n<ol>\n<li>readDICOMFile(\"../input/stage_1_train_images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", \n.     debug = TRUE)</li>\n<li>parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, \n.     flipud)</li>\n<li>stop(paste(\"Number of bytes in PixelData not specified; guess =\", \n.     guess))</li>\n</ol>\n\n<p>I think, this task is not meant to be solved in R....but I have to admit that I find Python somehow confusing... ;-)</p>\n\n<p>Daniel</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 383051,
          "author_name": "danielac",
          "author_url": "",
          "post_date": "09/07/2018 15:48:22",
          "content": "<p>I think I will try to convert the images to JPG or PNG using Python and extract the metadata with R. This is not very nice but it should work.\nThat of course leads to a new structure of the sets as the image itself and the metadata are stored in two different places...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 383133,
          "author_name": "bwhitcher",
          "author_url": "",
          "post_date": "09/07/2018 18:46:22",
          "content": "<p>David, would it be possible to expose in this thread the DICOM header fields from the test2$hdr data.frame?  There is a lot of information in the header, something there may help diagnose the issue.  The <code>pixelData = FALSE</code> setting will ignore the image.  </p>\n\n<p>Brandon</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385183,
          "author_name": "danielac",
          "author_url": "",
          "post_date": "09/10/2018 13:59:52",
          "content": "<p>I'm converting the files now with irfanview and a plugin. For now every picture has the size of about 1MB...and I cannot upload more than 50 files at one time to add as external data. So I think I will run some analysis on my computer...unless someone finds a way to view the images in R</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385842,
          "author_name": "gmobaz",
          "author_url": "",
          "post_date": "09/11/2018 17:22:08",
          "content": "<p>Hi Brandon! I posted some ideas you gave me the other day below. I didn´t know you were in Kaggle. Thanks again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 385869,
          "author_name": "bwhitcher",
          "author_url": "",
          "post_date": "09/11/2018 18:05:09",
          "content": "<p>Thanks, Guillermo.  I haven't visited the Kaggle site in years but your DICOM question piqued my interest :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 385835,
      "author_name": "gmobaz",
      "author_url": "",
      "post_date": "09/11/2018 17:06:11",
      "content": "<p>Hi DanielAC:\n you can easily read files with pydicom. I talked to Brandon Whitcher, and told me he is getting back into the medical image analysis field again,  but it will not happen for at least a few months. Hopefully he will upgrade <strong>oro.dicom</strong> which is a very good tool. He suggested using <strong>pydicom</strong> or <strong>DCMTK</strong>.  Using pydicom within R is very simple if you use library <strong>reticulate</strong>. You have to download pyton or anaconda and then you can install with pip in python or install in python within R. Here it goes an example; I downloaded anaconda:</p>\n\n<p>library(data.table)</p>\n\n<p>library(imager) #very good package</p>\n\n<p>library(oro.dicom)</p>\n\n<p>library(reticulate)</p>\n\n<h1>install py packages, including pydicom. Only neded one time to install in python</h1>\n\n<p>if(F){\n   conda_version()</p>\n\n<p>py_available(initialize = T)</p>\n\n<p>py_numpy_available(initialize = FALSE)</p>\n\n<p>py_install(\"pandas\")</p>\n\n<p>py_install(\"gdcm\")</p>\n\n<p>py_install(\"pydicom\")</p>\n\n<h1>some others you might need</h1>\n\n<p>py_install(\"keras\")</p>\n\n<p>py_install(\"tensorflow\")</p>\n\n<p>py_install(\"matplotlib\")\n} </p>\n\n<h3>#</h3>\n\n<p>conda_version()</p>\n\n<h1>import pydicom and use it</h1>\n\n<p>pydcm=import(\"pydicom\")</p>\n\n<p>trainfolder=paste0(getwd(),\"/input/stage_1_train_images/\")</p>\n\n<p>trainpatno=list.files(trainfolder)</p>\n\n<p>idx=1</p>\n\n<p>fname=paste0(trainfolder,trainpatno[idx])</p>\n\n<p>hdr=readDICOMFile(fname,pixelData=F)$hdr</p>\n\n<p>img=pydcm$read_file(fname)</p>\n\n<p>atribs=py_list_attributes(img)</p>\n\n<p>img=t(img$pixel_array)</p>\n\n<p>as.cimg(img)%&gt;%plot</p>\n\n<p>I hope it helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 385884,
          "author_name": "bwhitcher",
          "author_url": "",
          "post_date": "09/11/2018 18:19:46",
          "content": "<p>I have also played around a little bit with the data and here is an example of using DCMTK to uncompress all of the DICOM files provided in this challenge.  Then they can be read into R via <strong>oro.dicom</strong> no problem.  My suggestion would be to stack the images into a 3D array and convert the array into a NIfTI object using <strong>oro.nifti</strong>.  This would allow for easier manipulation both in R and via the filesystem.  Some example code:</p>\n\n<pre><code>library(oro.dicom)\n\nrsna &lt;- \"~/Kaggle/rsna_pneumonia_detection_challenge\"\ntestDir &lt;- \"stage_1_test_images\"\ndir.create(paste(file.path(rsna, testDir), \"uncompressed\", sep = \"_\"))\ntrainDir &lt;- sub(\"test\", \"train\", testDir)\ndir.create(paste(file.path(rsna, trainDir), \"uncompressed\", sep = \"_\"))\ndcmFile &lt;- file.path(rsna, trainDir, \"00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\")\ndcm &lt;- readDICOMFile(dcmFile, debug = TRUE) # fails\n# system(\"brew install dcmtk\")\ncmd &lt;- paste(\"dcmdjpeg\",\n             dcmFile,\n             sub(\"images\", \"images_uncompressed\", dcmFile))\nsystem(cmd)\ndcmFile &lt;- file.path(sub(\"images\", \"images_uncompressed\", dcmFile))\ndcm &lt;- readDICOMFile(dcmFile) # succeeds\nimage(t(dcm$img), col = grey(0:64/64), axes = FALSE)\n# now do this for all files in the test and training directories\ndcmList &lt;- c(list.files(file.path(rsna, testDir), full.names = TRUE),\n             list.files(file.path(rsna, trainDir), full.names = TRUE))\nfor (f in dcmList) {\n  cmd &lt;- paste(\"dcmdjpeg\", f, sub(\"images\", \"images_uncompressed\", f))\n  system(cmd)\n}\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "376996": "Hi,\n\nI tried to read the images in R using\nreadDICOMFile(\"../input/stage_1_train_images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\")\nbut all I get is an error:\nError in parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, : Number of bytes in PixelData not specified; guess = 1\n\nHow do I read the data correctly, i.e. how do I get the images and all the metadata included in that file?\n\nRegards\nDaniel",
    "378491": "push...",
    "380337": "Why there is not much people to use R for Kaggle competitions? I am a R user, looking for more R discussion",
    "380340": "Guys, it's still a good time to learn python ;)",
    "380477": "&gt; **Godfrey Cheung wrote**\n&gt; \n&gt; &gt; Why there is not much people to use R for Kaggle competitions? I am a R user, looking for more R discussion \n\nYou're not likely to find many R kernels and discussions on deep learning and particularly computer vision competitions. \n\nAs said by Henrique,  it's not too late to learn python ;)",
    "380812": ";)",
    "381077": "Hi Daniel,\nI'm in the same boat as you (trying to complete this very interesting exercise with R). The first 128 bits of each DCM file are zero's. This can be determined by using the debug=TRUE argument. \n\ntest1 &lt;- readDICOMFile(\"/Users/davem/OneDrive/Desktop/R Prog Files/MSDS692/train images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", debug=TRUE)\n\n*# First 128 bytes of DICOM header =\n[1] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[42] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[83] 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n[124] 00 00 00 00 00\n# DICM = TRUE*\n\nYou can use the 'boffset' argument to skip the first 128 bits but this only returns the header information.\n\ntest2 &lt;- readDICOMFile(\"/Users/davem/OneDrive/Desktop/R Prog Files/MSDS692/train images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", boffset = 128)\nnames(test2)\n\n[1] \"hdr\" \"img\"\n\nUnfortunately, the image data is NULL. I'm working through that. If you have any suggestions, I would appreciate them.\n\nDave",
    "381872": "Hi Dave,\n\ninteresting approach. Unfortunately I still get the same error. First I get those 128 \"00\"s and then some content or information, but in the end I get:\n\nError in parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, : Number of bytes in PixelData not specified; guess = 1\nTraceback:\n\n1. readDICOMFile(\"../input/stage_1_train_images/00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\", \n .     debug = TRUE)\n2. parsePixelData(fraw[(bstart + dcm$data.seek):fsize], hdr, endian, \n .     flipud)\n3. stop(paste(\"Number of bytes in PixelData not specified; guess =\", \n .     guess))\n\nI think, this task is not meant to be solved in R....but I have to admit that I find Python somehow confusing... ;-)\n\nDaniel",
    "383051": "I think I will try to convert the images to JPG or PNG using Python and extract the metadata with R. This is not very nice but it should work.\nThat of course leads to a new structure of the sets as the image itself and the metadata are stored in two different places...",
    "383133": "David, would it be possible to expose in this thread the DICOM header fields from the test2$hdr data.frame?  There is a lot of information in the header, something there may help diagnose the issue.  The `pixelData = FALSE` setting will ignore the image.  \n\nBrandon",
    "385183": "I'm converting the files now with irfanview and a plugin. For now every picture has the size of about 1MB...and I cannot upload more than 50 files at one time to add as external data. So I think I will run some analysis on my computer...unless someone finds a way to view the images in R",
    "385835": "Hi DanielAC:\n you can easily read files with pydicom. I talked to Brandon Whitcher, and told me he is getting back into the medical image analysis field again,  but it will not happen for at least a few months. Hopefully he will upgrade **oro.dicom** which is a very good tool. He suggested using **pydicom** or **DCMTK**.  Using pydicom within R is very simple if you use library **reticulate**. You have to download pyton or anaconda and then you can install with pip in python or install in python within R. Here it goes an example; I downloaded anaconda:\n\nlibrary(data.table)\n\nlibrary(imager) #very good package\n\nlibrary(oro.dicom)\n\nlibrary(reticulate)\n\n#install py packages, including pydicom. Only neded one time to install in python\n\nif(F){\n   conda_version()\n\n   py_available(initialize = T)\n\n   py_numpy_available(initialize = FALSE)\n\n   py_install(\"pandas\")\n\n   py_install(\"gdcm\")\n\n   py_install(\"pydicom\")\n\n#some others you might need\n\n   py_install(\"keras\")\n\n   py_install(\"tensorflow\")\n\n   py_install(\"matplotlib\")\n} \n####\nconda_version()\n#import pydicom and use it\npydcm=import(\"pydicom\")\n\ntrainfolder=paste0(getwd(),\"/input/stage_1_train_images/\")\n\ntrainpatno=list.files(trainfolder)\n\nidx=1\n\nfname=paste0(trainfolder,trainpatno[idx])\n\nhdr=readDICOMFile(fname,pixelData=F)$hdr\n\nimg=pydcm$read_file(fname)\n\natribs=py_list_attributes(img)\n\n\nimg=t(img$pixel_array)\n\nas.cimg(img)%&gt;%plot\n\n\n\nI hope it helps.",
    "385842": "Hi Brandon! I posted some ideas you gave me the other day below. I didn´t know you were in Kaggle. Thanks again.",
    "385869": "Thanks, Guillermo.  I haven't visited the Kaggle site in years but your DICOM question piqued my interest :)",
    "385884": "I have also played around a little bit with the data and here is an example of using DCMTK to uncompress all of the DICOM files provided in this challenge.  Then they can be read into R via **oro.dicom** no problem.  My suggestion would be to stack the images into a 3D array and convert the array into a NIfTI object using **oro.nifti**.  This would allow for easier manipulation both in R and via the filesystem.  Some example code:\n\n    library(oro.dicom)\n    \n    rsna &lt;- \"~/Kaggle/rsna_pneumonia_detection_challenge\"\n    testDir &lt;- \"stage_1_test_images\"\n    dir.create(paste(file.path(rsna, testDir), \"uncompressed\", sep = \"_\"))\n    trainDir &lt;- sub(\"test\", \"train\", testDir)\n    dir.create(paste(file.path(rsna, trainDir), \"uncompressed\", sep = \"_\"))\n    dcmFile &lt;- file.path(rsna, trainDir, \"00a85be6-6eb0-421d-8acf-ff2dc0007e8a.dcm\")\n    dcm &lt;- readDICOMFile(dcmFile, debug = TRUE) # fails\n    # system(\"brew install dcmtk\")\n    cmd &lt;- paste(\"dcmdjpeg\",\n                 dcmFile,\n                 sub(\"images\", \"images_uncompressed\", dcmFile))\n    system(cmd)\n    dcmFile &lt;- file.path(sub(\"images\", \"images_uncompressed\", dcmFile))\n    dcm &lt;- readDICOMFile(dcmFile) # succeeds\n    image(t(dcm$img), col = grey(0:64/64), axes = FALSE)\n    # now do this for all files in the test and training directories\n    dcmList &lt;- c(list.files(file.path(rsna, testDir), full.names = TRUE),\n                 list.files(file.path(rsna, trainDir), full.names = TRUE))\n    for (f in dcmList) {\n      cmd &lt;- paste(\"dcmdjpeg\", f, sub(\"images\", \"images_uncompressed\", f))\n      system(cmd)\n    }"
  },
  "source": "meta"
}