{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"This notebook takes a look at the data that will be used to train a model to predict Melanoma Skin Lessions. It focuses on work to understand/learn the DICOM data format as it it is used in the Kaggle ISIC 2020 competition.  It also shows work to visualize the embedded images in a more *natural* (RGB) format vs. the embedded `YBR_FULL_422` format that if displayed as is produces images that in the author's opinion do not look natural.\n\nThe Kaggle contest (https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion?sortBy=relevance&group=all&search=color&page=1&pageSize=20&category=all)has a very large image dataset. This notebook uses a small subset of the original dataset (a few images from the `test` directory)\n\nThis notebook does not include any modeling to predict melanoma. It only covers the work to open DICOM files, looking at the metadata, manipulating it if needed and displaying the data for visual analysis.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# References","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* `pydicom` https://pydicom.github.io/pydicom/stable/old/working_with_pixel_data.html#dataset-pixel-array\n* fast.ai V2. https://dev.fast.ai/\n* Fastai2 DICOM Starer. https://www.kaggle.com/avirdee/fastai2-dicom-starter\n* Choosing Colormaps in Matplotlib (`plt`). https://matplotlib.org/3.1.0/tutorials/colors/colormaps.html\n* YUV to RGB conversion. http://www.fourcc.org/fccyvrgb.php\n* DICOM is easy!. https://dicomiseasy.blogspot.com/2012/08/chapter-12-pixel-data.html\n* Convert PIL image to OpenCV Image. https://stackoverflow.com/questions/14134892/convert-image-from-pil-to-opencv-format\n* Convert OpenCV Image to PIL. https://stackoverflow.com/questions/13576161/convert-opencv-image-into-pil-image-in-python-for-use-with-zbar-library","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Background","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"DICOM Images (Digital Imaging and Communication in medicines) are the standard that is used in medicine to allow images from X-Ray, MRI, Cat Scan and others (along with metadadat) to be exchanged among medical professionals.\n\nIn python, DICOM images can be viewed and analyzed using the  `pydicom` package. If you have installed `fastai V2` then this package will have already be installed.\n\nThere is also a need to use a `PyTorch` computer vision package called `Kornia` however if using `fastai v2` this pakcage is also part of the distribution.\n\nThere are other packages that will be needed. They are installed in the`conda`environment used used in this project. To see what packages are part of this environment please see the `requirements.yml` file in the root directory of this project.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Initialize","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Load packages we need","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install --upgrade fastai2 > /dev/null","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from fastai2.basics import *\nfrom fastai2.callback.all import *\nfrom fastai2.vision.all import *\nfrom fastai2.medical.imaging import *","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import pydicom","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import cv2","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Get the files","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"source = Path(\"../input/siim-isic-melanoma-classification\")\nfiles = os.listdir(source)\nfiles","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = source/'train'\ntrain_files = get_dicom_files(train)\ntrain_files","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Pick one of the files and see what information is in the file","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"image = train_files[42]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dimg = dcmread(image)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Info for the image","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dimg","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Important fields:\n\n* **Pixel Data** is where the actual image is stored. The order of pixels encoded for each image plane is left to right, top to bottom, i.e., the upper left pixel (labeled 1,1) is encoded first\n\n* **Photometric Interpretation** This is the color space. From what I see there, these images are not in the usual RGB format, but they are in `YBR_FULL_422`. As I understand things today if I use `fast.ai v2` medical images to display the images, the will not look like what you expect a picture of a skin lession to look like.  I don't believe it matters when the model is trained, but later in this notebook after I extract th image I use `Open CV` to convert the color space to `RGB`\n\n* **Samples per Pixel** One for monochrome and 3 for RGB (or in this case for `YBR_FULL_422`) images","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Pydicom reads pixel data as the raw bytes found in the file and most of the time the `Pyxel Data` is not useable as read since the data can be stored in one of several formats:\n\n* Unsigned integers, or floats\n* There may be multiple image frames\n* There may be multiple planes per frame (i.e. RGB) and the order of the pixels may be different.\n\nsee the `pydicom` reference page for other possible formats.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"This is a function (I found this in this notebook in Kaggle: https://www.kaggle.com/avirdee/fastai2-dicom-starter) that will display an image and show choosen tags within the head of the DICOM, in this case PatientName, PatientID, PatientSex, BodyPartExamined and we can use Tranform from fastai2 that conveniently allows us to resize the image. Note that this function (as I understand things now) does not handle the display of the `YBR_FULL_422` images and when they are displayed, they will not look like you would expect a skim image to look like.  I will later in the notebook create a new function that will show the images with the *usual* colors (`RGB` color space)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_one_patient(file):\n    \"\"\" function to view patient image and choosen tags within the head of the DICOM\"\"\"\n    pat = dcmread(file)\n    print(f'patient Name: {pat.PatientName}')\n    print(f'Patient ID: {pat.PatientID}')\n    print(f'Patient age: {pat.PatientAge}')\n    print(f'Patient Sex: {pat.PatientSex}')\n    print(f'Body part: {pat.BodyPartExamined}')\n    trans = Transform(Resize(256))\n    dicom_create = PILDicom.create(file)\n    dicom_transform = trans(dicom_create)\n    return show_image(dicom_transform)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let,s Use `fast.ai v2` function to open and read a `DICOM1 file","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient = dcmread(image)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Look at the color space","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f'Photometric Interpretation: {patient.PhotometricInterpretation}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Get the image from the DICOM file as a PIL file","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_create = PILDicom.create(image)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Display the image. Since PIL (as I understand) only handles RGB images (or maybe there is another option that needs to be used that I know know about), the image does not look as one would excpet a skin picture","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_create.show(figsize=(6,6), cmap=plt.cm.gist_ncar)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I will use `OpenCV` to convert the `YBR_FULL_422` image to RGB. To do this I will extract from the `DICOM` file the image, which as we saw earlier is stored in `pixel_array`. We take the pixel data array and convert it to a `PIL` image (there is probably a better way to do the following stes, but I used this approach to to experiment, so I kept these steps)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"pil_image = Image.fromarray(dcmread(image).pixel_array)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now convert the PIL image to an OpenCV image. Note that the OpenCV image will be in BGR.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"open_cv_image = np.array(pil_image) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"According to https://dicomiseasy.blogspot.com/2012/08/chapter-12-pixel-data.html, `YBR_FULL_422` means that the color channels are in the YCbCr color space that is used in JPEG. The next line uses the OpenCV `cv2.cvtColor` to transform the color space.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"open_cv_image = cv2.cvtColor(open_cv_image, cv2.COLOR_YCrCb2BGR)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now that we have an OpenCV image in the right space, we convert it back to PIL to Display (note that I could have used OpenCV to display it, but since `fast.ai` uses PIL, I wanted to stay using `PIL`)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"pil_image = Image.fromarray(open_cv_image)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now when we display the image, it will look like you would expect a skin picture to look like.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.imshow(pil_image)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now that I know how to display the image embedded in the DICOM file in a natural format, I can create a new function to use in other notebooks rived from the show_one_patient in https://www.kaggle.com/avirdee/fastai2-dicom-starter","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_one_patient_RGB(file):\n    \"\"\" function to view patient image and choosen tags within the head of the DICOM\"\"\"\n    pat = dcmread(file)\n    print(f'patient Name: {pat.PatientName}')\n    print(f'Patient ID: {pat.PatientID}')\n    print(f'Patient age: {pat.PatientAge}')\n    print(f'Patient Sex: {pat.PatientSex}')\n    print(f'Body part: {pat.BodyPartExamined}')\n    trans = Transform(Resize(256))\n    \n    pil_image = Image.fromarray(dcmread(image).pixel_array)\n    # Not sure yet about the use of the following line. For now uncommented.\n    # pil_image = trans(pil_image)\n    open_cv_image = np.array(pil_image) \n    open_cv_image = cv2.cvtColor(open_cv_image, cv2.COLOR_YCrCb2BGR)\n    pil_image = Image.fromarray(open_cv_image)\n    # Could add a parameter to specify figize. For now it is set to (6,6)\n    \n    return show_image(pil_image, figsize=(6,6))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"show_one_patient_RGB(image)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As we can see, I have now the skin image in more *natural* colors or as I would expect to seem them from a photograph from the doctor's office.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}