{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Getting some basic idea of the CT-scans\n\nThe data contains CT-scans, which appear to be split into multiple slices. There must be some nice 3D software to combine them for the doctors. I am looking into the data in this kernel to get some idea what it is all about.\n\n# A look into the lungs\n\nThe CT-scan image data is given as 2D layers of a 3D image. They are generally of the chest area, so showing the internals of the lungs (as far as I can tell..). \n\nIn this kernel I just take a brief look at some of those images and their metadata. In my other [kernel](https://www.kaggle.com/donkeys/preprocessing-images-to-normalize-colors-and-sizes) I show how I preprocessed these. Here are two of those layered 3D images as animated gifs:\n\n![image1](https://i.imgur.com/KbL8C92.gif)\n\n![image2](https://i.imgur.com/TvPXWWg.gif)\n\nOne loops faster than the other since the scans have different numbers of frames.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"So, on with the show..","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pydicom\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nfrom PIL import Image\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport pydicom\nimport matplotlib.pyplot as plt\nfrom tqdm.auto import tqdm\nfrom PIL import Image\n\ntqdm.pandas()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"!ls","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls /kaggle/input/osic-pulmonary-fibrosis-progression","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls /kaggle/input/osic-pulmonary-fibrosispreprocessed/dataset","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train = pd.read_csv(\"/kaggle/input/osic-pulmonary-fibrosis-progression/train.csv\")\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_test = pd.read_csv(\"/kaggle/input/osic-pulmonary-fibrosis-progression/test.csv\")\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_meta = pd.read_csv(\"/kaggle/input/osic-pulmonary-fibrosispreprocessed/dataset/meta_files_train.csv\")\ndf_train_meta.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_test_meta = pd.read_csv(\"/kaggle/input/osic-pulmonary-fibrosispreprocessed/dataset/meta_files_test.csv\")\ndf_test_meta.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Patient Smoking Statuses","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train[\"SmokingStatus\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Patient Genders","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train[\"Sex\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Patient Ages","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train[\"Age\"].value_counts().sort_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Measurement counts by Patient","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train[\"Patient\"].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train[df_train[\"Patient\"] == \"ID00229637202260254240583\"]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls /kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430 | wc -l","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The above is interesting in showing that each patient has 6-10 measurement times in this dataset.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## DCM Files","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Data for a single patient appears to have multiple of the DCM files. It seems they are slices of the single CT-scan.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"    !ls /kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00229637202260254240583/","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Listing and Ordering the DCM files**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import glob\n\nfiles = glob.glob(\"/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00229637202260254240583/*.dcm\")\nfiles","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import re\n\nfiles.sort(key=lambda f: int(re.sub('\\D', '', f)))\nfiles","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Image Visuals","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"First a single CT layer:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_img(img_path):\n    ds = pydicom.dcmread(img_path)\n    im = Image.fromarray(ds.pixel_array)\n    #im = im.resize((DESIRED_SIZE,DESIRED_SIZE)) \n    #im.show()\n    plt.imshow(im, cmap=plt.cm.bone)\n    plt.show()\n    \n\nshow_img(files[0])\n#for file in files:\n#    show_img(file)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## DCM Metadata\n\nEach DCM file contains multiple metadata attributes:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import types\n\n#ds = pydicom.dcmread('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00229637202260254240583/1.dcm')\nds = pydicom.dcmread('/kaggle/input/osic-pulmonary-fibrosis-progression/train/ID00421637202311550012437/1.dcm')\nattrs = dir(ds)\nfor attr in attrs:\n    if attr.startswith(\"_\"):\n        continue\n    if attr == \"PixelData\" or attr == \"pixel_array\": \n        continue\n#            print(f\"{attr}\")\n    var_type = type(getattr(ds,attr))\n    if var_type == types.MethodType: \n        continue\n        #    print(f\"{attr}: {type(attr)}\")\n    #print(f\"{attr}: {var_type}\")\n    print(f\"{attr}: {getattr(ds,attr)}\")\n    ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Plotting CT-scan layers\n\nIf we plot all scan slices for a patient, in numerical order, and with associated coordinates, it shows the scan is sliced at regular intervals of 20 (whatever units it is..):","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import math\n\ndef plot_images(images, cols=3):\n#    plt.clf()\n#    plt.figure(figsize(14,8))\n    rows = math.ceil(len(images)/cols)\n    fig, ax = plt.subplots(nrows=rows, ncols=cols, figsize=(14,5*rows))\n\n    idx = 0\n    for row in ax:\n        for col in row:\n            if idx > len(images)-1:\n                break\n            img_path= images[idx]\n            ds = pydicom.dcmread(img_path)\n            im = Image.fromarray(ds.pixel_array)\n            col.imshow(im)\n            col.title.set_text(ds.ImagePositionPatient)\n            idx += 1\n    plt.show()\n\nplot_images(files)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the above plots, above each sub-plot are the \"coordinates\" of the slice inside the patient. The third number on each one is changing by 20 per layer. Is it the slice y-coordinate?","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Diffing Layers\n\nInitially I thought the DCM file might be showing a timeframe, but it is actually just slices of the same as I show above. So, I tried to diff the layers in any case. The most interesting part is maybe that they appear to be circular slices, with some static noise surrounding the actual \"beef\" in each layer:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import math\n\ndef plot_image_diffs(images, cols=3):\n#    plt.clf()\n#    plt.figure(figsize(14,8))\n    rows = math.ceil(len(images)/cols)\n    fig, ax = plt.subplots(nrows=rows, ncols=cols, figsize=(14,5*rows))\n\n    prev_img = None\n    idx = 0\n    for row in ax:\n        for col in row:\n            if idx > len(images)-1:\n                break\n            img_path= images[idx]\n            ds = pydicom.dcmread(img_path)\n            if prev_img is not None:\n                print(f\"no diff {idx}\")\n                diff = prev_img - ds.pixel_array\n                #diff = ds.pixel_array - prev_img\n                im = Image.fromarray(diff)\n            else:\n                print(f\"diff {idx}\")\n                im = Image.fromarray(ds.pixel_array)\n            col.imshow(im)\n            col.title.set_text(ds.ImagePositionPatient)\n            prev_img = ds.pixel_array\n            idx += 1\n    plt.show()\n\nplot_image_diffs(files)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Preprocessed files\n\nI have also uploaded a pre-processed dataset that reads all the .dcm files, takes out the medatadata for each, preprocessed each image to a matching color scaling and size, and writes them all out as .png files. Lets see here how that looks compared to the above. I will update this more later..","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import glob\n\nfiles = glob.glob(\"/kaggle/input/osic-pulmonary-fibrosispreprocessed/dataset/png/train/ID00229637202260254240583/*.png\")\nfiles","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import re\n\nfiles.sort(key=lambda f: int(re.sub('\\D', '', f)))\nfiles","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_png(img_path):\n    im = Image.open(img_path)\n    plt.imshow(im, cmap=plt.cm.bone)\n    plt.show()    \n\nshow_png(files[0])\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import math\n\ndef plot_images(images, cols=3):\n    rows = math.ceil(len(images)/cols)\n    fig, ax = plt.subplots(nrows=rows, ncols=cols, figsize=(14,5*rows))\n\n    idx = 0\n    for row in ax:\n        for col in row:\n            if idx > len(images)-1:\n                break\n            img_path= images[idx]\n            im = Image.open(img_path)\n            col.imshow(im, cmap=plt.cm.bone)\n            #col.title.set_text()\n            idx += 1\n    plt.show()\n\nplot_images(files)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import math\n\ndef plot_image_diffs(images, cols=3):\n    rows = math.ceil(len(images)/cols)\n    fig, ax = plt.subplots(nrows=rows, ncols=cols, figsize=(14,5*rows))\n\n    prev_img = None\n    idx = 0\n    for row in ax:\n        for col in row:\n            if idx > len(images)-1:\n                break\n            img_path= images[idx]\n            img = Image.open(img_path)\n            if prev_img is not None:\n                print(f\"no diff {idx}\")\n                img_data = np.asarray(img, dtype=np.uint8)\n                diff = prev_img - img_data\n                im = Image.fromarray(diff)\n            else:\n                print(f\"diff {idx}\")\n                im = img\n            col.imshow(im, cmap=plt.cm.bone)\n            #col.title.set_text(ds.ImagePositionPatient)\n            prev_img = np.asarray(img, dtype=np.uint8)\n            idx += 1\n    plt.show()\n\nplot_image_diffs(files)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_meta.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"# Selected Pictures\n\nAs I looked through the images when implementing the pre-processing, I collected some cases that seemed to represent potentially useful augmentation functions for training and evaluation. \n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_train_png(img_id, idx):\n    plt.figure(figsize=(12,8))\n    im = Image.open(f\"/kaggle/input/osic-pulmonary-fibrosispreprocessed/dataset/png/train/{img_id}/{idx}.png\")\n    plt.imshow(im, cmap=plt.cm.bone)\n    plt.show()    \n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Flat Shape 1\n\nMany images are quite flat, with seemingly the head pointing upwards. The teeth are often showing as a white arch on the top.\n\nHere is an example of teeth showing a bit:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00012637202177665765362\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Further along the layers, the same patient is a bit more rounded, I guess more in the middle of the chest:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00012637202177665765362\", 7)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Flat Shape 2\n\nThis one has the bones colored brighter white. Notice the upside down Y shape in the middle here. In the above figure the same structure seems to be upside down compared to this, which makes me think perhaps **vertical flip** could be a useful augmentation to experiment with?","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00019637202178323708467\", 2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00019637202178323708467\", 7)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Round Shape\n\nAs opposed to the two flatter images above, some are more rounded:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00020637202178344345685\", 1)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Detached Parts\n\nSometimes the \"parts\" of the body show as quite detached in some parts of the layers. Guess it is another artifact of body shape/position, device properties etc. Several images show some form of this, but most none.\n\nThis one has both sides detached:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00047637202184938901501\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This one shows one part:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00122637202216437668965\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Ray Storm\n\nThis one looks a bit like there was a burst of energy going from right to left. Not sure how to augment for such effects.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00076637202199015035026\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Shiny Objects\n\nBesides the above \"ray storm\", there were also multiple images with a shiny part emitting rays across. Here is one with such effect on top right corner.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00126637202218610655908\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Glowing Eyes\n\nNot sure what this person had in their eyes.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00133637202223847701934\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Static\n\nSome images had more static noise than others. So perhaps adding **noise** is a useful augmentation to try?","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00123637202217151272140\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Teeth\n\nI mentioned before how the teeth are sometimes very visible. This one was especially so:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00134637202223873059688\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Blurry\n\nSome patients are blurred more than others. Again, perhaps adding **blur** could be another useful augmentation? For example, this one is a bit fuzzy:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00139637202231703564336\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And this one is much more so:","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00264637202270643353440\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Tilt\n\nThe patients and devices are sometimes tilted and positioned in different ways. Some of this data is available in the image metadata, but rather than try to play with that, perhaps some level of **rotation** is another useful augmentation to try?","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"show_train_png(\"ID00196637202246668775836\", 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Overall, based on the above, I will play a bit with at least the following augmentations:\n\n- brigthness levels: even looking at the above images after I tried to \"normalize\" their color levels, some are still brighter than others. \n- Horizontal and vertical flip. Some images seem to show patients in different positions / angles both horizontally and vertically.\n- Adding noise. Some images had more static noise look than others. Not many very heavily, but some variation could be interesting to try.\n- Rotation levels. The angle of body changes in some images.\n- Dropping some parts of image. There is some variation in different parts being visible at different times.\n\nHow to best apply these to a 3D set of 2D images is another interesting question.. Or to model the 3D/2D aspect in general.\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}