{"cells":[{"metadata":{"_kg_hide-output":true,"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"!pip install --upgrade pip","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"!pip install -q tensorflow-io","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import os\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport random\nimport re\n\n#color\nfrom colorama import Fore, Back, Style\n\nimport tensorflow as tf\nimport tensorflow_io as tfio\n\n#read dicom\nimport pydicom\n\n#visialisation\nfrom PIL import Image\nfrom IPython.display import Image as show_gif\nimport scipy.misc\nimport matplotlib\n\n#Pandas Profiling\nimport pandas_profiling as pdp\n\n#For segmentation and 3d plotting\nimport skimage\nimport skimage.measure as measure\nfrom skimage.morphology import ball, disk, dilation, binary_erosion, remove_small_objects, erosion, closing, reconstruction, binary_closing\nfrom skimage.measure import label,regionprops, perimeter\nfrom skimage.morphology import binary_dilation, binary_opening\nfrom skimage.filters import roberts, sobel\nfrom skimage import measure, feature\nfrom skimage.segmentation import clear_border\nfrom skimage import data\nfrom scipy import ndimage as ndi\nfrom mpl_toolkits.mplot3d.art3d import Poly3DCollection\nimport scipy.misc\n\n# Settings for pretty nice plots\nplt.style.use('seaborn-bright')\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 align=\"center\"> <span style='font-size:30px;'>&#9672;</span> OSIC Pulmonary Fibrosis Competition </h1>","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### What is Idiopathic Pulmonary Fibrosis? \n\n<iframe width=\"560\" height=\"315\" src=\"https://www.youtube.com/embed/zxI4Hv6lNpA\" frameborder=\"0\" allow=\"accelerometer; autoplay; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen></iframe>\n\n\n### Purpose of this project: \n- <span style='font-size:15px;'> Predict a patient’s severity of decline in lung function based on a CT scan of their lungs. Lung function is assessed based on output from a spirometer, which measures the forced vital capacity (FVC), i.e. the volume of air exhaled. Specifically: </span>\n    - FVC - final 3 values for each patient (only these will be used for the final score)\n    - Confidence Value - a thread about what it is here. It is measured in ml and it's \"how confident\" are you about the the estimated FVC. High value in confidence means that you are off a lot from the actual FVC, while a very low (or even 0) confidence means you're very sure on the FVC. (if somebody thinks this is not correct, please text me, I might be wrong about this)\n    \n### Additional info:\n- <span style='font-size:15px;'> **FVC** - the recorded lung capacity in ml </span>\n\n- <span style='font-size:15px;'> **Percent** - a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics </span>\n- <span style='font-size:15px;'> **Weeks** - the relative number of weeks pre/post the baseline CT (may be negative) </span>\n\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<h3 align=\"center\"> <span style='font-size:30px;'>&#9672;</span> <strong> EDA </strong> <span style='font-size:30px;'>&#9672;</span> </h3>","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dataDir = '../input/osic-pulmonary-fibrosis-progression'\nworkDir = '../input/output/'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data = pd.read_csv(os.path.join(dataDir, 'train.csv'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Function to plot characteristics of patients\ndef plot_columns_data(df, col, bins):\n    plt.figure(figsize = (8, 4))\n    plt.hist(df[col], bins = bins)\n    plt.xlabel('{}'.format(col))\n    plt.ylabel('Count')\n    try:\n        plt.text(50, 450, r'$\\mu={},\\ \\sigma={}$'.format(df[col].unique().mean().round(2), df[col].unique().std().round(2)), fontsize = 12)\n    except:\n        pass\n    plt.show() ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(Fore.YELLOW + 'People\\'s ages, sorted: ', Style.RESET_ALL, sorted(train_data.Age.unique()))\nprint(Fore.YELLOW + 'Youngest person\\'s age:', Style.RESET_ALL, train_data.Age.unique().min())\nprint(Fore.YELLOW + 'Oldest person\\'s age  :', Style.RESET_ALL, train_data.Age.unique().max())\nprint(Fore.YELLOW + 'Mean age of people in the dataset:', Style.RESET_ALL, train_data.Age.unique().mean().round(2))\nprint(Fore.YELLOW + 'Person\\'s sex  : ', Style.RESET_ALL, train_data.Sex.unique())\nprint(Fore.YELLOW + 'Smoking status: ', Style.RESET_ALL, train_data.SmokingStatus.unique())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_columns_data(train_data, 'Age', 7)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_columns_data(train_data, 'Sex', 2)\ntrain_data.groupby(['Sex']).count()['Patient'].to_frame()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_columns_data(train_data, 'SmokingStatus', 3)\ntrain_data.groupby(['SmokingStatus']).count()['Patient'].to_frame()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('There are', Fore.YELLOW + '{}'.format(len(train_data[\"Patient\"].unique())), Style.RESET_ALL, 'unique patients in Train Data.', \"\\n\")\n\n# Recordings per Patient\ndata = train_data.groupby(\"Patient\")[\"Weeks\"].count().reset_index(drop=False)\n# Sort by Weeks\ndata = data.sort_values(['Weeks']).reset_index(drop=True)\nprint('Minimum recordings per patient:', Fore.YELLOW + '{}'.format(data[\"Weeks\"].min()), Style.RESET_ALL, \"\\n\" +\n      'Maximum recordings per patient:', Fore.YELLOW + '{}'.format(data[\"Weeks\"].max()), Style.RESET_ALL)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Min FVC value:', Fore.YELLOW + '{}'.format(train_data.FVC.unique().min()), Style.RESET_ALL)\nprint('Max FVC value:', Style.RESET_ALL, Fore.YELLOW + '{}'.format(train_data.FVC.unique().max()), Style.RESET_ALL)\nprint('Min Percent value:', Fore.MAGENTA + '{} %'.format(train_data.Percent.unique().min().round(2)), Style.RESET_ALL)\nprint('Max Percent value:', Fore.MAGENTA + '{} %'.format(train_data.Percent.unique().max().round(2)), Style.RESET_ALL)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Figure\nfig, (fvc, pct) = plt.subplots(1, 2, figsize = (16, 6))\n\nfvc = sns.distplot(train_data[\"FVC\"], ax=fvc, color = 'orange')\npct = sns.distplot(train_data[\"Percent\"], ax=pct, color = 'blue')\n\nfvc.set_title(\"FVC Distribution\", fontsize=16)\npct.set_title(\"Percent Distribution\", fontsize=16);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Min no. weeks pre(negative number)/post base CT scan:', Fore.YELLOW + '{}'.format(train_data['Weeks'].min()), Style.RESET_ALL)\nprint('Max no. weeks pre(negative number)/post base CT scan:', Fore.YELLOW + '{}'.format(train_data['Weeks'].max()), Style.RESET_ALL)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize = (16, 6))\n\nw = sns.distplot(train_data['Weeks'])\nplt.title(\"Number of weeks before/after the CT scan\", fontsize = 16)\nplt.xlabel(\"Weeks\", fontsize=14);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Let's check the correlations\nplt.figure(figsize=(16,6))\nsns.heatmap(train_data.corr(), \n            annot=True, \n            linewidths = .5,\n            square=True,\n            annot_kws={'size':12, 'weight': 'bold'},\n            center = 0,\n            cmap=sns.color_palette('bright'))\nplt.yticks(rotation = 0)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Since the Percent field approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics it is normal that we see higher correlation between these two features. As a whole the correlation between the features is close to 0.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_data = pd.read_csv(os.path.join(dataDir, 'test.csv'))\n# test_data.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_sub = pd.read_csv(os.path.join(dataDir, 'sample_submission.csv'))\n#sample_sub.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Explore the DICOM Files","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"trainDicomPath = '../input/osic-pulmonary-fibrosis-progression/train'\ntestDicomPath = '../input/osic-pulmonary-fibrosis-progression/test'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def getListOfFiles(dirName):\n    ''' For the given path, get the List of \n        all files in the directory tree \n    '''\n    # create a list of file and sub directories \n    # names in the given directory \n    listOfFile = os.listdir(dirName)\n    allFiles = list()\n    # Iterate over all the entries\n    for entry in listOfFile:\n        # Create full path\n        fullPath = os.path.join(dirName, entry)\n        # If entry is a directory then get the list of files in this directory \n        if os.path.isdir(fullPath):\n            allFiles = allFiles + getListOfFiles(fullPath)\n        else:\n            allFiles.append(fullPath)\n                \n    return allFiles   ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Get the list of all train files in directory tree at given path\nlistOfTrainDcmFiles = getListOfFiles(trainDicomPath)\nlen(listOfTrainDcmFiles)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Get the list of all test files in directory tree at given path\nlistOfTestDcmFiles = getListOfFiles(testDicomPath)\nlen(listOfTestDcmFiles)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Show random scans\nrandom.seed(22)\nfor file in random.sample(listOfTrainDcmFiles, 3):\n    dataset = pydicom.dcmread(file)\n    \n    print(Fore.MAGENTA + \"Patient id.......:\", Style.RESET_ALL, dataset.PatientID, \"\\n\" +\n          Fore.MAGENTA + \"Modality.........:\", Style.RESET_ALL, dataset.Modality, \"\\n\" +\n          Fore.MAGENTA + \"Rows.............:\", Style.RESET_ALL, dataset.Rows, \"\\n\" +\n          Fore.MAGENTA + \"Columns..........:\", Style.RESET_ALL, dataset.Columns)\n    \n    plt.imshow(dataset.pixel_array, cmap=plt.cm.bone)\n    plt.show()\n    ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### #A breath-in + hold for a Patient. \n#### *Credit: https://www.kaggle.com/andradaolteanu/pulmonary-fibrosis-competition-eda-dicom-prep*","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"patient_dir = \"../input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430\"\ndatasets = []\n\n# First Order the files in the dataset\nfiles = []\nfor dcm in list(os.listdir(patient_dir)):\n    files.append(dcm) \nfiles.sort(key=lambda f: int(re.sub('\\D', '', f)))\n\n# Read in the Dataset\nfor dcm in files:\n    path = patient_dir + \"/\" + dcm\n    datasets.append(pydicom.dcmread(path))\n\n# Plot the images\nfig=plt.figure(figsize=(16, 6))\ncolumns = 10\nrows = 3\n\nfor i in range(1, columns*rows +1):\n    img = datasets[i-1].pixel_array\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(img, cmap=plt.cm.bone)\n    plt.title(i, fontsize = 9)\n    plt.axis('off');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Create base director for Train .dcm files\ndirector = \"../input/osic-pulmonary-fibrosis-progression/train\"\n\n# Create path column with the path to each patient's CT\ntrain_data[\"Path\"] = director + \"/\" + train_data[\"Patient\"]\n\n# Create variable that shows how many CT scans each patient has\ntrain_data[\"CT_number\"] = 0\n\nfor k, path in enumerate(train_data[\"Path\"]):\n    train_data[\"CT_number\"][k] = len(os.listdir(path))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def create_gif(number_of_CT = 87):\n    \"\"\"Picks a patient at random and creates a GIF with their CT scans.\"\"\"\n    \n    # Select one of the patients\n    # patient = \"ID00007637202177411956430\"\n    patient = train_data[train_data[\"CT_number\"] == number_of_CT].sample(random_state=1)[\"Patient\"].values[0]\n    \n    # === READ IN .dcm FILES ===\n    patient_dir = \"../input/osic-pulmonary-fibrosis-progression/train/\" + patient\n    datasets = []\n\n    # First Order the files in the dataset\n    files = []\n    for dcm in list(os.listdir(patient_dir)):\n        files.append(dcm) \n    files.sort(key=lambda f: int(re.sub('\\D', '', f)))\n\n    # Read in the Dataset from the Patient path\n    for dcm in files:\n        path = patient_dir + \"/\" + dcm\n        datasets.append(pydicom.dcmread(path))\n        \n        \n    # === SAVE AS .png ===\n    # Create directory to save the png files\n    if os.path.isdir(f\"png_{patient}\") == False:\n        os.mkdir(f\"png_{patient}\")\n\n    # Save images to PNG\n    for i in range(len(datasets)):\n        img = datasets[i].pixel_array\n        matplotlib.image.imsave(f'png_{patient}/img_{i}.png', img)\n        \n        \n    # === CREATE GIF ===\n    # First Order the files in the dataset (again)\n    files = []\n    for png in list(os.listdir(f\"../working/png_{patient}\")):\n        files.append(png) \n    files.sort(key=lambda f: int(re.sub('\\D', '', f)))\n\n    # Create the frames\n    frames = []\n\n    # Create frames\n    for file in files:\n    #     print(\"../working/png_images/\" + name)\n        new_frame = Image.open(f\"../working/png_{patient}/\" + file)\n        frames.append(new_frame)\n\n    # Save into a GIF file that loops forever\n    frames[0].save(f'gif_{patient}.gif', format='GIF',\n                   append_images=frames[1:],\n                   save_all=True,\n                   duration=200, loop=0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"create_gif(number_of_CT=12)\ncreate_gif(number_of_CT=30)\ncreate_gif(number_of_CT=87)\n\nprint(\"First file len:\", len(os.listdir(\"../working/png_ID00165637202237320314458\")), \"\\n\" +\n      \"Second file len:\", len(os.listdir(\"../working/png_ID00199637202248141386743\")), \"\\n\" +\n      \"Third file len:\", len(os.listdir(\"../working/png_ID00340637202287399835821\")))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"show_gif(filename=\"./gif_ID00165637202237320314458.gif\", format='png', width=400, height=400)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"show_gif(filename=\"./gif_ID00340637202287399835821.gif\", format='png', width=400, height=400)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"show_gif(filename=\"./gif_ID00199637202248141386743.gif\", format='png', width=400, height=400)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Pandas Profiling","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"profile_train_df = pdp.ProfileReport(train_data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"profile_train_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Lung segmentation and 3d plot\n#### *Reference for this section: https://www.kaggle.com/arturscussel/lung-segmentation-and-candidate-points-generation*","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def read_ct_scan(folder_name):\n        # Read the slices from the dicom file\n        slices = [pydicom.dcmread(folder_name + filename) for filename in os.listdir(folder_name)]\n        \n        # Sort the dicom slices in their respective order\n        slices.sort(key=lambda x: int(x.InstanceNumber))\n        \n        # Get the pixel values for all the slices\n        slices = np.stack([s.pixel_array for s in slices])\n        slices[slices == -2000] = 0\n        return slices","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ct_scan_1 = read_ct_scan('../input/osic-pulmonary-fibrosis-progression/train/ID00007637202177411956430/') \nct_scan_2 = read_ct_scan('../input/osic-pulmonary-fibrosis-progression/train/ID00047637202184938901501/')\nct_scan_3 = read_ct_scan('../input/osic-pulmonary-fibrosis-progression/train/ID00012637202177665765362/')\nct_scan_4 = read_ct_scan('../input/osic-pulmonary-fibrosis-progression/train/ID00062637202188654068490/')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# ct_scan_2.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_ct_scan(scan):\n    f, plots = plt.subplots(int(scan.shape[0] / 20) + 1, 4, figsize=(10, 10))\n    for i in range(0, scan.shape[0], 5):\n        plots[int(i / 20), int((i % 20) / 5)].axis('off')\n        plots[int(i / 20), int((i % 20) / 5)].imshow(scan[i], cmap=plt.cm.bone) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_ct_scan(ct_scan_2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_segmented_lungs(im, plot=False):\n    \n    '''\n    This funtion segments the lungs from the given 2D slice.\n    '''\n    if plot == True:\n        f, plots = plt.subplots(8, 1, figsize=(5, 40))\n    '''\n    Step 1: Convert into a binary image. \n    '''\n    binary = im < 604\n    if plot == True:\n        plots[0].axis('off')\n        plots[0].imshow(binary, cmap=plt.cm.bone) \n    '''\n    Step 2: Remove the blobs connected to the border of the image.\n    '''\n    cleared = clear_border(binary)\n    if plot == True:\n        plots[1].axis('off')\n        plots[1].imshow(cleared, cmap=plt.cm.bone) \n    '''\n    Step 3: Label the image.\n    '''\n    label_image = label(cleared)\n    if plot == True:\n        plots[2].axis('off')\n        plots[2].imshow(label_image, cmap=plt.cm.bone) \n    '''\n    Step 4: Keep the labels with 2 largest areas.\n    '''\n    areas = [r.area for r in regionprops(label_image)]\n    areas.sort()\n    if len(areas) > 2:\n        for region in regionprops(label_image):\n            if region.area < areas[-2]:\n                for coordinates in region.coords:                \n                       label_image[coordinates[0], coordinates[1]] = 0\n    binary = label_image > 0\n    if plot == True:\n        plots[3].axis('off')\n        plots[3].imshow(binary, cmap=plt.cm.bone) \n    '''\n    Step 5: Erosion operation with a disk of radius 2. This operation is \n    seperate the lung nodules attached to the blood vessels.\n    '''\n    selem = disk(2)\n    binary = binary_erosion(binary, selem)\n    if plot == True:\n        plots[4].axis('off')\n        plots[4].imshow(binary, cmap=plt.cm.bone) \n    '''\n    Step 6: Closure operation with a disk of radius 10. This operation is \n    to keep nodules attached to the lung wall.\n    '''\n    selem = disk(10)\n    binary = binary_closing(binary, selem)\n    if plot == True:\n        plots[5].axis('off')\n        plots[5].imshow(binary, cmap=plt.cm.bone) \n    '''\n    Step 7: Fill in the small holes inside the binary mask of lungs.\n    '''\n    edges = roberts(binary)\n    binary = ndi.binary_fill_holes(edges)\n    if plot == True:\n        plots[6].axis('off')\n        plots[6].imshow(binary, cmap=plt.cm.bone) \n    '''\n    Step 8: Superimpose the binary mask on the input image.\n    '''\n    get_high_vals = binary == 0\n    im[get_high_vals] = 0\n    if plot == True:\n        plots[7].axis('off')\n        plots[7].imshow(im, cmap=plt.cm.gray) \n        \n    return im\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def segment_lung_from_ct_scan(ct_scan):\n    return np.asarray([get_segmented_lungs(slice) for slice in ct_scan])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def further_remove_noise(segmented_ct_scan):\n    selem = ball(2)\n    binary = binary_closing(segmented_ct_scan, selem)\n\n    label_scan = label(binary)\n\n    areas = [r.area for r in regionprops(label_scan)]\n    areas.sort()\n\n    for r in regionprops(label_scan):\n        max_x, max_y, max_z = 0, 0, 0\n        min_x, min_y, min_z = 1000, 1000, 1000\n    \n        for c in r.coords:\n            max_z = max(c[0], max_z)\n            max_y = max(c[1], max_y)\n            max_x = max(c[2], max_x)\n        \n            min_z = min(c[0], min_z)\n            min_y = min(c[1], min_y)\n            min_x = min(c[2], min_x)\n        if (min_z == max_z or min_y == max_y or min_x == max_x or r.area > areas[-3]):\n            for c in r.coords:\n                segmented_ct_scan[c[0], c[1], c[2]] = 0\n        else:\n            index = (max((max_x - min_x), (max_y - min_y), (max_z - min_z))) / (min((max_x - min_x), (max_y - min_y) , (max_z - min_z)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"segmented_ct_scan_1 = segment_lung_from_ct_scan(ct_scan_1)\nsegmented_ct_scan_1[segmented_ct_scan_1 < 604] = 0\nfurther_remove_noise(segmented_ct_scan_1)\n# plot_ct_scan(segmented_ct_scan_1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"segmented_ct_scan_2 = segment_lung_from_ct_scan(ct_scan_2)\nsegmented_ct_scan_2[segmented_ct_scan_2 < 604] = 0\nfurther_remove_noise(segmented_ct_scan_2)\n# plot_ct_scan(segmented_ct_scan_2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"segmented_ct_scan_3 = segment_lung_from_ct_scan(ct_scan_3)\nsegmented_ct_scan_3[segmented_ct_scan_3 < 604] = 0\nfurther_remove_noise(segmented_ct_scan_3)\n#plot_ct_scan(segmented_ct_scan_3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"segmented_ct_scan_4 = segment_lung_from_ct_scan(ct_scan_4)\nsegmented_ct_scan_4[segmented_ct_scan_4 < 604] = 0\nfurther_remove_noise(segmented_ct_scan_4)\n#plot_ct_scan(segmented_ct_scan_4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_3d(image):\n    \n    # Position the scan upright, \n    # so the head of the patient would be at the top facing the camera\n    p = image.transpose(2,1,0)\n    p = p[:,:,::-1]\n    \n    verts, faces, normals, values = measure.marching_cubes_lewiner(p)\n    \n    fig = plt.figure(figsize=(10, 10))\n    ax = fig.add_subplot(111, projection='3d')\n\n    #     ax.set_xlim(0, p.shape[0])\n    #     ax.set_ylim(0, p.shape[1])\n    #     ax.set_zlim(0, p.shape[2])\n    \n    ax.set_xlim(np.min(verts[:,0]), np.max(verts[:,0]))\n    ax.set_ylim(np.min(verts[:,1]), np.max(verts[:,1])) \n    ax.set_zlim(np.min(verts[:,2]), np.max(verts[:,2]))\n    \n    # Fancy indexing: `verts[faces]` to generate a collection of triangles\n    mesh = Poly3DCollection(verts[faces], alpha=0.1)\n    \n    face_color = [0.45, 0.45, 0.75]    \n    mesh.set_facecolor(face_color)\n    \n    mesh.set_edgecolor('k')\n    \n    ax.add_collection3d(mesh)\n    \n    plt.tight_layout()\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### **The normal FVC range for an adult is between 3.0 and 5.0 L.** (https://www.verywellhealth.com/forced-expiratory-capacity-measurement-914900)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_3d(segmented_ct_scan_1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data.loc[train_data.Patient == 'ID00007637202177411956430'].mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_3d(segmented_ct_scan_2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data.loc[train_data.Patient == 'ID00047637202184938901501'].mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_3d(segmented_ct_scan_3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data.loc[train_data.Patient == 'ID00012637202177665765362'].mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_3d(segmented_ct_scan_4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_data.loc[train_data.Patient == 'ID00062637202188654068490'].mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"A lung nodule (or mass) is a small abnormal area that is sometimes found during a CT scan of the chest. These scans are done for many reasons, such as part of lung cancer screening, or to check the lungs if you have symptoms. Most lung nodules seen on CT scans are not cancer. They are more often the result of old infections, scar tissue, or other causes. But tests are often needed to be sure a nodule is not cancer. [Reference](https://www.cancer.org/cancer/lung-cancer/detection-diagnosis-staging/lung-nodules.html#:~:text=A%20lung%20nodule%20(or%20mass,CT%20scans%20are%20not%20cancer.)) \n\nOn its [web page](https://www.uwhealth.org/thoracic-surgery/lung-nodules/11643#:~:text=Lung%20nodules%20are%20masses%20which,%2C%20diagnosed%2C%20and%20treated%20appropriately.) UW Health mentions that *Lung nodules ...can be early indicators of cancer, infection, or other lung diseases such as Pulmonary Fibrosis (PF)*.\n\n[This paper](https://www.journalpulmonology.org/en-novelties-in-imaging-in-pulmonary-articulo-S2531043719301795) stresses out the importance of CT scans for detecting Pulmonary fibrosis and precise measuring and characterizing of nodules. But it does not link between PF and nodules. I would appreciate if someone can give more information about it.\nIt can be seen that patients with FVC below 3L have many nodules, compared to the ones that have FVC in the normal range of 3 to 5L.\nI cannot claim this is representative result since I have tested it only for four randomly chosed IDs but it is worth checking if it is true.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"markdown","source":"TODO: Check correlations\n\nTODO: Convert categories to numbers\n\nTODO: Age wise smokers correlations","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}