{"cells":[{"metadata":{},"cell_type":"markdown","source":"### Importing the necessary libraries","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport pydicom as dcm\nimport time\nimport tensorflow as tf\nfrom matplotlib import pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## normalising the CT scans\nI wanted to normalise the CT scans to speed up training. At first, I wanted to compute the mean and std of the first image and then normalize the whole dataset using it; after all, a CT scan of a lung should be similar to another lung, right? But decided against it and decided that maybe computing the global mean and std might be worth the effort. So I wrote the following code.\n\nWhen I suspect the code is gonna take a long time I insert a print statement somewhere so that I know the code works. I coded the print statement so as to print the mean of images up to the patients that I opened their CT scans. I noticed that the the means starts out high and then goes down up to a negative value then goes up again, and the patern continues as the for loop goes on. I thought there was something wrong with my code.","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"code","source":"data_path = \"../input/osic-pulmonary-fibrosis-progression\"\ntrain_path = data_path+\"/train/\"\npatients = os.listdir(train_path)\ndef load_dcm(path):\n    dcm_data = dcm.dcmread(path)\n    pixels = dcm_data.pixel_array\n    mean, std = np.mean(pixels), np.std(pixels)\n    return mean, std\n    # return (pixels-mean)/std, dcm_data.get(\"ImagePositionPatient\")\nmean = []\nstd = []\nfor directory, _, files in os.walk(train_path):\n    for file in files:\n        try:\n            m, s = load_dcm(os.path.join(directory, file))\n            mean.append(m)\n            std.append(s)\n        except Exception:\n            print(os.path.join(directory, file))\n    print(np.mean(mean))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I tried to plot the distribution of the mean and turned out there are about two major clusters of means.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(12, 12))\nax.hist(mean, bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"ploting the std distribution yielded 3 clusters.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(12, 12))\nax.hist(std, bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So how do we go about this? do we normalize the the CT scans per cluster they belong to, or should we normalize each image with its own mean and std or do we just ignore this and just notmalize using the std and mean of the whole dataset?","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}