{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"# Pneumonia Detection Competition\n## Data Exploration\nWhat is pneumonia?\n\"Chest X-rays are currently the best available method for diagnosing pneumonia, playing a crucial role in clinical care and epidemiological studies. Pneumonia is responsible for more than 1 million hospitalizations and 50,000 deaths per year in the US alone.\" - [Link to Stanford ML Group Paper](https://stanfordmlgroup.github.io/projects/chexnet/)\n\n<img src=\"https://www.mayoclinic.org/-/media/671275b4a4e64a868f06eb8b18b002fa.jpg\" alt=\"drawing\" width=\"350\"/>"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pylab as plt\nimport pydicom\nimport os\nfrom os import listdir\nfrom os.path import isfile, join\n# print(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2f7f84831b2a8c7aef983aca9f81172d5e937333"},"cell_type":"markdown","source":"## Data Overview\n### Stage 1 Images - `stage_1_train_images.zip` and `stage_1_test_images.zip`\n- images for the current stage. Filenames are also patient names.\n\n### Stage 1 Labels - `stage_1_train_labels.csv` and Stage 1 Sample Submission `stage_1_sample_submission.csv`\n- Which provides the IDs for the test set, as well as a sample of what your submission should look like\n\n### Stage 1 Detailed Info - `stage_1_detailed_class_info.csv`\n- contains detailed information about the positive and negative classes in the training set, and may be used to build more nuanced models."},{"metadata":{"trusted":true,"_uuid":"7bf217d1ba26eb1a71701f2e6a5ceabfe35afe47"},"cell_type":"code","source":"# Images Example\ntrain_images_dir = '../input/stage_1_train_images/'\ntrain_images = [f for f in listdir(train_images_dir) if isfile(join(train_images_dir, f))]\ntest_images_dir = '../input/stage_1_test_images/'\ntest_images = [f for f in listdir(test_images_dir) if isfile(join(test_images_dir, f))]\nprint('5 Training images', train_images[:5]) # Print the first 5","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e29f78b7b7864b0598de8568cc4efc7c2022463e"},"cell_type":"code","source":"print('Number of train images:', len(train_images))\nprint('Number of test images:', len(test_images))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e374a680af821208159c00c18c662e0f48b82ef5"},"cell_type":"markdown","source":"## Plot a few training images from `stage_1_train_images.zip`"},{"metadata":{"trusted":true,"_uuid":"5245f0215e779539de41a2f2da38e35cc4307a30"},"cell_type":"code","source":"plt.style.use('default')\nfig=plt.figure(figsize=(20, 10))\ncolumns = 8; rows = 4\nfor i in range(1, columns*rows +1):\n    ds = pydicom.dcmread(train_images_dir + train_images[i])\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap=plt.cm.bone)\n    fig.add_subplot","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"39d546167d559bd02a6f60c1c8c2390e14e89352"},"cell_type":"markdown","source":"## Look at labels in `stage_1_train_labels.csv`"},{"metadata":{"trusted":true,"_uuid":"5da6f58f9ac66f036122027f1ab542f838fcc7d6"},"cell_type":"code","source":"train_labels = pd.read_csv('../input/stage_1_train_labels.csv')\ntrain_labels.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9a0d9430737f3355e2e9b39c6f831796acd4a300"},"cell_type":"markdown","source":"## Distribution of Positive Labels"},{"metadata":{"trusted":true,"_uuid":"e839869325afe418573215150de131478bb11a8f"},"cell_type":"code","source":"# Number of positive targets\nprint(round((8964 / (8964 + 20025)) * 100, 2), '% of the examples are positive')\npd.DataFrame(train_labels.groupby('Target')['patientId'].count())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bc9322d0c911ea1310cad969c98ab31ee0716428"},"cell_type":"code","source":"# Distribution of Target in Training Set\nplt.style.use('ggplot')\nplot = train_labels.groupby('Target') \\\n    .count()['patientId'] \\\n    .plot(kind='bar', figsize=(10,4), rot=0)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4f91219f2da00a6709a6bbbf094b4ba7e4b5d2c7"},"cell_type":"markdown","source":"## Size of the impacted area\nWe can make a new feature called \"area\" to the train labels data to see what the distribution of areas label look like."},{"metadata":{"trusted":true,"_uuid":"2ba5104921dccc040d02c4bcbf6dd4969ff45081"},"cell_type":"code","source":"plt.style.use('ggplot')\ntrain_labels['area'] = train_labels['width'] * train_labels['height']\nplot = train_labels['area'].plot(kind='hist',\n                          figsize=(10,4),\n                          bins=20,\n                          title='Distribution of Area within Image idenfitying a positive target')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"165b39d98bb44fd938a69179eef696a62411fa99"},"cell_type":"markdown","source":"# Plotting Boxes around Images\nThanks for plotting functions from @peterchang77 `https://www.kaggle.com/peterchang77/exploratory-data-analysis` !!!"},{"metadata":{"trusted":true,"_uuid":"1f89d74ed95de39ddb93aed96b82b381bb9f7473","_kg_hide-input":true},"cell_type":"code","source":"# Forked from `https://www.kaggle.com/peterchang77/exploratory-data-analysis`\ndef parse_data(df):\n    \"\"\"\n    Method to read a CSV file (Pandas dataframe) and parse the \n    data into the following nested dictionary:\n\n      parsed = {\n        \n        'patientId-00': {\n            'dicom': path/to/dicom/file,\n            'label': either 0 or 1 for normal or pnuemonia, \n            'boxes': list of box(es)\n        },\n        'patientId-01': {\n            'dicom': path/to/dicom/file,\n            'label': either 0 or 1 for normal or pnuemonia, \n            'boxes': list of box(es)\n        }, ...\n\n      }\n\n    \"\"\"\n    # --- Define lambda to extract coords in list [y, x, height, width]\n    extract_box = lambda row: [row['y'], row['x'], row['height'], row['width']]\n\n    parsed = {}\n    for n, row in df.iterrows():\n        # --- Initialize patient entry into parsed \n        pid = row['patientId']\n        if pid not in parsed:\n            parsed[pid] = {\n                'dicom': '../input/stage_1_train_images/%s.dcm' % pid,\n                'label': row['Target'],\n                'boxes': []}\n\n        # --- Add box if opacity is present\n        if parsed[pid]['label'] == 1:\n            parsed[pid]['boxes'].append(extract_box(row))\n\n    return parsed\n\nparsed = parse_data(train_labels)\n\ndef draw(data):\n    \"\"\"\n    Method to draw single patient with bounding box(es) if present \n\n    \"\"\"\n    # --- Open DICOM file\n    d = pydicom.read_file(data['dicom'])\n    im = d.pixel_array\n\n    # --- Convert from single-channel grayscale to 3-channel RGB\n    im = np.stack([im] * 3, axis=2)\n\n    # --- Add boxes with random color if present\n    for box in data['boxes']:\n        #rgb = np.floor(np.random.rand(3) * 256).astype('int')\n        rgb = [255, 251, 204] # Just use yellow\n        im = overlay_box(im=im, box=box, rgb=rgb, stroke=15)\n\n    plt.imshow(im, cmap=plt.cm.gist_gray)\n    plt.axis('off')\n\ndef overlay_box(im, box, rgb, stroke=2):\n    \"\"\"\n    Method to overlay single box on image\n\n    \"\"\"\n    # --- Convert coordinates to integers\n    box = [int(b) for b in box]\n    \n    # --- Extract coordinates\n    y1, x1, height, width = box\n    y2 = y1 + height\n    x2 = x1 + width\n\n    im[y1:y1 + stroke, x1:x2] = rgb\n    im[y2:y2 + stroke, x1:x2] = rgb\n    im[y1:y2, x1:x1 + stroke] = rgb\n    im[y1:y2, x2:x2 + stroke] = rgb\n\n    return im","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f589a0ceb6816af4e140f48a4eac0c945cc15150"},"cell_type":"code","source":"plt.style.use('default')\nfig=plt.figure(figsize=(20, 10))\ncolumns = 8; rows = 4\nfor i in range(1, columns*rows +1):\n    fig.add_subplot(rows, columns, i)\n    draw(parsed[train_labels['patientId'].unique()[i]])\n    fig.add_subplot","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b5152e8daad73f599aa4787b29db8517717ff338"},"cell_type":"markdown","source":"## A closer look at a Positive and Negative Example"},{"metadata":{"trusted":true,"_uuid":"04bde41d3df2d3f23a8f23a7b5ecbe2aa1f1e388"},"cell_type":"code","source":"fig=plt.figure(figsize=(20, 10))\ndraw(parsed[train_labels['patientId'].loc[20]])\nplt.show()\nfig=plt.figure(figsize=(20, 10))\ndraw(parsed[train_labels['patientId'].loc[10]])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d3937be9a76a488869e2e8a99c6a8caecaeabf67"},"cell_type":"markdown","source":"# EDA of Detailed Class Info"},{"metadata":{"trusted":true,"_uuid":"cba60817baa246cf91609dfefc6a5a5457c3bff1"},"cell_type":"code","source":"detailed_class_info = pd.read_csv('../input/stage_1_detailed_class_info.csv')\ndetailed_class_info.groupby('class').count()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b3c56ecaae1f2327cfcafb53461ab5138ce923e0"},"cell_type":"code","source":"plt.style.use('ggplot')\nplot = detailed_class_info.groupby('class').count().plot(kind='bar',\n                                                  rot=0,\n                                                  title='Count of Class Labels',\n                                                  figsize=(10,4))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e75d1e4c160e2df32b93b7544acb9e641ef7acf2"},"cell_type":"code","source":"count_labels_per_patient = detailed_class_info.groupby('patientId').count()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"79bf87bf0883fdc36adf1fe44a291f00464248b1"},"cell_type":"markdown","source":"# Images of Each Label Type"},{"metadata":{"trusted":true,"_uuid":"6e9e59423d0fa64a34ece255eeb95499d46b74d2"},"cell_type":"code","source":"opacity = detailed_class_info \\\n    .loc[detailed_class_info['class'] == 'Lung Opacity'] \\\n    .reset_index()\nnot_normal = detailed_class_info \\\n    .loc[detailed_class_info['class'] == 'No Lung Opacity / Not Normal'] \\\n    .reset_index()\nnormal = detailed_class_info \\\n    .loc[detailed_class_info['class'] == 'Normal'] \\\n    .reset_index()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"64714ed40a40b4beb311951b948b554ec36b6dbe"},"cell_type":"markdown","source":"## ** Lung Opacity** Examples"},{"metadata":{"trusted":true,"_uuid":"e467049de908af54f437910f58391b2d5b8a9dea"},"cell_type":"code","source":"plt.style.use('default')\nfig=plt.figure(figsize=(20, 10))\ncolumns = 8; rows = 4\nfor i in range(1, columns*rows +1):\n    fig.add_subplot(rows, columns, i)\n    draw(parsed[opacity['patientId'].unique()[i]])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"99c8e8d96374ee69c0396c6fa22fca406f4887fa"},"cell_type":"markdown","source":"## ** No Lung Opacity / Not Normal** Examples"},{"metadata":{"trusted":true,"_uuid":"22faf056244b7cb6a4db49d4de5d3b34c631e422"},"cell_type":"code","source":"plt.style.use('default')\nfig=plt.figure(figsize=(20, 10))\ncolumns = 8; rows = 4\nfor i in range(1, columns*rows +1):\n    fig.add_subplot(rows, columns, i)\n    draw(parsed[not_normal['patientId'].loc[i]])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"45fb88335a8fee4d63261031573a39a052ca8ed6"},"cell_type":"markdown","source":"## **Normal** Examples"},{"metadata":{"trusted":true,"_uuid":"43c9d72f33162e95b3872271c0a335923139f4f7"},"cell_type":"code","source":"plt.style.use('default')\nfig=plt.figure(figsize=(20, 10))\ncolumns = 8; rows = 4\nfor i in range(1, columns*rows +1):\n    fig.add_subplot(rows, columns, i)\n    draw(parsed[normal['patientId'].loc[i]])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"349a9b8464faeea65c8d8faf45b98399f6044caf"},"cell_type":"markdown","source":"# Side By Side Compare of Opacity/Not Normal/Normal"},{"metadata":{"trusted":true,"_uuid":"d6366f04cc4844386dd079234c99fd90bbaa00ba"},"cell_type":"code","source":"fig=plt.figure(figsize=(20, 10))\ncolumns = 3; rows = 1\nfig.add_subplot(rows, columns, 1).set_title(\"Normal\", fontsize=30)\ndraw(parsed[normal['patientId'].unique()[0]])\nfig.add_subplot(rows, columns, 2).set_title(\"Not Normal\", fontsize=30)\n# ax2.set_title(\"Not Normal\", fontsize=30)\ndraw(parsed[not_normal['patientId'].unique()[0]])\nfig.add_subplot(rows, columns, 3).set_title(\"Opacity\", fontsize=30)\n# ax3.set_title(\"Opacity\", fontsize=30)\ndraw(parsed[opacity['patientId'].unique()[0]])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a4f261a65730f713caeea73cd1e95636b55b5aff","_kg_hide-input":true},"cell_type":"code","source":"fig=plt.figure(figsize=(20, 10))\ncolumns = 3; rows = 1\nfig.add_subplot(rows, columns, 1).set_title(\"Normal\", fontsize=30)\ndraw(parsed[normal['patientId'].unique()[1]])\nfig.add_subplot(rows, columns, 2).set_title(\"Not Normal\", fontsize=30)\n# ax2.set_title(\"Not Normal\", fontsize=30)\ndraw(parsed[not_normal['patientId'].unique()[1]])\nfig.add_subplot(rows, columns, 3).set_title(\"Opacity\", fontsize=30)\n# ax3.set_title(\"Opacity\", fontsize=30)\ndraw(parsed[opacity['patientId'].unique()[1]])","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"e62c9df343a31202afd4163aa1c258a9594324e9"},"cell_type":"code","source":"fig=plt.figure(figsize=(20, 10))\ncolumns = 3; rows = 1\nfig.add_subplot(rows, columns, 1).set_title(\"Normal\", fontsize=30)\ndraw(parsed[normal['patientId'].unique()[2]])\nfig.add_subplot(rows, columns, 2).set_title(\"Not Normal\", fontsize=30)\n# ax2.set_title(\"Not Normal\", fontsize=30)\ndraw(parsed[not_normal['patientId'].unique()[2]])\nfig.add_subplot(rows, columns, 3).set_title(\"Opacity\", fontsize=30)\n# ax3.set_title(\"Opacity\", fontsize=30)\ndraw(parsed[opacity['patientId'].unique()[2]])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ad8f70ad1d0463fd81028e9d079d47d096658e92"},"cell_type":"markdown","source":"# Number of Labels Per Patientid\npatients have 0-4 labels, patients can have multiple labels.\n(Only including ones with at least on label)"},{"metadata":{"trusted":true,"_uuid":"76379bf96d8ce6c4aeca3f3066fb5fc35eacef75"},"cell_type":"code","source":"count_labels_per_patient.reset_index().groupby('class').count()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"943f026878050e96bac6c14323a20da9b6fc904a"},"cell_type":"code","source":"# Patients with 4 Labels\ncount_labels_per_patient.sort_values('class', ascending=False).head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ac469cabd9d92f863a8e1e4211beba9024d6ecb4"},"cell_type":"code","source":"detailed_class_info.loc[detailed_class_info['patientId'] == '7d674c82-5501-4730-92c5-d241fd6911e7']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"748222193ae6c39b118494ed0510b181eccc588c"},"cell_type":"markdown","source":"# Closer Look of Each Type"},{"metadata":{"trusted":true,"_uuid":"5dab035c855c12641e0663cc992637f2a22d0217"},"cell_type":"code","source":"fig=plt.figure(figsize=(20, 10))\nplt.suptitle('\"Lung Opacity\" Example', fontsize=16)\ndraw(parsed['7d674c82-5501-4730-92c5-d241fd6911e7'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dac1017acab9bc62cb6c77f523af04fed32875a0"},"cell_type":"code","source":"not_normal = detailed_class_info.loc[detailed_class_info['class'] == 'No Lung Opacity / Not Normal']\nnot_normal_example = not_normal['patientId']\nfig=plt.figure(figsize=(20, 10))\nplt.suptitle('\"No Lung Opacity / Not Normal\" Example', fontsize=16)\ndraw(parsed['019e035e-2f82-4c66-a198-57422a27925f'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7ecca40058e723164bac83098c76774616fc6972"},"cell_type":"code","source":"fig=plt.figure(figsize=(20, 10))\nplt.suptitle('\"Normal\" Example', fontsize=16)\ndraw(parsed['003d8fa0-6bf1-40ed-b54c-ac657f8495c5'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"68ea6d05b2433e2035d460f0a36a05984590dc42"},"cell_type":"markdown","source":"# What does the submission look like?"},{"metadata":{"trusted":true,"_uuid":"399af37e4e019e3bbe6c91f116f733376d77fb22"},"cell_type":"code","source":"pd.read_csv('../input/stage_1_sample_submission.csv').head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dc4b38cadcc47c00d7079c9375709c988f59a05f"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}