{"cells":[{"metadata":{"_uuid":"0cbce5cc636ba65b81d858c596e39aa532728780"},"cell_type":"markdown","source":"# Overview\nThe notebook aims to get a better feeling for the data and more importantly the distributions of values. We take the labels and combine them with the detailed class info and try and determine what the biggest challenges of the prediction might be. "},{"metadata":{"trusted":true,"_uuid":"ae448b0ed29194053d40ebd29b2fa03982468552"},"cell_type":"code","source":"%matplotlib inline\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pydicom\nimport pandas as pd\nfrom glob import glob\nimport os\nfrom matplotlib.patches import Rectangle\ndet_class_path = '../input/stage_1_detailed_class_info.csv'\nbbox_path = '../input/stage_1_train_labels.csv'\ndicom_dir = '../input/stage_1_train_images/'","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f4ff05105153200d3dc3f382b5a24b6e6dd80291"},"cell_type":"markdown","source":"# Detailed Class Info\nHere we show the image-level labels for the scans. The most interesting group here is the `No Lung Opacity / Not Normal` since they are cases that look like opacity but are not. So the first step might be to divide the test images into clear groups and then only perform the bounding box prediction on the suspicious images."},{"metadata":{"trusted":true,"_uuid":"a0a43187041d2773c035b62f68a9687811f9fcc2"},"cell_type":"code","source":"det_class_df = pd.read_csv(det_class_path)\nprint(det_class_df.shape[0], 'class infos loaded')\nprint(det_class_df['patientId'].value_counts().shape[0], 'patient cases')\ndet_class_df.groupby('class').size().plot.bar()\ndet_class_df.sample(3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4ca820c61b58fe03a613bce38c47786b8a8ef399"},"cell_type":"markdown","source":"# Load the Bounding Box Data\nHere we show the bounding boxes"},{"metadata":{"trusted":true,"_uuid":"cc03f5c79bee9f015c12dc8181f7997df584fbb8"},"cell_type":"code","source":"bbox_df = pd.read_csv(bbox_path)\nprint(bbox_df.shape[0], 'boxes loaded')\nprint(bbox_df['patientId'].value_counts().shape[0], 'patient cases')\nbbox_df.sample(3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c5cc20e2eb8e9cbecf4494fe5264397438fb73a1"},"cell_type":"markdown","source":"# Combine Boxes and Labels\nHere we bring the labels and the boxes together and now we can focus on how the boxes look on the images"},{"metadata":{"trusted":true,"_uuid":"ca18783acf240faea13a23fd06fba7b41f9ea71d"},"cell_type":"code","source":"# we first try a join and see that it doesn't work (we end up with too many boxes)\ncomb_bbox_df = pd.merge(bbox_df, det_class_df, how='inner', on='patientId')\nprint(comb_bbox_df.shape[0], 'combined cases')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"15ab554a46421da28b29d7932dcad07b5146203f"},"cell_type":"markdown","source":"## Concatenate\nWe have to concatenate the two datasets and then we get class and target information on each region"},{"metadata":{"trusted":true,"_uuid":"91447e135cd4525c0869692ab6a6269e5bc9f30f"},"cell_type":"code","source":"comb_bbox_df = pd.concat([bbox_df, \n                        det_class_df.drop('patientId',1)], 1)\nprint(comb_bbox_df.shape[0], 'combined cases')\ncomb_bbox_df.sample(3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6f29b7e35b8f879687bcfeadf8a93c708c94d9a7"},"cell_type":"markdown","source":"# Distribution of Boxes and Labels\nThe values below show the number of boxes and the patients that have that number. "},{"metadata":{"trusted":true,"_uuid":"569f38045b0359ae0635bca97dbbe600785ff565"},"cell_type":"code","source":"box_df = comb_bbox_df.groupby('patientId').\\\n    size().\\\n    reset_index(name='boxes')\ncomb_box_df = pd.merge(comb_bbox_df, box_df, on='patientId')\nbox_df.\\\n    groupby('boxes').\\\n    size().\\\n    reset_index(name='patients')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"277b277b5c63562afa5a66c0714a94bb139979ec"},"cell_type":"markdown","source":"# How are class and target related?\nI assume that all the `Target=1` values fall in the `Lung Opacity` class, but it doesn't hurt to check."},{"metadata":{"trusted":true,"_uuid":"2386013f7461e1407f1e782865228d146264bff0"},"cell_type":"code","source":"comb_bbox_df.groupby(['class', 'Target']).size().reset_index(name='Patient Count')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c5c58ddf2a9a7350650586865166ae0d5ed9a0cf"},"cell_type":"markdown","source":"# Images\nNow that we have the boxes and labels loaded we can examine a few images."},{"metadata":{"trusted":true,"_uuid":"46fe765772e24397a94295a70dbde46254c4cc8b"},"cell_type":"code","source":"image_df = pd.DataFrame({'path': glob(os.path.join(dicom_dir, '*.dcm'))})\nimage_df['patientId'] = image_df['path'].map(lambda x: os.path.splitext(os.path.basename(x))[0])\nprint(image_df.shape[0], 'images found')\nimg_pat_ids = set(image_df['patientId'].values.tolist())\nbox_pat_ids = set(comb_box_df['patientId'].values.tolist())\n# check to make sure there is no funny business\nassert img_pat_ids.union(box_pat_ids)==img_pat_ids, \"Patient IDs should be the same\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"49b1259692d6246459452f80ff2d7e1d4755d28d"},"cell_type":"code","source":"image_bbox_df = pd.merge(comb_box_df, \n                         image_df, \n                         on='patientId',\n                        how='left').sort_values('patientId')\nprint(image_bbox_df.shape[0], 'image bounding boxes')\nimage_bbox_df.head(5)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c6299cf5b16c13e9eb07c69a84a934fec4cf8b6f"},"cell_type":"markdown","source":"# Enrich the image fields\nWe have quite a bit of additional data in the DICOM header we can easily extract to help learn more about the patient like their age, view position and gender which can make the model much more precise"},{"metadata":{"trusted":true,"_uuid":"6fd8a408e2999759d3cc7500f7cb6d03d767103d"},"cell_type":"code","source":"from scipy.ndimage import zoom\nDCM_TAG_LIST = ['PatientAge', 'BodyPartExamined', 'ViewPosition', 'PatientSex']\ndef process_dicom(in_path, in_rows):\n    c_dicom = pydicom.read_file(in_path, stop_before_pixels=False)\n    tag_dict = {c_tag: getattr(c_dicom, c_tag, '') \n         for c_tag in DCM_TAG_LIST}\n    tag_dict['path'] = in_path\n    tag_dict['PatientAge'] = int(tag_dict['PatientAge'])\n    tag_dict['boxes'] = in_rows.shape[0]\n    tag_dict['class'] = in_rows['class'].iloc[0]\n    return tag_dict, c_dicom.pixel_array\ndef create_seg_image(base_img, box_rows, out_shape = (512, 512)):\n    c_size = base_img.shape\n    x_fact = out_shape[0]/c_size[0]\n    y_fact = out_shape[1]/c_size[1]\n    rs_img = zoom(base_img, (x_fact, y_fact))\n    mk_img = np.zeros(rs_img.shape, dtype=bool)\n    x_vec = box_rows['x'].map(lambda x: x*x_fact).values.astype(int)\n    y_vec = box_rows['y'].map(lambda y: y*y_fact).values.astype(int)\n    w_vec = box_rows['width'].map(lambda w: w*x_fact).values.astype(int)\n    h_vec = box_rows['height'].map(lambda h: h*y_fact).values.astype(int)\n    for x, y, w, h in zip(x_vec,\n                         y_vec, \n                         w_vec, \n                         h_vec):\n        mk_img[y:(y+h), x:(x+w)] = True\n    return rs_img, mk_img","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9ec768720bae3e01318e3585c3bf527f7b5d125d"},"cell_type":"code","source":"t_path, t_rows = next(iter(image_bbox_df.query('boxes>3').groupby(['path'])))\nd_info, d_img = process_dicom(t_path, t_rows)\nprint(d_info)\nfig, (ax0, ax1, ax2) = plt.subplots(1, 3, figsize = (10, 5))\n\nax0.imshow(d_img, cmap='bone')\nax0.set_title('Standard Overlay')\nfor _, c_row in t_rows.iterrows():\n    ax0.add_patch(Rectangle(xy=(c_row['x'], c_row['y']),\n                 width=c_row['width'],\n                 height=c_row['height'],\n                           alpha=5e-1))\n\nt_img, t_mask = create_seg_image(d_img, t_rows)\nax1.imshow(t_img)\nax1.set_title(\"Downscaled Image\")\nax2.imshow(t_mask)\nax2.set_title('Mask Image')\nt_rows","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4473489234288741452fd81ece375cd61fa0780b"},"cell_type":"code","source":"out_img_size = (512, 512)\nkeep_patients = 15000","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d4b779bd945a756343d07d52dd1fe653b165b10e"},"cell_type":"code","source":"balanced_patient_list = image_bbox_df.\\\n    query('Target==1')[['Target', 'path']].\\\n    drop_duplicates().\\\n    groupby('Target').\\\n    apply(lambda x: x.sample(keep_patients, random_state=2018, replace=True)).\\\n    reset_index(drop=True)['path'].values\nnp.random.seed(2018)\nall_groups = list(image_bbox_df[image_bbox_df['path'].isin(balanced_patient_list)].groupby(['path']))\nkeep_idx = np.random.choice(range(len(all_groups)), keep_patients)\nall_groups = [all_groups[i] for i in keep_idx]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"31ed9db8106aa30171b6abe500fed1784871972f"},"cell_type":"code","source":"import h5py\nfrom tqdm import tqdm_notebook\n\nwith h5py.File('train_segs.h5', 'w') as f:\n    image_out = f.create_dataset('image', \n                                 shape=(len(all_groups), out_img_size[0], out_img_size[1], 1),\n                                 dtype=np.uint8)\n    mask_out = f.create_dataset('mask', \n                                 shape=(len(all_groups), out_img_size[0], out_img_size[1], 1),\n                                 dtype=bool,\n                                 compression='gzip')\n    d_info, _ = process_dicom(t_path, t_rows)\n    \n    key_ds_out = {}\n    for k,v in d_info.items():\n        if isinstance(v, str):\n            key_ds_out[k] = f.create_dataset(k, \n                                 shape=(len(all_groups),),\n                                 dtype='S{}'.format(len(v)+2))\n        elif isinstance(v, int):\n            key_ds_out[k] = f.create_dataset(k, \n                                 shape=(len(all_groups),),\n                                 dtype=int)\n        else:\n            print('Unsupported key-type {}: {}'.format(type(v), v))\n    for i, (c_path, c_rows) in enumerate(tqdm_notebook(all_groups)):\n        c_info, c_raw_img = process_dicom(c_path, c_rows)\n        c_img, c_mask = create_seg_image(c_raw_img, c_rows, out_shape=out_img_size)\n        image_out[i, :, :, 0] = c_img\n        mask_out[i, :, :, 0] = c_mask\n        for k in key_ds_out.keys():\n            if k in c_info:\n                if isinstance(c_info[k], str):\n                    key_ds_out[k][i] = c_info[k].encode('ascii')\n                else:\n                    key_ds_out[k][i] = c_info[k]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ec702149a628bec091f6affdbb1a690fa38b5c78"},"cell_type":"code","source":"!ls -lh *.h5\nwith h5py.File('train_segs.h5', 'r') as f:\n    for k in f.keys():\n        print(k, f[k].shape, f[k].dtype)\n        print(k, f[k][0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ce24c231f3547a4e267c7a0baaa298b56b65f94a"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}