{"cells":[{"metadata":{"_uuid":"684c77cddc74e33287f16ed67c39b61a1078f435"},"cell_type":"markdown","source":"Let's start with exploring the dataset and see how the images look ?"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4879465a768b670cd74dbac99011f1fbaf86e9db"},"cell_type":"markdown","source":"Now I will import the data into a dataframe to be able to use it later"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"sample = pd.read_csv(\"../input/stage_1_sample_submission.csv\")\ntrain_labels  = pd.read_csv(\"../input/stage_1_train_labels.csv\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"919de4c1668cbba13367642044a634e05dfd8c9b"},"cell_type":"markdown","source":"We should know the number of data points in each of the two classes we have to determine which performance metric must be used."},{"metadata":{"trusted":true,"_uuid":"8f577e9b3e11b0bc589b3d31a20fa8fdef61d9e8"},"cell_type":"code","source":"print(\"number of points in the 0 Neg class is %i\" % (train_labels.drop_duplicates('patientId', keep = 'first'))[(train_labels.drop_duplicates('patientId', keep = 'first'))['Target'] == 0].shape[0])\nprint(\"number of points in the 1 pos class is %i\" % (train_labels.drop_duplicates('patientId', keep = 'first'))[(train_labels.drop_duplicates('patientId', keep = 'first'))['Target'] == 1].shape[0]) ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"79836414cad507021714d257e1f5c9f3ca361421"},"cell_type":"markdown","source":"I need to explore the number of data points in the dataset to know how many bounding boxes i have"},{"metadata":{"trusted":true,"_uuid":"e9707bcaa4a0772e1139291e2f0682ac634c9d9e"},"cell_type":"code","source":"print(\"The number of traning examples(data points) = %i \" % train_labels.shape[0])\nprint(\"The number of features we have = %i \" % train_labels.shape[1])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8333959b48ca3645f2f48d0153177332e4b50c5b"},"cell_type":"markdown","source":"After knowing the numbers of examples in each class it seems that the data is skewed that is one class is domenating the other in the RSNA dataset so I should use F1 score later instead of accuracy.\nNow i will explore some statistics using Dataframe.describe() ..."},{"metadata":{"trusted":true,"_uuid":"c1cae7fc23fb12d610e6ea1197bad2cec793ea0a"},"cell_type":"code","source":"train_labels.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dfa9630958e1b7b4b59233ff14157f35ea567f9d"},"cell_type":"markdown","source":"I will see how much nulls we have but since the data of the neg. class have nulls in the x , y and dimentions of the bounding boxes for the images, it seems that we will only need to drop them later when drawing."},{"metadata":{"trusted":true,"_uuid":"21ab951c9172deed4c2595f25905a28b48b10a07"},"cell_type":"code","source":"train_labels.isna().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6acd832973d4ff64e10067d863f7d28b7de97af0"},"cell_type":"code","source":"train_labels","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"009a2d3fbd0173ba28ffbe8621f045994db90a04"},"cell_type":"code","source":"import pydicom\n\nPathDicom = \"../input/stage_1_train_images/\"\nimages = []  # create an empty list\nfor dirName, subdirList, fileList in os.walk(PathDicom):\n    for filename in fileList:\n        images.append((os.path.join(dirName,filename),filename))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"0df7f73896b1768e606e7c3ac4ba551b443c4494"},"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.patches as patches\n\nf, ax = plt.subplots(2, 2, figsize=(25,20))\nimage_index = 0\nfor i in ax:\n    for j in i:\n        print(j)\n        data = pydicom.read_file(images[image_index][0])\n        print(\"/////////////////////////////////////////////////\\n\", data, \"\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\/n\" )\n        image = data.pixel_array\n        j.imshow(image) \n        rows = train_labels[train_labels[\"patientId\"].str.match(images[image_index][1][:-4])]\n        print( rows )\n        for index, row in rows.iterrows():\n            rect = patches.Rectangle((row['x'],row['y']),row['width'],row['height'],linewidth=5,edgecolor='b',facecolor='none')\n            j.add_patch(rect)\n        image_index += 1\n    ","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}