{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport cv2\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nsns.set()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Analyzing the Train dataset"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"path = '/kaggle/input/ranzcr-clip-catheter-line-classification/'\ntrain = pd.read_csv(path + 'train.csv')\ntrain.head(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Shape of Training File: ', train.shape)\nprint('Unique Patients: ', len(train['PatientID'].unique()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"classes_to_predict = ['ETT - Abnormal', 'ETT - Borderline','ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline',\n                      'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal','CVC - Borderline', 'CVC - Normal', \n                      'Swan Ganz Catheter Present']\n\nfor class_ in classes_to_predict:\n    number_of_positives = len(train[train[class_] == 1])\n    print(class_, '|', number_of_positives, '|', number_of_positives/len(train))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that the majority of classes are about normal insertions. However, we see that the percentage is summing up more than 100%, so maybe we can have more than on catheter inserted into someone at the same time, let's verify it."},{"metadata":{"trusted":true},"cell_type":"code","source":"train.set_index('StudyInstanceUID', inplace=True)\ntrain.drop(columns='PatientID', inplace=True)\ntrain['sum'] = train.sum(axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train[train['sum'] > 1].tail(4)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In fact, we have patients with more than one insertion and even more than one insertion of the same type, such as a Normal CVC and a Borderline CVC. Let's verify how this is distributed."},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(15,9))\nplt.title('Number of Catheters Inserted Distribution')\ntrain['sum'].hist()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, this csv has only the identification and the classes for each of the images on the train/ directory. As we could see, some images may have more than one positive (1) label and the 'Normal' labels are the majority of the classes, which creates an unbalance on the dataset as could be expected from this kind of problem."},{"metadata":{},"cell_type":"markdown","source":"# Analyzing the train annotations"},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '/kaggle/input/ranzcr-clip-catheter-line-classification/'\ntrain_annotations = pd.read_csv(path + 'train_annotations.csv')\ntrain_annotations.head(1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This seems to be the annotation with the segmentation of the tube inside each image. This could be useful at some point, let's see if all of the training data has this kind of annotation. As this appears to be a long format, let's also verify if we have more than one row per ID."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_annotations['StudyInstanceUID'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As expected, we indeed have more than one row per id."},{"metadata":{"trusted":true},"cell_type":"code","source":"annotations_ids = train_annotations['StudyInstanceUID'].unique()\nprint('Total number of training images: ', len(train.reset_index()['StudyInstanceUID'].unique()))\nprint('With annotations: ', len(train.reset_index()[train.reset_index()['StudyInstanceUID'].isin(annotations_ids)]['StudyInstanceUID'].unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, we have less than one third of the training examples with annotations."},{"metadata":{},"cell_type":"markdown","source":"## Visual inspection of the training set"},{"metadata":{"trusted":true},"cell_type":"code","source":"for class_ in classes_to_predict:\n    ids = train[train[class_] == 1].sample(3).index\n    \n    img = path + 'train/' + ids[0] + '.jpg'\n    img = cv2.imread(img)\n    plt.figure(figsize=(15,9))\n    plt.imshow(img)\n    plt.grid(None)\n    plt.title(class_)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}