{"cells":[{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">RANZCR CLiP - Catheter and Line Position Challenge</h1>\n\n![](https://storage.googleapis.com/kaggle-organizations/3765/thumbnail.png?r=270)"},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Table of Contents</h1>\n\n* [The Competition](#competition)\n\n* [Objective](#obj)\n\n* [Dataset](#ds)\n\n* [Importing Necessary Libraries](#libs)\n\n* [Training Dataset](#trainds)\n\n* [Image Annotations](#annot)\n\n* [Image Information](#image)\n\n* [Category Overlap](#overlap)"},{"metadata":{},"cell_type":"markdown","source":"<a id = \"competition\"></a>\n\n<h1 style=\"border:2px solid Purple;text-align:center\">The Competition</h1>\n\nSerious complications can occur as a result of malpositioned lines and tubes in patients. Doctors and nurses frequently use checklists for placement of lifesaving equipment to ensure they follow protocol in managing patients. Yet, these steps can be time consuming and are still prone to human error, especially in stressful situations when hospitals are at capacity.\n\nHospital patients can have catheters and lines inserted during the course of their admission and serious complications can arise if they are positioned incorrectly. Nasogastric tube malpositioning into the airways has been reported in up to 3% of cases, with up to 40% of these cases demonstrating complications [1-3]. Airway tube malposition in adult patients intubated outside the operating room is seen in up to 25% of cases [4,5]. The likelihood of complication is directly related to both the experience level and specialty of the proceduralist. Early recognition of malpositioned tubes is the key to preventing risky complications (even death), even more so now that millions of COVID-19 patients are in need of these tubes and lines.\n\nThe gold standard for the confirmation of line and tube positions are chest radiographs. However, a physician or radiologist must manually check these chest x-rays to verify that the lines and tubes are in the optimal position. Not only does this leave room for human error, but delays are also common as radiologists can be busy reporting other scans. Deep learning algorithms may be able to automatically detect malpositioned catheters and lines. Once alerted, clinicians can reposition or remove them to avoid life-threatening complications.\n\nThe Royal Australian and New Zealand College of Radiologists (RANZCR) is a not-for-profit professional organisation for clinical radiologists and radiation oncologists in Australia, New Zealand, and Singapore. The group is one of many medical organisations around the world (including the NHS) that recognizes malpositioned tubes and lines as preventable. RANZCR is helping design safety systems where such errors will be caught."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"<a id = \"obj\"></a>\n\n<h1 style=\"border:2px solid Purple;text-align:center\">Objective</h1>\n\nIn this competition, you’ll detect the presence and position of catheters and lines on chest x-rays. Use machine learning to train and test your model on 40,000 images to categorize a tube that is poorly placed.\n"},{"metadata":{},"cell_type":"markdown","source":"<a id = \"eval\"></a>\n\n<h1 style=\"border:2px solid Purple;text-align:center\">Evaluation Metric</h1>\n\nSubmissions are evaluated on area under the ROC curve between the predicted probability and the observed target.\n\nTo calculate the final score, AUC is calculated for each of the 11 labels, then averaged. The score is then the average of the individual AUCs of each predicted column.\n\nIn case you need a python code [Source](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203581#1115260)\n\n> np.mean([roc_auc_score(y_true[:, i], y_pred[:, i]) for i in range(11)])"},{"metadata":{},"cell_type":"markdown","source":"<a id = \"ds\"></a>\n\n<h1 style=\"border:2px solid Purple;text-align:center\">Dataset</h1>\n\nThe dataset has been labelled with a set of definitions to ensure consistency with labelling. The normal category includes lines that were appropriately positioned and did not require repositioning. The borderline category includes lines that would ideally require some repositioning but would in most cases still function adequately in their current position. The abnormal category included lines that required immediate repositioning."},{"metadata":{},"cell_type":"markdown","source":"<a id = \"libs\"></a>\n<h1 style=\"border:2px solid Purple;text-align:center\">Importing Necessary Libraries</h1>"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport cv2\nfrom ast import literal_eval\nfrom tqdm import tqdm\nimport os\n\nimport plotly_express as px\nimport plotly.graph_objects as go\nimport plotly.figure_factory as ff\nfrom plotly.subplots import make_subplots\nimport matplotlib.pyplot as plt\nfrom matplotlib_venn import *\n\nfrom plotly.offline import init_notebook_mode\ninit_notebook_mode()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id = \"trainds\"></a>\n<h1 style=\"border:2px solid Purple;text-align:center\">Training Dataset</h1>"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"BASE_DIR = \"../input/ranzcr-clip-catheter-line-classification/\"\ndf = pd.read_csv(BASE_DIR + 'train.csv')\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(f\"The size of the dataset is {df.shape[0]} and it contains {df['PatientID'].nunique()} patient ids\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Columns\n\n* **StudyInstanceUID** - unique ID for each image\n* **ETT - Abnormal** - endotracheal tube placement abnormal\n* **ETT - Borderline** - endotracheal tube placement borderline abnormal\n* **ETT - Normal** - endotracheal tube placement normal\n* **NGT - Abnormal** - nasogastric tube placement abnormal\n* **NGT - Borderline** - nasogastric tube placement borderline abnormal\n* **NGT - Incompletely Imaged** - nasogastric tube placement inconclusive due to imaging\n* **NGT - Normal** - nasogastric tube placement borderline normal\n* **CVC - Abnormal** - central venous catheter placement abnormal\n* **CVC - Borderline** - central venous catheter placement borderline abnormal\n* **CVC - Normal** - central venous catheter placement normal\n* **Swan Ganz Catheter Present**\n* **PatientID** - unique ID for each patient in the dataset"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"idCounts = df['PatientID'].value_counts().reset_index()\nidCounts.columns = ['PatientID', 'Number of Observations']\nidCounts = idCounts.sort_values(by = 'Number of Observations', ascending = False)\nidCounts.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig = px.histogram(idCounts, 'Number of Observations', title = 'Distribution of Number of Observations per PatientIDs', template = 'ggplot2')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most of the PatientID's have only one observation in the dataset, but there are some PatientIDs with as many as 100+ records in the training dataset. "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"categories = ['ETT - Abnormal', 'ETT - Borderline',\n       'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline',\n       'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal',\n       'CVC - Borderline', 'CVC - Normal','Swan Ganz Catheter Present']\ncategoryCounts = df[categories].sum(axis = 0).reset_index()\ncategoryCounts.columns = ['Malpositions', 'Number of Observations']","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig = px.bar(categoryCounts, y = 'Malpositions', x = 'Number of Observations', template = 'seaborn', text = 'Number of Observations', title = 'Line and tube positions')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"countsDf = {'Type' : [], 'Malposition'  : [], 'Num of Observations' : []}\nfor Type in ['Normal','Abnormal','Borderline']:\n    for malposition in ['ETT','NGT','CVC']:\n        colName = f'{malposition} - {Type}'\n        countsDf['Type'].append(Type)\n        countsDf['Malposition'].append(malposition)\n        val = df[colName].sum(axis = 0)\n        countsDf['Num of Observations'].append(val)\ncountsDf = pd.DataFrame(countsDf)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig = px.bar(countsDf, x = 'Num of Observations', y = 'Type', color = 'Malposition', barmode = 'stack', \n             color_discrete_map={'ETT' : '#a2885e', 'NGT' : '#e9cf87', 'CVC' : '#f1efd9'}, template = 'plotly_dark')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id = \"annot\"></a>\n<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations</h1>"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"annotations = pd.read_csv(BASE_DIR + 'train_annotations.csv')\nannotations['data'] = annotations['data'].apply(literal_eval)\nannotations.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"IMAGE_DIR_TRAIN = BASE_DIR + 'train/'\nIMAGE_DIR_TEST = BASE_DIR + 'test/'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def plot_image_and_annotations(image_uid, title,singleLabel = True):\n    image_path = IMAGE_DIR_TRAIN + image_uid + '.jpg'\n    data = annotations[annotations['StudyInstanceUID'] == image_uid]\n    if(singleLabel):\n        data = data[data['label'] == title]\n    data = data['data']\n    if(len(data) == 0):\n        print(title)\n        return\n    plt.figure(figsize=(10,6))\n\n    plt.subplot(1, 2,1)\n    img = plt.imread(image_path)\n    \n    print(f\"Image dimensions:  {img.shape[0],img.shape[1]}\")\n    print(f\"Maximum pixel value : {img.max():.1f} ; Minimum pixel value:{img.min():.1f}\")\n    print(f\"Mean value of the pixels : {img.mean():.1f} ; Standard deviation : {img.std():.1f}\")\n    \n    plt.imshow(img, cmap='gray')\n    plt.title('Actual Image')\n    plt.axis('off')\n\n    plt.subplot(1,2,2)\n    plt.imshow(img, cmap='gray')\n    plt.axis('off')\n    \n    for i in range(len(data)):\n        curr_data = data.values[i]\n        x_loc = [x[0] for x in curr_data]\n        y_loc = [x[1] for x in curr_data]\n        \n        plt.plot(x_loc, y_loc, linewidth = 5.0)\n    plt.tight_layout()\n    \n    plt.title('Annotated Image')\n    \n    plt.suptitle(title)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"imageIds = {}\nfor cat in categories:\n    imageIds[cat] = df[df[cat] == 1]['StudyInstanceUID'].to_list()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - ETT - Abnormal</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[0]][0], categories[0])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - ETT - Borderline</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[1]][2], categories[1])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - ETT - Normal</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[2]][-1], categories[2])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - NGT - Abnormal</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[3]][-1], categories[3])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - NGT - Borderline</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[4]][-1], categories[4])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - NGT - Incompletely Imaged</h1>"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[5]][-1], categories[5])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - NGT - Normal</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[6]][7], categories[6])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - CVC - Abnormal</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[7]][-1], categories[7])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - CVC - Borderline</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[8]][0], categories[8])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - CVC - Normal</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[9]][-1], categories[9])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<h1 style=\"border:2px solid Purple;text-align:center\">Image Annotations - Swan Ganz Catheter Present</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations(imageIds[categories[10]][5], categories[10])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id=\"image\"></a>\n<h1 style=\"border:2px solid Purple;text-align:center\">Image Info</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"image_width = []\nimage_height = []\nimage_channel_mean = []\nfor image_uid in tqdm(df['StudyInstanceUID'].to_list()):\n    image_path = IMAGE_DIR_TRAIN + image_uid + '.jpg'\n    img = cv2.imread(image_path,cv2.IMREAD_GRAYSCALE)\n    image_height.append(img.shape[0])\n    image_width.append(img.shape[1])\n    image_channel_mean.append(np.mean(img))\ndf['Image Height'] = image_height\ndf['Image Width'] = image_width\ndf['Total Pixels'] = df['Image Height'] * df['Image Width']","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"test_image_width = []\ntest_image_height = []\ntest_channel_mean = []\nfor image_name in tqdm(os.listdir(IMAGE_DIR_TEST)):\n    image_path = IMAGE_DIR_TEST + image_name\n    img = cv2.imread(image_path,cv2.IMREAD_GRAYSCALE)\n    test_image_height.append(img.shape[0])\n    test_image_width.append(img.shape[1])\n    test_channel_mean.append(np.mean(img))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"trace0 = go.Violin(x = image_height, name = 'Train Height')\ntrace1 = go.Violin(x = image_width, name = 'Train Width')\n\ntrace2 = go.Violin(x = test_image_width, name = 'Test Width')\ntrace3 = go.Violin(x = test_image_height, name = 'Test Height')\n\nfig = go.Figure([trace0,trace3, trace1, trace2])\nfig.update_layout(title = 'Image Height and Widths', template = 'ggplot2')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig = ff.create_distplot([image_channel_mean, test_channel_mean], group_labels=['Training Images', 'Testing Images'], colors = ['#cecece','#808080'])\nfig.update_layout(template = 'plotly_white', title_text = 'Channel Distribution')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<a id=\"overlap\"></a>\n<h1 style=\"border:2px solid Purple;text-align:center\">Category Overlap</h1>"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(6,6))\nfig = venn3([set(df[df[col] == 1]['StudyInstanceUID']) for col in ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal']], set_labels = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal'])\nplt.title('ETT Image Overlap')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(6,6))\nfig = venn3([set(df[df[col] == 1]['StudyInstanceUID']) for col in ['NGT - Abnormal', 'NGT - Borderline','NGT - Normal']], set_labels = ['NGT - Abnormal', 'NGT - Borderline','NGT - Normal'])\nplt.title('NGT Image Overlap')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(6,6))\nfig = venn3([set(df[df[col] == 1]['StudyInstanceUID']) for col in ['CVC - Abnormal', 'CVC - Borderline','CVC - Normal']], set_labels = ['CVC - Abnormal', 'CVC - Borderline','CVC - Normal'])\nplt.title('CVC Image Overlap')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### It must be noted that that the fact that we have an overlap in NGT and CVC categories does not mean that the images are mislabeled. \n\n> The labels reflect whether or not a particular finding is present or absent in the image, and are not mutually exclusive as there may be more than one line or tube of that type in the patient. For instance, it is not uncommon to have 2 or more CVCs in the same patient, where one could be in a Normal position but another could be in a Borderline or Abnormal position. This will explain the majority of the cases you see where there are more than one label for NGT or CVC. There may be a small minority of incorrectly labelled cases due to human error but this should be a small percentage.\n\nRead the following [discussion](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203372) to know more"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df[(df['CVC - Abnormal'] == 1) & (df['CVC - Normal'] == 1) & (df['CVC - Borderline'] == 1)]","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plot_image_and_annotations('1.2.826.0.1.3680043.8.498.74579763900891627023142390145044977755', 'Image with CVC Normal, Abnormal and Borderline Annotations',singleLabel=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Work in Progress"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}