{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport cv2\nimport os\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\nimport json\nfrom sklearn import preprocessing\n\nIMAGE_DIR_TRAIN = '/kaggle/input/ranzcr-clip-catheter-line-classification/train/'\nIMAGE_DIR_TEST = '/kaggle/input/ranzcr-clip-catheter-line-classification/test/'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## About this Notebook\n\nThere are good EDA Notebooks published to understand data of competition. Here there are any nice examples:\n\n* [Detailed resource notebook & EDA for beginners](https://www.kaggle.com/bipinkrishnan/detailed-resource-notebook-eda-for-beginners) by [Bipin Krishnan P](https://www.kaggle.com/bipinkrishnan). A good intro to understand basic concepts about the competition.\n* [RANZCR - Exploratory Data Analysis](https://www.kaggle.com/ihelon/ranzcr-exploratory-data-analysis) by [Yaroslav Isaienkov](https://www.kaggle.com/ihelon). Good annotations examples.\n* [RANZCR-CLiP : One Stop For All EDA Needs](https://www.kaggle.com/foolofatook/ranzcr-clip-one-stop-for-all-eda-needs) by [Aayush Jain](https://www.kaggle.com/foolofatook). Nice EDA. See category overlap section.\n\nHowever, we think that there are a couple of data that may be of interest to the competition."},{"metadata":{},"cell_type":"markdown","source":"## EDA"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"images = []\n\nfor filename in tqdm(os.listdir(IMAGE_DIR_TRAIN)):\n    file = IMAGE_DIR_TRAIN + filename\n    img = cv2.imread(file)\n    UID = os.path.splitext(filename)[0]\n    images.append([UID, 'steelblue', filename, img.shape[0], img.shape[1], img.shape[2]])\n\nfor filename in tqdm(os.listdir(IMAGE_DIR_TEST)):\n    file = IMAGE_DIR_TEST + filename\n    img = cv2.imread(file)\n    UID = os.path.splitext(filename)[0]\n    images.append([UID, 'blue', filename, img.shape[0], img.shape[1], img.shape[2]])\n\ndf_files = pd.DataFrame(images, columns=['UID','set','filename','h', 'w', 'c'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"le = preprocessing.LabelEncoder()\nle.fit([\n    'CVC - Abnormal',\n    'CVC - Borderline',\n    'CVC - Normal',\n    'ETT - Abnormal',\n    'ETT - Borderline',\n    'ETT - Normal',\n    'NGT - Abnormal',\n    'NGT - Borderline',\n    'NGT - Incompletely Imaged',\n    'NGT - Normal',\n    'Swan Ganz Catheter Present'\n])\n\nce = preprocessing.LabelEncoder()\nce.fit([\n    '.01',\n    '.02',\n    '.03',\n    '.101',\n    '.102',\n    '.103',\n    '.501',\n    '.502',\n    '.503',\n    '.504',\n    '.911'\n])\n\ndf_train = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\ndf_train.columns = ['UID','ETTA','ETTB','ETTN','NGTA','NGTB','NGTI','NGTN','CVCA','CVCB','CVCN','SGCP','PID']\ndf_annotations = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train_annotations.csv')\ndf_annotations.columns = ['UID','label','data']\ndf_annotations = df_annotations.join(df_files.set_index('UID'),how='inner',on='UID')\n\ndf_annotations['c'] = ce.inverse_transform(le.transform(df_annotations.label))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"df = pd.DataFrame()\nfor i, ann in tqdm(df_annotations.iterrows(), total=len(df_annotations)):\n    row_df = pd.DataFrame(json.loads(ann.data),columns=['x1', 'y1'])\n    row_df['c'] = ann.c\n    row_df['label'] = ann.label\n    row_df['x'] = row_df.x1/ann.w\n    row_df['y'] = row_df.y1/ann.h\n    df = df.append(row_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Image Sizes\n"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df_files[['h','w','set','filename']].groupby(by=['h','w','set'], as_index=False).count().plot(figsize=(20, 10), kind='scatter', x='w', y='h', s='filename', c='set', alpha=0.2)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Annotations\n\nWe want to show the areas in which the annotations are placed on average on an ideal x-ray. Below is an x-ray with some interesting anatomical elements."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 16))\nimg = plt.imread(\"../input/ranzcr-jpgs/Mediastinal_structures_on_chest_X-ray_annotated-1536x1245.jpg\")\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"(image credits https://www.radiologia2cero.com/dispositivos-en-la-rx-de-torax/)\n\nNow, try to make an imaginative effort and suppose the following images over the previous one."},{"metadata":{},"cell_type":"markdown","source":"#### All annotations"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df.plot(figsize=(16, 16), kind='scatter', x='x', y='y', c='c', s=100, alpha=.1)\nplt.xlim([0, 1])\nplt.ylim([0, 1])\nplt.gca().invert_yaxis()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### CVC annotations"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df.query(\"label in('CVC - Abnormal','CVC - Borderline', 'CVC - Normal')\").plot(figsize=(16, 16), kind='scatter', x='x', y='y', s=100, alpha=.1)\nplt.xlim([0, 1])\nplt.ylim([0, 1])\nplt.gca().invert_yaxis()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### ETT annotations"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df.query(\"label in('ETT - Abnormal','ETT - Borderline', 'ETT - Normal')\").plot(figsize=(16, 16), kind='scatter', x='x', y='y', s=100, alpha=.1)\nplt.xlim([0, 1])\nplt.ylim([0, 1])\nplt.gca().invert_yaxis()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### NGT annotations"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df.query(\"label in('NGT - Abnormal','NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal')\").plot(figsize=(16, 16), kind='scatter', x='x', y='y', s=100, alpha=.1)\nplt.xlim([0, 1])\nplt.ylim([0, 1])\nplt.gca().invert_yaxis()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Swan Ganz catheter annotations"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df.query(\"label in('Swan Ganz Catheter Present')\").plot(figsize=(16, 16), kind='scatter', x='x', y='y', s=100, alpha=.1)\nplt.xlim([0, 1])\nplt.ylim([0, 1])\nplt.gca().invert_yaxis()\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}