{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Edit: [Discussion](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203372) \n### This discussion concludes that the examples aren't contradictory, there are more than 1 tubes present instead of just 1.\n\n### Sorry if it caused any unnecessary confusion"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport plotly.graph_objects as go\nimport plotly.figure_factory as ff\nimport plotly.express as px","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(\"../input/ranzcr-clip-catheter-line-classification/train.csv\")\nsample_df = pd.read_csv(\"../input/ranzcr-clip-catheter-line-classification/sample_submission.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"Total Number of Rows in the Training Dataset: {train_df.shape[0]}\")\nprint(f\"Total Number of Columns in the Training Dataset: {train_df.shape[1]}\")\nprint()\nprint(f\"Total Number of Rows in the Testing Dataset: {sample_df.shape[0]}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **Let's Check if we have some samples with contradicting features**"},{"metadata":{"trusted":true},"cell_type":"code","source":"ETT_cols = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal']\nNGT_cols = ['NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal']\nCVC_cols = ['CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[ETT_cols].sum(axis=1).value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[NGT_cols].sum(axis=1).value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> ### **The samples with sum as 2 have '1' for more than one column which doesn't seem correct, Let's explore a bit more**"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[train_df[NGT_cols].sum(axis=1) == 2]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> ### **We can see that some samples have been classified into both NGT-Normal and NGT-Abnormal which definitely can't be possible**"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[CVC_cols].sum(axis=1).value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> ### Again we have samples with '1' for more than one column"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[train_df[CVC_cols].sum(axis=1) == 2]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> ### **There are many samples where we have more than 1 correct class which again looks like a case of misclassification**"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[train_df[CVC_cols].sum(axis=1) == 3]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> ### **These samples definitely have wrong labels as they are classified in all the 3 categories i.e. CVC-Abnormal, CVC-Normal and CVC-Borderline**"},{"metadata":{},"cell_type":"markdown","source":"# More Visualizations and Analysis to come soon! If you have any explaination about these discrepancies please do let me know"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}