{"cells":[{"metadata":{"trusted":true,"_uuid":"ae448b0ed29194053d40ebd29b2fa03982468552"},"cell_type":"code","source":"%matplotlib inline\nimport matplotlib.pyplot as plt\nimport pylab\nimport numpy as np\nimport pydicom\nimport pandas as pd\nfrom glob import glob\nimport os\n\ndatapath = '../input/'","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cecf54a047d75cd2de16cbc793198ff80641e3ab"},"cell_type":"markdown","source":"This notebook is a short version of my more detailed data exploration notebook here: [beginner-intro-to-lung-opacity-s1](https://www.kaggle.com/giuliasavorgnan/start-here-beginner-intro-to-lung-opacity-s1). I believe I identified a data leakage and wanted to share it asap. "},{"metadata":{"_uuid":"1e9f7c1d4802209d859267650bbe081010c85055"},"cell_type":"markdown","source":"# Data Leakage explanation\n\nThe \"view position\" of the radiograph image appears to be really important in terms of target split. This information is extracted from the metadata header of the radiograph images. \n\nFor ViewPosition=='AP' the Target [0,1] split is 0.62 - 0.37\nFor ViewPosition=='PA' the Target [0,1] split is 0.91 - 0.09\n\nAP means Anterior-Posterior, whereas PA means Posterior-Anterior. This [webpage](https://www.med-ed.virginia.edu/courses/rad/cxr/technique3chest.html) explains that \"Whenever possible the patient should be imaged in an upright PA position.  AP views are less useful and should be reserved for very ill patients who cannot stand erect\". One way to interpret this target unbalance is that patients that are imaged in an AP position are those that are more ill, and therefore more likely to have contracted pneumonia. Note that the absolute split between AP and PA images is about 50-50, so the above consideration is extremely significant. \n"},{"metadata":{"trusted":true,"_uuid":"fa5346f0e010257050379687a347a069a614f53a"},"cell_type":"code","source":"df_box = pd.read_csv(datapath+'stage_1_train_labels.csv')\nprint('Number of rows (unique boxes per patient) in main train dataset:', df_box.shape[0])\nprint('Number of unique patient IDs:', df_box['patientId'].nunique())\ndf_box.head(6)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f79065ff5ede4937b555d0dd3d3f9da486c83ab4"},"cell_type":"code","source":"df_aux = pd.read_csv(datapath+'stage_1_detailed_class_info.csv')\nprint('Number of rows in auxiliary dataset:', df_aux.shape[0])\nprint('Number of unique patient IDs:', df_aux['patientId'].nunique())\ndf_aux.head(6)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ca18783acf240faea13a23fd06fba7b41f9ea71d"},"cell_type":"code","source":"assert df_box['patientId'].values.tolist() == df_aux['patientId'].values.tolist(), 'PatientId columns are different.'\ndf_train = pd.concat([df_box, df_aux.drop(labels=['patientId'], axis=1)], axis=1)\ndf_train.head(6)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ae09da50fdcd33b96256cefb190171a79186adbc"},"cell_type":"code","source":"def get_dcm_data_per_patient(pId):\n    '''\n    Given one patient ID, \n    return the corresponding dicom data.\n    '''\n    return pydicom.read_file(datapath+'stage_1_train_images/'+pId+'.dcm')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8b8bb3a9d3f0fbea4f7c53adfab009b9c13fd25f"},"cell_type":"code","source":"def get_metadata_per_patient(pId, attribute):\n    '''\n    Given a patient ID, return useful meta-data from the corresponding dicom image header.\n    Return: \n    attribute value\n    '''\n    # get dicom image\n    dcmdata = get_dcm_data_per_patient(pId)\n    # extract attribute values\n    attribute_value = getattr(dcmdata, attribute)\n    return attribute_value","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a6c89433fb606b2a5303fe13ca58d793b181fae9"},"cell_type":"code","source":"# create list of attributes that we want to extract (manually edited after checking which attributes contained valuable information)\nattributes = ['PatientSex', 'PatientAge', 'ViewPosition']\nfor a in attributes:\n    df_train[a] = df_train['patientId'].apply(lambda x: get_metadata_per_patient(x, a))\n# convert patient age from string to numeric\ndf_train['PatientAge'] = df_train['PatientAge'].apply(pd.to_numeric, errors='coerce')\n# remove a few outliers\ndf_train['PatientAge'] = df_train['PatientAge'].apply(lambda x: x if x<120 else np.nan)\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9e76c9353cd060a87039fc0cc1c1c9f65ba9ac10"},"cell_type":"code","source":"# look at age statistics between positive and negative target groups\ndf_train.drop_duplicates('patientId').groupby('Target')['PatientAge'].describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3e34324b7a9ee1abf9e0b36ffc51f28f044def7c"},"cell_type":"code","source":"# look at gender statistics between positive and negative target groups\ndf_train.drop_duplicates('patientId').groupby(['PatientSex', 'Target']).size() / df_train.drop_duplicates('patientId').groupby(['PatientSex']).size()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"15bc5c4d8ce6a7eee5656508fe268e8af64af49c"},"cell_type":"code","source":"# look at patient position statistics between positive and negative target groups\ndf_train.drop_duplicates('patientId').groupby(['ViewPosition', 'Target']).size() / df_train.drop_duplicates('patientId').groupby(['ViewPosition']).size()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a0b303fec9db6ab5eac76fa8bf3feecac59340a2"},"cell_type":"code","source":"# absolute split of view position\ndf_train.groupby('ViewPosition').size()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}