{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction\n\nThis Kernel objective is to explore the dataset for SSIM-ISIC Melanoma Classification Challenge.  \n\nThe images are provided in DICOM format. These images are in the original format from the radiological scanner output, with high resolution, lossless compression.\n\nImages are also provided in JPEG and TFRecord format (in the jpeg and tfrecords directories, respectively). \n\nMetadata is also provided outside of the DICOM format, in CSV files. \n\nWe will explore both the metadata in DICOM format and in the CSV files.","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# Prepare for data exploration\n\n## Load packages","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import pandas as pd \nimport numpy as np\nimport os\nimport matplotlib\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm_notebook\nfrom matplotlib.patches import Rectangle\nimport pydicom as dcm\nimport seaborn as sns\n%matplotlib inline \nPATH = \"/kaggle/input/siim-isic-melanoma-classification/\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Load the Data\n\nLet's load the tabular data. There are three files:\n\n* Sample submission;  \n* Train;  \n* Test.  ","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"sample_submission_df = pd.read_csv(os.path.join(PATH,'sample_submission.csv'))\ntrain_df = pd.read_csv(os.path.join(PATH,'train.csv'))\ntest_df = pd.read_csv(os.path.join(PATH,'test.csv'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(f\"sample submission shape: {sample_submission_df.shape}\")\nprint(f\"train shape: {train_df.shape}\")\nprint(f\"test shape: {test_df.shape}\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"sample_submission_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"test_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Test and sample submission csv files have 10982 rows (samples).  \n\nIn train data there are 3 columns not present either in test or in sample submission csv. These columns are:\n\n* diagnosis;  \n* bening_malignant;  \n* target.  \n\nThe objective is to predict the probability (target) that the sample is malignant. \n\nLet's check now the train data in train and test folders.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_image_list = os.listdir(os.path.join(PATH, 'train'))\ntest_image_list = os.listdir(os.path.join(PATH, 'test'))\n\nprint(f\"train image_name list: {train_df.image_name.nunique()}\")\nprint(f\"train image list: {len(train_image_list)}\")\nprint(f\"test image_name list: {test_df.image_name.nunique()}\")\nprint(f\"test image list: {len(test_image_list)}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Missing data\n\n\nLet's check for missing data in the images list.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"trimmed_train_image_list = []\nfor img in train_image_list:\n    trimmed_train_image_list.append(img.split('.dcm')[0])\n    \ntrimmed_test_image_list = []\nfor img in test_image_list:\n    trimmed_test_image_list.append(img.split('.dcm')[0])  ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"intersect_train_train = (set(train_df.image_name.unique()) & set(trimmed_train_image_list))\nintersect_test_test = (set(test_df.image_name.unique()) & set(trimmed_test_image_list))\n\nprint(f\"image train (dcm) & train csv: {len(intersect_train_train)}\")\nprint(f\"image test (dcm) & test csv: {len(intersect_test_test)}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data exploration\n\n\nLet's explore the distribution of the data.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"We start with the distribution of patient id's.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"tr_patient_id = train_df.patient_id.nunique()\nte_patient_id = test_df.patient_id.nunique()\nlist_tr_patient_id = train_df.patient_id.unique()\nlist_te_patient_id = test_df.patient_id.unique()\nintersection = set(list_tr_patient_id) & set(list_te_patient_id)\nprint(f\"Unique patients in train: {tr_patient_id} and test: {te_patient_id}\")\nprint(f\"Patients in common in train and test: {len(intersection)}\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's group the images per patient to see what is the distribution of images per patient.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"tmp = train_df.groupby(['patient_id', 'target'])['image_name'].count()\ntr_df = pd.DataFrame(tmp).reset_index(); tr_df.columns = ['patient_id', 'target', 'images']\n\ntmp = test_df.groupby(['patient_id'])['image_name'].count()\nte_df = pd.DataFrame(tmp).reset_index(); te_df.columns = ['patient_id', 'images']\n\n\ntr_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"te_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure()\nfig, ax = plt.subplots(1,3,figsize=(16,6))\ng_tr0 = sns.distplot(tr_df.loc[tr_df.target==0, 'images'],kde=False,bins=50, color=\"green\",label='target = 0', ax=ax[0])\ng_tr1 = sns.distplot(tr_df.loc[tr_df.target==1, 'images'],kde=False,bins=50, color=\"red\",label='target = 1', ax=ax[1])\ng_te = sns.distplot(te_df['images'],kde=False,bins=50, color=\"blue\", label='columns', ax=ax[2])\ng_tr0.set_title('Number of images / patient - Train set\\n target = 0 (benign)')\ng_tr1.set_title('Number of images / patient - Train set\\n target = 1 (malignant)')\ng_te.set_title('Number of images / patient - Test set')\n\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can observe that the numberof images / patient in the case of malignant cases have a more compact distribution, maximum number of images / patient being 8, while in the case of patients with the result of exam benign the total number of images / patient might be up to 120. Of course, also in the case of patients with benign diagnosis the majority of cases will have just few images. In the same time, we can most probably conclude that the number of images / patient might be a good predictor for the value of target if the number of images is large. \n","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def plot_count(df, feature, title='', size=2, rotate_axis = False):\n    f, ax = plt.subplots(1,1, figsize=(3*size,2*size))\n    total = float(len(df))\n    sns.countplot(df[feature],order = df[feature].value_counts().index, palette='Set3')\n    plt.title(title)\n    if(rotate_axis):\n        ax.set_xticklabels(ax.get_xticklabels(), rotation=90)\n    for p in ax.patches:\n        height = p.get_height()\n        ax.text(p.get_x()+p.get_width()/2.,\n                height + 3,\n                '{:1.2f}%'.format(100*height/total),\n                ha=\"center\") \n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_count(train_df, 'sex', 'Patient sex (train) - data count and percent')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_count(test_df, 'sex', 'Patient sex (test) - data count and percent')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_count(test_df, 'age_approx', 'Patient approximate age (train) - data count and percent', size=4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_count(test_df, 'age_approx', 'Patient approximate age (test) - data count and percent', size=4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_count(train_df, 'anatom_site_general_challenge', 'Location of imaged site/anatomy part (train) - data count and percent', size=3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_count(test_df, 'anatom_site_general_challenge', 'Location of imaged site/anatomy part (test) - data count and percent', size=3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_count(train_df, 'diagnosis', 'Detailed diagnosis (train) - data count and percent', size=4, rotate_axis=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_count(train_df, 'benign_malignant', 'Indicator of malignancy of imaged lesion  (train) - data count and percent')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_count(train_df, 'target', 'Target value - 0: bening, 1: malignant (train)\\n data count and percent')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1,figsize=(16,6)) \ntmp = train_df.groupby('diagnosis')['target'].value_counts() \ndf = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index() \nsns.barplot(ax=ax,x = 'diagnosis', y='Exams',hue='target',data=df, palette='Set1') \nplt.title(\"Number of examinations grouped on Diagnosis and Target\") \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1,figsize=(16,6)) \ntmp = train_df.groupby('diagnosis')['benign_malignant'].value_counts() \ndf = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index() \nsns.barplot(ax=ax,x = 'diagnosis', y='Exams',hue='benign_malignant',data=df, palette='Set1') \nplt.title(\"Number of examinations grouped on Diagnosis and Benign/Malignant\") \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1,figsize=(16,6)) \ntmp = train_df.groupby('diagnosis')['sex'].value_counts() \ndf = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index() \nsns.barplot(ax=ax,x = 'diagnosis', y='Exams',hue='sex',data=df, palette='Set1') \nplt.title(\"Number of examinations grouped on Diagnosis and Sex\") \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1,figsize=(8,6)) \ntmp = train_df.groupby('sex')['target'].value_counts() \ndf = pd.DataFrame(data={'Exams': tmp.values}, index=tmp.index).reset_index() \nsns.barplot(ax=ax,x = 'sex', y='Exams',hue='target',data=df, palette='Set1') \nplt.title(\"Number of examinations grouped on Sex and Target\") \nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def plot_distribution_grouped(feature, feature_group, hist_flag=True):\n    fig, ax = plt.subplots(nrows=1,figsize=(12,6)) \n    for f in train_df[feature_group].unique():\n        df = train_df.loc[train_df[feature_group] == f]\n        sns.distplot(df[feature], hist=hist_flag, label=f)\n    plt.title(f'Data/image {feature} distribution, grouped by {feature_group}')\n    plt.legend()\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plot_distribution_grouped('age_approx', 'sex')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Images samples\n\n\nLet's visualize a part of the images, first from the training set.\n\nWe select only images corresponding to `target` = 1 (malignant).","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def show_dicom_images(data):\n    img_data = list(data.T.to_dict().values())\n    f, ax = plt.subplots(3,3, figsize=(16,18))\n    for i,data_row in enumerate(img_data):\n        patientImage = data_row['image_name']+'.dcm'\n        imagePath = os.path.join(PATH,\"train/\",patientImage)\n        data_row_img_data = dcm.read_file(imagePath)\n        modality = data_row_img_data.Modality\n        age = data_row_img_data.PatientAge\n        sex = data_row_img_data.PatientSex\n        data_row_img = dcm.dcmread(imagePath)\n        ax[i//3, i%3].imshow(data_row_img.pixel_array, cmap=plt.cm.gray) \n        ax[i//3, i%3].axis('off')\n        ax[i//3, i%3].set_title(f\"ID: {data_row['image_name']}\\nModality: {modality} Age: {age} Sex: {sex}\\nDiagnosis: {data_row['diagnosis']}\")\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"show_dicom_images(train_df[train_df['target']==1].sample(9))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's look now to few Dicom images with `target` = 0 (benign)","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"show_dicom_images(train_df[train_df['target']==0].sample(9))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let's select few images with diagnosis `melanoma`.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"show_dicom_images(train_df[train_df['diagnosis']=='melanoma'].sample(9))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And now let's look to few different anatomic parts. Next, `head/neck`.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"show_dicom_images(train_df[train_df['anatom_site_general_challenge']=='head/neck'].sample(9))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's look to anatomic part `torso`.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"show_dicom_images(train_df[train_df['anatom_site_general_challenge']=='torso'].sample(9))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Explore DICOM files\n\n\nLet's look to more details from the DICOM files.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_item = list(train_df[:3].T.to_dict().values())[0]\nsample_image_name = sample_item['image_name']\nsample_patient_sex = sample_item['sex']\nsample_patient_age = sample_item['age_approx']\nbody_part_examined = sample_item['anatom_site_general_challenge']\nimage_ID = sample_image_name +'.dcm'\ndicom_file_path = os.path.join(PATH,\"train/\",image_ID)\ndicom_file_dataset = dcm.read_file(dicom_file_path)\nprint(sample_image_name, sample_patient_sex, sample_patient_age,body_part_examined)\ndicom_file_dataset","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We observe that the information in the .csv file are aligned with the information in the corresponding DICOM file.\n\nAlso, we can observe that the information in the DICOM file was correctly anonymized for the challenge.  \n\nEven the `(0008, 1030) Study Description` is set to: `LO: 'ISIC 2020 Grand Challenge image'`\n\nVisualy inspecting the potential additional useful data in the DICOM file, we can spot (at the first glance) such informations like:  rows, columns, modality, body part examined.\n\n\nLet's create a function to extract key attributes from DICOM files.","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"def extract_DICOM_attributes(folder):\n    images = list(os.listdir(os.path.join(PATH, folder)))\n    df = pd.DataFrame()\n    for image in images:\n        image_name = image.split(\".\")[0]\n        dicom_file_path = os.path.join(PATH,folder,image)\n        dicom_file_dataset = dcm.read_file(dicom_file_path)\n        study_date = dicom_file_dataset.StudyDate\n        modality = dicom_file_dataset.Modality\n        age = dicom_file_dataset.PatientAge\n        sex = dicom_file_dataset.PatientSex\n        body_part_examined = dicom_file_dataset.BodyPartExamined\n        patient_orientation = dicom_file_dataset.PatientOrientation\n        photometric_interpretation = dicom_file_dataset.PhotometricInterpretation\n        rows = dicom_file_dataset.Rows\n        columns = dicom_file_dataset.Columns\n             \n        df = df.append(pd.DataFrame({'image_name': image_name, \n                        'dcm_modality': modality,'dcm_study_date':study_date, 'dcm_age': age, 'dcm_sex': sex,\n                        'dcm_body_part_examined': body_part_examined,'dcm_patient_orientation': patient_orientation,\n                        'dcm_photometric_interpretation': photometric_interpretation,\n                        'dcm_rows': rows, 'dcm_columns': columns}, index=[0]))\n    return df","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"tr_df = extract_DICOM_attributes('train')\ntrain_dicom_df = train_df.merge(tr_df, on='image_name')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"train_dicom_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"te_df = extract_DICOM_attributes('test')\ntest_dicom_df = test_df.merge(te_df, on='image_name')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_dicom_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Check rows and columns distribution","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure()\nfig, ax = plt.subplots(1,2,figsize=(12,6))\ng0 = sns.distplot(train_dicom_df['dcm_rows'],kde=False,bins=50, color=\"green\",label='rows', ax=ax[0])\ng1 = sns.distplot(train_dicom_df['dcm_columns'],kde=False,bins=50, color=\"magenta\", label='columns', ax=ax[1])\ng0.set_title('Train set - rows distribution')\ng1.set_title('Train set - columns distribution')\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure()\nfig, ax = plt.subplots(1,2,figsize=(12,6))\ng0 = sns.distplot(test_dicom_df['dcm_rows'],kde=False,bins=50, color=\"green\",label='rows', ax=ax[0])\ng1 = sns.distplot(test_dicom_df['dcm_columns'],kde=False,bins=50, color=\"magenta\", label='columns', ax=ax[1])\ng0.set_title('Test set - rows distribution')\ng1.set_title('Test set - columns distribution')\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}