{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Summary","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**[I. Generalities](#generalities)** -  Overview of the available data: format, number, values, description etc.\n\n**[II. Tabular data](#tabular_data)** - Exploration of the tabular data/metadata and their potential predictive power\n\n[II. a) Values, distributions & balancing](#vdb) - Inspection of each variable separately\n\n[II. b) Interactions between tabular data](#interactions_data) - Interactions between each variables\n   \n[II. c) Interaction at the patient level](#interactions_patient) - Information shared between different images of the same patient\n\n\n\n**[III. Image data](#images)** - A convenient tool to interactively show images by label\n\n**[IV. Wrap-up](#wrapup)** - Final comments and sum-up of the most interesting findings\n","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import os\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt \nplt.rcParams['figure.figsize']=(20,10)\nimport seaborn as sns\nsns.set_style(\"dark\")\nimport warnings\nwarnings.filterwarnings('ignore')\nfrom ipywidgets import interact","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div id=\"generalities\">\n    \n#  I. Generalities \n    \n</div>\n","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df = pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Distributions of the content of the images (diagnosis):","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df['diagnosis'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In this competition, we have both images and contextual data. The goal is to predict from them the existence of a malignant melanome.\n\n### Images are available in three formats:\n* **JPEG**\n* **DICOM**, a widely used format among the health data science practitionner\n* **TFRecord**, in this case data are resized to 1024 x 1024.\n    \n    \n### Contextual data are tabular data and include:\n* Information about the patient (the sex and the approximative age) \n* Information about the image (what part of the bodies is it taken from)\n* Detailed diagnostic (ie other classes, beside the binary melanoma/no melanoma dichotomy that is the frame of this competition). \n\nLet's stop a bit to know what are the definitions corresponding to those detailed classes. More information and images are provided in the links: \n* **[Nevus](https://dermnetnz.org/topics/mole/)**: a birthmark or a mole on the skin, especially a birthmark in the form of a raised red patch.\n* **[Melanoma](https://dermnetnz.org/topics/melanoma/)**: a tumour of melanin-forming cells, especially a malignant tumour associated with skin cancer.\n* **[Seborrheic keratosis](https://dermnetnz.org/topics/seborrhoeic-keratosis/)**: a non-cancerous (benign) skin tumour that originates from cells in the outer layer of the skin. Like liver spots, seborrheic keratoses are seen more often as people age\n* **[Lentigo NOS (Lentigo?)](https://dermnetnz.org/topics/lentigo/)**: A small pigmented spot on the skin with a clearly defined edge, surrounded by normal-appearing skin. It is a harmless (benign) hyperplasia of melanocytes which is linear in its spread.\n* **[Lichenoid keratosis](https://dermnetnz.org/topics/lichenoid-keratosis/)**:A usually small, solitary, inflamed macule or thin pigmented plaque. Multiple eruptive lichenoid keratoses in sun-exposed sites are also described. Their colour varies from an initial reddish brown to a greyish purple/brown as the lesion resolves several weeks or months later.\n* **[Solar lentigo](https://dermnetnz.org/topics/solar-lentigo/)**: Solar lentigo is a harmless patch of darkened skin. They are very common, especially in people over the age of 40 years.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"***","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<div id=\"tabular_data\">\n    \n# II. Tabular data\n    \n</div>\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<div id=\"vdb\">\n    \n## a) Values, distributions & balancing\n    \n</div>","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"categorical_cols = ['diagnosis','sex','anatom_site_general_challenge','benign_malignant','target']\n\nfig,ax = plt.subplots(2,(len(categorical_cols)+1)//2,figsize=(30,15))\n\nratio = {}\nfor i,col in enumerate(categorical_cols+['age_approx']):\n    ratio[col] = 100*df[col].value_counts(dropna=False)/df[col].value_counts(dropna=False).sum()\n    if i==5:\n        ax[1][2].hist(df['age_approx'])     \n    else:\n        ax[i%2][i//2].bar([str(x) for x in ratio[col].index],height=ratio[col])\n    ax[i%2][i//2].set_title(col,fontdict={'fontsize':20})\n    for tick in ax[i%2][i//2].get_xticklabels():\n        tick.set_rotation(90)\n        tick.set_fontsize('x-large')\n\nfig.suptitle('Distribution of each categorical data',y=1.02,fontsize=25)\nfig.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Takeaways:\n- Most of the categorical data are (strongly) imbalanced.\n- This is in particular the case of the target variable (near 98% of not malignant lesion). This is likely to be one of the difficulties of this competition. \n- If we look into the details of the target, in the \"diagnosis\" columns, we see that in most cases (80%), the status is \"unknown\", a probably very heterogeneous category. \n- Patients' sex is also slightly imbalanced, with a few more men. This data is almost always knwon. \n- Interestingly, the body part from which the picture is taken is also imbalanced, with more than 50% corresponding to the torso. It will be interesting to see how it can impact the result.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<div id=\"interactions_data\">\n    \n## b) Interactions between tabular data\n    \n</div>","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df_category = df.copy()\ndf_category['benign_malignant'] = (df_category['benign_malignant'] == 'malignant').apply(int)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Relation between target, begnin_malignant & diagnosis:","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(\"Agreement between 'malignant' value and target:\",100*((df['benign_malignant']=='malignant') == (df['target'])).sum()/len(df),'%')\nprint(\"Agreement between 'melanoma' value and target:\",100*(((df['diagnosis'] =='melanoma') ) == (df['target'])).sum()/len(df),'%')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Remark**: We see that the *benign_malignant* & *target* is actually the same information. *diagnosis* is a bit more complex, as it details all the other cases of not-malignant lesions. We thus do not did to keep benign_malignant for the next analysis.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Relation between biological data:","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"sex_c = df_category['sex'].value_counts()\nsex_c = sex_c.sort_index()\nfig,ax = plt.subplots() \nfor idx in sex_c.index:\n    sns.kdeplot(df.loc[df['sex']==idx,'age_approx'],shade=True,ax=ax)\nax.legend(sex_c.index)\nax.set_title('Age distribution for each sex',fontsize=20)\nres = ax.set_xlabel('approx_age')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"diag_c = df_category['diagnosis'].value_counts()\ndiag_c = diag_c.sort_index()\nfig,ax = plt.subplots() \nfor idx in diag_c.index:\n    sns.kdeplot(df.loc[df['diagnosis']==idx,'age_approx'],shade=True,ax=ax)\nax.legend(diag_c.index)\nax.set_title('Relation between age & diagnosis',fontsize=20)\nres = ax.set_xlabel('approx_age')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"ax = sns.catplot(x=\"sex\",\n            y=\"target\",\n            kind=\"bar\",\n            hue='age_approx',\n            data=df_category,\n            height=9, \n            aspect=2.3);\n_=ax.fig.suptitle('Influence of the age on likeliness that the diagnosis is melanoma',y=1.02, fontsize=20)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(\"Number of men > 90 in the dataset:\", len(df_category.loc[(df_category['age_approx']==90.0) & (df_category['sex']=='male'),'patient_id'].unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Takeaways: \n\n- There is not a huge difference between male and female age distributions. We should however note that since men are a bit more aged in this dataset, some slight \"sex effects\" could actually be explained by an \"age effect\". \n- As expected, diagnostic distribution differ with regard to the age: besides nevus and \"unknown\" lesions, all the lesion are more likely to appear later (40 years & more). \n- If we look more specifically at the target, we see that melanoma exist at any age, but it definitely becomes more likely with the age (> 40 years old and especially 70 years old). This correlation with age may be impacted by the sex (with for instance 40% of the diagnosis being melanoma for the most aged men - around 90 years old -  and far less for women), but this could also be an effect of data scarcity for some age range. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Relation between image location and other variables:","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"anatom_c = df_category['anatom_site_general_challenge'].value_counts()\nanatom_c = anatom_c.sort_index()\nfig,ax = plt.subplots() \nfor idx in anatom_c.index:\n    sns.kdeplot(df.loc[df['anatom_site_general_challenge']==idx,'age_approx'],shade=True,ax=ax)\nax.legend(anatom_c.index)\nax.set_title('Relation between age & anatomic site',fontsize=20)\nres = ax.set_xlabel('approx_age')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"ax = sns.catplot(x=\"sex\",\n            y=\"target\",\n            hue=\"anatom_site_general_challenge\",\n            kind=\"bar\",\n            data=df_category,\n            height=9, \n            aspect=2.3);\n_=ax.fig.suptitle('Influence of the image site on the target, by sex', y=1.02,fontsize=20)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(\"Number of female patient with an image of the oral/genital site:\", len(df.loc[(df_category['sex']==\"female\") & (df_category['anatom_site_general_challenge']==\"oral/genital\"),'patient_id'].unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Takeaways\n\nThe part of the body where the mole is located seems to have an important influence:\n- with head/neck moles usually more associated with malignant moles. \n- This also apparently dependent on the sex,as the latter effect is especially significan for men. \n- While this may be related to the relatively small sample we have, it seems taht for women, oral/genital moles are especially at risk.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<div id=\"interactions_patient\">\n    \n## c) Interaction at the patient level\n    \n</div>\n","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(\"Number of unique patients in the dataset:\", len(df_category['patient_id'].unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Actually, there is only 2056 different patients (for more than 30k pictures). The distribution of the number of images by patient can be seen below.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(20,10))\ndf_category.groupby('patient_id').count()['image_name'].hist(ax=ax,bins=100)\nax.set_title('Number of pictures by patient', fontsize=20)\nax.grid(False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Takeaways: \n\n- There are often more than one picture (and even more than 5) for each patient. Let's have a deeper dive into the relation that those images have between each other. ","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"ax= df_category.groupby('patient_id').aggregate({'age_approx':pd.Series.nunique}).hist()[0][0]\nax.set_title('Distribution of the number of different age_approx for each patient',fontsize=20)\nax.grid(False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Interestingly, same patient are often associated with more than one approximative age. While ths can be associated with errors in the dataset, this is also probably associatded with the fact that those images have been taken at different moments.\n\nAnother interesting question - partially influence by the previous - is the quantity of information carried by an image about another of the same patient. In other words, for a given image, could we use other images of the patient to get information on the diagnosis of the first image? ","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df_category_patient_with_melanoma = df_category.groupby('patient_id').apply(lambda x: 'melanoma' in set(x['diagnosis']))\nset_category_patient_with_melanoma = {x[0] for x  in df_category_patient_with_melanoma.iteritems() if x[1]}\ndf_category_patient_with_melanoma = df_category.loc[df_category['patient_id'].isin(set_category_patient_with_melanoma)]\n\n# We extract the id of the first image with melanoma for each patient with at least one melanoma \nfirst_melanoma_pictures = df_category_patient_with_melanoma.loc[df_category_patient_with_melanoma['diagnosis']=='melanoma'].drop_duplicates(['patient_id','diagnosis'],keep='first').index\n\n# We select all the other images\ndf_other_images_patient_with_melanoma = df_category_patient_with_melanoma.loc[df_category_patient_with_melanoma['diagnosis'] != 'melanoma']\ndf_other_images_patient_with_melanoma =  df_category_patient_with_melanoma.drop(first_melanoma_pictures)\ndf_other_images_patient_with_melanoma['image_set'] = 'Other images of patients with melanoma'\n\n\ndf_all_images = df_category.copy()\ndf_all_images['image_set'] = 'Full dataset'\n\n# We can concatenate all the images with only other images of patient with melanoma, in order to build comparative statics on the risk of occurence of melanoma\nconcatenated = pd.concat([df_all_images,\n          df_other_images_patient_with_melanoma])\n\nax = sns.catplot(x=\"sex\",\n            y=\"target\",\n            hue=\"image_set\",\n            kind=\"bar\",\n            data=concatenated,\n            height=8, \n            aspect=2.2);\n_=ax.fig.suptitle('Influence on the diagnosis of the existence of a melanoma elsewhere',fontsize=20,y=1.02)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Takeaways\n\n- Comparing all images vs images of patient with a melanoma (but not considering this image), we see that the latter category is more likely to have a melanoma.\n- This effect seems to be stable accross both sexes. \n- This is is an important information as it shows that prediction is likely to be improved using information (data or metadata) from other images.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<div id=\"images\">\n    \n# III. Image data\n\n    \n</div>","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"path_train_jpg = '../input/siim-isic-melanoma-classification/jpeg/train/'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The interactive cell below can be used to see the different types of diagnosis. It takes a few seconds to run. ","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"diagnosis_list = ['unknown', 'nevus','melanoma','seborrheic keratosis','lentigo NOS',\n 'lichenoid keratosis','solar lentigo','atypical melanocytic proliferation','cafe-au-lait macule']\ndef show_diagnosis(diagnosis='melanoma'):\n    assert diagnosis in diagnosis_list\n    fig, ax = plt.subplots(3,2,figsize=(10,10))\n    samples = df.loc[df['diagnosis']==diagnosis].sample(6,replace=True)['image_name']\n    ax = ax.ravel()\n    for j, name in enumerate(samples): \n        ax[j].imshow(plt.imread(path_train_jpg+name+'.jpg'))\n        ax[j].grid(False)\n        # Hide axes ticks\n        ax[j].set_xticks([])\n        ax[j].set_yticks([])\n    fig.suptitle(\"Diagnosis: \"+diagnosis,fontsize=20, y=0.95)\n\n    plt.show()\n    \nint = interact(show_diagnosis,diagnosis=diagnosis_list) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"***","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<div id=\"wrapup\">\n    \n# IV Wrap-up\n\n\nThis quick EDA provided some insights - mainly regarding tabular data/metadata, that will probably be useful to improve prediction performances on this dataset. Here are the most promising we found: \n\nInterestingly, the body part from which the picture is taken is also imbalanced, with more than 50% corresponding to the torso. It will be interesting to see how it can impact the result.\n\n- **Dataset is generally imbalanced**, and even strongly imbalanced when it comes to the target (98% of the diagnosis are not melanoma)\n- The more aged a person, the more likely the diagnosis for a skin lesion would be melanoma. This seems to be especially true for men. \n- **Melanoma diagnosis seems to be more likey for men** than for women (but further investigation are needed to understand if this is not only an \"age effect\"). \n- **The skin site is also an important parameter**, with head/neck & upper extremity more likely to result in a melanoma diagnosis. Here as well, there is apparently a \"sex effect\", with head/neck sites far more likely to be associated with melanoma diagnosis for men than for women.\n- In this dataset, same patients appear usually more than once, as multiple site of lesions are pictured. Furthermore, information a site is likely to be useful for another site since **people with a melanoma are apparently more likely to have another one melanoma on another site.**\n\n\nI'll had new analysis later. Don't hesitate to tell me if you have any remarks or ideas that could be explored to understand how to make the best use of those data ! \n    \n    \n</div>","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}