{"cells":[{"metadata":{},"cell_type":"markdown","source":"![](https://i.ytimg.com/vi/CF24dVuQImU/maxresdefault_live.jpg)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### This kernel is based on basic exploratory data analysis and augmentations.Please give me an upvote , if you like this notebook as this is my first exercise in deep learning.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### References\n\n* https://www.kaggle.com/nxrprime/siim-d3-eda-augmentations-and-resnext#seven\n* https://www.tensorflow.org/tutorials/structured_data/imbalanced_data#define_the_model_and_metrics\n* https://www.kaggle.com/parulpandey/melanoma-classification-eda-starter","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# What is Melanoma?","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* Melanoma is a type of skin cancer that occurs when pigment producing cells called melanocytes mutate and begin to divide uncontrollably.\n* Most pigment cells develop in the skin. Melanomas can develop anywhere on the skin, but certain areas are more at risk than others. \n* In men, it is most likely to affect the chest and back. In women, the legs are the most common site. Other common sites of melanoma include the face.\n* However, melanoma can also occur in the eyes and other parts of the body, including — on very rare occasions — the intestines.\n* When this happens, it can be difficult to treat, and the outlook may be poor. \n* Risk factors for melanoma include overexposure to the sun, having fair skin, and a family history of melanoma, among others.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Risk Factors","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* Research into the exact causes of melanoma is ongoing.\n* However, scientists do know that people with certain skin types are more prone to developing melanoma.\n\n**The following factors may also contribute to an increased risk of skin cancer:**\n\n1. A high density of freckles or a tendency to develop freckles following exposure to the sun.\n\n2. A high number of moles or five or more atypical moles\n\n3. The presence of actinic lentigines, also known as liver spots or age spots\n\n4. Pale skin that does not tan easily and tends to burn.\n\n5. Light eyes, red or light hair\n\n6. High sun exposure, particularly if it produces blistering sunburn, and if sun exposure is intermittent rather than regular older age\n\n7. family or personal history of melanoma\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## ABCDE examination","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* The ABCDE examination of moles is an important method for revealing potentially cancerous lesions. \n* It describes five simple characteristics to check for in a mole that can help a person either confirm or rule out melanoma:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![](https://discoverplasticsurgery.com/wp-content/uploads/2018/09/melanoma-risk-factors.jpg)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Objective","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* To identify melanoma in images of skin lesions. \n* In particular, you’ll use images within the same patient and determine which are likely to represent a melanoma. \n* Using patient-level contextual information may help the development of image analysis tools, which could better support clinical dermatologists.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Dataset Info","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* The dataset contains 33,126 dermoscopic training images of unique benign and malignant skin lesions from over 2,000 patients.\n* Each image is associated with one of these individuals using a unique patient identifier.\n* All malignant diagnoses have been confirmed via histopathology, and benign diagnoses have been confirmed using either expert agreement, longitudinal follow-up, or histopathology.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Images","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* The images are provided in three formats.\n\n1) DICOM (Digital Imaging and Communications in Medicine) is the international standard to transmit, store, retrieve, print, process, and display medical imaging information.This can be accessed using libraries like pydicom.\n\n2) JPEG\n\n3) TFRecord","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Import Libraries","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import os\n\nfrom os import listdir  #returns a list that gives the names of the entries in the directory\nfrom os.path import isfile,join\n\nimport pandas as pd\nimport numpy as np\nfrom numpy import math\nimport seaborn as sns\nsns.set(style='darkgrid')\nimport matplotlib.pyplot as plt\n%matplotlib inline\nplt.style.use('fivethirtyeight')\nplt.show()\n\n#Plotly\nimport plotly.express as px\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\nimport cufflinks\ncufflinks.go_offline()\ncufflinks.set_config_file(world_readable=True,theme='pearl')\n\n#To read a dicom image , we can use pydicom\nimport pydicom\n\n#Disable warnings\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nfrom tensorflow.keras.applications import ResNet50\nfrom keras.models import Sequential, Model,load_model\nfrom keras.layers import Flatten,Dense","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"DEVICE = 'GPU'\n\nif DEVICE == \"GPU\":\n    print(\"Num GPUs Available: \", len(tf.config.experimental.list_physical_devices('GPU')))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## List directories","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#The simplest way to get a list of entries in a directory is to use os.listdir()\n#Pass in the directory you need the entries\nos.listdir('../input/siim-isic-melanoma-classification')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We have two csv files,train.csv and test.csv\n* Two dicom image files, train and test\n* Tfrecords file\n* A jpeg file with train and test files\n* And sample submission file","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Csv files","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ntest = pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* 33126 Images and 8 columns","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Columns","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* **image_name** - unique identifier, points to filename of related DICOM image\n* **patient_id** - unique patient identifier\n* **sex** - the sex of the patient (when unknown, will be blank)\n* **age_approx** - approximate patient age at time of imaging\n* **anatom_site_general_challenge** - location of imaged site\n* **diagnosis** - detailed diagnosis information (train only)\n* **benign_malignant** - indicator of malignancy of imaged lesion\n* **target** - binarized version of the target variable","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Missing values","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#First we create a list of missing values by each feature\nmissing = list(train.isna().sum())\n\n#then we create a list of columns and their missing values as inner list to a separate list\nlst= []\ni=0\nfor col in train.columns:\n    insert_lst = [col,missing[i]]\n    lst.append(insert_lst)\n    i+=1\n\n#finally create a dataframe\nmissing_df = pd.DataFrame(data=lst,columns=['Column_Name','Missing_Values'])\n\nfig = px.bar(missing_df,x='Missing_Values',y='Column_Name',orientation='h',\n             text='Missing_Values',title='Missing values in train dataset')\nfig.update_traces(textposition='outside')\nfig.show()\n\n#Same thing for test file\nmissing = list(test.isna().sum())\n\nlst= []\ni=0\nfor col in test.columns:\n    insert_lst = [col,missing[i]]\n    lst.append(insert_lst)\n    i+=1\n\n#finally create a dataframe\nmissing_df = pd.DataFrame(data=lst,columns=['Column_Name','Missing_Values'])\n\nfig = px.bar(missing_df,x='Missing_Values',y='Column_Name',orientation='h',\n             text='Missing_Values',title='Missing values in test dataset')\nfig.update_traces(textposition='outside')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Three columns in train dataset with missing values\n* Anatom_site_general_challenge , sex and age_approx\n* One test column with missing value , Anatom_site_general_challenge","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### Sex","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* Now as sex has two unique values , male and female.It becomes difficult to impute missing values for this feature.\n* One method is to use mode . i.e male in this feature.\n* Another is to relate it with other variables.We'll try this method and see if we can find something. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# We separate the non nan values and nan values in separate dataframe.\n\nnot_null_sex = train[train['sex'].notnull()].reset_index(drop=True)\nnan_sex = train[train['sex'].isnull()].reset_index(drop=True)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"not_null_sex.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(15,6))\n\nfig1 = sns.countplot(data=not_null_sex,hue='sex',x='anatom_site_general_challenge')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Check the anatom site in missing values.\n\nnan_sex['anatom_site_general_challenge'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Two patients ['IP_9835712', 'IP_5205991'] (sex) and (age) is not given.All missing values are benign,diagnosis='unknown' and anatomy as seen above.\n* We relate it with other feature but we don't see any significant difference in both the sex\n* So we go with mode.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Compute missing value with mode of sex\n\ntrain['sex'].fillna(train['sex'].mode()[0],inplace=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Age","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['age_approx'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['age_approx'].median()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* The mode is 45 and median is 50.\n* It's best to use median to fill the missing values.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Compute missing values with median\n\ntrain['age_approx'].fillna(train['age_approx'].median(),inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['age_approx'].isna().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Anatom_site_general_challenge","execution_count":null},{"metadata":{"trusted":true},"cell_type":"markdown","source":"* There are six anatomy sites in our data.\n* If we see the mode , it is torso with 16845 values.\n* There are 527 missing values in Anatom_site_general_challenge.\n* So as there are more than 500 missing values in this feature , I will add another category of 'NK' i.e NotKnown as we can't predict what the anatomy site will be for the patient.\n* Test dataset also has one column with missing values i.e anatom_site_general_challenge\n* Now as we filled 'NK' inplace of missing values in training dataset , we'll do the same in test dataset","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train['anatom_site_general_challenge'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train['anatom_site_general_challenge'].fillna('NK',inplace=True)\ntest['anatom_site_general_challenge'].fillna('NK',inplace=True)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Let's check if there are any missing values left","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Train : {}'.format(train.isna().sum().sum()))\nprint('Test : {}'.format(test.isna().sum().sum()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# EDA on the above features","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## First , the target feature","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Our target feature has two categories** \n* **`Benign`**\n\n* **`Malignant`**","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![](https://chcsga.org/wp-content/uploads/2019/05/d.jpg**)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<img src=https://chcsga.org/wp-content/uploads/2019/05/d.jpg width=\"500\">","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig=plt.figure(figsize=(15,8))\n\nlabels = 'Benign','Malignant'\n\nbenign = train[train['benign_malignant']=='benign']\nmalignant = train[train['benign_malignant']=='malignant']\nsizes = [len(benign),len(malignant)]\n\ncolors= ['lightskyblue','red']\n#Plot\nplt.pie(sizes,labels=labels,colors=colors,autopct='%1.1f%%',shadow=True,startangle=140)\n\nplt.axis('equal');\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We have more benign cases than malignant.\n* About 98.2% are benign cases and only 1.8% malignant cases are there in train dataset.\n* We can clearly see , there is imbalance in class data.\n* This we need to keep in mind while model building.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### How many patients are there in the dataset.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"There are {} number of patients in our dataset.\".format(train['patient_id'].nunique()))\nprint(\"And there are total {} dicom images in the same dataset.\".format(train['image_name'].nunique()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# We groupby patient id and see the number of images wrt to each patient\n\nx = train.groupby(['patient_id'],as_index=False)['image_name'].count()\nx.sort_values(by=\"image_name\",ascending=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### So the maximum number of images for a patient is 115 and least is 2","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Sex","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"x = train.groupby(['sex'],as_index=False)['benign_malignant'].count()\nx = x.set_index('sex')\nx","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(data=train,x='sex',hue='benign_malignant');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# In the test dataset\n\nsns.countplot(data=test,x='sex');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* In both train and test datasets , males are more than females.\n* If we relate it to target in the train dataset , then in both gender , there are more number of benign cases.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Age","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<img src=https://www.cancerresearchuk.org/sites/default/files/cancer-stats/cases_crude_mf_allcancer_i17/cases_crude_mf_allcancer_i17.png width=\"1000\" height=\"700\">","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* The risk of melanoma increases as people age.\n* The average age of people when it is diagnosed is 65.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def create_dist(df,title):\n    fig = plt.figure(figsize=(15,6))\n\n    x= df[\"age_approx\"].value_counts(normalize=True).to_frame()\n    x = x.reset_index()\n    ax = sns.barplot(data=x,y='age_approx',x='index')\n    ax.set(xlabel='Age', ylabel='Percentage')\n    ax.set(title=title);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"create_dist(train,\"Age distribution in train dataset\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"create_dist(test,\"Age distribution in test dataset\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(15,6))\n\nax = sns.countplot(data=train,x='age_approx',hue='benign_malignant');\nax.set(title='Age vs Target');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* In the train dataset, Age follows a gaussian distribution and in test it's not the same.\n* But in both datasets, we have more number of middle aged patients.\n* So, till age 40 , there are no malignant cases in train dataset.And from age 45 to 75 , there are malignant cases.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Anatomy sites","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(train,y='anatom_site_general_challenge',height=500,width=800,color_discrete_sequence=['indianred'])\nfig.show()\n\nfig = px.histogram(train,x='anatom_site_general_challenge',color='benign_malignant',barmode='group',height=500,width=800)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Most of the cases in the dataset are on torso area and then on extremities ( upper and lower)\n* There are few cases on palms,soles,oral and genitals.\n* Only four locations in the body are having malignant cases , although the number is less (torso,extremity and head/neck)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(test,y='anatom_site_general_challenge',height=500,width=800,color_discrete_sequence=['indianred'],title='Similar case in test dataset too')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Skin lesions","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(train,y='diagnosis',height=500,width=800,color_discrete_sequence=['goldenrod'],title='Diagnoses skin lesions')\nfig.update_layout(uniformtext_minsize=8, uniformtext_mode='hide')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(train,x='diagnosis',color='benign_malignant',barmode='group',height=500,width=800)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So we have 9 different skin lesions out of which all are not cancerous.\n\n* **Nevus** :- A common pigmented skin lesion, usually developing during adulthood.In most cases, a nevus is benign and doesn't require treatment. Rarely, they turn into melanoma or other skin cancers. A nevus that changes shape, grows bigger or darkens should be evaluated for removal.\n\n* **Melanoma** :- We already know what this is.\n\n* **Seborrheic keratosis** :- A non-cancerous skin condition that appears as a waxy brown, black or tan growth.A seborrhoeic keratosis is one of the most common non-cancerous skin growths in older adults.\n\n* **Lentigo NOS** :- A typen of skin cancer that appears on your trunk, arms, and legs. Lentigo often starts at birth or during childhood. The spots can go away in time.\n\n* **Lichenoid keratosis** :- Lichenoid keratosis is a skin condition that typically occurs as a single, small, raised plaque, thickened area, or papule.This condition is harmless. However, in some cases lichenoid keratosis can be mistaken for other kinds of skin conditions, including skin cancers.\n\n* **Solar lentigo** :- Solar lentigo is caused by exposure to ultraviolet radiation from the sun. This type is common in people over age 40, but younger people can get it, too. It happens when UV radiation causes pigmented cells called melanocytes in the skin to multiply. Solar lentigo appears on sun-exposed areas of the body, like the face, hands, shoulders, and arms. The spots may grow over time.\n\n* **Cafe-au-lait macule** :- A café-au-lait macule is a common birthmark, presenting as a hyperpigmented skin patch with a sharp border and diameter of > 0.5 cm.\n\n* **Atypical melanocytic proliferation** :- Atypical Melanocytic lesions are irregular moles and skin spots that require further examination. The five visual characteristics used to identify an atypical melanocytic lesion are the same as the characteristics used to identify signs of invasive melanoma.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Let's look at some images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Create a separate images folder\ntrain_images_dir = '../input/siim-isic-melanoma-classification/train/'\ntrain_images = listdir(train_images_dir)\n\ntest_images_dir = '../input/siim-isic-melanoma-classification/test/'\ntest_images = listdir(test_images_dir)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Define a function to plot randomly sampled images using pydicom\n\ndef plot_images(df):\n    fig = plt.figure(figsize=(15,6))\n\n    for i in range(1,11):\n        image = df['image_name'][i]\n        ds = pydicom.dcmread(train_images_dir+image+'.dcm')\n        fig.add_subplot(2,5,i)\n        plt.imshow(ds.pixel_array)\n    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#We sample 11 rows from train dataset\nrandom = train.sample(n=11)\nrandom = random.reset_index(drop=True)\n\n#Plot the images\nplot_images(random)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We can see the images are of different sizes ,with different lighting conditions , different body parts.\n* These all things need to be considered for model building.\n* We need to perform scaling , resizing and some data augmentation techniques.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### As we are predicting benign and malignant cases, let's look at these images.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Benign images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Similary , we sample random benign images\nrandom = train[train['benign_malignant']=='benign'].sample(n=11)\nrandom = random.reset_index(drop=True)\n\nplot_images(random)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* In benign , we can see that it is concentrated and not spread out like me.\n* Also , the diameter is less.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Malignant images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Similary , we sample random malignant images\nrandom = train[train['benign_malignant']=='malignant'].sample(n=11)\nrandom = random.reset_index(drop=True)\n\nplot_images(random)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We can see the irregularities and change in shapes of the above images.\n* Also see the diameter and assymmetry.\n* Although, it is difficult to classify just looking at the images.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Skin lesions grow on various parts of the body.\nIn the anatomy_site feature,skin lesions on six location has been given.\nLet's look at each part and study the images.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"We have six locations where image is taken,we can study benign and malignant images in each of these sites","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#define a function for plotting anatomy sites\n\ndef plot_anatomy(target,anatomy_site):\n    anatomy = train[train['anatom_site_general_challenge']==anatomy_site]\n\n    fig = plt.figure(figsize=(15,6))\n    for i in range(0,4):\n        image = anatomy[anatomy['benign_malignant']==target].reset_index(drop=True)['image_name'][i]\n        ds = pydicom.dcmread(train_images_dir+image+'.dcm')\n        fig.add_subplot(2,4,i+1)\n        plt.imshow(ds.pixel_array)\n        plt.title(target)\n    plt.suptitle(anatomy_site)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_anatomy('benign','head/neck')\nplot_anatomy('malignant','head/neck')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_anatomy('benign','upper extremity')\nplot_anatomy('malignant','upper extremity')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_anatomy('benign','lower extremity')\nplot_anatomy('malignant','lower extremity')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_anatomy('benign','torso')\nplot_anatomy('malignant','torso')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_anatomy('benign','palms/soles')\nplot_anatomy('malignant','palms/soles')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_anatomy('benign','oral/genital')\nplot_anatomy('malignant','oral/genital')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Malignant images have some irregularities and also change in shapes and diameter.\n* Difference can be clearly seen in torso,palms,soles,genital and extremities sites.\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Different diagnosis of Skin lesions","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![](https://upload.wikimedia.org/wikipedia/commons/thumb/d/d3/Pie_chart_of_incidence_and_malignancy_of_pigmented_skin_lesions.png/800px-Pie_chart_of_incidence_and_malignancy_of_pigmented_skin_lesions.png)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* By looking at the above image, we se that Nevus,Keratosis,lentigo .These all are non=cancerous.\n* Melanoma is seen in red colour and malignant.\n* We'll study some of these lesions below for better understanding by looking at the images.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_diagnosis(skin_lesion):\n    fig = plt.figure(figsize=(12,6))\n\n    for i in range(0,6):\n        image = train[train['diagnosis']==skin_lesion].reset_index(drop=True)['image_name'][i]\n        ds = pydicom.dcmread(train_images_dir+image+'.dcm')\n        fig.add_subplot(2,3,i+1)\n        plt.imshow(ds.pixel_array)\n    plt.suptitle(skin_lesion.upper())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_diagnosis('nevus')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* From nevus , we see that the moles are concentrated and not spread out .Showing no signs of malignant.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_diagnosis('melanoma')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Melanoma as we all know is dangerous .The skin lesion is spread out , we can see the colour , the shape and irregularity.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_diagnosis('seborrheic keratosis')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Seborrheic keratosis is a noncancerous condition that can look a lot like melanoma.\n* The growths look waxy as if they are painted onto the body.\n* These do not typically cause symptons, but some people dislike the way they look.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_diagnosis('lentigo NOS')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_diagnosis('lichenoid keratosis')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* It looks like scaly, dry patches on the skin.\n*  Almost 90 percent of people with lichenoid keratosis will have just one lesion or spot on the skin.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_diagnosis('solar lentigo')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Solar lentigo appears on sun-exposed areas of the body, like the face, hands, shoulders, and arms. \n* The spots may grow over time. Solar lentigines are sometimes called liver spots or age spots.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### There is only one image for ('cafe-au-lait macule') and ('atypical melanocytic proliferation')","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(10,6))\n\nimage = train[train['diagnosis']=='cafe-au-lait macule'].reset_index(drop=True)['image_name'][0]\nds = pydicom.dcmread(train_images_dir+image+'.dcm')\nfig.add_subplot(1,2,1)\nplt.imshow(ds.pixel_array)\nplt.title('cafe-au-lait macule'.upper())\n\nimage = train[train['diagnosis']=='atypical melanocytic proliferation'].reset_index(drop=True)['image_name'][0]\nds = pydicom.dcmread(train_images_dir+image+'.dcm')\nfig.add_subplot(1,2,2)\nplt.imshow(ds.pixel_array)\nplt.title('atypical melanocytic proliferation'.upper());\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is all for EDA right now.The following model building is commented for now and I'll be soon updating it.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Model Building","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Using RepeatedKFold","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"* As our target data is highly imbalanced i.e 98.2% benign and 1.8% malignant.\n* We can use cross validation here for imbalanced data classification.\n* This ensures that the proportion of benign to malignant samples found in the original distribution is respected in all the folds.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#Import\nfrom sklearn.model_selection import RepeatedKFold\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def load_data_kfold(k):\n    #X = train['image_name']\n    #y = train['target']\n    \n    #X_train,X_val = tts(train_x, test_size=0.2, random_state=1234)\n\n    #y_train = np.array(y_train)\n    train_x = train[['image_name','target']]\n    train_x['image_name'] = train_x['image_name'].apply(lambda x: x + '.jpg')\n    folds = list(RepeatedKFold(n_splits=k, n_repeats=1, random_state=0).split(train_x))\n    \n    return folds,train_x\n\nk = 3\nfolds,train_x = load_data_kfold(k)\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"folds","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ResNet50 Model","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"'''METRICS = [\n      tf.keras.metrics.TruePositives(name='tp'),\n      tf.keras.metrics.FalsePositives(name='fp'),\n      tf.keras.metrics.TrueNegatives(name='tn'),\n      tf.keras.metrics.FalseNegatives(name='fn'), \n      tf.keras.metrics.BinaryAccuracy(name='accuracy'),\n      tf.keras.metrics.Precision(name='precision'),\n      tf.keras.metrics.Recall(name='recall'),\n      tf.keras.metrics.AUC(name='auc'),\n]'''\n\ndef get_model():\n    model =ResNet50(weights='imagenet',include_top=False,input_shape=(224,224,3))\n\n    for layer in model.layers:\n        layer.trainable = False \n\n    x=Flatten()(model.output)\n    output=Dense(1,activation='softmax')(x)\n\n    model = Model(model.input,output)\n    \n    model.compile(\n    'Adam',\n    loss='sparse_categorical_crossentropy',\n    metrics=['accuracy'],\n    )\n    \n    return model\n    ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Model Summary","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"model = get_model()\nmodel.summary()\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Train model on each fold","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Initial parameters","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"class Config:\n    BATCH_SIZE = 64\n    EPOCHS = 10\n    HEIGHT = 224\n    WIDTH = 224","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for j, (train_idx, val_idx) in enumerate(folds):\n    \n    print('\\nFold ',j)\n    print('///////////////////////////////////')\n    X_train_cv = train_x.iloc[train_idx]\n    #y_train_cv = y_train[train_idx]\n    X_valid_cv = train_x.iloc[val_idx]\n    #y_valid_cv= y_train[val_idx]\n    \n    #name_weights = \"final_model_fold\" + str(j) + \"_weights.h5\"\n    #callbacks = get_callbacks(name_weights = name_weights, patience_lr=10)\n\n    train_datagen=tf.keras.preprocessing.image.ImageDataGenerator(rescale=1./255, \n                         rotation_range=360,\n                         horizontal_flip=True,\n                         vertical_flip=True)\n    \n    train_generator=train_datagen.flow_from_dataframe(\n        dataframe=X_train_cv,\n        directory='../input/siim-isic-melanoma-classification/jpeg/train/',\n        x_col=\"image_name\",\n        y_col=\"target\",\n        class_mode=\"raw\",\n        batch_size=Config.BATCH_SIZE,\n        target_size=(Config.HEIGHT, Config.WIDTH))\n\n    validation_datagen = tf.keras.preprocessing.image.ImageDataGenerator(rescale=1./255)\n\n    valid_generator=validation_datagen.flow_from_dataframe(\n        dataframe=X_valid_cv,\n        directory='../input/siim-isic-melanoma-classification/jpeg/train/',\n        x_col=\"image_name\",\n        y_col=\"target\",\n        class_mode=\"raw\", \n        batch_size=Config.BATCH_SIZE,   \n        target_size=(Config.HEIGHT, Config.WIDTH))\n    \n    model = get_model()\n    \n    TRAINING_SIZE = len(train_generator)\n    VALIDATION_SIZE = len(valid_generator)\n    BATCH_SIZE = 64\n\n    compute_steps_per_epoch = lambda x: int(math.ceil(1. * x / BATCH_SIZE))\n    steps_per_epoch = compute_steps_per_epoch(TRAINING_SIZE)\n    validation_steps = compute_steps_per_epoch(VALIDATION_SIZE)\n    \n    history = model.fit_generator(generator=train_generator,\n                                        steps_per_epoch=steps_per_epoch,\n                                        validation_data=valid_generator,\n                                        validation_steps=validation_steps,\n                                        epochs=10,\n                                        verbose=1)\n    \n    #print(model.evaluate(X_valid_cv['image_name'], X_valid_cv['target']))\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We have got a pretty good validation accuracy of 96% after 3 folds.\n* I don't understand why the loss value is nan , I'll try to correct it in later update.\n* Next step is prediction on test dataset.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Evaluation on test dataset","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_x = test[['image_name']]\n\ntest_x['image_name'] = test_x['image_name'].apply(lambda x: x + '.jpg')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_datagen = tf.keras.preprocessing.image.ImageDataGenerator(rescale=1./255)\n\ntest_generator = test_datagen.flow_from_dataframe(  \n        dataframe=test_x,\n        directory = '../input/siim-isic-melanoma-classification/jpeg/test/',\n        x_col=\"image_name\",\n        batch_size=1,\n        class_mode=None,\n        shuffle=False,\n        target_size=(Config.HEIGHT, Config.WIDTH),\n        seed=0)\n\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"preds = model.predict_generator(test_generator,verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"predicted_class_indices = np.argmax(preds, axis = 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"predicted_class_indices","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(preds)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(predicted_class_indices)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Creating Submission file","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sub = pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/sample_submission.csv')\nsub","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub['target'] = predicted_class_indices","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"### Things to do","execution_count":null},{"metadata":{"trusted":true},"cell_type":"markdown","source":"* Data Augmentations\n* Training model with class weights\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Please give an upvote if you like this notebook :)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![](https://media0.giphy.com/media/wIVA0zh5pt0G5YtcAL/source.gif)","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}