{"cells":[{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div align='center'><font size=\"6\" color=\"#43C6DB\">SIIM-ISIC Melanoma Classification </font></div>\n<hr>\n\n![](https://higherlogicdownload.s3.amazonaws.com/SITCANCER/2c19e5a6-3adb-4d01-b46c-c01e11745b3a/UploadedImages/Education/Patient_CONNECT/Melanoma_Patient_Resource_Guide/Stages_of_Melnoma.jpg)\nsource:https://higherlogicdownload.s3.amazonaws.com/SITCANCER/2c19e5a6-3adb-4d01-b46c-c01e11745b3a/UploadedImages/Education/Patient_CONNECT/Melanoma_Patient_Resource_Guide/Stages_of_Melnoma.jpg\n\n\n## Objective\n\nIn this competition, the goal is to identify melanoma in images of skin lesions. There are images within the same patient and and the task is to determine which are likely to represent a melanoma. Value 0 denotes benign, and 1 indicates malignant. Using patient-level contextual information may help the development of image analysis tools, which could better support clinical dermatologists.\n\nMelanoma is a deadly disease, but if caught early, most melanomas can be cured with minor surgery. Image analysis tools that automate the diagnosis of melanoma will improve dermatologists' diagnostic accuracy. Better detection of melanoma has the opportunity to positively impact millions of people.\n\n\n## Evaluation Metric\n\n![](https://i.pinimg.com/736x/dd/aa/f1/ddaaf1a24e601be6085e8b2ef6c7d819.jpg)\nsource: https://i.pinimg.com/736x/dd/aa/f1/ddaaf1a24e601be6085e8b2ef6c7d819.jpg\n\n\n\nReferences: \n\nhttps://www.kaggle.com/parulpandey/melanoma-classification-eda-starter\n\nhttps://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"\n# Prepare workspace\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Upload Libraries","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"\nimport os\nimport pandas as pd \nimport numpy as np\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport statistics\nimport scipy.stats as stats\nimport statsmodels.api as sm\nfrom statsmodels.formula.api import logit\nfrom scipy.stats import chi2_contingency\nfrom scipy.stats import kurtosis \nfrom scipy.stats import skew\nfrom statistics import stdev \nimport seaborn as sns\nfrom tqdm import tqdm\nfrom PIL import Image\nimport warnings\nwarnings.filterwarnings('ignore')\nplt.style.use('fivethirtyeight')\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Upload data sets","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"\ntrain = pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ntest = pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Summarize Data","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Look at dimension of data set and types of each attribute\ntrain.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Summarize attribute distributions of the data frame\ntrain.describe(include='all')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.describe(include='all')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Take a peek at the first rows of the data\ntrain.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The number of unique patients is less than the total number of patients, so each patient have multiple images.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Exploratory Data Analysis","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Handling Missing Values","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check missing values both to numeric features and categorical features\ntrain.isnull().sum()/train.shape[0]*100","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.isnull().sum()/test.shape[0]*100","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Data Imputation\n# Input missing values with median or mode depending of features class\ntrain['sex'].fillna(train['sex'].mode()[0], inplace=True)\ntrain['age_approx'].fillna(train['age_approx'].median(), inplace=True)\ntrain['anatom_site_general_challenge'].fillna(train['anatom_site_general_challenge'].mode()[0], inplace=True)\ntest['anatom_site_general_challenge'].fillna(test['anatom_site_general_challenge'].mode()[0], inplace=True)\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Target Variable Analysis","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"The target variable is grouped into two classes: 0 = benign and 1 = malignant. Looking at the barplot, it's higly imbalanced.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Summarize the class distribution \ncount = pd.crosstab(index = train['target'], columns=\"count\")\npercentage = pd.crosstab(index = train['target'], columns=\"frequency\")/pd.crosstab(index = train['target'], columns=\"frequency\").sum()\npd.concat([count, percentage], axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plot the target variable\nax = sns.countplot(x=train['target'], data=train, order=[0,1]).set_title(\"Target Variable Distribution\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Numerical Features Analysis","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Univariate analysis looking at Standard Deviation, Skewness and Kurtosis for train set\n\nprint('\\nStandard Deviation :', stdev(train['age_approx']), \n      '\\nSkewness :', skew(train['age_approx']), \n        '\\nKurtosis :', kurtosis(train['age_approx']))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Univariate analysis looking at Standard Deviation, Skewness and Kurtosis for test set\n\nprint('\\nStandard Deviation :', stdev(test['age_approx']), \n      '\\nSkewness :', skew(test['age_approx']), \n        '\\nKurtosis :', kurtosis(test['age_approx']))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# graphical function for univariate analysis\n\ndef num_plot(dataframe, feature):\n    plt.figure(figsize=(15, 5))\n\n    # histogram\n    plt.subplot(1, 3, 1)\n    sns.distplot(train[feature], bins=30, color='g')\n    plt.title('Histogram')\n    # Q-Q plot\n    plt.subplot(1, 3, 2)\n    stats.probplot(train[feature], dist=\"norm\", plot=plt)\n    plt.ylabel('Variable quantiles')\n    # boxplot\n    plt.subplot(1, 3, 3)\n    x=train[feature]\n    sns.boxplot(x,linewidth=1.5, color='g')\n    plt.title('Boxplot')\n\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# train\n# age_approx\nnum_plot(train, 'age_approx')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# test\n# age_approx\nnum_plot(test, 'age_approx')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Feature selection with Kendall's Test\n\nalpha = 0.05\nvar = 'age_approx'\np = stats.kendalltau(train['target'],train[var])[1]\nif p <= alpha:\n    print('{0} Dependent (reject H0)'.format(var))\nelse:\n    print('{0} Independent (fail to reject H0)'.format(var))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Categorical Features Analysis","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Univariate analysis with frequency and barplots for train set\nsns.set( rc = {'figure.figsize': (5, 5)})\nfcat_tr = ['sex','anatom_site_general_challenge','diagnosis','benign_malignant']\n\nfor col in fcat_tr:\n    count = pd.crosstab(index = train[col], columns=\"count\")\n    percentage = pd.crosstab(index = train[col], columns=\"frequency\")/pd.crosstab(index = train[col], columns=\"frequency\").sum()\n    tab = pd.concat([count, percentage], axis=1)\n    plt.figure()\n    sns.countplot(x=train[col], data=train, palette=\"Set1\")\n    plt.xticks(rotation=45)\n    print(tab)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Univariate analysis with frequency and barplots for test set\nsns.set( rc = {'figure.figsize': (5, 5)})\nfcat_te = ['sex','anatom_site_general_challenge']\n\nfor col in fcat_te:\n    count = pd.crosstab(index = test[col], columns=\"count\")\n    percentage = pd.crosstab(index = test[col], columns=\"frequency\")/pd.crosstab(index = test[col], columns=\"frequency\").sum()\n    tab = pd.concat([count, percentage], axis=1)\n    plt.figure()\n    sns.countplot(x=test[col], data=test, palette=\"Set1\")\n    plt.xticks(rotation=45)\n    print(tab)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Bivariate analysis with barplots for train set\nsns.set( rc = {'figure.figsize': (5, 5)})\n\nfor col in fcat_tr:\n    plt.figure()\n    sns.countplot(x=train[col], hue=train['target'], data=train, palette=\"Set2\")\n    plt.xticks(rotation=45)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Feature Selection with Chi-Square Test \nalpha = 0.05\nfor var in fcat_tr:\n    X = train[var].astype(str)\n    Y = train['target'].astype(str)\n    dfObserved = pd.crosstab(Y,X)\n    chi2, p, dof, expected = stats.chi2_contingency(dfObserved.values)\n    if p <= alpha:\n    \tprint('{0} Dependent (reject H0)'.format(var))\n    else:\n        print('{0} Independent (fail to reject H0)'.format(var))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"### Visualising a keep of images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"img = []\nfor i, image_id in enumerate(tqdm(train['image_name'].head(20))):\n    im = Image.open(f'../input/siim-isic-melanoma-classification/jpeg/train/{image_id}.jpg')\n    im = im.resize((128, )*2)\n    img.append(im)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img[0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img[5]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img[10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img[15]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"### Visualizing Images with Benign lesions","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"benign = train[train['benign_malignant']=='benign']\nmalign = train[train['benign_malignant']=='malignant']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_b = []\nfor i, image_id in enumerate(tqdm(benign['image_name'].tail(20))):\n    im = Image.open(f'../input/siim-isic-melanoma-classification/jpeg/train/{image_id}.jpg')\n    im = im.resize((128, )*2)\n    img_b.append(im)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_b[1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_b[6]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_b[11]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_b[16]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Visualizing Images with Malignant lesions","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"img_m = []\nfor i, image_id in enumerate(tqdm(malign['image_name'].head(20))):\n    im = Image.open(f'../input/siim-isic-melanoma-classification/jpeg/train/{image_id}.jpg')\n    im = im.resize((128, )*2)\n    img_m.append(im)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_m[2]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img[7]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img[12]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img[17]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}