{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction\n\n# Melanoma Analysis Notebook\n\nHere is part of my notebook in GitHub repository do refer it and star it or fork it as per your convinience!!\nhttps://github.com/digs1998/Melanoma-Analysis","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Dataset\nI have made this project as part of [SIIM-ISIC-Melanoma-Classification competition](https://www.kaggle.com/c/siim-isic-melanoma-classification)\n\n# Process Flow of Project so far\n1. Importing Libraries\n2. Exploring the Dataset\n3. Data Visualisation using libraries like [plotly](https://plotly.com/python/)\n4. DICOM Image Preprocessing, libraries used were pydicom and PIL\n5. Creating Folds \n6. XGBoost based predictions\n\n***for DICOM Preprocessing I refered two notebooks***\n1. https://www.kaggle.com/nxrprime/siim-d3-js-eda-augmentations-model-seresunet\n2. https://www.kaggle.com/parulpandey/melanoma-classification-eda-starter","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Basic Introduction on Melanoma Cancer","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"from IPython.display import YouTubeVideo\nYouTubeVideo('hXYd0WRhzN4', width=800, height=450)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport pydicom, numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\nimport json\nfrom PIL import Image\nfrom IPython.display import display\nimport os\nimport plotly.graph_objects as px\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set()\nfrom plotly.offline import init_notebook_mode\nfrom plotly.offline import iplot\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.callbacks import EarlyStopping, LambdaCallback\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.layers import Dense, Conv2D, MaxPooling2D, Dropout, Flatten\nfrom tensorflow.keras.applications import DenseNet121\nfrom sklearn import metrics\nfrom sklearn.model_selection import StratifiedKFold\n\nimport cufflinks as cf\ncf.go_offline()\ncf.set_config_file(offline=False, world_readable=True)\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Loading Datasets","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\ntest_df = pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')\nsubs_df = pd.read_csv('../input/siim-isic-melanoma-classification/sample_submission.csv')\n\nimage_path = '../input/siim-isic-melanoma-classification/train/'\n\nprint('The size of training data : {}'.format(train_df.shape))\nprint('The size of testing data : {}'.format(test_df.shape))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"a = np.mean(train_df.target)\nprint('The Distribution of Training dataset : {}'.format(a))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"***Based on above value it is an binary classification type dataset***","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Visualizing Missing values","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize = (17,7))\npercent_missing = train_df.isnull().sum() / (train_df.shape[0])*100\npercent_missing.iplot(kind = 'bar',color ='blue')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Checking if Patient-ID were duplicated**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"The total patient ids are {}, from those the unique ids are {}\".format(train_df['patient_id'].count(),train_df['patient_id'].value_counts().shape[0] ))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"benign_gender = train_df.groupby(['benign_malignant']).count()['sex'].to_frame()\nbenign_gender.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Target and Age\n**Here I have used box-plot tool of seaborn library to visualize the distribution of age data in the target columns**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#target and age\nplt.figure(figsize = (17,7))\nsns.boxplot(x = train_df['target'], y = train_df['age_approx'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"feature_list = ['sex','age_approx','anatom_site_general_challenge'] \nfor i in feature_list: \n    train_df[i].value_counts(normalize=True).to_frame().iplot(kind='bar',\n                                                      yTitle='Percentage', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='blue',\n                                                      theme='pearl',\n                                                      bargap=0.8,\n                                                      gridcolor='white',                                                     \n                                                      title=f'<b>Distribution of {i} in train set.</b>')\n\n    test_df[i].value_counts(normalize=True).to_frame().iplot(kind='bar',\n                                                      yTitle='Percentage', \n                                                      linecolor='black', \n                                                      opacity=0.7,\n                                                      color='green',\n                                                      theme='pearl',\n                                                      bargap=0.8,\n                                                      gridcolor='white',                                                     \n                                                      title=f'<b>Distribution of {i} in test set.</b>')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Displaying Images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"im = train_df['image_name'].values\ndisplay(im)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(17,6))\n\nimage_dir = '../input/siim-isic-melanoma-classification/'\n\nimg = [np.random.choice(im + '.jpg') for i in range(10)]\nimg_dir = image_dir + '/jpeg/train'\n\nfor i in range(9):\n    plt.subplot(3,3, i+1)\n    images = plt.imread(os.path.join(img_dir, img[i]))\n    plt.imshow(images)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Display Benign and Malignant images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"malignant = train_df[train_df['benign_malignant']=='malignant']\nbenign = train_df[train_df['benign_malignant']=='benign']\n\nprint(malignant)\nprint(benign)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"malignant.head(5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"benign.head(5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Malignant**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"im_malignant = malignant['image_name'].values\nimage_dir = '../input/siim-isic-melanoma-classification/'\n\nimg = [np.random.choice(im_malignant + '.jpg') for i in range(10)]\nimg_dir = image_dir + '/jpeg/train'\nplt.figure(figsize=(17,17))\n\nfor i in range(9):\n    plt.subplot(3,3, i+1)\n    images = plt.imread(os.path.join(img_dir, img[i]))\n    plt.imshow(images)\n    plt.axis('off')\nplt.tight_layout()\nprint(\"Random Malignant Images are Displayed!!\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Benign**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"im_benign = benign['image_name'].values\nimage_dir = '../input/siim-isic-melanoma-classification/'\n\nimg = [np.random.choice(im_benign + '.jpg') for i in range(10)]\nimg_dir = image_dir + '/jpeg/train'\nplt.figure(figsize=(17,17))\n\nfor i in range(9):\n    plt.subplot(3,3, i+1)\n    images = plt.imread(os.path.join(img_dir, img[i]))\n    plt.imshow(images)\n    plt.axis('off')\nplt.tight_layout()\nprint('Random Benign Images are displayed!!')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# DICOM Images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.imshow(pydicom.dcmread(image_path + list(train_df['image_name'])[1] + '.dcm').pixel_array)\nplt.savefig('x.jpg')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Benign Images**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(17,17))\n\nfor i in range(9):\n    plt.subplot(3,3, i+1)\n    images = pydicom.dcmread(image_path + train_df[train_df['benign_malignant']=='benign']['image_name'][i] + '.dcm')\n    plt.imshow(images.pixel_array)\n    plt.axis('off')\nprint(\"--- Benign DICOM Images ---\")\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data Imbalance","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize = (10,10))\ndata = train_df.benign_malignant.value_counts()\ndata.iplot(kind = 'bar', color='blue', title = 'Data Imbalance')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Diagnosis and Age","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize = (20,15))\nsns.boxplot(x = train_df['diagnosis'], y = train_df['age_approx'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Histograms\nLet's visualize the image intensities using histogram ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**Benign Images**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"img = im_benign[0] + '.jpg'\nf = plt.figure(figsize=(15,10))\nf.add_subplot(1,2,1)\n\nimages = plt.imread(os.path.join(img_dir, img))\nplt.imshow(images, cmap='gray')\nplt.axis('off')\nplt.colorbar()\nplt.title('Benign Images')\n\nf.add_subplot(1,2,2)\n_= plt.hist(images[:, :, 0].ravel(), bins = 256, color = 'red', alpha = 0.5)\n_ = plt.hist(images[:, :, 1].ravel(), bins = 256, color = 'Green', alpha = 0.5)\n_ = plt.hist(images[:, :, 2].ravel(), bins = 256, color = 'Blue', alpha = 0.5)\n_ = plt.xlabel('Intensity Value')\n_ = plt.ylabel('Count')\n_ = plt.legend(['Red_Channel', 'Green_Channel', 'Blue_Channel'])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Malignant Images**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"img = im_malignant[0] + '.jpg'\nf = plt.figure(figsize=(15,10))\nf.add_subplot(1,2,1)\n\nimages = plt.imread(os.path.join(img_dir, img))\nplt.imshow(images, cmap='gray')\nplt.axis('off')\nplt.colorbar()\nplt.title('Malignant Images')\n\nf.add_subplot(1,2,2)\n_= plt.hist(images[:, :, 0].ravel(), bins = 256, color = 'red', alpha = 0.5)\n_ = plt.hist(images[:, :, 1].ravel(), bins = 256, color = 'Green', alpha = 0.5)\n_ = plt.hist(images[:, :, 2].ravel(), bins = 256, color = 'Blue', alpha = 0.5)\n_ = plt.xlabel('Intensity Value')\n_ = plt.ylabel('Count')\n_ = plt.legend(['Red_Channel', 'Green_Channel', 'Blue_Channel'])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Create Folds\nWe first create folds for Stratified Kfold model selection so that the ratio of positive to negative cases remain constant.\nYou can get more information from [here](https://www.youtube.com/watch?v=WaCFd-vL4HA)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn import model_selection","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#using the train.csv data\ntrain_df['kfold'] = -1\ntrain_df = train_df.sample(frac=1).reset_index(drop=True)\ny = train_df.target.values\nkf = model_selection.StratifiedKFold(n_splits=10)\n\nfor f, (t_, v_) in enumerate(kf.split(X=train_df, y=y)):\n    train_df.loc[v_, 'kfold'] = f\n\ntrain_df.to_csv(\"train_folds.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def train(fold):\n    train_path = \"../input/siic-isic-224x224-images/train/\"\n    df = pd.read_csv('/kaggle/working/train_folds.csv')\n    train_batch = 32\n    valid_batch = 16\n    epochs = 30\n    \n    df_train = df[df.kfold != fold].reset_index(drop=True)\n    df_valid = df[df.kfold == fold].reset_index(drop=True)\n\n    model = DenseNet121(pretrained=\"imagenet\", include_top=False)\n\n    train_aug = ImageDataGenerator(\n    featurewise_center=True,\n    featurewise_std_normalization=True,\n    rotation_range=20,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    horizontal_flip=True)\n    \n    valid_aug = ImageDataGenerator(\n    featurewise_center=True,\n    featurewise_std_normalization=True,\n    rotation_range=20,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    horizontal_flip=True)\n    \n\n    train_images = df_train.image_name.values.tolist()\n    train_images = [os.path.join(train_path, i + \".png\") for i in train_images]\n    train_targets = df_train.target.values\n\n    valid_images = df_valid.image_name.values.tolist()\n    valid_images = [os.path.join(training_path, i + \".png\") for i in valid_images]\n    valid_targets = df_valid.target.values\n\n    train_dataset = train_aug.flow_from_directory(\n        image_paths=train_images,\n        targets=train_targets,\n        resize=None\n    )\n    \n    \n\n    valid_dataset = valid_aug.flow_from_directory(\n        image_paths=valid_images,\n        targets=valid_targets,\n        resize=None\n     )\n\n    optimizer = tensorflow.keras.optimizers.Adam(model.parameters(), lr=1e-4)\n    scheduler = tensorflow.keras.callbacks.ReduceLROnPlateau(\n        monitor = 'val_loss',\n        patience=3,\n        min_lr = 0.001,\n        mode=\"max\"\n    )\n    es = EarlyStopping(patience=5, mode=\"max\")\n\n    for epoch in range(epochs):\n        train_loss = Engine.train(train_dataset, model, optimizer)\n        predictions, valid_loss = Engine.evaluate(\n             valid_dataset, model\n        )\n        predictions = np.vstack((predictions)).ravel()\n        auc = metrics.roc_auc_score(valid_targets, predictions)\n        print(f\"Epoch = {epoch}, AUC = {auc}\")\n        scheduler.step(auc)\n\n        es(auc, model, model_path=f\"model_fold_{fold}.bin\")\n        if es.early_stop:\n            print(\"Early stopping\")\n            break\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# XGBoost","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import xgboost as xgb","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['sex'] = train_df['sex'].fillna('na')\ntrain_df['age_approx'] = train_df['age_approx'].fillna(0)\ntrain_df['anatom_site_general_challenge'] = train_df['anatom_site_general_challenge'].fillna('na')\n\ntest_df['sex'] = test_df['sex'].fillna('na')\ntest_df['age_approx'] = test_df['age_approx'].fillna(0)\ntest_df['anatom_site_general_challenge'] = test_df['anatom_site_general_challenge'].fillna('na')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['sex'] = train_df['sex'].astype(\"category\").cat.codes +1\ntrain_df['anatom_site_general_challenge'] = train_df['anatom_site_general_challenge'].astype(\"category\").cat.codes +1\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['sex'] = test_df['sex'].astype(\"category\").cat.codes +1\ntest_df['anatom_site_general_challenge'] = test_df['anatom_site_general_challenge'].astype(\"category\").cat.codes +1\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x_train = train_df[['sex', 'age_approx','anatom_site_general_challenge']]\ny_train = train_df['target']\n\n\nx_test = test_df[['sex', 'age_approx','anatom_site_general_challenge']]\n#y_test = test_df['target']\n\n\ntrain_DMatrix = xgb.DMatrix(x_train, label= y_train)\ntest_DMatrix = xgb.DMatrix(x_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = xgb.XGBClassifier()\nmodel.fit(x_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"y_pred = model.predict(x_test)\ny_pred","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Neural Networks","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Submissions","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"subs_df.target = model.predict_proba(x_test)[:,1]\nsub_tabular = subs_df.copy()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"subs_df.to_csv('submissions.csv', index = False)\nprint('Successfull!!')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"subs_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"References\n\n1. https://www.kaggle.com/redwankarimsony/power-of-metadata-xgboost-cnn-ensemble/comments","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}