{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Le mélanome de la peau\n\n\n## Introduction\n\nUn mélanome de la peau est une maladie des cellules de la peau appelées mélanocytes. Il se développe à partir d’une cellule initialement normale qui se transforme et se multiplie de façon anarchique pour former une lésion appelée tumeur maligne.\n\nIl existe quatre principaux types de mélanome de la peau : le mélanome superficiel extensif, le mélanome de Dubreuilh, le mélanome nodulaire et le mélanome acrolentigineux. Le traitement de ces différents types de mélanome repose essentiellement sur la chirurgie qui sera adaptée à la topographie (c’est-à-dire à l’endroit où est situé le mélanome) et à la profondeur de la lésion."},{"metadata":{},"cell_type":"markdown","source":"## Formalisation des objectifs"},{"metadata":{},"cell_type":"markdown","source":"## Acquisition des données\n"},{"metadata":{},"cell_type":"markdown","source":"### Comprendre les données\n\nDans cette section, nous allons récupérer les données et étudier leurs distributions. Fait intéressant que les données contiennent à la fois des données tabulaires et des données d'image.\n\nNous continuerons en chargeant les métadonnées qui nous sont fournies. Les données de train ont 8 caractéristiques et 33126 observations,les données de test ont 5 caractéristiques et 10982 observations.\n\nL'ensemble de données de train se compose de:\n\n* nom de l'image -> le nom de fichier de l'image spécifique pour la rame\n* patient_id -> identifie le patient unique\n* sexe -> sexe du patient\n* age_approx -> âge approximatif du patient au moment de la numérisation\n* anatom_site_general_challenge -> emplacement du site d'analyse\n* diagnostic -> informations sur le diagnostic\n* benign_malignant - indique le résultat du scan s'il est malin ou bénin\n* target -> même que ci-dessus mais mieux pour la modélisation car c'est binaire"},{"metadata":{},"cell_type":"markdown","source":"### Analyser les données\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Setting color palette.\ncolor_black = [\n    '#333300','#666600','#999900','#CCCC00','#FFFF00','#FFFF33','#FFFF66','#FFFF99','#FFFFCC'\n]\n\ncolor_black_2 = [\n    '#FFD700','#BDB76B','#F0E68C','#EEE8AA','#FFEFD5','#FFDAB9','#330C73'\n]\n# Setting color palette.\norange_black = [\n    '#fdc029', '#df861d', '#FF6347', '#aa3d01', '#a30e15', '#800000', '#171820'\n]\ncolor_black_3=['#318ce7','#ff66cc']\nB = ['#C3073F','#1A1A1D']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"#Télécharger la bibliothèque nécessaire\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport pylab as pl\nfrom IPython import display\nimport seaborn as sns\nsns.set()\nimport re\nimport pydicom\nimport random\nimport torch\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms, models\nimport pandas as pd\n\n#import pytorch_lightning as pl\nfrom scipy.special import softmax\nfrom sklearn.model_selection import train_test_split, StratifiedKFold\nfrom sklearn.utils.class_weight import compute_class_weight\nfrom sklearn.metrics import roc_auc_score, auc\n\nfrom skimage.io import imread\nfrom PIL import Image\nimport plotly.express as px\n\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nfrom plotly.subplots import make_subplots\n\nimport os\nimport copy\n\nfrom albumentations import Compose, RandomCrop, Normalize,HorizontalFlip, Resize\nfrom albumentations import VerticalFlip, RGBShift, RandomBrightness\nfrom albumentations.core.transforms_interface import ImageOnlyTransform\nfrom albumentations.pytorch import ToTensor\n\nfrom tqdm.notebook import tqdm\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%matplotlib inline\nimport tensorflow as tf\nimport plotly.express as px\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm.notebook import tqdm\nfrom kaggle_datasets import KaggleDatasets\nfrom collections import Counter\nimport efficientnet.tfkeras as efn\nimport re\nfrom tensorflow.keras import layers as L\nimport sklearn\n\nsns.set_style(\"dark\")\nsns.set(rc={'figure.figsize':(12,8)})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"try:\n    # TPU detection. No parameters necessary if TPU_NAME environment variable is\n    # set: this is always the case on Kaggle.\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    # Default distribution strategy in Tensorflow. Works on CPU and single GPU.\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install -q efficientnet","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Reading the dataset\ndataset = pd.read_csv(\"../input/siim-isic-melanoma-classification/train.csv\")\ndataset","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data Exploration & Visualisations"},{"metadata":{},"cell_type":"markdown","source":"## Data Exploration"},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset.nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data Visualisations"},{"metadata":{"trusted":true},"cell_type":"code","source":"f, axes = plt.subplots(2, 2, figsize=(12,12))\nf.tight_layout() \nplt.subplots_adjust(left=0.01, wspace=0.6, hspace=0.4)\nsns.countplot(y=\"anatom_site_general_challenge\", data=dataset,  ax=axes[0][1])\nsns.countplot(y=\"diagnosis\", data=dataset,  ax=axes[0][0])\nsns.countplot(x='sex', data=dataset, ax=axes[1][0])\nsns.countplot(\"benign_malignant\", data=dataset,  ax=axes[1][1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.distplot(dataset['age_approx'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(\"target\", data=dataset)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset['target'].value_counts(normalize=True) * 10","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"diag = dataset.diagnosis.value_counts()\nfig = px.pie(diag,\n             values='diagnosis',\n             names=diag.index,\n             color_discrete_sequence=orange_black,\n             hole=.4)\nfig.update_traces(textinfo='percent+label', pull=0.05)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(16, 6))\nsns.boxplot(x='benign_malignant',\n            y='age_approx',\n            data=dataset,\n            hue='sex',\n            palette=color_black_2)\n\nplt.legend(loc='lower right')\n\nax.set_title('Scan Results by Age and Sex')\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> **IMBALANCED DATASET!**"},{"metadata":{},"cell_type":"markdown","source":"# Images Visualisations"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Showing a sample image\nimage = plt.imread('/kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_5766923.jpg')\nplt.imshow(image)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Benign Images"},{"metadata":{"trusted":true},"cell_type":"code","source":"w = 10\nh = 10\nfig = plt.figure(figsize=(15, 15))\ncolumns = 4\nrows = 4\n\n# ax permet d'accéder à la manipulation de chacun des sous-graphiques\nax = []\n\nfor i in range(columns*rows):\n    img = plt.imread('/kaggle/input/siim-isic-melanoma-classification/jpeg/train/'+dataset['image_name'][i]+'.jpg')\n    # créer un subplot et l'ajouter a ax\n    ax.append( fig.add_subplot(rows, columns, i+1) )\n    # Hide grid lines\n    ax[-1].grid(False)\n\n    # masquer axes ticks\n    ax[-1].set_xticks([])\n    ax[-1].set_yticks([])\n    ax[-1].set_title(dataset['benign_malignant'][i])  # set title\n    plt.imshow(img)\n\n\n\nplt.show()  # finally, render the plot","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"w = 10\nh = 10\nfig = plt.figure(figsize=(15, 15))\ncolumns = 4\nrows = 4\n\nax = []\n\nfor i in range(columns*rows):\n    img = plt.imread('/kaggle/input/siim-isic-melanoma-classification/jpeg/train/'+dataset.loc[dataset['target'] == 1]['image_name'].values[i]+'.jpg')\n    ax.append( fig.add_subplot(rows, columns, i+1) )\n    \n    ax[-1].grid(False)\n    ax[-1].set_xticks([])\n    ax[-1].set_yticks([])\n    ax[-1].set_title(dataset.loc[dataset['target'] == 1]['benign_malignant'].values[i])  # set title\n    plt.imshow(img)\n\n\n\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Cleaning Dataset"},{"metadata":{},"cell_type":"markdown","source":"## Removing NaN values"},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset.loc[dataset.isnull().any(axis=1)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* * Dans notre cas, on va supprimer les lignes avec des valeurs nan,"},{"metadata":{"trusted":true},"cell_type":"code","source":"dataset = dataset.dropna(axis=0)\ndataset.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Changing all column into categorical"},{"metadata":{"trusted":true},"cell_type":"code","source":"cleaned_dataset = dataset.copy()\ncleaned_dataset","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#La conversion d'une telle variable chaîne en une variable catégorielle économisera de la mémoire\ncleaned_dataset.sex = cleaned_dataset.sex.replace({'male':0, 'female':1})\ncleaned_dataset = cleaned_dataset.join(pd.get_dummies(cleaned_dataset.anatom_site_general_challenge))\ncleaned_dataset = cleaned_dataset.join(pd.get_dummies(cleaned_dataset.diagnosis))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.options.display.max_rows = 999\ncleaned_dataset = cleaned_dataset.reset_index()\ncleaned_dataset.head(35)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Creating the Training & Testing Dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"# For tf.dataset\nAUTO = tf.data.experimental.AUTOTUNE\n\n# Data access\nGCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/train*.tfrec')\nTEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/tfrecords/test*.tfrec')\n\nCLASSES = [0,1]   \nIMAGE_SIZE = [1024, 1024]\nBATCH_SIZE = 8 * strategy.num_replicas_in_sync","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def decode_image(image_data):\n    image = tf.image.decode_jpeg(image_data, channels=3)\n    image = tf.cast(image, tf.float32) / 255.0  # convertir image en floats entre[0, 1] range\n    image = tf.reshape(image, [*IMAGE_SIZE, 3]) # taille explicite de TPU\n    return image\n\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n       \n        \"target\": tf.io.FixedLenFeature([], tf.int64),  # shape [] means single element\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    label = tf.cast(example['target'], tf.int32)\n    \n    return image, label # returns a dataset of (image, label) pairs\n\ndef read_unlabeled_tfrecord(example):\n    UNLABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), # tf.string means bytestring\n        \"image_name\": tf.io.FixedLenFeature([], tf.string),  # shape [] means single element\n        # class is missing, this competitions's challenge is to predict flower classes for the test dataset\n    }\n    example = tf.io.parse_single_example(example, UNLABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    idnum = example['image_name']\n    return image, idnum # returns a dataset of image(s)\n\ndef load_dataset(filenames, labeled=True, ordered=False):\n   \n    ignore_order = tf.data.Options()\n    if not ordered:\n        ignore_order.experimental_deterministic = False \n\n    dataset = tf.data.TFRecordDataset(filenames, num_parallel_reads=AUTO) \n    dataset = dataset.with_options(ignore_order) \n    dataset = dataset.map(read_labeled_tfrecord if labeled else read_unlabeled_tfrecord, num_parallel_calls=AUTO)\n    return dataset\n\ndef data_augment(image, label):\n   \n    image = tf.image.random_flip_left_right(image)\n    image = tf.image.random_brightness(image, 0.1)\n    image = tf.image.random_flip_up_down(image)\n    \n    return image, label   \n\ndef get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() #elle doit se repeter pour plusieurs epochs\n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # \n    return dataset\n\ndef get_validation_dataset(ordered=False):\n    dataset = load_dataset(VALIDATION_FILENAMES, labeled=True, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.cache()\n    dataset = dataset.prefetch(AUTO) \n    return dataset\n\ndef get_test_dataset(ordered=False):\n    dataset = load_dataset(TEST_FILENAMES, labeled=False, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) # prélecture du lot suivant pendant l'entraînement (taille du tampon de prélecture automatique)\n    return dataset\n\ndef count_data_items(filenames):\n    n = [int(re.compile(r\"-([0-9]*)\\.\").search(filename).group(1)) for filename in filenames]\n    return np.sum(n)\n\nNUM_TRAINING_IMAGES = count_data_items(TRAINING_FILENAMES)\nNUM_TEST_IMAGES = count_data_items(TEST_FILENAMES)\nSTEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\nprint('Dataset: {} training images and {} unlabeled test images'.format(NUM_TRAINING_IMAGES,NUM_TEST_IMAGES))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"STEPS_PER_EPOCH = NUM_TRAINING_IMAGES // BATCH_SIZE\nEPOCHS = 5","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def build_lrfn(lr_start=0.00001, lr_max=0.0001, \n               lr_min=0.000001, lr_rampup_epochs=20, \n               lr_sustain_epochs=0, lr_exp_decay=.8):\n    lr_max = lr_max * strategy.num_replicas_in_sync\n\n    def lrfn(epoch):\n        if epoch < lr_rampup_epochs:\n            lr = (lr_max - lr_start) / lr_rampup_epochs * epoch + lr_start\n        elif epoch < lr_rampup_epochs + lr_sustain_epochs:\n            lr = lr_max\n        else:\n            lr = (lr_max - lr_min) * lr_exp_decay**(epoch - lr_rampup_epochs - lr_sustain_epochs) + lr_min\n        return lr\n    \n    return lrfn","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"with strategy.scope():\n    efficientnetb5_model = tf.keras.Sequential([\n        efn.EfficientNetB5(\n            input_shape=(*IMAGE_SIZE, 3),\n            #weights='imagenet',\n            weights='imagenet',\n            include_top=False\n        ),\n        L.GlobalAveragePooling2D(),\n        L.Dense(1024, activation = 'relu'), \n        L.Dropout(0.3), \n        L.Dense(512, activation= 'relu'), \n        L.Dropout(0.2), \n        L.Dense(256, activation='relu'), \n        L.Dropout(0.2), \n        L.Dense(128, activation='relu'), \n        L.Dropout(0.1), \n        L.Dense(1, activation='sigmoid')\n    ])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from tensorflow.keras import backend as K\n\n\n\ndef focal_loss(gamma=2., alpha=.25):\n\tdef focal_loss_fixed(y_true, y_pred):\n\t\tpt_1 = tf.where(tf.equal(y_true, 1), y_pred, tf.ones_like(y_pred))\n\t\tpt_0 = tf.where(tf.equal(y_true, 0), y_pred, tf.zeros_like(y_pred))\n\t\treturn -K.mean(alpha * K.pow(1. - pt_1, gamma) * K.log(pt_1)) - K.mean((1 - alpha) * K.pow(pt_0, gamma) * K.log(1. - pt_0))\n\treturn focal_loss_fixed","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"efficientnetb5_model.compile(\n    optimizer='adam',\n    loss = focal_loss(gamma=2., alpha=.25),\n    metrics=['binary_crossentropy', 'accuracy']\n)\nefficientnetb5_model.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lrfn = build_lrfn()\nlr_schedule = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#model.load_weights('../input/melenoma/model_weights.h5')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history = efficientnetb5_model.fit(\n    get_training_dataset(), \n    epochs=EPOCHS, \n    steps_per_epoch=STEPS_PER_EPOCH,\n    callbacks=[lr_schedule],\n    class_weight = {0:0.50899675,1: 28.28782609}\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# résumer l'historique pour plus de précision\nplt.plot(history.history['loss'])\nplt.title('model loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'val'], loc='upper left')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# résumer l'historique des pertes\n#mesure les performances d'un modèle de classification dont la sortie est une valeur de probabilité entre 0 et 1\nplt.plot(history.history['binary_crossentropy'])\nplt.title('model crossentropy')\nplt.ylabel('crossentropy')\nplt.xlabel('epoch')\nplt.legend(['train', 'val'], loc='upper left')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"efficientnetb5_model.save('complete_data_efficient_model.h5')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"efficientnetb5_model.save_weights('complete_data_efficient_weights.h5')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Generating the Predictions"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_ds = get_test_dataset(ordered=True)\ntest_images_ds = test_ds.map(lambda image, idnum: image)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"probabilities = efficientnetb5_model.predict(test_images_ds)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Generating submission.csv file...')\ntest_ids_ds = test_ds.map(lambda image, idnum: idnum).unbatch()\ntest_ids = next(iter(test_ids_ds.batch(NUM_TEST_IMAGES))).numpy().astype('U') # all in one batch","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_df = pd.DataFrame({'image_name': test_ids, 'target': np.concatenate(probabilities)})\npred_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub = pd.read_csv(\"/kaggle/input/siim-isic-melanoma-classification/sample_submission.csv\")\nsub","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del sub['target']\nsub = sub.merge(pred_df, on='image_name')\n#sub.to_csv('submission_label_smoothing.csv', index=False)\nsub.to_csv('complete_data.csv', index=False)\nsub.head()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}