{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Previous version\n\nIn version 30 of this notebook we submitted a solution that reached accuracy 87.3 in the public leaderboard. We used 299x299 pixels pic size and cut off the first 100 pixels on top of the pics to avoid the part of sky visible in 1% of the sample.\nIn this notebook we use another solution.[](http://)"},{"metadata":{},"cell_type":"markdown","source":"# Cassava competition\n\nIn this competition we aim at building a machine learning algorithm to classify different sort of Cassava plant diseases from the pictures of their sick leaves.\nCassava is an important provider of carbohydrates in Africa. Therefore, a fast method to identify its disease is of great importance in order to protect harvests. The data set is composed by images of healths and sick Cassava leaves, mostly taken by the farmers and classified by experts.\n\nIn particular, in the dataset we find pictures of four different sort of disease and pictures of healthy leaves. Let's do a first exploratory data analysis and visualise them.\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pdb\nimport os,re\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '../input/cassava-leaf-disease-classification'\ndf = pd.read_csv(path + '/train.csv')\ndf['path'] = path + '/train_images/' + df['image_id']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import json\n\ndisease_file_map = '../input/cassava-leaf-disease-classification/label_num_to_disease_map.json'\nf = open(disease_file_map)\nclasses_diseases = json.load(f)\nf.close()\n\nmap_classes ={int(i):j for i,j in classes_diseases.items()}\n\n#map_classes\n\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Pictures visualization\nLet's see some pictures for the different classes"},{"metadata":{"trusted":true},"cell_type":"code","source":"import cv2\ndef plot_images(df, boole, dict_labels):\n    df_class = df.loc[boole]\n    plt.figure(figsize=(15, 7))\n    for i, (idx, row) in enumerate(df_class[0:6].iterrows()):\n        image = cv2.imread(row['path'])\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        label = row['label']\n        disease = dict_labels[label]\n        #\n        ax = plt.subplot(2, 3, i + 1)\n        plt.imshow(image)\n        #print(image.shape)\n        title_str = str(label) + ':' + disease\n        plt.title(title_str)\n        plt.axis(\"off\")\n    plt.show()\n        ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Class 0: Cassava Bacterial Blight"},{"metadata":{"trusted":true},"cell_type":"code","source":"#select class 0\nplot_images(df, (df.label==0), map_classes)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Class 1: Cassava Brown Streak Disease"},{"metadata":{"trusted":true},"cell_type":"code","source":"#select class 1\nplot_images(df, (df.label==1), map_classes)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Note: some pictures show roots instead of leaves. We checked out eyeballing 2000 radom pics, and we found that only 6 pics represent roots. This fraction of outliers should not affect the training process, so no need to remove them."},{"metadata":{},"cell_type":"markdown","source":"### Class 2: Cassava Green Mottle"},{"metadata":{"trusted":true},"cell_type":"code","source":"#select class 2\nplot_images(df, (df.label==2), map_classes)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Class 3: Cassava Mosaic Disease"},{"metadata":{"trusted":true},"cell_type":"code","source":"#select class 3\nplot_images(df, (df.label==3), map_classes)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Class 4: Healthy"},{"metadata":{"trusted":true},"cell_type":"code","source":"#select class 4\nplot_images(df, (df.label==4), map_classes)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"#load necessary libraries\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import confusion_matrix, classification_report, accuracy_score\nfrom sklearn.utils import class_weight\n\nimport tensorflow as tf\nprint(tf.__version__)\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras import regularizers\nfrom tensorflow.keras.layers import Dense, Flatten, Conv2D, MaxPooling2D, GlobalAveragePooling2D, BatchNormalization, Dropout\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.preprocessing.image import load_img\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Are the classes balanced?\nLet's check how many pictures per class we have"},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(df['label'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It's clear that the data is not balanced and see class 3 having a larger number w.r.t other classes. However, after some tests we saw that this does not affect the results."},{"metadata":{},"cell_type":"markdown","source":"## Data split and data sets\nHere we split the data in training set, validation set, and test set. The latter will be used to evaluate the performance of the model on unseen data. The final data set to be predicted and submitted is called \"submission set\"."},{"metadata":{"trusted":true},"cell_type":"code","source":"#first split to get the test set\ndf_rest, df_test = train_test_split(df, test_size=0.1, random_state=2)\n#second split to get training and validation set\ndf_train, df_valid = train_test_split(df_rest, test_size=0.1, random_state=3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#here we prepared the class weights in case we want to use them\n#class_weights = class_weight.compute_class_weight('balanced', np.unique(df.label.values), df.label.values)\n#class_weights_dict = dict(zip(np.arange(5), class_weights))\n#class_weights_dict\n\n\n\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data sets and model preparation"},{"metadata":{"trusted":true},"cell_type":"code","source":"#prepare the tf.datasets\ntrain_ds = tf.data.Dataset.from_tensor_slices((df_train.path.values, df_train.label.values))\nvalid_ds = tf.data.Dataset.from_tensor_slices((df_valid.path.values, df_valid.label.values))\ntest_ds = tf.data.Dataset.from_tensor_slices((df_test.path.values, df_test.label.values))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#prepare the callbacks\ndef set_callbacks(path, optimizer):\n    #checkpoint path\n    checkpoint_path = path\n    metric = \"val_loss\"\n    #set checkpoint callback\n    checkpoint =tf.keras.callbacks.ModelCheckpoint(\n        filepath=checkpoint_path,\n        monitor=metric,\n        verbose=2,\n        save_best_only=True,\n        save_weights_only=True,\n        mode=\"auto\",\n        save_freq=\"epoch\"\n    )\n    #set reduce on plateau callback\n    ROP_callback = tf.keras.callbacks.ReduceLROnPlateau(\n        monitor=\"val_loss\",\n        factor=0.5,\n        patience=1,\n        min_delta=0.02,\n        min_lr=1e-6,\n        verbose=2\n    )\n    #set early stop\n    EarlyStop = tf.keras.callbacks.EarlyStopping(monitor = 'val_loss', min_delta = 0.001, \n                           patience = 6, mode = 'min', verbose = 2,\n                           restore_best_weights = True)\n    #set LR scheduler\n    def lr_function(epoch, lr):\n        #epoch = epoch_input//2\n        if epoch <= 3:\n            return lr * 0.8#tf.math.exp(0.8 * (-epoch))\n        else:\n            return np.max([1e-7,lr * 0.3])#tf.math.exp(0.5 * (3 - epoch))\n\n    LR_scheduler = tf.keras.callbacks.LearningRateScheduler(lr_function)\n    #set get LR value\n    def get_lr_metric(optimizer):\n        def lr(y_true, y_pred):\n            return optimizer.lr\n        return lr\n    lr_metric = get_lr_metric(optimizer)\n    ####\n    return checkpoint, ROP_callback, LR_scheduler, EarlyStop, lr_metric\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We decided (at the present version) to use frames of 512x512 pixels (unlike version 30)."},{"metadata":{"trusted":true},"cell_type":"code","source":"#set constant parameters\nSIZE_TARGET_X = 512\nSIZE_TARGET_Y = 512\nBATCH_SIZE =32\nAUTOTUNE = tf.data.experimental.AUTOTUNE\nINITIAL_LR = 0.005\nINITIAL_EPOCHS = 10\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here we istantiate the Xception model and upload the pre-trainied weights 'imagenet' that we added as input data. The top layer is left out."},{"metadata":{"trusted":true},"cell_type":"code","source":"path_weights = '../input/xception-weights/xception_weights_tf_dim_ordering_tf_kernels_notop.h5'\nbase_model = tf.keras.applications.xception.Xception(weights=path_weights, include_top=False)\nbase_model.summary()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Set the model and add a extra dense layer before the final top layer."},{"metadata":{"trusted":true},"cell_type":"code","source":"def set_model(base_model):\n    model = Sequential()\n    model.add(base_model)\n    model.add(GlobalAveragePooling2D())\n#    model.add(Flatten())\n    model.add(Dense(1024, activation = 'relu', kernel_regularizer='l2', kernel_initializer='he_uniform'))\n    model.add(BatchNormalization())\n    model.add(Dropout(0.5))\n    model.add(Dense(5, activation='softmax'))\n    return model\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data pre-processing and augmentation\nHere we define the functions for the pre-processing of the data sets. The augmentation of the training set is done by randomly cropping frames of 512x512 pixels from the original pictures of size 800x400 pixels and doing horizontal and vertical flip. The validation set undergo to the same augmentation process because this will be also applied to the final test and submission data sets.\n\n### Submission and test sets augmentation\nFew words about the treatment of the test and submission test sets. Because the crop with 512x512 pixels does not cover the whole picture, we decided to apply the preprocessing augmentation functions to the test and submission pictures and evaluate each pic for several times and then average the results. This method (borrowed from [this](https://www.kaggle.com/harveenchadha/effnetb4-tf-data-gpu-aug-5x-speedup-tta/data#Mixed-Precision-Training) notebook) reders better results. The same is applied to the submission set."},{"metadata":{"trusted":true},"cell_type":"code","source":"#####\ndef read_images(img_path,label):\n    img = tf.io.read_file(img_path)\n    img = tf.io.decode_jpeg(img, channels=3)\n    img = tf.cast(img, tf.float32)\n#    img = tf.image.crop_to_bounding_box(img, 200, 100, 400, 700)\n#    img = tf.image.adjust_hue(img, 0.8)\n    return img, label\n#####\ndef preprocess_train(img_path, label):\n    img, label = read_images(img_path, label)\n    img = tf.image.random_flip_left_right(img)\n    img = tf.image.random_flip_up_down(img)\n    img = tf.image.random_crop(img, size=[SIZE_TARGET_X,SIZE_TARGET_Y,3])\n    img = tf.image.random_brightness(img, max_delta=0.5)\n#    img = tf.image.random_saturation(img, 0.8, 1.2, seed=None)\n#    img = tf.image.random_contrast(img, 0.7, 1.)\n\n    img = img / 255.\n\n    return img, label\n#####\ndef preprocess_valid_test(img_path, label):\n    img, label = read_images(img_path, label)\n    img = tf.image.random_crop(img, size=[SIZE_TARGET_X,SIZE_TARGET_Y,3])\n#    img = tf.image.crop_to_bounding_box(img, 0, 0, SIZE_TARGET_X, SIZE_TARGET_Y)\n    img = img / 255.\n\n    return img, label\n#####\n\n\ntrain_ds_prep = (\n    train_ds\n    .shuffle(1000)\n    .map(preprocess_train, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n#\nvalid_ds_prep = (\n    valid_ds\n    .map(preprocess_valid_test, num_parallel_calls=AUTOTUNE)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTOTUNE)\n)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Compile and run the model\nFirst we train the model by freezing all the Xception model layers and setr as trainable only the extra dense layers added. Later, we will refine the training by setting as trainable some Xception layers too."},{"metadata":{"trusted":true},"cell_type":"code","source":"base_model.trainable = False\n\nmodel = set_model(base_model)\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"path_initial_weights = '/kaggle/working/weights/best_weights_inital'\n#compile the model\noptim = tf.optimizers.Adam(lr=INITIAL_LR)\n#set checkpoints\ncheckpoint, ROP_callback, LR_scheduler, EarlyStop, lr_metric = set_callbacks(path_initial_weights, optim)\n#optim = tf.keras.optimizers.SGD(learning_rate=INITIAL_LR, momentum=0.9, nesterov=True)\nmodel.compile(optimizer=optim, loss='sparse_categorical_crossentropy', metrics=['sparse_categorical_accuracy', lr_metric])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's train the model by using callbacks like early stop and \"Reduce On Plateau\" previously defined. The weights of the best found model are stored by the checkpoint callback as \"best_weights_initial\"."},{"metadata":{"trusted":true},"cell_type":"code","source":"history = model.fit(\n        train_ds_prep,\n        epochs=INITIAL_EPOCHS,\n        validation_data=valid_ds_prep,\n        batch_size=BATCH_SIZE,\n#        class_weight = class_weights_dict,\n        callbacks=[checkpoint, ROP_callback, EarlyStop])\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here we check the loss and accuracy curves."},{"metadata":{"trusted":true},"cell_type":"code","source":"#define a function to plot learning curves\ndef plot_loss_acc(history):\n    fig = plt.figure(figsize=(12, 6))\n\n    fig.add_subplot(121)\n\n    plt.plot(history.history['loss'])\n    plt.plot(history.history['val_loss'])\n    plt.title('loss vs. epochs')\n    plt.ylabel('Loss')\n    plt.xlabel('Epoch')\n    plt.legend(['Training', 'Validation'], loc='upper right')\n\n    fig.add_subplot(122)\n\n    plt.plot(history.history['sparse_categorical_accuracy'])\n    plt.plot(history.history['val_sparse_categorical_accuracy'])\n    plt.title('accuracy vs. epochs')\n    plt.ylabel('accuracy')\n    plt.xlabel('Epoch')\n    plt.legend(['Training', 'Validation'], loc='upper right')\n\n    plt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# plot learning curves\nplot_loss_acc(history)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The two curves shows some improvements. We stop the training early because we do not need to reach the top of the performance now. The model will be further trained later by using also the Xception layers. Here we only want to train a little the extra dense layers so that the backward gradients are not too high.\n\nNow we upload the previous found best weights and leave trainable the last 100 Xception layers."},{"metadata":{"trusted":true},"cell_type":"code","source":"\nbase_model.trainable = True\n\n## set as trainable the last 10 layers\nfor layer in base_model.layers[0:-100]:\n    layer.trainable = False\n\nmodel = set_model(base_model)\nmodel.load_weights(path_initial_weights)\n\n\n\n#set checkpoints\npath_final_weights = '/kaggle/working/weights/best_weights_final'\noptim = tf.optimizers.Adam(lr=INITIAL_LR/10.)\ncheckpoint, ROP_callback, LR_scheduler, EarlyStop, lr_metric = set_callbacks(path_final_weights, optim)\n#compile the model\nmodel.compile(optimizer=optim, loss='sparse_categorical_crossentropy', metrics=['sparse_categorical_accuracy', lr_metric])\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Train again the model..."},{"metadata":{"trusted":true},"cell_type":"code","source":"FINE_TUNE_EPOCHS = 30\nTOT_EPOCHS =  INITIAL_EPOCHS + FINE_TUNE_EPOCHS\n\nhistory_fine = model.fit(\n        train_ds_prep,\n        epochs=TOT_EPOCHS,\n        batch_size=BATCH_SIZE,\n        initial_epoch=history.epoch[-1],\n        validation_data=valid_ds_prep,\n#        class_weight = class_weights_dict,\n        callbacks=[checkpoint, ROP_callback, EarlyStop])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"...and check the loss and accuracy curves for this second training session"},{"metadata":{"trusted":true},"cell_type":"code","source":"# plot learning curves\nplot_loss_acc(history_fine)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Performance evaluation\nNow upload the best model weights and evaluate the performance on the test set that the model has never seen before."},{"metadata":{"trusted":true},"cell_type":"code","source":"#upload the best weights found\nmodel = set_model(base_model)\nmodel.load_weights(path_final_weights)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def predict_mean(data_set, n_times):\n    list_predict = []\n    #repeat the prediction n_times\n    for i in range(n_times):\n        #prepare the tf.dataset by pre-processing the data as done for the trainig set\n        #i.e. random cropping and flipping the pic\n        ds_prep = data_set.map(preprocess_train, num_parallel_calls=AUTOTUNE).batch(BATCH_SIZE)\n        #append the prediction to a list\n        pred = model.predict(ds_prep, workers=4, verbose=True)\n        list_predict.append(pred)\n    \n    #now average the predicted probabilities\n    predict_mean = np.mean(list_predict, axis=0)\n\n    #produce the final predited classes\n    predict_final = [np.argmax(i) for i in predict_mean]\n    predict_last = [np.argmax(i) for i in pred]\n    return predict_final, predict_last\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#compute the accuracy on the test set\nvalid_labels = df_test.label.values\n#the following row compute for 10 times the prediction on 10 different data sets with\n#generated with \"preprocess_train\" function\npredict_final, predict_last = predict_mean(test_ds, 10)\n#show one (last) accuracy\nprint('last accuracy score:{:.3f}'.format(accuracy_score(valid_labels, predict_last)))\n#show accuracy\nprint('mean accuracy score:{:.3f}'.format(accuracy_score(valid_labels, predict_final)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Visualize the confusion matrix..."},{"metadata":{"trusted":true},"cell_type":"code","source":"cfm = confusion_matrix(predict_final, valid_labels)\nax = sns.heatmap(cfm, annot=True, fmt=\"d\", cmap='coolwarm')\nax.set(xlabel='true labels', ylabel='predited label')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"...and print the classification report"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(classification_report(valid_labels, predict_final))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Submission"},{"metadata":{"trusted":true},"cell_type":"code","source":"#upload the submission file\ndf_submission = pd.read_csv(path + '/sample_submission.csv')\n#add a column with the path of the images\ndf_submission['path'] = path + '/test_images/' + df_submission['image_id']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#prepare the tf.dataset\nsubmission_ds = tf.data.Dataset.from_tensor_slices((df_submission.path.values, df_submission.label.values))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#predict the classes and put it into the 'label' column\npredict_submission, predict_last = predict_mean(submission_ds, 10)\ndf_submission['label'] = predict_submission","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#drop the column 'path'\ndf_submission.drop(columns='path', inplace=True)\n#write\ndf_submission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Acknowledgements\nI went through many good notebooks from which I have learned a lot. In particular I would like to cite some of them that inspired me and from which I borrowed some ideas. Thank you all!\n\n[Notebook by Ece Şimşek](https://www.kaggle.com/eceifter/xception-cassava-leaf-disease-classification/output)\n\n[Notebook by Harveen Singh Chadha](https://www.kaggle.com/harveenchadha/effnetb4-tf-data-gpu-aug-5x-speedup-tta/data#Mixed-Precision-Training)\n\n[Notebook by Maksym Shkliarevskyi](https://www.kaggle.com/maksymshkliarevskyi/cassava-leaf-disease-best-keras-cnn/notebook)"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}