{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# EfficientNetV2B0\n## Trained on APTOS19 + EyePACS15 ; Validation on Messidor-2 for generalization\n\n#### Quadratic Weighted Kappa (QWK):\nQuadratic Weighted Kappa (QWK, the greek letter $\\kappa$), also known as Cohen's Kappa, is the official evaluation metric in both 2015 and 2019 competitions. \nAccording to the [wikipedia article](https://en.wikipedia.org/wiki/Cohen%27s_kappa), we have\n> The definition of $\\kappa$ is:\n> $$\\kappa \\equiv \\frac{p_o - p_e}{1 - p_e}$$\n> where $p_o$ is the relative observed agreement among raters (identical to accuracy), and $p_e$ is the hypothetical probability of chance agreement, using the observed data to calculate the probabilities of each observer randomly seeing each category.\n\nThis metric is used because if we just using accuracy as the metric, it will give spurious results (because the data is unbalanced). The QWK is more fit to the problem.\n\n#### Ordinal Regression (OR):\n> In DR classification, the disease classes are not independent, meaning that the severity of the disease is not necessarily limited to a specific set of discrete classes. Therefore, an ordinal regression model is more appropriate because it can predict a continuous severity score, allowing for a more precise prediction of disease severity.\n---\n\nWe treat this as an Ordinal Regression Problem based on this litterature: \n- Mean absolute error as training metric instead of accuracy; \n> For an ordinal regression model, accuracy is not an appropriate metric for ordinal regression because it only measures the percentage of correctly classified samples and does not take into account the ordinal structure of the output. Instead, you can use metrics such as mean absolute error (MAE) or mean squared error (MSE) of the predicted scores compared to the true scores. \n- Mean Squared Error as loss function instead of Cross Entropy;  [This](https://ieeexplore.ieee.org/iel7/9342628/9342629/09342711.pdf) \nand [this (Presentation video)](https://www.youtube.com/watch?v=IvhzwhjFyFQ)\n\n\n#### Note#1: \nUnlike Binary Regression, we cannot use ROC, AUC, Spe, Sen nor CM :(\nTherefore we will attempt to use these plots to compare the models: \n- Learning curve: training and validation loss (e.g., MAE or MSE) as a function of the number of epochs. By comparing the learning curves of your models, you can assess which one is better at minimizing the loss function and which one is less prone to overfitting.\n- Scatter plot: A scatter plot shows the predicted severity level vs. the actual severity level for each sample in the validation set. By comparing the scatter plots of your models, you can assess which one has better predictions and how consistent the predictions are across the severity levels.\n- Box plot: A box plot shows the distribution of the predicted severity levels across different severity levels in the validation set. By comparing the box plots of your models, you can assess how well each model predicts the different severity levels and whether there are any systematic biases in the predictions.\n- Cumulative distribution function (CDF) plot: A CDF plot shows the cumulative distribution of the predicted severity levels across different severity levels in the validation set. By comparing the CDF plots of your models, you can assess how well each model predicts the different severity levels and how the predictions are distributed across the severity levels.\n\n#### Note#2:\nAlthough ROC curves are not appropriate for ordinal regression, you can use Cumulative Net Reclassification Improvement (NRI) or Receiver Operating Characteristic Integrated with Decision Analytic (ROC-DA) curves that are appropriate for ordinal regression. These curves allow you to evaluate how well each model performs at classifying samples across different severity levels.\n\n### Credits\n#### This series of notebooks is heavily inspired by FEDERICO RAIMONDI *brilliant* kernels (DenseNet and EfficientNet).\nPlease check his [Inference Kernel](https://www.kaggle.com/raimonds1993/aptos19-densenet-inference-old-new-data/data?scriptVersionId=17252732)!\n\nThank you!","metadata":{"jupyter":{"source_hidden":true}}},{"cell_type":"markdown","source":"#  EfficientNetV2-S Notebook \n","metadata":{}},{"cell_type":"code","source":"# To have reproducible results and compare them\nnr_seed = 2019\nimport numpy as np \nnp.random.seed(nr_seed)\nimport tensorflow as tf\nimport tensorflow_hub as hub\nfrom tensorflow.keras import layers, Sequential\nfrom tensorflow.keras.utils import plot_model\ntf.random.set_seed(nr_seed)","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:02:16.768650Z","iopub.execute_input":"2023-03-25T15:02:16.769077Z","iopub.status.idle":"2023-03-25T15:02:27.250357Z","shell.execute_reply.started":"2023-03-25T15:02:16.769028Z","shell.execute_reply":"2023-03-25T15:02:27.249061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import libraries\nimport json\nimport math\nfrom tqdm import tqdm, tqdm_notebook\nimport gc\nimport warnings\nimport os\n\nimport cv2\nfrom PIL import Image\n\nimport pandas as pd\nimport scipy\nimport matplotlib.pyplot as plt\n\nfrom keras import backend as K\nfrom keras import layers\nfrom keras.applications import EfficientNetV2S\nfrom keras.callbacks import Callback, ModelCheckpoint\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.models import Sequential\nfrom keras.optimizers import Adam\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import cohen_kappa_score, accuracy_score\n\nwarnings.filterwarnings(\"ignore\")\n\n%matplotlib inline","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T15:02:27.252706Z","iopub.execute_input":"2023-03-25T15:02:27.253442Z","iopub.status.idle":"2023-03-25T15:02:28.126443Z","shell.execute_reply.started":"2023-03-25T15:02:27.253401Z","shell.execute_reply":"2023-03-25T15:02:28.125049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Image size\nim_size = 224\n# Batch size\nBATCH_SIZE = 32","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:02:41.065306Z","iopub.execute_input":"2023-03-25T15:02:41.065774Z","iopub.status.idle":"2023-03-25T15:02:41.071816Z","shell.execute_reply.started":"2023-03-25T15:02:41.065721Z","shell.execute_reply":"2023-03-25T15:02:41.070338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading & Merging","metadata":{}},{"cell_type":"code","source":"new_train = pd.read_csv('../input/aptos2019-blindness-detection/train.csv')\nold_train = pd.read_csv('../input/diabetic-retinopathy-resized/trainLabels.csv')\n#print(new_train.shape)\n#print(old_train.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:02:52.244114Z","iopub.execute_input":"2023-03-25T15:02:52.244544Z","iopub.status.idle":"2023-03-25T15:02:52.296781Z","shell.execute_reply.started":"2023-03-25T15:02:52.244501Z","shell.execute_reply":"2023-03-25T15:02:52.295436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"old_train = old_train[['image','level']]\nold_train.columns = new_train.columns\nold_train.diagnosis.value_counts()\n\n# path columns\nnew_train['id_code'] = '../input/aptos2019-blindness-detection/train_images/' + new_train['id_code'].astype(str) + '.png'\nold_train['id_code'] = '../input/diabetic-retinopathy-resized/resized_train/resized_train/' + old_train['id_code'].astype(str) + '.jpeg'\n\ntrain_df = old_train.copy()\nval_df = new_train.copy()\n#train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:02:57.637002Z","iopub.execute_input":"2023-03-25T15:02:57.637448Z","iopub.status.idle":"2023-03-25T15:02:57.678635Z","shell.execute_reply.started":"2023-03-25T15:02:57.637406Z","shell.execute_reply":"2023-03-25T15:02:57.677326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's shuffle the datasets\ntrain_df = train_df.sample(frac=1).reset_index(drop=True)\nval_df = val_df.sample(frac=1).reset_index(drop=True)\n#print(train_df.shape)\n#print(val_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:03:00.583018Z","iopub.execute_input":"2023-03-25T15:03:00.583955Z","iopub.status.idle":"2023-03-25T15:03:00.601143Z","shell.execute_reply.started":"2023-03-25T15:03:00.583901Z","shell.execute_reply":"2023-03-25T15:03:00.599547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train - Valid split","metadata":{}},{"cell_type":"code","source":"# Not used in version 5 train_df, val_df = train_test_split(train_df, shuffle=True, stratify=train_df.diagnosis, test_size=0.2, random_state=2019)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T15:59:51.994098Z","iopub.execute_input":"2023-03-07T15:59:51.994638Z","iopub.status.idle":"2023-03-07T15:59:51.998522Z","shell.execute_reply.started":"2023-03-07T15:59:51.994446Z","shell.execute_reply":"2023-03-07T15:59:51.997603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_samples(df, columns=4, rows=3):\n    fig=plt.figure(figsize=(5*columns, 4*rows))\n\n    for i in range(columns*rows):\n        image_path = df.loc[i,'id_code']\n        image_id = df.loc[i,'diagnosis']\n        img = cv2.imread(f'{image_path}')\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n        #img = crop_image_from_gray(img)\n        img = cv2.resize(img, (im_size,im_size))\n        img = cv2.addWeighted(img,4,cv2.GaussianBlur(img, (0,0), im_size/40) ,-4 ,128)\n        \n        fig.add_subplot(rows, columns, i+1)\n        plt.title(image_id)\n        plt.imshow(img)\n    \n    plt.tight_layout()\n\ndisplay_samples(train_df)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T15:03:12.260085Z","iopub.execute_input":"2023-03-25T15:03:12.260532Z","iopub.status.idle":"2023-03-25T15:03:15.799340Z","shell.execute_reply.started":"2023-03-25T15:03:12.260494Z","shell.execute_reply":"2023-03-25T15:03:15.797558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Processing Images","metadata":{}},{"cell_type":"markdown","source":"Crop function: https://www.kaggle.com/ratthachat/aptos-updated-preprocessing-ben-s-cropping ","metadata":{}},{"cell_type":"code","source":"def crop_image1(img,tol=7):\n    #This function takes an input image img and a tolerance value tol. \n    #It creates a boolean mask where pixel values greater than tol are set to True, \n    #and all other pixel values are set to False.\n    #It returns the cropped image where rows and columns with all False values are removed. \n        \n    mask = img>tol\n    return img[np.ix_(mask.any(1),mask.any(0))]\n\ndef crop_image_from_gray(img,tol=7):\n    #crops an RGB image based on a grayscale version of the image and a tolerance value.\n    if img.ndim ==2:\n        mask = img>tol\n        return img[np.ix_(mask.any(1),mask.any(0))]\n    elif img.ndim==3:\n        gray_img = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n        mask = gray_img>tol\n        \n        check_shape = img[:,:,0][np.ix_(mask.any(1),mask.any(0))].shape[0]\n        if (check_shape == 0): # image is too dark so that we crop out everything,\n            return img # return original image\n        else:\n            img1=img[:,:,0][np.ix_(mask.any(1),mask.any(0))]\n            img2=img[:,:,1][np.ix_(mask.any(1),mask.any(0))]\n            img3=img[:,:,2][np.ix_(mask.any(1),mask.any(0))]\n            img = np.stack([img1,img2,img3],axis=-1)\n        return img\n\ndef preprocess_image(image_path, desired_size=224):\n    #reads, crops, resizes, and applies a weighted sum and Gaussian blur to an input image.\n    img = cv2.imread(image_path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    img = crop_image_from_gray(img)\n    img = cv2.resize(img, (desired_size,desired_size))\n    img = cv2.addWeighted(img,4,cv2.GaussianBlur(img, (0,0), desired_size/30) ,-4 ,128)\n    \n    return img\n\ndef preprocess_image_old(image_path, desired_size=224):\n    # reads, resizes, and applies a weighted sum and Gaussian blur to an input image.\n    img = cv2.imread(image_path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    #img = crop_image_from_gray(img)\n    img = cv2.resize(img, (desired_size,desired_size))\n    img = cv2.addWeighted(img,4,cv2.GaussianBlur(img, (0,0), desired_size/40) ,-4 ,128)\n    \n    return img","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:03:27.518524Z","iopub.execute_input":"2023-03-25T15:03:27.518971Z","iopub.status.idle":"2023-03-25T15:03:27.534438Z","shell.execute_reply.started":"2023-03-25T15:03:27.518931Z","shell.execute_reply":"2023-03-25T15:03:27.533477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__UPDATE:__ Here we are reading just the validation set. In order to use 224x224 images, we are going to load one bucket at a time only when needed. This will let our code run without memory-related errors.","metadata":{}},{"cell_type":"code","source":"# validation set\n# Set N to the number of rows in val_df\nN = val_df.shape[0]\n# Create an empty numpy array of dimensions (N, im_size, im_size, 3) to hold the preprocessed validation images\nx_val = np.empty((N, im_size, im_size, 3), dtype=np.uint8)\n# Loop over each image ID in the id_code column of the val_df dataframe\nfor i, image_id in enumerate(tqdm_notebook(val_df['id_code'])):\n    # Call the preprocess_image function to preprocess the image with the desired size im_size\n    preprocessed_image = preprocess_image(f'{image_id}', desired_size=im_size)\n    # Add the preprocessed image to the x_val numpy array\n    x_val[i, :, :, :] = preprocessed_image","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:03:44.293267Z","iopub.execute_input":"2023-03-25T15:03:44.294483Z","iopub.status.idle":"2023-03-25T15:16:47.919924Z","shell.execute_reply.started":"2023-03-25T15:03:44.294424Z","shell.execute_reply":"2023-03-25T15:16:47.917915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Assign \"diagnosis' values 0,1,2,3,4 from train and validation to y_train and y_val numpy arrays.\ny_train = train_df['diagnosis'].values\ny_val = val_df['diagnosis'].values\n\n#print(y_train.shape)\n#print(x_val.shape)\n#print(y_val.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:16:47.923097Z","iopub.execute_input":"2023-03-25T15:16:47.923529Z","iopub.status.idle":"2023-03-25T15:16:47.930585Z","shell.execute_reply.started":"2023-03-25T15:16:47.923486Z","shell.execute_reply":"2023-03-25T15:16:47.929501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating multilabels\n\nInstead of predicting a single label, we will change our target to be a multilabel problem; i.e., if the target is a certain class, then it encompasses all the classes before it. E.g. encoding a class 4 retinopathy would usually be `[0, 0, 0, 1]`, but in our case we will predict `[1, 1, 1, 1]`. For more details, please check out [Lex's kernel](https://www.kaggle.com/lextoumbourou/blindness-detection-resnet34-ordinal-targets).","metadata":{"jupyter":{"source_hidden":true}}},{"cell_type":"markdown","source":"y_train_multi = np.empty(y_train.shape, dtype=y_train.dtype)\ny_train_multi[:, 4] = y_train[:, 4]\n\nfor i in range(3, -1, -1):\n    y_train_multi[:, i] = np.logical_or(y_train[:, i], y_train_multi[:, i+1])\n\ny_val_multi = np.empty(y_val.shape, dtype=y_val.dtype)\ny_val_multi[:, 4] = y_val[:, 4]\n\nfor i in range(3, -1, -1):\n    y_val_multi[:, i] = np.logical_or(y_val[:, i], y_val_multi[:, i+1])\n\nprint(\"Y_train multi: {}\".format(y_train_multi.shape))\nprint(\"Y_val multi: {}\".format(y_val_multi.shape))","metadata":{"execution":{"iopub.status.busy":"2023-02-28T18:46:33.825579Z","iopub.execute_input":"2023-02-28T18:46:33.826046Z","iopub.status.idle":"2023-02-28T18:46:33.83561Z","shell.execute_reply.started":"2023-02-28T18:46:33.825983Z","shell.execute_reply":"2023-02-28T18:46:33.834857Z"},"jupyter":{"source_hidden":true}}},{"cell_type":"markdown","source":"y_train = y_train_multi\ny_val = y_val_multi","metadata":{"execution":{"iopub.status.busy":"2023-02-28T18:46:33.83678Z","iopub.execute_input":"2023-02-28T18:46:33.837242Z","iopub.status.idle":"2023-02-28T18:46:33.844853Z","shell.execute_reply.started":"2023-02-28T18:46:33.837193Z","shell.execute_reply":"2023-02-28T18:46:33.8441Z"},"jupyter":{"source_hidden":true}}},{"cell_type":"code","source":"# delete the uneeded df\ndel new_train\ndel old_train\ndel val_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:16:47.932337Z","iopub.execute_input":"2023-03-25T15:16:47.932766Z","iopub.status.idle":"2023-03-25T15:16:48.378464Z","shell.execute_reply.started":"2023-03-25T15:16:47.932727Z","shell.execute_reply":"2023-03-25T15:16:48.377247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating keras callback for QWK\n---\nThis is a class definition that inherits from the Keras Callback class. The purpose of this class is to calculate the quadratic weighted kappa (QWK) score for every epoch using the cohen_kappa_score function, and then save the model that achieves the highest QWK score during training.\nIn the on_epoch_end method, the validation data (X_val and y_val) are passed as inputs, and the model.predict method is called to generate predictions (y_pred). The predictions are then mapped to the closest integer label value based on the given coef values. The calculated QWK score is appended to the val_kappas list, and the score is printed to the console. If the current QWK score is the highest so far, the model is saved with the name model.h5.\nThis class is designed to work with a progressive learning approach, where training data is split into buckets and trained sequentially, so the QWK score is calculated based on the full validation set in order to consider the best kappa score among all the buckets.","metadata":{}},{"cell_type":"code","source":"class Metrics(Callback):\n\n    def __init__(self, validation_data):\n        super().__init__()\n        self.validation_data = validation_data\n        self.val_kappas = []\n\n    def on_epoch_end(self, epoch, logs={}):\n        X_val, y_val = self.validation_data\n        \n        y_pred = self.model.predict(X_val)\n        \n        coef = [0.5, 1.5, 2.5, 3.5]\n\n        for i, pred in enumerate(y_pred):\n            if pred < coef[0]:\n                y_pred[i] = 0\n            elif pred >= coef[0] and pred < coef[1]:\n                y_pred[i] = 1\n            elif pred >= coef[1] and pred < coef[2]:\n                y_pred[i] = 2\n            elif pred >= coef[2] and pred < coef[3]:\n                y_pred[i] = 3\n            else:\n                y_pred[i] = 4\n\n        _val_kappa = cohen_kappa_score(\n            y_val,\n            y_pred, \n            weights='quadratic'\n        )\n\n        self.val_kappas.append(_val_kappa)\n\n        print(f\"val_kappa: {_val_kappa:.4f}\")\n        \n        if _val_kappa == max(self.val_kappas):\n            print(\"Validation Kappa has improved. Saving model.\")\n            self.model.save('model.h5')\n\n        return\n","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:16:48.381042Z","iopub.execute_input":"2023-03-25T15:16:48.381509Z","iopub.status.idle":"2023-03-25T15:16:48.393387Z","shell.execute_reply.started":"2023-03-25T15:16:48.381450Z","shell.execute_reply":"2023-03-25T15:16:48.392180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Generator\n\nApply data augmentation techniques to images during training, including horizontal and vertical flipping, rotation up to 160 degrees, and zooming up to 35%.","metadata":{}},{"cell_type":"code","source":"def create_datagen():\n    return ImageDataGenerator(\n        horizontal_flip = True,\n        vertical_flip = True,\n        rotation_range = 160,\n        zoom_range=0.35\n    )","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:18:39.092010Z","iopub.execute_input":"2023-03-25T15:18:39.092827Z","iopub.status.idle":"2023-03-25T15:18:39.099780Z","shell.execute_reply.started":"2023-03-25T15:18:39.092774Z","shell.execute_reply":"2023-03-25T15:18:39.098047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model: EfficientNetV2S","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.layers import Dense, Dropout, GlobalAveragePooling2D, Lambda\nfrom keras.applications import EfficientNetV2S\n\n\nefficientnet = EfficientNetV2S(\n        include_top=False,\n        weights=\"imagenet\",\n        input_shape=(im_size,im_size,3)\n    )\n\ndef build_model():\n        model = tf.keras.models.Sequential()\n\n        model.add(efficientnet)\n        model.add(GlobalAveragePooling2D())\n        model.add(Dropout(0.5))\n        model.add(Dense(1, activation='linear'))\n        # model.add(Lambda(lambda x: x * 200.0)) output scaling\n\n        model.compile(\n            loss='mean_squared_error',\n            optimizer=tf.keras.optimizers.Adam(lr=0.0001),\n            metrics=['mae']\n        )\n\n        return model\nmodel = build_model()\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:19:12.037116Z","iopub.execute_input":"2023-03-25T15:19:12.037582Z","iopub.status.idle":"2023-03-25T15:19:23.935620Z","shell.execute_reply.started":"2023-03-25T15:19:12.037537Z","shell.execute_reply":"2023-03-25T15:19:23.934330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(\n    model, dpi=60,\n    show_shapes=True\n)","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:19:31.753436Z","iopub.execute_input":"2023-03-25T15:19:31.753858Z","iopub.status.idle":"2023-03-25T15:19:32.241478Z","shell.execute_reply.started":"2023-03-25T15:19:31.753819Z","shell.execute_reply":"2023-03-25T15:19:32.239928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training & Evaluation\n\nThis code sets up the parameters for a progressive learning method, where the training data is split into 8 buckets and the model is trained on each bucket for a specified number of epochs. The results of each training session are saved in \"results\" dataframe. \n","metadata":{}},{"cell_type":"code","source":"#train_df = train_df.reset_index(drop=True)\nnum_bucket = 10\ndiv = round(train_df.shape[0]/num_bucket)","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:19:51.731309Z","iopub.execute_input":"2023-03-25T15:19:51.732314Z","iopub.status.idle":"2023-03-25T15:19:51.739490Z","shell.execute_reply.started":"2023-03-25T15:19:51.732250Z","shell.execute_reply":"2023-03-25T15:19:51.738529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results = pd.DataFrame({\n                        'val_loss': [0.0],\n                        'val_mean_absolute_error': [0.0],\n                        'loss': [0.0], \n                        'mean_absolute_error': [0.0],\n                        'bucket': [0.0]\n                        })","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:19:55.051596Z","iopub.execute_input":"2023-03-25T15:19:55.052370Z","iopub.status.idle":"2023-03-25T15:19:55.061588Z","shell.execute_reply.started":"2023-03-25T15:19:55.052306Z","shell.execute_reply":"2023-03-25T15:19:55.060401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# I found that changing the nr. of epochs for each bucket helped in terms of performances\nepochs = [5,5,10,10,10,15,15,15,20,25]\nkappa_metrics = Metrics(validation_data=(x_val, y_val))\nkappa_metrics.val_kappas = []","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:19:59.217997Z","iopub.execute_input":"2023-03-25T15:19:59.218996Z","iopub.status.idle":"2023-03-25T15:19:59.225158Z","shell.execute_reply.started":"2023-03-25T15:19:59.218952Z","shell.execute_reply":"2023-03-25T15:19:59.223644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(0,num_bucket):\n    if i != (num_bucket-1):\n        print(\"Bucket Nr: {}\".format(i))\n        \n        N = train_df.iloc[i*div:(1+i)*div].shape[0]\n        x_train = np.empty((N, im_size, im_size, 3), dtype=np.uint8)\n        for j, image_id in enumerate(tqdm_notebook(train_df.iloc[i*div:(1+i)*div,0])):\n            x_train[j, :, :, :] = preprocess_image_old(f'{image_id}', desired_size = im_size)\n\n        data_generator = create_datagen().flow(x_train, y_train[i*div:(1+i)*div], batch_size=BATCH_SIZE)\n        history = model.fit(\n                        data_generator,\n                        steps_per_epoch=x_train.shape[0] / BATCH_SIZE,\n                        epochs=epochs[i],\n                        validation_data=(x_val, y_val),\n                        callbacks=[kappa_metrics],\n                        max_queue_size=2, \n                        workers=4,\n                        use_multiprocessing=True,\n                        )\n        \n        dic = history.history\n        df_model = pd.DataFrame(dic)\n        df_model['bucket'] = i\n    else:\n        print(\"Bucket Nr: {}\".format(i))\n        \n        N = train_df.iloc[i*div:].shape[0]\n        x_train = np.empty((N, im_size, im_size, 3), dtype=np.uint8)\n        for j, image_id in enumerate(tqdm_notebook(train_df.iloc[i*div:,0])):\n            x_train[j, :, :, :] = preprocess_image_old(f'{image_id}', desired_size = im_size)\n        data_generator = create_datagen().flow(x_train, y_train[i*div:], batch_size=BATCH_SIZE)\n        \n        history = model.fit(\n                        data_generator,\n                        steps_per_epoch=x_train.shape[0] / BATCH_SIZE,\n                        epochs=epochs[i],\n                        validation_data=(x_val, y_val),\n                        callbacks=[kappa_metrics],\n                        max_queue_size=2, \n                        workers=4,\n                        use_multiprocessing=True,\n                        )\n        \n        dic = history.history\n        df_model = pd.DataFrame(dic)\n        df_model['bucket'] = i\n\n    results = results.append(df_model)\n    \n    del data_generator\n    del x_train\n    gc.collect()\n    \n    print('-'*40)","metadata":{"execution":{"iopub.status.busy":"2023-03-25T15:20:06.350014Z","iopub.execute_input":"2023-03-25T15:20:06.350431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results = results.iloc[1:]\nresults['kappa'] = kappa_metrics.val_kappas\nresults = results.reset_index()\nresults = results.rename(index=str, columns={\"index\": \"epoch\"})\nprint(max(results.kappa))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results[['loss', 'val_loss']].plot()\nresults[['mean_absolute_error', 'val_mean_absolute_error']].plot()\nresults[['kappa']].plot()\nresults.to_csv('model_results.csv',index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# QWK Optimizer\nSee Abhishek optimizer [here](https://www.kaggle.com/abhishek/optimizer-for-quadratic-weighted-kappa)","metadata":{}},{"cell_type":"code","source":"class OptimizedRounder(object):\n    def __init__(self):\n        self.coef_ = 0\n\n    def _kappa_loss(self, coef, X, y):\n        X_p = np.copy(X)\n        for i, pred in enumerate(X_p):\n            if pred < coef[0]:\n                X_p[i] = 0\n            elif pred >= coef[0] and pred < coef[1]:\n                X_p[i] = 1\n            elif pred >= coef[1] and pred < coef[2]:\n                X_p[i] = 2\n            elif pred >= coef[2] and pred < coef[3]:\n                X_p[i] = 3\n            else:\n                X_p[i] = 4\n\n        ll = cohen_kappa_score(y, X_p, weights='quadratic')\n        return -ll\n\n    def fit(self, X, y):\n        loss_partial = partial(self._kappa_loss, X=X, y=y)\n        initial_coef = [0.5, 1.5, 2.5, 3.5]\n        self.coef_ = scipy.optimize.minimize(loss_partial, initial_coef, method='nelder-mead')\n        print(-loss_partial(self.coef_['x']))\n\n    def predict(self, X, coef):\n        X_p = np.copy(X)\n        for i, pred in enumerate(X_p):\n            if pred < coef[0]:\n                X_p[i] = 0\n            elif pred >= coef[0] and pred < coef[1]:\n                X_p[i] = 1\n            elif pred >= coef[1] and pred < coef[2]:\n                X_p[i] = 2\n            elif pred >= coef[2] and pred < coef[3]:\n                X_p[i] = 3\n            else:\n                X_p[i] = 4\n        return X_p\n\n    def coefficients(self):\n        return self.coef_['x']","metadata":{"execution":{"iopub.status.busy":"2023-03-07T18:50:25.239493Z","iopub.execute_input":"2023-03-07T18:50:25.240024Z","iopub.status.idle":"2023-03-07T18:50:25.271947Z","shell.execute_reply.started":"2023-03-07T18:50:25.239834Z","shell.execute_reply":"2023-03-07T18:50:25.270552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from functools import partial","metadata":{"execution":{"iopub.status.busy":"2023-03-07T18:50:25.273206Z","iopub.execute_input":"2023-03-07T18:50:25.273638Z","iopub.status.idle":"2023-03-07T18:50:25.277745Z","shell.execute_reply.started":"2023-03-07T18:50:25.273467Z","shell.execute_reply":"2023-03-07T18:50:25.276879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.load_weights('model.h5')\ny_val_pred = model.predict(x_val)\n\noptR = OptimizedRounder()\noptR.fit(y_val_pred, y_val)\ncoefficients = optR.coefficients()\nprint(f'Coefficients: {coefficients}')\ny_val_pred = optR.predict(y_val_pred, coefficients)\n\nscore = cohen_kappa_score(y_val_pred, y_val, weights='quadratic')\n\nprint('Optimized Validation QWK score: {}'.format(score))\nprint('Not Optimized Validation QWK score: {}'.format(max(results.kappa)))","metadata":{"execution":{"iopub.status.busy":"2023-03-07T18:50:25.279369Z","iopub.execute_input":"2023-03-07T18:50:25.279888Z","iopub.status.idle":"2023-03-07T18:50:41.492008Z","shell.execute_reply.started":"2023-03-07T18:50:25.279666Z","shell.execute_reply":"2023-03-07T18:50:41.490547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save final model","metadata":{}},{"cell_type":"code","source":"model.save('ResNet50_model.h5')","metadata":{"execution":{"iopub.status.busy":"2023-03-07T18:50:41.493316Z","iopub.execute_input":"2023-03-07T18:50:41.493613Z","iopub.status.idle":"2023-03-07T18:50:42.216185Z","shell.execute_reply.started":"2023-03-07T18:50:41.493567Z","shell.execute_reply":"2023-03-07T18:50:42.215334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, roc_curve, auc\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport statsmodels.api as sm\n\n# Load the model and predict on the validation data\nmodel.load_weights('model.h5')\ny_val_pred = model.predict(x_val)\n\n# Instantiate an OptimizedRounder object and fit on the validation data\noptR = OptimizedRounder()\noptR.fit(y_val_pred, y_val)\ncoefficients = optR.coefficients()\nprint(f'Coefficients: {coefficients}')\n\n# Get the predicted labels using the optimized coefficients\ny_val_pred_rounded = optR.predict(y_val_pred, coefficients)","metadata":{"execution":{"iopub.status.busy":"2023-03-07T19:00:23.409121Z","iopub.execute_input":"2023-03-07T19:00:23.409452Z","iopub.status.idle":"2023-03-07T19:00:36.753863Z","shell.execute_reply.started":"2023-03-07T19:00:23.409385Z","shell.execute_reply":"2023-03-07T19:00:36.753015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compute confusion matrix\ncm = confusion_matrix(y_val, y_val_pred_rounded)\n\n# Plot confusion matrix\nfig, ax = plt.subplots(figsize=(8, 6))\nsns.heatmap(cm, annot=True, cmap='Blues', fmt='g', ax=ax)\n\n# Add labels, title, and axis ticks\nax.set_xlabel('Predicted labels')\nax.set_ylabel('True labels')\nax.set_title('Confusion Matrix')\nax.xaxis.set_ticklabels(['No-DR', 'Mild NPDR', 'Moderate NPDR', 'Severe NPDR', 'PDR'])\nax.yaxis.set_ticklabels(['No-DR', 'Mild NPDR', 'Moderate NPDR', 'Severe NPDR', 'PDR'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-07T19:00:41.411215Z","iopub.execute_input":"2023-03-07T19:00:41.411551Z","iopub.status.idle":"2023-03-07T19:00:41.808494Z","shell.execute_reply.started":"2023-03-07T19:00:41.411491Z","shell.execute_reply":"2023-03-07T19:00:41.807518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Map 0 to 'No DR' and 1,2,3,4 to 'DR' in y_true\ny_true_binary = np.where(y_val == 0, 0, 1)\ny_val_pred_rounded_binary = np.where(y_val_pred_rounded == 0, 0, 1)\n\n# Compute ROC curve and AUC\nfpr, tpr, thresholds = roc_curve(y_true_binary, y_val_pred_rounded_binary, pos_label=1)\nroc_auc = auc(fpr, tpr)\n\nplt.plot(fpr, tpr, color='darkorange', lw=2, label='ROC curve (area = %0.2f)' % roc_auc)\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--')\nplt.xlim([0.0, 1.0])\nplt.ylim([0.0, 1.05])\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Receiver operating characteristic curve')\nplt.legend(loc=\"lower right\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-07T19:25:20.522101Z","iopub.execute_input":"2023-03-07T19:25:20.522411Z","iopub.status.idle":"2023-03-07T19:25:20.762741Z","shell.execute_reply.started":"2023-03-07T19:25:20.522357Z","shell.execute_reply":"2023-03-07T19:25:20.761954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compute confusion matrix\ncm = confusion_matrix(y_true_binary, y_val_pred_rounded_binary)\n\n# Calculate specificity and sensitivity\ntn, fp, fn, tp = cm.ravel()\nspecificity = tn / (tn + fp)\nsensitivity = tp / (tp + fn)\n\n# Display confusion matrix with percentages\nfig, ax = plt.subplots(figsize=(8, 6))\nsns.heatmap(cm, annot=True, cmap='Blues', fmt='g', ax=ax)\n\n# Add labels, title, and axis ticks\nax.set_xlabel('Predicted labels')\nax.set_ylabel('True labels')\nax.set_title('Confusion Matrix (Specificity={:.2f}, Sensitivity={:.2f})'.format(specificity, sensitivity))\nax.xaxis.set_ticklabels(['No-DR', 'DR'])\nax.yaxis.set_ticklabels(['No-DR', 'DR'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-07T19:31:23.705237Z","iopub.execute_input":"2023-03-07T19:31:23.705573Z","iopub.status.idle":"2023-03-07T19:31:23.973269Z","shell.execute_reply.started":"2023-03-07T19:31:23.705514Z","shell.execute_reply":"2023-03-07T19:31:23.972363Z"},"trusted":true},"execution_count":null,"outputs":[]}]}