{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CNN using transfer learning in Keras","metadata":{}},{"cell_type":"markdown","source":"## Abstraction\n\nWhat if we can detect cancer at early stage? In this competition, the challenge is to develop an model that can accurately identify metastatic cancer in small image patches extracted from larger digital pathology scans. The required task is to predict the probability, so it is not the classification but binary task.\n\nThis project goal is to develop a deep learning model using transfer learning to classify the image patches into positive or negative for metastatic cancer. We utilize the pretrained ResNet152 model as a feature extraction backbone and build a classifier on top of it, using some fully connected dense layers. The model is trained on a single GPU and evaluated using area under the ROC curve as the primary metric. We also monitor accuracy and validation loss during training to ensure optimal performance. We utilie hyper param tuning to get the appropriate learning rate for our model. After getting the final model, we make the submission file to get the final score of this kaggle's competition.\n\nThe link to the original competition is, https://www.kaggle.com/competitions/histopathologic-cancer-detection/overview. You can get the same dataset as we used in this notebook.\n\n\n**keywords**: binary classification, Keras, CNN, transfer learning, resnet152, image augumentation, hyper param tuning\n","metadata":{}},{"cell_type":"markdown","source":"### Import libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport os\n\nimport tensorflow as tf\nprint(tf.__version__)\nfrom tensorflow.keras import models, layers, mixed_precision\n\ntrain_dir = \"/kaggle/input/histopathologic-cancer-detection/train\"\ntest_dir = \"/kaggle/input/histopathologic-cancer-detection/test\"\n\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_global_policy(policy) \nprint('Compute dtype: %s' % policy.compute_dtype)\nprint('Variable dtype: %s' % policy.variable_dtype)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.execute_input":"2023-04-12T10:35:25.379268Z","iopub.status.busy":"2023-04-12T10:35:25.378754Z","iopub.status.idle":"2023-04-12T10:35:41.916008Z","shell.execute_reply":"2023-04-12T10:35:41.914734Z","shell.execute_reply.started":"2023-04-12T10:35:25.379212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Traing Number: \", len(os.listdir(train_dir)))\nprint(\"Test Number: \", len(os.listdir(test_dir)))","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:35:41.920417Z","iopub.status.busy":"2023-04-12T10:35:41.918579Z","iopub.status.idle":"2023-04-12T10:35:45.916678Z","shell.execute_reply":"2023-04-12T10:35:45.915422Z","shell.execute_reply.started":"2023-04-12T10:35:41.920369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We get the training dataframe for later image loading. The trainingset has 2 classes, 0 for no cancer, 1 for at least 1 cancer. 40 % of training datasets are cancer images.","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"/kaggle/input/histopathologic-cancer-detection/train_labels.csv\")\nprint(\"Data's target distribution ((1) label num/ (1 + 0) label num): \", len(df[df.label == 1]) / len(df))\ndf.head()","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:35:45.918914Z","iopub.status.busy":"2023-04-12T10:35:45.918474Z","iopub.status.idle":"2023-04-12T10:35:46.510336Z","shell.execute_reply":"2023-04-12T10:35:46.509102Z","shell.execute_reply.started":"2023-04-12T10:35:45.918873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.label = df.label.astype(str)\ndf.id = df.id + \".tif\"\nprint(df.info())\ndf.head()","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:35:46.514248Z","iopub.status.busy":"2023-04-12T10:35:46.513502Z","iopub.status.idle":"2023-04-12T10:35:46.728406Z","shell.execute_reply":"2023-04-12T10:35:46.726996Z","shell.execute_reply.started":"2023-04-12T10:35:46.514206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## EDA","metadata":{}},{"cell_type":"markdown","source":"The plot shows some training images. There are not obvious features that we find to classfiy which images indicate cancer or not. ","metadata":{}},{"cell_type":"code","source":"w = 10\nh = 10\nfig = plt.figure(figsize=(15, 15))\ncolumns = 10\nrows = 5\nfor i in range(1, columns*rows +1):\n    img = plt.imread(train_dir + \"/\" + df.iloc[i][\"id\"])\n    fig.add_subplot(rows, columns, i)\n    plt.axis(\"off\")\n    plt.title(df.iloc[i][\"label\"])\n    plt.imshow(img)\nplt.show()","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:35:46.731096Z","iopub.status.busy":"2023-04-12T10:35:46.730609Z","iopub.status.idle":"2023-04-12T10:35:49.238741Z","shell.execute_reply":"2023-04-12T10:35:49.236760Z","shell.execute_reply.started":"2023-04-12T10:35:46.731040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One image shape is `(width: 96, height: 96, color channel: 3)`.","metadata":{}},{"cell_type":"code","source":"plt.figure()\nimg = plt.imread(train_dir + \"/\" + df.iloc[0][\"id\"])\nprint(\"Image shape: \", img.shape)\nprint(\"Label: \", df.iloc[0][\"label\"])\nplt.imshow(img)\nplt.colorbar()\nplt.grid(False)\nplt.show()","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:35:49.242249Z","iopub.status.busy":"2023-04-12T10:35:49.241861Z","iopub.status.idle":"2023-04-12T10:35:49.755616Z","shell.execute_reply":"2023-04-12T10:35:49.754578Z","shell.execute_reply.started":"2023-04-12T10:35:49.242214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Belows codes are helper function we will use in model building and evaluating.","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\n\ndef get_train_val_generator(train_datagen, df, sample_frac=1.0, bs=64):\n    df = df.sample(frac=sample_frac, random_state=42)\n    \n    train_generator = train_datagen.flow_from_dataframe(dataframe=df,\n                                                       directory=train_dir,\n                                                       x_col=\"id\",\n                                                       y_col=\"label\",\n                                                       subset=\"training\",\n                                                       target_size=(96, 96),\n                                                       batch_size=bs,\n                                                       class_mode=\"binary\")\n    valid_generator = train_datagen.flow_from_dataframe(dataframe=df,\n                                                       directory=train_dir,\n                                                       x_col=\"id\",\n                                                       y_col=\"label\",\n                                                       subset=\"validation\",\n                                                       target_size=(96, 96),\n                                                       batch_size=bs,\n                                                       shuffle=False,\n                                                       class_mode=\"binary\")\n    return train_generator, valid_generator\n\ndef get_model(pretrained_model, preprocess_input):\n    inputs = tf.keras.Input(shape=(96, 96, 3))\n    # For feature extraction using transfer learning\n    x = preprocess_input(inputs)\n    x = pretrained_model(x) \n    # For classifier\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dropout(0.2)(x)\n    x = tf.keras.layers.Dense(64, activation=\"relu\")(x)\n    x = tf.keras.layers.BatchNormalization()(x)\n    x = tf.keras.layers.Dense(64, activation=\"relu\")(x)\n    x = tf.keras.layers.BatchNormalization()(x)\n    x = tf.keras.layers.Dense(1)(x)\n    outputs = tf.keras.layers.Activation(\"sigmoid\", dtype=\"float32\")(x)\n    \n    return tf.keras.Model(inputs, outputs)\n\ndef fit_model(model, train_generator, valid_generator, epochs=5, callbacks=[]):\n    return model.fit(train_generator,\n                    steps_per_epoch=train_generator.n//train_generator.batch_size,\n                    epochs=epochs,\n                    validation_data=valid_generator,\n                    validation_steps=valid_generator.n//valid_generator.batch_size,\n                    use_multiprocessing=True,\n                    workers=4,\n                    callbacks=callbacks)\n\ndef plt_performance(train, valid, title):\n    plt.figure(figsize=(10, 10))\n    plt.subplot(2, 1, 1)\n    plt.plot(train, label='Training')\n    plt.plot(valid, label='Validation')\n    plt.legend(loc='upper left')\n    plt.ylim([min(plt.ylim())-0.1,max(plt.ylim())+0.1])\n    plt.title(title)","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:35:49.757872Z","iopub.status.busy":"2023-04-12T10:35:49.757236Z","iopub.status.idle":"2023-04-12T10:35:49.792355Z","shell.execute_reply":"2023-04-12T10:35:49.791039Z","shell.execute_reply.started":"2023-04-12T10:35:49.757820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Load sample training and validation set","metadata":{}},{"cell_type":"markdown","source":"We get 30% sample of all training data so that we can iterate our experiments more faster. Then dataset is split into 2 parts, training and validating. We treat validation set to see the model' performance while training. ","metadata":{}},{"cell_type":"code","source":"train_datagen = ImageDataGenerator(validation_split=0.2)\ntrain_generator, valid_generator = get_train_val_generator(train_datagen, df, sample_frac=0.3)","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:35:49.802321Z","iopub.status.busy":"2023-04-12T10:35:49.798774Z","iopub.status.idle":"2023-04-12T10:39:56.071243Z","shell.execute_reply":"2023-04-12T10:39:56.070119Z","shell.execute_reply.started":"2023-04-12T10:35:49.802272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Building and Evaluation","metadata":{}},{"cell_type":"markdown","source":"### Compare pretrained model","metadata":{}},{"cell_type":"markdown","source":"The pretrained models will be used as feature extraction layers. We compare each model's initial validaiton loss, and conclude to use `Resnet152` as our base model. The efficent net model would also seem good. However, after some training, resnet would have better performance among all.","metadata":{}},{"cell_type":"code","source":"preprocess_mobile = tf.keras.applications.mobilenet_v2.preprocess_input\nmobilenet_v2 = tf.keras.applications.MobileNetV2(input_shape=(96, 96, 3), include_top=False, weights=\"imagenet\")\n\npreprocess_res = tf.keras.applications.resnet_v2.preprocess_input\nresnet_v2 = tf.keras.applications.ResNet152V2(input_shape=(96, 96, 3), include_top=False, weights=\"imagenet\")\n\npreprocess_incep = tf.keras.applications.inception_resnet_v2.preprocess_input\nincep_v2 = tf.keras.applications.InceptionResNetV2(input_shape=(96, 96, 3), include_top=False, weights=\"imagenet\")\n\npreprocess_dense = tf.keras.applications.densenet.preprocess_input\ndense = tf.keras.applications.DenseNet201(input_shape=(96, 96, 3), include_top=False, weights=\"imagenet\")\n\npreprocess_eff = tf.keras.applications.efficientnet.preprocess_input\neffnet_b2 = tf.keras.applications.EfficientNetB2(input_shape=(96, 96, 3), include_top=False, weights=\"imagenet\")","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:39:56.073334Z","iopub.status.busy":"2023-04-12T10:39:56.072949Z","iopub.status.idle":"2023-04-12T10:40:25.254717Z","shell.execute_reply":"2023-04-12T10:40:25.253489Z","shell.execute_reply.started":"2023-04-12T10:39:56.073290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models = [(mobilenet_v2, preprocess_mobile), (resnet_v2, preprocess_res), (incep_v2, preprocess_incep), (dense, preprocess_incep), (effnet_b2, preprocess_eff)]\nfor pretrained_model, preprocess in models:\n    model = get_model(pretrained_model, preprocess)\n    model.compile(optimizer=\"adam\", loss=\"binary_crossentropy\", metrics=[\"accuracy\"],)\n    val_loss, val_acc = model.evaluate(valid_generator)\n    print(\"\\nPretrained Model: \", pretrained_model.name)\n    print(\"Val Loss: \", val_loss)\n    print(\"Val Acc: \", val_acc)","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:40:25.258522Z","iopub.status.busy":"2023-04-12T10:40:25.258224Z","iopub.status.idle":"2023-04-12T10:45:21.379741Z","shell.execute_reply":"2023-04-12T10:45:21.378648Z","shell.execute_reply.started":"2023-04-12T10:40:25.258494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model building and evaluation","metadata":{}},{"cell_type":"markdown","source":"On top of resnet model, we put 2 fully connected dense layers for classifier. The loss is `BinaryCrossEntropy` since this is binary task, and put label smooothing to `0.1`. This smoothing formula is, `y_true * (1.0 - label_smoothing) + 0.5 * label_smoothing`. The optimizer is `Adam` with just quick learning rate. We tune this LR value at later part of this notebook. We chage the top 10 % resnet layer as trainable, holding other layers untrainable. It could increase the overfitting possibility, so we need to care for that. We see `Trainable params: 17,078,721` from model's summary, and we find this is good amount for this model and task. Thus, we keep it up and see the model's structure.","metadata":{}},{"cell_type":"code","source":"resnet_v2.trainable = True\nprint(\"Number of layers in the base net: \", len(resnet_v2.layers))\n\nfine_tune_at = round(len(resnet_v2.layers) * 0.9)\nprint(\"Mobile model would be trainable from \", fine_tune_at)\nfor l in resnet_v2.layers[:fine_tune_at]:\n    l.trainable = False\n    \nbase_lr = 3e-3\nmodel = get_model(resnet_v2, preprocess_res)\nmodel.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=base_lr),\n             loss=tf.keras.losses.BinaryCrossentropy(label_smoothing=0.1),\n             metrics=[\"accuracy\"],)\nprint(\"Model trainable param number: \", len(model.trainable_variables))\nmodel.summary()","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:51:49.815340Z","iopub.status.busy":"2023-04-12T10:51:49.813606Z","iopub.status.idle":"2023-04-12T10:51:51.511476Z","shell.execute_reply":"2023-04-12T10:51:51.510307Z","shell.execute_reply.started":"2023-04-12T10:51:49.815283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One important layer when we use transfer learning is pretrained model's preprocess input layer. In our model's case, it is `tf.keras.applications.resnet_v2.preprocess_input`. The pretrained model is trained on this preprocessd data, so we need to convert our data as the same way before feeding it into the model. Therefore, we apply the same preprocessing function on our input data to ensure that it matches the input format of the pretrained ResNet152 model.","metadata":{}},{"cell_type":"code","source":"tf.keras.utils.plot_model(model, show_shapes=True)","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:51:51.514805Z","iopub.status.busy":"2023-04-12T10:51:51.513671Z","iopub.status.idle":"2023-04-12T10:51:51.837988Z","shell.execute_reply":"2023-04-12T10:51:51.836824Z","shell.execute_reply.started":"2023-04-12T10:51:51.514765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model training","metadata":{}},{"cell_type":"markdown","source":"We first iterate 3 epochs to see if the model can traing our dataset. It should the model have about `90` % accuracy on validation set in first 2 epoch. Hoever, at 3 poch, the model would somewhat overfit and less generalize with lower performance. It indicates that we need to deal with that overfitting. We add some treatment, like adding dropout and batchnomalization, but still overfittting exits. Therefre we decided to add more randomized image data, applying image augmentation for our models genelization.","metadata":{}},{"cell_type":"code","source":"# Training\ndecay_steps = 20\nlr_decayed_fn = tf.keras.optimizers.schedules.CosineDecay(base_lr, decay_steps)\ncallbacks = [tf.keras.callbacks.LearningRateScheduler(lr_decayed_fn)]\nhistory = fit_model(model, train_generator, valid_generator, epochs=3, callbacks=callbacks)\n\n# Evaluating\nplt_performance(history.history[\"accuracy\"], history.history[\"val_accuracy\"], \"Train/Valid Accuracy\")\nplt_performance(history.history[\"loss\"], history.history[\"val_loss\"], \"Train/Valid Loss\")","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:52:07.276920Z","iopub.status.busy":"2023-04-12T10:52:07.275746Z","iopub.status.idle":"2023-04-12T10:58:59.429712Z","shell.execute_reply":"2023-04-12T10:58:59.428492Z","shell.execute_reply.started":"2023-04-12T10:52:07.276861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We apply some augmentation here at the stage of loading image datasets. Some random transformation would increase the training data and improve the model's validation performance. We didnot try test time augumentation(TTA), but it is also valid way to ease the overfitting.","metadata":{}},{"cell_type":"code","source":"train_datagen = ImageDataGenerator(validation_split=0.2,\n                                      vertical_flip=True,\n                                      horizontal_flip=True,\n                                      width_shift_range=0.2,\n                                      height_shift_range=0.2,\n                                      zoom_range=0.2,\n                                      )\ntrain_aug_generator, valid_aug_generator = get_train_val_generator(train_datagen, df, sample_frac=0.3)","metadata":{"execution":{"iopub.execute_input":"2023-04-12T10:58:59.433073Z","iopub.status.busy":"2023-04-12T10:58:59.432222Z","iopub.status.idle":"2023-04-12T11:00:59.228860Z","shell.execute_reply":"2023-04-12T11:00:59.227765Z","shell.execute_reply.started":"2023-04-12T10:58:59.433014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After loading augumented training dataset, we fit the model and see its performance. We can find the augumentation improve the model's validaiton accuracy with less overfitting.","metadata":{}},{"cell_type":"code","source":"# Training\nmodel = None\nmodel = get_model(effnet_b2, preprocess_eff)\nmodel.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=base_lr),\n             loss=tf.keras.losses.BinaryCrossentropy(label_smoothing=0.1),\n             metrics=[\"accuracy\"],)\nhistory = fit_model(model, train_aug_generator, valid_aug_generator, epochs=3, callbacks=callbacks)\n\n# Evaluating\nplt_performance(history.history[\"accuracy\"], history.history[\"val_accuracy\"], \"Train/Valid Accuracy\")\nplt_performance(history.history[\"loss\"], history.history[\"val_loss\"], \"Train/Valid Loss\")","metadata":{"execution":{"iopub.execute_input":"2023-04-12T11:00:59.230884Z","iopub.status.busy":"2023-04-12T11:00:59.230202Z","iopub.status.idle":"2023-04-12T11:15:04.079247Z","shell.execute_reply":"2023-04-12T11:15:04.077656Z","shell.execute_reply.started":"2023-04-12T11:00:59.230841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Hyper parameter tuning for better learning rate","metadata":{}},{"cell_type":"markdown","source":"So far we use some arbitray learning rate, but it could be better. We search more better LR using built in, `keras_tuner` and get the better one.","metadata":{}},{"cell_type":"code","source":"x_train, y_train = train_aug_generator.next()\nx_val, y_val = valid_aug_generator.next()","metadata":{"execution":{"iopub.execute_input":"2023-04-12T11:20:32.190923Z","iopub.status.busy":"2023-04-12T11:20:32.189652Z","iopub.status.idle":"2023-04-12T11:20:32.751346Z","shell.execute_reply":"2023-04-12T11:20:32.750204Z","shell.execute_reply.started":"2023-04-12T11:20:32.190862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import keras_tuner as kt\n\nclass MyHyperModel(kt.HyperModel):\n    def build(self, hp):\n        model = get_model(resnet_v2, preprocess_res)\n        model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=hp.Float('learning_rate', min_value = 1e-4, max_value =1e-2, sampling='log')),\n                     loss=tf.keras.losses.BinaryCrossentropy(label_smoothing=0.1),\n                     metrics=[\"accuracy\"],)\n\n        return model\n\ntuner = kt.RandomSearch(\n    MyHyperModel(),\n    objective='val_loss',\n    max_trials=10\n)\n\ntuner.search(x_train, y_train,\n             validation_data= (x_val,y_val), \n             epochs=10,\n             callbacks=[tf.keras.callbacks.EarlyStopping(patience=2)])","metadata":{"execution":{"iopub.execute_input":"2023-04-12T11:24:38.156351Z","iopub.status.busy":"2023-04-12T11:24:38.155620Z","iopub.status.idle":"2023-04-12T11:27:36.125366Z","shell.execute_reply":"2023-04-12T11:27:36.123964Z","shell.execute_reply.started":"2023-04-12T11:24:38.156313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_hps= tuner.get_best_hyperparameters(1)[0]\nprint(\"Best Learning Rate: \", best_hps.get('learning_rate'))","metadata":{"execution":{"iopub.execute_input":"2023-04-12T11:27:39.738396Z","iopub.status.busy":"2023-04-12T11:27:39.738021Z","iopub.status.idle":"2023-04-12T11:27:39.745427Z","shell.execute_reply":"2023-04-12T11:27:39.744372Z","shell.execute_reply.started":"2023-04-12T11:27:39.738362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model training with full training dataset","metadata":{}},{"cell_type":"markdown","source":"We test our final model on testset. We first load all training dataset and train our model with optimized hyper parameters. The fian model's accuracy are `0.95` on training set, and `0.93` on validation set.","metadata":{}},{"cell_type":"code","source":"train_full_generator, valid_full_generator = get_train_val_generator(train_datagen, df, sample_frac=1)","metadata":{"execution":{"iopub.execute_input":"2023-04-12T11:44:50.635483Z","iopub.status.busy":"2023-04-12T11:44:50.634553Z","iopub.status.idle":"2023-04-12T12:00:56.041438Z","shell.execute_reply":"2023-04-12T12:00:56.040329Z","shell.execute_reply.started":"2023-04-12T11:44:50.635442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_lr = 0.009\nmodel = None\nmodel = get_model(resnet_v2, preprocess_res)\nmodel.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=best_lr),\n             loss=tf.keras.losses.BinaryCrossentropy(label_smoothing=0.1),\n             metrics=[\"accuracy\"],)\n\nlr_decayed_fn = tf.keras.optimizers.schedules.CosineDecay(best_lr, 25)\ncallbacks = [tf.keras.callbacks.LearningRateScheduler(lr_decayed_fn), tf.keras.callbacks.EarlyStopping(patience=2)]\nhistory = fit_model(model, train_full_generator, valid_full_generator, epochs=12, callbacks=callbacks)\n\n# Evaluating\nplt_performance(history.history[\"accuracy\"], history.history[\"val_accuracy\"], \"Train/Valid Accuracy\")\nplt_performance(history.history[\"loss\"], history.history[\"val_loss\"], \"Train/Valid Loss\")","metadata":{"execution":{"iopub.execute_input":"2023-04-12T16:24:18.236637Z","iopub.status.busy":"2023-04-12T16:24:18.236029Z","iopub.status.idle":"2023-04-12T16:24:18.891018Z","shell.execute_reply":"2023-04-12T16:24:18.889976Z","shell.execute_reply.started":"2023-04-12T16:24:18.236598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model inference on testset","metadata":{}},{"cell_type":"markdown","source":"The code is for making submission file to get the final score on the original kaggle competition. ","metadata":{}},{"cell_type":"code","source":"df_test = pd.read_csv(\"/kaggle/input/histopathologic-cancer-detection/sample_submission.csv\")\ndf_test.id = df_test.id + \".tif\"\ntest_generator = ImageDataGenerator().flow_from_dataframe(dataframe=df_test,\n                                                        directory=test_dir,\n                                                        x_col=\"id\",\n                                                        y_col=None,\n                                                        target_size=(96, 96),\n                                                        batch_size=2,\n                                                        shuffle=False,\n                                                        class_mode=None)\n\ntest_generator.reset()\npreds = model.predict(test_generator,verbose=1)","metadata":{"execution":{"iopub.execute_input":"2023-04-12T15:32:12.820487Z","iopub.status.busy":"2023-04-12T15:32:12.818603Z","iopub.status.idle":"2023-04-12T16:19:18.351244Z","shell.execute_reply":"2023-04-12T16:19:18.349888Z","shell.execute_reply.started":"2023-04-12T15:32:12.820416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame()\nsubmission['id'] = df_test['id'].apply(lambda x: x.split('.')[0])\nsubmission['label'] = preds[:, 0]\nsubmission.to_csv('submission.csv', index=False)\nsubmission.head()","metadata":{"execution":{"iopub.execute_input":"2023-04-12T16:20:02.703246Z","iopub.status.busy":"2023-04-12T16:20:02.702434Z","iopub.status.idle":"2023-04-12T16:20:02.958119Z","shell.execute_reply":"2023-04-12T16:20:02.956853Z","shell.execute_reply.started":"2023-04-12T16:20:02.703184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Discussion","metadata":{}},{"cell_type":"markdown","source":"### Result\n\nWe got `0.9272` on private score, while `0.9499` on public score from kaggle's submission.","metadata":{}},{"cell_type":"markdown","source":"### Ky findings\n\nHere is what we learning in this project.\n\n* Final layer's activation function needs to match the task specific one. This project is binary task, requiring us to predict the probability not the label, so sigmoid is appropriate choise. Some other notebooks use softmax but this might be wrong.\n* Training takes a lot of time so we need to find a way to reduce it as much as possible. One way is subsampling the trainiing data while iterating and build the model. Another one is settin `use_multiprocessing=True, workers=4` on model.fit function to use multiple GPU.\n* When we use pretrained model, we also need to use its own preproces input as well. In our case, we use `tf.keras.applications.resnet_v2.preprocess_input` before feeding our data to `tf.keras.applications.ResNet152V2`. That is must step to ensure better performance.\n* Preventing overfit is necessary step, but we should take care of that only after that happend to us. At the first stege of model building, we should think about underfitting.\n* Batchnormalization is strong way to prevent overfiting, but it worked well more with dropout layer in this problem.\n* Image augumentation improve the overall accuracy with less overfitting, while taking more time to train the model.\n* Overall more layers and more data improve the model's accuracy. ","metadata":{}},{"cell_type":"markdown","source":"## Conclusion\n\nThe final model achieved a private score of 0.9272 and a public score of 0.9499 on Kaggle's competition. This project aimed to develop a deep learning model using transfer learning to detect metastatic cancer in small image patches extracted from larger digital pathology scans. The model utilized the ResNet152 model as a feature extraction backbone and fully connected dense layers as the classifier. The model was trained using area under the ROC curve as the primary metric, and accuracy and validation loss were monitored during training. Key findings from the project included the importance of matching the final layer's activation function to the task, using the pretrained model's preprocess input, preventing underfitting, utilizing batch normalization and dropout layers to prevent overfitting, and the effectiveness of image augmentation in improving accuracy while preventing overfitting. Overall, increasing the number of layers and data improved the model's accuracy.","metadata":{}},{"cell_type":"markdown","source":"## Referrences\n\n* tensorflow and keras tutorial: https://www.tensorflow.org/tutorials\n* pretrained keras model: https://keras.io/api/applications/\n* tip for speed training time up: https://analyticsindiamag.com/7-tricks-to-speed-up-the-training-of-a-neural-network/\n* keras hyper param tuning for LR: https://blog.paperspace.com/hyperparameter-optimization-with-keras-tuner/\n* multipe image plots: https://stackoverflow.com/questions/46615554/how-to-display-multiple-images-in-one-figure-correctly\n* kaggle notebook about building CNN from scrach: https://www.kaggle.com/code/hrmello/base-cnn-classification-from-scratch?scriptVersionId=7628679\n* tutorial for keras flow_from_dataframe: https://vijayabhaskar96.medium.com/tutorial-on-keras-flow-from-dataframe-1fd4493d237c","metadata":{}}]}