{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Testing TTA on unseeen data to not overfitt test set/train set"},{"metadata":{},"cell_type":"markdown","source":"**Version 36** : just passing testing my last model. The bug regarding the prediction of horizontal flip is due to the fact that magick exhaust all the ressources somehow and does not flip all images."},{"metadata":{},"cell_type":"markdown","source":"**NB, from version 21 onwards :** In the previous version of the notebook I was running the evaluation of the TTA on the training set. So, it is not really nice to run this kind of optimisation on a set that the model have already seen. Ideally I would like to use some hold-on data. Writting this I realized that I have actually used a seed so I can actually recollect the validation set from the training. I will use the validation data of the training of the eff net."},{"metadata":{},"cell_type":"markdown","source":"**NB, from version 23 onwards :** I have tried to use the dataset of the [previous competition](https://www.kaggle.com/c/cassava-disease). The goal was to see what series of transformation of input optimize the transformation. **Obvious caution :** this set and the test set must have different properties. Anyway, I got a C stack error, I will investigate why. Best bet : it is something related to something from the img format."},{"metadata":{},"cell_type":"markdown","source":"Looking at the winning solution of the [Plant Pathology 2020](https://www.kaggle.com/c/plant-pathology-2020-fgvc7/discussion/154056), the winner use test time augmentation. To put it simply, the goal is to use data augmentation on the different images when doing the inference, and then do an average of the prediction. Keras made it actually straightforward to illustrate it with code. Usually you wrotte something like : "},{"metadata":{"trusted":true},"cell_type":"code","source":"#data augmentation\n#datagen <- image_data_generator(\n#  rotation_range = 40,\n#  width_shift_range = 0.2,\n#  height_shift_range = 0.2,\n#  shear_range = 0.2,\n#  zoom_range = 0.5,\n#  horizontal_flip = TRUE,\n#  fill_mode = \"reflect\"\n#)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And you plug it into the image generator that generate the batch to train the model :"},{"metadata":{"trusted":true},"cell_type":"code","source":"#train_generator <- flow_images_from_dataframe(dataframe = train_labels, \n#                                              directory = image_path,\n#                                              generator = datagen,\n#                                              class_mode = \"other\",\n#                                              x_col = \"image_id\",\n#                                              y_col = c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\"),\n#                                              target_size = c(448, 448),\n#                                              batch_size=16)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Example [here](https://www.kaggle.com/cdk292/efficientnetb0-with-r-and-tf2-cyclic-lr). And while you do it into the train generator, you don't do it with the generator of batch for the validation set. So, what if you plug a datagen inside the test_generator ? You got prediction on altered images. Do it with several datagen and you get tta."},{"metadata":{},"cell_type":"markdown","source":"In this notebook I try several images augmentation on the training set (to not overfitt the test set), to see which one seems to improve the performance of the model. Code from [here](https://www.kaggle.com/cdk292/missclassified-pictures-and-bias-efficientnet-b0) and [here](https://www.kaggle.com/cdk292/efficientnet-b0-predict-w1th0ut-internet)."},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"library(tidyverse)\nlibrary(tensorflow)\ntf$executing_eagerly()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tensorflow::tf_version()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here I flex with my own version of keras. Basically, it is a fork with application wrapper for the efficient net."},{"metadata":{},"cell_type":"markdown","source":"**Disclaimer : I did not writte the code for the really handy applications wrappers.** It came [from this commit](https://github.com/rstudio/keras/commit/c406ec55f7bb2864ac58a17f963448810a531c18) for which the PR is hold until the fully release of tf 2.3, as stated [in this PR](https://github.com/rstudio/keras/pull/1097). I am not sure why the PR is closed."},{"metadata":{"trusted":true},"cell_type":"code","source":"#devtools::install_github(\"Cdk29/keras\", dependencies = FALSE)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"install.packages(\"../input/keras-cdk292/keras-master\", repos = NULL, type = \"source\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"library(keras)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"library(reticulate)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Constructing the model to load the weight "},{"metadata":{"trusted":true},"cell_type":"code","source":"smoothie <- function(y_true,y_pred){\n    tf$losses$categorical_crossentropy(y_true, y_pred, label_smoothing=0.3)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"library(reticulate) #for dict()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model <- load_model_tf(\"../input/fork-of-finetuning-efficientnetb7-label-smooth/EfficientB7_image_net_RE_fine_tuned/\", custom_objects = dict(smoothie = \"smoothie\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(model)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Predict"},{"metadata":{},"cell_type":"markdown","source":"### Train generator"},{"metadata":{"trusted":true},"cell_type":"code","source":"validation_set<-read_csv(\"../input/fork-of-finetuning-efficientnetb7-label-smooth/validation_set.csv\")\nhead(validation_set)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data <- read_csv(\"../input/cassava-leaf-disease-classification/train.csv\")\nhead(data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data<-data[which(data$image_id %in% validation_set$image_id),]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train <- data\ndim(data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path <-\"/kaggle/input/cassava-leaf-disease-classification/train_images/\"\n#With shuffle = FALSE to not mix images and got the right order in the predictions.\n\ntrain_generator <- flow_images_from_dataframe(dataframe = data, \n                                              directory = image_path,\n                                              class_mode = NULL,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"num_train_images<-as.numeric(dim(data)[1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred <- model %>% predict_generator(train_generator, steps=num_train_images)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Creating a function to count the accuracy and return informations on the prediction. Code hidden, just create global variable for cleaner output."},{"metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"trusted":true},"cell_type":"code","source":"pred<-as.data.frame(pred)\ncolnames(pred)<-c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\")\n\nlabel<-c()\nfor (row in 1:dim(pred)[1]){\n    label<-c(label, which(pred[row,]==max(pred[row,])))\n}\nlabel<-(label-1)\n  \nprediction<-data.frame(train$image_id, label)\ncolnames(prediction)<-c(\"image_id\", \"label\")\n    \nidx_img<-which(train$label!=prediction$label)\nmissclassified<-train[idx_img,]$label\nprint(table(missclassified))\nF1_train<<-MLmetrics::F1_Score(y_true=train$label, y_pred=label)\nAccu_train<<-MLmetrics::Accuracy(y_true=train$label, y_pred=label)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## TTA Datagen"},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy <- function(pred, num_train_images, train){\n    pred<-as.data.frame(pred)\n    colnames(pred)<-c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\")\n\n    label<-c()\n    for (row in 1:dim(pred)[1]){\n        label<-c(label, which(pred[row,]==max(pred[row,])))\n    }\n    label<-(label-1)\n    \n    prediction<-data.frame(train$image_id, label)\n    colnames(prediction)<-c(\"image_id\", \"label\")\n    \n    idx_img<-which(train$label!=prediction$label)\n    missclassified<-train[idx_img,]$label\n    print(table(missclassified))\n    print(\"Difference :\")\n    print(table(missclassified)-c(39, 48, 43, 53, 65))\n\n    print(\"F1 :\")\n    F1<-MLmetrics::F1_Score(y_true=train$label, y_pred=label)\n    print(F1) #0.6869757 score of the pred_tta on the training set\n    #print(\"vs 0.6869757 without tta\")\n    print(\"Accuracy :\")\n    Accuracy<-MLmetrics::Accuracy(y_true=train$label, y_pred=label)\n\n    print(Accuracy) #0.8938169\n    print(\"Improvements : \")\n    print(F1-F1_train)\n    print(Accuracy-Accu_train)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy(pred, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  zoom_range = c(1, 1 )\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_datagen <- function(datagen){\n    img_path<-\"/kaggle/input/cassava-leaf-disease-classification/test_images/2216849948.jpg\"\n\n    img <- image_load(img_path, target_size = c(448, 448))\n    img_array <- image_to_array(img)\n    img_array <- array_reshape(img_array, c(1, 448, 448, 3))\n    img_array<-img_array/255\n# Generated that will flow augmented images\n    augmentation_generator <- flow_images_from_data(\n      img_array, \n      generator = datagen, \n      batch_size = 1 \n    )\n    op <- par(mfrow = c(1, 1), pty = \"s\", mar = c(1, 0, 1, 0))\n    for (i in 1:1) {\n      batch <- generator_next(augmentation_generator)\n      plot(as.raster(batch[1,,,]))\n    }\n    par(op)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Testing zoom"},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  zoom_range = c(0.75, 0.75)\n)\nplot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tta_generator <- flow_images_from_dataframe(dataframe = data, \n                                              directory = image_path,\n                                              generator = datagen,\n                                              class_mode = NULL,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_tta_zoom <- model %>% predict_generator(tta_generator, steps=num_train_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy(pred_tta_zoom, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### What if adding with pred ? "},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy((pred_tta_zoom+pred), num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Width_shift_range"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_tta <- function(datagen, image_path, data){\n    tta_generator <- flow_images_from_dataframe(dataframe = data, \n                                              directory = image_path,\n                                              generator = datagen,\n                                              class_mode = NULL,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)\n    num_train_images<-as.numeric(dim(data)[1])\n    pred_tta <- model %>% predict_generator(tta_generator, steps=num_train_images)\n    pred_accuracy(pred_tta, num_train_images, train)\n    return(pred_tta)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n    width_shift_range = c(0.25, 0.25),\n)\nplot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred1<-test_tta(datagen, image_path, data) #keep ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n    width_shift_range = c(-0.25, -0.25),\n)\nplot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred2<-test_tta(datagen, image_path, data)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Both shift combined : "},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_tta_average<-pred1+pred2\npred_accuracy(pred_tta_average, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy(pred_tta_average+pred, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy((pred_tta_average/2)+pred, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# height_shift_range"},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n    height_shift_range = c(0.25, 0.25),\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_height1<-test_tta(datagen, image_path, data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n    height_shift_range = c(-0.25, -0.25),\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_height2<-test_tta(datagen, image_path, data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_tta_height<-pred_height1+pred_height2\npred_accuracy(pred_tta_height, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy((pred_tta_height+pred), num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy(((pred_tta_height/2)+pred), num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Shear"},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  shear_range = 20,\n)\nplot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_shear<-test_tta(datagen, image_path, data)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# brightness_range"},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  brightness_range = c(0.2, 0.2),\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_shear<-test_tta(datagen, image_path, data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  brightness_range = c(0.8, 0.8),\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_shear<-test_tta(datagen, image_path, data)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# rotation"},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  rotation_range = 45,\n)\nplot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_rotation<-test_tta(datagen, image_path, data)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Second rotation because of the randomness of the tta :"},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_rotation2<-test_tta(datagen, image_path, data)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## horizontal_flip"},{"metadata":{},"cell_type":"markdown","source":"Code from [here](https://tensorflow.rstudio.com/reference/keras/k_reverse/) and [this issue on the keras' github repository.](https://github.com/rstudio/keras/issues/225)"},{"metadata":{},"cell_type":"markdown","source":"**NB** : following datagen create an error (cf log version 16)."},{"metadata":{"trusted":true},"cell_type":"code","source":"#preprocess_input <- function(tensor){\n#  \n#    tensor<-keras::k_reverse(x=tensor, axes=2)\n#\n#  return(tensor)\n#}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#datagen <- image_data_generator(\n#   preprocessing_function = preprocess_input,\n#)\n#plot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path_horizontal <- \"/kaggle/input/dev-set-flipped/Horizontal_dev_set/Horizontal_dev_set/\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"horizontal_generator <- flow_images_from_dataframe(dataframe = data, \n                                              directory = image_path_horizontal,\n                                              class_mode = NULL,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_horizontal <- model %>% predict_generator(horizontal_generator, steps=num_train_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  zoom_range = c(0.75, 0.75)\n)\nplot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"horizontal_generator_zoom <- flow_images_from_dataframe(dataframe = data, \n                                              directory = image_path_horizontal,\n                                              generator = datagen,\n                                              class_mode = NULL,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_horizontal_zoom <- model %>% predict_generator(horizontal_generator_zoom, steps=num_train_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy(pred_horizontal_zoom, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## vertical_flip"},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path_vertical <- \"/kaggle/input/dev-set-flipped/Vertical_dev_set/Vertical_dev_set/\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"vertical_generator <- flow_images_from_dataframe(dataframe = data, \n                                              directory = image_path_vertical,\n                                              class_mode = NULL,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_vertical <- model %>% predict_generator(vertical_generator , steps=num_train_images)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Testing Accuracy on vertical and horizontal"},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy(pred_vertical, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"pred_accuracy((pred_vertical+pred), num_train_images, train)"},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy(pred_horizontal, num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy((pred_horizontal+pred), num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_accuracy((((pred_horizontal+pred_vertical)/2)+pred), num_train_images, train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Quick Conclusions"},{"metadata":{},"cell_type":"markdown","source":"I do realize that the downstream code is not really clear. It was just trial and error around the list structure of R, which is handy but painfull (like dictionnaries in Python) and quick design of repetitive actions around for loop. Just jump to the conclusion and go back to see the code if it interrest you."},{"metadata":{},"cell_type":"markdown","source":"**From version 27 :** I add somekind of grid search to ease my life.\n**From version 31 :** Plotting the second best scoring because of the randomness of some evaluation such as the rotation."},{"metadata":{"trusted":true},"cell_type":"code","source":"list_pred<-list(pred, pred_tta_zoom, pred_tta_height, pred_tta_average, pred_horizontal, pred_vertical,\n                pred_shear, pred_rotation, pred_rotation2, pred_horizontal_zoom)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"n <- length(list_pred)\nl <- rep(list(0:1), n)\n\ngrid<-expand.grid(l)\ncolnames(grid)<-c(\"pred\", \"pred_tta_zoom\", \"pred_tta_height\", \"pred_tta_average\", \"pred_horizontal\", \"pred_vertical\", \"pred_shear\", \n                  \"pred_rotation\", \"pred_rotation2\", \"pred_horizontal_zoom\")\nhead(grid,6)\ndim(grid)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Function to score the sum of the prediction for each TTA :"},{"metadata":{"trusted":true},"cell_type":"code","source":"return_pred_accuracy <- function(pred, num_train_images, train){\n    pred<-as.data.frame(pred)\n    colnames(pred)<-c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\")\n\n    label<-c()\n    for (row in 1:dim(pred)[1]){\n        label<-c(label, which(pred[row,]==max(pred[row,])))\n    }\n    label<-(label-1)\n    \n    prediction<-data.frame(train$image_id, label)\n    colnames(prediction)<-c(\"image_id\", \"label\")\n    Accuracy<-MLmetrics::Accuracy(y_true=train$label, y_pred=label)\n\n    return(Accuracy)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"accuracies<-c(0) #important, so the length of accuracies match the one of grid ! and so to read the correct","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Do the boring stuff :"},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"for (row in 2:dim(grid)[1]){\n    sum_pred<-0\n    for (i in which(grid[row,]==1)){\n      sum_pred <- sum_pred+list_pred[[i]]\n    }\n    #if (sum_pred!=0) {\n    accuracies<-c(accuracies, return_pred_accuracy(sum_pred, num_train_images, train))\n    #}\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Starting from the row 2 because one is full of zero. And the if condition would create ugly messages."},{"metadata":{"trusted":true},"cell_type":"code","source":"length(accuracies)\ndim(grid)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"max(accuracies)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"What are the best combinaisons ?"},{"metadata":{"trusted":true},"cell_type":"code","source":"grid[which(accuracies==max(accuracies)),]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"max(accuracies[accuracies!=max(accuracies)])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"grid[which(accuracies==max(accuracies[accuracies!=max(accuracies)])),]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Running prediction and Test generators"},{"metadata":{"trusted":true},"cell_type":"code","source":"test<-as.data.frame(list.files(\"/kaggle/input/cassava-leaf-disease-classification/test_images/\"))\ncolnames(test)<-\"image_id\"\ntest$image_id<-as.character(test$image_id)\nhead(test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path<-\"/kaggle/input/cassava-leaf-disease-classification/test_images\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_generator <- flow_images_from_dataframe(dataframe = test, \n                                              directory = image_path,\n                                              class_mode = NULL,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"num_test_images<-as.numeric(dim(test)[1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"num_test_images","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission<-read.csv(\"/kaggle/input/cassava-leaf-disease-classification/sample_submission.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Data generator"},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  zoom_range = c(0.75, 0.75)\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_generator_zoom <- flow_images_from_dataframe(dataframe = test, \n                                              directory = image_path,\n                                              class_mode = NULL,\n                                              generator = datagen,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  rotation_range = 45,\n)\nplot_datagen(datagen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_generator_rotation <- flow_images_from_dataframe(dataframe = test, \n                                              directory = image_path,\n                                              class_mode = NULL,\n                                              generator = datagen,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_test_generator_rotation <- model %>% predict_generator(test_generator_rotation, steps=num_test_images)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"No need to divide, the remaining of the code looks for the maximum."},{"metadata":{},"cell_type":"markdown","source":"## Adding horizontal flip"},{"metadata":{},"cell_type":"markdown","source":"The problem of the horizontal flip from the image data generator is that it is random. One way would be to writte something like this :"},{"metadata":{"trusted":true},"cell_type":"code","source":"preprocess_input <- function(tensor){\n\n    tensor<-keras::k_reverse(x=tensor, axes=2)\n\n  return(tensor)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- keras::image_data_generator(\n\n    preprocessing_function = preprocess_input,\n\n)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"But it create a C stak error : https://github.com/rstudio/keras/issues/1160. So I will just create a flipped version of each image and create a prediction on it."},{"metadata":{"trusted":true},"cell_type":"code","source":"library(magick)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dir.create(\"..//working/Horizontal/\")\npath_img<-\"../input/cassava-leaf-disease-classification/test_images/\"\nlist_img<-list.files(path_img)\nhead(list_img)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"horizontal_flip <- function(name) {\n  \n  img <- image_read(paste0(\"../input/cassava-leaf-disease-classification/test_images/\", name))\n  #print(img)\n  img <- image_flop(img)\n  #print(img)\n  infos<-image_info(img)\n  infos$width<infos$height\n  \n  if (infos$width<infos$height) {\n    img<-image_scale(img, \"x600\")\n  } else {\n    img<-image_scale(img, \"600\")\n  }\n  \n  image_write(img, path = paste0(\"..//working/Horizontal/\", name) , format = \"jpg\")\n\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Cannot flip every images without running out ressources."},{"metadata":{"trusted":true},"cell_type":"code","source":"#sapply(list_img, horizontal_flip)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test<-as.data.frame(list.files(\"/kaggle/input/cassava-leaf-disease-classification/test_images/\"))\ncolnames(test)<-\"image_id\"\ntest$image_id<-as.character(test$image_id)\nhead(test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"datagen <- image_data_generator(\n  zoom_range = c(0.75, 0.75)\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path_horizontal<-\"..//working/Horizontal/\"\ntest_generator_horizontal_zoom <- flow_images_from_dataframe(dataframe = test, \n                                              directory = image_path_horizontal,\n                                              class_mode = NULL,\n                                              generator = datagen,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(600, 600),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Out TTA"},{"metadata":{"trusted":true},"cell_type":"code","source":"pred <- model %>% predict_generator(test_generator, steps=num_test_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred_test_generator_zoom <- model %>% predict_generator(test_generator_zoom, steps=num_test_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred <- pred_test_generator_zoom + pred","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(pred)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred<-as.data.frame(pred)\ncolnames(pred)<-c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(pred)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label<-c()\nfor (row in 1:dim(pred)[1]){\n    label<-c(label, which(pred[row,]==max(pred[row,])))\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label<-(label-1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"prediction<-as.data.frame(cbind(\"id\", label))\ncolnames(prediction)<-c(\"image_id\", \"label\")\nprediction$label<-as.numeric(label)\nhead(prediction)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"prediction$image_id<-list.files(\"/kaggle/input/cassava-leaf-disease-classification/test_images/\")\n#prediction$image_id\nhead(prediction)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(pred)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"write.csv(prediction, file='submission.csv', row.names=FALSE, quote=FALSE)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Conclusion and Improvement of performance with TTA\n"},{"metadata":{},"cell_type":"markdown","source":"The Public Score of **the** EfficientNet used in this notebook is of 0.872 without TTA. I am writting this comment before the submit with all the TTA, but with some zoom added to the basic prediction I managed to reach 0.878. So I expect something similar at least."},{"metadata":{},"cell_type":"markdown","source":"The horizontal flip seems to be usefull but I need to write a specific implementation of it, since keras does not seems to allow to write the probability of the flip (so the flipping is random and not all the images are flipped. I have to import already flipped images since a datagen to flip them systematically create an error (cf version 16)."}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}