{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Efficientnet with R and Tf2"},{"metadata":{},"cell_type":"markdown","source":"This notebook contain the code for the learning rate finder for this [notebook](https://www.kaggle.com/cdk292/efficientnetb0-with-r-and-tf2-cyclic-lr). The notebook is forked from it for convenience, so a lot of commentary will be the same."},{"metadata":{},"cell_type":"markdown","source":"First I install tf 2.3 to be able to use the efficient net wrappers from keras."},{"metadata":{"trusted":true},"cell_type":"code","source":"reticulate::py_install(packages = \"tensorflow\", version = \"2.3.0\", pip=TRUE)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"library(tidyverse)\nlibrary(tensorflow)\ntf$executing_eagerly()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tensorflow::tf_version()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here I flex with my own version of keras. Basically, it is a fork with application wrapper for the efficient net."},{"metadata":{},"cell_type":"markdown","source":"**Disclaimer : I did not writte the code for the really handy applications wrappers.** It came [from this commit](https://github.com/rstudio/keras/commit/c406ec55f7bb2864ac58a17f963448810a531c18) for which the PR is hold until the fully release of tf 2.3, as stated [in this PR](https://github.com/rstudio/keras/pull/1097). I am not sure why the PR is closed."},{"metadata":{"trusted":true},"cell_type":"code","source":"devtools::install_github(\"Cdk29/keras\", dependencies = FALSE)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"library(keras)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels<-read_csv('/kaggle/input/cassava-leaf-disease-classification//train.csv')\nhead(labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"levels(as.factor(labels$label))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"idx0<-which(labels$label==0)\nidx1<-which(labels$label==1)\nidx2<-which(labels$label==2)\nidx3<-which(labels$label==3)\nidx4<-which(labels$label==4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels$CBB<-0\nlabels$CBSD<-0\nlabels$CGM<-0\nlabels$CMD<-0\nlabels$Healthy<-0","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels$CBB[idx0]<-1\nlabels$CBSD[idx1]<-1\nlabels$CGM[idx2]<-1\nlabels$CMD[idx3]<-1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"\"Would it have been easier to create a function to convert the labelling ?\" You may ask."},{"metadata":{"trusted":true},"cell_type":"code","source":"labels$Healthy[idx4]<-1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Probably."},{"metadata":{"trusted":true},"cell_type":"code","source":"labels$label<-NULL","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"set.seed(6)\n\nlabels <- labels  %>% mutate(id = row_number())#Check IDs\n\ntrain_labels <- labels  %>% sample_frac(.90)#Create test set\nval_labels <- anti_join(labels, train_labels, by = 'id')\ntrain_labels$id<-NULL\nval_labels$id<-NULL\n\nhead(train_labels)\nhead(val_labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(train_labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(val_labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path<-'/kaggle/input/cassava-leaf-disease-classification//train_images'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#data augmentation\ndatagen <- image_data_generator(\n  rotation_range = 40,\n  width_shift_range = 0.2,\n  height_shift_range = 0.2,\n  shear_range = 0.2,\n  zoom_range = 0.5,\n  horizontal_flip = TRUE,\n  fill_mode = \"reflect\"\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_path<-\"/kaggle/input/cassava-leaf-disease-classification/train_images/1000015157.jpg\"\n\nimg <- image_load(img_path, target_size = c(448, 448))\nimg_array <- image_to_array(img)\nimg_array <- array_reshape(img_array, c(1, 448, 448, 3))\nimg_array<-img_array/255\n# Generated that will flow augmented images\naugmentation_generator <- flow_images_from_data(\n  img_array, \n  generator = datagen, \n  batch_size = 1 \n)\nop <- par(mfrow = c(2, 2), pty = \"s\", mar = c(1, 0, 1, 0))\nfor (i in 1:4) {\n  batch <- generator_next(augmentation_generator)\n  plot(as.raster(batch[1,,,]))\n}\npar(op)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I create the flow_images_from_dataframe like in a previous competition, but maybe you can skip the conversion of the label into 1 and 0 and directly create train generator from the original label column of the dataframe."},{"metadata":{},"cell_type":"markdown","source":"The batch size is important for the structure of the curve of the learning rate finder. The bigger the batch size is the more easy to interpret is the curve for the learning rate finder."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_generator <- flow_images_from_dataframe(dataframe = train_labels, \n                                              directory = image_path,\n                                              generator = datagen,\n                                              class_mode = \"other\",\n                                              x_col = \"image_id\",\n                                              y_col = c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\"),\n                                              target_size = c(448, 448),\n                                              batch_size=32)\n\nvalidation_generator <- flow_images_from_dataframe(dataframe = val_labels, \n                                              directory = image_path,\n                                              class_mode = \"other\",\n                                              x_col = \"image_id\",\n                                              y_col = c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\"),\n                                              target_size = c(448, 448),\n                                              batch_size=16)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"conv_base<-keras::application_efficientnet_b0(weights = \"imagenet\", include_top = FALSE, input_shape = c(448, 448, 3))\n\nfreeze_weights(conv_base)\n\nmodel <- keras_model_sequential() %>%\n    conv_base %>% \n    layer_global_max_pooling_2d() %>% \n    layer_batch_normalization() %>% \n    layer_dropout(rate=0.5) %>%\n    layer_dense(units=5, activation=\"softmax\")\n\nsummary(model)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Learning rate finder"},{"metadata":{},"cell_type":"markdown","source":"A lot of the code around came from the blog [thecooldata](http://thecooldata.com/). The code has been modified to plot the loss in the y axis against the learning rate in a similar fashion to fast.ai. I don't like to use weighted average, I prefer simply plotting the loess of the point."},{"metadata":{"trusted":true},"cell_type":"code","source":"LogMetrics <- R6::R6Class(\"LogMetrics\",\n  inherit = KerasCallback,\n  public = list(\n    loss = NULL,\n    acc = NULL,\n    on_batch_end = function(batch, logs=list()) {\n      self$loss <- c(self$loss, logs[[\"loss\"]])\n      self$acc <- c(self$acc, logs[[\"acc\"]])\n    }\n))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"callback_lr_init <- function(logs){\n      iter <<- 0\n      lr_hist <<- c()\n      iter_hist <<- c()\n}\ncallback_lr_set <- function(batch, logs){\n      iter <<- iter + 1\n      LR <- l_rate[iter] # if number of iterations > l_rate values, make LR constant to last value\n      if(is.na(LR)) LR <- l_rate[length(l_rate)]\n      k_set_value(model$optimizer$lr, LR)\n}\ncallback_lr_log <- function(batch, logs){\n      lr_hist <<- c(lr_hist, k_get_value(model$optimizer$lr))\n      iter_hist <<- c(iter_hist, k_get_value(model$optimizer$iterations))\n}\ncallback_lr <- callback_lambda(on_train_begin=callback_lr_init, on_batch_begin=callback_lr_set)\n\ncallback_logger <- callback_lambda(on_batch_end=callback_lr_log)\n\ncallback_log_acc_lr <- LogMetrics$new()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Settings\n\nIt is also important to test for low value of learning rate decay, since a annhilation of the gradient is also a part of the learning rate finder."},{"metadata":{"trusted":true},"cell_type":"code","source":"lr0<-1e-12\n#lr_max<-0.01\nlr_max<-0.1\n\n#n_iteration :\nn <- 360 #from 120 to 360\nq<-(lr_max/lr0)^(1/(n-1))\n###Plot and lr finder\n\ni<-1:n\nl_rate<-lr0*(q^i) #formula from the blog article of fastai.\nplot(l_rate, type=\"b\", pch=16, cex=0.1, xlab=\"iteration\", ylab=\"learning rate\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model %>% compile(\n    optimizer=optimizer_rmsprop(lr=lr_max),\n    loss=\"categorical_crossentropy\",\n    metrics='categorical_accuracy'\n)\n\ncallback_list = list(callback_lr, callback_logger, callback_log_acc_lr)\n\nhistory <- model %>% fit_generator(\n    train_generator,\n    steps_per_epoch=n,\n    epochs = 1,\n    callbacks = callback_list,\n    validation_data = validation_generator,\n    validation_step=1\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data <- data.frame(\"Learning_rate\" = lr_hist, \"Loss\" = callback_log_acc_lr$loss)\nhead(data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x_breaks = 10^pretty(log10(data$Learning_rate))\nggplot(data, aes(x=Learning_rate, y=Loss)) + scale_x_log10(breaks=x_breaks) + geom_point() +  geom_smooth(span = 0.3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From the graph above we can see which learning rate trigger the biggest derivative of loss ;p"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}