{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Noisy Student weights "},{"metadata":{},"cell_type":"markdown","source":"I already made a kernel with an EfficientNetB0 [here](https://www.kaggle.com/cdk292/efficientnetb0-with-r-and-tf2-cyclic-lr). What is different in this one ?\n\nNo so much actually. I wanted to use the weights **noisy-student** such as in this notebook : https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods. In a nutshell, it is just differents initial weights than the one from imagenet for the model architecture.  It is an approach used by some winning competition of the **Plant Pathology 2020** and seems to perform better.\n\nIt turns out that in Python, keras does not allow to use them directly, and the application_ wrapper cannot use this weights (even if you try to load the h5 weights, you will have an error based on the numbers of layers). Python users usually install the package **efficientnet** to run efficientnet, in our current days. \n\nSo this notebook is essentially a technical demonstration to spare time to a R user. Here we install the python library with reticulate, and then use reticulate to import the library **efficientnet**.\n\nMore importantly, I cannot use the sequential API to create the model here. I use the more recent code from this [tutorial on the rstudio blog](https://tensorflow.rstudio.com/guide/tfhub/intro/) and write something like : model <- keras_model(input, output) instead of model <- keras_model_sequential() %>% ...\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"#reticulate::py_install(packages = \"tensorflow\", version = \"2.3.0\", pip=TRUE)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"library(tidyverse)\nlibrary(tensorflow)\ntf$executing_eagerly()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tensorflow::tf_version()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"library(keras)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels<-read_csv('/kaggle/input/cassava-leaf-disease-classification//train.csv')\nhead(labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"levels(as.factor(labels$label))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"idx0<-which(labels$label==0)\nidx1<-which(labels$label==1)\nidx2<-which(labels$label==2)\nidx3<-which(labels$label==3)\nidx4<-which(labels$label==4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels$CBB<-0\nlabels$CBSD<-0\nlabels$CGM<-0\nlabels$CMD<-0\nlabels$Healthy<-0","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels$CBB[idx0]<-1\nlabels$CBSD[idx1]<-1\nlabels$CGM[idx2]<-1\nlabels$CMD[idx3]<-1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"\"Would it have been easier to create a function to convert the labelling ?\" You may ask."},{"metadata":{"trusted":true},"cell_type":"code","source":"labels$Healthy[idx4]<-1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Probably."},{"metadata":{"trusted":true},"cell_type":"code","source":"#labels$label<-NULL","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(labels)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Following code is retaken from one fork of this notebook, [simple-convnet](https://www.kaggle.com/demetrypascal/simple-convnet), which use a better approach to create a validation set (not at random, but with stratification) (see version 31 for previous approach) :"},{"metadata":{"trusted":true},"cell_type":"code","source":"set.seed(6)\n\ntmp = splitstackshape::stratified(labels, c('label'), 0.9, bothSets = TRUE)\n\ntrain_labels = tmp[[1]]\nval_labels = tmp[[2]]\n\nwrite.csv(val_labels, file='validation_set.csv', row.names=FALSE, quote=FALSE)\n\n\ntrain_labels$label<-NULL\nval_labels$label<-NULL\n\nhead(train_labels)\nhead(val_labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(train_labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(val_labels)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path<-'/kaggle/input/cassava-leaf-disease-classification//train_images'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#data augmentation\ndatagen <- image_data_generator(\n  rotation_range = 40,\n  width_shift_range = 0.2,\n  height_shift_range = 0.2,\n  shear_range = 0.2,\n  zoom_range = 0.5,\n  horizontal_flip = TRUE,\n  fill_mode = \"reflect\"\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_path<-\"/kaggle/input/cassava-leaf-disease-classification/train_images/1000015157.jpg\"\n\nimg <- image_load(img_path, target_size = c(448, 448))\nimg_array <- image_to_array(img)\nimg_array <- array_reshape(img_array, c(1, 448, 448, 3))\nimg_array<-img_array/255\n# Generated that will flow augmented images\naugmentation_generator <- flow_images_from_data(\n  img_array, \n  generator = datagen, \n  batch_size = 1 \n)\nop <- par(mfrow = c(2, 2), pty = \"s\", mar = c(1, 0, 1, 0))\nfor (i in 1:4) {\n  batch <- generator_next(augmentation_generator)\n  plot(as.raster(batch[1,,,]))\n}\npar(op)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I create the flow_images_from_dataframe like in a previous competition, but maybe you can skip the conversion of the label into 1 and 0 and directly create train generator from the original label column of the dataframe."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_generator <- flow_images_from_dataframe(dataframe = train_labels, \n                                              directory = image_path,\n                                              generator = datagen,\n                                              class_mode = \"other\",\n                                              x_col = \"image_id\",\n                                              y_col = c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\"),\n                                              target_size = c(600, 600),\n                                              batch_size=8)\n\nvalidation_generator <- flow_images_from_dataframe(dataframe = val_labels, \n                                              directory = image_path,\n                                              class_mode = \"other\",\n                                              x_col = \"image_id\",\n                                              y_col = c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\"),\n                                              target_size = c(600, 600),\n                                              batch_size=8)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"library(reticulate)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"reticulate::py_install(\"efficientnet\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#reticulate::virtualenv_install(\"efficientnet\", envname = \"r-reticulate\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"efn <- import(\"efficientnet.keras\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-training-with-tpu-v2-pods/comments"},{"metadata":{"trusted":true},"cell_type":"code","source":"conv_base <- efn$EfficientNetB7(weights=\"noisy-student\", include_top=FALSE)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Input layer missing"},{"metadata":{"trusted":true},"cell_type":"code","source":"#conv_base","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"input <- layer_input(shape = c(600, 600, 3))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"freeze_weights(conv_base)\n\noutput <- input %>%\n    conv_base %>% \n    layer_global_max_pooling_2d() %>% \n    layer_batch_normalization() %>% \n    layer_dropout(rate=0.5) %>%\n    layer_dense(units=5, activation=\"softmax\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model <- keras_model(input, output)\n\nsummary(model)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Cyclical learning rate"},{"metadata":{},"cell_type":"markdown","source":"A lot of the code around came from the blog [\"the cool data\"](http://thecooldata.com/) if my memory is still correct. The idea to have a tail and the notion of annhilation of the gradient can be found [here, on The 1cycle policy](https://sgugger.github.io/the-1cycle-policy.html). The big difference is that I do not want to add an other vector of even lower learning rate at the end of the one generated by the function Cyclic_lr, it would force me to take it into account and create an other number of iteration for the compilation of the model. I prefer the approach of dividing more and more the last element of the cycle."},{"metadata":{"trusted":true},"cell_type":"code","source":"callback_lr_init <- function(logs){\n      iter <<- 0\n      lr_hist <<- c()\n      iter_hist <<- c()\n}\ncallback_lr_set <- function(batch, logs){\n      iter <<- iter + 1\n      LR <- l_rate[iter] # if number of iterations > l_rate values, make LR constant to last value\n      if(is.na(LR)) LR <- l_rate[length(l_rate)]\n      k_set_value(model$optimizer$lr, LR)\n}\n\ncallback_lr <- callback_lambda(on_train_begin=callback_lr_init, on_batch_begin=callback_lr_set)\n","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"####################\nCyclic_LR <- function(iteration=1:32000, base_lr=1e-5, max_lr=1e-3, step_size=2000, mode='triangular', gamma=1, scale_fn=NULL, scale_mode='cycle'){ # translated from python to R, original at: https://github.com/bckenstler/CLR/blob/master/clr_callback.py # This callback implements a cyclical learning rate policy (CLR). # The method cycles the learning rate between two boundaries with # some constant frequency, as detailed in this paper (https://arxiv.org/abs/1506.01186). # The amplitude of the cycle can be scaled on a per-iteration or per-cycle basis. # This class has three built-in policies, as put forth in the paper. # - \"triangular\": A basic triangular cycle w/ no amplitude scaling. # - \"triangular2\": A basic triangular cycle that scales initial amplitude by half each cycle. # - \"exp_range\": A cycle that scales initial amplitude by gamma**(cycle iterations) at each cycle iteration. # - \"sinus\": A sinusoidal form cycle # # Example # > clr <- Cyclic_LR(base_lr=0.001, max_lr=0.006, step_size=2000, mode='triangular', num_iterations=20000) # > plot(clr, cex=0.2)\n \n      # Class also supports custom scaling functions with function output max value of 1:\n      # > clr_fn <- function(x) 1/x # > clr <- Cyclic_LR(base_lr=0.001, max_lr=0.006, step_size=400, # scale_fn=clr_fn, scale_mode='cycle', num_iterations=20000) # > plot(clr, cex=0.2)\n \n      # # Arguments\n      #   iteration:\n      #       if is a number:\n      #           id of the iteration where: max iteration = epochs * (samples/batch)\n      #       if \"iteration\" is a vector i.e.: iteration=1:10000:\n      #           returns the whole sequence of lr as a vector\n      #   base_lr: initial learning rate which is the\n      #       lower boundary in the cycle.\n      #   max_lr: upper boundary in the cycle. Functionally,\n      #       it defines the cycle amplitude (max_lr - base_lr).\n      #       The lr at any cycle is the sum of base_lr\n      #       and some scaling of the amplitude; therefore \n      #       max_lr may not actually be reached depending on\n      #       scaling function.\n      #   step_size: number of training iterations per\n      #       half cycle. Authors suggest setting step_size\n      #       2-8 x training iterations in epoch.\n      #   mode: one of {triangular, triangular2, exp_range, sinus}.\n      #       Default 'triangular'.\n      #       Values correspond to policies detailed above.\n      #       If scale_fn is not None, this argument is ignored.\n      #   gamma: constant in 'exp_range' scaling function:\n      #       gamma**(cycle iterations)\n      #   scale_fn: Custom scaling policy defined by a single\n      #       argument lambda function, where \n      #       0 <= scale_fn(x) <= 1 for all x >= 0.\n      #       mode paramater is ignored \n      #   scale_mode: {'cycle', 'iterations'}.\n      #       Defines whether scale_fn is evaluated on \n      #       cycle number or cycle iterations (training\n      #       iterations since start of cycle). Default is 'cycle'.\n \n      ########\n      if(is.null(scale_fn)==TRUE){\n            if(mode=='triangular'){scale_fn <- function(x) 1; scale_mode <- 'cycle';}\n            if(mode=='triangular2'){scale_fn <- function(x) 1/(2^(x-1)); scale_mode <- 'cycle';}\n            if(mode=='exp_range'){scale_fn <- function(x) gamma^(x); scale_mode <- 'iterations';}\n            if(mode=='sinus'){scale_fn <- function(x) 0.5*(1+sin(x*pi/2)); scale_mode <- 'cycle';}\n            if(mode=='halfcosine'){scale_fn <- function(x) 0.5*(1+cos(x*pi)^2); scale_mode <- 'cycle';}\n      }\n      lr <- list()\n      if(is.vector(iteration)==TRUE){\n            for(iter in iteration){\n                  cycle <- floor(1 + (iter / (2*step_size)))\n                  x2 <- abs(iter/step_size-2 * cycle+1)\n                  if(scale_mode=='cycle') x <- cycle\n                  if(scale_mode=='iterations') x <- iter\n                  lr[[iter]] <- base_lr + (max_lr-base_lr) * max(0,(1-x2)) * scale_fn(x)\n            }\n      }\n      lr <- do.call(\"rbind\",lr)\n      return(as.vector(lr))\n}\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### The tail"},{"metadata":{},"cell_type":"markdown","source":"Okay, what is going on here ? Simple speaking I want the last steps to go several order of magnitude under the minimal learning rate, in a similar fashion to the fast.ai implementation. The most elegant way (without adding an other vector at the end) to do this is to divide the learning rate of the last steps by a number growing exponentially (to avoid a cut in the learning rate curve by dividing the number suddenly by 10). So we have a nice “tail” (see graphs below).\n\nOh there is no specific justifications for the exponent number. Just trial and error."},{"metadata":{"trusted":true},"cell_type":"code","source":"n=300\nnb_epochs <- 15","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tail <- 30 #annhilation of the gradient\ni<-1:tail\nl_rate_div<-1.1*(1.2^i) \nplot(l_rate_div, type=\"b\", pch=16, cex=0.1, xlab=\"iteration\", ylab=\"learning rate dividor\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"l_rate_cyclical <- Cyclic_LR(iteration=1:n, base_lr=1e-7, max_lr=1e-3, step_size=floor(n/2),\n                        mode='triangular', gamma=1, scale_fn=NULL, scale_mode='cycle')\n\nstart_tail <-length(l_rate_cyclical)-tail\nend_tail <- length(l_rate_cyclical)\nl_rate_cyclical[start_tail:end_tail] <- l_rate_cyclical[start_tail:end_tail]/l_rate_div","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"l_rate <- rep(l_rate_cyclical, nb_epochs)\n\nplot(l_rate_cyclical, type=\"b\", pch=16, xlab=\"iteration\", cex=0.2, ylab=\"learning rate\", col=\"grey50\")\n\nplot(l_rate, type=\"b\", pch=16, xlab=\"iteration\", cex=0.2, ylab=\"learning rate\", col=\"grey50\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"##### V7 onwards : adam"},{"metadata":{"trusted":true},"cell_type":"code","source":"model %>% compile(\n    optimizer=optimizer_adam(),\n    loss=\"categorical_crossentropy\",\n    metrics = \"categorical_accuracy\"\n)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The following code came from the tutorial of Keras \"tutorial_save_and_restore\"."},{"metadata":{"trusted":true},"cell_type":"code","source":"checkpoint_dir <- \"checkpoints\"\nunlink(checkpoint_dir, recursive = TRUE)\ndir.create(checkpoint_dir)\nfilepath <- file.path(checkpoint_dir, \"eff_net_weights.{epoch:02d}.hdf5\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"check_point_callback <- callback_model_checkpoint(\n  filepath = filepath,\n  save_weights_only = TRUE,\n  save_best_only = TRUE\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"callback_list<-list(callback_lr, check_point_callback ) #callback to update lr","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history <- model %>% fit_generator(\n    train_generator,\n    steps_per_epoch=n,\n    epochs = nb_epochs,\n    callbacks = callback_list, #callback to update cylic lr\n    validation_data = validation_generator,\n    validation_step=40\n)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot(history)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Fine tuning"},{"metadata":{},"cell_type":"markdown","source":"Here their is some code to load weight due to an update inside h5py. Details [here](https://www.kaggle.com/product-feedback/198871)."},{"metadata":{"trusted":true},"cell_type":"code","source":"library(reticulate)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#reticulate::virtualenv_remove(packages=\"h5py\", envname = \"r-reticulate\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#reticulate::virtualenv_install(\"h5py\", version = \"2.1.0\", envname = \"r-reticulate\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I don't want to load a model at random and take the risk to have the notebook fail compilation, but you can uncomment this one to load a model :"},{"metadata":{"trusted":true},"cell_type":"code","source":"#model %>% load_model_weights_hdf5(file.path(checkpoint_dir,\"eff_net_weights.04.hdf5\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So my understanding is that you cannot unfreeze from a specific part of the hub_layer, unlike what you can do with the conv_base CNN imported by a code like keras::application_resnet50(). This is not without problems, since it can degrade the really basic feature the networks has learned, but also create some memory troubles (lot of parameters to learn)."},{"metadata":{},"cell_type":"markdown","source":"In definitive I would say that the wrapper for efficient net and using a sequential model is far more effective and convenient."},{"metadata":{"trusted":true},"cell_type":"code","source":"#unfreeze_weights(hub_layer, from = 'block5a_expand_conv')\n#Error in py_get_attr_impl(x, name, silent): AttributeError: 'KerasLayer' object has no attribute 'layers'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"unfreeze_weights(conv_base, from = 'block5a_expand_conv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(model)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"nb_epochs<-30","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"l_rate_cyclical <- Cyclic_LR(iteration=1:n, base_lr=1e-7, max_lr=(1e-3/5), step_size=floor(n/2),\n                        mode='triangular', gamma=1, scale_fn=NULL, scale_mode='cycle')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"start_tail <-length(l_rate_cyclical)-tail\nend_tail <- length(l_rate_cyclical)\nl_rate_cyclical[start_tail:end_tail] <- l_rate_cyclical[start_tail:end_tail]/l_rate_div\n\nl_rate <- rep(l_rate_cyclical, nb_epochs)\n\nplot(l_rate, type=\"b\", pch=16, xlab=\"iteration\", cex=0.2, ylab=\"learning rate\", col=\"grey50\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model %>% compile(\n    optimizer=optimizer_adam(),\n    loss=\"categorical_crossentropy\",\n    metrics = \"categorical_accuracy\"\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"checkpoint_dir <- \"checkpoints_fine_tuned\"\nunlink(checkpoint_dir, recursive = TRUE)\ndir.create(checkpoint_dir)\nfilepath <- file.path(checkpoint_dir, \"fine_tuned_eff_net_weights.{epoch:02d}.hdf5\")\n\ncheck_point_callback <- callback_model_checkpoint(\n  filepath = filepath,\n  save_weights_only = TRUE,\n  save_best_only = TRUE\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"callback_list<-list(callback_lr, check_point_callback )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history <- model %>% fit_generator(\n    train_generator,\n    steps_per_epoch=n,\n    epochs = nb_epochs,\n    callbacks = callback_list, #callback to update cylic lr\n    validation_data = validation_generator,\n    validation_step=40\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot(history)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Save model "},{"metadata":{"trusted":true},"cell_type":"code","source":"model %>% save_model_tf(\"EfficientB7\")","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}