{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Missclassified pictures"},{"metadata":{},"cell_type":"markdown","source":"**Goal :** The goal of this notebook is to **mostly to find** which picture are badly classified by the model passed as input. Here the model is an EfficientNetB0 from this [notebook](https://www.kaggle.com/cdk292/efficientnetb0-with-r-and-tf2-cyclic-lr/)."},{"metadata":{},"cell_type":"markdown","source":"A also try to see if the EfficientNet is better to classify or missclassify one type of labels. To do so I am using a **Chi-Square Test**. There is certainly other biais to assess other than just the frequency of the labels correctly classified or missclassified (such as light exposure for example). They are not explored here."},{"metadata":{},"cell_type":"markdown","source":"I don't have ground truth citation or example to provide to support my approach, maybe it is wrong. Please tell me if so."},{"metadata":{},"cell_type":"markdown","source":"### Danger Zone \n\nShould we looks for the picture missclassified ? Recalling the course of the deep learning specialization of Andrew Ng, we could, to see if there is no caveat in the pictures that are missclassified. But there is also the danger of somehow introducing our own biais in the conception of the CNN, especially if we look at our dev set. "},{"metadata":{},"cell_type":"markdown","source":"First I install tf 2.3 to be able to use the efficient net wrappers from keras."},{"metadata":{"trusted":true},"cell_type":"code","source":"reticulate::py_install(packages = \"tensorflow\", version = \"2.3.0\", pip=TRUE)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"library(tidyverse)\nlibrary(tensorflow)\ntf$executing_eagerly()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tensorflow::tf_version()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here I flex with my own version of keras. Basically, it is a fork with application wrapper for the efficient net."},{"metadata":{},"cell_type":"markdown","source":"**Disclaimer : I did not writte the code for the really handy applications wrappers.** It came [from this commit](https://github.com/rstudio/keras/commit/c406ec55f7bb2864ac58a17f963448810a531c18) for which the PR is hold until the fully release of tf 2.3, as stated [in this PR](https://github.com/rstudio/keras/pull/1097). I am not sure why the PR is closed."},{"metadata":{"trusted":true},"cell_type":"code","source":"devtools::install_github(\"Cdk29/keras\", dependencies = FALSE)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"library(keras)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data<-read_csv('/kaggle/input/cassava-leaf-disease-classification//train.csv')\nhead(data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path<-'/kaggle/input/cassava-leaf-disease-classification//train_images'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"With shuffle = FALSE to not mix images and got the right order in the predictions."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_generator <- flow_images_from_dataframe(dataframe = data, \n                                              directory = image_path,\n                                              class_mode = NULL,\n                                              x_col = \"image_id\",\n                                              y_col = NULL,\n                                              target_size = c(448, 448),\n                                              shuffle = FALSE,\n                                              batch_size=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Loading the model structure"},{"metadata":{},"cell_type":"markdown","source":"This notebook is forked from the EfficientNetB0 with R and tf2, cyclic lr. "},{"metadata":{"trusted":true},"cell_type":"code","source":"conv_base<-keras::application_efficientnet_b0(weights = \"imagenet\", include_top = FALSE, input_shape = c(448, 448, 3))\n\nfreeze_weights(conv_base)\n\nmodel <- keras_model_sequential() %>%\n    conv_base %>% \n    layer_global_max_pooling_2d() %>% \n    layer_batch_normalization() %>% \n    layer_dropout(rate=0.5) %>%\n    layer_dense(units=5, activation=\"softmax\")\n\nsummary(model)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load the weight from a model"},{"metadata":{},"cell_type":"markdown","source":"The weight came from this [notebook](https://www.kaggle.com/cdk292/efficientnetb0-with-r-and-tf2-cyclic-lr/), passed as input of this notebook."},{"metadata":{"trusted":true},"cell_type":"code","source":"library(reticulate)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"reticulate::virtualenv_remove(packages=\"h5py\", envname = \"r-reticulate\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"reticulate::virtualenv_install(\"h5py\", version = \"2.1.0\", envname = \"r-reticulate\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"unfreeze_weights(conv_base, from = 'block5a_expand_conv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model %>% load_model_weights_hdf5(\"/kaggle/input/efficientnetb0-with-r-and-tf2-cyclic-lr/checkpoints_fine_tuned/fine_tuned_eff_net_weights.18.hdf5\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Predict"},{"metadata":{"trusted":true},"cell_type":"code","source":"num_test_images<-as.numeric(dim(data)[1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred <- model %>% predict_generator(train_generator, steps=num_test_images)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred<-as.data.frame(pred)\ncolnames(pred)<-c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label<-c()\nfor (row in 1:dim(pred)[1]){\n    label<-c(label, which(pred[row,]==max(pred[row,])))\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Comparison of predicted labels on train vs train label"},{"metadata":{},"cell_type":"markdown","source":"The goal here is to create a table with the labels of the classified and missclassified pictures and test if one proportion of label is different between the correctly classfied and not correctly classified pictures."},{"metadata":{"trusted":true},"cell_type":"code","source":"train<-read_csv('/kaggle/input/cassava-leaf-disease-classification//train.csv')\nhead(train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Correction to take into account the fact that there is five columns in pred :"},{"metadata":{"trusted":true},"cell_type":"code","source":"label<-(label-1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"prediction<-data.frame(train$image_id, label)\ncolnames(prediction)<-c(\"image_id\", \"label\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(prediction)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"length(which(train$label!=prediction$label))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Creating the list of missclassified pictures :"},{"metadata":{"trusted":true},"cell_type":"code","source":"idx_img<-which(train$label!=prediction$label)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Creating a table of missclassified labels :"},{"metadata":{"trusted":true},"cell_type":"code","source":"missclassified<-train[idx_img,]$label\ntable(missclassified)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"correctly_classified<-train[-idx_img,]$label\ntable(correctly_classified)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"missclassified<-c(489, 520, 671, 370, 609)\ncorrectly_classified<-c(598, 1669, 1715, 12788, 1968)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"proportion_df<-data.frame(correctly_classified, missclassified)\nproportion_df<-t(proportion_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"rownames(proportion_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"colnames(proportion_df)<-c(\"CBB\",\"CBSD\", \"CGM\", \"CMD\", \"Healthy\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"proportion_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Statistical test"},{"metadata":{},"cell_type":"markdown","source":"Correct me if I am wrong, but looking at the data (frequency table), one possibility we have is to test for an influence of the two variables (classified vs missclassified) using the chi-squared test."},{"metadata":{},"cell_type":"markdown","source":"I retake the definition of the test from [sthda](http://www.sthda.com/english/wiki/chi-square-test-of-independence-in-r) : *The chi-square test of independence is used to analyze the **frequency table** (i.e. contengency table) formed by **two categorical variables**. The chi-square test evaluates whether there is a significant association between the categories of the two variables.*"},{"metadata":{"trusted":true},"cell_type":"code","source":"res<-chisq.test(proportion_df)\nres","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"res$observed","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"What should be expected if there was not impact of the correcthness of the classification?"},{"metadata":{"trusted":true},"cell_type":"code","source":"res$expected","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Oh, actually I am surprised. At first glanced I focused on the results of CMD that are quite obvious (but does not tell us anything)."},{"metadata":{},"cell_type":"markdown","source":"What does it say about potential biais of our EfficientNet ? Potentially :\n\nIt seems to be bad at classifying correctly **CBB**, **CBSD** and **CGM**. The EfficientNet-B0 seems quite good to classify **CMD** and **Healthy**. And the **CMD** are the most abundant, explaining in a good part the really good results of the EfficientNet-B0."},{"metadata":{},"cell_type":"markdown","source":"## Solution ?"},{"metadata":{},"cell_type":"markdown","source":"In this notebook we looked there is one type of sample missclassified (one type of label). Turned out there is several. We did not check if this subset of missclassified images shared some particularity. Maybe there is. We could projecting them on a space reduction algorithm, like the [spectral clustering of this kernel](https://www.kaggle.com/foolofatook/starter-eda-cassava-leaf-disease), or an other one, to see if they cluster together in a unsupervised manner."},{"metadata":{},"cell_type":"markdown","source":"Maybe a good approach would be to make a ensemble with a CNN that is good (or biaised) to correctly classify CBB, CBSD and CGM. Some winning solution use ensembling between Efficientnet and Seresnextnet, [even if the winning solution didn't](https://www.kaggle.com/c/plant-pathology-2020-fgvc7/discussion/154056). Maybe train a model specifically for this labels. "},{"metadata":{},"cell_type":"markdown","source":"# About the Statistics"},{"metadata":{},"cell_type":"markdown","source":"We could go further and more in details of the analyses of the classification using the One-Proportion Z-Test, I believe. I concluded that the EfficientNet is slightly less good to classifiy some labels such as CBB by looking at the output of the Chi-test, but actually I did not try to test the significance of each ratio."},{"metadata":{},"cell_type":"markdown","source":"Is there better approach ? Probably. If you know one, tell me in the comments."}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}