{"cells":[{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-output":true,"_kg_hide-input":true},"cell_type":"code","source":"library(data.table)\nlibrary(keras)\nlibrary(dplyr)\nlibrary(abind)\nlibrary(ggplot2)\nlibrary(gridExtra)\nlibrary(readxl)\nlibrary(skimr)\n#library(tensorflow)\n\noptions(warn=-1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"![](https://storage.googleapis.com/kaggle-competitions/kaggle/23870/logos/header.png?t=2020-12-01-04-28-05)\n\nHospital patients can have catheters and lines inserted during the course of their admission and serious complications can arise if they are positioned incorrectly. Nasogastric tube malpositioning into the airways has been reported in up to 3% of cases, with up to 40% of these cases demonstrating complications [1-3]. Airway tube malposition in adult patients intubated outside the operating room is seen in up to 25% of cases [4,5]. The likelihood of complication is directly related to both the experience level and specialty of the proceduralist. Early recognition of malpositioned tubes is the key to preventing risky complications (even death), even more so now that millions of COVID-19 patients are in need of these tubes and lines.\n\nThe gold standard for the confirmation of line and tube positions are chest radiographs. However, a physician or radiologist must manually check these chest x-rays to verify that the lines and tubes are in the optimal position. Not only does this leave room for human error, but delays are also common as radiologists can be busy reporting other scans. Deep learning algorithms may be able to automatically detect malpositioned catheters and lines. Once alerted, clinicians can reposition or remove them to avoid life-threatening complications.\n\nThe Royal Australian and New Zealand College of Radiologists (RANZCR) is a not-for-profit professional organisation for clinical radiologists and radiation oncologists in Australia, New Zealand, and Singapore. The group is one of many medical organisations around the world (including the NHS) that recognizes malpositioned tubes and lines as preventable. RANZCR is helping design safety systems where such errors will be caught.\n\nIn this competition, we’ll detect the presence and position of catheters and lines on chest x-rays. Use machine learning to train and test your model on 40,000 images to categorize a tube that is poorly placed."},{"metadata":{},"cell_type":"markdown","source":"Setting the parameters:"},{"metadata":{"trusted":true},"cell_type":"code","source":"w = 350\nh = 350\nbs = 16\nfactor = .4\nreduceFrom = 2\nearly_stop = 4\nb1 = .89\nb2 = .99\nleRa = 0.0025\nclip = 5\neps = 19\nlabSm = .001","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Image Generators"},{"metadata":{"trusted":true},"cell_type":"code","source":"newImgs <- image_data_generator(\n  width_shift_range = 0.03,\n  height_shift_range = 0.03,\n  zoom_range = 0.05,\n  rotation_range = 5,\n  shear_range = 2/180 * pi,\n  fill_mode = \"reflect\",\n  horizontal_flip = FALSE,\n  vertical_flip = FALSE,\n  samplewise_center = FALSE,\n  samplewise_std_normalization = FALSE,\n  rescale = 1/255\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"TestGenerator <- image_data_generator(\n  samplewise_center = FALSE,\n  samplewise_std_normalization = FALSE,\n  rescale = 1/255\n)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Reading the data"},{"metadata":{"trusted":true},"cell_type":"code","source":"Y = fread(\"../input/ranzcr-clip-catheter-line-classification/train.csv\", data.table = F)\nY$StudyInstanceUID <- paste(Y$StudyInstanceUID, \".jpg\", sep = \"\")\nskim(Y)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Setting the path "},{"metadata":{"trusted":true},"cell_type":"code","source":"path <- \"../input/ranzcr-clip-catheter-line-classification/train\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"image read function."},{"metadata":{"trusted":true},"cell_type":"code","source":"IMG <- function(VizImgPath, h, w){\n    VizImg <- image_load(path = VizImgPath, target_size = c(w, h))\n    VizImg <- image_to_array(img = VizImg)\n    VizImg <- VizImg %>% \n        array_reshape(c(1, w, h, 3))#shape: (samples, rows, cols, channels) if data_format='channels_last' \n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"appending the reshaped images to the array imgDF"},{"metadata":{"trusted":true},"cell_type":"code","source":"imgs <- list.files(path)[1:10]\n\nfor(i in 1:length(imgs)){\n    if(i == 1){\n        imgDF <- IMG(paste(path, \"/\", imgs[i], sep =\"\"), w, h)\n    }else{\n        imgDF <- abind(imgDF, \n                       IMG(paste(path, \"/\", imgs[i], sep =\"\"), w, h), \n                       along = 1)\n    }\n}\n\ndim(imgDF)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Creating a flow from data (in RAM) with the data generator "},{"metadata":{"trusted":true},"cell_type":"code","source":"Viz_Flow <- \n    flow_images_from_data(\n      imgDF, \n      generator = newImgs, \n      batch_size = 10 \n    )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Using the random generator to plot some new images."},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 20, \n        repr.plot.height = 20)\n\npar(mfrow=c(5,5)) \nfor(i in 1:(5*5)){\n    r = as.integer(runif(1, 1, 10))\n    plot(as.raster(generator_next(Viz_Flow)[r,,,]))\n    }","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"rm(imgDF); invisible(gc())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Create test and train flows from data"},{"metadata":{},"cell_type":"markdown","source":"To save time, we use just a small part of the dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"z = runif(nrow(Y), 0, 1)\n\nY_tr <- Y[z <= 0.85,]\nY_te <- Y[z > 0.85,]\n\nhead(Y_tr, n = 3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Label Smoothing\nAlso possible with tf\\\\$losses$CategoricalCrossentropy(label_smoothing = labSm). However accuracy is not evaluated in the metrics."},{"metadata":{"trusted":true},"cell_type":"code","source":"for(i in 2:12){\n    Y_tr[,i] <- ifelse(Y_tr[,i] == 0, labSm, (1-labSm))\n}\n\nhead(Y_tr, n = 3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"creating the flows from the folder using a data frame."},{"metadata":{"trusted":true},"cell_type":"code","source":"labels <- colnames(Y)[2:12]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"flow_train <- flow_images_from_dataframe(dataframe = Y_tr,\n                                         directory = path,\n                                         x_col = \"StudyInstanceUID\",\n                                         class_mode = \"other\",\n                                         y_col = labels,\n                                         target_size = c(w, h),\n                                         batch_size = bs,\n                                         generator = newImgs\n                                         )\nflow_train$n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"flow_test <- flow_images_from_dataframe(dataframe = Y_te,\n                                        directory = path,\n                                        x_col = \"StudyInstanceUID\",\n                                        class_mode = \"other\",\n                                        y_col = labels,\n                                        target_size = c(w, h),\n                                        batch_size = bs,\n                                        generator = TestGenerator\n                                        )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Callbacks:"},{"metadata":{"trusted":true},"cell_type":"code","source":"reduce_lr <- callback_reduce_lr_on_plateau(factor = factor, \n                                           patience = reduceFrom, \n                                           verbose = 0, \n                                           mode = \"auto\", \n                                           monitor = \"val_accuracy\")\n\nstop <- callback_early_stopping(monitor = 'val_accuracy', \n                                patience = early_stop)\n\ncheck_point <- callback_model_checkpoint(\"model.h5\", \n                                         save_best_only = TRUE, \n                                         verbose = 0, \n                                         mode = \"auto\", \n                                         save_weights_only = TRUE)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Load model "},{"metadata":{"trusted":true},"cell_type":"code","source":"x <- load_model_hdf5(\"../input/testmodel/Xavg350.h5\", compile = FALSE)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Enlarge the Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"model <- keras_model_sequential() %>%\n  x %>% \n  layer_batch_normalization() %>% \n  layer_dropout(rate=0.5) %>%\n  layer_dense(units=11, \n              activation=\"sigmoid\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(model)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model %>% compile(\n  loss = loss_binary_crossentropy,\n  optimizer = optimizer_adam(beta_1 = b1, \n                             beta_2 = b2, \n                             clipvalue = clip,\n                             lr = leRa),\n  metrics = c('accuracy')\n)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Fitting with the flows"},{"metadata":{"trusted":true},"cell_type":"code","source":"assign(paste(\"history\", 1, sep = \"\"), \n       model %>% \n          fit_generator(\n              flow_train,\n              steps_per_epoch = flow_train$n %/% bs,\n              validation_data = flow_test,\n              epochs = eps,\n              verbose = 2,\n              callbacks = c(check_point,\n                            reduce_lr, \n                            stop\n                           )\n             )\n      )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let´s check the metrics"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 16, repr.plot.height = 9)\n\nmet <- as.data.frame(unlist(history1$metrics)) %>% rename(value = `unlist(history1$metrics)`)\nmet$metrics <- gsub(\"\\\\d+\", \"\", rownames(met))\nmet$epochs <- rep(1:(nrow(met)/5), times = 5)\n\nmet <- met %>% filter(metrics %in% c(\"accuracy\", \"val_accuracy\"))\n\np1 <- ggplot(met, aes(x=epochs, y=value, color=metrics)) +\n  geom_point(size=5, \n             shape=16) + \n  theme_minimal(base_size = 15) + \n  geom_smooth(method=loess, \n              aes(fill=metrics), \n              level = 0.5,\n              formula = y ~ x\n             ) +\n  scale_color_manual(values=c(\"#808080\", \"#3f94fb\")) + \n  scale_fill_manual(values = c(\"#808080\", \"#3f94fb\")) + \n  ylab(\"Accuracy\")\n\nmet <- as.data.frame(unlist(history1$metrics)) %>% rename(value = `unlist(history1$metrics)`)\nmet$metrics <- gsub(\"\\\\d+\", \"\", rownames(met))\nmet$epochs <- rep(1:(nrow(met)/5), times = 5)\n\nmet <- met %>% filter(metrics %in% c(\"loss\", \"val_loss\"))\n\np2 <- ggplot(met, aes(x=epochs, y=value, color=metrics)) +\n  geom_point(size=5, \n             shape=16) + \n  theme_minimal(base_size = 15) + \n  geom_smooth(method=loess, \n              aes(fill=metrics), \n              level = 0.5,\n              formula = y ~ x\n             ) +\n  scale_color_manual(values=c(\"#808080\", \"#3f94fb\")) + \n  scale_fill_manual(values = c(\"#808080\", \"#3f94fb\")) + \n  ylab(\"Loss\")\n\nmet <- as.data.frame(unlist(history1$metrics)) %>% rename(value = `unlist(history1$metrics)`)\nmet$metrics <- gsub(\"\\\\d+\", \"\", rownames(met))\nmet$epochs <- rep(1:(nrow(met)/5), times = 5)\n\nmet <- met %>% filter(metrics %in% c(\"lr\"))\n\np3 <- ggplot(met, aes(x=epochs, y=value, color=metrics)) +\n  geom_line(size=2, \n            linetype=\"dotted\", \n            color=\"#808080\") + \n  theme_minimal(base_size = 15) +\n  ylab(\"Learning Rate\")","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"grid.arrange(p1, p3, p2, p3, nrow = 2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Final Prediction"},{"metadata":{"trusted":true,"_kg_hide-input":false,"_kg_hide-output":false},"cell_type":"code","source":"sample_submission <- fread(\"../input/ranzcr-clip-catheter-line-classification/sample_submission.csv\")\nsample_submission$StudyInstanceUID <- paste(sample_submission$StudyInstanceUID, \".jpg\", sep = \"\")\nhead(sample_submission)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"flow_final <- flow_images_from_dataframe(dataframe = sample_submission, \n                                         directory = \"../input/ranzcr-clip-catheter-line-classification/test\",\n                                         class_mode = NULL,\n                                         x_col = \"StudyInstanceUID\",\n                                         y_col = NULL,\n                                         target_size = c(w, h),\n                                         shuffle = FALSE,\n                                         batch_size=1,\n                                         generator = TestGenerator\n                                         )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"load_model_weights_hdf5(model, \"model.h5\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred <- model %>% \n    predict_generator(flow_final, \n                      steps = nrow(sample_submission))\n\nhead(pred)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred <- pred %>% as.data.frame()\ncolnames(pred) = labels\npred$StudyInstanceUID <- flow_final$filenames\n\npred <- pred[,c(12,1:11)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Submit"},{"metadata":{"trusted":true},"cell_type":"code","source":"fwrite(pred, \"submission.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### <center>If you like the kernel, give it an upvote:</center>\n<center><a href=\"#top\" class=\"btn btn-info btn-lg active\" role=\"button\" aria-pressed=\"true\" style=\"color:white\" data-toggle=\"popover\">Go back to the TOP</a></center>"}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}