{"cells":[{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"# This R environment comes with many helpful analytics packages installed\n# It is defined by the kaggle/rstats Docker image: https://github.com/kaggle/docker-rstats\n# For example, here's a helpful package to load\n\nlibrary(tidyverse) # metapackage of all tidyverse packages\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nlist.files(path = \"../input\")\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\nlibrary(keras)\nlibrary(tidyverse)\nlibrary(tensorflow)\nlibrary(tictoc)\nlibrary(RColorBrewer)\n\nmaster_dataset_dir <- \"../input/siim-isic-melanoma-classification/\"\n\ntrain_csv_data <- read.csv(\"../input/siim-isic-melanoma-classification/train.csv\")\nimage_names <- as.matrix(train_csv_data[, 1])\n\ntest_csv_data <- read.csv(\"../input/siim-isic-melanoma-classification/test.csv\")\ntest_image_names <- as.matrix(test_csv_data[, 1])\n\n\ndir.create(\"../UnderSamplingData\")\nnew_data_location <- \"../UnderSamplingData\"\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the above section, I am creating a seprate folder so that I can copy some images to that folder. Rather than working on the entire training dataset, I am planning on doing undersampling because of the imbalanced dataset. Some basic eda will also be shown below.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"warning = FALSE\n\n# list of melanomas to modle\nmelanoma_list <- c(\"Yes\", \"No\")\n\n# number of output classes (i.e. fruits)\noutput_n <- length(melanoma_list)\n\n# The images are of various sizes and they are scaled down to 150x150\nimg_width <- 150\nimg_height <- 150\ntarget_size <- c(img_width, img_height)\n\n# RGB = 3 channels\nchannels <- 3\n\n# Get the data that are +ve for melanoma\nmelanoma_data <- train_csv_data %>% filter(target == 1)\n# Get the data that are -ve for melanoma\nno_melanoma_data <- train_csv_data %>% filter(target == 0) \n# Get the image names when is it +ve\nmelanoma_names <- as.matrix(melanoma_data[, 1])\n# Get the image names when is it -ve\nno_melanoma_names <- as.matrix(no_melanoma_data[, 1])\n\n\ncurrent_dir <- \"../input/siim-isic-melanoma-classification/jpeg/train\"\ncurrent_test_dir <- \"../input/siim-isic-melanoma-classification/jpeg/test\"\n\n# create train and test folders. This whole division should have been done in a more\n# intelligent way\nyes_dir <- file.path(new_data_location, \"Yes\")\ndir.create(yes_dir)\n\nno_dir <- file.path(new_data_location, \"No\")\ndir.create(no_dir)\n\ntemp_dir <- file.path(new_data_location, \"test\")\ndir.create(temp_dir)\ntest_dir <- file.path(temp_dir, \"test\")\ndir.create(test_dir)\n\nmelanoma_dir <- \"../UnderSamplingData/Yes/\"\nno_melanoma_dir <- \"../UnderSamplingData/No/\"\ntrain_dir <- \"../UnderSamplingData/\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Lets look at some basic EDA from the train.csv file\n\n# EDA\n\nThe list below shows the unqiue data in each column. In the train.csv data, there are $2056$ unique patients, $3$ genders (one being blank or unknown), $19$ age groups (including NA), $7$ location of images in the anatomy,  $9$ diagnosis information, and $2$  indicators of malignancy of imaged lesion (binary also available)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"##################################################################################################\n# List of unique data in each columns\n\nunique_data_list <- train_csv_data %>%\n  summarise(n_patients = n_distinct(patient_id),\n            n_gender = n_distinct(sex),\n            n_age = n_distinct(age_approx),\n            n_anatom_site = n_distinct(anatom_site_general_challenge),\n            n_diagnosis = n_distinct(diagnosis),\n            n_benign_malignant = n_distinct(benign_malignant),\n            n_target = n_distinct(target))\n\nunique_data_list","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Gender Distribution\n\nLets look at the gender distribution. The plot below shows that, in the train data, about $51.6\\%$ data belongs to male, and $48.2\\%$ for female. The rest is unknown as it is not provided. We can look at the patient ID's of these unkown observation to confirm whether they belong to any gender category.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"##################################################################################################\n# Sex Distribution\n#options(repr.plot.width = 10, repr.plot.height = 10) \npercent_sextype <- train_csv_data %>% group_by(sex) %>% count() %>% ungroup() %>% \n  arrange(desc(sex)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                      lab.pos = cumsum(Percent) - 0.5*Percent)\n\nggplot(data = percent_sextype, \n       aes(x = 2, y = Percent, fill = sex))+\n  geom_bar(stat = \"identity\")+\n  coord_polar(\"y\", start = 200) +\n  geom_text(aes(y = lab.pos, label = paste(Percent,\"%\", sep = \"\")), col = \"black\", size = 8) +\n  theme_void() +\n  scale_fill_brewer(palette = \"Oranges\")+ labs(fill = \"\")+ theme(legend.text=element_text(size=20)) + \n  xlim(.2,2.5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can furthur look into the indicators for malignancy based on the gender distribution as shown below. Since this is a highly imbalanced dataset, you can see that the malignant cases in each category are very less compared to benign cases.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"##################################################################################################\n# Sex + target Distribution\noptions(repr.plot.width = 10, repr.plot.height = 6)\nsex_benign_distribution <- train_csv_data %>% group_by(sex, benign_malignant) %>% summarise(count = n()) \nggplot(sex_benign_distribution, aes(benign_malignant, count)) + geom_bar(aes(fill = sex), position = \"dodge\", stat = \"identity\") + \n  scale_alpha_manual(values = c(0.6, 1)) +\n  facet_grid(. ~ sex) + xlab(\"\") +\n  scale_fill_brewer(palette = \"Set1\") +\n  theme_bw() + theme( strip.background  = element_blank(),\n                      panel.grid.major = element_blank(),\n                      panel.border = element_blank(),\n                      axis.ticks = element_blank(),\n                      panel.grid.minor.x=element_blank(),\n                      panel.grid.major.x=element_blank() ) +\n  theme(legend.position=\"bottom\") + labs(fill = \"\")+ theme(legend.text=element_text(size=18)) +theme(axis.text = element_text(size = 20))+\n  theme(strip.background = element_blank(), strip.text = element_blank()) + ylab(\"Count\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Furthur, looking for the data of location of the lesion in the anatomy for each gender, the plot below shows the distribution. For each category, approximately the distribution is quite similar for both the genders","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 20, repr.plot.height = 10)\n\nfemale_percent_anatom_side <- train_csv_data %>% filter(sex %in% (\"female\")) %>% \n  group_by(sex, anatom_site_general_challenge) %>% count() %>% ungroup() %>% \n  arrange(desc(anatom_site_general_challenge)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                                          lab.pos = cumsum(Percent) - 0.5*Percent)\n\nmale_percent_anatom_side <- train_csv_data %>% filter(sex %in% (\"male\")) %>% \n  group_by(sex, anatom_site_general_challenge) %>% count() %>% ungroup() %>% \n  arrange(desc(anatom_site_general_challenge)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                                          lab.pos = cumsum(Percent) - 0.5*Percent)\n\nanatom_sex_data <- rbind(female_percent_anatom_side, male_percent_anatom_side)\n\nggplot(anatom_sex_data, aes(sex, Percent, fill = anatom_site_general_challenge, color = sex)) +\n  geom_col(size=.8) +  scale_x_discrete(limits = c(\"\",\"male\", \"female\" )) + \n  scale_color_manual(values=c(\"black\", \"blue\")) +\n  geom_text(aes(y = lab.pos, label = paste(Percent,\"%\", sep = \"\")), col = \"black\", size = 8) +\n  scale_fill_brewer(palette = \"Oranges\") +\n  coord_polar(\"y\") + xlab(\"\") + ylab(\"Percent\") + labs(fill = \"\", color = \"\")+\n  theme(axis.title.x=element_blank(),\n        axis.title.y=element_blank(),\n        axis.text.x=element_blank(),\n        axis.ticks.x=element_blank(),        \n        axis.text.y=element_blank(),\n        axis.ticks.y=element_blank(),\n        panel.background = element_rect(fill = \"transparent\"), # bg of the panel\n        plot.background = element_rect(fill = \"transparent\", color = NA), # bg of the plot\n        panel.grid.major = element_blank(), # get rid of major grid\n        panel.grid.minor = element_blank() # get rid of minor grid\n  )+ theme(legend.text=element_text(size=20)) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Target Distribution\n\nSimilar analysis can be done for the target as well. Here are the results","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"percent_targettype <- train_csv_data %>% group_by(benign_malignant) %>% count() %>% ungroup() %>% \n  arrange(desc(benign_malignant)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                lab.pos = cumsum(Percent) - 0.5*Percent)\n\nggplot(data = percent_targettype, \n       aes(x = 2, y = Percent, fill = benign_malignant))+\n  geom_bar(stat = \"identity\")+\n  coord_polar(\"y\", start = 200) +\n  geom_text(aes(y = lab.pos, label = paste(Percent,\"%\", sep = \"\")), col = \"black\", size = 8) +\n  theme_void() +\n  scale_fill_brewer(palette = \"Greens\")+ labs(fill = \"\")+ theme(legend.text=element_text(size=20)) + \n  xlim(.2,2.5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Target -  Anatomy Distribution","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"##################################################################################################\n# Anatomy + target Distribution\n\noptions(repr.plot.width = 20, repr.plot.height = 10)\nbenign_percent_anatom_side <- train_csv_data %>% filter(benign_malignant %in% c(\"benign\")) %>% \n  group_by(benign_malignant, anatom_site_general_challenge) %>% count() %>% ungroup() %>% \n  arrange(desc(anatom_site_general_challenge)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                       lab.pos = cumsum(Percent) - 0.5*Percent)\n\nmalignant_percent_anatom_side <- train_csv_data %>% filter(benign_malignant %in% c(\"malignant\")) %>% \n  group_by(benign_malignant, anatom_site_general_challenge) %>% count() %>% ungroup() %>% \n  arrange(desc(anatom_site_general_challenge)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                                          lab.pos = cumsum(Percent) - 0.5*Percent)\n\nanatom_benign_malignant_data <- rbind(benign_percent_anatom_side, malignant_percent_anatom_side)\nggplot(anatom_benign_malignant_data, aes(benign_malignant, Percent, fill = anatom_site_general_challenge, color = benign_malignant)) +\n  geom_col(size=.8) +  scale_x_discrete(limits = c(\"\",\"benign\", \"malignant\" )) + \n  scale_color_manual(values=c(\"black\", \"blue\")) +\n  geom_text(aes(y = lab.pos, label = paste(Percent,\"%\", sep = \"\")), col = \"black\", size = 8) +\n  scale_fill_brewer(palette = \"Greens\") +\n  coord_polar(\"y\") + xlab(\"\") + ylab(\"Percent\") + labs(fill = \"\", color = \"\")+\n  theme(axis.title.x=element_blank(),\n        axis.title.y=element_blank(),\n        axis.text.x=element_blank(),\n        axis.ticks.x=element_blank(),        \n        axis.text.y=element_blank(),\n        axis.ticks.y=element_blank(),\n        panel.background = element_rect(fill = \"transparent\"), # bg of the panel\n        plot.background = element_rect(fill = \"transparent\", color = NA), # bg of the plot\n        panel.grid.major = element_blank(), # get rid of major grid\n        panel.grid.minor = element_blank() # get rid of minor grid\n  )+ theme(legend.text=element_text(size=20)) ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Age - Sex Distribution","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"##################################################################################################\n# Age + sex Distribution\noptions(repr.plot.width = 20, repr.plot.height = 10)\npercent_agetype <- train_csv_data %>% group_by(age_approx) %>% count() %>% ungroup() %>% \n  arrange(desc(age_approx)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                             lab.pos = cumsum(Percent) - 0.5*Percent)\npercent_agetype$age_approx <- as.factor(percent_agetype$age_approx)\n\nnb.cols <- 19\nmycolors <- colorRampPalette(brewer.pal(8, \"Set2\"))(nb.cols)\n\n\n# ggplot(data = percent_agetype,\n#        aes(x = 2, y = Percent, fill = age_approx))+\n#   geom_bar(stat = \"identity\")+\n#   coord_polar(\"y\", start = 200) +\n#   geom_text(aes(y = lab.pos, label = paste(Percent,\"%\", sep = \"\")), col = \"black\") +\n#   theme_void() +\n#   scale_fill_manual(values = mycolors) + labs(fill = \"\")+ theme(legend.text=element_text(size=12)) +\n#   xlim(.2,2.5)\n\n\nfemale_age_percent <- train_csv_data %>% filter(sex %in% (\"female\")) %>% \n  group_by(sex, age_approx) %>% count() %>% ungroup() %>% \n  arrange(desc(age_approx)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                                          lab.pos = cumsum(Percent) - 0.5*Percent)\n\nmale_age_percent <- train_csv_data %>% filter(sex %in% (\"male\")) %>% \n  group_by(sex, age_approx) %>% count() %>% ungroup() %>% \n  arrange(desc(age_approx)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                                          lab.pos = cumsum(Percent) - 0.5*Percent)\n\nage_sex_data <- rbind(female_age_percent, male_age_percent)\n\nage_sex_data$age_approx <- as.factor(age_sex_data$age_approx)\n\nggplot(age_sex_data, aes(sex, Percent, fill = age_approx, color = sex)) +\n  geom_col(size=.8) +  scale_x_discrete(limits = c(\"\",\"male\", \"female\" )) + \n  scale_color_manual(values=c(\"black\", \"blue\")) +\n  scale_fill_manual(values = mycolors) +\n  geom_text(aes(y = lab.pos, label = paste(Percent,\"%\", sep = \"\")), col = \"black\", size = 8) +\n  coord_polar(\"y\") + xlab(\"\") + ylab(\"Percent\") + labs(fill = \"\", color = \"\")+\n  theme(axis.title.x=element_blank(),\n        axis.title.y=element_blank(),\n        axis.text.x=element_blank(),\n        axis.ticks.x=element_blank(),        \n        axis.text.y=element_blank(),\n        axis.ticks.y=element_blank(),\n        panel.background = element_rect(fill = \"transparent\"), # bg of the panel\n        plot.background = element_rect(fill = \"transparent\", color = NA), # bg of the plot\n        panel.grid.major = element_blank(), # get rid of major grid\n        panel.grid.minor = element_blank() # get rid of minor grid\n  )+ theme(legend.text=element_text(size=20)) \n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Anatomy - Age Distribution","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"##################################################################################################\n# Anatomy + age Distribution\n\noptions(repr.plot.width = 20, repr.plot.height = 10)\nbenign_age_percent <- train_csv_data %>% filter(benign_malignant %in% c(\"benign\")) %>% \n  group_by(benign_malignant, age_approx) %>% count() %>% ungroup() %>% \n  arrange(desc(age_approx)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                                          lab.pos = cumsum(Percent) - 0.5*Percent)\n\nmalignant_age_percent <- train_csv_data %>% filter(benign_malignant %in% c(\"malignant\")) %>% \n  group_by(benign_malignant, age_approx) %>% count() %>% ungroup() %>% \n  arrange(desc(age_approx)) %>% mutate(Percent = round(n/sum(n), 3)*100,\n                                                          lab.pos = cumsum(Percent) - 0.5*Percent)\n\nanatom_age_data <- rbind(benign_age_percent, malignant_age_percent)\nanatom_age_data$age_approx <- as.factor(anatom_age_data$age_approx)\nggplot(anatom_age_data, aes(benign_malignant, Percent, fill = age_approx, color = benign_malignant)) +\n  geom_col(size=.8) +  scale_x_discrete(limits = c(\"\",\"benign\", \"malignant\" )) + \n  scale_color_manual(values=c(\"black\", \"blue\")) +\n  geom_text(aes(y = lab.pos, label = paste(Percent,\"%\", sep = \"\")), col = \"black\", size = 8) +\n  scale_fill_manual(values = mycolors) +\n  coord_polar(\"y\") + xlab(\"\") + ylab(\"Percent\") + labs(fill = \"\", color = \"\")+\n  theme(axis.title.x=element_blank(),\n        axis.title.y=element_blank(),\n        axis.text.x=element_blank(),\n        axis.ticks.x=element_blank(),        \n        axis.text.y=element_blank(),\n        axis.ticks.y=element_blank(),\n        panel.background = element_rect(fill = \"transparent\"), # bg of the panel\n        plot.background = element_rect(fill = \"transparent\", color = NA), # bg of the plot\n        panel.grid.major = element_blank(), # get rid of major grid\n        panel.grid.minor = element_blank() # get rid of minor grid\n  )+ theme(legend.text=element_text(size=20))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Patient Distibution","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"##################################################################################################\n# Patient Distribution\n\ncount_Patienttype <- train_csv_data %>% group_by(patient_id) %>% count()%>% arrange(desc(n)) \n\ntrain_csv_data %>%\n  dplyr::count(patient_id) %>%\n  ggplot(aes(n)) +\n  geom_histogram(bins = 50, color = \"black\", fill = \"orange\") +\n  ggtitle(\"\") + xlab(\"Number of records per patient\") + ylab(\"Count\")+theme(text = element_text(size = 20))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Modeling\n\nSince this is a highly imbalanced dataset, I am doing the first basic model using under sampling. In the code below, I am copying 584 melanoma images and 600 non-melanoma images into the train dataset. Then I use the basic workflow of the  Keras API, which is Define --> Compile --> Fit --> Evaluate --> Predict to predict the test dataset.\n\n### Image Loading","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Run this section only the first time\n# # \n\nno_melanoma_names <- as.matrix(sample(no_melanoma_names, 600))\n\nfor (i in 1:length(melanoma_names)) {\n   #copy vector\n   #setwd(current_dir)\n   name <- melanoma_names[i, 1]\n   filename <- paste(sep=\"\",name,\".jpg\")\n   file.copy(file.path(current_dir, filename),\n             file.path(melanoma_dir))\n }\n# #\n for (i in 1:length(no_melanoma_names)) {\n   #copy vector\n   #setwd(current_dir)\n   name <- no_melanoma_names[i, 1]\n   filename <- paste(sep=\"\",name,\".jpg\")\n   file.copy(file.path(current_dir, filename),\n             file.path(no_melanoma_dir))\n }\n\n for (i in 1:length(test_image_names)) {\n   #copy vector\n   #setwd(current_dir)\n   name <- test_image_names[i, 1]\n   filename <- paste(sep=\"\",name,\".jpg\")\n   file.copy(file.path(current_test_dir, filename),\n             file.path(test_dir))\n }\n\n# ### Data loading\n# \ntrain_data_gen = image_data_generator(\n  rescale = 1/255, #,\n  validation_split = 0.2\n)\n\n# # # Validation data shouldn't be augmented! But it should also be scaled.\nvalid_data_gen <- image_data_generator(\n  rescale = 1/255\n)\n\n# Validation data shouldn't be augmented! But it should also be scaled.\ntest_data_gen <- image_data_generator(\n  rescale = 1/255\n)\n# \n# # training images\ntrain_image_array_gen <- flow_images_from_directory(train_dir,\n                                                    train_data_gen,\n                                                    shuffle = TRUE,\n                                                    target_size = target_size,\n                                                    class_mode = \"binary\",\n                                                    classes = melanoma_list,\n                                                    subset = \"training\",\n                                                    seed = 50)\n\n\n# validation images\nvalid_image_array_gen <- flow_images_from_directory(train_dir,\n                                                    train_data_gen,\n                                                    shuffle = TRUE,\n                                                    target_size = target_size,\n                                                    class_mode = \"binary\",\n                                                    classes = melanoma_list,\n                                                    subset = \"validation\",\n                                                    seed = 50)\n# \n# \ntest_image_array_gen <- flow_images_from_directory(temp_dir,\n                                                   test_data_gen,\n                                                   shuffle = FALSE,\n                                                   target_size = target_size,\n                                                   class_mode = NULL,\n                                                   seed = 50)\n# \ncat(\"Number of images per class in training dataset:\")\ntable(factor(train_image_array_gen$classes))\n# \n# train_image_array_gen$class_indices\n# \ncat(\"Number of images per class in test dataset:\")\ntable(factor(test_image_array_gen$classes))\n# \n# test_image_array_gen$class_indices\n# \n# number of training samples\ntrain_samples <- train_image_array_gen$n\n# number of validation samples\nvalid_samples <- valid_image_array_gen$n\n# number of test samples\ntest_samples <- test_image_array_gen$n\n# \n# # define batch size and number of epochs\nbatch_size <- 32\nepochs <- 50","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Define","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# ########################################################################\n# # Convoluted network analysis\n# ########################################################################\nanalysis_convoluted_network <- keras_model_sequential() %>%\n  layer_conv_2d(filter = 2, kernel_size = c(3,3), #kernel_regularizer = regularizer_l1_l2(0.1),\n                activation  = \"relu\", input_shape = c(img_width, img_height, channels)) %>%\n  layer_max_pooling_2d(pool_size = c(2,2)) %>%\n  layer_conv_2d(filter = 4, kernel_size = c(3,3),  #kernel_regularizer = regularizer_l1_l2(0.1),\n                activation  = \"relu\") %>%\n  layer_max_pooling_2d(pool_size = c(2,2)) %>%\n  layer_conv_2d(filter = 8, kernel_size = c(3,3), # kernel_regularizer = regularizer_l1_l2(0.1),\n                activation  = \"relu\") %>%\n  layer_max_pooling_2d(pool_size = c(2,2)) %>%\n  layer_conv_2d(filter = 16, kernel_size = c(3,3), # kernel_regularizer = regularizer_l1_l2(0.01),\n                activation  = \"relu\") %>%\n  layer_max_pooling_2d(pool_size = c(2,2)) %>%\n  layer_flatten() %>%\n  layer_dense(100, activation = \"relu\") %>%\n  layer_dropout(0.5) %>%\n  layer_dense(1) %>%\n  layer_activation(\"sigmoid\")\n\nsummary(analysis_convoluted_network)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Compile\n\nThe code below shows the complie section. To judge the performance of the model, accuracy is chosen as the metric and binary_crossentropy loss between the labels and predictions are calculated. The optimizer used in this example is the the adam algorithm. Several others can also be tried in this section.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"analysis_convoluted_network %>% compile(\n  loss = \"binary_crossentropy\",\n  optimizer = \"adam\",\n  metrics = \"accuracy\"\n)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Fit\n\nThe section below trains the model using the fit_generator","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"hist_CNN <- analysis_convoluted_network %>% fit_generator(\n  # training data\n  train_image_array_gen,\n\n  # epochs\n  steps_per_epoch = as.integer(train_samples / batch_size),\n  epochs = epochs,\n\n  # validation data\n  validation_data = valid_image_array_gen,\n  validation_steps = as.integer(valid_samples / batch_size),\n\n  # print progress\n  verbose = 2,\n  callbacks = list(\n    # only needed for visualising with TensorBoard\n    callback_tensorboard(\"../USData/logs/run_a/\")\n  )\n)\n# \nplot(hist_CNN)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Evaluate\nNext, we use evaluate_generator to see how well our model performed","execution_count":null},{"metadata":{"trusted":true,"collapsed":true},"cell_type":"code","source":"metrics <- analysis_convoluted_network %>% evaluate_generator(valid_image_array_gen, steps = ceiling(valid_samples/batch_size))\nmetrics","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Predict\nUse predict_generator to predict the results on the test data set. When I ran this on my machine, I got the public score of 0.832","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"predictions <- analysis_convoluted_network %>% predict_generator(test_image_array_gen, steps = ceiling(test_samples/batch_size))\n# #We can compute the predicted class by taking the column with the highest probability, for example\ny_pred <- apply(predictions, 1, which.max) - 1\n\ny <- replicate(length(y_pred), 0)\nfor (i in 1 : length(y_pred)) {\n  y[i] <- predictions[i, 1]\n}\n\ny_1 <- 1 - y\nmy_submission <- data.frame(image_name = test_csv_data$image_name,target = y_1, stringsAsFactors = TRUE)\nwrite.csv(my_submission, 'mysubmission.csv', row.names = FALSE)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Future work\n\n### Oversampling\n### Data Augumentation\n### Pre-trained NN's","execution_count":null}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.6.0","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":4}