{"cells":[{"metadata":{},"cell_type":"markdown","source":"## About the Competition","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Welcome to this brand new competition on indentifying melanoma in lesion images! **Melanoma** is type of skin cancer, responsible for 75% of skin cancer deaths, despite being the least common skin cancer. As with other cancers, early and accurate detection—potentially aided by data science—can make treatment more effective.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Goal\nIn this competition, we are required to identify melanoma in images of skin lesions. We’ll be using images within the same patient and determine which are likely to represent a melanoma. Using patient-level contextual information may help the development of image analysis tools, which could better support clinical dermatologists.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Dataset\nThe dataset consists of images in  :\n\n+ DIOCOM format\n+ JPEG format in JPEG directory\n+ TFRecord format in tfrecords directory\n\nAdditionally, there are metadata information about images in train, test and submission file in CSV format.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Loading Required Packages","execution_count":null},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"library(tidyverse)\nlibrary(magick)\nlibrary(EBImage)\nlibrary(keras)\nlibrary(repr)\nlibrary(gridExtra)\nlibrary(gmodels)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Reading Data Files","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"train_path <- \"../input/siim-isic-melanoma-classification/jpeg/train\"\ntest_path <- \"../input/siim-isic-melanoma-classification/jpeg/test\"\ntrain_files <- list.files(\"../input/siim-isic-melanoma-classification/jpeg/train\")\ntest_files <- list.files(\"../input/siim-isic-melanoma-classification/jpeg/test\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train <- read.csv(\"../input/siim-isic-melanoma-classification/train.csv\")\ntest <- read.csv(\"../input/siim-isic-melanoma-classification/test.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data Exploration\nI will be using training set for my exploration.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"head(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"glimpse(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Following conclusions can be made based on the above outputs:\n+ *image_name* is unique id refering to each of the image in training set.\n+ *patient_id* refers to unique patient id. There are multiple images for the same patient.\n+ There are *33126* images present in the training set.\n+ Approx. age of patient being diagnosed is ranging from 0 to 90! There are images of babies as well.\n+ Out of 33126 patients, only 584 are malignant! Only 1.76%.\n+ There are few missing values in *sex*, *age_approx* and *anatom_site_general_challenge* variable.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Target Variable","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"prop.table(table(train$target))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It's highly unbalanced problem and proper care should be taken to handle this before modeling.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"options(repr.plot.width = 9, repr.plot.height = 6)\ntrain %>% ggplot(aes(target, fill = factor(target))) + geom_bar() + theme_minimal() + ggtitle(\"Distribution of Target Variable\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Patient ID","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Unique patients in train\nlength(unique(train$patient_id))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Unique patients in test\nlength(unique(test$patient_id))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 2056 unique patients in the training set while 690 unique patients in testing set.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Sex","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 6)\np1 <- train %>% ggplot(aes(sex, fill = sex)) + geom_bar() + theme_minimal() + ggtitle(\"Distribution of Sex in Training Data\")\np2 <- test %>% ggplot(aes(sex, fill = sex)) + geom_bar() + theme_minimal() + ggtitle(\"Distribution of Sex in Testing Data\")\ngrid.arrange(p1, p2, ncol = 2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"More patients are male compared to females. Testing set contains higher proportion of male patients. Let's compare target vs sex now.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"CrossTable(train$sex, train$target, prop.chisq = F, digits = 3 )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Higher proportion of males are affected by **Melanoma**.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Age Approx","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Let's see the summary of approx age variable in train and test set.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"summary(train$age_approx)\nsummary(test$age_approx)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Compare the distribution of approx age in train and test set.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"p1 <- train %>% ggplot(aes(age_approx)) + geom_histogram(fill = \"blue\") + theme_minimal() + \n    ggtitle(\"Distribution of Approx Age in Training Data\")\np2 <- test %>% ggplot(aes(age_approx)) + geom_histogram(fill = \"blue\") + theme_minimal() + \n    ggtitle(\"Distribution of Approx Age in Testing Data\")\ngrid.arrange(p1, p2, ncol = 2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Distribution across training and testing set are almost similar. Some patients have \"zero\" approx age. Let's zoom into that.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### How many babies?","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"train[which(train$age_approx == 0), ]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is only one baby patient whose torso was scanned. Fortunately, he was found healthy.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### Target Vs Age","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Let's see whether there is any difference between age of patients wrt our target variable.","execution_count":null},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Target Vs Age\noptions(repr.plot.width = 9, repr.plot.height = 6)\ntrain %>% ggplot(aes(factor(target), age_approx, fill = factor(target))) + geom_boxplot() + theme_minimal() + \n    ggtitle(\"Distribution of Approx Age wrt Target\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like older aged patients are more suseptible for developing **Melanoma**. Let's compare age and sex together with our target variable.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"#### Who are at High Risk?","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 6)\ntrain %>% ggplot(aes(factor(target), age_approx, fill = factor(target))) + geom_boxplot() + \n    theme_minimal() + facet_wrap(~sex) + \n    ggtitle(\"Distribution of Approx Age wrt Target and Sex\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Eventhough, more males are affected with **Melanoma**, females tend to get affected at comaparatively less age than males, in general. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Body Area Scanned","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 18, repr.plot.height = 6)\np1 <- train %>% group_by(anatom_site_general_challenge) %>% tally() %>% arrange(-n) %>% \n    ggplot(aes(x = reorder(anatom_site_general_challenge, n), y = n, fill = anatom_site_general_challenge)) + \n    geom_col() + xlab(\"anatom site general challenge\") + ylab(\"count\") + theme_minimal() + \n    ggtitle(\"Distribution of Body Area Scanned in Train Set\") + coord_flip()\n\np2 <- test %>% group_by(anatom_site_general_challenge) %>% tally() %>% arrange(-n) %>% \n    ggplot(aes(x = reorder(anatom_site_general_challenge, n), y = n, fill = anatom_site_general_challenge)) + \n    geom_col() + xlab(\"anatom site general challenge\") + ylab(\"count\") + theme_minimal() + \n    ggtitle(\"Distribution of Body Area Scanned in Test Set\") + coord_flip()\n\ngrid.arrange(p1, p2, ncol = 2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Diagnosis","execution_count":null},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 6)\ntrain %>% group_by(diagnosis) %>% tally() %>% arrange(-n) %>% \n    ggplot(aes(reorder(diagnosis, n), n, fill = diagnosis)) + geom_col() + theme_minimal() + \n    ylab(\"Count\") + xlab(\"Type of Diagnosis\") + ggtitle(\"Distribution of Diagnosis Type\") + coord_flip()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Exploring Images\nLet's explore few images from train and test sets. We will read 09 random sample of images from train and test set and explore them.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Train Images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Reading 09 random images and storing in a list\ntrain_img <- list()\n\nfor (img in sample(train_files, 9)){train_img[[img]] <- readImage(paste(train_path, img, sep = \"/\"))}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's see the information contained in each of the elements of the list.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_img[[1]]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that for each of the image in our list, we have information such as color, dimensions and numerical representation of each pixel of the image. Let's visualize these images from the list.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 14, repr.plot.height = 6)\npar(mfrow = c(3,3))\nfor (i in 1:9) {display(train_img[[i]])}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that the image size is varying for the images. Let's visualize the intensity of different color channels of the images. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 14, repr.plot.height = 9)\npar(mfrow = c(3,3))\nfor (i in 1:9) {hist(train_img[[i]])}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The above histogram shows the intensity of Red, Green & Blue colors for each of the train images in our list. Note that we have scaled these images while reading, hence intensity values are ranging from 0 to 1.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Test Images","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# Reading 09 random images and storing in a list\ntest_img <- list()\n\nfor (img in sample(test_files, 9)){test_img[[img]] <- readImage(paste(test_path, img, sep = \"/\"))}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_img[[1]]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 14, repr.plot.height = 6)\npar(mfrow = c(3,3))\nfor (i in 1:9) {display(test_img[[i]])}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that the image size is varying for the images. Let's visualize the intensity of different color channels of the images. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 14, repr.plot.height = 9)\npar(mfrow = c(3,3))\nfor (i in 1:9) {hist(test_img[[i]])}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The above histogram shows the intensity of Red, Green & Blue colors for each of the test images in our list.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Do upvote, if you liked the notebook..!!","execution_count":null}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}