{"cells":[{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"library(tidyverse)\nlibrary(magrittr)\nlibrary(fs)\nlibrary(jpeg)\nlibrary(grid)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load data"},{"metadata":{"trusted":true},"cell_type":"code","source":"dir_ls('../input/understanding_cloud_organization/train_images/') %>%\n    length() %>% message('number of train images: ', .)\ndir_ls('../input/understanding_cloud_organization/test_images/') %>%\n    length() %>% message('number of test images: ', .)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have 5546 training images and 3698 train images.\n\nLet's have a look at one of these..."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_img <- readJPEG('../input/understanding_cloud_organization/train_images/c9a1a80.jpg', native = TRUE)\n#train_img2 <- rasterGrob(train_img)\nplot(1:2, type='n'); rasterImage(train_img,1,1,2,2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dim(train_img)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Pre-processing\n\nWe can prepare the label data we have for the train images by separating filename and label into two separate columns and convert runlength data into start-pixel / pixel-runlength data pairs stored in a nested list column:"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_runlengths <- read_csv('../input/understanding_cloud_organization/train.csv')\n\nseparate_line_runs <- function(EncodedPixels) {\n    \n    # if no pixel values provided\n    if (is.na(EncodedPixels)) {\n        df <- tibble(\n            pixel_start = numeric(length = 0), \n            pixel_runline = numeric(length = 0))\n        return(df)\n    }\n    \n    # split into individual values\n    mat <- str_split(string = EncodedPixels, pattern = \" \", simplify = TRUE)\n    \n    # separate start index and runline length\n    pixel_start <- mat[seq(1, length(mat) - 1, 2)]\n    pixel_runline <- mat[seq(2, length(mat), 2)]\n    \n    # combine output into dataframe\n    df <- bind_cols(pixel_start = pixel_start, pixel_runline = pixel_runline)\n    return(df)\n}\n\ndf_train <- df_train_runlengths %>%\n    separate(col = Image_Label, into = c('image', 'label'), sep = '_') %>%\n    mutate(target = map(EncodedPixels, ~separate_line_runs(.x))) %>%\n    select(-EncodedPixels)\n\nglimpse(df_train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Seems that train.csv has multiple rows for every image, one per class. Some rows don't have bounding box information (e.g image 3), suggesting that rows are entered even if the class is not present in the image. Next, I'd like to understand how often we have only one or a few classes vs how often all four classes are present in the image."},{"metadata":{"trusted":true},"cell_type":"code","source":"# get pixel count per image per class\npixelcounts <- df_train %>%\n    mutate(n_pixels = map_dbl(\n        .x = target, \n        .f = ~.x %>%\n            pluck('pixel_runline') %>%\n            as.numeric() %>%\n            sum())) %>%\n    select(-target)\n\n# count and summarise classes per image\npixelcounts %>%\n    filter(n_pixels > 0) %>%\n    count(image) %>%\n    select(n) %>%\n    table()\n\n# count images per class\npixelcounts %>%\n    filter(n_pixels > 0) %>%\n    select(label) %>%\n    table()\n\n# count pixels per image class\npixelcounts %>%\n    group_by(label) %>%\n    summarise(n_pixels = sum(n_pixels))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Prediction\n\nSo it seems we have a pretty good balance in regards to images class. Also, most images have multiple classes. Class \"Sugar\" has the highest prevalence in terms of number of images. Assuming that is is the same between training and test dataset, what will be our score if we predict everyting as sugar?"},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission <- read_csv('../input/understanding_cloud_organization/sample_submission.csv')\n\nsubmission_all_sugar <- sample_submission %>%\n    separate(col = Image_Label, into = c('image', 'label'), sep = '_', remove = FALSE) %>%\n    # make runlength vector including all pixels\n    mutate(EncodedPixels = case_when(\n        label == \"Sugar\" ~ str_c(1 + seq(0, 1400 - 1, 1) * 2100, rep(2100, 1400), sep = \" \", collapse = \" \"),\n        TRUE ~ \" \")) %>%\n    select(-image, -label)\n\n# save submission file\nwrite_csv(submission_all_sugar, 'submission_all_sugar.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The above submission of \"all sugar\" scores at 0.393 so I'll use that as my baseline going forward."}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}