{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Label a sample of pictures\n\nWe would like to label pictures with bounding box. Unfortunately, with 51033 pictures, we avoid to label all of them. Instead, we would like to label a sample of them. This sample should represent the shape variability of our population. We will then stratify the sampling by specy.\n\n## Corrected dataset\n\nLet us load the corrected metadata.","metadata":{}},{"cell_type":"code","source":"library(dplyr)\nlibrary(readr)","metadata":{"execution":{"iopub.status.busy":"2022-03-17T09:59:27.929751Z","iopub.execute_input":"2022-03-17T09:59:27.932027Z","iopub.status.idle":"2022-03-17T09:59:28.154215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data <- \"/kaggle/input/happywhaleindividualtaxonomy/final.csv\" %>%\n    read_csv(col_types = 'c') %>%\n    rename(individual = individual_id, picture = image) %>%\n    select(family, genus, specy, individual, picture) %>%\n    as_tibble()\ndata","metadata":{"execution":{"iopub.status.busy":"2022-03-17T10:04:08.96133Z","iopub.execute_input":"2022-03-17T10:04:08.963119Z","iopub.status.idle":"2022-03-17T10:04:09.081711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Specy representation in the dataset","metadata":{}},{"cell_type":"code","source":"data %>%\n    group_by(specy) %>%\n    summarise(\n        individuals = n_distinct(individual),\n        pictures = n(),\n    )","metadata":{"execution":{"iopub.status.busy":"2022-03-17T11:27:03.808888Z","iopub.execute_input":"2022-03-17T11:27:03.810857Z","iopub.status.idle":"2022-03-17T11:27:03.851172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Indexing all observations\n\nIn order to get a representative sampling, we index all observations by specy and individual. This will be used afterwards to select files to label.","metadata":{}},{"cell_type":"code","source":"datarank <- data %>% \n    group_by(family, genus, specy) %>%\n    mutate(ids = row_number()) %>%\n    group_by(family, genus, specy, individual) %>%\n    mutate(idi = row_number()) %>%\n    arrange(ids, family, genus, specy, idi)\ndatarank","metadata":{"execution":{"iopub.status.busy":"2022-03-17T11:54:43.594139Z","iopub.execute_input":"2022-03-17T11:54:43.595664Z","iopub.status.idle":"2022-03-17T11:54:44.210155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"write_csv(datarank, \"datarank.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-03-17T11:40:40.90424Z","iopub.execute_input":"2022-03-17T11:40:40.90577Z","iopub.status.idle":"2022-03-17T11:40:41.076089Z"},"trusted":true},"execution_count":null,"outputs":[]}]}