{"cells":[{"metadata":{},"cell_type":"markdown","source":"# 💯 **Simple 92% Accurate Floor Model** 🏢\n\n\nIn this notebook, we will use the frequency of wifi bssid-RSSI combinations (only) to find the most likely floor.\n\nIt will proceed as follows:\n* Construct a floor-bssid-RSSI frequency table for each site\n* Give each floor 1 point for every bssid-RSSI pair that occurs both in the training set and the test path\n* Normalize by the total number of readings on each floor\n* Take the highest-voted floor.\n\nSimple! ✅"},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"library(tidyverse)\nlibrary(doParallel)\n\ncl <- makeCluster(detectCores())\ndoParallel::registerDoParallel(cl)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"For speed, we will use one of the already extracted datasets.\n\nI am using this one: https://www.kaggle.com/jiweiliu/wifi-label-encode . Feel free to upvote that notebook."},{"metadata":{"trusted":true},"cell_type":"code","source":"all_files <- list.files(\"../input/wifi-label-encode\", recursive = TRUE)\ntrain_files <- data.frame(filename = all_files[grepl(\"/train/\", all_files)]) %>%\n    separate(filename, into=c(\"dataset\",\"train_test\",\"site\",\"floor\",\"path\"), sep = \"/\", remove = F)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Part 1: Demonstrate the model"},{"metadata":{"trusted":true},"cell_type":"code","source":"# select the first site as an example\ntest_site <- train_files$site[1]\nsite_paths <- train_files %>% filter(site == test_site) %>% pull(path) %>% unique()\n\n# combine wifi readings from all files in this test_site into one data.frame\n#    be sure not eliminate the reading repeating at multiple timestamps\n# we also record which floor the readings are from\nall_wifi <- data.frame()\nfor(iter_path in site_paths){\n    this_path_wifi <- read.csv(paste0(\"../input/wifi-label-encode/\", train_files$filename[train_files$path == iter_path])) %>%\n        filter(last_timestamp >= (min(timestamp)-5000)) %>% select(-timestamp) %>% distinct()\n    all_wifi <- all_wifi %>% bind_rows(this_path_wifi %>% select(bssid, rssi) %>%\n                                       mutate(floor = train_files$floor[train_files$path == iter_path], path = iter_path)\n                                      )\n}\nall_wifi <- all_wifi %>% filter(rssi > -60)\nhead(all_wifi)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"set up CV folds to test the accuracy of our model"},{"metadata":{"trusted":true},"cell_type":"code","source":"nfolds <- 5\nset.seed(0)\nfolds <- sample(1:nfolds, length(site_paths), replace = T)\nthis_fold <- 1\n\ntrain_wifi <- all_wifi %>% filter(path %in% site_paths[folds != this_fold]) %>% select(-path)\ntest_wifi <- all_wifi %>% filter(path %in% site_paths[folds == this_fold]) %>% select(-floor)\n\nhead(test_wifi)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Match test set combinations to identical training records\n\nFor each wifi reading in our test set, match to all readings in the training set that have the identical bssid and rssi\nThen count how many times these bssid-rssi combinations occured on each floor"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_train_match <- inner_join(test_wifi, train_wifi, by = c(\"bssid\",\"rssi\")) %>%\n    group_by(path, bssid, rssi, floor) %>% summarise(n_obs = n(), .groups = \"drop\")\nhead(test_train_match)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Now take the most frequent floor for each path!"},{"metadata":{"trusted":true},"cell_type":"code","source":"overall_floor_freq <- table(train_wifi$floor) %>% data.frame() %>% setNames(c(\"floor\",\"freq\"))\n\nfloor_preds <- test_train_match %>% group_by(path, floor) %>%\n    summarise(n_obs = sum(n_obs), .groups=\"drop_last\") %>% \n    slice_max(n_obs/overall_floor_freq$freq[match(floor, overall_floor_freq$floor)]) %>%\n    left_join(train_files %>% select(path, floor), by = \"path\", suffix = c(\"_pred\",\"_actual\"))\nhead(floor_preds)\nprop.table(table(floor_preds$floor_pred == floor_preds$floor_actual))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Part 2: Run on all cases in training set\n\nNot bad!\n\nNow let's match every path to every other path. This is essentially the same as leave one out (LOO) validation."},{"metadata":{"trusted":true},"cell_type":"code","source":"# run for all sites\nloo_results <- foreach(test_site = (train_files$site %>% unique()), .combine = bind_rows, .packages = c(\"dplyr\")) %dopar% {\n    site_paths <- train_files %>% filter(site == test_site) %>% pull(path) %>% unique()\n\n    # load wifi readings from files\n    all_wifi <- data.frame()\n    for(iter_path in site_paths){\n        this_path_wifi <- read.csv(paste0(\"../input/wifi-label-encode/\", train_files$filename[train_files$path == iter_path])) %>%\n            filter(last_timestamp >= (min(timestamp)-5000)) %>% select(-timestamp) %>% distinct()\n        all_wifi <- all_wifi %>% bind_rows(this_path_wifi %>% select(bssid, rssi) %>%\n                                       mutate(floor = train_files$floor[train_files$path == iter_path], path = iter_path)\n                                      )\n    }\n    all_wifi <- all_wifi %>% filter(rssi > -60)\n    overall_floor_freq <- table(all_wifi$floor) %>% data.frame() %>% setNames(c(\"floor\",\"nreadings\"))\n    path_readings <- table(all_wifi$path) %>% data.frame() %>% setNames(c(\"path\",\"nreadings\"))\n    \n    # make predictions for each case\n    floor_preds <- all_wifi %>% left_join(all_wifi, by = c(\"bssid\",\"rssi\"), suffix = c(\"_test\",\"_train\")) %>%\n        filter(path_train != path_test) %>% group_by(path_test, floor_train) %>% summarise(nobservations = n(), .groups = \"drop\") %>%\n        mutate(train_readings = overall_floor_freq$nreadings[match(floor_train, overall_floor_freq$floor)] -\n                              path_readings$nreadings[match(path_test, path_readings$path)],\n               evidence = nobservations / train_readings, floor_pred = floor_train) %>%\n        group_by(path_test) %>% slice_max(evidence) %>%\n        left_join(train_files %>% select(path, floor), by = c(\"path_test\" = \"path\"))\n    \n    # send the results back to the main data.frame()\n    data.frame(site = test_site, npaths = length(site_paths), accuracy = mean(floor_preds$floor_pred == floor_preds$floor))\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"How 'bout them apples? 🍎😊"},{"metadata":{"trusted":true},"cell_type":"code","source":"loo_results\npaste(\"Accuracy = \", sum(loo_results$npaths * loo_results$accuracy, na.rm = T) /sum(loo_results$npaths))","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}