{"cells":[{"metadata":{},"cell_type":"markdown","source":"This is a starter notebook which allows the use of R in the competition through the reticulate package. Many of the ideas come from the notebook created here (https://www.kaggle.com/dmitriyguller/r-starter-notebook-with-15x-faster-prediction-loop) for the NFL Big Data Bowl competition. "},{"metadata":{},"cell_type":"markdown","source":"First, some setup code is needed to load packages and create the column specifications. "},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"library(tidyverse)\nlibrary(reticulate)\n\noptions(scipen = 999)\n\nriii <- import_from_path('competition','../input/riiid-test-answer-prediction/riiideducation/')\n\nenv <- riii$make_env()\niter <- env$iter_test()\n\npy_run_string(\"import io\")\npy_run_string(\"import pandas as pd\")\n\ncol_spec_test <- cols(row_id = col_double(), \n    timestamp = col_double(), \n    user_id = col_integer(), \n    content_id = col_integer(), \n    content_type_id = col_integer(),\n    task_container_id = col_integer(), \n    prior_question_elapsed_time = col_double(), \n    prior_question_had_explanation = col_logical())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"A placeholder function to generate valid predictions:"},{"metadata":{"trusted":true},"cell_type":"code","source":"generate_predictions <- function(test_df){\n    return(test_df %>% mutate(answered_correctly = 0.5) %>%\n        filter(content_type_id == 0) %>%\n        select(row_id, answered_correctly))\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is the main loop. It iterates through the available test data, calling generate_predictions for each new batch."},{"metadata":{"trusted":true},"cell_type":"code","source":"while(TRUE){\n    df <- iter_next(iter)\n    if (is.null(df)) break\n    \n    test_df = read_csv(df[[1]]$to_csv(index = FALSE), \n                       col_types = col_spec_test)\n\n    pred_df <- generate_predictions(test_df)\n    \n    py$pred_df_string <- format_csv(pred_df)\n    py_run_string(\"pred_df = pd.read_csv(io.StringIO(pred_df_string))\")\n    pred_df_pointer <- py_get_attr(py, \"pred_df\")\n\n    env$predict(pred_df_pointer)    \n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Modified from: https://www.kaggle.com/dmitriyguller/r-starter-notebook-with-15x-faster-prediction-loop"}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}