{"cells":[{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true},"cell_type":"code","source":"# This R environment comes with many helpful analytics packages installed\n# It is defined by the kaggle/rstats Docker image: https://github.com/kaggle/docker-rstats\n# For example, here's a helpful package to load\n\nlibrary(tidyverse) # metapackage of all tidyverse packages\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nlist.files(path = \"../input\")\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\n\ntheme_set(theme_bw())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Just doing some set up in this block. \n* Saving the file paths to vars\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"list.files(path = \"../input\")\nparent_path <- \"../input/riiid-test-answer-prediction/\"\ntrain_path <- paste0(parent_path, \"/train.csv\")\nlect_path <- paste0(parent_path, \"/lectures.csv\")\nques_path <- paste0(parent_path, \"/questions.csv\")\ntest_path <- paste0(parent_path, \"/example_test.csv\")\nsubm_path <- paste0(parent_path, \"/eample_sample_submission.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Loading data:"},{"metadata":{"trusted":true},"cell_type":"code","source":"train <- read_csv(train_path, n_max = 5e6)\nlects <- read_csv(lect_path)\nquest <- read_csv(ques_path)\n\nhead(train, 3)\nhead(lects, 3)\nhead(quest, 3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train %>% summary()\nrange(train$timestamp) / (1000 * 60 * 60 * 24 * 360) # to get a sense for how much time is represented by 5**6 rows of data\n\ntop_users <- train %>%\n    count(user_id) %>%\n    top_n(10)\n\ntop_users","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like have about 2.6 years worth of data loaded in the 5^6 rows. This neglects any such order of the timestamp in the train data set. At most, I can conclude that in the first 5^6 rows of data there is at least 1 user that has a recorded event 2.6 years after their completion of their first event with riiid.\n\nHow long do users spend on the site?... the average time between their first event and their last event."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_user_dur <- train %>%\n    group_by(user_id) %>%\n    summarise(dur = max(timestamp) / (1000 * 60 * 60 * 24))\n\ntrain_user_dur %>%\n    ggplot() +\n    geom_histogram(aes(x = dur))\n\ntrain_user_dur %>%\n    mutate(dur_rd = round(dur)) %>%\n    count(dur_rd) %>%\n    mutate(dur_pct = n / sum(n)) %>%\n    filter(dur_rd < 8) %>%\n    summarise(sum(dur_pct))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"40% of users are there for less than 1 day. 5% of users are there for about 1 day, 3% of users are there for about 2 days, 2% of users are there for about 3 days, and 2% of users are there for about 4 days.\n\n54% of their user base is there for 1 week. \n\nPercentages may not be truly indicative of actual user activity since the sample is not random, but rather a systematic slice of 5 billion recorded interactions. I would expect the number of questions answered by a user to correlate strongly with the amount of time they have spent on the platform.\n\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_user_ans <- train %>%\n    filter(content_type_id == 0) %>% # filtering for questions only\n    group_by(user_id) %>%\n    summarise(correct_ans = sum(answered_correctly), # number of correct answers given\n              ans = n()) # number of answers given \n\n(user_dur_ans <- inner_join(train_user_dur, train_user_ans, by = \"user_id\")) %>%\n    ggplot() +\n    geom_point(aes(x = dur, y = correct_ans / ans))\n\nuser_dur_ans %>%\n    filter(round(dur) == 0) %>%\n    ggplot() +\n    geom_histogram(aes(x = correct_ans/ans), fill = 'white', color = 'black')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As to have been expected, there is some realtionship between the amount of time spent on their platform and the overall correct answer percentage. For users that spend less than a day on their platform, there appears to be a bell shaped distribution in their correction_answer percentage. \n\nBased on the first plot, I would expect that user performance improves over time. "},{"metadata":{"trusted":true},"cell_type":"code","source":"train_user_ot <- train %>%\n    #filter(content_type_id == 0) %>%\n    group_by(user_id) %>%\n    summarise(correct_pct_1 = sum(ifelse(row_number() <= (n() / 2) & content_type_id == 0, answered_correctly, 0)) / sum(ifelse(row_number() <= (n() / 2), 1, 0)),\n              correct_pct_2 = sum(ifelse(row_number() >  (n() / 2) & content_type_id == 0, answered_correctly, 0)) / sum(ifelse(row_number() >  (n() / 2), 1, 0))) %>%\n    inner_join(user_dur_ans, by = \"user_id\")\n\ntrain_user_ot %>%\n    ggplot() +\n    geom_point(aes(x = correct_pct_2 - correct_pct_1, y = dur))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Wow, overall I did not expect to see a bell curve for performance in the first half of questions given to users versus their last half of questions. Could it be due to the fact that I am leaving in users that only answered 1 or 2 questions? Maybe... let's see."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_user_ot %>%\n    filter(ans <= 60) %>%\n    ggplot(aes(x = ans)) +\n    geom_bar()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most users seem to be answering 30 or so questions. So I'll filter for at least having 10 questions. "},{"metadata":{"trusted":true},"cell_type":"code","source":"train_user_ot %>%\n    filter(ans >= 10) %>%\n    ggplot() +\n    geom_point(aes(x = correct_pct_2 - correct_pct_1, y = dur))\n\ntrain_user_ot %>%\n    filter(ans >= 10) %>%\n    ggplot() +\n    geom_density(aes(x = round(correct_pct_2 - correct_pct_1, 2)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The amount of time a user spends overall vs the percentage of correct answers progress between the first half and the last half of their riiid ciriculum look fairly simiarly to the density of the density plot of the progress made between the first and second of half of their riiid ciriculum. That might indicate that most users do not change. So predictive accuracy would come form looking at the tails. \n\nHow does this progress appear against the number of lectures viewed."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_lect <- train %>%\n    group_by(user_id) %>%\n    summarise(lectures = n(),\n              lectures_1 = sum(ifelse(row_number() <= (n() / 2) & content_type_id == 1, 1, 0)),\n              lectures_2 = sum(ifelse(row_number() >  (n() / 2) & content_type_id == 1, 1, 0)))\n\ntrain_user_ot_lect <- inner_join(train_user_ot, train_lect, by = 'user_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_user_ot_lect %>%\n    mutate(ans_corr_chg = correct_pct_2 - correct_pct_1,\n           lec_view_chg = lectures_2 - lectures_1,\n           year_act_fct = factor(round(dur / 360))) %>%\n    ggplot() +\n    geom_point(aes(x = ans_corr_chg, y = lec_view_chg, color = year_act_fct), alpha = 0.25) +\n    facet_grid(year_act_fct~., scales = 'free')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In general, I am not seeing a strong a relationship between when children watch lectures and how they perform over time. "},{"metadata":{"trusted":true},"cell_type":"code","source":"train_corr_ans_ot <- train %>%\n    filter(content_type_id == 0) %>%\n    group_by(user_id) %>%\n    mutate(corr_pct = cumsum(answered_correctly) / row_number(),\n           day = ifelse(timestamp != 0, (timestamp - min(timestamp)) / (1000 * 60 * 60 * 24), 0),\n           day = round(day)) %>%\n    ungroup() %>%\n    group_by(user_id, day) %>%\n    summarise(corr_pct = mean(corr_pct))\n\n\nzero_dayers <- train_corr_ans_ot %>% \n    count(user_id) %>% \n    filter(n < 2) %>%\n    pull(user_id)\n\nnon_z_dayers <- train_corr_ans_ot %>% \n    count(user_id) %>% \n    filter(n >= 2) %>%\n    pull(user_id)\n\ntrain_corr_ans_ot %>%\n    filter(user_id %in% sample(non_z_dayers, 1000) & day > 3) %>%\n    ggplot() +\n    geom_line(aes(x = day, y = corr_pct), alpha = 0.5)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"\n\nLoaded a subset of the data. Just curious about an individual user for now. Which user do I have the most data about?"},{"metadata":{},"cell_type":"markdown","source":"I want to know:\n* how many different task container ids are there"},{"metadata":{"trusted":true},"cell_type":"code","source":"task_wts <- train %>%\n    count(task_container_id, sort = T)\n    \n\ndim(task_wts)[1] # number of unique ids","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I want to know:\n* which task ids do students tend to miss?"},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_acc <- train_filtered %>%\n    filter(answered_correctly != -1) %>%\n    group_by(task_container_id) %>%\n    summarise(percent_correct = mean(answered_correctly)) %>%\n    left_join(task_wts, by = \"task_container_id\") %>%\n    mutate(task_wts_pct = n/sum(n)) %>%\n    arrange(desc(percent_correct))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_acc %>% \n    ggplot() +\n    geom_density(aes(x = `n`)) +\n    scale_x_continuous(breaks = scales::pretty_breaks(n = 15))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(answer_acc)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_acc %>%\n    filter(n >= 80) %>%\n    ggplot() + \n    geom_jitter(aes(x = `task_wts_pct`, y = `percent_correct`))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(train)\ntrain %>%\n    filter(content_type_id == 1) %>%\n    summary()\n\ncid1 <- train %>%\n    filter(content_type_id == 1) %>%\n    pull(content_id)\n\ntrain %>%\n    filter(content_type_id == 0 & content_id %in% cid1) %>%\n    summary()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So questions are a part of lectures.\n\nIs there a relationship between if the user \"watched\" the lecture and if there answer is correct?\nI don't know if I know enough about the data set to answer that question right out. To start: How many are questions are there per lecture?\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"#(\ntrain_ans_lect <- train %>%\n        group_by(user_id, content_id) %>%\n        summarise(lectures = sum(content_type_id),\n                  questions  = length(content_type_id) - lectures,\n                  correct_answers = sum(ifelse(answered_correctly > 0, answered_correctly, 0)) / ifelse(questions > 0, questions, 1)) %>%\n        mutate(lecture_fct = factor(lectures))\n    #) %>%\n#train_ans_lect %>%\n    #ggplot() +\n    #geom_jitter(aes(x = `questions`, y = `correct_answers`, color = `lecture_fct`), alpha = 0.1, size = 1.5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_ans_lect %>%\n    ggplot() +\n    geom_jitter(aes(x = `questions`, y = `correct_answers`, color = `lecture_fct`), alpha = 0.25, size = 1.5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most people do not watch the lecture, but there are some that do, and then answer the subsequent questions. "},{"metadata":{"trusted":true},"cell_type":"code","source":"train_ans_lect %>%\n    group_by(content_id) %>%\n    summarise(questions = mean(questions),\n              lecture = mean(lectures)) %>%\n    ggplot() +\n    geom_jitter(aes(`questions`, `lecture`))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_ans_lect %>%\n    group_by(user_id) %>%\n    summarise(questions = mean(questions),\n              lecture = mean(lectures)) %>%\n    ggplot() +\n    geom_jitter(aes(`questions`, `lecture`))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"head(train_ans_lect)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_ans_lect %>%\n    filter(questions > 0) %>%\n    group_by(lecture_fct) %>%\n    summarise(questions = mean(questions),\n              cor_answr = mean(correct_answers)) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_ans_lect %>%\n    filter(questions > 0) %>%\n    group_by(user_id) %>%\n    summarise(questions = mean(questions),\n              cor_answr = mean(correct_answers),\n              lectures = mean(lectures)) %>%\n    ggplot() +\n    geom_point(aes(color = `questions`, y = `cor_answr`, x = `lectures`), alpha = 0.5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Between lectures and the mean percentage of correct answers, there appears to be some relationship. I think its worth being naive about the meaning of these statistics and just dive right in into some pytorch archetyping. "},{"metadata":{},"cell_type":"markdown","source":""}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}