{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Introduction\nThis notebook was created for my understanding of the competition data.  \nThe training data seems to huge(for me) in this competition.  \nLet's take a look at the data first."},{"metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"library(tidyverse) # metapackage of all tidyverse packages\nsuppressPackageStartupMessages(library(data.table))\nsuppressPackageStartupMessages(library(scales))\nsuppressPackageStartupMessages(library(gridExtra))\nsuppressPackageStartupMessages(library(psych))\n\n#ggplot setting\ntheme_set(theme_minimal() +\n         theme(plot.title = element_text(size = 17, face = \"bold\"),\n               plot.subtitle = element_text(size = 15, face = \"bold\"),\n               axis.text.x = element_text(size = 15),\n               axis.text.y = element_text(size = 15),\n               axis.title.x = element_text(size = 18),\n               axis.title.y = element_text(size = 18),\n               legend.text = element_text(size = 15),\n               legend.title = element_text(size = 15),\n               strip.text.x = element_text(size = 15)\n              )\n          )\n\n#Choose your favorite color\nmycol = c(\"turquoise\", \"lightslateblue\",\"mistyrose2\", \n          \"chartreuse2\", \"palegreen2\", \"lightskyblue\", \"grey60\", \"violetred1\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"#show_col(mycol)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data_path = \"../input/riiid-test-answer-prediction/\"\nlist.files(data_path)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Glance at our data\n\n### I would like to do EDA with all the data.\n### Let's load train data all!"},{"metadata":{"trusted":true},"cell_type":"code","source":"train <- fread(paste0(data_path, \"train.csv\"),\n               na.strings=c(\"\", \"NULL\"))\n\nquestions <- fread(paste0(data_path, \"questions.csv\"),\n               na.strings=c(\"\", \"NULL\"))\n\nlectures <- fread(paste0(data_path, \"lectures.csv\"),\n               na.strings=c(\"\", \"NULL\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**First, let's take a bird's eye view of the data.**"},{"metadata":{"trusted":true},"cell_type":"code","source":"glimpse(train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### train has 101,230,332 rows!!"},{"metadata":{"trusted":true},"cell_type":"code","source":"glimpse(questions)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"glimpse(lectures)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Check NA \n\n**I got a memory error when I checked it all by sapply function at once. Then I decided to calculate one column at a time.**"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"cat(\"train: number of NA in eatch column: \\n\")\nfor (i in names(train)){\n  tmp <- train[, colSums(sapply(.SD, is.na)), .SDcols = i] \n  cat(paste0(names(tmp), \" = \", tmp, \"\\n\"))\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **The last two features contain many NAs.**"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"cat(\"lectures: number of NA in eatch column: \\n\")\nfor (i in names(lectures)){\n  tmp <- lectures[, colSums(sapply(.SD, is.na)), .SDcols = i] \n  cat(paste0(names(tmp), \" = \", tmp, \"\\n\"))\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"cat(\"questions: number of NA in eatch column: \\n\")\nfor (i in names(questions)){\n  tmp <- questions[, colSums(sapply(.SD, is.na)), .SDcols = i] \n  cat(paste0(names(tmp), \" = \", tmp, \"\\n\"))\n}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **It seems that tags have only one NA.**"},{"metadata":{},"cell_type":"markdown","source":"## Let's check features in train using all data."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Simple function\n\ncreate_lf <- function(dt, st1){\n  rtn <- list()\n  n_levels <- dt[, .(N= .N), by = st1] %>% nrow()\n  dt_tmp <- dt[, .(N = .N), by = st1][order(-N)]\n  dt_tmp[, \":=\"(rate = N/nrow(train)*100, \n                rate_sum = cumsum(N/nrow(train)*100), \n                No = .I)]\n  setnames(dt_tmp, names(dt_tmp)[1], \"feature\") \n  #\n  rtn[[1]] <- n_levels\n  rtn[[2]] <- dt_tmp\n  \n  return(rtn)\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_cols <- names(train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ***row_id***"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"st1 = train_cols[1]\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **row_id is just row id.**\n* **Remove it as it is not used by EDA.**"},{"metadata":{"_kg_hide-output":true,"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train[, row_id := NULL]\ngc();gc()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_cols <- names(train)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ***timestamp***"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"st1 = train_cols[1]\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **There seem to be a lot of timestamps in the data**"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"g1 <- \n  lf[[2]][1:20, ] %>% \n  ggplot(aes(x = factor(feature), y=rate))+\n  geom_bar(stat = \"identity\", fill = mycol[1], alpha = 0.7) +\n  theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = paste0(\"top 20 of \", train_cols[2]),\n       y = \"rate[%]\",\n       subtitle = paste0(st1, \": in descending order of appearance frequency\"))\n\ng2 <-\n  lf[[2]][2:nrow(lf[[2]])] %>% count(N) %>%\n  ggplot(aes(x = N, y = n)) + \n  geom_bar(stat = \"identity\", fill = mycol[2], alpha = 0.7) +\n  scale_x_discrete(limits = factor(1:10))+\n  coord_cartesian(xlim = c(0, 10))+\n  labs(subtitle = \"Distribution of the number of data that each time stamp has\\n(without timestamp=0).\",\n      x = \"number of data that each time stamp has\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 10, repr.plot.height = 10)\ngrid.arrange(g1, g2, ncol = 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **The raw with 0 time stamp is the most, about 0.4% of the total.**\n* **Most time stamps appeared only once.**"},{"metadata":{},"cell_type":"markdown","source":"### Is this related to 'answered_correctly'?"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 5)\n\ng1 <- \ntrain[timestamp ==0, .N, by = \"answered_correctly\"] %>% mutate(sum = sum(N))%>%\n  mutate(rate = N/sum *100) %>% \n  ggplot(aes(x = factor(answered_correctly) %>% fct_rev(), y = rate)) +\n  geom_bar(stat = \"identity\", alpha = 0.7, fill = mycol[4])+\n  geom_text(aes(label = as.integer(rate)), vjust = 1, size = 6)+\n  labs(title = \"timestamp == 0\", x =  \"answered_correctly\", y = \"rate[%]\")\n\ng2 <- \ntrain[timestamp !=0, .N, by = \"answered_correctly\"] %>% mutate(sum = sum(N))%>%\n  mutate(rate = N/sum *100) %>% \n  ggplot(aes(x = factor(answered_correctly) %>% fct_rev(), y = rate)) +\n  geom_bar(stat = \"identity\", alpha = 0.7, fill = mycol[5])+\n  geom_text(aes(label = as.integer(rate)), vjust = 1, size = 6)+\n  labs(title = \"timestamp != 0\", x =  \"answered_correctly\", y = \"rate[%]\")\n\ngrid.arrange(g1, g2, ncol = 2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **The effect of timestamps on 'answered_correctly' seems small.**\n* **It seems that the correct answer rate is slightly higher for questions that can be answered immediately.**"},{"metadata":{},"cell_type":"markdown","source":"## ***user_id***"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"st1 = train_cols[2]\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"g1 <- \n  lf[[2]][1:20, ] %>% \n  ggplot(aes(x = reorder(factor(feature),-N), y=N))+\n  geom_bar(stat = \"identity\", fill = mycol[1], alpha = 0.7) +\n  theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = paste0(\"top 20 of \",st1),\n       y = \"number of data\",\n       subtitle = paste0(st1, \": in descending order of appearance frequency\"))\n\ng2 <-\n  lf[[2]] %>% count(N) %>%\n  ggplot(aes(x = N, y = n)) + \n  geom_bar(stat = \"identity\", fill = mycol[2], alpha = 0.7) +\n  scale_x_discrete(limits = factor(1:100), breaks = seq(0, 100, 5))+\n  coord_cartesian(xlim = c(0, 65))+\n  labs(subtitle = paste0(\"Distribution of appearance frequency on \", st1))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 10, repr.plot.height = 10)\ngrid.arrange(g1, g2, ncol = 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **There are a lot of user_ids with 30 apperances in data. Why?**\n* **Can user_id imbalances affect predictions?**"},{"metadata":{},"cell_type":"markdown","source":"**Out of curiosity, I extrace user_ids with 30 apperrances, and see the relationship with answered_correctly.**"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 5)\ntp1 <- lf[[2]][N == 30, feature]\n\ng1 <- \ntrain[user_id %chin% tp1, \"answered_correctly\"][, .N, by = answered_correctly]%>% \n  mutate(sum = sum(N))%>%\n  mutate(rate = N/sum *100) %>% \n  ggplot(aes(x = factor(answered_correctly) %>% fct_rev(), y = rate)) +\n  geom_bar(stat = \"identity\", alpha = 0.7, fill = mycol[4])+\n  geom_text(aes(label = as.integer(rate)), vjust = 1, size = 6)+\n  labs(title = \"user_id with 30 apperrances\", x =  \"answered_correctly\", y = \"rate[%]\")\n\ng2 <-\ntrain[, \"answered_correctly\"][, .N, by = answered_correctly]%>% \n  mutate(sum = sum(N))%>%\n  mutate(rate = N/sum *100) %>% \n  ggplot(aes(x = factor(answered_correctly) %>% fct_rev(), y = rate)) +\n  geom_bar(stat = \"identity\", alpha = 0.7, fill = mycol[5])+\n  geom_text(aes(label = as.integer(rate)), vjust = 1, size = 6)+\n  labs(title = \"all user_id\", x =  \"answered_correctly\", y = \"rate[%]\")\n\ngrid.arrange(g1, g2, ncol = 2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **The user_id that appears 30 times seems to have a high incorrect answer rate.**\n* **This can be noise. Is that so?**"},{"metadata":{"trusted":true},"cell_type":"code","source":"train[user_id %chin% tp1][, .N, by = task_container_id]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **It seems that task_container_id is 30 or less. I will look it up again later.**"},{"metadata":{},"cell_type":"markdown","source":"## ***content_id***"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"st1 = train_cols[3]\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"g1 <- \n  lf[[2]][1:20, ] %>% \n  ggplot(aes(x = reorder(factor(feature),-N), y=N))+\n  geom_bar(stat = \"identity\", fill = mycol[1], alpha = 0.7) +\n  theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = paste0(\"top 20 of \",st1),\n       y = \"number of data\",\n       subtitle = paste0(st1, \": in descending order of appearance frequency\"))\n\ng2 <- \n  lf[[2]] %>%\n  ggplot(aes(x = N))+\n  geom_histogram(breaks = seq(0, 200000, 100), fill = mycol[2], alpha = 0.7)+\n  coord_cartesian(xlim = c(0,30000))+\n  labs(subtitle = paste0(\"Distribution of appearance frequency on \", st1))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 10)\ngrid.arrange(g1, g2, ncol = 1)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"lf[[2]]$N %>% describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **The average number of appearances of content_id is 7435.**\n* **If you look closely at the number of data for each content_id, the peaks are divided.**"},{"metadata":{},"cell_type":"markdown","source":"### Let's also see the relationship with answered_correctly."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"dt_tmp <- train[answered_correctly != -1\n               ][, .N, by = .(content_id, answered_correctly)\n                ][order(content_id, answered_correctly)\n                 ][, .(sum_N = sum(N), correct = sum(N*answered_correctly)), by = content_id\n                  ][, .(correct_rate = correct/sum_N*100), by = content_id]\ndt_tmp2 <- lf[[2]][,c(\"feature\", \"N\")] %>% rename(content_id = feature) %>% as.data.table()\n\ndt_tmp[dt_tmp2, \":=\" (frequency = i.N), on = \"content_id\"]\n\ndt_tmp[, freq_bin := as.integer(frequency/100)]\ndt_tmp[, mean_correct_rate := mean(correct_rate), by = freq_bin]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 6)\ndt_tmp %>%\n  ggplot(aes(x = correct_rate)) +\n  geom_histogram(breaks = seq(0, 100, 2), fill = mycol[5], alpha = 0.7)+\n  labs(subtitle = \"Disribution of corret answer rate \", x = \"Correct answer rate for each content_id [%]\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **The correct answer rate for each content_id seems to be relatively high.**\n* **There seem to be a few challenges.**"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 6)\ndt_tmp %>%\n  ggplot(aes(x = freq_bin, y = mean_correct_rate))+\n  geom_point(color = ifelse(dt_tmp$mean_correct_rate > 75, mycol[8], mycol[2])) + \n  stat_smooth(mapping = aes(x = freq_bin, y = mean_correct_rate), \n              method = \"lm\", se = FALSE, \n              color = mycol[6], size = 1, linetype=\"dashed\")+\n  labs(x = \"bin: int(apperance frequency of content_id/100)\", \n      y = \"mean of the rate of answered_correctly = 1 \")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **Roughly speaking, the correct answer rate tends to decrease when there are many user interactions.**\n* **However, there seems to be a content_id with a specifically high correct rate.**"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"rm(dt_tmp, dt_tmp2)\ngc()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ***content_type_id***"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"st1 = train_cols[4]\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 6, repr.plot.height = 5)\n\nlf[[2]][1:lf[[1]], ] %>% \n  ggplot(aes(x = reorder(factor(feature),-N), y=rate))+\n  geom_bar(stat = \"identity\", fill = mycol[1], alpha = 0.7) +\n  #theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = paste0(\"\",st1),\n       y = \"rate[%]\",\n       subtitle = paste0(st1, \": appearance frequency\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **0 if the event was a question being posed to the user -> 98%**\n* **1 if the event was the user watching a lecture -> 2%**"},{"metadata":{},"cell_type":"markdown","source":"### Let's examine the correlation between 'questions' and 'lctures'."},{"metadata":{"trusted":true},"cell_type":"code","source":"s1 <- train[content_type_id == 0, content_id] %>% unique() %>% sort()\ns2 <- questions$question_id\nall.equal(s1, s2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### content_type_id = 0 is all associated with question_id."},{"metadata":{"trusted":true},"cell_type":"code","source":"s1 <- train[content_type_id == 1, content_id] %>% unique() %>% sort()\ns2 <- lectures$lecture_id\nall.equal(s1, s2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### There seems to be a difference with lectures.\n### Let's see the difference."},{"metadata":{"trusted":true},"cell_type":"code","source":"dt1 <- train[content_type_id == 1, \"content_id\"][order(content_id),] %>% \nmutate(train = \"Y\") %>% distinct() %>% rename(lecture_id = content_id) %>% as.data.table()\ndt2 <- data.table(lecture_id = lectures$lecture_id)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dt2[dt1, \":=\" (train = i.train), on = \"lecture_id\"][is.na(train),]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### id = 641, 4385 and 28098 are included in the 'lectures', but not in the 'train'.\n### Let's chack train."},{"metadata":{"trusted":true},"cell_type":"code","source":"train[content_id %chin% c(641, 4385, 28098)][, .N, by = .(content_id, content_type_id)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **lecture_id = 641 and 4385 had content_type_id = 0 in train.**\n* **It seems to be wrong to be classified as lectures**\n* **lecture_id = 64128098 dose not exist in train.**"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures[lecture_id %chin% c(641, 4385, 28098)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train[content_id %chin% c(641, 4385, 28098)][, .N, by = .(content_id, answered_correctly)]","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"rm(s1, s2, dt1, dt2)\ngc()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ***task_container_id***"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"st1 = train_cols[5]\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"g1 <- \n  lf[[2]][1:20, ] %>% \n  ggplot(aes(x = reorder(factor(feature),-N), y=rate))+\n  geom_bar(stat = \"identity\", fill = mycol[1], alpha = 0.7) +\n  theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = paste0(\"top 20 of \",st1),\n       y = \"rate[%]\",\n       subtitle = paste0(st1, \": in descending order of appearance frequency\"))\n\ng2 <- \n  lf[[2]] %>%\n  ggplot(aes(x = N))+\n  geom_histogram(breaks = seq(0, 200000, 100), fill = mycol[2], alpha = 0.7)+\n  coord_cartesian(xlim = c(0,20000))+\n  labs(subtitle = paste0(\"Distribution of appearance frequency on \", st1))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 10)\ngrid.arrange(g1, g2, ncol = 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lf[[2]][, bin:=as.integer(feature/100) ]\n\noptions(repr.plot.width = 12, repr.plot.height = 5)\nlf[[2]][, .(f_of_bin = sum(N)), by = bin ] %>%\n  ggplot(aes(x = bin, y = f_of_bin)) +\n  geom_bar(stat = \"identity\", fill = mycol[5])+\n  labs(x = \"bin: task_container_id/100\", y = \"frequency\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **The value and frequency of task_container_id are proportional.**"},{"metadata":{"_kg_hide-input":false,"trusted":true},"cell_type":"code","source":"train[task_container_id <= 500 & answered_correctly != -1\n      ][, .(answer_correct_rate =mean(answered_correctly)), \n        by =task_container_id\n        ] %>% \n  ggplot(aes(x = task_container_id, y = answer_correct_rate)) +\n  geom_point(color = mycol[8])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **There is a point where the answer_correct_rate is low in the part where the value of task_container_id is small.**\n* **As task_container_id rises, answer_correct_rate (the average of answered_correctly) seems to rise moderately.**"},{"metadata":{},"cell_type":"markdown","source":"## ***user_answer***"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"st1 = train_cols[6]\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"options(repr.plot.width = 6, repr.plot.height = 5)\n\nlf[[2]][1:lf[[1]], ] %>% \n  ggplot(aes(x = reorder(factor(feature),-N), y=rate))+\n  geom_bar(stat = \"identity\", fill = mycol[1], alpha = 0.7) +\n  #theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = paste0(\"\",st1),\n       y = \"rate[%]\",\n       subtitle = paste0(\"in descending order of the number of data\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ***answered_correctly***"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"st1 = train_cols[7]\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"options(repr.plot.width = 6, repr.plot.height = 5)\n\nlf[[2]][1:lf[[1]], ] %>% \n  ggplot(aes(x = reorder(factor(feature),-N), y=rate))+\n  geom_bar(stat = \"identity\", fill = mycol[1], alpha = 0.7) +\n  #theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = paste0(\"\",st1),\n       y = \"rate[%]\",\n       subtitle = paste0(\"In descending order of the number of data\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Let's see the relationship with user_answer."},{"metadata":{"trusted":true},"cell_type":"code","source":"train[, .SD,  .SDcols = c(\"user_answer\", \"answered_correctly\")\n     ][,.N, .(answered_correctly, user_answer)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Let's see if one user_id answers the same content multiple times.Take User_id499347415 as an example."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"gc()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"library(DT)\n\ntp1 <- \n  train[answered_correctly != -1\n      ][, .(N = .N, correct_rate = mean(answered_correctly)), \n        by = .(user_id, content_id)\n        ][order(-N)][1:10, user_id]\n\ntrain[user_id == tp1[1]][, .(N = .N, correct_rate = round(mean(answered_correctly),digits = 2)), \nby = .(content_id)][order(-N)] %>% datatable(caption = paste0(\"User_id: \", tp1[1]))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"gc()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **For example, it seems that the correct answer is given 83 times for content_id = 7857.**"},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"train <- train[, list(user_id, content_id,answered_correctly, \n                    prior_question_elapsed_time, \n                    prior_question_had_explanation)]\ngc()\ngc()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### For each content_id, the relationship between the number of times each user answered and the correct answer rate is aggregated."},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 7)\n\ntrain[answered_correctly != -1\n      ][, .(N = .N, correct_rate = mean(answered_correctly)), \n        by = .(user_id, content_id)] %>% \n  ggplot(aes(x = factor(N), y = correct_rate)) +\n  geom_boxplot(color = mycol[2], fill = mycol[6], alpha = 0.7)+\n  theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = \"number of times each user answered for each content_id\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **There are cases where the same user answers the correct answer many times for the same content.**\n* **As long as the number of answers is small, the correct answer rate seems to increase gradually.**"},{"metadata":{"trusted":true},"cell_type":"code","source":"train <- train[, list(answered_correctly, \n                    prior_question_elapsed_time, \n                    prior_question_had_explanation)]\ngc()\ngc()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ***prior_question_elapsed_time***\n\n### Let's draw a histogram at 1 second intervals with seconds as the unit."},{"metadata":{"trusted":true},"cell_type":"code","source":"options(repr.plot.width = 12, repr.plot.height = 5)\n\ntrain[!is.na(prior_question_elapsed_time), ] %>%\n  ggplot(aes(x = as.integer(prior_question_elapsed_time/1000)))+\n  geom_histogram(aes(y = ..density..), \n                 breaks = seq(0, 300, 1), fill = mycol[2],alpha = 0.7)+\n  scale_x_continuous(breaks = seq(0, 200, 10))+\n  coord_cartesian(xlim = c(0, 120))+\n  labs(x = \"prior_questionelapsed_time[s]\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **prior_questionelapsed_time seems to have a peak around 16 seconds.**"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"options(repr.plot.width = 10, repr.plot.height = 5)\n\ntrain[, c(\"answered_correctly\",  \"prior_question_elapsed_time\")\n      ][!is.na(prior_question_elapsed_time), \n       ][, .(mean = mean(prior_question_elapsed_time),\n                       sd = sd(prior_question_elapsed_time)), \n         by = answered_correctly] %>%\n  melt(measure.vars = c(\"mean\", \"sd\"),\n       variable.name = \"mean_sd\",\n       value.name = \"value\") %>%\n  ggplot(aes(x =mean_sd , y = value, fill = factor(answered_correctly)))+\n  geom_bar(stat = \"identity\", position = \"dodge\", alpha = 0.7) + \n  scale_fill_manual(values = c(mycol[4],mycol[5])) +\n  labs(fill = \"answer_correctly\", x = \"mean and sd of prior_question_elapsed_time\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **Mean and standard deviation do not seem to have anything to do with answer accuracy.**"},{"metadata":{},"cell_type":"markdown","source":"## ***prior_question_had_explanation***"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"st1 = \"prior_question_had_explanation\"\nlf <- create_lf(train, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"options(repr.plot.width = 6, repr.plot.height = 5)\n\nlf[[2]][1:lf[[1]], ] %>% \n  ggplot(aes(x = reorder(factor(feature),-N), y = rate))+\n  geom_bar(stat = \"identity\", fill = mycol[1], alpha = 0.7) +\n  #theme(axis.text.x = element_text(angle = 90, hjust = 1))+\n  labs(x = paste0(\"\",st1),\n       y = \"rate[%]\",\n       subtitle = paste0(\"In descending order of the number of data\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Let's check features in questions."},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"train <- fread(paste0(data_path, \"train.csv\"),\n               na.strings=c(\"\", \"NULL\"),\n              select = c(\"content_id\", \"content_type_id\", \"answered_correctly\"))\ngc()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"questions_col = names(questions)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Calculate the correct answer rate for each \"content_id\" from train."},{"metadata":{"_kg_hide-input":true,"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"dt_tmp <-\n  train[content_type_id == 0\n      ][, .(mean = mean(answered_correctly)), by = content_id]\nsetnames(dt_tmp, \"content_id\", \"question_id\")\n\nquestions[dt_tmp, \":=\" (correct_rate = i.mean), on = \"question_id\"]\n\nquestions_back <- questions %>% copy()\ngc()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ***part***"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"st1 = questions_col[4]\nlf <- create_lf(questions, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"questions[, N := .N, by =part]\n\n#\nquestions[, mean_correct_rate := mean(correct_rate), by = N]\n\noptions(repr.plot.width = 12, repr.plot.height = 5)\ng1<- \nquestions %>%\n  ggplot(aes(x = reorder(factor(part), mean_correct_rate), y = correct_rate)) +\n  geom_boxplot(color = mycol[2] ,fill = mycol[6], alpha = 0.7) +\n  scale_y_continuous(breaks = seq(0, 1, 0.2))+\n  labs(x = \"part\", y = \"answer correct rate\")\n\ng2 <-\nquestions %>% \n  ggplot(aes(x = factor(part), y = N)) +\n  stat_summary(fun=\"mean\", geom=\"bar\", fill = mycol[6], alpha = 0.7)+\n  #geom_bar(stat = \"identity\", fill = mycol[6], alpha = 0.7)+\n  labs(x = \"part\", y = \"number of data\")\n\ngrid.arrange(g1, g2, ncol = 2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"questions <- questions_back\nquestions[dt_tmp, \":=\" (correct_rate = i.mean), on = \"question_id\"]\ngc()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## ***tags***"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"st1 = questions_col[5]\nlf <- create_lf(questions, st1)\n\ncat(paste0(st1, \" has \", lf[[1]], \" unique values.\\n\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions[, list(question_id,tags)] %>% head(5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### tag seems to be a combination of some numbers"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"questions[, N := .N, by =tags]\n\nquestions[, tag_n := str_count(questions$tags, \" \") + 1]\nquestions[, tags_N := .N, by = tag_n]\n\n#\nquestions[, mean_correct_rate := mean(correct_rate), by = tag_n]\n\noptions(repr.plot.width = 12, repr.plot.height = 5)\n\ng1<- \nquestions %>%\n  ggplot(aes(x = reorder(factor(tag_n), mean_correct_rate), y = correct_rate)) +\n  geom_boxplot(color = mycol[2] ,fill = mycol[6], alpha = 0.7) +\n  scale_y_continuous(breaks = seq(0, 1, 0.2))+\n  labs(x = \"numer of values in tags\", y = \"answer correct rate\")\n\ng2 <-\nquestions %>% \n  ggplot(aes(x = factor(tag_n), y = tags_N)) +\n  stat_summary(fun=\"mean\", geom=\"bar\", fill = mycol[6], alpha = 0.7)+\n  #geom_bar(stat = \"identity\", fill = mycol[6], alpha = 0.7)+\n  labs(x = \"numer of values in tags\", y = \"number of data\")\n\ngrid.arrange(g1, g2, ncol = 2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* **Questiones associated with multiple tags have a higher accuracy rate.**\n  \n#### Let's take a look at the raw data"},{"metadata":{"trusted":true},"cell_type":"code","source":"questions[order(-mean_correct_rate)\n         ][1:20, list(question_id, tags,mean_correct_rate)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### I will continue a little more.\n#### What should I look for from lectus?"}],"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"3.6.3"}},"nbformat":4,"nbformat_minor":4}