{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# predict student performance with random forest\n\nprediction of student performance from game by random forest follow the recommendation of [[Gusthema and elliot robot ](http://)](https://www.kaggle.com/code/gusthema/student-performance-w-tensorflow-decision-forests/notebook). In the data manipulation section, we followed the structure of Gusthema and elliot robot  on our Notebook, after that we upload Training file to Kaggle environment.","metadata":{}},{"cell_type":"code","source":"# In RStudio, we use library(reticulate) to convert Python data to R data\n#seperate column variable in train_lable and merge to test dataset\ntrain_label<-read.csv(\"/kaggle/input/predict-student-performance-from-game-play/train_labels.csv\")\n\nlibrary(stringr); library(tidyverse)\ntrain_label$q<-sub(\".*_q(.*)\",\"\\\\1\",train_label$session_id)\ntrain_label$session_id <- sub(\"_.*\", \"\", train_label$session_id)\nglimpse(train_label)\n\nmerge_train1 <- merge(training1,train_label, by = \"session_id\")","metadata":{"execution":{"iopub.status.busy":"2023-05-17T10:17:01.812941Z","iopub.execute_input":"2023-05-17T10:17:01.814464Z","iopub.status.idle":"2023-05-17T10:17:02.337060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training model","metadata":{}},{"cell_type":"code","source":"#import library\nlibrary(reticulate); \nlibrary(tidyverse); library(dplyr); library(ggplot2); library(stringr); library(rsample);\nlibrary(tidymodels); library(lattice); library(caret); \nlibrary(yardstick); library(parsnip); library(recipes); library(ranger);\nlibrary(vip); library(workflows); library(modeldata) \n\n","metadata":{"execution":{"iopub.status.busy":"2023-05-18T10:20:45.720594Z","iopub.execute_input":"2023-05-18T10:20:45.724735Z","iopub.status.idle":"2023-05-18T10:20:45.782599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" **Import training data**","metadata":{}},{"cell_type":"code","source":"# training set\ndata_train <- read.csv(\"/kaggle/input/finaldataset/merged_train.csv\")\nglimpse(data_train)\n\n    \ndat_train <- data_train %>% select(-q)\ndat_train$correct <- as.factor(dat_train$correct)\nglimpse(dat_train)\n\nset.seed(123)\nsplit <- initial_split(dat_train, prop = 0.8)\ntrain <- training(split)\nvalid <- testing(split) #validation set\ndim(train)","metadata":{"execution":{"iopub.status.busy":"2023-05-18T10:22:34.295578Z","iopub.execute_input":"2023-05-18T10:22:34.297407Z","iopub.status.idle":"2023-05-18T10:22:49.149216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**data pre-processing**","metadata":{}},{"cell_type":"markdown","source":"using Random forest algorithm to predict correctness of student performance. The training model provide a high accuracy (0.702) and the level variable is the most important feature to predict student performane.","metadata":{}},{"cell_type":"code","source":"#pre-processing\n\ndat_preproc <- recipe(correct~., data = train) %>%\n prep(NULL) %>%\n juice()\n\nglimpse(dat_preproc, 10)\n\n#model specification\nrf <- ranger(correct ~ . ,\n            data = train,\n            importance = \"impurity\")\n\nvip(rf)\n\nrf","metadata":{"execution":{"iopub.status.idle":"2023-05-18T10:58:54.687373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**prediction**","metadata":{}},{"cell_type":"code","source":"#predict class for correct variable\n\npred_class <- predict(rf, valid, type=\"response\")$predictions\ntable(pred_class, valid$correct) %>% accuracy()","metadata":{"execution":{"iopub.status.busy":"2023-05-18T10:59:25.426711Z","iopub.execute_input":"2023-05-18T10:59:25.428282Z","iopub.status.idle":"2023-05-18T10:59:57.245162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**import testing data**\nwe summarise test dataset by 'session_id','level_group' and import that file into Kaggle enviroment as below.","metadata":{}},{"cell_type":"code","source":"test<-read.csv(\"/kaggle/input/data-test/testdat.csv\")\n\nnew_pred <- predict(rf, test, type=\"response\")$predictions\nnew_pred\ntest\n\ntest$correct <- new_pred\ntest","metadata":{"execution":{"iopub.status.busy":"2023-05-18T11:15:23.961513Z","iopub.execute_input":"2023-05-18T11:15:23.965326Z","iopub.status.idle":"2023-05-18T11:15:28.827634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission output","metadata":{}},{"cell_type":"code","source":"write.csv(test, \"/kaggle/working/submission.csv\", header = TRUE)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}