{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Riiid! Answer Correctness Prediction\n## Track knowledge states of 1M+ students in the wild\n\n\n![](https://www.riiid.co/assets/opengraph.png)"},{"metadata":{},"cell_type":"markdown","source":"In this competition, the challenge is to create algorithms for *Knowledge Tracing* the modeling of student knowledge over time. \n\nThe goal is to accurately predict how students will perform on future interactions;  We will predict whether students are able to answer their next questions correctly.\n"},{"metadata":{},"cell_type":"markdown","source":"## Presentation of the solution:\nThis is a simplistic solution that you can begin with, that gave good roc_auc results.\n\nThere are two assumptions for this heuristic model:\n\n1. Questions differ in difficulty. There are some questions easier than others.\n2. Checking the correctness of one's answer after responding to the question influences the performance on the next question.\n\n## Modeling of the solution:\n\nWe measure question difficulty by calculating the percentage of students who answered it correctly\n\nWe will do a **target-based Encoding** on the question difficulty and whether the student looked at the prior quesion explanation combined.\n\n\n\n"},{"metadata":{},"cell_type":"markdown","source":"# **Load Data** (whole Dataset)"},{"metadata":{},"cell_type":"markdown","source":"I am only going to use content_id and prior_question_had_explanation as features for this heuristic model.\n\nI am loading the whole dataset using datatable. ( Many thanks to @Vopani, check his notebook [here](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid)\n\n\n**{**  The only issue with this, is that I would rather read only three columns from the beggining. Instead of all dumping the whole table and then extracting the three columns. I'm not sure datatable in Python allows this. Can anyone help me with this? (Pandas allows this with the argument usecols in readcsv method)    **}**"},{"metadata":{"trusted":true},"cell_type":"code","source":"!pip install ../input/python-datatable/datatable-0.11.0-cp37-cp37m-manylinux2010_x86_64.whl","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import riiideducation\nimport pandas as pd\nimport numpy as np\nimport datatable as dt","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# saving the dataset in .jay (binary format)\ndt.fread(\"../input/riiid-test-answer-prediction/train.csv\").to_jay(\"train.jay\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\n# reading the dataset from .jay format\nimport datatable as dt\n\ntrain = dt.fread(\"train.jay\")\n\nprint(train.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n\ntrain = train.to_pandas()[['content_id','prior_question_had_explanation','answered_correctly']]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We don't need lectures data, we only need questions."},{"metadata":{"trusted":true},"cell_type":"code","source":"## Answer df is the train data with only info about answers i.e containing no lectures\nanswer_df = train.query('answered_correctly != -1')\ndel train","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_df.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"On average, 65% of questions are answered correctly.\n"},{"metadata":{},"cell_type":"markdown","source":"# Verification of the assumptions."},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_df[['content_id','answered_correctly']].groupby(['content_id']).agg('mean').plot()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The percentage of students answering correctly changes from a question to another.\n\nWe even note that that are two questions that have really low scores with percentages of students answering correctly lower than  0.1!\n\nTO DO:\n\n   Look into those questions."},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_df[['prior_question_had_explanation','answered_correctly']].groupby(['prior_question_had_explanation']).agg('mean').plot(kind='bar')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The bar plot above confirms that, on average, students perform better when they have checked the explanation of their previous answer."},{"metadata":{},"cell_type":"markdown","source":"# Heuristic model"},{"metadata":{"trusted":true},"cell_type":"code","source":"average_student_performance = answer_df.describe()['answered_correctly'][1]\naverage_student_performance","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"On average 65,72% of students perform well on questions.\n\nWe will use this statistic to fill in the missing question data on the dataset."},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\n## Calculate accuracy by content and explanation\nc_exp_acc = answer_df[['content_id','prior_question_had_explanation','answered_correctly']].groupby(['content_id','prior_question_had_explanation']).agg('mean').reset_index()\nc_exp_acc.columns = ['content_id','prior_question_had_explanation', 'content_explanation_acc']\nc_exp_acc.head()\ndel answer_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# You can only call make_env() once, so don't lose it!\nenv = riiideducation.make_env()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"iter_test = env.iter_test()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for (test_df, sample_prediction_df) in iter_test:\n\n    ## Add the calculated heuristics by content and question explanation\n    ## Then fill the missing values by the average question answer.\n    test_df = test_df.merge(c_exp_acc, how = 'left', on = ['content_id','prior_question_had_explanation'])\n    test_df['answered_correctly'] = test_df['content_explanation_acc'].fillna(average_student_performance)\n    \n\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}