{"cells":[{"metadata":{},"cell_type":"markdown","source":"This notebook is an EDA of the data from the [Riiid! Answer Correctness Prediction](https://www.kaggle.com/c/riiid-test-answer-prediction) "},{"metadata":{},"cell_type":"markdown","source":"> **Credit:** This notebook is forked and edited from [this kernel](https://www.kaggle.com/erikbruin/riiid-comprehensive-eda-baseline) by Erik Bruin."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\n# import matplotlib.style as style\n# style.use('fivethirtyeight')\n# import seaborn as sns\n# import os\n# from matplotlib.ticker import FuncFormatter\n# import gc  # garbage collection","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Get the data"},{"metadata":{},"cell_type":"markdown","source":"## train"},{"metadata":{},"cell_type":"markdown","source":"The data for this competition is relatively large, and it takes a lot of time to upload it. In [this kernel](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets/) by Rohan Rao the writer explores various file formats to efficiently store and access the data. They were uploaded to this notebook (all available [here](https://www.kaggle.com/rohanrao/riiid-train-data-multiple-formats)), and one of them, the gzipped pickle, is used."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"%%time\ntrain = pd.read_pickle(\"../input/riiid-train-data-multiple-formats/riiid_train.pkl.gzip\")\nprint(\"Train size:\", train.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Directly from the data description in the competition:\n* `row_id`: (int64) ID code for the row.\n* `timestamp`: (int64) the time in milliseconds between this user interaction and the first event completion from that user.\n* `user_id`: (int32) ID code for the user.\n* `content_id`: (int16) ID code for the user interaction\n* `content_type_id`: (int8) 0 if the event was a question being posed to the user, 1 if the event was the user watching a lecture.\n* `task_container_id`: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id.\n* `user_answer`: (int8) the user's answer to the question, if any. Read -1 as null, for lectures.\n* `answered_correctly`: (int8) if the user responded correctly. Read -1 as null, for lectures.\n* `prior_question_elapsed_time`: (float32) The average time in milliseconds it took a user to answer each question in the previous question bundle, ignoring any lectures in between. Is null for a user's first question bundle or lecture. Note that the time is the average time a user took to solve each question in the previous bundle.\n* `prior_question_had_explanation`: (bool) Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Typically the first several questions a user sees were part of an onboarding diagnostic test where they did not get any feedback\n\nThe train dataset is ordered by ascending user_id and ascending timestamp."},{"metadata":{},"cell_type":"markdown","source":"Memory analysis of the data (using `memory_usage(deep=True)`) reveals that `prior_question_had_explanation` is an object, and we cast it to Boolean."},{"metadata":{"trusted":true},"cell_type":"code","source":"train['prior_question_had_explanation'] = train['prior_question_had_explanation'].astype('boolean')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Sampling the data"},{"metadata":{},"cell_type":"markdown","source":"In order for the sampled data to represent the original data we must ensure that for each user all the corresponding transactions are taken."},{"metadata":{"trusted":true},"cell_type":"code","source":"user_interactions = train.user_id.value_counts()\nSAMPLE_SIZE = 100000\nsampled_users, n_sampled = [], 0\nwhile n_sampled < SAMPLE_SIZE:\n    user = user_interactions.sample(1)\n    user_id = user.index.values[0]\n    n_interactions = user.values[0]\n    sampled_users.append(user_id)\n    n_sampled += n_interactions\n#     print(user_id, n_interactions)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(sampled_users, n_sampled)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train.loc[train.user_id.isin(sampled_users)]\ntrain.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## General statistics"},{"metadata":{},"cell_type":"markdown","source":"How many users fo we have?"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.user_id.nunique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"What is the distruibution of the number of interactions?\n\n> Who are the users with thousands of interactions?"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.user_id.value_counts().plot.hist(bins=1000, xlim=[0, 1000])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"What is the balance between questions and lectures?"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.groupby('user_id')['content_type_id'].value_counts().unstack().fillna(0).median()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It is a big question whether the lectures have any influence."},{"metadata":{},"cell_type":"markdown","source":"Etc..."},{"metadata":{},"cell_type":"markdown","source":"# User exploration"},{"metadata":{},"cell_type":"markdown","source":"This step is **IMPORTANT** and should be repeated with many users!"},{"metadata":{"trusted":true},"cell_type":"code","source":"my_user = 2136150087","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> List of representative users:\n> * 1822813285 - 195+3 interactions, intensive 14 days + short session after 40 days, repeated 7 questions\n> * 453360579 - 438 + 4 interactions, 5 days in a row, repeated 91 questions twice and 17 questions three times. Closer look shows that this user answered questions 3363 & 3365 3 times on 3 different containers, and twice he repeated the same wrong answer.\n> * 1835864303 - 32 interactions, one of the containers had 12 interactions (usually 1-3), interactions span over a year."},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df = train.loc[train.user_id==my_user]\nuser_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df = user_df.sort_values(by='timestamp')\nuser_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Timeline"},{"metadata":{},"cell_type":"markdown","source":"Converting `timestamp` to days:"},{"metadata":{"trusted":true},"cell_type":"code","source":"ms_per_day = 24 * 60 * 60 * 1000\nprint(ms_per_day)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df.timestamp = user_df.timestamp / ms_per_day\nuser_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see the bunches of activities."},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df.timestamp.plot.hist(bins=20, xlabel='Days', density=True);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> **Note:** Some users have durations of more than a year..."},{"metadata":{},"cell_type":"markdown","source":"How are the activities divided between questions and lectures?"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(10, 4))\nax = fig.gca()\nax.plot(user_df.timestamp, \n        user_df.content_type_id, \n        '.--', ms=15, lw=1)\nax.set_xlabel('Days');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In general there are very few lectures."},{"metadata":{},"cell_type":"markdown","source":"### Questions vs. lectures"},{"metadata":{},"cell_type":"markdown","source":"How important is it to see the relevant lectures (based on tags)?\n* Lectures are relatively sparse, so it is very common to answer a question without watching any relevant lecture.\n* We define `ratio` to be the ratio of tags that have been watched out of the question's tags. It is not clear that `ratio` has any value. This may indicate the seriousness of a student. Perhaps the existence of lectures is more important than the actual tags."},{"metadata":{"trusted":true},"cell_type":"code","source":"answered_content = train.groupby(['content_id', 'answered_correctly']).size().unstack()\nanswered_questions = answered_content.loc[answered_content.loc[:, -1].isnull(), [0, 1]].fillna(0)\ncorrect_ratio = answered_questions.iloc[:, 1].divide(answered_questions.sum(axis=1))\ncorrect_ratio.plot.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> **Note:** Problematic questions\n> * questions that no one answers correctly\n> * interactions which took too less/much time to complete relative to other interactions with this question bundles"},{"metadata":{},"cell_type":"markdown","source":"> **Note:** With this statistics it makes sense to assign most questions an automatic Correct/Incorrect prediction based on the question history. Such questions (with ratio of 0 or 1) should be considered as noise..."},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df.content_type_id.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df.content_id.value_counts().value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df.answered_correctly.value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> **Note:** Something strange in the repetition of questions. user_id 453360579 had 91 questions answered twice and 17 questions with 3 trials. This indicate 91 + 17\\*2 = 125 mistakes. However this user has 144 incorrect answers..."},{"metadata":{},"cell_type":"markdown","source":"> **Note:** Are there \"confusing\" questions, which are more often than others successfully solved in the second trial?"},{"metadata":{},"cell_type":"markdown","source":"> **Note:** What about repeating questions? After some trials, the user will probably hit the right answer..."},{"metadata":{},"cell_type":"markdown","source":"## Questions and lectures"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"questions = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/questions.csv', index_col='question_id')\nlectures = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/lectures.csv', index_col='lecture_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"What are the so-called \"Parts\"? When following the link provided in the data description we find out that this relates to a test.\n\n> The TOEIC L&R uses an optically-scanned answer sheet. There are 200 questions to answer in two hours in Listening (approximately 45 minutes, 100 questions) and Reading (75 minutes, 100 questions). \n\nThe listening section consists of Part 1-4 (Listening Section (approx. 45 minutes, 100 questions)).\n\nThe reading section consists of Part 5-7 (Reading Section (75 minutes, 100 questions))."},{"metadata":{},"cell_type":"markdown","source":"> **Note:** One of the questions has no related tag so we remove it for now. "},{"metadata":{"trusted":true},"cell_type":"code","source":"questions = questions.loc[questions.tags.notnull()]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Does watching a specific lecture helps in answering the related questions (based on tags)?\n\nFor each question we evaluate the ratio of tags that were represented in the lectures history of the user."},{"metadata":{},"cell_type":"markdown","source":"Metadata for the lectures watched by users as they progress in their education.\n* `lecture_id`: foreign key for the train/test content_id column, when the content type is lecture (1).\n* `part`: top level category code for the lecture.\n* `tag`: one tag codes for the lecture. The meaning of the tags will not be provided, but these codes are sufficient for clustering the lectures together.\n* `type_of`: brief description of the core purpose of the lecture\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"user_questions = user_df.loc[~user_df.content_type_id]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_prev_lectures_tags(df, current_timestamp):\n    prev_df = df.loc[df.timestamp < current_timestamp]\n    lectures_ids = prev_df.loc[prev_df.content_type_id, 'content_id'].values\n    lectures_tags = lectures.loc[lectures_ids, 'tag'].values\n    return lectures_tags","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"get_prev_lectures_tags(user_df, 0.01)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def did_watch_tag_lecture(df):\n    did_he = []\n    for i in range(len(df)):\n        current_timestamp = df['timestamp'].iloc[i]\n        if df['content_type_id'].iloc[i]:  # lecture\n            did_he.append(None)\n        else:  # question\n            lectures_tags = set(get_prev_lectures_tags(df, current_timestamp))\n            question_tags = set(map(int, questions.loc[df['content_id'].iloc[i]].tags.split()))\n#             print(i, lectures_tags, question_tags)\n            did_he_ratio = len(lectures_tags & question_tags) / len(question_tags)\n            did_he.append(did_he_ratio)\n    return did_he","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> **Note:** the history of successful answers can be a good predictive for solving questions even if relevant tagged lectures are not present"},{"metadata":{"trusted":true},"cell_type":"code","source":"ratios = did_watch_tag_lecture(user_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df['watch_ratio'] = ratios\nuser_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_df.groupby(['watch_ratio', 'answered_correctly']).size()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Test data"},{"metadata":{},"cell_type":"markdown","source":"What is `env.itertest`?"},{"metadata":{},"cell_type":"markdown","source":"# Example test"},{"metadata":{"trusted":true},"cell_type":"code","source":"example_test = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/example_test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"example_test.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Thoughts about feature engineering\n"},{"metadata":{},"cell_type":"markdown","source":"What we would like to know for the prediction?\n* How are you in this part?\n* How is it going for you?\n* Do you try a question many times?\n* Does the user repeat questions? "},{"metadata":{"trusted":true},"cell_type":"code","source":"cols = ['user_id', 'answered_correctly', 'prior_question_had_explanation']\ntrain = train.loc[:, cols]\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train[train.answered_correctly != -1]\ntrain.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 2. Baseline model"},{"metadata":{"trusted":true},"cell_type":"code","source":"#this clears everything loaded in RAM, including the libraries\n%reset -f","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport riiideducation\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport matplotlib.style as style\nstyle.use('fivethirtyeight')\nimport seaborn as sns\nimport os\nimport lightgbm as lgb\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.preprocessing import LabelEncoder\nimport gc\nimport sys\npd.set_option('display.max_rows', None)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":false,"trusted":true},"cell_type":"code","source":"%%time\ncols_to_load = ['row_id', 'user_id', 'answered_correctly', 'content_id']\ntrain = pd.read_pickle(\"../input/riiid-train-data-multiple-formats/riiid_train.pkl.gzip\")[cols_to_load]\n# train['user_id'] = train['user_id'].astype('int64')\n# train['prior_question_had_explanation'] = train['prior_question_had_explanation'].astype('boolean')\n\nprint(\"Train size:\", train.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":false,"trusted":true},"cell_type":"code","source":"# %%time\n\n# questions = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/questions.csv')\n# lectures = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/lectures.csv')\n# example_test = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/example_test.csv')\n# example_sample_submission = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/example_sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Dropping lectures rows"},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train[train.answered_correctly != -1]\ntrain.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Dropping questions with extreme correct ratio."},{"metadata":{"trusted":true},"cell_type":"code","source":"answered_questions = train.groupby(['content_id', 'answered_correctly']).size().unstack()\ncorrect_ratio = answered_questions.iloc[:, 1].divide(answered_questions.sum(axis=1))\ncorrect_ratio.plot.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"easy_question_th = 0.95\nnormal_questions = correct_ratio.loc[correct_ratio < easy_question_th].index\ntrain = train.loc[train.content_id.isin(normal_questions)]\ntrain.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_train = train.iloc[:1000000]\ntrain_test = train.iloc[1000000:1200000]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"total_q = train_train.groupby('user_id').size()\nn_correct = train_train.groupby('user_id')['answered_correctly'].sum()\nratio_q = n_correct.divide(total_q)\ncurrent_user_data = pd.DataFrame({'n_questions': total_q, 'ratio_q': ratio_q})\ncurrent_user_data.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_test.loc['answered_correctly'] = 0.5\nfor idx, row in train_test.iterrows():\n    # TBD: update current_user_data, inc. n_questions & ratio_q\n    if row.user_id in current_user_data.index:\n        pred = current_user_data.loc[row.user_id, 'ratio_q']\n        print(pred)\n    else:\n        pred = 0.5\n\n    train_test.loc[idx, 'answered_correctly'] = pred","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_test","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Submission"},{"metadata":{"trusted":true},"cell_type":"code","source":"env = riiideducation.make_env()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"iter_test = env.iter_test()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for i, (test_df, sample_prediction_df) in enumerate(iter_test):\n    # Create target (all-0.5-)column\n    test_df['answered_correctly'] = 0.5\n    \n    # Making predictions\n    for idx, row in test_df.iterrows():\n        if row.user_id in current_user_data.index:\n            pred = current_user_data.loc[row.user_id, 'ratio_q']\n        else:\n            pred = 0.5\n        test_df.loc[idx, 'answered_correctly'] = pred    \n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n\n    # Updating knowledge based on latest batch\n    if i > 0:\n        prev_group_answers_correct = test_df.prior_group_answers_correct.iloc[0]\n        if isinstance(prev_group_answers_correct, str):\n            answers = map(int, prev_group_answers_correct.split())\n            answers = pd.DataFrame(answers, index=prev_test_df.index)\n    prev_test_df = test_df.copy()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}