{"cells":[{"metadata":{},"cell_type":"markdown","source":"The meaning of the features we are dealing with is very important, especially when they are  specific and few in quantity. In this post, I will dig deeper in some of the features. On the one hand I was confused with these features just by reading data description. On the other hand, I hope it can help you if you have the similar confusion. (This post may or may not update in future)\n\n* [Main data](#main-data)\n* [Lecture data](#lecture-data)\n* [Question data](#question-data)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', \n                       dtype={'row_id': 'int64', 'timestamp': 'int64', 'user_id': 'int32', 'content_id': 'int16', 'content_type_id': 'int8',\n                              'task_container_id': 'int16', 'user_answer': 'int8', 'answered_correctly': 'int8', 'prior_question_elapsed_time': 'float32', \n                             'prior_question_had_explanation': 'boolean',\n                             }\n                      )\nlectures_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/lectures.csv')\nquestions_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/questions.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Main Data\n### task_container_id\n\n``task_container_id`` is used to id which couple qustions/lectures are showed at the same time. As you can see below that all three questions happened at the same timestamp."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"train_df[40:43]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### user_answer\n``user_answer`` has only 5 unique values. Since the data description only says ``the user's answer to the question``, my opinion is that this value may indicate some kinds of categories or measurement of the answer. It's not very realistic that there are only 4 answers to all these questions. Now I said it. It is possible that all the questions are multiple choice/multiple response questions."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['user_answer'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Upon checking unique values of ``correct_answer``, every question has only one answer. If the questions are choice questions, they are multiple choice."},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df['correct_answer'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Lectures Data\n### tag\n\nThere is only one ``tag`` for each lecture."},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures_df['tag'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### part\n\n``part`` column has 7 unique ids which are compatible with the 7 test parts mentioned in this link [section of the TOEIC test](https://www.iibc-global.org/english/toeic/test/lr/about/format.html). It further comfirems that the questions are multiple choice questions, since it mentions ``Mark your answers on your answer sheet``."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"lectures_df['part'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Question Data\n### bundle_id\n\n``bundle_id`` is used to id which couple qustions are served together, as shown below. "},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df.loc[questions_df.duplicated('bundle_id', keep=False),:][:3]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now we choose ``question_id`` 1400 and find one of the users who take question 1400. Observing it's neighbour records, we can see questions belong the same bundle show at the same time. They have the same ``task_container_id``, but its value is different with ``bundle_id``."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"train_df.iloc[1023702:1023702+5,:]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's check one more user just to be safe. We can reach the same conclution from below data. Moreover, eventhough these two users received the same bundle, the values of `` task_container_id`` are different."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"train_df.iloc[94636590:94636590+5,:]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Please share with us with your exploration or confusion.   \nIf you find it helpful, please <font size='5'>upvote</font>."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}