{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Import required Libraries"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"%matplotlib inline\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport plotly.express as px","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load data"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', low_memory=False, nrows=10**5, index_col=0,\n                       dtype={'row_id': 'int64', 'timestamp': 'int64', 'user_id': 'int32', 'content_id': 'int16', 'content_type_id': 'int8',\n                              'task_container_id': 'int16', 'user_answer': 'int8', 'answered_correctly': 'int8', 'prior_question_elapsed_time': 'float32', \n                             'prior_question_had_explanation': 'boolean',\n                             }\n                      )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Initial exploration"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df['content_type_id'].value_counts().plot(kind='bar', title='Questions vs Lectures')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['content_type_id'].value_counts()/len(train_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We will try to analyze the questions that were answered, by filtering the records for questions alone"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df = train_df[train_df['content_type_id'] == 0]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's try to group some `user_id` and generate statistics"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df.groupby(['user_id', 'answered_correctly'])\\\n        .agg({'prior_question_elapsed_time':np.mean}).head(2000)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"train_df.groupby(['user_id', 'answered_correctly'])\\\n        .agg({'prior_question_elapsed_time':np.mean}).head(2000)\\\n        .groupby('user_id').agg(\n    {'prior_question_elapsed_time':lambda x: x.values[0]-x.values[1] if len(x) == 2 else x.values[0]})","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Above I have tried to generate a user level difference of time taken between correctly answered questions and incorrectly answered questions, if the students had taken less for answering correct if compared to the incorrect ones, the difference must be **negative** and vice-versa for the other scenario.\n\n_(Assuming 0 to be False and 1 to be True)_\n\nThis maybe helpful to get patterns from time taken for a question, We can see that there are some big differences in time elapsed for some users "},{"metadata":{},"cell_type":"markdown","source":"Let's say that each student maybe weak in some topics and stronger in some topics, identifying this can be helpful, to know if a student can answer a question from those topics correctly."},{"metadata":{},"cell_type":"markdown","source":"## Looking at the questions"},{"metadata":{"trusted":true},"cell_type":"code","source":"qdf = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/questions.csv')\nqdf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"qdf.info()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's check if the `bundle_id` and `question_id` columns are not same"},{"metadata":{"trusted":true},"cell_type":"code","source":"(qdf['bundle_id'] == qdf['question_id']).value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"No they are not, guess this was a bad assumption, since the `bundle_id` may have questions which are similar to one another or of the same topic of the lecture viewed."},{"metadata":{"trusted":true},"cell_type":"markdown","source":"Let's see the different types of tags for the questions"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"from collections import Counter","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"tags = []\nfor tag_str in qdf['tags'].values.tolist():\n    if not isinstance(tag_str, float):\n        tags.extend(tag_str.split(' '))\n\ncounts = dict(Counter(tags))\ncounts_df = pd.DataFrame.from_dict(counts, orient='index', columns=['Count']).reset_index()\ncounts_df = counts_df.rename({'index':'tag_id'}, axis='columns')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"counts_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig = px.bar(counts_df, x='tag_id', y='Count', title='Tag counts')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We should probably cluster this, but I don't know how"},{"metadata":{},"cell_type":"markdown","source":"# Looking at the `answered_correctly` as a time series problem for each user"},{"metadata":{},"cell_type":"markdown","source":"The competition description clearly states that we need to trace the knowledge of that particular student overtime to understand if he can answer the incoming question correctly."},{"metadata":{},"cell_type":"markdown","source":"Let's choose some well represented `user_id`s and try to look at the cumulative number of correctly answered questions for each of them overtime"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"uids = train_df['user_id'].value_counts().index.tolist()[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"random_user = train_df.loc[train_df['user_id'].isin(uids), ['user_id', 'timestamp', 'answered_correctly']]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"random_user['timestamp'] = random_user['timestamp']/1000","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"random_user.reset_index(drop=True, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"random_user['corr_cs'] = random_user.groupby('user_id').agg({'answered_correctly':np.cumsum})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig = px.line(random_user, x=\"timestamp\", y=\"corr_cs\", color='user_id')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"Some users have a long gap, it probably must be due to them watching lectures in the middle before moving on to the next bundle of questions"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}