{"cells":[{"metadata":{},"cell_type":"markdown","source":"Still in development.  \nUsed this [notebook](https://www.kaggle.com/ilialar/simple-eda-and-baseline) as inspiration."},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\n\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"for dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"NROWS=10**7","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Complete description"},{"metadata":{"trusted":true},"cell_type":"code","source":"dtype = {'row_id': 'int64', 'timestamp': 'int64', 'user_id': 'int32', 'content_id': 'int16', \n         'content_type_id': 'int8','task_container_id': 'int16', 'user_answer': 'int8', 'answered_correctly': 'int8', \n         'prior_question_elapsed_time': 'float32','prior_question_had_explanation': 'boolean',\n        }\n\nnrows = 10**7\ntrain = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', chunksize=nrows, dtype=dtype)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"j=0\nfor i in train:\n    print(i.info(null_counts=True))\n    j+=1\nj-=1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Number of rows: %f' % (j*nrows+len(i)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Analysis of questions ('content_type_id'=0)\n"},{"metadata":{},"cell_type":"markdown","source":"We can see that there are very few lectures compared to questions.  \nWe will assess the dataset without the rows containning missing rows"},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', nrows=NROWS, dtype=dtype)\ntrain = train.dropna()\ntrain = train[train['answered_correctly']!=-1]\ntrain","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"No missing values in the questions."},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train[train['content_type_id']==0]\ntrain","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Proportion of questions:%f' % (train.count().row_id/NROWS))\nprint('Proportion of unique questions:%f' % (len(train.content_id.unique())/len(train)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 1. timestamp column\n\nHistogram of timestamp"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.timestamp.hist(bins=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.timestamp.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Histogram of timestamp by user"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.iloc[:10**5].groupby('user_id').timestamp.hist(bins=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.iloc[:10**5].groupby('user_id').timestamp.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As expected, we can see different different behaviour for different users."},{"metadata":{},"cell_type":"markdown","source":"### 2. user_id column\n\nQuestions descriptions by user"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Average number of question per student: %f' % train.groupby('user_id').count().mean().row_id)\nprint('Standard deviation of number of question per student: %f' % train.groupby('user_id').count().std().row_id)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(train.groupby('user_id').count().describe().row_id)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Again, there are a lot of variations depending on the user and few users do many questions while most of them do few questions."},{"metadata":{"trusted":true},"cell_type":"code","source":"train.groupby('user_id').count().row_id.hist(bins=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.groupby('user_id').count().row_id.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 3. task_container_id column"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Proportion of unique task_container_id: %f' % (len(train.task_container_id.unique())/NROWS))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.groupby('task_container_id').count().describe().row_id","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 4. content_id column"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Percentage of questions that appears only once: %f' % ((train.groupby('content_id').count().row_id==1).mean()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Percentage of unique question: %f' % (len(train.groupby('content_id').count())/len(train)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.content_id.hist(bins=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.content_id.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Some questions appear more often than others and they are of course not ordered. Should be careful to not modelize this feature as ordered in the model."},{"metadata":{},"cell_type":"markdown","source":"### 5. prior_question_elapsed_time column"},{"metadata":{},"cell_type":"markdown","source":"Average time since the last question answered for a student and its standard deviation"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(train.groupby('user_id').mean().prior_question_elapsed_time.mean())\nprint(train.groupby('user_id').mean().prior_question_elapsed_time.std())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.prior_question_elapsed_time.hist(bins=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.prior_question_elapsed_time.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.iloc[:10**5].groupby('user_id').prior_question_elapsed_time.hist(bins=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.iloc[:10**5].groupby('user_id').prior_question_elapsed_time.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 5. prior_question_had_explanation column\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Percentage of questions that had an explanation: \\n%s' % (train.prior_question_had_explanation.value_counts()/NROWS))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Description of True prior_question_had_explanation per user: \\n%s' % (train.groupby(['user_id', 'prior_question_had_explanation']).count().row_id.reset_index()[train.groupby(['user_id', 'prior_question_had_explanation']).count().row_id.reset_index().prior_question_had_explanation==True].describe().row_id))\nprint('Number of user without True value: %d' % (len(train.user_id.unique())-len((train.groupby(['user_id', 'prior_question_had_explanation']).count().row_id.reset_index()[train.groupby(['user_id', 'prior_question_had_explanation']).count().row_id.reset_index().prior_question_had_explanation==True]))))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Description of False prior_question_had_explanation per user: \\n%s' % (train.groupby(['user_id', 'prior_question_had_explanation']).count().row_id.reset_index()[train.groupby(['user_id', 'prior_question_had_explanation']).count().row_id.reset_index().prior_question_had_explanation==False].describe().row_id))\nprint('Number of user without False value: %d' % (len(train.user_id.unique())-len((train.groupby(['user_id', 'prior_question_had_explanation']).count().row_id.reset_index()[train.groupby(['user_id', 'prior_question_had_explanation']).count().row_id.reset_index().prior_question_had_explanation==False]))))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 8. answered_correctly column"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Percentage of questions answered correctly: %f' % train['answered_correctly'].mean())\ntrain['answered_correctly'].hist(bins=10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Histograms of percentage of questions answered corretly by user"},{"metadata":{"trusted":true},"cell_type":"code","source":"count_answered_correctly_true_per_user = (train.groupby(['user_id', 'answered_correctly']).count().reset_index()[train.groupby(['user_id', 'answered_correctly']).count().reset_index().answered_correctly==1].set_index('user_id'))\nresults = train.groupby('user_id').count()\nresults.row_id = 0\nresults.loc[count_answered_correctly_true_per_user.index, 'row_id'] = count_answered_correctly_true_per_user.row_id","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(results.row_id/train.groupby('user_id').count().row_id).hist(bins=10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(results.row_id/train.groupby('user_id').count().row_id).hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"((results[results.row_id<50].row_id)/(train.groupby('user_id').count()[results.row_id<50].row_id)).hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(results[results.row_id>=50].row_id/train.groupby('user_id').count()[results.row_id>50].row_id).hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(results[results.row_id>=50].row_id/train.groupby('user_id').count()[results.row_id>1000].row_id).hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Analysis of lecture ('content_type_id'=1)"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', nrows=NROWS, dtype=dtype)\nlectures = lectures[lectures['content_type_id']==1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures.groupby('user_id').count().row_id.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Might be interesting to have a feature that keep informations about the lectures one student followed (number of lectures, type of lectures, did they do it since a long time?...)"},{"metadata":{"trusted":true},"cell_type":"code","source":"questions = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', nrows=NROWS, dtype=dtype)\nquestions = questions.dropna()\nquestions = questions[questions['answered_correctly']!=-1]\n\nlectures_count = questions.groupby('user_id').answered_correctly.mean()\nlectures_count.loc[:] = 0\nlectures_count.loc[lectures.groupby('user_id').count().index] = lectures.groupby('user_id').count().row_id\n\nplt.scatter(questions.groupby('user_id').answered_correctly.mean(), lectures_count)\nplt.xlabel(\"Correctness rate per student\")\nplt.ylabel(\"Number of lectures attended per student\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that students who attend lectures have a higher correctness rate."},{"metadata":{},"cell_type":"markdown","source":"# Questions"},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_type = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/questions.csv')\nquestions_type","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print((questions_type[(questions_type.question_id != questions_type.bundle_id)]))\nprint('\\nPercentage of questions served together: %f\\n' % (len(questions_type[(questions_type.question_id != questions_type.bundle_id)])/len(questions_type)))\nprint('Description of unique value of bundle_id:\\n%s\\n' % questions_type.bundle_id.value_counts().describe())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Description of unique value of part:\\n%s\\n' % questions_type.part.value_counts().describe())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"all_tags=[]\nfor j in [y.split() for y in questions_type['tags'].astype(str).values]:\n    for i in j:\n        all_tags.append(i)\nprint('Description of unique value of tags:\\n%s\\n' % pd.Series(all_tags).value_counts().describe())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Some tags appear very often while other don't. Very unbalanced."},{"metadata":{},"cell_type":"markdown","source":"# Lectures"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures_type = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/lectures.csv')\nlectures_type","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Description of unique value of part:\\n%s\\n' % lectures_type.part.value_counts().describe())\nprint('Description of unique value of tag: \\n%s\\n' % lectures_type.tag.value_counts().describe())\nprint('Value counts of the type of unique questions: \\n%s' %lectures_type.type_of.value_counts())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Very few intention and starter lectures. Should check if many students attended this type of lectures and should be careful the model doesn't overfit on this. Can apply the same reasoning with tag features/some tag appears only once."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}