{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Importing Libraries and data"},{"metadata":{},"cell_type":"markdown","source":"## Pls. Upvote if this helps. It motivates me."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nfrom collections import Counter\nfrom wordcloud import WordCloud","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir('../input/riiid-test-answer-prediction')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/riiid-test-answer-prediction/train.csv', nrows=1000000)\nlectures = pd.read_csv('../input/riiid-test-answer-prediction/lectures.csv')\nquestions = pd.read_csv('../input/riiid-test-answer-prediction/questions.csv')\nexample_test = pd.read_csv('../input/riiid-test-answer-prediction/example_test.csv')\nexample_sample_submission = pd.read_csv('../input/riiid-test-answer-prediction/example_sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Before anything else let's see what is our target"},{"metadata":{"trusted":true},"cell_type":"code","source":"example_sample_submission.head(2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"example_test.head(2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Just in case you are wondering why a example_test file then the reason is that the actual test file could only be seen from the time series api of the riid. rest info is here - https://www.kaggle.com/sohier/competition-api-detailed-introduction "},{"metadata":{},"cell_type":"markdown","source":"#### OK!! So we basically have to predict a probability for \"answered_correctly\" for a given \"row_id\" and \"group_num\" With that out of the way let's look at our training data"},{"metadata":{},"cell_type":"markdown","source":"#### Also the evaluation metric is area under the ROC curve."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"if len(train_df) == len(train_df.row_id.unique()):\n    print('row_id column is the key')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print('total train samples ', len(train_df))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Understanding The Problem"},{"metadata":{},"cell_type":"markdown","source":"#### So this is a binary classification and timeseries problem where you predict \"ansered_correctly\" by a student \"user_id\" over time \"timestamp\"."},{"metadata":{},"cell_type":"markdown","source":"we have been provided with `content_id` column which tells us user interestion with the content (not entirely sure wheather its like a book id fixed for that book for all user's or a id generated for each interaction but we will see it later on.), also we have a `content_type_id` column in which 0 means he answered a question and 1 means the content was a lecture and hence he watched a lecture. If he watched a lecture then there was no question hence we need not predict the `answered_correctly` and skip the row for prediction."},{"metadata":{},"cell_type":"markdown","source":"Next, we have `task_container_id` columns which means id's for set's of questions. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a `task_container_id`. ( A correlation/similarity is in a column `bundle_id` from `question.csv` file)"},{"metadata":{},"cell_type":"markdown","source":"Upnext, we have `user_answer` column where user gives a MCQ type answer say 1,2 or 3 etc. and whether that answer was correct or not is recorded in the next `answered_correctly` column which we need to predict."},{"metadata":{},"cell_type":"markdown","source":"Upnext, we have `prior_question_elapsed_time` column which tells how long it took a user to answer their previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Note that the time is the total time a user took to solve all the questions in the previous bundle. (so it depends on the number of questions that bundle had)."},{"metadata":{},"cell_type":"markdown","source":"At the last we have `prior_question_had_explanation` (bool) column which tells whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Typically the first several questions a user sees were part of an onboarding diagnostic test where they did not get any feedback. ( In case the questions are related then this can help much.)"},{"metadata":{},"cell_type":"markdown","source":"## Analizing the train.csv file"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"No of students = \", len(train_df['user_id'].unique()))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"distribution of number of samples per student\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"sns.set()\nfig = plt.figure(figsize=(15,6))\nfig = sns.kdeplot(train_df.groupby(by='user_id').count()['row_id'], shade=True, gridsize=50, color='g', legend=False)\nfig.figure.suptitle(\"User_id distribution\", fontsize = 20)\nplt.xlabel('User_id counts', fontsize=16)\nplt.ylabel('Probability', fontsize=16);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, most of the students have < 2K samples."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"How many question does each student attempt\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df = train_df[train_df['content_type_id'] == 0]\n\ndf = df.groupby(by='user_id').count()\n\nfig = plt.figure(figsize=(15,6))\nfig = sns.kdeplot(df['row_id'], shade=True, gridsize=50, color='r', legend=False)\nfig.figure.suptitle(\"User attempted questions distribution\", fontsize = 20)\nplt.xlabel('Questions counts', fontsize=16)\nplt.ylabel('Probability', fontsize=16)\nplt.legend(['Questions Attempted','Questions Correctly answered'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So they follow the same distribution. Majority of student have attempted < 2K questions."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"distribution of correct and incorrect and no answers\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"df = train_df[train_df['content_type_id'] == 0]\n\ndf2 = df[df['answered_correctly'] == 1]\ndf3 = df[df['answered_correctly'] == 0]\n\ndf2 = df2.groupby(by='user_id').count()\ndf3 = df3.groupby(by='user_id').count()\n\nfig = plt.figure(figsize=(15,6))\nfig = sns.kdeplot(df2['row_id'], shade=True, gridsize=50, color='b', legend=False)\nfig = sns.kdeplot(df3['row_id'], shade=True, gridsize=50, color='r', legend=False)\n\nfig.figure.suptitle(\"User attempted questions distribution\", fontsize = 20)\nplt.xlabel('Questions counts', fontsize=16)\nplt.ylabel('Probability', fontsize=16)\nplt.legend(['Correctly answered','Incorrectly answered'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So we can see that the probability of ansering correctly is mostly 2.5 times that of answering incorrectly."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"What is the distribution of students correctly answering a question ?\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"values = []\n\ndf = train_df[train_df['content_type_id'] == 0]\n\nfor group, frame in df.groupby(by='user_id'):\n    \n    value = len(frame[frame['answered_correctly'] == 1]) / len(frame)\n    values.append(value)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig = plt.figure(figsize=(15,6))\nfig = sns.kdeplot(values, shade=True, gridsize=50, color='b', legend=False)\nfig.figure.suptitle(\"User correctly answering distribution\", fontsize = 20)\nplt.xlabel('Percent Correct', fontsize=16)\n\nprint('MEAN: ', np.mean(values))\nprint(\"MAX: \", np.max(values))\nprint('MIN: ', np.min(values))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"the graph shows that the average accuracy of a student is near 54% in answering correctly. Some have scored perfect and some have scored zero. (We have outliars as getting all zeros is by probability as hard as getting all correct so we have to handle that)"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"What precent of students see explanations ?\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"values = []\n\ndf = train_df[train_df['content_type_id'] == 0]\n\nfor group, frame in df.groupby(by='user_id'):\n    \n    value = len(frame[frame['prior_question_had_explanation'] == True]) / len(frame)\n    values.append(value)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"px.histogram(values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print(\"Total task container id's having questions: \", len((train_df['task_container_id'][train_df['content_type_id'] == 0]).unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is a considerable amount of students who never watched prior explanations and yet answered correctly. Any further than this we will need to use timestamps or other files."},{"metadata":{},"cell_type":"markdown","source":"### Before we move into the real EDA which is with respect to time_stamp let's have a look at questions and lectures files"},{"metadata":{},"cell_type":"markdown","source":"## Questions file"},{"metadata":{"trusted":true},"cell_type":"code","source":"questions.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`question_id`: foreign key for the train/test content_id column, when the content type is question (0).\n\n`bundle_id`: code for which questions are served together.\n\n`correct_answer`: the answer to the question. Can be compared with the train user_answer column to check if the user was right.\n\n`part`: top level category code for the question.\n\n`tags`: one or more detailed tag codes for the question. The meaning of the tags will not be provided, but these codes are sufficient for clustering the questions together."},{"metadata":{"trusted":true},"cell_type":"code","source":"questions.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print('Total number of questions: ', len(questions['question_id'].unique()))\nprint(\"Total number of unique bundles: \", len(questions['bundle_id'].unique()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig = plt.figure(figsize=(10,6))\nfig = sns.countplot(questions.groupby(by='bundle_id').count()['question_id'])\nplt.xlabel('bundle_id')\nplt.title('Question in bundles');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Well I think that the majority of the question sets in task_container_id is from bundle 1"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"fig = plt.figure(figsize=(10,6))\nfig = sns.countplot(questions['correct_answer'])\nplt.xlabel('Answers')\nplt.title('Correct Answers distribution')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So the answers almost have a uniform distribution."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"Distribution of number of tags per question\")\nprint(\"I think of tags as subject includings like (maths, algebra, numbers. history etc. in encoded form)\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"no_of_tags = []\nfor i in questions['tags']:\n    value = len(str(i).strip().split(' '))\n    no_of_tags.append(value)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt.figure(figsize=(10,6))\nsns.countplot(no_of_tags)\nplt.xlabel('No of tags')\nplt.title('No of tags per question')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Quite Superising that the number of tags is not decreasing constantly but has a huge dip at 2."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"distribution of tags\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"total = []\n\nfor i in questions['tags']:\n    for j in str(i).strip().split(' '):\n        total.append(j)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"keys = set(total)\nfinal = {}\nfor i in keys:\n    final[i] = total.count(i)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"values = sorted(final.items(), key=lambda x: x[1], reverse=True)\nd = []\nfor i in values:\n    d.append(i[1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt.figure(figsize=(10,6))\npx.line(d, title='Tags distribution')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The distribution of tags is very skewed. Only 40 tags occur almost > 80% of time."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"tags = WordCloud().generate_from_frequencies(final)\npx.imshow(tags, title='Most frequent Tags')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## lectures file"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"`lecture_id`: foreign key for the train/test content_id column, when the content type is lecture (1).\n\n`part`: top level category code for the lecture. (This confuses me. If we have tag then why do we need part?)\n\n`tag`: one tag codes for the lecture. The meaning of the tags will not be provided, but these codes are sufficient for clustering the lectures together.\n\n`type_of`: brief description of the core purpose of the lecture"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print('Total no. of lectures: ', len(lectures['lecture_id'].unique()))\nprint('Only one tag per row: ', )\nprint('Total no. of tags in lecture: ', len(lectures['tag'].unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now comes the troubling part. Tags in lectures has no relation with tags in questions and hence they could be different altogether."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# distribution of lecture tags\n\ntotal = []\n\nfor i in lectures['tag']:\n    for j in str(i).strip().split(' '):\n        total.append(j)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"keys = set(total)\nfinal = {}\nfor i in keys:\n    final[i] = total.count(i)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"values = sorted(final.items(), key=lambda x: x[1], reverse=True)\nd = []\nfor i in values:\n    d.append(i[1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"plt.figure(figsize=(10,6))\npx.line(d, title='Tags distribution in lectures')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now that is some amazing pattern. The first idea is that the tags could be like subject title. Like lectures with high importance comes from a important chapter and that chapter may have only certen types of tags. Say math chapter may have (Maths, numbers, Algebra) as tags(encoded). "},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Most common tags\n\ntags = WordCloud().generate_from_frequencies(final)\npx.imshow(tags, title='Most frequent lecture Tags')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Looking at parts\nprint('Total type of parts: ', len(lectures.part.unique()))\nprint('Values of parts: ', lectures.part.unique())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Counts of parts\nplt.figure(figsize=(10,6))\nsns.countplot(lectures['part'])\nplt.title('Counts of Parts in Lectures');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# how many different unique tags does each part have or do they common tags as well ?\n\nno_unique_tags_l = []\nunique_tags_l = {}\ngroups = []\n\nfor group, frame in lectures.sort_values(by='part').groupby(by='part'):\n    \n    unique_tags = frame['tag'].unique()\n    no_unique_tags = len(unique_tags)\n    \n    unique_tags_l[group] = unique_tags\n    no_unique_tags_l.append(no_unique_tags)\n    groups.append(group)\n    \nno_unique_tags_l = pd.DataFrame(no_unique_tags_l, columns=['count'])\nno_unique_tags_l['group'] = groups","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Number of unique tags in each part ( here unique means internally part wise)\nplt.figure(figsize=(10,6))\nsns.barplot(x=no_unique_tags_l['group'], y=no_unique_tags_l['count'])\nplt.title('No. of unique tags in each part')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, it also follow the same distribution as above."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"final_unqiue = []\nparts = []\n\n\nfor part, array in unique_tags_l.items():\n    \n    unique_tags = []\n        \n    other_parts = list(unique_tags_l.keys())\n    final = set(other_parts)\n    final.remove(part)\n    \n    for j in final:\n        \n        for k in array:\n            \n            if k not in unique_tags_l[j]:\n                \n                unique_tags.append(k)\n    \n    final_unqiue.append(len(unique_tags))\n    parts.append(part)\n    \nfinal_unqiue = pd.DataFrame(final_unqiue, columns=['tags'])\nfinal_unqiue['part'] = parts","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# let's see how many tags are there in each part which are not in any other part\n\nplt.figure(figsize=(10,6))\nsns.barplot(x=no_unique_tags_l['group'], y=no_unique_tags_l['count'])\nplt.title('No. of unique tags in each part not in any other');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, we could rest assured that one tag comes only in one part."},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Finally let's look at type_of lecture\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"px.bar(lectures, x='type_of', color=lectures['type_of'], labels={'value':'type_of'}, title='Type of lectures distribution Overall')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"px.bar(lectures, x='type_of', color=lectures['type_of'], labels={'value':'type_of'}, title='Type of lectures distribution based on each part', facet_col='part')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Suprisingly intention belongs to only one part and starter is only in 2 parts."},{"metadata":{},"cell_type":"markdown","source":"## Now it's time to use the time stamp and see a few students from train file"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"we will see first 8 students for trends\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"no_students = 8\nscores = []\nuser_ids = []\nquestion_attempted_l = []\ncorrectly_answered_l = []\nprior_questions_explanations = []\n\nfor count, (group, frame) in enumerate(train_df.groupby(by='user_id')):\n    \n    if count == no_students:\n        break\n    \n    frame = frame.sort_values(by='timestamp')\n    \n    percentage = []\n    question_attempted = []\n    correctly_answered = []\n    explanations = []\n    attempted = 0\n    correct_answers = 0\n    explanation = 0\n    \n    df = frame[frame['content_type_id'] == 0]\n    \n    for answered_correctly, had_explanation in zip(df['answered_correctly'], df['prior_question_had_explanation']):\n        \n        attempted += 1\n        question_attempted.append(attempted)\n        \n        if answered_correctly == 1:\n            correct_answers += 1\n            \n        if had_explanation:\n            explanation += 1\n            \n        correctly_answered.append(correct_answers)\n            \n        percent = correct_answers / attempted * 100\n        percentage.append(percent)\n        explanations.append(explanation)\n        \n    \n    scores.append(percentage)\n    user_ids.append(group)\n    question_attempted_l.append(question_attempted)\n    correctly_answered_l.append(correctly_answered)\n    prior_questions_explanations.append(explanations)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Trend in attempted question and correctly answering\n\nplt.figure(figsize=(15,20))\n\nfor i in range(1,9):\n    plt.subplot(4,2,i)\n    plt.plot(question_attempted_l[i-1], question_attempted_l[i-1], label='Questions attempted')\n    plt.plot(question_attempted_l[i-1], correctly_answered_l[i-1], label='Questions correctly answered')\n    plt.plot(question_attempted_l[i-1], scores[i-1], label='Percentage correctly answered')\n    plt.plot(question_attempted_l[i-1], prior_questions_explanations[i-1], label='Prior_questions_explanations')\n    plt.legend()\n    plt.ylim(0,100)\n    plt.xlim(0,50)\n    plt.tight_layout(pad = 2)\n    plt.title(f'user_id: {user_ids[i-1]}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So much to see. So much trends and patterns. Well those who had prior explanation had better results. So the trend has many types. sudden spikes(+ve, -ve), consistency, continuous increment, decrement.<br>\nBad Students: Almost no one started watching explanations until they started performing bad."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Does students time spend on answering prior questions\n\nno_students = 8\ntime_spend_l = []\n\nfor count, (group, frame) in enumerate(train_df.groupby(by='user_id')):\n    \n    if count == no_students:\n        break\n    \n    frame = frame.sort_values(by='timestamp')\n    total_time_spends = []\n    time_spends = 0\n    \n    for time_spend in frame['prior_question_elapsed_time'][frame['content_type_id'] == 0]:\n        \n        if time_spend > 0:\n            time_spends += time_spend\n            total_time_spends.append(time_spends)\n        \n    \n    time_spend_l.append(total_time_spends)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"time_spend_l = np.array(time_spend_l)\nfor index, value in enumerate(time_spend_l):\n    time_spend_l[index] = np.array(time_spend_l[index]) / 10000","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"# Trend in time spend with percentage\n\nplt.figure(figsize=(15,20))\n\nfor i in range(1,9):\n    plt.subplot(4,2,i)\n    plt.plot(question_attempted_l[i-1], correctly_answered_l[i-1], label='Questions correctly answered')\n    plt.plot(question_attempted_l[i-1][1:], time_spend_l[i-1], label='time spend in 10000')\n    plt.plot(question_attempted_l[i-1], scores[i-1], label='Percentage correctly answered')\n    plt.legend()\n    plt.ylim(0,100)\n    plt.xlim(0,50)\n    plt.tight_layout(pad = 2)\n    plt.title(f'user_id: {user_ids[i-1]}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There is mostly a linear increase in prior question time elapsed."},{"metadata":{},"cell_type":"markdown","source":"## WORK IN PROGRESS"},{"metadata":{},"cell_type":"markdown","source":"## Pls. Upvote if this helps. It motivates me."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}