{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n        \nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nimport plotly.express as px\nfrom plotly.subplots import make_subplots\nimport plotly.graph_objs as go\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# train.csv"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"data_types_dict = {\n    'row_id': 'int64',\n    'timestamp': 'int64',\n    'user_id': 'int32',\n    'content_id': 'int16',\n    'content_type_id': 'int8',\n    'task_container_id': 'int16',\n    'user_answer': 'int8',\n    'answered_correctly': 'int8',\n    'prior_question_elapsed_time': 'float16',\n    'prior_question_had_explanation': 'boolean'\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv',\n                       low_memory=False,\n                       nrows=10**7,\n                       dtype=data_types_dict, \n                      )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# 欠損値\nprint('Part of missing values for every column')\nprint(train_df.isnull().sum() / len(train_df))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.describe().T","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cols = train_df.columns\n\nfor col in cols:\n    print(f'Unique values in {col} : {train_df[col].nunique()}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"categoricalカテゴリ変数\n\ncontent_type_id, user_answer , answered_correctly ,prior_question_had_explanation\n\nnow we can see that there are some very low integer we convert the columns content_type_id, user_answer , answered_correctly ,prior_question_had_explanation to categorical format when we train a model"},{"metadata":{},"cell_type":"markdown","source":"### timestamp"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['timestamp'].hist(bins=100);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"timestamp represents the time from the first user interaction to the current one. It is expected that the distribution looks like this."},{"metadata":{},"cell_type":"markdown","source":"timestamp・・・ユーザーとの対話からそのイベント終了までの時間\n\ntimestamp is defined as \"the time between this user interaction and the first event from that user\"\n\nprior_question_elapsed_time is defined as \"How long it took a user to answer their previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Note that the time is the total time a user took to solve all the questions in the previous bundle\"\n\nThe timestamp column shows when an activity is finished, not when it started. \n\nThe timestamp timer starts after first question is answered or lecture is finished.\n\nprior_question_elapsed_time timer starts when the user starts doing the previous question and it ends when the user moves to another question.\n\nmaybe timestamp is miliseconds. it cannot be seconds.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189351"},{"metadata":{"trusted":true},"cell_type":"code","source":"grouped_by_user_df = train_df.groupby('user_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"grouped_by_user_df.agg({'timestamp':'max'}).hist(bins=100);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"↑各ユーザーの最大のtimestampの分布・・・ほとんどのユーザーがすぐにプラットフォームを離れるようだ。\n\nThe distribution of the max timestamp for each user looks similar. It seems most users leave the platform quite soon (at least based on partial data we analyze)."},{"metadata":{},"cell_type":"markdown","source":"### Answered correctly\n ユーザーが正しく応答したかどうか。講義と質問がある。講義（lectures）の場合は、-1をnullとして読み取ります。質問の場合は、正答１、誤答０"},{"metadata":{"trusted":true},"cell_type":"code","source":"# 講義の割合  # 平均 -1 (True)の割合\n(train_df['answered_correctly'] == -1).mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['content_type_id'].value_counts().reset_index()\nds.columns = ['content_type_id', 'percent']\nds['percent'] /= len(train_df)\n\nfig = px.pie(\n    ds, \n    names='content_type_id', \n    values='percent', \n    title='Lecures & questions', \n    height=500, \n    width=600\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"trainデータの約2%は、「講義」である。→回答分析から除外する必要がある。\n\n2% of activities are lectures, we should exclude them for answers analysis."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_questions_only_df = train_df[train_df['answered_correctly'] != -1]\ntrain_questions_only_df['answered_correctly'].mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['answered_correctly'].value_counts().reset_index()\nds.columns = ['answered_correctly', 'percent_of_answers']\nds['percent_of_answers'] /= len(train_df)\nds = ds.sort_values(['percent_of_answers'])\n\nfig = px.pie(\n    ds, \n    names='answered_correctly', \n    values='percent_of_answers', \n    title='Percent of correct answers', \n    height=500, \n    width=600\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"correct = train_df[train_df.answered_correctly != -1].answered_correctly.value_counts()\n\nfig = plt.figure(figsize=(12,4))\n\ncorrect.plot.barh()\nplt.title(\"Questions answered correctly\")\nplt.xticks(rotation=0)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"平均して、ユーザーは最大６６％の質問に正しく答えている。 →ユーザーごとにどのくらい違うかも見てみる\n\nOn average users answer ~66% questions correctly. Let's look how it is different from user to user.\n\n「講義」を除外した、answered_correctlyをみてみると、１／３は質問に間違えている。\n\nWhen looking at the numbers of answered_correctly, we see the same number of missing answers. Without looking at the lecture interactions, we see about 1/3 of the questions was answered incorrectly."},{"metadata":{},"cell_type":"markdown","source":"### Answers by users"},{"metadata":{"trusted":true},"cell_type":"code","source":"grouped_by_user_df = train_questions_only_df.groupby('user_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# 回答率('mean')と回答数（'count'）で分ける\nuser_answers_df = grouped_by_user_df.agg({'answered_correctly': ['mean', 'count']})\nuser_answers_df[('answered_correctly', 'mean')].hist(bins=100); # bins = 棒の数","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_answers_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_answers_df[('answered_correctly', 'count')].hist(bins=100);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"(user_answers_df[('answered_correctly','count')]< 50).mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"ユーザーの54％が、50未満の質問に回答。 → すべてのユーザーを「初心者」と「アクティブユーザー」に分けてみる。\n\n54% of users answered less than 50 questions. Let's divide all users into novices and active users."},{"metadata":{"trusted":true},"cell_type":"code","source":"# 初心者の正答率\nuser_answers_df[user_answers_df[('answered_correctly', 'count')] < 50][('answered_correctly', 'mean')].mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_answers_df[user_answers_df[('answered_correctly', 'count')] < 50][('answered_correctly', 'mean')].hist(bins=100);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# アクティブユーザーの正答率\nuser_answers_df[user_answers_df[('answered_correctly', 'count')] >= 50][('answered_correctly', 'mean')].mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_answers_df[user_answers_df[('answered_correctly', 'count')] >= 50][('answered_correctly', 'mean')].hist(bins=100);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"アクティブユーザーは、初心者よりもはるかに優れている。\n\n全体の平均66%　しかし、平均ユーザースコアは、正解の全体の66％よりも低くなっている。→これは、ヘビーユーザーのスコアがさらに高くなることを意味する。\n\nWe can see that active users do much better than novices. But anyway average user score is lower than the overall % of correct answers. It means heavy users have even better scores. Let's look at them."},{"metadata":{"trusted":true},"cell_type":"code","source":"# ヘビーユーザーの割合 500以上questionを回答しているユーザーの割合\n(user_answers_df[('answered_correctly','count')] >= 500).mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# ヘビーユーザーの回答率の分布\nuser_answers_df[user_answers_df[('answered_correctly', 'count')] >= 500][('answered_correctly', 'mean')].hist(bins=100);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# ヘビーユーザーの正答率\nuser_answers_df[user_answers_df[('answered_correctly', 'count')] >= 500][('answered_correctly', 'mean')].mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.scatter(x = user_answers_df[('answered_correctly', 'count')], y = user_answers_df[('answered_correctly', 'mean')]);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### これまでのまとめ\n・Timestamp, ・アクティブユーザーの平均スコア, ・回答された質問の数、はベースラインの作成に役立ちそう。\n\nTimestamp, the average score for the active user, and the number of questions answered can be useful for baseline."},{"metadata":{},"cell_type":"markdown","source":"### Answers by content"},{"metadata":{"trusted":true},"cell_type":"code","source":"grouped_by_content_df = train_questions_only_df.groupby('content_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"content_answers_df = grouped_by_content_df.agg({'answered_correctly': ['mean', 'count']})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"content_answers_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"content_answers_df[('answered_correctly', 'count')].hist(bins=100);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"content_answers_df[('answered_correctly', 'mean')].hist(bins=100);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"質問(content_id)が異なれば、answered_correctlyも異なるため、ベースラインに使えそう。\n\nDifferent questions have different popularity and complexity, and it can also be used in the baseline."},{"metadata":{"trusted":true},"cell_type":"code","source":"content_answers_df[content_answers_df[('answered_correctly','count')]>50][('answered_correctly','mean')].hist(bins = 100);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Top 40 users by number of actions"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['user_id'].value_counts().reset_index()\nds.columns = ['user_id', 'count']\n\nds['user_id'] = ds['user_id'].astype(str) + '-'\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds.tail(40),\n    x='count',\n    y='user_id',\n    orientation='h', # horizontal bar char 横水平バー\n    title='Top40 users by number of actions',\n    height=900,\n    width=700\n)\n\nfig","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"一番多いユーザーで、15,871回データ出現している。"},{"metadata":{},"cell_type":"markdown","source":"### User action distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['user_id'].value_counts().reset_index()\nds.columns = ['user_id', 'count']\nds = ds.sort_values('user_id')\n\nfig = px.line(\n    ds, \n    x='user_id', \n    y='count', \n    title='User action distribution', \n    height=600, \n    width=900\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Top 40 most useful content_ids"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['content_id'].value_counts().reset_index()\nds.columns = ['content_id', 'count']\nds['content_id'] = ds['content_id'].astype(str) + '-'\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds.tail(40), \n    x='count', \n    y='content_id', \n    orientation='h', \n    title='Top40 most useful content_ids', \n    height=900, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"c_ids = train_df.content_id.value_counts()[:40]\n\nfig = plt.figure(figsize=(12,8))\n\nc_ids.plot.bar()\nplt.title(\"Top 40 most used content id's\")\nplt.xticks(rotation=90)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### content_id action distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['content_id'].value_counts().reset_index()\nds.columns = ['content_id', 'count']\nds = ds.sort_values('content_id')\n\nfig = px.line(\n    ds, \n    x='content_id', \n    y='count', \n    title='content_id action distribution', \n    height=600, \n    width=900\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Top 40 most useful task_container_id\n\ntask_container_id: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id.\n\n質問や講義のひとかたまりを表したIDコード\n\n例：　説明を見る前に、3つの質問を見たらそれらをtask_container_idとしてシェアしておく。\n\n⇨つまり、trainの「user_answer(ユーザーの回答)」だけではなく、(その答える時に「他の選択肢」も含んだ)がわかるIDカラム(外部キー)"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['task_container_id'].value_counts().reset_index()\nds.columns = ['task_container_id', 'count']\nds['task_container_id'] = ds['task_container_id'].astype(str) + '-'\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds.tail(40), \n    x='count', \n    y='task_container_id', \n    orientation='h', \n    title='Top 40 most useful task_container_ids', \n    height=900, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### task_container_id action distribution"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['task_container_id'].value_counts().reset_index()\nds.columns = ['task_container_id', 'count']\nds = ds.sort_values('task_container_id')\n\nfig = px.line(\n    ds, \n    x='task_container_id', \n    y='count', \n    title='task_container_id action distribution', \n    height=600, \n    width=800\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"task_container_id が小さいほど、データの出現回数も多い。"},{"metadata":{"trusted":true},"cell_type":"code","source":"# size() 全要素数を取得\ntask_id_correct = train_df[train_df.answered_correctly != -1].\\\ngroupby([\"task_container_id\", 'answered_correctly'], as_index=False).size()\n\ntask_id_correct","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"task_id_correct = task_id_correct.pivot(index='task_container_id',\\\n                                         columns='answered_correctly', values='size')\n\n# 正答率\ntask_id_correct['Percent Correct'] = round(task_id_correct.iloc[:,1]/(task_id_correct.iloc[:,0] + task_id_correct.iloc[:,1]),2)\n\n# %ごとに並び替え\ntask_id_correct = task_id_correct.sort_values(by = \"Percent Correct\", ascending = False)\n\ntask_id_correct","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = task_id_correct['Percent Correct'].value_counts().reset_index()\nds.columns = ['Percent Correct', 'count']\nds = ds.sort_values('Percent Correct')\n\nfig = px.line(\n    ds, \n    x='Percent Correct', \n    y='count', \n    title='Percent Correct action distribution of task_container_id', \n    height=600, \n    width=900\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"task_id_correct = train_df[train_df.answered_correctly != -1].\\\ngroupby([\"task_container_id\", 'answered_correctly'], as_index=False).size()\n\ntask_id_correct = task_id_correct.pivot(index='task_container_id',\\\n                                         columns='answered_correctly', values='size')\n\ntask_id_correct['Percent Correct'] = round(task_id_correct.iloc[:,1]/(task_id_correct.iloc[:,0] + task_id_correct.iloc[:,1]),2)\ntask_id_correct = task_id_correct.sort_values(by = \"Percent Correct\", ascending = False)\n\n# task_container_id - %\ntask_id_correct = task_id_correct.iloc[:,2]\n\ntask_id_correct = task_id_correct[:40]\n\nfig = plt.figure(figsize=(12,6))\ntask_id_correct.plot.bar()\nplt.title(\"Top 40 hardest batches of questions\")\nplt.xticks(rotation=90)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"回答が正解になりやすい質問たち(簡単な問題？）の割合\n\nyou can see the Top-40 of question batches with the highest percentage of questions answered correct."},{"metadata":{},"cell_type":"markdown","source":"### Percent of user answers for every option ユーザーが回答した「選択肢」の割合"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = train_df['user_answer'].value_counts().reset_index()\nds.columns = ['user_answer', 'percent_of_answers']\nds['percent_of_answers'] /= len(train_df)\nds = ds.sort_values(['percent_of_answers'])\n\nfig = px.bar(\n    ds, \n    x='user_answer', \n    y='percent_of_answers', \n    orientation='v', \n    title='Percent of user answers for every option', \n    height=500, \n    width=600\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"-1 は、lecture(講義)のため、null扱い"},{"metadata":{},"cell_type":"markdown","source":"### Percent of correct answers for every option 「選択肢」ごとの回答正解率\n\nこれまでのやつ\n\n* [全体] 平均して、ユーザーは最大６６％の質問に正しく答えている。\n* [初心者] 正解率 約48%\n* [アクティブユーザー] 正解率 約62%"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = make_subplots(rows=2, cols=3)\n\ntraces = [\n    go.Bar(\n        x=[-1, 0, 1], \n        y=[\n            len(train_df[(train_df['user_answer']==item) & (train_df['answered_correctly'] == -1)]),\n            len(train_df[(train_df['user_answer']==item) & (train_df['answered_correctly'] == 0)]),\n            len(train_df[(train_df['user_answer']==item) & (train_df['answered_correctly'] == 1)])\n        ], \n        name='Option: ' + str(item),\n        text = [\n            str(round(100 * len(train_df[(train_df['user_answer']==item) & (train_df['answered_correctly'] == -1)]) / len(train_df[(train_df['user_answer']==item)]), 2)) + '%',\n            str(round(100 * len(train_df[(train_df['user_answer']==item) & (train_df['answered_correctly'] == -0)]) / len(train_df[(train_df['user_answer']==item)]), 2)) + '%',\n            str(round(100 * len(train_df[(train_df['user_answer']==item) & (train_df['answered_correctly'] == 1)]) / len(train_df[(train_df['user_answer']==item)]), 2)) + '%',\n        ],\n        textposition='auto'\n    ) for item in train_df['user_answer'].unique().tolist()\n]\n\nfor i in range(len(traces)):\n    fig.append_trace(traces[i], (i // 3) + 1, (i % 3)  +1)\n\nfig.update_layout(\n    title_text='Percent of correct answers for every option',\n    height=600,\n    width=900\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"「選択肢」ごとの正答率\n\n* option0 65.93%\n* option1 64.93%\n* option2 66.95%\n* option3 66.00%\n* option-1 NaN"},{"metadata":{},"cell_type":"markdown","source":"### prior_question_elapsed_time distribution\n\nprior_question_elapsed_time: (float32) The average time in milliseconds it took a user to answer each question in the previous question bundle, ignoring any lectures in between. Is null for a user's first question bundle or lecture. Note that the time is the average time a user took to solve each question in the previous bundle.\n\n前の質問に回答してから、どのくらいミリセカンド秒経ったか。 NULL = 講義か初めての質問の場合。 このカラムは、「前の質問にどのくらい解決時間を要したか」の参考になる"},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(\n    train_df, \n    x=\"prior_question_elapsed_time\",\n    nbins=100,\n    width=700,\n    height=500,\n    title='prior_question_elapsed_time distribution'\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"10k = 10,000 miliseconds = 10 seconds(秒)\n\n分布見た感じ、16~25秒くらいがボリュームゾーン"},{"metadata":{},"cell_type":"markdown","source":"# Questions.csv"},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df = pd.read_csv('../input/riiid-test-answer-prediction/questions.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# 欠損値\nprint('Part of missing values for every column')\nprint(questions_df.isnull().sum() / len(questions_df))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"There are {len(questions_df['part'].unique())} different parts\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df['tags'].values[-1] # なんで最後の行のtagを取得してるのか？ → データの型を確認しているだけ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"unique_tags = set().union(*[y.split() for y in questions_df['tags'].astype(str).values])\n\nprint(f\"There are {len(unique_tags)} different tags\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# [question_id] content_type_idが質問(0)のとき、train/test content_id列の外部キー / [bundle_id] 質問と一緒に提供されるコード\n(questions_df['question_id'] != questions_df['bundle_id']).mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Number of correct answers per group 正解のナンバーの割合"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = questions_df['correct_answer'].value_counts().reset_index()\nds.columns = ['correct_answer', 'number_of_answers']\nds['correct_answer'] = ds['correct_answer'].astype(str) + '-'\nds = ds.sort_values(['number_of_answers'])\n\nfig = px.bar(\n    ds, \n    x='number_of_answers', \n    y='correct_answer', \n    orientation='h', \n    title='Number of correct answers per group', \n    height=400, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Parts distribution Partの分布"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = questions_df['part'].value_counts().reset_index()\nds.columns = ['part', 'count']\nds['part'] = ds['part'].astype(str) + '-'\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='part', \n    orientation='h', \n    title='Parts distribution', \n    height=500, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"part5のquestionが多い。"},{"metadata":{},"cell_type":"markdown","source":"### Number tags distribution tagsの数の分布\n\ntags: one or more detailed tag codes for the question. The meaning of the tags will not be provided, but these codes are sufficient for clustering the questions together.\n\ntags 意味は特にないけど、これらのコードは質問と一緒にクラスタリングする際に効果的らしい。"},{"metadata":{"trusted":true},"cell_type":"code","source":"# tagsの個数が何個あるかを示すカラム questions_df_copy\nquestions_df_copy = questions_df\nquestions_df_copy['tag'] = questions_df_copy['tags'].str.split(' ')\nquestions_df_copy = questions_df_copy.explode('tag')\nquestions_df_copy = pd.merge(questions_df_copy, questions_df_copy.groupby('question_id')['tag'].count().reset_index(), on='question_id')\nquestions_df_copy = questions_df_copy.drop(['tag_x'], axis=1)\nquestions_df_copy.columns = ['question_id', 'bundle_id', 'correct_answer', 'part', 'tags', 'tags_number']\nquestions_df_copy = questions_df_copy.drop_duplicates()\n\nquestions_df_copy","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = questions_df_copy['tags_number'].value_counts().reset_index()\nds.columns = ['tags_number', 'count']\nds['tags_number'] = ds['tags_number'].astype(str) + '-'\nds = ds.sort_values(['tags_number'])\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='tags_number', \n    orientation='h', \n    title='Number tags distribution', \n    height=400, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Top 40 most useful tags tagsの出現回数"},{"metadata":{"trusted":true},"cell_type":"code","source":"check = questions_df['tags'].str.split(' ').explode('tags').reset_index()\ncheck = check['tags'].value_counts().reset_index()\n\ncheck.columns = ['tag', 'count']\ncheck['tag'] = check['tag'].astype(str) + '-'\ncheck = check.sort_values(['count'])\n\nfig = px.bar(\n    check.tail(40), \n    x='count', \n    y='tag', \n    orientation='h', \n    title='Top 40 most useful tags', \n    height=900, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# lectures.csv\n\n講義内容の詳細データ"},{"metadata":{"trusted":true},"cell_type":"code","source":"lectures_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/lectures.csv')\nlectures_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# 欠損値\nprint('Part of missing values for every column')\nprint(lectures_df.isnull().sum() / len(lectures_df))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Top 40 lectures by number of tags 講義タグの数ランキング"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = lectures_df['tag'].value_counts().reset_index()\nds.columns = ['tag', 'count']\nds['tag'] = ds['tag'].astype(str) + '-'\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds.tail(40), \n    x='count', \n    y='tag', \n    orientation='h', \n    title='Top 40 lectures by number of tags', \n    height=800, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Parts distribution\n\npart: top level category code for the lecture."},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = lectures_df['part'].value_counts().reset_index()\nds.columns = ['part', 'count']\nds['part'] = ds['part'].astype(str) + '-'\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='part', \n    orientation='h', \n    title='Parts distribution', \n    height=500, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### type_of column distribution 講義(種類・内容)の内訳"},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = lectures_df['type_of'].value_counts().reset_index()\nds.columns = ['type_of', 'count']\nds = ds.sort_values(['count'])\n\nfig = px.bar(\n    ds, \n    x='count', \n    y='type_of', \n    orientation='h', \n    title='type_of column distribution', \n    height=500, \n    width=700\n)\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"corr = train_df.corr()\nmask = np.zeros_like(corr)\nmask[np.triu_indices_from(mask)] = True\nwith sns.axes_style(\"white\"):\n    f, ax = plt.subplots(figsize=(10, 10))\n    ax = sns.heatmap(corr,mask=mask,square=True,linewidths=.8,cmap=\"viridis\",annot=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.corr().style.background_gradient(cmap='Oranges')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"この中で高めな相関があったのは、[task_container_id - timestamp] と [content_type_id - answered_correctly]\n\nとはいえ、有効そうな相関はないって感じ。\n\nwe can see 2 correlations which have some high values:\n\n* task_container_id is correlated with the timestamp.\n\ntask_container_idは、ユーザーごとに単調に増えているから、timestampと相関が高めにでる。\n\n(task_container_id: Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id. Monotonically（単調に） increasing for each user.)\n\nThis might help us explain why it has a good correlation with the timestamp\n\n* content_type_id is correlated with answered_correctly\n(content_type_id: 0 if the event was a question being posed to the user, 1 if the event was the user watching a lecture.)\n\n0: 質問 1: 講義\n\n講義をみたユーザーの方が、回答率が高いのは想像つくから、これはいい相関である\n\nIf the user watched the lecture then chances of answering correctly increases so there is a good correlation I assume."},{"metadata":{},"cell_type":"markdown","source":"### prior_question_had_explanationと時間(timestamp)の流れ\n\nprior_question_had_explanation: (bool) Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Typically the first several questions a user sees were part of an onboarding diagnostic test where they did not get any feedback.\n\n### prior_question_had_explanation(ユーザーの質問回答後の反応)とanswaered_correctly(正解か)の関係を見る"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(20,12))\nsns.set_style('dark')\n\nmini_df = train_df.copy()\nmini_df = mini_df.sort_values(by=['timestamp'])\nmini_df = mini_df.drop_duplicates('timestamp')\n\n# Start\nmin_df = mini_df.head(1000)\nplt.subplot(3, 1, 1);\nsns.pointplot(x=min_df['timestamp'],y=min_df['prior_question_had_explanation'],hue= min_df['answered_correctly'],\n              linestyle='--',color='yellow',markers='x');\nplt.title('Start_time');\nplt.xticks([]);\nplt.yticks([0,1]);\n\n# Mid\nmid_df = mini_df[50000:51100]\nplt.subplot(3, 1, 2);\nsns.pointplot(x=mid_df['timestamp'],y=mid_df['prior_question_had_explanation'],hue= mid_df['answered_correctly'],\n              linestyle='--',color='orange',markers='x');\nplt.title('Middle_time');\nplt.xticks([]);\nplt.yticks([0,1]);\n\n# End\nmax_df = mini_df.tail(1000)\nplt.subplot(3, 1, 3);\nsns.pointplot(x=max_df['timestamp'],y=max_df['prior_question_had_explanation'],hue= max_df['answered_correctly'], \n              linestyle='--',color='red',markers='x');\nplt.title('End_time');\nplt.xticks([]);\nplt.yticks([0,1]);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"prior_question_had_explanation\n\n0: 回答したあと、ユーザーは無視している。 不真面目？　・・・　でも、回答に(簡単すぎて)正解したら、回答後に解説を確認しない。復習で、その質問に何回か答えていたら、解説を飛ばすよね。\n\n1: 回答したあと、ユーザーは解説を見ている。　真面目？\n\nprior_question_had_explanation\n\nユーザーが、(前の質問バンドルに\"答えた\"後or質問間の\"講義\"を無視した後)説明をみたか≒正しい反応をしたかどうか。※nullは、ユーザーにとって最初の質問or講義である。基本は、最初講義らしい。(通常、ユーザーに表示される最初のいくつかの質問は、フィードバックが得られなかった診断テストの一部)\n\n(bool) Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle, ignoring any lectures in between. The value is shared across a single question bundle, and is null for a user's first question bundle or lecture. Typically the first several questions a user sees were part of an onboarding diagnostic test where they did not get any feedback.\n\n* 【3つのプロットから推測できること】\n\nスタート: prior_explanation ない 0　・・・　最初の質問だから0だと思われる。・・・それにしても、結構回答を間違えている0ユーザーが多い。\n\nミドル: prior_explanation ない0　一部ある1　・・・　ほとんど正解1している。(解説を振り返らなくてもいいほど)簡単な問題なのか？　一部、きちんと解説を見てる(真面目!!)\n\nファイナル: prior_explanation ほとんど全部ある1　・・・　(まぁ、ほとんど正解1しているが、)みんな回答後に、(正解1しても)解説の説明を見ている！ (なんでだ？)\n\n→　timestamp時間が経つほど、ユーザーは回答後の解説を見ている(ほとんど正解1しているけど)・・・問題が難しいのかな？\n\n※　注意として、全ての質問に、「回答後の解説」がある訳ではない。(と思われる)\n\nThere are many things that can be inferred from the above 3 plots\n\n・First thing you can see that in the early stages there are no prior_explanation.\n\n・In the final stages you can see almost all had prior_explanation.\n\n・Notice that in starting time there are a lot of question that are answered incorrectly (marked by black x)\n\n・In the middle time session the questions that did not have prior explanation were answered wrong (look bottom of chart-2)\n\n・Final stages had nearly all answers correct"},{"metadata":{"trusted":true},"cell_type":"code","source":"# sns.set_style('white')\n# plt.figure(figsize=(10,6));\n# sns.set_style('whitegrid');\n# sns.scatterplot(x ='timestamp', y='prior_question_elapsed_time', data = train_df, hue='prior_question_had_explanation',alpha=0.8\n#                 ,linewidth=0,palette='viridis');\n# plt.legend(loc=\"best\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"ほとんどの質問が、解説つきである。　右下のプロットの塊は、おそらく大勢の生徒が同時に、「試験」をしたことを意味するだろう。\n\nFrom the above plot we can see most of the questions had an explanation. Also we can see near right bottom some points in groups. This maybe because a large number of students took their test at the same time.\n\nブルーのライン(x軸が0)は、timestampが0だから＝最初のやつはprior explanationsが0である決まりなので。\n\nAnother thing that we can notice is a faint blue line along the y-axis where x is 0. This is where the timestamp is 0 and there were no prior explanations.\n\nhttps://www.kaggle.com/nitindatta/eda-with-r3-id"},{"metadata":{"trusted":true},"cell_type":"code","source":"# plt.figure(figsize=(10,6));\n# sns.set_style('darkgrid');\n# sns.scatterplot(x = train_df['task_container_id'], y= train_df['prior_question_elapsed_time'], hue=train_df['user_id'],palette='plasma',linewidth=0, size=train_df['user_id'] ,alpha=1);\n# plt.legend(loc='best');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"user_answers_df.sort_values(('answered_correctly', 'count'), ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.scatter(x = user_answers_df[('answered_correctly', 'count')], y = user_answers_df[('answered_correctly', 'mean')]);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"how_good = train_df[train_df['answered_correctly'] != -1].groupby('user_id').mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize = (15,6))\n\nax = sns.distplot(how_good['answered_correctly'], color='darkcyan',bins=50)\n\nax.set_xlabel(\"Plot of the ratio of correct to incorrect answers by user\",fontsize=18)\nax.set_xlim(0,1)\n\nvalues = np.array([rec.get_height() for rec in ax.patches])\n\nnorm = plt.Normalize(values.min(), values.max())\n\ncolors = plt.cm.jet(norm(values))\n\nfor rec, col in zip(ax.patches, colors):\n    rec.set_color(col)\n\nplt.show();","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"The best score is: %.1f\" % (how_good['answered_correctly'].max()*100), \"%\")\nprint(\"The mean score is:  %.1f\" % (how_good['answered_correctly'].mean()*100), \"%\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 学生の数"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"No of students = \", len(train_df['user_id'].unique()))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 学生ごとにサンプルの数の分布を見る"},{"metadata":{"trusted":true},"cell_type":"code","source":"# distribution of number of samples per student\nsns.set()\nfig = plt.figure(figsize=(15,6))\nfig = sns.kdeplot(train_df.groupby(by='user_id').count()['row_id'], shade=True, gridsize=50, color='g', legend=False)\nfig.figure.suptitle(\"User_id distribution\", fontsize = 20)\nplt.xlabel('User_id counts', fontsize=16)\nplt.ylabel('Probability', fontsize=16);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"ほとんどの学生が、2000未満のデータを持っている。　train内で登場する回数を、ユーザーごとにカウント"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.groupby(by='user_id').count()['row_id'].sort_values()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 学生ごとに、どのくらい質問に回答を試みたか分布を見る"},{"metadata":{"trusted":true},"cell_type":"code","source":"# How many question does each student attempt\ndf = train_df[train_df['content_type_id'] == 0] #回答したやつ\n\ndf = df.groupby(by='user_id').count()\n\nfig = plt.figure(figsize=(15,6))\nfig = sns.kdeplot(df['row_id'], shade=True, gridsize=50, color='r', legend=False)\nfig.figure.suptitle(\"User attempted questions distribution\", fontsize = 20)\nplt.xlabel('Questions counts', fontsize=16)\nplt.ylabel('Probability', fontsize=16)\nplt.legend(['Questions Attempted','Questions Correctly answered'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"学生ごとにサンプルの数の分布と似ている。　ほとんどの学生が、2000未満の質問に回答している。"},{"metadata":{"trusted":true},"cell_type":"code","source":"# distribution of correct and incorrect and no answers\ndf = train_df[train_df['content_type_id'] == 0]\n\ndf2 = df[df['answered_correctly'] == 1]\ndf3 = df[df['answered_correctly'] == 0]\n\ndf2 = df2.groupby(by='user_id').count()\ndf3 = df3.groupby(by='user_id').count()\n\nfig = plt.figure(figsize=(15,6))\nfig = sns.kdeplot(df2['row_id'], shade=True, gridsize=50, color='b', legend=False)\nfig = sns.kdeplot(df3['row_id'], shade=True, gridsize=50, color='r', legend=False)\n\nfig.figure.suptitle(\"User attempted questions distribution\", fontsize = 20)\nplt.xlabel('Questions counts', fontsize=16)\nplt.ylabel('Probability', fontsize=16)\nplt.legend(['Correctly answered','Incorrectly answered'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"correctly answered の正解したユーザーの方が、全体的に質問に回答している割合が多いと見れる。"},{"metadata":{},"cell_type":"markdown","source":"### 学生ごとに、どのくらいのユーザーが解説を見ているか"},{"metadata":{"trusted":true},"cell_type":"code","source":"# What precent of students see explanations\n\nvalues = []\n\ndf = train_df[train_df['content_type_id'] == 0]\n\nfor group, frame in df.groupby(by='user_id'):\n    \n    value = len(frame[frame['prior_question_had_explanation'] == True]) / len(frame)\n    values.append(value)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,6))\nsns.distplot(values, kde=False)\nplt.title('Distribution if students who see x percent of explanations')\nplt.xlabel('Percent explanation seen out of attempted questions')\nplt.ylabel('Counts')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"割合的には、prior explanations(事前の説明)を見たことがなく、正解した生徒がかなりいます。これ以上の読み取りは、timestampの側面を考慮する必要がある。\n\nThere is a considerable amount of students who never watched prior explanations and yet answered correctly. Any further than this we will need to use timestamps or other files."},{"metadata":{"trusted":true},"cell_type":"code","source":"# distribution of tags\n\ntotal = []\n\nfor i in questions_df['tags']:\n    for j in str(i).strip().split(' '):\n        total.append(j)\n        \nkeys = set(total)\nfinal = {}\nfor i in keys:\n    final[i] = total.count(i)\n    \nvalues = sorted(final.items(), key=lambda x: x[1], reverse=True)\nd = []\nfor i in values:\n    d.append(i[1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(10,6))\npx.line(d, title='Tags distribution')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"タグの分布は、非常に偏っている。80％を超える割合で発生するタグは40個だけです。\n\nThe distribution of tags is very skewed. Only 40 tags occur almost > 80% of time. If we want to decrease the sparcity of our data we could use only the top 100 tags and it would be more than 95% of total tags with only 50% sparcity."},{"metadata":{"trusted":true},"cell_type":"code","source":"from wordcloud import WordCloud\n# Most commmon tags\ntags = WordCloud().generate_from_frequencies(final)\npx.imshow(tags, title='Most frequent Tags')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 適当に8人の学生の傾向を追ってみる。"},{"metadata":{"trusted":true},"cell_type":"code","source":"# we will see first 8 students for trends\nno_students = 8\nscores = []\nuser_ids = []\nquestion_attempted_l = []\ncorrectly_answered_l = []\nprior_questions_explanations = []\n\nfor count, (group, frame) in enumerate(train_df.groupby(by='user_id')):\n    \n    if count == no_students:\n        break\n    \n    frame = frame.sort_values(by='timestamp')\n    \n    percentage = []\n    question_attempted = []\n    correctly_answered = []\n    explanations = []\n    attempted = 0\n    correct_answers = 0\n    explanation = 0\n    \n    df = frame[frame['content_type_id'] == 0]\n    df = df.fillna(0)\n    \n    for answered_correctly, had_explanation in zip(df['answered_correctly'], df['prior_question_had_explanation']):\n        \n        attempted += 1\n        question_attempted.append(attempted)\n        \n        if answered_correctly == 1:\n            correct_answers += 1\n            \n        if had_explanation:\n            explanation += 1\n            \n        correctly_answered.append(correct_answers)\n            \n        percent = correct_answers / attempted * 100\n        percentage.append(percent)\n        explanations.append(explanation)\n        \n    \n    scores.append(percentage)\n    user_ids.append(group)\n    question_attempted_l.append(question_attempted)\n    correctly_answered_l.append(correctly_answered)\n    prior_questions_explanations.append(explanations)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Trend in attempted question and correctly answering\n\nplt.figure(figsize=(15,20))\n\nfor i in range(1,9):\n    plt.subplot(4,2,i)\n    plt.plot(question_attempted_l[i-1], question_attempted_l[i-1], label='Questions attempted')\n    plt.plot(question_attempted_l[i-1], correctly_answered_l[i-1], label='Questions correctly answered')\n    plt.plot(question_attempted_l[i-1], scores[i-1], label='Percentage correctly answered')\n    plt.plot(question_attempted_l[i-1], prior_questions_explanations[i-1], label='Prior_questions_explanations')\n    plt.legend()\n    plt.ylim(0,100)\n    plt.xlim(0,50)\n    plt.tight_layout(pad = 2)\n    plt.title(f'user_id: {user_ids[i-1]}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"非常に多くの傾向とパターンがある。\n\n質問の回答数が多くなるからといって、正答率が上がるとはいえない。\n\n\nSo much to see. So much trends and patterns. Well those who had prior explanation had better results. So the trend has many types. sudden spikes(+ve, -ve), consistency, continuous increment, decrement.\nBad Students: Almost no one started watching explanations until they started performing bad."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Does students time spend on answering prior questions\n\nno_students = 8\ntime_spend_l = []\n\nfor count, (group, frame) in enumerate(train_df.groupby(by='user_id')):\n    \n    if count == no_students:\n        break\n    \n    frame = frame.sort_values(by='timestamp')\n    total_time_spends = []\n    time_spends = 0\n    \n    for time_spend in frame['prior_question_elapsed_time'][frame['content_type_id'] == 0]:\n        \n        if time_spend > 0:\n            time_spends += time_spend\n            total_time_spends.append(time_spends)\n        \n    \n    time_spend_l.append(total_time_spends)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"time_spend_l = np.array(time_spend_l)\nfor index, value in enumerate(time_spend_l):\n    time_spend_l[index] = np.array(time_spend_l[index]) / 10000","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Trend in time spend with percentage\n\nplt.figure(figsize=(15,20))\n\nfor i in range(1,9):\n    plt.subplot(4,2,i)\n    plt.plot(question_attempted_l[i-1], correctly_answered_l[i-1], label='Questions correctly answered')\n    plt.plot(question_attempted_l[i-1][1:], time_spend_l[i-1], label='time spend in 10000')\n    plt.plot(question_attempted_l[i-1], scores[i-1], label='Percentage correctly answered')\n    plt.legend()\n    plt.ylim(0,100)\n    plt.xlim(0,50)\n    plt.tight_layout(pad = 2)\n    plt.title(f'user_id: {user_ids[i-1]}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"prior_question_elapsed_time は、前の質問の回答にどのくらいの時間を要したか。\n\nThere is mostly a linear increase in prior question time elapsed."},{"metadata":{},"cell_type":"markdown","source":"### answered_correctlyの上位と下位を比較したい。"},{"metadata":{},"cell_type":"markdown","source":"#### timestamp について\n\nIt is imprtant to remember that this is the time between this user interaction and the first event from that user. So starting time could be different for each user"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.groupby(['user_id'])['timestamp'].max().sort_values(ascending=False).head(20)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Some users have really huge activity time!"},{"metadata":{},"cell_type":"markdown","source":"#### content_id\nId of the content - question or lecture"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['content_id'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.loc[train_df['content_id'] == 6116]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.loc[train_df['content_id'] == 6116, 'user_answer'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"questions_df.loc[questions_df['question_id'] == 6116]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can see that a lot of people made mistakes answering this question."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}