{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Table of Contents\n\n1. [Overview]()\n2. [Imports]()\n3. [Loading Data]()\n4. [How Player Interacts With Notebook?]()\n5. [Notebook Sessions Bounds]()\n6. [Extracting Data]()\n7. [Data Exploration]()\n    1. [Notebook? I Don't Need It!]()\n    2. [Hypothesis Testing]()","metadata":{}},{"cell_type":"markdown","source":"# Overview","metadata":{}},{"cell_type":"markdown","source":"Joe's notebook is a crucial element in the game, as it enables players to keep track of clues and hints that help in their gaming progress. In this article, we will create a **dataset consisting of players' notebook sessions**  and investigate an interesting **hypothesis**  about this game feature. Specifically, we hypothesize that **players who use the notebook feature may perform better then others** when it comes to answering quiz questions.\n\n<img src=\"https://i.ibb.co/vk79SMt/1.png\" width=50%>","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:19:40.322858Z","iopub.execute_input":"2023-04-22T08:19:40.323333Z","iopub.status.idle":"2023-04-22T08:19:40.329710Z","shell.execute_reply.started":"2023-04-22T08:19:40.323285Z","shell.execute_reply":"2023-04-22T08:19:40.328636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading Data","metadata":{}},{"cell_type":"code","source":"labels = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\n\nlabels[['session_id', 'question']] = labels['session_id'].str.split('_', expand=True)\n\n# Needed for proper sorting\nlabels['question'] = labels['question'].str.slice(1)\nlabels['question'] = labels['question'].astype(int)\n\nlabels.head()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-04-22T06:26:59.575712Z","iopub.execute_input":"2023-04-22T06:26:59.576534Z","iopub.status.idle":"2023-04-22T06:27:01.719353Z","shell.execute_reply.started":"2023-04-22T06:26:59.576496Z","shell.execute_reply":"2023-04-22T06:27:01.718281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pivoted_questions = labels.pivot(columns='question', values='correct', index='session_id')\nstudents_score = pivoted_questions.iloc[:, 0:18].sum(axis=1)\nstudents_score = students_score.rename('student_score')\nstudents_score.head()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-04-22T06:27:01.724096Z","iopub.execute_input":"2023-04-22T06:27:01.734665Z","iopub.status.idle":"2023-04-22T06:27:02.053384Z","shell.execute_reply.started":"2023-04-22T06:27:01.734593Z","shell.execute_reply":"2023-04-22T06:27:02.052448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"use_columns = ['session_id',\n               'elapsed_time',\n               'event_name',\n               'name',\n               'level_group',\n               'page',\n               'fqid',]\n\ndtypes={'session_id':'category',\n        'index':np.int32,\n        'elapsed_time':np.int32,\n        'event_name':'category',\n        'name':'category',\n        'level_group':'category',\n        'page':'category',\n        'fqid':'category'}\n\ndf = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', \n                 dtype=dtypes, usecols=use_columns)\ndf.head()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-04-22T06:27:02.057837Z","iopub.execute_input":"2023-04-22T06:27:02.060072Z","iopub.status.idle":"2023-04-22T06:28:50.378090Z","shell.execute_reply.started":"2023-04-22T06:27:02.060027Z","shell.execute_reply":"2023-04-22T06:28:50.377008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# How Player Interacts With Notebook?","metadata":{}},{"cell_type":"code","source":"notebook_events = df[df['event_name'] == 'notebook_click']['name'].value_counts()\nnotebook_events","metadata":{"execution":{"iopub.status.busy":"2023-04-22T06:28:50.380261Z","iopub.execute_input":"2023-04-22T06:28:50.381208Z","iopub.status.idle":"2023-04-22T06:28:50.516148Z","shell.execute_reply.started":"2023-04-22T06:28:50.381171Z","shell.execute_reply":"2023-04-22T06:28:50.515254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see\n\n1. Player **opens** notebook (open event)\n2. Within it player can **click** everywhere on the screen (basic event)\n3. If clicks land on left or right arrow, a **page is switched** (prev or next event)\n4. In the end players **closes notebook** (close event)\n\nWhat I don't like here is that open and close events are not in 1 to 1 ration, there are less close events. Later, we should pay attention to this fact.","metadata":{}},{"cell_type":"markdown","source":"To create dataset, we find start and end of every notebook session to find specific information:\n\n1. **Session length**\n2. **Number of actions** per session\n    1. Basic clicks\n    2. Next clicks\n    3. Prev clicks","metadata":{}},{"cell_type":"markdown","source":"# Notebook Session Bounds","metadata":{}},{"cell_type":"code","source":"open_close = df[df['name'].isin(['open', 'close']) & (df['event_name'] == 'notebook_click')].copy()\nopen_close.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:23:48.066519Z","iopub.execute_input":"2023-04-22T08:23:48.066997Z","iopub.status.idle":"2023-04-22T08:23:49.375431Z","shell.execute_reply.started":"2023-04-22T08:23:48.066960Z","shell.execute_reply":"2023-04-22T08:23:49.374552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Every open should be followed by close, but as we remember that's not true. Let's drop every entry that is followed by the entry with the same name value. That way we leave only appropriate bounds (open -> close)","metadata":{}},{"cell_type":"code","source":"open_close['next_name'] = open_close['name'].shift(-1)\nopen_close = open_close[open_close['name'] != open_close['next_name']]\nopen_close.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:23:54.716691Z","iopub.execute_input":"2023-04-22T08:23:54.717445Z","iopub.status.idle":"2023-04-22T08:23:54.762577Z","shell.execute_reply.started":"2023-04-22T08:23:54.717404Z","shell.execute_reply":"2023-04-22T08:23:54.761721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"session_bounds = open_close.reset_index().rename(columns={'index':'df_index'})[['name', 'df_index']]\nsession_bounds = session_bounds.pivot(columns='name', values='df_index')\nsession_bounds['close'] = session_bounds['close'].shift(-1)\nsession_bounds = session_bounds.dropna()\nsession_bounds = session_bounds.astype(int)\nsession_bounds.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:23:55.510839Z","iopub.execute_input":"2023-04-22T08:23:55.511589Z","iopub.status.idle":"2023-04-22T08:23:55.815466Z","shell.execute_reply.started":"2023-04-22T08:23:55.511550Z","shell.execute_reply":"2023-04-22T08:23:55.814600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Number of notebook sessions - {session_bounds.shape[0]}')","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:23:56.262531Z","iopub.execute_input":"2023-04-22T08:23:56.263330Z","iopub.status.idle":"2023-04-22T08:23:56.268688Z","shell.execute_reply.started":"2023-04-22T08:23:56.263277Z","shell.execute_reply":"2023-04-22T08:23:56.267777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Extraction","metadata":{}},{"cell_type":"code","source":"def extract_notebook_session(session_bounds):\n    notebook_session = df.loc[session_bounds[0]:session_bounds[1]]\n    \n    entry = notebook_session['name'].value_counts()[['basic', 'prev', 'next']]\n    \n    entry['start'] = notebook_session['elapsed_time'].min()\n    entry['end'] = notebook_session['elapsed_time'].max()\n    \n    entry['session_id'] = notebook_session.iloc[0]['session_id']\n    entry['level_group'] = notebook_session.iloc[0]['level_group']\n    return entry","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:23:58.277383Z","iopub.execute_input":"2023-04-22T08:23:58.278126Z","iopub.status.idle":"2023-04-22T08:23:58.284077Z","shell.execute_reply.started":"2023-04-22T08:23:58.278086Z","shell.execute_reply":"2023-04-22T08:23:58.283047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# notebook_sessions = session_bounds.apply(extract_notebook_session, axis=1)\n# # Additional feature engineering\n# notebook_sessions['length'] = notebook_sessions['end'] - notebook_sessions['start']\n# # Except for open/close that are present in every session\n# notebook_sessions['actions'] = notebook_sessions[['basic', 'prev', 'next']].sum(axis=1)\n\n# notebook_sessions['page_switch'] = notebook_sessions[['prev', 'next']].sum(axis=1)\n\n# notebook_sessions.to_csv('notebook_sessions.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:23:59.153050Z","iopub.execute_input":"2023-04-22T08:23:59.153468Z","iopub.status.idle":"2023-04-22T08:23:59.158656Z","shell.execute_reply.started":"2023-04-22T08:23:59.153434Z","shell.execute_reply":"2023-04-22T08:23:59.157567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"notebook_sessions = pd.read_csv('/kaggle/input/notebook-sessions/notebook_sessions.csv')\nnotebook_sessions.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:23:59.309063Z","iopub.execute_input":"2023-04-22T08:23:59.309837Z","iopub.status.idle":"2023-04-22T08:23:59.458770Z","shell.execute_reply.started":"2023-04-22T08:23:59.309794Z","shell.execute_reply":"2023-04-22T08:23:59.457875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we load dataset, we **must restore initial list of session ids**, because there are no session ids who didn't use noteboo at all.","metadata":{}},{"cell_type":"code","source":"c = pd.CategoricalDtype(df['session_id'].unique())\nnotebook_sessions['session_id'] = notebook_sessions['session_id'].astype(str).astype(c)\nnotebook_sessions['session_id']","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:26:33.975464Z","iopub.execute_input":"2023-04-22T08:26:33.975939Z","iopub.status.idle":"2023-04-22T08:26:34.325310Z","shell.execute_reply.started":"2023-04-22T08:26:33.975904Z","shell.execute_reply":"2023-04-22T08:26:34.324077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Exploration","metadata":{}},{"cell_type":"markdown","source":"## Notebook? I Don't Need It!","metadata":{}},{"cell_type":"code","source":"notebook_opens = notebook_sessions.groupby('session_id').size()\nnotebook_opens.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:26:36.415264Z","iopub.execute_input":"2023-04-22T08:26:36.415715Z","iopub.status.idle":"2023-04-22T08:26:36.430665Z","shell.execute_reply.started":"2023-04-22T08:26:36.415678Z","shell.execute_reply":"2023-04-22T08:26:36.429679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"notebook_usage = (notebook_opens > 0).rename('used_notebook')\nnotebook_usage.value_counts(normalize=True)","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:26:37.610574Z","iopub.execute_input":"2023-04-22T08:26:37.611475Z","iopub.status.idle":"2023-04-22T08:26:37.623288Z","shell.execute_reply.started":"2023-04-22T08:26:37.611429Z","shell.execute_reply":"2023-04-22T08:26:37.622063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Someone doens't used notebook at all? But how that can be real? Every player at least ones open notebook in the beggining of the game?\n\n<img src=\"https://i.ibb.co/ypqKrbx/7.png\" width=50%>\n\nThe answer is simple - the first, when you pick up notebook it's counted as an object.","metadata":{}},{"cell_type":"code","source":"df.loc[25:27]","metadata":{"execution":{"iopub.status.busy":"2023-04-22T07:05:39.487988Z","iopub.execute_input":"2023-04-22T07:05:39.488706Z","iopub.status.idle":"2023-04-22T07:05:39.503670Z","shell.execute_reply.started":"2023-04-22T07:05:39.488647Z","shell.execute_reply":"2023-04-22T07:05:39.502656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Back to the ration, **11%** is a large per cent! In my opinion, some users don't interact with **notebook**, because it is **quite static**. The notes adds automatically as you progress throught out the game when new one added, notebook shakes a little bit. I suggest 2 solutions, that I find interesting:\n\n1. **Make notebook updates more visible**, for example by adding exclamation point near it.\n2. **Add typing or drawing function** for the notebook in some way. For instance player could type clues and tasks that he find before they add in the game the way they do. So game could also improve/check spelling.","metadata":{}},{"cell_type":"markdown","source":"## Hypothesis Testing","metadata":{}},{"cell_type":"markdown","source":"Now let's check our hypothesis! We assumed that **players who use notebook better deal with questions**. But I think we also should add users that used notebook feature a few times. To take an appropriate threshold let's take a look at distribution of notebook opens per user.","metadata":{}},{"cell_type":"code","source":"ax = sns.histplot(notebook_opens, binwidth=2)\nax.set_title('Notebook Opens Throughtout The Game')\nax.set(xlabel='Notebook Usage Per User')\n\nm = notebook_opens.mean()\nplt.axvline(m, color='k', linestyle='dashed', linewidth=1)\nmin_ylim, max_ylim = plt.ylim()\nplt.text(m*1.1, max_ylim*0.7, 'Mean: {:.2f}'.format(m))\n\nsns.despine(top=True, right=True)","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:28:58.603289Z","iopub.execute_input":"2023-04-22T08:28:58.603839Z","iopub.status.idle":"2023-04-22T08:28:58.904194Z","shell.execute_reply.started":"2023-04-22T08:28:58.603798Z","shell.execute_reply":"2023-04-22T08:28:58.903194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I take half of this value as a threshold - **5 notebook opens**","metadata":{}},{"cell_type":"code","source":"threshold = 5\nnotebook_usage = (notebook_opens > threshold).rename('used_notebook')\nnotebook_usage.value_counts(normalize=True)","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:29:32.325828Z","iopub.execute_input":"2023-04-22T08:29:32.326594Z","iopub.status.idle":"2023-04-22T08:29:32.336905Z","shell.execute_reply.started":"2023-04-22T08:29:32.326550Z","shell.execute_reply":"2023-04-22T08:29:32.336009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"used_score = pd.concat([notebook_usage, students_score], axis=1)\nused_score.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:29:32.562191Z","iopub.execute_input":"2023-04-22T08:29:32.562858Z","iopub.status.idle":"2023-04-22T08:29:32.787047Z","shell.execute_reply.started":"2023-04-22T08:29:32.562819Z","shell.execute_reply":"2023-04-22T08:29:32.785774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For t-test there are a few of conditions should be satisfied:\n\n1. **Distributions independence**: user either used or not used notebook so get into one group.\n2. **Dispersion homogeneity**: there is a pretty similar amount of variance within each group.\n\nIt's not required for distributions to be normally distributed for central limit theorem, so we don't bother. But still let's take a look at both group distribution, to check the 2 requirement.","metadata":{}},{"cell_type":"code","source":"ax = sns.histplot(data=used_score, x='student_score', hue='used_notebook', binwidth=1)\n\nax.set(xlabel='Quiz Score Per User', title='Comparison of Quiz Scores by Notebook Usage')\n\nplt.legend(title='Notebook Usage', labels=['A Lot (>5)', 'A Few (<=5)'])\nsns.despine(top=True, right=True)","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:30:09.778015Z","iopub.execute_input":"2023-04-22T08:30:09.778424Z","iopub.status.idle":"2023-04-22T08:30:10.155616Z","shell.execute_reply.started":"2023-04-22T08:30:09.778390Z","shell.execute_reply":"2023-04-22T08:30:10.154675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"used_score.groupby('used_notebook')['student_score'].std()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T08:30:11.004973Z","iopub.execute_input":"2023-04-22T08:30:11.006205Z","iopub.status.idle":"2023-04-22T08:30:11.015284Z","shell.execute_reply.started":"2023-04-22T08:30:11.006160Z","shell.execute_reply":"2023-04-22T08:30:11.014481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"They have pretty similar shapes, but not the same. That's why variance differ. I think this difference is small, so we suppose 2 condition is approximately true. What is interesting in this graph is the fact that **people who didn't use notebook feature tend to have a better score**. Let's check statistical significance!","metadata":{}},{"cell_type":"code","source":"from scipy.stats import ttest_ind\n\nused_a_lot = used_score[used_score['used_notebook']]['student_score']\nused_a_few = used_score[~used_score['used_notebook']]['student_score']\nttest_ind(used_a_lot, used_a_few)","metadata":{"execution":{"iopub.status.busy":"2023-04-22T07:42:42.081194Z","iopub.execute_input":"2023-04-22T07:42:42.081624Z","iopub.status.idle":"2023-04-22T07:42:42.097676Z","shell.execute_reply.started":"2023-04-22T07:42:42.081585Z","shell.execute_reply":"2023-04-22T07:42:42.096372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"P-value is far below 0.05 mark, so we can say that means indeed differ. Let's plot fact with confidence intervals!","metadata":{}},{"cell_type":"code","source":"ax = sns.catplot(data=used_score, x=\"used_notebook\", y=\"student_score\", kind=\"point\", order=[True, False])\n\nax.set(title='Mean Quiz Scores Across Notebook Usage', ylabel='Quiz Score', xlabel=None)\nax.set_xticklabels(['Used A Lot (>5)', 'Used A Few (<=5)'])","metadata":{"execution":{"iopub.status.busy":"2023-04-22T07:52:52.830757Z","iopub.execute_input":"2023-04-22T07:52:52.831156Z","iopub.status.idle":"2023-04-22T07:52:53.356429Z","shell.execute_reply.started":"2023-04-22T07:52:52.831122Z","shell.execute_reply":"2023-04-22T07:52:53.355319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I believe that the surprising outcome may be attributed to the **additional time that using the notebook requires**, leading to a **longer completion time**. In general, there is a known tendency for **players with higher scores to finish the game quicker** because they have a better understanding what to do next.\n\nHowever, the fact that players can use the notebook to answer questions should have balanced the odds. In theory, people who need hints and thus use the notebook should become more comfortable with it over time, allowing them to answer questions on the same level as those who didn't use it. In my opinion, **notebook feature is not implemented in its full educational power**.","metadata":{}},{"cell_type":"markdown","source":"# Wind-up ","metadata":{}},{"cell_type":"markdown","source":"So, what do we have in the end?\n\n- **Constructed a comprehensive dataset** consisting of information regarding every notebook session, which could provide valuable features for our models and insights into the effective use of Joe's notebook.\n- I provided my personal thoughts on how the notebook feature could be improved for educational purposes: **notebook typing or writing**.\n- **Investigated two distinct groups of quiz results**: those who used the notebook feature frequently and those who used it sparingly.\n- I personally tried to follow guidelines from the book, that I'm currently reading - **\"Storytelling with data\"** - to effectively communicate my findings.\n\nThank you for reading my notebook. If you have any questions, suggestions or ideas on this topic - leave them in the comments! In the next chapter we will continue exploring created dataset.","metadata":{}}]}