{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Table of contents\n1. [Notebook Overview](#notebook_overview)\n2. [Imports](#imports)\n3. [Loading Data](#loading_data)\n    1. [Loading Students Answers](#loading_students_answers)\n    2. [Evaluating Students Score](#evaluating_students_score)\n    3. [Loading Students Logs](#loading_students_logs)\n4. [Data Extraction](#data_extraction)\n    1. [Dialogue Compression](#dialogue_compression)\n    2. [Filtering Events](#filtering_events)\n    3. [Recap Observation Texts](#recap_observation_texts)\n    4. [Recap person dialogues](#recap_person_dialogues)\n    5. [Final Cleanups](#final_cleanups)\n5. [Data Exploration](#data_exploration)\n    1. [The Most Frequent Recap Text](#the_most_frequent_recap_text)\n    2. [Users Overall Recap Reading](#users_overall_recap_reading)\n    3. [Recap Influence On Quiz Results](#recap_influence_on_quiz_results)\n    4. [A Little Problem...](#a_little_problem)\n6. [Wind up](#wind_up)","metadata":{}},{"cell_type":"markdown","source":"# Notebook Overview <a name=\"notebook_overview\"></a>","metadata":{}},{"cell_type":"markdown","source":"In this Kaggle notebook, we explore the influence of recap dialogues with hints to progression in the game. We analyze the relationship between the recap texts and the quiz performance of players in the game to understand how these dialogues may influence overall performance on final quiz. The main hypothesis of this notebook is that players who use hints provided in recap texts a lot may have a lower quiz score compared to those who do not.","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://i.ibb.co/RYtmTsj/4.png\" width=50%>","metadata":{}},{"cell_type":"markdown","source":"# Imports <a name=\"imports\"></a>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:45:33.648401Z","iopub.execute_input":"2023-04-06T13:45:33.649151Z","iopub.status.idle":"2023-04-06T13:45:34.955527Z","shell.execute_reply.started":"2023-04-06T13:45:33.649108Z","shell.execute_reply":"2023-04-06T13:45:34.954373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading Data <a name=\"loading_data\"></a>","metadata":{}},{"cell_type":"markdown","source":"## Loading Students Answers <a name=\"loading_students_answers\"></a>","metadata":{}},{"cell_type":"code","source":"labels = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\n\nlabels[['session_id', 'question']] = labels['session_id'].str.split('_', expand=True)\n\n# Needed for proper sorting\nlabels['question'] = labels['question'].str.slice(1)\nlabels['question'] = labels['question'].astype(int)\n\nlabels.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-06T13:45:34.961295Z","iopub.execute_input":"2023-04-06T13:45:34.963837Z","iopub.status.idle":"2023-04-06T13:45:37.274831Z","shell.execute_reply.started":"2023-04-06T13:45:34.963775Z","shell.execute_reply":"2023-04-06T13:45:37.273828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluating Students Score <a name=\"evaluating_students_score\"></a>","metadata":{}},{"cell_type":"code","source":"pivoted_questions = labels.pivot(columns='question', values='correct', index='session_id')\nstudents_score = pivoted_questions.iloc[:, 0:18].sum(axis=1)\nstudents_score = students_score.rename('student_score')\nstudents_score.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-06T13:45:37.279092Z","iopub.execute_input":"2023-04-06T13:45:37.281260Z","iopub.status.idle":"2023-04-06T13:45:37.597805Z","shell.execute_reply.started":"2023-04-06T13:45:37.281215Z","shell.execute_reply":"2023-04-06T13:45:37.596308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Loading Students Logs <a name=\"loading_students_logs\"></a>","metadata":{}},{"cell_type":"code","source":"use_columns = ['session_id', \n               'elapsed_time',\n               'text',\n               'text_fqid',\n               'event_name']\ndtypes={'session_id':'category', \n        'text':'category',\n        'text_fqid':'category',\n        'elapsed_time':np.int32,\n        'event_name':'category'}\ndf = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', \n                 dtype=dtypes, usecols=use_columns)\ndf.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-06T13:45:37.600773Z","iopub.execute_input":"2023-04-06T13:45:37.601250Z","iopub.status.idle":"2023-04-06T13:47:21.223454Z","shell.execute_reply.started":"2023-04-06T13:45:37.601188Z","shell.execute_reply":"2023-04-06T13:47:21.222449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Extraction <a name=\"data_extraction\"></a>","metadata":{}},{"cell_type":"markdown","source":"## Dialogue compression <a name=\"dialogue_compression\"></a>","metadata":{}},{"cell_type":"markdown","source":"In my previous notebook (and eventually deleted notebook), I made a mistake by working with raw text, which can vary from one version to another (more about it in this [notebook](https://www.kaggle.com/code/steubk/meetings-are-boring-the-notebook)). While it may be tempting to try to convert all dialogue to a single version (for example dry version), this approach has its limitations. There may be dialogue that does not have an equivalent and is therefore not represented in the above notebook.\n\nA better approach is to use the text_fqid column. This allows us to uniquely identify where the dialogue occurs, when it occurs, which characters are involved, and the common meaning across all versions. This approach is more robust.","metadata":{}},{"cell_type":"code","source":"df.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:21.224662Z","iopub.execute_input":"2023-04-06T13:47:21.225317Z","iopub.status.idle":"2023-04-06T13:47:21.238304Z","shell.execute_reply.started":"2023-04-06T13:47:21.225279Z","shell.execute_reply":"2023-04-06T13:47:21.237365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The first 10 entries, except for the first one, are part of the initial conversation with gramps. However, the data is currently presented in multiple rows that belong to a single dialogue. To simplify the data and make it easier to analyze, we can group together all the values that have the same text_fqid and appear consecutively in the column. This will allow us to create a detailed user dialogue history.","metadata":{}},{"cell_type":"code","source":"df['in_the_same_dialogue'] = df['text_fqid'].shift()\ndf['in_the_same_dialogue'] = df['text_fqid'] == df['in_the_same_dialogue']\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:21.239476Z","iopub.execute_input":"2023-04-06T13:47:21.240406Z","iopub.status.idle":"2023-04-06T13:47:21.392161Z","shell.execute_reply.started":"2023-04-06T13:47:21.240370Z","shell.execute_reply":"2023-04-06T13:47:21.391017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are no session collisions in the data. Each session begins with the text_fqid value 'tunic.historicalsociety.closet.intro', and does not end with this value. Therefore, the first entry of each session will have a value of 'False' for the column in_the_same_dialogue.","metadata":{}},{"cell_type":"code","source":"dialogue_sequence = df[~df['in_the_same_dialogue']]\ndialogue_sequence = dialogue_sequence[~dialogue_sequence['text_fqid'].isna()]\ndialogue_sequence.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:21.393516Z","iopub.execute_input":"2023-04-06T13:47:21.393958Z","iopub.status.idle":"2023-04-06T13:47:22.294996Z","shell.execute_reply.started":"2023-04-06T13:47:21.393903Z","shell.execute_reply":"2023-04-06T13:47:22.293818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We first perform compression on the dataset. After compression, we then remove any null values and apply the necessary filtering. This order is chosen because navigation clicks always precede different observations and person clicks, as explained in this notebook - [event meanings](https://www.kaggle.com/code/shashwatraman/meaning-of-each-event-name-and-eda)","metadata":{}},{"cell_type":"markdown","source":"## Filtering events <a name=\"filtering_events\"></a>","metadata":{}},{"cell_type":"markdown","source":"To investigate when the players is lost in the game, we look at the dialogue that guide him. As examples there are two pictures below.\n\nThis dialogue comes from person_click event (Jo asks gramps again)\n\n<img src=\"https://i.ibb.co/wK1WLHy/1.png\">\n\nAnd this is the example of observation_click event (Jo tries to leave the room with active task)\n\n<img src=\"https://i.ibb.co/TT6Dhtq/2.png\">\n\nLet's filter them.","metadata":{}},{"cell_type":"code","source":"dialogue_sequence = dialogue_sequence[(dialogue_sequence['event_name'] == 'observation_click') | \\\n                                      (dialogue_sequence['event_name'] == 'person_click')]\ndialogue_sequence = dialogue_sequence.drop(columns=['text', 'in_the_same_dialogue', 'elapsed_time'], errors='ignore')\ndialogue_sequence.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:22.296329Z","iopub.execute_input":"2023-04-06T13:47:22.296650Z","iopub.status.idle":"2023-04-06T13:47:22.359147Z","shell.execute_reply.started":"2023-04-06T13:47:22.296619Z","shell.execute_reply":"2023-04-06T13:47:22.357850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Recap observation texts <a name=\"recap_observation_texts\"></a>","metadata":{}},{"cell_type":"markdown","source":"Upon manual inspection of all text_fqid values associated with observation clicks, I noticed that all the relevant ids typically include the term 'block' in their definitions. This is quite intuitive.","metadata":{}},{"cell_type":"code","source":"observations = df[df['event_name'] == 'observation_click']['text_fqid'].unique()\nrecap_observations = []\nfor observation in observations:\n    if('block' in observation):\n        recap_observations.append(observation)\nrecap_observations","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:22.360609Z","iopub.execute_input":"2023-04-06T13:47:22.361021Z","iopub.status.idle":"2023-04-06T13:47:22.414040Z","shell.execute_reply.started":"2023-04-06T13:47:22.360964Z","shell.execute_reply":"2023-04-06T13:47:22.412985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Recap person dialogues <a name=\"recap_person_dialogues\"></a>","metadata":{}},{"cell_type":"markdown","source":"Upon manual examination of all the text_fqid values associated with person clicks, I noticed that the most relevant ids typically contain the terms 'recap' or 'lost' in their definitions. This makes sence too.","metadata":{}},{"cell_type":"code","source":"dialogues = df[df['event_name'] == 'person_click']['text_fqid'].unique()\nrecap_dialogues = []\nfor dialogue in dialogues:\n    if('recap' in dialogue or 'lost' in dialogue):\n        recap_dialogues.append(dialogue)\nrecap_dialogues","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:22.417743Z","iopub.execute_input":"2023-04-06T13:47:22.418193Z","iopub.status.idle":"2023-04-06T13:47:22.725695Z","shell.execute_reply.started":"2023-04-06T13:47:22.418157Z","shell.execute_reply":"2023-04-06T13:47:22.724565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's leave only important text-fqid values.","metadata":{}},{"cell_type":"code","source":"dialogue_sequence = dialogue_sequence[dialogue_sequence['text_fqid'].isin(recap_observations) | \\\n                    dialogue_sequence['text_fqid'].isin(recap_dialogues)]\ndialogue_sequence.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:22.727065Z","iopub.execute_input":"2023-04-06T13:47:22.727500Z","iopub.status.idle":"2023-04-06T13:47:22.758852Z","shell.execute_reply.started":"2023-04-06T13:47:22.727464Z","shell.execute_reply":"2023-04-06T13:47:22.757867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Final Cleanups <a name=\"final_cleanups\"></a>","metadata":{}},{"cell_type":"markdown","source":"The last step is manually removing recap dialogue that doesn't provide any useful information on game progression.","metadata":{}},{"cell_type":"code","source":"all_recaps = list(dialogue_sequence['text_fqid'].unique())\nfor recap in all_recaps:\n    texts = list(df[df['text_fqid'] == recap]['text'].unique())\n    print(recap)\n    print(texts)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-04-06T13:47:22.760460Z","iopub.execute_input":"2023-04-06T13:47:22.761150Z","iopub.status.idle":"2023-04-06T13:47:23.287997Z","shell.execute_reply.started":"2023-04-06T13:47:22.761103Z","shell.execute_reply":"2023-04-06T13:47:23.286742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dialogue_sequence = dialogue_sequence[\n                    (dialogue_sequence['text_fqid'] != 'tunic.wildlife.center.expert.recap') & \n                    (dialogue_sequence['text_fqid'] != 'tunic.historicalsociety.entry.wells.flag_recap') &\n                    (dialogue_sequence['text_fqid'] != 'tunic.historicalsociety.frontdesk.block_magnify')]","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:23.289364Z","iopub.execute_input":"2023-04-06T13:47:23.290235Z","iopub.status.idle":"2023-04-06T13:47:23.299852Z","shell.execute_reply.started":"2023-04-06T13:47:23.290194Z","shell.execute_reply":"2023-04-06T13:47:23.298535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following text_fqid values associated with recap dialogues were manually removed as they did not provide any useful information on game progression:\n\n- 'tunic.wildlife.center.expert.recap': This dialogue was simply an expert thanking the player for their help and did not offer any hints.\n- 'tunic.historicalsociety.entry.wells.flag_recap': Wells hurries the player to visit the wild center, but did not mention it.\n- 'tunic.historicalsociety.frontdesk.block_magnify': This dialogue appears when the player tries to pick up the magnifying glass before talking to the archivist.","metadata":{}},{"cell_type":"markdown","source":"# Data Exploration <a name=\"data_exploration\"></a>","metadata":{}},{"cell_type":"markdown","source":"# The Most Frequent Recap Text <a name=\"the_most_frequent_recap_text\"></a>","metadata":{}},{"cell_type":"code","source":"players_number = df['session_id'].nunique()\nplayers_number","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:23.301211Z","iopub.execute_input":"2023-04-06T13:47:23.301547Z","iopub.status.idle":"2023-04-06T13:47:23.444191Z","shell.execute_reply.started":"2023-04-06T13:47:23.301514Z","shell.execute_reply":"2023-04-06T13:47:23.443085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"recap_frequency = dialogue_sequence.groupby(['event_name', 'text_fqid'], observed=True).size()\nrecap_frequency = recap_frequency/players_number\nrecap_frequency = recap_frequency.reset_index()\nrecap_frequency = recap_frequency.rename(columns={0:'mean_freq'})\nrecap_frequency.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:23.445561Z","iopub.execute_input":"2023-04-06T13:47:23.446006Z","iopub.status.idle":"2023-04-06T13:47:23.470567Z","shell.execute_reply.started":"2023-04-06T13:47:23.445970Z","shell.execute_reply":"2023-04-06T13:47:23.469422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_frequent = recap_frequency.sort_values(by='mean_freq', ascending=False)\ntop_frequent = top_frequent.reset_index(drop=True).groupby('event_name').head(5)\ntop_frequent","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:23.471882Z","iopub.execute_input":"2023-04-06T13:47:23.472301Z","iopub.status.idle":"2023-04-06T13:47:23.488258Z","shell.execute_reply.started":"2023-04-06T13:47:23.472267Z","shell.execute_reply":"2023-04-06T13:47:23.487047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are the most frequent recap fqids in the game. As we can see from the index column, the leaders come from characters' dialogues, i.e. person_clicks. Let's look at their texts.","metadata":{}},{"cell_type":"code","source":"top_recaps = list(top_frequent['text_fqid'])\nfor recap in top_recaps:\n    texts = list(df[df['text_fqid'] == recap]['text'].unique())\n    print(recap)\n    print(texts)","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:23.489516Z","iopub.execute_input":"2023-04-06T13:47:23.489836Z","iopub.status.idle":"2023-04-06T13:47:23.682521Z","shell.execute_reply.started":"2023-04-06T13:47:23.489803Z","shell.execute_reply":"2023-04-06T13:47:23.681256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These texts have hints where to go next, and can serve as useful indicators that the player may be lost and in need of guidance.","metadata":{}},{"cell_type":"markdown","source":"## Users Overall Recap Reading <a name=\"users_overall_recap_reading\"></a>","metadata":{}},{"cell_type":"code","source":"session_event_recap = dialogue_sequence.groupby(['session_id', 'event_name']) \\\n                                    .size() \\\n                                    .reset_index() \\\n                                    .rename(columns={0:'recap_reading'})\n\nsession_event_recap = session_event_recap[(session_event_recap['event_name'] == 'observation_click') | \\\n                (session_event_recap['event_name'] == 'person_click')]\nsession_event_recap.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:23.683869Z","iopub.execute_input":"2023-04-06T13:47:23.684282Z","iopub.status.idle":"2023-04-06T13:47:24.659591Z","shell.execute_reply.started":"2023-04-06T13:47:23.684246Z","shell.execute_reply":"2023-04-06T13:47:24.658353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"session_recap = session_event_recap.groupby('session_id')['recap_reading'].sum()\nsession_recap","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:24.660995Z","iopub.execute_input":"2023-04-06T13:47:24.661406Z","iopub.status.idle":"2023-04-06T13:47:24.701712Z","shell.execute_reply.started":"2023-04-06T13:47:24.661372Z","shell.execute_reply":"2023-04-06T13:47:24.700390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.histplot(x=session_recap, binwidth=1)","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:24.704239Z","iopub.execute_input":"2023-04-06T13:47:24.704614Z","iopub.status.idle":"2023-04-06T13:47:25.150844Z","shell.execute_reply.started":"2023-04-06T13:47:24.704578Z","shell.execute_reply":"2023-04-06T13:47:25.150004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It appears that players typically read around 7 recap dialogues to progress in the game. Given the prevalence of text associated with person_click events, it is reasonable to assume that these events have a significant influence on this distribution.","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(session_event_recap, col='event_name', col_order=['observation_click', 'person_click'])\n\ng.map(sns.histplot, 'recap_reading', binwidth=2)\n\nsession_event_recap.groupby('event_name', observed=True)['recap_reading'].mean()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:25.152123Z","iopub.execute_input":"2023-04-06T13:47:25.152629Z","iopub.status.idle":"2023-04-06T13:47:25.674302Z","shell.execute_reply.started":"2023-04-06T13:47:25.152596Z","shell.execute_reply":"2023-04-06T13:47:25.673115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To enhance the clarity and coherence of the text, we can rephrase it as follows:\n\nThe prevalence of non-zero values in person_click dialogues greatly affects the overall sum distribution. There are two possible reasons for this:\n\n- Players may require hints from characters about where to go next in the game, despite the availability of a notebook for this purpose.\n- Players may simply be curious and try to explore all available dialogues in the game. For instance, during my first playthrough of the game, I tried to tap every dialogue to see what the character would say.","metadata":{}},{"cell_type":"markdown","source":"## Recap Influence On Quiz Results <a name=\"recap_influence_on_quiz_results\"></a>","metadata":{}},{"cell_type":"code","source":"recap_score = pd.concat([session_recap, students_score], axis=1)\nrecap_score.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:25.675970Z","iopub.execute_input":"2023-04-06T13:47:25.676312Z","iopub.status.idle":"2023-04-06T13:47:25.859646Z","shell.execute_reply.started":"2023-04-06T13:47:25.676278Z","shell.execute_reply":"2023-04-06T13:47:25.858819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(data=recap_score.groupby('student_score')['recap_reading'].mean())","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:25.861258Z","iopub.execute_input":"2023-04-06T13:47:25.861583Z","iopub.status.idle":"2023-04-06T13:47:26.063814Z","shell.execute_reply.started":"2023-04-06T13:47:25.861542Z","shell.execute_reply":"2023-04-06T13:47:26.062994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although the relationship between the number of recap dialogues read and quiz results seems clear, we need to be mindful of the sample size for low scores. Before drawing any conclusions, let's examine the distribution of scores.","metadata":{}},{"cell_type":"code","source":"students_score.value_counts().sort_values().head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:26.065116Z","iopub.execute_input":"2023-04-06T13:47:26.065699Z","iopub.status.idle":"2023-04-06T13:47:26.073683Z","shell.execute_reply.started":"2023-04-06T13:47:26.065651Z","shell.execute_reply":"2023-04-06T13:47:26.072786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distribution of scores reveals that there are a few players who scored poorly, which could significantly affect the accuracy of the mean recap_reading. As a result, the first few dots on the chart may appear to be oddly placed. To take in account for this problem, we need to add 95% confidence intervals for the means to the chart, which will allow us to better visualize the accuracy of the means, particularly for low scores where the sample size may be small.","metadata":{}},{"cell_type":"code","source":"sns.pointplot(data=recap_score, x='student_score', y='recap_reading')","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:26.075090Z","iopub.execute_input":"2023-04-06T13:47:26.075876Z","iopub.status.idle":"2023-04-06T13:47:27.004841Z","shell.execute_reply.started":"2023-04-06T13:47:26.075833Z","shell.execute_reply":"2023-04-06T13:47:27.003689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The chart with confidence intervals provides a better visualization of the accuracy of the means for different scores. The large intervals for the lower scores are due to the small sample size, as I said. However, as we move to higher scores, the intervals become narrower, which indicates a more precise estimation of the mean recap_reading.\n\nIn addition, we can create separate charts for person_click and observation_click events.","metadata":{}},{"cell_type":"code","source":"event_recap_score = session_event_recap.merge(students_score.to_frame(), on='session_id')\nevent_recap_score.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:27.006504Z","iopub.execute_input":"2023-04-06T13:47:27.007783Z","iopub.status.idle":"2023-04-06T13:47:27.054371Z","shell.execute_reply.started":"2023-04-06T13:47:27.007735Z","shell.execute_reply":"2023-04-06T13:47:27.053077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.FacetGrid(event_recap_score,\n                  col='event_name', \n                  col_order=['observation_click', 'person_click'],\n                  sharey=False)\n\ng.map(sns.pointplot, 'student_score', 'recap_reading')","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:47:27.056516Z","iopub.execute_input":"2023-04-06T13:47:27.057015Z","iopub.status.idle":"2023-04-06T13:47:28.917444Z","shell.execute_reply.started":"2023-04-06T13:47:27.056956Z","shell.execute_reply":"2023-04-06T13:47:28.916378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The relationship between score and recap_reading holds true for both the person_click and observation_click events. However, the chart for the observation_click events has a smaller slope compared to the chart for person_click events. This can be attributed to the fact that initially recap observation clicks are rarer than recap dialogues.","metadata":{}},{"cell_type":"markdown","source":"## A Little Problem... <a name=\"a_little_problem\"></a>","metadata":{}},{"cell_type":"markdown","source":"What I found interesting in our conclusion, but didn't investigated yet, is the following:","metadata":{}},{"cell_type":"code","source":"texts = df[(~df['text'].isna()) & (df['text'] != 'undefined')]\nreading = texts.groupby('session_id').size()\nreading_score = pd.concat([reading, students_score], axis=1).rename(columns={0:'reading_count'})\nreading_score.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-06T13:54:33.056947Z","iopub.execute_input":"2023-04-06T13:54:33.057377Z","iopub.status.idle":"2023-04-06T13:54:33.885904Z","shell.execute_reply.started":"2023-04-06T13:54:33.057319Z","shell.execute_reply":"2023-04-06T13:54:33.884944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.scatterplot(data=reading_score.groupby('student_score')['reading_count'].mean())","metadata":{"execution":{"iopub.status.busy":"2023-04-06T13:54:37.772436Z","iopub.execute_input":"2023-04-06T13:54:37.772863Z","iopub.status.idle":"2023-04-06T13:54:37.970337Z","shell.execute_reply.started":"2023-04-06T13:54:37.772828Z","shell.execute_reply":"2023-04-06T13:54:37.969295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's possible that our conclusion about the effectiveness of recap dialogues is influenced by this general tendency. Alternatively, it could be that people with low scores read more recap dialogues thus read more in general.\n\nThe mean for recap dialogue reading in every score group is relatively small, so it doesn't support the second scenario. Nonetheless, we can't exclude the possibility of the first scenario.\n\nHow to deal with this problem? Is it even a problem? Share your opinion in the comments, I'll try to investigate this problem on my own by reviewing the topics of correlation and causation! 🕵️‍","metadata":{}},{"cell_type":"markdown","source":"# Wind up 👨‍💻 <a name=\"wind_up\"></a> ","metadata":{}},{"cell_type":"markdown","source":"So, what do we have in the end?\n\n- The number of recap dialogues viewed by players affects their quiz score. The smaller the number of recap dialogues viewed, the higher the score. This suggests that players are not simply viewing a large number of recap dialogues out of curiosity, but rather because they need the hints provided in the dialogues.\n- A little problem with correlation and causation, that I'm confused with?\n- The new way of working with dialogues, by creating text_fquids history\n- New, especially for me, graph - pointplot that already implements confidence intervals, that are crucial when sample sizes are low.\n- Two great features that you can use for you model that you can try implement yourself to see improvements\n\nThank you for reading my EDA. If you have any questions, suggestions or ideas, please leave them in the comments, they will be very useful.","metadata":{}}]}