{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Table of contents\n1. [Imports](#imports)\n2. [Loading Data](#loading_data)\n     1. [Loading Students Answers](#loading_students_answers)\n     2. [Loading Students Logs](#loading_students_logs)\n3. [Time Between Events](#time_between_events)\n    1. [Feature Creation](#feature_creation)\n    2. [Session Collision Fix](#session_collision_fix)\n4. [Player Reading Time](#player_reading_time)\n    1. [Text Extraction](#text_extraction)\n    2. [Session Aggregation](#session_aggregation)\n    3. [Distribution Plots](#distribution_plots)\n5. [Influence of reading time on quiz results](#influence_of_reading_time_on_quiz_results)\n    1. [Hypothesis Statement](#hypothesis_statement)\n    2. [Groups Distribution](#groups_distribution)\n    3. [P-hacking](#p-hacking)\n    4. [Questions For Research](#questions_for_research)\n    5. [Hypothesis Testing](#hypothesis_testing)\n    6. [Tests Conclusions](#tests_conclusions)\n6. [Wind-up](#wind_up)","metadata":{}},{"cell_type":"markdown","source":"# Imports <a name=\"imports\"></a>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport scipy.stats as stats","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:09:57.478642Z","iopub.execute_input":"2023-03-24T10:09:57.479937Z","iopub.status.idle":"2023-03-24T10:09:57.485079Z","shell.execute_reply.started":"2023-03-24T10:09:57.479888Z","shell.execute_reply":"2023-03-24T10:09:57.483868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading Data <a name=\"loading_data\"></a>","metadata":{}},{"cell_type":"markdown","source":"## Loading Students Answers <a name=\"loading_students_answers\"></a>","metadata":{}},{"cell_type":"code","source":"labels = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\nlabels.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:09:57.517639Z","iopub.execute_input":"2023-03-24T10:09:57.518061Z","iopub.status.idle":"2023-03-24T10:09:57.810745Z","shell.execute_reply.started":"2023-03-24T10:09:57.518025Z","shell.execute_reply":"2023-03-24T10:09:57.809708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All session id info comes in format we\"session_id + question number\"\nFirst, we separate column by \"_\" delimeter","metadata":{}},{"cell_type":"code","source":"labels[['session_id', 'question']] = labels['session_id'].str.split('_', expand=True)\n\n# Needed for proper sorting\nlabels['question'] = labels['question'].str.slice(1)\nlabels['question'] = labels['question'].astype(int)\n\nlabels.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:09:57.812752Z","iopub.execute_input":"2023-03-24T10:09:57.813082Z","iopub.status.idle":"2023-03-24T10:09:59.127854Z","shell.execute_reply.started":"2023-03-24T10:09:57.813050Z","shell.execute_reply":"2023-03-24T10:09:59.126743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Loading Students Logs <a name=\"loading_students_logs\"></a>","metadata":{}},{"cell_type":"code","source":"use_columns = ['session_id', \n               'index', \n               'elapsed_time', \n               'text']\ndtypes={'session_id':'category', \n'elapsed_time':np.int32,\n    'event_name':'category',\n    'name':'category',\n    'level':np.uint8,\n    'page':'category',\n    'room_coor_x':np.float32,\n    'room_coor_y':np.float32,\n    'screen_coor_x':np.float32,\n    'screen_coor_y':np.float32,\n    'hover_duration':np.float32,\n     'text':'category',\n     'fqid':'category',\n     'room_fqid':'category',\n     'text_fqid':'category',\n     'fullscreen':'category',\n     'hq':'category',\n     'music':'category',\n     'level_group':'category'}\ndf = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', \n                 dtype=dtypes, usecols=use_columns)","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:09:59.129372Z","iopub.execute_input":"2023-03-24T10:09:59.129727Z","iopub.status.idle":"2023-03-24T10:11:35.881356Z","shell.execute_reply.started":"2023-03-24T10:09:59.129694Z","shell.execute_reply":"2023-03-24T10:11:35.880179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we use memory optimizations from this [notebook](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384359) and load only columns relevant to our research.","metadata":{}},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:35.883963Z","iopub.execute_input":"2023-03-24T10:11:35.884290Z","iopub.status.idle":"2023-03-24T10:11:35.895235Z","shell.execute_reply.started":"2023-03-24T10:11:35.884259Z","shell.execute_reply":"2023-03-24T10:11:35.894397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:35.896600Z","iopub.execute_input":"2023-03-24T10:11:35.897203Z","iopub.status.idle":"2023-03-24T10:11:35.936958Z","shell.execute_reply.started":"2023-03-24T10:11:35.897169Z","shell.execute_reply":"2023-03-24T10:11:35.936079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Time between events <a name=\"time_between_events\"></a>","metadata":{}},{"cell_type":"markdown","source":"## Feature Creation  <a name=\"feature_creation\"></a>","metadata":{}},{"cell_type":"markdown","source":"Before filtering out events with text, we need to calculate time elapsed from the previous event till the current. We assume this as a reading time for those, who contain text. To do so, we first shift up time columns, so now every row has access to the previous time.","metadata":{}},{"cell_type":"code","source":"df['previous_elapsed_time'] = df['elapsed_time'].shift()\ndf['time_from_previous'] = df['elapsed_time'] - df['previous_elapsed_time']\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:35.938229Z","iopub.execute_input":"2023-03-24T10:11:35.938776Z","iopub.status.idle":"2023-03-24T10:11:36.317384Z","shell.execute_reply.started":"2023-03-24T10:11:35.938732Z","shell.execute_reply":"2023-03-24T10:11:36.316257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see in the example below, there are a collisions between sessions that cause big negative numbers.","metadata":{}},{"cell_type":"code","source":"df.iloc[880:882]","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:36.318627Z","iopub.execute_input":"2023-03-24T10:11:36.318972Z","iopub.status.idle":"2023-03-24T10:11:36.332629Z","shell.execute_reply.started":"2023-03-24T10:11:36.318939Z","shell.execute_reply":"2023-03-24T10:11:36.331678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Session Collision Fix <a name=\"session_collision_fix\"></a>","metadata":{}},{"cell_type":"markdown","source":"To fix this we shift up the same way session_id value and verify is previous row from the current session or not,","metadata":{}},{"cell_type":"code","source":"df['previous_session_id'] = df['session_id'].shift()\ndf['is_from_previous'] = df['session_id'] == df['previous_session_id']\ndf['is_from_previous'] = df['is_from_previous'].map({False:np.NaN, True:1})\ndf['time_from_previous'] = df['time_from_previous'] * df['is_from_previous']\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:36.333798Z","iopub.execute_input":"2023-03-24T10:11:36.334781Z","iopub.status.idle":"2023-03-24T10:11:39.611836Z","shell.execute_reply.started":"2023-03-24T10:11:36.334742Z","shell.execute_reply":"2023-03-24T10:11:39.610802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.iloc[880:882]","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:39.613153Z","iopub.execute_input":"2023-03-24T10:11:39.613776Z","iopub.status.idle":"2023-03-24T10:11:39.628186Z","shell.execute_reply.started":"2023-03-24T10:11:39.613738Z","shell.execute_reply":"2023-03-24T10:11:39.627136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Problem is solved, let's finish the work by removing redutant columns","metadata":{}},{"cell_type":"code","source":"df = df.drop(columns=['previous_elapsed_time','previous_session_id','is_from_previous'])","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:39.632061Z","iopub.execute_input":"2023-03-24T10:11:39.632673Z","iopub.status.idle":"2023-03-24T10:11:40.505390Z","shell.execute_reply.started":"2023-03-24T10:11:39.632620Z","shell.execute_reply":"2023-03-24T10:11:40.504371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:40.506641Z","iopub.execute_input":"2023-03-24T10:11:40.507179Z","iopub.status.idle":"2023-03-24T10:11:40.519887Z","shell.execute_reply.started":"2023-03-24T10:11:40.507145Z","shell.execute_reply":"2023-03-24T10:11:40.518713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see there is still negative numbers, that can be a result of users' fast performing actions. Event that comes after another event (according to the index column) has smaller elapsed time. The problem is described [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384342#2134312)","metadata":{}},{"cell_type":"markdown","source":"# Player Reading time <a name=\"player_reading_time\"></a>","metadata":{}},{"cell_type":"markdown","source":"## Text Extraction <a name=\"text_extraction\"></a>","metadata":{}},{"cell_type":"markdown","source":"There are two values that coressponds to no text that I have found: NaN and undefined. Let's filter them out.","metadata":{}},{"cell_type":"code","source":"nan_text_filter = df['text'].isna()\nundefined_text_filter = (df['text'] == 'undefined')","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:40.521417Z","iopub.execute_input":"2023-03-24T10:11:40.521769Z","iopub.status.idle":"2023-03-24T10:11:40.550341Z","shell.execute_reply.started":"2023-03-24T10:11:40.521735Z","shell.execute_reply":"2023-03-24T10:11:40.549151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'''Text missing values\nNan values: {nan_text_filter.sum()}\nUndefined values: {undefined_text_filter.sum()}\n''')","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:40.551763Z","iopub.execute_input":"2023-03-24T10:11:40.552772Z","iopub.status.idle":"2023-03-24T10:11:40.617147Z","shell.execute_reply.started":"2023-03-24T10:11:40.552724Z","shell.execute_reply":"2023-03-24T10:11:40.615952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"texts = df[~nan_text_filter & ~undefined_text_filter]","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:40.618510Z","iopub.execute_input":"2023-03-24T10:11:40.618953Z","iopub.status.idle":"2023-03-24T10:11:41.122328Z","shell.execute_reply.started":"2023-03-24T10:11:40.618915Z","shell.execute_reply":"2023-03-24T10:11:41.120883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:41.123735Z","iopub.execute_input":"2023-03-24T10:11:41.124088Z","iopub.status.idle":"2023-03-24T10:11:41.131445Z","shell.execute_reply.started":"2023-03-24T10:11:41.124050Z","shell.execute_reply":"2023-03-24T10:11:41.130402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When all texts are gathered, we can gather median reading time for each of the students. We use median as our desriptive statistic, because of outliers caused by different phenomena like AFK jumps in elapsed time and not chronological index series ([article](https://www.kaggle.com/code/abaojiang/eda-on-game-progress/notebook?scriptVersionId=120133716) about this problem).","metadata":{}},{"cell_type":"markdown","source":"## Session Aggregation <a name=\"session_aggregation\"></a>","metadata":{}},{"cell_type":"code","source":"median_reading_time = texts.groupby('session_id')['time_from_previous'].median()\nmedian_reading_time = median_reading_time.rename('median_reading_time')\nmedian_reading_time","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:41.133130Z","iopub.execute_input":"2023-03-24T10:11:41.133892Z","iopub.status.idle":"2023-03-24T10:11:41.537461Z","shell.execute_reply.started":"2023-03-24T10:11:41.133841Z","shell.execute_reply":"2023-03-24T10:11:41.536511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Distribution Plots <a name=\"distribution_plots\"></a>","metadata":{}},{"cell_type":"code","source":"ax = sns.histplot(median_reading_time)\n\nplt.xlabel('User mean reading time')","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:41.538912Z","iopub.execute_input":"2023-03-24T10:11:41.539511Z","iopub.status.idle":"2023-03-24T10:11:41.998118Z","shell.execute_reply.started":"2023-03-24T10:11:41.539475Z","shell.execute_reply":"2023-03-24T10:11:41.997064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This looks like a normal distribution with a little skew to the left. Let's take a look at QQ-plot against normal distribution.","metadata":{}},{"cell_type":"code","source":"# Calculating z-score to compare our distibution with standart normal distibution\nm = median_reading_time.mean()\nsd = median_reading_time.std()\nscored_reading_time = (median_reading_time - m)/sd\nscored_reading_time","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:41.999576Z","iopub.execute_input":"2023-03-24T10:11:42.000250Z","iopub.status.idle":"2023-03-24T10:11:42.037791Z","shell.execute_reply.started":"2023-03-24T10:11:42.000212Z","shell.execute_reply":"2023-03-24T10:11:42.036895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"stats.probplot(scored_reading_time, dist=\"norm\", plot=plt)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:42.039178Z","iopub.execute_input":"2023-03-24T10:11:42.039775Z","iopub.status.idle":"2023-03-24T10:11:42.259329Z","shell.execute_reply.started":"2023-03-24T10:11:42.039733Z","shell.execute_reply":"2023-03-24T10:11:42.258202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see around the median everything is locally fine, but as we go away from median there are  more bigger values on the tails. I think this degree is similarity is enough for future tests.","metadata":{}},{"cell_type":"markdown","source":"# Influence of reading time on quiz <a name=\"influence_of_reading_time_on_quiz_results\"></a>","metadata":{}},{"cell_type":"markdown","source":"## Hypothesis statement <a name=\"hypothesis_statement\"></a>","metadata":{}},{"cell_type":"markdown","source":"### The longer the player reads the text, dialogues in the game, the better he will deal with the questions","metadata":{}},{"cell_type":"markdown","source":"## Groups Distribution <a name=\"groups_distribution\"></a>","metadata":{}},{"cell_type":"markdown","source":"To test this hypothesis we need to combine reading time data and questions data.","metadata":{}},{"cell_type":"code","source":"reading_questions = labels.join(other=median_reading_time, on='session_id')\nreading_questions.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:42.260828Z","iopub.execute_input":"2023-03-24T10:11:42.261123Z","iopub.status.idle":"2023-03-24T10:11:42.468028Z","shell.execute_reply.started":"2023-03-24T10:11:42.261093Z","shell.execute_reply":"2023-03-24T10:11:42.466561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correct_answers_dist = reading_questions.groupby(['question', 'correct'])['session_id'].count()\ncorrect_answers_dist = correct_answers_dist.reset_index()\ncorrect_answers_dist.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:42.469742Z","iopub.execute_input":"2023-03-24T10:11:42.470098Z","iopub.status.idle":"2023-03-24T10:11:42.530505Z","shell.execute_reply.started":"2023-03-24T10:11:42.470063Z","shell.execute_reply":"2023-03-24T10:11:42.529416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = sns.barplot(data=correct_answers_dist, x='question', y='session_id', hue='correct')\nax.set(xlabel='Question number', ylabel='Count')","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:42.531939Z","iopub.execute_input":"2023-03-24T10:11:42.532324Z","iopub.status.idle":"2023-03-24T10:11:43.051980Z","shell.execute_reply.started":"2023-03-24T10:11:42.532290Z","shell.execute_reply":"2023-03-24T10:11:43.050643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For some questions groups are unbalanced which can troublesome for our analysis, but t-test formula do aware of sample sizes, so I think it's okay.","metadata":{}},{"cell_type":"markdown","source":"![t-test formula](https://media.geeksforgeeks.org/wp-content/uploads/formujla.jpg)","metadata":{}},{"cell_type":"code","source":"g = sns.FacetGrid(reading_questions, col=\"question\", hue='correct', height=2, col_wrap=6)\ng.map(sns.histplot, \"median_reading_time\")\ng.add_legend()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:17:51.836545Z","iopub.execute_input":"2023-03-24T10:17:51.837106Z","iopub.status.idle":"2023-03-24T10:18:05.468115Z","shell.execute_reply.started":"2023-03-24T10:17:51.837060Z","shell.execute_reply":"2023-03-24T10:18:05.467151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see median reading time among different questions. As I think we can't tell really much from this graph because both groups in each graph have similar shape (they differ only in size due to easy or hard questions) that looks like normal distiribution. ","metadata":{}},{"cell_type":"markdown","source":"## P-hacking <a name=\"p-hacking\"></a>","metadata":{}},{"cell_type":"markdown","source":"To check our hypothesis we could use t-test on each of all 18 questions. But here's the catch that I came up during investigation of exisiting statistical tests, that involve p-value - it's a ***p-value hacking***! You can view more about it in this [StatQuest video](https://www.youtube.com/watch?v=HDCOUXE3HMM&t=2s). But in short, a lot of p-value tests can lead to false-positives. If I understood the problem properly, to avoid it we need to\n1. Minimize number of qustions that we research\n2. Take smaller threshold","metadata":{}},{"cell_type":"code","source":"p = 0.0001","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.327501Z","iopub.execute_input":"2023-03-24T10:11:56.327851Z","iopub.status.idle":"2023-03-24T10:11:56.332121Z","shell.execute_reply.started":"2023-03-24T10:11:56.327818Z","shell.execute_reply":"2023-03-24T10:11:56.331228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To find out which question are worth to test we filter questions where difference between mean of correct and incorrect groups is larger than 50.","metadata":{}},{"cell_type":"code","source":"min_difference = 50","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.333842Z","iopub.execute_input":"2023-03-24T10:11:56.334208Z","iopub.status.idle":"2023-03-24T10:11:56.348949Z","shell.execute_reply.started":"2023-03-24T10:11:56.334157Z","shell.execute_reply":"2023-03-24T10:11:56.347775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Questions For Research <a name=\"questions_for_research\"></a>","metadata":{}},{"cell_type":"code","source":"def correctness_difference(statistic):\n    question_difference = reading_questions.groupby(['question', 'correct'])[\"median_reading_time\"] \\\n                                .agg(statistic)\n    question_difference = question_difference.reset_index()\n    question_difference = question_difference.pivot(columns='correct', \n                                                    index='question', \n                                                    values='median_reading_time')\n    question_difference['difference'] = question_difference[1] - question_difference[0]\n    question_difference = question_difference.drop(columns=[0,1])\n    return question_difference","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.350857Z","iopub.execute_input":"2023-03-24T10:11:56.351774Z","iopub.status.idle":"2023-03-24T10:11:56.363359Z","shell.execute_reply.started":"2023-03-24T10:11:56.351714Z","shell.execute_reply":"2023-03-24T10:11:56.362192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correctness_mean_difference = correctness_difference('mean')\ncorrectness_mean_difference.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.365116Z","iopub.execute_input":"2023-03-24T10:11:56.365898Z","iopub.status.idle":"2023-03-24T10:11:56.411202Z","shell.execute_reply.started":"2023-03-24T10:11:56.365858Z","shell.execute_reply":"2023-03-24T10:11:56.410370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's take questions with diffirence between group higher than 50 (in both direction, so abs is required)","metadata":{}},{"cell_type":"code","source":"questions_for_research_filter = abs(correctness_mean_difference['difference']) > min_difference\nquestions_for_research = correctness_mean_difference[questions_for_research_filter]\nquestions_for_research = list(questions_for_research.index)\nquestions_for_research","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.416228Z","iopub.execute_input":"2023-03-24T10:11:56.416951Z","iopub.status.idle":"2023-03-24T10:11:56.427046Z","shell.execute_reply.started":"2023-03-24T10:11:56.416910Z","shell.execute_reply":"2023-03-24T10:11:56.425799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Hypothesis testing <a name=\"hypothesis_testing\"></a>","metadata":{}},{"cell_type":"markdown","source":"Requiremenets for t-testing:\n1. Our 2 distributions are independent 😃 (user answers either correct or incorrect, and % of users that play game multiple times is incredibly small)\n2. Both distributions are normally distributed 😐 (with a little skew, like an initial distibution)\n3. Have a similar amount of variance within each group or dispersion homogeneity (check is below) 😕","metadata":{}},{"cell_type":"markdown","source":"Check of third requirement:","metadata":{}},{"cell_type":"code","source":"correctness_std_difference = correctness_difference('std')\ncorrectness_std_difference.loc[questions_for_research]","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.428725Z","iopub.execute_input":"2023-03-24T10:11:56.429053Z","iopub.status.idle":"2023-03-24T10:11:56.471467Z","shell.execute_reply.started":"2023-03-24T10:11:56.429022Z","shell.execute_reply":"2023-03-24T10:11:56.470211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although, not all requirements perfectly met, we can try to use t-test and see the results.","metadata":{}},{"cell_type":"markdown","source":"To start the test we assume the ***null hypothesis***: means of studied variable in both groups are actually the same. Then we calculate t-criteria and p-value. P-value is a probability to get such and more different means like above, given that they actually are equal. If it's below threshold, we can tell that it's ***likely*** (not 100% true) that this two groups indeed differ in their reading_time","metadata":{}},{"cell_type":"code","source":"def question_t_test(question_number):\n    question = reading_questions[reading_questions['question'] == question_number]\n    \n    question_incorrect = question[question['correct'] == 0]['median_reading_time']\n    question_correct = question[question['correct'] == 1]['median_reading_time']\n    \n    p_value = stats.ttest_ind(question_incorrect, question_correct).pvalue\n    return p_value\n\nquestions_with_difference = []\n\nfor question in questions_for_research:\n    p_value = question_t_test(question)\n    if(p_value < p):\n        questions_with_difference.append(question)\n    print(f\"#{question} question \\nIt's p-value - {p_value} \\nSignificant difference - {p_value < p}\")","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.473135Z","iopub.execute_input":"2023-03-24T10:11:56.473492Z","iopub.status.idle":"2023-03-24T10:11:56.521189Z","shell.execute_reply.started":"2023-03-24T10:11:56.473459Z","shell.execute_reply":"2023-03-24T10:11:56.520327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are questions where reading of dialogues, texts do matter","metadata":{}},{"cell_type":"code","source":"questions_with_difference","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.522283Z","iopub.execute_input":"2023-03-24T10:11:56.522972Z","iopub.status.idle":"2023-03-24T10:11:56.528789Z","shell.execute_reply.started":"2023-03-24T10:11:56.522935Z","shell.execute_reply":"2023-03-24T10:11:56.527944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For these 6 we can reject null hypothesis.","metadata":{}},{"cell_type":"markdown","source":"## Tests Conclusions <a name=\"tests_conclusions\"></a>","metadata":{}},{"cell_type":"markdown","source":"Let's take a look of difference direction (according to data more reading can be bad in some cases)","metadata":{}},{"cell_type":"markdown","source":"I had an idea to add 95% confidence intervals, but they are too narrow relative to the graph size, but I left the code for creating it below.","metadata":{}},{"cell_type":"code","source":"for question in questions_with_difference:\n    for correct in [0,1]:\n        group_data = reading_questions[(reading_questions['question'] == question) & \\\n                                            (reading_questions['correct'] == correct)]\n        m = group_data['median_reading_time'].mean()\n        n = group_data.shape[0]\n        sd = group_data['median_reading_time'].std()/((n)**1/2)\n        r = stats.t.ppf(q=.95,df=n-1)\n        lower_bound = m - r * sd\n        upper_bound = m + r * sd\n        print(lower_bound,upper_bound)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-24T10:11:56.530067Z","iopub.execute_input":"2023-03-24T10:11:56.530367Z","iopub.status.idle":"2023-03-24T10:11:56.590751Z","shell.execute_reply.started":"2023-03-24T10:11:56.530337Z","shell.execute_reply":"2023-03-24T10:11:56.589526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"question_correct_means = reading_questions[reading_questions['question'].isin(questions_with_difference)]\nquestion_correct_means = question_correct_means.groupby(['question', 'correct'])[\"median_reading_time\"].mean()\nquestion_correct_means = question_correct_means.reset_index()\nquestion_correct_means.head() ","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:11:56.592077Z","iopub.execute_input":"2023-03-24T10:11:56.594634Z","iopub.status.idle":"2023-03-24T10:11:56.630837Z","shell.execute_reply.started":"2023-03-24T10:11:56.594595Z","shell.execute_reply":"2023-03-24T10:11:56.629548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.barplot(data=question_correct_means , x='question', y='median_reading_time', hue='correct')\nplt.legend(loc='lower right', title='correct')","metadata":{"execution":{"iopub.status.busy":"2023-03-24T10:17:23.272613Z","iopub.execute_input":"2023-03-24T10:17:23.273008Z","iopub.status.idle":"2023-03-24T10:17:23.520204Z","shell.execute_reply.started":"2023-03-24T10:17:23.272973Z","shell.execute_reply":"2023-03-24T10:17:23.518834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For three questions mean for correct answers is larger than in incorrect group (questions 2, 4, 14)\nOn the other hand, for three remaining questions (13, 16, 17) this is lower than in the incorrect group, although not so much for question 16 and 17.","metadata":{}},{"cell_type":"markdown","source":"# Wind-up <a name=\"wind_up\"></a>","metadata":{}},{"cell_type":"markdown","source":"So my final opinion is that reading mostly is higher among those who answered correcly, or doesn't differ, beside only one odd question 13. Maybe hints for this question are hidden in the notebook texts (for which we don't have access) and less in the game dialogues,for which we calculated reading time. Apart for all these tests and hypothesis, we get a nice feature for model, that you can try implement yourself to see improvements (or decrease in efficiency, but I hope that's not the case 😆)","metadata":{}},{"cell_type":"markdown","source":"Thank you for reading my EDA. If you have any questions, suggestions or ideas, please leave them in the comments, they will be very useful, because it's my first serious notebook. Specifically, I would like to know your opinion on statistical tests usage: was it proper and was my conclusion right?","metadata":{}},{"cell_type":"markdown","source":"#  <a name=\"closure_questions_feedback\"></a>","metadata":{}}]}