{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"> Student Performance from Game Play - Train Labels - EDA</h1>\n<h2><center> ","metadata":{}},{"cell_type":"markdown","source":"In this competition we need to predict how players perform when asked 18 questions. Based on [PJMATHEMATICIAN](https://www.kaggle.com/pjmathematician) [analysis](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796), the questions are most probably the quizzes that arise during the game itself.  In this Notebook, we will investigate these questions. All the results here are based on the train labels data.\n\nI would like to thank  [PJMATHEMATICIAN](https://www.kaggle.com/pjmathematician) for his patient study of the game which I use here.","metadata":{}},{"cell_type":"markdown","source":"# 1. How hard are the questions?","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom colorama import init as colorama_init\nfrom colorama import Fore\nfrom colorama import Style\nfrom sklearn.metrics import f1_score,accuracy_score\npd.set_option(\"display.max_columns\", None)\npd.set_option(\"display.max_rows\", 400)\n\n#load train_targets\ntargets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]))\ntargets['q'] = targets.session_id.apply(lambda x: int(x.split('_')[-1][1:]))\n\n#Groupby questions\nstudent_nb=targets['session'].nunique()\npass_rate=pd.DataFrame(targets.groupby('q').agg('correct').sum()/student_nb).reset_index()\npass_rate.columns=['Question','Pass_Rate']\npass_rate=pass_rate.sort_values('Pass_Rate',ascending=False)\n\n#Seaborn barplot\nplt.rcParams['figure.figsize']=(15,5)\nsns.barplot(x=\"Question\", y=\"Pass_Rate\", data=pass_rate,order=pass_rate.Question)\nplt.title('Question difficulty', fontsize=16)\nplt.xlabel('Questions',fontsize=14)\nplt.ylabel('Pass Rate',fontsize=14)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T02:32:40.512057Z","iopub.execute_input":"2023-03-25T02:32:40.512981Z","iopub.status.idle":"2023-03-25T02:32:41.778303Z","shell.execute_reply.started":"2023-03-25T02:32:40.512937Z","shell.execute_reply":"2023-03-25T02:32:41.777365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some questions are very easy and some are hard.","metadata":{}},{"cell_type":"code","source":"pass_rate=pass_rate.sort_values('Pass_Rate').reset_index(drop=True)\neasiest_q = pass_rate.loc[17,'Question']\nbest_pass_rate=pass_rate.loc[17,'Pass_Rate']\nhardest_q = pass_rate.loc[0,'Question']\nworst_pass_rate=pass_rate.loc[0,'Pass_Rate']\nmean_pass_rate=pass_rate['Pass_Rate'].mean()\nprint (f'Question {Fore.GREEN}{Style.BRIGHT}{easiest_q}{Style.RESET_ALL} is the easiest. {Fore.GREEN}{Style.BRIGHT}{best_pass_rate*100:.1f}%{Style.RESET_ALL} of the players got it right!')\nprint (f'Question {Fore.GREEN}{Style.BRIGHT}{hardest_q}{Style.RESET_ALL} is the hardest. Only {Fore.GREEN}{Style.BRIGHT}{worst_pass_rate*100:.1f}%{Style.RESET_ALL} of the players got it right!')\nprint (f'In average,{Fore.GREEN}{Style.BRIGHT}{mean_pass_rate*100:.1f}%{Style.RESET_ALL} of the answers are correct.')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T02:32:41.780011Z","iopub.execute_input":"2023-03-25T02:32:41.780410Z","iopub.status.idle":"2023-03-25T02:32:41.790256Z","shell.execute_reply.started":"2023-03-25T02:32:41.780379Z","shell.execute_reply":"2023-03-25T02:32:41.789433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on [PJMATHEMATICIAN](https://www.kaggle.com/pjmathematician) [analysis](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384796) ","metadata":{}},{"cell_type":"markdown","source":"The **easiest question** is :\n* Q_2 : \"How do I know the shirt isn't a basketball jersey?\"\n* A_2 : Slip is from 1916 but basketball team started in 1976 [Notebook second page, second item]","metadata":{}},{"cell_type":"markdown","source":"The **hardest question** is:\n\n* Q_13 : \"I think Wells is going to have Teddy stuffed and mounted!\"\n* A_13 : Taxidermy [Notebook seventh page, first item]","metadata":{}},{"cell_type":"markdown","source":"# 1. How hard are the groups?","metadata":{}},{"cell_type":"markdown","source":"The questions are submitted in groups. There are 3 groups (note that the names of the groups are based on the game levels, not the question numbers).\n* Group 0_4: Question 1 to 3 (3 questions)\n* Group 5_12: Question 4 to 13 (10 questions)\n* Group 13_22: Question 14 to 18 (5 questions)","metadata":{}},{"cell_type":"code","source":"#Add column grp to targets \ndef get_group(q):\n    if q<=3: \n        grp = '0-4'\n    elif q<=13: \n        grp = '5-12'\n    elif q<=18: \n        grp = '13-22'\n    return grp\ndef group_nb_q(q):\n    if q<=3: \n        return 3\n    elif q<=13: \n        return 10\n    elif q<=18: \n        return 5\n    return grp\n\ntargets['grp']=targets['q'].apply(lambda x: get_group(x))\ntargets['grp_nb']=targets['q'].apply(lambda x: group_nb_q(x))\ntargets['correct_grp']=targets['correct']/targets['grp_nb']\n\n#groupby grp\npass_rate_grp=pd.DataFrame(targets[['grp','correct_grp']].groupby('grp').agg('correct_grp').sum()/student_nb).reset_index()\npass_rate_grp.columns=['Group','Pass_Rate']\n#pass_rate_grp.sort_values(by=['Group'])\n#pass_rate_grp['new_idx']=range(1,4)\n#pass_rate_grp.loc[pass_rate_grp['Group']=='13-22','new_idx'] = 4\n#pass_rate_grp=pass_rate_grp.sort_values(\"new_idx\").drop('new_idx', axis=1)\n#pass_rate_grp=pass_rate_grp.sort_values('Group').reset_index()\n\n\n#seaborn barplot\nsns.barplot(x=\"Group\", y=\"Pass_Rate\", data=pass_rate_grp)\nplt.title('Group difficulty', fontsize=16)\nplt.xlabel('Groups',fontsize=14)\nplt.ylabel('Pass Rate',fontsize=14)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T02:32:41.791672Z","iopub.execute_input":"2023-03-25T02:32:41.792310Z","iopub.status.idle":"2023-03-25T02:32:42.381848Z","shell.execute_reply.started":"2023-03-25T02:32:41.792278Z","shell.execute_reply":"2023-03-25T02:32:42.381084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The groups have different level of difficulty: \n* Group 0-4   : Pass Rate 88%\n* Group 13-22 : Pass Rate 71%\n* Group 5-12  : Pass Rate 65%\n\n","metadata":{}},{"cell_type":"markdown","source":"# 3. Player Performance Distribution","metadata":{}},{"cell_type":"markdown","source":"Here we look at the Performance over the 18 questions for each session and look at the distribution.","metadata":{}},{"cell_type":"code","source":"#Groupby Student (session)\nstudent_performance=pd.DataFrame(targets.groupby('session').agg({\n    'correct': 'sum'})).reset_index()\nstud_perf_distrib = pd.DataFrame(student_performance.groupby('correct').agg({\n    'session': 'count'})).reset_index()\nstud_perf_distrib['session']=stud_perf_distrib['session']/student_nb\nstud_perf_distrib.columns=['Score','Density']\n\n#Seaborn barplot\nsns.barplot(x=\"Score\", y=\"Density\", data=stud_perf_distrib)\nplt.title('Players Performance Distribution', fontsize=16)\nplt.xlabel('Score',fontsize=14)\nplt.ylabel('Density',fontsize=14)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T02:32:42.383716Z","iopub.execute_input":"2023-03-25T02:32:42.384315Z","iopub.status.idle":"2023-03-25T02:32:42.666862Z","shell.execute_reply.started":"2023-03-25T02:32:42.384281Z","shell.execute_reply":"2023-03-25T02:32:42.666089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"density_1=stud_perf_distrib.loc[0,'Density']\ndensity_18=stud_perf_distrib.loc[17,'Density']\ndensity_14=stud_perf_distrib.loc[13,'Density']\n\nprint (f'{Fore.GREEN}{Style.BRIGHT}{density_1*100:.2f}%{Style.RESET_ALL} of the players got only{Fore.GREEN}{Style.BRIGHT} 1 question{Style.RESET_ALL} right!')\nprint (f'{Fore.GREEN}{Style.BRIGHT}{density_18*100:.2f}%{Style.RESET_ALL} of the players got {Fore.GREEN}{Style.BRIGHT}all questions{Style.RESET_ALL} right! Congratulations!')\nprint (f'{Fore.GREEN}{Style.BRIGHT}{density_14*100:.2f}%{Style.RESET_ALL} of the players got {Fore.GREEN}{Style.BRIGHT}14 questions{Style.RESET_ALL} right. This is the mode.')\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T02:32:42.668100Z","iopub.execute_input":"2023-03-25T02:32:42.668710Z","iopub.status.idle":"2023-03-25T02:32:42.674915Z","shell.execute_reply.started":"2023-03-25T02:32:42.668677Z","shell.execute_reply":"2023-03-25T02:32:42.674120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Questions Pairs Accuracy Score","metadata":{}},{"cell_type":"markdown","source":"We would like to compare the players performance for a pair of questions. In a previous version, I was using the Pearson correlation.\nHowever, when dealing with binary vectors, the Pearson correlation is not a very good tool as there are only four possible outcomes ((0,0),(0,1),(1,0),(1,1)).\nA better way to compare questions is to use the Accuracy Score (proportion of pairs for which value are equals) between 2 questions.","metadata":{}},{"cell_type":"markdown","source":"We first create a Dataframe with sessions as rows and questions as columns. Each cell indicates if during the session (row) the question (column) was answered correctly.","metadata":{}},{"cell_type":"code","source":"session_question_df=targets[['session','q','correct']]\n\nsession_question = session_question_df[session_question_df['q']==1]\nsession_question=session_question[['session','correct']]\nsession_question.columns=['session',f'q{1}']\nfor q in range(2,19):\n    \n    session_question_q = session_question_df[session_question_df['q']==q]\n    session_question_q=session_question_q[['session','correct']]\n    session_question_q.columns=['session',f'q{q}']\n    session_question=pd.merge(session_question,session_question_q,how='left',on='session')\n\nsession_question","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T02:32:42.676414Z","iopub.execute_input":"2023-03-25T02:32:42.677028Z","iopub.status.idle":"2023-03-25T02:32:42.880588Z","shell.execute_reply.started":"2023-03-25T02:32:42.676982Z","shell.execute_reply":"2023-03-25T02:32:42.879775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now calculate the Accuracy Score for each pair of questions.","metadata":{}},{"cell_type":"code","source":"session_question=session_question.drop('session',axis=1)\ncorr=session_question.corr()\nfor i in range (1,19):\n    accs=[]\n    for j in range (1,19):\n        accs.append(accuracy_score(session_question[f'q{i}'],session_question[f'q{j}']))\n    corr[f'q{i}']=accs\n\nnp.fill_diagonal(corr.values, np.nan)\nf = plt.figure(figsize=(19, 15))\nplt.matshow(corr, fignum=f.number)\nplt.xticks(range(session_question.shape[1]), session_question.columns, fontsize=14, rotation=45)\nplt.yticks(range(session_question.shape[1]), session_question.columns, fontsize=14)\ncb = plt.colorbar()\ncb.ax.tick_params(labelsize=14)\nplt.title('Question Pairs accuracy Score', fontsize=16)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-25T02:32:42.881972Z","iopub.execute_input":"2023-03-25T02:32:42.882598Z","iopub.status.idle":"2023-03-25T02:32:44.047740Z","shell.execute_reply.started":"2023-03-25T02:32:42.882564Z","shell.execute_reply":"2023-03-25T02:32:44.046833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Question 13** stands apart and is weakly connected to the other questions. It is also the most difficult question.\n\n\nQ_13 : \"I think Wells is going to have Teddy stuffed and mounted!\"\n\nA_13 : Taxidermy [Notebook seventh page, first item]","metadata":{}},{"cell_type":"markdown","source":"\n**Questions 2, 3, 12 and 18** form a strongly connected group. This makes sense as they are the easiest.\n\nQ_2 : \"How do I know the shirt isn't a basketball jersey?\"\n\nA_2 : Slip is from 1916 but basketball team started in 1976 [Notebook second page, second item]\n\nQ_3 : \"Why does it matter that the basketball team started in 1974?\"\n\nA_3 : Slip is from 1916 [Notebook second page, first item]\n\nQ_12 : \"How can I connect the coffee to Wells?\"\n\nA_12 : Wells was spotted at Bean Town [Notebook fourth page, first item]\n\nQ_18 : \"What else did I find for the exhibit?\"\n\nA_18 : Images of the first Earth day [Notebook fourteenth page, first item]","metadata":{}}]}