{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Predict Train Means Baseline\nIn Kaggle's Predict Student Performance Competition, we are asked to predict whether a student gets each of 18 questions correct or incorrect. In this notebook we compute the mean of correctness for each of the 18 questions from train data. Then we use this mean to predict 1s and 0s for each of the test users per question.\n\nWe fork and edit the basic submission template provided by Kaggle [here][1]. The idea of finding an optimal threshold (to convert prob to pred) was first done by Mayukh Bhattacharyya [here][2]\n\n[1]: https://www.kaggle.com/code/philculliton/basic-submission-demo\n[2]: https://www.kaggle.com/code/mayukh18/question-mean-optimum-baseline","metadata":{}},{"cell_type":"code","source":"import pandas as pd, numpy as np","metadata":{"papermill":{"duration":0.023295,"end_time":"2022-06-03T21:13:10.412151","exception":false,"start_time":"2022-06-03T21:13:10.388856","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-08T02:47:12.629354Z","iopub.execute_input":"2023-02-08T02:47:12.630031Z","iopub.status.idle":"2023-02-08T02:47:12.657226Z","shell.execute_reply.started":"2023-02-08T02:47:12.629938Z","shell.execute_reply":"2023-02-08T02:47:12.655573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute Train Means","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/student-performance-and-game-play/train_labels.csv')\nprint( train.shape )\ntrain['q'] = train['session_id'].apply(lambda x: int(x.split('_')[-1][1:]) )\ndisplay( train.sample(5) )","metadata":{"execution":{"iopub.status.busy":"2023-02-08T02:47:12.659437Z","iopub.execute_input":"2023-02-08T02:47:12.659781Z","iopub.status.idle":"2023-02-08T02:47:13.041278Z","shell.execute_reply.started":"2023-02-08T02:47:12.659752Z","shell.execute_reply":"2023-02-08T02:47:13.040366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"question_means = train.groupby('q').correct.agg('mean').to_dict()\nquestion_means","metadata":{"execution":{"iopub.status.busy":"2023-02-08T02:47:13.044368Z","iopub.execute_input":"2023-02-08T02:47:13.045241Z","iopub.status.idle":"2023-02-08T02:47:13.060365Z","shell.execute_reply.started":"2023-02-08T02:47:13.045207Z","shell.execute_reply":"2023-02-08T02:47:13.059586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute Validation Score\nThe competition metric is `f1_score` macro. We need to convert probability predictions into predictions of 1s and 0s. To do this, we find a threshold. Then when `prob > threhsold` we will predict 1 else 0. Let's search all thresholds and find which threshold is best.\n\n(Note: this technique of finding the optimal threshold was first done by Mayukh Bhattacharyya [here][1]. My earlier notebooks used weird (and suboptimal) ways to convert probabilities into predictions. Using a threshold is the correct way to convert probabilties into predictions)\n\n[1]: https://www.kaggle.com/code/mayukh18/question-mean-optimum-baseline","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import f1_score\n\ntrain['m'] = train.q.map(question_means)\ntrain.sample(5)","metadata":{"execution":{"iopub.status.busy":"2023-02-08T02:47:42.467472Z","iopub.execute_input":"2023-02-08T02:47:42.467852Z","iopub.status.idle":"2023-02-08T02:47:42.489869Z","shell.execute_reply.started":"2023-02-08T02:47:42.467822Z","shell.execute_reply":"2023-02-08T02:47:42.488790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# FIND BEST THRESHOLD TO CONVERT PROBS INTO 1s AND 0s\nscores = []; thresholds = []\nbest_score = 0; best_threshold = 0\n\nfor threshold in np.arange(0.4,0.81,0.01):\n    print(f'{threshold:.02f}, ',end='')\n    train['p'] = (train.m > threshold).astype('int')\n    m = f1_score(train.correct.values, train.p.values, average='macro')   \n    scores.append(m)\n    thresholds.append(threshold)\n    if m>best_score:\n        best_score = m\n        best_threshold = threshold","metadata":{"execution":{"iopub.status.busy":"2023-02-08T02:49:25.690785Z","iopub.execute_input":"2023-02-08T02:49:25.691210Z","iopub.status.idle":"2023-02-08T02:49:27.669793Z","shell.execute_reply.started":"2023-02-08T02:49:25.691181Z","shell.execute_reply":"2023-02-08T02:49:27.667814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# PLOT THRESHOLD VS. F1_SCORE\nplt.figure(figsize=(20,5))\nplt.plot(thresholds,scores,'-o',color='blue')\nplt.scatter([best_threshold], [best_score], color='blue', s=300, alpha=1)\nplt.xlabel('Threshold',size=14)\nplt.ylabel('Validation F1 Score',size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score:.3f} at Best Threshold = {best_threshold:.3}',size=18)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-08T02:49:35.979745Z","iopub.execute_input":"2023-02-08T02:49:35.980155Z","iopub.status.idle":"2023-02-08T02:49:36.206518Z","shell.execute_reply.started":"2023-02-08T02:49:35.980124Z","shell.execute_reply":"2023-02-08T02:49:36.205749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predict Train Means","metadata":{}},{"cell_type":"code","source":"import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2023-02-07T17:53:31.735753Z","iopub.execute_input":"2023-02-07T17:53:31.736327Z","iopub.status.idle":"2023-02-07T17:53:31.745596Z","shell.execute_reply.started":"2023-02-07T17:53:31.736268Z","shell.execute_reply":"2023-02-07T17:53:31.744472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"counter = 0\n# The API will deliver two dataframes in this specific order,\n# for every session+level grouping (one group per session for each checkpoint)\nfor (sample_submission, test) in iter_test:\n    if counter==0:\n        display(sample_submission.head())\n        display(test.head())\n        print(test.shape)\n        \n    ## users make predictions here using the test data\n    for index,row in sample_submission.iterrows():\n        q = int( row['session_id'].split('_')[-1][1:] )\n        p = int( question_means[q]>best_threshold )\n        sample_submission.loc[index,'correct'] = p\n    \n    ## env.predict appends the session+level sample_submission to the overall\n    ## submission\n    env.predict(sample_submission)\n    counter += 1","metadata":{"papermill":{"duration":0.337707,"end_time":"2022-06-03T21:13:10.798069","exception":false,"start_time":"2022-06-03T21:13:10.460362","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-07T17:53:31.747896Z","iopub.execute_input":"2023-02-07T17:53:31.748282Z","iopub.status.idle":"2023-02-07T17:53:31.842050Z","shell.execute_reply.started":"2023-02-07T17:53:31.748248Z","shell.execute_reply":"2023-02-07T17:53:31.840147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## the end result is a submission file containing all test session predictions\n#! head submission.csv\ndf = pd.read_csv('submission.csv')\nprint('Sample submission shape:', df.shape )\nprint('Sample submission average prediction:', df.correct.mean() )\ndf.head()","metadata":{"papermill":{"duration":0.767504,"end_time":"2022-06-03T21:13:11.572788","exception":false,"start_time":"2022-06-03T21:13:10.805284","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-07T17:53:31.843381Z","iopub.execute_input":"2023-02-07T17:53:31.843673Z","iopub.status.idle":"2023-02-07T17:53:31.864300Z","shell.execute_reply.started":"2023-02-07T17:53:31.843648Z","shell.execute_reply":"2023-02-07T17:53:31.862871Z"},"trusted":true},"execution_count":null,"outputs":[]}]}