{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"We fork and edit the basic submission template provided by Kaggle [here][1], which builds on work referenced [here][2]. The idea of finding an optimal threshold (to convert prob to pred) was first done by Mayukh Bhattacharyya [here][3]\n\n[1]: https://www.kaggle.com/code/cdeotte/predict-train-mean-baseline-0-648\n[2]: https://www.kaggle.com/code/philculliton/basic-submission-demo\n[3]: https://www.kaggle.com/code/mayukh18/question-mean-optimum-baseline\n\nIn the approach pursued here, the goal is to include all earlier scores for the same session ID, and also all the average success scores over all earlier sessions, and then submit the resultant table to a learner suitable for tabular data.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-14T00:26:15.803537Z","iopub.execute_input":"2023-02-14T00:26:15.80405Z","iopub.status.idle":"2023-02-14T00:26:15.837139Z","shell.execute_reply.started":"2023-02-14T00:26:15.803935Z","shell.execute_reply":"2023-02-14T00:26:15.83628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pip install -Uqq fastbook","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:15.838949Z","iopub.execute_input":"2023-02-14T00:26:15.83943Z","iopub.status.idle":"2023-02-14T00:26:31.222348Z","shell.execute_reply.started":"2023-02-14T00:26:15.839397Z","shell.execute_reply":"2023-02-14T00:26:31.221176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -lt /kaggle/input/","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:31.223911Z","iopub.execute_input":"2023-02-14T00:26:31.224585Z","iopub.status.idle":"2023-02-14T00:26:32.306064Z","shell.execute_reply.started":"2023-02-14T00:26:31.224547Z","shell.execute_reply":"2023-02-14T00:26:32.304959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\nprint( train.shape )\ntrain['q'] = train['session_id'].apply(lambda x: int(x.split('_')[-1][1:]) )\ndisplay( train.sample(25) )","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:32.308253Z","iopub.execute_input":"2023-02-14T00:26:32.308595Z","iopub.status.idle":"2023-02-14T00:26:32.771364Z","shell.execute_reply.started":"2023-02-14T00:26:32.308562Z","shell.execute_reply":"2023-02-14T00:26:32.770489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.describe()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:32.772624Z","iopub.execute_input":"2023-02-14T00:26:32.772929Z","iopub.status.idle":"2023-02-14T00:26:32.811746Z","shell.execute_reply.started":"2023-02-14T00:26:32.772902Z","shell.execute_reply":"2023-02-14T00:26:32.810534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train)","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:32.814157Z","iopub.execute_input":"2023-02-14T00:26:32.814521Z","iopub.status.idle":"2023-02-14T00:26:32.825403Z","shell.execute_reply.started":"2023-02-14T00:26:32.814489Z","shell.execute_reply":"2023-02-14T00:26:32.824313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train.dropna())","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:32.827829Z","iopub.execute_input":"2023-02-14T00:26:32.828172Z","iopub.status.idle":"2023-02-14T00:26:32.863976Z","shell.execute_reply.started":"2023-02-14T00:26:32.828142Z","shell.execute_reply":"2023-02-14T00:26:32.86312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"No missing data. Very good.","metadata":{}},{"cell_type":"code","source":"question_means = train.groupby('q').correct.agg('mean').to_dict()\nquestion_means","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:32.864969Z","iopub.execute_input":"2023-02-14T00:26:32.865927Z","iopub.status.idle":"2023-02-14T00:26:32.879858Z","shell.execute_reply.started":"2023-02-14T00:26:32.865893Z","shell.execute_reply":"2023-02-14T00:26:32.878828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute Validation Score\nThe competition metric is `f1_score` macro. We need to convert probability predictions into predictions of 1s and 0s. To do this, we find a threshold. Then when `prob > threhsold` we will predict 1 else 0. Let's search all thresholds and find which threshold is best.\n\n(Note: this technique of finding the optimal threshold was first done by Mayukh Bhattacharyya [here][1]. My earlier notebooks used weird (and suboptimal) ways to convert probabilities into predictions. Using a threshold is the correct way to convert probabilties into predictions)\n\n[1]: https://www.kaggle.com/code/mayukh18/question-mean-optimum-baseline","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import f1_score\n\ntrain['m'] = train.q.map(question_means)\ntrain.sample(5)","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:32.881094Z","iopub.execute_input":"2023-02-14T00:26:32.881509Z","iopub.status.idle":"2023-02-14T00:26:33.411048Z","shell.execute_reply.started":"2023-02-14T00:26:32.881479Z","shell.execute_reply":"2023-02-14T00:26:33.410293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import copy\nN_QUESTIONS = 18\n#N_correct = np.zeros((N_QUESTIONS,))\n#N_total = np.zeros((N_QUESTIONS,))\nsession_data = {} # keyed by session and having 18\n\nsessions = []\n\nfor index, row in train.iterrows():\n    #print(row, '\\n')\n    elems = row['session_id'].split('_')\n    q = int( elems[-1][1:] )\n    session = int( elems[0] )\n    sessions.append(session)\n    if not(session in session_data.keys()):\n        session_data[session] = np.zeros((N_QUESTIONS,)).astype(int)\n    thiscorrect = row[\"correct\"]\n    session_data[session][q-1] = int(thiscorrect)\n    #N_correct[q] += thiscorrect\n# now, add columns for the results of all PREVIOUS questions\n\nprev_questions = np.zeros((len(train), N_QUESTIONS)).astype(int)\ni = 0\ni_y = 0\nfor index, row in train.iterrows():\n    #print(row, '\\n')\n    elems = row['session_id'].split('_')\n    q = int( elems[-1][1:] )\n    session = int( elems[0] )\n\n    prev_questions[i-1, :(q-1)] = copy.copy(session_data[session][:(q-1)])\n    if q-1 > 3:\n        i_y += 1\n        if i_y < 20:\n            print(q, i-1, row['session_id'], session, prev_questions[i-1,:]) #session_data[session][:(max([q+1, N_QUESTIONS-1]))])\n    i += 1\nprint(i)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:33.414263Z","iopub.execute_input":"2023-02-14T00:26:33.414984Z","iopub.status.idle":"2023-02-14T00:26:57.90385Z","shell.execute_reply.started":"2023-02-14T00:26:33.414951Z","shell.execute_reply":"2023-02-14T00:26:57.90279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks OK -- for the displayed rows pertaining to question 5 we see what the prior track record was for the 4 earlier questions -- clearly we would expect someone who got all earlier questions correct to do better on question 5 than those with a poorer track record. (We also see that everything in column 5 and greater is -1, meaning 'not-available', so that is correct, too.)\n\nAlso note we have kept a list of the session id's without the questions in the list called 'sessions' and we will use this later to make sure that all the questions corresponding to a given session are placed in either the training set, or else, all the questions will be placed in the validation set.","metadata":{}},{"cell_type":"code","source":"prev_df = pd.DataFrame(data=prev_questions, index=train.index, columns=['q' + str(q+1) for q in range(N_QUESTIONS)])","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:57.904971Z","iopub.execute_input":"2023-02-14T00:26:57.905256Z","iopub.status.idle":"2023-02-14T00:26:57.91075Z","shell.execute_reply.started":"2023-02-14T00:26:57.90523Z","shell.execute_reply":"2023-02-14T00:26:57.91Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.concat([train, prev_df], axis=1, copy=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:57.911875Z","iopub.execute_input":"2023-02-14T00:26:57.912874Z","iopub.status.idle":"2023-02-14T00:26:57.961103Z","shell.execute_reply.started":"2023-02-14T00:26:57.912843Z","shell.execute_reply":"2023-02-14T00:26:57.960241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.iloc[47116:47122]","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:57.962527Z","iopub.execute_input":"2023-02-14T00:26:57.963518Z","iopub.status.idle":"2023-02-14T00:26:57.984969Z","shell.execute_reply.started":"2023-02-14T00:26:57.963477Z","shell.execute_reply":"2023-02-14T00:26:57.984165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So we again see that for q5, the earlier results are visible (but the rest are blank, i.e. set to -1).","metadata":{}},{"cell_type":"code","source":"train.columns","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:57.988201Z","iopub.execute_input":"2023-02-14T00:26:57.988829Z","iopub.status.idle":"2023-02-14T00:26:57.994291Z","shell.execute_reply.started":"2023-02-14T00:26:57.988792Z","shell.execute_reply":"2023-02-14T00:26:57.993403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#hide\n! [ -e /content ] && pip install -Uqq fastbook\nimport fastbook\nfastbook.setup_book()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:26:57.995344Z","iopub.execute_input":"2023-02-14T00:26:57.99584Z","iopub.status.idle":"2023-02-14T00:27:02.275219Z","shell.execute_reply.started":"2023-02-14T00:26:57.995811Z","shell.execute_reply":"2023-02-14T00:27:02.273891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fastai.tabular.all import *","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:02.27664Z","iopub.execute_input":"2023-02-14T00:27:02.276946Z","iopub.status.idle":"2023-02-14T00:27:02.288542Z","shell.execute_reply.started":"2023-02-14T00:27:02.27691Z","shell.execute_reply":"2023-02-14T00:27:02.287577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We must be sure to cluster all the data belongingn to a given session to the trade dataset, or else, the validation set. The simplest way is to sorth the datay by sessionId, and then make sure N_valid is some multiple of N_questions.\n\nHowever, in order to sort the id's we need to change \"...._q2\"  to \"....q02\" and so forth, otherwise q11 will be placed earlier than the corresponding q2 and q3.\n","metadata":{}},{"cell_type":"code","source":"uniquesessions = list(set(sessions))","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:02.290102Z","iopub.execute_input":"2023-02-14T00:27:02.290721Z","iopub.status.idle":"2023-02-14T00:27:02.31078Z","shell.execute_reply.started":"2023-02-14T00:27:02.290676Z","shell.execute_reply":"2023-02-14T00:27:02.309798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import random as rn\n\nPCT_VALID = 0.2\n\nvalidsessions = rn.sample(uniquesessions, int(np.round(PCT_VALID * len(uniquesessions))))","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:02.312305Z","iopub.execute_input":"2023-02-14T00:27:02.313095Z","iopub.status.idle":"2023-02-14T00:27:02.323659Z","shell.execute_reply.started":"2023-02-14T00:27:02.313052Z","shell.execute_reply":"2023-02-14T00:27:02.322725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid_idx = []\nfor index, row in train.iterrows():\n    #print(row, '\\n')\n    elems = row['session_id'].split('_')\n    q = int( elems[-1][1:] )\n    session = int( elems[0] )\n    if session in validsessions:\n        valid_idx.append(index)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:02.325072Z","iopub.execute_input":"2023-02-14T00:27:02.325687Z","iopub.status.idle":"2023-02-14T00:27:23.2766Z","shell.execute_reply.started":"2023-02-14T00:27:02.325648Z","shell.execute_reply":"2023-02-14T00:27:23.275653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(N_QUESTIONS):\n    train[\"s\" + str(i+1)] = -np.log(train[\"m\"]) * train[\"q\" + str(i+1)]\n","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.277947Z","iopub.execute_input":"2023-02-14T00:27:23.278348Z","iopub.status.idle":"2023-02-14T00:27:23.357097Z","shell.execute_reply.started":"2023-02-14T00:27:23.278318Z","shell.execute_reply":"2023-02-14T00:27:23.356269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.columns","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.358207Z","iopub.execute_input":"2023-02-14T00:27:23.358498Z","iopub.status.idle":"2023-02-14T00:27:23.364779Z","shell.execute_reply.started":"2023-02-14T00:27:23.358471Z","shell.execute_reply":"2023-02-14T00:27:23.364067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['correct'].head()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.365736Z","iopub.execute_input":"2023-02-14T00:27:23.366419Z","iopub.status.idle":"2023-02-14T00:27:23.377508Z","shell.execute_reply.started":"2023-02-14T00:27:23.366391Z","shell.execute_reply":"2023-02-14T00:27:23.376834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fastai.tabular.all import *","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.378374Z","iopub.execute_input":"2023-02-14T00:27:23.379157Z","iopub.status.idle":"2023-02-14T00:27:23.387119Z","shell.execute_reply.started":"2023-02-14T00:27:23.379125Z","shell.execute_reply":"2023-02-14T00:27:23.386204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_names = []\ncont_names = [  's' + str(i+1) for i in range(N_QUESTIONS-1)]  #list(train.columns[1:-1])\ny_names = \"correct\"","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.388085Z","iopub.execute_input":"2023-02-14T00:27:23.388338Z","iopub.status.idle":"2023-02-14T00:27:23.398153Z","shell.execute_reply.started":"2023-02-14T00:27:23.388314Z","shell.execute_reply":"2023-02-14T00:27:23.397371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"notused = []\nfor i in train.columns:\n    if not(i in cont_names) and i != 'correct':\n        notused.append(i)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.399193Z","iopub.execute_input":"2023-02-14T00:27:23.39963Z","iopub.status.idle":"2023-02-14T00:27:23.409041Z","shell.execute_reply.started":"2023-02-14T00:27:23.399603Z","shell.execute_reply":"2023-02-14T00:27:23.408224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x = train.drop(notused, axis=1)\n#x = x.astype(float)\nx[\"correct\"] = x[\"correct\"].astype(int)\nx.describe()\n\n","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.410222Z","iopub.execute_input":"2023-02-14T00:27:23.410508Z","iopub.status.idle":"2023-02-14T00:27:23.633343Z","shell.execute_reply.started":"2023-02-14T00:27:23.410483Z","shell.execute_reply":"2023-02-14T00:27:23.632626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.63435Z","iopub.execute_input":"2023-02-14T00:27:23.634954Z","iopub.status.idle":"2023-02-14T00:27:23.658371Z","shell.execute_reply.started":"2023-02-14T00:27:23.634922Z","shell.execute_reply":"2023-02-14T00:27:23.657582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"procs = [Categorify, FillMissing, Normalize]\ndls = TabularDataLoaders.from_df(x, procs=procs, cat_names=cat_names, cont_names=cont_names, \n                                 y_names=\"correct\", valid_idx=valid_idx, bs=64,metrics=f1_score)\nlearn = tabular_learner(dls)","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.663541Z","iopub.execute_input":"2023-02-14T00:27:23.663939Z","iopub.status.idle":"2023-02-14T00:27:23.915395Z","shell.execute_reply.started":"2023-02-14T00:27:23.663892Z","shell.execute_reply":"2023-02-14T00:27:23.914495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"newcols = list(x.columns[1:]) + [x.columns[0]]\n\nx = x[newcols]\nx.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.91655Z","iopub.execute_input":"2023-02-14T00:27:23.91687Z","iopub.status.idle":"2023-02-14T00:27:23.950761Z","shell.execute_reply.started":"2023-02-14T00:27:23.916843Z","shell.execute_reply":"2023-02-14T00:27:23.94996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.fit_one_cycle(5) #8","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:27:23.95174Z","iopub.execute_input":"2023-02-14T00:27:23.952554Z","iopub.status.idle":"2023-02-14T00:30:30.68249Z","shell.execute_reply.started":"2023-02-14T00:27:23.952521Z","shell.execute_reply":"2023-02-14T00:30:30.681605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%debug","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:30:30.686884Z","iopub.execute_input":"2023-02-14T00:30:30.688898Z","iopub.status.idle":"2023-02-14T00:30:30.695394Z","shell.execute_reply.started":"2023-02-14T00:30:30.688861Z","shell.execute_reply":"2023-02-14T00:30:30.694505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# for the rest, we shall follow https://www.kaggle.com/code/hjhrgov/predict-train-mean-baseline","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:30:30.709525Z","iopub.execute_input":"2023-02-14T00:30:30.710428Z","iopub.status.idle":"2023-02-14T00:30:30.720607Z","shell.execute_reply.started":"2023-02-14T00:30:30.710395Z","shell.execute_reply":"2023-02-14T00:30:30.719395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predict Train Means","metadata":{}},{"cell_type":"code","source":"import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:30:30.721821Z","iopub.execute_input":"2023-02-14T00:30:30.722226Z","iopub.status.idle":"2023-02-14T00:30:30.74872Z","shell.execute_reply.started":"2023-02-14T00:30:30.722196Z","shell.execute_reply":"2023-02-14T00:30:30.747603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"counter = 0\n\n\n\n# The API will deliver two dataframes in this specific order,\n# for every session+level grouping (one group per session for each checkpoint)\nfor (sample_submission, test) in iter_test:\n    if counter==0:\n        display(sample_submission.head())\n        display(test.head())\n        print(test.shape)\n        \n        \n\nsession_data = {} # keyed by session and having 18\n\nsessions = []\n\ni = 0\nfor (sample_submission, test) in iter_test:\n    if i == 0:\n        print(sample_submission, test, '\\n')\n    i += 1\n    elems = test['session_id'].split('_')\n    q = int( elems[-1][1:] )\n    session = int( elems[0] )\n    sessions.append(session)\n    if not(session in session_data.keys()):\n        session_data[session] = np.zeros((N_QUESTIONS,)).astype(int)\n    thiscorrect = row[\"correct\"]\n    session_data[session][q-1] = int(thiscorrect)\n    #N_correct[q] += thiscorrect\n# now, add columns for the results of all PREVIOUS questions\n\nprev_questions = np.zeros((len(train), N_QUESTIONS)).astype(int)\ni = 0\ni_y = 0\nfor (sample_submission, test) in iter_test:\n    #print(row, '\\n')\n    elems = row['session_id'].split('_')\n    q = int( elems[-1][1:] )\n    session = int( elems[0] )\n\n    prev_questions[i-1, :(q-1)] = copy.copy(session_data[session][:(q-1)])\n    if q-1 > 3:\n        i_y += 1\n        if i_y < 20:\n            print(q, i-1, row['session_id'], session, prev_questions[i-1,:]) #session_data[session][:(max([q+1, N_QUESTIONS-1]))])\n    i += 1\nprint(i)\n        \n        \n        \n        \nfor index, row in train.iterrows():\n    #print(row, '\\n')        \n    ## users make predictions here using the test data\n    for index,row in sample_submission.iterrows():\n        elems = row['session_id'].split('_')\n        q = int( elems[-1][1:] )\n        session = int( elems[0] )\n        p = int( question_means[q]>best_threshold )\n        sample_submission.loc[index,'correct'] = p\n\n    \n    ## env.predict appends the session+level sample_submission to the overall\n    ## submission\n    env.predict(sample_submission)\n    counter += 1","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:30:30.750311Z","iopub.execute_input":"2023-02-14T00:30:30.750676Z","iopub.status.idle":"2023-02-14T00:30:31.062689Z","shell.execute_reply.started":"2023-02-14T00:30:30.750642Z","shell.execute_reply":"2023-02-14T00:30:31.0615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## the end result is a submission file containing all test session predictions\n#! head submission.csv\ndf = pd.read_csv('submission.csv')\nprint('Sample submission shape:', df.shape )\nprint('Sample submission average prediction:', df.correct.mean() )\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T00:30:31.06359Z","iopub.status.idle":"2023-02-14T00:30:31.064592Z","shell.execute_reply.started":"2023-02-14T00:30:31.064323Z","shell.execute_reply":"2023-02-14T00:30:31.064349Z"},"trusted":true},"execution_count":null,"outputs":[]}]}