{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# XGBoost\n- Train GroupKFold models for each of the 18 questions. CV score is 0.68. We infer test using one of our KFold models. We can improve our CV and LB by engineering more features for our xgboost and/or trying different models (like other ML models and/or RNN and/or Transformer). Also we can improve our LB by using more KFold models OR training one model using all data (and the hyperparameters that we found from our KFold cross validation).\n\n- To avoid memory error read train data in chunks and feature engineering in chunks. Note that another way to avoid memory error is to use two notebooks. Train models in one notebook that has 32GB RAM (and save models), and then submit the required 8GB RAM notebook (with loaded models) as a second notebook. (Discussion [here][1]).\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218","metadata":{"papermill":{"duration":0.005932,"end_time":"2023-02-07T00:59:58.147501","exception":false,"start_time":"2023-02-07T00:59:58.141569","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### Links\n- https://github.com/pandas-dev/pandas/issues/11562\n- https://scikit-learn.org/stable/modules/neural_networks_supervised.html\n- https://docs.scipy.org/doc/scipy/reference/stats.html\n- https://stackoverflow.com/questions/34674797/xgboost-xgbclassifier-defaults-in-python\n- https://xgboost.readthedocs.io/en/stable/parameter.html\n- https://stackoverflow.com/questions/66097701/how-can-i-fix-this-warning-in-xgboost","metadata":{}},{"cell_type":"markdown","source":"- creating two notebooks. Put the train.csv feature engineering and model training in the first GPU notebook. Then after training XGB Classifier model use -\n### model.save_model(f'XGB_question_{t}.xgb')\n\n- Upload these models to a Kaggle dataset. Then put the inference code in a second CPU notebook. And load model with\n\n### model = XGBClassifier()\n### model.load_model(f'XGB_question_{t}.xgb')\n\n### RAPIDS cuDF Feature Engineer\n- In the first notebook to utilize cuDF, we can change import pandas as pd into the following - \n### import cudf as pd\n\n- convert Pandas to RAPIDS cuDF. Use the two codes below\n\n### targets = pd.read_csv('train_labels.csv')\n### targets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )\n### targets['q'] = pd.to_numeric( targets.session_id.str.split('_q').list.get(1) )\n\n- And because GroupKFold wants Pandas:\n\n### P1 = df.iloc[:,0].to_pandas()\n### P2 = df.index.to_pandas()\n### for i, (train_index, test_index) in enumerate(gkf.split(X=P1, groups=P2)):","metadata":{}},{"cell_type":"markdown","source":"### sklearn.model_selection.GroupKFold\n- K-fold iterator variant with non-overlapping groups.\n- Each group will appear exactly once in the test set across all folds \n\n\n### KFold\n- K-Folds cross-validator . Provides train/test indices to split data in train/test sets. Split dataset into k consecutive folds (without shuffling by default).\n\n### numpy.ceil\n- Return the ceiling of the input, element-wise.","metadata":{}},{"cell_type":"code","source":"import pandas as pd, numpy as np, gc\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score","metadata":{"papermill":{"duration":1.027875,"end_time":"2023-02-07T00:59:59.180261","exception":false,"start_time":"2023-02-07T00:59:58.152386","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T11:35:10.085507Z","iopub.execute_input":"2023-06-20T11:35:10.086335Z","iopub.status.idle":"2023-06-20T11:35:11.289050Z","shell.execute_reply.started":"2023-06-20T11:35:10.086235Z","shell.execute_reply":"2023-06-20T11:35:11.288079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Train Data and Labels\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202","metadata":{"papermill":{"duration":0.004542,"end_time":"2023-02-07T00:59:59.189777","exception":false,"start_time":"2023-02-07T00:59:59.185235","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# READ USER ID ONLY\ntmp = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\",usecols=[0])\ntmp = tmp.groupby('session_id').session_id.agg('count')\n\n# COMPUTE READS AND SKIPS\nPIECES = 10\nCHUNK = int( np.ceil(len(tmp)/PIECES) )\n\nreads = []\nskips = [0]\nfor k in range(PIECES):\n    a = k*CHUNK\n    b = (k+1)*CHUNK\n    if b>len(tmp): b=len(tmp)\n    r = tmp.iloc[a:b].sum()\n    reads.append(r)\n    skips.append(skips[-1]+r)\n    \nprint(f'To avoid memory error, we will read train in {PIECES} pieces of sizes:')\nprint(reads)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-06-20T12:10:00.721635Z","iopub.execute_input":"2023-06-20T12:10:00.722414Z","iopub.status.idle":"2023-06-20T12:10:53.906954Z","shell.execute_reply.started":"2023-06-20T12:10:00.722373Z","shell.execute_reply":"2023-06-20T12:10:53.906133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', nrows=reads[0])\nprint('Train size of first piece:', train.shape )\ntrain.head()","metadata":{"papermill":{"duration":59.284316,"end_time":"2023-02-07T01:00:58.478743","exception":false,"start_time":"2023-02-07T00:59:59.194427","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:11:31.331484Z","iopub.execute_input":"2023-06-20T12:11:31.332203Z","iopub.status.idle":"2023-06-20T12:11:37.679841Z","shell.execute_reply.started":"2023-06-20T12:11:31.332163Z","shell.execute_reply":"2023-06-20T12:11:37.678871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )\ntargets['q'] = targets.session_id.apply(lambda x: int(x.split('_')[-1][1:]) )\nprint( targets.shape )\ntargets.head()","metadata":{"papermill":{"duration":0.598155,"end_time":"2023-02-07T01:00:59.082015","exception":false,"start_time":"2023-02-07T01:00:58.48386","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:11:44.676815Z","iopub.execute_input":"2023-06-20T12:11:44.677434Z","iopub.status.idle":"2023-06-20T12:11:45.700451Z","shell.execute_reply.started":"2023-06-20T12:11:44.677398Z","shell.execute_reply":"2023-06-20T12:11:45.699619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineer\n- Create basic aggregate features. Try creating more features to boost CV and LB! The idea for EVENTS feature is from [here][1]\n\n[1]: https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data","metadata":{"papermill":{"duration":0.005196,"end_time":"2023-02-07T01:00:59.092865","exception":false,"start_time":"2023-02-07T01:00:59.087669","status":"completed"},"tags":[]}},{"cell_type":"code","source":"CATS = ['event_name', 'fqid', 'room_fqid', 'text']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n\n# https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data\nEVENTS = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']","metadata":{"papermill":{"duration":0.014685,"end_time":"2023-02-07T01:00:59.112856","exception":false,"start_time":"2023-02-07T01:00:59.098171","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:11:49.815908Z","iopub.execute_input":"2023-06-20T12:11:49.816303Z","iopub.status.idle":"2023-06-20T12:11:49.822025Z","shell.execute_reply.started":"2023-06-20T12:11:49.816269Z","shell.execute_reply":"2023-06-20T12:11:49.821087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def feature_engineer(train):\n    \n#     dfs = []\n#     for c in CATS:\n#         tmp = train.groupby(['session_id','level_group'])[c].agg('nunique')\n#         tmp.name = tmp.name + '_nunique'\n#         dfs.append(tmp)\n#     for c in NUMS:\n#         tmp = train.groupby(['session_id','level_group'])[c].agg('mean')\n#         tmp.name = tmp.name + '_mean'\n#         dfs.append(tmp)\n#     for c in NUMS:\n#         tmp = train.groupby(['session_id','level_group'])[c].agg('std')\n#         tmp.name = tmp.name + '_std'\n#         dfs.append(tmp)\n#     for c in EVENTS: \n#         train[c] = (train.event_name == c).astype('int8')\n#     for c in EVENTS + ['elapsed_time']:\n#         tmp = train.groupby(['session_id','level_group'])[c].agg('sum')\n#         tmp.name = tmp.name + '_sum'\n#         dfs.append(tmp)\n#     train = train.drop(EVENTS,axis=1)\n        \n#     df = pd.concat(dfs,axis=1)\n#     df = df.fillna(-1)\n#     df = df.reset_index()\n#     df = df.set_index('session_id')\n#     return df","metadata":{"papermill":{"duration":0.017716,"end_time":"2023-02-07T01:00:59.136021","exception":false,"start_time":"2023-02-07T01:00:59.118305","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T05:35:47.478427Z","iopub.execute_input":"2023-06-20T05:35:47.479444Z","iopub.status.idle":"2023-06-20T05:35:47.492030Z","shell.execute_reply.started":"2023-06-20T05:35:47.479410Z","shell.execute_reply":"2023-06-20T05:35:47.490821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import scipy\nfrom scipy import stats\ndef feature_engineer(train):\n    \n    dfs = []\n    for c in CATS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('nunique')\n        tmp.name = tmp.name + '_nunique'\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('mean')\n        tmp.name = tmp.name + '_mean'\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('median')\n        tmp.name = tmp.name + '_median'\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('std')\n        tmp.name = tmp.name + '_std'\n        dfs.append(tmp)\n    for c in EVENTS: \n        train[c] = (train.event_name == c).astype('int8')\n    for c in EVENTS + ['elapsed_time']:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('sum')\n        tmp.name = tmp.name + '_sum'\n        dfs.append(tmp)\n    train = train.drop(EVENTS,axis=1)\n        \n    df = pd.concat(dfs,axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df = df.set_index('session_id')\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-06-20T12:11:52.186054Z","iopub.execute_input":"2023-06-20T12:11:52.186794Z","iopub.status.idle":"2023-06-20T12:11:52.198347Z","shell.execute_reply.started":"2023-06-20T12:11:52.186751Z","shell.execute_reply":"2023-06-20T12:11:52.197336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# PROCESS TRAIN DATA IN PIECES\nall_pieces = []\nprint(f'Processing train as {PIECES} pieces to avoid memory error... ')\nfor k in range(PIECES):\n    print(k,', ',end='')\n    SKIPS = 0\n    if k>0: SKIPS = range(1,skips[k]+1)\n    train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv',\n                        nrows=reads[k], skiprows=SKIPS)\n    df = feature_engineer(train)\n    all_pieces.append(df)\n    \n# CONCATENATE ALL PIECES\nprint('\\n')\ndel train; gc.collect()\ndf = pd.concat(all_pieces, axis=0)\nprint('Shape of all train data after feature engineering:', df.shape )\ndf.head()","metadata":{"papermill":{"duration":34.516494,"end_time":"2023-02-07T01:01:33.658043","exception":false,"start_time":"2023-02-07T01:00:59.141549","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:12:28.206156Z","iopub.execute_input":"2023-06-20T12:12:28.206807Z","iopub.status.idle":"2023-06-20T12:18:13.770381Z","shell.execute_reply.started":"2023-06-20T12:12:28.206771Z","shell.execute_reply":"2023-06-20T12:18:13.769553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train XGBoost Model\n- We train one model for each of 18 questions. Furthermore, we use data from `level_groups = '0-4'` to train model for questions 1-3, and `level groups '5-12'` to train questions 4 thru 13 and `level groups '13-22'` to train questions 14 thru 18. Because this is the data we get (to predict corresponding questions) from Kaggle's inference API during test inference. We can improve our model by saving a user's previous data from earlier `level_groups` and using that to predict future `level_groups`.","metadata":{"papermill":{"duration":0.00565,"end_time":"2023-02-07T01:01:33.669525","exception":false,"start_time":"2023-02-07T01:01:33.663875","status":"completed"},"tags":[]}},{"cell_type":"code","source":"FEATURES = [c for c in df.columns if c != 'level_group']\nprint('We will train with', len(FEATURES) ,'features')\nALL_USERS = df.index.unique()\nprint('We will train with', len(ALL_USERS) ,'users info')","metadata":{"papermill":{"duration":0.014699,"end_time":"2023-02-07T01:01:33.689953","exception":false,"start_time":"2023-02-07T01:01:33.675254","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:21:30.601026Z","iopub.execute_input":"2023-06-20T12:21:30.601928Z","iopub.status.idle":"2023-06-20T12:21:30.612266Z","shell.execute_reply.started":"2023-06-20T12:21:30.601869Z","shell.execute_reply":"2023-06-20T12:21:30.611025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gkf = GroupKFold(n_splits=4)\noof = pd.DataFrame(data=np.zeros((len(ALL_USERS),18)), index=ALL_USERS)\nmodels = {}\n\n# COMPUTE CV SCORE WITH 5 GROUP K FOLD\nfor i, (train_index, test_index) in enumerate(gkf.split(X=df, groups=df.index)):\n    print('#'*25)\n    print('### Fold',i+1)\n    print('#'*25)\n    \n    xgb_params = {\n    'objective' : 'binary:logitraw',\n    'eval_metric':'logloss',\n    'sampling_method':'uniform',\n    'learning_rate': 0.85,\n    'max_depth': 6,\n    'n_estimators': 1000,\n    'early_stopping_rounds': 46,\n    'tree_method':'approx',\n    'subsample':0.5,\n    'colsample_bytree': 0.6,\n    'use_label_encoder' : False}\n    \n    # ITERATE THRU QUESTIONS 1 THRU 18\n    for t in range(1,19):\n        \n        # USE THIS TRAIN DATA WITH THESE QUESTIONS\n        if t<=3: grp = '0-4'\n        elif t<=13: grp = '5-12'\n        elif t<=22: grp = '13-22'\n            \n        # TRAIN DATA\n        train_x = df.iloc[train_index]\n        train_x = train_x.loc[train_x.level_group == grp]\n        train_users = train_x.index.values\n        train_y = targets.loc[targets.q==t].set_index('session').loc[train_users]\n        \n        # VALID DATA\n        valid_x = df.iloc[test_index]\n        valid_x = valid_x.loc[valid_x.level_group == grp]\n        valid_users = valid_x.index.values\n        valid_y = targets.loc[targets.q==t].set_index('session').loc[valid_users]\n        \n        # TRAIN MODEL        \n        clf =  XGBClassifier(**xgb_params)\n        clf.fit(train_x[FEATURES].astype('float32'), train_y['correct'],\n                eval_set=[ (valid_x[FEATURES].astype('float32'), valid_y['correct']) ],\n                verbose=0)\n        print(f'{t}({clf.best_ntree_limit}), ',end='')\n        \n        # SAVE MODEL, PREDICT VALID OOF\n        models[f'{grp}_{t}'] = clf\n        oof.loc[valid_users, t-1] = clf.predict_proba(valid_x[FEATURES].astype('float32'))[:,1]\n        \n    print()","metadata":{"papermill":{"duration":69.877213,"end_time":"2023-02-07T01:02:43.57299","exception":false,"start_time":"2023-02-07T01:01:33.695777","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:25:49.645186Z","iopub.execute_input":"2023-06-20T12:25:49.645885Z","iopub.status.idle":"2023-06-20T12:29:29.982989Z","shell.execute_reply.started":"2023-06-20T12:25:49.645846Z","shell.execute_reply":"2023-06-20T12:29:29.981993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute CV Score\n- We need to convert prediction probabilities into `1s` and `0s`. The competition metric is F1 Score which is the harmonic mean of precision and recall. Let's find the optimal threshold for `p > threshold` when to predict `1` and when to predict `0` to maximize F1 Score.","metadata":{"papermill":{"duration":0.011241,"end_time":"2023-02-07T01:02:43.59638","exception":false,"start_time":"2023-02-07T01:02:43.585139","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# PUT TRUE LABELS INTO DATAFRAME WITH 18 COLUMNS\ntrue = oof.copy()\nfor k in range(18):\n    # GET TRUE LABELS\n    tmp = targets.loc[targets.q == k+1].set_index('session').loc[ALL_USERS]\n    true[k] = tmp.correct.values","metadata":{"execution":{"iopub.status.busy":"2023-06-20T12:30:18.626688Z","iopub.execute_input":"2023-06-20T12:30:18.627491Z","iopub.status.idle":"2023-06-20T12:30:18.750767Z","shell.execute_reply.started":"2023-06-20T12:30:18.627452Z","shell.execute_reply":"2023-06-20T12:30:18.749923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# FIND BEST THRESHOLD TO CONVERT PROBS INTO 1s AND 0s\nscores = []; thresholds = []\nbest_score = 0; best_threshold = 0\n\nfor threshold in np.arange(0.4,0.81,0.01):\n    print(f'{threshold:.02f}, ',end='')\n    preds = (oof.values.reshape((-1))>threshold).astype('int')\n    m = f1_score(true.values.reshape((-1)), preds, average='macro')   \n    scores.append(m)\n    thresholds.append(threshold)\n    if m>best_score:\n        best_score = m\n        best_threshold = threshold","metadata":{"execution":{"iopub.status.busy":"2023-06-20T12:30:21.201205Z","iopub.execute_input":"2023-06-20T12:30:21.201958Z","iopub.status.idle":"2023-06-20T12:30:29.580941Z","shell.execute_reply.started":"2023-06-20T12:30:21.201919Z","shell.execute_reply":"2023-06-20T12:30:29.579694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# PLOT THRESHOLD VS. F1_SCORE\nplt.figure(figsize=(20,5))\nplt.plot(thresholds,scores,'-o',color='blue')\nplt.scatter([best_threshold], [best_score], color='blue', s=300, alpha=1)\nplt.xlabel('Threshold',size=14)\nplt.ylabel('Validation F1 Score',size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score:.3f} at Best Threshold = {best_threshold:.3}',size=18)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T12:30:45.017252Z","iopub.execute_input":"2023-06-20T12:30:45.017654Z","iopub.status.idle":"2023-06-20T12:30:45.227171Z","shell.execute_reply.started":"2023-06-20T12:30:45.017621Z","shell.execute_reply":"2023-06-20T12:30:45.225986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('When using optimal threshold...')\nfor k in range(18):\n        \n    # COMPUTE F1 SCORE PER QUESTION\n    m = f1_score(true[k].values, (oof[k].values>best_threshold).astype('int'), average='macro')\n    print(f'Q{k}: F1 =',m)\n    \n# COMPUTE F1 SCORE OVERALL\nm = f1_score(true.values.reshape((-1)), (oof.values.reshape((-1))>best_threshold).astype('int'), average='macro')\nprint('==> Overall F1 =',m)","metadata":{"papermill":{"duration":0.771134,"end_time":"2023-02-07T01:02:44.378465","exception":false,"start_time":"2023-02-07T01:02:43.607331","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:30:48.096590Z","iopub.execute_input":"2023-06-20T12:30:48.097010Z","iopub.status.idle":"2023-06-20T12:30:48.491800Z","shell.execute_reply.started":"2023-06-20T12:30:48.096978Z","shell.execute_reply":"2023-06-20T12:30:48.490560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Infer Test Data","metadata":{"papermill":{"duration":0.011075,"end_time":"2023-02-07T01:02:44.400918","exception":false,"start_time":"2023-02-07T01:02:44.389843","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# IMPORT KAGGLE API\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n# CLEAR MEMORY\nimport gc\ndel targets, df, oof, true\n_ = gc.collect()","metadata":{"papermill":{"duration":0.052132,"end_time":"2023-02-07T01:02:44.464739","exception":false,"start_time":"2023-02-07T01:02:44.412607","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:01:09.252100Z","iopub.execute_input":"2023-06-20T12:01:09.252884Z","iopub.status.idle":"2023-06-20T12:01:09.403844Z","shell.execute_reply.started":"2023-06-20T12:01:09.252842Z","shell.execute_reply":"2023-06-20T12:01:09.402581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"limits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    \n    # FEATURE ENGINEER TEST DATA\n    df = feature_engineer(test)\n    \n    # INFER TEST DATA\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n    \n    env.predict(sample_submission)","metadata":{"papermill":{"duration":1.002014,"end_time":"2023-02-07T01:02:45.47927","exception":false,"start_time":"2023-02-07T01:02:44.477256","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:30:52.756173Z","iopub.execute_input":"2023-06-20T12:30:52.756606Z","iopub.status.idle":"2023-06-20T12:30:52.764976Z","shell.execute_reply.started":"2023-06-20T12:30:52.756572Z","shell.execute_reply":"2023-06-20T12:30:52.763793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA submission.csv","metadata":{"papermill":{"duration":0.011427,"end_time":"2023-02-07T01:02:45.502331","exception":false,"start_time":"2023-02-07T01:02:45.490904","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df = pd.read_csv('submission.csv')\nprint( df.shape )\ndf.head()","metadata":{"papermill":{"duration":0.027432,"end_time":"2023-02-07T01:02:45.541022","exception":false,"start_time":"2023-02-07T01:02:45.51359","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:30:56.286162Z","iopub.execute_input":"2023-06-20T12:30:56.286588Z","iopub.status.idle":"2023-06-20T12:30:56.299722Z","shell.execute_reply.started":"2023-06-20T12:30:56.286556Z","shell.execute_reply":"2023-06-20T12:30:56.298973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df.correct.mean())","metadata":{"papermill":{"duration":0.020233,"end_time":"2023-02-07T01:02:45.57314","exception":false,"start_time":"2023-02-07T01:02:45.552907","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-20T12:30:59.386391Z","iopub.execute_input":"2023-06-20T12:30:59.387027Z","iopub.status.idle":"2023-06-20T12:30:59.392602Z","shell.execute_reply.started":"2023-06-20T12:30:59.386989Z","shell.execute_reply":"2023-06-20T12:30:59.391534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}