{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# XGBoost Baseline - LB 0.678\nIn this notebook we present a XGBoost baseline. We train GroupKFold models for each of the 18 questions. Our CV score is 0.678. We infer test using one of our KFold models. We can improve our CV and LB by engineering more features for our xgboost and/or trying different models (like other ML models and/or RNN and/or Transformer). Also we can improve our LB by using more KFold models OR training one model using all data (and the hyperparameters that we found from our KFold cross validation).\n\n**UPDATE** On March 20 2023, Kaggle doubled the size of train data. Therefore we updated this notebook to avoid memory error. We accomplish this by reading train data in chunks and feature engineering in chunks. Note that another way to avoid memory error is to use two notebooks. Train models in one notebook that has 32GB RAM (and save models), and then submit the required 8GB RAM notebook (with loaded models) as a second notebook. (Discussion [here][1]).\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218","metadata":{"papermill":{"duration":0.005932,"end_time":"2023-02-07T00:59:58.147501","exception":false,"start_time":"2023-02-07T00:59:58.141569","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import pandas as pd, numpy as np, gc\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score","metadata":{"papermill":{"duration":1.027875,"end_time":"2023-02-07T00:59:59.180261","exception":false,"start_time":"2023-02-07T00:59:58.152386","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T02:58:31.328461Z","iopub.execute_input":"2023-06-05T02:58:31.329350Z","iopub.status.idle":"2023-06-05T02:58:32.721562Z","shell.execute_reply.started":"2023-06-05T02:58:31.329247Z","shell.execute_reply":"2023-06-05T02:58:32.720513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This section of the code imports the necessary libraries for the rest of the script.\n\npandas and numpy are fundamental libraries for data manipulation and computation in Python.\ngc is the garbage collector, used to free up memory.\nKFold and GroupKFold from sklearn.model_selection are used for splitting the dataset into different groups for cross-validation.\nXGBClassifier is an implementation of the gradient boosting algorithm from the XGBoost library, used for the machine learning model.\nf1_score from sklearn.metrics is used to evaluate the performance of the machine learning model, which is a balance between precision and recall.","metadata":{}},{"cell_type":"markdown","source":"# Load Train Data and Labels\nOn March 20 2023, Kaggle doubled the size of train data (discussion [here][1]). The train data is now 4.7GB! To avoid memory error, we will read the train data in as 10 pieces and feature engineer each piece before reading the next piece. This works because feature engineering shrinks the size of each piece.\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202","metadata":{"papermill":{"duration":0.004542,"end_time":"2023-02-07T00:59:59.189777","exception":false,"start_time":"2023-02-07T00:59:59.185235","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# READ USER ID ONLY\ntmp = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\",usecols=[0])\ntmp = tmp.groupby('session_id').session_id.agg('count')\n\n# COMPUTE READS AND SKIPS\nPIECES = 10\nCHUNK = int( np.ceil(len(tmp)/PIECES) )\n\nreads = []\nskips = [0]\nfor k in range(PIECES):\n    a = k*CHUNK\n    b = (k+1)*CHUNK\n    if b>len(tmp): b=len(tmp)\n    r = tmp.iloc[a:b].sum()\n    reads.append(r)\n    skips.append(skips[-1]+r)\n    \nprint(f'To avoid memory error, we will read train in {PIECES} pieces of sizes:')\nprint(reads)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-06-05T02:59:25.862623Z","iopub.execute_input":"2023-06-05T02:59:25.863057Z","iopub.status.idle":"2023-06-05T03:00:47.077068Z","shell.execute_reply.started":"2023-06-05T02:59:25.863024Z","shell.execute_reply":"2023-06-05T03:00:47.076157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This section reads in the 'session_id' column from the train.csv file and counts the number of occurrences of each 'session_id'.\n\nHere, the train.csv file is split into 10 pieces (PIECES = 10). \nThe number of rows (CHUNK) in each piece is calculated by taking the total number of rows (len(tmp)) and dividing by 10, rounding up if necessary (np.ceil). For each piece, it calculates the number of rows to read (r) and the number of rows to skip before reading. The number of rows to read for each piece is stored in the reads list, and the cumulative number of rows to skip before each piece is stored in the skips list.\n\nThis prints the sizes of the pieces that the train.csv file will be read in. This is done to avoid memory errors by not reading in the entire file at once.\n","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv', nrows=reads[0])\nprint('Train size of first piece:', train.shape )\ntrain.head()","metadata":{"papermill":{"duration":59.284316,"end_time":"2023-02-07T01:00:58.478743","exception":false,"start_time":"2023-02-07T00:59:59.194427","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:01:00.930628Z","iopub.execute_input":"2023-06-05T03:01:00.931036Z","iopub.status.idle":"2023-06-05T03:01:09.882176Z","shell.execute_reply.started":"2023-06-05T03:01:00.931003Z","shell.execute_reply":"2023-06-05T03:01:09.878270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This reads in the first reads[0] rows from the train.csv file, and then prints the shape of the resulting DataFrame. The .head() function displays the first few rows of this DataFrame.\n\nLastly, the targets are loaded:\n","metadata":{}},{"cell_type":"code","source":"targets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )\ntargets['q'] = targets.session_id.apply(lambda x: int(x.split('_')[-1][1:]) )\nprint( targets.shape )\ntargets.head()","metadata":{"papermill":{"duration":0.598155,"end_time":"2023-02-07T01:00:59.082015","exception":false,"start_time":"2023-02-07T01:00:58.48386","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:01:15.527367Z","iopub.execute_input":"2023-06-05T03:01:15.527786Z","iopub.status.idle":"2023-06-05T03:01:17.063362Z","shell.execute_reply.started":"2023-06-05T03:01:15.527751Z","shell.execute_reply":"2023-06-05T03:01:17.062410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This section reads the train_labels.csv file into a DataFrame called targets. The file train_labels.csv presumably contains the correct answers or targets for the training data.\n\nThen, it creates two new columns in the targets DataFrame: session and q. The session column is created by applying a lambda function to the session_id column that splits the session_id on the underscore character ('') and converts the first part to an integer. Similarly, the q column is created by applying a lambda function to the session_id column that splits the session_id on the underscore character (''), takes the last part, skips the first character of this part ('q'), and converts the rest to an integer.\n\nLastly, it prints the shape of the targets DataFrame and displays the first few rows using the .head() function.","metadata":{}},{"cell_type":"markdown","source":"# Feature Engineer\nWe create basic aggregate features. Try creating more features to boost CV and LB! The idea for EVENTS feature is from [here][1]\n\n[1]: https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data","metadata":{"papermill":{"duration":0.005196,"end_time":"2023-02-07T01:00:59.092865","exception":false,"start_time":"2023-02-07T01:00:59.087669","status":"completed"},"tags":[]}},{"cell_type":"code","source":"CATS = ['event_name', 'fqid', 'room_fqid', 'text']\nNUMS = ['elapsed_time','level','page','room_coor_x', 'room_coor_y', \n        'screen_coor_x', 'screen_coor_y', 'hover_duration']\n\n# https://www.kaggle.com/code/kimtaehun/lightgbm-baseline-with-aggregated-log-data\nEVENTS = ['navigate_click','person_click','cutscene_click','object_click',\n          'map_hover','notification_click','map_click','observation_click',\n          'checkpoint']","metadata":{"papermill":{"duration":0.014685,"end_time":"2023-02-07T01:00:59.112856","exception":false,"start_time":"2023-02-07T01:00:59.098171","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:01:41.773057Z","iopub.execute_input":"2023-06-05T03:01:41.773606Z","iopub.status.idle":"2023-06-05T03:01:41.781875Z","shell.execute_reply.started":"2023-06-05T03:01:41.773559Z","shell.execute_reply":"2023-06-05T03:01:41.780407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_engineer(train):\n    \n    # New feature engineering steps\n    train['elapsed_time_change'] = train.groupby('session_id')['elapsed_time'].diff()\n    train['elapsed_time_change'] = train['elapsed_time_change'].fillna(0)\n    train['level_page_ratio'] = train['level'] / train['page']\n    train['level_page_ratio'] = train['level_page_ratio'].replace([np.inf, -np.inf], np.nan)\n    train['level_page_ratio'] = train['level_page_ratio'].fillna(train['level_page_ratio'].mean())\n\n    dfs = []\n    for c in CATS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('nunique')\n        tmp.name = tmp.name + '_nunique'\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('mean')\n        tmp.name = tmp.name + '_mean'\n        dfs.append(tmp)\n    for c in NUMS:\n        tmp = train.groupby(['session_id','level_group'])[c].agg('std')\n        tmp.name = tmp.name + '_std'\n        dfs.append(tmp)\n    for c in EVENTS: \n        train[c] = (train.event_name == c).astype('int8')\n    for c in EVENTS + ['elapsed_time', 'elapsed_time_change', 'level_page_ratio']:  # Include new features\n        tmp = train.groupby(['session_id','level_group'])[c].agg('sum')\n        tmp.name = tmp.name + '_sum'\n        dfs.append(tmp)\n    train = train.drop(EVENTS,axis=1)\n        \n    df = pd.concat(dfs,axis=1)\n    df = df.fillna(-1)\n    df = df.reset_index()\n    df = df.set_index('session_id')\n    \n    return df\n\n","metadata":{"papermill":{"duration":0.017716,"end_time":"2023-02-07T01:00:59.136021","exception":false,"start_time":"2023-02-07T01:00:59.118305","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:26:17.459353Z","iopub.execute_input":"2023-06-05T03:26:17.459790Z","iopub.status.idle":"2023-06-05T03:26:17.477554Z","shell.execute_reply.started":"2023-06-05T03:26:17.459756Z","shell.execute_reply":"2023-06-05T03:26:17.476310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# PROCESS TRAIN DATA IN PIECES\nall_pieces = []\nprint(f'Processing train as {PIECES} pieces to avoid memory error... ')\nfor k in range(PIECES):\n    print(k,', ',end='')\n    SKIPS = 0\n    if k>0: SKIPS = range(1,skips[k]+1)\n    train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv',\n                        nrows=reads[k], skiprows=SKIPS)\n    df = feature_engineer(train)\n    all_pieces.append(df)\n    \n# CONCATENATE ALL PIECES\nprint('\\n')\ndel train; gc.collect()\ndf = pd.concat(all_pieces, axis=0)\nprint('Shape of all train data after feature engineering:', df.shape )\ndf.head()","metadata":{"papermill":{"duration":34.516494,"end_time":"2023-02-07T01:01:33.658043","exception":false,"start_time":"2023-02-07T01:00:59.141549","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:26:18.725496Z","iopub.execute_input":"2023-06-05T03:26:18.725935Z","iopub.status.idle":"2023-06-05T03:34:46.983338Z","shell.execute_reply.started":"2023-06-05T03:26:18.725900Z","shell.execute_reply":"2023-06-05T03:34:46.982130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code is creating more advanced features and is working with a larger set of categorical, numerical, and event-related features. Here's what it does:\n\nDefine feature categories: The block begins by specifying the names of different types of features in the data. It separates them into categorical (CATS), numerical (NUMS), and event-related (EVENTS) features.\n\nDefine feature_engineer function: This function generates new features by applying various aggregations to the existing features. These include: the number of unique values (nunique) for each categorical feature, the mean and standard deviation for each numerical feature, and the sum for each event-related feature and 'elapsed_time'. It also creates new binary features indicating whether a particular event occurred. The function then concatenates all these new features into a single DataFrame (df), fills any missing values with -1, resets the index, and sets 'session_id' as the new index.\n\nProcess the data in pieces: To avoid running out of memory, the training data is read in and processed in pieces. The code uses a loop to read each piece of the data, apply the feature_engineer function, and append the result to a list. The pieces are then concatenated into a single DataFrame (df).\n\nPrint the shape of the data and the first few rows: Finally, the code prints the shape of the resulting DataFrame and displays the first few rows to check that the feature engineering was performed correctly.\n\nThe wall time at the end shows how long the entire operation took, which can be useful for benchmarking and optimization.\n\nThe output is a DataFrame where each row corresponds to a session (as indicated by 'session_id'), and the columns represent the new features. Each feature is an aggregation over the corresponding feature in the original data, grouped by 'session_id' and 'level_group'. The '_mean', '_std', '_nunique', and '_sum' suffixes indicate the type of aggregation performed. The event-related features are binary and indicate whether a particular event occurred. The 'elapsed_time_sum' feature represents the total elapsed time for each session.","metadata":{}},{"cell_type":"markdown","source":"# Train XGBoost Model\nWe train one model for each of 18 questions. Furthermore, we use data from `level_groups = '0-4'` to train model for questions 1-3, and `level groups '5-12'` to train questions 4 thru 13 and `level groups '13-22'` to train questions 14 thru 18. Because this is the data we get (to predict corresponding questions) from Kaggle's inference API during test inference. We can improve our model by saving a user's previous data from earlier `level_groups` and using that to predict future `level_groups`.","metadata":{"papermill":{"duration":0.00565,"end_time":"2023-02-07T01:01:33.669525","exception":false,"start_time":"2023-02-07T01:01:33.663875","status":"completed"},"tags":[]}},{"cell_type":"code","source":"FEATURES = [c for c in df.columns if c != 'level_group']\nprint('We will train with', len(FEATURES) ,'features')\nALL_USERS = df.index.unique()\nprint('We will train with', len(ALL_USERS) ,'users info')","metadata":{"papermill":{"duration":0.014699,"end_time":"2023-02-07T01:01:33.689953","exception":false,"start_time":"2023-02-07T01:01:33.675254","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:35:01.425944Z","iopub.execute_input":"2023-06-05T03:35:01.427410Z","iopub.status.idle":"2023-06-05T03:35:01.436882Z","shell.execute_reply.started":"2023-06-05T03:35:01.427358Z","shell.execute_reply":"2023-06-05T03:35:01.435664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gkf = GroupKFold(n_splits=5)\noof = pd.DataFrame(data=np.zeros((len(ALL_USERS),18)), index=ALL_USERS)\nmodels = {}\n\n# COMPUTE CV SCORE WITH 5 GROUP K FOLD\nfor i, (train_index, test_index) in enumerate(gkf.split(X=df, groups=df.index)):\n    print('#'*25)\n    print('### Fold',i+1)\n    print('#'*25)\n    \n    xgb_params = {\n        'objective' : 'binary:logistic',\n        'eval_metric':'logloss',\n        'learning_rate': 0.01,  # Lower learning rate\n        'max_depth': 6,  # Increase max depth\n        'n_estimators': 2000,  # Increase number of estimators\n        'early_stopping_rounds': 100,  # Increase early stopping rounds\n        'tree_method':'hist',\n        'subsample':0.8,\n        'colsample_bytree': 0.6,  # Decrease column subsampling\n        'use_label_encoder' : False\n    }\n    \n    # ITERATE THRU QUESTIONS 1 THRU 18\n    for t in range(1,19):\n        \n        # USE THIS TRAIN DATA WITH THESE QUESTIONS\n        if t<=3: grp = '0-4'\n        elif t<=13: grp = '5-12'\n        elif t<=22: grp = '13-22'\n            \n        # TRAIN DATA\n        train_x = df.iloc[train_index]\n        train_x = train_x.loc[train_x.level_group == grp]\n        train_users = train_x.index.values\n        train_y = targets.loc[targets.q==t].set_index('session').loc[train_users]\n        \n        # VALID DATA\n        valid_x = df.iloc[test_index]\n        valid_x = valid_x.loc[valid_x.level_group == grp]\n        valid_users = valid_x.index.values\n        valid_y = targets.loc[targets.q==t].set_index('session').loc[valid_users]\n        \n        # TRAIN MODEL        \n        clf =  XGBClassifier(**xgb_params)\n        clf.fit(train_x[FEATURES].astype('float32'), train_y['correct'],\n                eval_set=[ (valid_x[FEATURES].astype('float32'), valid_y['correct']) ],\n                verbose=0)\n        print(f'{t}({clf.best_ntree_limit}), ',end='')\n        \n        # SAVE MODEL, PREDICT VALID OOF\n        models[f'{grp}_{t}'] = clf\n        oof.loc[valid_users, t-1] = clf.predict_proba(valid_x[FEATURES].astype('float32'))[:,1]\n        \n    print()","metadata":{"papermill":{"duration":69.877213,"end_time":"2023-02-07T01:02:43.57299","exception":false,"start_time":"2023-02-07T01:01:33.695777","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:35:48.556464Z","iopub.execute_input":"2023-06-05T03:35:48.557293Z","iopub.status.idle":"2023-06-05T03:48:01.774997Z","shell.execute_reply.started":"2023-06-05T03:35:48.557250Z","shell.execute_reply":"2023-06-05T03:48:01.773946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute CV Score\nWe need to convert prediction probabilities into `1s` and `0s`. The competition metric is F1 Score which is the harmonic mean of precision and recall. Let's find the optimal threshold for `p > threshold` when to predict `1` and when to predict `0` to maximize F1 Score.","metadata":{"papermill":{"duration":0.011241,"end_time":"2023-02-07T01:02:43.59638","exception":false,"start_time":"2023-02-07T01:02:43.585139","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# PUT TRUE LABELS INTO DATAFRAME WITH 18 COLUMNS\ntrue = oof.copy()\nfor k in range(18):\n    # GET TRUE LABELS\n    tmp = targets.loc[targets.q == k+1].set_index('session').loc[ALL_USERS]\n    true[k] = tmp.correct.values","metadata":{"execution":{"iopub.status.busy":"2023-06-05T03:48:10.640390Z","iopub.execute_input":"2023-06-05T03:48:10.640815Z","iopub.status.idle":"2023-06-05T03:48:10.778538Z","shell.execute_reply.started":"2023-06-05T03:48:10.640781Z","shell.execute_reply":"2023-06-05T03:48:10.777204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# FIND BEST THRESHOLD TO CONVERT PROBS INTO 1s AND 0s\nscores = []; thresholds = []\nbest_score = 0; best_threshold = 0\n\nfor threshold in np.arange(0.4,0.81,0.01):\n    print(f'{threshold:.02f}, ',end='')\n    preds = (oof.values.reshape((-1))>threshold).astype('int')\n    m = f1_score(true.values.reshape((-1)), preds, average='macro')   \n    scores.append(m)\n    thresholds.append(threshold)\n    if m>best_score:\n        best_score = m\n        best_threshold = threshold","metadata":{"execution":{"iopub.status.busy":"2023-06-05T03:48:14.029956Z","iopub.execute_input":"2023-06-05T03:48:14.030401Z","iopub.status.idle":"2023-06-05T03:48:22.282593Z","shell.execute_reply.started":"2023-06-05T03:48:14.030365Z","shell.execute_reply":"2023-06-05T03:48:22.281210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# PLOT THRESHOLD VS. F1_SCORE\nplt.figure(figsize=(20,5))\nplt.plot(thresholds,scores,'-o',color='blue')\nplt.scatter([best_threshold], [best_score], color='blue', s=300, alpha=1)\nplt.xlabel('Threshold',size=14)\nplt.ylabel('Validation F1 Score',size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score:.3f} at Best Threshold = {best_threshold:.3}',size=18)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T03:48:30.765267Z","iopub.execute_input":"2023-06-05T03:48:30.765685Z","iopub.status.idle":"2023-06-05T03:48:31.101441Z","shell.execute_reply.started":"2023-06-05T03:48:30.765651Z","shell.execute_reply":"2023-06-05T03:48:31.100468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('When using optimal threshold...')\nfor k in range(18):\n        \n    # COMPUTE F1 SCORE PER QUESTION\n    m = f1_score(true[k].values, (oof[k].values>best_threshold).astype('int'), average='macro')\n    print(f'Q{k}: F1 =',m)\n    \n# COMPUTE F1 SCORE OVERALL\nm = f1_score(true.values.reshape((-1)), (oof.values.reshape((-1))>best_threshold).astype('int'), average='macro')\nprint('==> Overall F1 =',m)","metadata":{"papermill":{"duration":0.771134,"end_time":"2023-02-07T01:02:44.378465","exception":false,"start_time":"2023-02-07T01:02:43.607331","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:48:35.930325Z","iopub.execute_input":"2023-06-05T03:48:35.930743Z","iopub.status.idle":"2023-06-05T03:48:36.333937Z","shell.execute_reply.started":"2023-06-05T03:48:35.930708Z","shell.execute_reply":"2023-06-05T03:48:36.332828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Infer Test Data","metadata":{"papermill":{"duration":0.011075,"end_time":"2023-02-07T01:02:44.400918","exception":false,"start_time":"2023-02-07T01:02:44.389843","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# IMPORT KAGGLE API\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n# CLEAR MEMORY\nimport gc\ndel targets, df, oof, true\n_ = gc.collect()","metadata":{"papermill":{"duration":0.052132,"end_time":"2023-02-07T01:02:44.464739","exception":false,"start_time":"2023-02-07T01:02:44.412607","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:48:41.471411Z","iopub.execute_input":"2023-06-05T03:48:41.471835Z","iopub.status.idle":"2023-06-05T03:48:41.684090Z","shell.execute_reply.started":"2023-06-05T03:48:41.471802Z","shell.execute_reply":"2023-06-05T03:48:41.682922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"limits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    \n    # FEATURE ENGINEER TEST DATA\n    df = feature_engineer(test)\n    \n    # INFER TEST DATA\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n    \n    env.predict(sample_submission)","metadata":{"papermill":{"duration":1.002014,"end_time":"2023-02-07T01:02:45.47927","exception":false,"start_time":"2023-02-07T01:02:44.477256","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:48:43.197227Z","iopub.execute_input":"2023-06-05T03:48:43.197927Z","iopub.status.idle":"2023-06-05T03:48:44.674771Z","shell.execute_reply.started":"2023-06-05T03:48:43.197887Z","shell.execute_reply":"2023-06-05T03:48:44.673635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA submission.csv","metadata":{"papermill":{"duration":0.011427,"end_time":"2023-02-07T01:02:45.502331","exception":false,"start_time":"2023-02-07T01:02:45.490904","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df = pd.read_csv('submission.csv')\nprint( df.shape )\ndf.head()","metadata":{"papermill":{"duration":0.027432,"end_time":"2023-02-07T01:02:45.541022","exception":false,"start_time":"2023-02-07T01:02:45.51359","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:49:27.804843Z","iopub.execute_input":"2023-06-05T03:49:27.805337Z","iopub.status.idle":"2023-06-05T03:49:27.828363Z","shell.execute_reply.started":"2023-06-05T03:49:27.805297Z","shell.execute_reply":"2023-06-05T03:49:27.826748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df.correct.mean())","metadata":{"papermill":{"duration":0.020233,"end_time":"2023-02-07T01:02:45.57314","exception":false,"start_time":"2023-02-07T01:02:45.552907","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-05T03:49:29.750346Z","iopub.execute_input":"2023-06-05T03:49:29.750759Z","iopub.status.idle":"2023-06-05T03:49:29.757503Z","shell.execute_reply.started":"2023-06-05T03:49:29.750724Z","shell.execute_reply":"2023-06-05T03:49:29.756162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}