{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Modified from  https://www.kaggle.com/code/takanashihumbert/magic-bingo-train-part-lb-0-687\n- Please upvote the above original notebook if you use this notebook\n- The following comments/credits were shown in the above notebook\n--------------------\n\nI'm glad to share my plan. This is based on the landmark notebook by Chris [Here](https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676). Thanks to him and all the participants who share their ideas.\nMy inference part is [Here](https://www.kaggle.com/code/takanashihumbert/magic-bingo-inference-part-lb-0-687)\n\n--------------------\n\n### For CSC532/621 Machine Learning class:\n- Take note of the differences between this current notebook and the source\n- Document your notebook thoroughly (use ChatGPT/search engines as necessary but put the content in quotations and cite the source)\n- Only share and discuss with your teammates who are officially on your Kaggle team\n- Use the following to keep track of your runs (feel free to add additional columns):\n\n## Changelog\n\n|Version | Description | LB | Note |\n| --: | -- | -- | -- |\n| V0 | Chris [Here](https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676) | `0.676` |baseline (xgboost)|\n| V1 | Source [Notebook](https://www.kaggle.com/code/takanashihumbert/magic-bingo-train-part-lb-0-687)| `0.687` |source baseline (xgboost for each question)|\n| V2 |  |  |catboost|\n\n","metadata":{}},{"cell_type":"markdown","source":"# **Import Libraries**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport gc\nimport pickle\nimport polars as pl\nfrom sklearn.model_selection import GroupKFold\nfrom catboost import CatBoostClassifier, Pool\nfrom sklearn.metrics import f1_score\nfrom tqdm.notebook import tqdm\nfrom collections import defaultdict\nimport warnings\nfrom itertools import combinations\n\nwarnings.filterwarnings('ignore')\npd.set_option(\"display.max_columns\", None)\npd.set_option(\"display.max_rows\", 200)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:41:57.957502Z","iopub.execute_input":"2023-02-26T18:41:57.958544Z","iopub.status.idle":"2023-02-26T18:41:58.825426Z","shell.execute_reply.started":"2023-02-26T18:41:57.958501Z","shell.execute_reply":"2023-02-26T18:41:58.824377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet below creates a list **train_histories** containing three **tf.RaggedTensor** objects. Each **tf.RaggedTensor** represents a \"history\" of unique events played during game sessions in a particular \"level_group\" category.\n\nTo create these tensors, the code first subsets **train_df** to only include rows with a \"level_group\" value equal to the current iteration index **i**. Then, the \"session_ix\" and \"unique_ix\" columns are selected and converted to **np.int64** data type using **astype()**.\n\nThese two columns are then used to create a **tf.RaggedTensor** using the **tf.RaggedTensor.from_value_rowids()** function. The **values** argument is set to a tensor of **unique_ix** values, and the **value_rowids** argument is set to a tensor of **session_ix** values. The **nrows** argument is set to **num_session**, which specifies the number of unique game sessions.\n\nFinally, the resulting **tf.RaggedTensor** is appended to the **train_histories** list, and the process repeats for the remaining two \"level_group\" categories.","metadata":{}},{"cell_type":"code","source":"targets = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv')\ntargets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]))\ntargets['q'] = targets.session_id.apply(lambda x: int(x.split('_')[-1][1:]))\nprint(targets.shape)","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:41:58.831744Z","iopub.execute_input":"2023-02-26T18:41:58.832511Z","iopub.status.idle":"2023-02-26T18:41:59.383222Z","shell.execute_reply.started":"2023-02-26T18:41:58.832471Z","shell.execute_reply":"2023-02-26T18:41:59.381166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet below is using PyPolars (a fast and efficient open-source data analysis library for Python) to read a large dataset from a parquet file, filter it based on the **level_group** column, and create separate dataframes **df1**, **df2**, and **df3** for each level group.\n\nSpecifically, the code is doing the following:\n\n1. **pl.read_parquet(\"/kaggle/input/game-play-train-parquet/train.parquet\")**: reads the data from the parquet file into a PyPolars DataFrame.\n\n2. **.drop([\"fullscreen\", \"hq\", \"music\"])**: drops the columns \"fullscreen\", \"hq\", and \"music\" from the DataFrame.\n\n3. **.with_columns(columns)**: creates new columns in the DataFrame based on the columns list.\n\n4. **df1 = df.filter(pl.col(\"level_group\")=='0-4')**: filters the DataFrame to create a new DataFrame **df1** that only contains rows with a **level_group** value of '0-4'.\n\n5. **df2 = df.filter(pl.col(\"level_group\")=='5-12')**: filters the DataFrame to create a new DataFrame **df2** that only contains rows with a **level_group** value of '5-12'.\n\n6. **df3 = df.filter(pl.col(\"level_group\")=='13-22')**: filters the DataFrame to create a new DataFrame **df3** that only contains rows with a **level_group** value of '13-22'.","metadata":{}},{"cell_type":"code","source":"%%time\ncolumns = [\n    pl.col(\"page\").cast(pl.Float32),\n    (\n        (pl.col(\"elapsed_time\") - pl.col(\"elapsed_time\").shift(1)) # time used for each action\n         .fill_null(0)\n         .clip(0, 1e9)\n         .over([\"session_id\", \"level_group\"])\n         .alias(\"elapsed_time_diff\")\n    ),\n    (\n        (pl.col(\"screen_coor_x\") - pl.col(\"screen_coor_x\").shift(1)) # location x changed for click \n         .abs()\n         .over([\"session_id\", \"level_group\"])\n    ),\n    (\n        (pl.col(\"screen_coor_y\") - pl.col(\"screen_coor_y\").shift(1)) # location y changed for click \n         .abs()\n         .over([\"session_id\", \"level_group\"])\n    ),\n    pl.col(\"fqid\").fill_null(\"fqid_None\"),\n    pl.col(\"text_fqid\").fill_null(\"text_fqid_None\")\n]\n\ndf = (pl.read_parquet(\"/kaggle/input/game-play-train-parquet/train.parquet\")\n      .drop([\"fullscreen\", \"hq\", \"music\"])\n      .with_columns(columns))\n\ndf1 = df.filter(pl.col(\"level_group\")=='0-4')\ndf2 = df.filter(pl.col(\"level_group\")=='5-12')\ndf3 = df.filter(pl.col(\"level_group\")=='13-22')","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:41:59.385647Z","iopub.execute_input":"2023-02-26T18:41:59.386738Z","iopub.status.idle":"2023-02-26T18:42:14.146843Z","shell.execute_reply.started":"2023-02-26T18:41:59.386693Z","shell.execute_reply":"2023-02-26T18:42:14.14588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code below is creating three dataframes (**df1**, **df2**, and **df3**) by filtering rows from a larger dataframe read from a Parquet file. The filter criteria is based on the **level_group** column of the dataframe, which has three possible values: **'0-4'**, **'5-12'**, and **'13-22'**.\n\nEach of the three dataframes contains columns related to game play events such as the name of the event, the name of the object interacted with, and various numerical values related to time and location. The code defines CATS as a list of categorical feature columns ('**event_name**', '**name**', '**fqid**', '**room_fqid**', '**text**', '**text_fqid**', and '**level**') and **NUMS** as a list of numerical feature columns ('**page**', '**room_coor_x**', '**room_coor_y**', '**screen_coor_x**', '**screen_coor_y**', '**hover_duration**', and '**elapsed_time_diff**'). The code also defines two lists of possible values for the '**event_name**' and '**name**' categorical columns (**event_name_feature** and **name_feature**).","metadata":{}},{"cell_type":"code","source":"CATS = ['event_name', 'name','fqid', 'room_fqid', 'text', 'text_fqid', 'level']\n\nevent_name_feature = ['cutscene_click', 'person_click', 'navigate_click',\n       'observation_click', 'notification_click', 'object_click',\n       'object_hover', 'map_hover', 'map_click', 'checkpoint',\n       'notebook_click']\n\nname_feature = ['basic', 'undefined', 'close', 'open', 'prev', 'next']\n\nNUMS = [ \n    'page', \n    'room_coor_x', \n    'room_coor_y', \n    'screen_coor_x', \n    'screen_coor_y', \n    'hover_duration', \n    'elapsed_time_diff'\n]","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:42:14.151367Z","iopub.execute_input":"2023-02-26T18:42:14.152502Z","iopub.status.idle":"2023-02-26T18:42:14.163627Z","shell.execute_reply.started":"2023-02-26T18:42:14.152463Z","shell.execute_reply":"2023-02-26T18:42:14.162563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are some useful features:\n* event numbers for each sessions\n* average, minimum and maximum time consumed for each 'event_name' and 'name'\n* features about 'bingo'(when users successfully click the correct place and finish the phased games), which can be translated as comprehension and deductive ability.","metadata":{}},{"cell_type":"markdown","source":"This function below appears to be implementing feature engineering on a given dataset. It takes in several parameters, including **x** which is the dataset to be processed, **grp** which is a string indicating the age group of the users, **use_extra** which is a boolean indicating whether to include additional features or not, and **feature_suffix** which is a string used as a suffix for each generated feature name.\n\nThe function first defines a list of aggregation functions to be applied on the dataset, including count of sessions, number of unique values for categorical variables, mean, minimum, and maximum values for numerical variables, and mean, minimum, and maximum elapsed time for different event and name features. The function then groups the dataset by session_id, applies the defined aggregations, and returns a Pandas dataframe.\n\nIf **use_extra** is True, additional aggregation functions are defined and applied on the dataset based on the age group indicated by **grp**. These additional features are then joined with the previously generated features before returning the final dataframe.","metadata":{}},{"cell_type":"code","source":"def feature_engineer(x, grp, use_extra, feature_suffix):\n        \n    aggs = [\n        pl.col(\"index\").count().alias(f\"session_number_{feature_suffix}\"),\n        *[pl.col(c).drop_nulls().n_unique().alias(f\"{c}_unique_{feature_suffix}\") for c in CATS],\n        *[pl.col(c).mean().alias(f\"{c}_mean_{feature_suffix}\") for c in NUMS],\n        *[pl.col(c).min().alias(f\"{c}_min_{feature_suffix}\") for c in NUMS],\n        *[pl.col(c).max().alias(f\"{c}_max_{feature_suffix}\") for c in NUMS],\n        *[pl.col(\"elapsed_time_diff\").filter(pl.col(\"event_name\")==c).mean().alias(f\"{c}_ET_mean_{feature_suffix}\") for c in event_name_feature],\n        *[pl.col(\"elapsed_time_diff\").filter(pl.col(\"event_name\")==c).max().alias(f\"{c}_ET_max_{feature_suffix}\") for c in event_name_feature],\n        *[pl.col(\"elapsed_time_diff\").filter(pl.col(\"event_name\")==c).min().alias(f\"{c}_ET_min_{feature_suffix}\") for c in event_name_feature],\n        *[pl.col(\"elapsed_time_diff\").filter(pl.col(\"name\")==c).mean().alias(f\"{c}_ET_mean_{feature_suffix}\") for c in name_feature],\n        *[pl.col(\"elapsed_time_diff\").filter(pl.col(\"name\")==c).max().alias(f\"{c}_ET_max_{feature_suffix}\") for c in name_feature],\n        *[pl.col(\"elapsed_time_diff\").filter(pl.col(\"name\")==c).min().alias(f\"{c}_ET_min_{feature_suffix}\") for c in name_feature],\n    ]\n    \n    df = x.groupby([\"session_id\"], maintain_order=True).agg(aggs).sort(\"session_id\")\n    \n    if use_extra:\n        if grp=='5-12':\n            aggs = [\n                pl.col(\"elapsed_time\").filter((pl.col(\"text\")==\"Here's the log book.\")|(pl.col(\"fqid\")=='logbook.page.bingo')).apply(lambda s: s.max()-s.min()).alias(\"logbook_bingo_duration\"),\n                pl.col(\"index\").filter((pl.col(\"text\")==\"Here's the log book.\")|(pl.col(\"fqid\")=='logbook.page.bingo')).apply(lambda s: s.max()-s.min()).alias(\"logbook_bingo_indexCount\"),\n                pl.col(\"elapsed_time\").filter(((pl.col(\"event_name\")=='navigate_click')&(pl.col(\"fqid\")=='reader'))|(pl.col(\"fqid\")==\"reader.paper2.bingo\")).apply(lambda s: s.max()-s.min()).alias(\"reader_bingo_duration\"),\n                pl.col(\"index\").filter(((pl.col(\"event_name\")=='navigate_click')&(pl.col(\"fqid\")=='reader'))|(pl.col(\"fqid\")==\"reader.paper2.bingo\")).apply(lambda s: s.max()-s.min()).alias(\"reader_bingo_indexCount\"),\n                pl.col(\"elapsed_time\").filter(((pl.col(\"event_name\")=='navigate_click')&(pl.col(\"fqid\")=='journals'))|(pl.col(\"fqid\")==\"journals.pic_2.bingo\")).apply(lambda s: s.max()-s.min()).alias(\"journals_bingo_duration\"),\n                pl.col(\"index\").filter(((pl.col(\"event_name\")=='navigate_click')&(pl.col(\"fqid\")=='journals'))|(pl.col(\"fqid\")==\"journals.pic_2.bingo\")).apply(lambda s: s.max()-s.min()).alias(\"journals_bingo_indexCount\"),\n            ]\n            tmp = x.groupby([\"session_id\"], maintain_order=True).agg(aggs).sort(\"session_id\")\n            df = df.join(tmp, on=\"session_id\", how='left')\n\n        if grp=='13-22':\n            aggs = [\n                pl.col(\"elapsed_time\").filter(((pl.col(\"event_name\")=='navigate_click')&(pl.col(\"fqid\")=='reader_flag'))|(pl.col(\"fqid\")==\"tunic.library.microfiche.reader_flag.paper2.bingo\")).apply(lambda s: s.max()-s.min() if s.len()>0 else 0).alias(\"reader_flag_duration\"),\n                pl.col(\"index\").filter(((pl.col(\"event_name\")=='navigate_click')&(pl.col(\"fqid\")=='reader_flag'))|(pl.col(\"fqid\")==\"tunic.library.microfiche.reader_flag.paper2.bingo\")).apply(lambda s: s.max()-s.min() if s.len()>0 else 0).alias(\"reader_flag_indexCount\"),\n                pl.col(\"elapsed_time\").filter(((pl.col(\"event_name\")=='navigate_click')&(pl.col(\"fqid\")=='journals_flag'))|(pl.col(\"fqid\")==\"journals_flag.pic_0.bingo\")).apply(lambda s: s.max()-s.min() if s.len()>0 else 0).alias(\"journalsFlag_bingo_duration\"),\n                pl.col(\"index\").filter(((pl.col(\"event_name\")=='navigate_click')&(pl.col(\"fqid\")=='journals_flag'))|(pl.col(\"fqid\")==\"journals_flag.pic_0.bingo\")).apply(lambda s: s.max()-s.min() if s.len()>0 else 0).alias(\"journalsFlag_bingo_indexCount\")\n            ]\n            tmp = x.groupby([\"session_id\"], maintain_order=True).agg(aggs).sort(\"session_id\")\n            df = df.join(tmp, on=\"session_id\", how='left')\n        \n    return df.to_pandas()","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:42:14.165348Z","iopub.execute_input":"2023-02-26T18:42:14.166146Z","iopub.status.idle":"2023-02-26T18:42:14.206339Z","shell.execute_reply.started":"2023-02-26T18:42:14.16611Z","shell.execute_reply":"2023-02-26T18:42:14.20513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"this function below apper to timing how long it takes to perform feature engineering on three different dataframes (**df1**, **df2**, and **df3**) using the **feature_engineer** function with different parameters (**grp='0-4'**, **grp='5-12'**, and **grp='13-22'**) and printing out a message when each dataframe is done.","metadata":{}},{"cell_type":"code","source":"%%time\ndf1 = feature_engineer(df1, grp='0-4', use_extra=True, feature_suffix='')\nprint('df1 done')\ndf2 = feature_engineer(df2, grp='5-12', use_extra=True, feature_suffix='')\nprint('df2 done')\ndf3 = feature_engineer(df3, grp='13-22', use_extra=True, feature_suffix='')\nprint('df3 done')","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:42:14.212275Z","iopub.execute_input":"2023-02-26T18:42:14.214925Z","iopub.status.idle":"2023-02-26T18:42:28.790035Z","shell.execute_reply.started":"2023-02-26T18:42:14.214885Z","shell.execute_reply":"2023-02-26T18:42:28.788891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Remove some redundant features","metadata":{}},{"cell_type":"markdown","source":"the code below is checking for null values in each dataframe (df1, df2, df3) and dropping columns with more than 90% null values. Then, it is also checking for columns with only one unique value in each dataframe and appending them to the drop list.\n\nThe progress of the loops is printed using tqdm.","metadata":{}},{"cell_type":"code","source":"null1 = df1.isnull().sum().sort_values(ascending=False) / len(df1)\nnull2 = df2.isnull().sum().sort_values(ascending=False) / len(df1)\nnull3 = df3.isnull().sum().sort_values(ascending=False) / len(df1)\n\ndrop1 = list(null1[null1>0.9].index)\ndrop2 = list(null2[null2>0.9].index)\ndrop3 = list(null3[null3>0.9].index)\nprint(len(drop1), len(drop2), len(drop3))\n\nfor col in tqdm(df1.columns):\n    if df1[col].nunique()==1:\n        print(col)\n        drop1.append(col)\nprint(\"*********df1 DONE*********\")\nfor col in tqdm(df2.columns):\n    if df2[col].nunique()==1:\n        print(col)\n        drop2.append(col)\nprint(\"*********df2 DONE*********\")\nfor col in tqdm(df3.columns):\n    if df3[col].nunique()==1:\n        print(col)\n        drop3.append(col)\nprint(\"*********df3 DONE*********\")","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:42:28.791875Z","iopub.execute_input":"2023-02-26T18:42:28.792647Z","iopub.status.idle":"2023-02-26T18:42:28.968832Z","shell.execute_reply.started":"2023-02-26T18:42:28.792606Z","shell.execute_reply":"2023-02-26T18:42:28.96766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now all dataframes have their session_id column set as index.","metadata":{}},{"cell_type":"code","source":"df1 = df1.set_index('session_id')\ndf2 = df2.set_index('session_id')\ndf3 = df3.set_index('session_id')","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:42:28.970374Z","iopub.execute_input":"2023-02-26T18:42:28.972071Z","iopub.status.idle":"2023-02-26T18:42:28.993923Z","shell.execute_reply.started":"2023-02-26T18:42:28.972029Z","shell.execute_reply":"2023-02-26T18:42:28.992958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Prepared the data and selected the relevant features for training. It's great to see that you have a large number of features for each data set, and that you will be training on all users in the data.","metadata":{}},{"cell_type":"code","source":"FEATURES1 = [c for c in df1.columns if c not in drop1+['level_group']]\nFEATURES2 = [c for c in df2.columns if c not in drop2+['level_group']]\nFEATURES3 = [c for c in df3.columns if c not in drop3+['level_group']]\nprint('We will train with', len(FEATURES1), len(FEATURES2), len(FEATURES3) ,'features')\nALL_USERS = df1.index.unique()\nprint('We will train with', len(ALL_USERS) ,'users info')","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:42:28.995231Z","iopub.execute_input":"2023-02-26T18:42:28.997192Z","iopub.status.idle":"2023-02-26T18:42:29.006883Z","shell.execute_reply.started":"2023-02-26T18:42:28.997159Z","shell.execute_reply":"2023-02-26T18:42:29.004668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code below seems to be training a CatBoost model for each of the 18 questions in the dataset, using the features selected previously for each question. It uses GroupKFold cross-validation to split the data into training and validation sets, ensuring that all sessions for the same user are kept in the same fold.\n\nFor each question, the code prints the question number, the number of selected features, and the cross-validation results for each of the 10 folds. It also saves the best iteration for each fold and each question, the feature importance for each fold and each question, and the out-of-fold predictions for the validation set.\n\nFinally, it prints the top 10 features with the highest importance for each question.","metadata":{}},{"cell_type":"code","source":"%%time\ngkf = GroupKFold(n_splits=10)\noof_cb = pd.DataFrame(data=np.zeros((len(ALL_USERS),18)), index=ALL_USERS, columns=[f'meta_{i}' for i in range(1, 19)])\n#models = {}\nbest_iteration_cb = defaultdict(list)\nimportance_dict = {}\n\n# ITERATE THRU QUESTIONS 1 THRU 18\nfor t in range(1,19):\n\n    # USE THIS TRAIN DATA WITH THESE QUESTIONS\n    if t<=3: \n        grp = '0-4'\n        df = df1\n        FEATURES = FEATURES1\n    elif t<=13: \n        grp = '5-12'\n        df = df2\n        FEATURES = FEATURES2\n    elif t<=22: \n        grp = '13-22'\n        df = df3\n        FEATURES = FEATURES3\n        \n    print('#'*25)\n    print('### question', t, 'with features', len(FEATURES))\n    print('#'*25)\n    \n    cb_params = {\n        'iterations': 9999,\n        'early_stopping_rounds': 90,\n        'depth': 4,\n        'learning_rate': 0.05,\n        'loss_function': \"Logloss\",\n        'random_seed': 0,\n        'metric_period': 1,\n        'subsample': 0.8,\n        'colsample_bylevel': 0.4,\n        'verbose': 0,\n    }\n\n    feature_importance_df = pd.DataFrame()\n    # COMPUTE CV SCORE WITH 5 GROUP K FOLD\n    for i, (train_index, test_index) in enumerate(gkf.split(X=df, groups=df.index)):\n        \n        # TRAIN DATA\n        train_x = df.iloc[train_index]\n        train_users = train_x.index.values\n        train_y = targets.loc[targets.q==t].set_index('session').loc[train_users]\n        \n        # VALID DATA\n        valid_x = df.iloc[test_index]\n        valid_users = valid_x.index.values\n        valid_y = targets.loc[targets.q==t].set_index('session').loc[valid_users]\n        \n        # TRAIN MODEL        \n        clf = CatBoostClassifier(**cb_params)\n        clf.fit(train_x[FEATURES].astype('float32'), train_y['correct'],\n                eval_set=[(valid_x[FEATURES].astype('float32'), valid_y['correct'])],\n                verbose=0)\n        print(i+1, ', ', end='')\n        best_iteration_cb[str(t)].append(clf.best_iteration_)\n        \n        fold_importance_df = pd.DataFrame()\n        fold_importance_df[\"feature\"] = FEATURES\n        fold_importance_df[\"importance\"] = clf.get_feature_importance()\n        fold_importance_df[\"fold\"] = i + 1\n        feature_importance_df = pd.concat([feature_importance_df, fold_importance_df], axis=0)\n        \n        # SAVE MODEL, PREDICT VALID OOF\n        oof_cb.loc[valid_users, f'meta_{t}'] = clf.predict_proba(valid_x[FEATURES].astype('float32'))[:,1]\n            \n\n    print()\n    feature_importance_df = feature_importance_df.groupby(['feature'])['importance'].agg(['mean']).sort_values(by='mean', ascending=False)\n    display(feature_importance_df.head(10))","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:42:29.009121Z","iopub.execute_input":"2023-02-26T18:42:29.009875Z","iopub.status.idle":"2023-02-26T18:44:15.790001Z","shell.execute_reply.started":"2023-02-26T18:42:29.009837Z","shell.execute_reply":"2023-02-26T18:44:15.788891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code below is finding the best threshold to convert the predicted probabilities from the CatBoost model into 1s and 0s. It does this by iterating over a range of thresholds from 0.4 to 0.81 in steps of 0.005, and for each threshold, it computes the F1 score between the true labels and the predicted labels (after applying the threshold). The threshold that gives the highest F1 score is chosen as the best threshold. Finally, the code plots the threshold vs. the F1 score, highlighting the best threshold with a blue dot.\n\n\n\n","metadata":{}},{"cell_type":"code","source":"true = oof_cb.copy()\nfor i in range(1, 19):\n    # GET TRUE LABELS\n    tmp = targets.loc[targets.q==i].set_index('session').loc[ALL_USERS]\n    true[f'meta_{i}'] = tmp.correct.values\n\n# FIND BEST THRESHOLD TO CONVERT PROBS INTO 1s AND 0s\nscores = []; thresholds = []\nbest_score_cb = 0; best_threshold_cb = 0\n\nfor threshold in np.arange(0.4,0.81,0.005):\n    print(f'{threshold:.03f}, ',end='')\n    preds = (oof_cb.values.reshape((-1))>threshold).astype('int')\n    m = f1_score(true.values.reshape((-1)), preds, average='macro')   \n    scores.append(m)\n    thresholds.append(threshold)\n    if m>best_score_cb:\n        best_score_cb = m\n        best_threshold_cb = threshold\n\n# PLOT THRESHOLD VS. F1_SCORE\nplt.figure(figsize=(20,5))\nplt.plot(thresholds,scores,'-o',color='blue')\nplt.scatter([best_threshold_cb], [best_score_cb], color='blue', s=300, alpha=1)\nplt.xlabel('Threshold',size=14)\nplt.ylabel('Validation F1 Score',size=14)\nplt.title(f'Threshold vs. F1_Score with Best F1_Score = {best_score_cb:.5f} at Best Threshold = {best_threshold_cb:.4}',size=18)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:44:15.791729Z","iopub.execute_input":"2023-02-26T18:44:15.792109Z","iopub.status.idle":"2023-02-26T18:44:22.066981Z","shell.execute_reply.started":"2023-02-26T18:44:15.792061Z","shell.execute_reply":"2023-02-26T18:44:22.065836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we use full data to retrain our models with the best iterations, and save them. ","metadata":{}},{"cell_type":"markdown","source":"This code cell below is training a CatBoostClassifier model for each question from 1 to 18. The training data used depends on which question is being trained for, as determined by the if/elif statements. The optimal number of estimators for each model is determined by taking the median of the best iterations found in a previous step. The CatBoostClassifier is initialized with some hyperparameters, and the model is trained using the training data and the hyperparameters.\n\nThe trained models are saved to disk as 'CB_question[t].cb', where [t] is the question number.","metadata":{}},{"cell_type":"code","source":"%%time\n# ITERATE THRU QUESTIONS 1 THRU 18\nfor t in range(1,19):\n\n    # USE THIS TRAIN DATA WITH THESE QUESTIONS\n    if t<=3: \n        grp = '0-4'\n        df = df1\n        FEATURES = FEATURES1\n    elif t<=13: \n        grp = '5-12'\n        df = df2\n        FEATURES = FEATURES2\n    elif t<=22: \n        grp = '13-22'\n        df = df3\n        FEATURES = FEATURES3\n    \n    n_estimators = int(np.median(best_iteration_cb[str(t)]) + 1)\n    cb_params = {\n        'iterations': n_estimators,\n        'depth': 4,\n        'learning_rate': 0.05,\n        'loss_function': \"Logloss\",\n        'random_seed': 0,\n        'metric_period': 1,\n        'subsample': 0.8,\n        'colsample_bylevel': 0.4,\n        'verbose': 0,\n    }\n    \n    print('#'*25)\n    print(f'### question {t} features {len(FEATURES)}')\n        \n    # TRAIN DATA\n    train_users = df.index.values\n    train_y = targets.loc[targets.q==t].set_index('session').loc[train_users]\n\n    # TRAIN MODEL        \n    clf =  CatBoostClassifier(**cb_params)\n    clf.fit(df[FEATURES].astype('float32'), train_y['correct'], verbose=0)\n    clf.save_model(f'CB_question{t}.cb')\n    \n    print()","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:44:22.068214Z","iopub.execute_input":"2023-02-26T18:44:22.071068Z","iopub.status.idle":"2023-02-26T18:44:33.679011Z","shell.execute_reply.started":"2023-02-26T18:44:22.071028Z","shell.execute_reply":"2023-02-26T18:44:33.677929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We save features names as dict for each questions","metadata":{}},{"cell_type":"code","source":"importance_dict = {}\nfor t in range(1, 19):\n    if t<=3: \n        importance_dict[str(t)] = FEATURES1\n    elif t<=13: \n        importance_dict[str(t)] = FEATURES2\n    elif t<=22:\n        importance_dict[str(t)] = FEATURES3\n\nf_save = open('importance_dict.pkl', 'wb')\npickle.dump(importance_dict, f_save)\nf_save.close()","metadata":{"execution":{"iopub.status.busy":"2023-02-26T18:44:33.683022Z","iopub.execute_input":"2023-02-26T18:44:33.683343Z","iopub.status.idle":"2023-02-26T18:44:33.692112Z","shell.execute_reply.started":"2023-02-26T18:44:33.683313Z","shell.execute_reply":"2023-02-26T18:44:33.691045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Document with ChatGPT**","metadata":{}}]}