{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CatBoost with ptls embeddings on coordinates [0.698 on val] ","metadata":{"id":"wel9fZYkzgE7"}},{"cell_type":"markdown","source":"Hi there! Thank you for visiting my notebook.\n\n### Summary \nI found a nice library Pytorch Lifestream and my goal was to create CatBoost model and add not only categorical and numerical features but also embedding features.\n\n### I created embedding fetures using \n1.   Coordinates (I firstly encoded coordinates to patches and linked every event to one of 50 rectangle patches on display)\n2.   Event name (label encoded to one of 11 unique categories)\n3.   Time diff between previous and current elapsed time to illustrate event duration\n\n### I obtained model wich has shown 0.698 on validation and I created an inference. \nBut when I uploaded models to run inference then I saw that the size of 18 models is 8.5 Gb and I couldn't run it as is with 8 Gb RAM notebook of competition\n\n* In this notebook you will find source code to construct dataframe with embeddings using only coordinates patches (I think it will be possible to run in inference as is) but you easily could also add other fetures to the embeddings using this code as a template. The result of this notebook is a **res_emb dataframe** which then in train notebook is merged with the **main feature engineered dataframe** to train CatBoost model.\n\n**Note 1** Firstly I upploaded to this notebook a Colab notebook with executed cells (to save notebook with cells output). Kaggle notebook usually allows to save imported notebook with the historical outputs. But not this time.. Therefore a half of sells doesn't have output there but I attached original notebook with cells output in the input as well 'jw-emb-small.pt' file with embeddings caclulation.\n\n**Note 2** You will have OOM failure running notebook there. I recommend to run embedding training with at least 25 Gb of RAM (more is better). I had GPU with 80 Gb RAM \n\n* You can find source code to train CatBoost model on dataframe with embeddings in this notebook https://www.kaggle.com/ivanisaev/catboost-embed-model-train. It takes near 12 howrs to train and size of 18 CatBoost models is about 8 Gb as I said. But I attached models for the first fold in dataset and use only one fold to speed up computation. If you will use 5 folds you will get 0.698 instead of 0.697\n\n* You can find template for the inference with embeddings there https://www.kaggle.com/ivanisaev/embed-model-inference\n\n\n### How you can use this stuff\n1.   For example you can make embeddings not per each level but per level group or add modells with embeddings only for particular question. \n2.   You can also use embedding models only for a part (for example half of questions for which it outperforms other open models). Bellow are scores comapring to Catboost Mix (+/-/= sign in the tail is comparison of nodel with embeddings to Catboost Mix)\n\nQ0: F1 = 0.6820389661884507 +\n\nQ1: F1 = 0.4964824307581739 -\n\nQ2: F1 = 0.5134095523025964 -\n\nQ3: F1 = 0.6824019565074062 +\n\nQ4: F1 = 0.6332403813506928 +!\n\nQ5: F1 = 0.6437609220449432 +\n\nQ6: F1 = 0.6203295875015424 -\n\nQ7: F1 = 0.561351851437172 -\n\nQ8: F1 = 0.6315313558621829 =\n\nQ9: F1 = 0.5753142329226386 +\n\nQ10: F1 = 0.613984855671378 +\n\nQ11: F1 = 0.5175918772436235 -\n\nQ12: F1 = 0.46237961000051414 +\n\nQ13: F1 = 0.6353788380168102 =\n\nQ14: F1 = 0.5963746272434363 +\n\nQ15: F1 = 0.49890687086950786 +\n\nQ16: F1 = 0.5530371819235946 -\n\nQ17: F1 = 0.49148964784124965 =","metadata":{"id":"5nf1-vZ8GZ-X"}},{"cell_type":"markdown","source":"### Data Loading","metadata":{"id":"vWmwuvuBzgFA"}},{"cell_type":"code","source":"import pandas as pd, numpy as np, gc\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score","metadata":{"id":"OA0vttrsz4UQ","outputId":"2b689c41-7aa1-474d-b9f2-d1e02958c251","execution":{"iopub.status.busy":"2023-05-29T19:50:11.678306Z","iopub.execute_input":"2023-05-29T19:50:11.678773Z","iopub.status.idle":"2023-05-29T19:50:12.452011Z","shell.execute_reply.started":"2023-05-29T19:50:11.678731Z","shell.execute_reply":"2023-05-29T19:50:12.451092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd, numpy as np, gc\nfrom sklearn.model_selection import KFold, GroupKFold\nfrom xgboost import XGBClassifier\nfrom sklearn.metrics import f1_score\n# seed everything\n# torch.manual_seed(101)\nnp.random.seed(101)\n\n# READ USER ID ONLY\ntmp = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train.csv\",usecols=[0])\ntmp = tmp.groupby('session_id').session_id.agg('count')\n\n# COMPUTE READS AND SKIPS\nPIECES = 10\nCHUNK = int( np.ceil(len(tmp)/PIECES) )\n\nreads = []\nskips = [0]\nfor k in range(PIECES):\n    a = k*CHUNK\n    b = (k+1)*CHUNK\n    if b>len(tmp): b=len(tmp)\n    r = tmp.iloc[a:b].sum()\n    reads.append(r)\n    skips.append(skips[-1]+r)\n    \nprint(f'To avoid memory error, we will read train in {PIECES} pieces of sizes:')\nprint(reads)","metadata":{"id":"G60OfjXozgFB","outputId":"32df98b6-03af-4395-839a-4c5111845645","execution":{"iopub.status.busy":"2023-05-29T19:50:12.453912Z","iopub.execute_input":"2023-05-29T19:50:12.454229Z","iopub.status.idle":"2023-05-29T19:51:22.762599Z","shell.execute_reply.started":"2023-05-29T19:50:12.454204Z","shell.execute_reply":"2023-05-29T19:51:22.761481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# tmp = pd.read_csv(\"/content/drive/MyDrive/ref-predict-student-performance-from-game-play/train.csv\",usecols=['room_coor_x','room_coor_y'])\n# max_height = tmp['room_coor_x'].max()\n# min_height = tmp['room_coor_x'].min()\n# max_width = tmp['room_coor_y'].max()\n# min_width = tmp['room_coor_y'].min()\n# print(max_height, min_height, max_width, min_width)\n# del tmp\n# gc.collect()\nmax_height = 1261.7737454550663 \nmin_height = -1992.3545688360275 \nmax_width = 543.6164243795992 \nmin_width = -918.1623490877204\n","metadata":{"id":"b5Fk3NqYzgFC","outputId":"04acbbf1-ffd0-4401-fb75-f8ad806f8fb9","execution":{"iopub.status.busy":"2023-05-29T19:51:22.764081Z","iopub.execute_input":"2023-05-29T19:51:22.764511Z","iopub.status.idle":"2023-05-29T19:51:22.770435Z","shell.execute_reply.started":"2023-05-29T19:51:22.764478Z","shell.execute_reply":"2023-05-29T19:51:22.769269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_px = 8\nnum_py = 6\n# Calculate patch index given a coordinate (x,y)\ndef get_patch_index(x, y, min_image_height, max_image_height, min_image_width, max_image_width, num_patches_height = 8, num_patches_width = 6):\n    if np.isnan(x):\n        return num_px * num_py + 1\n    patch_height = (max_image_height - min_image_height + 1) / num_patches_height\n    patch_width = (max_image_width - min_image_width + 1) / num_patches_width\n    patch_x = (x - min_image_height) // patch_height\n    patch_y = (y - min_image_width) // patch_width\n    patch_index = patch_x * num_patches_width + patch_y + 1\n    return patch_index","metadata":{"id":"gGHzEWV-zgFF","execution":{"iopub.status.busy":"2023-05-29T19:51:22.773637Z","iopub.execute_input":"2023-05-29T19:51:22.773979Z","iopub.status.idle":"2023-05-29T19:51:22.783883Z","shell.execute_reply.started":"2023-05-29T19:51:22.773950Z","shell.execute_reply":"2023-05-29T19:51:22.782972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# PROCESS TRAIN DATA IN PIECES\ngps = []\nprint(f'Processing train as {PIECES} pieces to avoid memory error... ')\nfor k in range(PIECES):\n    print(k,', ',end='')\n    SKIPS = 0\n    if k>0: SKIPS = range(1,skips[k]+1)\n    train = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv',\n                        usecols = ['session_id', 'level_group', 'level', 'elapsed_time', 'room_coor_x', 'room_coor_y'],\n                        nrows=reads[k], skiprows=SKIPS)\n    for _, session in train.groupby('session_id'):\n        for _, gp in session.groupby('level_group'):\n            gp['patch'] = gp.apply(lambda row : get_patch_index(row['room_coor_x'], row['room_coor_y'], min_height, max_height, min_width, max_width, num_px, num_py), axis = 1)\n            gps.append(gp)\n    \n# CONCATENATE ALL PIECES\nprint('\\n')\ndel train; gc.collect()\nnum_patches = num_px * num_py + 2\ndf = pd.concat(gps)\nprint('Shape of all train data after adding patch:', df.shape )","metadata":{"id":"aCDPn_13zgFF","outputId":"19899de6-5470-4345-f477-1753da3ea0f8","execution":{"iopub.status.busy":"2023-05-29T19:51:22.784966Z","iopub.execute_input":"2023-05-29T19:51:22.785746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create embeddings","metadata":{"id":"zrdtUARJ2Owq"}},{"cell_type":"code","source":"!pip install pytorch-lifestream\n!pip install \"torch<2\"\n!pip install -U \"pytorch-lightning<2\"\n!pip install -U \"torchvision<0.15.1\"","metadata":{"id":"QOVLbwyi2NCG","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip freeze | grep torch","metadata":{"id":"dXqt_uem2NUb","outputId":"65ed80f1-fe02-406c-c19b-ac1f66cc7abc","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%load_ext autoreload\n%autoreload 2\n\n# import logging\nimport torch\nimport pytorch_lightning as pl\n# import warnings\n\n# warnings.filterwarnings('ignore')\n# logging.getLogger(\"pytorch_lightning\").setLevel(logging.ERROR)","metadata":{"id":"pxgcWAvp2NXT","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[\"sid_level\"] = df[\"session_id\"].astype(str) + ' ' + df[\"level\"].astype(str)\ndf = df.drop([\"session_id\", \"level\", 'room_coor_x' ,'room_coor_y', 'level_group'], axis = 1)","metadata":{"id":"hEe_oq7M26np","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from ptls.preprocessing import PandasDataPreprocessor\n\npreprocessor = PandasDataPreprocessor(\n    col_id= 'sid_level',\n    col_event_time='elapsed_time',\n    event_time_transformation='none',\n    cols_category= ['patch'],\n    return_records=True,\n)","metadata":{"id":"pcSNj0G42902","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ntrain = preprocessor.fit_transform(df)","metadata":{"id":"LCA7Kt_x3Abt","outputId":"0efb9970-548c-4adb-d6ff-41b4eb86cf72","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pickle\n\nwith open('/kaggle/working/preprocessor.p', 'wb') as f:\n    pickle.dump(preprocessor, f)","metadata":{"id":"-MimcRgC3Cm9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = sorted(train, key=lambda x: x['sid_level'])","metadata":{"id":"dqAHVLAO3DWk","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from functools import partial\nfrom ptls.nn import TrxEncoder, RnnSeqEncoder\nfrom ptls.frames.coles import CoLESModule\n\ntrx_encoder_params = dict(\n    embeddings_noise=0.003,\n    embeddings={\n        'patch': {'in': 55, 'out': 55},\n    },\n)\n\nseq_encoder = RnnSeqEncoder(\n    trx_encoder=TrxEncoder(**trx_encoder_params),\n    hidden_size=256,\n    type='gru',\n)\n\nmodel = CoLESModule(\n    seq_encoder=seq_encoder,\n    optimizer_partial=partial(torch.optim.Adam, lr=0.001),\n    lr_scheduler_partial=partial(torch.optim.lr_scheduler.StepLR, step_size=30, gamma=0.9),\n)","metadata":{"id":"lkvmCFEC3DTw","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from ptls.data_load.datasets import MemoryMapDataset\nfrom ptls.data_load.iterable_processing import SeqLenFilter\nfrom ptls.frames.coles import ColesDataset\nfrom ptls.frames.coles.split_strategy import SampleSlices\nfrom ptls.frames import PtlsDataModule\n\ntrain_dl = PtlsDataModule(\n    train_data=ColesDataset(\n        MemoryMapDataset(\n            data=train,\n            i_filters=[\n                SeqLenFilter(min_seq_len=25),\n            ],\n        ),\n        splitter=SampleSlices(\n            split_count=5,\n            cnt_min=25,\n            cnt_max=200,\n        ),\n    ),\n    train_num_workers=16,\n    train_batch_size=256,\n)","metadata":{"id":"az9NvDB23DRH","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nimport pytorch_lightning as pl\n\nimport logging\n\ntrainer = pl.Trainer(\n    max_epochs=15,\n    gpus=1 if torch.cuda.is_available() else 0,\n    enable_progress_bar=False,\n)","metadata":{"id":"Pri5_4De3O6M","outputId":"f3b39177-ff03-48e1-b0ab-7e27e1d1ddfe","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nprint(f'logger.version = {trainer.logger.version}')\ntrainer.fit(model, train_dl)\nprint(trainer.logged_metrics)","metadata":{"id":"Vup31uN93RdA","outputId":"2f6838df-173e-4c13-f89e-3ed4d6843d71","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.save(seq_encoder.state_dict(), \"/content/drive/MyDrive/ref-predict-student-performance-from-game-play/jw-emb-small.pt\")","metadata":{"id":"daJn5Liw3TpQ","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create dataframe with embeddings","metadata":{"id":"5gHVsI4R613v"}},{"cell_type":"code","source":"from ptls.preprocessing import PandasDataPreprocessor\nimport torch\nfrom ptls.frames.supervised import SequenceToTarget\nfrom ptls.nn import TrxEncoder, RnnSeqEncoder\nimport pytorch_lightning as pl\nfrom ptls.data_load.datasets import inference_data_loader\n\nseq_encoder.load_state_dict(torch.load('/kaggle/input/jw-emb-smallpt/jw-emb-small.pt'))\nmodel = SequenceToTarget(seq_encoder)\nmodel.eval();\n\ntrainer = pl.Trainer(gpus=1 if torch.cuda.is_available() else 0)\ntest_dl = inference_data_loader(train, num_workers=0, batch_size=256)\ntest_embeds = torch.vstack(trainer.predict(model, test_dl))\n\ntest_df_patched = pd.DataFrame(data=test_embeds, columns=[f'embed_{i}' for i in range(test_embeds.shape[1])])\ntest_df_patched['sid_level'] = [x['sid_level'] for x in train]\n\nemeb_df = pd.DataFrame(np.arange(len(test_df_patched)))\nemeb_df['embeddings'] = test_df_patched.iloc[:, :-1].values.tolist()\nemeb_df = emeb_df.drop([0], axis=1)\nemeb_df['sid_level'] = test_df_patched['sid_level']\n\nemeb_df['session_id'], emeb_df['level'] = emeb_df['sid_level'].str.split(' ', 1).str\nemeb_df = emeb_df.drop('sid_level', axis = 1)\nemeb_df['level'] = emeb_df['level'].astype(int)\nemeb_df = emeb_df.sort_values(by=['session_id', 'level'])\n\nemeb_df['level'] = emeb_df['level'].astype(int)\nemeb_df = emeb_df.sort_values(by=['session_id', 'level'])\n\nres_emb = emeb_df.pivot(index='session_id', columns='level', values='embeddings').reset_index()\nres_emb = res_emb.set_index('session_id').reset_index().rename_axis(None, axis=1)\ncols = [f'level{x}_emb' for x in range(23)]\ncols.insert(0, 'session_id')\nres_emb.columns = cols\nemb_features = list(res_emb.columns)\nemb_features.remove('session_id')\n\nres_emb['session_id'] = res_emb['session_id'].astype(int)\nres_emb['level7_emb'] = res_emb['level7_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level15_emb'] = res_emb['level15_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level20_emb'] = res_emb['level20_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level21_emb'] = res_emb['level21_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\n\nres_emb['level0_emb'] = res_emb['level0_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level1_emb'] = res_emb['level1_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level2_emb'] = res_emb['level2_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level3_emb'] = res_emb['level3_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level4_emb'] = res_emb['level4_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level5_emb'] = res_emb['level5_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level6_emb'] = res_emb['level6_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level8_emb'] = res_emb['level8_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level9_emb'] = res_emb['level9_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level10_emb'] = res_emb['level10_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level11_emb'] = res_emb['level11_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level12_emb'] = res_emb['level12_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level13_emb'] = res_emb['level13_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level14_emb'] = res_emb['level14_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level16_emb'] = res_emb['level16_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level17_emb'] = res_emb['level17_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level18_emb'] = res_emb['level18_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level19_emb'] = res_emb['level19_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\nres_emb['level22_emb'] = res_emb['level22_emb'].apply(lambda d: d if isinstance(d, list) else [float(0)]*256)\n\n# df_wemb1 = df1.merge(res_emb, on = 'session_id')\n# df_wemb1 = df_wemb.set_index('session_id')\n# df2 = df2.merge(res_emb, on = 'session_id')\n# df2 = df2.set_index('session_id')\n# df3 = df3.merge(res_emb, on = 'session_id')\n# df3 = df3.set_index('session_id')","metadata":{"id":"5poRbDd-6f-L","outputId":"41f7acc7-9179-4714-ce1d-8356fee5abea","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"res_emb.to_csv('/kaggle/working/res_emb.csv', index = False)\n# You can find res_emb.csv in the /kaggle/working/res-emb directory of this notebook\n","metadata":{"id":"7kw9UupuCkjD","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Merge res_emb with main feature engineered df and train model\n* You can find source code to do that in this notebook https://www.kaggle.com/ivanisaev/catboost-embed-model-train\nThe only difference -- it uses the first version of embeddings (with patch, even name and elapsed time diff). You can load this dataset from the https://www.kaggle.com/ivanisaev/catboost-embed-model-train notebook attachement and run pretrained models out of the box\n\n* You can find template for the inference with embeddings there https://www.kaggle.com/ivanisaev/embed-model-inference","metadata":{"id":"5xO6WfdPDe6T"}}]}