{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this notebook we will train an `XGBoost Ranker` on the GPU and perform prediction.\n\nTraining with varied architectures and ensembling (please see [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions) for a tutorial on ensembling) can offer you a significant jump on the LB!\n\nTraining with `XGBoost` however offers more additional advantages. In comparison to `LGBM`, `XGBoost` allows you to train with the following objectives (`LGBM` gives you access to a single loss only for ranking, training with different objectives is a great way of improving your ensemble!):\n* `rank:pairwise`\n* `rank:ndcg`\n* `rank:map`\n\nOn top of that, we will train on the GPU! 🔥 GPU can offer a significant speed-up. You can train more and bigger models in a shorter amount of time. However, when training on the GPU with large amounts of tabular data, you can easily run into problems (how to load the data onto the GPU for processing in chunks, how to manage memory).\n\nAs we want to focus on feature engineering and training lets offload all the low level, tedious considerations to the `Merlin Framework`!\n\nIn this notebook, we will introduce the entire pipeline. We will preprocess our data on the GPU using a library specifically designed for tabular data preprocessing, `NVTabular`. We will then proceed to train our `XGBoost` model with `Merlin Models`. In the background  we will leverage `dask_cuda` and distributed training to optimize the use of available GPU RAM, but we will let the libraries handle all that! No additional configuration will be required from us.\n\nLet's get started!\n\n## Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Libraries installation\n\nWe will need a couple of libraries that do not come preinstalled on the Kaggle VM. Let's install them here.","metadata":{}},{"cell_type":"code","source":"%%capture\n\n!pip install nvtabular==1.3.3 merlin-models polars merlin-core==v0.4.0 dask_cuda","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:05:59.135162Z","iopub.execute_input":"2022-11-28T04:05:59.135804Z","iopub.status.idle":"2022-11-28T04:07:43.864267Z","shell.execute_reply.started":"2022-11-28T04:05:59.135703Z","shell.execute_reply":"2022-11-28T04:07:43.862961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Processing","metadata":{}},{"cell_type":"markdown","source":"We will briefly preprocess our data using polars. After that step, we will hand it over to `NVTabular` to tag our data (so that our model will know where to find the information it needs for training).","metadata":{}},{"cell_type":"code","source":"from nvtabular import *\nfrom merlin.schema.tags import Tags\nimport polars as pl\nimport xgboost as xgb\n\nfrom merlin.core.utils import Distributed\nfrom merlin.models.xgb import XGBoost\nfrom nvtabular.ops import AddTags","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:07:43.868869Z","iopub.execute_input":"2022-11-28T04:07:43.869191Z","iopub.status.idle":"2022-11-28T04:07:47.576923Z","shell.execute_reply.started":"2022-11-28T04:07:43.869161Z","shell.execute_reply":"2022-11-28T04:07:47.575967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pl.read_parquet('../input/otto-train-and-test-data-for-local-validation/test.parquet')\ntrain_labels = pl.read_parquet('../input/otto-train-and-test-data-for-local-validation/test_labels.parquet')\n\ndef add_action_num_reverse_chrono(df):\n    return df.select([\n        pl.col('*'),\n        pl.col('session').cumcount().reverse().over('session').alias('action_num_reverse_chrono')\n    ])\n\ndef add_session_length(df):\n    return df.select([\n        pl.col('*'),\n        pl.col('session').count().over('session').alias('session_length')\n    ])\n\ndef add_log_recency_score(df):\n    linear_interpolation = 0.1 + ((1-0.1) / (df['session_length']-1)) * (df['session_length']-df['action_num_reverse_chrono']-1)\n    return df.with_columns(pl.Series(2**linear_interpolation - 1).alias('log_recency_score')).fill_nan(1)\n\ndef add_type_weighted_log_recency_score(df):\n    type_weights = {0:1, 1:6, 2:3}\n    type_weighted_log_recency_score = pl.Series(df['type'].apply(lambda x: type_weights[x]) * df['log_recency_score'])\n    return df.with_column(type_weighted_log_recency_score.alias('type_weighted_log_recency_score'))\n\ndef apply(df, pipeline):\n    for f in pipeline:\n        df = f(df)\n    return df\n\npipeline = [add_action_num_reverse_chrono, add_session_length, add_log_recency_score, add_type_weighted_log_recency_score]\n\ntrain = apply(train, pipeline)\n\ntype2id = {\"clicks\": 0, \"carts\": 1, \"orders\": 2}\n\ntrain_labels = train_labels.explode('ground_truth').with_columns([\n    pl.col('ground_truth').alias('aid'),\n    pl.col('type').apply(lambda x: type2id[x])\n])[['session', 'type', 'aid']]\n\ntrain_labels = train_labels.with_columns([\n    pl.col('session').cast(pl.datatypes.Int32),\n    pl.col('type').cast(pl.datatypes.UInt8),\n    pl.col('aid').cast(pl.datatypes.Int32)\n])\n\ntrain_labels = train_labels.with_column(pl.lit(1).alias('gt'))\n\ntrain = train.join(train_labels, how='left', on=['session', 'type', 'aid']).with_column(pl.col('gt').fill_null(0))","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:07:47.577974Z","iopub.execute_input":"2022-11-28T04:07:47.578640Z","iopub.status.idle":"2022-11-28T04:07:59.861642Z","shell.execute_reply.started":"2022-11-28T04:07:47.578604Z","shell.execute_reply":"2022-11-28T04:07:59.860512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us now define the preprocessing steps we would like to apply to our data.","metadata":{}},{"cell_type":"code","source":"train_ds = Dataset(train.to_pandas())\n\nfeature_cols = ['aid', 'type','action_num_reverse_chrono', 'session_length', 'log_recency_score', 'type_weighted_log_recency_score']\ntarget = ['gt'] >> AddTags([Tags.TARGET])\nqid_column = ['session'] >>  AddTags([Tags.USER_ID]) # we will use sessions as a query ID column\n                                                     # in XGBoost parlance this a way of grouping together for training\n                                                     # when training with LGBM we had to calculate session lengths, but here the model does all the work for us!","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:07:59.866528Z","iopub.execute_input":"2022-11-28T04:07:59.868321Z","iopub.status.idle":"2022-11-28T04:08:02.172448Z","shell.execute_reply.started":"2022-11-28T04:07:59.868274Z","shell.execute_reply":"2022-11-28T04:08:02.171493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Having defined the preprocessing steps, we can now apply them to our data. The preprocessing is going to run on the GPU!","metadata":{}},{"cell_type":"code","source":"wf = Workflow(feature_cols + target + qid_column)\ntrain_processed = wf.fit_transform(train_ds)","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:08:02.173972Z","iopub.execute_input":"2022-11-28T04:08:02.174367Z","iopub.status.idle":"2022-11-28T04:08:02.235884Z","shell.execute_reply.started":"2022-11-28T04:08:02.174328Z","shell.execute_reply":"2022-11-28T04:08:02.235045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model training","metadata":{}},{"cell_type":"code","source":"ranker = XGBoost(train_processed.schema, objective='rank:pairwise')","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:08:02.237319Z","iopub.execute_input":"2022-11-28T04:08:02.237680Z","iopub.status.idle":"2022-11-28T04:08:02.242565Z","shell.execute_reply.started":"2022-11-28T04:08:02.237645Z","shell.execute_reply":"2022-11-28T04:08:02.241670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `Distributed` context manager will start a dask cudf cluster of us. A Dask cluster will be able to better manage memory usage for us. Normally, setting it up would be quite tedious -- here, we get all the benefits with a single line of Python code!","metadata":{}},{"cell_type":"code","source":"# version mismatch doesn't result in a loss of functionality here for us\n# it stems from the versions of libraries that the Kaggle vm comes preinstalled with\n\nwith Distributed():\n    ranker.fit(train_processed)","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:08:02.244079Z","iopub.execute_input":"2022-11-28T04:08:02.244759Z","iopub.status.idle":"2022-11-28T04:09:21.884227Z","shell.execute_reply.started":"2022-11-28T04:08:02.244725Z","shell.execute_reply":"2022-11-28T04:09:21.883060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have now trained our model! Let's predict on test!","metadata":{}},{"cell_type":"markdown","source":"# Predict on test data","metadata":{}},{"cell_type":"markdown","source":"Let's load our test set, process it and predict on it.","metadata":{}},{"cell_type":"code","source":"test = pl.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')\ntest = apply(test, pipeline)\ntest_ds = Dataset(test.to_pandas())\n\nwf = wf.remove_inputs(['gt']) # we don't have ground truth information in test!\n\ntest_ds_transformed = wf.transform(test_ds)","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:09:21.889747Z","iopub.execute_input":"2022-11-28T04:09:21.890152Z","iopub.status.idle":"2022-11-28T04:09:31.400720Z","shell.execute_reply.started":"2022-11-28T04:09:21.890104Z","shell.execute_reply":"2022-11-28T04:09:31.399733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's output the predictions","metadata":{}},{"cell_type":"code","source":"test_preds = ranker.booster.predict(xgb.DMatrix(test_ds_transformed.compute()))","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:09:31.401973Z","iopub.execute_input":"2022-11-28T04:09:31.404540Z","iopub.status.idle":"2022-11-28T04:09:33.874448Z","shell.execute_reply.started":"2022-11-28T04:09:31.404507Z","shell.execute_reply":"2022-11-28T04:09:33.873266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create submission","metadata":{}},{"cell_type":"code","source":"test = test.with_columns(pl.Series(name='score', values=test_preds))\ntest_predictions = test.sort(['session', 'score'], reverse=True).groupby('session').agg([\n    pl.col('aid').limit(20).list()\n])","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:09:33.881890Z","iopub.execute_input":"2022-11-28T04:09:33.882210Z","iopub.status.idle":"2022-11-28T04:09:35.157756Z","shell.execute_reply.started":"2022-11-28T04:09:33.882180Z","shell.execute_reply":"2022-11-28T04:09:35.156744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"session_types = []\nlabels = []\n\nfor session, preds in zip(test_predictions['session'].to_numpy(), test_predictions['aid'].to_numpy()):\n    l = ' '.join(str(p) for p in preds)\n    for session_type in ['clicks', 'carts', 'orders']:\n        labels.append(l)\n        session_types.append(f'{session}_{session_type}')","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:09:35.159845Z","iopub.execute_input":"2022-11-28T04:09:35.160472Z","iopub.status.idle":"2022-11-28T04:09:44.757392Z","shell.execute_reply.started":"2022-11-28T04:09:35.160432Z","shell.execute_reply":"2022-11-28T04:09:44.755681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pl.DataFrame({'session_type': session_types, 'labels': labels})\nsubmission.write_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-28T04:09:44.758738Z","iopub.execute_input":"2022-11-28T04:09:44.759395Z","iopub.status.idle":"2022-11-28T04:09:47.602908Z","shell.execute_reply.started":"2022-11-28T04:09:44.759344Z","shell.execute_reply":"2022-11-28T04:09:47.601399Z"},"trusted":true},"execution_count":null,"outputs":[]}]}