{
  "id": 386497,
  "title": "3rd place - Benny's part - XGB Reranker with Transformers+GRU input - LB0.601",
  "url": "/competitions/otto-recommender-system/discussion/386497",
  "author_name": "",
  "post_date": "2023-02-13T11:35:41.002498900Z",
  "votes": 18,
  "comment_count": 2,
  "views": 0,
  "content": "<p>thanks to Otto and Kaggle to organize the competition. The competition and dataset were well organized without any leakage or shake up. I competed in previous RecSys competition, but this was my first Kaggle competition. It was great to work with <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> . You can read <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013\" target=\"_blank\">Chris's solution</a> and <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382975\" target=\"_blank\">Theo's solution</a>. I learned a lot during this competition, again.</p>\n<h1>Summary</h1>\n<p>Each of us developed their own model and we ensembled our final models by adding the ranks. I will focus on my model. Similar to many solutions, I used a tree-based reranker (XGBoost). I will focus my write up about:</p>\n<ul>\n<li>Short Description of Reranker</li>\n<li>Data Split</li>\n<li>Transformer4Rec</li>\n<li>GRU</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Local CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Chris Baseline</td>\n<td>0.567</td>\n<td>0.575</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 4th week truncated)</td>\n<td>0.576</td>\n<td>0.583</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 20% 4th week truncated</td>\n<td>0.585</td>\n<td>0.591</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 20% 4th week truncated + activity history</td>\n<td>0.587</td>\n<td>0.594</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 100% 4th week truncated + activity history</td>\n<td>0.592</td>\n<td>0.598</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 100% 4th week truncated + activity history + Word2Vec</td>\n<td>0.593</td>\n<td>0.599</td>\n</tr>\n<tr>\n<td>(I will explain the models in the post)</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h1>Reranker - LB 0.601</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F582742%2F3ed3ae4f60217e6622235469056e439a%2Fmodel.png?generation=1676287838345268&amp;alt=media\" alt=\"\"></p>\n<p>My final model is a similar pipeline as <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">proposed by Chris</a> and described by many top solution. I will briefly summarize my solution:</p>\n<ol>\n<li>Feature Engineering:</li>\n</ol>\n<ul>\n<li>Generating CoVisitation Matrix (7 different ones)</li>\n<li>After merging teams, I added Chris’ CoVisitation Matrix (only 10)</li>\n<li>Training a Transformer4Rec model to predict scores (next item prediction)</li>\n<li>Training a GRU model to predict scores (next item prediction)</li>\n<li>Train a <a href=\"https://www.kaggle.com/code/duuuscha/train-submit-word2vec-optimized-hparams\" target=\"_blank\">Word2Vec model</a> generating embeddings</li>\n<li>Session Features - length of session, day, time, etc.</li>\n<li>Candidate Features - how many clicks, addtocarts, orders on same day, previous 7 days, previous 13 days, etc.</li>\n</ul>\n<ol>\n<li>Generating Candidates:<br>\nI used a different candidate generation process pre target. I noticed increasing the number of candidates improved local CV and LB score, but the clicks dataset is too large and I haven’t enough time to refactor my pipeline to support more candidates for clicks.</li>\n</ol>\n<p>Clicks: Union of Top80 Candidates per session for each of ~6 different CoVisitation Matrix + items which were in the session history<br>\nCarts+Orders: Union of Top80 Candidates per session all CoVisitation Matrix + session history and keep only candidates, which were generated from at least 2 different sources</p>\n<p>Note: Increasing the number of candidates for Carts+Clicks improved my score from 0.599 to 0.601</p>\n<ol>\n<li>Add Features/Scores for session x candidate pairs:<br>\nAs <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382975\" target=\"_blank\">Theo wrote</a> - how to generate unique scores?<br>\nI used different combinations between - max, mean, weighted mean by rank or linear weights, sum, count to generate the final score. For example, I calculate the cosine similarity between all viewed items in a session and the candidate and then groupby session, candidate with aggregating the cosine similarity with max, mean, weighted mean.</li>\n</ol>\n<p>I had around ~160 features and around avg of 100-150 candidates per session.</p>\n<ol>\n<li>Training a XGBoost model</li>\n</ol>\n<p>Some additional notes:</p>\n<ul>\n<li>I added hierarchical predictions. I used all sessions that have a click target and NO carts nor order targets to train a XGBoost model and add the click probability as an input to my carts and orders model. Similarly I added carts probability to the orders model. I used the hierarchical prediction only for orders.</li>\n<li>Although I had already added Transformer4Rec scores, adding Word2Vec embeddings improved my LB score</li>\n<li>I tried NN reranker but it didnt improve my local CV. I should have tried to ensemble with XGBoost reranker.</li>\n<li>I tried different feature selection methods. However, removing features decreased my local CV score.</li>\n<li>The pipeline was implemented with <a href=\"https://rapids.ai/\" target=\"_blank\">RAPIDs cuDF</a> using a single GPU with 32GB memory. I ran some code in parallel by using multiple workers with each 1x GPU. </li>\n</ul>\n<h1>Data Split</h1>\n<p>The main question for me was: how to split the data? Which data should be used for feature engineering and which data should be used to train the reranker?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F582742%2F2d46eb5384967948c6a7068306f31a85%2Fdatasplit.png?generation=1676288119588241&amp;alt=media\" alt=\"\"></p>\n<h3>XGBoost Reranker: trained on 4th week truncated - LB 0.584</h3>\n<p>I used Radek's local CV strategy and run the host's dataset myself to split the training dataset into 3 full weeks and 4th week of truncated sessions. I used week 1-3, truncated 4th week (without labels) and test dataset to engineer features (e.g. co-visitation matrix). I trained the XGBoost reranker on the 4th week truncated sessions, where the session data are used to generate a X dataframe and the labels are the targets.</p>\n<h3>XGBoost Reranker: trained on 20% 4th week truncated - LB 0.591</h3>\n<p>I noticed that I lose too much information by truncating all sessions of the 4th week. One strategy is to generate features by running the pipeline for train dataset and submission dataset with different inputs. Train dataset uses the truncated versions and submission dataset uses the full train dataset. In previous RecSys competition, I had bad experience because the feature distribution can shift. I wanted that my train and submission dataset uses features generated from the same dataset (e.g. same co-visitation matrix). </p>\n<p>I decided to split the 4th week into 5 folds by session ID. For one fold, I will truncate the sessions and the remaining 4 folds are based on the original train dataset. Therefore, I have more data for generating the features and enough data for training my XGBoost model.</p>\n<h3>XGBoost Reranker: trained on 100% 4th week truncated + activity history - LB 0.598</h3>\n<p>I ran the pipeline of <em>XGBoost Reranker: trained on 20% 4th week truncated</em> for each fold - having 100% of the 4th week as truncated sessions, which can be used for training the Reranker. Each fold is generated with full week 1-3, 80% of full 4th week, 20% of truncated 4th week and test sessions.  </p>\n<p>It significantly improved my LB score. Unfortunately, running experiments became really slow. Running the pipeline for each fold was complex - managing the files, executing the pipeline and computation time.</p>\n<h1>Transformer4Rec</h1>\n<p>I started with transformer-based architectures to predict the next item using <a href=\"https://github.com/NVIDIA-Merlin/Transformers4Rec/\" target=\"_blank\">Transformer4Rec</a>. It provided great results without a complicated pipeline of feature engineering, candidate generation and reranker. Some observations:</p>\n<ul>\n<li>Increasing embedding width improved local CV (maximum possible was 256)</li>\n<li>Using Dot-Product (weight-tying) as the output layer of the full item catalog was inefficient. I implemented a negative sampling strategy, which significantly reduced training time. I will create a PR for Transformer4Rec</li>\n<li>I optimized ~25 hyperparameters</li>\n</ul>\n<p>Unfortunately, at the time, I truncated the full 4th week, removing a lot of data. I did not split sessions into subsessions.</p>\n<h1>GRU</h1>\n<p>After I added the Word2Vec embeddings to my reranker, I wanted to use a GRU model to learn an aggregation function of embeddings -&gt; final score. I trained the GRU model with pre-trained word2vec embeddings and random initialized embeddings. I created sub datasets by excluding the last interactions of a session as the target and keeping the remaining interactions as input data. I repeated the process multiple times. I used negative sampling strategies to reduce training time.</p>",
  "messages": [
    {
      "id": "2142205",
      "postDate": "02/13/2023 11:35:41",
      "content": "<p>thanks to Otto and Kaggle to organize the competition. The competition and dataset were well organized without any leakage or shake up. I competed in previous RecSys competition, but this was my first Kaggle competition. It was great to work with <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> . You can read <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013\" target=\"_blank\">Chris's solution</a> and <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382975\" target=\"_blank\">Theo's solution</a>. I learned a lot during this competition, again.</p>\n<h1>Summary</h1>\n<p>Each of us developed their own model and we ensembled our final models by adding the ranks. I will focus on my model. Similar to many solutions, I used a tree-based reranker (XGBoost). I will focus my write up about:</p>\n<ul>\n<li>Short Description of Reranker</li>\n<li>Data Split</li>\n<li>Transformer4Rec</li>\n<li>GRU</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Local CV</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Chris Baseline</td>\n<td>0.567</td>\n<td>0.575</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 4th week truncated)</td>\n<td>0.576</td>\n<td>0.583</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 20% 4th week truncated</td>\n<td>0.585</td>\n<td>0.591</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 20% 4th week truncated + activity history</td>\n<td>0.587</td>\n<td>0.594</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 100% 4th week truncated + activity history</td>\n<td>0.592</td>\n<td>0.598</td>\n</tr>\n<tr>\n<td>XGBoost Reranker: trained on 100% 4th week truncated + activity history + Word2Vec</td>\n<td>0.593</td>\n<td>0.599</td>\n</tr>\n<tr>\n<td>(I will explain the models in the post)</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h1>Reranker - LB 0.601</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F582742%2F3ed3ae4f60217e6622235469056e439a%2Fmodel.png?generation=1676287838345268&amp;alt=media\" alt=\"\"></p>\n<p>My final model is a similar pipeline as <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">proposed by Chris</a> and described by many top solution. I will briefly summarize my solution:</p>\n<ol>\n<li>Feature Engineering:</li>\n</ol>\n<ul>\n<li>Generating CoVisitation Matrix (7 different ones)</li>\n<li>After merging teams, I added Chris’ CoVisitation Matrix (only 10)</li>\n<li>Training a Transformer4Rec model to predict scores (next item prediction)</li>\n<li>Training a GRU model to predict scores (next item prediction)</li>\n<li>Train a <a href=\"https://www.kaggle.com/code/duuuscha/train-submit-word2vec-optimized-hparams\" target=\"_blank\">Word2Vec model</a> generating embeddings</li>\n<li>Session Features - length of session, day, time, etc.</li>\n<li>Candidate Features - how many clicks, addtocarts, orders on same day, previous 7 days, previous 13 days, etc.</li>\n</ul>\n<ol>\n<li>Generating Candidates:<br>\nI used a different candidate generation process pre target. I noticed increasing the number of candidates improved local CV and LB score, but the clicks dataset is too large and I haven’t enough time to refactor my pipeline to support more candidates for clicks.</li>\n</ol>\n<p>Clicks: Union of Top80 Candidates per session for each of ~6 different CoVisitation Matrix + items which were in the session history<br>\nCarts+Orders: Union of Top80 Candidates per session all CoVisitation Matrix + session history and keep only candidates, which were generated from at least 2 different sources</p>\n<p>Note: Increasing the number of candidates for Carts+Clicks improved my score from 0.599 to 0.601</p>\n<ol>\n<li>Add Features/Scores for session x candidate pairs:<br>\nAs <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382975\" target=\"_blank\">Theo wrote</a> - how to generate unique scores?<br>\nI used different combinations between - max, mean, weighted mean by rank or linear weights, sum, count to generate the final score. For example, I calculate the cosine similarity between all viewed items in a session and the candidate and then groupby session, candidate with aggregating the cosine similarity with max, mean, weighted mean.</li>\n</ol>\n<p>I had around ~160 features and around avg of 100-150 candidates per session.</p>\n<ol>\n<li>Training a XGBoost model</li>\n</ol>\n<p>Some additional notes:</p>\n<ul>\n<li>I added hierarchical predictions. I used all sessions that have a click target and NO carts nor order targets to train a XGBoost model and add the click probability as an input to my carts and orders model. Similarly I added carts probability to the orders model. I used the hierarchical prediction only for orders.</li>\n<li>Although I had already added Transformer4Rec scores, adding Word2Vec embeddings improved my LB score</li>\n<li>I tried NN reranker but it didnt improve my local CV. I should have tried to ensemble with XGBoost reranker.</li>\n<li>I tried different feature selection methods. However, removing features decreased my local CV score.</li>\n<li>The pipeline was implemented with <a href=\"https://rapids.ai/\" target=\"_blank\">RAPIDs cuDF</a> using a single GPU with 32GB memory. I ran some code in parallel by using multiple workers with each 1x GPU. </li>\n</ul>\n<h1>Data Split</h1>\n<p>The main question for me was: how to split the data? Which data should be used for feature engineering and which data should be used to train the reranker?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F582742%2F2d46eb5384967948c6a7068306f31a85%2Fdatasplit.png?generation=1676288119588241&amp;alt=media\" alt=\"\"></p>\n<h3>XGBoost Reranker: trained on 4th week truncated - LB 0.584</h3>\n<p>I used Radek's local CV strategy and run the host's dataset myself to split the training dataset into 3 full weeks and 4th week of truncated sessions. I used week 1-3, truncated 4th week (without labels) and test dataset to engineer features (e.g. co-visitation matrix). I trained the XGBoost reranker on the 4th week truncated sessions, where the session data are used to generate a X dataframe and the labels are the targets.</p>\n<h3>XGBoost Reranker: trained on 20% 4th week truncated - LB 0.591</h3>\n<p>I noticed that I lose too much information by truncating all sessions of the 4th week. One strategy is to generate features by running the pipeline for train dataset and submission dataset with different inputs. Train dataset uses the truncated versions and submission dataset uses the full train dataset. In previous RecSys competition, I had bad experience because the feature distribution can shift. I wanted that my train and submission dataset uses features generated from the same dataset (e.g. same co-visitation matrix). </p>\n<p>I decided to split the 4th week into 5 folds by session ID. For one fold, I will truncate the sessions and the remaining 4 folds are based on the original train dataset. Therefore, I have more data for generating the features and enough data for training my XGBoost model.</p>\n<h3>XGBoost Reranker: trained on 100% 4th week truncated + activity history - LB 0.598</h3>\n<p>I ran the pipeline of <em>XGBoost Reranker: trained on 20% 4th week truncated</em> for each fold - having 100% of the 4th week as truncated sessions, which can be used for training the Reranker. Each fold is generated with full week 1-3, 80% of full 4th week, 20% of truncated 4th week and test sessions.  </p>\n<p>It significantly improved my LB score. Unfortunately, running experiments became really slow. Running the pipeline for each fold was complex - managing the files, executing the pipeline and computation time.</p>\n<h1>Transformer4Rec</h1>\n<p>I started with transformer-based architectures to predict the next item using <a href=\"https://github.com/NVIDIA-Merlin/Transformers4Rec/\" target=\"_blank\">Transformer4Rec</a>. It provided great results without a complicated pipeline of feature engineering, candidate generation and reranker. Some observations:</p>\n<ul>\n<li>Increasing embedding width improved local CV (maximum possible was 256)</li>\n<li>Using Dot-Product (weight-tying) as the output layer of the full item catalog was inefficient. I implemented a negative sampling strategy, which significantly reduced training time. I will create a PR for Transformer4Rec</li>\n<li>I optimized ~25 hyperparameters</li>\n</ul>\n<p>Unfortunately, at the time, I truncated the full 4th week, removing a lot of data. I did not split sessions into subsessions.</p>\n<h1>GRU</h1>\n<p>After I added the Word2Vec embeddings to my reranker, I wanted to use a GRU model to learn an aggregation function of embeddings -&gt; final score. I trained the GRU model with pre-trained word2vec embeddings and random initialized embeddings. I created sub datasets by excluding the last interactions of a session as the target and keeping the remaining interactions as input data. I repeated the process multiple times. I used negative sampling strategies to reduce training time.</p>",
      "rawMarkdown": "thanks to Otto and Kaggle to organize the competition. The competition and dataset were well organized without any leakage or shake up. I competed in previous RecSys competition, but this was my first Kaggle competition. It was great to work with @cdeotte @titericz and @theoviel . You can read [Chris's solution](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013) and [Theo's solution](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382975). I learned a lot during this competition, again.\n\n# Summary\n\nEach of us developed their own model and we ensembled our final models by adding the ranks. I will focus on my model. Similar to many solutions, I used a tree-based reranker (XGBoost). I will focus my write up about:\n- Short Description of Reranker\n- Data Split\n- Transformer4Rec\n- GRU\n\n| Model | Local CV | Public LB |\n| --- | --- | --- |\n| Chris Baseline | 0.567 | 0.575 |\n| XGBoost Reranker: trained on 4th week truncated) | 0.576 | 0.583 |\n| XGBoost Reranker: trained on 20% 4th week truncated | 0.585 | 0.591 |\n| XGBoost Reranker: trained on 20% 4th week truncated + activity history | 0.587 | 0.594 |\n| XGBoost Reranker: trained on 100% 4th week truncated + activity history | 0.592 | 0.598 |\n| XGBoost Reranker: trained on 100% 4th week truncated + activity history + Word2Vec | 0.593 | 0.599 |\n(I will explain the models in the post)\n\n# Reranker - LB 0.601\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F582742%2F3ed3ae4f60217e6622235469056e439a%2Fmodel.png?generation=1676287838345268&alt=media)\n\nMy final model is a similar pipeline as [proposed by Chris](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210) and described by many top solution. I will briefly summarize my solution:\n\n1. Feature Engineering:\n- Generating CoVisitation Matrix (7 different ones)\n- After merging teams, I added Chris’ CoVisitation Matrix (only 10)\n- Training a Transformer4Rec model to predict scores (next item prediction)\n- Training a GRU model to predict scores (next item prediction)\n- Train a [Word2Vec model](https://www.kaggle.com/code/duuuscha/train-submit-word2vec-optimized-hparams) generating embeddings\n- Session Features - length of session, day, time, etc.\n- Candidate Features - how many clicks, addtocarts, orders on same day, previous 7 days, previous 13 days, etc.\n\n2. Generating Candidates:\nI used a different candidate generation process pre target. I noticed increasing the number of candidates improved local CV and LB score, but the clicks dataset is too large and I haven’t enough time to refactor my pipeline to support more candidates for clicks.\n\nClicks: Union of Top80 Candidates per session for each of ~6 different CoVisitation Matrix + items which were in the session history\nCarts+Orders: Union of Top80 Candidates per session all CoVisitation Matrix + session history and keep only candidates, which were generated from at least 2 different sources\n\nNote: Increasing the number of candidates for Carts+Clicks improved my score from 0.599 to 0.601\n\n3. Add Features/Scores for session x candidate pairs:\nAs [Theo wrote](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382975) - how to generate unique scores?\nI used different combinations between - max, mean, weighted mean by rank or linear weights, sum, count to generate the final score. For example, I calculate the cosine similarity between all viewed items in a session and the candidate and then groupby session, candidate with aggregating the cosine similarity with max, mean, weighted mean.\n\nI had around ~160 features and around avg of 100-150 candidates per session.\n\n4. Training a XGBoost model\n\nSome additional notes:\n- I added hierarchical predictions. I used all sessions that have a click target and NO carts nor order targets to train a XGBoost model and add the click probability as an input to my carts and orders model. Similarly I added carts probability to the orders model. I used the hierarchical prediction only for orders.\n- Although I had already added Transformer4Rec scores, adding Word2Vec embeddings improved my LB score\n- I tried NN reranker but it didnt improve my local CV. I should have tried to ensemble with XGBoost reranker.\n- I tried different feature selection methods. However, removing features decreased my local CV score.\n- The pipeline was implemented with [RAPIDs cuDF](https://rapids.ai/) using a single GPU with 32GB memory. I ran some code in parallel by using multiple workers with each 1x GPU. \n\n# Data Split\nThe main question for me was: how to split the data? Which data should be used for feature engineering and which data should be used to train the reranker?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F582742%2F2d46eb5384967948c6a7068306f31a85%2Fdatasplit.png?generation=1676288119588241&alt=media)\n\n### XGBoost Reranker: trained on 4th week truncated - LB 0.584\nI used Radek's local CV strategy and run the host's dataset myself to split the training dataset into 3 full weeks and 4th week of truncated sessions. I used week 1-3, truncated 4th week (without labels) and test dataset to engineer features (e.g. co-visitation matrix). I trained the XGBoost reranker on the 4th week truncated sessions, where the session data are used to generate a X dataframe and the labels are the targets.\n\n### XGBoost Reranker: trained on 20% 4th week truncated - LB 0.591\nI noticed that I lose too much information by truncating all sessions of the 4th week. One strategy is to generate features by running the pipeline for train dataset and submission dataset with different inputs. Train dataset uses the truncated versions and submission dataset uses the full train dataset. In previous RecSys competition, I had bad experience because the feature distribution can shift. I wanted that my train and submission dataset uses features generated from the same dataset (e.g. same co-visitation matrix). \n\nI decided to split the 4th week into 5 folds by session ID. For one fold, I will truncate the sessions and the remaining 4 folds are based on the original train dataset. Therefore, I have more data for generating the features and enough data for training my XGBoost model.\n\n### XGBoost Reranker: trained on 100% 4th week truncated + activity history - LB 0.598\nI ran the pipeline of *XGBoost Reranker: trained on 20% 4th week truncated* for each fold - having 100% of the 4th week as truncated sessions, which can be used for training the Reranker. Each fold is generated with full week 1-3, 80% of full 4th week, 20% of truncated 4th week and test sessions.  \n\nIt significantly improved my LB score. Unfortunately, running experiments became really slow. Running the pipeline for each fold was complex - managing the files, executing the pipeline and computation time.\n\n# Transformer4Rec \n\nI started with transformer-based architectures to predict the next item using [Transformer4Rec](https://github.com/NVIDIA-Merlin/Transformers4Rec/). It provided great results without a complicated pipeline of feature engineering, candidate generation and reranker. Some observations:\n- Increasing embedding width improved local CV (maximum possible was 256)\n- Using Dot-Product (weight-tying) as the output layer of the full item catalog was inefficient. I implemented a negative sampling strategy, which significantly reduced training time. I will create a PR for Transformer4Rec\n- I optimized ~25 hyperparameters\n\nUnfortunately, at the time, I truncated the full 4th week, removing a lot of data. I did not split sessions into subsessions.\n\n# GRU\n\nAfter I added the Word2Vec embeddings to my reranker, I wanted to use a GRU model to learn an aggregation function of embeddings -> final score. I trained the GRU model with pre-trained word2vec embeddings and random initialized embeddings. I created sub datasets by excluding the last interactions of a session as the target and keeping the remaining interactions as input data. I repeated the process multiple times. I used negative sampling strategies to reduce training time.",
      "votes": null
    },
    {
      "id": "2145268",
      "postDate": "02/14/2023 21:53:46",
      "content": "<p>I uploaded my code to github: <a href=\"https://github.com/bschifferer/Kaggle-Otto-Comp\" target=\"_blank\">https://github.com/bschifferer/Kaggle-Otto-Comp</a></p>",
      "rawMarkdown": "I uploaded my code to github: https://github.com/bschifferer/Kaggle-Otto-Comp",
      "votes": null
    },
    {
      "id": "2243399",
      "postDate": "05/02/2023 21:55:48",
      "content": "<p>Hi Benny, thanks so much for sharing the amazing code! I was learning your code and had a question:<br>\nIn this file <code>01e_FE_Transformer/02_T4R_Train-v2.ipynb</code>, <code>from custom_t4r import *</code>. Whereas I didn't see where the custom_t4r is, could you help with this? Appreciate your reply!</p>",
      "rawMarkdown": "Hi Benny, thanks so much for sharing the amazing code! I was learning your code and had a question:\nIn this file `01e_FE_Transformer/02_T4R_Train-v2.ipynb`, `from custom_t4r import *`. Whereas I didn't see where the custom_t4r is, could you help with this? Appreciate your reply!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2145268,
      "author_name": "benediktschifferer",
      "author_url": "",
      "post_date": "02/14/2023 21:53:46",
      "content": "<p>I uploaded my code to github: <a href=\"https://github.com/bschifferer/Kaggle-Otto-Comp\" target=\"_blank\">https://github.com/bschifferer/Kaggle-Otto-Comp</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2243399,
      "author_name": "sherrypomi",
      "author_url": "",
      "post_date": "05/02/2023 21:55:48",
      "content": "<p>Hi Benny, thanks so much for sharing the amazing code! I was learning your code and had a question:<br>\nIn this file <code>01e_FE_Transformer/02_T4R_Train-v2.ipynb</code>, <code>from custom_t4r import *</code>. Whereas I didn't see where the custom_t4r is, could you help with this? Appreciate your reply!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2142205": "thanks to Otto and Kaggle to organize the competition. The competition and dataset were well organized without any leakage or shake up. I competed in previous RecSys competition, but this was my first Kaggle competition. It was great to work with @cdeotte @titericz and @theoviel . You can read [Chris's solution](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013) and [Theo's solution](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382975). I learned a lot during this competition, again.\n\n# Summary\n\nEach of us developed their own model and we ensembled our final models by adding the ranks. I will focus on my model. Similar to many solutions, I used a tree-based reranker (XGBoost). I will focus my write up about:\n- Short Description of Reranker\n- Data Split\n- Transformer4Rec\n- GRU\n\n| Model | Local CV | Public LB |\n| --- | --- | --- |\n| Chris Baseline | 0.567 | 0.575 |\n| XGBoost Reranker: trained on 4th week truncated) | 0.576 | 0.583 |\n| XGBoost Reranker: trained on 20% 4th week truncated | 0.585 | 0.591 |\n| XGBoost Reranker: trained on 20% 4th week truncated + activity history | 0.587 | 0.594 |\n| XGBoost Reranker: trained on 100% 4th week truncated + activity history | 0.592 | 0.598 |\n| XGBoost Reranker: trained on 100% 4th week truncated + activity history + Word2Vec | 0.593 | 0.599 |\n(I will explain the models in the post)\n\n# Reranker - LB 0.601\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F582742%2F3ed3ae4f60217e6622235469056e439a%2Fmodel.png?generation=1676287838345268&alt=media)\n\nMy final model is a similar pipeline as [proposed by Chris](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210) and described by many top solution. I will briefly summarize my solution:\n\n1. Feature Engineering:\n- Generating CoVisitation Matrix (7 different ones)\n- After merging teams, I added Chris’ CoVisitation Matrix (only 10)\n- Training a Transformer4Rec model to predict scores (next item prediction)\n- Training a GRU model to predict scores (next item prediction)\n- Train a [Word2Vec model](https://www.kaggle.com/code/duuuscha/train-submit-word2vec-optimized-hparams) generating embeddings\n- Session Features - length of session, day, time, etc.\n- Candidate Features - how many clicks, addtocarts, orders on same day, previous 7 days, previous 13 days, etc.\n\n2. Generating Candidates:\nI used a different candidate generation process pre target. I noticed increasing the number of candidates improved local CV and LB score, but the clicks dataset is too large and I haven’t enough time to refactor my pipeline to support more candidates for clicks.\n\nClicks: Union of Top80 Candidates per session for each of ~6 different CoVisitation Matrix + items which were in the session history\nCarts+Orders: Union of Top80 Candidates per session all CoVisitation Matrix + session history and keep only candidates, which were generated from at least 2 different sources\n\nNote: Increasing the number of candidates for Carts+Clicks improved my score from 0.599 to 0.601\n\n3. Add Features/Scores for session x candidate pairs:\nAs [Theo wrote](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382975) - how to generate unique scores?\nI used different combinations between - max, mean, weighted mean by rank or linear weights, sum, count to generate the final score. For example, I calculate the cosine similarity between all viewed items in a session and the candidate and then groupby session, candidate with aggregating the cosine similarity with max, mean, weighted mean.\n\nI had around ~160 features and around avg of 100-150 candidates per session.\n\n4. Training a XGBoost model\n\nSome additional notes:\n- I added hierarchical predictions. I used all sessions that have a click target and NO carts nor order targets to train a XGBoost model and add the click probability as an input to my carts and orders model. Similarly I added carts probability to the orders model. I used the hierarchical prediction only for orders.\n- Although I had already added Transformer4Rec scores, adding Word2Vec embeddings improved my LB score\n- I tried NN reranker but it didnt improve my local CV. I should have tried to ensemble with XGBoost reranker.\n- I tried different feature selection methods. However, removing features decreased my local CV score.\n- The pipeline was implemented with [RAPIDs cuDF](https://rapids.ai/) using a single GPU with 32GB memory. I ran some code in parallel by using multiple workers with each 1x GPU. \n\n# Data Split\nThe main question for me was: how to split the data? Which data should be used for feature engineering and which data should be used to train the reranker?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F582742%2F2d46eb5384967948c6a7068306f31a85%2Fdatasplit.png?generation=1676288119588241&alt=media)\n\n### XGBoost Reranker: trained on 4th week truncated - LB 0.584\nI used Radek's local CV strategy and run the host's dataset myself to split the training dataset into 3 full weeks and 4th week of truncated sessions. I used week 1-3, truncated 4th week (without labels) and test dataset to engineer features (e.g. co-visitation matrix). I trained the XGBoost reranker on the 4th week truncated sessions, where the session data are used to generate a X dataframe and the labels are the targets.\n\n### XGBoost Reranker: trained on 20% 4th week truncated - LB 0.591\nI noticed that I lose too much information by truncating all sessions of the 4th week. One strategy is to generate features by running the pipeline for train dataset and submission dataset with different inputs. Train dataset uses the truncated versions and submission dataset uses the full train dataset. In previous RecSys competition, I had bad experience because the feature distribution can shift. I wanted that my train and submission dataset uses features generated from the same dataset (e.g. same co-visitation matrix). \n\nI decided to split the 4th week into 5 folds by session ID. For one fold, I will truncate the sessions and the remaining 4 folds are based on the original train dataset. Therefore, I have more data for generating the features and enough data for training my XGBoost model.\n\n### XGBoost Reranker: trained on 100% 4th week truncated + activity history - LB 0.598\nI ran the pipeline of *XGBoost Reranker: trained on 20% 4th week truncated* for each fold - having 100% of the 4th week as truncated sessions, which can be used for training the Reranker. Each fold is generated with full week 1-3, 80% of full 4th week, 20% of truncated 4th week and test sessions.  \n\nIt significantly improved my LB score. Unfortunately, running experiments became really slow. Running the pipeline for each fold was complex - managing the files, executing the pipeline and computation time.\n\n# Transformer4Rec \n\nI started with transformer-based architectures to predict the next item using [Transformer4Rec](https://github.com/NVIDIA-Merlin/Transformers4Rec/). It provided great results without a complicated pipeline of feature engineering, candidate generation and reranker. Some observations:\n- Increasing embedding width improved local CV (maximum possible was 256)\n- Using Dot-Product (weight-tying) as the output layer of the full item catalog was inefficient. I implemented a negative sampling strategy, which significantly reduced training time. I will create a PR for Transformer4Rec\n- I optimized ~25 hyperparameters\n\nUnfortunately, at the time, I truncated the full 4th week, removing a lot of data. I did not split sessions into subsessions.\n\n# GRU\n\nAfter I added the Word2Vec embeddings to my reranker, I wanted to use a GRU model to learn an aggregation function of embeddings -> final score. I trained the GRU model with pre-trained word2vec embeddings and random initialized embeddings. I created sub datasets by excluding the last interactions of a session as the target and keeping the remaining interactions as input data. I repeated the process multiple times. I used negative sampling strategies to reduce training time.",
    "2145268": "I uploaded my code to github: https://github.com/bschifferer/Kaggle-Otto-Comp",
    "2243399": "Hi Benny, thanks so much for sharing the amazing code! I was learning your code and had a question:\nIn this file `01e_FE_Transformer/02_T4R_Train-v2.ipynb`, `from custom_t4r import *`. Whereas I didn't see where the custom_t4r is, could you help with this? Appreciate your reply!"
  },
  "source": "meta"
}