{
  "id": 382924,
  "title": "20th place solution (Transformer inside)",
  "url": "/competitions/otto-recommender-system/writeups/mikhail-kamenshchikov-20th-place-solution-transfor",
  "author_name": "",
  "post_date": "2023-02-08T09:48:49.647Z",
  "votes": 29,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, I want to thank OTTO for organising great competition. It's was very challenging and clean :) <br>\nSecondly, I want to thank all people who share their thoughts and ideas in kernels or forums. It helps a lot.<br>\nI will try to put my code on GitHub in a few days, need to clean up some deadline madness there.</p>\n<h2>TL;DR</h2>\n<ul>\n<li>use 3 types of co-occurence matrices (somewhat similar ideas to public notebooks, but different implementation and details)</li>\n<li>Bert MLM</li>\n<li>Matrix Factorization</li>\n<li>Catboost + PairLogitPairwise</li>\n</ul>\n<h3>LB Progression (public)</h3>\n<ul>\n<li>0.576 - history + 60 candidates all-to-all co-occurence matrix + lightgbm ranker</li>\n<li>0.579 - same, but 200 candidates</li>\n<li>0.583 - added a bunch of item features (conversions, popularity etc.)</li>\n<li><strong>0.593 - optimized co-occurence matrix</strong> (switched to sessions, added time weighing)</li>\n<li>0.595 - add buy2buy features and Transfomer candidates, switch to catboost</li>\n<li>0.597 - add buy2buy and type weighted co-occurence candidates, use more data for training (16/32 chunks)</li>\n<li>0.598 - use full data for training</li>\n<li>0.599 - use different candidates configs for clicks/buys, add MF candidates and scores</li>\n<li>0.600 - use different transformers</li>\n<li>0.601  - use x1.5 more candidates from each source for carts/orders</li>\n</ul>\n<h2>Candidates retrieval</h2>\n<p>I use <strong>max-recall@200</strong> as a retrieval quality measure, but candidates from different sources can be more or less common with user history, so it seems more fair to outer join candidates with history to calculate max-recall. <strong>0.598 LB is achievable only with co-visitation candidates</strong></p>\n<p>I used different combinations of candidates for clicks and carts/orders.</p>\n<pre><code>clicks:\n    history_rank: \n    cooc_rank: \n    buy2buy_rank: \n    cooc_tw_rank: \n    mfc_rank: \n    transformer_rank: \n  carts/oders: \n    history_rank: \n    cooc_rank: \n    cooc_tw_rank: \n    buy2buy_rank: \n    mfc_rank: \n    transformer_rank: \n</code></pre>\n<h3>Co-visitation (all-to-all)</h3>\n<ul>\n<li>Use actual user’s “sessions” - consecutive series of events, if there’s a gap &gt; 900 seconds, it’s another session.</li>\n<li>Use exponential time weighing (more distant events are less significant)  0.99995^(abs(ts.x - ts.y))</li>\n<li>inverse rank weighing for user history events</li>\n</ul>\n<p>Hyperparameters like session gap, time base and rank weight function were optimized with optuna, so <strong>max-recall@200</strong> was like <strong>67.6</strong> for this method.</p>\n<h3>Co-visitation (type weighted)</h3>\n<ol>\n<li>Use 1 day gap with exponential time weighing (no “sessions”)</li>\n<li>10x weight for carts, 3x for orders</li>\n</ol>\n<h3>Co-visitation (buy2buy)</h3>\n<ol>\n<li>Use 2 weeks gap, exponential time weighing</li>\n<li>Use only carts and orders to calculate stats</li>\n</ol>\n<h3>Transformer (small BERT)</h3>\n<p>It’s was hugely inspired by the <a href=\"https://github.com/Chubasik/yacup_recsys_2022\" target=\"_blank\">winning solution</a> (by <a href=\"https://www.kaggle.com/chubasik\" target=\"_blank\">@chubasik</a>) of recent Yandex.Cup Recsys track (I took <a href=\"https://github.com/greenwolf-nsk/yandex-cup-2022-recsys\" target=\"_blank\">2nd place</a> there with classical 2-stage approach).</p>\n<p>The idea is to train Masked Language Model, and then predict the “fake” last masked item in user session. Also, I fed action types (click, order, cart) as token_type_ids (which is mainly used for context separation in NLP tasks).</p>\n<p>I trained MLM on train sessions, used 500k most carted items, <strong>max-recall@200</strong> was like <strong>0.66.</strong> This quality can be achieved with 3 epochs on full data, but it takes very long to train (7 hours on A100 GPU), so my experiments were very limited.</p>\n<p>Adding this source of candidates gave boost of 0.002 in local CV, and this model score was second most valuable feature for ranking model. However, LB change was less then 0.001, and I spend last two weeks figuring out what went wrong. Finally, I trained two different models (train_no_val + val and train + test), this probably helped a little, but CV-LB gap was still bigger than before.<br>\nThere's 3 epochs training from scratch, with metrics every 0.5 epochs<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172966%2F916debefa57949c714c09d06c2ac6f9e%2F2023-02-01%20%2021.49.24.png?generation=1675263027888603&amp;alt=media\" alt=\"3 epoch training (eval every 0.5 epochs)\"></p>\n<h3>Matrix Factorization</h3>\n<p>Matrix factorization on Pytorch with BPR-like loss and hard negative sampling. It achieves <strong>0.665 max-recall@200</strong>. There’s probably a huge space for improvement in weighing events by type and time and negative sampling strategies. Training time is <strong>7h</strong> in total for 20 epochs with AdamW optimizer. It almost made no difference to CV/LB score, but MF score was also one of the strongest features.</p>\n<h2>Features</h2>\n<p>Best model uses around 200 features.</p>\n<h3>User</h3>\n<ul>\n<li>counters by type, normalized counter, time-based features</li>\n<li>number of “sessions”, avg session length</li>\n<li>avg/min/max/std “popularity” of item in user history</li>\n</ul>\n<h3>Item</h3>\n<ul>\n<li>item popularity by type - counters and ranks</li>\n<li>tried some derivatives to detect “trending” items, but they didn’t work for me</li>\n<li>item click/cart/order conversion rates</li>\n</ul>\n<h3>User-Item</h3>\n<ul>\n<li>interaction stats with item (number of clicks/carts, last timestamp)</li>\n<li>all features from co-visit matrices and statistics (mean/min/max/std for score/rank/normalized score)</li>\n<li>score from MF model</li>\n<li>statistics on MF item-item similarity with user history</li>\n</ul>\n<h2>Ranker</h2>\n<p>I found out that <code>Catboost</code> with <code>PairLogitPairwise</code> loss is the best option for my data and final score is achievable without ensembling. Inference is fast (1-1.5h), but not as fast as LightGBM/XGBoost with cuml.ForestInference (thanks <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a> for the clue).</p>\n<p>Summary:</p>\n<ul>\n<li>3-fold CV</li>\n<li>Catboost, PairLogitPairwise, 5000 iterations (2-4 minutes per fold on A100)</li>\n<li>separate models for each target (drop sessions w/o target + 20% random downsampling)</li>\n<li>LightGBM / XGBoost give slightly worse results (with lambdarank objectives), ensembling makes no differences</li>\n</ul>\n<h2>Pipeline &amp; technical details</h2>\n<p>I used mostly CUDF for data preparation and feature engineering. GPU memory is a bottleneck here, so I split data in 32 chunks.</p>\n<p>From the beginning I tried to implement robust pipeline with DVC, and it worked well until the last days of the competition, when I decided to increase candidates count from avg 200 to avg 300 :)</p>\n<p>One of the features of DVC is that it keeps track of parameters and changes, and stores results in cache. For example, if you want to experiment with some source of candidates, others won’t be recalculated. And you could return to previous state of data because of cache, but it’s not practical when you’re dealing with big amounts of data.<br>\nHere's how pipeline looks (below is almost the same part for test):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172966%2Fa51cd925311210365102fb8bb8e7211e%2F2023-02-01%20%2021.54.43.png?generation=1675263329789877&amp;alt=media\" alt=\"\"></p>\n<h2>Takeaways &amp; fails</h2>\n<ul>\n<li>try to keep data size as low as possible when actively trying ideas (firstly, I went from 60 to 200 candidates for 0.576 → 0.579 boost, then i went from 1/3 to full data for 0.596 → 0.597).</li>\n<li>as it’s multi-objective recommender system, I tried to use predictions of carts models as a feature for orders, but it did not work</li>\n<li>ensembling did not work after 0.6 LB. I used inverse rank averaging different rankers (e.g. catboost &amp; lightgbm) on different candidates setups.</li>\n<li>computational and personal time investment in Transformer models was not great in terms of leaderboard score, but knowledge I got is priceless. Also, it was the first time I tried Weight &amp; Biases for DL experiments tracking, and it’s awesome.</li>\n<li>this comp is hard, many ideas just don’t work. I think it took x3 time and effort than H&amp;M with almost the same LB position</li>\n</ul>",
  "messages": [
    {
      "id": "2125268",
      "postDate": "02/01/2023 15:01:50",
      "content": "<p>First of all, I want to thank OTTO for organising great competition. It's was very challenging and clean :) <br>\nSecondly, I want to thank all people who share their thoughts and ideas in kernels or forums. It helps a lot.<br>\nI will try to put my code on GitHub in a few days, need to clean up some deadline madness there.</p>\n<h2>TL;DR</h2>\n<ul>\n<li>use 3 types of co-occurence matrices (somewhat similar ideas to public notebooks, but different implementation and details)</li>\n<li>Bert MLM</li>\n<li>Matrix Factorization</li>\n<li>Catboost + PairLogitPairwise</li>\n</ul>\n<h3>LB Progression (public)</h3>\n<ul>\n<li>0.576 - history + 60 candidates all-to-all co-occurence matrix + lightgbm ranker</li>\n<li>0.579 - same, but 200 candidates</li>\n<li>0.583 - added a bunch of item features (conversions, popularity etc.)</li>\n<li><strong>0.593 - optimized co-occurence matrix</strong> (switched to sessions, added time weighing)</li>\n<li>0.595 - add buy2buy features and Transfomer candidates, switch to catboost</li>\n<li>0.597 - add buy2buy and type weighted co-occurence candidates, use more data for training (16/32 chunks)</li>\n<li>0.598 - use full data for training</li>\n<li>0.599 - use different candidates configs for clicks/buys, add MF candidates and scores</li>\n<li>0.600 - use different transformers</li>\n<li>0.601  - use x1.5 more candidates from each source for carts/orders</li>\n</ul>\n<h2>Candidates retrieval</h2>\n<p>I use <strong>max-recall@200</strong> as a retrieval quality measure, but candidates from different sources can be more or less common with user history, so it seems more fair to outer join candidates with history to calculate max-recall. <strong>0.598 LB is achievable only with co-visitation candidates</strong></p>\n<p>I used different combinations of candidates for clicks and carts/orders.</p>\n<pre><code>clicks:\n    history_rank: \n    cooc_rank: \n    buy2buy_rank: \n    cooc_tw_rank: \n    mfc_rank: \n    transformer_rank: \n  carts/oders: \n    history_rank: \n    cooc_rank: \n    cooc_tw_rank: \n    buy2buy_rank: \n    mfc_rank: \n    transformer_rank: \n</code></pre>\n<h3>Co-visitation (all-to-all)</h3>\n<ul>\n<li>Use actual user’s “sessions” - consecutive series of events, if there’s a gap &gt; 900 seconds, it’s another session.</li>\n<li>Use exponential time weighing (more distant events are less significant)  0.99995^(abs(ts.x - ts.y))</li>\n<li>inverse rank weighing for user history events</li>\n</ul>\n<p>Hyperparameters like session gap, time base and rank weight function were optimized with optuna, so <strong>max-recall@200</strong> was like <strong>67.6</strong> for this method.</p>\n<h3>Co-visitation (type weighted)</h3>\n<ol>\n<li>Use 1 day gap with exponential time weighing (no “sessions”)</li>\n<li>10x weight for carts, 3x for orders</li>\n</ol>\n<h3>Co-visitation (buy2buy)</h3>\n<ol>\n<li>Use 2 weeks gap, exponential time weighing</li>\n<li>Use only carts and orders to calculate stats</li>\n</ol>\n<h3>Transformer (small BERT)</h3>\n<p>It’s was hugely inspired by the <a href=\"https://github.com/Chubasik/yacup_recsys_2022\" target=\"_blank\">winning solution</a> (by <a href=\"https://www.kaggle.com/chubasik\" target=\"_blank\">@chubasik</a>) of recent Yandex.Cup Recsys track (I took <a href=\"https://github.com/greenwolf-nsk/yandex-cup-2022-recsys\" target=\"_blank\">2nd place</a> there with classical 2-stage approach).</p>\n<p>The idea is to train Masked Language Model, and then predict the “fake” last masked item in user session. Also, I fed action types (click, order, cart) as token_type_ids (which is mainly used for context separation in NLP tasks).</p>\n<p>I trained MLM on train sessions, used 500k most carted items, <strong>max-recall@200</strong> was like <strong>0.66.</strong> This quality can be achieved with 3 epochs on full data, but it takes very long to train (7 hours on A100 GPU), so my experiments were very limited.</p>\n<p>Adding this source of candidates gave boost of 0.002 in local CV, and this model score was second most valuable feature for ranking model. However, LB change was less then 0.001, and I spend last two weeks figuring out what went wrong. Finally, I trained two different models (train_no_val + val and train + test), this probably helped a little, but CV-LB gap was still bigger than before.<br>\nThere's 3 epochs training from scratch, with metrics every 0.5 epochs<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172966%2F916debefa57949c714c09d06c2ac6f9e%2F2023-02-01%20%2021.49.24.png?generation=1675263027888603&amp;alt=media\" alt=\"3 epoch training (eval every 0.5 epochs)\"></p>\n<h3>Matrix Factorization</h3>\n<p>Matrix factorization on Pytorch with BPR-like loss and hard negative sampling. It achieves <strong>0.665 max-recall@200</strong>. There’s probably a huge space for improvement in weighing events by type and time and negative sampling strategies. Training time is <strong>7h</strong> in total for 20 epochs with AdamW optimizer. It almost made no difference to CV/LB score, but MF score was also one of the strongest features.</p>\n<h2>Features</h2>\n<p>Best model uses around 200 features.</p>\n<h3>User</h3>\n<ul>\n<li>counters by type, normalized counter, time-based features</li>\n<li>number of “sessions”, avg session length</li>\n<li>avg/min/max/std “popularity” of item in user history</li>\n</ul>\n<h3>Item</h3>\n<ul>\n<li>item popularity by type - counters and ranks</li>\n<li>tried some derivatives to detect “trending” items, but they didn’t work for me</li>\n<li>item click/cart/order conversion rates</li>\n</ul>\n<h3>User-Item</h3>\n<ul>\n<li>interaction stats with item (number of clicks/carts, last timestamp)</li>\n<li>all features from co-visit matrices and statistics (mean/min/max/std for score/rank/normalized score)</li>\n<li>score from MF model</li>\n<li>statistics on MF item-item similarity with user history</li>\n</ul>\n<h2>Ranker</h2>\n<p>I found out that <code>Catboost</code> with <code>PairLogitPairwise</code> loss is the best option for my data and final score is achievable without ensembling. Inference is fast (1-1.5h), but not as fast as LightGBM/XGBoost with cuml.ForestInference (thanks <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a> for the clue).</p>\n<p>Summary:</p>\n<ul>\n<li>3-fold CV</li>\n<li>Catboost, PairLogitPairwise, 5000 iterations (2-4 minutes per fold on A100)</li>\n<li>separate models for each target (drop sessions w/o target + 20% random downsampling)</li>\n<li>LightGBM / XGBoost give slightly worse results (with lambdarank objectives), ensembling makes no differences</li>\n</ul>\n<h2>Pipeline &amp; technical details</h2>\n<p>I used mostly CUDF for data preparation and feature engineering. GPU memory is a bottleneck here, so I split data in 32 chunks.</p>\n<p>From the beginning I tried to implement robust pipeline with DVC, and it worked well until the last days of the competition, when I decided to increase candidates count from avg 200 to avg 300 :)</p>\n<p>One of the features of DVC is that it keeps track of parameters and changes, and stores results in cache. For example, if you want to experiment with some source of candidates, others won’t be recalculated. And you could return to previous state of data because of cache, but it’s not practical when you’re dealing with big amounts of data.<br>\nHere's how pipeline looks (below is almost the same part for test):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172966%2Fa51cd925311210365102fb8bb8e7211e%2F2023-02-01%20%2021.54.43.png?generation=1675263329789877&amp;alt=media\" alt=\"\"></p>\n<h2>Takeaways &amp; fails</h2>\n<ul>\n<li>try to keep data size as low as possible when actively trying ideas (firstly, I went from 60 to 200 candidates for 0.576 → 0.579 boost, then i went from 1/3 to full data for 0.596 → 0.597).</li>\n<li>as it’s multi-objective recommender system, I tried to use predictions of carts models as a feature for orders, but it did not work</li>\n<li>ensembling did not work after 0.6 LB. I used inverse rank averaging different rankers (e.g. catboost &amp; lightgbm) on different candidates setups.</li>\n<li>computational and personal time investment in Transformer models was not great in terms of leaderboard score, but knowledge I got is priceless. Also, it was the first time I tried Weight &amp; Biases for DL experiments tracking, and it’s awesome.</li>\n<li>this comp is hard, many ideas just don’t work. I think it took x3 time and effort than H&amp;M with almost the same LB position</li>\n</ul>",
      "rawMarkdown": "First of all, I want to thank OTTO for organising great competition. It's was very challenging and clean :) \nSecondly, I want to thank all people who share their thoughts and ideas in kernels or forums. It helps a lot.\nI will try to put my code on GitHub in a few days, need to clean up some deadline madness there.\n\n\n## TL;DR\n\n- use 3 types of co-occurence matrices (somewhat similar ideas to public notebooks, but different implementation and details)\n- Bert MLM\n- Matrix Factorization\n- Catboost + PairLogitPairwise\n\n### LB Progression (public)\n\n- 0.576 - history + 60 candidates all-to-all co-occurence matrix + lightgbm ranker\n- 0.579 - same, but 200 candidates\n- 0.583 - added a bunch of item features (conversions, popularity etc.)\n- **0.593 - optimized co-occurence matrix** (switched to sessions, added time weighing)\n- 0.595 - add buy2buy features and Transfomer candidates, switch to catboost\n- 0.597 - add buy2buy and type weighted co-occurence candidates, use more data for training (16/32 chunks)\n- 0.598 - use full data for training\n- 0.599 - use different candidates configs for clicks/buys, add MF candidates and scores\n- 0.600 - use different transformers\n- 0.601  - use x1.5 more candidates from each source for carts/orders\n\n\n## Candidates retrieval\n\nI use **max-recall@200** as a retrieval quality measure, but candidates from different sources can be more or less common with user history, so it seems more fair to outer join candidates with history to calculate max-recall. **0.598 LB is achievable only with co-visitation candidates**\n\nI used different combinations of candidates for clicks and carts/orders.\n\n```python\nclicks:\n    history_rank: 100\n    cooc_rank: 200\n    buy2buy_rank: 0\n    cooc_tw_rank: 0\n    mfc_rank: 100\n    transformer_rank: 50\n  carts/oders: # ~190 candidates on avg, 0.685+ WR\n    history_rank: 100\n    cooc_rank: 100\n    cooc_tw_rank: 100\n    buy2buy_rank: 100\n    mfc_rank: 50\n    transformer_rank: 50\n```\n\n### Co-visitation (all-to-all)\n\n- Use actual user’s “sessions” - consecutive series of events, if there’s a gap > 900 seconds, it’s another session.\n- Use exponential time weighing (more distant events are less significant)  0.99995^(abs(ts.x - ts.y))\n- inverse rank weighing for user history events\n\nHyperparameters like session gap, time base and rank weight function were optimized with optuna, so **max-recall@200** was like **67.6** for this method.\n\n### Co-visitation (type weighted)\n\n1. Use 1 day gap with exponential time weighing (no “sessions”)\n2. 10x weight for carts, 3x for orders\n\n### Co-visitation (buy2buy)\n\n1. Use 2 weeks gap, exponential time weighing\n2. Use only carts and orders to calculate stats\n\n### Transformer (small BERT)\n\nIt’s was hugely inspired by the [winning solution](https://github.com/Chubasik/yacup_recsys_2022) (by @chubasik) of recent Yandex.Cup Recsys track (I took [2nd place](https://github.com/greenwolf-nsk/yandex-cup-2022-recsys) there with classical 2-stage approach).\n\nThe idea is to train Masked Language Model, and then predict the “fake” last masked item in user session. Also, I fed action types (click, order, cart) as token_type_ids (which is mainly used for context separation in NLP tasks).\n\nI trained MLM on train sessions, used 500k most carted items, **max-recall@200** was like **0.66.** This quality can be achieved with 3 epochs on full data, but it takes very long to train (7 hours on A100 GPU), so my experiments were very limited.\n\nAdding this source of candidates gave boost of 0.002 in local CV, and this model score was second most valuable feature for ranking model. However, LB change was less then 0.001, and I spend last two weeks figuring out what went wrong. Finally, I trained two different models (train_no_val + val and train + test), this probably helped a little, but CV-LB gap was still bigger than before.\nThere's 3 epochs training from scratch, with metrics every 0.5 epochs\n![3 epoch training (eval every 0.5 epochs)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172966%2F916debefa57949c714c09d06c2ac6f9e%2F2023-02-01%20%2021.49.24.png?generation=1675263027888603&alt=media)\n\n\n\n### Matrix Factorization\n\nMatrix factorization on Pytorch with BPR-like loss and hard negative sampling. It achieves **0.665 max-recall@200**. There’s probably a huge space for improvement in weighing events by type and time and negative sampling strategies. Training time is **7h** in total for 20 epochs with AdamW optimizer. It almost made no difference to CV/LB score, but MF score was also one of the strongest features.\n\n\n## Features\n\nBest model uses around 200 features.\n\n### User\n\n- counters by type, normalized counter, time-based features\n- number of “sessions”, avg session length\n- avg/min/max/std “popularity” of item in user history\n\n### Item\n\n- item popularity by type - counters and ranks\n- tried some derivatives to detect “trending” items, but they didn’t work for me\n- item click/cart/order conversion rates\n\n### User-Item\n\n- interaction stats with item (number of clicks/carts, last timestamp)\n- all features from co-visit matrices and statistics (mean/min/max/std for score/rank/normalized score)\n- score from MF model\n- statistics on MF item-item similarity with user history\n\n\n## Ranker\n\nI found out that `Catboost` with `PairLogitPairwise` loss is the best option for my data and final score is achievable without ensembling. Inference is fast (1-1.5h), but not as fast as LightGBM/XGBoost with cuml.ForestInference (thanks @buumoo for the clue).\n\nSummary:\n- 3-fold CV\n- Catboost, PairLogitPairwise, 5000 iterations (2-4 minutes per fold on A100)\n- separate models for each target (drop sessions w/o target + 20% random downsampling)\n- LightGBM / XGBoost give slightly worse results (with lambdarank objectives), ensembling makes no differences\n\n## Pipeline & technical details\n\nI used mostly CUDF for data preparation and feature engineering. GPU memory is a bottleneck here, so I split data in 32 chunks.\n\nFrom the beginning I tried to implement robust pipeline with DVC, and it worked well until the last days of the competition, when I decided to increase candidates count from avg 200 to avg 300 :)\n\nOne of the features of DVC is that it keeps track of parameters and changes, and stores results in cache. For example, if you want to experiment with some source of candidates, others won’t be recalculated. And you could return to previous state of data because of cache, but it’s not practical when you’re dealing with big amounts of data.\nHere's how pipeline looks (below is almost the same part for test):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172966%2Fa51cd925311210365102fb8bb8e7211e%2F2023-02-01%20%2021.54.43.png?generation=1675263329789877&alt=media)\n\n\n## Takeaways & fails\n\n- try to keep data size as low as possible when actively trying ideas (firstly, I went from 60 to 200 candidates for 0.576 → 0.579 boost, then i went from 1/3 to full data for 0.596 → 0.597).\n- as it’s multi-objective recommender system, I tried to use predictions of carts models as a feature for orders, but it did not work\n- ensembling did not work after 0.6 LB. I used inverse rank averaging different rankers (e.g. catboost & lightgbm) on different candidates setups.\n- computational and personal time investment in Transformer models was not great in terms of leaderboard score, but knowledge I got is priceless. Also, it was the first time I tried Weight & Biases for DL experiments tracking, and it’s awesome.\n- this comp is hard, many ideas just don’t work. I think it took x3 time and effort than H&M with almost the same LB position",
      "votes": null
    },
    {
      "id": "2125346",
      "postDate": "02/01/2023 15:48:08",
      "content": "<p>Thanks for your sharing. I did same thing at the last day </p>\n<blockquote>\n  <p>use predictions of carts models as a feature for orders, but it did not work</p>\n</blockquote>\n<ol>\n<li>carts models as a feature for orders not works - same result and recall@20 dropped around 5e-4</li>\n<li>orders models as a feature for carts - improved a bit, around 2e-4</li>\n</ol>",
      "rawMarkdown": "Thanks for your sharing. I did same thing at the last day \n\n> use predictions of carts models as a feature for orders, but it did not work\n\n1. carts models as a feature for orders not works - same result and recall@20 dropped around 5e-4\n2. orders models as a feature for carts - improved a bit, around 2e-4",
      "votes": null
    },
    {
      "id": "2125363",
      "postDate": "02/01/2023 16:03:19",
      "content": "<p>I used click and carts as features to pred orders and click and order as features to pred carts. Both didn't work. But I test it in local cv, it improve by ~ 0.0008. confused.👀</p>",
      "rawMarkdown": "I used click and carts as features to pred orders and click and order as features to pred carts. Both didn't work. But I test it in local cv, it improve by ~ 0.0008. confused.👀",
      "votes": null
    },
    {
      "id": "2125385",
      "postDate": "02/01/2023 16:16:30",
      "content": "<p>I tested clicks as features of carts &amp; orders at early, got same result. But I made a mistake at that moment, my kfold prediction only include the session which has at least one true label, so there is some leakage. Anyway, the improvement is quite limited as exp #2</p>",
      "rawMarkdown": "I tested clicks as features of carts & orders at early, got same result. But I made a mistake at that moment, my kfold prediction only include the session which has at least one true label, so there is some leakage. Anyway, the improvement is quite limited as exp #2",
      "votes": null
    },
    {
      "id": "2125829",
      "postDate": "02/01/2023 23:27:24",
      "content": "<p>Great solution! Thanks for the insights into Transformer, I also tried my own implementation of BERT4Rec with some modifications but unfortunately it didn't performed well, so now I have some food for thoughts what I do wrong.</p>",
      "rawMarkdown": "Great solution! Thanks for the insights into Transformer, I also tried my own implementation of BERT4Rec with some modifications but unfortunately it didn't performed well, so now I have some food for thoughts what I do wrong.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2125346,
      "author_name": "gongbi",
      "author_url": "",
      "post_date": "02/01/2023 15:48:08",
      "content": "<p>Thanks for your sharing. I did same thing at the last day </p>\n<blockquote>\n  <p>use predictions of carts models as a feature for orders, but it did not work</p>\n</blockquote>\n<ol>\n<li>carts models as a feature for orders not works - same result and recall@20 dropped around 5e-4</li>\n<li>orders models as a feature for carts - improved a bit, around 2e-4</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2125363,
          "author_name": "hookman",
          "author_url": "",
          "post_date": "02/01/2023 16:03:19",
          "content": "<p>I used click and carts as features to pred orders and click and order as features to pred carts. Both didn't work. But I test it in local cv, it improve by ~ 0.0008. confused.👀</p>",
          "votes": null,
          "replies": [
            {
              "id": 2125385,
              "author_name": "gongbi",
              "author_url": "",
              "post_date": "02/01/2023 16:16:30",
              "content": "<p>I tested clicks as features of carts &amp; orders at early, got same result. But I made a mistake at that moment, my kfold prediction only include the session which has at least one true label, so there is some leakage. Anyway, the improvement is quite limited as exp #2</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2125829,
      "author_name": "sirpantene",
      "author_url": "",
      "post_date": "02/01/2023 23:27:24",
      "content": "<p>Great solution! Thanks for the insights into Transformer, I also tried my own implementation of BERT4Rec with some modifications but unfortunately it didn't performed well, so now I have some food for thoughts what I do wrong.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2125268": "First of all, I want to thank OTTO for organising great competition. It's was very challenging and clean :) \nSecondly, I want to thank all people who share their thoughts and ideas in kernels or forums. It helps a lot.\nI will try to put my code on GitHub in a few days, need to clean up some deadline madness there.\n\n\n## TL;DR\n\n- use 3 types of co-occurence matrices (somewhat similar ideas to public notebooks, but different implementation and details)\n- Bert MLM\n- Matrix Factorization\n- Catboost + PairLogitPairwise\n\n### LB Progression (public)\n\n- 0.576 - history + 60 candidates all-to-all co-occurence matrix + lightgbm ranker\n- 0.579 - same, but 200 candidates\n- 0.583 - added a bunch of item features (conversions, popularity etc.)\n- **0.593 - optimized co-occurence matrix** (switched to sessions, added time weighing)\n- 0.595 - add buy2buy features and Transfomer candidates, switch to catboost\n- 0.597 - add buy2buy and type weighted co-occurence candidates, use more data for training (16/32 chunks)\n- 0.598 - use full data for training\n- 0.599 - use different candidates configs for clicks/buys, add MF candidates and scores\n- 0.600 - use different transformers\n- 0.601  - use x1.5 more candidates from each source for carts/orders\n\n\n## Candidates retrieval\n\nI use **max-recall@200** as a retrieval quality measure, but candidates from different sources can be more or less common with user history, so it seems more fair to outer join candidates with history to calculate max-recall. **0.598 LB is achievable only with co-visitation candidates**\n\nI used different combinations of candidates for clicks and carts/orders.\n\n```python\nclicks:\n    history_rank: 100\n    cooc_rank: 200\n    buy2buy_rank: 0\n    cooc_tw_rank: 0\n    mfc_rank: 100\n    transformer_rank: 50\n  carts/oders: # ~190 candidates on avg, 0.685+ WR\n    history_rank: 100\n    cooc_rank: 100\n    cooc_tw_rank: 100\n    buy2buy_rank: 100\n    mfc_rank: 50\n    transformer_rank: 50\n```\n\n### Co-visitation (all-to-all)\n\n- Use actual user’s “sessions” - consecutive series of events, if there’s a gap > 900 seconds, it’s another session.\n- Use exponential time weighing (more distant events are less significant)  0.99995^(abs(ts.x - ts.y))\n- inverse rank weighing for user history events\n\nHyperparameters like session gap, time base and rank weight function were optimized with optuna, so **max-recall@200** was like **67.6** for this method.\n\n### Co-visitation (type weighted)\n\n1. Use 1 day gap with exponential time weighing (no “sessions”)\n2. 10x weight for carts, 3x for orders\n\n### Co-visitation (buy2buy)\n\n1. Use 2 weeks gap, exponential time weighing\n2. Use only carts and orders to calculate stats\n\n### Transformer (small BERT)\n\nIt’s was hugely inspired by the [winning solution](https://github.com/Chubasik/yacup_recsys_2022) (by @chubasik) of recent Yandex.Cup Recsys track (I took [2nd place](https://github.com/greenwolf-nsk/yandex-cup-2022-recsys) there with classical 2-stage approach).\n\nThe idea is to train Masked Language Model, and then predict the “fake” last masked item in user session. Also, I fed action types (click, order, cart) as token_type_ids (which is mainly used for context separation in NLP tasks).\n\nI trained MLM on train sessions, used 500k most carted items, **max-recall@200** was like **0.66.** This quality can be achieved with 3 epochs on full data, but it takes very long to train (7 hours on A100 GPU), so my experiments were very limited.\n\nAdding this source of candidates gave boost of 0.002 in local CV, and this model score was second most valuable feature for ranking model. However, LB change was less then 0.001, and I spend last two weeks figuring out what went wrong. Finally, I trained two different models (train_no_val + val and train + test), this probably helped a little, but CV-LB gap was still bigger than before.\nThere's 3 epochs training from scratch, with metrics every 0.5 epochs\n![3 epoch training (eval every 0.5 epochs)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172966%2F916debefa57949c714c09d06c2ac6f9e%2F2023-02-01%20%2021.49.24.png?generation=1675263027888603&alt=media)\n\n\n\n### Matrix Factorization\n\nMatrix factorization on Pytorch with BPR-like loss and hard negative sampling. It achieves **0.665 max-recall@200**. There’s probably a huge space for improvement in weighing events by type and time and negative sampling strategies. Training time is **7h** in total for 20 epochs with AdamW optimizer. It almost made no difference to CV/LB score, but MF score was also one of the strongest features.\n\n\n## Features\n\nBest model uses around 200 features.\n\n### User\n\n- counters by type, normalized counter, time-based features\n- number of “sessions”, avg session length\n- avg/min/max/std “popularity” of item in user history\n\n### Item\n\n- item popularity by type - counters and ranks\n- tried some derivatives to detect “trending” items, but they didn’t work for me\n- item click/cart/order conversion rates\n\n### User-Item\n\n- interaction stats with item (number of clicks/carts, last timestamp)\n- all features from co-visit matrices and statistics (mean/min/max/std for score/rank/normalized score)\n- score from MF model\n- statistics on MF item-item similarity with user history\n\n\n## Ranker\n\nI found out that `Catboost` with `PairLogitPairwise` loss is the best option for my data and final score is achievable without ensembling. Inference is fast (1-1.5h), but not as fast as LightGBM/XGBoost with cuml.ForestInference (thanks @buumoo for the clue).\n\nSummary:\n- 3-fold CV\n- Catboost, PairLogitPairwise, 5000 iterations (2-4 minutes per fold on A100)\n- separate models for each target (drop sessions w/o target + 20% random downsampling)\n- LightGBM / XGBoost give slightly worse results (with lambdarank objectives), ensembling makes no differences\n\n## Pipeline & technical details\n\nI used mostly CUDF for data preparation and feature engineering. GPU memory is a bottleneck here, so I split data in 32 chunks.\n\nFrom the beginning I tried to implement robust pipeline with DVC, and it worked well until the last days of the competition, when I decided to increase candidates count from avg 200 to avg 300 :)\n\nOne of the features of DVC is that it keeps track of parameters and changes, and stores results in cache. For example, if you want to experiment with some source of candidates, others won’t be recalculated. And you could return to previous state of data because of cache, but it’s not practical when you’re dealing with big amounts of data.\nHere's how pipeline looks (below is almost the same part for test):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172966%2Fa51cd925311210365102fb8bb8e7211e%2F2023-02-01%20%2021.54.43.png?generation=1675263329789877&alt=media)\n\n\n## Takeaways & fails\n\n- try to keep data size as low as possible when actively trying ideas (firstly, I went from 60 to 200 candidates for 0.576 → 0.579 boost, then i went from 1/3 to full data for 0.596 → 0.597).\n- as it’s multi-objective recommender system, I tried to use predictions of carts models as a feature for orders, but it did not work\n- ensembling did not work after 0.6 LB. I used inverse rank averaging different rankers (e.g. catboost & lightgbm) on different candidates setups.\n- computational and personal time investment in Transformer models was not great in terms of leaderboard score, but knowledge I got is priceless. Also, it was the first time I tried Weight & Biases for DL experiments tracking, and it’s awesome.\n- this comp is hard, many ideas just don’t work. I think it took x3 time and effort than H&M with almost the same LB position",
    "2125346": "Thanks for your sharing. I did same thing at the last day \n\n> use predictions of carts models as a feature for orders, but it did not work\n\n1. carts models as a feature for orders not works - same result and recall@20 dropped around 5e-4\n2. orders models as a feature for carts - improved a bit, around 2e-4",
    "2125363": "I used click and carts as features to pred orders and click and order as features to pred carts. Both didn't work. But I test it in local cv, it improve by ~ 0.0008. confused.👀",
    "2125385": "I tested clicks as features of carts & orders at early, got same result. But I made a mistake at that moment, my kfold prediction only include the session which has at least one true label, so there is some leakage. Anyway, the improvement is quite limited as exp #2",
    "2125829": "Great solution! Thanks for the insights into Transformer, I also tried my own implementation of BERT4Rec with some modifications but unfortunately it didn't performed well, so now I have some food for thoughts what I do wrong."
  },
  "source": "meta"
}