{
  "id": 382771,
  "title": "20th Place Solution",
  "url": "/competitions/otto-recommender-system/writeups/kicchotto-20th-place-solution",
  "author_name": "",
  "post_date": "2023-02-11T23:37:58.030Z",
  "votes": 90,
  "comment_count": 31,
  "views": 0,
  "content": "<h1>Acknowledgements</h1>\n<p>First, I would like to thank the organizers and those who shared knowledge. Especially I would like to thank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for posting so many codes and discussions. I learned a lot from you. </p>\n<p>Considering the controversy about the possible cheaters, I will publish the entire code after everything is finalized and I will just share the ideas in this solution. (Although <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382440#2123460\" target=\"_blank\">Will says sharing code is not something we should withhold</a>)</p>\n<p><br></p>\n<h1>Solution Summary</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2F438095d4971fd2b40113d40cf1d96fc1%2FScreen%20Shot%202023-02-01%20at%209.28.52.png?generation=1675211352995855&amp;alt=media\" alt=\"\"></p>\n<p>My code is here: <a href=\"https://github.com/kiccho1101/kaggle-otto2\" target=\"_blank\">https://github.com/kiccho1101/kaggle-otto2</a></p>\n<p><br></p>\n<h1>CV Strategy</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2F728b7a7f2c1ef5eeec8207ac1e7f870f%2Frecommend-data-split.svg?generation=1675210198356258&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>I used Radek’s CV strategy</li>\n<li>Although CV data was good for the local experiments, it took quite a long time to run experiments so I used only 1/20 sessions by random sampling. As a result, the average time that takes to run 1 experiment became about 10~30mins which allowed me to run many experiments quickly.</li>\n<li>Even if I sampled the data, the CV-LB score correlation was very stable.</li>\n<li>The data used for each stage is below</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Step</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Candidate Generation</td>\n<td>TrainData + ValidDataA</td>\n<td>TrainData + ValidDataA + ValidDataB + TestDataA</td>\n</tr>\n<tr>\n<td>User Feature Creation</td>\n<td>ValidDataA</td>\n<td>TestDataA</td>\n</tr>\n<tr>\n<td>Item Feature Creation</td>\n<td>TrainData + ValidDataA</td>\n<td>TrainData + ValidDataA + ValidDataB + TestDataA</td>\n</tr>\n<tr>\n<td>User-Item Feature Creation</td>\n<td>ValidDataA</td>\n<td>TestDataA</td>\n</tr>\n<tr>\n<td>Re-Ranking</td>\n<td>ValidDataA</td>\n<td>TestDataA</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<h1>1st Stage - Candidate Generation</h1>\n<p><br></p>\n<h3>word2vec</h3>\n<ul>\n<li>Trained word2vec model with aid sequences</li>\n<li>Retrieve top-k aids by word2vec embeddings (Used faiss-gpu to speed up)</li>\n</ul>\n<p><br></p>\n<h3>CoVis</h3>\n<ul>\n<li>I used Chris’s co-visitation matrix to generate candidates</li>\n</ul>\n<p><br></p>\n<h3>Item MF (Item Matrix Factorization)</h3>\n<pre><code>class ItemMFModel(nn.Module):\n    def __init__(self, n_aid: int, n_factors: int):\n        super().__init__()\n        self.criterion = BPRLoss()\n        self.n_factors = n_factors\n        self.n_aid = n_aid\n        self.aid_embeddings = nn.Embedding(self.n_aid, self.n_factors)\n\n        initrange = 1.0 / self.n_factors\n        nn.init.uniform_(self.aid_embeddings.weight.data, -initrange, initrange)\n\n    def forward(self, aid_x, aid_y):\n        aid_x = self.aid_embeddings(aid_x)\n        aid_y = self.aid_embeddings(aid_y)\n        return (aid_x * aid_y).sum(dim=1)\n\n    def calc_loss(self, aid_x, aid_y, size_x, size_y):\n        rand_idx = torch.randperm(aid_y.size(0))\n        output_pos = self.forward(aid_x, aid_y)\n        output_neg = self.forward(aid_x, aid_y[rand_idx])\n        loss = self.criterion(output_pos, output_neg)\n        return loss\n</code></pre>\n<ul>\n<li>Trained aid_embeddings with BPR loss to make co-occurring embeddings become similar</li>\n<li>What worked<ul>\n<li>Multiply the inverse of item_size by loss (Removing popularity bias)</li>\n<li>Multiply the inverse of ts_diff by loss (The closer the co-occur timing is, the more similar the embeddings become)</li></ul></li>\n</ul>\n<p><br></p>\n<h3>User MF (User Matrix Factorization)</h3>\n<pre><code>class UserMFModel(nn.Module):\n    def __init__(self, n_session: int, n_aid: int, n_factors: int):\n        super().__init__()\n        self.n_factors = n_factors\n        self.n_session = n_session\n        self.n_aid = n_aid\n\n        self.session_embeddings = nn.Embedding(self.n_session, self.n_factors)\n        self.aid_embeddings = nn.Embedding(self.n_aid, self.n_factors)\n\n        self.criterion = BPRLoss()\n\n        initrange = 1.0 / self.n_factors\n        nn.init.uniform_(self.session_embeddings.weight.data, -initrange, initrange)\n        nn.init.uniform_(self.aid_embeddings.weight.data, -initrange, initrange)\n\n    def forward(self, session, aid, aid_size):\n        session_emb = self.session_embeddings(session)\n        aid_emb = self.aid_embeddings(aid)\n        return (session_emb * aid_emb).sum(dim=1)\n\n    def calc_loss(self, session, aid, aid_size):\n        rand_idx = torch.randperm(aid.size(0))\n        output_pos = self.forward(session, aid)\n        output_neg = self.forward(session, aid[rand_idx])\n        loss = self.criterion(output_pos, output_neg)\n        return loss\n</code></pre>\n<ul>\n<li>Trained session_embeddings and aid_embeddings with BPR loss</li>\n<li>What worked<ul>\n<li>Multiply the inverse of item_size by loss (Removing popularity bias)</li>\n<li>Multiply the inverse of ts_diff by loss (The closer the co-occur timing is, the more similar the embeddings become)</li></ul></li>\n</ul>\n<p><br></p>\n<h3>Item CF</h3>\n<ul>\n<li>Implemented item cf with polars</li>\n<li>Calculated the similarity weights for each item-item pair and retrieved candidates by getting the most similar items based on sum/min/max/mean of weights</li>\n<li>What worked<ul>\n<li>Multiply the inverse of item_size by weight (Removing popularity bias)</li>\n<li>Multiply the inverse of ts_diff by weight (The closer the co-occur timing is, the bigger the weight becomes)</li>\n<li>Multiply trend coefficient (The more ts is recent, the bigger the weight becomes)</li></ul></li>\n</ul>\n<p><br></p>\n<h3>User CF</h3>\n<ul>\n<li>Implemented in the same way as item cf</li>\n</ul>\n<p><br></p>\n<h1>2nd Stage - Re-Ranking</h1>\n<ul>\n<li>Feature<ul>\n<li>Created ~200 features in total</li>\n<li>pl.col(’ts’).agg([mean,min,max,std]).over({’session’ or ‘aid’})</li>\n<li>candidate_selected_(count, rank)</li>\n<li>candidate_(score, rank, selected)</li>\n<li>(inter, click, cart, order)_hour_mean.over({’session’ or ‘aid’})</li>\n<li>(inter, click, cart, order)_count.over({’session’ or ‘aid’})@(3d, 7d, 14d, 21d)</li>\n<li>aid_multi_(inter,click,cart,order)_prob</li></ul></li>\n<li>Model<ul>\n<li>LightGBM Ranker (lambdarank)</li>\n<li>CatBoost Classifier (Logloss)</li>\n<li>CatBoost Ranker (YetiRank)</li></ul></li>\n</ul>\n<p><br></p>\n<h1>What did not work well</h1>\n<ul>\n<li>Transformers</li>\n<li>GRU</li>\n<li>CDAE</li>\n<li>RecVAE</li>\n<li>Implicit(ALS, BPR)</li>\n<li>Clustering by item embeddings</li>\n<li>Popular items</li>\n<li>Stacking</li>\n<li>Pseudo Labeling</li>\n</ul>\n<p><br></p>\n<h1>Lessons Learned</h1>\n<ul>\n<li>polars is all you need to create features<ul>\n<li>At first, I created features with pandas and cudf but I switched to polars during the competition because polars is fast, memory efficient, and the syntax is easy to understand.</li></ul></li>\n<li>Creating a good baseline is very important<ul>\n<li>It took me 1~2 month to create a baseline pipeline</li>\n<li>Thanks to Chris’s great discussion post, I was able to create a 2-stage baseline from scratch (I will publish the code on GitHub soon)</li></ul></li>\n<li>Feature Store is useful<ul>\n<li>I stored intermediate files to parquet files, which saved me to run experiments quickly in a more reproducible way.</li></ul></li>\n</ul>\n<p><br></p>\n<h1>Score Timeline</h1>\n<table>\n<thead>\n<tr>\n<th>CV (1/20 sampled)</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.536</td>\n<td>0.569</td>\n<td>0.5799</td>\n<td>2-Stage Baseline</td>\n</tr>\n<tr>\n<td>0.539</td>\n<td>0.572</td>\n<td>0.5830</td>\n<td>Added Item2Vec ItemMF</td>\n</tr>\n<tr>\n<td>0.547</td>\n<td>0.581</td>\n<td>0.5930</td>\n<td>Added ItemCF UserMF</td>\n</tr>\n<tr>\n<td>0.550</td>\n<td>0.585</td>\n<td>0.5970</td>\n<td>Added variation to ItemCF</td>\n</tr>\n<tr>\n<td>0.552</td>\n<td>0.587</td>\n<td>0.5985</td>\n<td>Ensemble(LightGBM + CatBoost)</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<h1>Feature Importance</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2Febea15581ccf1d44b29d37db8e493fc5%2F68747470733a2f2f71696974612d696d6167652d73746f72652e73332e61702d6e6f727468656173742d312e616d617a6f6e6177732e636f6d2f302f3235393431372f31346465363937652d373162652d396462362d663337362d3066643762336537333630322e706e67.png?generation=1675210821369636&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2124330",
      "postDate": "02/01/2023 00:10:44",
      "content": "<h1>Acknowledgements</h1>\n<p>First, I would like to thank the organizers and those who shared knowledge. Especially I would like to thank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for posting so many codes and discussions. I learned a lot from you. </p>\n<p>Considering the controversy about the possible cheaters, I will publish the entire code after everything is finalized and I will just share the ideas in this solution. (Although <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/382440#2123460\" target=\"_blank\">Will says sharing code is not something we should withhold</a>)</p>\n<p><br></p>\n<h1>Solution Summary</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2F438095d4971fd2b40113d40cf1d96fc1%2FScreen%20Shot%202023-02-01%20at%209.28.52.png?generation=1675211352995855&amp;alt=media\" alt=\"\"></p>\n<p>My code is here: <a href=\"https://github.com/kiccho1101/kaggle-otto2\" target=\"_blank\">https://github.com/kiccho1101/kaggle-otto2</a></p>\n<p><br></p>\n<h1>CV Strategy</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2F728b7a7f2c1ef5eeec8207ac1e7f870f%2Frecommend-data-split.svg?generation=1675210198356258&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>I used Radek’s CV strategy</li>\n<li>Although CV data was good for the local experiments, it took quite a long time to run experiments so I used only 1/20 sessions by random sampling. As a result, the average time that takes to run 1 experiment became about 10~30mins which allowed me to run many experiments quickly.</li>\n<li>Even if I sampled the data, the CV-LB score correlation was very stable.</li>\n<li>The data used for each stage is below</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Step</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Candidate Generation</td>\n<td>TrainData + ValidDataA</td>\n<td>TrainData + ValidDataA + ValidDataB + TestDataA</td>\n</tr>\n<tr>\n<td>User Feature Creation</td>\n<td>ValidDataA</td>\n<td>TestDataA</td>\n</tr>\n<tr>\n<td>Item Feature Creation</td>\n<td>TrainData + ValidDataA</td>\n<td>TrainData + ValidDataA + ValidDataB + TestDataA</td>\n</tr>\n<tr>\n<td>User-Item Feature Creation</td>\n<td>ValidDataA</td>\n<td>TestDataA</td>\n</tr>\n<tr>\n<td>Re-Ranking</td>\n<td>ValidDataA</td>\n<td>TestDataA</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<h1>1st Stage - Candidate Generation</h1>\n<p><br></p>\n<h3>word2vec</h3>\n<ul>\n<li>Trained word2vec model with aid sequences</li>\n<li>Retrieve top-k aids by word2vec embeddings (Used faiss-gpu to speed up)</li>\n</ul>\n<p><br></p>\n<h3>CoVis</h3>\n<ul>\n<li>I used Chris’s co-visitation matrix to generate candidates</li>\n</ul>\n<p><br></p>\n<h3>Item MF (Item Matrix Factorization)</h3>\n<pre><code>class ItemMFModel(nn.Module):\n    def __init__(self, n_aid: int, n_factors: int):\n        super().__init__()\n        self.criterion = BPRLoss()\n        self.n_factors = n_factors\n        self.n_aid = n_aid\n        self.aid_embeddings = nn.Embedding(self.n_aid, self.n_factors)\n\n        initrange = 1.0 / self.n_factors\n        nn.init.uniform_(self.aid_embeddings.weight.data, -initrange, initrange)\n\n    def forward(self, aid_x, aid_y):\n        aid_x = self.aid_embeddings(aid_x)\n        aid_y = self.aid_embeddings(aid_y)\n        return (aid_x * aid_y).sum(dim=1)\n\n    def calc_loss(self, aid_x, aid_y, size_x, size_y):\n        rand_idx = torch.randperm(aid_y.size(0))\n        output_pos = self.forward(aid_x, aid_y)\n        output_neg = self.forward(aid_x, aid_y[rand_idx])\n        loss = self.criterion(output_pos, output_neg)\n        return loss\n</code></pre>\n<ul>\n<li>Trained aid_embeddings with BPR loss to make co-occurring embeddings become similar</li>\n<li>What worked<ul>\n<li>Multiply the inverse of item_size by loss (Removing popularity bias)</li>\n<li>Multiply the inverse of ts_diff by loss (The closer the co-occur timing is, the more similar the embeddings become)</li></ul></li>\n</ul>\n<p><br></p>\n<h3>User MF (User Matrix Factorization)</h3>\n<pre><code>class UserMFModel(nn.Module):\n    def __init__(self, n_session: int, n_aid: int, n_factors: int):\n        super().__init__()\n        self.n_factors = n_factors\n        self.n_session = n_session\n        self.n_aid = n_aid\n\n        self.session_embeddings = nn.Embedding(self.n_session, self.n_factors)\n        self.aid_embeddings = nn.Embedding(self.n_aid, self.n_factors)\n\n        self.criterion = BPRLoss()\n\n        initrange = 1.0 / self.n_factors\n        nn.init.uniform_(self.session_embeddings.weight.data, -initrange, initrange)\n        nn.init.uniform_(self.aid_embeddings.weight.data, -initrange, initrange)\n\n    def forward(self, session, aid, aid_size):\n        session_emb = self.session_embeddings(session)\n        aid_emb = self.aid_embeddings(aid)\n        return (session_emb * aid_emb).sum(dim=1)\n\n    def calc_loss(self, session, aid, aid_size):\n        rand_idx = torch.randperm(aid.size(0))\n        output_pos = self.forward(session, aid)\n        output_neg = self.forward(session, aid[rand_idx])\n        loss = self.criterion(output_pos, output_neg)\n        return loss\n</code></pre>\n<ul>\n<li>Trained session_embeddings and aid_embeddings with BPR loss</li>\n<li>What worked<ul>\n<li>Multiply the inverse of item_size by loss (Removing popularity bias)</li>\n<li>Multiply the inverse of ts_diff by loss (The closer the co-occur timing is, the more similar the embeddings become)</li></ul></li>\n</ul>\n<p><br></p>\n<h3>Item CF</h3>\n<ul>\n<li>Implemented item cf with polars</li>\n<li>Calculated the similarity weights for each item-item pair and retrieved candidates by getting the most similar items based on sum/min/max/mean of weights</li>\n<li>What worked<ul>\n<li>Multiply the inverse of item_size by weight (Removing popularity bias)</li>\n<li>Multiply the inverse of ts_diff by weight (The closer the co-occur timing is, the bigger the weight becomes)</li>\n<li>Multiply trend coefficient (The more ts is recent, the bigger the weight becomes)</li></ul></li>\n</ul>\n<p><br></p>\n<h3>User CF</h3>\n<ul>\n<li>Implemented in the same way as item cf</li>\n</ul>\n<p><br></p>\n<h1>2nd Stage - Re-Ranking</h1>\n<ul>\n<li>Feature<ul>\n<li>Created ~200 features in total</li>\n<li>pl.col(’ts’).agg([mean,min,max,std]).over({’session’ or ‘aid’})</li>\n<li>candidate_selected_(count, rank)</li>\n<li>candidate_(score, rank, selected)</li>\n<li>(inter, click, cart, order)_hour_mean.over({’session’ or ‘aid’})</li>\n<li>(inter, click, cart, order)_count.over({’session’ or ‘aid’})@(3d, 7d, 14d, 21d)</li>\n<li>aid_multi_(inter,click,cart,order)_prob</li></ul></li>\n<li>Model<ul>\n<li>LightGBM Ranker (lambdarank)</li>\n<li>CatBoost Classifier (Logloss)</li>\n<li>CatBoost Ranker (YetiRank)</li></ul></li>\n</ul>\n<p><br></p>\n<h1>What did not work well</h1>\n<ul>\n<li>Transformers</li>\n<li>GRU</li>\n<li>CDAE</li>\n<li>RecVAE</li>\n<li>Implicit(ALS, BPR)</li>\n<li>Clustering by item embeddings</li>\n<li>Popular items</li>\n<li>Stacking</li>\n<li>Pseudo Labeling</li>\n</ul>\n<p><br></p>\n<h1>Lessons Learned</h1>\n<ul>\n<li>polars is all you need to create features<ul>\n<li>At first, I created features with pandas and cudf but I switched to polars during the competition because polars is fast, memory efficient, and the syntax is easy to understand.</li></ul></li>\n<li>Creating a good baseline is very important<ul>\n<li>It took me 1~2 month to create a baseline pipeline</li>\n<li>Thanks to Chris’s great discussion post, I was able to create a 2-stage baseline from scratch (I will publish the code on GitHub soon)</li></ul></li>\n<li>Feature Store is useful<ul>\n<li>I stored intermediate files to parquet files, which saved me to run experiments quickly in a more reproducible way.</li></ul></li>\n</ul>\n<p><br></p>\n<h1>Score Timeline</h1>\n<table>\n<thead>\n<tr>\n<th>CV (1/20 sampled)</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.536</td>\n<td>0.569</td>\n<td>0.5799</td>\n<td>2-Stage Baseline</td>\n</tr>\n<tr>\n<td>0.539</td>\n<td>0.572</td>\n<td>0.5830</td>\n<td>Added Item2Vec ItemMF</td>\n</tr>\n<tr>\n<td>0.547</td>\n<td>0.581</td>\n<td>0.5930</td>\n<td>Added ItemCF UserMF</td>\n</tr>\n<tr>\n<td>0.550</td>\n<td>0.585</td>\n<td>0.5970</td>\n<td>Added variation to ItemCF</td>\n</tr>\n<tr>\n<td>0.552</td>\n<td>0.587</td>\n<td>0.5985</td>\n<td>Ensemble(LightGBM + CatBoost)</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<h1>Feature Importance</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2Febea15581ccf1d44b29d37db8e493fc5%2F68747470733a2f2f71696974612d696d6167652d73746f72652e73332e61702d6e6f727468656173742d312e616d617a6f6e6177732e636f6d2f302f3235393431372f31346465363937652d373162652d396462362d663337362d3066643762336537333630322e706e67.png?generation=1675210821369636&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "# Acknowledgements\nFirst, I would like to thank the organizers and those who shared knowledge. Especially I would like to thank @cdeotte and @radek1 for posting so many codes and discussions. I learned a lot from you. \n\nConsidering the controversy about the possible cheaters, I will publish the entire code after everything is finalized and I will just share the ideas in this solution. (Although [Will says sharing code is not something we should withhold](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382440#2123460))\n\n\n<br>\n# Solution Summary\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2F438095d4971fd2b40113d40cf1d96fc1%2FScreen%20Shot%202023-02-01%20at%209.28.52.png?generation=1675211352995855&alt=media)\n\nMy code is here: https://github.com/kiccho1101/kaggle-otto2\n\n<br>\n# CV Strategy\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2F728b7a7f2c1ef5eeec8207ac1e7f870f%2Frecommend-data-split.svg?generation=1675210198356258&alt=media)\n- I used Radek’s CV strategy\n- Although CV data was good for the local experiments, it took quite a long time to run experiments so I used only 1/20 sessions by random sampling. As a result, the average time that takes to run 1 experiment became about 10~30mins which allowed me to run many experiments quickly.\n- Even if I sampled the data, the CV-LB score correlation was very stable.\n- The data used for each stage is below\n\n| Step                       | CV                     | LB                                              |\n|----------------------------|------------------------|-------------------------------------------------|\n| Candidate Generation       | TrainData + ValidDataA | TrainData + ValidDataA + ValidDataB + TestDataA |\n| User Feature Creation      | ValidDataA             | TestDataA                                       |\n| Item Feature Creation      | TrainData + ValidDataA | TrainData + ValidDataA + ValidDataB + TestDataA |\n| User-Item Feature Creation | ValidDataA             | TestDataA                                       |\n| Re-Ranking                 | ValidDataA             | TestDataA                                       |\n\n<br>\n# 1st Stage - Candidate Generation\n\n<br>\n### word2vec\n- Trained word2vec model with aid sequences\n- Retrieve top-k aids by word2vec embeddings (Used faiss-gpu to speed up)\n\n<br>\n### CoVis\n- I used Chris’s co-visitation matrix to generate candidates\n\n<br>\n### Item MF (Item Matrix Factorization)\n```\nclass ItemMFModel(nn.Module):\n    def __init__(self, n_aid: int, n_factors: int):\n        super().__init__()\n        self.criterion = BPRLoss()\n        self.n_factors = n_factors\n        self.n_aid = n_aid\n        self.aid_embeddings = nn.Embedding(self.n_aid, self.n_factors)\n\n        initrange = 1.0 / self.n_factors\n        nn.init.uniform_(self.aid_embeddings.weight.data, -initrange, initrange)\n\n    def forward(self, aid_x, aid_y):\n        aid_x = self.aid_embeddings(aid_x)\n        aid_y = self.aid_embeddings(aid_y)\n        return (aid_x * aid_y).sum(dim=1)\n\n    def calc_loss(self, aid_x, aid_y, size_x, size_y):\n        rand_idx = torch.randperm(aid_y.size(0))\n        output_pos = self.forward(aid_x, aid_y)\n        output_neg = self.forward(aid_x, aid_y[rand_idx])\n        loss = self.criterion(output_pos, output_neg)\n        return loss\n```\n- Trained aid_embeddings with BPR loss to make co-occurring embeddings become similar\n- What worked\n    - Multiply the inverse of item_size by loss (Removing popularity bias)\n    - Multiply the inverse of ts_diff by loss (The closer the co-occur timing is, the more similar the embeddings become)\n\n<br>\n### User MF (User Matrix Factorization)\n```\nclass UserMFModel(nn.Module):\n    def __init__(self, n_session: int, n_aid: int, n_factors: int):\n        super().__init__()\n        self.n_factors = n_factors\n        self.n_session = n_session\n        self.n_aid = n_aid\n\n        self.session_embeddings = nn.Embedding(self.n_session, self.n_factors)\n        self.aid_embeddings = nn.Embedding(self.n_aid, self.n_factors)\n\n        self.criterion = BPRLoss()\n\n        initrange = 1.0 / self.n_factors\n        nn.init.uniform_(self.session_embeddings.weight.data, -initrange, initrange)\n        nn.init.uniform_(self.aid_embeddings.weight.data, -initrange, initrange)\n\n    def forward(self, session, aid, aid_size):\n        session_emb = self.session_embeddings(session)\n        aid_emb = self.aid_embeddings(aid)\n        return (session_emb * aid_emb).sum(dim=1)\n\n    def calc_loss(self, session, aid, aid_size):\n        rand_idx = torch.randperm(aid.size(0))\n        output_pos = self.forward(session, aid)\n        output_neg = self.forward(session, aid[rand_idx])\n        loss = self.criterion(output_pos, output_neg)\n        return loss\n```\n- Trained session_embeddings and aid_embeddings with BPR loss\n- What worked\n    - Multiply the inverse of item_size by loss (Removing popularity bias)\n    - Multiply the inverse of ts_diff by loss (The closer the co-occur timing is, the more similar the embeddings become)\n\n<br>\n### Item CF\n- Implemented item cf with polars\n- Calculated the similarity weights for each item-item pair and retrieved candidates by getting the most similar items based on sum/min/max/mean of weights\n- What worked\n    - Multiply the inverse of item_size by weight (Removing popularity bias)\n    - Multiply the inverse of ts_diff by weight (The closer the co-occur timing is, the bigger the weight becomes)\n    - Multiply trend coefficient (The more ts is recent, the bigger the weight becomes)\n\n<br>\n### User CF\n- Implemented in the same way as item cf\n\n<br>\n# 2nd Stage - Re-Ranking\n- Feature\n    - Created ~200 features in total\n    - pl.col(’ts’).agg([mean,min,max,std]).over({’session’ or ‘aid’})\n    - candidate_selected_(count, rank)\n    - candidate_(score, rank, selected)\n    - (inter, click, cart, order)_hour_mean.over({’session’ or ‘aid’})\n    - (inter, click, cart, order)_count.over({’session’ or ‘aid’})@(3d, 7d, 14d, 21d)\n    - aid_multi_(inter,click,cart,order)_prob\n- Model\n    - LightGBM Ranker (lambdarank)\n    - CatBoost Classifier (Logloss)\n    - CatBoost Ranker (YetiRank)\n\n<br>\n# What did not work well\n- Transformers\n- GRU\n- CDAE\n- RecVAE\n- Implicit(ALS, BPR)\n- Clustering by item embeddings\n- Popular items\n- Stacking\n- Pseudo Labeling\n\n<br>\n# Lessons Learned\n- polars is all you need to create features\n    - At first, I created features with pandas and cudf but I switched to polars during the competition because polars is fast, memory efficient, and the syntax is easy to understand.\n- Creating a good baseline is very important\n    - It took me 1~2 month to create a baseline pipeline\n    - Thanks to Chris’s great discussion post, I was able to create a 2-stage baseline from scratch (I will publish the code on GitHub soon)\n- Feature Store is useful\n    - I stored intermediate files to parquet files, which saved me to run experiments quickly in a more reproducible way.\n\n<br>\n# Score Timeline\n| CV (1/20 sampled) | CV    | Public LB | Description                   |\n|-------------------|-------|-----------|-------------------------------|\n| 0.536             | 0.569 | 0.5799     | 2-Stage Baseline              |\n| 0.539             | 0.572 | 0.5830     | Added Item2Vec ItemMF         |\n| 0.547             | 0.581 | 0.5930     | Added ItemCF UserMF           |\n| 0.550             | 0.585 | 0.5970     | Added variation to ItemCF     |\n| 0.552             | 0.587 | 0.5985    | Ensemble(LightGBM + CatBoost) |\n\n<br>\n# Feature Importance\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2Febea15581ccf1d44b29d37db8e493fc5%2F68747470733a2f2f71696974612d696d6167652d73746f72652e73332e61702d6e6f727468656173742d312e616d617a6f6e6177732e636f6d2f302f3235393431372f31346465363937652d373162652d396462362d663337362d3066643762336537333630322e706e67.png?generation=1675210821369636&alt=media)",
      "votes": null
    },
    {
      "id": "2124332",
      "postDate": "02/01/2023 00:13:28",
      "content": "<p>Don't forget to link to your solution from the team page.</p>",
      "rawMarkdown": "Don't forget to link to your solution from the team page.",
      "votes": null
    },
    {
      "id": "2124334",
      "postDate": "02/01/2023 00:14:47",
      "content": "<p>I linked it. Thanks!</p>",
      "rawMarkdown": "I linked it. Thanks!",
      "votes": null
    },
    {
      "id": "2124342",
      "postDate": "02/01/2023 00:25:10",
      "content": "<p>Thanks for sharing. May I know what's your public lb only using retrieve methods (i.e. 1st Stage - Candidate Generation)?</p>",
      "rawMarkdown": "Thanks for sharing. May I know what's your public lb only using retrieve methods (i.e. 1st Stage - Candidate Generation)?",
      "votes": null
    },
    {
      "id": "2124346",
      "postDate": "02/01/2023 00:31:17",
      "content": "<p><a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a> Thanks for sharing such detailed insights and lessons :)</p>",
      "rawMarkdown": "kiccho11 Thanks for sharing such detailed insights and lessons :)",
      "votes": null
    },
    {
      "id": "2124358",
      "postDate": "02/01/2023 00:35:12",
      "content": "<p>Thanks a lot for sharing the details, especially the Score Timeline part. I was looking for an understanding like that for a long time, how to break the 579 barrier.</p>\n<p>I also took two months to build a decent baseline, but differently from you I could not evolve much from that.</p>\n<p>A lot of techniques on this solution. Congrats!</p>",
      "rawMarkdown": "Thanks a lot for sharing the details, especially the Score Timeline part. I was looking for an understanding like that for a long time, how to break the 579 barrier.\n\nI also took two months to build a decent baseline, but differently from you I could not evolve much from that.\n\nA lot of techniques on this solution. Congrats!",
      "votes": null
    },
    {
      "id": "2124359",
      "postDate": "02/01/2023 00:35:50",
      "content": "<p>I didn't submit with only retrieval methods but on my local cv, 0.5596(without re-ranking) and 0.585(with re-ranking)</p>",
      "rawMarkdown": "I didn't submit with only retrieval methods but on my local cv, 0.5596(without re-ranking) and 0.585(with re-ranking)",
      "votes": null
    },
    {
      "id": "2124367",
      "postDate": "02/01/2023 00:42:52",
      "content": "<p>Thanks for sharing! May I ask how do you merge all candidates from a bunch of strategy?</p>",
      "rawMarkdown": "Thanks for sharing! May I ask how do you merge all candidates from a bunch of strategy?",
      "votes": null
    },
    {
      "id": "2124369",
      "postDate": "02/01/2023 00:47:02",
      "content": "<p>Well-documented introduction! Thanks for sharing, quick question, how many candidate do you use in the end based on so many recall strategies and how to distribute the number of candidate from each retrieval method? Thanks for your reply! </p>",
      "rawMarkdown": "Well-documented introduction! Thanks for sharing, quick question, how many candidate do you use in the end based on so many recall strategies and how to distribute the number of candidate from each retrieval method? Thanks for your reply!",
      "votes": null
    },
    {
      "id": "2124380",
      "postDate": "02/01/2023 00:59:26",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>! 🥳 And thank you for sharing this great write-up! 🙂</p>",
      "rawMarkdown": "congrats @kiccho11! 🥳 And thank you for sharing this great write-up! 🙂",
      "votes": null
    },
    {
      "id": "2124382",
      "postDate": "02/01/2023 01:05:43",
      "content": "<p>We generated average 300 candidates for each user. For the topk numbers of each reacll method, I could not come up with a good way so I tuned them heuristically just like hyper parameters.</p>",
      "rawMarkdown": "We generated average 300 candidates for each user. For the topk numbers of each reacll method, I could not come up with a good way so I tuned them heuristically just like hyper parameters.",
      "votes": null
    },
    {
      "id": "2124385",
      "postDate": "02/01/2023 01:08:31",
      "content": "<p>We concatenated all candidates vertically and dropped duplicates. Dropping duplicates is memory-heavy so we ran this by chunk. </p>",
      "rawMarkdown": "We concatenated all candidates vertically and dropped duplicates. Dropping duplicates is memory-heavy so we ran this by chunk.",
      "votes": null
    },
    {
      "id": "2124405",
      "postDate": "02/01/2023 01:39:26",
      "content": "<p>very useful, thanks for sharing!  congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a> </p>",
      "rawMarkdown": "very useful, thanks for sharing!  congrats @kiccho11",
      "votes": null
    },
    {
      "id": "2124448",
      "postDate": "02/01/2023 02:25:38",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>, implicit's BPR works well in my side to make the similarity between the candidate items &amp; user session as features.</p>",
      "rawMarkdown": "Congratulations @kiccho11, implicit's BPR works well in my side to make the similarity between the candidate items & user session as features.",
      "votes": null
    },
    {
      "id": "2124514",
      "postDate": "02/01/2023 03:54:03",
      "content": "<p>Congratz on your solo gold medal! I will try BPR for late submission.</p>",
      "rawMarkdown": "Congratz on your solo gold medal! I will try BPR for late submission.",
      "votes": null
    },
    {
      "id": "2124546",
      "postDate": "02/01/2023 04:19:11",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>! Amazing solution! Thanks for sharing this write up.</p>",
      "rawMarkdown": "Congrats @kiccho11! Amazing solution! Thanks for sharing this write up.",
      "votes": null
    },
    {
      "id": "2124547",
      "postDate": "02/01/2023 04:19:46",
      "content": "<p>Congratulations! It seems MF and CF worked well for you guys.</p>",
      "rawMarkdown": "Congratulations! It seems MF and CF worked well for you guys.",
      "votes": null
    },
    {
      "id": "2124559",
      "postDate": "02/01/2023 04:30:51",
      "content": "<p>Yes. ItemCF was the best recall method and ItemMF was the next for us</p>",
      "rawMarkdown": "Yes. ItemCF was the best recall method and ItemMF was the next for us",
      "votes": null
    },
    {
      "id": "2124597",
      "postDate": "02/01/2023 05:03:47",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a> ! and thanks for sharing!</p>\n<p>Do you mind elaborating on these features?</p>\n<ul>\n<li>candidate_selected_(count, rank)</li>\n<li>candidate_(score, rank, selected)</li>\n<li>and the meaning \"inter\" here, (inter, click, cart, order)_hour_mean.over({’session’ or ‘aid’})</li>\n</ul>",
      "rawMarkdown": "Congrats @kiccho11 ! and thanks for sharing!\n\nDo you mind elaborating on these features?\n- candidate_selected_(count, rank)\n- candidate_(score, rank, selected)\n- and the meaning \"inter\" here, (inter, click, cart, order)_hour_mean.over({’session’ or ‘aid’})",
      "votes": null
    },
    {
      "id": "2124599",
      "postDate": "02/01/2023 05:04:41",
      "content": "<p>Great job <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>! Your solution is truly impressive. Thank you for sharing this well-written article with us.</p>",
      "rawMarkdown": "Great job @kiccho11! Your solution is truly impressive. Thank you for sharing this well-written article with us.",
      "votes": null
    },
    {
      "id": "2124708",
      "postDate": "02/01/2023 06:57:19",
      "content": "<ul>\n<li>candidate_selected_count: By how many recall methods the item is recalled</li>\n<li>candidate_selected_rank: The rank of candidate_selected_count over each session</li>\n<li>candidate_score: The score of each candidate generation method. (for example, for ItemMF the score is cosine similarity)</li>\n<li>candidate_rank: The rank of candidate_score over each session</li>\n<li>candidate_selected: 0/1 feature that indicates if the item is recalled by each recall method</li>\n<li>\"inter\" is either click/cart/order. If user did either of click/cart/order action, inter is 1, if not 0 (inter stands for interaction)</li>\n</ul>",
      "rawMarkdown": "candidate_selected_count: By how many recall methods the item is recalled\n- candidate_selected_rank: The rank of candidate_selected_count over each session\n- candidate_score: The score of each candidate generation method. (for example, for ItemMF the score is cosine similarity)\n- candidate_rank: The rank of candidate_score over each session\n- candidate_selected: 0/1 feature that indicates if the item is recalled by each recall method\n- \"inter\" is either click/cart/order. If user did either of click/cart/order action, inter is 1, if not 0 (inter stands for interaction)",
      "votes": null
    },
    {
      "id": "2124976",
      "postDate": "02/01/2023 10:39:52",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>, and thanks for sharing the article describing your solution. I have a question related to the graphs showing the feature's importance. What algorithm did you use to determine the values presented there?</p>",
      "rawMarkdown": "Congrats @kiccho11, and thanks for sharing the article describing your solution. I have a question related to the graphs showing the feature's importance. What algorithm did you use to determine the values presented there?",
      "votes": null
    },
    {
      "id": "2125929",
      "postDate": "02/02/2023 02:36:52",
      "content": "<p>The values are from <code>model.feature_importance</code> of LightGBM.</p>",
      "rawMarkdown": "The values are from `model.feature_importance` of LightGBM.",
      "votes": null
    },
    {
      "id": "2126044",
      "postDate": "02/02/2023 04:15:35",
      "content": "<p>thanks for your kind reply!</p>",
      "rawMarkdown": "thanks for your kind reply!",
      "votes": null
    },
    {
      "id": "2126048",
      "postDate": "02/02/2023 04:21:29",
      "content": "<p>Got it! Thanks for your reply!</p>",
      "rawMarkdown": "Got it! Thanks for your reply!",
      "votes": null
    },
    {
      "id": "2126716",
      "postDate": "02/02/2023 12:13:22",
      "content": "<p>Thanks for your reply. </p>",
      "rawMarkdown": "Thanks for your reply.",
      "votes": null
    },
    {
      "id": "2136504",
      "postDate": "02/09/2023 11:24:10",
      "content": "<p>Hi Kiccho</p>\n<p>Thanks for the write up. Will you be sharing code</p>",
      "rawMarkdown": "Hi Kiccho\n\nThanks for the write up. Will you be sharing code",
      "votes": null
    },
    {
      "id": "2137401",
      "postDate": "02/10/2023 00:25:28",
      "content": "<p>Thanks for reminding. I’m doing some refactoring on my code and will publish soon. </p>",
      "rawMarkdown": "Thanks for reminding. I’m doing some refactoring on my code and will publish soon.",
      "votes": null
    },
    {
      "id": "2140622",
      "postDate": "02/11/2023 23:35:48",
      "content": "<p>I uploaded my code to Github. Have fun!<br>\n<a href=\"https://github.com/kiccho1101/kaggle-otto2\" target=\"_blank\">https://github.com/kiccho1101/kaggle-otto2</a></p>",
      "rawMarkdown": "I uploaded my code to Github. Have fun!\nhttps://github.com/kiccho1101/kaggle-otto2",
      "votes": null
    },
    {
      "id": "2141803",
      "postDate": "02/13/2023 05:19:13",
      "content": "<p>wow, thanks. I think ItemCF would be a good way for me to start. And do you have some other potential way to improve? <br>\nThanks again for sharing! </p>",
      "rawMarkdown": "wow, thanks. I think ItemCF would be a good way for me to start. And do you have some other potential way to improve? \nThanks again for sharing!",
      "votes": null
    },
    {
      "id": "2187993",
      "postDate": "03/19/2023 08:17:11",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a> ! and thanks for sharing! I have problems when i run your code ’ <strong>evaluate_util</strong>' it says ‘<strong>to_numpy not supported for dtype: List/ Boolean</strong>'. I wonder if you have the same problems when u run the code.</p>",
      "rawMarkdown": "Congrats @kiccho11 ! and thanks for sharing! I have problems when i run your code ’ **evaluate_util**' it says ‘**to_numpy not supported for dtype: List/ Boolean**'. I wonder if you have the same problems when u run the code.",
      "votes": null
    },
    {
      "id": "2190761",
      "postDate": "03/21/2023 13:16:00",
      "content": "<p>Hello，kiccho！Thanks for sharing.I recently had trouble running your code on github, what is the Last Inter method in the recall phase?</p>",
      "rawMarkdown": "Hello，kiccho！Thanks for sharing.I recently had trouble running your code on github, what is the Last Inter method in the recall phase?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2124332,
      "author_name": "kaggleqrdl",
      "author_url": "",
      "post_date": "02/01/2023 00:13:28",
      "content": "<p>Don't forget to link to your solution from the team page.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124334,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/01/2023 00:14:47",
          "content": "<p>I linked it. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2124342,
      "author_name": "hookman",
      "author_url": "",
      "post_date": "02/01/2023 00:25:10",
      "content": "<p>Thanks for sharing. May I know what's your public lb only using retrieve methods (i.e. 1st Stage - Candidate Generation)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124359,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/01/2023 00:35:50",
          "content": "<p>I didn't submit with only retrieval methods but on my local cv, 0.5596(without re-ranking) and 0.585(with re-ranking)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2124346,
      "author_name": "minhbtnguyen",
      "author_url": "",
      "post_date": "02/01/2023 00:31:17",
      "content": "<p><a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a> Thanks for sharing such detailed insights and lessons :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124358,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "02/01/2023 00:35:12",
      "content": "<p>Thanks a lot for sharing the details, especially the Score Timeline part. I was looking for an understanding like that for a long time, how to break the 579 barrier.</p>\n<p>I also took two months to build a decent baseline, but differently from you I could not evolve much from that.</p>\n<p>A lot of techniques on this solution. Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124367,
      "author_name": "huyduong7101",
      "author_url": "",
      "post_date": "02/01/2023 00:42:52",
      "content": "<p>Thanks for sharing! May I ask how do you merge all candidates from a bunch of strategy?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124385,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/01/2023 01:08:31",
          "content": "<p>We concatenated all candidates vertically and dropped duplicates. Dropping duplicates is memory-heavy so we ran this by chunk. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2124369,
      "author_name": "sherrypomi",
      "author_url": "",
      "post_date": "02/01/2023 00:47:02",
      "content": "<p>Well-documented introduction! Thanks for sharing, quick question, how many candidate do you use in the end based on so many recall strategies and how to distribute the number of candidate from each retrieval method? Thanks for your reply! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2124382,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/01/2023 01:05:43",
          "content": "<p>We generated average 300 candidates for each user. For the topk numbers of each reacll method, I could not come up with a good way so I tuned them heuristically just like hyper parameters.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2126048,
              "author_name": "sherrypomi",
              "author_url": "",
              "post_date": "02/02/2023 04:21:29",
              "content": "<p>Got it! Thanks for your reply!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2124380,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "02/01/2023 00:59:26",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>! 🥳 And thank you for sharing this great write-up! 🙂</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124405,
      "author_name": "elsabc",
      "author_url": "",
      "post_date": "02/01/2023 01:39:26",
      "content": "<p>very useful, thanks for sharing!  congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124448,
      "author_name": "gongbi",
      "author_url": "",
      "post_date": "02/01/2023 02:25:38",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>, implicit's BPR works well in my side to make the similarity between the candidate items &amp; user session as features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124514,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/01/2023 03:54:03",
          "content": "<p>Congratz on your solo gold medal! I will try BPR for late submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2124546,
      "author_name": "ravishah1",
      "author_url": "",
      "post_date": "02/01/2023 04:19:11",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>! Amazing solution! Thanks for sharing this write up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124547,
      "author_name": "nlztrk",
      "author_url": "",
      "post_date": "02/01/2023 04:19:46",
      "content": "<p>Congratulations! It seems MF and CF worked well for you guys.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2124559,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/01/2023 04:30:51",
          "content": "<p>Yes. ItemCF was the best recall method and ItemMF was the next for us</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2124597,
      "author_name": "ajisamudra",
      "author_url": "",
      "post_date": "02/01/2023 05:03:47",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a> ! and thanks for sharing!</p>\n<p>Do you mind elaborating on these features?</p>\n<ul>\n<li>candidate_selected_(count, rank)</li>\n<li>candidate_(score, rank, selected)</li>\n<li>and the meaning \"inter\" here, (inter, click, cart, order)_hour_mean.over({’session’ or ‘aid’})</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2124708,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/01/2023 06:57:19",
          "content": "<ul>\n<li>candidate_selected_count: By how many recall methods the item is recalled</li>\n<li>candidate_selected_rank: The rank of candidate_selected_count over each session</li>\n<li>candidate_score: The score of each candidate generation method. (for example, for ItemMF the score is cosine similarity)</li>\n<li>candidate_rank: The rank of candidate_score over each session</li>\n<li>candidate_selected: 0/1 feature that indicates if the item is recalled by each recall method</li>\n<li>\"inter\" is either click/cart/order. If user did either of click/cart/order action, inter is 1, if not 0 (inter stands for interaction)</li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2126044,
              "author_name": "ajisamudra",
              "author_url": "",
              "post_date": "02/02/2023 04:15:35",
              "content": "<p>thanks for your kind reply!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2124599,
      "author_name": "tahamhaider",
      "author_url": "",
      "post_date": "02/01/2023 05:04:41",
      "content": "<p>Great job <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>! Your solution is truly impressive. Thank you for sharing this well-written article with us.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124976,
      "author_name": "balaganiarz0",
      "author_url": "",
      "post_date": "02/01/2023 10:39:52",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a>, and thanks for sharing the article describing your solution. I have a question related to the graphs showing the feature's importance. What algorithm did you use to determine the values presented there?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125929,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/02/2023 02:36:52",
          "content": "<p>The values are from <code>model.feature_importance</code> of LightGBM.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2126716,
              "author_name": "balaganiarz0",
              "author_url": "",
              "post_date": "02/02/2023 12:13:22",
              "content": "<p>Thanks for your reply. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2136504,
      "author_name": "purohit",
      "author_url": "",
      "post_date": "02/09/2023 11:24:10",
      "content": "<p>Hi Kiccho</p>\n<p>Thanks for the write up. Will you be sharing code</p>",
      "votes": null,
      "replies": [
        {
          "id": 2137401,
          "author_name": "kiccho11",
          "author_url": "",
          "post_date": "02/10/2023 00:25:28",
          "content": "<p>Thanks for reminding. I’m doing some refactoring on my code and will publish soon. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2140622,
      "author_name": "kiccho11",
      "author_url": "",
      "post_date": "02/11/2023 23:35:48",
      "content": "<p>I uploaded my code to Github. Have fun!<br>\n<a href=\"https://github.com/kiccho1101/kaggle-otto2\" target=\"_blank\">https://github.com/kiccho1101/kaggle-otto2</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2141803,
      "author_name": "giranntu",
      "author_url": "",
      "post_date": "02/13/2023 05:19:13",
      "content": "<p>wow, thanks. I think ItemCF would be a good way for me to start. And do you have some other potential way to improve? <br>\nThanks again for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2187993,
      "author_name": "kossin",
      "author_url": "",
      "post_date": "03/19/2023 08:17:11",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/kiccho11\" target=\"_blank\">@kiccho11</a> ! and thanks for sharing! I have problems when i run your code ’ <strong>evaluate_util</strong>' it says ‘<strong>to_numpy not supported for dtype: List/ Boolean</strong>'. I wonder if you have the same problems when u run the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2190761,
      "author_name": "yxy991028",
      "author_url": "",
      "post_date": "03/21/2023 13:16:00",
      "content": "<p>Hello，kiccho！Thanks for sharing.I recently had trouble running your code on github, what is the Last Inter method in the recall phase?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2124330": "# Acknowledgements\nFirst, I would like to thank the organizers and those who shared knowledge. Especially I would like to thank @cdeotte and @radek1 for posting so many codes and discussions. I learned a lot from you. \n\nConsidering the controversy about the possible cheaters, I will publish the entire code after everything is finalized and I will just share the ideas in this solution. (Although [Will says sharing code is not something we should withhold](https://www.kaggle.com/competitions/otto-recommender-system/discussion/382440#2123460))\n\n\n<br>\n# Solution Summary\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2F438095d4971fd2b40113d40cf1d96fc1%2FScreen%20Shot%202023-02-01%20at%209.28.52.png?generation=1675211352995855&alt=media)\n\nMy code is here: https://github.com/kiccho1101/kaggle-otto2\n\n<br>\n# CV Strategy\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2F728b7a7f2c1ef5eeec8207ac1e7f870f%2Frecommend-data-split.svg?generation=1675210198356258&alt=media)\n- I used Radek’s CV strategy\n- Although CV data was good for the local experiments, it took quite a long time to run experiments so I used only 1/20 sessions by random sampling. As a result, the average time that takes to run 1 experiment became about 10~30mins which allowed me to run many experiments quickly.\n- Even if I sampled the data, the CV-LB score correlation was very stable.\n- The data used for each stage is below\n\n| Step                       | CV                     | LB                                              |\n|----------------------------|------------------------|-------------------------------------------------|\n| Candidate Generation       | TrainData + ValidDataA | TrainData + ValidDataA + ValidDataB + TestDataA |\n| User Feature Creation      | ValidDataA             | TestDataA                                       |\n| Item Feature Creation      | TrainData + ValidDataA | TrainData + ValidDataA + ValidDataB + TestDataA |\n| User-Item Feature Creation | ValidDataA             | TestDataA                                       |\n| Re-Ranking                 | ValidDataA             | TestDataA                                       |\n\n<br>\n# 1st Stage - Candidate Generation\n\n<br>\n### word2vec\n- Trained word2vec model with aid sequences\n- Retrieve top-k aids by word2vec embeddings (Used faiss-gpu to speed up)\n\n<br>\n### CoVis\n- I used Chris’s co-visitation matrix to generate candidates\n\n<br>\n### Item MF (Item Matrix Factorization)\n```\nclass ItemMFModel(nn.Module):\n    def __init__(self, n_aid: int, n_factors: int):\n        super().__init__()\n        self.criterion = BPRLoss()\n        self.n_factors = n_factors\n        self.n_aid = n_aid\n        self.aid_embeddings = nn.Embedding(self.n_aid, self.n_factors)\n\n        initrange = 1.0 / self.n_factors\n        nn.init.uniform_(self.aid_embeddings.weight.data, -initrange, initrange)\n\n    def forward(self, aid_x, aid_y):\n        aid_x = self.aid_embeddings(aid_x)\n        aid_y = self.aid_embeddings(aid_y)\n        return (aid_x * aid_y).sum(dim=1)\n\n    def calc_loss(self, aid_x, aid_y, size_x, size_y):\n        rand_idx = torch.randperm(aid_y.size(0))\n        output_pos = self.forward(aid_x, aid_y)\n        output_neg = self.forward(aid_x, aid_y[rand_idx])\n        loss = self.criterion(output_pos, output_neg)\n        return loss\n```\n- Trained aid_embeddings with BPR loss to make co-occurring embeddings become similar\n- What worked\n    - Multiply the inverse of item_size by loss (Removing popularity bias)\n    - Multiply the inverse of ts_diff by loss (The closer the co-occur timing is, the more similar the embeddings become)\n\n<br>\n### User MF (User Matrix Factorization)\n```\nclass UserMFModel(nn.Module):\n    def __init__(self, n_session: int, n_aid: int, n_factors: int):\n        super().__init__()\n        self.n_factors = n_factors\n        self.n_session = n_session\n        self.n_aid = n_aid\n\n        self.session_embeddings = nn.Embedding(self.n_session, self.n_factors)\n        self.aid_embeddings = nn.Embedding(self.n_aid, self.n_factors)\n\n        self.criterion = BPRLoss()\n\n        initrange = 1.0 / self.n_factors\n        nn.init.uniform_(self.session_embeddings.weight.data, -initrange, initrange)\n        nn.init.uniform_(self.aid_embeddings.weight.data, -initrange, initrange)\n\n    def forward(self, session, aid, aid_size):\n        session_emb = self.session_embeddings(session)\n        aid_emb = self.aid_embeddings(aid)\n        return (session_emb * aid_emb).sum(dim=1)\n\n    def calc_loss(self, session, aid, aid_size):\n        rand_idx = torch.randperm(aid.size(0))\n        output_pos = self.forward(session, aid)\n        output_neg = self.forward(session, aid[rand_idx])\n        loss = self.criterion(output_pos, output_neg)\n        return loss\n```\n- Trained session_embeddings and aid_embeddings with BPR loss\n- What worked\n    - Multiply the inverse of item_size by loss (Removing popularity bias)\n    - Multiply the inverse of ts_diff by loss (The closer the co-occur timing is, the more similar the embeddings become)\n\n<br>\n### Item CF\n- Implemented item cf with polars\n- Calculated the similarity weights for each item-item pair and retrieved candidates by getting the most similar items based on sum/min/max/mean of weights\n- What worked\n    - Multiply the inverse of item_size by weight (Removing popularity bias)\n    - Multiply the inverse of ts_diff by weight (The closer the co-occur timing is, the bigger the weight becomes)\n    - Multiply trend coefficient (The more ts is recent, the bigger the weight becomes)\n\n<br>\n### User CF\n- Implemented in the same way as item cf\n\n<br>\n# 2nd Stage - Re-Ranking\n- Feature\n    - Created ~200 features in total\n    - pl.col(’ts’).agg([mean,min,max,std]).over({’session’ or ‘aid’})\n    - candidate_selected_(count, rank)\n    - candidate_(score, rank, selected)\n    - (inter, click, cart, order)_hour_mean.over({’session’ or ‘aid’})\n    - (inter, click, cart, order)_count.over({’session’ or ‘aid’})@(3d, 7d, 14d, 21d)\n    - aid_multi_(inter,click,cart,order)_prob\n- Model\n    - LightGBM Ranker (lambdarank)\n    - CatBoost Classifier (Logloss)\n    - CatBoost Ranker (YetiRank)\n\n<br>\n# What did not work well\n- Transformers\n- GRU\n- CDAE\n- RecVAE\n- Implicit(ALS, BPR)\n- Clustering by item embeddings\n- Popular items\n- Stacking\n- Pseudo Labeling\n\n<br>\n# Lessons Learned\n- polars is all you need to create features\n    - At first, I created features with pandas and cudf but I switched to polars during the competition because polars is fast, memory efficient, and the syntax is easy to understand.\n- Creating a good baseline is very important\n    - It took me 1~2 month to create a baseline pipeline\n    - Thanks to Chris’s great discussion post, I was able to create a 2-stage baseline from scratch (I will publish the code on GitHub soon)\n- Feature Store is useful\n    - I stored intermediate files to parquet files, which saved me to run experiments quickly in a more reproducible way.\n\n<br>\n# Score Timeline\n| CV (1/20 sampled) | CV    | Public LB | Description                   |\n|-------------------|-------|-----------|-------------------------------|\n| 0.536             | 0.569 | 0.5799     | 2-Stage Baseline              |\n| 0.539             | 0.572 | 0.5830     | Added Item2Vec ItemMF         |\n| 0.547             | 0.581 | 0.5930     | Added ItemCF UserMF           |\n| 0.550             | 0.585 | 0.5970     | Added variation to ItemCF     |\n| 0.552             | 0.587 | 0.5985    | Ensemble(LightGBM + CatBoost) |\n\n<br>\n# Feature Importance\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F770256%2Febea15581ccf1d44b29d37db8e493fc5%2F68747470733a2f2f71696974612d696d6167652d73746f72652e73332e61702d6e6f727468656173742d312e616d617a6f6e6177732e636f6d2f302f3235393431372f31346465363937652d373162652d396462362d663337362d3066643762336537333630322e706e67.png?generation=1675210821369636&alt=media)",
    "2124332": "Don't forget to link to your solution from the team page.",
    "2124334": "I linked it. Thanks!",
    "2124342": "Thanks for sharing. May I know what's your public lb only using retrieve methods (i.e. 1st Stage - Candidate Generation)?",
    "2124346": "kiccho11 Thanks for sharing such detailed insights and lessons :)",
    "2124358": "Thanks a lot for sharing the details, especially the Score Timeline part. I was looking for an understanding like that for a long time, how to break the 579 barrier.\n\nI also took two months to build a decent baseline, but differently from you I could not evolve much from that.\n\nA lot of techniques on this solution. Congrats!",
    "2124359": "I didn't submit with only retrieval methods but on my local cv, 0.5596(without re-ranking) and 0.585(with re-ranking)",
    "2124367": "Thanks for sharing! May I ask how do you merge all candidates from a bunch of strategy?",
    "2124369": "Well-documented introduction! Thanks for sharing, quick question, how many candidate do you use in the end based on so many recall strategies and how to distribute the number of candidate from each retrieval method? Thanks for your reply!",
    "2124380": "congrats @kiccho11! 🥳 And thank you for sharing this great write-up! 🙂",
    "2124382": "We generated average 300 candidates for each user. For the topk numbers of each reacll method, I could not come up with a good way so I tuned them heuristically just like hyper parameters.",
    "2124385": "We concatenated all candidates vertically and dropped duplicates. Dropping duplicates is memory-heavy so we ran this by chunk.",
    "2124405": "very useful, thanks for sharing!  congrats @kiccho11",
    "2124448": "Congratulations @kiccho11, implicit's BPR works well in my side to make the similarity between the candidate items & user session as features.",
    "2124514": "Congratz on your solo gold medal! I will try BPR for late submission.",
    "2124546": "Congrats @kiccho11! Amazing solution! Thanks for sharing this write up.",
    "2124547": "Congratulations! It seems MF and CF worked well for you guys.",
    "2124559": "Yes. ItemCF was the best recall method and ItemMF was the next for us",
    "2124597": "Congrats @kiccho11 ! and thanks for sharing!\n\nDo you mind elaborating on these features?\n- candidate_selected_(count, rank)\n- candidate_(score, rank, selected)\n- and the meaning \"inter\" here, (inter, click, cart, order)_hour_mean.over({’session’ or ‘aid’})",
    "2124599": "Great job @kiccho11! Your solution is truly impressive. Thank you for sharing this well-written article with us.",
    "2124708": "candidate_selected_count: By how many recall methods the item is recalled\n- candidate_selected_rank: The rank of candidate_selected_count over each session\n- candidate_score: The score of each candidate generation method. (for example, for ItemMF the score is cosine similarity)\n- candidate_rank: The rank of candidate_score over each session\n- candidate_selected: 0/1 feature that indicates if the item is recalled by each recall method\n- \"inter\" is either click/cart/order. If user did either of click/cart/order action, inter is 1, if not 0 (inter stands for interaction)",
    "2124976": "Congrats @kiccho11, and thanks for sharing the article describing your solution. I have a question related to the graphs showing the feature's importance. What algorithm did you use to determine the values presented there?",
    "2125929": "The values are from `model.feature_importance` of LightGBM.",
    "2126044": "thanks for your kind reply!",
    "2126048": "Got it! Thanks for your reply!",
    "2126716": "Thanks for your reply.",
    "2136504": "Hi Kiccho\n\nThanks for the write up. Will you be sharing code",
    "2137401": "Thanks for reminding. I’m doing some refactoring on my code and will publish soon.",
    "2140622": "I uploaded my code to Github. Have fun!\nhttps://github.com/kiccho1101/kaggle-otto2",
    "2141803": "wow, thanks. I think ItemCF would be a good way for me to start. And do you have some other potential way to improve? \nThanks again for sharing!",
    "2187993": "Congrats @kiccho11 ! and thanks for sharing! I have problems when i run your code ’ **evaluate_util**' it says ‘**to_numpy not supported for dtype: List/ Boolean**'. I wonder if you have the same problems when u run the code.",
    "2190761": "Hello，kiccho！Thanks for sharing.I recently had trouble running your code on github, what is the Last Inter method in the recall phase?"
  },
  "source": "meta"
}