{
  "id": 324127,
  "title": "9th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/saber-9th-place-solution",
  "author_name": "",
  "post_date": "2022-05-10T08:35:00.148497800Z",
  "votes": 50,
  "comment_count": 16,
  "views": 0,
  "content": "<p>It's too long since I last see the competition about recommendations on Kaggle. So great thanks to the Kaggle team and H&amp;M for hosting such an interesting competition. The high correlation between CV and LB, big training data for deep data mining, no leaking of test information, friendly staff and participants, all of these make the successful competition.</p>\n<h1>Summary</h1>\n<p>The key idea of my solution will be illustrated here for fast reading purposes.</p>\n<h5>Validation strategy</h5>\n<pre><code>train = transactions.loc[transactions[\"t_dat\"] &lt; \"2020-09-16\"]\nvalid = transactions.loc[transactions[\"t_dat\"] &gt;= \"2020-09-16\"]\n</code></pre>\n<p>It is better to use k-fold validation in <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919\" target=\"_blank\">this post</a> to avoid overfitting the last week if you have enough computing power.</p>\n<p>But from my experience in this competition, LB score and CV score are highly correlated using last week only as a validation set.</p>\n<h5>Recall</h5>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SAR</td>\n<td><a href=\"https://github.com/microsoft/recommenders/blob/main/examples/00_quick_start/sar_movielens.ipynb\" target=\"_blank\">notebook</a></td>\n</tr>\n<tr>\n<td>Trending</td>\n<td><a href=\"https://www.kaggle.com/code/byfone/h-m-trending-products-weekly\" target=\"_blank\">notebook</a></td>\n</tr>\n<tr>\n<td>Most popular items based on age</td>\n<td>/</td>\n</tr>\n</tbody>\n</table>\n<h5>Ranking</h5>\n<p>LightGBM trained with 20 weeks of data using around 300 features</p>\n<h5>Ensemble</h5>\n<p>I have no time to train all 5 models with k fold split on the train set, only 3 of them are completed before the deadline. So my final submission is based on 1 model trained on all data samples and the other 3 models trained based on k fold split mentioned before (4 models in total).</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>public score</th>\n<th>private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>one model</td>\n<td>0.0338</td>\n<td>0.0341</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>0.0343</td>\n<td>0.0346</td>\n</tr>\n</tbody>\n</table>\n<h5>Optimization</h5>\n<p>Using cudf and Forest Inference Library from Rapids for fast feature generation and fast tree inference</p>\n<h5>Hardware</h5>\n<table>\n<thead>\n<tr>\n<th>Hardware</th>\n<th>value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CPU</td>\n<td>32core</td>\n</tr>\n<tr>\n<td>Memory</td>\n<td>128G</td>\n</tr>\n<tr>\n<td>GPU</td>\n<td>V100 32G</td>\n</tr>\n</tbody>\n</table>\n<h1>Data Analysis</h1>\n<ol>\n<li>We can find 90%+ items in the validation set also in the last 30 days</li>\n</ol>\n<pre><code>train_condition = (trans[\"t_dat\"] &gt;= \"2020-08-16\") &amp; (trans[\"t_dat\"] &lt; \"2020-09-16\")\nitems_in_valid = trans.loc[trans[\"t_dat\"] &gt;= \"2020-09-16\", \"article_id\"].unique()\nitems_last_month = trans.loc[train_condition, \"article_id\"].unique()\nprint(len(set(items_in_valid) &amp; set(items_last_month)) / len(set(items_in_valid)))\n# 0.92\n</code></pre>\n<ol>\n<li>Approximately 50% of customers in the validation set have transaction records in the last 30 days</li>\n</ol>\n<pre><code>train_condition = (trans[\"t_dat\"] &gt;= \"2020-08-16\") &amp; (trans[\"t_dat\"] &lt; \"2020-09-16\")\nusers_in_valid = trans.loc[trans[\"t_dat\"] &gt;= \"2020-09-16\", \"customer_id\"].unique()\nusers_last_month = trans.loc[train_condition, \"customer_id\"].unique()\nprint(len(set(users_in_valid) &amp; set(users_last_month)) / len(set(users_in_valid)))\n# 0.47\n</code></pre>\n<p>Considering these results, we can focus on recommending items bought in the last 30 days, especially for those customers who have transaction records.</p>\n<h1>Recall</h1>\n<p>Only mapk of each recall method on the local validation set is recorded. Most of these ideas are coming from public notebooks or discussions and I will give the link if so.</p>\n<h5>Most popular items last 7 days</h5>\n<p>Simply sorting items by transaction records. This is also the basic rule for those customers has no age data</p>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.00814889</td>\n</tr>\n</tbody>\n</table>\n<h5>Most popular itmes last 7 days by age</h5>\n<p>We are doing personalized recommendations, so it's better to use some customers' features based on the rule for recall. Unfortunately, only 2 features seem valuable to me at first glance, age, and postal code. It's impossible to get the location feature from the postal code since it is encoded. Finally, the only choice is age.</p>\n<pre><code>bins = [0, 18, 22, 28, 35, 45, 55, 65, 200]\nlast_week_transactions[\"age_bin\"] = pd.cut(user[\"age\"], bins=bins)\nlast_week_transactions_popular = last_week_transactions.groupby(\"age_bin\").size()\n</code></pre>\n<p>Bins are simply tuned with last week's data. The mapk score is slightly higher than only using the most popular items.</p>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.0081488</td>\n</tr>\n<tr>\n<td>Top 12 items by age</td>\n<td>0.0091866</td>\n</tr>\n</tbody>\n</table>\n<h5>Consider items that customers have bought before</h5>\n<p>The first method is just adding items that customers have bought in the last 4 weeks. Similar to <a href=\"https://www.kaggle.com/code/hengzheng/time-is-our-best-friend-v2\" target=\"_blank\">this notebook</a></p>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.0081488</td>\n</tr>\n<tr>\n<td>Top 12 items by age</td>\n<td>0.0091866</td>\n</tr>\n<tr>\n<td>bought last 4 week +Top 12 items by age</td>\n<td>0.0245722</td>\n</tr>\n</tbody>\n</table>\n<h5>Use Trending</h5>\n<p>There is a significant increase in mapk. But since we know only half of the customers have transaction history in the last 30 days, we can do better. According to <a href=\"https://www.kaggle.com/code/byfone/h-m-trending-products-weekly\" target=\"_blank\">this notebook</a>, we have a wonderful method to compute the score of each item to all customers without the limitation of 30 days.</p>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.0081488</td>\n</tr>\n<tr>\n<td>Top 12 items by age</td>\n<td>0.0091866</td>\n</tr>\n<tr>\n<td>bought last 4 week +Top 12 items by age</td>\n<td>0.0245722</td>\n</tr>\n<tr>\n<td>Trending + Top 12 items by age</td>\n<td>0.0255676</td>\n</tr>\n</tbody>\n</table>\n<h5>Recommend with SAR</h5>\n<p>SAR is a kind of item-based collaborative filtering. As mentioned in <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316205\" target=\"_blank\">this post</a>, we can get 0.27+ with SAR. Here is the <a href=\"https://github.com/microsoft/recommenders/blob/main/examples/00_quick_start/sar_movielens.ipynb\" target=\"_blank\">example</a> from the Microsoft official repo. My SAR is trained with the last 30 days. Adding more data will make the performance worse in my case.</p>\n<table>\n<thead>\n<tr>\n<th>experiment</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.0081488</td>\n</tr>\n<tr>\n<td>Top 12 items by age</td>\n<td>0.0091866</td>\n</tr>\n<tr>\n<td>bought last 4 week +Top 12 items by age</td>\n<td>0.0245722</td>\n</tr>\n<tr>\n<td>Trending + Top 12 items by age</td>\n<td>0.0255676</td>\n</tr>\n<tr>\n<td>SAR + Trending + Top 12 items by age</td>\n<td>0.0278576</td>\n</tr>\n</tbody>\n</table>\n<h5>Overall recall strategy</h5>\n<p>Let's say we are going to recommend K items for each customer. My Final strategy for the recall is</p>\n<ol>\n<li>Recommend with SAR if the customer has a transaction in the last 30 days</li>\n<li>If the number of candidates is less than K, pick items from the trending method (remove items that have already been in the first step)</li>\n<li>If the number of candidates is still less than K, fill with most popular items by age (using most popular items overall if the age data is missing)</li>\n<li>Each recall method has a score and we will use them as features in the ranking phase.</li>\n</ol>\n<h1>Ranking</h1>\n<p>LightGBM classifier is used as a ranking model. There is no significant difference between classifier and lambda ranker. But from my previous experience, the classifier will give a slightly higher performance in general. So I haven't tried lambda ranker in this competition, but it's a good way for ensembling.</p>\n<h5>Aggregation features</h5>\n<p>Nothing special here, just group by user/item, (user, item) and do some aggregation.</p>\n<table>\n<thead>\n<tr>\n<th>Feature</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>user feature</td>\n<td>average prices of items this customer has bought in the last 7/30/90 days</td>\n</tr>\n<tr>\n<td>item feature</td>\n<td>average age of customers who have bought this item in the last 7/30/90 days</td>\n</tr>\n<tr>\n<td>time</td>\n<td>how many days after the customer's last transaction record</td>\n</tr>\n<tr>\n<td>interaction</td>\n<td>whether the customer has bought the current item before</td>\n</tr>\n<tr>\n<td>others</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<h5>Embedding feature</h5>\n<p>We can treat the transaction history of a customer as a sentence, then use some NLP technique to pretrain the embedding of customer/item. My solution using word2vec to train the embedding and simply averaging the item embedding of a customer transaction history as one of the customer features (be care of leaking problems, do not use items in the current week)</p>\n<h5>CV vs LB</h5>\n<p>train with the last 20 weeks and generate 100 candidates for each customer (a bug was fixed in the submission phase several hours before the deadline, so the gap should be smaller)</p>\n<table>\n<thead>\n<tr>\n<th>validation</th>\n<th>score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>local cv</td>\n<td>0.0363</td>\n</tr>\n<tr>\n<td>public score</td>\n<td>0.0310</td>\n</tr>\n<tr>\n<td>private score</td>\n<td>0.0312</td>\n</tr>\n</tbody>\n</table>\n<h1>Optimization</h1>\n<p>Feature engineering with cudf is 30X faster than using pandas in my case. It allows me to generate more candidates and get the feature in the submission phase. Load LightGBM model with Forest Inference Library and inference with GPU is 40X faster than using 32 core CPUs.</p>\n<h1>Final Train &amp; Inference</h1>\n<ul>\n<li>Train</li>\n</ul>\n<p>Generate 200 candidates for each customer, keeping all positive samples and random choosing half of the negative samples in the last 20 weeks</p>\n<ul>\n<li>Inference</li>\n</ul>\n<p>Generate 400 candidates for each customer, making prediction scores with 4 LightGBM models. Averaging them as the final score and getting top12 items</p>",
  "messages": [
    {
      "id": "1783261",
      "postDate": "05/10/2022 08:35:00",
      "content": "<p>It's too long since I last see the competition about recommendations on Kaggle. So great thanks to the Kaggle team and H&amp;M for hosting such an interesting competition. The high correlation between CV and LB, big training data for deep data mining, no leaking of test information, friendly staff and participants, all of these make the successful competition.</p>\n<h1>Summary</h1>\n<p>The key idea of my solution will be illustrated here for fast reading purposes.</p>\n<h5>Validation strategy</h5>\n<pre><code>train = transactions.loc[transactions[\"t_dat\"] &lt; \"2020-09-16\"]\nvalid = transactions.loc[transactions[\"t_dat\"] &gt;= \"2020-09-16\"]\n</code></pre>\n<p>It is better to use k-fold validation in <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919\" target=\"_blank\">this post</a> to avoid overfitting the last week if you have enough computing power.</p>\n<p>But from my experience in this competition, LB score and CV score are highly correlated using last week only as a validation set.</p>\n<h5>Recall</h5>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>link</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SAR</td>\n<td><a href=\"https://github.com/microsoft/recommenders/blob/main/examples/00_quick_start/sar_movielens.ipynb\" target=\"_blank\">notebook</a></td>\n</tr>\n<tr>\n<td>Trending</td>\n<td><a href=\"https://www.kaggle.com/code/byfone/h-m-trending-products-weekly\" target=\"_blank\">notebook</a></td>\n</tr>\n<tr>\n<td>Most popular items based on age</td>\n<td>/</td>\n</tr>\n</tbody>\n</table>\n<h5>Ranking</h5>\n<p>LightGBM trained with 20 weeks of data using around 300 features</p>\n<h5>Ensemble</h5>\n<p>I have no time to train all 5 models with k fold split on the train set, only 3 of them are completed before the deadline. So my final submission is based on 1 model trained on all data samples and the other 3 models trained based on k fold split mentioned before (4 models in total).</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>public score</th>\n<th>private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>one model</td>\n<td>0.0338</td>\n<td>0.0341</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>0.0343</td>\n<td>0.0346</td>\n</tr>\n</tbody>\n</table>\n<h5>Optimization</h5>\n<p>Using cudf and Forest Inference Library from Rapids for fast feature generation and fast tree inference</p>\n<h5>Hardware</h5>\n<table>\n<thead>\n<tr>\n<th>Hardware</th>\n<th>value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CPU</td>\n<td>32core</td>\n</tr>\n<tr>\n<td>Memory</td>\n<td>128G</td>\n</tr>\n<tr>\n<td>GPU</td>\n<td>V100 32G</td>\n</tr>\n</tbody>\n</table>\n<h1>Data Analysis</h1>\n<ol>\n<li>We can find 90%+ items in the validation set also in the last 30 days</li>\n</ol>\n<pre><code>train_condition = (trans[\"t_dat\"] &gt;= \"2020-08-16\") &amp; (trans[\"t_dat\"] &lt; \"2020-09-16\")\nitems_in_valid = trans.loc[trans[\"t_dat\"] &gt;= \"2020-09-16\", \"article_id\"].unique()\nitems_last_month = trans.loc[train_condition, \"article_id\"].unique()\nprint(len(set(items_in_valid) &amp; set(items_last_month)) / len(set(items_in_valid)))\n# 0.92\n</code></pre>\n<ol>\n<li>Approximately 50% of customers in the validation set have transaction records in the last 30 days</li>\n</ol>\n<pre><code>train_condition = (trans[\"t_dat\"] &gt;= \"2020-08-16\") &amp; (trans[\"t_dat\"] &lt; \"2020-09-16\")\nusers_in_valid = trans.loc[trans[\"t_dat\"] &gt;= \"2020-09-16\", \"customer_id\"].unique()\nusers_last_month = trans.loc[train_condition, \"customer_id\"].unique()\nprint(len(set(users_in_valid) &amp; set(users_last_month)) / len(set(users_in_valid)))\n# 0.47\n</code></pre>\n<p>Considering these results, we can focus on recommending items bought in the last 30 days, especially for those customers who have transaction records.</p>\n<h1>Recall</h1>\n<p>Only mapk of each recall method on the local validation set is recorded. Most of these ideas are coming from public notebooks or discussions and I will give the link if so.</p>\n<h5>Most popular items last 7 days</h5>\n<p>Simply sorting items by transaction records. This is also the basic rule for those customers has no age data</p>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.00814889</td>\n</tr>\n</tbody>\n</table>\n<h5>Most popular itmes last 7 days by age</h5>\n<p>We are doing personalized recommendations, so it's better to use some customers' features based on the rule for recall. Unfortunately, only 2 features seem valuable to me at first glance, age, and postal code. It's impossible to get the location feature from the postal code since it is encoded. Finally, the only choice is age.</p>\n<pre><code>bins = [0, 18, 22, 28, 35, 45, 55, 65, 200]\nlast_week_transactions[\"age_bin\"] = pd.cut(user[\"age\"], bins=bins)\nlast_week_transactions_popular = last_week_transactions.groupby(\"age_bin\").size()\n</code></pre>\n<p>Bins are simply tuned with last week's data. The mapk score is slightly higher than only using the most popular items.</p>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.0081488</td>\n</tr>\n<tr>\n<td>Top 12 items by age</td>\n<td>0.0091866</td>\n</tr>\n</tbody>\n</table>\n<h5>Consider items that customers have bought before</h5>\n<p>The first method is just adding items that customers have bought in the last 4 weeks. Similar to <a href=\"https://www.kaggle.com/code/hengzheng/time-is-our-best-friend-v2\" target=\"_blank\">this notebook</a></p>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.0081488</td>\n</tr>\n<tr>\n<td>Top 12 items by age</td>\n<td>0.0091866</td>\n</tr>\n<tr>\n<td>bought last 4 week +Top 12 items by age</td>\n<td>0.0245722</td>\n</tr>\n</tbody>\n</table>\n<h5>Use Trending</h5>\n<p>There is a significant increase in mapk. But since we know only half of the customers have transaction history in the last 30 days, we can do better. According to <a href=\"https://www.kaggle.com/code/byfone/h-m-trending-products-weekly\" target=\"_blank\">this notebook</a>, we have a wonderful method to compute the score of each item to all customers without the limitation of 30 days.</p>\n<table>\n<thead>\n<tr>\n<th>Recall method</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.0081488</td>\n</tr>\n<tr>\n<td>Top 12 items by age</td>\n<td>0.0091866</td>\n</tr>\n<tr>\n<td>bought last 4 week +Top 12 items by age</td>\n<td>0.0245722</td>\n</tr>\n<tr>\n<td>Trending + Top 12 items by age</td>\n<td>0.0255676</td>\n</tr>\n</tbody>\n</table>\n<h5>Recommend with SAR</h5>\n<p>SAR is a kind of item-based collaborative filtering. As mentioned in <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316205\" target=\"_blank\">this post</a>, we can get 0.27+ with SAR. Here is the <a href=\"https://github.com/microsoft/recommenders/blob/main/examples/00_quick_start/sar_movielens.ipynb\" target=\"_blank\">example</a> from the Microsoft official repo. My SAR is trained with the last 30 days. Adding more data will make the performance worse in my case.</p>\n<table>\n<thead>\n<tr>\n<th>experiment</th>\n<th>mapk</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Top 12 items</td>\n<td>0.0081488</td>\n</tr>\n<tr>\n<td>Top 12 items by age</td>\n<td>0.0091866</td>\n</tr>\n<tr>\n<td>bought last 4 week +Top 12 items by age</td>\n<td>0.0245722</td>\n</tr>\n<tr>\n<td>Trending + Top 12 items by age</td>\n<td>0.0255676</td>\n</tr>\n<tr>\n<td>SAR + Trending + Top 12 items by age</td>\n<td>0.0278576</td>\n</tr>\n</tbody>\n</table>\n<h5>Overall recall strategy</h5>\n<p>Let's say we are going to recommend K items for each customer. My Final strategy for the recall is</p>\n<ol>\n<li>Recommend with SAR if the customer has a transaction in the last 30 days</li>\n<li>If the number of candidates is less than K, pick items from the trending method (remove items that have already been in the first step)</li>\n<li>If the number of candidates is still less than K, fill with most popular items by age (using most popular items overall if the age data is missing)</li>\n<li>Each recall method has a score and we will use them as features in the ranking phase.</li>\n</ol>\n<h1>Ranking</h1>\n<p>LightGBM classifier is used as a ranking model. There is no significant difference between classifier and lambda ranker. But from my previous experience, the classifier will give a slightly higher performance in general. So I haven't tried lambda ranker in this competition, but it's a good way for ensembling.</p>\n<h5>Aggregation features</h5>\n<p>Nothing special here, just group by user/item, (user, item) and do some aggregation.</p>\n<table>\n<thead>\n<tr>\n<th>Feature</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>user feature</td>\n<td>average prices of items this customer has bought in the last 7/30/90 days</td>\n</tr>\n<tr>\n<td>item feature</td>\n<td>average age of customers who have bought this item in the last 7/30/90 days</td>\n</tr>\n<tr>\n<td>time</td>\n<td>how many days after the customer's last transaction record</td>\n</tr>\n<tr>\n<td>interaction</td>\n<td>whether the customer has bought the current item before</td>\n</tr>\n<tr>\n<td>others</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<h5>Embedding feature</h5>\n<p>We can treat the transaction history of a customer as a sentence, then use some NLP technique to pretrain the embedding of customer/item. My solution using word2vec to train the embedding and simply averaging the item embedding of a customer transaction history as one of the customer features (be care of leaking problems, do not use items in the current week)</p>\n<h5>CV vs LB</h5>\n<p>train with the last 20 weeks and generate 100 candidates for each customer (a bug was fixed in the submission phase several hours before the deadline, so the gap should be smaller)</p>\n<table>\n<thead>\n<tr>\n<th>validation</th>\n<th>score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>local cv</td>\n<td>0.0363</td>\n</tr>\n<tr>\n<td>public score</td>\n<td>0.0310</td>\n</tr>\n<tr>\n<td>private score</td>\n<td>0.0312</td>\n</tr>\n</tbody>\n</table>\n<h1>Optimization</h1>\n<p>Feature engineering with cudf is 30X faster than using pandas in my case. It allows me to generate more candidates and get the feature in the submission phase. Load LightGBM model with Forest Inference Library and inference with GPU is 40X faster than using 32 core CPUs.</p>\n<h1>Final Train &amp; Inference</h1>\n<ul>\n<li>Train</li>\n</ul>\n<p>Generate 200 candidates for each customer, keeping all positive samples and random choosing half of the negative samples in the last 20 weeks</p>\n<ul>\n<li>Inference</li>\n</ul>\n<p>Generate 400 candidates for each customer, making prediction scores with 4 LightGBM models. Averaging them as the final score and getting top12 items</p>",
      "rawMarkdown": "It's too long since I last see the competition about recommendations on Kaggle. So great thanks to the Kaggle team and H&M for hosting such an interesting competition. The high correlation between CV and LB, big training data for deep data mining, no leaking of test information, friendly staff and participants, all of these make the successful competition.\n\n# Summary\n\nThe key idea of my solution will be illustrated here for fast reading purposes.\n\n##### Validation strategy\n\n```python\ntrain = transactions.loc[transactions[\"t_dat\"] < \"2020-09-16\"]\nvalid = transactions.loc[transactions[\"t_dat\"] >= \"2020-09-16\"]\n```\n\nIt is better to use k-fold validation in [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919) to avoid overfitting the last week if you have enough computing power.\n\nBut from my experience in this competition, LB score and CV score are highly correlated using last week only as a validation set.\n\n\n##### Recall\n\n| Recall method                   | link                                                                                                        |\n| ------------------------------- | ----------------------------------------------------------------------------------------------------------- |\n| SAR                             | [notebook](https://github.com/microsoft/recommenders/blob/main/examples/00_quick_start/sar_movielens.ipynb) |\n| Trending                        | [notebook](https://www.kaggle.com/code/byfone/h-m-trending-products-weekly)                                 |\n| Most popular items based on age | /                                                                                                           |\n\n\n##### Ranking\n\nLightGBM trained with 20 weeks of data using around 300 features\n\n\n##### Ensemble\n\nI have no time to train all 5 models with k fold split on the train set, only 3 of them are completed before the deadline. So my final submission is based on 1 model trained on all data samples and the other 3 models trained based on k fold split mentioned before (4 models in total).\n\n| Model     | public score | private score |\n| --------- | ------------ | ------------- |\n| one model | 0.0338       | 0.0341        |\n| ensemble  | 0.0343       | 0.0346        |\n\n\n##### Optimization\n\nUsing cudf and Forest Inference Library from Rapids for fast feature generation and fast tree inference\n\n\n##### Hardware\n\n| Hardware | value    |\n| -------- | -------- |\n| CPU      | 32core   |\n| Memory   | 128G     |\n| GPU      | V100 32G |\n\n\n# Data Analysis\n\n1. We can find 90%+ items in the validation set also in the last 30 days\n\n```python\ntrain_condition = (trans[\"t_dat\"] >= \"2020-08-16\") & (trans[\"t_dat\"] < \"2020-09-16\")\nitems_in_valid = trans.loc[trans[\"t_dat\"] >= \"2020-09-16\", \"article_id\"].unique()\nitems_last_month = trans.loc[train_condition, \"article_id\"].unique()\nprint(len(set(items_in_valid) & set(items_last_month)) / len(set(items_in_valid)))\n# 0.92\n```\n\n2. Approximately 50% of customers in the validation set have transaction records in the last 30 days\n\n```python\ntrain_condition = (trans[\"t_dat\"] >= \"2020-08-16\") & (trans[\"t_dat\"] < \"2020-09-16\")\nusers_in_valid = trans.loc[trans[\"t_dat\"] >= \"2020-09-16\", \"customer_id\"].unique()\nusers_last_month = trans.loc[train_condition, \"customer_id\"].unique()\nprint(len(set(users_in_valid) & set(users_last_month)) / len(set(users_in_valid)))\n# 0.47\n```\n\nConsidering these results, we can focus on recommending items bought in the last 30 days, especially for those customers who have transaction records.\n\n# Recall\n\nOnly mapk of each recall method on the local validation set is recorded. Most of these ideas are coming from public notebooks or discussions and I will give the link if so.\n\n##### Most popular items last 7 days\n\nSimply sorting items by transaction records. This is also the basic rule for those customers has no age data\n\n| Recall method | mapk       |\n| ------------- | ---------- |\n| Top 12 items  | 0.00814889 |\n\n##### Most popular itmes last 7 days by age\n\nWe are doing personalized recommendations, so it's better to use some customers' features based on the rule for recall. Unfortunately, only 2 features seem valuable to me at first glance, age, and postal code. It's impossible to get the location feature from the postal code since it is encoded. Finally, the only choice is age.\n\n```python\nbins = [0, 18, 22, 28, 35, 45, 55, 65, 200]\nlast_week_transactions[\"age_bin\"] = pd.cut(user[\"age\"], bins=bins)\nlast_week_transactions_popular = last_week_transactions.groupby(\"age_bin\").size()\n```\n\nBins are simply tuned with last week's data. The mapk score is slightly higher than only using the most popular items.\n\n| Recall method       | mapk      |\n| ------------------- | --------- |\n| Top 12 items        | 0.0081488 |\n| Top 12 items by age | 0.0091866 |\n\n##### Consider items that customers have bought before\n\nThe first method is just adding items that customers have bought in the last 4 weeks. Similar to [this notebook](https://www.kaggle.com/code/hengzheng/time-is-our-best-friend-v2)\n\n| Recall method                           | mapk      |\n| --------------------------------------- | --------- |\n| Top 12 items                            | 0.0081488 |\n| Top 12 items by age                     | 0.0091866 |\n| bought last 4 week +Top 12 items by age | 0.0245722 |\n\n##### Use Trending\n\nThere is a significant increase in mapk. But since we know only half of the customers have transaction history in the last 30 days, we can do better. According to [this notebook](https://www.kaggle.com/code/byfone/h-m-trending-products-weekly), we have a wonderful method to compute the score of each item to all customers without the limitation of 30 days.\n\n\n| Recall method                           | mapk      |\n| --------------------------------------- | --------- |\n| Top 12 items                            | 0.0081488 |\n| Top 12 items by age                     | 0.0091866 |\n| bought last 4 week +Top 12 items by age | 0.0245722 |\n| Trending + Top 12 items by age          | 0.0255676 |\n\n##### Recommend with SAR\n\nSAR is a kind of item-based collaborative filtering. As mentioned in [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316205), we can get 0.27+ with SAR. Here is the [example](https://github.com/microsoft/recommenders/blob/main/examples/00_quick_start/sar_movielens.ipynb) from the Microsoft official repo. My SAR is trained with the last 30 days. Adding more data will make the performance worse in my case.\n\n| experiment                              | mapk      |\n| --------------------------------------- | --------- |\n| Top 12 items                            | 0.0081488 |\n| Top 12 items by age                     | 0.0091866 |\n| bought last 4 week +Top 12 items by age | 0.0245722 |\n| Trending + Top 12 items by age          | 0.0255676 |\n| SAR + Trending + Top 12 items by age    | 0.0278576 |\n\n##### Overall recall strategy\n\nLet's say we are going to recommend K items for each customer. My Final strategy for the recall is\n1. Recommend with SAR if the customer has a transaction in the last 30 days\n2. If the number of candidates is less than K, pick items from the trending method (remove items that have already been in the first step)\n3. If the number of candidates is still less than K, fill with most popular items by age (using most popular items overall if the age data is missing)\n4. Each recall method has a score and we will use them as features in the ranking phase.\n\n# Ranking\n\nLightGBM classifier is used as a ranking model. There is no significant difference between classifier and lambda ranker. But from my previous experience, the classifier will give a slightly higher performance in general. So I haven't tried lambda ranker in this competition, but it's a good way for ensembling.\n\n##### Aggregation features\n\nNothing special here, just group by user/item, (user, item) and do some aggregation.\n| Feature      | Description                                                                 |\n| ------------ | --------------------------------------------------------------------------- |\n| user feature | average prices of items this customer has bought in the last 7/30/90 days   |\n| item feature | average age of customers who have bought this item in the last 7/30/90 days |\n| time         | how many days after the customer's last transaction record                  |\n| interaction  | whether the customer has bought the current item before                     |\n| others       | ...                                                                         |\n\n##### Embedding feature\n\nWe can treat the transaction history of a customer as a sentence, then use some NLP technique to pretrain the embedding of customer/item. My solution using word2vec to train the embedding and simply averaging the item embedding of a customer transaction history as one of the customer features (be care of leaking problems, do not use items in the current week)\n\n##### CV vs LB\n\ntrain with the last 20 weeks and generate 100 candidates for each customer (a bug was fixed in the submission phase several hours before the deadline, so the gap should be smaller)\n\n| validation    | score  |\n| ------------- | ------ |\n| local cv      | 0.0363 |\n| public score  | 0.0310 |\n| private score | 0.0312 |\n\n# Optimization\n\nFeature engineering with cudf is 30X faster than using pandas in my case. It allows me to generate more candidates and get the feature in the submission phase. Load LightGBM model with Forest Inference Library and inference with GPU is 40X faster than using 32 core CPUs.\n\n\n# Final Train & Inference\n\n- Train\n\nGenerate 200 candidates for each customer, keeping all positive samples and random choosing half of the negative samples in the last 20 weeks\n\n- Inference\n\nGenerate 400 candidates for each customer, making prediction scores with 4 LightGBM models. Averaging them as the final score and getting top12 items",
      "votes": null
    },
    {
      "id": "1783546",
      "postDate": "05/10/2022 13:19:02",
      "content": "<p>Great job!</p>",
      "rawMarkdown": "Great job!",
      "votes": null
    },
    {
      "id": "1783648",
      "postDate": "05/10/2022 14:52:53",
      "content": "<p>Well done! I'm glad SAR worked for you as well… we knew it was a good way to generate candidates but unfortunately we did something wrong in our pipeline and we couldn't improve our score.</p>",
      "rawMarkdown": "Well done! I'm glad SAR worked for you as well... we knew it was a good way to generate candidates but unfortunately we did something wrong in our pipeline and we couldn't improve our score.",
      "votes": null
    },
    {
      "id": "1783653",
      "postDate": "05/10/2022 14:58:18",
      "content": "<p>Thanks for posting and congratz for gold!<br>\nThere's something I don't understood from \"Overall recall strategy\"</p>\n<blockquote>\n  <p>Let's say we are going to recommend K items for each customer. My Final strategy for the recall is</p>\n  <ol>\n  <li>Recommend with SAR if the customer has a transaction in the last 30 days</li>\n  <li>If the number of candidates is less than K, pick items from the trending method (remove items that have already been in the first step)</li>\n  </ol>\n</blockquote>\n<p>Can you cite when \"If the number of candidates is less than K\" could happen?</p>",
      "rawMarkdown": "Thanks for posting and congratz for gold!\nThere's something I don't understood from \"Overall recall strategy\"\n\n> Let's say we are going to recommend K items for each customer. My Final strategy for the recall is\n1. Recommend with SAR if the customer has a transaction in the last 30 days\n2. If the number of candidates is less than K, pick items from the trending method (remove items that have already been in the first step)\n\nCan you cite when \"If the number of candidates is less than K\" could happen?",
      "votes": null
    },
    {
      "id": "1783679",
      "postDate": "05/10/2022 15:21:56",
      "content": "<p>Congratz!!!</p>",
      "rawMarkdown": "Congratz!!!",
      "votes": null
    },
    {
      "id": "1783686",
      "postDate": "05/10/2022 15:32:47",
      "content": "<p>Great job !!!</p>",
      "rawMarkdown": "Great job !!!",
      "votes": null
    },
    {
      "id": "1783788",
      "postDate": "05/10/2022 17:11:15",
      "content": "<p>Thanks for this great explanation, great job!</p>",
      "rawMarkdown": "Thanks for this great explanation, great job!",
      "votes": null
    },
    {
      "id": "1784205",
      "postDate": "05/11/2022 02:54:51",
      "content": "<p>There is a threshold for each recall method tuned with my local validation set. If the prediction score of SAR/trending method is less than a value (tuned locally), these items will be removed from my candidates. So actually SAR won't recall K items for each customer especially for those with few transactions (score will be lower).</p>",
      "rawMarkdown": "There is a threshold for each recall method tuned with my local validation set. If the prediction score of SAR/trending method is less than a value (tuned locally), these items will be removed from my candidates. So actually SAR won't recall K items for each customer especially for those with few transactions (score will be lower).",
      "votes": null
    },
    {
      "id": "1784214",
      "postDate": "05/11/2022 03:03:11",
      "content": "<p>Thank you for pointing out the performance of SAR in the early stage, this method not only generates good candidates but also outputs the score which can be used as a golden feature for my ranking model. Hope you can get a better result in the next RecSys competition!</p>",
      "rawMarkdown": "Thank you for pointing out the performance of SAR in the early stage, this method not only generates good candidates but also outputs the score which can be used as a golden feature for my ranking model. Hope you can get a better result in the next RecSys competition!",
      "votes": null
    },
    {
      "id": "1785882",
      "postDate": "05/12/2022 13:19:31",
      "content": "<p>Great Work! Thanks for sharing.</p>\n<blockquote>\n  <p>Generate 200 candidates for each customer, keeping all positive samples and random choosing half of the negative samples in the last 20 weeks</p>\n</blockquote>\n<p>Does it mean that if get 50 positive candidates and 150 negative candidates , label 50 all positive as 1 and  75 negative as 0?<br>\nMy idea are very similarity with you especially on  Data Analysis (90%,50%).But I have no ability to implement because I'm a newbie here.Could you share your code on github ? If convenient, it will help a lot of people like me. I will appreciate it so much.BTY,I found you are in shanghai,and I'm studying there now.</p>",
      "rawMarkdown": "Great Work! Thanks for sharing.\n> Generate 200 candidates for each customer, keeping all positive samples and random choosing half of the negative samples in the last 20 weeks\n\n\nDoes it mean that if get 50 positive candidates and 150 negative candidates , label 50 all positive as 1 and  75 negative as 0?\nMy idea are very similarity with you especially on  Data Analysis (90%,50%).But I have no ability to implement because I'm a newbie here.Could you share your code on github ? If convenient, it will help a lot of people like me. I will appreciate it so much.BTY,I found you are in shanghai,and I'm studying there now.",
      "votes": null
    },
    {
      "id": "1786577",
      "postDate": "05/13/2022 03:28:00",
      "content": "<blockquote>\n  <p>Does it mean that if get 50 positive candidates and 150 negative candidates , label 50 all positive as 1 and 75 negative as 0?</p>\n</blockquote>\n<p>Yes, do this properly for each customer because the number of candidates for each customer is different. There may be a better method but I have no time to focus on this.</p>\n<blockquote>\n  <p>Could you share your code on github ?</p>\n</blockquote>\n<p>I have no plan to release my messy code so far. I may try to clean my code and add some comments if I have time to do that.</p>",
      "rawMarkdown": "> Does it mean that if get 50 positive candidates and 150 negative candidates , label 50 all positive as 1 and 75 negative as 0?\n\nYes, do this properly for each customer because the number of candidates for each customer is different. There may be a better method but I have no time to focus on this.\n\n> Could you share your code on github ?\n\nI have no plan to release my messy code so far. I may try to clean my code and add some comments if I have time to do that.",
      "votes": null
    },
    {
      "id": "1790493",
      "postDate": "05/15/2022 02:50:06",
      "content": "<p>Thansk for sharing, I have one problem: <br>\nin the SAR recall model, how do you control the <code>[&lt;Event Weight&gt;]</code> in a interaction?</p>",
      "rawMarkdown": "Thansk for sharing, I have one problem: \nin the SAR recall model, how do you control the `[<Event Weight>]` in a interaction?",
      "votes": null
    },
    {
      "id": "1790659",
      "postDate": "05/15/2022 07:12:43",
      "content": "<p>A good job! Could you share the sorted importance of features used in lightgbm? Besides, How long is the embedding dimension?</p>",
      "rawMarkdown": "A good job! Could you share the sorted importance of features used in lightgbm? Besides, How long is the embedding dimension?",
      "votes": null
    },
    {
      "id": "1790798",
      "postDate": "05/15/2022 10:34:45",
      "content": "<blockquote>\n  <p>Each event type can be assigned a different weight, for example, we might assign a “buy” event a weight of 10, while a “view” event might only have a weight of 1.</p>\n</blockquote>\n<p>As shown in MS official example, the event weight shows the importance of this interaction. Thus, I give the same value to every behavior since all of them are \"buy\".</p>",
      "rawMarkdown": "> Each event type can be assigned a different weight, for example, we might assign a “buy” event a weight of 10, while a “view” event might only have a weight of 1.\n\nAs shown in MS official example, the event weight shows the importance of this interaction. Thus, I give the same value to every behavior since all of them are \"buy\".",
      "votes": null
    },
    {
      "id": "1790803",
      "postDate": "05/15/2022 10:49:57",
      "content": "<blockquote>\n  <p>Could you share the sorted importance of features used in lightgbm?</p>\n</blockquote>\n<p>As is expected, the most important features in top 10 are related to repurchase such as how many days since the customer last bought this item/product and how many times this customer bought this item/product. </p>\n<blockquote>\n  <p>How long is the embedding dimension?</p>\n</blockquote>\n<p>I have tried sizes 32 and 64 and the performance of size 64 is slightly better than size 32 on my local CV. But I finally use size 32 considering the limitation of my hardware.</p>",
      "rawMarkdown": "> Could you share the sorted importance of features used in lightgbm?\n\nAs is expected, the most important features in top 10 are related to repurchase such as how many days since the customer last bought this item/product and how many times this customer bought this item/product. \n\n> How long is the embedding dimension?\n\nI have tried sizes 32 and 64 and the performance of size 64 is slightly better than size 32 on my local CV. But I finally use size 32 considering the limitation of my hardware.",
      "votes": null
    },
    {
      "id": "1790902",
      "postDate": "05/15/2022 13:04:08",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1791013",
      "postDate": "05/15/2022 15:10:35",
      "content": "<p>Thanks for your reply, learned a lot!</p>",
      "rawMarkdown": "Thanks for your reply, learned a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783546,
      "author_name": "xiamaozi11",
      "author_url": "",
      "post_date": "05/10/2022 13:19:02",
      "content": "<p>Great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783648,
      "author_name": "igormunizims",
      "author_url": "",
      "post_date": "05/10/2022 14:52:53",
      "content": "<p>Well done! I'm glad SAR worked for you as well… we knew it was a good way to generate candidates but unfortunately we did something wrong in our pipeline and we couldn't improve our score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784214,
          "author_name": "yzheng21",
          "author_url": "",
          "post_date": "05/11/2022 03:03:11",
          "content": "<p>Thank you for pointing out the performance of SAR in the early stage, this method not only generates good candidates but also outputs the score which can be used as a golden feature for my ranking model. Hope you can get a better result in the next RecSys competition!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783653,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "05/10/2022 14:58:18",
      "content": "<p>Thanks for posting and congratz for gold!<br>\nThere's something I don't understood from \"Overall recall strategy\"</p>\n<blockquote>\n  <p>Let's say we are going to recommend K items for each customer. My Final strategy for the recall is</p>\n  <ol>\n  <li>Recommend with SAR if the customer has a transaction in the last 30 days</li>\n  <li>If the number of candidates is less than K, pick items from the trending method (remove items that have already been in the first step)</li>\n  </ol>\n</blockquote>\n<p>Can you cite when \"If the number of candidates is less than K\" could happen?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784205,
          "author_name": "yzheng21",
          "author_url": "",
          "post_date": "05/11/2022 02:54:51",
          "content": "<p>There is a threshold for each recall method tuned with my local validation set. If the prediction score of SAR/trending method is less than a value (tuned locally), these items will be removed from my candidates. So actually SAR won't recall K items for each customer especially for those with few transactions (score will be lower).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783679,
      "author_name": "renxingkai",
      "author_url": "",
      "post_date": "05/10/2022 15:21:56",
      "content": "<p>Congratz!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783686,
      "author_name": "mohamedaliguirat",
      "author_url": "",
      "post_date": "05/10/2022 15:32:47",
      "content": "<p>Great job !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783788,
      "author_name": "ryanrose1",
      "author_url": "",
      "post_date": "05/10/2022 17:11:15",
      "content": "<p>Thanks for this great explanation, great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1785882,
      "author_name": "jiahongxie",
      "author_url": "",
      "post_date": "05/12/2022 13:19:31",
      "content": "<p>Great Work! Thanks for sharing.</p>\n<blockquote>\n  <p>Generate 200 candidates for each customer, keeping all positive samples and random choosing half of the negative samples in the last 20 weeks</p>\n</blockquote>\n<p>Does it mean that if get 50 positive candidates and 150 negative candidates , label 50 all positive as 1 and  75 negative as 0?<br>\nMy idea are very similarity with you especially on  Data Analysis (90%,50%).But I have no ability to implement because I'm a newbie here.Could you share your code on github ? If convenient, it will help a lot of people like me. I will appreciate it so much.BTY,I found you are in shanghai,and I'm studying there now.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1786577,
          "author_name": "yzheng21",
          "author_url": "",
          "post_date": "05/13/2022 03:28:00",
          "content": "<blockquote>\n  <p>Does it mean that if get 50 positive candidates and 150 negative candidates , label 50 all positive as 1 and 75 negative as 0?</p>\n</blockquote>\n<p>Yes, do this properly for each customer because the number of candidates for each customer is different. There may be a better method but I have no time to focus on this.</p>\n<blockquote>\n  <p>Could you share your code on github ?</p>\n</blockquote>\n<p>I have no plan to release my messy code so far. I may try to clean my code and add some comments if I have time to do that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1790493,
      "author_name": "cherrizhu",
      "author_url": "",
      "post_date": "05/15/2022 02:50:06",
      "content": "<p>Thansk for sharing, I have one problem: <br>\nin the SAR recall model, how do you control the <code>[&lt;Event Weight&gt;]</code> in a interaction?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1790798,
          "author_name": "yzheng21",
          "author_url": "",
          "post_date": "05/15/2022 10:34:45",
          "content": "<blockquote>\n  <p>Each event type can be assigned a different weight, for example, we might assign a “buy” event a weight of 10, while a “view” event might only have a weight of 1.</p>\n</blockquote>\n<p>As shown in MS official example, the event weight shows the importance of this interaction. Thus, I give the same value to every behavior since all of them are \"buy\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1791013,
          "author_name": "cherrizhu",
          "author_url": "",
          "post_date": "05/15/2022 15:10:35",
          "content": "<p>Thanks for your reply, learned a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1790659,
      "author_name": "dengniewei",
      "author_url": "",
      "post_date": "05/15/2022 07:12:43",
      "content": "<p>A good job! Could you share the sorted importance of features used in lightgbm? Besides, How long is the embedding dimension?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1790803,
          "author_name": "yzheng21",
          "author_url": "",
          "post_date": "05/15/2022 10:49:57",
          "content": "<blockquote>\n  <p>Could you share the sorted importance of features used in lightgbm?</p>\n</blockquote>\n<p>As is expected, the most important features in top 10 are related to repurchase such as how many days since the customer last bought this item/product and how many times this customer bought this item/product. </p>\n<blockquote>\n  <p>How long is the embedding dimension?</p>\n</blockquote>\n<p>I have tried sizes 32 and 64 and the performance of size 64 is slightly better than size 32 on my local CV. But I finally use size 32 considering the limitation of my hardware.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1790902,
          "author_name": "dengniewei",
          "author_url": "",
          "post_date": "05/15/2022 13:04:08",
          "content": "<p>Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1783261": "It's too long since I last see the competition about recommendations on Kaggle. So great thanks to the Kaggle team and H&M for hosting such an interesting competition. The high correlation between CV and LB, big training data for deep data mining, no leaking of test information, friendly staff and participants, all of these make the successful competition.\n\n# Summary\n\nThe key idea of my solution will be illustrated here for fast reading purposes.\n\n##### Validation strategy\n\n```python\ntrain = transactions.loc[transactions[\"t_dat\"] < \"2020-09-16\"]\nvalid = transactions.loc[transactions[\"t_dat\"] >= \"2020-09-16\"]\n```\n\nIt is better to use k-fold validation in [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919) to avoid overfitting the last week if you have enough computing power.\n\nBut from my experience in this competition, LB score and CV score are highly correlated using last week only as a validation set.\n\n\n##### Recall\n\n| Recall method                   | link                                                                                                        |\n| ------------------------------- | ----------------------------------------------------------------------------------------------------------- |\n| SAR                             | [notebook](https://github.com/microsoft/recommenders/blob/main/examples/00_quick_start/sar_movielens.ipynb) |\n| Trending                        | [notebook](https://www.kaggle.com/code/byfone/h-m-trending-products-weekly)                                 |\n| Most popular items based on age | /                                                                                                           |\n\n\n##### Ranking\n\nLightGBM trained with 20 weeks of data using around 300 features\n\n\n##### Ensemble\n\nI have no time to train all 5 models with k fold split on the train set, only 3 of them are completed before the deadline. So my final submission is based on 1 model trained on all data samples and the other 3 models trained based on k fold split mentioned before (4 models in total).\n\n| Model     | public score | private score |\n| --------- | ------------ | ------------- |\n| one model | 0.0338       | 0.0341        |\n| ensemble  | 0.0343       | 0.0346        |\n\n\n##### Optimization\n\nUsing cudf and Forest Inference Library from Rapids for fast feature generation and fast tree inference\n\n\n##### Hardware\n\n| Hardware | value    |\n| -------- | -------- |\n| CPU      | 32core   |\n| Memory   | 128G     |\n| GPU      | V100 32G |\n\n\n# Data Analysis\n\n1. We can find 90%+ items in the validation set also in the last 30 days\n\n```python\ntrain_condition = (trans[\"t_dat\"] >= \"2020-08-16\") & (trans[\"t_dat\"] < \"2020-09-16\")\nitems_in_valid = trans.loc[trans[\"t_dat\"] >= \"2020-09-16\", \"article_id\"].unique()\nitems_last_month = trans.loc[train_condition, \"article_id\"].unique()\nprint(len(set(items_in_valid) & set(items_last_month)) / len(set(items_in_valid)))\n# 0.92\n```\n\n2. Approximately 50% of customers in the validation set have transaction records in the last 30 days\n\n```python\ntrain_condition = (trans[\"t_dat\"] >= \"2020-08-16\") & (trans[\"t_dat\"] < \"2020-09-16\")\nusers_in_valid = trans.loc[trans[\"t_dat\"] >= \"2020-09-16\", \"customer_id\"].unique()\nusers_last_month = trans.loc[train_condition, \"customer_id\"].unique()\nprint(len(set(users_in_valid) & set(users_last_month)) / len(set(users_in_valid)))\n# 0.47\n```\n\nConsidering these results, we can focus on recommending items bought in the last 30 days, especially for those customers who have transaction records.\n\n# Recall\n\nOnly mapk of each recall method on the local validation set is recorded. Most of these ideas are coming from public notebooks or discussions and I will give the link if so.\n\n##### Most popular items last 7 days\n\nSimply sorting items by transaction records. This is also the basic rule for those customers has no age data\n\n| Recall method | mapk       |\n| ------------- | ---------- |\n| Top 12 items  | 0.00814889 |\n\n##### Most popular itmes last 7 days by age\n\nWe are doing personalized recommendations, so it's better to use some customers' features based on the rule for recall. Unfortunately, only 2 features seem valuable to me at first glance, age, and postal code. It's impossible to get the location feature from the postal code since it is encoded. Finally, the only choice is age.\n\n```python\nbins = [0, 18, 22, 28, 35, 45, 55, 65, 200]\nlast_week_transactions[\"age_bin\"] = pd.cut(user[\"age\"], bins=bins)\nlast_week_transactions_popular = last_week_transactions.groupby(\"age_bin\").size()\n```\n\nBins are simply tuned with last week's data. The mapk score is slightly higher than only using the most popular items.\n\n| Recall method       | mapk      |\n| ------------------- | --------- |\n| Top 12 items        | 0.0081488 |\n| Top 12 items by age | 0.0091866 |\n\n##### Consider items that customers have bought before\n\nThe first method is just adding items that customers have bought in the last 4 weeks. Similar to [this notebook](https://www.kaggle.com/code/hengzheng/time-is-our-best-friend-v2)\n\n| Recall method                           | mapk      |\n| --------------------------------------- | --------- |\n| Top 12 items                            | 0.0081488 |\n| Top 12 items by age                     | 0.0091866 |\n| bought last 4 week +Top 12 items by age | 0.0245722 |\n\n##### Use Trending\n\nThere is a significant increase in mapk. But since we know only half of the customers have transaction history in the last 30 days, we can do better. According to [this notebook](https://www.kaggle.com/code/byfone/h-m-trending-products-weekly), we have a wonderful method to compute the score of each item to all customers without the limitation of 30 days.\n\n\n| Recall method                           | mapk      |\n| --------------------------------------- | --------- |\n| Top 12 items                            | 0.0081488 |\n| Top 12 items by age                     | 0.0091866 |\n| bought last 4 week +Top 12 items by age | 0.0245722 |\n| Trending + Top 12 items by age          | 0.0255676 |\n\n##### Recommend with SAR\n\nSAR is a kind of item-based collaborative filtering. As mentioned in [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316205), we can get 0.27+ with SAR. Here is the [example](https://github.com/microsoft/recommenders/blob/main/examples/00_quick_start/sar_movielens.ipynb) from the Microsoft official repo. My SAR is trained with the last 30 days. Adding more data will make the performance worse in my case.\n\n| experiment                              | mapk      |\n| --------------------------------------- | --------- |\n| Top 12 items                            | 0.0081488 |\n| Top 12 items by age                     | 0.0091866 |\n| bought last 4 week +Top 12 items by age | 0.0245722 |\n| Trending + Top 12 items by age          | 0.0255676 |\n| SAR + Trending + Top 12 items by age    | 0.0278576 |\n\n##### Overall recall strategy\n\nLet's say we are going to recommend K items for each customer. My Final strategy for the recall is\n1. Recommend with SAR if the customer has a transaction in the last 30 days\n2. If the number of candidates is less than K, pick items from the trending method (remove items that have already been in the first step)\n3. If the number of candidates is still less than K, fill with most popular items by age (using most popular items overall if the age data is missing)\n4. Each recall method has a score and we will use them as features in the ranking phase.\n\n# Ranking\n\nLightGBM classifier is used as a ranking model. There is no significant difference between classifier and lambda ranker. But from my previous experience, the classifier will give a slightly higher performance in general. So I haven't tried lambda ranker in this competition, but it's a good way for ensembling.\n\n##### Aggregation features\n\nNothing special here, just group by user/item, (user, item) and do some aggregation.\n| Feature      | Description                                                                 |\n| ------------ | --------------------------------------------------------------------------- |\n| user feature | average prices of items this customer has bought in the last 7/30/90 days   |\n| item feature | average age of customers who have bought this item in the last 7/30/90 days |\n| time         | how many days after the customer's last transaction record                  |\n| interaction  | whether the customer has bought the current item before                     |\n| others       | ...                                                                         |\n\n##### Embedding feature\n\nWe can treat the transaction history of a customer as a sentence, then use some NLP technique to pretrain the embedding of customer/item. My solution using word2vec to train the embedding and simply averaging the item embedding of a customer transaction history as one of the customer features (be care of leaking problems, do not use items in the current week)\n\n##### CV vs LB\n\ntrain with the last 20 weeks and generate 100 candidates for each customer (a bug was fixed in the submission phase several hours before the deadline, so the gap should be smaller)\n\n| validation    | score  |\n| ------------- | ------ |\n| local cv      | 0.0363 |\n| public score  | 0.0310 |\n| private score | 0.0312 |\n\n# Optimization\n\nFeature engineering with cudf is 30X faster than using pandas in my case. It allows me to generate more candidates and get the feature in the submission phase. Load LightGBM model with Forest Inference Library and inference with GPU is 40X faster than using 32 core CPUs.\n\n\n# Final Train & Inference\n\n- Train\n\nGenerate 200 candidates for each customer, keeping all positive samples and random choosing half of the negative samples in the last 20 weeks\n\n- Inference\n\nGenerate 400 candidates for each customer, making prediction scores with 4 LightGBM models. Averaging them as the final score and getting top12 items",
    "1783546": "Great job!",
    "1783648": "Well done! I'm glad SAR worked for you as well... we knew it was a good way to generate candidates but unfortunately we did something wrong in our pipeline and we couldn't improve our score.",
    "1783653": "Thanks for posting and congratz for gold!\nThere's something I don't understood from \"Overall recall strategy\"\n\n> Let's say we are going to recommend K items for each customer. My Final strategy for the recall is\n1. Recommend with SAR if the customer has a transaction in the last 30 days\n2. If the number of candidates is less than K, pick items from the trending method (remove items that have already been in the first step)\n\nCan you cite when \"If the number of candidates is less than K\" could happen?",
    "1783679": "Congratz!!!",
    "1783686": "Great job !!!",
    "1783788": "Thanks for this great explanation, great job!",
    "1784205": "There is a threshold for each recall method tuned with my local validation set. If the prediction score of SAR/trending method is less than a value (tuned locally), these items will be removed from my candidates. So actually SAR won't recall K items for each customer especially for those with few transactions (score will be lower).",
    "1784214": "Thank you for pointing out the performance of SAR in the early stage, this method not only generates good candidates but also outputs the score which can be used as a golden feature for my ranking model. Hope you can get a better result in the next RecSys competition!",
    "1785882": "Great Work! Thanks for sharing.\n> Generate 200 candidates for each customer, keeping all positive samples and random choosing half of the negative samples in the last 20 weeks\n\n\nDoes it mean that if get 50 positive candidates and 150 negative candidates , label 50 all positive as 1 and  75 negative as 0?\nMy idea are very similarity with you especially on  Data Analysis (90%,50%).But I have no ability to implement because I'm a newbie here.Could you share your code on github ? If convenient, it will help a lot of people like me. I will appreciate it so much.BTY,I found you are in shanghai,and I'm studying there now.",
    "1786577": "> Does it mean that if get 50 positive candidates and 150 negative candidates , label 50 all positive as 1 and 75 negative as 0?\n\nYes, do this properly for each customer because the number of candidates for each customer is different. There may be a better method but I have no time to focus on this.\n\n> Could you share your code on github ?\n\nI have no plan to release my messy code so far. I may try to clean my code and add some comments if I have time to do that.",
    "1790493": "Thansk for sharing, I have one problem: \nin the SAR recall model, how do you control the `[<Event Weight>]` in a interaction?",
    "1790659": "A good job! Could you share the sorted importance of features used in lightgbm? Besides, How long is the embedding dimension?",
    "1790798": "> Each event type can be assigned a different weight, for example, we might assign a “buy” event a weight of 10, while a “view” event might only have a weight of 1.\n\nAs shown in MS official example, the event weight shows the importance of this interaction. Thus, I give the same value to every behavior since all of them are \"buy\".",
    "1790803": "> Could you share the sorted importance of features used in lightgbm?\n\nAs is expected, the most important features in top 10 are related to repurchase such as how many days since the customer last bought this item/product and how many times this customer bought this item/product. \n\n> How long is the embedding dimension?\n\nI have tried sizes 32 and 64 and the performance of size 64 is slightly better than size 32 on my local CV. But I finally use size 32 considering the limitation of my hardware.",
    "1790902": "Thanks for sharing.",
    "1791013": "Thanks for your reply, learned a lot!"
  },
  "source": "meta"
}