{
  "id": 324207,
  "title": "13th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/pksha-don-t-let-recommend-13th-place-solution",
  "author_name": "",
  "post_date": "2022-05-10T14:11:08.660360400Z",
  "votes": 28,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Before going on to our solution, I would like to thank Kaggle staff and H&amp;M team, alongside with all competitors on the leaderboard for competing with us in this competition. And of course, special thanks to <a href=\"https://www.kaggle.com/A.Sato\" target=\"_blank\">@A.Sato</a>, <a href=\"https://www.kaggle.com/taksai\" target=\"_blank\">@taksai</a>, and <a href=\"https://www.kaggle.com/tomo20180402\" target=\"_blank\">@tomo20180402</a> who are my teammates, for their amazing effort. We are honored to win the gold medal!</p>\n<h1>Overall Strategy</h1>\n<p>The overall strategy of our solution is based on four steps.</p>\n<ol>\n<li>Candidate Selection</li>\n<li>Feature Engineering</li>\n<li>Ranking</li>\n<li>Ensembling</li>\n</ol>\n<h1>Key Points</h1>\n<p>We trained three LightGBM ranker models with different candidate selection methodologies, and finally used an ensemble technique to boost our score.<br>\nThe key idea of our solution was utilizing item2vec embeddings, which is an algorithm inspired by word2vec.<br>\nArticle embeddings were retrieved using item2vec embeddings as they are, and customer embeddings were created by taking an average among all dimensions of all past purchased articles for each customer.</p>\n<h1>Step1: Candidate(Article) Selection</h1>\n<p>In the candidate selection phase, we tried to sample positive and negative samples for each customer. The candidate selection methods were different for each of the three models, and the diversity of candidates between our models leveraged our scores after ensembling.</p>\n<h2>Model1</h2>\n<p>The ensemble method for Model1 was a concatenation of 4 strategies. Past 14 weeks were used to train the model.</p>\n<ol>\n<li>recent 50 purchased articles for each customer</li>\n<li>top 20 products with the same product code of each article in 1.</li>\n<li>last week top 100 articles of all customers for each sales_channel_id group</li>\n<li>last week top 30 of all customers for each [sales_channel_id, index_group_no] group</li>\n</ol>\n<h2>Model2</h2>\n<p>Aside from recently purchased items (and items with the same product code of recently purchased items), my main model solely relies on the local popularity among the nearest 10,000 customers. 1.I obtained user's latent expression by getting the mean of the item2vec embedded expression (32 dimentions) of the items that the customer purchased in the past, 2.then selected 10,000 nearest customers for EACH customer based on the cosine similarity of user latent expression, 3.then picked top 300 last week popular items among 10,000 nearest customers as candidates.  <br>\n <br><br>\nAs for single model, this approach got the best score among three models(Private LB 330-ish).</p>\n<h2>Model3</h2>\n<p>Firstly, I created 4 types of clusters for each customer based on the techniques bellow.<br>\n(Article-based embeddings were averaged by each dimension of past bought articles to create customer embeddings)</p>\n<ul>\n<li>Swin-Transformer Image embeddings</li>\n<li>Age of each customer</li>\n<li>item2vec based article embeddings</li>\n<li>postal_code2vec (inspired by item2vec) embeddings</li>\n</ul>\n<p>\nAfter the clustering process, each customer will belong to 4 different clusters. <br>\nFor each cluster, I created an article ranking for the past 30 days, and made a weighted overall ranking depending on which clusters each customer belonged to.<br>\nRecent bought article_ids and product_codes were weighted heavier according to the customer.\nFrom the overall ranking, Model3 took the top 350 articles per user.\n</p>\n<h1>Step2: Feature Engineering</h1>\n<p>For each of the candidates, we created features as bellow.</p>\n<ul>\n<li>deep learning based latent expressions<ul>\n<li>Articles embedded expresion by BERT4REC</li>\n<li>Articles embedded expresion by LightGCN</li>\n<li>Detailed description's latent expression by Roberta (plus PCA for dimensionality reduction)</li>\n<li>Swin-Transformer based image features</li></ul></li>\n<li>cosine similarity of the article and the customer based on swin-transformer embeddings and item2vec embeddings</li>\n<li>number of past purchases/interval_days by the user for the candidate article_id, product_code, etc</li>\n<li>number of past purchases/interval_days by the cluster(clustered using item2vec) of the relevant user for the candidate article_id, product_code, etc</li>\n<li>article sales forecasts (forecasted by LightGBM regression model) for the corresponding week</li>\n</ul>\n<h1>Step3: Model</h1>\n<p>All three models were trained by LightGBMRaker.<br>\nModel1 was trained with 14 weeks,and model2/3 were only trained by using candidates from the last 1 week.</p>\n<h1>Step4: Ensembling</h1>\n<p>Ensembling was simply based by using the methodology used in many of the public notebooks (with a slight change).</p>\n<pre><code>def cust_blend(dt, W = [1,1,1], base=3): # base option is added\n\n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['model1'].split())\n    REC.append(dt['model2'].split())\n    REC.append(dt['model3'].split())\n\n\n    res = {}\n    for M in range(len(REC)):\n        if type(REC[M]) == list:\n            for n, v in enumerate(REC[M]):\n                if v in res:\n                    res[v] += (W[M]/(n+base))\n                else:\n                    res[v] = (W[M]/(n+base))\n\n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n\n    return ' '.join(res[:12])\n</code></pre>\n<h1>Failures</h1>\n<ul>\n<li>All neural network driven approaches discussed in the academic paper failed for us. Recbole and recommenders gave us marginal boost in this competition.</li>\n<li>Increasing training weeks worsened the local scores for model2 and 3.</li>\n<li>catboost (It's score was worse compared to LightGBM)</li>\n<li>general MLP</li>\n</ul>",
  "messages": [
    {
      "id": "1783606",
      "postDate": "05/10/2022 14:11:08",
      "content": "<p>Before going on to our solution, I would like to thank Kaggle staff and H&amp;M team, alongside with all competitors on the leaderboard for competing with us in this competition. And of course, special thanks to <a href=\"https://www.kaggle.com/A.Sato\" target=\"_blank\">@A.Sato</a>, <a href=\"https://www.kaggle.com/taksai\" target=\"_blank\">@taksai</a>, and <a href=\"https://www.kaggle.com/tomo20180402\" target=\"_blank\">@tomo20180402</a> who are my teammates, for their amazing effort. We are honored to win the gold medal!</p>\n<h1>Overall Strategy</h1>\n<p>The overall strategy of our solution is based on four steps.</p>\n<ol>\n<li>Candidate Selection</li>\n<li>Feature Engineering</li>\n<li>Ranking</li>\n<li>Ensembling</li>\n</ol>\n<h1>Key Points</h1>\n<p>We trained three LightGBM ranker models with different candidate selection methodologies, and finally used an ensemble technique to boost our score.<br>\nThe key idea of our solution was utilizing item2vec embeddings, which is an algorithm inspired by word2vec.<br>\nArticle embeddings were retrieved using item2vec embeddings as they are, and customer embeddings were created by taking an average among all dimensions of all past purchased articles for each customer.</p>\n<h1>Step1: Candidate(Article) Selection</h1>\n<p>In the candidate selection phase, we tried to sample positive and negative samples for each customer. The candidate selection methods were different for each of the three models, and the diversity of candidates between our models leveraged our scores after ensembling.</p>\n<h2>Model1</h2>\n<p>The ensemble method for Model1 was a concatenation of 4 strategies. Past 14 weeks were used to train the model.</p>\n<ol>\n<li>recent 50 purchased articles for each customer</li>\n<li>top 20 products with the same product code of each article in 1.</li>\n<li>last week top 100 articles of all customers for each sales_channel_id group</li>\n<li>last week top 30 of all customers for each [sales_channel_id, index_group_no] group</li>\n</ol>\n<h2>Model2</h2>\n<p>Aside from recently purchased items (and items with the same product code of recently purchased items), my main model solely relies on the local popularity among the nearest 10,000 customers. 1.I obtained user's latent expression by getting the mean of the item2vec embedded expression (32 dimentions) of the items that the customer purchased in the past, 2.then selected 10,000 nearest customers for EACH customer based on the cosine similarity of user latent expression, 3.then picked top 300 last week popular items among 10,000 nearest customers as candidates.  <br>\n <br><br>\nAs for single model, this approach got the best score among three models(Private LB 330-ish).</p>\n<h2>Model3</h2>\n<p>Firstly, I created 4 types of clusters for each customer based on the techniques bellow.<br>\n(Article-based embeddings were averaged by each dimension of past bought articles to create customer embeddings)</p>\n<ul>\n<li>Swin-Transformer Image embeddings</li>\n<li>Age of each customer</li>\n<li>item2vec based article embeddings</li>\n<li>postal_code2vec (inspired by item2vec) embeddings</li>\n</ul>\n<p>\nAfter the clustering process, each customer will belong to 4 different clusters. <br>\nFor each cluster, I created an article ranking for the past 30 days, and made a weighted overall ranking depending on which clusters each customer belonged to.<br>\nRecent bought article_ids and product_codes were weighted heavier according to the customer.\nFrom the overall ranking, Model3 took the top 350 articles per user.\n</p>\n<h1>Step2: Feature Engineering</h1>\n<p>For each of the candidates, we created features as bellow.</p>\n<ul>\n<li>deep learning based latent expressions<ul>\n<li>Articles embedded expresion by BERT4REC</li>\n<li>Articles embedded expresion by LightGCN</li>\n<li>Detailed description's latent expression by Roberta (plus PCA for dimensionality reduction)</li>\n<li>Swin-Transformer based image features</li></ul></li>\n<li>cosine similarity of the article and the customer based on swin-transformer embeddings and item2vec embeddings</li>\n<li>number of past purchases/interval_days by the user for the candidate article_id, product_code, etc</li>\n<li>number of past purchases/interval_days by the cluster(clustered using item2vec) of the relevant user for the candidate article_id, product_code, etc</li>\n<li>article sales forecasts (forecasted by LightGBM regression model) for the corresponding week</li>\n</ul>\n<h1>Step3: Model</h1>\n<p>All three models were trained by LightGBMRaker.<br>\nModel1 was trained with 14 weeks,and model2/3 were only trained by using candidates from the last 1 week.</p>\n<h1>Step4: Ensembling</h1>\n<p>Ensembling was simply based by using the methodology used in many of the public notebooks (with a slight change).</p>\n<pre><code>def cust_blend(dt, W = [1,1,1], base=3): # base option is added\n\n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['model1'].split())\n    REC.append(dt['model2'].split())\n    REC.append(dt['model3'].split())\n\n\n    res = {}\n    for M in range(len(REC)):\n        if type(REC[M]) == list:\n            for n, v in enumerate(REC[M]):\n                if v in res:\n                    res[v] += (W[M]/(n+base))\n                else:\n                    res[v] = (W[M]/(n+base))\n\n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n\n    return ' '.join(res[:12])\n</code></pre>\n<h1>Failures</h1>\n<ul>\n<li>All neural network driven approaches discussed in the academic paper failed for us. Recbole and recommenders gave us marginal boost in this competition.</li>\n<li>Increasing training weeks worsened the local scores for model2 and 3.</li>\n<li>catboost (It's score was worse compared to LightGBM)</li>\n<li>general MLP</li>\n</ul>",
      "rawMarkdown": "Before going on to our solution, I would like to thank Kaggle staff and H&M team, alongside with all competitors on the leaderboard for competing with us in this competition. And of course, special thanks to @A.Sato, @taksai, and @tomo20180402 who are my teammates, for their amazing effort. We are honored to win the gold medal!\n \n# Overall Strategy\nThe overall strategy of our solution is based on four steps.\n\n1. Candidate Selection\n2. Feature Engineering\n3. Ranking\n4. Ensembling\n\n# Key Points\nWe trained three LightGBM ranker models with different candidate selection methodologies, and finally used an ensemble technique to boost our score.\nThe key idea of our solution was utilizing item2vec embeddings, which is an algorithm inspired by word2vec.\nArticle embeddings were retrieved using item2vec embeddings as they are, and customer embeddings were created by taking an average among all dimensions of all past purchased articles for each customer.\n \n# Step1: Candidate(Article) Selection\nIn the candidate selection phase, we tried to sample positive and negative samples for each customer. The candidate selection methods were different for each of the three models, and the diversity of candidates between our models leveraged our scores after ensembling.\n \n## Model1\nThe ensemble method for Model1 was a concatenation of 4 strategies. Past 14 weeks were used to train the model.\n  1. recent 50 purchased articles for each customer\n  2. top 20 products with the same product code of each article in 1.\n  3. last week top 100 articles of all customers for each sales_channel_id group\n  4. last week top 30 of all customers for each [sales_channel_id, index_group_no] group\n \n## Model2\nAside from recently purchased items (and items with the same product code of recently purchased items), my main model solely relies on the local popularity among the nearest 10,000 customers. 1.I obtained user's latent expression by getting the mean of the item2vec embedded expression (32 dimentions) of the items that the customer purchased in the past, 2.then selected 10,000 nearest customers for EACH customer based on the cosine similarity of user latent expression, 3.then picked top 300 last week popular items among 10,000 nearest customers as candidates.  \n <br />\nAs for single model, this approach got the best score among three models(Private LB 330-ish).\n \n## Model3\nFirstly, I created 4 types of clusters for each customer based on the techniques bellow.\n(Article-based embeddings were averaged by each dimension of past bought articles to create customer embeddings)\n- Swin-Transformer Image embeddings\n- Age of each customer\n- item2vec based article embeddings\n- postal_code2vec (inspired by item2vec) embeddings\n<p>\nAfter the clustering process, each customer will belong to 4 different clusters. <br />\nFor each cluster, I created an article ranking for the past 30 days, and made a weighted overall ranking depending on which clusters each customer belonged to.<br />\nRecent bought article_ids and product_codes were weighted heavier according to the customer.\nFrom the overall ranking, Model3 took the top 350 articles per user.\n</p>\n \n# Step2: Feature Engineering\nFor each of the candidates, we created features as bellow.\n- deep learning based latent expressions\n    - Articles embedded expresion by BERT4REC\n    - Articles embedded expresion by LightGCN\n    - Detailed description's latent expression by Roberta (plus PCA for dimensionality reduction)\n    - Swin-Transformer based image features\n- cosine similarity of the article and the customer based on swin-transformer embeddings and item2vec embeddings\n- number of past purchases/interval_days by the user for the candidate article_id, product_code, etc\n- number of past purchases/interval_days by the cluster(clustered using item2vec) of the relevant user for the candidate article_id, product_code, etc\n- article sales forecasts (forecasted by LightGBM regression model) for the corresponding week\n \n# Step3: Model\nAll three models were trained by LightGBMRaker.\nModel1 was trained with 14 weeks,and model2/3 were only trained by using candidates from the last 1 week.\n \n# Step4: Ensembling\nEnsembling was simply based by using the methodology used in many of the public notebooks (with a slight change).\n\n```\ndef cust_blend(dt, W = [1,1,1], base=3): # base option is added\n    \n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['model1'].split())\n    REC.append(dt['model2'].split())\n    REC.append(dt['model3'].split())\n\n\n    res = {}\n    for M in range(len(REC)):\n        if type(REC[M]) == list:\n            for n, v in enumerate(REC[M]):\n                if v in res:\n                    res[v] += (W[M]/(n+base))\n                else:\n                    res[v] = (W[M]/(n+base))\n    \n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n    \n    return ' '.join(res[:12])\n```\n \n# Failures\n- All neural network driven approaches discussed in the academic paper failed for us. Recbole and recommenders gave us marginal boost in this competition.\n- Increasing training weeks worsened the local scores for model2 and 3.\n- catboost (It's score was worse compared to LightGBM)\n- general MLP",
      "votes": null
    },
    {
      "id": "1783663",
      "postDate": "05/10/2022 15:03:10",
      "content": "<p>Awesome result! For the second model could I ask what kind of library did you use to obtain item2vec embedded expression for the items that the customer purchased in the past? The idea seems so simple and awesome! :)</p>",
      "rawMarkdown": "Awesome result! For the second model could I ask what kind of library did you use to obtain item2vec embedded expression for the items that the customer purchased in the past? The idea seems so simple and awesome! :)",
      "votes": null
    },
    {
      "id": "1784148",
      "postDate": "05/11/2022 01:08:59",
      "content": "<p>Thank you!<br>\nThe library we used is called <a href=\"https://radimrehurek.com/gensim/\" target=\"_blank\">gensim</a></p>",
      "rawMarkdown": "Thank you!\nThe library we used is called [gensim](https://radimrehurek.com/gensim/)",
      "votes": null
    },
    {
      "id": "1785013",
      "postDate": "05/11/2022 17:46:46",
      "content": "<p>Congr <a href=\"https://www.kaggle.com/tomotomo5\" target=\"_blank\">@tomotomo5</a>, keep growing.</p>",
      "rawMarkdown": "Congr @tomotomo5, keep growing.",
      "votes": null
    },
    {
      "id": "1785610",
      "postDate": "05/12/2022 08:58:49",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/muhammedtausif\" target=\"_blank\">@muhammedtausif</a> </p>",
      "rawMarkdown": "Thank you @muhammedtausif",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783663,
      "author_name": "simonelis",
      "author_url": "",
      "post_date": "05/10/2022 15:03:10",
      "content": "<p>Awesome result! For the second model could I ask what kind of library did you use to obtain item2vec embedded expression for the items that the customer purchased in the past? The idea seems so simple and awesome! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784148,
          "author_name": "tomotomo5",
          "author_url": "",
          "post_date": "05/11/2022 01:08:59",
          "content": "<p>Thank you!<br>\nThe library we used is called <a href=\"https://radimrehurek.com/gensim/\" target=\"_blank\">gensim</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1785013,
      "author_name": "muhammedtausif",
      "author_url": "",
      "post_date": "05/11/2022 17:46:46",
      "content": "<p>Congr <a href=\"https://www.kaggle.com/tomotomo5\" target=\"_blank\">@tomotomo5</a>, keep growing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1785610,
          "author_name": "tomotomo5",
          "author_url": "",
          "post_date": "05/12/2022 08:58:49",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/muhammedtausif\" target=\"_blank\">@muhammedtausif</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1783606": "Before going on to our solution, I would like to thank Kaggle staff and H&M team, alongside with all competitors on the leaderboard for competing with us in this competition. And of course, special thanks to @A.Sato, @taksai, and @tomo20180402 who are my teammates, for their amazing effort. We are honored to win the gold medal!\n \n# Overall Strategy\nThe overall strategy of our solution is based on four steps.\n\n1. Candidate Selection\n2. Feature Engineering\n3. Ranking\n4. Ensembling\n\n# Key Points\nWe trained three LightGBM ranker models with different candidate selection methodologies, and finally used an ensemble technique to boost our score.\nThe key idea of our solution was utilizing item2vec embeddings, which is an algorithm inspired by word2vec.\nArticle embeddings were retrieved using item2vec embeddings as they are, and customer embeddings were created by taking an average among all dimensions of all past purchased articles for each customer.\n \n# Step1: Candidate(Article) Selection\nIn the candidate selection phase, we tried to sample positive and negative samples for each customer. The candidate selection methods were different for each of the three models, and the diversity of candidates between our models leveraged our scores after ensembling.\n \n## Model1\nThe ensemble method for Model1 was a concatenation of 4 strategies. Past 14 weeks were used to train the model.\n  1. recent 50 purchased articles for each customer\n  2. top 20 products with the same product code of each article in 1.\n  3. last week top 100 articles of all customers for each sales_channel_id group\n  4. last week top 30 of all customers for each [sales_channel_id, index_group_no] group\n \n## Model2\nAside from recently purchased items (and items with the same product code of recently purchased items), my main model solely relies on the local popularity among the nearest 10,000 customers. 1.I obtained user's latent expression by getting the mean of the item2vec embedded expression (32 dimentions) of the items that the customer purchased in the past, 2.then selected 10,000 nearest customers for EACH customer based on the cosine similarity of user latent expression, 3.then picked top 300 last week popular items among 10,000 nearest customers as candidates.  \n <br />\nAs for single model, this approach got the best score among three models(Private LB 330-ish).\n \n## Model3\nFirstly, I created 4 types of clusters for each customer based on the techniques bellow.\n(Article-based embeddings were averaged by each dimension of past bought articles to create customer embeddings)\n- Swin-Transformer Image embeddings\n- Age of each customer\n- item2vec based article embeddings\n- postal_code2vec (inspired by item2vec) embeddings\n<p>\nAfter the clustering process, each customer will belong to 4 different clusters. <br />\nFor each cluster, I created an article ranking for the past 30 days, and made a weighted overall ranking depending on which clusters each customer belonged to.<br />\nRecent bought article_ids and product_codes were weighted heavier according to the customer.\nFrom the overall ranking, Model3 took the top 350 articles per user.\n</p>\n \n# Step2: Feature Engineering\nFor each of the candidates, we created features as bellow.\n- deep learning based latent expressions\n    - Articles embedded expresion by BERT4REC\n    - Articles embedded expresion by LightGCN\n    - Detailed description's latent expression by Roberta (plus PCA for dimensionality reduction)\n    - Swin-Transformer based image features\n- cosine similarity of the article and the customer based on swin-transformer embeddings and item2vec embeddings\n- number of past purchases/interval_days by the user for the candidate article_id, product_code, etc\n- number of past purchases/interval_days by the cluster(clustered using item2vec) of the relevant user for the candidate article_id, product_code, etc\n- article sales forecasts (forecasted by LightGBM regression model) for the corresponding week\n \n# Step3: Model\nAll three models were trained by LightGBMRaker.\nModel1 was trained with 14 weeks,and model2/3 were only trained by using candidates from the last 1 week.\n \n# Step4: Ensembling\nEnsembling was simply based by using the methodology used in many of the public notebooks (with a slight change).\n\n```\ndef cust_blend(dt, W = [1,1,1], base=3): # base option is added\n    \n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['model1'].split())\n    REC.append(dt['model2'].split())\n    REC.append(dt['model3'].split())\n\n\n    res = {}\n    for M in range(len(REC)):\n        if type(REC[M]) == list:\n            for n, v in enumerate(REC[M]):\n                if v in res:\n                    res[v] += (W[M]/(n+base))\n                else:\n                    res[v] = (W[M]/(n+base))\n    \n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n    \n    return ' '.join(res[:12])\n```\n \n# Failures\n- All neural network driven approaches discussed in the academic paper failed for us. Recbole and recommenders gave us marginal boost in this competition.\n- Increasing training weeks worsened the local scores for model2 and 3.\n- catboost (It's score was worse compared to LightGBM)\n- general MLP",
    "1783663": "Awesome result! For the second model could I ask what kind of library did you use to obtain item2vec embedded expression for the items that the customer purchased in the past? The idea seems so simple and awesome! :)",
    "1784148": "Thank you!\nThe library we used is called [gensim](https://radimrehurek.com/gensim/)",
    "1785013": "Congr @tomotomo5, keep growing.",
    "1785610": "Thank you @muhammedtausif"
  },
  "source": "meta"
}