{
  "id": 372976,
  "title": "Winning solutions summary from H&M Recommendations comp",
  "url": "/competitions/otto-recommender-system/discussion/372976",
  "author_name": "",
  "post_date": "2022-12-19T01:56:10.101959900Z",
  "votes": 40,
  "comment_count": 1,
  "views": 0,
  "content": "<h4>Winning solutions summary from H&amp;M Recommendations comp</h4>\n<p>Good morning, </p>\n<blockquote>\n  <p>Happy 4AM to you too! </p>\n</blockquote>\n<p>The followings are summaries of the winning solutions to the recommendation engine competition from earlier this year - H&amp;M Fashion Recommendation.<br>\nUse this to get ideas and as an index of the previous winning solutions. </p>\n<p>Enjoy! </p>\n<hr>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070\" target=\"_blank\">1st place</a> - By <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">SENKIN13</a></h4>\n<ul>\n<li>Candidate generation and feature engineering key to high accuracy</li>\n<li>Used retrieval strategies and GBDT model</li>\n<li>Focus on improving single lightgbm model</li>\n<li>Used negative downsampling and TreeLite to optimize model</li>\n<li>Ensemble of 5 lightgbm and 7 catboost classifiers</li>\n<li>Used label encoding and feature store to optimize performance</li>\n<li>Trained on desktop computers with 128G RAM and 300G RAM GCP instances (But can it run crysis?)</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324197\" target=\"_blank\">2nd place</a> - Written By <a href=\"https://www.kaggle.com/wht1996\" target=\"_blank\">(⊙﹏⊙)</a></h4>\n<ul>\n<li>Used LGB model to select 130 product candidates for each user based on user and product basic features, user-product combination features, age-product combination features, user-product repurchase features, and higher-order combinatorial features</li>\n<li>The solution also included itemCF feature in model</li>\n<li>Solution used a variety of recall strategies including customer and article attributes, image similarities, cooccurrences, and random graph walk over item-user graph</li>\n<li>Features for solution included customer/article attributes, similarity measures, and streaks </li>\n<li>Trained models using LGBMRanker with lambdarankmap objective</li>\n<li>Deep learning models and product embeddings did not work well in the solution</li>\n<li>Final model was an ensemble of single models from all three team members</li>\n</ul>\n<p><strong>Features:</strong></p>\n<ul>\n<li>User basic features: including num, price, sales_channel_id.</li>\n<li>Product basic features: statistics based on each attribute of the product, including times, price, age, sales_channel_id, FN, Active, club_member_status, fashion_news_frequency, last purchase time, average purchase interval.</li>\n<li>User product combination features: statistics based on each attribute of the product, including num, time, sales_channel_id, last purchase time, and average purchase interval.</li>\n<li>Age product combination features: products popularity under each age group.</li>\n<li>User product repurchase features: whether user will repurchase product and whether product will be repurchased.</li>\n<li>Higher-order combinatorial features: For example, predict when the user will next purchase the product.</li>\n<li>itemCF feature: calculate the similarity of each item through itemCF, and then calculate the score that whether user will buy the product.</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324129\" target=\"_blank\">3rd place</a> - Written By <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">SIRIUS</a></h4>\n<ul>\n<li>Pipeline consists of candidate generation and ranking</li>\n<li>Used features about recall strategy (whether item was recalled and its rank under strategy) to boost score from 0.02855 to 0.03262</li>\n<li>Used BPR matrix factorization for user2item similarity to boost score from 0.03363 to 0.03510</li>\n<li>Features for ranking included: user and item static attributes, counts of user-item interactions, similarities between user and item, item popularity, recall strategy features</li>\n<li>Models used included LightGBM, XGBoost, and CatBoost</li>\n<li>Recalled popular items, items recently purchased by user, relative items, popular items under user attributes, and items with similar prod_name to recent purchases</li>\n<li>Samples for ranking were generated by labeling items purchased in target week as 1 and those not purchased as 0, with downsampling to address imbalance</li>\n<li>Features for ranking also included: day diff of user-item interactions, jaccard similarity of item attributes, txt and image similarity, weekly trend score</li>\n<li>Final model used an average of LightGBM, XGBoost, and CatBoost scores</li>\n<li>Model was trained on transactions from previous weeks and used the BPR matrix factorization for user2item similarity</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324094\" target=\"_blank\">4th place</a> - Written By <a href=\"https://www.kaggle.com/hongweizhang\" target=\"_blank\">Hongwei Zhang</a></h4>\n<ul>\n<li>Used a pipeline consisting of generating candidates using recall models and building a ranking model to rank candidates within a customer</li>\n<li>Developed 4 recall models: item-to-item collaborative filtering, repurchasing latest products, popular ranking, and Two Tower MMoE</li>\n<li>Used lightgbm with lambda-rank objective for ranking model and attempted to implement DCN model but didn't have enough time to tune it</li>\n<li>Handled cold-start users and items by using demographic features and extracting features from product text and image</li>\n<li>Combined previous submissions using h-m-ensembling-how-to method for final submission</li>\n<li>Focused on improving performance of Two Tower MMoE model and used gating network to improve learning for recent active and non active customers</li>\n<li>Extracted TF-IDF features from product description and cluster products using SVD and K-means for text features</li>\n<li>Extracted image vectors using pre-trained model and cluster products using PCA and K-means for image features</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324098\" target=\"_blank\">5th place</a> - Written By <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">HAO</a></h4>\n<ul>\n<li>Focused on recall methods using article embeddings and faiss package</li>\n<li>Used 21 recall methods to create candidates</li>\n<li>Created embeddings using swin transformer, SentenceTransformer, tfidf, and word2vec</li>\n<li>Calibrated recall methods using shared method</li>\n<li>Divided features into articles features, customer features, and customer-article features to accelerate computation time</li>\n<li>Used lightgbm ranker model for training</li>\n<li>Used different days window to create aggregated stat features for articles and customers, as well as customer-article cross aggregated stat features and cosine similarity features using embeddings</li>\n<li>Calibrated recall methods to account for different positive ratios</li>\n<li>Accelerated computation time by dividing features into articles, customer, and customer-article</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324075\" target=\"_blank\">6th place</a> - Written By <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">HAO</a></h4>\n<ul>\n<li>Solution split into two parts: recall and rank</li>\n<li>Recall methods included: recent items purchased by user, item-based collaborative filtering, tags associated with items, and hot items for different age groups</li>\n<li>Important features included: counts of user/item purchases, time gaps between purchases, item discounts, tf-idf scores, collaborative filtering scores, and categorical features for lightgbm</li>\n<li>Best single model was catboost with cv score of 0.00403 and lb score of 0.0341, ensemble of multiple models achieved cv score of 0.0412 and lb score of 0.0348</li>\n<li>Giba's solution included using sales popularity of last 2 weeks to create a binary dataset and training model on week 90-103, validating on week 104</li>\n<li>Best features included probability of purchase and time since last purchase</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324185\" target=\"_blank\">8th place</a> - Written By <a href=\"https://www.kaggle.com/askaky\" target=\"_blank\">KAZUKI</a></h4>\n<ul>\n<li>Consists of three main steps: generating item candidates for each customer, creating features, and learning to rank with LGBMRanker</li>\n<li>The candidates are selected based on repurchased items, same product code, user-based collaborative filtering, most popular items, and most popular items grouped by age and sales channel</li>\n<li>Customer features were created but did not improve the model</li>\n<li>Article features used include article attributes (excluding article ID and product code to prevent overfitting), sales in the last N days/weeks, release date, mean price, repurchase statistics, and product code statistics</li>\n<li>Customer x Article features include the number of purchases with the same attributes as the candidate item, purchase date with the same attributes, difference between mean customer and article price, and score from user-based collaborative filtering</li>\n<li>The model used is LGBMRanker, trained on a 7-fold cross-validation with the last 3 weeks of data used as the training set and the preceding 98-104 weeks used as the validation set. The predictions from each fold are then combined using weighted ensemble.</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324127\" target=\"_blank\">9th place</a> - Written By <a href=\"https://www.kaggle.com/yzheng21\" target=\"_blank\">SABER</a></h4>\n<ul>\n<li>Validation strategy involved using the last week of data as a validation set, but k-fold validation may be more effective if there is sufficient computing power</li>\n<li>The solution used recall methods including SAR, Trending, and recommendations based on age, as well as a LightGBM model trained with 20 weeks of data and around 300 features</li>\n<li>The final submission was based on an ensemble of 4 models, 3 of which were trained using k-fold split on the training data</li>\n<li>The solution was optimized using cudf and the Forest Inference Library from Rapids for fast feature generation and tree inference</li>\n<li>Hardware used included a 32-core CPU, 128GB of memory, and a V100 32GB GPU</li>\n<li>Analysis of the data found that 90% of items in the validation set were also present in the previous 30 days and about 50% of customers in the validation set had transaction records in the previous 30 days</li>\n<li>Recall methods used included recommending the most popular items from the previous 7 days, personalized recommendations based on age, and a combination of the two.</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324223\" target=\"_blank\">10th place</a> - Written By <a href=\"https://www.kaggle.com/ferdinandlimburg\" target=\"_blank\">FLG2K</a></h4>\n<ul>\n<li>Experienced good correlation between cross-validation and leaderboard scores</li>\n<li>Avoided information leakage by training a full set of models every week</li>\n<li>Used sales prediction model to detect products that are phasing in or out</li>\n<li>Used LightFM for recommendation and item-to-item similarity</li>\n<li>Final model was an ensemble of four LightGBM models</li>\n<li>Used a desktop with 12 cores and 64GB of RAM for optimization (Not a lot on this competition)</li>\n<li>Stored features liberally and trained LightGBM model from HDF5</li>\n<li>Improved final model by ~1% compared to the single best model</li>\n</ul>",
  "messages": [
    {
      "id": "2069443",
      "postDate": "12/19/2022 01:56:10",
      "content": "<h4>Winning solutions summary from H&amp;M Recommendations comp</h4>\n<p>Good morning, </p>\n<blockquote>\n  <p>Happy 4AM to you too! </p>\n</blockquote>\n<p>The followings are summaries of the winning solutions to the recommendation engine competition from earlier this year - H&amp;M Fashion Recommendation.<br>\nUse this to get ideas and as an index of the previous winning solutions. </p>\n<p>Enjoy! </p>\n<hr>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070\" target=\"_blank\">1st place</a> - By <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">SENKIN13</a></h4>\n<ul>\n<li>Candidate generation and feature engineering key to high accuracy</li>\n<li>Used retrieval strategies and GBDT model</li>\n<li>Focus on improving single lightgbm model</li>\n<li>Used negative downsampling and TreeLite to optimize model</li>\n<li>Ensemble of 5 lightgbm and 7 catboost classifiers</li>\n<li>Used label encoding and feature store to optimize performance</li>\n<li>Trained on desktop computers with 128G RAM and 300G RAM GCP instances (But can it run crysis?)</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324197\" target=\"_blank\">2nd place</a> - Written By <a href=\"https://www.kaggle.com/wht1996\" target=\"_blank\">(⊙﹏⊙)</a></h4>\n<ul>\n<li>Used LGB model to select 130 product candidates for each user based on user and product basic features, user-product combination features, age-product combination features, user-product repurchase features, and higher-order combinatorial features</li>\n<li>The solution also included itemCF feature in model</li>\n<li>Solution used a variety of recall strategies including customer and article attributes, image similarities, cooccurrences, and random graph walk over item-user graph</li>\n<li>Features for solution included customer/article attributes, similarity measures, and streaks </li>\n<li>Trained models using LGBMRanker with lambdarankmap objective</li>\n<li>Deep learning models and product embeddings did not work well in the solution</li>\n<li>Final model was an ensemble of single models from all three team members</li>\n</ul>\n<p><strong>Features:</strong></p>\n<ul>\n<li>User basic features: including num, price, sales_channel_id.</li>\n<li>Product basic features: statistics based on each attribute of the product, including times, price, age, sales_channel_id, FN, Active, club_member_status, fashion_news_frequency, last purchase time, average purchase interval.</li>\n<li>User product combination features: statistics based on each attribute of the product, including num, time, sales_channel_id, last purchase time, and average purchase interval.</li>\n<li>Age product combination features: products popularity under each age group.</li>\n<li>User product repurchase features: whether user will repurchase product and whether product will be repurchased.</li>\n<li>Higher-order combinatorial features: For example, predict when the user will next purchase the product.</li>\n<li>itemCF feature: calculate the similarity of each item through itemCF, and then calculate the score that whether user will buy the product.</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324129\" target=\"_blank\">3rd place</a> - Written By <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">SIRIUS</a></h4>\n<ul>\n<li>Pipeline consists of candidate generation and ranking</li>\n<li>Used features about recall strategy (whether item was recalled and its rank under strategy) to boost score from 0.02855 to 0.03262</li>\n<li>Used BPR matrix factorization for user2item similarity to boost score from 0.03363 to 0.03510</li>\n<li>Features for ranking included: user and item static attributes, counts of user-item interactions, similarities between user and item, item popularity, recall strategy features</li>\n<li>Models used included LightGBM, XGBoost, and CatBoost</li>\n<li>Recalled popular items, items recently purchased by user, relative items, popular items under user attributes, and items with similar prod_name to recent purchases</li>\n<li>Samples for ranking were generated by labeling items purchased in target week as 1 and those not purchased as 0, with downsampling to address imbalance</li>\n<li>Features for ranking also included: day diff of user-item interactions, jaccard similarity of item attributes, txt and image similarity, weekly trend score</li>\n<li>Final model used an average of LightGBM, XGBoost, and CatBoost scores</li>\n<li>Model was trained on transactions from previous weeks and used the BPR matrix factorization for user2item similarity</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324094\" target=\"_blank\">4th place</a> - Written By <a href=\"https://www.kaggle.com/hongweizhang\" target=\"_blank\">Hongwei Zhang</a></h4>\n<ul>\n<li>Used a pipeline consisting of generating candidates using recall models and building a ranking model to rank candidates within a customer</li>\n<li>Developed 4 recall models: item-to-item collaborative filtering, repurchasing latest products, popular ranking, and Two Tower MMoE</li>\n<li>Used lightgbm with lambda-rank objective for ranking model and attempted to implement DCN model but didn't have enough time to tune it</li>\n<li>Handled cold-start users and items by using demographic features and extracting features from product text and image</li>\n<li>Combined previous submissions using h-m-ensembling-how-to method for final submission</li>\n<li>Focused on improving performance of Two Tower MMoE model and used gating network to improve learning for recent active and non active customers</li>\n<li>Extracted TF-IDF features from product description and cluster products using SVD and K-means for text features</li>\n<li>Extracted image vectors using pre-trained model and cluster products using PCA and K-means for image features</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324098\" target=\"_blank\">5th place</a> - Written By <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">HAO</a></h4>\n<ul>\n<li>Focused on recall methods using article embeddings and faiss package</li>\n<li>Used 21 recall methods to create candidates</li>\n<li>Created embeddings using swin transformer, SentenceTransformer, tfidf, and word2vec</li>\n<li>Calibrated recall methods using shared method</li>\n<li>Divided features into articles features, customer features, and customer-article features to accelerate computation time</li>\n<li>Used lightgbm ranker model for training</li>\n<li>Used different days window to create aggregated stat features for articles and customers, as well as customer-article cross aggregated stat features and cosine similarity features using embeddings</li>\n<li>Calibrated recall methods to account for different positive ratios</li>\n<li>Accelerated computation time by dividing features into articles, customer, and customer-article</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324075\" target=\"_blank\">6th place</a> - Written By <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">HAO</a></h4>\n<ul>\n<li>Solution split into two parts: recall and rank</li>\n<li>Recall methods included: recent items purchased by user, item-based collaborative filtering, tags associated with items, and hot items for different age groups</li>\n<li>Important features included: counts of user/item purchases, time gaps between purchases, item discounts, tf-idf scores, collaborative filtering scores, and categorical features for lightgbm</li>\n<li>Best single model was catboost with cv score of 0.00403 and lb score of 0.0341, ensemble of multiple models achieved cv score of 0.0412 and lb score of 0.0348</li>\n<li>Giba's solution included using sales popularity of last 2 weeks to create a binary dataset and training model on week 90-103, validating on week 104</li>\n<li>Best features included probability of purchase and time since last purchase</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324185\" target=\"_blank\">8th place</a> - Written By <a href=\"https://www.kaggle.com/askaky\" target=\"_blank\">KAZUKI</a></h4>\n<ul>\n<li>Consists of three main steps: generating item candidates for each customer, creating features, and learning to rank with LGBMRanker</li>\n<li>The candidates are selected based on repurchased items, same product code, user-based collaborative filtering, most popular items, and most popular items grouped by age and sales channel</li>\n<li>Customer features were created but did not improve the model</li>\n<li>Article features used include article attributes (excluding article ID and product code to prevent overfitting), sales in the last N days/weeks, release date, mean price, repurchase statistics, and product code statistics</li>\n<li>Customer x Article features include the number of purchases with the same attributes as the candidate item, purchase date with the same attributes, difference between mean customer and article price, and score from user-based collaborative filtering</li>\n<li>The model used is LGBMRanker, trained on a 7-fold cross-validation with the last 3 weeks of data used as the training set and the preceding 98-104 weeks used as the validation set. The predictions from each fold are then combined using weighted ensemble.</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324127\" target=\"_blank\">9th place</a> - Written By <a href=\"https://www.kaggle.com/yzheng21\" target=\"_blank\">SABER</a></h4>\n<ul>\n<li>Validation strategy involved using the last week of data as a validation set, but k-fold validation may be more effective if there is sufficient computing power</li>\n<li>The solution used recall methods including SAR, Trending, and recommendations based on age, as well as a LightGBM model trained with 20 weeks of data and around 300 features</li>\n<li>The final submission was based on an ensemble of 4 models, 3 of which were trained using k-fold split on the training data</li>\n<li>The solution was optimized using cudf and the Forest Inference Library from Rapids for fast feature generation and tree inference</li>\n<li>Hardware used included a 32-core CPU, 128GB of memory, and a V100 32GB GPU</li>\n<li>Analysis of the data found that 90% of items in the validation set were also present in the previous 30 days and about 50% of customers in the validation set had transaction records in the previous 30 days</li>\n<li>Recall methods used included recommending the most popular items from the previous 7 days, personalized recommendations based on age, and a combination of the two.</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324223\" target=\"_blank\">10th place</a> - Written By <a href=\"https://www.kaggle.com/ferdinandlimburg\" target=\"_blank\">FLG2K</a></h4>\n<ul>\n<li>Experienced good correlation between cross-validation and leaderboard scores</li>\n<li>Avoided information leakage by training a full set of models every week</li>\n<li>Used sales prediction model to detect products that are phasing in or out</li>\n<li>Used LightFM for recommendation and item-to-item similarity</li>\n<li>Final model was an ensemble of four LightGBM models</li>\n<li>Used a desktop with 12 cores and 64GB of RAM for optimization (Not a lot on this competition)</li>\n<li>Stored features liberally and trained LightGBM model from HDF5</li>\n<li>Improved final model by ~1% compared to the single best model</li>\n</ul>",
      "rawMarkdown": "#### Winning solutions summary from H&M Recommendations comp\n\nGood morning, \n\n> Happy 4AM to you too! \n\nThe followings are summaries of the winning solutions to the recommendation engine competition from earlier this year - H&M Fashion Recommendation.\nUse this to get ideas and as an index of the previous winning solutions. \n\nEnjoy! \n\n_____\n\n\n#### [1st place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070) - By [SENKIN13](https://www.kaggle.com/senkin13)\n\n- Candidate generation and feature engineering key to high accuracy\n- Used retrieval strategies and GBDT model\n- Focus on improving single lightgbm model\n- Used negative downsampling and TreeLite to optimize model\n- Ensemble of 5 lightgbm and 7 catboost classifiers\n- Used label encoding and feature store to optimize performance\n- Trained on desktop computers with 128G RAM and 300G RAM GCP instances (But can it run crysis?)\n\n\n\n#### [2nd place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324197) - Written By [(⊙﹏⊙)](https://www.kaggle.com/wht1996)\n\n- Used LGB model to select 130 product candidates for each user based on user and product basic features, user-product combination features, age-product combination features, user-product repurchase features, and higher-order combinatorial features\n- The solution also included itemCF feature in model\n- Solution used a variety of recall strategies including customer and article attributes, image similarities, cooccurrences, and random graph walk over item-user graph\n- Features for solution included customer/article attributes, similarity measures, and streaks \n- Trained models using LGBMRanker with lambdarankmap objective\n- Deep learning models and product embeddings did not work well in the solution\n- Final model was an ensemble of single models from all three team members\n\n**Features:**\n- User basic features: including num, price, sales_channel_id.\n- Product basic features: statistics based on each attribute of the product, including times, price, age, sales_channel_id, FN, Active, club_member_status, fashion_news_frequency, last purchase time, average purchase interval.\n- User product combination features: statistics based on each attribute of the product, including num, time, sales_channel_id, last purchase time, and average purchase interval.\n- Age product combination features: products popularity under each age group.\n- User product repurchase features: whether user will repurchase product and whether product will be repurchased.\n- Higher-order combinatorial features: For example, predict when the user will next purchase the product.\n- itemCF feature: calculate the similarity of each item through itemCF, and then calculate the score that whether user will buy the product.\n\n\n\n#### [3rd place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324129) - Written By [SIRIUS](https://www.kaggle.com/sirius81)\n\n- Pipeline consists of candidate generation and ranking\n- Used features about recall strategy (whether item was recalled and its rank under strategy) to boost score from 0.02855 to 0.03262\n- Used BPR matrix factorization for user2item similarity to boost score from 0.03363 to 0.03510\n- Features for ranking included: user and item static attributes, counts of user-item interactions, similarities between user and item, item popularity, recall strategy features\n- Models used included LightGBM, XGBoost, and CatBoost\n- Recalled popular items, items recently purchased by user, relative items, popular items under user attributes, and items with similar prod_name to recent purchases\n- Samples for ranking were generated by labeling items purchased in target week as 1 and those not purchased as 0, with downsampling to address imbalance\n- Features for ranking also included: day diff of user-item interactions, jaccard similarity of item attributes, txt and image similarity, weekly trend score\n- Final model used an average of LightGBM, XGBoost, and CatBoost scores\n- Model was trained on transactions from previous weeks and used the BPR matrix factorization for user2item similarity\n\n\n\n#### [4th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324094) - Written By [Hongwei Zhang](https://www.kaggle.com/hongweizhang)\n\n- Used a pipeline consisting of generating candidates using recall models and building a ranking model to rank candidates within a customer\n- Developed 4 recall models: item-to-item collaborative filtering, repurchasing latest products, popular ranking, and Two Tower MMoE\n- Used lightgbm with lambda-rank objective for ranking model and attempted to implement DCN model but didn't have enough time to tune it\n- Handled cold-start users and items by using demographic features and extracting features from product text and image\n- Combined previous submissions using h-m-ensembling-how-to method for final submission\n- Focused on improving performance of Two Tower MMoE model and used gating network to improve learning for recent active and non active customers\n- Extracted TF-IDF features from product description and cluster products using SVD and K-means for text features\n- Extracted image vectors using pre-trained model and cluster products using PCA and K-means for image features\n\n\n\n#### [5th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324098) - Written By [HAO](https://www.kaggle.com/lihaorocky)\n\n- Focused on recall methods using article embeddings and faiss package\n- Used 21 recall methods to create candidates\n- Created embeddings using swin transformer, SentenceTransformer, tfidf, and word2vec\n- Calibrated recall methods using shared method\n- Divided features into articles features, customer features, and customer-article features to accelerate computation time\n- Used lightgbm ranker model for training\n- Used different days window to create aggregated stat features for articles and customers, as well as customer-article cross aggregated stat features and cosine similarity features using embeddings\n- Calibrated recall methods to account for different positive ratios\n- Accelerated computation time by dividing features into articles, customer, and customer-article\n\n\n\n#### [6th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324075) - Written By [HAO](https://www.kaggle.com/lihaorocky)\n\n- Solution split into two parts: recall and rank\n- Recall methods included: recent items purchased by user, item-based collaborative filtering, tags associated with items, and hot items for different age groups\n- Important features included: counts of user/item purchases, time gaps between purchases, item discounts, tf-idf scores, collaborative filtering scores, and categorical features for lightgbm\n- Best single model was catboost with cv score of 0.00403 and lb score of 0.0341, ensemble of multiple models achieved cv score of 0.0412 and lb score of 0.0348\n- Giba's solution included using sales popularity of last 2 weeks to create a binary dataset and training model on week 90-103, validating on week 104\n- Best features included probability of purchase and time since last purchase\n\n\n\n#### [8th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324185) - Written By [KAZUKI](https://www.kaggle.com/askaky)\n\n- Consists of three main steps: generating item candidates for each customer, creating features, and learning to rank with LGBMRanker\n- The candidates are selected based on repurchased items, same product code, user-based collaborative filtering, most popular items, and most popular items grouped by age and sales channel\n- Customer features were created but did not improve the model\n- Article features used include article attributes (excluding article ID and product code to prevent overfitting), sales in the last N days/weeks, release date, mean price, repurchase statistics, and product code statistics\n- Customer x Article features include the number of purchases with the same attributes as the candidate item, purchase date with the same attributes, difference between mean customer and article price, and score from user-based collaborative filtering\n- The model used is LGBMRanker, trained on a 7-fold cross-validation with the last 3 weeks of data used as the training set and the preceding 98-104 weeks used as the validation set. The predictions from each fold are then combined using weighted ensemble.\n\n\n\n#### [9th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324127) - Written By [SABER](https://www.kaggle.com/yzheng21)\n\n\n- Validation strategy involved using the last week of data as a validation set, but k-fold validation may be more effective if there is sufficient computing power\n- The solution used recall methods including SAR, Trending, and recommendations based on age, as well as a LightGBM model trained with 20 weeks of data and around 300 features\n- The final submission was based on an ensemble of 4 models, 3 of which were trained using k-fold split on the training data\n- The solution was optimized using cudf and the Forest Inference Library from Rapids for fast feature generation and tree inference\n- Hardware used included a 32-core CPU, 128GB of memory, and a V100 32GB GPU\n- Analysis of the data found that 90% of items in the validation set were also present in the previous 30 days and about 50% of customers in the validation set had transaction records in the previous 30 days\n- Recall methods used included recommending the most popular items from the previous 7 days, personalized recommendations based on age, and a combination of the two.\n\n\n\n#### [10th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324223) - Written By [FLG2K](https://www.kaggle.com/ferdinandlimburg)\n\n\n- Experienced good correlation between cross-validation and leaderboard scores\n- Avoided information leakage by training a full set of models every week\n- Used sales prediction model to detect products that are phasing in or out\n- Used LightFM for recommendation and item-to-item similarity\n- Final model was an ensemble of four LightGBM models\n- Used a desktop with 12 cores and 64GB of RAM for optimization (Not a lot on this competition)\n- Stored features liberally and trained LightGBM model from HDF5\n- Improved final model by ~1% compared to the single best model",
      "votes": null
    },
    {
      "id": "2070115",
      "postDate": "12/19/2022 16:16:27",
      "content": "<p>This is great, I wonder if other recommendation competition solutions have been so heavily dominated by LightGBM models? </p>",
      "rawMarkdown": "This is great, I wonder if other recommendation competition solutions have been so heavily dominated by LightGBM models?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2070115,
      "author_name": "ethanmd0519",
      "author_url": "",
      "post_date": "12/19/2022 16:16:27",
      "content": "<p>This is great, I wonder if other recommendation competition solutions have been so heavily dominated by LightGBM models? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2069443": "#### Winning solutions summary from H&M Recommendations comp\n\nGood morning, \n\n> Happy 4AM to you too! \n\nThe followings are summaries of the winning solutions to the recommendation engine competition from earlier this year - H&M Fashion Recommendation.\nUse this to get ideas and as an index of the previous winning solutions. \n\nEnjoy! \n\n_____\n\n\n#### [1st place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324070) - By [SENKIN13](https://www.kaggle.com/senkin13)\n\n- Candidate generation and feature engineering key to high accuracy\n- Used retrieval strategies and GBDT model\n- Focus on improving single lightgbm model\n- Used negative downsampling and TreeLite to optimize model\n- Ensemble of 5 lightgbm and 7 catboost classifiers\n- Used label encoding and feature store to optimize performance\n- Trained on desktop computers with 128G RAM and 300G RAM GCP instances (But can it run crysis?)\n\n\n\n#### [2nd place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324197) - Written By [(⊙﹏⊙)](https://www.kaggle.com/wht1996)\n\n- Used LGB model to select 130 product candidates for each user based on user and product basic features, user-product combination features, age-product combination features, user-product repurchase features, and higher-order combinatorial features\n- The solution also included itemCF feature in model\n- Solution used a variety of recall strategies including customer and article attributes, image similarities, cooccurrences, and random graph walk over item-user graph\n- Features for solution included customer/article attributes, similarity measures, and streaks \n- Trained models using LGBMRanker with lambdarankmap objective\n- Deep learning models and product embeddings did not work well in the solution\n- Final model was an ensemble of single models from all three team members\n\n**Features:**\n- User basic features: including num, price, sales_channel_id.\n- Product basic features: statistics based on each attribute of the product, including times, price, age, sales_channel_id, FN, Active, club_member_status, fashion_news_frequency, last purchase time, average purchase interval.\n- User product combination features: statistics based on each attribute of the product, including num, time, sales_channel_id, last purchase time, and average purchase interval.\n- Age product combination features: products popularity under each age group.\n- User product repurchase features: whether user will repurchase product and whether product will be repurchased.\n- Higher-order combinatorial features: For example, predict when the user will next purchase the product.\n- itemCF feature: calculate the similarity of each item through itemCF, and then calculate the score that whether user will buy the product.\n\n\n\n#### [3rd place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324129) - Written By [SIRIUS](https://www.kaggle.com/sirius81)\n\n- Pipeline consists of candidate generation and ranking\n- Used features about recall strategy (whether item was recalled and its rank under strategy) to boost score from 0.02855 to 0.03262\n- Used BPR matrix factorization for user2item similarity to boost score from 0.03363 to 0.03510\n- Features for ranking included: user and item static attributes, counts of user-item interactions, similarities between user and item, item popularity, recall strategy features\n- Models used included LightGBM, XGBoost, and CatBoost\n- Recalled popular items, items recently purchased by user, relative items, popular items under user attributes, and items with similar prod_name to recent purchases\n- Samples for ranking were generated by labeling items purchased in target week as 1 and those not purchased as 0, with downsampling to address imbalance\n- Features for ranking also included: day diff of user-item interactions, jaccard similarity of item attributes, txt and image similarity, weekly trend score\n- Final model used an average of LightGBM, XGBoost, and CatBoost scores\n- Model was trained on transactions from previous weeks and used the BPR matrix factorization for user2item similarity\n\n\n\n#### [4th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324094) - Written By [Hongwei Zhang](https://www.kaggle.com/hongweizhang)\n\n- Used a pipeline consisting of generating candidates using recall models and building a ranking model to rank candidates within a customer\n- Developed 4 recall models: item-to-item collaborative filtering, repurchasing latest products, popular ranking, and Two Tower MMoE\n- Used lightgbm with lambda-rank objective for ranking model and attempted to implement DCN model but didn't have enough time to tune it\n- Handled cold-start users and items by using demographic features and extracting features from product text and image\n- Combined previous submissions using h-m-ensembling-how-to method for final submission\n- Focused on improving performance of Two Tower MMoE model and used gating network to improve learning for recent active and non active customers\n- Extracted TF-IDF features from product description and cluster products using SVD and K-means for text features\n- Extracted image vectors using pre-trained model and cluster products using PCA and K-means for image features\n\n\n\n#### [5th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324098) - Written By [HAO](https://www.kaggle.com/lihaorocky)\n\n- Focused on recall methods using article embeddings and faiss package\n- Used 21 recall methods to create candidates\n- Created embeddings using swin transformer, SentenceTransformer, tfidf, and word2vec\n- Calibrated recall methods using shared method\n- Divided features into articles features, customer features, and customer-article features to accelerate computation time\n- Used lightgbm ranker model for training\n- Used different days window to create aggregated stat features for articles and customers, as well as customer-article cross aggregated stat features and cosine similarity features using embeddings\n- Calibrated recall methods to account for different positive ratios\n- Accelerated computation time by dividing features into articles, customer, and customer-article\n\n\n\n#### [6th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324075) - Written By [HAO](https://www.kaggle.com/lihaorocky)\n\n- Solution split into two parts: recall and rank\n- Recall methods included: recent items purchased by user, item-based collaborative filtering, tags associated with items, and hot items for different age groups\n- Important features included: counts of user/item purchases, time gaps between purchases, item discounts, tf-idf scores, collaborative filtering scores, and categorical features for lightgbm\n- Best single model was catboost with cv score of 0.00403 and lb score of 0.0341, ensemble of multiple models achieved cv score of 0.0412 and lb score of 0.0348\n- Giba's solution included using sales popularity of last 2 weeks to create a binary dataset and training model on week 90-103, validating on week 104\n- Best features included probability of purchase and time since last purchase\n\n\n\n#### [8th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324185) - Written By [KAZUKI](https://www.kaggle.com/askaky)\n\n- Consists of three main steps: generating item candidates for each customer, creating features, and learning to rank with LGBMRanker\n- The candidates are selected based on repurchased items, same product code, user-based collaborative filtering, most popular items, and most popular items grouped by age and sales channel\n- Customer features were created but did not improve the model\n- Article features used include article attributes (excluding article ID and product code to prevent overfitting), sales in the last N days/weeks, release date, mean price, repurchase statistics, and product code statistics\n- Customer x Article features include the number of purchases with the same attributes as the candidate item, purchase date with the same attributes, difference between mean customer and article price, and score from user-based collaborative filtering\n- The model used is LGBMRanker, trained on a 7-fold cross-validation with the last 3 weeks of data used as the training set and the preceding 98-104 weeks used as the validation set. The predictions from each fold are then combined using weighted ensemble.\n\n\n\n#### [9th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324127) - Written By [SABER](https://www.kaggle.com/yzheng21)\n\n\n- Validation strategy involved using the last week of data as a validation set, but k-fold validation may be more effective if there is sufficient computing power\n- The solution used recall methods including SAR, Trending, and recommendations based on age, as well as a LightGBM model trained with 20 weeks of data and around 300 features\n- The final submission was based on an ensemble of 4 models, 3 of which were trained using k-fold split on the training data\n- The solution was optimized using cudf and the Forest Inference Library from Rapids for fast feature generation and tree inference\n- Hardware used included a 32-core CPU, 128GB of memory, and a V100 32GB GPU\n- Analysis of the data found that 90% of items in the validation set were also present in the previous 30 days and about 50% of customers in the validation set had transaction records in the previous 30 days\n- Recall methods used included recommending the most popular items from the previous 7 days, personalized recommendations based on age, and a combination of the two.\n\n\n\n#### [10th place](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324223) - Written By [FLG2K](https://www.kaggle.com/ferdinandlimburg)\n\n\n- Experienced good correlation between cross-validation and leaderboard scores\n- Avoided information leakage by training a full set of models every week\n- Used sales prediction model to detect products that are phasing in or out\n- Used LightFM for recommendation and item-to-item similarity\n- Final model was an ensemble of four LightGBM models\n- Used a desktop with 12 cores and 64GB of RAM for optimization (Not a lot on this competition)\n- Stored features liberally and trained LightGBM model from HDF5\n- Improved final model by ~1% compared to the single best model",
    "2070115": "This is great, I wonder if other recommendation competition solutions have been so heavily dominated by LightGBM models?"
  },
  "source": "meta"
}