{
  "id": 324075,
  "title": "6th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/hard2rec-6th-place-solution",
  "author_name": "",
  "post_date": "2022-05-12T15:48:27.593Z",
  "votes": 71,
  "comment_count": 24,
  "views": 0,
  "content": "<p>First of all, thanks to competition organizers for hosting this interesting competition and my great teammates(@Giba, <a href=\"https://www.kaggle.com/qysx\" target=\"_blank\">@qysx</a>, <a href=\"https://www.kaggle.com/hyd\" target=\"_blank\">@hyd</a>). Secondly, thanks to the kaggle community, I have learned a lot from those great kernels and disscussions.</p>\n<p>In this competition, we were asked to recommend products to users based on their historical purchasing behavior. To solve this problem, our solution splits into two parts: recall&amp;rank. Good strategy for generating candidates and feature engineering are both important to win a good place.</p>\n<h3>Recall</h3>\n<p>We used several methods to generate different candidates, in order to improve the coverage of positive samples(both user and user-item).</p>\n<ol>\n<li>u2i: items that the user recently purchased</li>\n<li>i2i: item based collaborative filtering</li>\n<li>u2tag2i: tag can be \"product_code\", \"product_type_no\", \"department_no\", \"section_no\"</li>\n<li>hot: generate hot items for different age_bins's user</li>\n</ol>\n<h3>Features</h3>\n<p>Most important features are as follows,</p>\n<ol>\n<li>count: user, item, user-item based, it's important to use all historical data to generate the features of user.</li>\n<li>gap: user, item, user-item first/last purchase time to now. For item's first purchased time, it may  represent the \"release time\" of this item. So we can use \"user\" as key to calculate the statistics of \"release time\" , it may help us to get user's preference for new or old items.</li>\n<li>items' discount: use max/mean to represent the common price of the item. use the price for each row to calculate the discount of the item.So we can get user's preference for discounted items and the number of days which a item is on sale recently.</li>\n<li>tfidf: set user's historical purchased article/product_code/product_type_no to a sentence, then use tfidf+svd to generate feature. </li>\n<li>item_sim: collaborative filtering score of i2i, calculate the items between the candidates and user's historical purchased.</li>\n<li>set categorical_features(\"department_no\", \"product_type_no\", etc.) for lightgbm will improve cv score about 0.0005~0.0008. </li>\n</ol>\n<h3>Model</h3>\n<p>We use last 9 weeks for training(last week for validation).Our best single model is catboost with cv score(0.00403) and lb score(0.0341). We use several models to ensemble(cv: 0.0412, lb: 0.0348). And we use the kernels(<a href=\"url\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/titericz/h-m-ensembling-how-to\" target=\"_blank\">https://www.kaggle.com/code/titericz/h-m-ensembling-how-to</a> )Giba shared before to handle the different candidates list when blending.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>objective</th>\n<th>average num of candidates for each user</th>\n<th>cv map@12</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>catboost</td>\n<td>binary</td>\n<td>120</td>\n<td>0.0403</td>\n</tr>\n<tr>\n<td>lightgbm</td>\n<td>binary</td>\n<td>120</td>\n<td>0.0395</td>\n</tr>\n<tr>\n<td>catboost</td>\n<td>binary</td>\n<td>220</td>\n<td>0.0402</td>\n</tr>\n<tr>\n<td>lightgbm</td>\n<td>binary</td>\n<td>220</td>\n<td>0.0396</td>\n</tr>\n<tr>\n<td>lightgbm</td>\n<td>lambdarank</td>\n<td>220</td>\n<td>0.0400</td>\n</tr>\n<tr>\n<td>lightgbm</td>\n<td>lambdarank</td>\n<td>1000</td>\n<td>0.0381</td>\n</tr>\n</tbody>\n</table>\n<p>The result of our best cv is the best one on PB. Trust your local cv is really importance in kaggle.</p>\n<h3>Giba's solution</h3>\n<p>Set around 10k article_id candidate list based in sales popularity of last 2 weeks. Build a binary dataset with all combinations of customer_id x 10k article_id . Build that huge dataset by batches and using cudf to speedup the entire process. Create a binary trainset from weeks 90 to 104. Train model on week 90-103, validate on week 104. To train the model select random 200, 300 or 500 negative candidates for each customer_id. Blend models trained using LightGBM lambdarank, LightGBM xendcg, XGBoost  lambdarank and CatBoost ranker. Best features based in the probability of purchase each article in a given period range and time passed since last purchase.<br>\nMore details on <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324278\" target=\"_blank\">Giba Solution</a>.</p>\n<h3>To be continued</h3>",
  "messages": [
    {
      "id": "1782959",
      "postDate": "05/10/2022 03:02:53",
      "content": "<p>First of all, thanks to competition organizers for hosting this interesting competition and my great teammates(@Giba, <a href=\"https://www.kaggle.com/qysx\" target=\"_blank\">@qysx</a>, <a href=\"https://www.kaggle.com/hyd\" target=\"_blank\">@hyd</a>). Secondly, thanks to the kaggle community, I have learned a lot from those great kernels and disscussions.</p>\n<p>In this competition, we were asked to recommend products to users based on their historical purchasing behavior. To solve this problem, our solution splits into two parts: recall&amp;rank. Good strategy for generating candidates and feature engineering are both important to win a good place.</p>\n<h3>Recall</h3>\n<p>We used several methods to generate different candidates, in order to improve the coverage of positive samples(both user and user-item).</p>\n<ol>\n<li>u2i: items that the user recently purchased</li>\n<li>i2i: item based collaborative filtering</li>\n<li>u2tag2i: tag can be \"product_code\", \"product_type_no\", \"department_no\", \"section_no\"</li>\n<li>hot: generate hot items for different age_bins's user</li>\n</ol>\n<h3>Features</h3>\n<p>Most important features are as follows,</p>\n<ol>\n<li>count: user, item, user-item based, it's important to use all historical data to generate the features of user.</li>\n<li>gap: user, item, user-item first/last purchase time to now. For item's first purchased time, it may  represent the \"release time\" of this item. So we can use \"user\" as key to calculate the statistics of \"release time\" , it may help us to get user's preference for new or old items.</li>\n<li>items' discount: use max/mean to represent the common price of the item. use the price for each row to calculate the discount of the item.So we can get user's preference for discounted items and the number of days which a item is on sale recently.</li>\n<li>tfidf: set user's historical purchased article/product_code/product_type_no to a sentence, then use tfidf+svd to generate feature. </li>\n<li>item_sim: collaborative filtering score of i2i, calculate the items between the candidates and user's historical purchased.</li>\n<li>set categorical_features(\"department_no\", \"product_type_no\", etc.) for lightgbm will improve cv score about 0.0005~0.0008. </li>\n</ol>\n<h3>Model</h3>\n<p>We use last 9 weeks for training(last week for validation).Our best single model is catboost with cv score(0.00403) and lb score(0.0341). We use several models to ensemble(cv: 0.0412, lb: 0.0348). And we use the kernels(<a href=\"url\" target=\"_blank\"></a><a href=\"https://www.kaggle.com/code/titericz/h-m-ensembling-how-to\" target=\"_blank\">https://www.kaggle.com/code/titericz/h-m-ensembling-how-to</a> )Giba shared before to handle the different candidates list when blending.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>objective</th>\n<th>average num of candidates for each user</th>\n<th>cv map@12</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>catboost</td>\n<td>binary</td>\n<td>120</td>\n<td>0.0403</td>\n</tr>\n<tr>\n<td>lightgbm</td>\n<td>binary</td>\n<td>120</td>\n<td>0.0395</td>\n</tr>\n<tr>\n<td>catboost</td>\n<td>binary</td>\n<td>220</td>\n<td>0.0402</td>\n</tr>\n<tr>\n<td>lightgbm</td>\n<td>binary</td>\n<td>220</td>\n<td>0.0396</td>\n</tr>\n<tr>\n<td>lightgbm</td>\n<td>lambdarank</td>\n<td>220</td>\n<td>0.0400</td>\n</tr>\n<tr>\n<td>lightgbm</td>\n<td>lambdarank</td>\n<td>1000</td>\n<td>0.0381</td>\n</tr>\n</tbody>\n</table>\n<p>The result of our best cv is the best one on PB. Trust your local cv is really importance in kaggle.</p>\n<h3>Giba's solution</h3>\n<p>Set around 10k article_id candidate list based in sales popularity of last 2 weeks. Build a binary dataset with all combinations of customer_id x 10k article_id . Build that huge dataset by batches and using cudf to speedup the entire process. Create a binary trainset from weeks 90 to 104. Train model on week 90-103, validate on week 104. To train the model select random 200, 300 or 500 negative candidates for each customer_id. Blend models trained using LightGBM lambdarank, LightGBM xendcg, XGBoost  lambdarank and CatBoost ranker. Best features based in the probability of purchase each article in a given period range and time passed since last purchase.<br>\nMore details on <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324278\" target=\"_blank\">Giba Solution</a>.</p>\n<h3>To be continued</h3>",
      "rawMarkdown": "First of all, thanks to competition organizers for hosting this interesting competition and my great teammates(@Giba, @qysx, @hyd). Secondly, thanks to the kaggle community, I have learned a lot from those great kernels and disscussions.\n\nIn this competition, we were asked to recommend products to users based on their historical purchasing behavior. To solve this problem, our solution splits into two parts: recall&rank. Good strategy for generating candidates and feature engineering are both important to win a good place.\n\n### Recall ###\nWe used several methods to generate different candidates, in order to improve the coverage of positive samples(both user and user-item).\n1. u2i: items that the user recently purchased\n2. i2i: item based collaborative filtering\n3. u2tag2i: tag can be \"product_code\", \"product_type_no\", \"department_no\", \"section_no\"\n4. hot: generate hot items for different age_bins's user\n\n### Features ###\nMost important features are as follows,\n1. count: user, item, user-item based, it's important to use all historical data to generate the features of user.\n2. gap: user, item, user-item first/last purchase time to now. For item's first purchased time, it may  represent the \"release time\" of this item. So we can use \"user\" as key to calculate the statistics of \"release time\" , it may help us to get user's preference for new or old items.\n3. items' discount: use max/mean to represent the common price of the item. use the price for each row to calculate the discount of the item.So we can get user's preference for discounted items and the number of days which a item is on sale recently.\n4. tfidf: set user's historical purchased article/product_code/product_type_no to a sentence, then use tfidf+svd to generate feature. \n5. item_sim: collaborative filtering score of i2i, calculate the items between the candidates and user's historical purchased.\n6. set categorical_features(\"department_no\", \"product_type_no\", etc.) for lightgbm will improve cv score about 0.0005~0.0008. \n\n### Model ###\nWe use last 9 weeks for training(last week for validation).Our best single model is catboost with cv score(0.00403) and lb score(0.0341). We use several models to ensemble(cv: 0.0412, lb: 0.0348). And we use the kernels([https://www.kaggle.com/code/titericz/h-m-ensembling-how-to ](url))Giba shared before to handle the different candidates list when blending.\n\n| model | objective| average num of candidates for each user | cv map@12|\n| --- | --- |---|---|\n| catboost | binary | 120 |0.0403 |\n| lightgbm | binary | 120|0.0395 |\n| catboost| binary| 220|0.0402 |\n|lightgbm| binary| 220|0.0396 |\n|lightgbm|lambdarank|220|0.0400 |\n|lightgbm|lambdarank|1000|0.0381|\n\nThe result of our best cv is the best one on PB. Trust your local cv is really importance in kaggle.\n\n### Giba's solution ###\nSet around 10k article_id candidate list based in sales popularity of last 2 weeks. Build a binary dataset with all combinations of customer_id x 10k article_id . Build that huge dataset by batches and using cudf to speedup the entire process. Create a binary trainset from weeks 90 to 104. Train model on week 90-103, validate on week 104. To train the model select random 200, 300 or 500 negative candidates for each customer_id. Blend models trained using LightGBM lambdarank, LightGBM xendcg, XGBoost  lambdarank and CatBoost ranker. Best features based in the probability of purchase each article in a given period range and time passed since last purchase.\nMore details on [Giba Solution](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324278).\n\n\n\n### To be continued ###",
      "votes": null
    },
    {
      "id": "1782967",
      "postDate": "05/10/2022 03:13:04",
      "content": "<p>Congrats for another gold medal !</p>",
      "rawMarkdown": "Congrats for another gold medal !",
      "votes": null
    },
    {
      "id": "1782968",
      "postDate": "05/10/2022 03:15:37",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a> <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a>.  Nice write up.</p>",
      "rawMarkdown": "Congrats @chenxin1991 @juzqyxs.  Nice write up.",
      "votes": null
    },
    {
      "id": "1782971",
      "postDate": "05/10/2022 03:20:46",
      "content": "<p>Great respect for all my teammates and thanks.</p>",
      "rawMarkdown": "Great respect for all my teammates and thanks.",
      "votes": null
    },
    {
      "id": "1783013",
      "postDate": "05/10/2022 04:07:09",
      "content": "<p>This is my first team up since 2017. I am deeply impressed by efforts of my teammates. Great finish and looking forward to next game.</p>",
      "rawMarkdown": "This is my first team up since 2017. I am deeply impressed by efforts of my teammates. Great finish and looking forward to next game.",
      "votes": null
    },
    {
      "id": "1783030",
      "postDate": "05/10/2022 04:21:41",
      "content": "<p>Congratulations for finishing in the money. Consistently you have improved your model's performance over the months which is great 👌</p>",
      "rawMarkdown": "Congratulations for finishing in the money. Consistently you have improved your model's performance over the months which is great 👌",
      "votes": null
    },
    {
      "id": "1783214",
      "postDate": "05/10/2022 07:49:52",
      "content": "<p><a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a> wow congratulations!</p>",
      "rawMarkdown": "chenxin1991 wow congratulations!",
      "votes": null
    },
    {
      "id": "1783296",
      "postDate": "05/10/2022 09:11:00",
      "content": "<p>Congrats for this great win and explanation !<br>\nCould you explain how did you use a binary classifier instead of a ranking model ? Like what was the target y representing ?</p>",
      "rawMarkdown": "Congrats for this great win and explanation !\nCould you explain how did you use a binary classifier instead of a ranking model ? Like what was the target y representing ?",
      "votes": null
    },
    {
      "id": "1783302",
      "postDate": "05/10/2022 09:25:11",
      "content": "<p>Because the cv and lb score of binary classifier is much better than ranking model in our case.</p>",
      "rawMarkdown": "Because the cv and lb score of binary classifier is much better than ranking model in our case.",
      "votes": null
    },
    {
      "id": "1783331",
      "postDate": "05/10/2022 09:54:02",
      "content": "<p><a href=\"https://www.kaggle.com/souamesannis\" target=\"_blank\">@souamesannis</a> Target y represents whether the user purchase the item(candidates) or not in the week which need to predict.</p>",
      "rawMarkdown": "souamesannis Target y represents whether the user purchase the item(candidates) or not in the week which need to predict.",
      "votes": null
    },
    {
      "id": "1783358",
      "postDate": "05/10/2022 10:34:18",
      "content": "<p>Congrats ,GM is on the way</p>",
      "rawMarkdown": "Congrats ,GM is on the way",
      "votes": null
    },
    {
      "id": "1783662",
      "postDate": "05/10/2022 15:03:04",
      "content": "<p>Congrats and thanks for sharing! Could you share some experience about the team work, plz？</p>",
      "rawMarkdown": "Congrats and thanks for sharing! Could you share some experience about the team work, plz？",
      "votes": null
    },
    {
      "id": "1783925",
      "postDate": "05/10/2022 19:26:18",
      "content": "<p>Amazing work as always, keep it up!</p>",
      "rawMarkdown": "Amazing work as always, keep it up!",
      "votes": null
    },
    {
      "id": "1784130",
      "postDate": "05/11/2022 00:28:56",
      "content": "<p>Thank for sharing! I've learned a lot of from yours.</p>",
      "rawMarkdown": "Thank for sharing! I've learned a lot of from yours.",
      "votes": null
    },
    {
      "id": "1784582",
      "postDate": "05/11/2022 09:28:55",
      "content": "<p>How do you compute tf-idf efficiently? </p>",
      "rawMarkdown": "How do you compute tf-idf efficiently?",
      "votes": null
    },
    {
      "id": "1784924",
      "postDate": "05/11/2022 15:35:51",
      "content": "<p>just use TfidfVectorizer in sklearn.</p>",
      "rawMarkdown": "just use TfidfVectorizer in sklearn.",
      "votes": null
    },
    {
      "id": "1785395",
      "postDate": "05/12/2022 04:14:36",
      "content": "<p>Use concatenation of user's historically purchased article/product_code/product_type_no as term, and concatenation of article's article/product_code/product_type_no as document?</p>",
      "rawMarkdown": "Use concatenation of user's historically purchased article/product_code/product_type_no as term, and concatenation of article's article/product_code/product_type_no as document?",
      "votes": null
    },
    {
      "id": "1785538",
      "postDate": "05/12/2022 07:17:31",
      "content": "<p>user's historically purchased article/product_code/product_type_no for user tfidf, article who bought by history users for article tfidf.</p>",
      "rawMarkdown": "user's historically purchased article/product_code/product_type_no for user tfidf, article who bought by history users for article tfidf.",
      "votes": null
    },
    {
      "id": "1786793",
      "postDate": "05/13/2022 09:35:35",
      "content": "<p>amazing work and keep it up</p>",
      "rawMarkdown": "amazing work and keep it up",
      "votes": null
    },
    {
      "id": "1786947",
      "postDate": "05/13/2022 13:20:11",
      "content": "<p>could you explain how to use <code>article who bought by history users</code> to generate the article tfidf?</p>",
      "rawMarkdown": "could you explain how to use `article who bought by history users` to generate the article tfidf?",
      "votes": null
    },
    {
      "id": "1786974",
      "postDate": "05/13/2022 13:40:45",
      "content": "<p><a href=\"https://www.kaggle.com/cherrizhu\" target=\"_blank\">@cherrizhu</a> <br>\n<code>docs = transactions.groupby(['customer_id'])['article_id'].apply(lambda x: ' '.join(list(x))).reset_index()</code><br>\n<code>enc = tfidf()</code><br>\n<code>enc.fit(docs)</code><br>\nsomething like this.</p>",
      "rawMarkdown": "cherrizhu \n`docs = transactions.groupby(['customer_id'])['article_id'].apply(lambda x: ' '.join(list(x))).reset_index()`\n`enc = tfidf()`\n`enc.fit(docs)`\nsomething like this.",
      "votes": null
    },
    {
      "id": "1787128",
      "postDate": "05/13/2022 15:37:19",
      "content": "<p>Thanks for your reply!!<br>\nI have understood the use of <code>tfidf</code>,  and what's the effect of <code>svd</code>?<br>\nTo generate the user / article embedding based of <code>tfidf</code> sequence ?</p>",
      "rawMarkdown": "Thanks for your reply!!\nI have understood the use of `tfidf`,  and what's the effect of `svd`?\nTo generate the user / article embedding based of `tfidf` sequence ?",
      "votes": null
    },
    {
      "id": "1787147",
      "postDate": "05/13/2022 15:49:32",
      "content": "<p>yes,just transform tfidf vectors to dense embedding.</p>",
      "rawMarkdown": "yes,just transform tfidf vectors to dense embedding.",
      "votes": null
    },
    {
      "id": "1789581",
      "postDate": "05/14/2022 03:08:16",
      "content": "<p>thank you very much for your reply!!!</p>",
      "rawMarkdown": "thank you very much for your reply!!!",
      "votes": null
    },
    {
      "id": "1789768",
      "postDate": "05/14/2022 07:54:56",
      "content": "<p>Very interesting! Thanks for sharing!</p>",
      "rawMarkdown": "Very interesting! Thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1782967,
      "author_name": "hengzheng",
      "author_url": "",
      "post_date": "05/10/2022 03:13:04",
      "content": "<p>Congrats for another gold medal !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1782968,
      "author_name": "zhoumichael",
      "author_url": "",
      "post_date": "05/10/2022 03:15:37",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a> <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a>.  Nice write up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1782971,
      "author_name": "juzqyxs",
      "author_url": "",
      "post_date": "05/10/2022 03:20:46",
      "content": "<p>Great respect for all my teammates and thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783013,
      "author_name": "hydantess",
      "author_url": "",
      "post_date": "05/10/2022 04:07:09",
      "content": "<p>This is my first team up since 2017. I am deeply impressed by efforts of my teammates. Great finish and looking forward to next game.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783030,
      "author_name": "tarique7",
      "author_url": "",
      "post_date": "05/10/2022 04:21:41",
      "content": "<p>Congratulations for finishing in the money. Consistently you have improved your model's performance over the months which is great 👌</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783214,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/10/2022 07:49:52",
      "content": "<p><a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a> wow congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783296,
      "author_name": "souamesannis",
      "author_url": "",
      "post_date": "05/10/2022 09:11:00",
      "content": "<p>Congrats for this great win and explanation !<br>\nCould you explain how did you use a binary classifier instead of a ranking model ? Like what was the target y representing ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783302,
          "author_name": "juzqyxs",
          "author_url": "",
          "post_date": "05/10/2022 09:25:11",
          "content": "<p>Because the cv and lb score of binary classifier is much better than ranking model in our case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783331,
          "author_name": "chenxin1991",
          "author_url": "",
          "post_date": "05/10/2022 09:54:02",
          "content": "<p><a href=\"https://www.kaggle.com/souamesannis\" target=\"_blank\">@souamesannis</a> Target y represents whether the user purchase the item(candidates) or not in the week which need to predict.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783358,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "05/10/2022 10:34:18",
      "content": "<p>Congrats ,GM is on the way</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783662,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "05/10/2022 15:03:04",
      "content": "<p>Congrats and thanks for sharing! Could you share some experience about the team work, plz？</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783925,
      "author_name": "",
      "author_url": "",
      "post_date": "05/10/2022 19:26:18",
      "content": "<p>Amazing work as always, keep it up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784130,
      "author_name": "arti1117",
      "author_url": "",
      "post_date": "05/11/2022 00:28:56",
      "content": "<p>Thank for sharing! I've learned a lot of from yours.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784582,
      "author_name": "homoalways",
      "author_url": "",
      "post_date": "05/11/2022 09:28:55",
      "content": "<p>How do you compute tf-idf efficiently? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1784924,
          "author_name": "juzqyxs",
          "author_url": "",
          "post_date": "05/11/2022 15:35:51",
          "content": "<p>just use TfidfVectorizer in sklearn.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1785395,
          "author_name": "homoalways",
          "author_url": "",
          "post_date": "05/12/2022 04:14:36",
          "content": "<p>Use concatenation of user's historically purchased article/product_code/product_type_no as term, and concatenation of article's article/product_code/product_type_no as document?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1785538,
          "author_name": "juzqyxs",
          "author_url": "",
          "post_date": "05/12/2022 07:17:31",
          "content": "<p>user's historically purchased article/product_code/product_type_no for user tfidf, article who bought by history users for article tfidf.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786947,
          "author_name": "cherrizhu",
          "author_url": "",
          "post_date": "05/13/2022 13:20:11",
          "content": "<p>could you explain how to use <code>article who bought by history users</code> to generate the article tfidf?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786974,
          "author_name": "chenxin1991",
          "author_url": "",
          "post_date": "05/13/2022 13:40:45",
          "content": "<p><a href=\"https://www.kaggle.com/cherrizhu\" target=\"_blank\">@cherrizhu</a> <br>\n<code>docs = transactions.groupby(['customer_id'])['article_id'].apply(lambda x: ' '.join(list(x))).reset_index()</code><br>\n<code>enc = tfidf()</code><br>\n<code>enc.fit(docs)</code><br>\nsomething like this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1787128,
          "author_name": "cherrizhu",
          "author_url": "",
          "post_date": "05/13/2022 15:37:19",
          "content": "<p>Thanks for your reply!!<br>\nI have understood the use of <code>tfidf</code>,  and what's the effect of <code>svd</code>?<br>\nTo generate the user / article embedding based of <code>tfidf</code> sequence ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1787147,
          "author_name": "juzqyxs",
          "author_url": "",
          "post_date": "05/13/2022 15:49:32",
          "content": "<p>yes,just transform tfidf vectors to dense embedding.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1789581,
          "author_name": "cherrizhu",
          "author_url": "",
          "post_date": "05/14/2022 03:08:16",
          "content": "<p>thank you very much for your reply!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1786793,
      "author_name": "poolpy11",
      "author_url": "",
      "post_date": "05/13/2022 09:35:35",
      "content": "<p>amazing work and keep it up</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1789768,
      "author_name": "rafiaaa",
      "author_url": "",
      "post_date": "05/14/2022 07:54:56",
      "content": "<p>Very interesting! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1782959": "First of all, thanks to competition organizers for hosting this interesting competition and my great teammates(@Giba, @qysx, @hyd). Secondly, thanks to the kaggle community, I have learned a lot from those great kernels and disscussions.\n\nIn this competition, we were asked to recommend products to users based on their historical purchasing behavior. To solve this problem, our solution splits into two parts: recall&rank. Good strategy for generating candidates and feature engineering are both important to win a good place.\n\n### Recall ###\nWe used several methods to generate different candidates, in order to improve the coverage of positive samples(both user and user-item).\n1. u2i: items that the user recently purchased\n2. i2i: item based collaborative filtering\n3. u2tag2i: tag can be \"product_code\", \"product_type_no\", \"department_no\", \"section_no\"\n4. hot: generate hot items for different age_bins's user\n\n### Features ###\nMost important features are as follows,\n1. count: user, item, user-item based, it's important to use all historical data to generate the features of user.\n2. gap: user, item, user-item first/last purchase time to now. For item's first purchased time, it may  represent the \"release time\" of this item. So we can use \"user\" as key to calculate the statistics of \"release time\" , it may help us to get user's preference for new or old items.\n3. items' discount: use max/mean to represent the common price of the item. use the price for each row to calculate the discount of the item.So we can get user's preference for discounted items and the number of days which a item is on sale recently.\n4. tfidf: set user's historical purchased article/product_code/product_type_no to a sentence, then use tfidf+svd to generate feature. \n5. item_sim: collaborative filtering score of i2i, calculate the items between the candidates and user's historical purchased.\n6. set categorical_features(\"department_no\", \"product_type_no\", etc.) for lightgbm will improve cv score about 0.0005~0.0008. \n\n### Model ###\nWe use last 9 weeks for training(last week for validation).Our best single model is catboost with cv score(0.00403) and lb score(0.0341). We use several models to ensemble(cv: 0.0412, lb: 0.0348). And we use the kernels([https://www.kaggle.com/code/titericz/h-m-ensembling-how-to ](url))Giba shared before to handle the different candidates list when blending.\n\n| model | objective| average num of candidates for each user | cv map@12|\n| --- | --- |---|---|\n| catboost | binary | 120 |0.0403 |\n| lightgbm | binary | 120|0.0395 |\n| catboost| binary| 220|0.0402 |\n|lightgbm| binary| 220|0.0396 |\n|lightgbm|lambdarank|220|0.0400 |\n|lightgbm|lambdarank|1000|0.0381|\n\nThe result of our best cv is the best one on PB. Trust your local cv is really importance in kaggle.\n\n### Giba's solution ###\nSet around 10k article_id candidate list based in sales popularity of last 2 weeks. Build a binary dataset with all combinations of customer_id x 10k article_id . Build that huge dataset by batches and using cudf to speedup the entire process. Create a binary trainset from weeks 90 to 104. Train model on week 90-103, validate on week 104. To train the model select random 200, 300 or 500 negative candidates for each customer_id. Blend models trained using LightGBM lambdarank, LightGBM xendcg, XGBoost  lambdarank and CatBoost ranker. Best features based in the probability of purchase each article in a given period range and time passed since last purchase.\nMore details on [Giba Solution](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324278).\n\n\n\n### To be continued ###",
    "1782967": "Congrats for another gold medal !",
    "1782968": "Congrats @chenxin1991 @juzqyxs.  Nice write up.",
    "1782971": "Great respect for all my teammates and thanks.",
    "1783013": "This is my first team up since 2017. I am deeply impressed by efforts of my teammates. Great finish and looking forward to next game.",
    "1783030": "Congratulations for finishing in the money. Consistently you have improved your model's performance over the months which is great 👌",
    "1783214": "chenxin1991 wow congratulations!",
    "1783296": "Congrats for this great win and explanation !\nCould you explain how did you use a binary classifier instead of a ranking model ? Like what was the target y representing ?",
    "1783302": "Because the cv and lb score of binary classifier is much better than ranking model in our case.",
    "1783331": "souamesannis Target y represents whether the user purchase the item(candidates) or not in the week which need to predict.",
    "1783358": "Congrats ,GM is on the way",
    "1783662": "Congrats and thanks for sharing! Could you share some experience about the team work, plz？",
    "1783925": "Amazing work as always, keep it up!",
    "1784130": "Thank for sharing! I've learned a lot of from yours.",
    "1784582": "How do you compute tf-idf efficiently?",
    "1784924": "just use TfidfVectorizer in sklearn.",
    "1785395": "Use concatenation of user's historically purchased article/product_code/product_type_no as term, and concatenation of article's article/product_code/product_type_no as document?",
    "1785538": "user's historically purchased article/product_code/product_type_no for user tfidf, article who bought by history users for article tfidf.",
    "1786793": "amazing work and keep it up",
    "1786947": "could you explain how to use `article who bought by history users` to generate the article tfidf?",
    "1786974": "cherrizhu \n`docs = transactions.groupby(['customer_id'])['article_id'].apply(lambda x: ' '.join(list(x))).reset_index()`\n`enc = tfidf()`\n`enc.fit(docs)`\nsomething like this.",
    "1787128": "Thanks for your reply!!\nI have understood the use of `tfidf`,  and what's the effect of `svd`?\nTo generate the user / article embedding based of `tfidf` sequence ?",
    "1787147": "yes,just transform tfidf vectors to dense embedding.",
    "1789581": "thank you very much for your reply!!!",
    "1789768": "Very interesting! Thanks for sharing!"
  },
  "source": "meta"
}