{
  "id": 324129,
  "title": "3rd place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/324129",
  "author_name": "sirius",
  "post_date": "2022-05-10T08:46:49.397000",
  "votes": 96,
  "comment_count": 40,
  "views": 0,
  "content": "<p>At First, thank kaggle staff and h&amp;m team for organizing this interesting competition. As a new kaggler (but an old RECSYSer), I enjoy this game very much and very very happy to get this place. And, thank all the competitors and the community, from whom I learned very very much, which is a bigger gain for me than the $8000🤓. At last, I must thank the covid-19 that isolated me at home for two months😂, which forced me to focus on the competition in the spare time because of no outting activities, especially during May day holiday.</p>\n<p>As my pipeline is very similar to others' (candidate generation + ranking) and I notice most part of my solution (either the recall methods or the ranking features/models) has been used and described in other guy's threads, I will just briefly introduce something different and useful in my case here. (LB scores in the following are all single model Private scores.)</p>\n<ul>\n<li><strong>Features about recall strategy</strong>:  whether this article is recalled by <code>strategy_name</code> and the rank of this article under the <code>strategy_name</code> (because I've ordered the candidate under each recall strategy by some metric, the rank num can reflact the relevance). These features, plus an expansion of candidates num for each user from dozens to hundred, boost my LB score <strong>from 0.02855 to 0.03262</strong>, which changes the color of medal from silver to near gold. In addition, if I only increase the recall num and don't adding the recall features, the CV score is very very poor. </li>\n<li><strong>BPR matrix factorization</strong>: The similarity of user and article is important for ranking. Features about item2item similarity, such as count of bought together and word2vec, are commonly used in most competitors' model and those also boost my score a lot, but what improves my model most is the user2item similarity obtained from BPR matrix factorization. This BPR model is trained with all the transactions before the target week (I've trained one BPR for each week) using <a href=\"https://implicit.readthedocs.io/en/latest/bpr.html\" target=\"_blank\">implicit</a>. The auc of BPR similarity is ~0.720, while the auc of the whole ranking model is ~0.806 and the best auc of other single feature is ~0.680. At last, this single similarity feature boost my LB score <strong>from 0.03363 to 0.03510</strong>, which takes me from gold medal area to the prize area.</li>\n</ul>\n<p>Any questions are welcome. Thank you for reading!</p>\n<p>-----------------------UPDATE 2022.05.15------------------------------</p>\n<h1>Recalling</h1>\n<ul>\n<li>the most popular items</li>\n<li>items that the user recently purchased</li>\n<li>relative items to what the user recently purchased</li>\n<li>popular items under user's attributes (age, gender and so on)</li>\n<li>items which have the same prod_name with what the user recently purchased</li>\n</ul>\n<h1>Ranking</h1>\n<h3>Samples</h3>\n<ul>\n<li>The generated candidates which the customer didn't purchase in the target week are labeled with 0 and those which the customer purchased are labeled with 1</li>\n<li>Because of big imbalance (pos:neg is about 1:300), downsampling is adopted. I just keep the negative samples with amount of <code>30*len(pos_samples)</code> </li>\n<li>The pos samples are only the ones which are recalled in the candidates and the other items that the user purchased in actual are not included.</li>\n</ul>\n<h3>Features</h3>\n<ul>\n<li>user and item static attributes</li>\n<li>counts of user-item and user-item attributes in the last 7 days, 30 days and all transactions. </li>\n<li>day diff of user-item and user-item attributes last seen in the transactions;</li>\n<li>similarities between user and item:<ul>\n<li>BPR matrix factorization user2item similarity</li>\n<li>word2vec item2item similarity</li>\n<li>jaccard similarity between the attributes of items that user purchased and the attributes of the target item</li>\n<li>txt similarity and image similarity by the embeddings obtained from public pre-trained models</li></ul></li>\n<li>item popularity:<ul>\n<li>counts of item/item attributes purchased  in the last 7 days, 30 days and all transactions;</li>\n<li>time weighted popularity, where the weight is 1/days of now minus purchased date</li>\n<li>weekly trend score from <a href=\"https://www.kaggle.com/code/byfone/h-m-trending-products-weekly\" target=\"_blank\">this kernal</a></li></ul></li>\n<li>features about recall strategy:<ul>\n<li>whether this article is recalled by <code>strategy_name</code> </li>\n<li>the rank of this article under the <code>strategy_name</code></li></ul></li>\n</ul>\n<h3>Models</h3>\n<p>The metric of each model is very similar. The final version is the average score of 16 models which were trained with different hyperparameters (learning rate, max_depth and even random seed).</p>\n<ul>\n<li>lightgbm</li>\n<li>catboost</li>\n<li>xgboost</li>\n</ul>",
  "messages": [
    {
      "id": 1783269,
      "postDate": "2022-05-10T08:46:49.397Z",
      "content": "<p>At First, thank kaggle staff and h&amp;m team for organizing this interesting competition. As a new kaggler (but an old RECSYSer), I enjoy this game very much and very very happy to get this place. And, thank all the competitors and the community, from whom I learned very very much, which is a bigger gain for me than the $8000🤓. At last, I must thank the covid-19 that isolated me at home for two months😂, which forced me to focus on the competition in the spare time because of no outting activities, especially during May day holiday.</p>\n<p>As my pipeline is very similar to others' (candidate generation + ranking) and I notice most part of my solution (either the recall methods or the ranking features/models) has been used and described in other guy's threads, I will just briefly introduce something different and useful in my case here. (LB scores in the following are all single model Private scores.)</p>\n<ul>\n<li><strong>Features about recall strategy</strong>:  whether this article is recalled by <code>strategy_name</code> and the rank of this article under the <code>strategy_name</code> (because I've ordered the candidate under each recall strategy by some metric, the rank num can reflact the relevance). These features, plus an expansion of candidates num for each user from dozens to hundred, boost my LB score <strong>from 0.02855 to 0.03262</strong>, which changes the color of medal from silver to near gold. In addition, if I only increase the recall num and don't adding the recall features, the CV score is very very poor. </li>\n<li><strong>BPR matrix factorization</strong>: The similarity of user and article is important for ranking. Features about item2item similarity, such as count of bought together and word2vec, are commonly used in most competitors' model and those also boost my score a lot, but what improves my model most is the user2item similarity obtained from BPR matrix factorization. This BPR model is trained with all the transactions before the target week (I've trained one BPR for each week) using <a href=\"https://implicit.readthedocs.io/en/latest/bpr.html\" target=\"_blank\">implicit</a>. The auc of BPR similarity is ~0.720, while the auc of the whole ranking model is ~0.806 and the best auc of other single feature is ~0.680. At last, this single similarity feature boost my LB score <strong>from 0.03363 to 0.03510</strong>, which takes me from gold medal area to the prize area.</li>\n</ul>\n<p>Any questions are welcome. Thank you for reading!</p>\n<p>-----------------------UPDATE 2022.05.15------------------------------</p>\n<h1>Recalling</h1>\n<ul>\n<li>the most popular items</li>\n<li>items that the user recently purchased</li>\n<li>relative items to what the user recently purchased</li>\n<li>popular items under user's attributes (age, gender and so on)</li>\n<li>items which have the same prod_name with what the user recently purchased</li>\n</ul>\n<h1>Ranking</h1>\n<h3>Samples</h3>\n<ul>\n<li>The generated candidates which the customer didn't purchase in the target week are labeled with 0 and those which the customer purchased are labeled with 1</li>\n<li>Because of big imbalance (pos:neg is about 1:300), downsampling is adopted. I just keep the negative samples with amount of <code>30*len(pos_samples)</code> </li>\n<li>The pos samples are only the ones which are recalled in the candidates and the other items that the user purchased in actual are not included.</li>\n</ul>\n<h3>Features</h3>\n<ul>\n<li>user and item static attributes</li>\n<li>counts of user-item and user-item attributes in the last 7 days, 30 days and all transactions. </li>\n<li>day diff of user-item and user-item attributes last seen in the transactions;</li>\n<li>similarities between user and item:<ul>\n<li>BPR matrix factorization user2item similarity</li>\n<li>word2vec item2item similarity</li>\n<li>jaccard similarity between the attributes of items that user purchased and the attributes of the target item</li>\n<li>txt similarity and image similarity by the embeddings obtained from public pre-trained models</li></ul></li>\n<li>item popularity:<ul>\n<li>counts of item/item attributes purchased  in the last 7 days, 30 days and all transactions;</li>\n<li>time weighted popularity, where the weight is 1/days of now minus purchased date</li>\n<li>weekly trend score from <a href=\"https://www.kaggle.com/code/byfone/h-m-trending-products-weekly\" target=\"_blank\">this kernal</a></li></ul></li>\n<li>features about recall strategy:<ul>\n<li>whether this article is recalled by <code>strategy_name</code> </li>\n<li>the rank of this article under the <code>strategy_name</code></li></ul></li>\n</ul>\n<h3>Models</h3>\n<p>The metric of each model is very similar. The final version is the average score of 16 models which were trained with different hyperparameters (learning rate, max_depth and even random seed).</p>\n<ul>\n<li>lightgbm</li>\n<li>catboost</li>\n<li>xgboost</li>\n</ul>",
      "rawMarkdown": "At First, thank kaggle staff and h&m team for organizing this interesting competition. As a new kaggler (but an old RECSYSer), I enjoy this game very much and very very happy to get this place. And, thank all the competitors and the community, from whom I learned very very much, which is a bigger gain for me than the $8000🤓. At last, I must thank the covid-19 that isolated me at home for two months😂, which forced me to focus on the competition in the spare time because of no outting activities, especially during May day holiday.\n\nAs my pipeline is very similar to others' (candidate generation + ranking) and I notice most part of my solution (either the recall methods or the ranking features/models) has been used and described in other guy's threads, I will just briefly introduce something different and useful in my case here. (LB scores in the following are all single model Private scores.)\n- **Features about recall strategy**:  whether this article is recalled by `strategy_name` and the rank of this article under the `strategy_name` (because I've ordered the candidate under each recall strategy by some metric, the rank num can reflact the relevance). These features, plus an expansion of candidates num for each user from dozens to hundred, boost my LB score **from 0.02855 to 0.03262**, which changes the color of medal from silver to near gold. In addition, if I only increase the recall num and don't adding the recall features, the CV score is very very poor. \n- **BPR matrix factorization**: The similarity of user and article is important for ranking. Features about item2item similarity, such as count of bought together and word2vec, are commonly used in most competitors' model and those also boost my score a lot, but what improves my model most is the user2item similarity obtained from BPR matrix factorization. This BPR model is trained with all the transactions before the target week (I've trained one BPR for each week) using [implicit](https://implicit.readthedocs.io/en/latest/bpr.html). The auc of BPR similarity is ~0.720, while the auc of the whole ranking model is ~0.806 and the best auc of other single feature is ~0.680. At last, this single similarity feature boost my LB score **from 0.03363 to 0.03510**, which takes me from gold medal area to the prize area.\n\nAny questions are welcome. Thank you for reading!\n\n-----------------------UPDATE 2022.05.15------------------------------\n# Recalling\n\n- the most popular items\n- items that the user recently purchased\n- relative items to what the user recently purchased\n- popular items under user's attributes (age, gender and so on)\n- items which have the same prod_name with what the user recently purchased\n\n# Ranking\n### Samples\n- The generated candidates which the customer didn't purchase in the target week are labeled with 0 and those which the customer purchased are labeled with 1\n- Because of big imbalance (pos:neg is about 1:300), downsampling is adopted. I just keep the negative samples with amount of `30*len(pos_samples)` \n- The pos samples are only the ones which are recalled in the candidates and the other items that the user purchased in actual are not included.\n### Features\n\n- user and item static attributes\n- counts of user-item and user-item attributes in the last 7 days, 30 days and all transactions. \n-  day diff of user-item and user-item attributes last seen in the transactions;\n- similarities between user and item:\n    - BPR matrix factorization user2item similarity\n    - word2vec item2item similarity\n    - jaccard similarity between the attributes of items that user purchased and the attributes of the target item\n    - txt similarity and image similarity by the embeddings obtained from public pre-trained models\n- item popularity:\n    - counts of item/item attributes purchased  in the last 7 days, 30 days and all transactions;\n    - time weighted popularity, where the weight is 1/days of now minus purchased date\n    - weekly trend score from [this kernal](https://www.kaggle.com/code/byfone/h-m-trending-products-weekly)\n- features about recall strategy:\n    - whether this article is recalled by `strategy_name` \n    - the rank of this article under the `strategy_name`\n### Models\nThe metric of each model is very similar. The final version is the average score of 16 models which were trained with different hyperparameters (learning rate, max_depth and even random seed).\n- lightgbm\n- catboost\n- xgboost\n\n\n\n\n\n\n\n",
      "votes": 96
    },
    {
      "id": 1783714,
      "postDate": "2022-05-10T16:06:24.883Z",
      "content": "<p><strong>Features about recall strategy</strong>:</p>\n<p>So if I understand correctly, you have 2 features for <em>each</em> strategy you used, one binary feature for whether the candidate was retrieved using that strategy, and another for the \"ranking\" of that candidate within that strategy?</p>\n<p>Is the \"ranking\" some measure relevant to the strategy (i.e. recent sales in popularity retrieval), or is it a simple ranking (i.e. #11 of 30 candidates retrieved for that customer using that strategy)?</p>",
      "rawMarkdown": "**Features about recall strategy**:\n\nSo if I understand correctly, you have 2 features for *each* strategy you used, one binary feature for whether the candidate was retrieved using that strategy, and another for the \"ranking\" of that candidate within that strategy?\n\nIs the \"ranking\" some measure relevant to the strategy (i.e. recent sales in popularity retrieval), or is it a simple ranking (i.e. #11 of 30 candidates retrieved for that customer using that strategy)?",
      "votes": 3,
      "replies": [
        {
          "id": 1783733,
          "postDate": "2022-05-10T16:15:40.073Z",
          "content": "<ol>\n<li>Yes</li>\n<li>Both the ranking features you described was used in my final model. But If only talking about the boosting from 0.02855 to 0.0326, it's the latter one（ranking num） contributing most. In that submission, features about relevance measuring was not complete, and CV score was poor before adding the features about ranking num.</li>\n</ol>",
          "rawMarkdown": "1. Yes\n2. Both the ranking features you described was used in my final model. But If only talking about the boosting from 0.02855 to 0.0326, it's the latter one（ranking num） contributing most. In that submission, features about relevance measuring was not complete, and CV score was poor before adding the features about ranking num.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1783288,
      "postDate": "2022-05-10T09:05:19.170Z",
      "content": "<p>Congrats on the solo gold <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> </p>\n<p>Thanks for sharing the secret sauce on BPR similarity.</p>",
      "rawMarkdown": "Congrats on the solo gold @sirius81 \n\nThanks for sharing the secret sauce on BPR similarity.",
      "votes": 3
    },
    {
      "id": 1783621,
      "postDate": "2022-05-10T14:28:23.273Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> for top3 and solo prize/gold! Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats @sirius81 for top3 and solo prize/gold! Thanks for sharing your solution.",
      "votes": 4
    },
    {
      "id": 1783529,
      "postDate": "2022-05-10T13:03:22.920Z",
      "content": "<p>Congrats and nice work with BPR. I also used BPR from implicit package but only to generate limited candidates and not to score user-item pairs. I had some ideas to factorize user and item state but not to just use embeddings trained with BPR. This is very nice and clean idea.</p>",
      "rawMarkdown": "Congrats and nice work with BPR. I also used BPR from implicit package but only to generate limited candidates and not to score user-item pairs. I had some ideas to factorize user and item state but not to just use embeddings trained with BPR. This is very nice and clean idea.",
      "votes": 4,
      "replies": [
        {
          "id": 1783612,
          "postDate": "2022-05-10T14:21:06.230Z",
          "content": "<p>Congrats on your 2nd place. Thank you for this comment and lots of valuable and impressive discussion during the competition!</p>",
          "rawMarkdown": "Congrats on your 2nd place. Thank you for this comment and lots of valuable and impressive discussion during the competition!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1786451,
      "postDate": "2022-05-12T22:48:01.013Z",
      "content": "<p>Congrats on the solo win, <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a>!!! And thank you for a wonderful write-up focusing on the key points! 🔥</p>\n<p>Could I please ask you how did you frame the competition and what did you do for duplicates?</p>\n<p>I went from all purchases to baskets (purchases per customer per week) and I removed duplicates. Not sure if that was the best approach? How did you handle this? Do you ever predict duplicated purchases?</p>\n<p>Also, if I could please ask you, how did you train your model? Do you train your model on some number of weeks, use the last week for validation and then predict on the test set? Or do you use the last week for validation to find hyperparams but then retrain the model on the full set of data (including the last week) to make predictions for submission?</p>\n<p>Thank you so much for all your help! 🙂🙏</p>",
      "rawMarkdown": "Congrats on the solo win, @sirius81!!! And thank you for a wonderful write-up focusing on the key points! 🔥\n\nCould I please ask you how did you frame the competition and what did you do for duplicates?\n\nI went from all purchases to baskets (purchases per customer per week) and I removed duplicates. Not sure if that was the best approach? How did you handle this? Do you ever predict duplicated purchases?\n\nAlso, if I could please ask you, how did you train your model? Do you train your model on some number of weeks, use the last week for validation and then predict on the test set? Or do you use the last week for validation to find hyperparams but then retrain the model on the full set of data (including the last week) to make predictions for submission?\n\nThank you so much for all your help! 🙂🙏",
      "votes": 1,
      "replies": [
        {
          "id": 1790970,
          "postDate": "2022-05-15T14:14:12.777Z",
          "content": "<p>Thank you!</p>\n<ol>\n<li>No, there aren't duplicated (customer_id, article_id) in my prediction. A customer may purchase the duplicated article in the next week (such as three pairs of sock), but it's really hard to predict the amount of duplication. </li>\n<li>The latter one.</li>\n</ol>",
          "rawMarkdown": "Thank you!\n1. No, there aren't duplicated (customer_id, article_id) in my prediction. A customer may purchase the duplicated article in the next week (such as three pairs of sock), but it's really hard to predict the amount of duplication. \n2. The latter one.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1783429,
      "postDate": "2022-05-10T11:54:34.983Z",
      "content": "<p>Congrats! btw, I am curious about some details about your solution:</p>\n<ol>\n<li>how many weeks(or days) during the ranking phrase?</li>\n<li>what'r your ranking model?</li>\n</ol>",
      "rawMarkdown": "Congrats! btw, I am curious about some details about your solution:\n1. how many weeks(or days) during the ranking phrase?\n2. what'r your ranking model?",
      "votes": 1,
      "replies": [
        {
          "id": 1783467,
          "postDate": "2022-05-10T12:20:02.720Z",
          "content": "<p>Thanks!</p>\n<ol>\n<li>18 weeks</li>\n<li>Xgboost, Catboost and Lightgbm give similar CV scores, so the final submission is a mix of these models.</li>\n</ol>",
          "rawMarkdown": "Thanks!\n1. 18 weeks\n2. Xgboost, Catboost and Lightgbm give similar CV scores, so the final submission is a mix of these models.",
          "votes": 1
        },
        {
          "id": 1783483,
          "postDate": "2022-05-10T12:29:57.503Z",
          "content": "<p>thanks for your sharing. btw, for the every week in the 18weeks, we need run all the methods to create candidate? so all the methods are run by 18+1(valid)+1(test)=20times?</p>",
          "rawMarkdown": "thanks for your sharing. btw, for the every week in the 18weeks, we need run all the methods to create candidate? so all the methods are run by 18+1(valid)+1(test)=20times?"
        },
        {
          "id": 1783497,
          "postDate": "2022-05-10T12:40:04.157Z",
          "content": "<p>Yes. Totally right!</p>",
          "rawMarkdown": "Yes. Totally right!",
          "votes": 1
        },
        {
          "id": 1784747,
          "postDate": "2022-05-11T12:36:56.890Z",
          "content": "<p>thanks for your reply, for my information, how the performance improve by enlarging the training week? I just use one week for train, and I get 0.0277 on lb. </p>",
          "rawMarkdown": "thanks for your reply, for my information, how the performance improve by enlarging the training week? I just use one week for train, and I get 0.0277 on lb. "
        },
        {
          "id": 1784864,
          "postDate": "2022-05-11T14:43:13.700Z",
          "content": "<p>That depends. More data can reduce overfitting. In my case, 18weeks vs 2weeks is about 0.0005-0.0010 up</p>",
          "rawMarkdown": "That depends. More data can reduce overfitting. In my case, 18weeks vs 2weeks is about 0.0005-0.0010 up",
          "votes": 1
        }
      ]
    },
    {
      "id": 1783871,
      "postDate": "2022-05-10T18:42:29.930Z",
      "content": "<p>This is an excellent solution! Well done on achieving a top 3 finish.</p>\n<p>It's great to see that you are using features from the recall strategy to improve your ranking model. This is something that I haven't seen before and it is definitely something that could be helpful for other competitors.</p>\n<p>The use of BPR matrix factorization is also interesting and I can see how it would improve the performance of the model. Overall, this is a very impressive solution and well deserving of a top 3 finish.</p>",
      "rawMarkdown": "This is an excellent solution! Well done on achieving a top 3 finish.\n\nIt's great to see that you are using features from the recall strategy to improve your ranking model. This is something that I haven't seen before and it is definitely something that could be helpful for other competitors.\n\nThe use of BPR matrix factorization is also interesting and I can see how it would improve the performance of the model. Overall, this is a very impressive solution and well deserving of a top 3 finish.",
      "votes": 2
    },
    {
      "id": 1783362,
      "postDate": "2022-05-10T10:35:56.443Z",
      "content": "<p>Congrats! your user2item similarity is very strong.</p>",
      "rawMarkdown": "Congrats! your user2item similarity is very strong.",
      "votes": 2
    },
    {
      "id": 1783283,
      "postDate": "2022-05-10T08:59:07.970Z",
      "content": "<p>Congrats! Thanks for sharing. BPR is really interesting.</p>",
      "rawMarkdown": "Congrats! Thanks for sharing. BPR is really interesting.",
      "votes": 2
    },
    {
      "id": 2404477,
      "postDate": "2023-08-23T09:21:50.400Z",
      "content": "<p>congrats on the placing!  A truly fascinating way to solve this problem.</p>",
      "rawMarkdown": "congrats on the placing!  A truly fascinating way to solve this problem."
    },
    {
      "id": 2369006,
      "postDate": "2023-08-01T13:45:30.283Z",
      "content": "<p>I know it's been a year since your last update, but could you explain how did you generate negative samples for ranker training?</p>",
      "rawMarkdown": "I know it's been a year since your last update, but could you explain how did you generate negative samples for ranker training?"
    },
    {
      "id": 2034571,
      "postDate": "2022-11-18T08:47:00.960Z",
      "content": "<p>Great Job! Congrats!</p>\n<p>I am curious about how to calculate \"word2vec item2item similarity\" as a feature ?  Between a candidate artical and an user.  Look forward your feedback.</p>",
      "rawMarkdown": "Great Job! Congrats!\n\nI am curious about how to calculate \"word2vec item2item similarity\" as a feature ?  Between a candidate artical and an user.  Look forward your feedback.",
      "replies": [
        {
          "id": 2036212,
          "postDate": "2022-11-19T14:56:27.243Z",
          "content": "<p>It's about the similarity of the candidate artical and the user-purchased articals.</p>\n<p>For example, if a user has purchased [a1, a2, a3] in the history, \"word2vec item2item similarity\" is the mean/sum/max of [cosine_sim(a1, candidate), cosine_sim(a2, candidate), cosine_sim(a3, candidate)].</p>",
          "rawMarkdown": "It's about the similarity of the candidate artical and the user-purchased articals.\n\nFor example, if a user has purchased [a1, a2, a3] in the history, \"word2vec item2item similarity\" is the mean/sum/max of [cosine_sim(a1, candidate), cosine_sim(a2, candidate), cosine_sim(a3, candidate)].",
          "votes": 1
        }
      ]
    },
    {
      "id": 1790755,
      "postDate": "2022-05-15T09:42:47.723Z",
      "content": "<p>By the way, are you from ecnu? I graduated from ecnu last year and I'm now at renmin university.</p>",
      "rawMarkdown": "By the way, are you from ecnu? I graduated from ecnu last year and I'm now at renmin university.",
      "replies": [
        {
          "id": 1790972,
          "postDate": "2022-05-15T14:17:03.607Z",
          "content": "<p>Yes! Got the Master degree 4 years ago in ECNU🤝</p>",
          "rawMarkdown": "Yes! Got the Master degree 4 years ago in ECNU🤝"
        },
        {
          "id": 1791405,
          "postDate": "2022-05-16T02:03:38.507Z",
          "content": "<p>🤝Learn from you.</p>",
          "rawMarkdown": "🤝Learn from you."
        }
      ]
    },
    {
      "id": 1790751,
      "postDate": "2022-05-15T09:40:58.330Z",
      "content": "<p>Simple ideas that work well, great job! A question: what machine do you use(how much memory), and do you use cpu or gpu to process features?</p>",
      "rawMarkdown": "Simple ideas that work well, great job! A question: what machine do you use(how much memory), and do you use cpu or gpu to process features?",
      "replies": [
        {
          "id": 1790971,
          "postDate": "2022-05-15T14:16:02.837Z",
          "content": "<p>Thanks!</p>\n<ol>\n<li>120GB</li>\n<li>Feature processing is only on cpu and model training is on gpu</li>\n</ol>",
          "rawMarkdown": "Thanks!\n1. 120GB\n2. Feature processing is only on cpu and model training is on gpu",
          "votes": 2
        }
      ]
    },
    {
      "id": 1786498,
      "postDate": "2022-05-13T00:50:04.630Z",
      "content": "<p>Congrats! Thanks for sharing your solution</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution\n"
    },
    {
      "id": 1784542,
      "postDate": "2022-05-11T08:35:17.470Z",
      "content": "<p>Congrats for your solo gold! It's very impressive work.<br>\nI have some questions for your great work in more detail.</p>\n<ol>\n<li><p>I think you used various candidate generation strategies including MFBPR model. I also considered about to use weekly MFBPR candidates that are trained with the data before target week. So you considered all interactions right before the target weeks and trained for 18 weeks right? And do you mean item2user similarity means dot product of the trained user / item embedding vectors?</p></li>\n<li><p>What was the different candidate generation strategies? I want to know about the exact strategies what you used for last submission.</p></li>\n<li><p>In the ranking step, the rank feature of the candidates per strategy was important and is there any other useful features? I want to know about the exact input setting for LightGBM, including negative sampling for labels or something.</p></li>\n</ol>\n<p>Again, thanks for your sharing.</p>",
      "rawMarkdown": "Congrats for your solo gold! It's very impressive work.\nI have some questions for your great work in more detail.\n\n1. I think you used various candidate generation strategies including MFBPR model. I also considered about to use weekly MFBPR candidates that are trained with the data before target week. So you considered all interactions right before the target weeks and trained for 18 weeks right? And do you mean item2user similarity means dot product of the trained user / item embedding vectors?\n\n2. What was the different candidate generation strategies? I want to know about the exact strategies what you used for last submission.\n\n3. In the ranking step, the rank feature of the candidates per strategy was important and is there any other useful features? I want to know about the exact input setting for LightGBM, including negative sampling for labels or something.\n\nAgain, thanks for your sharing.",
      "replies": [
        {
          "id": 1784851,
          "postDate": "2022-05-11T14:38:08.883Z",
          "content": "<p>Thanks！<br>\n1.Yes. But I haven't generate candidates using bpr, just being a feature in the ranking model (didn't have more spare time and felt enough high recall rate ~18％)</p>\n<p>2&amp;3. Busy with work now. May be updated this weekend</p>",
          "rawMarkdown": "Thanks！\n1.Yes. But I haven't generate candidates using bpr, just being a feature in the ranking model (didn't have more spare time and felt enough high recall rate ~18％)\n\n2&3. Busy with work now. May be updated this weekend",
          "votes": 1
        },
        {
          "id": 1790969,
          "postDate": "2022-05-15T14:12:10.043Z",
          "content": "<p>updated Q2&amp;Q3  <a href=\"https://www.kaggle.com/hyungeunjo\" target=\"_blank\">@hyungeunjo</a> </p>",
          "rawMarkdown": "updated Q2&Q3  @hyungeunjo ",
          "votes": 1
        },
        {
          "id": 1791410,
          "postDate": "2022-05-16T02:10:03.593Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> ! great update, I was looking forward to it. It's very helpful👍</p>\n<ol>\n<li><p>So, in ranking part, did you set the target label in binary setting?(etc, if that candidate actually interacted with target date, it becomes 1 and 0 otherwise). I'm very appreciate if you explain more detail about the label setting(I'm the first time to use LightGBM things).</p></li>\n<li><p>And I'm not clear with the Ranking part - Features 2, 3.. If you don't mind, could you add some detail about calculation?</p></li>\n</ol>\n<p>Thanks !</p>",
          "rawMarkdown": "Thanks @sirius81 ! great update, I was looking forward to it. It's very helpful👍\n\n1. So, in ranking part, did you set the target label in binary setting?(etc, if that candidate actually interacted with target date, it becomes 1 and 0 otherwise). I'm very appreciate if you explain more detail about the label setting(I'm the first time to use LightGBM things).\n\n2. And I'm not clear with the Ranking part - Features 2, 3.. If you don't mind, could you add some detail about calculation?\n\nThanks !"
        },
        {
          "id": 1791481,
          "postDate": "2022-05-16T03:46:42.570Z",
          "content": "<ol>\n<li>Updated <code>Ranking-Samples</code></li>\n<li>Feature 2 is like <code>transaction_7days.groupby(['customer_id', 'article_id']).size()</code>,  meaning the count of the user purchasing the target item in the last 7 days; <br>\nFeature 3 is like <code>transaction.groupby(['customer_id', 'article_id']).min('days')</code>, meaning how many days are since the user purchased the target article  in last time.</li>\n</ol>",
          "rawMarkdown": "1. Updated `Ranking-Samples`\n2. Feature 2 is like `transaction_7days.groupby(['customer_id', 'article_id']).size()`,  meaning the count of the user purchasing the target item in the last 7 days; \nFeature 3 is like `transaction.groupby(['customer_id', 'article_id']).min('days')`, meaning how many days are since the user purchased the target article  in last time.",
          "votes": 1
        },
        {
          "id": 1791600,
          "postDate": "2022-05-16T07:01:24.153Z",
          "content": "<p>Thanks for your reply!</p>",
          "rawMarkdown": "Thanks for your reply!"
        },
        {
          "id": 1800851,
          "postDate": "2022-05-25T09:02:51.180Z",
          "content": "<p>Hello, <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> !<br>\nCould you give me some details about feature of the different recall strategies? If the candidates are came from strategies : 1. weekly popular items 2. previously purchased item, what metric did you used for that strategies to rank them?</p>",
          "rawMarkdown": "Hello, @sirius81 !\nCould you give me some details about feature of the different recall strategies? If the candidates are came from strategies : 1. weekly popular items 2. previously purchased item, what metric did you used for that strategies to rank them?"
        },
        {
          "id": 1803744,
          "postDate": "2022-05-28T05:53:45.880Z",
          "content": "<p>Both used the score described in <a href=\"https://www.kaggle.com/code/hervind/h-m-faster-trending-products-weekly\" target=\"_blank\">this kernal</a>. I only adjusted some parameters a very little. <a href=\"https://www.kaggle.com/hyungeunjo\" target=\"_blank\">@hyungeunjo</a> </p>",
          "rawMarkdown": "Both used the score described in [this kernal](https://www.kaggle.com/code/hervind/h-m-faster-trending-products-weekly). I only adjusted some parameters a very little. @hyungeunjo ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1783526,
      "postDate": "2022-05-10T12:59:40.630Z",
      "content": "<p>Thanks for the readup! Congratz on solo gold!<br>\nCan you explain this part, please?</p>\n<blockquote>\n  <p>These features, plus an expansion of candidates num for each user from dozens to hundred, </p>\n</blockquote>\n<p>What do you mean by expansion of candidates num for each user from dozens to hundred?</p>",
      "rawMarkdown": "Thanks for the readup! Congratz on solo gold!\nCan you explain this part, please?\n> These features, plus an expansion of candidates num for each user from dozens to hundred, \n\nWhat do you mean by expansion of candidates num for each user from dozens to hundred?",
      "replies": [
        {
          "id": 1783599,
          "postDate": "2022-05-10T14:05:21.820Z",
          "content": "<p>Thank you. My pipeline consists of two part: candidate generation and ranking. At the beginning, I generated dozens of candidate articles for each user. Then, for a higher recall rate, I tried to increase the recall num of each candidate strategy, such as popular items 30-&gt;100, itemcf 20-&gt;50 and so on. As a result, the total count of candidates for each user increased to hundreds and the recall rates increased from ~10% to ~18%.</p>",
          "rawMarkdown": "Thank you. My pipeline consists of two part: candidate generation and ranking. At the beginning, I generated dozens of candidate articles for each user. Then, for a higher recall rate, I tried to increase the recall num of each candidate strategy, such as popular items 30->100, itemcf 20->50 and so on. As a result, the total count of candidates for each user increased to hundreds and the recall rates increased from ~10% to ~18%.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1786731,
      "postDate": "2022-05-13T07:59:51.427Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1785312,
      "postDate": "2022-05-12T02:40:03.500Z",
      "content": "<p>Thank you for sharing strategy!</p>",
      "rawMarkdown": "Thank you for sharing strategy!"
    },
    {
      "id": 1785293,
      "postDate": "2022-05-12T02:07:15.407Z",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !"
    }
  ],
  "comments": [
    {
      "id": 1783714,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-05-10T16:06:24.883000",
      "content": "<p><strong>Features about recall strategy</strong>:</p>\n<p>So if I understand correctly, you have 2 features for <em>each</em> strategy you used, one binary feature for whether the candidate was retrieved using that strategy, and another for the \"ranking\" of that candidate within that strategy?</p>\n<p>Is the \"ranking\" some measure relevant to the strategy (i.e. recent sales in popularity retrieval), or is it a simple ranking (i.e. #11 of 30 candidates retrieved for that customer using that strategy)?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1783733,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-10T16:15:40.073000",
          "content": "<ol>\n<li>Yes</li>\n<li>Both the ranking features you described was used in my final model. But If only talking about the boosting from 0.02855 to 0.0326, it's the latter one（ranking num） contributing most. In that submission, features about relevance measuring was not complete, and CV score was poor before adding the features about ranking num.</li>\n</ol>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1783288,
      "author_name": "SRK",
      "author_url": "",
      "post_date": "2022-05-10T09:05:19.170000",
      "content": "<p>Congrats on the solo gold <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> </p>\n<p>Thanks for sharing the secret sauce on BPR similarity.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1783621,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2022-05-10T14:28:23.273000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> for top3 and solo prize/gold! Thanks for sharing your solution.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1783529,
      "author_name": "Paweł Jankiewicz",
      "author_url": "",
      "post_date": "2022-05-10T13:03:22.920000",
      "content": "<p>Congrats and nice work with BPR. I also used BPR from implicit package but only to generate limited candidates and not to score user-item pairs. I had some ideas to factorize user and item state but not to just use embeddings trained with BPR. This is very nice and clean idea.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1783612,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-10T14:21:06.230000",
          "content": "<p>Congrats on your 2nd place. Thank you for this comment and lots of valuable and impressive discussion during the competition!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1786451,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-05-12T22:48:01.013000",
      "content": "<p>Congrats on the solo win, <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a>!!! And thank you for a wonderful write-up focusing on the key points! 🔥</p>\n<p>Could I please ask you how did you frame the competition and what did you do for duplicates?</p>\n<p>I went from all purchases to baskets (purchases per customer per week) and I removed duplicates. Not sure if that was the best approach? How did you handle this? Do you ever predict duplicated purchases?</p>\n<p>Also, if I could please ask you, how did you train your model? Do you train your model on some number of weeks, use the last week for validation and then predict on the test set? Or do you use the last week for validation to find hyperparams but then retrain the model on the full set of data (including the last week) to make predictions for submission?</p>\n<p>Thank you so much for all your help! 🙂🙏</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1790970,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-15T14:14:12.777000",
          "content": "<p>Thank you!</p>\n<ol>\n<li>No, there aren't duplicated (customer_id, article_id) in my prediction. A customer may purchase the duplicated article in the next week (such as three pairs of sock), but it's really hard to predict the amount of duplication. </li>\n<li>The latter one.</li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1783429,
      "author_name": "biubiuG",
      "author_url": "",
      "post_date": "2022-05-10T11:54:34.983000",
      "content": "<p>Congrats! btw, I am curious about some details about your solution:</p>\n<ol>\n<li>how many weeks(or days) during the ranking phrase?</li>\n<li>what'r your ranking model?</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1783467,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-10T12:20:02.720000",
          "content": "<p>Thanks!</p>\n<ol>\n<li>18 weeks</li>\n<li>Xgboost, Catboost and Lightgbm give similar CV scores, so the final submission is a mix of these models.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1783483,
          "author_name": "biubiuG",
          "author_url": "",
          "post_date": "2022-05-10T12:29:57.503000",
          "content": "<p>thanks for your sharing. btw, for the every week in the 18weeks, we need run all the methods to create candidate? so all the methods are run by 18+1(valid)+1(test)=20times?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1783497,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-10T12:40:04.157000",
          "content": "<p>Yes. Totally right!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1784747,
          "author_name": "biubiuG",
          "author_url": "",
          "post_date": "2022-05-11T12:36:56.890000",
          "content": "<p>thanks for your reply, for my information, how the performance improve by enlarging the training week? I just use one week for train, and I get 0.0277 on lb. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1784864,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-11T14:43:13.700000",
          "content": "<p>That depends. More data can reduce overfitting. In my case, 18weeks vs 2weeks is about 0.0005-0.0010 up</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1783871,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-10T18:42:29.930000",
      "content": "<p>This is an excellent solution! Well done on achieving a top 3 finish.</p>\n<p>It's great to see that you are using features from the recall strategy to improve your ranking model. This is something that I haven't seen before and it is definitely something that could be helpful for other competitors.</p>\n<p>The use of BPR matrix factorization is also interesting and I can see how it would improve the performance of the model. Overall, this is a very impressive solution and well deserving of a top 3 finish.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1783362,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "2022-05-10T10:35:56.443000",
      "content": "<p>Congrats! your user2item similarity is very strong.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1783283,
      "author_name": "Ethan",
      "author_url": "",
      "post_date": "2022-05-10T08:59:07.970000",
      "content": "<p>Congrats! Thanks for sharing. BPR is really interesting.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2404477,
      "author_name": "Omar Abdulbagi",
      "author_url": "",
      "post_date": "2023-08-23T09:21:50.400000",
      "content": "<p>congrats on the placing!  A truly fascinating way to solve this problem.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2369006,
      "author_name": "Bartosz Pietrzak",
      "author_url": "",
      "post_date": "2023-08-01T13:45:30.283000",
      "content": "<p>I know it's been a year since your last update, but could you explain how did you generate negative samples for ranker training?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2034571,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2022-11-18T08:47:00.960000",
      "content": "<p>Great Job! Congrats!</p>\n<p>I am curious about how to calculate \"word2vec item2item similarity\" as a feature ?  Between a candidate artical and an user.  Look forward your feedback.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2036212,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-11-19T14:56:27.243000",
          "content": "<p>It's about the similarity of the candidate artical and the user-purchased articals.</p>\n<p>For example, if a user has purchased [a1, a2, a3] in the history, \"word2vec item2item similarity\" is the mean/sum/max of [cosine_sim(a1, candidate), cosine_sim(a2, candidate), cosine_sim(a3, candidate)].</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1790755,
      "author_name": "Homoalways",
      "author_url": "",
      "post_date": "2022-05-15T09:42:47.723000",
      "content": "<p>By the way, are you from ecnu? I graduated from ecnu last year and I'm now at renmin university.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1790972,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-15T14:17:03.607000",
          "content": "<p>Yes! Got the Master degree 4 years ago in ECNU🤝</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1791405,
          "author_name": "Homoalways",
          "author_url": "",
          "post_date": "2022-05-16T02:03:38.507000",
          "content": "<p>🤝Learn from you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1790751,
      "author_name": "Homoalways",
      "author_url": "",
      "post_date": "2022-05-15T09:40:58.330000",
      "content": "<p>Simple ideas that work well, great job! A question: what machine do you use(how much memory), and do you use cpu or gpu to process features?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1790971,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-15T14:16:02.837000",
          "content": "<p>Thanks!</p>\n<ol>\n<li>120GB</li>\n<li>Feature processing is only on cpu and model training is on gpu</li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1786498,
      "author_name": "Raj_B_IK",
      "author_url": "",
      "post_date": "2022-05-13T00:50:04.630000",
      "content": "<p>Congrats! Thanks for sharing your solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1784542,
      "author_name": "Hyungeun Jo",
      "author_url": "",
      "post_date": "2022-05-11T08:35:17.470000",
      "content": "<p>Congrats for your solo gold! It's very impressive work.<br>\nI have some questions for your great work in more detail.</p>\n<ol>\n<li><p>I think you used various candidate generation strategies including MFBPR model. I also considered about to use weekly MFBPR candidates that are trained with the data before target week. So you considered all interactions right before the target weeks and trained for 18 weeks right? And do you mean item2user similarity means dot product of the trained user / item embedding vectors?</p></li>\n<li><p>What was the different candidate generation strategies? I want to know about the exact strategies what you used for last submission.</p></li>\n<li><p>In the ranking step, the rank feature of the candidates per strategy was important and is there any other useful features? I want to know about the exact input setting for LightGBM, including negative sampling for labels or something.</p></li>\n</ol>\n<p>Again, thanks for your sharing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1784851,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-11T14:38:08.883000",
          "content": "<p>Thanks！<br>\n1.Yes. But I haven't generate candidates using bpr, just being a feature in the ranking model (didn't have more spare time and felt enough high recall rate ~18％)</p>\n<p>2&amp;3. Busy with work now. May be updated this weekend</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1790969,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-15T14:12:10.043000",
          "content": "<p>updated Q2&amp;Q3  <a href=\"https://www.kaggle.com/hyungeunjo\" target=\"_blank\">@hyungeunjo</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1791410,
          "author_name": "Hyungeun Jo",
          "author_url": "",
          "post_date": "2022-05-16T02:10:03.593000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> ! great update, I was looking forward to it. It's very helpful👍</p>\n<ol>\n<li><p>So, in ranking part, did you set the target label in binary setting?(etc, if that candidate actually interacted with target date, it becomes 1 and 0 otherwise). I'm very appreciate if you explain more detail about the label setting(I'm the first time to use LightGBM things).</p></li>\n<li><p>And I'm not clear with the Ranking part - Features 2, 3.. If you don't mind, could you add some detail about calculation?</p></li>\n</ol>\n<p>Thanks !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1791481,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-16T03:46:42.570000",
          "content": "<ol>\n<li>Updated <code>Ranking-Samples</code></li>\n<li>Feature 2 is like <code>transaction_7days.groupby(['customer_id', 'article_id']).size()</code>,  meaning the count of the user purchasing the target item in the last 7 days; <br>\nFeature 3 is like <code>transaction.groupby(['customer_id', 'article_id']).min('days')</code>, meaning how many days are since the user purchased the target article  in last time.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1791600,
          "author_name": "Hyungeun Jo",
          "author_url": "",
          "post_date": "2022-05-16T07:01:24.153000",
          "content": "<p>Thanks for your reply!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1800851,
          "author_name": "Hyungeun Jo",
          "author_url": "",
          "post_date": "2022-05-25T09:02:51.180000",
          "content": "<p>Hello, <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> !<br>\nCould you give me some details about feature of the different recall strategies? If the candidates are came from strategies : 1. weekly popular items 2. previously purchased item, what metric did you used for that strategies to rank them?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1803744,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-28T05:53:45.880000",
          "content": "<p>Both used the score described in <a href=\"https://www.kaggle.com/code/hervind/h-m-faster-trending-products-weekly\" target=\"_blank\">this kernal</a>. I only adjusted some parameters a very little. <a href=\"https://www.kaggle.com/hyungeunjo\" target=\"_blank\">@hyungeunjo</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1783526,
      "author_name": "Igor Kuivjogi Fernandes",
      "author_url": "",
      "post_date": "2022-05-10T12:59:40.630000",
      "content": "<p>Thanks for the readup! Congratz on solo gold!<br>\nCan you explain this part, please?</p>\n<blockquote>\n  <p>These features, plus an expansion of candidates num for each user from dozens to hundred, </p>\n</blockquote>\n<p>What do you mean by expansion of candidates num for each user from dozens to hundred?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1783599,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2022-05-10T14:05:21.820000",
          "content": "<p>Thank you. My pipeline consists of two part: candidate generation and ranking. At the beginning, I generated dozens of candidate articles for each user. Then, for a higher recall rate, I tried to increase the recall num of each candidate strategy, such as popular items 30-&gt;100, itemcf 20-&gt;50 and so on. As a result, the total count of candidates for each user increased to hundreds and the recall rates increased from ~10% to ~18%.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1786731,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-13T07:59:51.427000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1785312,
      "author_name": "Making TARS",
      "author_url": "",
      "post_date": "2022-05-12T02:40:03.500000",
      "content": "<p>Thank you for sharing strategy!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1785293,
      "author_name": "Dark Wang",
      "author_url": "",
      "post_date": "2022-05-12T02:07:15.407000",
      "content": "<p>Thanks for sharing !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1783269": "At First, thank kaggle staff and h&m team for organizing this interesting competition. As a new kaggler (but an old RECSYSer), I enjoy this game very much and very very happy to get this place. And, thank all the competitors and the community, from whom I learned very very much, which is a bigger gain for me than the $8000🤓. At last, I must thank the covid-19 that isolated me at home for two months😂, which forced me to focus on the competition in the spare time because of no outting activities, especially during May day holiday.\n\nAs my pipeline is very similar to others' (candidate generation + ranking) and I notice most part of my solution (either the recall methods or the ranking features/models) has been used and described in other guy's threads, I will just briefly introduce something different and useful in my case here. (LB scores in the following are all single model Private scores.)\n- **Features about recall strategy**:  whether this article is recalled by `strategy_name` and the rank of this article under the `strategy_name` (because I've ordered the candidate under each recall strategy by some metric, the rank num can reflact the relevance). These features, plus an expansion of candidates num for each user from dozens to hundred, boost my LB score **from 0.02855 to 0.03262**, which changes the color of medal from silver to near gold. In addition, if I only increase the recall num and don't adding the recall features, the CV score is very very poor. \n- **BPR matrix factorization**: The similarity of user and article is important for ranking. Features about item2item similarity, such as count of bought together and word2vec, are commonly used in most competitors' model and those also boost my score a lot, but what improves my model most is the user2item similarity obtained from BPR matrix factorization. This BPR model is trained with all the transactions before the target week (I've trained one BPR for each week) using [implicit](https://implicit.readthedocs.io/en/latest/bpr.html). The auc of BPR similarity is ~0.720, while the auc of the whole ranking model is ~0.806 and the best auc of other single feature is ~0.680. At last, this single similarity feature boost my LB score **from 0.03363 to 0.03510**, which takes me from gold medal area to the prize area.\n\nAny questions are welcome. Thank you for reading!\n\n-----------------------UPDATE 2022.05.15------------------------------\n# Recalling\n\n- the most popular items\n- items that the user recently purchased\n- relative items to what the user recently purchased\n- popular items under user's attributes (age, gender and so on)\n- items which have the same prod_name with what the user recently purchased\n\n# Ranking\n### Samples\n- The generated candidates which the customer didn't purchase in the target week are labeled with 0 and those which the customer purchased are labeled with 1\n- Because of big imbalance (pos:neg is about 1:300), downsampling is adopted. I just keep the negative samples with amount of `30*len(pos_samples)` \n- The pos samples are only the ones which are recalled in the candidates and the other items that the user purchased in actual are not included.\n### Features\n\n- user and item static attributes\n- counts of user-item and user-item attributes in the last 7 days, 30 days and all transactions. \n-  day diff of user-item and user-item attributes last seen in the transactions;\n- similarities between user and item:\n    - BPR matrix factorization user2item similarity\n    - word2vec item2item similarity\n    - jaccard similarity between the attributes of items that user purchased and the attributes of the target item\n    - txt similarity and image similarity by the embeddings obtained from public pre-trained models\n- item popularity:\n    - counts of item/item attributes purchased  in the last 7 days, 30 days and all transactions;\n    - time weighted popularity, where the weight is 1/days of now minus purchased date\n    - weekly trend score from [this kernal](https://www.kaggle.com/code/byfone/h-m-trending-products-weekly)\n- features about recall strategy:\n    - whether this article is recalled by `strategy_name` \n    - the rank of this article under the `strategy_name`\n### Models\nThe metric of each model is very similar. The final version is the average score of 16 models which were trained with different hyperparameters (learning rate, max_depth and even random seed).\n- lightgbm\n- catboost\n- xgboost\n\n\n\n\n\n\n\n",
    "1783714": "**Features about recall strategy**:\n\nSo if I understand correctly, you have 2 features for *each* strategy you used, one binary feature for whether the candidate was retrieved using that strategy, and another for the \"ranking\" of that candidate within that strategy?\n\nIs the \"ranking\" some measure relevant to the strategy (i.e. recent sales in popularity retrieval), or is it a simple ranking (i.e. #11 of 30 candidates retrieved for that customer using that strategy)?",
    "1783288": "Congrats on the solo gold @sirius81 \n\nThanks for sharing the secret sauce on BPR similarity.",
    "1783621": "Congrats @sirius81 for top3 and solo prize/gold! Thanks for sharing your solution.",
    "1783529": "Congrats and nice work with BPR. I also used BPR from implicit package but only to generate limited candidates and not to score user-item pairs. I had some ideas to factorize user and item state but not to just use embeddings trained with BPR. This is very nice and clean idea.",
    "1786451": "Congrats on the solo win, @sirius81!!! And thank you for a wonderful write-up focusing on the key points! 🔥\n\nCould I please ask you how did you frame the competition and what did you do for duplicates?\n\nI went from all purchases to baskets (purchases per customer per week) and I removed duplicates. Not sure if that was the best approach? How did you handle this? Do you ever predict duplicated purchases?\n\nAlso, if I could please ask you, how did you train your model? Do you train your model on some number of weeks, use the last week for validation and then predict on the test set? Or do you use the last week for validation to find hyperparams but then retrain the model on the full set of data (including the last week) to make predictions for submission?\n\nThank you so much for all your help! 🙂🙏",
    "1783429": "Congrats! btw, I am curious about some details about your solution:\n1. how many weeks(or days) during the ranking phrase?\n2. what'r your ranking model?",
    "1783871": "This is an excellent solution! Well done on achieving a top 3 finish.\n\nIt's great to see that you are using features from the recall strategy to improve your ranking model. This is something that I haven't seen before and it is definitely something that could be helpful for other competitors.\n\nThe use of BPR matrix factorization is also interesting and I can see how it would improve the performance of the model. Overall, this is a very impressive solution and well deserving of a top 3 finish.",
    "1783362": "Congrats! your user2item similarity is very strong.",
    "1783283": "Congrats! Thanks for sharing. BPR is really interesting.",
    "2404477": "congrats on the placing!  A truly fascinating way to solve this problem.",
    "2369006": "I know it's been a year since your last update, but could you explain how did you generate negative samples for ranker training?",
    "2034571": "Great Job! Congrats!\n\nI am curious about how to calculate \"word2vec item2item similarity\" as a feature ?  Between a candidate artical and an user.  Look forward your feedback.",
    "1790755": "By the way, are you from ecnu? I graduated from ecnu last year and I'm now at renmin university.",
    "1790751": "Simple ideas that work well, great job! A question: what machine do you use(how much memory), and do you use cpu or gpu to process features?",
    "1786498": "Congrats! Thanks for sharing your solution\n",
    "1784542": "Congrats for your solo gold! It's very impressive work.\nI have some questions for your great work in more detail.\n\n1. I think you used various candidate generation strategies including MFBPR model. I also considered about to use weekly MFBPR candidates that are trained with the data before target week. So you considered all interactions right before the target weeks and trained for 18 weeks right? And do you mean item2user similarity means dot product of the trained user / item embedding vectors?\n\n2. What was the different candidate generation strategies? I want to know about the exact strategies what you used for last submission.\n\n3. In the ranking step, the rank feature of the candidates per strategy was important and is there any other useful features? I want to know about the exact input setting for LightGBM, including negative sampling for labels or something.\n\nAgain, thanks for your sharing.",
    "1783526": "Thanks for the readup! Congratz on solo gold!\nCan you explain this part, please?\n> These features, plus an expansion of candidates num for each user from dozens to hundred, \n\nWhat do you mean by expansion of candidates num for each user from dozens to hundred?",
    "1786731": "",
    "1785312": "Thank you for sharing strategy!",
    "1785293": "Thanks for sharing !"
  }
}