{
  "id": 324098,
  "title": "5th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/hao-5th-place-solution",
  "author_name": "",
  "post_date": "2022-05-10T06:51:21.963Z",
  "votes": 64,
  "comment_count": 36,
  "views": 0,
  "content": "<p>First of all, thanks to the competition organizers to host such an interesting competition and my great teammates <a href=\"https://www.kaggle.com/rendongltt\" target=\"_blank\">@rendongltt</a> (who is a huge fan of online-shopping, which gives us a lot of help😂) and <a href=\"https://www.kaggle.com/jiaqizhang35\" target=\"_blank\">@jiaqizhang35</a> (who is my friend, without whom I didn't even join this competition). And also I would like to thank <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>, without his great posts, we will not reach this far. And also you (and also a lot of other nice guys, who love sharing) motivated me to do more sharing with the community.</p>\n<p><strong>Overview</strong><br>\nOverall we love this competition so much because it did provide infinite ways to solve this problem especially the possible recall methods. Talking to feature engineering, even though we're quite confident about our skills, but we may have no shot to get a gold medal compared to other grandmasters. So we focused ourselves on the recall methods from the very start. And our recall methods is deeply dependent on article embeddings (image, txt, word2vec, tfidf) and the amazing python package faiss to find the similar articles and customers. As to the feature engineering, we didn't spend too much time on it but using different days window to create aggreated stat features of articles and customers seperatly and customer-article cross aggreated stat features and different cosine similarity feaures using embeddings aforementioned. </p>\n<p><strong>Recall methods</strong><br>\nThe recall methods comprised of those parts:</p>\n<ol>\n<li>Customer's last buckets(and same product code to customers' last bucket's articles), recent bought articles</li>\n<li>U2I: User based collaborative filtering. </li>\n<li>I2i: Item based collaborative filtering</li>\n<li>word2vec: to predict customer's next possible purchase articles after last buckets</li>\n<li>Articles bought together. I modified this nice <a href=\"https://www.kaggle.com/code/titericz/article-id-pairs-in-3s-using-cudf\" target=\"_blank\">notebook</a> from <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a></li>\n<li>Popular articles with different customers' attributes like age bin, gender etc.<br>\nSo overall we use 21 recall methods to create article candidates.</li>\n</ol>\n<p><strong>Embedding methods</strong><br>\nWe created article embdding using different data and methods</p>\n<ol>\n<li>For article images, we use swin transformer to extract the embedding features.</li>\n<li>For the article texts, first we concatnate all text values of each article and then I use SentenceTransformer to extract the embeddings</li>\n<li>Using this great <a href=\"https://www.kaggle.com/code/aerdem4/h-m-rapids-article2vec\" target=\"_blank\">notebook</a> from  <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> you can easily get article embeddings using tfidf method</li>\n<li>The last but also the most useful method to get article embeddings are word2vec method, for which we use gensim to ease the process.</li>\n</ol>\n<p><strong>Recall method calibration</strong><br>\nBecause different recall methods have different positive ratio, so it's better to use some method to calibrate all the recall methods. I use the method which I shared <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672#:~:text=I%20would%20sugguest,goes%20like%20this.\" target=\"_blank\">here</a> to do the preprocess.</p>\n<p><strong>Feature engineering</strong><br>\nNothing special about this part but I found some method which will hugely accelerate the computing time for feature computing. <br>\nBasicly we divide the features into 3 parts:</p>\n<ol>\n<li>articles features, which will be constant to all customers once you fixed the overvation time</li>\n<li>similarly, customer features will be contant to all candidate articles once you fixed the overvation date</li>\n<li>customer-article features, like cosine similarity between each pair of customer and article, customers' bought history aggreated features of this particular article.<br>\nFor part 1 and 2, you could do the computation and save them once and merged with the features of part 3 later. This method really save me a lot of time especially during inference time.</li>\n</ol>\n<p><strong>Modeling</strong> <br>\nBecause we spend most of the time to devolop recall strategies, we use only lightgbm ranker model for training and don't even have time to try other methods like catboost and xgboost. So basically our score is single model's score.</p>\n<p><strong>Remaining question</strong><br>\nWe have one question that we still couldn't solve it right now, which is about the word2vec model. We use all transaction data to train the word2vec model, which actually leaked the future info a little bit. <br>\nImagne that for observation date 20200915 you compute the customer-article cosine similarity using the word2vec model trained using all transaction data. The word2vec embedings of two article ids aid_1, aid_2, which maybe mainly appeared duing week 20200916-20200922, are used on the observation date 20200915. <br>\nThe effect is everytime I use those embeddings to create candidates or compute features, the local CV will improve a lot while the LB improves only a little bit. For my final ranking model, my map@12 score for last week is 0.0441, but the LB is only 0.0350, the gap is much larger than the ones mentioned by others. But if we don't use this feature, both CV and LB score are dropping. So we don't really know how to do it to make the most use of this method. What do you think?</p>",
  "messages": [
    {
      "id": "1783085",
      "postDate": "05/10/2022 05:48:06",
      "content": "<p>First of all, thanks to the competition organizers to host such an interesting competition and my great teammates <a href=\"https://www.kaggle.com/rendongltt\" target=\"_blank\">@rendongltt</a> (who is a huge fan of online-shopping, which gives us a lot of help😂) and <a href=\"https://www.kaggle.com/jiaqizhang35\" target=\"_blank\">@jiaqizhang35</a> (who is my friend, without whom I didn't even join this competition). And also I would like to thank <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>, without his great posts, we will not reach this far. And also you (and also a lot of other nice guys, who love sharing) motivated me to do more sharing with the community.</p>\n<p><strong>Overview</strong><br>\nOverall we love this competition so much because it did provide infinite ways to solve this problem especially the possible recall methods. Talking to feature engineering, even though we're quite confident about our skills, but we may have no shot to get a gold medal compared to other grandmasters. So we focused ourselves on the recall methods from the very start. And our recall methods is deeply dependent on article embeddings (image, txt, word2vec, tfidf) and the amazing python package faiss to find the similar articles and customers. As to the feature engineering, we didn't spend too much time on it but using different days window to create aggreated stat features of articles and customers seperatly and customer-article cross aggreated stat features and different cosine similarity feaures using embeddings aforementioned. </p>\n<p><strong>Recall methods</strong><br>\nThe recall methods comprised of those parts:</p>\n<ol>\n<li>Customer's last buckets(and same product code to customers' last bucket's articles), recent bought articles</li>\n<li>U2I: User based collaborative filtering. </li>\n<li>I2i: Item based collaborative filtering</li>\n<li>word2vec: to predict customer's next possible purchase articles after last buckets</li>\n<li>Articles bought together. I modified this nice <a href=\"https://www.kaggle.com/code/titericz/article-id-pairs-in-3s-using-cudf\" target=\"_blank\">notebook</a> from <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a></li>\n<li>Popular articles with different customers' attributes like age bin, gender etc.<br>\nSo overall we use 21 recall methods to create article candidates.</li>\n</ol>\n<p><strong>Embedding methods</strong><br>\nWe created article embdding using different data and methods</p>\n<ol>\n<li>For article images, we use swin transformer to extract the embedding features.</li>\n<li>For the article texts, first we concatnate all text values of each article and then I use SentenceTransformer to extract the embeddings</li>\n<li>Using this great <a href=\"https://www.kaggle.com/code/aerdem4/h-m-rapids-article2vec\" target=\"_blank\">notebook</a> from  <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> you can easily get article embeddings using tfidf method</li>\n<li>The last but also the most useful method to get article embeddings are word2vec method, for which we use gensim to ease the process.</li>\n</ol>\n<p><strong>Recall method calibration</strong><br>\nBecause different recall methods have different positive ratio, so it's better to use some method to calibrate all the recall methods. I use the method which I shared <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672#:~:text=I%20would%20sugguest,goes%20like%20this.\" target=\"_blank\">here</a> to do the preprocess.</p>\n<p><strong>Feature engineering</strong><br>\nNothing special about this part but I found some method which will hugely accelerate the computing time for feature computing. <br>\nBasicly we divide the features into 3 parts:</p>\n<ol>\n<li>articles features, which will be constant to all customers once you fixed the overvation time</li>\n<li>similarly, customer features will be contant to all candidate articles once you fixed the overvation date</li>\n<li>customer-article features, like cosine similarity between each pair of customer and article, customers' bought history aggreated features of this particular article.<br>\nFor part 1 and 2, you could do the computation and save them once and merged with the features of part 3 later. This method really save me a lot of time especially during inference time.</li>\n</ol>\n<p><strong>Modeling</strong> <br>\nBecause we spend most of the time to devolop recall strategies, we use only lightgbm ranker model for training and don't even have time to try other methods like catboost and xgboost. So basically our score is single model's score.</p>\n<p><strong>Remaining question</strong><br>\nWe have one question that we still couldn't solve it right now, which is about the word2vec model. We use all transaction data to train the word2vec model, which actually leaked the future info a little bit. <br>\nImagne that for observation date 20200915 you compute the customer-article cosine similarity using the word2vec model trained using all transaction data. The word2vec embedings of two article ids aid_1, aid_2, which maybe mainly appeared duing week 20200916-20200922, are used on the observation date 20200915. <br>\nThe effect is everytime I use those embeddings to create candidates or compute features, the local CV will improve a lot while the LB improves only a little bit. For my final ranking model, my map@12 score for last week is 0.0441, but the LB is only 0.0350, the gap is much larger than the ones mentioned by others. But if we don't use this feature, both CV and LB score are dropping. So we don't really know how to do it to make the most use of this method. What do you think?</p>",
      "rawMarkdown": "First of all, thanks to the competition organizers to host such an interesting competition and my great teammates @rendongltt (who is a huge fan of online-shopping, which gives us a lot of help😂) and @jiaqizhang35 (who is my friend, without whom I didn't even join this competition). And also I would like to thank @paweljankiewicz, without his great posts, we will not reach this far. And also you (and also a lot of other nice guys, who love sharing) motivated me to do more sharing with the community.\n\n**Overview**\nOverall we love this competition so much because it did provide infinite ways to solve this problem especially the possible recall methods. Talking to feature engineering, even though we're quite confident about our skills, but we may have no shot to get a gold medal compared to other grandmasters. So we focused ourselves on the recall methods from the very start. And our recall methods is deeply dependent on article embeddings (image, txt, word2vec, tfidf) and the amazing python package faiss to find the similar articles and customers. As to the feature engineering, we didn't spend too much time on it but using different days window to create aggreated stat features of articles and customers seperatly and customer-article cross aggreated stat features and different cosine similarity feaures using embeddings aforementioned. \n\n**Recall methods**\nThe recall methods comprised of those parts:\n1. Customer's last buckets(and same product code to customers' last bucket's articles), recent bought articles\n2. U2I: User based collaborative filtering. \n3. I2i: Item based collaborative filtering\n4. word2vec: to predict customer's next possible purchase articles after last buckets\n5. Articles bought together. I modified this nice [notebook](https://www.kaggle.com/code/titericz/article-id-pairs-in-3s-using-cudf) from @titericz\n6. Popular articles with different customers' attributes like age bin, gender etc.\nSo overall we use 21 recall methods to create article candidates.\n\n**Embedding methods**\nWe created article embdding using different data and methods\n1. For article images, we use swin transformer to extract the embedding features.\n2. For the article texts, first we concatnate all text values of each article and then I use SentenceTransformer to extract the embeddings\n3. Using this great [notebook](https://www.kaggle.com/code/aerdem4/h-m-rapids-article2vec) from  @aerdem4 you can easily get article embeddings using tfidf method\n4. The last but also the most useful method to get article embeddings are word2vec method, for which we use gensim to ease the process.\n\n**Recall method calibration**\nBecause different recall methods have different positive ratio, so it's better to use some method to calibrate all the recall methods. I use the method which I shared [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672#:~:text=I%20would%20sugguest,goes%20like%20this.) to do the preprocess.\n\n**Feature engineering**\nNothing special about this part but I found some method which will hugely accelerate the computing time for feature computing. \nBasicly we divide the features into 3 parts:\n1. articles features, which will be constant to all customers once you fixed the overvation time\n2. similarly, customer features will be contant to all candidate articles once you fixed the overvation date\n3. customer-article features, like cosine similarity between each pair of customer and article, customers' bought history aggreated features of this particular article.\nFor part 1 and 2, you could do the computation and save them once and merged with the features of part 3 later. This method really save me a lot of time especially during inference time.\n\n**Modeling** \nBecause we spend most of the time to devolop recall strategies, we use only lightgbm ranker model for training and don't even have time to try other methods like catboost and xgboost. So basically our score is single model's score.\n\n**Remaining question**\nWe have one question that we still couldn't solve it right now, which is about the word2vec model. We use all transaction data to train the word2vec model, which actually leaked the future info a little bit. \nImagne that for observation date 20200915 you compute the customer-article cosine similarity using the word2vec model trained using all transaction data. The word2vec embedings of two article ids aid_1, aid_2, which maybe mainly appeared duing week 20200916-20200922, are used on the observation date 20200915. \nThe effect is everytime I use those embeddings to create candidates or compute features, the local CV will improve a lot while the LB improves only a little bit. For my final ranking model, my map@12 score for last week is 0.0441, but the LB is only 0.0350, the gap is much larger than the ones mentioned by others. But if we don't use this feature, both CV and LB score are dropping. So we don't really know how to do it to make the most use of this method. What do you think?",
      "votes": null
    },
    {
      "id": "1783095",
      "postDate": "05/10/2022 06:00:31",
      "content": "<p>Congratulations!<br>\nI learned a lot from you during the competition!<br>\nThank you!</p>",
      "rawMarkdown": "Congratulations!\nI learned a lot from you during the competition!\nThank you!",
      "votes": null
    },
    {
      "id": "1783098",
      "postDate": "05/10/2022 06:02:17",
      "content": "<p>You are welcome! You also did well. Congrats!</p>",
      "rawMarkdown": "You are welcome! You also did well. Congrats!",
      "votes": null
    },
    {
      "id": "1783109",
      "postDate": "05/10/2022 06:14:08",
      "content": "<p>Actually, use all transaction data to train word2vec embedding is leaked, but the relationship of different items may be represented,.Like the competition hosted by Wechat in 2021.</p>",
      "rawMarkdown": "Actually, use all transaction data to train word2vec embedding is leaked, but the relationship of different items may be represented,.Like the competition hosted by Wechat in 2021.",
      "votes": null
    },
    {
      "id": "1783119",
      "postDate": "05/10/2022 06:19:40",
      "content": "<p>Yes, I just wonder how other teams deal with this problem to make the CV-LB gap much smaller than mine.</p>",
      "rawMarkdown": "Yes, I just wonder how other teams deal with this problem to make the CV-LB gap much smaller than mine.",
      "votes": null
    },
    {
      "id": "1783121",
      "postDate": "05/10/2022 06:22:35",
      "content": "<p>We do not use all transaction data, only train word2vec in history transactions, so the cv score not boosted.</p>",
      "rawMarkdown": "We do not use all transaction data, only train word2vec in history transactions, so the cv score not boosted.",
      "votes": null
    },
    {
      "id": "1783126",
      "postDate": "05/10/2022 06:25:00",
      "content": "<p>Got it. Thanks! Then you must train a lot of word2vec models.</p>",
      "rawMarkdown": "Got it. Thanks! Then you must train a lot of word2vec models.",
      "votes": null
    },
    {
      "id": "1783131",
      "postDate": "05/10/2022 06:30:55",
      "content": "<p>Because many word2vec models were trained, the distribution may not in same space, that could be the reason why cv not boosted I guess.</p>",
      "rawMarkdown": "Because many word2vec models were trained, the distribution may not in same space, that could be the reason why cv not boosted I guess.",
      "votes": null
    },
    {
      "id": "1783132",
      "postDate": "05/10/2022 06:32:26",
      "content": "<p>Congratulations! Thank you for your sharing. I have learned a lot from you during the competition!</p>",
      "rawMarkdown": "Congratulations! Thank you for your sharing. I have learned a lot from you during the competition!",
      "votes": null
    },
    {
      "id": "1783133",
      "postDate": "05/10/2022 06:34:18",
      "content": "<p>Glad to hear that. And you also did well! Congrats!</p>",
      "rawMarkdown": "Glad to hear that. And you also did well! Congrats!",
      "votes": null
    },
    {
      "id": "1783135",
      "postDate": "05/10/2022 06:37:20",
      "content": "<p>Without your comments we couldn't have built our model.<br>\nThank you for your contributions to the many discussions.</p>",
      "rawMarkdown": "Without your comments we couldn't have built our model.\nThank you for your contributions to the many discussions.",
      "votes": null
    },
    {
      "id": "1783144",
      "postDate": "05/10/2022 06:54:33",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nCongratulations for your gold medal!!<br>\nAnd thank you for your answering my question in discussion during the competition, which improved my place a lot!<br>\nCould you tell me what is “cosine similarity between each pair of customer and article” in details?</p>",
      "rawMarkdown": "lihaorocky \nCongratulations for your gold medal!!\nAnd thank you for your answering my question in discussion during the competition, which improved my place a lot!\nCould you tell me what is “cosine similarity between each pair of customer and article” in details?",
      "votes": null
    },
    {
      "id": "1783159",
      "postDate": "05/10/2022 07:04:51",
      "content": "<p>Glad to hear that!<br>\nSo you have a lot of candidate articles for each customer. And after you get the embedding of each article, you could compute the customer's embedding represention based on the customer's transaction history(aggregate the articles' embeddings to compute the mean/median embedding). Then you have the embedding of the customer and the embedding of the candidate article, you could calculate the vector cosine distance (also cosine similarity) of the customer and the article using dot(vec1, vec2) / (norm(vec1) * norm(vec2)), which represent how close between the customer and the article.</p>",
      "rawMarkdown": "Glad to hear that!\nSo you have a lot of candidate articles for each customer. And after you get the embedding of each article, you could compute the customer's embedding represention based on the customer's transaction history(aggregate the articles' embeddings to compute the mean/median embedding). Then you have the embedding of the customer and the embedding of the candidate article, you could calculate the vector cosine distance (also cosine similarity) of the customer and the article using dot(vec1, vec2) / (norm(vec1) * norm(vec2)), which represent how close between the customer and the article.",
      "votes": null
    },
    {
      "id": "1783229",
      "postDate": "05/10/2022 08:03:02",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Congratulations!!</p>",
      "rawMarkdown": "lihaorocky Congratulations!!",
      "votes": null
    },
    {
      "id": "1783385",
      "postDate": "05/10/2022 11:02:34",
      "content": "<p>Congrats!If you don't use leakage features,should get higher accuracy.</p>",
      "rawMarkdown": "Congrats!If you don't use leakage features,should get higher accuracy.",
      "votes": null
    },
    {
      "id": "1783470",
      "postDate": "05/10/2022 12:22:41",
      "content": "<p>Congratulations!!! very impressive for your method to calibrate！</p>",
      "rawMarkdown": "Congratulations!!! very impressive for your method to calibrate！",
      "votes": null
    },
    {
      "id": "1783671",
      "postDate": "05/10/2022 15:12:48",
      "content": "<p>Congrats and thanks for sharing! <br>\nAbout the word2vec and other pretrain embedding methods, I just train one model for each target week using the transactions before the week, and the result seems no leak.</p>",
      "rawMarkdown": "Congrats and thanks for sharing! \nAbout the word2vec and other pretrain embedding methods, I just train one model for each target week using the transactions before the week, and the result seems no leak.",
      "votes": null
    },
    {
      "id": "1783684",
      "postDate": "05/10/2022 15:31:18",
      "content": "<p>Thanks for sharing your method. Guess I was too lazy to even think about training more than one word2vec model to avoid the leak. And also congrats to your solo 3rd place finish. Your score improving progress during last weeks are quite impressive!</p>",
      "rawMarkdown": "Thanks for sharing your method. Guess I was too lazy to even think about training more than one word2vec model to avoid the leak. And also congrats to your solo 3rd place finish. Your score improving progress during last weeks are quite impressive!",
      "votes": null
    },
    {
      "id": "1783703",
      "postDate": "05/10/2022 15:52:55",
      "content": "<p>Saying the last week progressing, I must thank the covid-19 that isolated me at home for two months😂, which forced me to focus on the competition in the spare time, especially during May day holiday. Just says every coin has two sides.</p>",
      "rawMarkdown": "Saying the last week progressing, I must thank the covid-19 that isolated me at home for two months😂, which forced me to focus on the competition in the spare time, especially during May day holiday. Just says every coin has two sides.",
      "votes": null
    },
    {
      "id": "1783768",
      "postDate": "05/10/2022 16:48:01",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> - I was wondering what you meant when you shared during the competition that Customer/Article similarity features helped a lot.</p>\n<p>I thought of using mean of customer's article embeddings, but I thought that it's probably meaningless - once you average together different articles' embeddings, you lose the information.</p>\n<p>Guess I was wrong :)</p>",
      "rawMarkdown": "lihaorocky - I was wondering what you meant when you shared during the competition that Customer/Article similarity features helped a lot.\n\nI thought of using mean of customer's article embeddings, but I thought that it's probably meaningless - once you average together different articles' embeddings, you lose the information.\n\nGuess I was wrong :)",
      "votes": null
    },
    {
      "id": "1783770",
      "postDate": "05/10/2022 16:49:19",
      "content": "<p>what text are you referring to when you reference \"word2vec\"?</p>",
      "rawMarkdown": "what text are you referring to when you reference \"word2vec\"?",
      "votes": null
    },
    {
      "id": "1783772",
      "postDate": "05/10/2022 16:51:51",
      "content": "<p>Is that also what you used for user-based collaborative filtering?</p>",
      "rawMarkdown": "Is that also what you used for user-based collaborative filtering?",
      "votes": null
    },
    {
      "id": "1784128",
      "postDate": "05/11/2022 00:27:10",
      "content": "<p>For word2vec, I mean using customers' time-ordered purchase article sequence to train a word2vec model getting the embeddings of articles.</p>",
      "rawMarkdown": "For word2vec, I mean using customers' time-ordered purchase article sequence to train a word2vec model getting the embeddings of articles.",
      "votes": null
    },
    {
      "id": "1784139",
      "postDate": "05/11/2022 00:42:38",
      "content": "<p>Yes, those features help a lot. And I don't think it's meaningless. We all know every customer has his/her own preference for styles(articles). And if our representation embeddings of articles are not bad, they have the ability to distinguish the styles. For example, if our embedding is 1-dimensional vector, a set of articles' embedding have values like [0.89], [0.90], [0.902]… Calculating average of those embeddings, you get a value around [0.9]. For another style set of articles, if one customer frequently bought a lot, after aggregating the embeddings of those articles, you get the mean embedding vector [0.2]. Well now we know where the separating capability of this method comes from.<br>\nIt may lose some info, but the noise info is also filtered and we get the main taste of the customer.</p>",
      "rawMarkdown": "Yes, those features help a lot. And I don't think it's meaningless. We all know every customer has his/her own preference for styles(articles). And if our representation embeddings of articles are not bad, they have the ability to distinguish the styles. For example, if our embedding is 1-dimensional vector, a set of articles' embedding have values like [0.89], [0.90], [0.902]... Calculating average of those embeddings, you get a value around [0.9]. For another style set of articles, if one customer frequently bought a lot, after aggregating the embeddings of those articles, you get the mean embedding vector [0.2]. Well now we know where the separating capability of this method comes from.\nIt may lose some info, but the noise info is also filtered and we get the main taste of the customer.",
      "votes": null
    },
    {
      "id": "1784142",
      "postDate": "05/11/2022 00:54:38",
      "content": "<p>How \"lucky\" you are… 😂 Anyway you did really well and all the best staying at home. </p>",
      "rawMarkdown": "How \"lucky\" you are... 😂 Anyway you did really well and all the best staying at home.",
      "votes": null
    },
    {
      "id": "1784257",
      "postDate": "05/11/2022 04:13:22",
      "content": "<p>Congratulations!<br>\nIf you use all the data to build word2vec, that's definitely data leak. Other possible sources of data leakage are in building collaborative filtering. Did you use all the data to build U2I and I2I? </p>",
      "rawMarkdown": "Congratulations!\nIf you use all the data to build word2vec, that's definitely data leak. Other possible sources of data leakage are in building collaborative filtering. Did you use all the data to build U2I and I2I?",
      "votes": null
    },
    {
      "id": "1784268",
      "postDate": "05/11/2022 04:32:31",
      "content": "<p>No, for the U2I and I2I I control the data usage very well. </p>",
      "rawMarkdown": "No, for the U2I and I2I I control the data usage very well.",
      "votes": null
    },
    {
      "id": "1784494",
      "postDate": "05/11/2022 07:58:24",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a><br>\nWow!!<br>\nI didn't come up with such a way to get customer embedding!<br>\nThank you.</p>",
      "rawMarkdown": "lihaorocky\nWow!!\nI didn't come up with such a way to get customer embedding!\nThank you.",
      "votes": null
    },
    {
      "id": "1784614",
      "postDate": "05/11/2022 10:07:21",
      "content": "<p>I guess it's item2vec? <a href=\"https://arxiv.org/abs/1603.04259\" target=\"_blank\">https://arxiv.org/abs/1603.04259</a></p>",
      "rawMarkdown": "I guess it's item2vec? https://arxiv.org/abs/1603.04259",
      "votes": null
    },
    {
      "id": "1784623",
      "postDate": "05/11/2022 10:29:59",
      "content": "<p>Never heard of \"item2vec\" before, but it seems a promising approach. I just use gensim training a word2vec model with the article_id sequence to get the article_ids' embeddings. </p>",
      "rawMarkdown": "Never heard of \"item2vec\" before, but it seems a promising approach. I just use gensim training a word2vec model with the article_id sequence to get the article_ids' embeddings.",
      "votes": null
    },
    {
      "id": "1785340",
      "postDate": "05/12/2022 03:31:38",
      "content": "<p>\"Item2vec\" is just word2vec applied to items. The only difference is that they treat articles each customer buy as a set instead of sequence(loss computed over all pairs of items in the set). I guess this is the difference with your approach.</p>",
      "rawMarkdown": "\"Item2vec\" is just word2vec applied to items. The only difference is that they treat articles each customer buy as a set instead of sequence(loss computed over all pairs of items in the set). I guess this is the difference with your approach.",
      "votes": null
    },
    {
      "id": "1785345",
      "postDate": "05/12/2022 03:39:27",
      "content": "<p>Thank you for sharing this. Sounds great. I will definitely try it next time.</p>",
      "rawMarkdown": "Thank you for sharing this. Sounds great. I will definitely try it next time.",
      "votes": null
    },
    {
      "id": "1790740",
      "postDate": "05/15/2022 09:25:01",
      "content": "<p>\"articles features, which will be constant to all customers once you fixed the overvation time\"means \"use transactions before overvation time to bulid features, and the features will be constant no matter the occurrence time  of transactions?</p>",
      "rawMarkdown": "\"articles features, which will be constant to all customers once you fixed the overvation time\"means \"use transactions before overvation time to bulid features, and the features will be constant no matter the occurrence time  of transactions?",
      "votes": null
    },
    {
      "id": "1790883",
      "postDate": "05/15/2022 12:48:55",
      "content": "<p>Imagine you have candidates A1,A2 for customer C1,C2,C3 on observation datetime D1, you could first calculate features for articles A1,A2,A3 and the article feature part for pairs (C1, A1), (C2, A1), (C3, A1) will be constant although the customers are different for those pairs, which will save quite an amount of calculation time. </p>",
      "rawMarkdown": "Imagine you have candidates A1,A2 for customer C1,C2,C3 on observation datetime D1, you could first calculate features for articles A1,A2,A3 and the article feature part for pairs (C1, A1), (C2, A1), (C3, A1) will be constant although the customers are different for those pairs, which will save quite an amount of calculation time.",
      "votes": null
    },
    {
      "id": "1792012",
      "postDate": "05/16/2022 14:55:59",
      "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> </p>",
      "rawMarkdown": "thanks for sharing @lihaorocky",
      "votes": null
    },
    {
      "id": "1792884",
      "postDate": "05/17/2022 11:48:28",
      "content": "<p>Congratulations!<br>\nI have a very simple question to ask, how do you divide the time to get the data source for your ranking model? For example, train the recall model with the training set of the first week, then get the recall results of the second week, and use the recall results to train the ranking model. After that, train the recall model with the data from the first two weeks, then get the recall results from the third week, and use the recall results to train the ranking model …… </p>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratulations!\nI have a very simple question to ask, how do you divide the time to get the data source for your ranking model? For example, train the recall model with the training set of the first week, then get the recall results of the second week, and use the recall results to train the ranking model. After that, train the recall model with the data from the first two weeks, then get the recall results from the third week, and use the recall results to train the ranking model ...... \n\nThanks!",
      "votes": null
    },
    {
      "id": "1792990",
      "postDate": "05/17/2022 13:53:35",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Congratsss!!</p>",
      "rawMarkdown": "lihaorocky Congratsss!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783095,
      "author_name": "zakopur0",
      "author_url": "",
      "post_date": "05/10/2022 06:00:31",
      "content": "<p>Congratulations!<br>\nI learned a lot from you during the competition!<br>\nThank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783098,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/10/2022 06:02:17",
          "content": "<p>You are welcome! You also did well. Congrats!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783135,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "05/10/2022 06:37:20",
          "content": "<p>Without your comments we couldn't have built our model.<br>\nThank you for your contributions to the many discussions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783109,
      "author_name": "juzqyxs",
      "author_url": "",
      "post_date": "05/10/2022 06:14:08",
      "content": "<p>Actually, use all transaction data to train word2vec embedding is leaked, but the relationship of different items may be represented,.Like the competition hosted by Wechat in 2021.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783119,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/10/2022 06:19:40",
          "content": "<p>Yes, I just wonder how other teams deal with this problem to make the CV-LB gap much smaller than mine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783121,
          "author_name": "juzqyxs",
          "author_url": "",
          "post_date": "05/10/2022 06:22:35",
          "content": "<p>We do not use all transaction data, only train word2vec in history transactions, so the cv score not boosted.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783126,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/10/2022 06:25:00",
          "content": "<p>Got it. Thanks! Then you must train a lot of word2vec models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783131,
          "author_name": "juzqyxs",
          "author_url": "",
          "post_date": "05/10/2022 06:30:55",
          "content": "<p>Because many word2vec models were trained, the distribution may not in same space, that could be the reason why cv not boosted I guess.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783132,
      "author_name": "alexkim2",
      "author_url": "",
      "post_date": "05/10/2022 06:32:26",
      "content": "<p>Congratulations! Thank you for your sharing. I have learned a lot from you during the competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783133,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/10/2022 06:34:18",
          "content": "<p>Glad to hear that. And you also did well! Congrats!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783144,
      "author_name": "hanejiyuto",
      "author_url": "",
      "post_date": "05/10/2022 06:54:33",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nCongratulations for your gold medal!!<br>\nAnd thank you for your answering my question in discussion during the competition, which improved my place a lot!<br>\nCould you tell me what is “cosine similarity between each pair of customer and article” in details?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783159,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/10/2022 07:04:51",
          "content": "<p>Glad to hear that!<br>\nSo you have a lot of candidate articles for each customer. And after you get the embedding of each article, you could compute the customer's embedding represention based on the customer's transaction history(aggregate the articles' embeddings to compute the mean/median embedding). Then you have the embedding of the customer and the embedding of the candidate article, you could calculate the vector cosine distance (also cosine similarity) of the customer and the article using dot(vec1, vec2) / (norm(vec1) * norm(vec2)), which represent how close between the customer and the article.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783768,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/10/2022 16:48:01",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> - I was wondering what you meant when you shared during the competition that Customer/Article similarity features helped a lot.</p>\n<p>I thought of using mean of customer's article embeddings, but I thought that it's probably meaningless - once you average together different articles' embeddings, you lose the information.</p>\n<p>Guess I was wrong :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783772,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/10/2022 16:51:51",
          "content": "<p>Is that also what you used for user-based collaborative filtering?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1784139,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/11/2022 00:42:38",
          "content": "<p>Yes, those features help a lot. And I don't think it's meaningless. We all know every customer has his/her own preference for styles(articles). And if our representation embeddings of articles are not bad, they have the ability to distinguish the styles. For example, if our embedding is 1-dimensional vector, a set of articles' embedding have values like [0.89], [0.90], [0.902]… Calculating average of those embeddings, you get a value around [0.9]. For another style set of articles, if one customer frequently bought a lot, after aggregating the embeddings of those articles, you get the mean embedding vector [0.2]. Well now we know where the separating capability of this method comes from.<br>\nIt may lose some info, but the noise info is also filtered and we get the main taste of the customer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1784494,
          "author_name": "hanejiyuto",
          "author_url": "",
          "post_date": "05/11/2022 07:58:24",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a><br>\nWow!!<br>\nI didn't come up with such a way to get customer embedding!<br>\nThank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783229,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/10/2022 08:03:02",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Congratulations!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783385,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "05/10/2022 11:02:34",
      "content": "<p>Congrats!If you don't use leakage features,should get higher accuracy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783470,
      "author_name": "biubiug",
      "author_url": "",
      "post_date": "05/10/2022 12:22:41",
      "content": "<p>Congratulations!!! very impressive for your method to calibrate！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783671,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "05/10/2022 15:12:48",
      "content": "<p>Congrats and thanks for sharing! <br>\nAbout the word2vec and other pretrain embedding methods, I just train one model for each target week using the transactions before the week, and the result seems no leak.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783684,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/10/2022 15:31:18",
          "content": "<p>Thanks for sharing your method. Guess I was too lazy to even think about training more than one word2vec model to avoid the leak. And also congrats to your solo 3rd place finish. Your score improving progress during last weeks are quite impressive!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783703,
          "author_name": "sirius81",
          "author_url": "",
          "post_date": "05/10/2022 15:52:55",
          "content": "<p>Saying the last week progressing, I must thank the covid-19 that isolated me at home for two months😂, which forced me to focus on the competition in the spare time, especially during May day holiday. Just says every coin has two sides.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1784142,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/11/2022 00:54:38",
          "content": "<p>How \"lucky\" you are… 😂 Anyway you did really well and all the best staying at home. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783770,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/10/2022 16:49:19",
      "content": "<p>what text are you referring to when you reference \"word2vec\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784128,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/11/2022 00:27:10",
          "content": "<p>For word2vec, I mean using customers' time-ordered purchase article sequence to train a word2vec model getting the embeddings of articles.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1784614,
          "author_name": "homoalways",
          "author_url": "",
          "post_date": "05/11/2022 10:07:21",
          "content": "<p>I guess it's item2vec? <a href=\"https://arxiv.org/abs/1603.04259\" target=\"_blank\">https://arxiv.org/abs/1603.04259</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1784623,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/11/2022 10:29:59",
          "content": "<p>Never heard of \"item2vec\" before, but it seems a promising approach. I just use gensim training a word2vec model with the article_id sequence to get the article_ids' embeddings. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1785340,
          "author_name": "homoalways",
          "author_url": "",
          "post_date": "05/12/2022 03:31:38",
          "content": "<p>\"Item2vec\" is just word2vec applied to items. The only difference is that they treat articles each customer buy as a set instead of sequence(loss computed over all pairs of items in the set). I guess this is the difference with your approach.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1785345,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/12/2022 03:39:27",
          "content": "<p>Thank you for sharing this. Sounds great. I will definitely try it next time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1784257,
      "author_name": "jialunhe",
      "author_url": "",
      "post_date": "05/11/2022 04:13:22",
      "content": "<p>Congratulations!<br>\nIf you use all the data to build word2vec, that's definitely data leak. Other possible sources of data leakage are in building collaborative filtering. Did you use all the data to build U2I and I2I? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1784268,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/11/2022 04:32:31",
          "content": "<p>No, for the U2I and I2I I control the data usage very well. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1790740,
      "author_name": "dengniewei",
      "author_url": "",
      "post_date": "05/15/2022 09:25:01",
      "content": "<p>\"articles features, which will be constant to all customers once you fixed the overvation time\"means \"use transactions before overvation time to bulid features, and the features will be constant no matter the occurrence time  of transactions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1790883,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/15/2022 12:48:55",
          "content": "<p>Imagine you have candidates A1,A2 for customer C1,C2,C3 on observation datetime D1, you could first calculate features for articles A1,A2,A3 and the article feature part for pairs (C1, A1), (C2, A1), (C3, A1) will be constant although the customers are different for those pairs, which will save quite an amount of calculation time. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1792012,
      "author_name": "fajarwibowo",
      "author_url": "",
      "post_date": "05/16/2022 14:55:59",
      "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1792884,
      "author_name": "gaorecmaster",
      "author_url": "",
      "post_date": "05/17/2022 11:48:28",
      "content": "<p>Congratulations!<br>\nI have a very simple question to ask, how do you divide the time to get the data source for your ranking model? For example, train the recall model with the training set of the first week, then get the recall results of the second week, and use the recall results to train the ranking model. After that, train the recall model with the data from the first two weeks, then get the recall results from the third week, and use the recall results to train the ranking model …… </p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1792990,
      "author_name": "adithyaawati",
      "author_url": "",
      "post_date": "05/17/2022 13:53:35",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Congratsss!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1783085": "First of all, thanks to the competition organizers to host such an interesting competition and my great teammates @rendongltt (who is a huge fan of online-shopping, which gives us a lot of help😂) and @jiaqizhang35 (who is my friend, without whom I didn't even join this competition). And also I would like to thank @paweljankiewicz, without his great posts, we will not reach this far. And also you (and also a lot of other nice guys, who love sharing) motivated me to do more sharing with the community.\n\n**Overview**\nOverall we love this competition so much because it did provide infinite ways to solve this problem especially the possible recall methods. Talking to feature engineering, even though we're quite confident about our skills, but we may have no shot to get a gold medal compared to other grandmasters. So we focused ourselves on the recall methods from the very start. And our recall methods is deeply dependent on article embeddings (image, txt, word2vec, tfidf) and the amazing python package faiss to find the similar articles and customers. As to the feature engineering, we didn't spend too much time on it but using different days window to create aggreated stat features of articles and customers seperatly and customer-article cross aggreated stat features and different cosine similarity feaures using embeddings aforementioned. \n\n**Recall methods**\nThe recall methods comprised of those parts:\n1. Customer's last buckets(and same product code to customers' last bucket's articles), recent bought articles\n2. U2I: User based collaborative filtering. \n3. I2i: Item based collaborative filtering\n4. word2vec: to predict customer's next possible purchase articles after last buckets\n5. Articles bought together. I modified this nice [notebook](https://www.kaggle.com/code/titericz/article-id-pairs-in-3s-using-cudf) from @titericz\n6. Popular articles with different customers' attributes like age bin, gender etc.\nSo overall we use 21 recall methods to create article candidates.\n\n**Embedding methods**\nWe created article embdding using different data and methods\n1. For article images, we use swin transformer to extract the embedding features.\n2. For the article texts, first we concatnate all text values of each article and then I use SentenceTransformer to extract the embeddings\n3. Using this great [notebook](https://www.kaggle.com/code/aerdem4/h-m-rapids-article2vec) from  @aerdem4 you can easily get article embeddings using tfidf method\n4. The last but also the most useful method to get article embeddings are word2vec method, for which we use gensim to ease the process.\n\n**Recall method calibration**\nBecause different recall methods have different positive ratio, so it's better to use some method to calibrate all the recall methods. I use the method which I shared [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672#:~:text=I%20would%20sugguest,goes%20like%20this.) to do the preprocess.\n\n**Feature engineering**\nNothing special about this part but I found some method which will hugely accelerate the computing time for feature computing. \nBasicly we divide the features into 3 parts:\n1. articles features, which will be constant to all customers once you fixed the overvation time\n2. similarly, customer features will be contant to all candidate articles once you fixed the overvation date\n3. customer-article features, like cosine similarity between each pair of customer and article, customers' bought history aggreated features of this particular article.\nFor part 1 and 2, you could do the computation and save them once and merged with the features of part 3 later. This method really save me a lot of time especially during inference time.\n\n**Modeling** \nBecause we spend most of the time to devolop recall strategies, we use only lightgbm ranker model for training and don't even have time to try other methods like catboost and xgboost. So basically our score is single model's score.\n\n**Remaining question**\nWe have one question that we still couldn't solve it right now, which is about the word2vec model. We use all transaction data to train the word2vec model, which actually leaked the future info a little bit. \nImagne that for observation date 20200915 you compute the customer-article cosine similarity using the word2vec model trained using all transaction data. The word2vec embedings of two article ids aid_1, aid_2, which maybe mainly appeared duing week 20200916-20200922, are used on the observation date 20200915. \nThe effect is everytime I use those embeddings to create candidates or compute features, the local CV will improve a lot while the LB improves only a little bit. For my final ranking model, my map@12 score for last week is 0.0441, but the LB is only 0.0350, the gap is much larger than the ones mentioned by others. But if we don't use this feature, both CV and LB score are dropping. So we don't really know how to do it to make the most use of this method. What do you think?",
    "1783095": "Congratulations!\nI learned a lot from you during the competition!\nThank you!",
    "1783098": "You are welcome! You also did well. Congrats!",
    "1783109": "Actually, use all transaction data to train word2vec embedding is leaked, but the relationship of different items may be represented,.Like the competition hosted by Wechat in 2021.",
    "1783119": "Yes, I just wonder how other teams deal with this problem to make the CV-LB gap much smaller than mine.",
    "1783121": "We do not use all transaction data, only train word2vec in history transactions, so the cv score not boosted.",
    "1783126": "Got it. Thanks! Then you must train a lot of word2vec models.",
    "1783131": "Because many word2vec models were trained, the distribution may not in same space, that could be the reason why cv not boosted I guess.",
    "1783132": "Congratulations! Thank you for your sharing. I have learned a lot from you during the competition!",
    "1783133": "Glad to hear that. And you also did well! Congrats!",
    "1783135": "Without your comments we couldn't have built our model.\nThank you for your contributions to the many discussions.",
    "1783144": "lihaorocky \nCongratulations for your gold medal!!\nAnd thank you for your answering my question in discussion during the competition, which improved my place a lot!\nCould you tell me what is “cosine similarity between each pair of customer and article” in details?",
    "1783159": "Glad to hear that!\nSo you have a lot of candidate articles for each customer. And after you get the embedding of each article, you could compute the customer's embedding represention based on the customer's transaction history(aggregate the articles' embeddings to compute the mean/median embedding). Then you have the embedding of the customer and the embedding of the candidate article, you could calculate the vector cosine distance (also cosine similarity) of the customer and the article using dot(vec1, vec2) / (norm(vec1) * norm(vec2)), which represent how close between the customer and the article.",
    "1783229": "lihaorocky Congratulations!!",
    "1783385": "Congrats!If you don't use leakage features,should get higher accuracy.",
    "1783470": "Congratulations!!! very impressive for your method to calibrate！",
    "1783671": "Congrats and thanks for sharing! \nAbout the word2vec and other pretrain embedding methods, I just train one model for each target week using the transactions before the week, and the result seems no leak.",
    "1783684": "Thanks for sharing your method. Guess I was too lazy to even think about training more than one word2vec model to avoid the leak. And also congrats to your solo 3rd place finish. Your score improving progress during last weeks are quite impressive!",
    "1783703": "Saying the last week progressing, I must thank the covid-19 that isolated me at home for two months😂, which forced me to focus on the competition in the spare time, especially during May day holiday. Just says every coin has two sides.",
    "1783768": "lihaorocky - I was wondering what you meant when you shared during the competition that Customer/Article similarity features helped a lot.\n\nI thought of using mean of customer's article embeddings, but I thought that it's probably meaningless - once you average together different articles' embeddings, you lose the information.\n\nGuess I was wrong :)",
    "1783770": "what text are you referring to when you reference \"word2vec\"?",
    "1783772": "Is that also what you used for user-based collaborative filtering?",
    "1784128": "For word2vec, I mean using customers' time-ordered purchase article sequence to train a word2vec model getting the embeddings of articles.",
    "1784139": "Yes, those features help a lot. And I don't think it's meaningless. We all know every customer has his/her own preference for styles(articles). And if our representation embeddings of articles are not bad, they have the ability to distinguish the styles. For example, if our embedding is 1-dimensional vector, a set of articles' embedding have values like [0.89], [0.90], [0.902]... Calculating average of those embeddings, you get a value around [0.9]. For another style set of articles, if one customer frequently bought a lot, after aggregating the embeddings of those articles, you get the mean embedding vector [0.2]. Well now we know where the separating capability of this method comes from.\nIt may lose some info, but the noise info is also filtered and we get the main taste of the customer.",
    "1784142": "How \"lucky\" you are... 😂 Anyway you did really well and all the best staying at home.",
    "1784257": "Congratulations!\nIf you use all the data to build word2vec, that's definitely data leak. Other possible sources of data leakage are in building collaborative filtering. Did you use all the data to build U2I and I2I?",
    "1784268": "No, for the U2I and I2I I control the data usage very well.",
    "1784494": "lihaorocky\nWow!!\nI didn't come up with such a way to get customer embedding!\nThank you.",
    "1784614": "I guess it's item2vec? https://arxiv.org/abs/1603.04259",
    "1784623": "Never heard of \"item2vec\" before, but it seems a promising approach. I just use gensim training a word2vec model with the article_id sequence to get the article_ids' embeddings.",
    "1785340": "\"Item2vec\" is just word2vec applied to items. The only difference is that they treat articles each customer buy as a set instead of sequence(loss computed over all pairs of items in the set). I guess this is the difference with your approach.",
    "1785345": "Thank you for sharing this. Sounds great. I will definitely try it next time.",
    "1790740": "\"articles features, which will be constant to all customers once you fixed the overvation time\"means \"use transactions before overvation time to bulid features, and the features will be constant no matter the occurrence time  of transactions?",
    "1790883": "Imagine you have candidates A1,A2 for customer C1,C2,C3 on observation datetime D1, you could first calculate features for articles A1,A2,A3 and the article feature part for pairs (C1, A1), (C2, A1), (C3, A1) will be constant although the customers are different for those pairs, which will save quite an amount of calculation time.",
    "1792012": "thanks for sharing @lihaorocky",
    "1792884": "Congratulations!\nI have a very simple question to ask, how do you divide the time to get the data source for your ranking model? For example, train the recall model with the training set of the first week, then get the recall results of the second week, and use the recall results to train the ranking model. After that, train the recall model with the data from the first two weeks, then get the recall results from the third week, and use the recall results to train the ranking model ...... \n\nThanks!",
    "1792990": "lihaorocky Congratsss!!"
  },
  "source": "meta"
}