{
  "id": 316386,
  "title": "Setup Local CV – for Modeling",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386",
  "author_name": "",
  "post_date": "2022-04-01T16:28:10.515030400Z",
  "votes": 77,
  "comment_count": 25,
  "views": 0,
  "content": "<h2>TLDR;</h2>\n<h2>Using ML modeling for ranking?</h2>\n<h2>Adjust CV to avoid overfitting.</h2>\n<p>In <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919\" target=\"_blank\">this post</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> shares how to setup local cv for this competition.</p>\n<p>This works well for many public notebooks, but if move on to using ML models for ranking candidates, you need a little more.</p>\n<p>Let me explain: </p>\n<p>In <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">this popular post</a>, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> discusses the two stages often taken for this kind of problem:  </p>\n<ol>\n<li>First you find promising “candidates” – customer/item pairs that are likely to have happened.  </li>\n<li>Then you train a model to “rank”, for each customer, how likely each of the candidates are.    </li>\n</ol>\n<p>For submission, you take the 12 highest-ranked candidates for each customer.  </p>\n<p>As he mentions, candidate selection doesn’t necessarily need to involve ML. Many public high-scoring notebooks go a step further – not only do they select candidates based on hard-coded methods, but the ranking is also done without ML, by sorting on certain columns.</p>\n<p>Let’s look at the data like this:   <br>\n(instead of using dates, we’ll use weeks, where week #105 = LB data)<br>\n<a href=\"https://postimg.cc/QH2BWBfz\" target=\"_blank\"><img src=\"https://i.postimg.cc/9QMyKZZc/handm-data-lb.jpg\" alt=\"handm-data-lb.jpg\"></a></p>\n<p>This is what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>’s cv strategy looks like:</p>\n<p><a href=\"https://postimg.cc/V05x1kM9\" target=\"_blank\"><img src=\"https://i.postimg.cc/mkSBYcgn/handm-cdeotte-cv.jpg\" alt=\"handm-cdeotte-cv.jpg\"></a></p>\n<p>For each fold, the “train” data is where we get the features, and the “validation” week is where we get the ground truth – the labels.</p>\n<p>We decide on a strategy for selecting candidates and ranking them using features from the train data, and evaluate how well they do by looking at the labels from the validation week.  </p>\n<p><strong>But if we’re designing our algorithm based on the actual labels, why aren't we overfitting?</strong></p>\n<p>The answer is that hard-coded algorithms aren’t very flexible, so we don’t have to worry much about overfitting. (in ML terms, you could say they have high bias and low variance)</p>\n<p>Once we move to a ML model like LGBMRanker, overfitting is going to be a big problem.</p>\n<p>For modeling, we need a CV that looks like this:<br>\n<a href=\"https://postimg.cc/v4Zr3xrZ\" target=\"_blank\"><img src=\"https://i.postimg.cc/cLQF1Q6w/handm-my-cv.jpg\" alt=\"handm-my-cv.jpg\"></a></p>\n<p>This way we can tune the model on the training features/labels, and do evaluation using the validation data/labels.</p>\n<p>Could be this is simple for many people, but it took me a little while to wrap my head around it.<br>\nHope it’s helpful for you!</p>",
  "messages": [
    {
      "id": "1742269",
      "postDate": "04/01/2022 16:28:10",
      "content": "<h2>TLDR;</h2>\n<h2>Using ML modeling for ranking?</h2>\n<h2>Adjust CV to avoid overfitting.</h2>\n<p>In <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919\" target=\"_blank\">this post</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> shares how to setup local cv for this competition.</p>\n<p>This works well for many public notebooks, but if move on to using ML models for ranking candidates, you need a little more.</p>\n<p>Let me explain: </p>\n<p>In <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">this popular post</a>, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> discusses the two stages often taken for this kind of problem:  </p>\n<ol>\n<li>First you find promising “candidates” – customer/item pairs that are likely to have happened.  </li>\n<li>Then you train a model to “rank”, for each customer, how likely each of the candidates are.    </li>\n</ol>\n<p>For submission, you take the 12 highest-ranked candidates for each customer.  </p>\n<p>As he mentions, candidate selection doesn’t necessarily need to involve ML. Many public high-scoring notebooks go a step further – not only do they select candidates based on hard-coded methods, but the ranking is also done without ML, by sorting on certain columns.</p>\n<p>Let’s look at the data like this:   <br>\n(instead of using dates, we’ll use weeks, where week #105 = LB data)<br>\n<a href=\"https://postimg.cc/QH2BWBfz\" target=\"_blank\"><img src=\"https://i.postimg.cc/9QMyKZZc/handm-data-lb.jpg\" alt=\"handm-data-lb.jpg\"></a></p>\n<p>This is what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>’s cv strategy looks like:</p>\n<p><a href=\"https://postimg.cc/V05x1kM9\" target=\"_blank\"><img src=\"https://i.postimg.cc/mkSBYcgn/handm-cdeotte-cv.jpg\" alt=\"handm-cdeotte-cv.jpg\"></a></p>\n<p>For each fold, the “train” data is where we get the features, and the “validation” week is where we get the ground truth – the labels.</p>\n<p>We decide on a strategy for selecting candidates and ranking them using features from the train data, and evaluate how well they do by looking at the labels from the validation week.  </p>\n<p><strong>But if we’re designing our algorithm based on the actual labels, why aren't we overfitting?</strong></p>\n<p>The answer is that hard-coded algorithms aren’t very flexible, so we don’t have to worry much about overfitting. (in ML terms, you could say they have high bias and low variance)</p>\n<p>Once we move to a ML model like LGBMRanker, overfitting is going to be a big problem.</p>\n<p>For modeling, we need a CV that looks like this:<br>\n<a href=\"https://postimg.cc/v4Zr3xrZ\" target=\"_blank\"><img src=\"https://i.postimg.cc/cLQF1Q6w/handm-my-cv.jpg\" alt=\"handm-my-cv.jpg\"></a></p>\n<p>This way we can tune the model on the training features/labels, and do evaluation using the validation data/labels.</p>\n<p>Could be this is simple for many people, but it took me a little while to wrap my head around it.<br>\nHope it’s helpful for you!</p>",
      "rawMarkdown": "## TLDR; \n## Using ML modeling for ranking? \n## Adjust CV to avoid overfitting.\n\nIn [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919), @cdeotte shares how to setup local cv for this competition.\n\nThis works well for many public notebooks, but if move on to using ML models for ranking candidates, you need a little more.\n\nLet me explain: \n\nIn [this popular post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288), @paweljankiewicz discusses the two stages often taken for this kind of problem:  \n1. First you find promising “candidates” – customer/item pairs that are likely to have happened.  \n2. Then you train a model to “rank”, for each customer, how likely each of the candidates are.    \n\nFor submission, you take the 12 highest-ranked candidates for each customer.  \n\nAs he mentions, candidate selection doesn’t necessarily need to involve ML. Many public high-scoring notebooks go a step further – not only do they select candidates based on hard-coded methods, but the ranking is also done without ML, by sorting on certain columns.\n\nLet’s look at the data like this:   \n(instead of using dates, we’ll use weeks, where week #105 = LB data)\n[![handm-data-lb.jpg](https://i.postimg.cc/9QMyKZZc/handm-data-lb.jpg)](https://postimg.cc/QH2BWBfz)\n\n\nThis is what @cdeotte’s cv strategy looks like:\n\n[![handm-cdeotte-cv.jpg](https://i.postimg.cc/mkSBYcgn/handm-cdeotte-cv.jpg)](https://postimg.cc/V05x1kM9)\n\nFor each fold, the “train” data is where we get the features, and the “validation” week is where we get the ground truth – the labels.\n\nWe decide on a strategy for selecting candidates and ranking them using features from the train data, and evaluate how well they do by looking at the labels from the validation week.  \n\n**But if we’re designing our algorithm based on the actual labels, why aren't we overfitting?**\n\nThe answer is that hard-coded algorithms aren’t very flexible, so we don’t have to worry much about overfitting. (in ML terms, you could say they have high bias and low variance)\n\nOnce we move to a ML model like LGBMRanker, overfitting is going to be a big problem.\n\nFor modeling, we need a CV that looks like this:\n[![handm-my-cv.jpg](https://i.postimg.cc/cLQF1Q6w/handm-my-cv.jpg)](https://postimg.cc/v4Zr3xrZ)\n\nThis way we can tune the model on the training features/labels, and do evaluation using the validation data/labels.\n\nCould be this is simple for many people, but it took me a little while to wrap my head around it.\nHope it’s helpful for you!",
      "votes": null
    },
    {
      "id": "1742679",
      "postDate": "04/02/2022 07:05:23",
      "content": "<p>Thanks     👍</p>",
      "rawMarkdown": "Thanks     👍",
      "votes": null
    },
    {
      "id": "1743019",
      "postDate": "04/02/2022 14:09:57",
      "content": "<p>Infomative! Thank you</p>",
      "rawMarkdown": "Infomative! Thank you",
      "votes": null
    },
    {
      "id": "1744358",
      "postDate": "04/03/2022 22:57:44",
      "content": "<p>Good summary. In my case my train features are generated before 2020-04-01 and move week by week until 2020-09-22. I noticed that the ranking model doesn't need much data. For example to train the ranking model I use only labels generated from last 5 months. I think there are some gains to be made when using more data for training but I would need to invest to computation. It is too easy to just add RAM much harder to optimize the code and one's thinking.</p>",
      "rawMarkdown": "Good summary. In my case my train features are generated before 2020-04-01 and move week by week until 2020-09-22. I noticed that the ranking model doesn't need much data. For example to train the ranking model I use only labels generated from last 5 months. I think there are some gains to be made when using more data for training but I would need to invest to computation. It is too easy to just add RAM much harder to optimize the code and one's thinking.",
      "votes": null
    },
    {
      "id": "1744780",
      "postDate": "04/04/2022 09:57:58",
      "content": "<p>Thanks alot , Helpful</p>",
      "rawMarkdown": "Thanks alot , Helpful",
      "votes": null
    },
    {
      "id": "1744988",
      "postDate": "04/04/2022 13:38:12",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>!</p>\n<p>So it sounds like I got it a bit wrong - whereas I've put each train/validation split as a separate \"fold\", to be trained/evaluated separately, I should instead be using them to generate training data with labels, then concatenate all the data/labels together, and train on everything together.</p>\n<p>Is that right?</p>",
      "rawMarkdown": "Thanks, @paweljankiewicz!\n\nSo it sounds like I got it a bit wrong - whereas I've put each train/validation split as a separate \"fold\", to be trained/evaluated separately, I should instead be using them to generate training data with labels, then concatenate all the data/labels together, and train on everything together.\n\nIs that right?",
      "votes": null
    },
    {
      "id": "1745075",
      "postDate": "04/04/2022 15:17:06",
      "content": "<p>Yes, I think you can concatenate more than 1 week for training your ranking model. I, kind of assumed, this from the beginning. The model for submission I train without validation on all the weeks up to 09/22 and it improves by about 0.0005 (compared to training set without the last week). I haven't done exhaustive testing. It would be kind of interesting to see the effect on the leaderboard/validation of adding more and more weeks in reverse order.</p>\n<p>My explanation for this is that when you are training your ranking model all the strategies you are using are kind of relative: bestsellers are bestsellers whichever period you are analysing. It may be that H&amp;M used different recommendation strategies throughout the provided data which may totally skew the strategies. Which would be my biggest concern in this competition - to what extent we are creating something new vs recreating H&amp;M recommendation strategy?</p>",
      "rawMarkdown": "Yes, I think you can concatenate more than 1 week for training your ranking model. I, kind of assumed, this from the beginning. The model for submission I train without validation on all the weeks up to 09/22 and it improves by about 0.0005 (compared to training set without the last week). I haven't done exhaustive testing. It would be kind of interesting to see the effect on the leaderboard/validation of adding more and more weeks in reverse order.\n\nMy explanation for this is that when you are training your ranking model all the strategies you are using are kind of relative: bestsellers are bestsellers whichever period you are analysing. It may be that H&M used different recommendation strategies throughout the provided data which may totally skew the strategies. Which would be my biggest concern in this competition - to what extent we are creating something new vs recreating H&M recommendation strategy?",
      "votes": null
    },
    {
      "id": "1745089",
      "postDate": "04/04/2022 15:29:11",
      "content": "<p>Makes sense, thanks!</p>",
      "rawMarkdown": "Makes sense, thanks!",
      "votes": null
    },
    {
      "id": "1745601",
      "postDate": "04/05/2022 04:54:11",
      "content": "<p>Interesting stuff, thanks for sharing 🙌</p>",
      "rawMarkdown": "Interesting stuff, thanks for sharing 🙌",
      "votes": null
    },
    {
      "id": "1746251",
      "postDate": "04/05/2022 15:46:14",
      "content": "<p>Thank you for your work. <br>\nI have a question. What is train label?   Does train label mean wether  a customer bought article or not?</p>",
      "rawMarkdown": "Thank you for your work. \nI have a question. What is train label?   Does train label mean wether  a customer bought article or not?",
      "votes": null
    },
    {
      "id": "1746295",
      "postDate": "04/05/2022 16:33:28",
      "content": "<p>Yes, label means whether the customer bought the article or not. That's what our model is trying to predict.</p>\n<p>Train labels are the labels for the training data, that we're training the model on.<br>\nValid labels are the labels for the validation data, which we'll use to evaluate how well our trained model's predictions were.</p>",
      "rawMarkdown": "Yes, label means whether the customer bought the article or not. That's what our model is trying to predict.\n\nTrain labels are the labels for the training data, that we're training the model on.\nValid labels are the labels for the validation data, which we'll use to evaluate how well our trained model's predictions were.",
      "votes": null
    },
    {
      "id": "1746323",
      "postDate": "04/05/2022 16:57:40",
      "content": "<p>Thank you for your answer. <br>\nI have one more question. How did you generate negative labels?<br>\nI think that labels are 0 or 1. For example, Fold 0, positive  train label  for the customer  is the article the customer bought in week 103. What is negative label for the customer.</p>",
      "rawMarkdown": "Thank you for your answer. \nI have one more question. How did you generate negative labels?\nI think that labels are 0 or 1. For example, Fold 0, positive  train label  for the customer  is the article the customer bought in week 103. What is negative label for the customer.",
      "votes": null
    },
    {
      "id": "1746424",
      "postDate": "04/05/2022 18:50:07",
      "content": "<p>It's any of the candidates you selected (from step #1), that the customer did not buy in week 103.<br>\nSee <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">this post</a> and the comments there for a more detailed explanation.</p>",
      "rawMarkdown": "It's any of the candidates you selected (from step #1), that the customer did not buy in week 103.\nSee [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288) and the comments there for a more detailed explanation.",
      "votes": null
    },
    {
      "id": "1746687",
      "postDate": "04/06/2022 02:41:01",
      "content": "<p>Thank you so much</p>",
      "rawMarkdown": "Thank you so much",
      "votes": null
    },
    {
      "id": "1747165",
      "postDate": "04/06/2022 12:42:07",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!",
      "votes": null
    },
    {
      "id": "1747630",
      "postDate": "04/06/2022 20:19:04",
      "content": "<p>Interesting stuff. helpful</p>",
      "rawMarkdown": "Interesting stuff. helpful",
      "votes": null
    },
    {
      "id": "1748140",
      "postDate": "04/07/2022 10:48:23",
      "content": "<p>Shocked by 5 months… I use only last 2 weeks transaction to build the training samples (separately). Have you compared the cv scores of different amount of training samples? </p>",
      "rawMarkdown": "Shocked by 5 months... I use only last 2 weeks transaction to build the training samples (separately). Have you compared the cv scores of different amount of training samples?",
      "votes": null
    },
    {
      "id": "1748416",
      "postDate": "04/07/2022 15:11:19",
      "content": "<p><a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> I checked and used 6 months and it worsened the result, I checked 4 months and it worsened the results too. So I guess I'm using the right amount of transactions :D. And this is only for the ranking model, for many features I'm actually using all the transactions.</p>",
      "rawMarkdown": "sirius81 I checked and used 6 months and it worsened the result, I checked 4 months and it worsened the results too. So I guess I'm using the right amount of transactions :D. And this is only for the ranking model, for many features I'm actually using all the transactions.",
      "votes": null
    },
    {
      "id": "1748467",
      "postDate": "04/07/2022 16:06:07",
      "content": "<p>Thanks for the information, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> .<br>\nWhen you say \"my train features are generated before 2020-04-01 and move week by week until 2020-09-22\", what kind of features you mean?</p>\n<p>For instance:<br>\nFor week 100, your features would be: \"bought_article_next_week? (0 or 1)\" indicating if this article was bought by this customer in the week 101.<br>\nAnd the same goes on for the week pairs (X: 101, y: 102), (X: 102, y: 103), … (…)</p>\n<p>Is this kind of feature you mean?</p>",
      "rawMarkdown": "Thanks for the information, @paweljankiewicz .\nWhen you say \"my train features are generated before 2020-04-01 and move week by week until 2020-09-22\", what kind of features you mean?\n\nFor instance:\nFor week 100, your features would be: \"bought_article_next_week? (0 or 1)\" indicating if this article was bought by this customer in the week 101.\nAnd the same goes on for the week pairs (X: 101, y: 102), (X: 102, y: 103), ... (...)\n\nIs this kind of feature you mean?",
      "votes": null
    },
    {
      "id": "1772379",
      "postDate": "04/30/2022 07:33:30",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> , can you explain the difference between training separately vs concatenating all labels at the end? I'm having a bit of trouble understanding how labels can be concatenated in the first place. </p>\n<p>ex. train on week 0-50, predict on week 51; <br>\ntrain 1-51, predict on week 52;<br>\nthen average the results of the two predictions?</p>",
      "rawMarkdown": "Hey @jacob34 , can you explain the difference between training separately vs concatenating all labels at the end? I'm having a bit of trouble understanding how labels can be concatenated in the first place. \n\nex. train on week 0-50, predict on week 51; \ntrain 1-51, predict on week 52;\nthen average the results of the two predictions?",
      "votes": null
    },
    {
      "id": "1773341",
      "postDate": "05/01/2022 04:22:22",
      "content": "<p><a href=\"https://www.kaggle.com/kennyxie\" target=\"_blank\">@kennyxie</a> - you're describing ensembling models - which can also help.  </p>\n<p>But this is the difference between separate and concatenating:</p>\n<h2>Separate</h2>\n<p><strong>Fold #1</strong></p>\n<ol>\n<li><p>using data from weeks 0-50, generate candidates and features for recommendations for week 51</p></li>\n<li><p>the labels are whether the purchase happened in week 51</p></li>\n<li><p>train ranker on those features/labels (model #1)  </p></li>\n<li><p>using data from weeks 0-51, generate candidates and features for recommendations for week 52</p></li>\n<li><p>use model #1 to rank your candidates </p></li>\n<li><p>the labels are whether the purchase happened in week 52</p></li>\n<li><p>Compare you predictions to the labels to get your evaluation score</p></li>\n</ol>\n<p><strong>Fold #2</strong></p>\n<ol>\n<li><p>using data from weeks 0-51, generate candidates and features for recommendations for week 52</p></li>\n<li><p>the labels are whether the purchase happened in week 52</p></li>\n<li><p>train ranker on those features/labels (model #2)</p></li>\n<li><p>using data from weeks 0-52, generate candidates and features for recommendations for week 53</p></li>\n<li><p>use model #2 to rank your candidates </p></li>\n<li><p>the labels are whether the purchase happened in week 53</p></li>\n<li><p>Compare you predictions to the labels to get your evaluation score</p></li>\n</ol>\n<h2>Together</h2>\n<ol>\n<li><p>using data from weeks 0-50, generate candidates and features for recommendations for week 51</p></li>\n<li><p>the labels are whether the purchase happened in week 51</p></li>\n<li><p>using data from weeks 0-51, generate candidates and features for recommendations for week 52</p></li>\n<li><p>the labels are whether the purchase happened in week 52</p></li>\n<li><p>Concatenate the candidates w/ features together, and the labels together</p></li>\n<li><p>Train ranker on the combined features/labels</p></li>\n<li><p>using data from weeks 0-52, generate candidates and features for recommendations for week 53</p></li>\n<li><p>use model #2 to rank your candidates </p></li>\n<li><p>the labels are whether the purchase happened in week 53</p></li>\n<li><p>Compare you predictions to the labels to get your evaluation score</p></li>\n</ol>",
      "rawMarkdown": "kennyxie - you're describing ensembling models - which can also help.  \n\nBut this is the difference between separate and concatenating:\n\n## Separate\n**Fold #1**\n1. using data from weeks 0-50, generate candidates and features for recommendations for week 51\n2. the labels are whether the purchase happened in week 51\n3. train ranker on those features/labels (model #1)  \n\n5. using data from weeks 0-51, generate candidates and features for recommendations for week 52\n6. use model #1 to rank your candidates \n7. the labels are whether the purchase happened in week 52\n8. Compare you predictions to the labels to get your evaluation score\n\n**Fold #2**\n1. using data from weeks 0-51, generate candidates and features for recommendations for week 52\n2. the labels are whether the purchase happened in week 52\n3. train ranker on those features/labels (model #2)\n\n4. using data from weeks 0-52, generate candidates and features for recommendations for week 53\n5. use model #2 to rank your candidates \n6. the labels are whether the purchase happened in week 53\n7. Compare you predictions to the labels to get your evaluation score\n\n## Together\n1. using data from weeks 0-50, generate candidates and features for recommendations for week 51\n2. the labels are whether the purchase happened in week 51\n3. using data from weeks 0-51, generate candidates and features for recommendations for week 52\n4. the labels are whether the purchase happened in week 52\n5. Concatenate the candidates w/ features together, and the labels together\n6. Train ranker on the combined features/labels\n\n7. using data from weeks 0-52, generate candidates and features for recommendations for week 53\n8. use model #2 to rank your candidates \n9. the labels are whether the purchase happened in week 53\n10. Compare you predictions to the labels to get your evaluation score",
      "votes": null
    },
    {
      "id": "1774762",
      "postDate": "05/02/2022 12:32:57",
      "content": "<p>Would streaming data be more beneficial here? Or it still scrapes the roof?</p>",
      "rawMarkdown": "Would streaming data be more beneficial here? Or it still scrapes the roof?",
      "votes": null
    },
    {
      "id": "1778635",
      "postDate": "05/05/2022 14:33:51",
      "content": "<p>Thanks a lot my friend, it's very useful!</p>",
      "rawMarkdown": "Thanks a lot my friend, it's very useful!",
      "votes": null
    },
    {
      "id": "1778645",
      "postDate": "05/05/2022 14:46:55",
      "content": "<p><a href=\"https://www.kaggle.com/tiger0\" target=\"_blank\">@tiger0</a> - thanks for the feedback!</p>",
      "rawMarkdown": "tiger0 - thanks for the feedback!",
      "votes": null
    },
    {
      "id": "1779691",
      "postDate": "05/06/2022 18:00:25",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> I've been running my CVs seperately since Im still a bit confused on how one concatenates the candidates with features + labels together from seperate weeks. I think the main thing I'm getting from your explanation is that you generate/train LGBMranker with multiple sets of candidates instead of just 1?</p>",
      "rawMarkdown": "jacob34 I've been running my CVs seperately since Im still a bit confused on how one concatenates the candidates with features + labels together from seperate weeks. I think the main thing I'm getting from your explanation is that you generate/train LGBMranker with multiple sets of candidates instead of just 1?",
      "votes": null
    },
    {
      "id": "1779844",
      "postDate": "05/06/2022 22:02:37",
      "content": "<p>Yes, that's right.</p>",
      "rawMarkdown": "Yes, that's right.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1742679,
      "author_name": "sunyuri",
      "author_url": "",
      "post_date": "04/02/2022 07:05:23",
      "content": "<p>Thanks     👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1743019,
      "author_name": "samsonlo",
      "author_url": "",
      "post_date": "04/02/2022 14:09:57",
      "content": "<p>Infomative! Thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1744358,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "04/03/2022 22:57:44",
      "content": "<p>Good summary. In my case my train features are generated before 2020-04-01 and move week by week until 2020-09-22. I noticed that the ranking model doesn't need much data. For example to train the ranking model I use only labels generated from last 5 months. I think there are some gains to be made when using more data for training but I would need to invest to computation. It is too easy to just add RAM much harder to optimize the code and one's thinking.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1744988,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/04/2022 13:38:12",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>!</p>\n<p>So it sounds like I got it a bit wrong - whereas I've put each train/validation split as a separate \"fold\", to be trained/evaluated separately, I should instead be using them to generate training data with labels, then concatenate all the data/labels together, and train on everything together.</p>\n<p>Is that right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1745075,
          "author_name": "paweljankiewicz",
          "author_url": "",
          "post_date": "04/04/2022 15:17:06",
          "content": "<p>Yes, I think you can concatenate more than 1 week for training your ranking model. I, kind of assumed, this from the beginning. The model for submission I train without validation on all the weeks up to 09/22 and it improves by about 0.0005 (compared to training set without the last week). I haven't done exhaustive testing. It would be kind of interesting to see the effect on the leaderboard/validation of adding more and more weeks in reverse order.</p>\n<p>My explanation for this is that when you are training your ranking model all the strategies you are using are kind of relative: bestsellers are bestsellers whichever period you are analysing. It may be that H&amp;M used different recommendation strategies throughout the provided data which may totally skew the strategies. Which would be my biggest concern in this competition - to what extent we are creating something new vs recreating H&amp;M recommendation strategy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1745089,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/04/2022 15:29:11",
          "content": "<p>Makes sense, thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1748140,
          "author_name": "sirius81",
          "author_url": "",
          "post_date": "04/07/2022 10:48:23",
          "content": "<p>Shocked by 5 months… I use only last 2 weeks transaction to build the training samples (separately). Have you compared the cv scores of different amount of training samples? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1748416,
          "author_name": "paweljankiewicz",
          "author_url": "",
          "post_date": "04/07/2022 15:11:19",
          "content": "<p><a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> I checked and used 6 months and it worsened the result, I checked 4 months and it worsened the results too. So I guess I'm using the right amount of transactions :D. And this is only for the ranking model, for many features I'm actually using all the transactions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1748467,
          "author_name": "igorkf",
          "author_url": "",
          "post_date": "04/07/2022 16:06:07",
          "content": "<p>Thanks for the information, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> .<br>\nWhen you say \"my train features are generated before 2020-04-01 and move week by week until 2020-09-22\", what kind of features you mean?</p>\n<p>For instance:<br>\nFor week 100, your features would be: \"bought_article_next_week? (0 or 1)\" indicating if this article was bought by this customer in the week 101.<br>\nAnd the same goes on for the week pairs (X: 101, y: 102), (X: 102, y: 103), … (…)</p>\n<p>Is this kind of feature you mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1772379,
          "author_name": "kennyxie",
          "author_url": "",
          "post_date": "04/30/2022 07:33:30",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> , can you explain the difference between training separately vs concatenating all labels at the end? I'm having a bit of trouble understanding how labels can be concatenated in the first place. </p>\n<p>ex. train on week 0-50, predict on week 51; <br>\ntrain 1-51, predict on week 52;<br>\nthen average the results of the two predictions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1773341,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/01/2022 04:22:22",
          "content": "<p><a href=\"https://www.kaggle.com/kennyxie\" target=\"_blank\">@kennyxie</a> - you're describing ensembling models - which can also help.  </p>\n<p>But this is the difference between separate and concatenating:</p>\n<h2>Separate</h2>\n<p><strong>Fold #1</strong></p>\n<ol>\n<li><p>using data from weeks 0-50, generate candidates and features for recommendations for week 51</p></li>\n<li><p>the labels are whether the purchase happened in week 51</p></li>\n<li><p>train ranker on those features/labels (model #1)  </p></li>\n<li><p>using data from weeks 0-51, generate candidates and features for recommendations for week 52</p></li>\n<li><p>use model #1 to rank your candidates </p></li>\n<li><p>the labels are whether the purchase happened in week 52</p></li>\n<li><p>Compare you predictions to the labels to get your evaluation score</p></li>\n</ol>\n<p><strong>Fold #2</strong></p>\n<ol>\n<li><p>using data from weeks 0-51, generate candidates and features for recommendations for week 52</p></li>\n<li><p>the labels are whether the purchase happened in week 52</p></li>\n<li><p>train ranker on those features/labels (model #2)</p></li>\n<li><p>using data from weeks 0-52, generate candidates and features for recommendations for week 53</p></li>\n<li><p>use model #2 to rank your candidates </p></li>\n<li><p>the labels are whether the purchase happened in week 53</p></li>\n<li><p>Compare you predictions to the labels to get your evaluation score</p></li>\n</ol>\n<h2>Together</h2>\n<ol>\n<li><p>using data from weeks 0-50, generate candidates and features for recommendations for week 51</p></li>\n<li><p>the labels are whether the purchase happened in week 51</p></li>\n<li><p>using data from weeks 0-51, generate candidates and features for recommendations for week 52</p></li>\n<li><p>the labels are whether the purchase happened in week 52</p></li>\n<li><p>Concatenate the candidates w/ features together, and the labels together</p></li>\n<li><p>Train ranker on the combined features/labels</p></li>\n<li><p>using data from weeks 0-52, generate candidates and features for recommendations for week 53</p></li>\n<li><p>use model #2 to rank your candidates </p></li>\n<li><p>the labels are whether the purchase happened in week 53</p></li>\n<li><p>Compare you predictions to the labels to get your evaluation score</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1774762,
          "author_name": "dmytrovalantsevich",
          "author_url": "",
          "post_date": "05/02/2022 12:32:57",
          "content": "<p>Would streaming data be more beneficial here? Or it still scrapes the roof?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1779691,
          "author_name": "kennyxie",
          "author_url": "",
          "post_date": "05/06/2022 18:00:25",
          "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> I've been running my CVs seperately since Im still a bit confused on how one concatenates the candidates with features + labels together from seperate weeks. I think the main thing I'm getting from your explanation is that you generate/train LGBMranker with multiple sets of candidates instead of just 1?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1779844,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/06/2022 22:02:37",
          "content": "<p>Yes, that's right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1744780,
      "author_name": "ammarabulfeilat",
      "author_url": "",
      "post_date": "04/04/2022 09:57:58",
      "content": "<p>Thanks alot , Helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1745601,
      "author_name": "niekvanderzwaag",
      "author_url": "",
      "post_date": "04/05/2022 04:54:11",
      "content": "<p>Interesting stuff, thanks for sharing 🙌</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1746251,
      "author_name": "oriori4244",
      "author_url": "",
      "post_date": "04/05/2022 15:46:14",
      "content": "<p>Thank you for your work. <br>\nI have a question. What is train label?   Does train label mean wether  a customer bought article or not?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1746295,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/05/2022 16:33:28",
          "content": "<p>Yes, label means whether the customer bought the article or not. That's what our model is trying to predict.</p>\n<p>Train labels are the labels for the training data, that we're training the model on.<br>\nValid labels are the labels for the validation data, which we'll use to evaluate how well our trained model's predictions were.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1746323,
          "author_name": "oriori4244",
          "author_url": "",
          "post_date": "04/05/2022 16:57:40",
          "content": "<p>Thank you for your answer. <br>\nI have one more question. How did you generate negative labels?<br>\nI think that labels are 0 or 1. For example, Fold 0, positive  train label  for the customer  is the article the customer bought in week 103. What is negative label for the customer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1746424,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/05/2022 18:50:07",
          "content": "<p>It's any of the candidates you selected (from step #1), that the customer did not buy in week 103.<br>\nSee <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">this post</a> and the comments there for a more detailed explanation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1746687,
          "author_name": "oriori4244",
          "author_url": "",
          "post_date": "04/06/2022 02:41:01",
          "content": "<p>Thank you so much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1747165,
      "author_name": "nyuzuto",
      "author_url": "",
      "post_date": "04/06/2022 12:42:07",
      "content": "<p>thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1747630,
      "author_name": "gazu468",
      "author_url": "",
      "post_date": "04/06/2022 20:19:04",
      "content": "<p>Interesting stuff. helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1778635,
      "author_name": "tiger0",
      "author_url": "",
      "post_date": "05/05/2022 14:33:51",
      "content": "<p>Thanks a lot my friend, it's very useful!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1778645,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/05/2022 14:46:55",
          "content": "<p><a href=\"https://www.kaggle.com/tiger0\" target=\"_blank\">@tiger0</a> - thanks for the feedback!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1742269": "## TLDR; \n## Using ML modeling for ranking? \n## Adjust CV to avoid overfitting.\n\nIn [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308919), @cdeotte shares how to setup local cv for this competition.\n\nThis works well for many public notebooks, but if move on to using ML models for ranking candidates, you need a little more.\n\nLet me explain: \n\nIn [this popular post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288), @paweljankiewicz discusses the two stages often taken for this kind of problem:  \n1. First you find promising “candidates” – customer/item pairs that are likely to have happened.  \n2. Then you train a model to “rank”, for each customer, how likely each of the candidates are.    \n\nFor submission, you take the 12 highest-ranked candidates for each customer.  \n\nAs he mentions, candidate selection doesn’t necessarily need to involve ML. Many public high-scoring notebooks go a step further – not only do they select candidates based on hard-coded methods, but the ranking is also done without ML, by sorting on certain columns.\n\nLet’s look at the data like this:   \n(instead of using dates, we’ll use weeks, where week #105 = LB data)\n[![handm-data-lb.jpg](https://i.postimg.cc/9QMyKZZc/handm-data-lb.jpg)](https://postimg.cc/QH2BWBfz)\n\n\nThis is what @cdeotte’s cv strategy looks like:\n\n[![handm-cdeotte-cv.jpg](https://i.postimg.cc/mkSBYcgn/handm-cdeotte-cv.jpg)](https://postimg.cc/V05x1kM9)\n\nFor each fold, the “train” data is where we get the features, and the “validation” week is where we get the ground truth – the labels.\n\nWe decide on a strategy for selecting candidates and ranking them using features from the train data, and evaluate how well they do by looking at the labels from the validation week.  \n\n**But if we’re designing our algorithm based on the actual labels, why aren't we overfitting?**\n\nThe answer is that hard-coded algorithms aren’t very flexible, so we don’t have to worry much about overfitting. (in ML terms, you could say they have high bias and low variance)\n\nOnce we move to a ML model like LGBMRanker, overfitting is going to be a big problem.\n\nFor modeling, we need a CV that looks like this:\n[![handm-my-cv.jpg](https://i.postimg.cc/cLQF1Q6w/handm-my-cv.jpg)](https://postimg.cc/v4Zr3xrZ)\n\nThis way we can tune the model on the training features/labels, and do evaluation using the validation data/labels.\n\nCould be this is simple for many people, but it took me a little while to wrap my head around it.\nHope it’s helpful for you!",
    "1742679": "Thanks     👍",
    "1743019": "Infomative! Thank you",
    "1744358": "Good summary. In my case my train features are generated before 2020-04-01 and move week by week until 2020-09-22. I noticed that the ranking model doesn't need much data. For example to train the ranking model I use only labels generated from last 5 months. I think there are some gains to be made when using more data for training but I would need to invest to computation. It is too easy to just add RAM much harder to optimize the code and one's thinking.",
    "1744780": "Thanks alot , Helpful",
    "1744988": "Thanks, @paweljankiewicz!\n\nSo it sounds like I got it a bit wrong - whereas I've put each train/validation split as a separate \"fold\", to be trained/evaluated separately, I should instead be using them to generate training data with labels, then concatenate all the data/labels together, and train on everything together.\n\nIs that right?",
    "1745075": "Yes, I think you can concatenate more than 1 week for training your ranking model. I, kind of assumed, this from the beginning. The model for submission I train without validation on all the weeks up to 09/22 and it improves by about 0.0005 (compared to training set without the last week). I haven't done exhaustive testing. It would be kind of interesting to see the effect on the leaderboard/validation of adding more and more weeks in reverse order.\n\nMy explanation for this is that when you are training your ranking model all the strategies you are using are kind of relative: bestsellers are bestsellers whichever period you are analysing. It may be that H&M used different recommendation strategies throughout the provided data which may totally skew the strategies. Which would be my biggest concern in this competition - to what extent we are creating something new vs recreating H&M recommendation strategy?",
    "1745089": "Makes sense, thanks!",
    "1745601": "Interesting stuff, thanks for sharing 🙌",
    "1746251": "Thank you for your work. \nI have a question. What is train label?   Does train label mean wether  a customer bought article or not?",
    "1746295": "Yes, label means whether the customer bought the article or not. That's what our model is trying to predict.\n\nTrain labels are the labels for the training data, that we're training the model on.\nValid labels are the labels for the validation data, which we'll use to evaluate how well our trained model's predictions were.",
    "1746323": "Thank you for your answer. \nI have one more question. How did you generate negative labels?\nI think that labels are 0 or 1. For example, Fold 0, positive  train label  for the customer  is the article the customer bought in week 103. What is negative label for the customer.",
    "1746424": "It's any of the candidates you selected (from step #1), that the customer did not buy in week 103.\nSee [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288) and the comments there for a more detailed explanation.",
    "1746687": "Thank you so much",
    "1747165": "thank you!",
    "1747630": "Interesting stuff. helpful",
    "1748140": "Shocked by 5 months... I use only last 2 weeks transaction to build the training samples (separately). Have you compared the cv scores of different amount of training samples?",
    "1748416": "sirius81 I checked and used 6 months and it worsened the result, I checked 4 months and it worsened the results too. So I guess I'm using the right amount of transactions :D. And this is only for the ranking model, for many features I'm actually using all the transactions.",
    "1748467": "Thanks for the information, @paweljankiewicz .\nWhen you say \"my train features are generated before 2020-04-01 and move week by week until 2020-09-22\", what kind of features you mean?\n\nFor instance:\nFor week 100, your features would be: \"bought_article_next_week? (0 or 1)\" indicating if this article was bought by this customer in the week 101.\nAnd the same goes on for the week pairs (X: 101, y: 102), (X: 102, y: 103), ... (...)\n\nIs this kind of feature you mean?",
    "1772379": "Hey @jacob34 , can you explain the difference between training separately vs concatenating all labels at the end? I'm having a bit of trouble understanding how labels can be concatenated in the first place. \n\nex. train on week 0-50, predict on week 51; \ntrain 1-51, predict on week 52;\nthen average the results of the two predictions?",
    "1773341": "kennyxie - you're describing ensembling models - which can also help.  \n\nBut this is the difference between separate and concatenating:\n\n## Separate\n**Fold #1**\n1. using data from weeks 0-50, generate candidates and features for recommendations for week 51\n2. the labels are whether the purchase happened in week 51\n3. train ranker on those features/labels (model #1)  \n\n5. using data from weeks 0-51, generate candidates and features for recommendations for week 52\n6. use model #1 to rank your candidates \n7. the labels are whether the purchase happened in week 52\n8. Compare you predictions to the labels to get your evaluation score\n\n**Fold #2**\n1. using data from weeks 0-51, generate candidates and features for recommendations for week 52\n2. the labels are whether the purchase happened in week 52\n3. train ranker on those features/labels (model #2)\n\n4. using data from weeks 0-52, generate candidates and features for recommendations for week 53\n5. use model #2 to rank your candidates \n6. the labels are whether the purchase happened in week 53\n7. Compare you predictions to the labels to get your evaluation score\n\n## Together\n1. using data from weeks 0-50, generate candidates and features for recommendations for week 51\n2. the labels are whether the purchase happened in week 51\n3. using data from weeks 0-51, generate candidates and features for recommendations for week 52\n4. the labels are whether the purchase happened in week 52\n5. Concatenate the candidates w/ features together, and the labels together\n6. Train ranker on the combined features/labels\n\n7. using data from weeks 0-52, generate candidates and features for recommendations for week 53\n8. use model #2 to rank your candidates \n9. the labels are whether the purchase happened in week 53\n10. Compare you predictions to the labels to get your evaluation score",
    "1774762": "Would streaming data be more beneficial here? Or it still scrapes the roof?",
    "1778635": "Thanks a lot my friend, it's very useful!",
    "1778645": "tiger0 - thanks for the feedback!",
    "1779691": "jacob34 I've been running my CVs seperately since Im still a bit confused on how one concatenates the candidates with features + labels together from seperate weeks. I think the main thing I'm getting from your explanation is that you generate/train LGBMranker with multiple sets of candidates instead of just 1?",
    "1779844": "Yes, that's right."
  },
  "source": "meta"
}