{
  "id": 317374,
  "title": "LGBMRanker - evaluation metric not improving",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374",
  "author_name": "Clear n' Simple",
  "post_date": "2022-04-06T18:52:13.758000",
  "votes": 43,
  "comment_count": 53,
  "views": 0,
  "content": "<p><strong>This thread has (happily) evolved from the topic title to a more general sharing by <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> and <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> (both very high-ranking on the LB) on their methods in this competition - see the sub-threads below.</strong></p>\n<p><strong>Update per the original post:\nFigured it out - there was silent leakage going on, when I was merging the labels in.\nSee <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081\" target=\"_blank\">this post</a> for details.</strong></p>\n<p>I'm using the default training loss (lambdarank), and have tried both the default evaluation metric (NDCG) as well as MAP, but the score deteriorates as training progresses, even on the training set.</p>\n<p>How is it possible that the loss is not correlated with the evaluation metric?<br>\nWhat does one do in such a case?</p>",
  "messages": [
    {
      "id": 1747563,
      "postDate": "2022-04-06T18:52:13.760Z",
      "content": "<p><strong>This thread has (happily) evolved from the topic title to a more general sharing by <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> and <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> (both very high-ranking on the LB) on their methods in this competition - see the sub-threads below.</strong></p>\n<p><strong>Update per the original post:\nFigured it out - there was silent leakage going on, when I was merging the labels in.\nSee <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081\" target=\"_blank\">this post</a> for details.</strong></p>\n<p>I'm using the default training loss (lambdarank), and have tried both the default evaluation metric (NDCG) as well as MAP, but the score deteriorates as training progresses, even on the training set.</p>\n<p>How is it possible that the loss is not correlated with the evaluation metric?<br>\nWhat does one do in such a case?</p>",
      "rawMarkdown": "**This thread has (happily) evolved from the topic title to a more general sharing by @lihaorocky and @paweljankiewicz (both very high-ranking on the LB) on their methods in this competition - see the sub-threads below.**\n\n**Update per the original post:\nFigured it out - there was silent leakage going on, when I was merging the labels in.\nSee [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081) for details.**\n\n\nI'm using the default training loss (lambdarank), and have tried both the default evaluation metric (NDCG) as well as MAP, but the score deteriorates as training progresses, even on the training set.\n\nHow is it possible that the loss is not correlated with the evaluation metric?\nWhat does one do in such a case?",
      "votes": 43
    },
    {
      "id": 1747770,
      "postDate": "2022-04-07T02:17:49.330Z",
      "content": "<p>I would suggest you doing the follow steps:<br>\n1: Say you have a column in your feature dataframe called \"target_week\" to represent the following 7-days window. Then sort the dataframe <br>\n<code>df_feat_train.sort_values(['target_week', 'customer_id'], ignore_index=True, inplace=True)</code><br>\n2: After splitting the training and validation dataset, try: <br>\n<code>train_baskets = df_feat_train.loc[train_idx, :].groupby(['target_week', 'customer_id'])['article_id'].count().values</code> <br>\nto get the train baskets. And do the same for the validation dataset.<br>\n3: Lower the learning rate, in my case, the training loss is early-stopped in sort of 150 training rounds with learning_rate 0.03 when using only one week for training and one week for validation.<br>\nHope it helps.</p>",
      "rawMarkdown": "I would suggest you doing the follow steps:\n1: Say you have a column in your feature dataframe called \"target_week\" to represent the following 7-days window. Then sort the dataframe \n`df_feat_train.sort_values(['target_week', 'customer_id'], ignore_index=True, inplace=True)`\n2: After splitting the training and validation dataset, try: \n`train_baskets = df_feat_train.loc[train_idx, :].groupby(['target_week', 'customer_id'])['article_id'].count().values` \nto get the train baskets. And do the same for the validation dataset.\n3: Lower the learning rate, in my case, the training loss is early-stopped in sort of 150 training rounds with learning_rate 0.03 when using only one week for training and one week for validation.\nHope it helps.",
      "votes": 16,
      "replies": [
        {
          "id": 1748395,
          "postDate": "2022-04-07T14:49:14.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> - thank you for replying, and for sharing.</p>\n<p>I'm only using one target week as of now, so I don't need to group by target_week, but I'm doing exactly that.<br>\nMy code is:</p>\n<pre><code>features_df = features_df.sort_values(\"customer_id\").reset_index(drop=True)\ngroups = list(features_df.groupby(\"customer_id\")[\"article_id\"].count())\n</code></pre>\n<p>I can't get the metric to improve even for 2 training rounds.</p>\n<p>Maybe I just need to move away from prototyping and adding more features and data…</p>",
          "rawMarkdown": "@lihaorocky - thank you for replying, and for sharing.\n\nI'm only using one target week as of now, so I don't need to group by target_week, but I'm doing exactly that.\nMy code is:\n\n```\nfeatures_df = features_df.sort_values(\"customer_id\").reset_index(drop=True)\ngroups = list(features_df.groupby(\"customer_id\")[\"article_id\"].count())\n```\n\nI can't get the metric to improve even for 2 training rounds.\n\nMaybe I just need to move away from prototyping and adding more features and data..."
        },
        {
          "id": 1749561,
          "postDate": "2022-04-08T17:45:52.487Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . Thanks for sharing<br>\nAs I understood <code>target_week</code> is a binary relevance (1 or 0) to denote if customer bought or not an article, right?<br>\nIf this is the case, there's any reason to sort values by <code>target_week</code>? This way seems you are putting 0's first in the whole dataset. </p>",
          "rawMarkdown": "Hi, @lihaorocky . Thanks for sharing\nAs I understood `target_week` is a binary relevance (1 or 0) to denote if customer bought or not an article, right?\nIf this is the case, there's any reason to sort values by `target_week`? This way seems you are putting 0's first in the whole dataset. "
        },
        {
          "id": 1749961,
          "postDate": "2022-04-09T07:12:12.373Z",
          "content": "<p>Not really. I use target_week column to represent the target week period. For example, on an observation date 20200915, your target week will be 20200916-20200922, I give it a label 0. And for an observation date 20200908, the target week will be 20200909-20200915, I give it a label 1 etc. This way you will get lots of target_weeks 0 1 2 3 4 5….I use the data before observation date to create candidate samples and sample features, while using the data during target week to create sample labels(bought or not). I sort the samples based on target_week and customer_id because I have a lot of weeks' samples to be input to LightGBMRanker. So it need to be sorted to create baskets.</p>",
          "rawMarkdown": "Not really. I use target_week column to represent the target week period. For example, on an observation date 20200915, your target week will be 20200916-20200922, I give it a label 0. And for an observation date 20200908, the target week will be 20200909-20200915, I give it a label 1 etc. This way you will get lots of target_weeks 0 1 2 3 4 5....I use the data before observation date to create candidate samples and sample features, while using the data during target week to create sample labels(bought or not). I sort the samples based on target_week and customer_id because I have a lot of weeks' samples to be input to LightGBMRanker. So it need to be sorted to create baskets.",
          "votes": 10
        },
        {
          "id": 1750082,
          "postDate": "2022-04-09T10:03:30.087Z",
          "content": "<p>Hi， <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . Thanks for sharing<br>\nI‘d like to ask ， \"target_weeks\" this column just for building pos/neg samples ?</p>\n<p>And in the sorting model, I only use some existing features on the item side and user side. I'm a little confused about the sorting feature and don't know how to construct it, can you give me some advice?</p>",
          "rawMarkdown": "Hi， @lihaorocky . Thanks for sharing\nI‘d like to ask ， \"target_weeks\" this column just for building pos/neg samples ?\n\nAnd in the sorting model, I only use some existing features on the item side and user side. I'm a little confused about the sorting feature and don't know how to construct it, can you give me some advice?"
        },
        {
          "id": 1750110,
          "postDate": "2022-04-09T10:33:36.200Z",
          "content": "<p>For your first question, yes, and it also eases filtering training and validation samples based on it.<br>\nWell, there are a lot of ways to construct features. Besides the features of customer and article alone, you can measure the similarity between customers and articles. I will not talk too much about this but I would say it will give huge performance improvement.</p>",
          "rawMarkdown": "For your first question, yes, and it also eases filtering training and validation samples based on it.\nWell, there are a lot of ways to construct features. Besides the features of customer and article alone, you can measure the similarity between customers and articles. I will not talk too much about this but I would say it will give huge performance improvement.",
          "votes": 6
        },
        {
          "id": 1750165,
          "postDate": "2022-04-09T11:39:55.357Z",
          "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>. Seems that LGBMRanker is limited to a relevance of 31, so your <code>target_weeks</code> lies in the range [0, max(31, n_weeks)]? </p>",
          "rawMarkdown": "Thanks for the answer @lihaorocky. Seems that LGBMRanker is limited to a relevance of 31, so your `target_weeks` lies in the range [0, max(31, n_weeks)]? "
        },
        {
          "id": 1750183,
          "postDate": "2022-04-09T11:58:04.067Z",
          "content": "<p>Limited by the Ram size, my current score is using only 5 weeks for training. I'm trying adding more weeks.</p>",
          "rawMarkdown": "Limited by the Ram size, my current score is using only 5 weeks for training. I'm trying adding more weeks.",
          "votes": 6
        },
        {
          "id": 1750196,
          "postDate": "2022-04-09T12:11:13.747Z",
          "content": "<p>I got it. Thanks for the infos <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . You helped me a lot.</p>",
          "rawMarkdown": "I got it. Thanks for the infos @lihaorocky . You helped me a lot."
        },
        {
          "id": 1750922,
          "postDate": "2022-04-10T08:17:31.963Z",
          "content": "<p>Thanks for your sharing, it really helps me a lot! But I have a question, do you use the popularity features of articles for LGBMRanker training? Because I find the importance of this feature is much higher than other features, such as the similarity between customers and articles, and it reduces the score of my final prediction, I wonder if you meet this situation, thank you! <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> </p>",
          "rawMarkdown": "Thanks for your sharing, it really helps me a lot! But I have a question, do you use the popularity features of articles for LGBMRanker training? Because I find the importance of this feature is much higher than other features, such as the similarity between customers and articles, and it reduces the score of my final prediction, I wonder if you meet this situation, thank you! @lihaorocky ",
          "votes": 2
        },
        {
          "id": 1750950,
          "postDate": "2022-04-10T08:44:56.767Z",
          "content": "<p>Yes, I did use. I think the way you generated this kind of features (popularity of articles, similarity of articles and customers) are very important. And it's also depending on the strategy you created candidate pairs. In my case, the similarity features are always the most important ones.</p>",
          "rawMarkdown": "Yes, I did use. I think the way you generated this kind of features (popularity of articles, similarity of articles and customers) are very important. And it's also depending on the strategy you created candidate pairs. In my case, the similarity features are always the most important ones.",
          "votes": 4
        },
        {
          "id": 1751089,
          "postDate": "2022-04-10T11:04:11.340Z",
          "content": "<p>Thanks for your reply, it's really helpful!</p>",
          "rawMarkdown": "Thanks for your reply, it's really helpful!"
        },
        {
          "id": 1754264,
          "postDate": "2022-04-13T14:02:16.813Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nI'm bit confused on how to split train/validation sets after generating the dataset.<br>\nTo create train baskets you do:</p>\n<pre><code>train_baskets = df_feat_train.loc[train_idx, :].groupby(['target_week', 'customer_id'])['article_id'].count().values\n</code></pre>\n<p>In this case, <code>train_idx</code> should be created in a group split way?<br>\nI mean, all rows from a customer must be in train <strong>OR</strong> in val or you're doing just a holdout (e.g. 80/20) split?</p>",
          "rawMarkdown": "Hi @lihaorocky \nI'm bit confused on how to split train/validation sets after generating the dataset.\nTo create train baskets you do:\n```\ntrain_baskets = df_feat_train.loc[train_idx, :].groupby(['target_week', 'customer_id'])['article_id'].count().values\n```\nIn this case, `train_idx` should be created in a group split way?\nI mean, all rows from a customer must be in train **OR** in val or you're doing just a holdout (e.g. 80/20) split?"
        },
        {
          "id": 1754289,
          "postDate": "2022-04-13T14:14:56.937Z",
          "content": "<p>I split the train/validation sets based on time. Specifically I split them based on the value of target_week(0 as validation set and &gt;1 as train set).</p>",
          "rawMarkdown": "I split the train/validation sets based on time. Specifically I split them based on the value of target_week(0 as validation set and >1 as train set).",
          "votes": 3
        },
        {
          "id": 1754302,
          "postDate": "2022-04-13T14:27:21.540Z",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . This really makes sense…I was overcomplicating.</p>",
          "rawMarkdown": "Thank you very much @lihaorocky . This really makes sense...I was overcomplicating."
        },
        {
          "id": 1754358,
          "postDate": "2022-04-13T15:32:25.403Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>  ,thanks for your reply.<br>\nPlease allow me to ask two more questions</p>\n<ol>\n<li><p>During the recall phase, did you use multiple recalls?</p></li>\n<li><p>If each user recalls nearly 50 articles, then more than 1 million users will have more than 50 million recall data, and then I am sorting these recall data. But I am a little confused now, how to solve such a large amount of data?</p></li>\n</ol>",
          "rawMarkdown": "Hi @lihaorocky  ,thanks for your reply.\nPlease allow me to ask two more questions\n\n1. During the recall phase, did you use multiple recalls?\n\n2. If each user recalls nearly 50 articles, then more than 1 million users will have more than 50 million recall data, and then I am sorting these recall data. But I am a little confused now, how to solve such a large amount of data?",
          "votes": 1
        },
        {
          "id": 1754366,
          "postDate": "2022-04-13T15:47:18.590Z",
          "content": "<p>1: Yes, I use multiple recalls<br>\n2: As others posted in some discussions, if you are using LGBMRanker you can drop the users' recalls if they haven't any purchase in next 7 days or your recalls don't contains any real purchase that will really happen, which will hugely reduce the number of candidates to be input to your ranking model.</p>\n<p>As to the inference stage, you can't use the two above mentioned tricks to reduce candidates to be predicted. And yes, there will be \"more than 50 million recall data\", but you can generate features and predict for these recalls batch by batch(for example in each batch you generate features and predict for 1 million recalls), which actually doesn't cost too much time. In my case, predicting for 160 million recalls cost overall 2 hours including generating features and model prediction.</p>",
          "rawMarkdown": "1: Yes, I use multiple recalls\n2: As others posted in some discussions, if you are using LGBMRanker you can drop the users' recalls if they haven't any purchase in next 7 days or your recalls don't contains any real purchase that will really happen, which will hugely reduce the number of candidates to be input to your ranking model.\n \nAs to the inference stage, you can't use the two above mentioned tricks to reduce candidates to be predicted. And yes, there will be \"more than 50 million recall data\", but you can generate features and predict for these recalls batch by batch(for example in each batch you generate features and predict for 1 million recalls), which actually doesn't cost too much time. In my case, predicting for 160 million recalls cost overall 2 hours including generating features and model prediction.",
          "votes": 6
        },
        {
          "id": 1754547,
          "postDate": "2022-04-13T18:52:16.883Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> could be as well describing my process. I have 1000 batches for the submission and the prediction is the slowest part - it takes about 5 hours. You did well on the candidate generation - I think I have close to 1 billion candidates.</p>",
          "rawMarkdown": "@lihaorocky could be as well describing my process. I have 1000 batches for the submission and the prediction is the slowest part - it takes about 5 hours. You did well on the candidate generation - I think I have close to 1 billion candidates.",
          "votes": 5
        },
        {
          "id": 1754567,
          "postDate": "2022-04-13T19:02:55.870Z",
          "content": "<p>If you don't mind:</p>\n<ul>\n<li>How much time for generating your candidates?</li>\n</ul>",
          "rawMarkdown": "If you don't mind:\n- How much time for generating your candidates?",
          "votes": 2
        },
        {
          "id": 1754580,
          "postDate": "2022-04-13T19:10:54.183Z",
          "content": "<p>I'm using Rust for all the feature generation. Overall it takes somewhere around 2.5 hours on 32 threads. I think I could reduce it to 1 hour if I parallelized more calculations.</p>",
          "rawMarkdown": "I'm using Rust for all the feature generation. Overall it takes somewhere around 2.5 hours on 32 threads. I think I could reduce it to 1 hour if I parallelized more calculations.",
          "votes": 2
        },
        {
          "id": 1756464,
          "postDate": "2022-04-15T14:38:38.147Z",
          "content": "<p><a href=\"https://www.kaggle.com/panadam\" target=\"_blank\">@panadam</a> <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nI have a question about your interesting discussion.<br>\nWhat does 'recalls' exactly mean?</p>",
          "rawMarkdown": "@panadam @lihaorocky \nI have a question about your interesting discussion.\nWhat does 'recalls' exactly mean?",
          "votes": 5
        },
        {
          "id": 1756490,
          "postDate": "2022-04-15T15:03:25.190Z",
          "content": "<p>In recomendations system, normally it consists two stages: recall and ranking. <br>\nRecalls here means finding some candidate articles one customers will have interations with. In this stage, you want to contains as many as articles one customers will bought(in this competition) using as few as candidates. You can develop multiple strategies to creating candidate articles which are mentioned in a lot of posts.<br>\nIn second stage, you want to rank those candidates based on the possibility being purchased buy customers.</p>",
          "rawMarkdown": "In recomendations system, normally it consists two stages: recall and ranking. \nRecalls here means finding some candidate articles one customers will have interations with. In this stage, you want to contains as many as articles one customers will bought(in this competition) using as few as candidates. You can develop multiple strategies to creating candidate articles which are mentioned in a lot of posts.\nIn second stage, you want to rank those candidates based on the possibility being purchased buy customers.",
          "votes": 6
        },
        {
          "id": 1756512,
          "postDate": "2022-04-15T15:23:34.630Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nThanks for reply!<br>\nThen, I have interpreted 'multiple recalls' means that you used multiple strategies to make candidates in the SAME exam. How do you compare possibilities from different strategies? I guess, do you use standarization?</p>",
          "rawMarkdown": "@lihaorocky \nThanks for reply!\nThen, I have interpreted 'multiple recalls' means that you used multiple strategies to make candidates in the SAME exam. How do you compare possibilities from different strategies? I guess, do you use standarization?",
          "votes": 3
        },
        {
          "id": 1756517,
          "postDate": "2022-04-15T15:27:11.353Z",
          "content": "<p>Actually, I give the  choice mostly to the ranking model. He (or she) is much smarter than me.😂</p>",
          "rawMarkdown": "Actually, I give the  choice mostly to the ranking model. He (or she) is much smarter than me.😂",
          "votes": 3
        },
        {
          "id": 1756539,
          "postDate": "2022-04-15T15:37:31.610Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nAI overtakes humans :D<br>\nDo you suggest to give ['possibility', 'strategy', …(other features)] to ranking model?</p>",
          "rawMarkdown": "@lihaorocky \nAI overtakes humans :D\nDo you suggest to give ['possibility', 'strategy', ...(other features)] to ranking model?",
          "votes": 3
        },
        {
          "id": 1756549,
          "postDate": "2022-04-15T15:49:04.177Z",
          "content": "<p>Yes, for example when you develop the strategy based on article's hotness. The article hottness can be measured easily using some method like purchased count in previous week. I use some thing like this. But from the performance comparison using or not using those measurement feature in ranking model, the improvement of adding those features is quite small. That's why I say, mostly I leave it to the ranking model. </p>",
          "rawMarkdown": "Yes, for example when you develop the strategy based on article's hotness. The article hottness can be measured easily using some method like purchased count in previous week. I use some thing like this. But from the performance comparison using or not using those measurement feature in ranking model, the improvement of adding those features is quite small. That's why I say, mostly I leave it to the ranking model. ",
          "votes": 4
        },
        {
          "id": 1756568,
          "postDate": "2022-04-15T16:11:19.640Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> thanks for your shares. <br>\nThere's a way to know if the strategy can generate good candidates?</p>",
          "rawMarkdown": "Hi @lihaorocky thanks for your shares. \nThere's a way to know if the strategy can generate good candidates?",
          "votes": 1
        },
        {
          "id": 1756588,
          "postDate": "2022-04-15T16:31:02.050Z",
          "content": "<p>Yes, for example on an observation date 20200915, from one of your candidate generating strategy you create 1 million pairs(customer_id, article_id, target_week), and from the next 7-days transations window 20200916~20200922, you can label those pairs with 1(purchase happend) and 0(purchase not happen). Say totally you have 50 thousands actually happend pairs. Then the \"matching ratio\" will be 50000 / 1000000 = 0.05. Clearly if keeping the number of total candidate pairs fixed, the bigger the numebr actully happend purchase pairs the better. So you can use this above mentioned ratio to measure how good you strategy is.</p>",
          "rawMarkdown": "Yes, for example on an observation date 20200915, from one of your candidate generating strategy you create 1 million pairs(customer_id, article_id, target_week), and from the next 7-days transations window 20200916~20200922, you can label those pairs with 1(purchase happend) and 0(purchase not happen). Say totally you have 50 thousands actually happend pairs. Then the \"matching ratio\" will be 50000 / 1000000 = 0.05. Clearly if keeping the number of total candidate pairs fixed, the bigger the numebr actully happend purchase pairs the better. So you can use this above mentioned ratio to measure how good you strategy is.",
          "votes": 5
        },
        {
          "id": 1756645,
          "postDate": "2022-04-15T18:04:45.777Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nThank you very much for your explanation. This totally makes sense</p>",
          "rawMarkdown": "Hi @lihaorocky \nThank you very much for your explanation. This totally makes sense",
          "votes": 1
        },
        {
          "id": 1756828,
          "postDate": "2022-04-15T23:15:51.023Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nAfter all, my understanding of ranking model improved very well thanks to your clear explanation.<br>\nLet's do our best together!</p>",
          "rawMarkdown": "@lihaorocky \nAfter all, my understanding of ranking model improved very well thanks to your clear explanation.\nLet's do our best together!",
          "votes": 3
        },
        {
          "id": 1759075,
          "postDate": "2022-04-18T10:20:49.903Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Hi, roughly what percentage of actual purchases can you generate in the recall stage, with 160 million candidates? </p>",
          "rawMarkdown": "@lihaorocky Hi, roughly what percentage of actual purchases can you generate in the recall stage, with 160 million candidates? ",
          "votes": 1
        },
        {
          "id": 1759099,
          "postDate": "2022-04-18T10:41:58.393Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1759101,
          "postDate": "2022-04-18T10:42:50.867Z",
          "content": "<p>Well, the 160 million candidates is in the prediction stage, which mean the actually purchases is not known until 20200929. If you mean what the percentage ratio of candidates in training stage is, different recall strategies acts very differently, which range from 0.3~0.001.</p>",
          "rawMarkdown": "Well, the 160 million candidates is in the prediction stage, which mean the actually purchases is not known until 20200929. If you mean what the percentage ratio of candidates in training stage is, different recall strategies acts very differently, which range from 0.3~0.001.",
          "votes": 4
        },
        {
          "id": 1759122,
          "postDate": "2022-04-18T11:04:24.343Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> I think he meant the recall of real sales from the candidates altogether. So if there was 300k transactions in the last week how many you can generate using your candidates. I must check this too.</p>",
          "rawMarkdown": "@lihaorocky I think he meant the recall of real sales from the candidates altogether. So if there was 300k transactions in the last week how many you can generate using your candidates. I must check this too.",
          "votes": 2
        },
        {
          "id": 1759145,
          "postDate": "2022-04-18T11:23:58.870Z",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> Yes that's what I mean. <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> 30% is striking, I haven't been woking on this that much but saw a post saying that he could hardly exceed 7.5%.  <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314458\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314458</a></p>",
          "rawMarkdown": "@paweljankiewicz Yes that's what I mean. @lihaorocky 30% is striking, I haven't been woking on this that much but saw a post saying that he could hardly exceed 7.5%.  https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314458",
          "votes": 1
        },
        {
          "id": 1759148,
          "postDate": "2022-04-18T11:27:12.257Z",
          "content": "<p>I just realized you mean recall. In my case, it's about 25%.</p>",
          "rawMarkdown": "I just realized you mean recall. In my case, it's about 25%.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1755632,
      "postDate": "2022-04-14T19:29:27.600Z",
      "content": "<p>Hi. I'm thinking how much your CV boosted after ranking the set of candidates with LGBMRanker or other model. <br>\nCurrently I'm getting just +0.0015 boost, so I think something is wrong.</p>\n<p>Could you guys <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> share this if you don't mind?</p>",
      "rawMarkdown": "Hi. I'm thinking how much your CV boosted after ranking the set of candidates with LGBMRanker or other model. \nCurrently I'm getting just +0.0015 boost, so I think something is wrong.\n\nCould you guys @paweljankiewicz @lihaorocky @jacob34 share this if you don't mind?",
      "votes": 1,
      "replies": [
        {
          "id": 1755675,
          "postDate": "2022-04-14T21:04:12.523Z",
          "content": "<p>There is no CV before the ranking. Ranking model is a must. I have more than 1000 candidate items for each customer. How else would I be able to use them without the ranker? Single strategies to generate candidates are pretty useless on their own. You need to have multiple candidate generation strategies like I explained in the other thread.</p>",
          "rawMarkdown": "There is no CV before the ranking. Ranking model is a must. I have more than 1000 candidate items for each customer. How else would I be able to use them without the ranker? Single strategies to generate candidates are pretty useless on their own. You need to have multiple candidate generation strategies like I explained in the other thread.",
          "votes": 2
        },
        {
          "id": 1755683,
          "postDate": "2022-04-14T21:23:21.233Z",
          "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>. <br>\nI understood that we need to have many candidates generation strategies, but there's some points not clear to me yet.</p>\n<ol>\n<li><p>How do I know if a candidate strategy is good? I suppose that good strategy would generate more 1's (item was bought by the customer in the target week) than a strategy with bad candidates. For instance, good strategy A generated 1000 candidates (85% is label=0, 15% is label=1), meanwhile strategy B generated 1000 candidates (95% is label=0, 5% is label=1)</p></li>\n<li><p>How should we assign the label? Here's a pseudo code:</p></li>\n</ol>\n<pre><code>for each row in candidates:\n    if this candidate item was bought by this customer in the next week:\n       label = 1\n    else:\n       label = 0\n</code></pre>\n<p>Is this right?</p>\n<p>Sorry for the (possible) obvious questions…this is the first time we're using a ranking model.</p>",
          "rawMarkdown": "Thanks for the answer @paweljankiewicz. \nI understood that we need to have many candidates generation strategies, but there's some points not clear to me yet.\n1. How do I know if a candidate strategy is good? I suppose that good strategy would generate more 1's (item was bought by the customer in the target week) than a strategy with bad candidates. For instance, good strategy A generated 1000 candidates (85% is label=0, 15% is label=1), meanwhile strategy B generated 1000 candidates (95% is label=0, 5% is label=1)\n\n2. How should we assign the label? Here's a pseudo code:\n```\nfor each row in candidates:\n    if this candidate item was bought by this customer in the next week:\n       label = 1\n    else:\n       label = 0\n```\nIs this right?\n\nSorry for the (possible) obvious questions...this is the first time we're using a ranking model.\n",
          "votes": 1
        },
        {
          "id": 1756598,
          "postDate": "2022-04-15T16:41:47.153Z",
          "content": "<p>No worries. Ranking model are quite helpful in many situations.</p>\n<p>The pseudo code to generate the training data would be something like this for 2 strategies</p>\n<pre><code>all_data = []\nfor query_id, customer in enumerate(customers):\n      candidates_a = generate_candidates_a(customer)\n      candidates_b = generate_candidates_b(customer)\n      all_candidates = union(candidates_a, candidates_b)   # you generate a list of all unique candidates\n      customer_df = pd.DataFrame({\n            query_id: query_id,\n            customer_id: [customer_id]*len(all_candidates),\n            article_id: all_candidates,\n            bought_next_week: all_candidates.isin(customer_sales_next_week.get(customer, [])), # this is your label\n            article_in_strategy_a: all_candidates.isin(candidates_a),\n            article_in_strategy_b: all_candidates.isin(candidates_b),\n      })\n      all_data.append(customer_df)\n</code></pre>\n<p>A dataframe like this for each customer is a query for a ranking model. It is like a separate part of your data that is evaluated by the ranking model. Unlike some other methods like regression and classification you want to properly rank items withing customer query. </p>\n<p>So going back to your first question - the ranking model sees the articles, whether they were bought or not and from which strategies they came so it can assess how good the candidate is based on this information. So you want to have precise strategies without adding too many negative observations.</p>",
          "rawMarkdown": "No worries. Ranking model are quite helpful in many situations.\n\nThe pseudo code to generate the training data would be something like this for 2 strategies\n\n```python\nall_data = []\nfor query_id, customer in enumerate(customers):\n      candidates_a = generate_candidates_a(customer)\n      candidates_b = generate_candidates_b(customer)\n      all_candidates = union(candidates_a, candidates_b)   # you generate a list of all unique candidates\n      customer_df = pd.DataFrame({\n            query_id: query_id,\n            customer_id: [customer_id]*len(all_candidates),\n            article_id: all_candidates,\n            bought_next_week: all_candidates.isin(customer_sales_next_week.get(customer, [])), # this is your label\n            article_in_strategy_a: all_candidates.isin(candidates_a),\n            article_in_strategy_b: all_candidates.isin(candidates_b),\n      })\n      all_data.append(customer_df)\n```\n\nA dataframe like this for each customer is a query for a ranking model. It is like a separate part of your data that is evaluated by the ranking model. Unlike some other methods like regression and classification you want to properly rank items withing customer query. \n\nSo going back to your first question - the ranking model sees the articles, whether they were bought or not and from which strategies they came so it can assess how good the candidate is based on this information. So you want to have precise strategies without adding too many negative observations.",
          "votes": 9
        },
        {
          "id": 1756641,
          "postDate": "2022-04-15T17:54:01.930Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nThank you very much for your help. I would say that is not a pseudo code, but a 100%-real code 😄<br>\nNow it's more clear how to combine multiple candidates. You rock</p>",
          "rawMarkdown": "Hi @paweljankiewicz \nThank you very much for your help. I would say that is not a pseudo code, but a 100%-real code 😄\nNow it's more clear how to combine multiple candidates. You rock",
          "votes": 1
        },
        {
          "id": 1756698,
          "postDate": "2022-04-15T18:45:31.240Z",
          "content": "<p>In the first line: <code>for query_id, customer in enumerate(customers)</code>,  <br>\n<code>customers</code> is a list of customers from current week or customers from the next week? </p>",
          "rawMarkdown": "In the first line: `for query_id, customer in enumerate(customers)`,  \n`customers` is a list of customers from current week or customers from the next week? ",
          "votes": 1
        },
        {
          "id": 1756735,
          "postDate": "2022-04-15T19:37:35.723Z",
          "content": "<p>Current week. Labels are from next week.</p>",
          "rawMarkdown": "Current week. Labels are from next week.",
          "votes": 3
        },
        {
          "id": 1756750,
          "postDate": "2022-04-15T19:54:36.323Z",
          "content": "<p>Thanks for your help <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> </p>",
          "rawMarkdown": "Thanks for your help @paweljankiewicz "
        },
        {
          "id": 1756875,
          "postDate": "2022-04-16T02:35:23.467Z",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> Thanks for your sharing! But I have a little confusion about your pseudo code. In my previous idea, I union the articles from all strategies, and add other features (like article popularity, and similarity between articles and customers) to these articles, but in your code, you just point out what strategies these articles came from. Do you think it is necessary to add other features?</p>",
          "rawMarkdown": "@paweljankiewicz Thanks for your sharing! But I have a little confusion about your pseudo code. In my previous idea, I union the articles from all strategies, and add other features (like article popularity, and similarity between articles and customers) to these articles, but in your code, you just point out what strategies these articles came from. Do you think it is necessary to add other features?",
          "votes": 2
        },
        {
          "id": 1757146,
          "postDate": "2022-04-16T10:06:20.433Z",
          "content": "<p><a href=\"https://www.kaggle.com/rock139\" target=\"_blank\">@rock139</a> Of course you need to add other features as well as I explained in the other thread. This was only a question about how to combine the strategies into training data.</p>",
          "rawMarkdown": "@rock139 Of course you need to add other features as well as I explained in the other thread. This was only a question about how to combine the strategies into training data.",
          "votes": 2
        },
        {
          "id": 1757159,
          "postDate": "2022-04-16T10:29:50.023Z",
          "content": "<p>Thanks for your help!</p>",
          "rawMarkdown": "Thanks for your help!"
        }
      ]
    },
    {
      "id": 1759500,
      "postDate": "2022-04-18T17:14:40.880Z",
      "content": "<p>Update:</p>\n<p>I've found that after deteriorating for a few steps, the evaluation metric does start improving (although it never gets as good as what it was after the first step).<br>\nAt that point, it does appear to correlate with the competition metric, and across cv/LB.<br>\nSo I'm working with that.</p>\n<p>I still don't know why it's acting this way originally, but I'm assuming it's some artifact of slight difference between loss function and evaluation metrics that is showing because of the small size of my training set and weak features.</p>",
      "rawMarkdown": "Update:\n\nI've found that after deteriorating for a few steps, the evaluation metric does start improving (although it never gets as good as what it was after the first step).\nAt that point, it does appear to correlate with the competition metric, and across cv/LB.\nSo I'm working with that.\n\nI still don't know why it's acting this way originally, but I'm assuming it's some artifact of slight difference between loss function and evaluation metrics that is showing because of the small size of my training set and weak features.\n",
      "votes": 2
    },
    {
      "id": 1777718,
      "postDate": "2022-05-04T18:30:10.893Z",
      "content": "<p>Update:</p>\n<p>Figured it out - there was silent leakage going on, when I was merging the labels in.<br>\nSee <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081\" target=\"_blank\">this post</a> for details.</p>",
      "rawMarkdown": "Update:\n\nFigured it out - there was silent leakage going on, when I was merging the labels in.\nSee [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081) for details."
    },
    {
      "id": 1760807,
      "postDate": "2022-04-19T14:22:19.347Z",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>  well, thanks for your reply! I guess I have only one question. I didn't recall any candidates, I just use week 1 as label and data in week 5, 4, 3, 2 as candidates for week 1, which means I label data in week 1 as positive and data in week 5, 4, 3, 2 as negative after remove samples which customer_id not appear in week 1.   So, do you think this is why my candidates is bad? <br>\nThank you again!</p>",
      "rawMarkdown": "@lihaorocky  well, thanks for your reply! I guess I have only one question. I didn't recall any candidates, I just use week 1 as label and data in week 5, 4, 3, 2 as candidates for week 1, which means I label data in week 1 as positive and data in week 5, 4, 3, 2 as negative after remove samples which customer_id not appear in week 1.   So, do you think this is why my candidates is bad? \nThank you again!"
    },
    {
      "id": 1760730,
      "postDate": "2022-04-19T13:43:53.160Z",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>  Hi, cound you explain more detailed about how do you create the training and testing set?  From my understanding, for example, there are 5, 4, 3, 2, 1, 0 weeks, and you just need take the week 1 data as label, to label data before week 1 (such as week 5, 4, 3, 2), this can be  used as train data. And take the week 0 data to label data before week 0, this can be used as valid data. Is that right?</p>\n<p>If that's right, then I come with same problem, my NDCG deteriorates as training progresses…</p>",
      "rawMarkdown": "@lihaorocky  Hi, cound you explain more detailed about how do you create the training and testing set?  From my understanding, for example, there are 5, 4, 3, 2, 1, 0 weeks, and you just need take the week 1 data as label, to label data before week 1 (such as week 5, 4, 3, 2), this can be  used as train data. And take the week 0 data to label data before week 0, this can be used as valid data. Is that right?\n\nIf that's right, then I come with same problem, my NDCG deteriorates as training progresses...",
      "replies": [
        {
          "id": 1760772,
          "postDate": "2022-04-19T14:04:51.783Z",
          "content": "<p>For your question, yes, I think your description was correct.<br>\nAnd as to the problem you mentioned, I didn't meet this problem. But if you did all other things right, I guess it happend maybe because of the quality of the candidates you generated, the amount of training dataset is not big enough, or the features you created are too few or not good enough. Just be patient and good luck.</p>",
          "rawMarkdown": "For your question, yes, I think your description was correct.\nAnd as to the problem you mentioned, I didn't meet this problem. But if you did all other things right, I guess it happend maybe because of the quality of the candidates you generated, the amount of training dataset is not big enough, or the features you created are too few or not good enough. Just be patient and good luck.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1758929,
      "postDate": "2022-04-18T07:12:05.363Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1747770,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2022-04-07T02:17:49.330000",
      "content": "<p>I would suggest you doing the follow steps:<br>\n1: Say you have a column in your feature dataframe called \"target_week\" to represent the following 7-days window. Then sort the dataframe <br>\n<code>df_feat_train.sort_values(['target_week', 'customer_id'], ignore_index=True, inplace=True)</code><br>\n2: After splitting the training and validation dataset, try: <br>\n<code>train_baskets = df_feat_train.loc[train_idx, :].groupby(['target_week', 'customer_id'])['article_id'].count().values</code> <br>\nto get the train baskets. And do the same for the validation dataset.<br>\n3: Lower the learning rate, in my case, the training loss is early-stopped in sort of 150 training rounds with learning_rate 0.03 when using only one week for training and one week for validation.<br>\nHope it helps.</p>",
      "votes": 16,
      "replies": [
        {
          "id": 1748395,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-07T14:49:14.867000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> - thank you for replying, and for sharing.</p>\n<p>I'm only using one target week as of now, so I don't need to group by target_week, but I'm doing exactly that.<br>\nMy code is:</p>\n<pre><code>features_df = features_df.sort_values(\"customer_id\").reset_index(drop=True)\ngroups = list(features_df.groupby(\"customer_id\")[\"article_id\"].count())\n</code></pre>\n<p>I can't get the metric to improve even for 2 training rounds.</p>\n<p>Maybe I just need to move away from prototyping and adding more features and data…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1749561,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-08T17:45:52.487000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . Thanks for sharing<br>\nAs I understood <code>target_week</code> is a binary relevance (1 or 0) to denote if customer bought or not an article, right?<br>\nIf this is the case, there's any reason to sort values by <code>target_week</code>? This way seems you are putting 0's first in the whole dataset. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1749961,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-09T07:12:12.373000",
          "content": "<p>Not really. I use target_week column to represent the target week period. For example, on an observation date 20200915, your target week will be 20200916-20200922, I give it a label 0. And for an observation date 20200908, the target week will be 20200909-20200915, I give it a label 1 etc. This way you will get lots of target_weeks 0 1 2 3 4 5….I use the data before observation date to create candidate samples and sample features, while using the data during target week to create sample labels(bought or not). I sort the samples based on target_week and customer_id because I have a lot of weeks' samples to be input to LightGBMRanker. So it need to be sorted to create baskets.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1750082,
          "author_name": "Pan Adam",
          "author_url": "",
          "post_date": "2022-04-09T10:03:30.087000",
          "content": "<p>Hi， <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . Thanks for sharing<br>\nI‘d like to ask ， \"target_weeks\" this column just for building pos/neg samples ?</p>\n<p>And in the sorting model, I only use some existing features on the item side and user side. I'm a little confused about the sorting feature and don't know how to construct it, can you give me some advice?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1750110,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-09T10:33:36.200000",
          "content": "<p>For your first question, yes, and it also eases filtering training and validation samples based on it.<br>\nWell, there are a lot of ways to construct features. Besides the features of customer and article alone, you can measure the similarity between customers and articles. I will not talk too much about this but I would say it will give huge performance improvement.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1750165,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-09T11:39:55.357000",
          "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>. Seems that LGBMRanker is limited to a relevance of 31, so your <code>target_weeks</code> lies in the range [0, max(31, n_weeks)]? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1750183,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-09T11:58:04.067000",
          "content": "<p>Limited by the Ram size, my current score is using only 5 weeks for training. I'm trying adding more weeks.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1750196,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-09T12:11:13.747000",
          "content": "<p>I got it. Thanks for the infos <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . You helped me a lot.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1750922,
          "author_name": "Rock",
          "author_url": "",
          "post_date": "2022-04-10T08:17:31.963000",
          "content": "<p>Thanks for your sharing, it really helps me a lot! But I have a question, do you use the popularity features of articles for LGBMRanker training? Because I find the importance of this feature is much higher than other features, such as the similarity between customers and articles, and it reduces the score of my final prediction, I wonder if you meet this situation, thank you! <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1750950,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-10T08:44:56.767000",
          "content": "<p>Yes, I did use. I think the way you generated this kind of features (popularity of articles, similarity of articles and customers) are very important. And it's also depending on the strategy you created candidate pairs. In my case, the similarity features are always the most important ones.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1751089,
          "author_name": "Rock",
          "author_url": "",
          "post_date": "2022-04-10T11:04:11.340000",
          "content": "<p>Thanks for your reply, it's really helpful!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1754264,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-13T14:02:16.813000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nI'm bit confused on how to split train/validation sets after generating the dataset.<br>\nTo create train baskets you do:</p>\n<pre><code>train_baskets = df_feat_train.loc[train_idx, :].groupby(['target_week', 'customer_id'])['article_id'].count().values\n</code></pre>\n<p>In this case, <code>train_idx</code> should be created in a group split way?<br>\nI mean, all rows from a customer must be in train <strong>OR</strong> in val or you're doing just a holdout (e.g. 80/20) split?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1754289,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-13T14:14:56.937000",
          "content": "<p>I split the train/validation sets based on time. Specifically I split them based on the value of target_week(0 as validation set and &gt;1 as train set).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1754302,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-13T14:27:21.540000",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . This really makes sense…I was overcomplicating.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1754358,
          "author_name": "Pan Adam",
          "author_url": "",
          "post_date": "2022-04-13T15:32:25.403000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>  ,thanks for your reply.<br>\nPlease allow me to ask two more questions</p>\n<ol>\n<li><p>During the recall phase, did you use multiple recalls?</p></li>\n<li><p>If each user recalls nearly 50 articles, then more than 1 million users will have more than 50 million recall data, and then I am sorting these recall data. But I am a little confused now, how to solve such a large amount of data?</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1754366,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-13T15:47:18.590000",
          "content": "<p>1: Yes, I use multiple recalls<br>\n2: As others posted in some discussions, if you are using LGBMRanker you can drop the users' recalls if they haven't any purchase in next 7 days or your recalls don't contains any real purchase that will really happen, which will hugely reduce the number of candidates to be input to your ranking model.</p>\n<p>As to the inference stage, you can't use the two above mentioned tricks to reduce candidates to be predicted. And yes, there will be \"more than 50 million recall data\", but you can generate features and predict for these recalls batch by batch(for example in each batch you generate features and predict for 1 million recalls), which actually doesn't cost too much time. In my case, predicting for 160 million recalls cost overall 2 hours including generating features and model prediction.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1754547,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-13T18:52:16.883000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> could be as well describing my process. I have 1000 batches for the submission and the prediction is the slowest part - it takes about 5 hours. You did well on the candidate generation - I think I have close to 1 billion candidates.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1754567,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-13T19:02:55.870000",
          "content": "<p>If you don't mind:</p>\n<ul>\n<li>How much time for generating your candidates?</li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1754580,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-13T19:10:54.183000",
          "content": "<p>I'm using Rust for all the feature generation. Overall it takes somewhere around 2.5 hours on 32 threads. I think I could reduce it to 1 hour if I parallelized more calculations.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1756464,
          "author_name": "Y-Haneji",
          "author_url": "",
          "post_date": "2022-04-15T14:38:38.147000",
          "content": "<p><a href=\"https://www.kaggle.com/panadam\" target=\"_blank\">@panadam</a> <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nI have a question about your interesting discussion.<br>\nWhat does 'recalls' exactly mean?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1756490,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-15T15:03:25.190000",
          "content": "<p>In recomendations system, normally it consists two stages: recall and ranking. <br>\nRecalls here means finding some candidate articles one customers will have interations with. In this stage, you want to contains as many as articles one customers will bought(in this competition) using as few as candidates. You can develop multiple strategies to creating candidate articles which are mentioned in a lot of posts.<br>\nIn second stage, you want to rank those candidates based on the possibility being purchased buy customers.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1756512,
          "author_name": "Y-Haneji",
          "author_url": "",
          "post_date": "2022-04-15T15:23:34.630000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nThanks for reply!<br>\nThen, I have interpreted 'multiple recalls' means that you used multiple strategies to make candidates in the SAME exam. How do you compare possibilities from different strategies? I guess, do you use standarization?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1756517,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-15T15:27:11.353000",
          "content": "<p>Actually, I give the  choice mostly to the ranking model. He (or she) is much smarter than me.😂</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1756539,
          "author_name": "Y-Haneji",
          "author_url": "",
          "post_date": "2022-04-15T15:37:31.610000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nAI overtakes humans :D<br>\nDo you suggest to give ['possibility', 'strategy', …(other features)] to ranking model?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1756549,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-15T15:49:04.177000",
          "content": "<p>Yes, for example when you develop the strategy based on article's hotness. The article hottness can be measured easily using some method like purchased count in previous week. I use some thing like this. But from the performance comparison using or not using those measurement feature in ranking model, the improvement of adding those features is quite small. That's why I say, mostly I leave it to the ranking model. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1756568,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-15T16:11:19.640000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> thanks for your shares. <br>\nThere's a way to know if the strategy can generate good candidates?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1756588,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-15T16:31:02.050000",
          "content": "<p>Yes, for example on an observation date 20200915, from one of your candidate generating strategy you create 1 million pairs(customer_id, article_id, target_week), and from the next 7-days transations window 20200916~20200922, you can label those pairs with 1(purchase happend) and 0(purchase not happen). Say totally you have 50 thousands actually happend pairs. Then the \"matching ratio\" will be 50000 / 1000000 = 0.05. Clearly if keeping the number of total candidate pairs fixed, the bigger the numebr actully happend purchase pairs the better. So you can use this above mentioned ratio to measure how good you strategy is.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1756645,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-15T18:04:45.777000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nThank you very much for your explanation. This totally makes sense</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1756828,
          "author_name": "Y-Haneji",
          "author_url": "",
          "post_date": "2022-04-15T23:15:51.023000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <br>\nAfter all, my understanding of ranking model improved very well thanks to your clear explanation.<br>\nLet's do our best together!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1759075,
          "author_name": "Homoalways",
          "author_url": "",
          "post_date": "2022-04-18T10:20:49.903000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Hi, roughly what percentage of actual purchases can you generate in the recall stage, with 160 million candidates? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1759099,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-04-18T10:41:58.393000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1759101,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-18T10:42:50.867000",
          "content": "<p>Well, the 160 million candidates is in the prediction stage, which mean the actually purchases is not known until 20200929. If you mean what the percentage ratio of candidates in training stage is, different recall strategies acts very differently, which range from 0.3~0.001.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1759122,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-18T11:04:24.343000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> I think he meant the recall of real sales from the candidates altogether. So if there was 300k transactions in the last week how many you can generate using your candidates. I must check this too.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1759145,
          "author_name": "Homoalways",
          "author_url": "",
          "post_date": "2022-04-18T11:23:58.870000",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> Yes that's what I mean. <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> 30% is striking, I haven't been woking on this that much but saw a post saying that he could hardly exceed 7.5%.  <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314458\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/314458</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1759148,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-18T11:27:12.257000",
          "content": "<p>I just realized you mean recall. In my case, it's about 25%.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1755632,
      "author_name": "Igor Kuivjogi Fernandes",
      "author_url": "",
      "post_date": "2022-04-14T19:29:27.600000",
      "content": "<p>Hi. I'm thinking how much your CV boosted after ranking the set of candidates with LGBMRanker or other model. <br>\nCurrently I'm getting just +0.0015 boost, so I think something is wrong.</p>\n<p>Could you guys <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> share this if you don't mind?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1755675,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-14T21:04:12.523000",
          "content": "<p>There is no CV before the ranking. Ranking model is a must. I have more than 1000 candidate items for each customer. How else would I be able to use them without the ranker? Single strategies to generate candidates are pretty useless on their own. You need to have multiple candidate generation strategies like I explained in the other thread.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1755683,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-14T21:23:21.233000",
          "content": "<p>Thanks for the answer <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>. <br>\nI understood that we need to have many candidates generation strategies, but there's some points not clear to me yet.</p>\n<ol>\n<li><p>How do I know if a candidate strategy is good? I suppose that good strategy would generate more 1's (item was bought by the customer in the target week) than a strategy with bad candidates. For instance, good strategy A generated 1000 candidates (85% is label=0, 15% is label=1), meanwhile strategy B generated 1000 candidates (95% is label=0, 5% is label=1)</p></li>\n<li><p>How should we assign the label? Here's a pseudo code:</p></li>\n</ol>\n<pre><code>for each row in candidates:\n    if this candidate item was bought by this customer in the next week:\n       label = 1\n    else:\n       label = 0\n</code></pre>\n<p>Is this right?</p>\n<p>Sorry for the (possible) obvious questions…this is the first time we're using a ranking model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1756598,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-15T16:41:47.153000",
          "content": "<p>No worries. Ranking model are quite helpful in many situations.</p>\n<p>The pseudo code to generate the training data would be something like this for 2 strategies</p>\n<pre><code>all_data = []\nfor query_id, customer in enumerate(customers):\n      candidates_a = generate_candidates_a(customer)\n      candidates_b = generate_candidates_b(customer)\n      all_candidates = union(candidates_a, candidates_b)   # you generate a list of all unique candidates\n      customer_df = pd.DataFrame({\n            query_id: query_id,\n            customer_id: [customer_id]*len(all_candidates),\n            article_id: all_candidates,\n            bought_next_week: all_candidates.isin(customer_sales_next_week.get(customer, [])), # this is your label\n            article_in_strategy_a: all_candidates.isin(candidates_a),\n            article_in_strategy_b: all_candidates.isin(candidates_b),\n      })\n      all_data.append(customer_df)\n</code></pre>\n<p>A dataframe like this for each customer is a query for a ranking model. It is like a separate part of your data that is evaluated by the ranking model. Unlike some other methods like regression and classification you want to properly rank items withing customer query. </p>\n<p>So going back to your first question - the ranking model sees the articles, whether they were bought or not and from which strategies they came so it can assess how good the candidate is based on this information. So you want to have precise strategies without adding too many negative observations.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1756641,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-15T17:54:01.930000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nThank you very much for your help. I would say that is not a pseudo code, but a 100%-real code 😄<br>\nNow it's more clear how to combine multiple candidates. You rock</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1756698,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-15T18:45:31.240000",
          "content": "<p>In the first line: <code>for query_id, customer in enumerate(customers)</code>,  <br>\n<code>customers</code> is a list of customers from current week or customers from the next week? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1756735,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-15T19:37:35.723000",
          "content": "<p>Current week. Labels are from next week.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1756750,
          "author_name": "Igor Kuivjogi Fernandes",
          "author_url": "",
          "post_date": "2022-04-15T19:54:36.323000",
          "content": "<p>Thanks for your help <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1756875,
          "author_name": "Rock",
          "author_url": "",
          "post_date": "2022-04-16T02:35:23.467000",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> Thanks for your sharing! But I have a little confusion about your pseudo code. In my previous idea, I union the articles from all strategies, and add other features (like article popularity, and similarity between articles and customers) to these articles, but in your code, you just point out what strategies these articles came from. Do you think it is necessary to add other features?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1757146,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-16T10:06:20.433000",
          "content": "<p><a href=\"https://www.kaggle.com/rock139\" target=\"_blank\">@rock139</a> Of course you need to add other features as well as I explained in the other thread. This was only a question about how to combine the strategies into training data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1757159,
          "author_name": "Rock",
          "author_url": "",
          "post_date": "2022-04-16T10:29:50.023000",
          "content": "<p>Thanks for your help!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1759500,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-04-18T17:14:40.880000",
      "content": "<p>Update:</p>\n<p>I've found that after deteriorating for a few steps, the evaluation metric does start improving (although it never gets as good as what it was after the first step).<br>\nAt that point, it does appear to correlate with the competition metric, and across cv/LB.<br>\nSo I'm working with that.</p>\n<p>I still don't know why it's acting this way originally, but I'm assuming it's some artifact of slight difference between loss function and evaluation metrics that is showing because of the small size of my training set and weak features.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1777718,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-05-04T18:30:10.893000",
      "content": "<p>Update:</p>\n<p>Figured it out - there was silent leakage going on, when I was merging the labels in.<br>\nSee <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081\" target=\"_blank\">this post</a> for details.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1760807,
      "author_name": "Zhisheng Jiang",
      "author_url": "",
      "post_date": "2022-04-19T14:22:19.347000",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>  well, thanks for your reply! I guess I have only one question. I didn't recall any candidates, I just use week 1 as label and data in week 5, 4, 3, 2 as candidates for week 1, which means I label data in week 1 as positive and data in week 5, 4, 3, 2 as negative after remove samples which customer_id not appear in week 1.   So, do you think this is why my candidates is bad? <br>\nThank you again!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1760730,
      "author_name": "Zhisheng Jiang",
      "author_url": "",
      "post_date": "2022-04-19T13:43:53.160000",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>  Hi, cound you explain more detailed about how do you create the training and testing set?  From my understanding, for example, there are 5, 4, 3, 2, 1, 0 weeks, and you just need take the week 1 data as label, to label data before week 1 (such as week 5, 4, 3, 2), this can be  used as train data. And take the week 0 data to label data before week 0, this can be used as valid data. Is that right?</p>\n<p>If that's right, then I come with same problem, my NDCG deteriorates as training progresses…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1760772,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2022-04-19T14:04:51.783000",
          "content": "<p>For your question, yes, I think your description was correct.<br>\nAnd as to the problem you mentioned, I didn't meet this problem. But if you did all other things right, I guess it happend maybe because of the quality of the candidates you generated, the amount of training dataset is not big enough, or the features you created are too few or not good enough. Just be patient and good luck.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1758929,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-18T07:12:05.363000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1747563": "**This thread has (happily) evolved from the topic title to a more general sharing by @lihaorocky and @paweljankiewicz (both very high-ranking on the LB) on their methods in this competition - see the sub-threads below.**\n\n**Update per the original post:\nFigured it out - there was silent leakage going on, when I was merging the labels in.\nSee [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081) for details.**\n\n\nI'm using the default training loss (lambdarank), and have tried both the default evaluation metric (NDCG) as well as MAP, but the score deteriorates as training progresses, even on the training set.\n\nHow is it possible that the loss is not correlated with the evaluation metric?\nWhat does one do in such a case?",
    "1747770": "I would suggest you doing the follow steps:\n1: Say you have a column in your feature dataframe called \"target_week\" to represent the following 7-days window. Then sort the dataframe \n`df_feat_train.sort_values(['target_week', 'customer_id'], ignore_index=True, inplace=True)`\n2: After splitting the training and validation dataset, try: \n`train_baskets = df_feat_train.loc[train_idx, :].groupby(['target_week', 'customer_id'])['article_id'].count().values` \nto get the train baskets. And do the same for the validation dataset.\n3: Lower the learning rate, in my case, the training loss is early-stopped in sort of 150 training rounds with learning_rate 0.03 when using only one week for training and one week for validation.\nHope it helps.",
    "1755632": "Hi. I'm thinking how much your CV boosted after ranking the set of candidates with LGBMRanker or other model. \nCurrently I'm getting just +0.0015 boost, so I think something is wrong.\n\nCould you guys @paweljankiewicz @lihaorocky @jacob34 share this if you don't mind?",
    "1759500": "Update:\n\nI've found that after deteriorating for a few steps, the evaluation metric does start improving (although it never gets as good as what it was after the first step).\nAt that point, it does appear to correlate with the competition metric, and across cv/LB.\nSo I'm working with that.\n\nI still don't know why it's acting this way originally, but I'm assuming it's some artifact of slight difference between loss function and evaluation metrics that is showing because of the small size of my training set and weak features.\n",
    "1777718": "Update:\n\nFigured it out - there was silent leakage going on, when I was merging the labels in.\nSee [this post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081) for details.",
    "1760807": "@lihaorocky  well, thanks for your reply! I guess I have only one question. I didn't recall any candidates, I just use week 1 as label and data in week 5, 4, 3, 2 as candidates for week 1, which means I label data in week 1 as positive and data in week 5, 4, 3, 2 as negative after remove samples which customer_id not appear in week 1.   So, do you think this is why my candidates is bad? \nThank you again!",
    "1760730": "@lihaorocky  Hi, cound you explain more detailed about how do you create the training and testing set?  From my understanding, for example, there are 5, 4, 3, 2, 1, 0 weeks, and you just need take the week 1 data as label, to label data before week 1 (such as week 5, 4, 3, 2), this can be  used as train data. And take the week 0 data to label data before week 0, this can be used as valid data. Is that right?\n\nIf that's right, then I come with same problem, my NDCG deteriorates as training progresses...",
    "1758929": ""
  }
}