{
  "id": 321672,
  "title": "Multiple Recall Methods---evaluation metric by GBMRanker not improving",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672",
  "author_name": "",
  "post_date": "2022-04-28T04:29:47.306963600Z",
  "votes": 18,
  "comment_count": 21,
  "views": 0,
  "content": "<p>When I use multiple recall methods, the evaluation metrics (Map12) by GBMRanker get worse than  by using only one recall method. I guess that when I generate more positive samples by multiply recall methods, I also get more negative samples. It may make GBMRanker harder to select the top-12 articles. What can I do in such a case? Should I generate more features for my candidates?</p>",
  "messages": [
    {
      "id": "1770188",
      "postDate": "04/28/2022 04:29:47",
      "content": "<p>When I use multiple recall methods, the evaluation metrics (Map12) by GBMRanker get worse than  by using only one recall method. I guess that when I generate more positive samples by multiply recall methods, I also get more negative samples. It may make GBMRanker harder to select the top-12 articles. What can I do in such a case? Should I generate more features for my candidates?</p>",
      "rawMarkdown": "When I use multiple recall methods, the evaluation metrics (Map12) by GBMRanker get worse than  by using only one recall method. I guess that when I generate more positive samples by multiply recall methods, I also get more negative samples. It may make GBMRanker harder to select the top-12 articles. What can I do in such a case? Should I generate more features for my candidates?",
      "votes": null
    },
    {
      "id": "1770202",
      "postDate": "04/28/2022 04:49:44",
      "content": "<p>I would sugguest you controlling the postive ratio of each recall method. <br>\nFor example, from recall method M1, for each customer_id your create 120 articles sorted by ranking score, from which the top ranking candidates have high possibility to be true postive and overall the postive ratio of M1 is 0.003. And from recall method M2, you also create 120 articles for each customer_id and overall the positive ratio of M2 is 0.001. <br>\nIn this case, you can control the ranking order of each strategy to make the positive ratio of each strategy is above a threshold. say, for M1 you select top 32 candidates of each customer_id to make the overall positive ratio is just above 0.005 and for M2 you select top 24 candidates of each customer_id to meet the condition of overall postive ratio is above 0.005. Then combining the two recall methods will have no problem, at least in my case, it goes like this.</p>",
      "rawMarkdown": "I would sugguest you controlling the postive ratio of each recall method. \nFor example, from recall method M1, for each customer_id your create 120 articles sorted by ranking score, from which the top ranking candidates have high possibility to be true postive and overall the postive ratio of M1 is 0.003. And from recall method M2, you also create 120 articles for each customer_id and overall the positive ratio of M2 is 0.001. \nIn this case, you can control the ranking order of each strategy to make the positive ratio of each strategy is above a threshold. say, for M1 you select top 32 candidates of each customer_id to make the overall positive ratio is just above 0.005 and for M2 you select top 24 candidates of each customer_id to meet the condition of overall postive ratio is above 0.005. Then combining the two recall methods will have no problem, at least in my case, it goes like this.",
      "votes": null
    },
    {
      "id": "1770273",
      "postDate": "04/28/2022 06:09:54",
      "content": "<p>Thanks for your reply!</p>",
      "rawMarkdown": "Thanks for your reply!",
      "votes": null
    },
    {
      "id": "1770295",
      "postDate": "04/28/2022 06:45:12",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>, sorry to disturb you again😂. .As to the inference stage, did you use all the data to generate the candidate set.(about 1.3 million users). Limited by my memory, I only use 5 weeks data to generate the candidate set (about 300k users) in the the inference stage.</p>",
      "rawMarkdown": "Hi, @lihaorocky, sorry to disturb you again😂. .As to the inference stage, did you use all the data to generate the candidate set.(about 1.3 million users). Limited by my memory, I only use 5 weeks data to generate the candidate set (about 300k users) in the the inference stage.",
      "votes": null
    },
    {
      "id": "1770319",
      "postDate": "04/28/2022 07:09:16",
      "content": "<p>For different recall methods, I use different amount of data. For example, for last buckets a customer bought, I use all weeks of data. But for recent popular articles, I use only recent N weeks of data.</p>",
      "rawMarkdown": "For different recall methods, I use different amount of data. For example, for last buckets a customer bought, I use all weeks of data. But for recent popular articles, I use only recent N weeks of data.",
      "votes": null
    },
    {
      "id": "1770326",
      "postDate": "04/28/2022 07:15:44",
      "content": "<p>Thanks for your sharing <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>!  For your first answers, does that mean I should use pre-ranking model (GBMRanker) to generate ranking scores for candidates to control the positive ratio of each strategy. And then the selected candidates by pre-ranking model are ranking again by GBMRanker?</p>",
      "rawMarkdown": "Thanks for your sharing @lihaorocky!  For your first answers, does that mean I should use pre-ranking model (GBMRanker) to generate ranking scores for candidates to control the positive ratio of each strategy. And then the selected candidates by pre-ranking model are ranking again by GBMRanker?",
      "votes": null
    },
    {
      "id": "1770330",
      "postDate": "04/28/2022 07:21:35",
      "content": "<p>No, I don't think it's necessary. In my case, for different recalls method, I use different simple rules to pre-rank the candidates for each customer. For example, when using recent popular articles, the selling amount of those articles we choose are different, we can use this to rank the articles, etc.</p>",
      "rawMarkdown": "No, I don't think it's necessary. In my case, for different recalls method, I use different simple rules to pre-rank the candidates for each customer. For example, when using recent popular articles, the selling amount of those articles we choose are different, we can use this to rank the articles, etc.",
      "votes": null
    },
    {
      "id": "1770346",
      "postDate": "04/28/2022 07:30:50",
      "content": "<p>I got it. Thank you very much for your explanation. </p>",
      "rawMarkdown": "I got it. Thank you very much for your explanation.",
      "votes": null
    },
    {
      "id": "1770467",
      "postDate": "04/28/2022 09:33:20",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>. I still fell confused about the sampling for negative sample. In the inference stage, there is no labels for us to control the positive ratio. Should I use all candidates to the pre-trained GBMRanker model? For example, there are 120 articles for recall method M1 and 120 articles for recall method M2. The total 240 articles for each customer_id should be put into the Ranking model in the inference stage. But the GBMRanker model is trained on the negative sampling data. Will the difference between the training set with negative sampling and the test set without sampling influence the final MAP12 evaluation metric.</p>",
      "rawMarkdown": "Hi, @lihaorocky. I still fell confused about the sampling for negative sample. In the inference stage, there is no labels for us to control the positive ratio. Should I use all candidates to the pre-trained GBMRanker model? For example, there are 120 articles for recall method M1 and 120 articles for recall method M2. The total 240 articles for each customer_id should be put into the Ranking model in the inference stage. But the GBMRanker model is trained on the negative sampling data. Will the difference between the training set with negative sampling and the test set without sampling influence the final MAP12 evaluation metric.",
      "votes": null
    },
    {
      "id": "1770555",
      "postDate": "04/28/2022 11:13:14",
      "content": "<p>Should I use all candidates to the pre-trained GBMRanker model?<br>\nNo. Like I said, before training, you could control the positive ratio of each recall method using some tricks. After this step, for each recall method, you get a threshold value, say topN. So if you have 120 articles for your recall method and after the selection, you find positive ratio of \"topN\" of the 120 articles are above your pre-defined ratio, say 0.003. Then you use only topN of articles from this recall method for training. And also you select topN of articles for each customer from the recall method during inference stage(because the threshold topN is decided before training, so you don't need to know the label of test data) to do prediction.<br>\nHope it helps.</p>",
      "rawMarkdown": "Should I use all candidates to the pre-trained GBMRanker model?\nNo. Like I said, before training, you could control the positive ratio of each recall method using some tricks. After this step, for each recall method, you get a threshold value, say topN. So if you have 120 articles for your recall method and after the selection, you find positive ratio of \"topN\" of the 120 articles are above your pre-defined ratio, say 0.003. Then you use only topN of articles from this recall method for training. And also you select topN of articles for each customer from the recall method during inference stage(because the threshold topN is decided before training, so you don't need to know the label of test data) to do prediction.\nHope it helps.",
      "votes": null
    },
    {
      "id": "1770574",
      "postDate": "04/28/2022 11:34:43",
      "content": "<p>This completely solved my doubts！I learn a lot from your reply. Thanks again.</p>",
      "rawMarkdown": "This completely solved my doubts！I learn a lot from your reply. Thanks again.",
      "votes": null
    },
    {
      "id": "1771716",
      "postDate": "04/29/2022 13:54:41",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> - thank you for your generous, clear and helpful sharing, in this and other threads!</p>\n<p>Two questions:</p>\n<p>#1<br>\nI imagine there's a tradeoff between the positive threshold ratio and the recall. <br>\nHow to you choose how to balance it?</p>\n<p>#2<br>\nAnd what do you do with recent popular items - I find the positivity ratio is much lower there, even for only 12 items.</p>",
      "rawMarkdown": "lihaorocky - thank you for your generous, clear and helpful sharing, in this and other threads!\n\nTwo questions:\n\n\\#1\nI imagine there's a tradeoff between the positive threshold ratio and the recall. \nHow to you choose how to balance it?\n\n\\#2\nAnd what do you do with recent popular items - I find the positivity ratio is much lower there, even for only 12 items.",
      "votes": null
    },
    {
      "id": "1771734",
      "postDate": "04/29/2022 14:08:56",
      "content": "<p>For your first question.<br>\nYes, you are absolutely right. And honestly speaking I'm still doing experiments about it. I didn't submit the recent results yet. But from the local CV score when changing the threshold from 0.007 to 0.005, the local map@12 increased about 0.0003 stably. I will keep lowering the threshold to see what maybe the best trade-off is. <br>\nFor your second question.<br>\nIf you calculate the popularity of articles overall, I guess, yes, the positive ratio will be very low. But if you calculate the hierarchic popularity combining with customers' age, gender etc. you will have a different view about it.</p>\n<p>Hope it helps.</p>",
      "rawMarkdown": "For your first question.\nYes, you are absolutely right. And honestly speaking I'm still doing experiments about it. I didn't submit the recent results yet. But from the local CV score when changing the threshold from 0.007 to 0.005, the local map@12 increased about 0.0003 stably. I will keep lowering the threshold to see what maybe the best trade-off is. \nFor your second question.\nIf you calculate the popularity of articles overall, I guess, yes, the positive ratio will be very low. But if you calculate the hierarchic popularity combining with customers' age, gender etc. you will have a different view about it.\n\nHope it helps.",
      "votes": null
    },
    {
      "id": "1771768",
      "postDate": "04/29/2022 14:53:00",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>. Really thank you for your helpful insight!<br>\nBTW, I have question about your comment:<br>\nThe meaning of  <strong><em>select top 32 candidates of each customer_id</em></strong>  is label \"1\"(purchased) to generated candiates, tag their week as validation week, and feed to ranking model to train?<br>\n(Where 32/total generated candidates is equal to threshold)</p>",
      "rawMarkdown": "Hi, @lihaorocky. Really thank you for your helpful insight!\nBTW, I have question about your comment:\nThe meaning of  ***select top 32 candidates of each customer_id***  is label \"1\"(purchased) to generated candiates, tag their week as validation week, and feed to ranking model to train?\n(Where 32/total generated candidates is equal to threshold)",
      "votes": null
    },
    {
      "id": "1771816",
      "postDate": "04/29/2022 15:50:35",
      "content": "<p>It helps a lot, thank you!</p>\n<p>You must have a very strong model/features if you're able to get good ranking results from a pool with just .5% threshold.</p>\n<p>I'm struggling with a much higher concentration of positives.</p>",
      "rawMarkdown": "It helps a lot, thank you!\n\nYou must have a very strong model/features if you're able to get good ranking results from a pool with just .5% threshold.\n\nI'm struggling with a much higher concentration of positives.",
      "votes": null
    },
    {
      "id": "1771848",
      "postDate": "04/29/2022 16:02:54",
      "content": "<p>Thank you for this answer <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> , your perspective on how to calibrate each strategy is very interesting</p>",
      "rawMarkdown": "Thank you for this answer @lihaorocky , your perspective on how to calibrate each strategy is very interesting",
      "votes": null
    },
    {
      "id": "1771851",
      "postDate": "04/29/2022 16:03:58",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>. Thanks for your sharing. I have another question. Different recall strategies may generate duplicate articles. In my case, for different recall methods ,it's about 40% when I combining  candidates generated by  all recall methods. After I drop this duplicate articles, the positive ratio decreases under the  the threshold. (eg.0.005), although I ensure the positive ratio of each recall method is above this threshold. Should I sample the combined candidates again?</p>",
      "rawMarkdown": "Hi, @lihaorocky. Thanks for your sharing. I have another question. Different recall strategies may generate duplicate articles. In my case, for different recall methods ,it's about 40% when I combining  candidates generated by  all recall methods. After I drop this duplicate articles, the positive ratio decreases under the  the threshold. (eg.0.005), although I ensure the positive ratio of each recall method is above this threshold. Should I sample the combined candidates again?",
      "votes": null
    },
    {
      "id": "1771870",
      "postDate": "04/29/2022 16:12:42",
      "content": "<p>I just simply outer merge them after selecting the ranking thresholds to do the filtering. I think it's OK one candidate is shared by multiple recall methods and it's beneficial to keep it rather drop it.</p>",
      "rawMarkdown": "I just simply outer merge them after selecting the ranking thresholds to do the filtering. I think it's OK one candidate is shared by multiple recall methods and it's beneficial to keep it rather drop it.",
      "votes": null
    },
    {
      "id": "1772159",
      "postDate": "04/29/2022 22:57:47",
      "content": "<p>It's an interesting trick. I'll try it. Thanks.👍</p>",
      "rawMarkdown": "It's an interesting trick. I'll try it. Thanks.👍",
      "votes": null
    },
    {
      "id": "1772365",
      "postDate": "04/30/2022 07:16:46",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a>, can you explain me about how to calculate positive threshold ratio and it’s meaning?<br>\nAnd I’m not sure how to label generated candidates for train. Is it okay to label “1” to generated candidates, and tag their week as validation week?</p>",
      "rawMarkdown": "Hi @jacob34, can you explain me about how to calculate positive threshold ratio and it’s meaning?\nAnd I’m not sure how to label generated candidates for train. Is it okay to label “1” to generated candidates, and tag their week as validation week?",
      "votes": null
    },
    {
      "id": "1772802",
      "postDate": "04/30/2022 15:29:20",
      "content": "<p><a href=\"https://www.kaggle.com/qizhengpan\" target=\"_blank\">@qizhengpan</a> , well done , the best ☝️☝️☝️</p>",
      "rawMarkdown": "qizhengpan , well done , the best ☝️☝️☝️",
      "votes": null
    },
    {
      "id": "1773190",
      "postDate": "04/30/2022 23:58:19",
      "content": "<p>I think you are hurting your performance by generating more negative samples with multiple recall methods. I would recommend you try controlling the number of positive examples that each method generates. Try creating a smaller set of articles for each customer_id and make sure the top ranked candidates are all positives.</p>",
      "rawMarkdown": "I think you are hurting your performance by generating more negative samples with multiple recall methods. I would recommend you try controlling the number of positive examples that each method generates. Try creating a smaller set of articles for each customer_id and make sure the top ranked candidates are all positives.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1770202,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "04/28/2022 04:49:44",
      "content": "<p>I would sugguest you controlling the postive ratio of each recall method. <br>\nFor example, from recall method M1, for each customer_id your create 120 articles sorted by ranking score, from which the top ranking candidates have high possibility to be true postive and overall the postive ratio of M1 is 0.003. And from recall method M2, you also create 120 articles for each customer_id and overall the positive ratio of M2 is 0.001. <br>\nIn this case, you can control the ranking order of each strategy to make the positive ratio of each strategy is above a threshold. say, for M1 you select top 32 candidates of each customer_id to make the overall positive ratio is just above 0.005 and for M2 you select top 24 candidates of each customer_id to meet the condition of overall postive ratio is above 0.005. Then combining the two recall methods will have no problem, at least in my case, it goes like this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1770273,
          "author_name": "qizhengpan",
          "author_url": "",
          "post_date": "04/28/2022 06:09:54",
          "content": "<p>Thanks for your reply!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1770295,
          "author_name": "qizhengpan",
          "author_url": "",
          "post_date": "04/28/2022 06:45:12",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>, sorry to disturb you again😂. .As to the inference stage, did you use all the data to generate the candidate set.(about 1.3 million users). Limited by my memory, I only use 5 weeks data to generate the candidate set (about 300k users) in the the inference stage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1770319,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "04/28/2022 07:09:16",
          "content": "<p>For different recall methods, I use different amount of data. For example, for last buckets a customer bought, I use all weeks of data. But for recent popular articles, I use only recent N weeks of data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1770326,
          "author_name": "qizhengpan",
          "author_url": "",
          "post_date": "04/28/2022 07:15:44",
          "content": "<p>Thanks for your sharing <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>!  For your first answers, does that mean I should use pre-ranking model (GBMRanker) to generate ranking scores for candidates to control the positive ratio of each strategy. And then the selected candidates by pre-ranking model are ranking again by GBMRanker?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1770330,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "04/28/2022 07:21:35",
          "content": "<p>No, I don't think it's necessary. In my case, for different recalls method, I use different simple rules to pre-rank the candidates for each customer. For example, when using recent popular articles, the selling amount of those articles we choose are different, we can use this to rank the articles, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1770346,
          "author_name": "qizhengpan",
          "author_url": "",
          "post_date": "04/28/2022 07:30:50",
          "content": "<p>I got it. Thank you very much for your explanation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1770467,
          "author_name": "qizhengpan",
          "author_url": "",
          "post_date": "04/28/2022 09:33:20",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>. I still fell confused about the sampling for negative sample. In the inference stage, there is no labels for us to control the positive ratio. Should I use all candidates to the pre-trained GBMRanker model? For example, there are 120 articles for recall method M1 and 120 articles for recall method M2. The total 240 articles for each customer_id should be put into the Ranking model in the inference stage. But the GBMRanker model is trained on the negative sampling data. Will the difference between the training set with negative sampling and the test set without sampling influence the final MAP12 evaluation metric.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1770555,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "04/28/2022 11:13:14",
          "content": "<p>Should I use all candidates to the pre-trained GBMRanker model?<br>\nNo. Like I said, before training, you could control the positive ratio of each recall method using some tricks. After this step, for each recall method, you get a threshold value, say topN. So if you have 120 articles for your recall method and after the selection, you find positive ratio of \"topN\" of the 120 articles are above your pre-defined ratio, say 0.003. Then you use only topN of articles from this recall method for training. And also you select topN of articles for each customer from the recall method during inference stage(because the threshold topN is decided before training, so you don't need to know the label of test data) to do prediction.<br>\nHope it helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1770574,
          "author_name": "qizhengpan",
          "author_url": "",
          "post_date": "04/28/2022 11:34:43",
          "content": "<p>This completely solved my doubts！I learn a lot from your reply. Thanks again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771716,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/29/2022 13:54:41",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> - thank you for your generous, clear and helpful sharing, in this and other threads!</p>\n<p>Two questions:</p>\n<p>#1<br>\nI imagine there's a tradeoff between the positive threshold ratio and the recall. <br>\nHow to you choose how to balance it?</p>\n<p>#2<br>\nAnd what do you do with recent popular items - I find the positivity ratio is much lower there, even for only 12 items.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771734,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "04/29/2022 14:08:56",
          "content": "<p>For your first question.<br>\nYes, you are absolutely right. And honestly speaking I'm still doing experiments about it. I didn't submit the recent results yet. But from the local CV score when changing the threshold from 0.007 to 0.005, the local map@12 increased about 0.0003 stably. I will keep lowering the threshold to see what maybe the best trade-off is. <br>\nFor your second question.<br>\nIf you calculate the popularity of articles overall, I guess, yes, the positive ratio will be very low. But if you calculate the hierarchic popularity combining with customers' age, gender etc. you will have a different view about it.</p>\n<p>Hope it helps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771768,
          "author_name": "lionsheep24",
          "author_url": "",
          "post_date": "04/29/2022 14:53:00",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>. Really thank you for your helpful insight!<br>\nBTW, I have question about your comment:<br>\nThe meaning of  <strong><em>select top 32 candidates of each customer_id</em></strong>  is label \"1\"(purchased) to generated candiates, tag their week as validation week, and feed to ranking model to train?<br>\n(Where 32/total generated candidates is equal to threshold)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771816,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/29/2022 15:50:35",
          "content": "<p>It helps a lot, thank you!</p>\n<p>You must have a very strong model/features if you're able to get good ranking results from a pool with just .5% threshold.</p>\n<p>I'm struggling with a much higher concentration of positives.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771848,
          "author_name": "lorenzopagliaro01",
          "author_url": "",
          "post_date": "04/29/2022 16:02:54",
          "content": "<p>Thank you for this answer <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> , your perspective on how to calibrate each strategy is very interesting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771851,
          "author_name": "qizhengpan",
          "author_url": "",
          "post_date": "04/29/2022 16:03:58",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a>. Thanks for your sharing. I have another question. Different recall strategies may generate duplicate articles. In my case, for different recall methods ,it's about 40% when I combining  candidates generated by  all recall methods. After I drop this duplicate articles, the positive ratio decreases under the  the threshold. (eg.0.005), although I ensure the positive ratio of each recall method is above this threshold. Should I sample the combined candidates again?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1771870,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "04/29/2022 16:12:42",
          "content": "<p>I just simply outer merge them after selecting the ranking thresholds to do the filtering. I think it's OK one candidate is shared by multiple recall methods and it's beneficial to keep it rather drop it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1772159,
          "author_name": "qizhengpan",
          "author_url": "",
          "post_date": "04/29/2022 22:57:47",
          "content": "<p>It's an interesting trick. I'll try it. Thanks.👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1772365,
          "author_name": "lionsheep24",
          "author_url": "",
          "post_date": "04/30/2022 07:16:46",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a>, can you explain me about how to calculate positive threshold ratio and it’s meaning?<br>\nAnd I’m not sure how to label generated candidates for train. Is it okay to label “1” to generated candidates, and tag their week as validation week?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1772802,
      "author_name": "alexandrparkhomenko",
      "author_url": "",
      "post_date": "04/30/2022 15:29:20",
      "content": "<p><a href=\"https://www.kaggle.com/qizhengpan\" target=\"_blank\">@qizhengpan</a> , well done , the best ☝️☝️☝️</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1773190,
      "author_name": "",
      "author_url": "",
      "post_date": "04/30/2022 23:58:19",
      "content": "<p>I think you are hurting your performance by generating more negative samples with multiple recall methods. I would recommend you try controlling the number of positive examples that each method generates. Try creating a smaller set of articles for each customer_id and make sure the top ranked candidates are all positives.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1770188": "When I use multiple recall methods, the evaluation metrics (Map12) by GBMRanker get worse than  by using only one recall method. I guess that when I generate more positive samples by multiply recall methods, I also get more negative samples. It may make GBMRanker harder to select the top-12 articles. What can I do in such a case? Should I generate more features for my candidates?",
    "1770202": "I would sugguest you controlling the postive ratio of each recall method. \nFor example, from recall method M1, for each customer_id your create 120 articles sorted by ranking score, from which the top ranking candidates have high possibility to be true postive and overall the postive ratio of M1 is 0.003. And from recall method M2, you also create 120 articles for each customer_id and overall the positive ratio of M2 is 0.001. \nIn this case, you can control the ranking order of each strategy to make the positive ratio of each strategy is above a threshold. say, for M1 you select top 32 candidates of each customer_id to make the overall positive ratio is just above 0.005 and for M2 you select top 24 candidates of each customer_id to meet the condition of overall postive ratio is above 0.005. Then combining the two recall methods will have no problem, at least in my case, it goes like this.",
    "1770273": "Thanks for your reply!",
    "1770295": "Hi, @lihaorocky, sorry to disturb you again😂. .As to the inference stage, did you use all the data to generate the candidate set.(about 1.3 million users). Limited by my memory, I only use 5 weeks data to generate the candidate set (about 300k users) in the the inference stage.",
    "1770319": "For different recall methods, I use different amount of data. For example, for last buckets a customer bought, I use all weeks of data. But for recent popular articles, I use only recent N weeks of data.",
    "1770326": "Thanks for your sharing @lihaorocky!  For your first answers, does that mean I should use pre-ranking model (GBMRanker) to generate ranking scores for candidates to control the positive ratio of each strategy. And then the selected candidates by pre-ranking model are ranking again by GBMRanker?",
    "1770330": "No, I don't think it's necessary. In my case, for different recalls method, I use different simple rules to pre-rank the candidates for each customer. For example, when using recent popular articles, the selling amount of those articles we choose are different, we can use this to rank the articles, etc.",
    "1770346": "I got it. Thank you very much for your explanation.",
    "1770467": "Hi, @lihaorocky. I still fell confused about the sampling for negative sample. In the inference stage, there is no labels for us to control the positive ratio. Should I use all candidates to the pre-trained GBMRanker model? For example, there are 120 articles for recall method M1 and 120 articles for recall method M2. The total 240 articles for each customer_id should be put into the Ranking model in the inference stage. But the GBMRanker model is trained on the negative sampling data. Will the difference between the training set with negative sampling and the test set without sampling influence the final MAP12 evaluation metric.",
    "1770555": "Should I use all candidates to the pre-trained GBMRanker model?\nNo. Like I said, before training, you could control the positive ratio of each recall method using some tricks. After this step, for each recall method, you get a threshold value, say topN. So if you have 120 articles for your recall method and after the selection, you find positive ratio of \"topN\" of the 120 articles are above your pre-defined ratio, say 0.003. Then you use only topN of articles from this recall method for training. And also you select topN of articles for each customer from the recall method during inference stage(because the threshold topN is decided before training, so you don't need to know the label of test data) to do prediction.\nHope it helps.",
    "1770574": "This completely solved my doubts！I learn a lot from your reply. Thanks again.",
    "1771716": "lihaorocky - thank you for your generous, clear and helpful sharing, in this and other threads!\n\nTwo questions:\n\n\\#1\nI imagine there's a tradeoff between the positive threshold ratio and the recall. \nHow to you choose how to balance it?\n\n\\#2\nAnd what do you do with recent popular items - I find the positivity ratio is much lower there, even for only 12 items.",
    "1771734": "For your first question.\nYes, you are absolutely right. And honestly speaking I'm still doing experiments about it. I didn't submit the recent results yet. But from the local CV score when changing the threshold from 0.007 to 0.005, the local map@12 increased about 0.0003 stably. I will keep lowering the threshold to see what maybe the best trade-off is. \nFor your second question.\nIf you calculate the popularity of articles overall, I guess, yes, the positive ratio will be very low. But if you calculate the hierarchic popularity combining with customers' age, gender etc. you will have a different view about it.\n\nHope it helps.",
    "1771768": "Hi, @lihaorocky. Really thank you for your helpful insight!\nBTW, I have question about your comment:\nThe meaning of  ***select top 32 candidates of each customer_id***  is label \"1\"(purchased) to generated candiates, tag their week as validation week, and feed to ranking model to train?\n(Where 32/total generated candidates is equal to threshold)",
    "1771816": "It helps a lot, thank you!\n\nYou must have a very strong model/features if you're able to get good ranking results from a pool with just .5% threshold.\n\nI'm struggling with a much higher concentration of positives.",
    "1771848": "Thank you for this answer @lihaorocky , your perspective on how to calibrate each strategy is very interesting",
    "1771851": "Hi, @lihaorocky. Thanks for your sharing. I have another question. Different recall strategies may generate duplicate articles. In my case, for different recall methods ,it's about 40% when I combining  candidates generated by  all recall methods. After I drop this duplicate articles, the positive ratio decreases under the  the threshold. (eg.0.005), although I ensure the positive ratio of each recall method is above this threshold. Should I sample the combined candidates again?",
    "1771870": "I just simply outer merge them after selecting the ranking thresholds to do the filtering. I think it's OK one candidate is shared by multiple recall methods and it's beneficial to keep it rather drop it.",
    "1772159": "It's an interesting trick. I'll try it. Thanks.👍",
    "1772365": "Hi @jacob34, can you explain me about how to calculate positive threshold ratio and it’s meaning?\nAnd I’m not sure how to label generated candidates for train. Is it okay to label “1” to generated candidates, and tag their week as validation week?",
    "1772802": "qizhengpan , well done , the best ☝️☝️☝️",
    "1773190": "I think you are hurting your performance by generating more negative samples with multiple recall methods. I would recommend you try controlling the number of positive examples that each method generates. Try creating a smaller set of articles for each customer_id and make sure the top ranked candidates are all positives."
  },
  "source": "meta"
}