{
  "id": 316205,
  "title": "Best single model",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/316205",
  "author_name": "",
  "post_date": "2022-03-31T19:17:57.231188500Z",
  "votes": 29,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hey folks,</p>\n<p>Starting a common thread<br>\nWhat is your current best single model?</p>\n<p>Ours:<br>\nCollaborative filtering similarity-based (SAR) + most popular items for cold-start users </p>\n<p>CV: 0.272<br>\nLB: 0.230</p>",
  "messages": [
    {
      "id": "1741395",
      "postDate": "03/31/2022 19:17:57",
      "content": "<p>Hey folks,</p>\n<p>Starting a common thread<br>\nWhat is your current best single model?</p>\n<p>Ours:<br>\nCollaborative filtering similarity-based (SAR) + most popular items for cold-start users </p>\n<p>CV: 0.272<br>\nLB: 0.230</p>",
      "rawMarkdown": "Hey folks,\n\nStarting a common thread\nWhat is your current best single model?\n\nOurs:\nCollaborative filtering similarity-based (SAR) + most popular items for cold-start users \n\nCV: 0.272\nLB: 0.230",
      "votes": null
    },
    {
      "id": "1741873",
      "postDate": "04/01/2022 08:25:23",
      "content": "<p>good job!</p>",
      "rawMarkdown": "good job!",
      "votes": null
    },
    {
      "id": "1744745",
      "postDate": "04/04/2022 09:13:28",
      "content": "<p>Good job!!</p>",
      "rawMarkdown": "Good job!!",
      "votes": null
    },
    {
      "id": "1744839",
      "postDate": "04/04/2022 11:09:28",
      "content": "<p>good job! This is beyond my imagination</p>",
      "rawMarkdown": "good job! This is beyond my imagination",
      "votes": null
    },
    {
      "id": "1746769",
      "postDate": "04/06/2022 04:49:06",
      "content": "<p>Method: LGBMRanker <br>\nCV: 0.0406<br>\nLB: 0.0323</p>",
      "rawMarkdown": "Method: LGBMRanker \nCV: 0.0406\nLB: 0.0323",
      "votes": null
    },
    {
      "id": "1747223",
      "postDate": "04/06/2022 13:55:08",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1753816",
      "postDate": "04/13/2022 05:53:21",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> 大佬能留一个联系方式么<br>\n想和你组队交流一下</p>",
      "rawMarkdown": "lihaorocky 大佬能留一个联系方式么\n想和你组队交流一下",
      "votes": null
    },
    {
      "id": "1754481",
      "postDate": "04/13/2022 18:04:41",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> it is so funny. I have the exact same numbers. CV around 0.0406 and leaderboard between 0.0323 - 0.0340. I thought that I my CV is faulty but seeing you're number it is probably good.</p>",
      "rawMarkdown": "lihaorocky it is so funny. I have the exact same numbers. CV around 0.0406 and leaderboard between 0.0323 - 0.0340. I thought that I my CV is faulty but seeing you're number it is probably good.",
      "votes": null
    },
    {
      "id": "1754768",
      "postDate": "04/14/2022 02:09:02",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> Even though, I'm still not very confident about it. Every time I have a huge increase with my CV score, but the LB score will improve only a little bit. And right now, my CV score is around 0.0425, but the LB score is just a little bit above 0.033. With my LB score increasing, the gap between CV and LB is also getting bigger and bigger, which is very frustrating.😂</p>",
      "rawMarkdown": "paweljankiewicz Even though, I'm still not very confident about it. Every time I have a huge increase with my CV score, but the LB score will improve only a little bit. And right now, my CV score is around 0.0425, but the LB score is just a little bit above 0.033. With my LB score increasing, the gap between CV and LB is also getting bigger and bigger, which is very frustrating.😂",
      "votes": null
    },
    {
      "id": "1755039",
      "postDate": "04/14/2022 08:59:02",
      "content": "<p>Yeah I noticed this as well that my CV score stopped corresponding very well to the LB. I think CV can be inflated because of different ratio of cold users on the LB or some marketing campaigns? 2020-09 is definitely quite strange. Not having negative observations has a lot of disadvantages. Imagine that on 2020-09-23 H&amp;M created a huge marketing campaign for some products that could affect the sales.</p>",
      "rawMarkdown": "Yeah I noticed this as well that my CV score stopped corresponding very well to the LB. I think CV can be inflated because of different ratio of cold users on the LB or some marketing campaigns? 2020-09 is definitely quite strange. Not having negative observations has a lot of disadvantages. Imagine that on 2020-09-23 H&M created a huge marketing campaign for some products that could affect the sales.",
      "votes": null
    },
    {
      "id": "1756134",
      "postDate": "04/15/2022 09:15:02",
      "content": "<p>handbuilt basically same model as urs:</p>\n<p>CV:0.0268 (using last week as validation)<br>\nLB:0.0242</p>\n<p>Might try some lgbmranker next couple of weeks to see how that pans out, I've read the lambdarank paper and understand how it works and how to use it but it seems like I would need an ungodly amount of memory.</p>\n<p>Not sure how so many people are getting CVs of 0.04 and LBs above 0.03, I feel like I've squeezed the lemon as much as is possible.<br>\nI have a hunch that they might've stumbled upon the way that H&amp;M does their ad campaigns themselves with the lgbmranker.</p>\n<p>Update:</p>\n<p>did first attempt with lgbm somewhat surprised by result because I am only using a 16gb kaggle notebook and my cv perfectly alines with my lb. To me this seems like it is really just replicating H&amp;Ms strategy</p>\n<p>CV:0.0298 (using last week as validation)<br>\nLB:0.0298</p>\n<p>In fact, thinking about it, it should be inevitable that a machine learned model will fit its parameters to reproduce the strategy by H&amp;M, which makes these types of competitions ultimately useless for the company if they don't make the test set such that there wasn't anything advertised to the customers. And even then, building an unbiased model will be very difficult because of the heavy influence of the recommendations in the past on the customer's choices.</p>",
      "rawMarkdown": "handbuilt basically same model as urs:\n\nCV:0.0268 (using last week as validation)\nLB:0.0242\n\nMight try some lgbmranker next couple of weeks to see how that pans out, I've read the lambdarank paper and understand how it works and how to use it but it seems like I would need an ungodly amount of memory.\n\nNot sure how so many people are getting CVs of 0.04 and LBs above 0.03, I feel like I've squeezed the lemon as much as is possible.\nI have a hunch that they might've stumbled upon the way that H&M does their ad campaigns themselves with the lgbmranker.\n\nUpdate:\n\ndid first attempt with lgbm somewhat surprised by result because I am only using a 16gb kaggle notebook and my cv perfectly alines with my lb. To me this seems like it is really just replicating H&Ms strategy\n\nCV:0.0298 (using last week as validation)\nLB:0.0298\n\nIn fact, thinking about it, it should be inevitable that a machine learned model will fit its parameters to reproduce the strategy by H&M, which makes these types of competitions ultimately useless for the company if they don't make the test set such that there wasn't anything advertised to the customers. And even then, building an unbiased model will be very difficult because of the heavy influence of the recommendations in the past on the customer's choices.",
      "votes": null
    },
    {
      "id": "1756369",
      "postDate": "04/15/2022 12:54:03",
      "content": "<blockquote>\n  <p>I have a hunch that they might've stumbled upon the way that H&amp;M does their ad campaigns themselves with the lgbmranker.</p>\n</blockquote>\n<p>The way you have written this you assume that there is some trick to get a good score and you diminish the hard work people put into this competition. There is probably an overlap between H&amp;M recommendation strategies and top positions strategies but the results are pretty much very random. Honestly I'm not looking for any trick just iterate as much as possible. My full process takes about 14-15 hours right now so I can afford to do 1 experiment per day mainly centered around creating strategies to generate candidates.</p>\n<blockquote>\n  <p>Not sure how so many people are getting CVs of 0.04 and LBs above 0.03, I feel like I've squeezed the lemon as much as is possible.</p>\n</blockquote>\n<p>I have 1.5+ billion candidates for submission. The dataframe with the candidates on 2020-09-22 approaches 200gb compressed right now. Recommendation problems are notorious for creating very big datasets. i recommend you start working on good ways to generate decent candidates.</p>",
      "rawMarkdown": "> I have a hunch that they might've stumbled upon the way that H&M does their ad campaigns themselves with the lgbmranker.\n\nThe way you have written this you assume that there is some trick to get a good score and you diminish the hard work people put into this competition. There is probably an overlap between H&M recommendation strategies and top positions strategies but the results are pretty much very random. Honestly I'm not looking for any trick just iterate as much as possible. My full process takes about 14-15 hours right now so I can afford to do 1 experiment per day mainly centered around creating strategies to generate candidates.\n\n> Not sure how so many people are getting CVs of 0.04 and LBs above 0.03, I feel like I've squeezed the lemon as much as is possible.\n\nI have 1.5+ billion candidates for submission. The dataframe with the candidates on 2020-09-22 approaches 200gb compressed right now. Recommendation problems are notorious for creating very big datasets. i recommend you start working on good ways to generate decent candidates.",
      "votes": null
    },
    {
      "id": "1756436",
      "postDate": "04/15/2022 14:14:38",
      "content": "<p>alright  thanks for the tips<br>\ndidn't mean it to sound diminishing or that there is some unfair \"trick\" , i just have terrible communication skills ^^</p>",
      "rawMarkdown": "alright  thanks for the tips\ndidn't mean it to sound diminishing or that there is some unfair \"trick\" , i just have terrible communication skills ^^",
      "votes": null
    },
    {
      "id": "1756527",
      "postDate": "04/15/2022 15:30:52",
      "content": "<p>Do you mind sharing if you have a mostly different set of candidates for each customer while training or if you have the same set of candidates for everyone while training? </p>\n<p>I imagine it would be something like go week by week and use like popular items of previous week(s) based on gender and specific buying habits based on gender and age and such?</p>",
      "rawMarkdown": "Do you mind sharing if you have a mostly different set of candidates for each customer while training or if you have the same set of candidates for everyone while training? \n\nI imagine it would be something like go week by week and use like popular items of previous week(s) based on gender and specific buying habits based on gender and age and such?",
      "votes": null
    },
    {
      "id": "1756559",
      "postDate": "04/15/2022 15:55:25",
      "content": "<p>Some strategies use \"global\" popularity, some of them use other features of the user. It is a mix.</p>\n<blockquote>\n  <p>I imagine it would be something like go week by week and use like popular items of previous week(s) based on gender and specific buying habits based on gender and age and such?</p>\n</blockquote>\n<p>This is exactly right.</p>",
      "rawMarkdown": "Some strategies use \"global\" popularity, some of them use other features of the user. It is a mix.\n\n> I imagine it would be something like go week by week and use like popular items of previous week(s) based on gender and specific buying habits based on gender and age and such?\n\nThis is exactly right.",
      "votes": null
    },
    {
      "id": "1761513",
      "postDate": "04/20/2022 01:27:55",
      "content": "<p>good  job！</p>",
      "rawMarkdown": "good  job！",
      "votes": null
    },
    {
      "id": "1764235",
      "postDate": "04/22/2022 09:25:14",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> 1.5+ billion is an amazing num! If you don't mind, could I know whether you use all the candidates to build the training samples? </p>",
      "rawMarkdown": "paweljankiewicz 1.5+ billion is an amazing num! If you don't mind, could I know whether you use all the candidates to build the training samples?",
      "votes": null
    },
    {
      "id": "1775097",
      "postDate": "05/02/2022 17:44:05",
      "content": "<p>thanks for the tips again, did my first attempt with the lgbmranker and i could get really good results already by only using 50 or so candidates on average</p>",
      "rawMarkdown": "thanks for the tips again, did my first attempt with the lgbmranker and i could get really good results already by only using 50 or so candidates on average",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1741873,
      "author_name": "xuxiaodong",
      "author_url": "",
      "post_date": "04/01/2022 08:25:23",
      "content": "<p>good job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1744745,
      "author_name": "edogru",
      "author_url": "",
      "post_date": "04/04/2022 09:13:28",
      "content": "<p>Good job!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1744839,
      "author_name": "shizhan233",
      "author_url": "",
      "post_date": "04/04/2022 11:09:28",
      "content": "<p>good job! This is beyond my imagination</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1746769,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "04/06/2022 04:49:06",
      "content": "<p>Method: LGBMRanker <br>\nCV: 0.0406<br>\nLB: 0.0323</p>",
      "votes": null,
      "replies": [
        {
          "id": 1747223,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "04/06/2022 13:55:08",
          "content": "<p>Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1754481,
          "author_name": "paweljankiewicz",
          "author_url": "",
          "post_date": "04/13/2022 18:04:41",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> it is so funny. I have the exact same numbers. CV around 0.0406 and leaderboard between 0.0323 - 0.0340. I thought that I my CV is faulty but seeing you're number it is probably good.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1754768,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "04/14/2022 02:09:02",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> Even though, I'm still not very confident about it. Every time I have a huge increase with my CV score, but the LB score will improve only a little bit. And right now, my CV score is around 0.0425, but the LB score is just a little bit above 0.033. With my LB score increasing, the gap between CV and LB is also getting bigger and bigger, which is very frustrating.😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1755039,
          "author_name": "paweljankiewicz",
          "author_url": "",
          "post_date": "04/14/2022 08:59:02",
          "content": "<p>Yeah I noticed this as well that my CV score stopped corresponding very well to the LB. I think CV can be inflated because of different ratio of cold users on the LB or some marketing campaigns? 2020-09 is definitely quite strange. Not having negative observations has a lot of disadvantages. Imagine that on 2020-09-23 H&amp;M created a huge marketing campaign for some products that could affect the sales.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1753816,
      "author_name": "zhangxueren",
      "author_url": "",
      "post_date": "04/13/2022 05:53:21",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> 大佬能留一个联系方式么<br>\n想和你组队交流一下</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1756134,
      "author_name": "kaggledoer",
      "author_url": "",
      "post_date": "04/15/2022 09:15:02",
      "content": "<p>handbuilt basically same model as urs:</p>\n<p>CV:0.0268 (using last week as validation)<br>\nLB:0.0242</p>\n<p>Might try some lgbmranker next couple of weeks to see how that pans out, I've read the lambdarank paper and understand how it works and how to use it but it seems like I would need an ungodly amount of memory.</p>\n<p>Not sure how so many people are getting CVs of 0.04 and LBs above 0.03, I feel like I've squeezed the lemon as much as is possible.<br>\nI have a hunch that they might've stumbled upon the way that H&amp;M does their ad campaigns themselves with the lgbmranker.</p>\n<p>Update:</p>\n<p>did first attempt with lgbm somewhat surprised by result because I am only using a 16gb kaggle notebook and my cv perfectly alines with my lb. To me this seems like it is really just replicating H&amp;Ms strategy</p>\n<p>CV:0.0298 (using last week as validation)<br>\nLB:0.0298</p>\n<p>In fact, thinking about it, it should be inevitable that a machine learned model will fit its parameters to reproduce the strategy by H&amp;M, which makes these types of competitions ultimately useless for the company if they don't make the test set such that there wasn't anything advertised to the customers. And even then, building an unbiased model will be very difficult because of the heavy influence of the recommendations in the past on the customer's choices.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1756369,
          "author_name": "paweljankiewicz",
          "author_url": "",
          "post_date": "04/15/2022 12:54:03",
          "content": "<blockquote>\n  <p>I have a hunch that they might've stumbled upon the way that H&amp;M does their ad campaigns themselves with the lgbmranker.</p>\n</blockquote>\n<p>The way you have written this you assume that there is some trick to get a good score and you diminish the hard work people put into this competition. There is probably an overlap between H&amp;M recommendation strategies and top positions strategies but the results are pretty much very random. Honestly I'm not looking for any trick just iterate as much as possible. My full process takes about 14-15 hours right now so I can afford to do 1 experiment per day mainly centered around creating strategies to generate candidates.</p>\n<blockquote>\n  <p>Not sure how so many people are getting CVs of 0.04 and LBs above 0.03, I feel like I've squeezed the lemon as much as is possible.</p>\n</blockquote>\n<p>I have 1.5+ billion candidates for submission. The dataframe with the candidates on 2020-09-22 approaches 200gb compressed right now. Recommendation problems are notorious for creating very big datasets. i recommend you start working on good ways to generate decent candidates.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1756436,
          "author_name": "kaggledoer",
          "author_url": "",
          "post_date": "04/15/2022 14:14:38",
          "content": "<p>alright  thanks for the tips<br>\ndidn't mean it to sound diminishing or that there is some unfair \"trick\" , i just have terrible communication skills ^^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1756527,
          "author_name": "kaggledoer",
          "author_url": "",
          "post_date": "04/15/2022 15:30:52",
          "content": "<p>Do you mind sharing if you have a mostly different set of candidates for each customer while training or if you have the same set of candidates for everyone while training? </p>\n<p>I imagine it would be something like go week by week and use like popular items of previous week(s) based on gender and specific buying habits based on gender and age and such?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1756559,
          "author_name": "paweljankiewicz",
          "author_url": "",
          "post_date": "04/15/2022 15:55:25",
          "content": "<p>Some strategies use \"global\" popularity, some of them use other features of the user. It is a mix.</p>\n<blockquote>\n  <p>I imagine it would be something like go week by week and use like popular items of previous week(s) based on gender and specific buying habits based on gender and age and such?</p>\n</blockquote>\n<p>This is exactly right.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764235,
          "author_name": "sirius81",
          "author_url": "",
          "post_date": "04/22/2022 09:25:14",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> 1.5+ billion is an amazing num! If you don't mind, could I know whether you use all the candidates to build the training samples? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1775097,
          "author_name": "kaggledoer",
          "author_url": "",
          "post_date": "05/02/2022 17:44:05",
          "content": "<p>thanks for the tips again, did my first attempt with the lgbmranker and i could get really good results already by only using 50 or so candidates on average</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1761513,
      "author_name": "pangzi233",
      "author_url": "",
      "post_date": "04/20/2022 01:27:55",
      "content": "<p>good  job！</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1741395": "Hey folks,\n\nStarting a common thread\nWhat is your current best single model?\n\nOurs:\nCollaborative filtering similarity-based (SAR) + most popular items for cold-start users \n\nCV: 0.272\nLB: 0.230",
    "1741873": "good job!",
    "1744745": "Good job!!",
    "1744839": "good job! This is beyond my imagination",
    "1746769": "Method: LGBMRanker \nCV: 0.0406\nLB: 0.0323",
    "1747223": "Thanks for sharing.",
    "1753816": "lihaorocky 大佬能留一个联系方式么\n想和你组队交流一下",
    "1754481": "lihaorocky it is so funny. I have the exact same numbers. CV around 0.0406 and leaderboard between 0.0323 - 0.0340. I thought that I my CV is faulty but seeing you're number it is probably good.",
    "1754768": "paweljankiewicz Even though, I'm still not very confident about it. Every time I have a huge increase with my CV score, but the LB score will improve only a little bit. And right now, my CV score is around 0.0425, but the LB score is just a little bit above 0.033. With my LB score increasing, the gap between CV and LB is also getting bigger and bigger, which is very frustrating.😂",
    "1755039": "Yeah I noticed this as well that my CV score stopped corresponding very well to the LB. I think CV can be inflated because of different ratio of cold users on the LB or some marketing campaigns? 2020-09 is definitely quite strange. Not having negative observations has a lot of disadvantages. Imagine that on 2020-09-23 H&M created a huge marketing campaign for some products that could affect the sales.",
    "1756134": "handbuilt basically same model as urs:\n\nCV:0.0268 (using last week as validation)\nLB:0.0242\n\nMight try some lgbmranker next couple of weeks to see how that pans out, I've read the lambdarank paper and understand how it works and how to use it but it seems like I would need an ungodly amount of memory.\n\nNot sure how so many people are getting CVs of 0.04 and LBs above 0.03, I feel like I've squeezed the lemon as much as is possible.\nI have a hunch that they might've stumbled upon the way that H&M does their ad campaigns themselves with the lgbmranker.\n\nUpdate:\n\ndid first attempt with lgbm somewhat surprised by result because I am only using a 16gb kaggle notebook and my cv perfectly alines with my lb. To me this seems like it is really just replicating H&Ms strategy\n\nCV:0.0298 (using last week as validation)\nLB:0.0298\n\nIn fact, thinking about it, it should be inevitable that a machine learned model will fit its parameters to reproduce the strategy by H&M, which makes these types of competitions ultimately useless for the company if they don't make the test set such that there wasn't anything advertised to the customers. And even then, building an unbiased model will be very difficult because of the heavy influence of the recommendations in the past on the customer's choices.",
    "1756369": "> I have a hunch that they might've stumbled upon the way that H&M does their ad campaigns themselves with the lgbmranker.\n\nThe way you have written this you assume that there is some trick to get a good score and you diminish the hard work people put into this competition. There is probably an overlap between H&M recommendation strategies and top positions strategies but the results are pretty much very random. Honestly I'm not looking for any trick just iterate as much as possible. My full process takes about 14-15 hours right now so I can afford to do 1 experiment per day mainly centered around creating strategies to generate candidates.\n\n> Not sure how so many people are getting CVs of 0.04 and LBs above 0.03, I feel like I've squeezed the lemon as much as is possible.\n\nI have 1.5+ billion candidates for submission. The dataframe with the candidates on 2020-09-22 approaches 200gb compressed right now. Recommendation problems are notorious for creating very big datasets. i recommend you start working on good ways to generate decent candidates.",
    "1756436": "alright  thanks for the tips\ndidn't mean it to sound diminishing or that there is some unfair \"trick\" , i just have terrible communication skills ^^",
    "1756527": "Do you mind sharing if you have a mostly different set of candidates for each customer while training or if you have the same set of candidates for everyone while training? \n\nI imagine it would be something like go week by week and use like popular items of previous week(s) based on gender and specific buying habits based on gender and age and such?",
    "1756559": "Some strategies use \"global\" popularity, some of them use other features of the user. It is a mix.\n\n> I imagine it would be something like go week by week and use like popular items of previous week(s) based on gender and specific buying habits based on gender and age and such?\n\nThis is exactly right.",
    "1761513": "good  job！",
    "1764235": "paweljankiewicz 1.5+ billion is an amazing num! If you don't mind, could I know whether you use all the candidates to build the training samples?",
    "1775097": "thanks for the tips again, did my first attempt with the lgbmranker and i could get really good results already by only using 50 or so candidates on average"
  },
  "source": "meta"
}