{
  "id": 307288,
  "title": "Addressing common questions and what the competition is really about",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288",
  "author_name": "Paweł Jankiewicz",
  "post_date": "2022-02-13T16:40:50.791000",
  "votes": 431,
  "comment_count": 78,
  "views": 0,
  "content": "<p>Just to give you some background - my team won RecSys 2019 so I'm pretty familiar with problems like these. Now that we have this out of the way…</p>\n<p>To simplify the discussion I propose to call the set of articles that a customer bought on a given date a <strong>basket</strong>. And then this problem can be really summarized as predict next basket.</p>\n<p>General approach towards problems like this could be as follows. At least mine will be.</p>\n<h2>1. For each customer generate candidates that they can buy.</h2>\n<p>These candidates are your negative and positive examples. I'm saying this because someone on the forum asked where are the negative observations. And the answer is you generate them. </p>\n<p>The strategies to generate those candidates are probably the most important aspect of this competition. The strategies can be based on:</p>\n<ul>\n<li>item to item similarities between customer previous baskets - so image similarity is not out of the question as somebody suggested</li>\n<li>user based collaborative filtering</li>\n<li>last baskets (with the hope that the user will buy them again)</li>\n<li>model based predictions with customers without the transaction history</li>\n<li>etc</li>\n</ul>\n<p>Coming up with different strategies is a fun challenge :). I cannot imagine that a winning solution will be based on a closed form model representing one strategy. Once you have all the different strategies ready for each customer then it is time to create a huge table with columns like these:</p>\n<ul>\n<li>observation_date - the date for which you are making predictions - submission should be based on all the data before 2020-09-22 (my birthday btw).</li>\n<li>customer_id</li>\n<li>article_id</li>\n<li>bought in the next week (this is your label)</li>\n<li>customer static attributes</li>\n<li>customer dynamic attributes</li>\n<li>article static attributes</li>\n<li>article dynamic attributes (maybe item popularity and such)</li>\n<li>whether the article exist in the strategies and which ones</li>\n<li>strategy score (whatever it is)</li>\n<li>many other features that relate the item to the user</li>\n</ul>\n<p>Once you have a table like this it all becomes very familiar. There were many competitions involving click through predictions. They have almost an identical structure the only difference is the source of observations.</p>\n<h2>2. Build a ranking model that ranks the items within Customer</h2>\n<p>Ranking is a different problem than classification but in general classification models can be used as well. My favourite ranking model is LightGBM (we used it in RecSys 2019) <a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html</a>.</p>\n<p>The customer id can be treated as a query in ranking model because you are interested in optimizing the scores at a customer level.</p>\n<h2>General tips</h2>\n<ol>\n<li><p>Sample the data as much as possible - most of the competition you should spend validating your features on the sample. Increase the sample you are using when you are no longer seeing the correlation between local validation and the leaderboard.</p></li>\n<li><p>Test different ideas quickly. Fast local validation is key here because you must do as much experiments as possible.</p></li>\n</ol>\n<p>Good luck :)</p>\n<p>Update 1. Added code for sampling here <a href=\"https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples\" target=\"_blank\">https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples</a></p>",
  "messages": [
    {
      "id": 1688473,
      "postDate": "2022-02-13T16:40:50.790Z",
      "content": "<p>Just to give you some background - my team won RecSys 2019 so I'm pretty familiar with problems like these. Now that we have this out of the way…</p>\n<p>To simplify the discussion I propose to call the set of articles that a customer bought on a given date a <strong>basket</strong>. And then this problem can be really summarized as predict next basket.</p>\n<p>General approach towards problems like this could be as follows. At least mine will be.</p>\n<h2>1. For each customer generate candidates that they can buy.</h2>\n<p>These candidates are your negative and positive examples. I'm saying this because someone on the forum asked where are the negative observations. And the answer is you generate them. </p>\n<p>The strategies to generate those candidates are probably the most important aspect of this competition. The strategies can be based on:</p>\n<ul>\n<li>item to item similarities between customer previous baskets - so image similarity is not out of the question as somebody suggested</li>\n<li>user based collaborative filtering</li>\n<li>last baskets (with the hope that the user will buy them again)</li>\n<li>model based predictions with customers without the transaction history</li>\n<li>etc</li>\n</ul>\n<p>Coming up with different strategies is a fun challenge :). I cannot imagine that a winning solution will be based on a closed form model representing one strategy. Once you have all the different strategies ready for each customer then it is time to create a huge table with columns like these:</p>\n<ul>\n<li>observation_date - the date for which you are making predictions - submission should be based on all the data before 2020-09-22 (my birthday btw).</li>\n<li>customer_id</li>\n<li>article_id</li>\n<li>bought in the next week (this is your label)</li>\n<li>customer static attributes</li>\n<li>customer dynamic attributes</li>\n<li>article static attributes</li>\n<li>article dynamic attributes (maybe item popularity and such)</li>\n<li>whether the article exist in the strategies and which ones</li>\n<li>strategy score (whatever it is)</li>\n<li>many other features that relate the item to the user</li>\n</ul>\n<p>Once you have a table like this it all becomes very familiar. There were many competitions involving click through predictions. They have almost an identical structure the only difference is the source of observations.</p>\n<h2>2. Build a ranking model that ranks the items within Customer</h2>\n<p>Ranking is a different problem than classification but in general classification models can be used as well. My favourite ranking model is LightGBM (we used it in RecSys 2019) <a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html\" target=\"_blank\">https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html</a>.</p>\n<p>The customer id can be treated as a query in ranking model because you are interested in optimizing the scores at a customer level.</p>\n<h2>General tips</h2>\n<ol>\n<li><p>Sample the data as much as possible - most of the competition you should spend validating your features on the sample. Increase the sample you are using when you are no longer seeing the correlation between local validation and the leaderboard.</p></li>\n<li><p>Test different ideas quickly. Fast local validation is key here because you must do as much experiments as possible.</p></li>\n</ol>\n<p>Good luck :)</p>\n<p>Update 1. Added code for sampling here <a href=\"https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples\" target=\"_blank\">https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples</a></p>",
      "rawMarkdown": "Just to give you some background - my team won RecSys 2019 so I'm pretty familiar with problems like these. Now that we have this out of the way...\n\nTo simplify the discussion I propose to call the set of articles that a customer bought on a given date a **basket**. And then this problem can be really summarized as predict next basket.\n\nGeneral approach towards problems like this could be as follows. At least mine will be.\n\n## 1. For each customer generate candidates that they can buy. \n\nThese candidates are your negative and positive examples. I'm saying this because someone on the forum asked where are the negative observations. And the answer is you generate them. \n\nThe strategies to generate those candidates are probably the most important aspect of this competition. The strategies can be based on:\n- item to item similarities between customer previous baskets - so image similarity is not out of the question as somebody suggested\n- user based collaborative filtering\n- last baskets (with the hope that the user will buy them again)\n- model based predictions with customers without the transaction history\n- etc\n\nComing up with different strategies is a fun challenge :). I cannot imagine that a winning solution will be based on a closed form model representing one strategy. Once you have all the different strategies ready for each customer then it is time to create a huge table with columns like these:\n\n- observation_date - the date for which you are making predictions - submission should be based on all the data before 2020-09-22 (my birthday btw).\n- customer_id\n- article_id\n- bought in the next week (this is your label)\n- customer static attributes\n- customer dynamic attributes\n- article static attributes\n- article dynamic attributes (maybe item popularity and such)\n- whether the article exist in the strategies and which ones\n- strategy score (whatever it is)\n- many other features that relate the item to the user\n\nOnce you have a table like this it all becomes very familiar. There were many competitions involving click through predictions. They have almost an identical structure the only difference is the source of observations.\n\n## 2. Build a ranking model that ranks the items within Customer \n\nRanking is a different problem than classification but in general classification models can be used as well. My favourite ranking model is LightGBM (we used it in RecSys 2019) https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html.\n\nThe customer id can be treated as a query in ranking model because you are interested in optimizing the scores at a customer level.\n\n## General tips\n\n1. Sample the data as much as possible - most of the competition you should spend validating your features on the sample. Increase the sample you are using when you are no longer seeing the correlation between local validation and the leaderboard.\n\n2. Test different ideas quickly. Fast local validation is key here because you must do as much experiments as possible.\n\nGood luck :)\n\nUpdate 1. Added code for sampling here https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples",
      "votes": 430
    },
    {
      "id": 1688506,
      "postDate": "2022-02-13T17:03:05.520Z",
      "content": "<p>For anyone that might not be familiar <a href=\"https://recsys.acm.org\" target=\"_blank\">RecSys</a> is one of the most respected competitions in the Recommender space. </p>\n<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> I was trying to find your winning solution but I could only find the press release content. Is it available anywhere? </p>\n<p>TIA! :)</p>",
      "rawMarkdown": "For anyone that might not be familiar [RecSys](https://recsys.acm.org) is one of the most respected competitions in the Recommender space. \n\n@paweljankiewicz I was trying to find your winning solution but I could only find the press release content. Is it available anywhere? \n\nTIA! :)",
      "votes": 9,
      "replies": [
        {
          "id": 1688518,
          "postDate": "2022-02-13T17:12:18.343Z",
          "content": "<p>Yeah here it is - <a href=\"https://github.com/logicai-io/recsys2019\" target=\"_blank\">https://github.com/logicai-io/recsys2019</a>.</p>",
          "rawMarkdown": "Yeah here it is - https://github.com/logicai-io/recsys2019.",
          "votes": 16
        },
        {
          "id": 1695896,
          "postDate": "2022-02-18T12:47:09.383Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> , i try to download the data in Recsys 2019 </p>\n<ul>\n<li>by code in recsys2019/data/download_data.sh. while it did not run. this is the response:<br>\n<code>--2022-02-18 20:38:52--  https://storage.googleapis.com/logicai-recsys2019/trivagoRecSysChallengeData2019_v2.zip\n正在解析主机 storage.googleapis.com (storage.googleapis.com)... 142.250.204.112, 142.250.204.48, 142.250.66.144, ...\n正在连接 storage.googleapis.com (storage.googleapis.com)|142.250.204.112|:443... 已连接。\n已发出 HTTP 请求，正在等待回应... 404 Not Found\n2022-02-18 20:38:54 错误 404：Not Found。</code></li>\n<li>through this url (<a href=\"https://recsys.trivago.cloud/challenge/dataset/\" target=\"_blank\">https://recsys.trivago.cloud/challenge/dataset/</a>)<br>\nbut it also failed.<br>\nCould I download the data in other way?<br>\nThank you so much</li>\n</ul>",
          "rawMarkdown": "Hi, @paweljankiewicz , i try to download the data in Recsys 2019 \n- by code in recsys2019/data/download_data.sh. while it did not run. this is the response:\n`--2022-02-18 20:38:52--  https://storage.googleapis.com/logicai-recsys2019/trivagoRecSysChallengeData2019_v2.zip\n正在解析主机 storage.googleapis.com (storage.googleapis.com)... 142.250.204.112, 142.250.204.48, 142.250.66.144, ...\n正在连接 storage.googleapis.com (storage.googleapis.com)|142.250.204.112|:443... 已连接。\n已发出 HTTP 请求，正在等待回应... 404 Not Found\n2022-02-18 20:38:54 错误 404：Not Found。`\n- through this url (https://recsys.trivago.cloud/challenge/dataset/)\nbut it also failed.\nCould I download the data in other way?\nThank you so much",
          "votes": 1
        },
        {
          "id": 1696209,
          "postDate": "2022-02-18T16:53:56.520Z",
          "content": "<p>I don't have the original dataset. Also the code is written in such a way that it needs close to 500gb of ram to train the models.</p>",
          "rawMarkdown": "I don't have the original dataset. Also the code is written in such a way that it needs close to 500gb of ram to train the models.",
          "votes": 5
        },
        {
          "id": 1696214,
          "postDate": "2022-02-18T16:57:10.007Z",
          "content": "<p>For anyone else surprised by:</p>\n<blockquote>\n  <p>Also the code is written in such a way that it needs close to 500gb of ram to train the models.</p>\n</blockquote>\n<p>When I had the opportunity to do an interview for another <a href=\"https://www.youtube.com/watch?v=W3aWEXqIkWk\" target=\"_blank\">RecSys winning team</a>, I learned this is quite a common number for such problems. 😅</p>\n<p>The embedding sizes based on number of users and recommendations just eats up A LOT of memory by the nature of their size</p>",
          "rawMarkdown": "For anyone else surprised by:\n\n> Also the code is written in such a way that it needs close to 500gb of ram to train the models.\n\nWhen I had the opportunity to do an interview for another [RecSys winning team](https://www.youtube.com/watch?v=W3aWEXqIkWk), I learned this is quite a common number for such problems. 😅\n\nThe embedding sizes based on number of users and recommendations just eats up A LOT of memory by the nature of their size",
          "votes": 7
        },
        {
          "id": 1696389,
          "postDate": "2022-02-18T19:52:49.823Z",
          "content": "<p>BTW I found the data here on Kaggle (but don't seem to have the link to it now), it didn't seem to be available anywhere else.</p>\n<p>But I got the most mileage out of reading the code in the repo, didn't get around to doing much with the generated data.</p>",
          "rawMarkdown": "BTW I found the data here on Kaggle (but don't seem to have the link to it now), it didn't seem to be available anywhere else.\n\nBut I got the most mileage out of reading the code in the repo, didn't get around to doing much with the generated data.",
          "votes": 1
        },
        {
          "id": 1696410,
          "postDate": "2022-02-18T20:07:22.843Z",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I believe <a href=\"https://www.kaggle.com/pranavmahajan725/trivagorecsyschallengedata2019\" target=\"_blank\">this</a> is the link to the dataset?</p>",
          "rawMarkdown": "@radek1 I believe [this](https://www.kaggle.com/pranavmahajan725/trivagorecsyschallengedata2019) is the link to the dataset?",
          "votes": 2
        },
        {
          "id": 1696492,
          "postDate": "2022-02-18T22:30:40.160Z",
          "content": "<p>yes, you found it! :) I think it was that one</p>\n<p>with the data I downloaded the preprocessing steps ran without issue</p>",
          "rawMarkdown": "yes, you found it! :) I think it was that one\n\nwith the data I downloaded the preprocessing steps ran without issue"
        },
        {
          "id": 1696542,
          "postDate": "2022-02-18T23:16:10.543Z",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> the most important concept in RecSys 2019 was the notion of accumulators classes which have 2 methods:</p>\n<ul>\n<li>update - updates the state of accumulator</li>\n<li>get - extracts the state of accumulator</li>\n</ul>\n<p>When you process events by date you first extract the information from the accumulator and then update it with new information. This way there is no possibility to leak any info. The accumulators themselves can track information about the item, customer, item+customer, other groups of information. Really anything you need.</p>\n<p>The code could be summarized as follows:</p>\n<pre><code>for transaction in transactions:\n      for accumulator in accumulators:\n             get features from accumulator - really important that it is before update\n             update the accumulator\n</code></pre>\n<p>There is an opportunity to parallelize the extraction by reversing the loop so you can process and extract features from accumulators in parallel.</p>",
          "rawMarkdown": "@radek1 @init27 the most important concept in RecSys 2019 was the notion of accumulators classes which have 2 methods:\n\n- update - updates the state of accumulator\n- get - extracts the state of accumulator\n\nWhen you process events by date you first extract the information from the accumulator and then update it with new information. This way there is no possibility to leak any info. The accumulators themselves can track information about the item, customer, item+customer, other groups of information. Really anything you need.\n\nThe code could be summarized as follows:\n\n```python\nfor transaction in transactions:\n      for accumulator in accumulators:\n             get features from accumulator - really important that it is before update\n             update the accumulator\n```\n\nThere is an opportunity to parallelize the extraction by reversing the loop so you can process and extract features from accumulators in parallel.",
          "votes": 4
        },
        {
          "id": 1696572,
          "postDate": "2022-02-19T00:18:28.483Z",
          "content": "<p>Ah, I see! <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> thank you very much for sharing this, this is wonderful! 😊</p>\n<p>If I am reading this right, the accumulators should be anything -- items of the given type purchased by a customer up to this point, count of transactions, etc.</p>\n<p>I have completely not thought of processing data like this, this is very cool, eager to give it a try!</p>\n<p>So far, I have been attempting to generate features (and candidate transactions) using dataframes, groupby, shift, and merging. But that takes a ridiculous amount of RAM and the code is very hard to reason about and prone to having bugs!</p>\n<p>Would love to go back to straight-up Python for more of that :) </p>\n<p>I am not sure I understand the parallelization bit, where we reverse the loop? There is a first pass through the dataset that we cannot parallelize, to compute the accumulators. And we can only parallelize adding features to records in the second pass? (split the df into n_chunks, apply features from accumulator to each chunk, combine?) I am thinking I might be missing something here :) I am also thinking we probably want the accumulators to be simple so that they can only be applied as they go through the list of data in chronological order -- I guess I am really lost on the parallelization bit :) </p>\n<p>Once we get all these features, do we do feature selection at all? Or do we just see how correlated columns are and drop the ones that don't seem to add a lot of signal?</p>\n<p>I saw in the repo you shared that it seems models were run on different train sets? (maybe I misread this). Would you grab some numbers of columns and train different models on different subsets? I am wondering if I didn't misread this as I think there is a way to make lgbm randomly sample a subset of columns to grow each tree, but maybe such an approach is nice for blending?</p>\n<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> your posts are legendary and this thread is an absolute treasure trove of information for people just getting started with recsys, thank you so much for all your help!!! 🙏</p>",
          "rawMarkdown": "Ah, I see! @paweljankiewicz thank you very much for sharing this, this is wonderful! 😊\n\nIf I am reading this right, the accumulators should be anything -- items of the given type purchased by a customer up to this point, count of transactions, etc.\n\nI have completely not thought of processing data like this, this is very cool, eager to give it a try!\n\nSo far, I have been attempting to generate features (and candidate transactions) using dataframes, groupby, shift, and merging. But that takes a ridiculous amount of RAM and the code is very hard to reason about and prone to having bugs!\n\nWould love to go back to straight-up Python for more of that :) \n\nI am not sure I understand the parallelization bit, where we reverse the loop? There is a first pass through the dataset that we cannot parallelize, to compute the accumulators. And we can only parallelize adding features to records in the second pass? (split the df into n_chunks, apply features from accumulator to each chunk, combine?) I am thinking I might be missing something here :) I am also thinking we probably want the accumulators to be simple so that they can only be applied as they go through the list of data in chronological order -- I guess I am really lost on the parallelization bit :) \n\nOnce we get all these features, do we do feature selection at all? Or do we just see how correlated columns are and drop the ones that don't seem to add a lot of signal?\n\nI saw in the repo you shared that it seems models were run on different train sets? (maybe I misread this). Would you grab some numbers of columns and train different models on different subsets? I am wondering if I didn't misread this as I think there is a way to make lgbm randomly sample a subset of columns to grow each tree, but maybe such an approach is nice for blending?\n\n@paweljankiewicz your posts are legendary and this thread is an absolute treasure trove of information for people just getting started with recsys, thank you so much for all your help!!! 🙏",
          "votes": 1
        },
        {
          "id": 1696593,
          "postDate": "2022-02-19T00:33:44.940Z",
          "content": "<p>I think I might have understood it now! :)</p>\n<p>We don't reverse the loop in the sense that we run from end to finish, we just reverse it by moving the inside of the loop outside, like this:</p>\n<pre><code>for accumulator in accumulators:\n  for transaction in transactions:\n    run acc and apply features\n</code></pre>\n<p>We can walk the column/dataframe in parallel by multiple accumulators and then combine the results!</p>\n<p>I would love to try this with something like <code>dask</code> and <code>delayed</code>, I suspect this might be a nice way to work around the limitations of <code>dask's</code> API, really curious how well this could work :)</p>",
          "rawMarkdown": "I think I might have understood it now! :)\n\nWe don't reverse the loop in the sense that we run from end to finish, we just reverse it by moving the inside of the loop outside, like this:\n\n```\nfor accumulator in accumulators:\n  for transaction in transactions:\n    run acc and apply features\n```\n\nWe can walk the column/dataframe in parallel by multiple accumulators and then combine the results!\n\nI would love to try this with something like `dask` and `delayed`, I suspect this might be a nice way to work around the limitations of `dask's` API, really curious how well this could work :)",
          "votes": 1
        },
        {
          "id": 1697962,
          "postDate": "2022-02-20T02:37:40.973Z",
          "content": "<p>Great work!  Respect to all of you!</p>",
          "rawMarkdown": "Great work!  Respect to all of you!"
        },
        {
          "id": 1717998,
          "postDate": "2022-03-10T12:21:07.620Z",
          "content": "<p>So is it true that if I don't have a machine with lots of memory, I cannot become one of the winners😂</p>",
          "rawMarkdown": "So is it true that if I don't have a machine with lots of memory, I cannot become one of the winners😂"
        }
      ]
    },
    {
      "id": 1776281,
      "postDate": "2022-05-03T21:35:19.013Z",
      "content": "<p>Hi, thank you so much for this post, it is very useful. I had one question about the generation of the positive examples you use to train the ranking model. Do you use the same method for finding these as with the final test set which means there are missing positive examples or are you including all positive examples as they are known in the training data?</p>",
      "rawMarkdown": "Hi, thank you so much for this post, it is very useful. I had one question about the generation of the positive examples you use to train the ranking model. Do you use the same method for finding these as with the final test set which means there are missing positive examples or are you including all positive examples as they are known in the training data?",
      "votes": 3,
      "replies": [
        {
          "id": 1776316,
          "postDate": "2022-05-03T22:24:40.767Z",
          "content": "<p>That's a good question. I tried adding all positives but the score didn't improve. I'm using only positives which are associated with some strategy.</p>",
          "rawMarkdown": "That's a good question. I tried adding all positives but the score didn't improve. I'm using only positives which are associated with some strategy.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1778497,
      "postDate": "2022-05-05T11:20:42.293Z",
      "content": "<p>Thank you for your posting. This post is very helpful.<br>\nI would like to ask you positive-negative sample rate during training process.<br>\nYou said that about 1000 candidates are collected for each customers in your strategy.<br>\nBut, when training model, this is unbalanced-data. I think down-sampling is necessary.<br>\nIf you don't mind, please tell me your positive-negative rate.</p>",
      "rawMarkdown": "Thank you for your posting. This post is very helpful.\nI would like to ask you positive-negative sample rate during training process.\nYou said that about 1000 candidates are collected for each customers in your strategy.\nBut, when training model, this is unbalanced-data. I think down-sampling is necessary.\nIf you don't mind, please tell me your positive-negative rate.",
      "votes": 1,
      "replies": [
        {
          "id": 1778588,
          "postDate": "2022-05-05T13:30:48.177Z",
          "content": "<p>You are right. At the time of writing I thought that increasing the number of candidates can be good. Then I got to a point where adding more candidates doesn't help. I have way too many candidates right now. During training I take only 5-10% of negative examples. <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> commented that he is using only 200 candidates which makes sense. I think most of the true positives that we are able to predict are just popular items. With 5% downsampling I have about 4% positives.</p>",
          "rawMarkdown": "You are right. At the time of writing I thought that increasing the number of candidates can be good. Then I got to a point where adding more candidates doesn't help. I have way too many candidates right now. During training I take only 5-10% of negative examples. @lihaorocky commented that he is using only 200 candidates which makes sense. I think most of the true positives that we are able to predict are just popular items. With 5% downsampling I have about 4% positives.",
          "votes": 3
        },
        {
          "id": 1778653,
          "postDate": "2022-05-05T14:58:26.153Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> ;) I'm optimizing number of candidates with Bayesian Optimization. For every candidate generation method I have a <code>k</code> parameter determining how many articles can be returned. Then I optimize  Precision@50 for a baseline model (no reranking model) where the <code>k's</code> are hyperparameters. After a while you can clearly see which generators should be switched off or limited.</p>",
          "rawMarkdown": "Hi @paweljankiewicz ;) I'm optimizing number of candidates with Bayesian Optimization. For every candidate generation method I have a `k` parameter determining how many articles can be returned. Then I optimize ~~Recall@50~~ Precision@50 for a baseline model (no reranking model) where the `k's` are hyperparameters. After a while you can clearly see which generators should be switched off or limited."
        },
        {
          "id": 1779241,
          "postDate": "2022-05-06T08:23:02.900Z",
          "content": "<p>Interesting approach. I think that this competition shows very well that most cf recommendation algorithms are almost useless in this setting. It is more like a symbolic recommendation system where you don't rely on particular items but rather pools of items created from recent transaction history. So recommendation can be based on:</p>\n<p>0.5<em>top_50_articles_department + 0.5</em>most_similar_articles_to_the_last_basket_that_are_within_1000_most_popular_articles</p>\n<p>There are some obvious flaws in the competition like guessing the availability of items but it is a valuable experience nonetheless.</p>",
          "rawMarkdown": "Interesting approach. I think that this competition shows very well that most cf recommendation algorithms are almost useless in this setting. It is more like a symbolic recommendation system where you don't rely on particular items but rather pools of items created from recent transaction history. So recommendation can be based on:\n\n0.5*top_50_articles_department + 0.5*most_similar_articles_to_the_last_basket_that_are_within_1000_most_popular_articles\n\nThere are some obvious flaws in the competition like guessing the availability of items but it is a valuable experience nonetheless.\n\n",
          "votes": 3
        },
        {
          "id": 1779443,
          "postDate": "2022-05-06T12:57:13.977Z",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> - thanks for sharing that!</p>\n<p>What does <code>top_50_articles_department</code> mean?</p>",
          "rawMarkdown": "@paweljankiewicz - thanks for sharing that!\n\nWhat does `top_50_articles_department` mean?"
        },
        {
          "id": 1779470,
          "postDate": "2022-05-06T13:25:40.517Z",
          "content": "<p>It is just an example of some strategy. In this case it would be something like: take popular items from some specific item category that the user has in the historical transactions. </p>",
          "rawMarkdown": "It is just an example of some strategy. In this case it would be something like: take popular items from some specific item category that the user has in the historical transactions. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1767890,
      "postDate": "2022-04-25T18:37:44.290Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>, very very helpful post. Thanks </p>\n<p>Just one question, what do you mean with “model based predictions with customers without the transaction history” ?</p>",
      "rawMarkdown": "Hi @paweljankiewicz, very very helpful post. Thanks \n\nJust one question, what do you mean with “model based predictions with customers without the transaction history” ?\n",
      "votes": 1
    },
    {
      "id": 1761945,
      "postDate": "2022-04-20T10:22:19.840Z",
      "content": "<p>Hi Pawl, </p>\n<p>thanks for your informative post. I would like to ask you what do you mean by dynamic/static attributes for customer and articles? if illustrated with an example will be more helpful. Thanks!</p>",
      "rawMarkdown": "Hi Pawl, \n\nthanks for your informative post. I would like to ask you what do you mean by dynamic/static attributes for customer and articles? if illustrated with an example will be more helpful. Thanks!",
      "votes": 1
    },
    {
      "id": 1742077,
      "postDate": "2022-04-01T12:44:09.713Z",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nThanks for great idea!<br>\nI have a question.<br>\nWhat does observation_date exactly mean? Please give me some example.</p>",
      "rawMarkdown": "@paweljankiewicz \nThanks for great idea!\nI have a question.\nWhat does observation_date exactly mean? Please give me some example.",
      "votes": 1,
      "replies": [
        {
          "id": 1744357,
          "postDate": "2022-04-03T22:50:51.283Z",
          "content": "<p>Observation date is the date you treat as \"now\" so you can use all the data before this date. So for validation you need to take a date like 2020-09-15. For submission the observation date is of course the last available date 2020-09-22.</p>",
          "rawMarkdown": "Observation date is the date you treat as \"now\" so you can use all the data before this date. So for validation you need to take a date like 2020-09-15. For submission the observation date is of course the last available date 2020-09-22.",
          "votes": 2
        },
        {
          "id": 1750050,
          "postDate": "2022-04-09T09:14:06.753Z",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nThanks! I have completely understood.</p>",
          "rawMarkdown": "@paweljankiewicz \nThanks! I have completely understood."
        }
      ]
    },
    {
      "id": 1700846,
      "postDate": "2022-02-22T10:00:18.293Z",
      "content": "<p>Thank you for sharing your idea. This is really a great approach. I am getting \"run out of memory\" problem. I am using my local PC which has 16GB RAM. Can you suggest how to deal with this situation. Thank You! </p>",
      "rawMarkdown": "Thank you for sharing your idea. This is really a great approach. I am getting \"run out of memory\" problem. I am using my local PC which has 16GB RAM. Can you suggest how to deal with this situation. Thank You! ",
      "votes": 1,
      "replies": [
        {
          "id": 1700865,
          "postDate": "2022-02-22T10:18:45.013Z",
          "content": "<p>16GB RAM is probably not enough to comfortably work with this dataset - at least the approach I suggested needs a lot more memory. You can experiment on a smaller dataset <a href=\"https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples\" target=\"_blank\">https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples</a>. Even using 5% sample should give you similar performance to 100% on this data. </p>",
          "rawMarkdown": "16GB RAM is probably not enough to comfortably work with this dataset - at least the approach I suggested needs a lot more memory. You can experiment on a smaller dataset https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples. Even using 5% sample should give you similar performance to 100% on this data. ",
          "votes": 5
        },
        {
          "id": 1701108,
          "postDate": "2022-02-22T14:24:37.200Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        },
        {
          "id": 1704464,
          "postDate": "2022-02-25T14:47:55.280Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> :)</p>\n<p>Thank you for sharing</p>\n<p>You might find this code useful to make a stratified Sample using many columns.</p>\n<p><a href=\"https://github.com/flaboss/python_stratified_sampling/blob/master/stratifiedSample.py\" target=\"_blank\">https://github.com/flaboss/python_stratified_sampling/blob/master/stratifiedSample.py</a></p>\n<p>All the best </p>",
          "rawMarkdown": "Hi @paweljankiewicz :)\n\nThank you for sharing\n\nYou might find this code useful to make a stratified Sample using many columns.\n\nhttps://github.com/flaboss/python_stratified_sampling/blob/master/stratifiedSample.py\n\nAll the best ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1698165,
      "postDate": "2022-02-20T07:42:19.007Z",
      "content": "<p>I am new to the recommender system, do you have any references for me? Thanks.</p>",
      "rawMarkdown": "I am new to the recommender system, do you have any references for me? Thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 1699049,
          "postDate": "2022-02-20T21:36:41.370Z",
          "content": "<p>I suggest to start with the most famous datasets for recommendations - something like movielens - <a href=\"https://grouplens.org/datasets/movielens/\" target=\"_blank\">https://grouplens.org/datasets/movielens/</a>. Follow some tutorial using these datasets. This competition is on the harder spectrum because it is not a very typical dataset. I remember watching a part of this course <a href=\"https://www.coursera.org/specializations/recommender-systems\" target=\"_blank\">https://www.coursera.org/specializations/recommender-systems</a> and it was decent.</p>",
          "rawMarkdown": "I suggest to start with the most famous datasets for recommendations - something like movielens - https://grouplens.org/datasets/movielens/. Follow some tutorial using these datasets. This competition is on the harder spectrum because it is not a very typical dataset. I remember watching a part of this course https://www.coursera.org/specializations/recommender-systems and it was decent.",
          "votes": 12
        }
      ]
    },
    {
      "id": 1696455,
      "postDate": "2022-02-18T21:03:01.053Z",
      "content": "<p>how could you create negative samples? </p>",
      "rawMarkdown": "how could you create negative samples? ",
      "votes": 1,
      "replies": [
        {
          "id": 1696519,
          "postDate": "2022-02-18T22:52:07.610Z",
          "content": "<p>As I have written there are many ways to create negative samples. For example items in the previous basket can be a source of negative samples. Popular items can be too. There are countless of examples.</p>",
          "rawMarkdown": "As I have written there are many ways to create negative samples. For example items in the previous basket can be a source of negative samples. Popular items can be too. There are countless of examples.",
          "votes": 2
        },
        {
          "id": 1715321,
          "postDate": "2022-03-07T21:50:41.483Z",
          "content": "<p>Thanks for the inspiring post <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> It is very helpful especially on the negative sampling. <br>\nI have another question about negative sampling. Shall we sample from items which are not similar to positive samples as negative samples? I am asking because it may confuse the model if we sample popular or previously purchased items as negative samples. They could be potentially positive samples in the test set. On the other hand, they could represent an interest shift of the customers. Do you have suggestions from your experience? Thank you ahead. </p>",
          "rawMarkdown": "Thanks for the inspiring post @paweljankiewicz It is very helpful especially on the negative sampling. \nI have another question about negative sampling. Shall we sample from items which are not similar to positive samples as negative samples? I am asking because it may confuse the model if we sample popular or previously purchased items as negative samples. They could be potentially positive samples in the test set. On the other hand, they could represent an interest shift of the customers. Do you have suggestions from your experience? Thank you ahead. "
        },
        {
          "id": 1717420,
          "postDate": "2022-03-09T23:19:48.607Z",
          "content": "<p>He said 'use items in the previous basket', it might mean 'user other customer's basket as the negative samples for the current customer', make sense? I think it's ok.</p>\n<p>But what if these two customers bought the same items, does it confuse the model? </p>",
          "rawMarkdown": "He said 'use items in the previous basket', it might mean 'user other customer's basket as the negative samples for the current customer', make sense? I think it's ok.\n\nBut what if these two customers bought the same items, does it confuse the model? "
        }
      ]
    },
    {
      "id": 1707473,
      "postDate": "2022-02-28T13:49:40.003Z",
      "content": "<p>Pawel, thanks so much for putting this together, this is awesome! 🙏</p>\n<p>I have a question for you about search strategies. Is there any reason why, instead of separately implementing a search strategy and inserting the result as a column in the training set, it couldn't also work to just directly add the variables that the strategy depends on into the training set? </p>\n<p>So for example if strategy S is dependent on 10 variables x1, x2… x10, instead of adding S as a column to the training set, adding x1 - x10. Although I suppose this would only be reasonable for simpler strategies and wouldn't work for strategies like image detection.</p>",
      "rawMarkdown": "Pawel, thanks so much for putting this together, this is awesome! 🙏\n\nI have a question for you about search strategies. Is there any reason why, instead of separately implementing a search strategy and inserting the result as a column in the training set, it couldn't also work to just directly add the variables that the strategy depends on into the training set? \n\nSo for example if strategy S is dependent on 10 variables x1, x2... x10, instead of adding S as a column to the training set, adding x1 - x10. Although I suppose this would only be reasonable for simpler strategies and wouldn't work for strategies like image detection.",
      "votes": 2,
      "replies": [
        {
          "id": 1707500,
          "postDate": "2022-02-28T14:14:09.507Z",
          "content": "<p>This is a good question. I think there are several problems</p>\n<p>1) For each customer you must consider 100k items if you really want to include all variables with which the strategy was created.<br>\n2) What would you do in a case that you want to relate a set of items from the history with those candidates. The strategy can really on more complex feature set than x1,…,x10.</p>\n<p>A search strategy is a way to reduce the number of candidates that you consider for each customer.</p>",
          "rawMarkdown": "This is a good question. I think there are several problems\n\n1) For each customer you must consider 100k items if you really want to include all variables with which the strategy was created.\n2) What would you do in a case that you want to relate a set of items from the history with those candidates. The strategy can really on more complex feature set than x1,...,x10.\n\nA search strategy is a way to reduce the number of candidates that you consider for each customer.",
          "votes": 1
        },
        {
          "id": 1707630,
          "postDate": "2022-02-28T16:23:25.903Z",
          "content": "<p>That makes sense to me. Thanks!</p>",
          "rawMarkdown": "That makes sense to me. Thanks!"
        }
      ]
    },
    {
      "id": 1698993,
      "postDate": "2022-02-20T20:11:23.767Z",
      "content": "<p>As I keep working on this and learning about it, I keep coming to this thread and rereading it. This is such a wealth of information!</p>\n<p>Thank you again <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>!</p>\n<p>What blew my mind today was using ‘customer_id’ as a query! I thought we might want to query on [‘customer_id’, ‘observation_date’] to essentially query on baskets?</p>\n<p>Also, the list of transactions can be a good staying set of positive examples? Meaning, from the post it might seem like we need to do something extra (‘the label indicates whether the article will be purchased next week’) but this is essentially what the list of transactions does? It is an observation date and if we add an indicator column such as ‘purchased’ we will have everything that we need? I am just worried I might be missing something subtle here.</p>",
      "rawMarkdown": "As I keep working on this and learning about it, I keep coming to this thread and rereading it. This is such a wealth of information!\n\nThank you again @paweljankiewicz!\n\nWhat blew my mind today was using ‘customer_id’ as a query! I thought we might want to query on [‘customer_id’, ‘observation_date’] to essentially query on baskets?\n\nAlso, the list of transactions can be a good staying set of positive examples? Meaning, from the post it might seem like we need to do something extra (‘the label indicates whether the article will be purchased next week’) but this is essentially what the list of transactions does? It is an observation date and if we add an indicator column such as ‘purchased’ we will have everything that we need? I am just worried I might be missing something subtle here.",
      "votes": 2,
      "replies": [
        {
          "id": 1699032,
          "postDate": "2022-02-20T21:09:57.727Z",
          "content": "<p>No you are right for more than 1 observation date the query should be customer_id + observation_dt. Because I have everything sorted by date I didn't notice that I'm implicitly including observation_dt in the query.</p>",
          "rawMarkdown": "No you are right for more than 1 observation date the query should be customer_id + observation_dt. Because I have everything sorted by date I didn't notice that I'm implicitly including observation_dt in the query.",
          "votes": 2
        },
        {
          "id": 1699040,
          "postDate": "2022-02-20T21:18:58.783Z",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I think you are missing something else. When you generate candidates, it is not 100% sure that you will find any transactions. Let's say your only strategy is to include the last basket items. If you have a customer which didn't buy anything twice then you can discard whole customer + observation_dt. The way the ranking models work they need positive examples in the query. This is a source of an important optimization in my opinion.</p>\n<p>The process should be like this: you generate a candidates set something like <code>Dict[Strategy, List[ArticleId]]</code>. A unique set of ArticleId is your number of observations that you have to consider for the customer. </p>\n<p>As I have said for training purposes you can discard all the queries without positive examples. But for validation and submission you must include all of your candidates.</p>",
          "rawMarkdown": "@radek1 I think you are missing something else. When you generate candidates, it is not 100% sure that you will find any transactions. Let's say your only strategy is to include the last basket items. If you have a customer which didn't buy anything twice then you can discard whole customer + observation_dt. The way the ranking models work they need positive examples in the query. This is a source of an important optimization in my opinion.\n\nThe process should be like this: you generate a candidates set something like `Dict[Strategy, List[ArticleId]]`. A unique set of ArticleId is your number of observations that you have to consider for the customer. \n\nAs I have said for training purposes you can discard all the queries without positive examples. But for validation and submission you must include all of your candidates.",
          "votes": 3
        },
        {
          "id": 1699125,
          "postDate": "2022-02-21T00:08:45.827Z",
          "content": "<p>Thank you so much for these additional details, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>! 😊 All this is of great help</p>",
          "rawMarkdown": "Thank you so much for these additional details, @paweljankiewicz! 😊 All this is of great help"
        }
      ]
    },
    {
      "id": 1758919,
      "postDate": "2022-04-18T06:50:27.390Z",
      "content": "<p>nice work!!</p>",
      "rawMarkdown": "nice work!!"
    },
    {
      "id": 1757936,
      "postDate": "2022-04-17T06:43:41.253Z",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nInspired by you, I have been working on rank model these days.<br>\nI have a question about mixing several strategy to make candidates. How can we put \"strategy score\" of candidate from different strategy together? Or make a column for each strategy?<br>\nTake an example of mixing \"recent transaction strategy\"(strategy1) and \"LSTM strategy\"(strategy2). I can simply make candidates of [customer, article, strategy1: bool, strategy2: bool] columns. <br>\nBUT, I want to use features from strategies, like \"transaction count\" from strategy1 or \"LSTM's probability\" from strategy2. I think this is \"strategy score\".</p>",
      "rawMarkdown": "@paweljankiewicz \nInspired by you, I have been working on rank model these days.\nI have a question about mixing several strategy to make candidates. How can we put \"strategy score\" of candidate from different strategy together? Or make a column for each strategy?\nTake an example of mixing \"recent transaction strategy\"(strategy1) and \"LSTM strategy\"(strategy2). I can simply make candidates of [customer, article, strategy1: bool, strategy2: bool] columns. \nBUT, I want to use features from strategies, like \"transaction count\" from strategy1 or \"LSTM's probability\" from strategy2. I think this is \"strategy score\".\n",
      "replies": [
        {
          "id": 1758243,
          "postDate": "2022-04-17T13:40:48.827Z",
          "content": "<p>I think it is worth having boolean flags for each strategy and of course you can use additional features from the strategy like counts and probabilities. The only thing you have to watch for is data leakage, so make sure you use only the data before the observation date to train your models.</p>",
          "rawMarkdown": "I think it is worth having boolean flags for each strategy and of course you can use additional features from the strategy like counts and probabilities. The only thing you have to watch for is data leakage, so make sure you use only the data before the observation date to train your models.",
          "votes": 2
        },
        {
          "id": 1758408,
          "postDate": "2022-04-17T16:43:06.213Z",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nThanks for reply!<br>\nI understand that to avoid leakage is important.<br>\nI have adopted naive strategies, but will develop some elaborate strategies from now on and mix them.</p>",
          "rawMarkdown": "@paweljankiewicz \nThanks for reply!\nI understand that to avoid leakage is important.\nI have adopted naive strategies, but will develop some elaborate strategies from now on and mix them."
        }
      ]
    },
    {
      "id": 1747328,
      "postDate": "2022-04-06T15:18:31.213Z",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> - I'm using LGBMRanker, and printing out the ndcg at each step, and I'm finding that as the model trains, it's getting worse, even on the training set!</p>\n<p>Same when using <code>map</code>.</p>",
      "rawMarkdown": "@paweljankiewicz - I'm using LGBMRanker, and printing out the ndcg at each step, and I'm finding that as the model trains, it's getting worse, even on the training set!\n\nSame when using `map`.",
      "replies": [
        {
          "id": 1747656,
          "postDate": "2022-04-06T20:58:00.873Z",
          "content": "<p>That's really strange. You probably have some sort of a bug. Are you sure you are providing a proper group parameter? I can vouch for LGBMRanker. </p>",
          "rawMarkdown": "That's really strange. You probably have some sort of a bug. Are you sure you are providing a proper group parameter? I can vouch for LGBMRanker. ",
          "votes": 1
        },
        {
          "id": 1747673,
          "postDate": "2022-04-06T21:36:46.240Z",
          "content": "<p>We had the same problem and so far I couldn't find a solution for it. Not sure but I think it might be related to the way we are building our dataset which Idk if it's right. We are considering 1 week only as the label and all data before this week as train. Maybe there is a better way to carry time information than that.</p>",
          "rawMarkdown": "We had the same problem and so far I couldn't find a solution for it. Not sure but I think it might be related to the way we are building our dataset which Idk if it's right. We are considering 1 week only as the label and all data before this week as train. Maybe there is a better way to carry time information than that.\n\n\n"
        },
        {
          "id": 1748364,
          "postDate": "2022-04-07T14:11:57.113Z",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> - thanks for letting me know that it shouldn't be happening, I'll keep digging until I find out what's wrong.</p>\n<p>I went back and checked, the group parameter looks good.</p>\n<p>Maybe it's because I haven't added customer features yet, only article and customer-article ones?</p>",
          "rawMarkdown": "@paweljankiewicz - thanks for letting me know that it shouldn't be happening, I'll keep digging until I find out what's wrong.\n\nI went back and checked, the group parameter looks good.\n\nMaybe it's because I haven't added customer features yet, only article and customer-article ones?"
        },
        {
          "id": 1748384,
          "postDate": "2022-04-07T14:37:02.650Z",
          "content": "<p><a href=\"https://www.kaggle.com/igormunizims\" target=\"_blank\">@igormunizims</a>, I'm also only considering one week as label for starters.</p>\n<p>I don't see why that should be a problem, though.<br>\nWhen you have very little data, you overfit quickly, but it should still improve on the training set. </p>",
          "rawMarkdown": "@igormunizims, I'm also only considering one week as label for starters.\n\nI don't see why that should be a problem, though.\nWhen you have very little data, you overfit quickly, but it should still improve on the training set. ",
          "votes": 1
        },
        {
          "id": 1748428,
          "postDate": "2022-04-07T15:27:42.110Z",
          "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> there are some optimizations you can do when you are training the ranking model. For example you can consider only the queries with at least 1 positive example. For the ranking model the queries with all negative observations are useless. If you didn't remove them maybe you can check if this is the reason why the ranking model is behaving in a strange way. It may be that after removing all such observations the number of observations is really low and it does indeed is overfiiting a lot.</p>\n<p>Unfortunately when you are predicting for submission you must use all the observations which is painful (in my case I need to predict 100gb of compressed dataframes and it is the slowest part of the process).</p>",
          "rawMarkdown": "@jacob34 there are some optimizations you can do when you are training the ranking model. For example you can consider only the queries with at least 1 positive example. For the ranking model the queries with all negative observations are useless. If you didn't remove them maybe you can check if this is the reason why the ranking model is behaving in a strange way. It may be that after removing all such observations the number of observations is really low and it does indeed is overfiiting a lot.\n\nUnfortunately when you are predicting for submission you must use all the observations which is painful (in my case I need to predict 100gb of compressed dataframes and it is the slowest part of the process).",
          "votes": 5
        },
        {
          "id": 1748487,
          "postDate": "2022-04-07T16:29:31.913Z",
          "content": "<p>Yes, I did follow your advice on that.</p>\n<p>For a single fold, I have 390k rows, for 12k customers, with 15k positives.<br>\nI realize it may be overfitting, but evaluation metric should still go down on the training set, no?</p>\n<p>Wow - 100gb is massive.</p>",
          "rawMarkdown": "Yes, I did follow your advice on that.\n\nFor a single fold, I have 390k rows, for 12k customers, with 15k positives.\nI realize it may be overfitting, but evaluation metric should still go down on the training set, no?\n\nWow - 100gb is massive.",
          "votes": 1
        },
        {
          "id": 1748580,
          "postDate": "2022-04-07T18:15:37.703Z",
          "content": "<p>Yes. It should go down. Are the results when you are checking the metric manually the same?</p>",
          "rawMarkdown": "Yes. It should go down. Are the results when you are checking the metric manually the same?",
          "votes": 1
        },
        {
          "id": 1759512,
          "postDate": "2022-04-18T17:30:53.483Z",
          "content": "<p>Yes, checking manually gives me the same.</p>\n<p>I find that after a few steps, it starts improving again (although never getting as good at it originally was), and at that point, the different evaluation metrics correlate, and cv correlates with LB, so I'm working with that.</p>\n<p>Still not sure why it's happening, but I suspect it's because of weak features and small training set size.</p>",
          "rawMarkdown": "Yes, checking manually gives me the same.\n\nI find that after a few steps, it starts improving again (although never getting as good at it originally was), and at that point, the different evaluation metrics correlate, and cv correlates with LB, so I'm working with that.\n\nStill not sure why it's happening, but I suspect it's because of weak features and small training set size."
        },
        {
          "id": 1777731,
          "postDate": "2022-05-04T18:35:47.290Z",
          "content": "<p>I finally figured it out - it was a nasty, silent leakage of the label, in the form of positive exampled getting put earlier than negative ones.</p>\n<p>The details are <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "I finally figured it out - it was a nasty, silent leakage of the label, in the form of positive exampled getting put earlier than negative ones.\n\nThe details are [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081)."
        }
      ]
    },
    {
      "id": 1720857,
      "postDate": "2022-03-13T07:52:42.817Z",
      "content": "<p>Thank you very much! But I have a question. Sampling 5% data could keep the original data distribution?</p>",
      "rawMarkdown": "Thank you very much! But I have a question. Sampling 5% data could keep the original data distribution?"
    },
    {
      "id": 1718932,
      "postDate": "2022-03-11T09:58:46.230Z",
      "content": "<p>Great tips, Thanks a lot !  How do you think of transformer based model in this competition?<br>\njust like riiid 2021。<br>\nDoes the Magnitudes of users 、items 、orders and dates  satisfy ？</p>",
      "rawMarkdown": "Great tips, Thanks a lot !  How do you think of transformer based model in this competition?\njust like riiid 2021。\nDoes the Magnitudes of users 、items 、orders and dates  satisfy ？"
    },
    {
      "id": 1716637,
      "postDate": "2022-03-09T08:43:21.993Z",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> , thanks for the great post and learned alot…<br>\nwanted to know how do we emphasize on ranking/relevance of articles for the customer <br>\nbecause there are lot customers who are returning or rebuying the same stuff. So,, how do we differentiate.</p>",
      "rawMarkdown": "@paweljankiewicz , thanks for the great post and learned alot...\nwanted to know how do we emphasize on ranking/relevance of articles for the customer \nbecause there are lot customers who are returning or rebuying the same stuff. So,, how do we differentiate.",
      "replies": [
        {
          "id": 1720877,
          "postDate": "2022-03-13T08:19:09.790Z",
          "content": "<p>I think ranking of relevance of articles to customers may take care of that.</p>",
          "rawMarkdown": "I think ranking of relevance of articles to customers may take care of that."
        }
      ]
    },
    {
      "id": 1702159,
      "postDate": "2022-02-23T11:29:44.417Z",
      "content": "<p>My apologies in advance but is it fair to tie a recommendation to purchase? Is it a given fact that if the Customer made a purchase, the attribution should go to the recommendation. Is it more of a prediction than a recommendation?</p>",
      "rawMarkdown": "My apologies in advance but is it fair to tie a recommendation to purchase? Is it a given fact that if the Customer made a purchase, the attribution should go to the recommendation. Is it more of a prediction than a recommendation?"
    },
    {
      "id": 1699304,
      "postDate": "2022-02-21T05:13:50.970Z",
      "content": "<p>this was very helpful</p>",
      "rawMarkdown": "this was very helpful\n"
    },
    {
      "id": 1691296,
      "postDate": "2022-02-15T10:38:22.437Z",
      "content": "<p>Thank you for these great tips, this is my first competition in Kaggle (I have experience in smaller scale competitions in another platform). I wanted to know what do you mean by sampling data as much as possible ? Do you mean we should resample data every time we experiment a new thing ? Are there any notebooks or tutorials that talk about this in more detail ? </p>\n<p>Thank you again :)</p>",
      "rawMarkdown": "Thank you for these great tips, this is my first competition in Kaggle (I have experience in smaller scale competitions in another platform). I wanted to know what do you mean by sampling data as much as possible ? Do you mean we should resample data every time we experiment a new thing ? Are there any notebooks or tutorials that talk about this in more detail ? \n\nThank you again :)",
      "replies": [
        {
          "id": 1691508,
          "postDate": "2022-02-15T12:52:03.397Z",
          "content": "<p>One sample is enough for quick tests. Added the code to sample the data <a href=\"https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples\" target=\"_blank\">https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples</a></p>",
          "rawMarkdown": "One sample is enough for quick tests. Added the code to sample the data https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples",
          "votes": 1
        },
        {
          "id": 1692185,
          "postDate": "2022-02-15T22:10:10.660Z",
          "content": "<p>Thank you so much, but the notebook is empty, I think it did not run :/</p>",
          "rawMarkdown": "Thank you so much, but the notebook is empty, I think it did not run :/"
        },
        {
          "id": 1693237,
          "postDate": "2022-02-16T14:38:29.617Z",
          "content": "<p>It should be good now.</p>",
          "rawMarkdown": "It should be good now."
        }
      ]
    },
    {
      "id": 1728123,
      "postDate": "2022-03-18T15:51:18.240Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1728274,
          "postDate": "2022-03-18T18:21:30.870Z",
          "content": "<p>Ad 1.</p>\n<p>In my setting I propose to convert the problem to a table where you have a list of item candidates and you mark the sales as either 0 or 1 depending whether the item was sold in the next week.</p>\n<p>Imagine you want to create negative samples as a list of 1000 most popular items in the last week. Some of those items will be bought by the customers.</p>\n<p>So the table you need to create is:</p>\n<ul>\n<li>customer_id</li>\n<li>label = whether it was bought or not</li>\n<li>article_id from the list of 1000 most popular items</li>\n</ul>\n<p>I'm not saying this is the best approach here but it is more or less what I'm doing. This is a technique where you create a set of explicit negative items. There are also techniques for implicit recommendations but all of them assume some sort of a strategy to generate negative samples.</p>\n<p>Ad 2.</p>\n<p>Please see sampling process here <a href=\"https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples\" target=\"_blank\">https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples</a>.</p>",
          "rawMarkdown": "Ad 1.\n\nIn my setting I propose to convert the problem to a table where you have a list of item candidates and you mark the sales as either 0 or 1 depending whether the item was sold in the next week.\n\nImagine you want to create negative samples as a list of 1000 most popular items in the last week. Some of those items will be bought by the customers.\n\nSo the table you need to create is:\n- customer_id\n- label = whether it was bought or not\n- article_id from the list of 1000 most popular items\n\nI'm not saying this is the best approach here but it is more or less what I'm doing. This is a technique where you create a set of explicit negative items. There are also techniques for implicit recommendations but all of them assume some sort of a strategy to generate negative samples.\n\nAd 2.\n\nPlease see sampling process here https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples.",
          "votes": 10
        },
        {
          "id": 1748764,
          "postDate": "2022-04-08T00:35:29.250Z",
          "content": "<p>I hava a question. </p>\n<p>Did you mean I need to create 1000 most popular items  for each customer and label = whether it was bought or not by each customer?  Should the length of train table  be len(nunique(customer))×1000?</p>",
          "rawMarkdown": "I hava a question. \n\n Did you mean I need to create 1000 most popular items  for each customer and label = whether it was bought or not by each customer?  Should the length of train table  be len(nunique(customer))×1000?",
          "votes": 1
        },
        {
          "id": 1749400,
          "postDate": "2022-04-08T14:38:42.747Z",
          "content": "<p>Pretty much yes, except for training you can discard customers without transactions. Only for making the submission predictions you need to consider nunique(customer)×1000 candidates.</p>",
          "rawMarkdown": "Pretty much yes, except for training you can discard customers without transactions. Only for making the submission predictions you need to consider nunique(customer)×1000 candidates.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1690140,
      "postDate": "2022-02-14T18:09:42.267Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1690153,
          "postDate": "2022-02-14T18:26:49.887Z",
          "content": "<p>\"articles\" are products in french</p>",
          "rawMarkdown": "\"articles\" are products in french",
          "votes": 4
        }
      ]
    },
    {
      "id": 1782309,
      "postDate": "2022-05-09T13:13:07.437Z",
      "content": "<p>Thank you for sharing, learn a lot!</p>",
      "rawMarkdown": "Thank you for sharing, learn a lot!"
    },
    {
      "id": 1763307,
      "postDate": "2022-04-21T12:25:22.520Z",
      "content": "<p>It is very nice! thanks for you advice!</p>",
      "rawMarkdown": "It is very nice! thanks for you advice!"
    },
    {
      "id": 1762021,
      "postDate": "2022-04-20T11:57:43.520Z",
      "content": "<p>Thank you for sharing! nice work!</p>",
      "rawMarkdown": "Thank you for sharing! nice work!"
    },
    {
      "id": 1743427,
      "postDate": "2022-04-03T01:26:36.553Z",
      "content": "<p>thanks for your sampling code</p>",
      "rawMarkdown": "thanks for your sampling code"
    },
    {
      "id": 1723072,
      "postDate": "2022-03-15T05:27:28.673Z",
      "content": "<p>Thank you for sharing your great idea.</p>",
      "rawMarkdown": "Thank you for sharing your great idea."
    }
  ],
  "comments": [
    {
      "id": 1688506,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2022-02-13T17:03:05.520000",
      "content": "<p>For anyone that might not be familiar <a href=\"https://recsys.acm.org\" target=\"_blank\">RecSys</a> is one of the most respected competitions in the Recommender space. </p>\n<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> I was trying to find your winning solution but I could only find the press release content. Is it available anywhere? </p>\n<p>TIA! :)</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1688518,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-13T17:12:18.343000",
          "content": "<p>Yeah here it is - <a href=\"https://github.com/logicai-io/recsys2019\" target=\"_blank\">https://github.com/logicai-io/recsys2019</a>.</p>",
          "votes": 16,
          "replies": []
        },
        {
          "id": 1695896,
          "author_name": "chuan",
          "author_url": "",
          "post_date": "2022-02-18T12:47:09.383000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> , i try to download the data in Recsys 2019 </p>\n<ul>\n<li>by code in recsys2019/data/download_data.sh. while it did not run. this is the response:<br>\n<code>--2022-02-18 20:38:52--  https://storage.googleapis.com/logicai-recsys2019/trivagoRecSysChallengeData2019_v2.zip\n正在解析主机 storage.googleapis.com (storage.googleapis.com)... 142.250.204.112, 142.250.204.48, 142.250.66.144, ...\n正在连接 storage.googleapis.com (storage.googleapis.com)|142.250.204.112|:443... 已连接。\n已发出 HTTP 请求，正在等待回应... 404 Not Found\n2022-02-18 20:38:54 错误 404：Not Found。</code></li>\n<li>through this url (<a href=\"https://recsys.trivago.cloud/challenge/dataset/\" target=\"_blank\">https://recsys.trivago.cloud/challenge/dataset/</a>)<br>\nbut it also failed.<br>\nCould I download the data in other way?<br>\nThank you so much</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1696209,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-18T16:53:56.520000",
          "content": "<p>I don't have the original dataset. Also the code is written in such a way that it needs close to 500gb of ram to train the models.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1696214,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-18T16:57:10.007000",
          "content": "<p>For anyone else surprised by:</p>\n<blockquote>\n  <p>Also the code is written in such a way that it needs close to 500gb of ram to train the models.</p>\n</blockquote>\n<p>When I had the opportunity to do an interview for another <a href=\"https://www.youtube.com/watch?v=W3aWEXqIkWk\" target=\"_blank\">RecSys winning team</a>, I learned this is quite a common number for such problems. 😅</p>\n<p>The embedding sizes based on number of users and recommendations just eats up A LOT of memory by the nature of their size</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1696389,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-02-18T19:52:49.823000",
          "content": "<p>BTW I found the data here on Kaggle (but don't seem to have the link to it now), it didn't seem to be available anywhere else.</p>\n<p>But I got the most mileage out of reading the code in the repo, didn't get around to doing much with the generated data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1696410,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-18T20:07:22.843000",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I believe <a href=\"https://www.kaggle.com/pranavmahajan725/trivagorecsyschallengedata2019\" target=\"_blank\">this</a> is the link to the dataset?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1696492,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-02-18T22:30:40.160000",
          "content": "<p>yes, you found it! :) I think it was that one</p>\n<p>with the data I downloaded the preprocessing steps ran without issue</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1696542,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-18T23:16:10.543000",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> the most important concept in RecSys 2019 was the notion of accumulators classes which have 2 methods:</p>\n<ul>\n<li>update - updates the state of accumulator</li>\n<li>get - extracts the state of accumulator</li>\n</ul>\n<p>When you process events by date you first extract the information from the accumulator and then update it with new information. This way there is no possibility to leak any info. The accumulators themselves can track information about the item, customer, item+customer, other groups of information. Really anything you need.</p>\n<p>The code could be summarized as follows:</p>\n<pre><code>for transaction in transactions:\n      for accumulator in accumulators:\n             get features from accumulator - really important that it is before update\n             update the accumulator\n</code></pre>\n<p>There is an opportunity to parallelize the extraction by reversing the loop so you can process and extract features from accumulators in parallel.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1696572,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-02-19T00:18:28.483000",
          "content": "<p>Ah, I see! <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> thank you very much for sharing this, this is wonderful! 😊</p>\n<p>If I am reading this right, the accumulators should be anything -- items of the given type purchased by a customer up to this point, count of transactions, etc.</p>\n<p>I have completely not thought of processing data like this, this is very cool, eager to give it a try!</p>\n<p>So far, I have been attempting to generate features (and candidate transactions) using dataframes, groupby, shift, and merging. But that takes a ridiculous amount of RAM and the code is very hard to reason about and prone to having bugs!</p>\n<p>Would love to go back to straight-up Python for more of that :) </p>\n<p>I am not sure I understand the parallelization bit, where we reverse the loop? There is a first pass through the dataset that we cannot parallelize, to compute the accumulators. And we can only parallelize adding features to records in the second pass? (split the df into n_chunks, apply features from accumulator to each chunk, combine?) I am thinking I might be missing something here :) I am also thinking we probably want the accumulators to be simple so that they can only be applied as they go through the list of data in chronological order -- I guess I am really lost on the parallelization bit :) </p>\n<p>Once we get all these features, do we do feature selection at all? Or do we just see how correlated columns are and drop the ones that don't seem to add a lot of signal?</p>\n<p>I saw in the repo you shared that it seems models were run on different train sets? (maybe I misread this). Would you grab some numbers of columns and train different models on different subsets? I am wondering if I didn't misread this as I think there is a way to make lgbm randomly sample a subset of columns to grow each tree, but maybe such an approach is nice for blending?</p>\n<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> your posts are legendary and this thread is an absolute treasure trove of information for people just getting started with recsys, thank you so much for all your help!!! 🙏</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1696593,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-02-19T00:33:44.940000",
          "content": "<p>I think I might have understood it now! :)</p>\n<p>We don't reverse the loop in the sense that we run from end to finish, we just reverse it by moving the inside of the loop outside, like this:</p>\n<pre><code>for accumulator in accumulators:\n  for transaction in transactions:\n    run acc and apply features\n</code></pre>\n<p>We can walk the column/dataframe in parallel by multiple accumulators and then combine the results!</p>\n<p>I would love to try this with something like <code>dask</code> and <code>delayed</code>, I suspect this might be a nice way to work around the limitations of <code>dask's</code> API, really curious how well this could work :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1697962,
          "author_name": "chuan",
          "author_url": "",
          "post_date": "2022-02-20T02:37:40.973000",
          "content": "<p>Great work!  Respect to all of you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1717998,
          "author_name": "Homoalways",
          "author_url": "",
          "post_date": "2022-03-10T12:21:07.620000",
          "content": "<p>So is it true that if I don't have a machine with lots of memory, I cannot become one of the winners😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1776281,
      "author_name": "Daan Le",
      "author_url": "",
      "post_date": "2022-05-03T21:35:19.013000",
      "content": "<p>Hi, thank you so much for this post, it is very useful. I had one question about the generation of the positive examples you use to train the ranking model. Do you use the same method for finding these as with the final test set which means there are missing positive examples or are you including all positive examples as they are known in the training data?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1776316,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-05-03T22:24:40.767000",
          "content": "<p>That's a good question. I tried adding all positives but the score didn't improve. I'm using only positives which are associated with some strategy.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1778497,
      "author_name": "Ryota",
      "author_url": "",
      "post_date": "2022-05-05T11:20:42.293000",
      "content": "<p>Thank you for your posting. This post is very helpful.<br>\nI would like to ask you positive-negative sample rate during training process.<br>\nYou said that about 1000 candidates are collected for each customers in your strategy.<br>\nBut, when training model, this is unbalanced-data. I think down-sampling is necessary.<br>\nIf you don't mind, please tell me your positive-negative rate.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1778588,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-05-05T13:30:48.177000",
          "content": "<p>You are right. At the time of writing I thought that increasing the number of candidates can be good. Then I got to a point where adding more candidates doesn't help. I have way too many candidates right now. During training I take only 5-10% of negative examples. <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> commented that he is using only 200 candidates which makes sense. I think most of the true positives that we are able to predict are just popular items. With 5% downsampling I have about 4% positives.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1778653,
          "author_name": "Piotr Gabrys",
          "author_url": "",
          "post_date": "2022-05-05T14:58:26.153000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> ;) I'm optimizing number of candidates with Bayesian Optimization. For every candidate generation method I have a <code>k</code> parameter determining how many articles can be returned. Then I optimize  Precision@50 for a baseline model (no reranking model) where the <code>k's</code> are hyperparameters. After a while you can clearly see which generators should be switched off or limited.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1779241,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-05-06T08:23:02.900000",
          "content": "<p>Interesting approach. I think that this competition shows very well that most cf recommendation algorithms are almost useless in this setting. It is more like a symbolic recommendation system where you don't rely on particular items but rather pools of items created from recent transaction history. So recommendation can be based on:</p>\n<p>0.5<em>top_50_articles_department + 0.5</em>most_similar_articles_to_the_last_basket_that_are_within_1000_most_popular_articles</p>\n<p>There are some obvious flaws in the competition like guessing the availability of items but it is a valuable experience nonetheless.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1779443,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-05-06T12:57:13.977000",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> - thanks for sharing that!</p>\n<p>What does <code>top_50_articles_department</code> mean?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1779470,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-05-06T13:25:40.517000",
          "content": "<p>It is just an example of some strategy. In this case it would be something like: take popular items from some specific item category that the user has in the historical transactions. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1767890,
      "author_name": "Fnoa",
      "author_url": "",
      "post_date": "2022-04-25T18:37:44.290000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>, very very helpful post. Thanks </p>\n<p>Just one question, what do you mean with “model based predictions with customers without the transaction history” ?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1761945,
      "author_name": "MohammedAlsayed",
      "author_url": "",
      "post_date": "2022-04-20T10:22:19.840000",
      "content": "<p>Hi Pawl, </p>\n<p>thanks for your informative post. I would like to ask you what do you mean by dynamic/static attributes for customer and articles? if illustrated with an example will be more helpful. Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1742077,
      "author_name": "Y-Haneji",
      "author_url": "",
      "post_date": "2022-04-01T12:44:09.713000",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nThanks for great idea!<br>\nI have a question.<br>\nWhat does observation_date exactly mean? Please give me some example.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1744357,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-03T22:50:51.283000",
          "content": "<p>Observation date is the date you treat as \"now\" so you can use all the data before this date. So for validation you need to take a date like 2020-09-15. For submission the observation date is of course the last available date 2020-09-22.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1750050,
          "author_name": "Y-Haneji",
          "author_url": "",
          "post_date": "2022-04-09T09:14:06.753000",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nThanks! I have completely understood.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1700846,
      "author_name": "Pankaj Kumar",
      "author_url": "",
      "post_date": "2022-02-22T10:00:18.293000",
      "content": "<p>Thank you for sharing your idea. This is really a great approach. I am getting \"run out of memory\" problem. I am using my local PC which has 16GB RAM. Can you suggest how to deal with this situation. Thank You! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1700865,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-22T10:18:45.013000",
          "content": "<p>16GB RAM is probably not enough to comfortably work with this dataset - at least the approach I suggested needs a lot more memory. You can experiment on a smaller dataset <a href=\"https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples\" target=\"_blank\">https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples</a>. Even using 5% sample should give you similar performance to 100% on this data. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1701108,
          "author_name": "Pankaj Kumar",
          "author_url": "",
          "post_date": "2022-02-22T14:24:37.200000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1704464,
          "author_name": "Faisal Alsrheed",
          "author_url": "",
          "post_date": "2022-02-25T14:47:55.280000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> :)</p>\n<p>Thank you for sharing</p>\n<p>You might find this code useful to make a stratified Sample using many columns.</p>\n<p><a href=\"https://github.com/flaboss/python_stratified_sampling/blob/master/stratifiedSample.py\" target=\"_blank\">https://github.com/flaboss/python_stratified_sampling/blob/master/stratifiedSample.py</a></p>\n<p>All the best </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1698165,
      "author_name": "Mahamadi NIKIEMA",
      "author_url": "",
      "post_date": "2022-02-20T07:42:19.007000",
      "content": "<p>I am new to the recommender system, do you have any references for me? Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1699049,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-20T21:36:41.370000",
          "content": "<p>I suggest to start with the most famous datasets for recommendations - something like movielens - <a href=\"https://grouplens.org/datasets/movielens/\" target=\"_blank\">https://grouplens.org/datasets/movielens/</a>. Follow some tutorial using these datasets. This competition is on the harder spectrum because it is not a very typical dataset. I remember watching a part of this course <a href=\"https://www.coursera.org/specializations/recommender-systems\" target=\"_blank\">https://www.coursera.org/specializations/recommender-systems</a> and it was decent.</p>",
          "votes": 12,
          "replies": []
        }
      ]
    },
    {
      "id": 1696455,
      "author_name": "Jianing ",
      "author_url": "",
      "post_date": "2022-02-18T21:03:01.053000",
      "content": "<p>how could you create negative samples? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1696519,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-18T22:52:07.610000",
          "content": "<p>As I have written there are many ways to create negative samples. For example items in the previous basket can be a source of negative samples. Popular items can be too. There are countless of examples.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1715321,
          "author_name": "HongweiLuan",
          "author_url": "",
          "post_date": "2022-03-07T21:50:41.483000",
          "content": "<p>Thanks for the inspiring post <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> It is very helpful especially on the negative sampling. <br>\nI have another question about negative sampling. Shall we sample from items which are not similar to positive samples as negative samples? I am asking because it may confuse the model if we sample popular or previously purchased items as negative samples. They could be potentially positive samples in the test set. On the other hand, they could represent an interest shift of the customers. Do you have suggestions from your experience? Thank you ahead. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1717420,
          "author_name": "AI Diffusion",
          "author_url": "",
          "post_date": "2022-03-09T23:19:48.607000",
          "content": "<p>He said 'use items in the previous basket', it might mean 'user other customer's basket as the negative samples for the current customer', make sense? I think it's ok.</p>\n<p>But what if these two customers bought the same items, does it confuse the model? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1707473,
      "author_name": "Ethan Kershner",
      "author_url": "",
      "post_date": "2022-02-28T13:49:40.003000",
      "content": "<p>Pawel, thanks so much for putting this together, this is awesome! 🙏</p>\n<p>I have a question for you about search strategies. Is there any reason why, instead of separately implementing a search strategy and inserting the result as a column in the training set, it couldn't also work to just directly add the variables that the strategy depends on into the training set? </p>\n<p>So for example if strategy S is dependent on 10 variables x1, x2… x10, instead of adding S as a column to the training set, adding x1 - x10. Although I suppose this would only be reasonable for simpler strategies and wouldn't work for strategies like image detection.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1707500,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-28T14:14:09.507000",
          "content": "<p>This is a good question. I think there are several problems</p>\n<p>1) For each customer you must consider 100k items if you really want to include all variables with which the strategy was created.<br>\n2) What would you do in a case that you want to relate a set of items from the history with those candidates. The strategy can really on more complex feature set than x1,…,x10.</p>\n<p>A search strategy is a way to reduce the number of candidates that you consider for each customer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1707630,
          "author_name": "Ethan Kershner",
          "author_url": "",
          "post_date": "2022-02-28T16:23:25.903000",
          "content": "<p>That makes sense to me. Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1698993,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-02-20T20:11:23.767000",
      "content": "<p>As I keep working on this and learning about it, I keep coming to this thread and rereading it. This is such a wealth of information!</p>\n<p>Thank you again <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>!</p>\n<p>What blew my mind today was using ‘customer_id’ as a query! I thought we might want to query on [‘customer_id’, ‘observation_date’] to essentially query on baskets?</p>\n<p>Also, the list of transactions can be a good staying set of positive examples? Meaning, from the post it might seem like we need to do something extra (‘the label indicates whether the article will be purchased next week’) but this is essentially what the list of transactions does? It is an observation date and if we add an indicator column such as ‘purchased’ we will have everything that we need? I am just worried I might be missing something subtle here.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1699032,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-20T21:09:57.727000",
          "content": "<p>No you are right for more than 1 observation date the query should be customer_id + observation_dt. Because I have everything sorted by date I didn't notice that I'm implicitly including observation_dt in the query.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1699040,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-20T21:18:58.783000",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I think you are missing something else. When you generate candidates, it is not 100% sure that you will find any transactions. Let's say your only strategy is to include the last basket items. If you have a customer which didn't buy anything twice then you can discard whole customer + observation_dt. The way the ranking models work they need positive examples in the query. This is a source of an important optimization in my opinion.</p>\n<p>The process should be like this: you generate a candidates set something like <code>Dict[Strategy, List[ArticleId]]</code>. A unique set of ArticleId is your number of observations that you have to consider for the customer. </p>\n<p>As I have said for training purposes you can discard all the queries without positive examples. But for validation and submission you must include all of your candidates.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1699125,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-02-21T00:08:45.827000",
          "content": "<p>Thank you so much for these additional details, <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>! 😊 All this is of great help</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1758919,
      "author_name": "WENDY LIU",
      "author_url": "",
      "post_date": "2022-04-18T06:50:27.390000",
      "content": "<p>nice work!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1757936,
      "author_name": "Y-Haneji",
      "author_url": "",
      "post_date": "2022-04-17T06:43:41.253000",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nInspired by you, I have been working on rank model these days.<br>\nI have a question about mixing several strategy to make candidates. How can we put \"strategy score\" of candidate from different strategy together? Or make a column for each strategy?<br>\nTake an example of mixing \"recent transaction strategy\"(strategy1) and \"LSTM strategy\"(strategy2). I can simply make candidates of [customer, article, strategy1: bool, strategy2: bool] columns. <br>\nBUT, I want to use features from strategies, like \"transaction count\" from strategy1 or \"LSTM's probability\" from strategy2. I think this is \"strategy score\".</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1758243,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-17T13:40:48.827000",
          "content": "<p>I think it is worth having boolean flags for each strategy and of course you can use additional features from the strategy like counts and probabilities. The only thing you have to watch for is data leakage, so make sure you use only the data before the observation date to train your models.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1758408,
          "author_name": "Y-Haneji",
          "author_url": "",
          "post_date": "2022-04-17T16:43:06.213000",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> <br>\nThanks for reply!<br>\nI understand that to avoid leakage is important.<br>\nI have adopted naive strategies, but will develop some elaborate strategies from now on and mix them.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1747328,
      "author_name": "Clear n' Simple",
      "author_url": "",
      "post_date": "2022-04-06T15:18:31.213000",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> - I'm using LGBMRanker, and printing out the ndcg at each step, and I'm finding that as the model trains, it's getting worse, even on the training set!</p>\n<p>Same when using <code>map</code>.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1747656,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-06T20:58:00.873000",
          "content": "<p>That's really strange. You probably have some sort of a bug. Are you sure you are providing a proper group parameter? I can vouch for LGBMRanker. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1747673,
          "author_name": "IgorMuniz",
          "author_url": "",
          "post_date": "2022-04-06T21:36:46.240000",
          "content": "<p>We had the same problem and so far I couldn't find a solution for it. Not sure but I think it might be related to the way we are building our dataset which Idk if it's right. We are considering 1 week only as the label and all data before this week as train. Maybe there is a better way to carry time information than that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1748364,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-07T14:11:57.113000",
          "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> - thanks for letting me know that it shouldn't be happening, I'll keep digging until I find out what's wrong.</p>\n<p>I went back and checked, the group parameter looks good.</p>\n<p>Maybe it's because I haven't added customer features yet, only article and customer-article ones?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1748384,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-07T14:37:02.650000",
          "content": "<p><a href=\"https://www.kaggle.com/igormunizims\" target=\"_blank\">@igormunizims</a>, I'm also only considering one week as label for starters.</p>\n<p>I don't see why that should be a problem, though.<br>\nWhen you have very little data, you overfit quickly, but it should still improve on the training set. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1748428,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-07T15:27:42.110000",
          "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> there are some optimizations you can do when you are training the ranking model. For example you can consider only the queries with at least 1 positive example. For the ranking model the queries with all negative observations are useless. If you didn't remove them maybe you can check if this is the reason why the ranking model is behaving in a strange way. It may be that after removing all such observations the number of observations is really low and it does indeed is overfiiting a lot.</p>\n<p>Unfortunately when you are predicting for submission you must use all the observations which is painful (in my case I need to predict 100gb of compressed dataframes and it is the slowest part of the process).</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1748487,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-07T16:29:31.913000",
          "content": "<p>Yes, I did follow your advice on that.</p>\n<p>For a single fold, I have 390k rows, for 12k customers, with 15k positives.<br>\nI realize it may be overfitting, but evaluation metric should still go down on the training set, no?</p>\n<p>Wow - 100gb is massive.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1748580,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-07T18:15:37.703000",
          "content": "<p>Yes. It should go down. Are the results when you are checking the metric manually the same?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1759512,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-04-18T17:30:53.483000",
          "content": "<p>Yes, checking manually gives me the same.</p>\n<p>I find that after a few steps, it starts improving again (although never getting as good at it originally was), and at that point, the different evaluation metrics correlate, and cv correlates with LB, so I'm working with that.</p>\n<p>Still not sure why it's happening, but I suspect it's because of weak features and small training set size.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1777731,
          "author_name": "Clear n' Simple",
          "author_url": "",
          "post_date": "2022-05-04T18:35:47.290000",
          "content": "<p>I finally figured it out - it was a nasty, silent leakage of the label, in the form of positive exampled getting put earlier than negative ones.</p>\n<p>The details are <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081\" target=\"_blank\">here</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1720857,
      "author_name": "G!",
      "author_url": "",
      "post_date": "2022-03-13T07:52:42.817000",
      "content": "<p>Thank you very much! But I have a question. Sampling 5% data could keep the original data distribution?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1718932,
      "author_name": "ctscsu",
      "author_url": "",
      "post_date": "2022-03-11T09:58:46.230000",
      "content": "<p>Great tips, Thanks a lot !  How do you think of transformer based model in this competition?<br>\njust like riiid 2021。<br>\nDoes the Magnitudes of users 、items 、orders and dates  satisfy ？</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1716637,
      "author_name": "Varun Jagannath",
      "author_url": "",
      "post_date": "2022-03-09T08:43:21.993000",
      "content": "<p><a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> , thanks for the great post and learned alot…<br>\nwanted to know how do we emphasize on ranking/relevance of articles for the customer <br>\nbecause there are lot customers who are returning or rebuying the same stuff. So,, how do we differentiate.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1720877,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-03-13T08:19:09.790000",
          "content": "<p>I think ranking of relevance of articles to customers may take care of that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1702159,
      "author_name": "AtulVerma",
      "author_url": "",
      "post_date": "2022-02-23T11:29:44.417000",
      "content": "<p>My apologies in advance but is it fair to tie a recommendation to purchase? Is it a given fact that if the Customer made a purchase, the attribution should go to the recommendation. Is it more of a prediction than a recommendation?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1699304,
      "author_name": "kwalpratham",
      "author_url": "",
      "post_date": "2022-02-21T05:13:50.970000",
      "content": "<p>this was very helpful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1691296,
      "author_name": "Mohamed Annis SOUAMES",
      "author_url": "",
      "post_date": "2022-02-15T10:38:22.437000",
      "content": "<p>Thank you for these great tips, this is my first competition in Kaggle (I have experience in smaller scale competitions in another platform). I wanted to know what do you mean by sampling data as much as possible ? Do you mean we should resample data every time we experiment a new thing ? Are there any notebooks or tutorials that talk about this in more detail ? </p>\n<p>Thank you again :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1691508,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-15T12:52:03.397000",
          "content": "<p>One sample is enough for quick tests. Added the code to sample the data <a href=\"https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples\" target=\"_blank\">https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1692185,
          "author_name": "Mohamed Annis SOUAMES",
          "author_url": "",
          "post_date": "2022-02-15T22:10:10.660000",
          "content": "<p>Thank you so much, but the notebook is empty, I think it did not run :/</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1693237,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-02-16T14:38:29.617000",
          "content": "<p>It should be good now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1728123,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-18T15:51:18.240000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1728274,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-03-18T18:21:30.870000",
          "content": "<p>Ad 1.</p>\n<p>In my setting I propose to convert the problem to a table where you have a list of item candidates and you mark the sales as either 0 or 1 depending whether the item was sold in the next week.</p>\n<p>Imagine you want to create negative samples as a list of 1000 most popular items in the last week. Some of those items will be bought by the customers.</p>\n<p>So the table you need to create is:</p>\n<ul>\n<li>customer_id</li>\n<li>label = whether it was bought or not</li>\n<li>article_id from the list of 1000 most popular items</li>\n</ul>\n<p>I'm not saying this is the best approach here but it is more or less what I'm doing. This is a technique where you create a set of explicit negative items. There are also techniques for implicit recommendations but all of them assume some sort of a strategy to generate negative samples.</p>\n<p>Ad 2.</p>\n<p>Please see sampling process here <a href=\"https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples\" target=\"_blank\">https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples</a>.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1748764,
          "author_name": "oriori4244",
          "author_url": "",
          "post_date": "2022-04-08T00:35:29.250000",
          "content": "<p>I hava a question. </p>\n<p>Did you mean I need to create 1000 most popular items  for each customer and label = whether it was bought or not by each customer?  Should the length of train table  be len(nunique(customer))×1000?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1749400,
          "author_name": "Paweł Jankiewicz",
          "author_url": "",
          "post_date": "2022-04-08T14:38:42.747000",
          "content": "<p>Pretty much yes, except for training you can discard customers without transactions. Only for making the submission predictions you need to consider nunique(customer)×1000 candidates.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1690140,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-14T18:09:42.267000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1690153,
          "author_name": "Ala Eddine Ayadi",
          "author_url": "",
          "post_date": "2022-02-14T18:26:49.887000",
          "content": "<p>\"articles\" are products in french</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1782309,
      "author_name": "Lukan",
      "author_url": "",
      "post_date": "2022-05-09T13:13:07.437000",
      "content": "<p>Thank you for sharing, learn a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1763307,
      "author_name": "Zhicheng Wang",
      "author_url": "",
      "post_date": "2022-04-21T12:25:22.520000",
      "content": "<p>It is very nice! thanks for you advice!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1762021,
      "author_name": "amandus-kim",
      "author_url": "",
      "post_date": "2022-04-20T11:57:43.520000",
      "content": "<p>Thank you for sharing! nice work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1743427,
      "author_name": "stallone",
      "author_url": "",
      "post_date": "2022-04-03T01:26:36.553000",
      "content": "<p>thanks for your sampling code</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1723072,
      "author_name": "Tact Suetsugu",
      "author_url": "",
      "post_date": "2022-03-15T05:27:28.673000",
      "content": "<p>Thank you for sharing your great idea.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1688473": "Just to give you some background - my team won RecSys 2019 so I'm pretty familiar with problems like these. Now that we have this out of the way...\n\nTo simplify the discussion I propose to call the set of articles that a customer bought on a given date a **basket**. And then this problem can be really summarized as predict next basket.\n\nGeneral approach towards problems like this could be as follows. At least mine will be.\n\n## 1. For each customer generate candidates that they can buy. \n\nThese candidates are your negative and positive examples. I'm saying this because someone on the forum asked where are the negative observations. And the answer is you generate them. \n\nThe strategies to generate those candidates are probably the most important aspect of this competition. The strategies can be based on:\n- item to item similarities between customer previous baskets - so image similarity is not out of the question as somebody suggested\n- user based collaborative filtering\n- last baskets (with the hope that the user will buy them again)\n- model based predictions with customers without the transaction history\n- etc\n\nComing up with different strategies is a fun challenge :). I cannot imagine that a winning solution will be based on a closed form model representing one strategy. Once you have all the different strategies ready for each customer then it is time to create a huge table with columns like these:\n\n- observation_date - the date for which you are making predictions - submission should be based on all the data before 2020-09-22 (my birthday btw).\n- customer_id\n- article_id\n- bought in the next week (this is your label)\n- customer static attributes\n- customer dynamic attributes\n- article static attributes\n- article dynamic attributes (maybe item popularity and such)\n- whether the article exist in the strategies and which ones\n- strategy score (whatever it is)\n- many other features that relate the item to the user\n\nOnce you have a table like this it all becomes very familiar. There were many competitions involving click through predictions. They have almost an identical structure the only difference is the source of observations.\n\n## 2. Build a ranking model that ranks the items within Customer \n\nRanking is a different problem than classification but in general classification models can be used as well. My favourite ranking model is LightGBM (we used it in RecSys 2019) https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRanker.html.\n\nThe customer id can be treated as a query in ranking model because you are interested in optimizing the scores at a customer level.\n\n## General tips\n\n1. Sample the data as much as possible - most of the competition you should spend validating your features on the sample. Increase the sample you are using when you are no longer seeing the correlation between local validation and the leaderboard.\n\n2. Test different ideas quickly. Fast local validation is key here because you must do as much experiments as possible.\n\nGood luck :)\n\nUpdate 1. Added code for sampling here https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples",
    "1688506": "For anyone that might not be familiar [RecSys](https://recsys.acm.org) is one of the most respected competitions in the Recommender space. \n\n@paweljankiewicz I was trying to find your winning solution but I could only find the press release content. Is it available anywhere? \n\nTIA! :)",
    "1776281": "Hi, thank you so much for this post, it is very useful. I had one question about the generation of the positive examples you use to train the ranking model. Do you use the same method for finding these as with the final test set which means there are missing positive examples or are you including all positive examples as they are known in the training data?",
    "1778497": "Thank you for your posting. This post is very helpful.\nI would like to ask you positive-negative sample rate during training process.\nYou said that about 1000 candidates are collected for each customers in your strategy.\nBut, when training model, this is unbalanced-data. I think down-sampling is necessary.\nIf you don't mind, please tell me your positive-negative rate.",
    "1767890": "Hi @paweljankiewicz, very very helpful post. Thanks \n\nJust one question, what do you mean with “model based predictions with customers without the transaction history” ?\n",
    "1761945": "Hi Pawl, \n\nthanks for your informative post. I would like to ask you what do you mean by dynamic/static attributes for customer and articles? if illustrated with an example will be more helpful. Thanks!",
    "1742077": "@paweljankiewicz \nThanks for great idea!\nI have a question.\nWhat does observation_date exactly mean? Please give me some example.",
    "1700846": "Thank you for sharing your idea. This is really a great approach. I am getting \"run out of memory\" problem. I am using my local PC which has 16GB RAM. Can you suggest how to deal with this situation. Thank You! ",
    "1698165": "I am new to the recommender system, do you have any references for me? Thanks.",
    "1696455": "how could you create negative samples? ",
    "1707473": "Pawel, thanks so much for putting this together, this is awesome! 🙏\n\nI have a question for you about search strategies. Is there any reason why, instead of separately implementing a search strategy and inserting the result as a column in the training set, it couldn't also work to just directly add the variables that the strategy depends on into the training set? \n\nSo for example if strategy S is dependent on 10 variables x1, x2... x10, instead of adding S as a column to the training set, adding x1 - x10. Although I suppose this would only be reasonable for simpler strategies and wouldn't work for strategies like image detection.",
    "1698993": "As I keep working on this and learning about it, I keep coming to this thread and rereading it. This is such a wealth of information!\n\nThank you again @paweljankiewicz!\n\nWhat blew my mind today was using ‘customer_id’ as a query! I thought we might want to query on [‘customer_id’, ‘observation_date’] to essentially query on baskets?\n\nAlso, the list of transactions can be a good staying set of positive examples? Meaning, from the post it might seem like we need to do something extra (‘the label indicates whether the article will be purchased next week’) but this is essentially what the list of transactions does? It is an observation date and if we add an indicator column such as ‘purchased’ we will have everything that we need? I am just worried I might be missing something subtle here.",
    "1758919": "nice work!!",
    "1757936": "@paweljankiewicz \nInspired by you, I have been working on rank model these days.\nI have a question about mixing several strategy to make candidates. How can we put \"strategy score\" of candidate from different strategy together? Or make a column for each strategy?\nTake an example of mixing \"recent transaction strategy\"(strategy1) and \"LSTM strategy\"(strategy2). I can simply make candidates of [customer, article, strategy1: bool, strategy2: bool] columns. \nBUT, I want to use features from strategies, like \"transaction count\" from strategy1 or \"LSTM's probability\" from strategy2. I think this is \"strategy score\".\n",
    "1747328": "@paweljankiewicz - I'm using LGBMRanker, and printing out the ndcg at each step, and I'm finding that as the model trains, it's getting worse, even on the training set!\n\nSame when using `map`.",
    "1720857": "Thank you very much! But I have a question. Sampling 5% data could keep the original data distribution?",
    "1718932": "Great tips, Thanks a lot !  How do you think of transformer based model in this competition?\njust like riiid 2021。\nDoes the Magnitudes of users 、items 、orders and dates  satisfy ？",
    "1716637": "@paweljankiewicz , thanks for the great post and learned alot...\nwanted to know how do we emphasize on ranking/relevance of articles for the customer \nbecause there are lot customers who are returning or rebuying the same stuff. So,, how do we differentiate.",
    "1702159": "My apologies in advance but is it fair to tie a recommendation to purchase? Is it a given fact that if the Customer made a purchase, the attribution should go to the recommendation. Is it more of a prediction than a recommendation?",
    "1699304": "this was very helpful\n",
    "1691296": "Thank you for these great tips, this is my first competition in Kaggle (I have experience in smaller scale competitions in another platform). I wanted to know what do you mean by sampling data as much as possible ? Do you mean we should resample data every time we experiment a new thing ? Are there any notebooks or tutorials that talk about this in more detail ? \n\nThank you again :)",
    "1728123": "",
    "1690140": "",
    "1782309": "Thank you for sharing, learn a lot!",
    "1763307": "It is very nice! thanks for you advice!",
    "1762021": "Thank you for sharing! nice work!",
    "1743427": "thanks for your sampling code",
    "1723072": "Thank you for sharing your great idea."
  }
}