{
  "id": 324278,
  "title": "6th Place - Giba's Part Solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/324278",
  "author_name": "",
  "post_date": "2022-05-10T20:33:47.444148200Z",
  "votes": 49,
  "comment_count": 12,
  "views": 0,
  "content": "<p>This competition was an amazing experience: tabular+image+text, no leaks, big dataset, lots of public sharing and a good challenge to solve. Thanks H&amp;M and organizers for it and congrats to all prize winners and gold, silver and bronze medalists!</p>\n<p>My solution since the beggining was trying to build something generic enough that could be used in future competitions. So I believe its a little different from other top solutions. </p>\n<p>Before merging teams with <a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a>, <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> and <a href=\"https://www.kaggle.com/hydantess\" target=\"_blank\">@hydantess</a> I worked in a pipeline using RAPIDS to convert the problem into a binary classification and generate features at the same time. <br>\nThe challenge comprisses in recommending around 100k articles to around 1.3 millions customers, more especifically which articles from that 100k list each user is going to buy next week. So the problem is extremelly sparse. </p>\n<p>One way to convert the challenge to binary is to use brute force and construct one row for each customer and article combination. It would generate a dataset of 1.3M customer x 100k articles ~ 130 billions rows dataset for each week in trainset. This is insanely huge to build and store, instead I filtered only 10k articles based in the top10k frequency of sales from the previous 2 weeks, and used that smaller articles list to build the dataset for each customer. So my retrieval list was around 10k articles, same list for all customers each week. So each customer ended having a dataset of 10k rows, one row per article. Doing that using Rapids cudf is just a matter of minutes for all customers for each week. Pandas would had taken days. I ended using weeks 90 to 103 to train my models and week 104 (last week) to validate. Since datasets size and memory grows fast I made my pipeline process everything by batches of 1000 users and store partial results to disk. I ended having more than 1.5TB of datasets in disk at some point of the competition. </p>\n<p>Features that worked best for me were based in article purchase counts, probability of purchase by day, week and 2*weeks. Also good features are based in the period(number of days) since last purchase for each article or properties of the articles. I didn't used text or images in my solution.</p>\n<p>To train the models I picked all positive rows for each customer, but sampled randomly the negative rows to have 200, 300 and 500 depending on the dataset. I ended training 3 models: LightGBM(lambdarank), XGBoost(lambdarank + xendcg) and Catboost(Yetirank). <br>\nThe final solution is an ensemble of the 4 models + 3 different negative sample datasets + bagging with models trained over different periods (weeks 95-103,  100-103, 102-103).</p>\n<p>That blend by itself can score in Public~0.0323 and Private~0.0319.</p>\n<p>Things that didn't worked:</p>\n<ul>\n<li>Train model in high unbalanced datasets (large number of negatives)</li>\n<li>Train models using binary loss.</li>\n<li>Deep learning or sequential models.</li>\n</ul>",
  "messages": [
    {
      "id": "1783994",
      "postDate": "05/10/2022 20:33:47",
      "content": "<p>This competition was an amazing experience: tabular+image+text, no leaks, big dataset, lots of public sharing and a good challenge to solve. Thanks H&amp;M and organizers for it and congrats to all prize winners and gold, silver and bronze medalists!</p>\n<p>My solution since the beggining was trying to build something generic enough that could be used in future competitions. So I believe its a little different from other top solutions. </p>\n<p>Before merging teams with <a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a>, <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> and <a href=\"https://www.kaggle.com/hydantess\" target=\"_blank\">@hydantess</a> I worked in a pipeline using RAPIDS to convert the problem into a binary classification and generate features at the same time. <br>\nThe challenge comprisses in recommending around 100k articles to around 1.3 millions customers, more especifically which articles from that 100k list each user is going to buy next week. So the problem is extremelly sparse. </p>\n<p>One way to convert the challenge to binary is to use brute force and construct one row for each customer and article combination. It would generate a dataset of 1.3M customer x 100k articles ~ 130 billions rows dataset for each week in trainset. This is insanely huge to build and store, instead I filtered only 10k articles based in the top10k frequency of sales from the previous 2 weeks, and used that smaller articles list to build the dataset for each customer. So my retrieval list was around 10k articles, same list for all customers each week. So each customer ended having a dataset of 10k rows, one row per article. Doing that using Rapids cudf is just a matter of minutes for all customers for each week. Pandas would had taken days. I ended using weeks 90 to 103 to train my models and week 104 (last week) to validate. Since datasets size and memory grows fast I made my pipeline process everything by batches of 1000 users and store partial results to disk. I ended having more than 1.5TB of datasets in disk at some point of the competition. </p>\n<p>Features that worked best for me were based in article purchase counts, probability of purchase by day, week and 2*weeks. Also good features are based in the period(number of days) since last purchase for each article or properties of the articles. I didn't used text or images in my solution.</p>\n<p>To train the models I picked all positive rows for each customer, but sampled randomly the negative rows to have 200, 300 and 500 depending on the dataset. I ended training 3 models: LightGBM(lambdarank), XGBoost(lambdarank + xendcg) and Catboost(Yetirank). <br>\nThe final solution is an ensemble of the 4 models + 3 different negative sample datasets + bagging with models trained over different periods (weeks 95-103,  100-103, 102-103).</p>\n<p>That blend by itself can score in Public~0.0323 and Private~0.0319.</p>\n<p>Things that didn't worked:</p>\n<ul>\n<li>Train model in high unbalanced datasets (large number of negatives)</li>\n<li>Train models using binary loss.</li>\n<li>Deep learning or sequential models.</li>\n</ul>",
      "rawMarkdown": "This competition was an amazing experience: tabular+image+text, no leaks, big dataset, lots of public sharing and a good challenge to solve. Thanks H&M and organizers for it and congrats to all prize winners and gold, silver and bronze medalists!\n\nMy solution since the beggining was trying to build something generic enough that could be used in future competitions. So I believe its a little different from other top solutions. \n\nBefore merging teams with @chenxin1991, @juzqyxs and @hydantess I worked in a pipeline using RAPIDS to convert the problem into a binary classification and generate features at the same time. \nThe challenge comprisses in recommending around 100k articles to around 1.3 millions customers, more especifically which articles from that 100k list each user is going to buy next week. So the problem is extremelly sparse. \n\nOne way to convert the challenge to binary is to use brute force and construct one row for each customer and article combination. It would generate a dataset of 1.3M customer x 100k articles ~ 130 billions rows dataset for each week in trainset. This is insanely huge to build and store, instead I filtered only 10k articles based in the top10k frequency of sales from the previous 2 weeks, and used that smaller articles list to build the dataset for each customer. So my retrieval list was around 10k articles, same list for all customers each week. So each customer ended having a dataset of 10k rows, one row per article. Doing that using Rapids cudf is just a matter of minutes for all customers for each week. Pandas would had taken days. I ended using weeks 90 to 103 to train my models and week 104 (last week) to validate. Since datasets size and memory grows fast I made my pipeline process everything by batches of 1000 users and store partial results to disk. I ended having more than 1.5TB of datasets in disk at some point of the competition. \n\nFeatures that worked best for me were based in article purchase counts, probability of purchase by day, week and 2*weeks. Also good features are based in the period(number of days) since last purchase for each article or properties of the articles. I didn't used text or images in my solution.\n\nTo train the models I picked all positive rows for each customer, but sampled randomly the negative rows to have 200, 300 and 500 depending on the dataset. I ended training 3 models: LightGBM(lambdarank), XGBoost(lambdarank + xendcg) and Catboost(Yetirank). \nThe final solution is an ensemble of the 4 models + 3 different negative sample datasets + bagging with models trained over different periods (weeks 95-103,  100-103, 102-103).\n\nThat blend by itself can score in Public~0.0323 and Private~0.0319.\n\nThings that didn't worked:\n- Train model in high unbalanced datasets (large number of negatives)\n- Train models using binary loss.\n- Deep learning or sequential models.",
      "votes": null
    },
    {
      "id": "1784012",
      "postDate": "05/10/2022 20:57:04",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": null
    },
    {
      "id": "1784039",
      "postDate": "05/10/2022 21:55:16",
      "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> congratulations! Thanks for sharing your process</p>",
      "rawMarkdown": "titericz congratulations! Thanks for sharing your process",
      "votes": null
    },
    {
      "id": "1784077",
      "postDate": "05/10/2022 23:20:09",
      "content": "<p>This is a great post and an impressive solution. It's very thorough and explains the thinking behind the approach, as well as what worked and didn't work. This is exactly the kind of post that makes Kaggle so valuable - congratulations on a successful competition!</p>",
      "rawMarkdown": "This is a great post and an impressive solution. It's very thorough and explains the thinking behind the approach, as well as what worked and didn't work. This is exactly the kind of post that makes Kaggle so valuable - congratulations on a successful competition!",
      "votes": null
    },
    {
      "id": "1784796",
      "postDate": "05/11/2022 13:42:00",
      "content": "<p>Congrats!<br>\nI'm very glad that your solution similar to mine because other top solutions are quite different. XD<br>\nBut I do not afford use such amount of data and could not approach you at all…<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152\" target=\"_blank\">FYI. my solution</a></p>",
      "rawMarkdown": "Congrats!\nI'm very glad that your solution similar to mine because other top solutions are quite different. XD\nBut I do not afford use such amount of data and could not approach you at all...\n[FYI. my solution](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152)",
      "votes": null
    },
    {
      "id": "1784868",
      "postDate": "05/11/2022 14:45:56",
      "content": "<blockquote>\n  <p>It would generate a dataset of 1.3M customer x 100k articles ~ 130 billions rows dataset for each week in trainset. This is insanely huge to build and store.</p>\n</blockquote>\n<p>Whereas…</p>\n<blockquote>\n  <p>instead I filtered only 10k articles… I ended using weeks 90 to 103 to train my models…</p>\n</blockquote>\n<p>1.3M customers X 10k articles X 14 weeks = 182 Billion rows - now that's perfectly reasonable!</p>\n<p>Classic Giba!!!</p>",
      "rawMarkdown": "> It would generate a dataset of 1.3M customer x 100k articles ~ 130 billions rows dataset for each week in trainset. This is insanely huge to build and store.\n\nWhereas...\n\n> instead I filtered only 10k articles... I ended using weeks 90 to 103 to train my models...\n\n1.3M customers X 10k articles X 14 weeks = 182 Billion rows - now that's perfectly reasonable!\n\nClassic Giba!!!",
      "votes": null
    },
    {
      "id": "1784870",
      "postDate": "05/11/2022 14:47:48",
      "content": "<p>Didn't understand what you mean by:</p>\n<blockquote>\n  <p>probability of purchase by day, week and 2*weeks</p>\n</blockquote>",
      "rawMarkdown": "Didn't understand what you mean by:\n\n> probability of purchase by day, week and 2*weeks",
      "votes": null
    },
    {
      "id": "1785107",
      "postDate": "05/11/2022 19:42:33",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> indeed for test set we have 1.3M customers, but actually for each week there are around 60k customers only, so: 60k customers X 10k articles X 14 weeks = 8.4 Billion rows \"only\".</p>",
      "rawMarkdown": "jacob34 indeed for test set we have 1.3M customers, but actually for each week there are around 60k customers only, so: 60k customers X 10k articles X 14 weeks = 8.4 Billion rows \"only\".",
      "votes": null
    },
    {
      "id": "1785472",
      "postDate": "05/12/2022 06:03:35",
      "content": "<p>Huge congrats! 🥳</p>\n<p>I love this approach and the way you handled the complexity of this problem 🙂 I put together a starter pack some time ago <a href=\"https://github.com/radekosmulski/personalized_fashion_recs\" target=\"_blank\">here</a> but as I continued to work on the code I got crushed by the idea to separate the work into three separate workflows</p>\n<ul>\n<li>tuning the LGBM model</li>\n<li>local validation</li>\n<li>training on the entirety of the available data and making a submission</li>\n</ul>\n<p>This was just unmanageable.</p>\n<p>When you mention you used the last week for local validation, does it mean that you had a single end-to-end pipeline per model? That you trained on all weeks minus the last one, used that week for local validation, and then submitted based on the candidates you generated for the test set?</p>",
      "rawMarkdown": "Huge congrats! 🥳\n\nI love this approach and the way you handled the complexity of this problem 🙂 I put together a starter pack some time ago [here](https://github.com/radekosmulski/personalized_fashion_recs) but as I continued to work on the code I got crushed by the idea to separate the work into three separate workflows\n- tuning the LGBM model\n- local validation\n- training on the entirety of the available data and making a submission\n\nThis was just unmanageable.\n\nWhen you mention you used the last week for local validation, does it mean that you had a single end-to-end pipeline per model? That you trained on all weeks minus the last one, used that week for local validation, and then submitted based on the candidates you generated for the test set?",
      "votes": null
    },
    {
      "id": "1786209",
      "postDate": "05/12/2022 17:22:59",
      "content": "<p>basically I divide the (sales per item) / (total sales per day)</p>",
      "rawMarkdown": "basically I divide the (sales per item) / (total sales per day)",
      "votes": null
    },
    {
      "id": "1786210",
      "postDate": "05/12/2022 17:24:11",
      "content": "<p>yes, for your question.</p>",
      "rawMarkdown": "yes, for your question.",
      "votes": null
    },
    {
      "id": "1786432",
      "postDate": "05/12/2022 22:21:47",
      "content": "<p>Thank you very much for your reply, <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, appreciate it!!!</p>\n<p>Could I please also ask you how you handled multiple purchases?</p>\n<p>I made the following simplifications -- I map purchases in a week to baskets per customer and remove duplicates. But I am not sure if there isn't a better way to handle this?</p>\n<p>How did you go about this? Do you ever predict duplicate items as recommendations?</p>\n<p>I think this competition might be quite fun to try things on (especially with late submission being available). But would be really great if I could identify what the simplest framing is that makes sense, hence me asking!</p>\n<p>Thank you for all your help!</p>",
      "rawMarkdown": "Thank you very much for your reply, @titericz, appreciate it!!!\n\nCould I please also ask you how you handled multiple purchases?\n\nI made the following simplifications -- I map purchases in a week to baskets per customer and remove duplicates. But I am not sure if there isn't a better way to handle this?\n\nHow did you go about this? Do you ever predict duplicate items as recommendations?\n\nI think this competition might be quite fun to try things on (especially with late submission being available). But would be really great if I could identify what the simplest framing is that makes sense, hence me asking!\n\nThank you for all your help!",
      "votes": null
    },
    {
      "id": "1787284",
      "postDate": "05/13/2022 18:32:35",
      "content": "<p>I removed duplicated rows of items purchased in the same week by the same customer. but I added a new column saying how many items were purchased.</p>",
      "rawMarkdown": "I removed duplicated rows of items purchased in the same week by the same customer. but I added a new column saying how many items were purchased.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1784012,
      "author_name": "igormunizims",
      "author_url": "",
      "post_date": "05/10/2022 20:57:04",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784039,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/10/2022 21:55:16",
      "content": "<p><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> congratulations! Thanks for sharing your process</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784077,
      "author_name": "",
      "author_url": "",
      "post_date": "05/10/2022 23:20:09",
      "content": "<p>This is a great post and an impressive solution. It's very thorough and explains the thinking behind the approach, as well as what worked and didn't work. This is exactly the kind of post that makes Kaggle so valuable - congratulations on a successful competition!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784796,
      "author_name": "iwatatakuya",
      "author_url": "",
      "post_date": "05/11/2022 13:42:00",
      "content": "<p>Congrats!<br>\nI'm very glad that your solution similar to mine because other top solutions are quite different. XD<br>\nBut I do not afford use such amount of data and could not approach you at all…<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152\" target=\"_blank\">FYI. my solution</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784868,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/11/2022 14:45:56",
      "content": "<blockquote>\n  <p>It would generate a dataset of 1.3M customer x 100k articles ~ 130 billions rows dataset for each week in trainset. This is insanely huge to build and store.</p>\n</blockquote>\n<p>Whereas…</p>\n<blockquote>\n  <p>instead I filtered only 10k articles… I ended using weeks 90 to 103 to train my models…</p>\n</blockquote>\n<p>1.3M customers X 10k articles X 14 weeks = 182 Billion rows - now that's perfectly reasonable!</p>\n<p>Classic Giba!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1785107,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "05/11/2022 19:42:33",
          "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> indeed for test set we have 1.3M customers, but actually for each week there are around 60k customers only, so: 60k customers X 10k articles X 14 weeks = 8.4 Billion rows \"only\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1784870,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/11/2022 14:47:48",
      "content": "<p>Didn't understand what you mean by:</p>\n<blockquote>\n  <p>probability of purchase by day, week and 2*weeks</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1786209,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "05/12/2022 17:22:59",
          "content": "<p>basically I divide the (sales per item) / (total sales per day)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1785472,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "05/12/2022 06:03:35",
      "content": "<p>Huge congrats! 🥳</p>\n<p>I love this approach and the way you handled the complexity of this problem 🙂 I put together a starter pack some time ago <a href=\"https://github.com/radekosmulski/personalized_fashion_recs\" target=\"_blank\">here</a> but as I continued to work on the code I got crushed by the idea to separate the work into three separate workflows</p>\n<ul>\n<li>tuning the LGBM model</li>\n<li>local validation</li>\n<li>training on the entirety of the available data and making a submission</li>\n</ul>\n<p>This was just unmanageable.</p>\n<p>When you mention you used the last week for local validation, does it mean that you had a single end-to-end pipeline per model? That you trained on all weeks minus the last one, used that week for local validation, and then submitted based on the candidates you generated for the test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1786210,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "05/12/2022 17:24:11",
          "content": "<p>yes, for your question.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786432,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "05/12/2022 22:21:47",
          "content": "<p>Thank you very much for your reply, <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, appreciate it!!!</p>\n<p>Could I please also ask you how you handled multiple purchases?</p>\n<p>I made the following simplifications -- I map purchases in a week to baskets per customer and remove duplicates. But I am not sure if there isn't a better way to handle this?</p>\n<p>How did you go about this? Do you ever predict duplicate items as recommendations?</p>\n<p>I think this competition might be quite fun to try things on (especially with late submission being available). But would be really great if I could identify what the simplest framing is that makes sense, hence me asking!</p>\n<p>Thank you for all your help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1787284,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "05/13/2022 18:32:35",
          "content": "<p>I removed duplicated rows of items purchased in the same week by the same customer. but I added a new column saying how many items were purchased.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1783994": "This competition was an amazing experience: tabular+image+text, no leaks, big dataset, lots of public sharing and a good challenge to solve. Thanks H&M and organizers for it and congrats to all prize winners and gold, silver and bronze medalists!\n\nMy solution since the beggining was trying to build something generic enough that could be used in future competitions. So I believe its a little different from other top solutions. \n\nBefore merging teams with @chenxin1991, @juzqyxs and @hydantess I worked in a pipeline using RAPIDS to convert the problem into a binary classification and generate features at the same time. \nThe challenge comprisses in recommending around 100k articles to around 1.3 millions customers, more especifically which articles from that 100k list each user is going to buy next week. So the problem is extremelly sparse. \n\nOne way to convert the challenge to binary is to use brute force and construct one row for each customer and article combination. It would generate a dataset of 1.3M customer x 100k articles ~ 130 billions rows dataset for each week in trainset. This is insanely huge to build and store, instead I filtered only 10k articles based in the top10k frequency of sales from the previous 2 weeks, and used that smaller articles list to build the dataset for each customer. So my retrieval list was around 10k articles, same list for all customers each week. So each customer ended having a dataset of 10k rows, one row per article. Doing that using Rapids cudf is just a matter of minutes for all customers for each week. Pandas would had taken days. I ended using weeks 90 to 103 to train my models and week 104 (last week) to validate. Since datasets size and memory grows fast I made my pipeline process everything by batches of 1000 users and store partial results to disk. I ended having more than 1.5TB of datasets in disk at some point of the competition. \n\nFeatures that worked best for me were based in article purchase counts, probability of purchase by day, week and 2*weeks. Also good features are based in the period(number of days) since last purchase for each article or properties of the articles. I didn't used text or images in my solution.\n\nTo train the models I picked all positive rows for each customer, but sampled randomly the negative rows to have 200, 300 and 500 depending on the dataset. I ended training 3 models: LightGBM(lambdarank), XGBoost(lambdarank + xendcg) and Catboost(Yetirank). \nThe final solution is an ensemble of the 4 models + 3 different negative sample datasets + bagging with models trained over different periods (weeks 95-103,  100-103, 102-103).\n\nThat blend by itself can score in Public~0.0323 and Private~0.0319.\n\nThings that didn't worked:\n- Train model in high unbalanced datasets (large number of negatives)\n- Train models using binary loss.\n- Deep learning or sequential models.",
    "1784012": "Congratulations and thanks for sharing!",
    "1784039": "titericz congratulations! Thanks for sharing your process",
    "1784077": "This is a great post and an impressive solution. It's very thorough and explains the thinking behind the approach, as well as what worked and didn't work. This is exactly the kind of post that makes Kaggle so valuable - congratulations on a successful competition!",
    "1784796": "Congrats!\nI'm very glad that your solution similar to mine because other top solutions are quite different. XD\nBut I do not afford use such amount of data and could not approach you at all...\n[FYI. my solution](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152)",
    "1784868": "> It would generate a dataset of 1.3M customer x 100k articles ~ 130 billions rows dataset for each week in trainset. This is insanely huge to build and store.\n\nWhereas...\n\n> instead I filtered only 10k articles... I ended using weeks 90 to 103 to train my models...\n\n1.3M customers X 10k articles X 14 weeks = 182 Billion rows - now that's perfectly reasonable!\n\nClassic Giba!!!",
    "1784870": "Didn't understand what you mean by:\n\n> probability of purchase by day, week and 2*weeks",
    "1785107": "jacob34 indeed for test set we have 1.3M customers, but actually for each week there are around 60k customers only, so: 60k customers X 10k articles X 14 weeks = 8.4 Billion rows \"only\".",
    "1785472": "Huge congrats! 🥳\n\nI love this approach and the way you handled the complexity of this problem 🙂 I put together a starter pack some time ago [here](https://github.com/radekosmulski/personalized_fashion_recs) but as I continued to work on the code I got crushed by the idea to separate the work into three separate workflows\n- tuning the LGBM model\n- local validation\n- training on the entirety of the available data and making a submission\n\nThis was just unmanageable.\n\nWhen you mention you used the last week for local validation, does it mean that you had a single end-to-end pipeline per model? That you trained on all weeks minus the last one, used that week for local validation, and then submitted based on the candidates you generated for the test set?",
    "1786209": "basically I divide the (sales per item) / (total sales per day)",
    "1786210": "yes, for your question.",
    "1786432": "Thank you very much for your reply, @titericz, appreciate it!!!\n\nCould I please also ask you how you handled multiple purchases?\n\nI made the following simplifications -- I map purchases in a week to baskets per customer and remove duplicates. But I am not sure if there isn't a better way to handle this?\n\nHow did you go about this? Do you ever predict duplicate items as recommendations?\n\nI think this competition might be quite fun to try things on (especially with late submission being available). But would be really great if I could identify what the simplest framing is that makes sense, hence me asking!\n\nThank you for all your help!",
    "1787284": "I removed duplicated rows of items purchased in the same week by the same customer. but I added a new column saying how many items were purchased."
  },
  "source": "meta"
}