{
  "id": 324076,
  "title": "52nd place with 20 minute Kaggle Notebook",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/clear-n-simple-52nd-place-with-20-minute-kaggle-no",
  "author_name": "",
  "post_date": "2022-05-15T18:27:58.397Z",
  "votes": 75,
  "comment_count": 24,
  "views": 0,
  "content": "<h3>I was able to generate my final submission - including candidate retrieval, feature engineering, and training/predicting with two models - in just 20 minutes.</h3>\n<p>Partially, it was from using cudf for whatever I could, and it's really amazing how much time it saves.</p>\n<p>But I think the main reason was that I used a light-weight candidate retrieval method that gave me a good result without needing many candidates per customer, and without needing collaborative filtering or image/text processing.</p>\n<p>I had 33 million candidates (which comes out to only ~25 per customer on average).</p>\n<p>I started with <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s notebooks - <a href=\"https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this\" target=\"_blank\">customers who bought this frequently bought this</a> and <a href=\"https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">Recommend Items Purchased Together</a>.</p>\n<p>I refactored his first notebook in <a href=\"https://www.kaggle.com/code/jacob34/cdeotte-pairs-in-3-minutes\" target=\"_blank\">this notebook</a>, so that I could get the same result in 3 minutes.</p>\n<p>Then I made several changes:</p>\n<ol>\n<li>I had it only look for \"pairs\" from articles that were sold in the past week</li>\n<li>I had it find 5 \"pairs\" for each article, instead of just 1</li>\n<li>I ignored any article that wasn't sold within the past week</li>\n<li>I ignored any \"pair\" purchased by less than 2 customers</li>\n</ol>\n<p>We get a very strong retrieval method, a sort of manual, time-aware collaborative filtering method - what customers with similar purchase interests were purchased <em>in the past week</em> - so it includes trend information as well.</p>\n<p>For the second part, I followed his method of recommending \"pairs\" of items a customer had purchased in their last few weeks of history.</p>\n<p>And I also kept features to feed to LGBMRanker about:</p>\n<ul>\n<li>the strength of the \"match\" (i.e. how many/what percentage of customers the \"pair\" was based on)</li>\n<li>the \"source\" of the pair (i.e. how recently the original article was purchased, how many times it was purchased)</li>\n</ul>\n<p>Of course, I also generated the regular, standard candidates: past purchases, and 12 most popular items (I generated those based on age group).</p>\n<p>I think this retrieval method was the main reason I was able to get a good result with relatively few candidates/resources.</p>\n<h3>Update - code</h3>\n<p>Some people asked for it, so I made a clean notebook (moved nearly all code to the repo).<br>\nHere's <a href=\"https://www.kaggle.com/jacob34/clear-n-simple-final-shared-notebook\" target=\"_blank\">the notebook</a>, here's <a href=\"https://github.com/JacobCP/kaggle-handm-helpers\" target=\"_blank\">the repo</a>, and here's the <a href=\"https://www.kaggle.com/datasets/jacob34/handmhelpers\" target=\"_blank\">kaggle dataset</a> that syncs with the repo.</p>\n<p>The notebook runs in 20 minutes, and gets the same private LB score (it varies a bit per run).</p>\n<p>My workflow was:</p>\n<ol>\n<li>I'd work with a function in the notebook (for example, a candidate generator)</li>\n<li>Once the function was working, I'd move it to the repo and point the notebook there.</li>\n<li>Any parameters of the function (for example, candidate threshold), I'd add to my <code>params</code> dictionary in the notebook.</li>\n<li>Each function had a <code>**kwargs</code> argument, so I could pass the same single <code>params</code> dictionary to every function, and the function would use the arguments it needed.</li>\n<li>If I needed to revisit a function, or work with the values generated in middle, I'd move the function/s back to the notebook. </li>\n</ol>\n<h3>Update #2 - cuML's FIL</h3>\n<p>Afterwards, I refactored my code to use cuML's FIL for inference, and that brought notebook time down from 20 minutes to just 12 minutes!</p>",
  "messages": [
    {
      "id": "1782960",
      "postDate": "05/10/2022 03:02:55",
      "content": "<h3>I was able to generate my final submission - including candidate retrieval, feature engineering, and training/predicting with two models - in just 20 minutes.</h3>\n<p>Partially, it was from using cudf for whatever I could, and it's really amazing how much time it saves.</p>\n<p>But I think the main reason was that I used a light-weight candidate retrieval method that gave me a good result without needing many candidates per customer, and without needing collaborative filtering or image/text processing.</p>\n<p>I had 33 million candidates (which comes out to only ~25 per customer on average).</p>\n<p>I started with <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s notebooks - <a href=\"https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this\" target=\"_blank\">customers who bought this frequently bought this</a> and <a href=\"https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">Recommend Items Purchased Together</a>.</p>\n<p>I refactored his first notebook in <a href=\"https://www.kaggle.com/code/jacob34/cdeotte-pairs-in-3-minutes\" target=\"_blank\">this notebook</a>, so that I could get the same result in 3 minutes.</p>\n<p>Then I made several changes:</p>\n<ol>\n<li>I had it only look for \"pairs\" from articles that were sold in the past week</li>\n<li>I had it find 5 \"pairs\" for each article, instead of just 1</li>\n<li>I ignored any article that wasn't sold within the past week</li>\n<li>I ignored any \"pair\" purchased by less than 2 customers</li>\n</ol>\n<p>We get a very strong retrieval method, a sort of manual, time-aware collaborative filtering method - what customers with similar purchase interests were purchased <em>in the past week</em> - so it includes trend information as well.</p>\n<p>For the second part, I followed his method of recommending \"pairs\" of items a customer had purchased in their last few weeks of history.</p>\n<p>And I also kept features to feed to LGBMRanker about:</p>\n<ul>\n<li>the strength of the \"match\" (i.e. how many/what percentage of customers the \"pair\" was based on)</li>\n<li>the \"source\" of the pair (i.e. how recently the original article was purchased, how many times it was purchased)</li>\n</ul>\n<p>Of course, I also generated the regular, standard candidates: past purchases, and 12 most popular items (I generated those based on age group).</p>\n<p>I think this retrieval method was the main reason I was able to get a good result with relatively few candidates/resources.</p>\n<h3>Update - code</h3>\n<p>Some people asked for it, so I made a clean notebook (moved nearly all code to the repo).<br>\nHere's <a href=\"https://www.kaggle.com/jacob34/clear-n-simple-final-shared-notebook\" target=\"_blank\">the notebook</a>, here's <a href=\"https://github.com/JacobCP/kaggle-handm-helpers\" target=\"_blank\">the repo</a>, and here's the <a href=\"https://www.kaggle.com/datasets/jacob34/handmhelpers\" target=\"_blank\">kaggle dataset</a> that syncs with the repo.</p>\n<p>The notebook runs in 20 minutes, and gets the same private LB score (it varies a bit per run).</p>\n<p>My workflow was:</p>\n<ol>\n<li>I'd work with a function in the notebook (for example, a candidate generator)</li>\n<li>Once the function was working, I'd move it to the repo and point the notebook there.</li>\n<li>Any parameters of the function (for example, candidate threshold), I'd add to my <code>params</code> dictionary in the notebook.</li>\n<li>Each function had a <code>**kwargs</code> argument, so I could pass the same single <code>params</code> dictionary to every function, and the function would use the arguments it needed.</li>\n<li>If I needed to revisit a function, or work with the values generated in middle, I'd move the function/s back to the notebook. </li>\n</ol>\n<h3>Update #2 - cuML's FIL</h3>\n<p>Afterwards, I refactored my code to use cuML's FIL for inference, and that brought notebook time down from 20 minutes to just 12 minutes!</p>",
      "rawMarkdown": "### I was able to generate my final submission - including candidate retrieval, feature engineering, and training/predicting with two models - in just 20 minutes.\n\nPartially, it was from using cudf for whatever I could, and it's really amazing how much time it saves.\n\nBut I think the main reason was that I used a light-weight candidate retrieval method that gave me a good result without needing many candidates per customer, and without needing collaborative filtering or image/text processing.\n\nI had 33 million candidates (which comes out to only ~25 per customer on average).\n\nI started with @cdeotte's notebooks - [customers who bought this frequently bought this](https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this) and [Recommend Items Purchased Together](https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021).\n\nI refactored his first notebook in [this notebook](https://www.kaggle.com/code/jacob34/cdeotte-pairs-in-3-minutes), so that I could get the same result in 3 minutes.\n\nThen I made several changes:\n1. I had it only look for \"pairs\" from articles that were sold in the past week\n2. I had it find 5 \"pairs\" for each article, instead of just 1\n3. I ignored any article that wasn't sold within the past week\n4. I ignored any \"pair\" purchased by less than 2 customers\n\nWe get a very strong retrieval method, a sort of manual, time-aware collaborative filtering method - what customers with similar purchase interests were purchased *in the past week* - so it includes trend information as well.\n\nFor the second part, I followed his method of recommending \"pairs\" of items a customer had purchased in their last few weeks of history.\n\nAnd I also kept features to feed to LGBMRanker about:\n- the strength of the \"match\" (i.e. how many/what percentage of customers the \"pair\" was based on)\n- the \"source\" of the pair (i.e. how recently the original article was purchased, how many times it was purchased)\n\nOf course, I also generated the regular, standard candidates: past purchases, and 12 most popular items (I generated those based on age group).\n\nI think this retrieval method was the main reason I was able to get a good result with relatively few candidates/resources.\n\n### Update - code\nSome people asked for it, so I made a clean notebook (moved nearly all code to the repo).\nHere's [the notebook](https://www.kaggle.com/jacob34/clear-n-simple-final-shared-notebook), here's [the repo](https://github.com/JacobCP/kaggle-handm-helpers), and here's the [kaggle dataset](https://www.kaggle.com/datasets/jacob34/handmhelpers) that syncs with the repo.\n\nThe notebook runs in 20 minutes, and gets the same private LB score (it varies a bit per run).\n\nMy workflow was:\n1. I'd work with a function in the notebook (for example, a candidate generator)\n2. Once the function was working, I'd move it to the repo and point the notebook there.\n3. Any parameters of the function (for example, candidate threshold), I'd add to my `params` dictionary in the notebook.\n4. Each function had a `**kwargs` argument, so I could pass the same single `params` dictionary to every function, and the function would use the arguments it needed.\n5. If I needed to revisit a function, or work with the values generated in middle, I'd move the function/s back to the notebook. \n\n### Update #2 - cuML's FIL\nAfterwards, I refactored my code to use cuML's FIL for inference, and that brought notebook time down from 20 minutes to just 12 minutes!",
      "votes": null
    },
    {
      "id": "1782963",
      "postDate": "05/10/2022 03:08:24",
      "content": "<p>good work ! thanks for sharing !</p>",
      "rawMarkdown": "good work ! thanks for sharing !",
      "votes": null
    },
    {
      "id": "1782982",
      "postDate": "05/10/2022 03:34:18",
      "content": "<p>Congratulations Clear n' Simple! Great job finishing solo Silver finish 52 out of 3000 teams. Impressive results for just focusing on pairs of purchased items. Great job achieving super fast speed ups!</p>\n<p>Thanks for sharing helpful discussions and helpful notebooks!</p>",
      "rawMarkdown": "Congratulations Clear n' Simple! Great job finishing solo Silver finish 52 out of 3000 teams. Impressive results for just focusing on pairs of purchased items. Great job achieving super fast speed ups!\n\nThanks for sharing helpful discussions and helpful notebooks!",
      "votes": null
    },
    {
      "id": "1782989",
      "postDate": "05/10/2022 03:38:27",
      "content": "<p>You are the best! Thanks for sharing! I learnt a lot from your notebook!</p>",
      "rawMarkdown": "You are the best! Thanks for sharing! I learnt a lot from your notebook!",
      "votes": null
    },
    {
      "id": "1782990",
      "postDate": "05/10/2022 03:39:48",
      "content": "<p>your Job is epic</p>\n<p>I learn a lot from you, thanks for all the sharing in this competition </p>",
      "rawMarkdown": "your Job is epic\n\nI learn a lot from you, thanks for all the sharing in this competition",
      "votes": null
    },
    {
      "id": "1783047",
      "postDate": "05/10/2022 04:49:23",
      "content": "<p>Congrats! Will you share the code for your pipeline including candidates retrieval, training and prediction?</p>",
      "rawMarkdown": "Congrats! Will you share the code for your pipeline including candidates retrieval, training and prediction?",
      "votes": null
    },
    {
      "id": "1783153",
      "postDate": "05/10/2022 06:59:06",
      "content": "<p>Well done! And mate, I would <em>love</em> to able to see your code</p>",
      "rawMarkdown": "Well done! And mate, I would _love_ to able to see your code",
      "votes": null
    },
    {
      "id": "1783442",
      "postDate": "05/10/2022 11:59:48",
      "content": "<p>Congratz! Awesome work. Thanks for sharing so many insights.</p>",
      "rawMarkdown": "Congratz! Awesome work. Thanks for sharing so many insights.",
      "votes": null
    },
    {
      "id": "1783602",
      "postDate": "05/10/2022 14:07:39",
      "content": "<p>Thanks, Chris, means a lot to me!</p>\n<p>Thanks for coming out of retirement for this 😉.</p>\n<p>More seriously, I was happy you were taking part in the competition, and was hoping to continue to benefit from following your sharing, especially with your recent wins in recommendation comps, but I guess you didn't end up having the time for it.</p>",
      "rawMarkdown": "Thanks, Chris, means a lot to me!\n\nThanks for coming out of retirement for this 😉.\n\nMore seriously, I was happy you were taking part in the competition, and was hoping to continue to benefit from following your sharing, especially with your recent wins in recommendation comps, but I guess you didn't end up having the time for it.",
      "votes": null
    },
    {
      "id": "1783604",
      "postDate": "05/10/2022 14:09:22",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hervind\" target=\"_blank\">@hervind</a>!</p>\n<p>I took inspiration from your notebook on the ability to speed things up with proper code design.</p>",
      "rawMarkdown": "Thanks @hervind!\n\nI took inspiration from your notebook on the ability to speed things up with proper code design.",
      "votes": null
    },
    {
      "id": "1783616",
      "postDate": "05/10/2022 14:25:50",
      "content": "<p>Thanks!<br>\nCongratulations to you on your solo silver!</p>",
      "rawMarkdown": "Thanks!\nCongratulations to you on your solo silver!",
      "votes": null
    },
    {
      "id": "1783881",
      "postDate": "05/10/2022 18:52:01",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/marcogorelli\" target=\"_blank\">@marcogorelli</a>!</p>\n<p>I've cleaned up my notebook for you, and have updated my post with links to the code/repo/dataset.<br>\nQuestions welcome!</p>",
      "rawMarkdown": "Thanks @marcogorelli!\n\nI've cleaned up my notebook for you, and have updated my post with links to the code/repo/dataset.\nQuestions welcome!",
      "votes": null
    },
    {
      "id": "1783884",
      "postDate": "05/10/2022 18:53:06",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/xfffrank\" target=\"_blank\">@xfffrank</a>!</p>\n<p>I've updated my post with links to a clean notebook and my repo/dataset.</p>",
      "rawMarkdown": "Thanks, @xfffrank!\n\nI've updated my post with links to a clean notebook and my repo/dataset.",
      "votes": null
    },
    {
      "id": "1783890",
      "postDate": "05/10/2022 18:54:57",
      "content": "<p>legend, thanks a tonne!</p>",
      "rawMarkdown": "legend, thanks a tonne!",
      "votes": null
    },
    {
      "id": "1784576",
      "postDate": "05/11/2022 09:21:12",
      "content": "<p>Congrats and thanks for sharing a great work!<br>\nI'm interested in what score we could achieve only using a kaggle notebook, and yours seems to be the best…!</p>",
      "rawMarkdown": "Congrats and thanks for sharing a great work!\nI'm interested in what score we could achieve only using a kaggle notebook, and yours seems to be the best...!",
      "votes": null
    },
    {
      "id": "1784678",
      "postDate": "05/11/2022 11:24:04",
      "content": "<p>Congratulations on a simple solution. </p>",
      "rawMarkdown": "Congratulations on a simple solution.",
      "votes": null
    },
    {
      "id": "1785543",
      "postDate": "05/12/2022 07:35:11",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> <br>\nCongratulations for your achievement!<br>\nTo complete such a resource-consuming competition in the kaggle environment is really challenging. </p>\n<p>I have a question. I read some code in your repository. I felt it was well structured and like the package. Did you write almost all the code during the competition, or use some kind of your own pipeline to reuse?</p>",
      "rawMarkdown": "jacob34 \nCongratulations for your achievement!\nTo complete such a resource-consuming competition in the kaggle environment is really challenging. \n\nI have a question. I read some code in your repository. I felt it was well structured and like the package. Did you write almost all the code during the competition, or use some kind of your own pipeline to reuse?",
      "votes": null
    },
    {
      "id": "1785913",
      "postDate": "05/12/2022 13:41:06",
      "content": "<p>Thanks for the feedback, <a href=\"https://www.kaggle.com/hanejiyuto\" target=\"_blank\">@hanejiyuto</a>!</p>\n<p>I wrote practically all of it for this competition.<br>\nAlthough I do hope to reuse parts of it in the future.</p>",
      "rawMarkdown": "Thanks for the feedback, @hanejiyuto!\n\nI wrote practically all of it for this competition.\nAlthough I do hope to reuse parts of it in the future.",
      "votes": null
    },
    {
      "id": "1786576",
      "postDate": "05/13/2022 03:26:52",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> <br>\nOh, it is great!</p>",
      "rawMarkdown": "jacob34 \nOh, it is great!",
      "votes": null
    },
    {
      "id": "1789861",
      "postDate": "05/14/2022 09:47:47",
      "content": "<p>Legend, mate, I've worked through your notebook, and it's really impressive - thanks a tonne for posting!</p>\n<p>For generating the set for LGBMRanker, my understanding is that if you're predicting on some week (say, 103), then you:</p>\n<ul>\n<li>generate candidates, using only information from strictly before week 103</li>\n<li>any of those candidates were actually bought in 103, then they get a label of 1. Else, they get a label of 0</li>\n<li>if a customer bought an article in week 103, but that article didn't appear in that customer's candidates, then that article won't be used for training</li>\n</ul>\n<p>Is this correct? If so, then the last part seems a bit different to what I've understood from <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/309220\" target=\"_blank\">the other LGBMRanker example which was posted</a>, in which all the articles which a customer bought get a label of 1.</p>",
      "rawMarkdown": "Legend, mate, I've worked through your notebook, and it's really impressive - thanks a tonne for posting!\n\nFor generating the set for LGBMRanker, my understanding is that if you're predicting on some week (say, 103), then you:\n- generate candidates, using only information from strictly before week 103\n- any of those candidates were actually bought in 103, then they get a label of 1. Else, they get a label of 0\n- if a customer bought an article in week 103, but that article didn't appear in that customer's candidates, then that article won't be used for training\n\nIs this correct? If so, then the last part seems a bit different to what I've understood from [the other LGBMRanker example which was posted](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/309220), in which all the articles which a customer bought get a label of 1.",
      "votes": null
    },
    {
      "id": "1789949",
      "postDate": "05/14/2022 11:25:56",
      "content": "<p>Congrats:)</p>",
      "rawMarkdown": "Congrats:)",
      "votes": null
    },
    {
      "id": "1790573",
      "postDate": "05/15/2022 05:11:40",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": null
    },
    {
      "id": "1790956",
      "postDate": "05/15/2022 14:02:06",
      "content": "<p>Pleasure!</p>\n<p>Yes, that's correct.</p>\n<p>The reasoning is, we want to train the model to do well on the week of the LB (week 105 in my code) - for that week, we'll only have the candidates we generate.</p>\n<p>You'll want to see <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324076\" target=\"_blank\">this comment</a> from <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>.</p>",
      "rawMarkdown": "Pleasure!\n\nYes, that's correct.\n\nThe reasoning is, we want to train the model to do well on the week of the LB (week 105 in my code) - for that week, we'll only have the candidates we generate.\n\nYou'll want to see [this comment](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324076) from @paweljankiewicz.",
      "votes": null
    },
    {
      "id": "1791215",
      "postDate": "05/15/2022 18:28:04",
      "content": "<h3>Update:</h3>\n<p>By using cuML's FIL, notebook run time went down from 20 minutes to 12 minutes!</p>",
      "rawMarkdown": "### Update:\n\nBy using cuML's FIL, notebook run time went down from 20 minutes to 12 minutes!",
      "votes": null
    },
    {
      "id": "1792008",
      "postDate": "05/16/2022 14:54:50",
      "content": "<p>Congratz! Awesome work. Thanks for sharing so many insights.</p>",
      "rawMarkdown": "Congratz! Awesome work. Thanks for sharing so many insights.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1782963,
      "author_name": "xuxiaodong",
      "author_url": "",
      "post_date": "05/10/2022 03:08:24",
      "content": "<p>good work ! thanks for sharing !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783616,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/10/2022 14:25:50",
          "content": "<p>Thanks!<br>\nCongratulations to you on your solo silver!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1782982,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/10/2022 03:34:18",
      "content": "<p>Congratulations Clear n' Simple! Great job finishing solo Silver finish 52 out of 3000 teams. Impressive results for just focusing on pairs of purchased items. Great job achieving super fast speed ups!</p>\n<p>Thanks for sharing helpful discussions and helpful notebooks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783602,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/10/2022 14:07:39",
          "content": "<p>Thanks, Chris, means a lot to me!</p>\n<p>Thanks for coming out of retirement for this 😉.</p>\n<p>More seriously, I was happy you were taking part in the competition, and was hoping to continue to benefit from following your sharing, especially with your recent wins in recommendation comps, but I guess you didn't end up having the time for it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1782989,
      "author_name": "eggachecat",
      "author_url": "",
      "post_date": "05/10/2022 03:38:27",
      "content": "<p>You are the best! Thanks for sharing! I learnt a lot from your notebook!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1782990,
      "author_name": "hervind",
      "author_url": "",
      "post_date": "05/10/2022 03:39:48",
      "content": "<p>your Job is epic</p>\n<p>I learn a lot from you, thanks for all the sharing in this competition </p>",
      "votes": null,
      "replies": [
        {
          "id": 1783604,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/10/2022 14:09:22",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hervind\" target=\"_blank\">@hervind</a>!</p>\n<p>I took inspiration from your notebook on the ability to speed things up with proper code design.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783047,
      "author_name": "xfffrank",
      "author_url": "",
      "post_date": "05/10/2022 04:49:23",
      "content": "<p>Congrats! Will you share the code for your pipeline including candidates retrieval, training and prediction?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783884,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/10/2022 18:53:06",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/xfffrank\" target=\"_blank\">@xfffrank</a>!</p>\n<p>I've updated my post with links to a clean notebook and my repo/dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783153,
      "author_name": "marcogorelli",
      "author_url": "",
      "post_date": "05/10/2022 06:59:06",
      "content": "<p>Well done! And mate, I would <em>love</em> to able to see your code</p>",
      "votes": null,
      "replies": [
        {
          "id": 1783881,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/10/2022 18:52:01",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/marcogorelli\" target=\"_blank\">@marcogorelli</a>!</p>\n<p>I've cleaned up my notebook for you, and have updated my post with links to the code/repo/dataset.<br>\nQuestions welcome!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1783890,
          "author_name": "marcogorelli",
          "author_url": "",
          "post_date": "05/10/2022 18:54:57",
          "content": "<p>legend, thanks a tonne!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783442,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "05/10/2022 11:59:48",
      "content": "<p>Congratz! Awesome work. Thanks for sharing so many insights.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784576,
      "author_name": "negoto",
      "author_url": "",
      "post_date": "05/11/2022 09:21:12",
      "content": "<p>Congrats and thanks for sharing a great work!<br>\nI'm interested in what score we could achieve only using a kaggle notebook, and yours seems to be the best…!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784678,
      "author_name": "muhammedtausif",
      "author_url": "",
      "post_date": "05/11/2022 11:24:04",
      "content": "<p>Congratulations on a simple solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1785543,
      "author_name": "hanejiyuto",
      "author_url": "",
      "post_date": "05/12/2022 07:35:11",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> <br>\nCongratulations for your achievement!<br>\nTo complete such a resource-consuming competition in the kaggle environment is really challenging. </p>\n<p>I have a question. I read some code in your repository. I felt it was well structured and like the package. Did you write almost all the code during the competition, or use some kind of your own pipeline to reuse?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1785913,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/12/2022 13:41:06",
          "content": "<p>Thanks for the feedback, <a href=\"https://www.kaggle.com/hanejiyuto\" target=\"_blank\">@hanejiyuto</a>!</p>\n<p>I wrote practically all of it for this competition.<br>\nAlthough I do hope to reuse parts of it in the future.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786576,
          "author_name": "hanejiyuto",
          "author_url": "",
          "post_date": "05/13/2022 03:26:52",
          "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> <br>\nOh, it is great!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1789861,
      "author_name": "marcogorelli",
      "author_url": "",
      "post_date": "05/14/2022 09:47:47",
      "content": "<p>Legend, mate, I've worked through your notebook, and it's really impressive - thanks a tonne for posting!</p>\n<p>For generating the set for LGBMRanker, my understanding is that if you're predicting on some week (say, 103), then you:</p>\n<ul>\n<li>generate candidates, using only information from strictly before week 103</li>\n<li>any of those candidates were actually bought in 103, then they get a label of 1. Else, they get a label of 0</li>\n<li>if a customer bought an article in week 103, but that article didn't appear in that customer's candidates, then that article won't be used for training</li>\n</ul>\n<p>Is this correct? If so, then the last part seems a bit different to what I've understood from <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/309220\" target=\"_blank\">the other LGBMRanker example which was posted</a>, in which all the articles which a customer bought get a label of 1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1790956,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/15/2022 14:02:06",
          "content": "<p>Pleasure!</p>\n<p>Yes, that's correct.</p>\n<p>The reasoning is, we want to train the model to do well on the week of the LB (week 105 in my code) - for that week, we'll only have the candidates we generate.</p>\n<p>You'll want to see <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324076\" target=\"_blank\">this comment</a> from <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1789949,
      "author_name": "happylive",
      "author_url": "",
      "post_date": "05/14/2022 11:25:56",
      "content": "<p>Congrats:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1790573,
      "author_name": "ambujnayanmishra",
      "author_url": "",
      "post_date": "05/15/2022 05:11:40",
      "content": "<p>congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1791215,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/15/2022 18:28:04",
      "content": "<h3>Update:</h3>\n<p>By using cuML's FIL, notebook run time went down from 20 minutes to 12 minutes!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1792008,
      "author_name": "fajarwibowo",
      "author_url": "",
      "post_date": "05/16/2022 14:54:50",
      "content": "<p>Congratz! Awesome work. Thanks for sharing so many insights.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1782960": "### I was able to generate my final submission - including candidate retrieval, feature engineering, and training/predicting with two models - in just 20 minutes.\n\nPartially, it was from using cudf for whatever I could, and it's really amazing how much time it saves.\n\nBut I think the main reason was that I used a light-weight candidate retrieval method that gave me a good result without needing many candidates per customer, and without needing collaborative filtering or image/text processing.\n\nI had 33 million candidates (which comes out to only ~25 per customer on average).\n\nI started with @cdeotte's notebooks - [customers who bought this frequently bought this](https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this) and [Recommend Items Purchased Together](https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021).\n\nI refactored his first notebook in [this notebook](https://www.kaggle.com/code/jacob34/cdeotte-pairs-in-3-minutes), so that I could get the same result in 3 minutes.\n\nThen I made several changes:\n1. I had it only look for \"pairs\" from articles that were sold in the past week\n2. I had it find 5 \"pairs\" for each article, instead of just 1\n3. I ignored any article that wasn't sold within the past week\n4. I ignored any \"pair\" purchased by less than 2 customers\n\nWe get a very strong retrieval method, a sort of manual, time-aware collaborative filtering method - what customers with similar purchase interests were purchased *in the past week* - so it includes trend information as well.\n\nFor the second part, I followed his method of recommending \"pairs\" of items a customer had purchased in their last few weeks of history.\n\nAnd I also kept features to feed to LGBMRanker about:\n- the strength of the \"match\" (i.e. how many/what percentage of customers the \"pair\" was based on)\n- the \"source\" of the pair (i.e. how recently the original article was purchased, how many times it was purchased)\n\nOf course, I also generated the regular, standard candidates: past purchases, and 12 most popular items (I generated those based on age group).\n\nI think this retrieval method was the main reason I was able to get a good result with relatively few candidates/resources.\n\n### Update - code\nSome people asked for it, so I made a clean notebook (moved nearly all code to the repo).\nHere's [the notebook](https://www.kaggle.com/jacob34/clear-n-simple-final-shared-notebook), here's [the repo](https://github.com/JacobCP/kaggle-handm-helpers), and here's the [kaggle dataset](https://www.kaggle.com/datasets/jacob34/handmhelpers) that syncs with the repo.\n\nThe notebook runs in 20 minutes, and gets the same private LB score (it varies a bit per run).\n\nMy workflow was:\n1. I'd work with a function in the notebook (for example, a candidate generator)\n2. Once the function was working, I'd move it to the repo and point the notebook there.\n3. Any parameters of the function (for example, candidate threshold), I'd add to my `params` dictionary in the notebook.\n4. Each function had a `**kwargs` argument, so I could pass the same single `params` dictionary to every function, and the function would use the arguments it needed.\n5. If I needed to revisit a function, or work with the values generated in middle, I'd move the function/s back to the notebook. \n\n### Update #2 - cuML's FIL\nAfterwards, I refactored my code to use cuML's FIL for inference, and that brought notebook time down from 20 minutes to just 12 minutes!",
    "1782963": "good work ! thanks for sharing !",
    "1782982": "Congratulations Clear n' Simple! Great job finishing solo Silver finish 52 out of 3000 teams. Impressive results for just focusing on pairs of purchased items. Great job achieving super fast speed ups!\n\nThanks for sharing helpful discussions and helpful notebooks!",
    "1782989": "You are the best! Thanks for sharing! I learnt a lot from your notebook!",
    "1782990": "your Job is epic\n\nI learn a lot from you, thanks for all the sharing in this competition",
    "1783047": "Congrats! Will you share the code for your pipeline including candidates retrieval, training and prediction?",
    "1783153": "Well done! And mate, I would _love_ to able to see your code",
    "1783442": "Congratz! Awesome work. Thanks for sharing so many insights.",
    "1783602": "Thanks, Chris, means a lot to me!\n\nThanks for coming out of retirement for this 😉.\n\nMore seriously, I was happy you were taking part in the competition, and was hoping to continue to benefit from following your sharing, especially with your recent wins in recommendation comps, but I guess you didn't end up having the time for it.",
    "1783604": "Thanks @hervind!\n\nI took inspiration from your notebook on the ability to speed things up with proper code design.",
    "1783616": "Thanks!\nCongratulations to you on your solo silver!",
    "1783881": "Thanks @marcogorelli!\n\nI've cleaned up my notebook for you, and have updated my post with links to the code/repo/dataset.\nQuestions welcome!",
    "1783884": "Thanks, @xfffrank!\n\nI've updated my post with links to a clean notebook and my repo/dataset.",
    "1783890": "legend, thanks a tonne!",
    "1784576": "Congrats and thanks for sharing a great work!\nI'm interested in what score we could achieve only using a kaggle notebook, and yours seems to be the best...!",
    "1784678": "Congratulations on a simple solution.",
    "1785543": "jacob34 \nCongratulations for your achievement!\nTo complete such a resource-consuming competition in the kaggle environment is really challenging. \n\nI have a question. I read some code in your repository. I felt it was well structured and like the package. Did you write almost all the code during the competition, or use some kind of your own pipeline to reuse?",
    "1785913": "Thanks for the feedback, @hanejiyuto!\n\nI wrote practically all of it for this competition.\nAlthough I do hope to reuse parts of it in the future.",
    "1786576": "jacob34 \nOh, it is great!",
    "1789861": "Legend, mate, I've worked through your notebook, and it's really impressive - thanks a tonne for posting!\n\nFor generating the set for LGBMRanker, my understanding is that if you're predicting on some week (say, 103), then you:\n- generate candidates, using only information from strictly before week 103\n- any of those candidates were actually bought in 103, then they get a label of 1. Else, they get a label of 0\n- if a customer bought an article in week 103, but that article didn't appear in that customer's candidates, then that article won't be used for training\n\nIs this correct? If so, then the last part seems a bit different to what I've understood from [the other LGBMRanker example which was posted](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/309220), in which all the articles which a customer bought get a label of 1.",
    "1789949": "Congrats:)",
    "1790573": "congratulations",
    "1790956": "Pleasure!\n\nYes, that's correct.\n\nThe reasoning is, we want to train the model to do well on the week of the LB (week 105 in my code) - for that week, we'll only have the candidates we generate.\n\nYou'll want to see [this comment](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324076) from @paweljankiewicz.",
    "1791215": "### Update:\n\nBy using cuML's FIL, notebook run time went down from 20 minutes to 12 minutes!",
    "1792008": "Congratz! Awesome work. Thanks for sharing so many insights."
  },
  "source": "meta"
}