{
  "id": 324185,
  "title": "8th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/kazuki-8th-place-solution",
  "author_name": "",
  "post_date": "2022-05-10T13:09:34.211062200Z",
  "votes": 39,
  "comment_count": 22,
  "views": 0,
  "content": "<p>First of all, Thanks to Kaggle and H&amp;M staff for hosting this interesting competition!<br>\nI am very happy to get solo gold medal and Kaggle Master!<br>\n I'll share my solution.</p>\n<h2>Overview</h2>\n<p>The basic strategy is the same as the other top winners, with the following three steps.</p>\n<ol>\n<li>generate item candidates for each customers</li>\n<li>create features</li>\n<li>learning to rank with LGBMRanker</li>\n</ol>\n<h2>Candidates</h2>\n<p>I select top 100 candidates for all candidate type.</p>\n<table>\n<thead>\n<tr>\n<th>candidate type</th>\n<th>ranking method</th>\n<th>cv</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>repurchased</td>\n<td>purchase date</td>\n<td>0.029</td>\n</tr>\n<tr>\n<td>same product code</td>\n<td>purchase date</td>\n<td>0.014</td>\n</tr>\n<tr>\n<td>user based cf</td>\n<td>sales in the last week</td>\n<td>0.024</td>\n</tr>\n<tr>\n<td>most popular</td>\n<td>sales in the last week</td>\n<td>0.017</td>\n</tr>\n<tr>\n<td>most popular categorized by age</td>\n<td>sales in the last week</td>\n<td>0.018</td>\n</tr>\n<tr>\n<td>most popular categorized by sales channel id</td>\n<td>sales in the last week</td>\n<td>0.016</td>\n</tr>\n</tbody>\n</table>\n<h2>Features</h2>\n<h4>Customer features</h4>\n<p>I created various customer features, but did not work.</p>\n<h4>Article features</h4>\n<ul>\n<li>article attributes ( except for article_id, product_cd ,causing overfitting )</li>\n<li>sales in the last N days or N weeks</li>\n<li>mean sales channel id</li>\n<li>release date</li>\n<li>mean price</li>\n<li>repurchase statistics</li>\n<li>product code statistics</li>\n<li>rank in the candidate type</li>\n</ul>\n<h4>Customer x Article features</h4>\n<ul>\n<li>number of purchases with the same attributes as the candidate item</li>\n<li>purchase date with the same attributes as the candidate item</li>\n<li>difference between mean customer price and mean article price</li>\n<li>score by user based cf</li>\n</ul>\n<h2>Model</h2>\n<p>I user LGBMRanker model and train model per CV.<br>\nI make weighted ensemble of model predictions per CV</p>\n<ul>\n<li>valid set : 98 -104 week ( 7 folds )</li>\n<li>training set : last 3 weeks from valid set</li>\n</ul>",
  "messages": [
    {
      "id": "1783534",
      "postDate": "05/10/2022 13:09:34",
      "content": "<p>First of all, Thanks to Kaggle and H&amp;M staff for hosting this interesting competition!<br>\nI am very happy to get solo gold medal and Kaggle Master!<br>\n I'll share my solution.</p>\n<h2>Overview</h2>\n<p>The basic strategy is the same as the other top winners, with the following three steps.</p>\n<ol>\n<li>generate item candidates for each customers</li>\n<li>create features</li>\n<li>learning to rank with LGBMRanker</li>\n</ol>\n<h2>Candidates</h2>\n<p>I select top 100 candidates for all candidate type.</p>\n<table>\n<thead>\n<tr>\n<th>candidate type</th>\n<th>ranking method</th>\n<th>cv</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>repurchased</td>\n<td>purchase date</td>\n<td>0.029</td>\n</tr>\n<tr>\n<td>same product code</td>\n<td>purchase date</td>\n<td>0.014</td>\n</tr>\n<tr>\n<td>user based cf</td>\n<td>sales in the last week</td>\n<td>0.024</td>\n</tr>\n<tr>\n<td>most popular</td>\n<td>sales in the last week</td>\n<td>0.017</td>\n</tr>\n<tr>\n<td>most popular categorized by age</td>\n<td>sales in the last week</td>\n<td>0.018</td>\n</tr>\n<tr>\n<td>most popular categorized by sales channel id</td>\n<td>sales in the last week</td>\n<td>0.016</td>\n</tr>\n</tbody>\n</table>\n<h2>Features</h2>\n<h4>Customer features</h4>\n<p>I created various customer features, but did not work.</p>\n<h4>Article features</h4>\n<ul>\n<li>article attributes ( except for article_id, product_cd ,causing overfitting )</li>\n<li>sales in the last N days or N weeks</li>\n<li>mean sales channel id</li>\n<li>release date</li>\n<li>mean price</li>\n<li>repurchase statistics</li>\n<li>product code statistics</li>\n<li>rank in the candidate type</li>\n</ul>\n<h4>Customer x Article features</h4>\n<ul>\n<li>number of purchases with the same attributes as the candidate item</li>\n<li>purchase date with the same attributes as the candidate item</li>\n<li>difference between mean customer price and mean article price</li>\n<li>score by user based cf</li>\n</ul>\n<h2>Model</h2>\n<p>I user LGBMRanker model and train model per CV.<br>\nI make weighted ensemble of model predictions per CV</p>\n<ul>\n<li>valid set : 98 -104 week ( 7 folds )</li>\n<li>training set : last 3 weeks from valid set</li>\n</ul>",
      "rawMarkdown": "First of all, Thanks to Kaggle and H&M staff for hosting this interesting competition!\nI am very happy to get solo gold medal and Kaggle Master!\n I'll share my solution.\n\n## Overview\nThe basic strategy is the same as the other top winners, with the following three steps.\n1. generate item candidates for each customers\n2. create features\n3. learning to rank with LGBMRanker\n\n\n## Candidates\nI select top 100 candidates for all candidate type.\n| candidate type | ranking method | cv |\n| --- | --- | --- |\n| repurchased | purchase date | 0.029 |\n| same product code | purchase date | 0.014 |\n| user based cf | sales in the last week | 0.024 |\n| most popular | sales in the last week | 0.017 |\n| most popular categorized by age | sales in the last week | 0.018 |\n| most popular categorized by sales channel id | sales in the last week | 0.016 |\n\n\n## Features\n#### Customer features\nI created various customer features, but did not work.\n\n#### Article features\n- article attributes ( except for article_id, product_cd ,causing overfitting )\n- sales in the last N days or N weeks\n- mean sales channel id\n- release date\n- mean price\n- repurchase statistics\n- product code statistics\n- rank in the candidate type\n\n#### Customer x Article features\n- number of purchases with the same attributes as the candidate item\n- purchase date with the same attributes as the candidate item\n- difference between mean customer price and mean article price\n- score by user based cf\n\n\n## Model\nI user LGBMRanker model and train model per CV.\nI make weighted ensemble of model predictions per CV\n- valid set : 98 -104 week ( 7 folds )\n- training set : last 3 weeks from valid set",
      "votes": null
    },
    {
      "id": "1783622",
      "postDate": "05/10/2022 14:28:32",
      "content": "<p>Nice work~  And curious about the meaning of <code>Kazuki**2-aragaki</code>and <code>Kazuki**2+shimacos*1e-3</code>. It seems that Kazuki is a mascot which could bring good luck to the followers😊</p>",
      "rawMarkdown": "Nice work~  And curious about the meaning of `Kazuki**2-aragaki `and `Kazuki**2+shimacos*1e-3`. It seems that Kazuki is a mascot which could bring good luck to the followers😊",
      "votes": null
    },
    {
      "id": "1783927",
      "postDate": "05/10/2022 19:27:23",
      "content": "<p>Thanks for this!</p>",
      "rawMarkdown": "Thanks for this!",
      "votes": null
    },
    {
      "id": "1784041",
      "postDate": "05/10/2022 21:56:42",
      "content": "<p>thanks you for sharing this!</p>",
      "rawMarkdown": "thanks you for sharing this!",
      "votes": null
    },
    {
      "id": "1784131",
      "postDate": "05/11/2022 00:31:19",
      "content": "<p><a href=\"https://www.kaggle.com/askaky\" target=\"_blank\">@askaky</a> congratulations on the result. Are there any customer features you would have generated in retrospect?</p>",
      "rawMarkdown": "askaky congratulations on the result. Are there any customer features you would have generated in retrospect?",
      "votes": null
    },
    {
      "id": "1784254",
      "postDate": "05/11/2022 04:07:49",
      "content": "<p>Thanks! By chance my first name is Kazuki and team <code>Kazuki**2+shimacos*1e-3</code> has two members named Kazuki. I don't know if <code>Kazuki**2-aragaki</code> has the name Kazuki.</p>",
      "rawMarkdown": "Thanks! By chance my first name is Kazuki and team `Kazuki**2+shimacos*1e-3` has two members named Kazuki. I don't know if `Kazuki**2-aragaki` has the name Kazuki.",
      "votes": null
    },
    {
      "id": "1784265",
      "postDate": "05/11/2022 04:24:08",
      "content": "<p>Thanks! I would have liked to embed customer purchase history in LSTM, etc.</p>",
      "rawMarkdown": "Thanks! I would have liked to embed customer purchase history in LSTM, etc.",
      "votes": null
    },
    {
      "id": "1784285",
      "postDate": "05/11/2022 05:01:23",
      "content": "<p>Thanks for sharing ur solution and it is well explained.</p>\n<p>I have a question here, what is <code>repurchased</code>? Is it meaning items that user bought in previous?</p>\n<p>Congrats.</p>",
      "rawMarkdown": "Thanks for sharing ur solution and it is well explained.\n\nI have a question here, what is `repurchased `? Is it meaning items that user bought in previous?\n\nCongrats.",
      "votes": null
    },
    {
      "id": "1784420",
      "postDate": "05/11/2022 07:03:40",
      "content": "<p>Thanks! Yes, you're right.</p>",
      "rawMarkdown": "Thanks! Yes, you're right.",
      "votes": null
    },
    {
      "id": "1784589",
      "postDate": "05/11/2022 09:39:44",
      "content": "<p>Thanks :) </p>",
      "rawMarkdown": "Thanks :)",
      "votes": null
    },
    {
      "id": "1784778",
      "postDate": "05/11/2022 13:06:06",
      "content": "<p>Thanks for sharing you solution and congratulations for solo gold!! <br>\nCould you please share with us what was the model used for <code>user based cf</code> candidate strategy?</p>",
      "rawMarkdown": "Thanks for sharing you solution and congratulations for solo gold!! \nCould you please share with us what was the model used for `user based cf` candidate strategy?",
      "votes": null
    },
    {
      "id": "1785406",
      "postDate": "05/12/2022 04:32:49",
      "content": "<p>Thanks! <br>\nMy user base cf strategy is very simple. For each item, extract a group of customers who purchased target item and calculate last 1 week group trend. Sum group trends for each customer according to their purchase history.</p>",
      "rawMarkdown": "Thanks! \nMy user base cf strategy is very simple. For each item, extract a group of customers who purchased target item and calculate last 1 week group trend. Sum group trends for each customer according to their purchase history.",
      "votes": null
    },
    {
      "id": "1785478",
      "postDate": "05/12/2022 06:12:25",
      "content": "<p>That is a really cool approach <a href=\"https://www.kaggle.com/askaky\" target=\"_blank\">@askaky</a>, huge congrats! 🥳</p>\n<p>If I am reading your post right, if I were to generate candidates for each customer from only items they purchased before, and rank them by purchased date, that would give me a cv of 0.029? </p>\n<p>In this approach, what did you do for customers that didn't have any purchases before, or had less than 12 purchases?</p>",
      "rawMarkdown": "That is a really cool approach @askaky, huge congrats! 🥳\n\nIf I am reading your post right, if I were to generate candidates for each customer from only items they purchased before, and rank them by purchased date, that would give me a cv of 0.029? \n\nIn this approach, what did you do for customers that didn't have any purchases before, or had less than 12 purchases?",
      "votes": null
    },
    {
      "id": "1785658",
      "postDate": "05/12/2022 09:59:35",
      "content": "<p>Thanks for your comment!</p>\n<blockquote>\n  <p>If I am reading your post right, if I were to generate candidates for each customer from only items they purchased before, and rank them by purchased date, that would give me a cv of 0.029? </p>\n</blockquote>\n<p>Yes, you're right.</p>\n<blockquote>\n  <p>In this approach, what did you do for customers that didn't have any purchases before, or had less than 12 purchases?</p>\n</blockquote>\n<p>Such customers have no prediction or less than 12 prediction. If we fill in blank predictions with most popular candidates, cv score is up.</p>",
      "rawMarkdown": "Thanks for your comment!\n\n> If I am reading your post right, if I were to generate candidates for each customer from only items they purchased before, and rank them by purchased date, that would give me a cv of 0.029? \n\nYes, you're right.\n\n> In this approach, what did you do for customers that didn't have any purchases before, or had less than 12 purchases?\n\nSuch customers have no prediction or less than 12 prediction. If we fill in blank predictions with most popular candidates, cv score is up.",
      "votes": null
    },
    {
      "id": "1785919",
      "postDate": "05/12/2022 13:47:05",
      "content": "<p>Thanks for sharing that!</p>\n<blockquote>\n  <p>Sum group trends for each customer  </p>\n</blockquote>\n<p>Can you explain how you do that?<br>\nI was doing something similar, and that was the part I struggled with.</p>",
      "rawMarkdown": "Thanks for sharing that!\n\n> Sum group trends for each customer  \n\nCan you explain how you do that?\nI was doing something similar, and that was the part I struggled with.",
      "votes": null
    },
    {
      "id": "1785947",
      "postDate": "05/12/2022 14:07:24",
      "content": "<p>Hi. I have another questions if you don't mind.</p>\n<p>1 - What you consider \"most popular\"? Is the most purchased items? <br>\nIn this candidate strategy you used \"sales in the last week\" as ranking method. <br>\nFor instance, for target_week = 105, most popular would be this as follows?</p>\n<pre><code># if using 12 candidates for this strategy\ntrain[train['week'] == 104]['article_id'].value_counts().index.tolist()[:12]\n</code></pre>\n<p>2 - You used 6 strategies. When you say \"I select top 100 candidates for all candidate type.\" you mean you ended with 600 candidates (100 for each candidate strategy) or you mean 100 candidates after binding them all?</p>\n<p>Thanks for your help.</p>",
      "rawMarkdown": "Hi. I have another questions if you don't mind.\n\n1 - What you consider \"most popular\"? Is the most purchased items? \nIn this candidate strategy you used \"sales in the last week\" as ranking method. \nFor instance, for target_week = 105, most popular would be this as follows?\n```\n# if using 12 candidates for this strategy\ntrain[train['week'] == 104]['article_id'].value_counts().index.tolist()[:12]\n```\n\n2 - You used 6 strategies. When you say \"I select top 100 candidates for all candidate type.\" you mean you ended with 600 candidates (100 for each candidate strategy) or you mean 100 candidates after binding them all?\n\nThanks for your help.",
      "votes": null
    },
    {
      "id": "1786060",
      "postDate": "05/12/2022 15:25:23",
      "content": "<p>For example, if customer X purchased three items A,B,C, I generated candidates for customer X as follows.</p>\n<ol>\n<li>extract three customer groups A,B,C that has purchased each item</li>\n<li>calculate sales in the last 1 week for each customer group A,B,C</li>\n<li>simply add sales A,B and C, and select top 100 as candidates</li>\n</ol>\n<p>Does this answer your question?</p>",
      "rawMarkdown": "For example, if customer X purchased three items A,B,C, I generated candidates for customer X as follows.\n1. extract three customer groups A,B,C that has purchased each item\n2. calculate sales in the last 1 week for each customer group A,B,C\n3. simply add sales A,B and C, and select top 100 as candidates\n\nDoes this answer your question?",
      "votes": null
    },
    {
      "id": "1786076",
      "postDate": "05/12/2022 15:34:36",
      "content": "<pre><code>1 - What you consider \"most popular\"? Is the most purchased items?\nIn this candidate strategy you used \"sales in the last week\" as ranking method.\nFor instance, for target_week = 105, most popular would be this as follows?\n\n# if using 12 candidates for this strategy\ntrain[train['week'] == 104]['article_id'].value_counts().index.tolist()[:12]\n</code></pre>\n<p>Yes, you're right.</p>\n<blockquote>\n  <p>2 - You used 6 strategies. When you say \"I select top 100 candidates for all candidate type.\" you mean you ended with 600 candidates (100 for each candidate strategy) or you mean 100 candidates after binding them all?</p>\n</blockquote>\n<p>600 candidates is correct. The actual number of candidates will be much smaller 600 because of the overlap between strategies.</p>",
      "rawMarkdown": "```\n1 - What you consider \"most popular\"? Is the most purchased items?\nIn this candidate strategy you used \"sales in the last week\" as ranking method.\nFor instance, for target_week = 105, most popular would be this as follows?\n\n# if using 12 candidates for this strategy\ntrain[train['week'] == 104]['article_id'].value_counts().index.tolist()[:12]\n```\nYes, you're right.\n\n> 2 - You used 6 strategies. When you say \"I select top 100 candidates for all candidate type.\" you mean you ended with 600 candidates (100 for each candidate strategy) or you mean 100 candidates after binding them all?\n\n600 candidates is correct. The actual number of candidates will be much smaller 600 because of the overlap between strategies.",
      "votes": null
    },
    {
      "id": "1786082",
      "postDate": "05/12/2022 15:39:56",
      "content": "<p>Thanks for your answer.    <br>\nI don't know if I understood correctly:</p>\n<blockquote>\n  <p>The actual number of candidates will be much smaller 600 because of the overlap between strategies.</p>\n</blockquote>\n<p>This means that you remove overlap? I mean, you bind all candidates for a customer and take the uniques like <code>set(candidates_strategy_1, candidates_strategy_2, ...)</code> for each customer?</p>",
      "rawMarkdown": "Thanks for your answer.    \nI don't know if I understood correctly:\n> The actual number of candidates will be much smaller 600 because of the overlap between strategies.\n\nThis means that you remove overlap? I mean, you bind all candidates for a customer and take the uniques like `set(candidates_strategy_1, candidates_strategy_2, ...)` for each customer?",
      "votes": null
    },
    {
      "id": "1786107",
      "postDate": "05/12/2022 15:54:58",
      "content": "<p>Yes, I remove overlap candidates for each customer.</p>",
      "rawMarkdown": "Yes, I remove overlap candidates for each customer.",
      "votes": null
    },
    {
      "id": "1786154",
      "postDate": "05/12/2022 16:35:23",
      "content": "<p>It does, thank you!</p>\n<p>I did step 1 and 2 the same as you.</p>\n<p>But for step 3, I took top X from each customer group A, B, C, and used all of them as candidates.</p>\n<p>Problem was that customers with more items would have more candidates than others customers this way, and it wasn't clear to me how to rank them if I wanted to take top X of the total.<br>\nAnother problem was that I wasn't weighting bigger groups more than smaller groups.</p>\n<p>Your way seems to solve that!</p>",
      "rawMarkdown": "It does, thank you!\n\nI did step 1 and 2 the same as you.\n\nBut for step 3, I took top X from each customer group A, B, C, and used all of them as candidates.\n\nProblem was that customers with more items would have more candidates than others customers this way, and it wasn't clear to me how to rank them if I wanted to take top X of the total.\nAnother problem was that I wasn't weighting bigger groups more than smaller groups.\n\nYour way seems to solve that!",
      "votes": null
    },
    {
      "id": "1790717",
      "postDate": "05/15/2022 08:45:48",
      "content": "<p>What is the mean of “product code statistics”？How do yo build it?</p>",
      "rawMarkdown": "What is the mean of “product code statistics”？How do yo build it?",
      "votes": null
    },
    {
      "id": "1791884",
      "postDate": "05/16/2022 12:50:34",
      "content": "<p>I created “product code statistics” features as follows.</p>\n<ul>\n<li>number of product code multi-purchased  / number of product code purchased( for each customer id )</li>\n<li>number of candidate item's product code purchased ( for each customer id )</li>\n<li>mean number of candidate item's product code purchased ( for each product code )</li>\n</ul>",
      "rawMarkdown": "I created “product code statistics” features as follows.\n- number of product code multi-purchased  / number of product code purchased( for each customer id )\n- number of candidate item's product code purchased ( for each customer id )\n- mean number of candidate item's product code purchased ( for each product code )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783622,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "05/10/2022 14:28:32",
      "content": "<p>Nice work~  And curious about the meaning of <code>Kazuki**2-aragaki</code>and <code>Kazuki**2+shimacos*1e-3</code>. It seems that Kazuki is a mascot which could bring good luck to the followers😊</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784254,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/11/2022 04:07:49",
          "content": "<p>Thanks! By chance my first name is Kazuki and team <code>Kazuki**2+shimacos*1e-3</code> has two members named Kazuki. I don't know if <code>Kazuki**2-aragaki</code> has the name Kazuki.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783927,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/10/2022 19:27:23",
      "content": "<p>Thanks for this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784041,
      "author_name": "akitrc",
      "author_url": "",
      "post_date": "05/10/2022 21:56:42",
      "content": "<p>thanks you for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784131,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/11/2022 00:31:19",
      "content": "<p><a href=\"https://www.kaggle.com/askaky\" target=\"_blank\">@askaky</a> congratulations on the result. Are there any customer features you would have generated in retrospect?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784265,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/11/2022 04:24:08",
          "content": "<p>Thanks! I would have liked to embed customer purchase history in LSTM, etc.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1784285,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "05/11/2022 05:01:23",
      "content": "<p>Thanks for sharing ur solution and it is well explained.</p>\n<p>I have a question here, what is <code>repurchased</code>? Is it meaning items that user bought in previous?</p>\n<p>Congrats.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784420,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/11/2022 07:03:40",
          "content": "<p>Thanks! Yes, you're right.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1784589,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "05/11/2022 09:39:44",
          "content": "<p>Thanks :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1784778,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "05/11/2022 13:06:06",
      "content": "<p>Thanks for sharing you solution and congratulations for solo gold!! <br>\nCould you please share with us what was the model used for <code>user based cf</code> candidate strategy?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1785406,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/12/2022 04:32:49",
          "content": "<p>Thanks! <br>\nMy user base cf strategy is very simple. For each item, extract a group of customers who purchased target item and calculate last 1 week group trend. Sum group trends for each customer according to their purchase history.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1785919,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/12/2022 13:47:05",
          "content": "<p>Thanks for sharing that!</p>\n<blockquote>\n  <p>Sum group trends for each customer  </p>\n</blockquote>\n<p>Can you explain how you do that?<br>\nI was doing something similar, and that was the part I struggled with.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786060,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/12/2022 15:25:23",
          "content": "<p>For example, if customer X purchased three items A,B,C, I generated candidates for customer X as follows.</p>\n<ol>\n<li>extract three customer groups A,B,C that has purchased each item</li>\n<li>calculate sales in the last 1 week for each customer group A,B,C</li>\n<li>simply add sales A,B and C, and select top 100 as candidates</li>\n</ol>\n<p>Does this answer your question?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786154,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/12/2022 16:35:23",
          "content": "<p>It does, thank you!</p>\n<p>I did step 1 and 2 the same as you.</p>\n<p>But for step 3, I took top X from each customer group A, B, C, and used all of them as candidates.</p>\n<p>Problem was that customers with more items would have more candidates than others customers this way, and it wasn't clear to me how to rank them if I wanted to take top X of the total.<br>\nAnother problem was that I wasn't weighting bigger groups more than smaller groups.</p>\n<p>Your way seems to solve that!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1785478,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "05/12/2022 06:12:25",
      "content": "<p>That is a really cool approach <a href=\"https://www.kaggle.com/askaky\" target=\"_blank\">@askaky</a>, huge congrats! 🥳</p>\n<p>If I am reading your post right, if I were to generate candidates for each customer from only items they purchased before, and rank them by purchased date, that would give me a cv of 0.029? </p>\n<p>In this approach, what did you do for customers that didn't have any purchases before, or had less than 12 purchases?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1785658,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/12/2022 09:59:35",
          "content": "<p>Thanks for your comment!</p>\n<blockquote>\n  <p>If I am reading your post right, if I were to generate candidates for each customer from only items they purchased before, and rank them by purchased date, that would give me a cv of 0.029? </p>\n</blockquote>\n<p>Yes, you're right.</p>\n<blockquote>\n  <p>In this approach, what did you do for customers that didn't have any purchases before, or had less than 12 purchases?</p>\n</blockquote>\n<p>Such customers have no prediction or less than 12 prediction. If we fill in blank predictions with most popular candidates, cv score is up.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1785947,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "05/12/2022 14:07:24",
      "content": "<p>Hi. I have another questions if you don't mind.</p>\n<p>1 - What you consider \"most popular\"? Is the most purchased items? <br>\nIn this candidate strategy you used \"sales in the last week\" as ranking method. <br>\nFor instance, for target_week = 105, most popular would be this as follows?</p>\n<pre><code># if using 12 candidates for this strategy\ntrain[train['week'] == 104]['article_id'].value_counts().index.tolist()[:12]\n</code></pre>\n<p>2 - You used 6 strategies. When you say \"I select top 100 candidates for all candidate type.\" you mean you ended with 600 candidates (100 for each candidate strategy) or you mean 100 candidates after binding them all?</p>\n<p>Thanks for your help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1786076,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/12/2022 15:34:36",
          "content": "<pre><code>1 - What you consider \"most popular\"? Is the most purchased items?\nIn this candidate strategy you used \"sales in the last week\" as ranking method.\nFor instance, for target_week = 105, most popular would be this as follows?\n\n# if using 12 candidates for this strategy\ntrain[train['week'] == 104]['article_id'].value_counts().index.tolist()[:12]\n</code></pre>\n<p>Yes, you're right.</p>\n<blockquote>\n  <p>2 - You used 6 strategies. When you say \"I select top 100 candidates for all candidate type.\" you mean you ended with 600 candidates (100 for each candidate strategy) or you mean 100 candidates after binding them all?</p>\n</blockquote>\n<p>600 candidates is correct. The actual number of candidates will be much smaller 600 because of the overlap between strategies.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786082,
          "author_name": "igorkf",
          "author_url": "",
          "post_date": "05/12/2022 15:39:56",
          "content": "<p>Thanks for your answer.    <br>\nI don't know if I understood correctly:</p>\n<blockquote>\n  <p>The actual number of candidates will be much smaller 600 because of the overlap between strategies.</p>\n</blockquote>\n<p>This means that you remove overlap? I mean, you bind all candidates for a customer and take the uniques like <code>set(candidates_strategy_1, candidates_strategy_2, ...)</code> for each customer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786107,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/12/2022 15:54:58",
          "content": "<p>Yes, I remove overlap candidates for each customer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1790717,
      "author_name": "dengniewei",
      "author_url": "",
      "post_date": "05/15/2022 08:45:48",
      "content": "<p>What is the mean of “product code statistics”？How do yo build it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1791884,
          "author_name": "askaky",
          "author_url": "",
          "post_date": "05/16/2022 12:50:34",
          "content": "<p>I created “product code statistics” features as follows.</p>\n<ul>\n<li>number of product code multi-purchased  / number of product code purchased( for each customer id )</li>\n<li>number of candidate item's product code purchased ( for each customer id )</li>\n<li>mean number of candidate item's product code purchased ( for each product code )</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1783534": "First of all, Thanks to Kaggle and H&M staff for hosting this interesting competition!\nI am very happy to get solo gold medal and Kaggle Master!\n I'll share my solution.\n\n## Overview\nThe basic strategy is the same as the other top winners, with the following three steps.\n1. generate item candidates for each customers\n2. create features\n3. learning to rank with LGBMRanker\n\n\n## Candidates\nI select top 100 candidates for all candidate type.\n| candidate type | ranking method | cv |\n| --- | --- | --- |\n| repurchased | purchase date | 0.029 |\n| same product code | purchase date | 0.014 |\n| user based cf | sales in the last week | 0.024 |\n| most popular | sales in the last week | 0.017 |\n| most popular categorized by age | sales in the last week | 0.018 |\n| most popular categorized by sales channel id | sales in the last week | 0.016 |\n\n\n## Features\n#### Customer features\nI created various customer features, but did not work.\n\n#### Article features\n- article attributes ( except for article_id, product_cd ,causing overfitting )\n- sales in the last N days or N weeks\n- mean sales channel id\n- release date\n- mean price\n- repurchase statistics\n- product code statistics\n- rank in the candidate type\n\n#### Customer x Article features\n- number of purchases with the same attributes as the candidate item\n- purchase date with the same attributes as the candidate item\n- difference between mean customer price and mean article price\n- score by user based cf\n\n\n## Model\nI user LGBMRanker model and train model per CV.\nI make weighted ensemble of model predictions per CV\n- valid set : 98 -104 week ( 7 folds )\n- training set : last 3 weeks from valid set",
    "1783622": "Nice work~  And curious about the meaning of `Kazuki**2-aragaki `and `Kazuki**2+shimacos*1e-3`. It seems that Kazuki is a mascot which could bring good luck to the followers😊",
    "1783927": "Thanks for this!",
    "1784041": "thanks you for sharing this!",
    "1784131": "askaky congratulations on the result. Are there any customer features you would have generated in retrospect?",
    "1784254": "Thanks! By chance my first name is Kazuki and team `Kazuki**2+shimacos*1e-3` has two members named Kazuki. I don't know if `Kazuki**2-aragaki` has the name Kazuki.",
    "1784265": "Thanks! I would have liked to embed customer purchase history in LSTM, etc.",
    "1784285": "Thanks for sharing ur solution and it is well explained.\n\nI have a question here, what is `repurchased `? Is it meaning items that user bought in previous?\n\nCongrats.",
    "1784420": "Thanks! Yes, you're right.",
    "1784589": "Thanks :)",
    "1784778": "Thanks for sharing you solution and congratulations for solo gold!! \nCould you please share with us what was the model used for `user based cf` candidate strategy?",
    "1785406": "Thanks! \nMy user base cf strategy is very simple. For each item, extract a group of customers who purchased target item and calculate last 1 week group trend. Sum group trends for each customer according to their purchase history.",
    "1785478": "That is a really cool approach @askaky, huge congrats! 🥳\n\nIf I am reading your post right, if I were to generate candidates for each customer from only items they purchased before, and rank them by purchased date, that would give me a cv of 0.029? \n\nIn this approach, what did you do for customers that didn't have any purchases before, or had less than 12 purchases?",
    "1785658": "Thanks for your comment!\n\n> If I am reading your post right, if I were to generate candidates for each customer from only items they purchased before, and rank them by purchased date, that would give me a cv of 0.029? \n\nYes, you're right.\n\n> In this approach, what did you do for customers that didn't have any purchases before, or had less than 12 purchases?\n\nSuch customers have no prediction or less than 12 prediction. If we fill in blank predictions with most popular candidates, cv score is up.",
    "1785919": "Thanks for sharing that!\n\n> Sum group trends for each customer  \n\nCan you explain how you do that?\nI was doing something similar, and that was the part I struggled with.",
    "1785947": "Hi. I have another questions if you don't mind.\n\n1 - What you consider \"most popular\"? Is the most purchased items? \nIn this candidate strategy you used \"sales in the last week\" as ranking method. \nFor instance, for target_week = 105, most popular would be this as follows?\n```\n# if using 12 candidates for this strategy\ntrain[train['week'] == 104]['article_id'].value_counts().index.tolist()[:12]\n```\n\n2 - You used 6 strategies. When you say \"I select top 100 candidates for all candidate type.\" you mean you ended with 600 candidates (100 for each candidate strategy) or you mean 100 candidates after binding them all?\n\nThanks for your help.",
    "1786060": "For example, if customer X purchased three items A,B,C, I generated candidates for customer X as follows.\n1. extract three customer groups A,B,C that has purchased each item\n2. calculate sales in the last 1 week for each customer group A,B,C\n3. simply add sales A,B and C, and select top 100 as candidates\n\nDoes this answer your question?",
    "1786076": "```\n1 - What you consider \"most popular\"? Is the most purchased items?\nIn this candidate strategy you used \"sales in the last week\" as ranking method.\nFor instance, for target_week = 105, most popular would be this as follows?\n\n# if using 12 candidates for this strategy\ntrain[train['week'] == 104]['article_id'].value_counts().index.tolist()[:12]\n```\nYes, you're right.\n\n> 2 - You used 6 strategies. When you say \"I select top 100 candidates for all candidate type.\" you mean you ended with 600 candidates (100 for each candidate strategy) or you mean 100 candidates after binding them all?\n\n600 candidates is correct. The actual number of candidates will be much smaller 600 because of the overlap between strategies.",
    "1786082": "Thanks for your answer.    \nI don't know if I understood correctly:\n> The actual number of candidates will be much smaller 600 because of the overlap between strategies.\n\nThis means that you remove overlap? I mean, you bind all candidates for a customer and take the uniques like `set(candidates_strategy_1, candidates_strategy_2, ...)` for each customer?",
    "1786107": "Yes, I remove overlap candidates for each customer.",
    "1786154": "It does, thank you!\n\nI did step 1 and 2 the same as you.\n\nBut for step 3, I took top X from each customer group A, B, C, and used all of them as candidates.\n\nProblem was that customers with more items would have more candidates than others customers this way, and it wasn't clear to me how to rank them if I wanted to take top X of the total.\nAnother problem was that I wasn't weighting bigger groups more than smaller groups.\n\nYour way seems to solve that!",
    "1790717": "What is the mean of “product code statistics”？How do yo build it?",
    "1791884": "I created “product code statistics” features as follows.\n- number of product code multi-purchased  / number of product code purchased( for each customer id )\n- number of candidate item's product code purchased ( for each customer id )\n- mean number of candidate item's product code purchased ( for each product code )"
  },
  "source": "meta"
}