{
  "id": 324085,
  "title": "23th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/zkmrd-23th-place-solution",
  "author_name": "",
  "post_date": "2022-05-10T04:23:46.260Z",
  "votes": 39,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congratulations to all the winners, and thanks so much for hosting such an interesting competition!<br>\nAnd thanks for my teammate <a href=\"https://www.kaggle.com/zakopur0\" target=\"_blank\">@zakopur0</a> <a href=\"https://www.kaggle.com/irrohas\" target=\"_blank\">@irrohas</a> <a href=\"https://www.kaggle.com/negoto\" target=\"_blank\">@negoto</a> <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a> .<br>\nThis competitions is very hard for us.  I'll share our team ZKMRD solution.</p>\n<h3>summary</h3>\n<p><img src=\"https://user-images.githubusercontent.com/43205304/167539304-0851ae6a-15c6-4944-999a-11a27453ee6e.png\"></p>\n<h3>Candidates</h3>\n<p>We generate item candidates for each customers and each week by multi strategies.</p>\n<ul>\n<li>Most popular </li>\n<li>Past purchased</li>\n<li>User based CF(Collaborative Filtering)</li>\n<li>Item based CF</li>\n<li>Different color</li>\n</ul>\n<h3>Features</h3>\n<p>This was a big part of the reason why we were able to increase our score in the last 2week.</p>\n<p><strong>Customer static attributes</strong></p>\n<ul>\n<li>basic attributes in customers.csv</li>\n</ul>\n<p><strong>Customer dynamic attributes</strong></p>\n<ul>\n<li>purchase count</li>\n<li>last purchase flag</li>\n<li>Days/Week since last purchase</li>\n<li>price/discount of purchased items</li>\n<li>mean sales channel</li>\n<li>purchase rate by item segment</li>\n<li>repurchase rate</li>\n</ul>\n<p><strong>Article static attributes</strong></p>\n<ul>\n<li>basic attributes in articles.csv</li>\n<li>Bert sentence vector</li>\n</ul>\n<p><strong>Article dynamic attributes</strong></p>\n<ul>\n<li>trend value</li>\n<li>weekly popular ranking</li>\n<li>purchase count (1day, 2day ago, last week..)</li>\n<li>purchase count by segment (age, item group)</li>\n<li>popular rank by segment (age, item group)</li>\n<li>mean sales channel</li>\n</ul>\n<p><strong>CF features</strong></p>\n<ul>\n<li>score by item based CF</li>\n<li>score/rank by LightGCN (GCN based CF model)</li>\n</ul>\n<h3>Dataset</h3>\n<p>It was difficult to create a good dataset, but we were able to do so with the help of <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> 's comments. Many thanks.</p>\n<p>Key points about dataset</p>\n<ul>\n<li>Use only candidate examples generated, not all positive examples</li>\n<li>Use customers with at least one positive example in the candidates</li>\n<li>Remove not for sales now items from candidates</li>\n</ul>\n<h3>Model</h3>\n<p>Our team use LGBMRanker model. <br>\nSetting is very simple. Nothing special.</p>\n<h3>CV</h3>\n<p>The CV strategy is illustrated in the figure above.</p>\n<ul>\n<li>create feature and candidate: Past week from train</li>\n<li>train: 100w-103w</li>\n<li>valid: 104w</li>\n</ul>\n<p>When creating submission, we use 100-104w as train.</p>\n<h3>Post Processing</h3>\n<p>We apply two kind of post processing to all model.</p>\n<ol>\n<li>Use age based most popular items as predictions of cold start customers.</li>\n<li>Remove the low offline sales ratio items from the customers who prefer offline.</li>\n</ol>\n<h3>Ensemble</h3>\n<p>Blend 6 model predictions which is created by the different candidates and features.<br>\nEnsemble is effective for our model. LB score increased by approx. 0.001.</p>",
  "messages": [
    {
      "id": "1783017",
      "postDate": "05/10/2022 04:08:24",
      "content": "<p>Congratulations to all the winners, and thanks so much for hosting such an interesting competition!<br>\nAnd thanks for my teammate <a href=\"https://www.kaggle.com/zakopur0\" target=\"_blank\">@zakopur0</a> <a href=\"https://www.kaggle.com/irrohas\" target=\"_blank\">@irrohas</a> <a href=\"https://www.kaggle.com/negoto\" target=\"_blank\">@negoto</a> <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a> .<br>\nThis competitions is very hard for us.  I'll share our team ZKMRD solution.</p>\n<h3>summary</h3>\n<p><img src=\"https://user-images.githubusercontent.com/43205304/167539304-0851ae6a-15c6-4944-999a-11a27453ee6e.png\"></p>\n<h3>Candidates</h3>\n<p>We generate item candidates for each customers and each week by multi strategies.</p>\n<ul>\n<li>Most popular </li>\n<li>Past purchased</li>\n<li>User based CF(Collaborative Filtering)</li>\n<li>Item based CF</li>\n<li>Different color</li>\n</ul>\n<h3>Features</h3>\n<p>This was a big part of the reason why we were able to increase our score in the last 2week.</p>\n<p><strong>Customer static attributes</strong></p>\n<ul>\n<li>basic attributes in customers.csv</li>\n</ul>\n<p><strong>Customer dynamic attributes</strong></p>\n<ul>\n<li>purchase count</li>\n<li>last purchase flag</li>\n<li>Days/Week since last purchase</li>\n<li>price/discount of purchased items</li>\n<li>mean sales channel</li>\n<li>purchase rate by item segment</li>\n<li>repurchase rate</li>\n</ul>\n<p><strong>Article static attributes</strong></p>\n<ul>\n<li>basic attributes in articles.csv</li>\n<li>Bert sentence vector</li>\n</ul>\n<p><strong>Article dynamic attributes</strong></p>\n<ul>\n<li>trend value</li>\n<li>weekly popular ranking</li>\n<li>purchase count (1day, 2day ago, last week..)</li>\n<li>purchase count by segment (age, item group)</li>\n<li>popular rank by segment (age, item group)</li>\n<li>mean sales channel</li>\n</ul>\n<p><strong>CF features</strong></p>\n<ul>\n<li>score by item based CF</li>\n<li>score/rank by LightGCN (GCN based CF model)</li>\n</ul>\n<h3>Dataset</h3>\n<p>It was difficult to create a good dataset, but we were able to do so with the help of <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> 's comments. Many thanks.</p>\n<p>Key points about dataset</p>\n<ul>\n<li>Use only candidate examples generated, not all positive examples</li>\n<li>Use customers with at least one positive example in the candidates</li>\n<li>Remove not for sales now items from candidates</li>\n</ul>\n<h3>Model</h3>\n<p>Our team use LGBMRanker model. <br>\nSetting is very simple. Nothing special.</p>\n<h3>CV</h3>\n<p>The CV strategy is illustrated in the figure above.</p>\n<ul>\n<li>create feature and candidate: Past week from train</li>\n<li>train: 100w-103w</li>\n<li>valid: 104w</li>\n</ul>\n<p>When creating submission, we use 100-104w as train.</p>\n<h3>Post Processing</h3>\n<p>We apply two kind of post processing to all model.</p>\n<ol>\n<li>Use age based most popular items as predictions of cold start customers.</li>\n<li>Remove the low offline sales ratio items from the customers who prefer offline.</li>\n</ol>\n<h3>Ensemble</h3>\n<p>Blend 6 model predictions which is created by the different candidates and features.<br>\nEnsemble is effective for our model. LB score increased by approx. 0.001.</p>",
      "rawMarkdown": "Congratulations to all the winners, and thanks so much for hosting such an interesting competition!\nAnd thanks for my teammate @zakopur0 @irrohas @negoto @dehokanta .\nThis competitions is very hard for us.  I'll share our team ZKMRD solution.\n\n### summary\n<img src=\"https://user-images.githubusercontent.com/43205304/167539304-0851ae6a-15c6-4944-999a-11a27453ee6e.png\" width=\"1080px\">\n\n### Candidates\nWe generate item candidates for each customers and each week by multi strategies.\n- Most popular \n- Past purchased\n- User based CF(Collaborative Filtering)\n- Item based CF\n- Different color\n\n### Features\nThis was a big part of the reason why we were able to increase our score in the last 2week.\n\n**Customer static attributes**\n- basic attributes in customers.csv\n\n**Customer dynamic attributes**\n- purchase count\n- last purchase flag\n- Days/Week since last purchase\n- price/discount of purchased items\n- mean sales channel\n- purchase rate by item segment\n- repurchase rate\n\n**Article static attributes**\n- basic attributes in articles.csv\n- Bert sentence vector\n\n**Article dynamic attributes**\n- trend value\n- weekly popular ranking\n- purchase count (1day, 2day ago, last week..)\n- purchase count by segment (age, item group)\n- popular rank by segment (age, item group)\n- mean sales channel\n\n**CF features**\n- score by item based CF\n- score/rank by LightGCN (GCN based CF model)\n\n\n### Dataset\nIt was difficult to create a good dataset, but we were able to do so with the help of @paweljankiewicz and @lihaorocky 's comments. Many thanks.\n\nKey points about dataset\n- Use only candidate examples generated, not all positive examples\n- Use customers with at least one positive example in the candidates\n- Remove not for sales now items from candidates\n\n### Model\nOur team use LGBMRanker model. \nSetting is very simple. Nothing special.\n\n### CV\nThe CV strategy is illustrated in the figure above.\n- create feature and candidate: Past week from train\n- train: 100w-103w\n- valid: 104w\n\nWhen creating submission, we use 100-104w as train.\n\n\n### Post Processing\nWe apply two kind of post processing to all model.\n1. Use age based most popular items as predictions of cold start customers.\n2. Remove the low offline sales ratio items from the customers who prefer offline.\n\n### Ensemble\nBlend 6 model predictions which is created by the different candidates and features.\nEnsemble is effective for our model. LB score increased by approx. 0.001.",
      "votes": null
    },
    {
      "id": "1783864",
      "postDate": "05/10/2022 18:34:46",
      "content": "<p>Thanks for sharing!</p>\n<p>Beautiful visual - what did you use to make it?</p>",
      "rawMarkdown": "Thanks for sharing!\n\nBeautiful visual - what did you use to make it?",
      "votes": null
    },
    {
      "id": "1783911",
      "postDate": "05/10/2022 19:15:05",
      "content": "<p>Nice solution! Thanks for sharing.</p>",
      "rawMarkdown": "Nice solution! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1784066",
      "postDate": "05/10/2022 22:56:44",
      "content": "<p>Thanks! I use Google Slides and get icon from here.<br>\n<a href=\"https://icooon-mono.com/?lang=en\" target=\"_blank\">https://icooon-mono.com/?lang=en</a></p>",
      "rawMarkdown": "Thanks! I use Google Slides and get icon from here.\nhttps://icooon-mono.com/?lang=en",
      "votes": null
    },
    {
      "id": "1784127",
      "postDate": "05/11/2022 00:22:50",
      "content": "<p><a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> congratulations! I found the visual helpful to understanding your process</p>",
      "rawMarkdown": "kuto0633 congratulations! I found the visual helpful to understanding your process",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783864,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/10/2022 18:34:46",
      "content": "<p>Thanks for sharing!</p>\n<p>Beautiful visual - what did you use to make it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784066,
          "author_name": "kuto0633",
          "author_url": "",
          "post_date": "05/10/2022 22:56:44",
          "content": "<p>Thanks! I use Google Slides and get icon from here.<br>\n<a href=\"https://icooon-mono.com/?lang=en\" target=\"_blank\">https://icooon-mono.com/?lang=en</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783911,
      "author_name": "",
      "author_url": "",
      "post_date": "05/10/2022 19:15:05",
      "content": "<p>Nice solution! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784127,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/11/2022 00:22:50",
      "content": "<p><a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> congratulations! I found the visual helpful to understanding your process</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1783017": "Congratulations to all the winners, and thanks so much for hosting such an interesting competition!\nAnd thanks for my teammate @zakopur0 @irrohas @negoto @dehokanta .\nThis competitions is very hard for us.  I'll share our team ZKMRD solution.\n\n### summary\n<img src=\"https://user-images.githubusercontent.com/43205304/167539304-0851ae6a-15c6-4944-999a-11a27453ee6e.png\" width=\"1080px\">\n\n### Candidates\nWe generate item candidates for each customers and each week by multi strategies.\n- Most popular \n- Past purchased\n- User based CF(Collaborative Filtering)\n- Item based CF\n- Different color\n\n### Features\nThis was a big part of the reason why we were able to increase our score in the last 2week.\n\n**Customer static attributes**\n- basic attributes in customers.csv\n\n**Customer dynamic attributes**\n- purchase count\n- last purchase flag\n- Days/Week since last purchase\n- price/discount of purchased items\n- mean sales channel\n- purchase rate by item segment\n- repurchase rate\n\n**Article static attributes**\n- basic attributes in articles.csv\n- Bert sentence vector\n\n**Article dynamic attributes**\n- trend value\n- weekly popular ranking\n- purchase count (1day, 2day ago, last week..)\n- purchase count by segment (age, item group)\n- popular rank by segment (age, item group)\n- mean sales channel\n\n**CF features**\n- score by item based CF\n- score/rank by LightGCN (GCN based CF model)\n\n\n### Dataset\nIt was difficult to create a good dataset, but we were able to do so with the help of @paweljankiewicz and @lihaorocky 's comments. Many thanks.\n\nKey points about dataset\n- Use only candidate examples generated, not all positive examples\n- Use customers with at least one positive example in the candidates\n- Remove not for sales now items from candidates\n\n### Model\nOur team use LGBMRanker model. \nSetting is very simple. Nothing special.\n\n### CV\nThe CV strategy is illustrated in the figure above.\n- create feature and candidate: Past week from train\n- train: 100w-103w\n- valid: 104w\n\nWhen creating submission, we use 100-104w as train.\n\n\n### Post Processing\nWe apply two kind of post processing to all model.\n1. Use age based most popular items as predictions of cold start customers.\n2. Remove the low offline sales ratio items from the customers who prefer offline.\n\n### Ensemble\nBlend 6 model predictions which is created by the different candidates and features.\nEnsemble is effective for our model. LB score increased by approx. 0.001.",
    "1783864": "Thanks for sharing!\n\nBeautiful visual - what did you use to make it?",
    "1783911": "Nice solution! Thanks for sharing.",
    "1784066": "Thanks! I use Google Slides and get icon from here.\nhttps://icooon-mono.com/?lang=en",
    "1784127": "kuto0633 congratulations! I found the visual helpful to understanding your process"
  },
  "source": "meta"
}