{
  "id": 324350,
  "title": "26th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/324350",
  "author_name": "Sumire",
  "post_date": "2022-05-11T07:39:00.438000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle team, H&amp;M staff and all competitors.<br>\nIt was so interesting and greatly well-designed competition (with a little shake) that I was excited to join.<br>\nIt is grateful that our team (<a href=\"https://www.kaggle.com/ruiknemui\" target=\"_blank\">@ruik</a> and me) got the highest rank in private leaderboard, but we also feel \"the wall\" to Gold medal.<br>\nWe'd like to keep efforts until we become Master.<br>\nAnyway, let me share our solution. I'm so happy if you give comments or advises for our approach.</p>\n<h2>Pipeline</h2>\n<p>Not so different from others, very simple.</p>\n<ol>\n<li>Generate base candidates for each customer from their historical transactions in last N(=96) weeks. Transactions of articles which don't appear in historical transactions are not included in this time, so this approach is based on <strong>\"What will they repurchase?\"</strong></li>\n<li>Add candidates</li>\n<li>Create features</li>\n<li>Fit/predict by LGBMRanker</li>\n<li>Postprocess</li>\n</ol>\n<h2>How to add candidates</h2>\n<ul>\n<li>Calculate article embeddings using word2vec as item2vec, then select N(=48) candidates whose embeddings have high similarity with predicted vector as \"next purchase\".</li>\n<li>Add candidates have <strong>product codes</strong> as same as those of \"pair\" with each customer's last N(=3) transactions. (thanks to <a href=\"https://www.kaggle.com/code/titericz/article-id-pairs-in-3s-using-cudf\" target=\"_blank\">Mr.Giba's notebook</a>)<br>\nPlease note that these method are applied to customers <strong>who have enough historical transactions</strong>.</li>\n</ul>\n<h2>Create features</h2>\n<p>Not so different from winners.</p>\n<ul>\n<li><strong>article features</strong><ul>\n<li># of sales in last N week (for not only article itself but also attributes which article is belonged to, like \"department_name\") + that segmented by age</li>\n<li>max/min/median price purchased (smoothing with global median)</li>\n<li>online purchased rate</li>\n<li># of elapsed days from last purchase among all customers</li>\n<li># of elapsed days from article's first sell</li></ul></li>\n<li><strong>customer features</strong><ul>\n<li>age: binned and adjusted (10 if age &lt; 20, 60 if age &gt;= 60)</li>\n<li>\"club_member_status\",\"fashion_news_frequency\"</li>\n<li>attribute: estimated from index_code of articles they purchased (ex. if a customer usually purchases \"Ladieswear\", she is estimated as \"Lady\")</li>\n<li>max/min/median price purchased (smoothing with global median)</li>\n<li># of elapsed days from customer's last purchase</li></ul></li>\n<li><strong>customer x article features</strong><ul>\n<li># of elapsed days from last purchase for each article/customer combination</li>\n<li># of each article/customer pair appeared in historical transactions</li>\n<li>purchased rate(transactions) of article for each age/attribute</li>\n<li>ratio of article's price(max/min/median) to customer's price(max/min/median) (ex. if ratio of article's price min to customer's price max is over 1, we can estimate <strong>\"he has no budget to purchase the item.\"</strong>)</li></ul></li>\n</ul>\n<h2>Fit/predict by LGBMRanker</h2>\n<p>We divide customers to 3 folds, 30/35/35% (actually, at first aiming to conduct experiments quickly).<br>\nNext, we train and predict LGBMRanker for each fold of customer. Then apply weighted averaging (0.3/0.35/0.35) to prediction.<br>\nTrain/valid period is only 1 set, valid set is last 1 week and train set is last 3 weeks before valid.<br>\nWe retrain model with last 3 weeks data when predicting test after parameter tuning with optuna.</p>\n<h2>Postprocess</h2>\n<p>We apply postprocess to prediction of <strong>customers who have no or less than 12 candidates (transactions) in last N(=96) weeks.</strong></p>\n<ul>\n<li>We predict top 12 sales in last 1 week among customers with no candidates in last N(=96) weeks to those customers (+0.0002)</li>\n<li>We predict top 12-K sales in last 1 week for each age group to customers with K(&lt;12) candidates (+0.0001)</li>\n</ul>\n<p>These method has improved score about 0.0003. (slight, but important improvement)</p>\n<h2>What did not work</h2>\n<ul>\n<li>Candidates<ul>\n<li>Using pictures (adding candidates which have most similar pictures and have the same \"product_type_no\", \"index_group_no\")</li>\n<li>Using detail description (as same as using pictures)</li></ul></li>\n</ul>\n<p>I think what these method did not work is due to my mistake in implementation. When adding these candidates, it looks leaked. I have realized I have to make much more efforts to utilize ideas.</p>\n<ul>\n<li><p>Features</p>\n<ul>\n<li>FN, active, postal_code (even count frequency)</li>\n<li>price feat, online or offline purchased</li>\n<li>on sale label: it would be because most articles are on sale for a few days suddenly, not continously</li></ul></li>\n<li><p>Model</p>\n<ul>\n<li>LGBMClassifier (predict purchase or not): In fact, it worked well when LB is under 0.03. Score has jumped up from 0.0274 to 0.0287. Mysterious…</li></ul></li>\n<li><p>Postprocess</p>\n<ul>\n<li>Replace top sales to \"less confident\" prediction: I did not have enough time to try, so there would be better way…</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 1784466,
      "postDate": "2022-05-11T07:39:00.440Z",
      "content": "<p>First of all, thanks to Kaggle team, H&amp;M staff and all competitors.<br>\nIt was so interesting and greatly well-designed competition (with a little shake) that I was excited to join.<br>\nIt is grateful that our team (<a href=\"https://www.kaggle.com/ruiknemui\" target=\"_blank\">@ruik</a> and me) got the highest rank in private leaderboard, but we also feel \"the wall\" to Gold medal.<br>\nWe'd like to keep efforts until we become Master.<br>\nAnyway, let me share our solution. I'm so happy if you give comments or advises for our approach.</p>\n<h2>Pipeline</h2>\n<p>Not so different from others, very simple.</p>\n<ol>\n<li>Generate base candidates for each customer from their historical transactions in last N(=96) weeks. Transactions of articles which don't appear in historical transactions are not included in this time, so this approach is based on <strong>\"What will they repurchase?\"</strong></li>\n<li>Add candidates</li>\n<li>Create features</li>\n<li>Fit/predict by LGBMRanker</li>\n<li>Postprocess</li>\n</ol>\n<h2>How to add candidates</h2>\n<ul>\n<li>Calculate article embeddings using word2vec as item2vec, then select N(=48) candidates whose embeddings have high similarity with predicted vector as \"next purchase\".</li>\n<li>Add candidates have <strong>product codes</strong> as same as those of \"pair\" with each customer's last N(=3) transactions. (thanks to <a href=\"https://www.kaggle.com/code/titericz/article-id-pairs-in-3s-using-cudf\" target=\"_blank\">Mr.Giba's notebook</a>)<br>\nPlease note that these method are applied to customers <strong>who have enough historical transactions</strong>.</li>\n</ul>\n<h2>Create features</h2>\n<p>Not so different from winners.</p>\n<ul>\n<li><strong>article features</strong><ul>\n<li># of sales in last N week (for not only article itself but also attributes which article is belonged to, like \"department_name\") + that segmented by age</li>\n<li>max/min/median price purchased (smoothing with global median)</li>\n<li>online purchased rate</li>\n<li># of elapsed days from last purchase among all customers</li>\n<li># of elapsed days from article's first sell</li></ul></li>\n<li><strong>customer features</strong><ul>\n<li>age: binned and adjusted (10 if age &lt; 20, 60 if age &gt;= 60)</li>\n<li>\"club_member_status\",\"fashion_news_frequency\"</li>\n<li>attribute: estimated from index_code of articles they purchased (ex. if a customer usually purchases \"Ladieswear\", she is estimated as \"Lady\")</li>\n<li>max/min/median price purchased (smoothing with global median)</li>\n<li># of elapsed days from customer's last purchase</li></ul></li>\n<li><strong>customer x article features</strong><ul>\n<li># of elapsed days from last purchase for each article/customer combination</li>\n<li># of each article/customer pair appeared in historical transactions</li>\n<li>purchased rate(transactions) of article for each age/attribute</li>\n<li>ratio of article's price(max/min/median) to customer's price(max/min/median) (ex. if ratio of article's price min to customer's price max is over 1, we can estimate <strong>\"he has no budget to purchase the item.\"</strong>)</li></ul></li>\n</ul>\n<h2>Fit/predict by LGBMRanker</h2>\n<p>We divide customers to 3 folds, 30/35/35% (actually, at first aiming to conduct experiments quickly).<br>\nNext, we train and predict LGBMRanker for each fold of customer. Then apply weighted averaging (0.3/0.35/0.35) to prediction.<br>\nTrain/valid period is only 1 set, valid set is last 1 week and train set is last 3 weeks before valid.<br>\nWe retrain model with last 3 weeks data when predicting test after parameter tuning with optuna.</p>\n<h2>Postprocess</h2>\n<p>We apply postprocess to prediction of <strong>customers who have no or less than 12 candidates (transactions) in last N(=96) weeks.</strong></p>\n<ul>\n<li>We predict top 12 sales in last 1 week among customers with no candidates in last N(=96) weeks to those customers (+0.0002)</li>\n<li>We predict top 12-K sales in last 1 week for each age group to customers with K(&lt;12) candidates (+0.0001)</li>\n</ul>\n<p>These method has improved score about 0.0003. (slight, but important improvement)</p>\n<h2>What did not work</h2>\n<ul>\n<li>Candidates<ul>\n<li>Using pictures (adding candidates which have most similar pictures and have the same \"product_type_no\", \"index_group_no\")</li>\n<li>Using detail description (as same as using pictures)</li></ul></li>\n</ul>\n<p>I think what these method did not work is due to my mistake in implementation. When adding these candidates, it looks leaked. I have realized I have to make much more efforts to utilize ideas.</p>\n<ul>\n<li><p>Features</p>\n<ul>\n<li>FN, active, postal_code (even count frequency)</li>\n<li>price feat, online or offline purchased</li>\n<li>on sale label: it would be because most articles are on sale for a few days suddenly, not continously</li></ul></li>\n<li><p>Model</p>\n<ul>\n<li>LGBMClassifier (predict purchase or not): In fact, it worked well when LB is under 0.03. Score has jumped up from 0.0274 to 0.0287. Mysterious…</li></ul></li>\n<li><p>Postprocess</p>\n<ul>\n<li>Replace top sales to \"less confident\" prediction: I did not have enough time to try, so there would be better way…</li></ul></li>\n</ul>",
      "rawMarkdown": "First of all, thanks to Kaggle team, H&M staff and all competitors.\nIt was so interesting and greatly well-designed competition (with a little shake) that I was excited to join.\nIt is grateful that our team ([@ruik](https://www.kaggle.com/ruiknemui) and me) got the highest rank in private leaderboard, but we also feel \"the wall\" to Gold medal.\nWe'd like to keep efforts until we become Master.\nAnyway, let me share our solution. I'm so happy if you give comments or advises for our approach.\n\n## Pipeline\nNot so different from others, very simple.\n1. Generate base candidates for each customer from their historical transactions in last N(=96) weeks. Transactions of articles which don't appear in historical transactions are not included in this time, so this approach is based on **\"What will they repurchase?\"**\n2. Add candidates\n3. Create features\n4. Fit/predict by LGBMRanker\n5. Postprocess\n\n## How to add candidates\n* Calculate article embeddings using word2vec as item2vec, then select N(=48) candidates whose embeddings have high similarity with predicted vector as \"next purchase\".\n* Add candidates have **product codes** as same as those of \"pair\" with each customer's last N(=3) transactions. (thanks to [Mr.Giba's notebook](https://www.kaggle.com/code/titericz/article-id-pairs-in-3s-using-cudf))\nPlease note that these method are applied to customers **who have enough historical transactions**.\n\n## Create features\nNot so different from winners.\n* **article features**\n  - # of sales in last N week (for not only article itself but also attributes which article is belonged to, like \"department_name\") + that segmented by age\n  - max/min/median price purchased (smoothing with global median)\n  - online purchased rate\n  - # of elapsed days from last purchase among all customers\n  - # of elapsed days from article's first sell\n* **customer features**\n  - age: binned and adjusted (10 if age < 20, 60 if age >= 60)\n  - \"club_member_status\",\"fashion_news_frequency\"\n  - attribute: estimated from index_code of articles they purchased (ex. if a customer usually purchases \"Ladieswear\", she is estimated as \"Lady\")\n  - max/min/median price purchased (smoothing with global median)\n  - # of elapsed days from customer's last purchase\n* **customer x article features**\n  - # of elapsed days from last purchase for each article/customer combination\n  - # of each article/customer pair appeared in historical transactions\n  - purchased rate(transactions) of article for each age/attribute\n  - ratio of article's price(max/min/median) to customer's price(max/min/median) (ex. if ratio of article's price min to customer's price max is over 1, we can estimate **\"he has no budget to purchase the item.\"**)\n\n## Fit/predict by LGBMRanker\nWe divide customers to 3 folds, 30/35/35% (actually, at first aiming to conduct experiments quickly).\nNext, we train and predict LGBMRanker for each fold of customer. Then apply weighted averaging (0.3/0.35/0.35) to prediction.\nTrain/valid period is only 1 set, valid set is last 1 week and train set is last 3 weeks before valid.\nWe retrain model with last 3 weeks data when predicting test after parameter tuning with optuna.\n\n## Postprocess\nWe apply postprocess to prediction of **customers who have no or less than 12 candidates (transactions) in last N(=96) weeks.**\n* We predict top 12 sales in last 1 week among customers with no candidates in last N(=96) weeks to those customers (+0.0002)\n* We predict top 12-K sales in last 1 week for each age group to customers with K(<12) candidates (+0.0001)\n\nThese method has improved score about 0.0003. (slight, but important improvement)\n\n## What did not work\n\n* Candidates\n  - Using pictures (adding candidates which have most similar pictures and have the same \"product_type_no\", \"index_group_no\")\n  - Using detail description (as same as using pictures)\n\nI think what these method did not work is due to my mistake in implementation. When adding these candidates, it looks leaked. I have realized I have to make much more efforts to utilize ideas.\n\n* Features\n  - FN, active, postal_code (even count frequency)\n  - price feat, online or offline purchased\n  - on sale label: it would be because most articles are on sale for a few days suddenly, not continously\n\n* Model\n  - LGBMClassifier (predict purchase or not): In fact, it worked well when LB is under 0.03. Score has jumped up from 0.0274 to 0.0287. Mysterious...\n\n* Postprocess\n  - Replace top sales to \"less confident\" prediction: I did not have enough time to try, so there would be better way...",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1784466": "First of all, thanks to Kaggle team, H&M staff and all competitors.\nIt was so interesting and greatly well-designed competition (with a little shake) that I was excited to join.\nIt is grateful that our team ([@ruik](https://www.kaggle.com/ruiknemui) and me) got the highest rank in private leaderboard, but we also feel \"the wall\" to Gold medal.\nWe'd like to keep efforts until we become Master.\nAnyway, let me share our solution. I'm so happy if you give comments or advises for our approach.\n\n## Pipeline\nNot so different from others, very simple.\n1. Generate base candidates for each customer from their historical transactions in last N(=96) weeks. Transactions of articles which don't appear in historical transactions are not included in this time, so this approach is based on **\"What will they repurchase?\"**\n2. Add candidates\n3. Create features\n4. Fit/predict by LGBMRanker\n5. Postprocess\n\n## How to add candidates\n* Calculate article embeddings using word2vec as item2vec, then select N(=48) candidates whose embeddings have high similarity with predicted vector as \"next purchase\".\n* Add candidates have **product codes** as same as those of \"pair\" with each customer's last N(=3) transactions. (thanks to [Mr.Giba's notebook](https://www.kaggle.com/code/titericz/article-id-pairs-in-3s-using-cudf))\nPlease note that these method are applied to customers **who have enough historical transactions**.\n\n## Create features\nNot so different from winners.\n* **article features**\n  - # of sales in last N week (for not only article itself but also attributes which article is belonged to, like \"department_name\") + that segmented by age\n  - max/min/median price purchased (smoothing with global median)\n  - online purchased rate\n  - # of elapsed days from last purchase among all customers\n  - # of elapsed days from article's first sell\n* **customer features**\n  - age: binned and adjusted (10 if age < 20, 60 if age >= 60)\n  - \"club_member_status\",\"fashion_news_frequency\"\n  - attribute: estimated from index_code of articles they purchased (ex. if a customer usually purchases \"Ladieswear\", she is estimated as \"Lady\")\n  - max/min/median price purchased (smoothing with global median)\n  - # of elapsed days from customer's last purchase\n* **customer x article features**\n  - # of elapsed days from last purchase for each article/customer combination\n  - # of each article/customer pair appeared in historical transactions\n  - purchased rate(transactions) of article for each age/attribute\n  - ratio of article's price(max/min/median) to customer's price(max/min/median) (ex. if ratio of article's price min to customer's price max is over 1, we can estimate **\"he has no budget to purchase the item.\"**)\n\n## Fit/predict by LGBMRanker\nWe divide customers to 3 folds, 30/35/35% (actually, at first aiming to conduct experiments quickly).\nNext, we train and predict LGBMRanker for each fold of customer. Then apply weighted averaging (0.3/0.35/0.35) to prediction.\nTrain/valid period is only 1 set, valid set is last 1 week and train set is last 3 weeks before valid.\nWe retrain model with last 3 weeks data when predicting test after parameter tuning with optuna.\n\n## Postprocess\nWe apply postprocess to prediction of **customers who have no or less than 12 candidates (transactions) in last N(=96) weeks.**\n* We predict top 12 sales in last 1 week among customers with no candidates in last N(=96) weeks to those customers (+0.0002)\n* We predict top 12-K sales in last 1 week for each age group to customers with K(<12) candidates (+0.0001)\n\nThese method has improved score about 0.0003. (slight, but important improvement)\n\n## What did not work\n\n* Candidates\n  - Using pictures (adding candidates which have most similar pictures and have the same \"product_type_no\", \"index_group_no\")\n  - Using detail description (as same as using pictures)\n\nI think what these method did not work is due to my mistake in implementation. When adding these candidates, it looks leaked. I have realized I have to make much more efforts to utilize ideas.\n\n* Features\n  - FN, active, postal_code (even count frequency)\n  - price feat, online or offline purchased\n  - on sale label: it would be because most articles are on sale for a few days suddenly, not continously\n\n* Model\n  - LGBMClassifier (predict purchase or not): In fact, it worked well when LB is under 0.03. Score has jumped up from 0.0274 to 0.0287. Mysterious...\n\n* Postprocess\n  - Replace top sales to \"less confident\" prediction: I did not have enough time to try, so there would be better way..."
  }
}