{
  "id": 324411,
  "title": "30th Solution Intro and Personal Summary",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/yannick-30th-solution-intro-and-personal-summary",
  "author_name": "",
  "post_date": "2022-05-11T16:15:09.997209600Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, thanks for the efforts of HM and Kaggle platform, who give us this wonderful experience~</p>\n<p>Let me introduce my solution, and try to point out the shortages compared with the solutions with better performance.</p>\n<ul>\n<li><p>0.Solution Framework<br>\nI used two stage strategy, recall and rank, just like most of other competitors.</p></li>\n<li><p>1. Local valid dataset generation<br>\nI only used the last week to create my local valid dataset<br>\nThe number of unique customer_id in last week only about 60k, but the size of total customer is 1300k in the full dataset, apparently, it will limited generalization of ranking step.</p></li>\n<li><p>2. Recall stage<br>\nI only rely on a single itemCF algorithm to do the recall work. I also used the popularity of items as an extra weight, which could boost mAP score helpfully.<br>\nOther competitors usually leverage multi-way recall strategy. Besides itemCF, u2i, methods like graph model, SVD, history item are also included in their recall stage. <strong>I guess that is the main reason why my final rank is only 30th.</strong><br>\nActually, I tried deepMatch, but it not worked well. I guess the default “last item train strategy” is not suitable for HM dataset. I believe the reason is a lot of customers had long gap between the last purchasing and their prior purchasing, which made it difficult to train the item embeddings effectively. If modify the “last item train strategy”, it could be more useful, but I do not have extra time to do so.<br>\nMy final recall stage score around 0.028 in LB</p></li>\n<li><p>3.Rank stage<br>\nI use 5 folds lightGBM with lambda rank in this part. <br>\nI tried NN models like DIN and deepFM, but they did not beat LGB.<br>\nMy feature engineering focused on the interactions between customers and items. It included item frequency, gap of purchasing time and category feature statistic (type , color, department, etc ).<br>\nIt is worth noting that, I also used vector which is produced from image files in rank part to evaluate the similarity between history and candidate items, but it doesn’t bring improvement largely.</p></li>\n</ul>\n<p><strong>My Rank stage got 0.031 in LB finally, it only brings extra 0.003 mAP score to recall part.\nI believe the local dataset generation strategy limited the power of ranking models.</strong><br>\nLeaders of LB usually used 6 weeks or more data as trainset. Apparently, more data could bring extra benefits to enhance generalization of rank part .</p>\n<ul>\n<li>4.Team<br>\nI have to say, a good team is valuable. It is quite fatiguing for finish this competition by myself.<br>\nI highly recommend everyone should try to paly Kaggle with others (If you don’t mind solo gold, haha). It not only may bring higher rank, the encouragement from team members is necessary at any time.</li>\n</ul>",
  "messages": [
    {
      "id": "1784955",
      "postDate": "05/11/2022 16:15:09",
      "content": "<p>First of all, thanks for the efforts of HM and Kaggle platform, who give us this wonderful experience~</p>\n<p>Let me introduce my solution, and try to point out the shortages compared with the solutions with better performance.</p>\n<ul>\n<li><p>0.Solution Framework<br>\nI used two stage strategy, recall and rank, just like most of other competitors.</p></li>\n<li><p>1. Local valid dataset generation<br>\nI only used the last week to create my local valid dataset<br>\nThe number of unique customer_id in last week only about 60k, but the size of total customer is 1300k in the full dataset, apparently, it will limited generalization of ranking step.</p></li>\n<li><p>2. Recall stage<br>\nI only rely on a single itemCF algorithm to do the recall work. I also used the popularity of items as an extra weight, which could boost mAP score helpfully.<br>\nOther competitors usually leverage multi-way recall strategy. Besides itemCF, u2i, methods like graph model, SVD, history item are also included in their recall stage. <strong>I guess that is the main reason why my final rank is only 30th.</strong><br>\nActually, I tried deepMatch, but it not worked well. I guess the default “last item train strategy” is not suitable for HM dataset. I believe the reason is a lot of customers had long gap between the last purchasing and their prior purchasing, which made it difficult to train the item embeddings effectively. If modify the “last item train strategy”, it could be more useful, but I do not have extra time to do so.<br>\nMy final recall stage score around 0.028 in LB</p></li>\n<li><p>3.Rank stage<br>\nI use 5 folds lightGBM with lambda rank in this part. <br>\nI tried NN models like DIN and deepFM, but they did not beat LGB.<br>\nMy feature engineering focused on the interactions between customers and items. It included item frequency, gap of purchasing time and category feature statistic (type , color, department, etc ).<br>\nIt is worth noting that, I also used vector which is produced from image files in rank part to evaluate the similarity between history and candidate items, but it doesn’t bring improvement largely.</p></li>\n</ul>\n<p><strong>My Rank stage got 0.031 in LB finally, it only brings extra 0.003 mAP score to recall part.\nI believe the local dataset generation strategy limited the power of ranking models.</strong><br>\nLeaders of LB usually used 6 weeks or more data as trainset. Apparently, more data could bring extra benefits to enhance generalization of rank part .</p>\n<ul>\n<li>4.Team<br>\nI have to say, a good team is valuable. It is quite fatiguing for finish this competition by myself.<br>\nI highly recommend everyone should try to paly Kaggle with others (If you don’t mind solo gold, haha). It not only may bring higher rank, the encouragement from team members is necessary at any time.</li>\n</ul>",
      "rawMarkdown": "First of all, thanks for the efforts of HM and Kaggle platform, who give us this wonderful experience~\n\nLet me introduce my solution, and try to point out the shortages compared with the solutions with better performance.\n\n- 0.Solution Framework\nI used two stage strategy, recall and rank, just like most of other competitors.\n\n- 1. Local valid dataset generation\nI only used the last week to create my local valid dataset\nThe number of unique customer_id in last week only about 60k, but the size of total customer is 1300k in the full dataset, apparently, it will limited generalization of ranking step.\n\n- 2. Recall stage\nI only rely on a single itemCF algorithm to do the recall work. I also used the popularity of items as an extra weight, which could boost mAP score helpfully.\nOther competitors usually leverage multi-way recall strategy. Besides itemCF, u2i, methods like graph model, SVD, history item are also included in their recall stage. **I guess that is the main reason why my final rank is only 30th.**\nActually, I tried deepMatch, but it not worked well. I guess the default “last item train strategy” is not suitable for HM dataset. I believe the reason is a lot of customers had long gap between the last purchasing and their prior purchasing, which made it difficult to train the item embeddings effectively. If modify the “last item train strategy”, it could be more useful, but I do not have extra time to do so.\nMy final recall stage score around 0.028 in LB\n\n- 3.Rank stage\nI use 5 folds lightGBM with lambda rank in this part. \nI tried NN models like DIN and deepFM, but they did not beat LGB.\nMy feature engineering focused on the interactions between customers and items. It included item frequency, gap of purchasing time and category feature statistic (type , color, department, etc ).\nIt is worth noting that, I also used vector which is produced from image files in rank part to evaluate the similarity between history and candidate items, but it doesn’t bring improvement largely.\n\n**My Rank stage got 0.031 in LB finally, it only brings extra 0.003 mAP score to recall part.\nI believe the local dataset generation strategy limited the power of ranking models.**\nLeaders of LB usually used 6 weeks or more data as trainset. Apparently, more data could bring extra benefits to enhance generalization of rank part .\n\n\n- 4.Team\nI have to say, a good team is valuable. It is quite fatiguing for finish this competition by myself.\nI highly recommend everyone should try to paly Kaggle with others (If you don’t mind solo gold, haha). It not only may bring higher rank, the encouragement from team members is necessary at any time.",
      "votes": null
    },
    {
      "id": "1785410",
      "postDate": "05/12/2022 04:39:15",
      "content": "<p><a href=\"https://www.kaggle.com/breezeyuner\" target=\"_blank\">@breezeyuner</a> thanks for the good advice and for posting your detailed approach</p>",
      "rawMarkdown": "breezeyuner thanks for the good advice and for posting your detailed approach",
      "votes": null
    },
    {
      "id": "1786017",
      "postDate": "05/12/2022 15:01:05",
      "content": "<p>Thank you for sharing your experience and solution summary. It is helpful to read about the strategies others used in this competition. It sounds like you did a lot of work and it is impressive that you placed 30th overall. I agree that a good team is valuable and it sounds like it was a tough competition. Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing your experience and solution summary. It is helpful to read about the strategies others used in this competition. It sounds like you did a lot of work and it is impressive that you placed 30th overall. I agree that a good team is valuable and it sounds like it was a tough competition. Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1792088",
      "postDate": "05/16/2022 16:51:35",
      "content": "<p>Thanks for sharing!</p>\n<p>0.028 is an amazing score without any ranking algorithm!<br>\nCould you share more details on the itemCF algorithm implementation, and how you combined it with popularity weights?</p>",
      "rawMarkdown": "Thanks for sharing!\n\n0.028 is an amazing score without any ranking algorithm!\nCould you share more details on the itemCF algorithm implementation, and how you combined it with popularity weights?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1785410,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/12/2022 04:39:15",
      "content": "<p><a href=\"https://www.kaggle.com/breezeyuner\" target=\"_blank\">@breezeyuner</a> thanks for the good advice and for posting your detailed approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1786017,
      "author_name": "",
      "author_url": "",
      "post_date": "05/12/2022 15:01:05",
      "content": "<p>Thank you for sharing your experience and solution summary. It is helpful to read about the strategies others used in this competition. It sounds like you did a lot of work and it is impressive that you placed 30th overall. I agree that a good team is valuable and it sounds like it was a tough competition. Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1792088,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "05/16/2022 16:51:35",
      "content": "<p>Thanks for sharing!</p>\n<p>0.028 is an amazing score without any ranking algorithm!<br>\nCould you share more details on the itemCF algorithm implementation, and how you combined it with popularity weights?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1784955": "First of all, thanks for the efforts of HM and Kaggle platform, who give us this wonderful experience~\n\nLet me introduce my solution, and try to point out the shortages compared with the solutions with better performance.\n\n- 0.Solution Framework\nI used two stage strategy, recall and rank, just like most of other competitors.\n\n- 1. Local valid dataset generation\nI only used the last week to create my local valid dataset\nThe number of unique customer_id in last week only about 60k, but the size of total customer is 1300k in the full dataset, apparently, it will limited generalization of ranking step.\n\n- 2. Recall stage\nI only rely on a single itemCF algorithm to do the recall work. I also used the popularity of items as an extra weight, which could boost mAP score helpfully.\nOther competitors usually leverage multi-way recall strategy. Besides itemCF, u2i, methods like graph model, SVD, history item are also included in their recall stage. **I guess that is the main reason why my final rank is only 30th.**\nActually, I tried deepMatch, but it not worked well. I guess the default “last item train strategy” is not suitable for HM dataset. I believe the reason is a lot of customers had long gap between the last purchasing and their prior purchasing, which made it difficult to train the item embeddings effectively. If modify the “last item train strategy”, it could be more useful, but I do not have extra time to do so.\nMy final recall stage score around 0.028 in LB\n\n- 3.Rank stage\nI use 5 folds lightGBM with lambda rank in this part. \nI tried NN models like DIN and deepFM, but they did not beat LGB.\nMy feature engineering focused on the interactions between customers and items. It included item frequency, gap of purchasing time and category feature statistic (type , color, department, etc ).\nIt is worth noting that, I also used vector which is produced from image files in rank part to evaluate the similarity between history and candidate items, but it doesn’t bring improvement largely.\n\n**My Rank stage got 0.031 in LB finally, it only brings extra 0.003 mAP score to recall part.\nI believe the local dataset generation strategy limited the power of ranking models.**\nLeaders of LB usually used 6 weeks or more data as trainset. Apparently, more data could bring extra benefits to enhance generalization of rank part .\n\n\n- 4.Team\nI have to say, a good team is valuable. It is quite fatiguing for finish this competition by myself.\nI highly recommend everyone should try to paly Kaggle with others (If you don’t mind solo gold, haha). It not only may bring higher rank, the encouragement from team members is necessary at any time.",
    "1785410": "breezeyuner thanks for the good advice and for posting your detailed approach",
    "1786017": "Thank you for sharing your experience and solution summary. It is helpful to read about the strategies others used in this competition. It sounds like you did a lot of work and it is impressive that you placed 30th overall. I agree that a good team is valuable and it sounds like it was a tough competition. Thank you for sharing!",
    "1792088": "Thanks for sharing!\n\n0.028 is an amazing score without any ranking algorithm!\nCould you share more details on the itemCF algorithm implementation, and how you combined it with popularity weights?"
  },
  "source": "meta"
}