{
  "id": 324595,
  "title": "17th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/apolo-17th-place-solution",
  "author_name": "",
  "post_date": "2022-05-12T12:15:04.457Z",
  "votes": 37,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First of all, thank you to the competition organizers for a great competition.</p>\n<p>I am very disappointed to have missed the gold medal, but I gained a lot of experience through this competition!</p>\n<p>I will share my solution, including the thought process that led me to try that solution and the methods that did not work.</p>\n<h1>Overview</h1>\n<p>Since GBDT model was used in the top solutions of the past recommendation competitions on Kaggle (like <a href=\"https://www.kaggle.com/competitions/instacart-market-basket-analysis\" target=\"_blank\">this</a>), I tried GBDT model early in the competition and got good scores.</p>\n<p>I did not use the ranking models such as LGBMRanker and CatboostRanker (loss: YetiRank) because solving this task as a classification problem provided better results. Finally, I used Catboost, which has a high prediction speed.</p>\n<h1>Data Handling</h1>\n<p>I used only the last week of transaction data (2020-09-16 to 2020-09-22) for validation.</p>\n<p>I created features using non-validation data (~2020-09-15) and created 0/1 flag based on whether or not the item was purchased during the validation period (2020-09-16 to 2020-09-22).</p>\n<p><strong>CV Strategy</strong><br>\n5fold / GroupKfold(customer_id)</p>\n<h1>Candidate Generation</h1>\n<p>I generated a total of 27,341,620 candidates for 68,984 customers.</p>\n<ul>\n<li>top 300 popular items*</li>\n<li>previously purchased items</li>\n<li>items with the same product code as previously purchased items</li>\n</ul>\n<p>(*) I used the top 300 popular items from 2020-09-16 to 2020-09-22 instead of from 2020-09-09 to 2020-09-15. As a result, the difference between LB and CV was relatively small compared to other competitors. (CV: 0.0385, Public LB: 0.0326, Private LB: 0.0330)</p>\n<p>For candidate generation, I tried the following other methods, but they did not work.</p>\n<ul>\n<li><br>\ntop 300 popular items by age or gender</li>\n<li><br>\nitems with the same <code>department_no</code> , <code>section_no</code> and <code>product_type_no</code> as previously purchased items<br>\n(This idea comes from the fact that 13.8% of customers purchased items with the same <code>department_no</code>, <code>section_no</code> and <code>product_type_no</code> within 3 weeks.)<br>\n(Reference: <a href=\"https://www.kaggle.com/code/lichtlab/do-customers-buy-the-same-products-again\" target=\"_blank\">this notebook</a>)</li>\n</ul>\n<h1>Feature Engineering</h1>\n<h5>Customer features</h5>\n<ul>\n<li>customer attributes (<code>age</code>, <code>gender</code> etc)</li>\n<li>number of repeats of the same item in the past by each customer</li>\n<li>average purchase span / difference between average purchase span and last purchase date</li>\n<li>mean price / max price</li>\n<li>mean channel_id by each customer</li>\n</ul>\n<h5>Article features</h5>\n<ul>\n<li>article attributes (<code>product_code</code>, <code>section_no</code> etc)</li>\n<li>number of sales per week (week0~week5)</li>\n<li>number of sales per day (day0~day6)</li>\n<li>mean channel_id by each article</li>\n</ul>\n<h5>Customer x Article features</h5>\n<ul>\n<li>last purchase date with the same attributes as the candidate item</li>\n<li>difference between mean channel_id by each customer and by each article</li>\n<li>purchase rate of the candidate item by customers of the same age</li>\n<li>percentage of items with the same attributes as the candidate item among the customer's past purchases</li>\n<li>CF features (reference: <a href=\"https://www.kaggle.com/code/poteman/hm-item-cf\" target=\"_blank\">this notebook</a>)</li>\n</ul>\n<p>By adding this CF features, the score improved to CV: 0.0385, Public LB: 0.0326.<br>\nHowever, when I created too many CF features with slightly different conditions, for some reason the CV went up but the LB went down. (CV: 0.0397, Public LB: 0.0319)</p>\n<h1>Model</h1>\n<p>I trained all candidates in one Catboost model.<br>\nI also tried training a different model for each reason for candidate generation, but the performance did not improve.</p>\n<h5>Hyperparameter Tuning</h5>\n<p>Since this data was imbalanced (positive samples: 27272479, negative samples: 69141), I set <code>scale_pos_weight</code> parameter.</p>\n<p>If I follow the <a href=\"https://catboost.ai/en/docs/references/training-parameters/common#scale_pos_weight\" target=\"_blank\">official documentation</a>, it would be 394(=27272479/69141), but to avoid overfit, I set it to 100. Setting <code>scale_pos_weight</code> contributed significantly to the result.</p>\n<p>Since the number of candidates and the number of positive samples differ from customer to customer, I tried to set the weight accordingly, but it did not work.</p>\n<h1>Ensemble</h1>\n<p>I ensembled several models using the code in <a href=\"https://www.kaggle.com/code/titericz/h-m-ensembling-how-to\" target=\"_blank\">this notebook</a>. Weight was all set to 1 to avoid overfit to the LB.</p>\n<h1>Post Process</h1>\n<p>Of the 1,371,980 customers in sample_submission.csv, 9699 customers have no data in transactions_train.csv. </p>\n<p>For the following reasons, I assumed that all of these 9699 customers would have made a purchase during the prediction week.</p>\n<ul>\n<li>5572 customers purchased items for the first time in week0.</li>\n<li>the data shows that a discount sale has been held on the last Saturday in September for the past two years, and the number of customers was higher compared to other weeks.</li>\n</ul>\n<p>Although customer attributes and number of recent sales were included in the features of Catboost model, the prediction results were not convincing. So I overwrote the predicted results for new customers with the Top 12 most popular items.</p>\n<ul>\n<li><br>\nLooking at last year's data, the percentage of online purchases increased significantly on the last Saturday in September due to a discount sale.<br>\nTherefore, I tried post-processing with the hypothesis that if the similar discount sale was held on the last Saturday of September in 2020, products with an originally high ratio of online purchases would be sold well, but it did not work.</li>\n</ul>\n<h1>Environment</h1>\n<p>All code was run on Colab Pro+.</p>\n<hr>\n<p>I am still learning machine learning, so if you have any suggestions or advice, I would be glad to hear them :)</p>",
  "messages": [
    {
      "id": "1785763",
      "postDate": "05/12/2022 11:48:35",
      "content": "<p>First of all, thank you to the competition organizers for a great competition.</p>\n<p>I am very disappointed to have missed the gold medal, but I gained a lot of experience through this competition!</p>\n<p>I will share my solution, including the thought process that led me to try that solution and the methods that did not work.</p>\n<h1>Overview</h1>\n<p>Since GBDT model was used in the top solutions of the past recommendation competitions on Kaggle (like <a href=\"https://www.kaggle.com/competitions/instacart-market-basket-analysis\" target=\"_blank\">this</a>), I tried GBDT model early in the competition and got good scores.</p>\n<p>I did not use the ranking models such as LGBMRanker and CatboostRanker (loss: YetiRank) because solving this task as a classification problem provided better results. Finally, I used Catboost, which has a high prediction speed.</p>\n<h1>Data Handling</h1>\n<p>I used only the last week of transaction data (2020-09-16 to 2020-09-22) for validation.</p>\n<p>I created features using non-validation data (~2020-09-15) and created 0/1 flag based on whether or not the item was purchased during the validation period (2020-09-16 to 2020-09-22).</p>\n<p><strong>CV Strategy</strong><br>\n5fold / GroupKfold(customer_id)</p>\n<h1>Candidate Generation</h1>\n<p>I generated a total of 27,341,620 candidates for 68,984 customers.</p>\n<ul>\n<li>top 300 popular items*</li>\n<li>previously purchased items</li>\n<li>items with the same product code as previously purchased items</li>\n</ul>\n<p>(*) I used the top 300 popular items from 2020-09-16 to 2020-09-22 instead of from 2020-09-09 to 2020-09-15. As a result, the difference between LB and CV was relatively small compared to other competitors. (CV: 0.0385, Public LB: 0.0326, Private LB: 0.0330)</p>\n<p>For candidate generation, I tried the following other methods, but they did not work.</p>\n<ul>\n<li><br>\ntop 300 popular items by age or gender</li>\n<li><br>\nitems with the same <code>department_no</code> , <code>section_no</code> and <code>product_type_no</code> as previously purchased items<br>\n(This idea comes from the fact that 13.8% of customers purchased items with the same <code>department_no</code>, <code>section_no</code> and <code>product_type_no</code> within 3 weeks.)<br>\n(Reference: <a href=\"https://www.kaggle.com/code/lichtlab/do-customers-buy-the-same-products-again\" target=\"_blank\">this notebook</a>)</li>\n</ul>\n<h1>Feature Engineering</h1>\n<h5>Customer features</h5>\n<ul>\n<li>customer attributes (<code>age</code>, <code>gender</code> etc)</li>\n<li>number of repeats of the same item in the past by each customer</li>\n<li>average purchase span / difference between average purchase span and last purchase date</li>\n<li>mean price / max price</li>\n<li>mean channel_id by each customer</li>\n</ul>\n<h5>Article features</h5>\n<ul>\n<li>article attributes (<code>product_code</code>, <code>section_no</code> etc)</li>\n<li>number of sales per week (week0~week5)</li>\n<li>number of sales per day (day0~day6)</li>\n<li>mean channel_id by each article</li>\n</ul>\n<h5>Customer x Article features</h5>\n<ul>\n<li>last purchase date with the same attributes as the candidate item</li>\n<li>difference between mean channel_id by each customer and by each article</li>\n<li>purchase rate of the candidate item by customers of the same age</li>\n<li>percentage of items with the same attributes as the candidate item among the customer's past purchases</li>\n<li>CF features (reference: <a href=\"https://www.kaggle.com/code/poteman/hm-item-cf\" target=\"_blank\">this notebook</a>)</li>\n</ul>\n<p>By adding this CF features, the score improved to CV: 0.0385, Public LB: 0.0326.<br>\nHowever, when I created too many CF features with slightly different conditions, for some reason the CV went up but the LB went down. (CV: 0.0397, Public LB: 0.0319)</p>\n<h1>Model</h1>\n<p>I trained all candidates in one Catboost model.<br>\nI also tried training a different model for each reason for candidate generation, but the performance did not improve.</p>\n<h5>Hyperparameter Tuning</h5>\n<p>Since this data was imbalanced (positive samples: 27272479, negative samples: 69141), I set <code>scale_pos_weight</code> parameter.</p>\n<p>If I follow the <a href=\"https://catboost.ai/en/docs/references/training-parameters/common#scale_pos_weight\" target=\"_blank\">official documentation</a>, it would be 394(=27272479/69141), but to avoid overfit, I set it to 100. Setting <code>scale_pos_weight</code> contributed significantly to the result.</p>\n<p>Since the number of candidates and the number of positive samples differ from customer to customer, I tried to set the weight accordingly, but it did not work.</p>\n<h1>Ensemble</h1>\n<p>I ensembled several models using the code in <a href=\"https://www.kaggle.com/code/titericz/h-m-ensembling-how-to\" target=\"_blank\">this notebook</a>. Weight was all set to 1 to avoid overfit to the LB.</p>\n<h1>Post Process</h1>\n<p>Of the 1,371,980 customers in sample_submission.csv, 9699 customers have no data in transactions_train.csv. </p>\n<p>For the following reasons, I assumed that all of these 9699 customers would have made a purchase during the prediction week.</p>\n<ul>\n<li>5572 customers purchased items for the first time in week0.</li>\n<li>the data shows that a discount sale has been held on the last Saturday in September for the past two years, and the number of customers was higher compared to other weeks.</li>\n</ul>\n<p>Although customer attributes and number of recent sales were included in the features of Catboost model, the prediction results were not convincing. So I overwrote the predicted results for new customers with the Top 12 most popular items.</p>\n<ul>\n<li><br>\nLooking at last year's data, the percentage of online purchases increased significantly on the last Saturday in September due to a discount sale.<br>\nTherefore, I tried post-processing with the hypothesis that if the similar discount sale was held on the last Saturday of September in 2020, products with an originally high ratio of online purchases would be sold well, but it did not work.</li>\n</ul>\n<h1>Environment</h1>\n<p>All code was run on Colab Pro+.</p>\n<hr>\n<p>I am still learning machine learning, so if you have any suggestions or advice, I would be glad to hear them :)</p>",
      "rawMarkdown": "First of all, thank you to the competition organizers for a great competition.\n\nI am very disappointed to have missed the gold medal, but I gained a lot of experience through this competition!\n\nI will share my solution, including the thought process that led me to try that solution and the methods that did not work.\n\n# Overview\n\nSince GBDT model was used in the top solutions of the past recommendation competitions on Kaggle (like [this](https://www.kaggle.com/competitions/instacart-market-basket-analysis)), I tried GBDT model early in the competition and got good scores.\n\nI did not use the ranking models such as LGBMRanker and CatboostRanker (loss: YetiRank) because solving this task as a classification problem provided better results. Finally, I used Catboost, which has a high prediction speed.\n\n# Data Handling\n\nI used only the last week of transaction data (2020-09-16 to 2020-09-22) for validation.\n\nI created features using non-validation data (~2020-09-15) and created 0/1 flag based on whether or not the item was purchased during the validation period (2020-09-16 to 2020-09-22).\n\n**CV Strategy**\n5fold / GroupKfold(customer_id)\n\n# Candidate Generation\n\nI generated a total of 27,341,620 candidates for 68,984 customers.\n\n- top 300 popular items*\n- previously purchased items\n- items with the same product code as previously purchased items\n\n(*) I used the top 300 popular items from 2020-09-16 to 2020-09-22 instead of from 2020-09-09 to 2020-09-15. As a result, the difference between LB and CV was relatively small compared to other competitors. (CV: 0.0385, Public LB: 0.0326, Private LB: 0.0330)\n\nFor candidate generation, I tried the following other methods, but they did not work.\n\n- <u>Unused idea1</u>\ntop 300 popular items by age or gender\n- <u>Unused idea2</u>\nitems with the same `department_no` , `section_no` and `product_type_no` as previously purchased items\n(This idea comes from the fact that 13.8% of customers purchased items with the same `department_no`, `section_no` and `product_type_no` within 3 weeks.)\n(Reference: [this notebook](https://www.kaggle.com/code/lichtlab/do-customers-buy-the-same-products-again))\n\n# Feature Engineering\n\n##### Customer features\n\n- customer attributes (`age`, `gender` etc)\n- number of repeats of the same item in the past by each customer\n- average purchase span / difference between average purchase span and last purchase date\n- mean price / max price\n- mean channel_id by each customer\n\n##### Article features\n\n- article attributes (`product_code`, `section_no` etc)\n- number of sales per week (week0~week5)\n- number of sales per day (day0~day6)\n- mean channel_id by each article\n\n##### Customer x Article features\n\n- last purchase date with the same attributes as the candidate item\n- difference between mean channel_id by each customer and by each article\n- purchase rate of the candidate item by customers of the same age\n- percentage of items with the same attributes as the candidate item among the customer's past purchases\n- CF features (reference: [this notebook](https://www.kaggle.com/code/poteman/hm-item-cf))\n\nBy adding this CF features, the score improved to CV: 0.0385, Public LB: 0.0326.\nHowever, when I created too many CF features with slightly different conditions, for some reason the CV went up but the LB went down. (CV: 0.0397, Public LB: 0.0319)\n\n# Model\n\nI trained all candidates in one Catboost model.\nI also tried training a different model for each reason for candidate generation, but the performance did not improve.\n\n##### Hyperparameter Tuning\n\nSince this data was imbalanced (positive samples: 27272479, negative samples: 69141), I set `scale_pos_weight` parameter.\n\nIf I follow the [official documentation](https://catboost.ai/en/docs/references/training-parameters/common#scale_pos_weight), it would be 394(=27272479/69141), but to avoid overfit, I set it to 100. Setting `scale_pos_weight` contributed significantly to the result.\n\nSince the number of candidates and the number of positive samples differ from customer to customer, I tried to set the weight accordingly, but it did not work.\n\n# Ensemble\n\nI ensembled several models using the code in [this notebook](https://www.kaggle.com/code/titericz/h-m-ensembling-how-to). Weight was all set to 1 to avoid overfit to the LB.\n\n# Post Process\n\nOf the 1,371,980 customers in sample_submission.csv, 9699 customers have no data in transactions_train.csv. \n\nFor the following reasons, I assumed that all of these 9699 customers would have made a purchase during the prediction week.\n\n- 5572 customers purchased items for the first time in week0.\n- the data shows that a discount sale has been held on the last Saturday in September for the past two years, and the number of customers was higher compared to other weeks.\n\nAlthough customer attributes and number of recent sales were included in the features of Catboost model, the prediction results were not convincing. So I overwrote the predicted results for new customers with the Top 12 most popular items.\n\n- <u>Unused idea</u>\nLooking at last year's data, the percentage of online purchases increased significantly on the last Saturday in September due to a discount sale.\nTherefore, I tried post-processing with the hypothesis that if the similar discount sale was held on the last Saturday of September in 2020, products with an originally high ratio of online purchases would be sold well, but it did not work.\n\n# Environment\nAll code was run on Colab Pro+.\n\n--------------------------------------------------\n\nI am still learning machine learning, so if you have any suggestions or advice, I would be glad to hear them :)",
      "votes": null
    },
    {
      "id": "1787019",
      "postDate": "05/13/2022 14:14:46",
      "content": "<p>17th place is a great achievement! Well done on implementing a solution that achieved a high score.</p>\n<p>It sounds like you've gained a lot of valuable experience from this competition. I'm glad you're sharing your thoughts and solutions - it will be helpful for other Kagglers who are just starting out.</p>\n<p>Looking forward to seeing more of your posts in the future!</p>",
      "rawMarkdown": "17th place is a great achievement! Well done on implementing a solution that achieved a high score.\n\nIt sounds like you've gained a lot of valuable experience from this competition. I'm glad you're sharing your thoughts and solutions - it will be helpful for other Kagglers who are just starting out.\n\nLooking forward to seeing more of your posts in the future!",
      "votes": null
    },
    {
      "id": "1787204",
      "postDate": "05/13/2022 16:57:47",
      "content": "<p>Hi. Thanks for sharing and congratulations for silver (almost gold) medal!   <br>\nI'm sure that you will get a gold soon.  </p>\n<p>I have a question:</p>\n<blockquote>\n  <p>I generated a total of 27,341,620 candidates for 68,984 customers.</p>\n</blockquote>\n<p>This means that you generated candidates for all customers with purchased in week 105 (2020-09-16 ~ 2020-09-22)?<br>\nBecause in this week there are exactly 68984 unique customers.</p>",
      "rawMarkdown": "Hi. Thanks for sharing and congratulations for silver (almost gold) medal!   \nI'm sure that you will get a gold soon.  \n\nI have a question:\n> I generated a total of 27,341,620 candidates for 68,984 customers.\n\nThis means that you generated candidates for all customers with purchased in week 105 (2020-09-16 ~ 2020-09-22)?\nBecause in this week there are exactly 68984 unique customers.",
      "votes": null
    },
    {
      "id": "1787265",
      "postDate": "05/13/2022 18:04:42",
      "content": "<p>hi, anyway this is really good achievement. Can you share your code than we can learn more? It is really hard to learn for a new beginner if I only get your idea. Really appreciate!</p>",
      "rawMarkdown": "hi, anyway this is really good achievement. Can you share your code than we can learn more? It is really hard to learn for a new beginner if I only get your idea. Really appreciate!",
      "votes": null
    },
    {
      "id": "1792806",
      "postDate": "05/17/2022 10:15:41",
      "content": "<p>Exactly!<br>\nSince the evaluation in LB is only done for customers who purchased between 2020-09-23 ~ 2020-09-29, I only trained on customers who purchased in week 105 (2020-09-16 ~ 2020-09-22).</p>",
      "rawMarkdown": "Exactly!\nSince the evaluation in LB is only done for customers who purchased between 2020-09-23 ~ 2020-09-29, I only trained on customers who purchased in week 105 (2020-09-16 ~ 2020-09-22).",
      "votes": null
    },
    {
      "id": "1792810",
      "postDate": "05/17/2022 10:19:38",
      "content": "<p>Thank you so much.<br>\nI will actively post discussions and codes in the future!</p>",
      "rawMarkdown": "Thank you so much.\nI will actively post discussions and codes in the future!",
      "votes": null
    },
    {
      "id": "1792811",
      "postDate": "05/17/2022 10:26:12",
      "content": "<p>I'll upload the code soon.</p>",
      "rawMarkdown": "I'll upload the code soon.",
      "votes": null
    },
    {
      "id": "1792934",
      "postDate": "05/17/2022 12:30:50",
      "content": "<p>Hi, congratz</p>",
      "rawMarkdown": "Hi, congratz",
      "votes": null
    },
    {
      "id": "1792939",
      "postDate": "05/17/2022 12:38:46",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    },
    {
      "id": "1819961",
      "postDate": "06/14/2022 09:06:30",
      "content": "<p>at the end you used one single model to predict ? or rather every user_id had their own model ?</p>",
      "rawMarkdown": "at the end you used one single model to predict ? or rather every user_id had their own model ?",
      "votes": null
    },
    {
      "id": "2019721",
      "postDate": "11/06/2022 22:27:24",
      "content": "<p>Hello, really great work, thank you for share the code too! <br>\nI have a couple of questions: </p>\n<p>How to generate gender.pickle, cf_score_v3.pickle, cf_score_v2.pickle, cf_score_{feature}.pickle, cf_score_article_with_channel.pickle?<br>\nRegarding CATEGORICAL_COLS and DROP_COLS which values do I have to comment and which not? Thank you!</p>",
      "rawMarkdown": "Hello, really great work, thank you for share the code too! \nI have a couple of questions: \n\nHow to generate gender.pickle, cf_score_v3.pickle, cf_score_v2.pickle, cf_score_{feature}.pickle, cf_score_article_with_channel.pickle?\nRegarding CATEGORICAL_COLS and DROP_COLS which values do I have to comment and which not? Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1787019,
      "author_name": "",
      "author_url": "",
      "post_date": "05/13/2022 14:14:46",
      "content": "<p>17th place is a great achievement! Well done on implementing a solution that achieved a high score.</p>\n<p>It sounds like you've gained a lot of valuable experience from this competition. I'm glad you're sharing your thoughts and solutions - it will be helpful for other Kagglers who are just starting out.</p>\n<p>Looking forward to seeing more of your posts in the future!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1792810,
          "author_name": "shkanda",
          "author_url": "",
          "post_date": "05/17/2022 10:19:38",
          "content": "<p>Thank you so much.<br>\nI will actively post discussions and codes in the future!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1787204,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "05/13/2022 16:57:47",
      "content": "<p>Hi. Thanks for sharing and congratulations for silver (almost gold) medal!   <br>\nI'm sure that you will get a gold soon.  </p>\n<p>I have a question:</p>\n<blockquote>\n  <p>I generated a total of 27,341,620 candidates for 68,984 customers.</p>\n</blockquote>\n<p>This means that you generated candidates for all customers with purchased in week 105 (2020-09-16 ~ 2020-09-22)?<br>\nBecause in this week there are exactly 68984 unique customers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1792806,
          "author_name": "shkanda",
          "author_url": "",
          "post_date": "05/17/2022 10:15:41",
          "content": "<p>Exactly!<br>\nSince the evaluation in LB is only done for customers who purchased between 2020-09-23 ~ 2020-09-29, I only trained on customers who purchased in week 105 (2020-09-16 ~ 2020-09-22).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1787265,
      "author_name": "aliciaworld",
      "author_url": "",
      "post_date": "05/13/2022 18:04:42",
      "content": "<p>hi, anyway this is really good achievement. Can you share your code than we can learn more? It is really hard to learn for a new beginner if I only get your idea. Really appreciate!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1792811,
          "author_name": "shkanda",
          "author_url": "",
          "post_date": "05/17/2022 10:26:12",
          "content": "<p>I'll upload the code soon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1792934,
      "author_name": "ashoka98228",
      "author_url": "",
      "post_date": "05/17/2022 12:30:50",
      "content": "<p>Hi, congratz</p>",
      "votes": null,
      "replies": [
        {
          "id": 1792939,
          "author_name": "shkanda",
          "author_url": "",
          "post_date": "05/17/2022 12:38:46",
          "content": "<p>Thank you so much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1819961,
      "author_name": "huhulimaguli",
      "author_url": "",
      "post_date": "06/14/2022 09:06:30",
      "content": "<p>at the end you used one single model to predict ? or rather every user_id had their own model ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2019721,
      "author_name": "arcangelopisa",
      "author_url": "",
      "post_date": "11/06/2022 22:27:24",
      "content": "<p>Hello, really great work, thank you for share the code too! <br>\nI have a couple of questions: </p>\n<p>How to generate gender.pickle, cf_score_v3.pickle, cf_score_v2.pickle, cf_score_{feature}.pickle, cf_score_article_with_channel.pickle?<br>\nRegarding CATEGORICAL_COLS and DROP_COLS which values do I have to comment and which not? Thank you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1785763": "First of all, thank you to the competition organizers for a great competition.\n\nI am very disappointed to have missed the gold medal, but I gained a lot of experience through this competition!\n\nI will share my solution, including the thought process that led me to try that solution and the methods that did not work.\n\n# Overview\n\nSince GBDT model was used in the top solutions of the past recommendation competitions on Kaggle (like [this](https://www.kaggle.com/competitions/instacart-market-basket-analysis)), I tried GBDT model early in the competition and got good scores.\n\nI did not use the ranking models such as LGBMRanker and CatboostRanker (loss: YetiRank) because solving this task as a classification problem provided better results. Finally, I used Catboost, which has a high prediction speed.\n\n# Data Handling\n\nI used only the last week of transaction data (2020-09-16 to 2020-09-22) for validation.\n\nI created features using non-validation data (~2020-09-15) and created 0/1 flag based on whether or not the item was purchased during the validation period (2020-09-16 to 2020-09-22).\n\n**CV Strategy**\n5fold / GroupKfold(customer_id)\n\n# Candidate Generation\n\nI generated a total of 27,341,620 candidates for 68,984 customers.\n\n- top 300 popular items*\n- previously purchased items\n- items with the same product code as previously purchased items\n\n(*) I used the top 300 popular items from 2020-09-16 to 2020-09-22 instead of from 2020-09-09 to 2020-09-15. As a result, the difference between LB and CV was relatively small compared to other competitors. (CV: 0.0385, Public LB: 0.0326, Private LB: 0.0330)\n\nFor candidate generation, I tried the following other methods, but they did not work.\n\n- <u>Unused idea1</u>\ntop 300 popular items by age or gender\n- <u>Unused idea2</u>\nitems with the same `department_no` , `section_no` and `product_type_no` as previously purchased items\n(This idea comes from the fact that 13.8% of customers purchased items with the same `department_no`, `section_no` and `product_type_no` within 3 weeks.)\n(Reference: [this notebook](https://www.kaggle.com/code/lichtlab/do-customers-buy-the-same-products-again))\n\n# Feature Engineering\n\n##### Customer features\n\n- customer attributes (`age`, `gender` etc)\n- number of repeats of the same item in the past by each customer\n- average purchase span / difference between average purchase span and last purchase date\n- mean price / max price\n- mean channel_id by each customer\n\n##### Article features\n\n- article attributes (`product_code`, `section_no` etc)\n- number of sales per week (week0~week5)\n- number of sales per day (day0~day6)\n- mean channel_id by each article\n\n##### Customer x Article features\n\n- last purchase date with the same attributes as the candidate item\n- difference between mean channel_id by each customer and by each article\n- purchase rate of the candidate item by customers of the same age\n- percentage of items with the same attributes as the candidate item among the customer's past purchases\n- CF features (reference: [this notebook](https://www.kaggle.com/code/poteman/hm-item-cf))\n\nBy adding this CF features, the score improved to CV: 0.0385, Public LB: 0.0326.\nHowever, when I created too many CF features with slightly different conditions, for some reason the CV went up but the LB went down. (CV: 0.0397, Public LB: 0.0319)\n\n# Model\n\nI trained all candidates in one Catboost model.\nI also tried training a different model for each reason for candidate generation, but the performance did not improve.\n\n##### Hyperparameter Tuning\n\nSince this data was imbalanced (positive samples: 27272479, negative samples: 69141), I set `scale_pos_weight` parameter.\n\nIf I follow the [official documentation](https://catboost.ai/en/docs/references/training-parameters/common#scale_pos_weight), it would be 394(=27272479/69141), but to avoid overfit, I set it to 100. Setting `scale_pos_weight` contributed significantly to the result.\n\nSince the number of candidates and the number of positive samples differ from customer to customer, I tried to set the weight accordingly, but it did not work.\n\n# Ensemble\n\nI ensembled several models using the code in [this notebook](https://www.kaggle.com/code/titericz/h-m-ensembling-how-to). Weight was all set to 1 to avoid overfit to the LB.\n\n# Post Process\n\nOf the 1,371,980 customers in sample_submission.csv, 9699 customers have no data in transactions_train.csv. \n\nFor the following reasons, I assumed that all of these 9699 customers would have made a purchase during the prediction week.\n\n- 5572 customers purchased items for the first time in week0.\n- the data shows that a discount sale has been held on the last Saturday in September for the past two years, and the number of customers was higher compared to other weeks.\n\nAlthough customer attributes and number of recent sales were included in the features of Catboost model, the prediction results were not convincing. So I overwrote the predicted results for new customers with the Top 12 most popular items.\n\n- <u>Unused idea</u>\nLooking at last year's data, the percentage of online purchases increased significantly on the last Saturday in September due to a discount sale.\nTherefore, I tried post-processing with the hypothesis that if the similar discount sale was held on the last Saturday of September in 2020, products with an originally high ratio of online purchases would be sold well, but it did not work.\n\n# Environment\nAll code was run on Colab Pro+.\n\n--------------------------------------------------\n\nI am still learning machine learning, so if you have any suggestions or advice, I would be glad to hear them :)",
    "1787019": "17th place is a great achievement! Well done on implementing a solution that achieved a high score.\n\nIt sounds like you've gained a lot of valuable experience from this competition. I'm glad you're sharing your thoughts and solutions - it will be helpful for other Kagglers who are just starting out.\n\nLooking forward to seeing more of your posts in the future!",
    "1787204": "Hi. Thanks for sharing and congratulations for silver (almost gold) medal!   \nI'm sure that you will get a gold soon.  \n\nI have a question:\n> I generated a total of 27,341,620 candidates for 68,984 customers.\n\nThis means that you generated candidates for all customers with purchased in week 105 (2020-09-16 ~ 2020-09-22)?\nBecause in this week there are exactly 68984 unique customers.",
    "1787265": "hi, anyway this is really good achievement. Can you share your code than we can learn more? It is really hard to learn for a new beginner if I only get your idea. Really appreciate!",
    "1792806": "Exactly!\nSince the evaluation in LB is only done for customers who purchased between 2020-09-23 ~ 2020-09-29, I only trained on customers who purchased in week 105 (2020-09-16 ~ 2020-09-22).",
    "1792810": "Thank you so much.\nI will actively post discussions and codes in the future!",
    "1792811": "I'll upload the code soon.",
    "1792934": "Hi, congratz",
    "1792939": "Thank you so much!",
    "1819961": "at the end you used one single model to predict ? or rather every user_id had their own model ?",
    "2019721": "Hello, really great work, thank you for share the code too! \nI have a couple of questions: \n\nHow to generate gender.pickle, cf_score_v3.pickle, cf_score_v2.pickle, cf_score_{feature}.pickle, cf_score_article_with_channel.pickle?\nRegarding CATEGORICAL_COLS and DROP_COLS which values do I have to comment and which not? Thank you!"
  },
  "source": "meta"
}