{
  "id": 325407,
  "title": "H&M 22th solution and code (myaun part) ",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/325407",
  "author_name": "",
  "post_date": "2022-05-16T09:48:35.335645200Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is part of the 22nd solution (sorry for the split sharing by each member)</p>\n<ul>\n<li>members: <a href=\"https://www.kaggle.com/ganchan\" target=\"_blank\">@ganchan</a> <a href=\"https://www.kaggle.com/NARI\" target=\"_blank\">@NARI</a> <a href=\"https://www.kaggle.com/Moro\" target=\"_blank\">@Moro</a> <a href=\"https://www.kaggle.com/minguin\" target=\"_blank\">@minguin</a>, thanks~</li>\n<li>ganchan's solution: <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152</a></li>\n<li>moro's solution: <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224</a></li>\n</ul>\n<p>I simply share part of my solution (my single model and ensemble model) and the code.</p>\n<h1>Single model</h1>\n<ul>\n<li><p><a href=\"https://github.com/haradai1262/h-and-m-personalized-fashion-recommendations/blob/master/exp/cat_v4-3.py\" target=\"_blank\">Code</a></p></li>\n<li><p>Score</p>\n<ul>\n<li>CV: 0.0365</li>\n<li>Public LB: 0.03072</li>\n<li>Private LB: 0.03048</li></ul></li>\n<li><p>This model is forked from ganchan’s model</p>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152\" target=\"_blank\">Part of 22nd solution - single LGBM (Private:0.03038)</a></p></li>\n<li><p>diff</p>\n<ul>\n<li><p>dataset for training is ten weeks (ganchan’s model is four weeks)</p>\n<ul>\n<li><p>I use learning weights that get smaller the further away from the latest date</p></li>\n<li><p>each week, I select candidates from top sales of 15000 items in the previous four weeks.<br>\n<a href=\"https://ibb.co/zNybH8f\"><img src=\"https://i.ibb.co/V37DpHB/sampling.png\" alt=\"sampling\"></a></p></li>\n<li><p>(and I randomly selected 50 articles as the negative sample for each customer in the same manner as ganchan)</p></li></ul></li>\n<li><p>use Catboost Ranker (ganchan’s model is LGBM Classifier)</p>\n<ul>\n<li>time-series splits, train:~9/15, valid: 9/16~9/22</li></ul></li></ul></li></ul></li>\n</ul>\n<h1>Ensemble model</h1>\n<ul>\n<li><p><a href=\"https://github.com/haradai1262/h-and-m-personalized-fashion-recommendations/blob/master/exp/stacking_nb_6.ipynb\" target=\"_blank\">Code</a></p></li>\n<li><p>Score</p>\n<ul>\n<li>CV: 0.0407</li>\n<li>Public LB: 0.03296</li>\n<li>Private LB: 0.03294</li></ul></li>\n<li><p>Candidates</p>\n<ul>\n<li>extract top 24 predicted items from each model<ul>\n<li>we use 13 models, which are members’ best single models and their variants</li>\n<li>public LB of their models: 0.0307 ~ 0.0294</li></ul></li></ul></li>\n<li><p>Features</p>\n<ul>\n<li><p>Predict features (1st-stage model outputs)</p>\n<ul>\n<li><p>predicted value</p></li>\n<li><p>predicted value normilized per customer</p>\n<pre><code>preds[\"pred_norm\"] = preds.groupby('customer_id')[\"pred\"].transform(lambda x: (x - x.mean()) / x.std())\n</code></pre></li>\n<li><p>rank based on predicted value per customer</p>\n<pre><code>preds[\"pred_rank\"] = preds.groupby(\"customer_id\")[\"pred\"].rank(ascending=False, method=\"dense\")\n</code></pre></li></ul></li>\n<li><p>aggregate predict features (group by models)</p>\n<ul>\n<li>count, sum, max, min</li></ul></li>\n<li><p>article attributes</p></li>\n<li><p>customer attributes</p></li>\n<li><p>article buy count {yesterday, last week, last month, last year, time-decayed count}</p></li></ul></li>\n<li><p>Model</p>\n<ul>\n<li>CatboostRanker (Groupkfold, group=customer, k=10)</li></ul></li>\n<li><p>My note</p>\n<ul>\n<li>our ensemble effectiveness (only public LB, easy summary not strict)<ul>\n<li>0.0307: single</li>\n<li>0.0318 add <a href=\"https://www.kaggle.com/ganchan\" target=\"_blank\">@ganchan</a> <a href=\"https://www.kaggle.com/NARI\" target=\"_blank\">@NARI</a></li>\n<li>0.0322 add <a href=\"https://www.kaggle.com/minguin\" target=\"_blank\">@minguin</a></li>\n<li>0.0329 add <a href=\"https://www.kaggle.com/Moro\" target=\"_blank\">@Moro</a></li>\n<li>great effect on the ensemble! but more effective fusion methods might have suited.</li></ul></li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1791744",
      "postDate": "05/16/2022 09:48:35",
      "content": "<p>This is part of the 22nd solution (sorry for the split sharing by each member)</p>\n<ul>\n<li>members: <a href=\"https://www.kaggle.com/ganchan\" target=\"_blank\">@ganchan</a> <a href=\"https://www.kaggle.com/NARI\" target=\"_blank\">@NARI</a> <a href=\"https://www.kaggle.com/Moro\" target=\"_blank\">@Moro</a> <a href=\"https://www.kaggle.com/minguin\" target=\"_blank\">@minguin</a>, thanks~</li>\n<li>ganchan's solution: <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152</a></li>\n<li>moro's solution: <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224</a></li>\n</ul>\n<p>I simply share part of my solution (my single model and ensemble model) and the code.</p>\n<h1>Single model</h1>\n<ul>\n<li><p><a href=\"https://github.com/haradai1262/h-and-m-personalized-fashion-recommendations/blob/master/exp/cat_v4-3.py\" target=\"_blank\">Code</a></p></li>\n<li><p>Score</p>\n<ul>\n<li>CV: 0.0365</li>\n<li>Public LB: 0.03072</li>\n<li>Private LB: 0.03048</li></ul></li>\n<li><p>This model is forked from ganchan’s model</p>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152\" target=\"_blank\">Part of 22nd solution - single LGBM (Private:0.03038)</a></p></li>\n<li><p>diff</p>\n<ul>\n<li><p>dataset for training is ten weeks (ganchan’s model is four weeks)</p>\n<ul>\n<li><p>I use learning weights that get smaller the further away from the latest date</p></li>\n<li><p>each week, I select candidates from top sales of 15000 items in the previous four weeks.<br>\n<a href=\"https://ibb.co/zNybH8f\"><img src=\"https://i.ibb.co/V37DpHB/sampling.png\" alt=\"sampling\"></a></p></li>\n<li><p>(and I randomly selected 50 articles as the negative sample for each customer in the same manner as ganchan)</p></li></ul></li>\n<li><p>use Catboost Ranker (ganchan’s model is LGBM Classifier)</p>\n<ul>\n<li>time-series splits, train:~9/15, valid: 9/16~9/22</li></ul></li></ul></li></ul></li>\n</ul>\n<h1>Ensemble model</h1>\n<ul>\n<li><p><a href=\"https://github.com/haradai1262/h-and-m-personalized-fashion-recommendations/blob/master/exp/stacking_nb_6.ipynb\" target=\"_blank\">Code</a></p></li>\n<li><p>Score</p>\n<ul>\n<li>CV: 0.0407</li>\n<li>Public LB: 0.03296</li>\n<li>Private LB: 0.03294</li></ul></li>\n<li><p>Candidates</p>\n<ul>\n<li>extract top 24 predicted items from each model<ul>\n<li>we use 13 models, which are members’ best single models and their variants</li>\n<li>public LB of their models: 0.0307 ~ 0.0294</li></ul></li></ul></li>\n<li><p>Features</p>\n<ul>\n<li><p>Predict features (1st-stage model outputs)</p>\n<ul>\n<li><p>predicted value</p></li>\n<li><p>predicted value normilized per customer</p>\n<pre><code>preds[\"pred_norm\"] = preds.groupby('customer_id')[\"pred\"].transform(lambda x: (x - x.mean()) / x.std())\n</code></pre></li>\n<li><p>rank based on predicted value per customer</p>\n<pre><code>preds[\"pred_rank\"] = preds.groupby(\"customer_id\")[\"pred\"].rank(ascending=False, method=\"dense\")\n</code></pre></li></ul></li>\n<li><p>aggregate predict features (group by models)</p>\n<ul>\n<li>count, sum, max, min</li></ul></li>\n<li><p>article attributes</p></li>\n<li><p>customer attributes</p></li>\n<li><p>article buy count {yesterday, last week, last month, last year, time-decayed count}</p></li></ul></li>\n<li><p>Model</p>\n<ul>\n<li>CatboostRanker (Groupkfold, group=customer, k=10)</li></ul></li>\n<li><p>My note</p>\n<ul>\n<li>our ensemble effectiveness (only public LB, easy summary not strict)<ul>\n<li>0.0307: single</li>\n<li>0.0318 add <a href=\"https://www.kaggle.com/ganchan\" target=\"_blank\">@ganchan</a> <a href=\"https://www.kaggle.com/NARI\" target=\"_blank\">@NARI</a></li>\n<li>0.0322 add <a href=\"https://www.kaggle.com/minguin\" target=\"_blank\">@minguin</a></li>\n<li>0.0329 add <a href=\"https://www.kaggle.com/Moro\" target=\"_blank\">@Moro</a></li>\n<li>great effect on the ensemble! but more effective fusion methods might have suited.</li></ul></li></ul></li>\n</ul>",
      "rawMarkdown": "This is part of the 22nd solution (sorry for the split sharing by each member)\n\n- members: @ganchan @NARI @Moro @minguin, thanks~\n- ganchan's solution: [https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152)\n- moro's solution: [https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224)\n\nI simply share part of my solution (my single model and ensemble model) and the code.\n\n# Single model\n\n- [Code](https://github.com/haradai1262/h-and-m-personalized-fashion-recommendations/blob/master/exp/cat_v4-3.py)\n- Score\n    - CV: 0.0365\n    - Public LB: 0.03072\n    - Private LB: 0.03048\n- This model is forked from ganchan’s model\n    - [Part of 22nd solution - single LGBM (Private:0.03038)](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152)\n    - diff\n        - dataset for training is ten weeks (ganchan’s model is four weeks)\n            - I use learning weights that get smaller the further away from the latest date\n            - each week, I select candidates from top sales of 15000 items in the previous four weeks.\n                <a href=\"https://ibb.co/zNybH8f\"><img src=\"https://i.ibb.co/V37DpHB/sampling.png\" alt=\"sampling\" border=\"0\"></a>\n                \n            - (and I randomly selected 50 articles as the negative sample for each customer in the same manner as ganchan)\n        - use Catboost Ranker (ganchan’s model is LGBM Classifier)\n            - time-series splits, train:~9/15, valid: 9/16~9/22\n\n# Ensemble model\n\n- [Code](https://github.com/haradai1262/h-and-m-personalized-fashion-recommendations/blob/master/exp/stacking_nb_6.ipynb)\n- Score\n    - CV: 0.0407\n    - Public LB: 0.03296\n    - Private LB: 0.03294\n- Candidates\n    - extract top 24 predicted items from each model\n        - we use 13 models, which are members’ best single models and their variants\n        - public LB of their models: 0.0307 ~ 0.0294\n- Features\n    - Predict features (1st-stage model outputs)\n        - predicted value\n        - predicted value normilized per customer\n            \n            ```python\n            preds[\"pred_norm\"] = preds.groupby('customer_id')[\"pred\"].transform(lambda x: (x - x.mean()) / x.std())\n            ```\n            \n        - rank based on predicted value per customer\n            \n            ```python\n            preds[\"pred_rank\"] = preds.groupby(\"customer_id\")[\"pred\"].rank(ascending=False, method=\"dense\")\n            ```\n    - aggregate predict features (group by models)\n        - count, sum, max, min\n    - article attributes\n    - customer attributes\n    - article buy count {yesterday, last week, last month, last year, time-decayed count}\n- Model\n    - CatboostRanker (Groupkfold, group=customer, k=10)\n- My note\n    - our ensemble effectiveness (only public LB, easy summary not strict)\n        - 0.0307: single\n        - 0.0318 add @ganchan @NARI\n        - 0.0322 add @minguin\n        - 0.0329 add @Moro\n        - great effect on the ensemble! but more effective fusion methods might have suited.",
      "votes": null
    },
    {
      "id": "1791953",
      "postDate": "05/16/2022 13:52:54",
      "content": "<p>Great work! This is a really well-thought-out solution and I love the details you've included. It's clear that a lot of time and effort went into this and it shows.</p>",
      "rawMarkdown": "Great work! This is a really well-thought-out solution and I love the details you've included. It's clear that a lot of time and effort went into this and it shows.",
      "votes": null
    },
    {
      "id": "1973781",
      "postDate": "10/05/2022 20:35:14",
      "content": "<p>Great work and thank you for sharing! How did you generate the customer_age_gorup.csv? </p>",
      "rawMarkdown": "Great work and thank you for sharing! How did you generate the customer_age_gorup.csv?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1791953,
      "author_name": "",
      "author_url": "",
      "post_date": "05/16/2022 13:52:54",
      "content": "<p>Great work! This is a really well-thought-out solution and I love the details you've included. It's clear that a lot of time and effort went into this and it shows.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1973781,
      "author_name": "arcangelopisa",
      "author_url": "",
      "post_date": "10/05/2022 20:35:14",
      "content": "<p>Great work and thank you for sharing! How did you generate the customer_age_gorup.csv? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1791744": "This is part of the 22nd solution (sorry for the split sharing by each member)\n\n- members: @ganchan @NARI @Moro @minguin, thanks~\n- ganchan's solution: [https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152)\n- moro's solution: [https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224)\n\nI simply share part of my solution (my single model and ensemble model) and the code.\n\n# Single model\n\n- [Code](https://github.com/haradai1262/h-and-m-personalized-fashion-recommendations/blob/master/exp/cat_v4-3.py)\n- Score\n    - CV: 0.0365\n    - Public LB: 0.03072\n    - Private LB: 0.03048\n- This model is forked from ganchan’s model\n    - [Part of 22nd solution - single LGBM (Private:0.03038)](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/324152)\n    - diff\n        - dataset for training is ten weeks (ganchan’s model is four weeks)\n            - I use learning weights that get smaller the further away from the latest date\n            - each week, I select candidates from top sales of 15000 items in the previous four weeks.\n                <a href=\"https://ibb.co/zNybH8f\"><img src=\"https://i.ibb.co/V37DpHB/sampling.png\" alt=\"sampling\" border=\"0\"></a>\n                \n            - (and I randomly selected 50 articles as the negative sample for each customer in the same manner as ganchan)\n        - use Catboost Ranker (ganchan’s model is LGBM Classifier)\n            - time-series splits, train:~9/15, valid: 9/16~9/22\n\n# Ensemble model\n\n- [Code](https://github.com/haradai1262/h-and-m-personalized-fashion-recommendations/blob/master/exp/stacking_nb_6.ipynb)\n- Score\n    - CV: 0.0407\n    - Public LB: 0.03296\n    - Private LB: 0.03294\n- Candidates\n    - extract top 24 predicted items from each model\n        - we use 13 models, which are members’ best single models and their variants\n        - public LB of their models: 0.0307 ~ 0.0294\n- Features\n    - Predict features (1st-stage model outputs)\n        - predicted value\n        - predicted value normilized per customer\n            \n            ```python\n            preds[\"pred_norm\"] = preds.groupby('customer_id')[\"pred\"].transform(lambda x: (x - x.mean()) / x.std())\n            ```\n            \n        - rank based on predicted value per customer\n            \n            ```python\n            preds[\"pred_rank\"] = preds.groupby(\"customer_id\")[\"pred\"].rank(ascending=False, method=\"dense\")\n            ```\n    - aggregate predict features (group by models)\n        - count, sum, max, min\n    - article attributes\n    - customer attributes\n    - article buy count {yesterday, last week, last month, last year, time-decayed count}\n- Model\n    - CatboostRanker (Groupkfold, group=customer, k=10)\n- My note\n    - our ensemble effectiveness (only public LB, easy summary not strict)\n        - 0.0307: single\n        - 0.0318 add @ganchan @NARI\n        - 0.0322 add @minguin\n        - 0.0329 add @Moro\n        - great effect on the ensemble! but more effective fusion methods might have suited.",
    "1791953": "Great work! This is a really well-thought-out solution and I love the details you've included. It's clear that a lot of time and effort went into this and it shows.",
    "1973781": "Great work and thank you for sharing! How did you generate the customer_age_gorup.csv?"
  },
  "source": "meta"
}