{
  "id": 324152,
  "title": "Part of 22nd solution - single LGBM (Private:0.03038)",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/h-m-m-n-m-g-m-part-of-22nd-solution-single-lgbm-pr",
  "author_name": "",
  "post_date": "2022-05-10T13:29:56.377Z",
  "votes": 45,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I'd like to thank this great competition's hosts and congrats winners and all medalists!</p>\n<p>Super thanks <a href=\"https://www.kaggle.com/myaun\" target=\"_blank\">@myaun</a> <a href=\"https://www.kaggle.com/NARI\" target=\"_blank\">@NARI</a> <a href=\"https://www.kaggle.com/Moro\" target=\"_blank\">@Moro</a> <a href=\"https://www.kaggle.com/minguin\" target=\"_blank\">@minguin</a> for giving great ideas and solution as a team.</p>\n<p>In this post, I would show you my part of solution.</p>\n<p>I think my solution is quite simple but takes time to train and infer.</p>\n<p>My solution is published as Kaggle notebook.<br>\nIf you like it ,please post comments and upvote. <br>\n<a href=\"https://www.kaggle.com/code/iwatatakuya/22nd-place-lgbm-model-single-train\" target=\"_blank\">22nd-place-lgbm-model-single-train</a><br>\n<a href=\"https://www.kaggle.com/code/iwatatakuya/22nd-place-lgbm-model-single-infer\" target=\"_blank\">22nd-place-lgbm-model-single-infer</a></p>\n<ol>\n<li><p>environment and validation strategy<br>\nI used mainly Google Colab Pro+ because I need GPU (cudf, cuml) for fast training and inference with heavy data processing.<br>\nIt takes about 15 hours to train and 7 hours to infer (for validation data) and 100 hours to infer (for test data) with 1 GPU.<br>\nGoogle Colab Pro+'s multi-session made it 2 times faster but still it takes too much time.<br>\nI used 2020/9/16~9/22 data for validation and used ~2020/9/15 for training.</p></li>\n<li><p>Dataset for training<br>\nMaybe this part is most characteristic in my solution.<br>\nFirst, I made dataset of customers who bought some items in 2020/9/9~9/15, 9/2~9/8, 8/25~9/1 and 8/18~8/24.<br>\nThe primary key are customer_id and article_id and each pair has target column with 1 (bought) or 0 (did not buy).<br>\nWe cannot use all articles which customer did not buy because they are too much quantity, so I randomly selected 30 articles as negative sample for each customer.<br>\nI think features I made are not so special such as n_buy for customer and article.<br>\nPlease see my notebook for detail.</p></li>\n<li><p>Training<br>\nI used LGBM Classifier (not Ranker).<br>\nIn my case, Classifier was better than Ranker.</p></li>\n<li><p>Inference<br>\nI scored each pair of customer and item and sorted to predict items which the customer seems to buy.<br>\nCandidate items are top sales 15,000 items in 2020/8/18~2020/9/15.<br>\nCandidate items are same among customers.<br>\nSuch large number of candidates makes inference slower but it effects score significantly.</p></li>\n</ol>\n<p>Thank you for reading this far.<br>\nQuestions are welcome!</p>",
  "messages": [
    {
      "id": "1783426",
      "postDate": "05/10/2022 11:49:49",
      "content": "<p>I'd like to thank this great competition's hosts and congrats winners and all medalists!</p>\n<p>Super thanks <a href=\"https://www.kaggle.com/myaun\" target=\"_blank\">@myaun</a> <a href=\"https://www.kaggle.com/NARI\" target=\"_blank\">@NARI</a> <a href=\"https://www.kaggle.com/Moro\" target=\"_blank\">@Moro</a> <a href=\"https://www.kaggle.com/minguin\" target=\"_blank\">@minguin</a> for giving great ideas and solution as a team.</p>\n<p>In this post, I would show you my part of solution.</p>\n<p>I think my solution is quite simple but takes time to train and infer.</p>\n<p>My solution is published as Kaggle notebook.<br>\nIf you like it ,please post comments and upvote. <br>\n<a href=\"https://www.kaggle.com/code/iwatatakuya/22nd-place-lgbm-model-single-train\" target=\"_blank\">22nd-place-lgbm-model-single-train</a><br>\n<a href=\"https://www.kaggle.com/code/iwatatakuya/22nd-place-lgbm-model-single-infer\" target=\"_blank\">22nd-place-lgbm-model-single-infer</a></p>\n<ol>\n<li><p>environment and validation strategy<br>\nI used mainly Google Colab Pro+ because I need GPU (cudf, cuml) for fast training and inference with heavy data processing.<br>\nIt takes about 15 hours to train and 7 hours to infer (for validation data) and 100 hours to infer (for test data) with 1 GPU.<br>\nGoogle Colab Pro+'s multi-session made it 2 times faster but still it takes too much time.<br>\nI used 2020/9/16~9/22 data for validation and used ~2020/9/15 for training.</p></li>\n<li><p>Dataset for training<br>\nMaybe this part is most characteristic in my solution.<br>\nFirst, I made dataset of customers who bought some items in 2020/9/9~9/15, 9/2~9/8, 8/25~9/1 and 8/18~8/24.<br>\nThe primary key are customer_id and article_id and each pair has target column with 1 (bought) or 0 (did not buy).<br>\nWe cannot use all articles which customer did not buy because they are too much quantity, so I randomly selected 30 articles as negative sample for each customer.<br>\nI think features I made are not so special such as n_buy for customer and article.<br>\nPlease see my notebook for detail.</p></li>\n<li><p>Training<br>\nI used LGBM Classifier (not Ranker).<br>\nIn my case, Classifier was better than Ranker.</p></li>\n<li><p>Inference<br>\nI scored each pair of customer and item and sorted to predict items which the customer seems to buy.<br>\nCandidate items are top sales 15,000 items in 2020/8/18~2020/9/15.<br>\nCandidate items are same among customers.<br>\nSuch large number of candidates makes inference slower but it effects score significantly.</p></li>\n</ol>\n<p>Thank you for reading this far.<br>\nQuestions are welcome!</p>",
      "rawMarkdown": "I'd like to thank this great competition's hosts and congrats winners and all medalists!\n\nSuper thanks @myaun @NARI @Moro @minguin for giving great ideas and solution as a team.\n\n\nIn this post, I would show you my part of solution.\n\nI think my solution is quite simple but takes time to train and infer.\n\nMy solution is published as Kaggle notebook.\nIf you like it ,please post comments and upvote. \n[22nd-place-lgbm-model-single-train](https://www.kaggle.com/code/iwatatakuya/22nd-place-lgbm-model-single-train)\n[22nd-place-lgbm-model-single-infer](https://www.kaggle.com/code/iwatatakuya/22nd-place-lgbm-model-single-infer)\n\n0. environment and validation strategy\nI used mainly Google Colab Pro+ because I need GPU (cudf, cuml) for fast training and inference with heavy data processing.\nIt takes about 15 hours to train and 7 hours to infer (for validation data) and 100 hours to infer (for test data) with 1 GPU.\nGoogle Colab Pro+'s multi-session made it 2 times faster but still it takes too much time.\nI used 2020/9/16~9/22 data for validation and used ~2020/9/15 for training.\n\n1. Dataset for training\nMaybe this part is most characteristic in my solution.\nFirst, I made dataset of customers who bought some items in 2020/9/9~9/15, 9/2~9/8, 8/25~9/1 and 8/18~8/24.\nThe primary key are customer_id and article_id and each pair has target column with 1 (bought) or 0 (did not buy).\nWe cannot use all articles which customer did not buy because they are too much quantity, so I randomly selected 30 articles as negative sample for each customer.\nI think features I made are not so special such as n_buy for customer and article.\nPlease see my notebook for detail.\n\n2. Training\nI used LGBM Classifier (not Ranker).\nIn my case, Classifier was better than Ranker.\n\n3. Inference\nI scored each pair of customer and item and sorted to predict items which the customer seems to buy.\nCandidate items are top sales 15,000 items in 2020/8/18~2020/9/15.\nCandidate items are same among customers.\nSuch large number of candidates makes inference slower but it effects score significantly.\n\n\nThank you for reading this far.\nQuestions are welcome!",
      "votes": null
    },
    {
      "id": "1783799",
      "postDate": "05/10/2022 17:22:08",
      "content": "<p>Congratulations on silver medal!</p>\n<blockquote>\n  <p>Training<br>\n  I used LGBM Classifier (not Ranker).<br>\n  In my case, Classifier was better than Ranker.</p>\n</blockquote>\n<p>Could you please share with us the CV/LB for Classifier VS Ranker? <br>\nI'm curious about that difference</p>",
      "rawMarkdown": "Congratulations on silver medal!\n> Training\nI used LGBM Classifier (not Ranker).\nIn my case, Classifier was better than Ranker.\n\nCould you please share with us the CV/LB for Classifier VS Ranker? \nI'm curious about that difference",
      "votes": null
    },
    {
      "id": "1783810",
      "postDate": "05/10/2022 17:38:14",
      "content": "<p>Nice, well done! And thanks for sharing code!</p>",
      "rawMarkdown": "Nice, well done! And thanks for sharing code!",
      "votes": null
    },
    {
      "id": "1784016",
      "postDate": "05/10/2022 21:05:58",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "rawMarkdown": "Congrats! Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "1784213",
      "postDate": "05/11/2022 03:01:56",
      "content": "<p>Congrats! Thanks for sharing great work.<br>\nCan you share the code of the calculating articles2vec part?</p>",
      "rawMarkdown": "Congrats! Thanks for sharing great work.\nCan you share the code of the calculating articles2vec part?",
      "votes": null
    },
    {
      "id": "1784423",
      "postDate": "05/11/2022 07:08:17",
      "content": "<p>Thank you for your comment!<br>\nI used public notebook's articles2vec<br>\n<a href=\"https://www.kaggle.com/code/aerdem4/h-m-rapids-article2vec\" target=\"_blank\">h-m-rapids-article2vec</a></p>",
      "rawMarkdown": "Thank you for your comment!\nI used public notebook's articles2vec\n[h-m-rapids-article2vec](https://www.kaggle.com/code/aerdem4/h-m-rapids-article2vec)",
      "votes": null
    },
    {
      "id": "1784427",
      "postDate": "05/11/2022 07:12:41",
      "content": "<p>Thank you for your comments!<br>\nIn my case, as below.<br>\nClassifier-&gt; CV:0.03559  Public:0.0303<br>\nRanker-&gt;CV:0.03497 Public: not submitted<br>\nI did not submit ranker because it takes tooo  long time to make a submission file…</p>",
      "rawMarkdown": "Thank you for your comments!\nIn my case, as below.\nClassifier-> CV:0.03559  Public:0.0303\nRanker->CV:0.03497 Public: not submitted\nI did not submit ranker because it takes tooo  long time to make a submission file...",
      "votes": null
    },
    {
      "id": "1784444",
      "postDate": "05/11/2022 07:27:16",
      "content": "<p>Great works! Thank for your sharing :)</p>",
      "rawMarkdown": "Great works! Thank for your sharing :)",
      "votes": null
    },
    {
      "id": "1785426",
      "postDate": "05/12/2022 05:01:26",
      "content": "<p>Your python code is very beautiful. Thank you for sharing!</p>",
      "rawMarkdown": "Your python code is very beautiful. Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1792745",
      "postDate": "05/17/2022 08:56:40",
      "content": "<p>Congrats! Simple but effective!</p>",
      "rawMarkdown": "Congrats! Simple but effective!",
      "votes": null
    },
    {
      "id": "1966315",
      "postDate": "10/01/2022 21:05:14",
      "content": "<p>Congrats and thank you for share!<br>\nCould you please share how did you generate those pkl files? Thank you </p>\n<p>list_model = pd.read_pickle(f\"../input/hmmodel/{iter_train}<em>models</em>{idx_file}_{day_start_val.date()}.pkl\")</p>",
      "rawMarkdown": "Congrats and thank you for share!\nCould you please share how did you generate those pkl files? Thank you \n\nlist_model = pd.read_pickle(f\"../input/hmmodel/{iter_train}_models_{idx_file}_{day_start_val.date()}.pkl\")",
      "votes": null
    },
    {
      "id": "2058604",
      "postDate": "12/08/2022 04:23:00",
      "content": "<p>Thanks and it helps a lot! From your description: Such large number of candidates makes inference slower but it effects score significantly -- I am wondering how does the number of candidate influence the score quantitatively? What's your thought using a some retrieval stage with various methods? Thanks!</p>",
      "rawMarkdown": "Thanks and it helps a lot! From your description: Such large number of candidates makes inference slower but it effects score significantly -- I am wondering how does the number of candidate influence the score quantitatively? What's your thought using a some retrieval stage with various methods? Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783799,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "05/10/2022 17:22:08",
      "content": "<p>Congratulations on silver medal!</p>\n<blockquote>\n  <p>Training<br>\n  I used LGBM Classifier (not Ranker).<br>\n  In my case, Classifier was better than Ranker.</p>\n</blockquote>\n<p>Could you please share with us the CV/LB for Classifier VS Ranker? <br>\nI'm curious about that difference</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784427,
          "author_name": "iwatatakuya",
          "author_url": "",
          "post_date": "05/11/2022 07:12:41",
          "content": "<p>Thank you for your comments!<br>\nIn my case, as below.<br>\nClassifier-&gt; CV:0.03559  Public:0.0303<br>\nRanker-&gt;CV:0.03497 Public: not submitted<br>\nI did not submit ranker because it takes tooo  long time to make a submission file…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1783810,
      "author_name": "marcogorelli",
      "author_url": "",
      "post_date": "05/10/2022 17:38:14",
      "content": "<p>Nice, well done! And thanks for sharing code!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784016,
      "author_name": "subhambanga",
      "author_url": "",
      "post_date": "05/10/2022 21:05:58",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784213,
      "author_name": "hyungeunjo",
      "author_url": "",
      "post_date": "05/11/2022 03:01:56",
      "content": "<p>Congrats! Thanks for sharing great work.<br>\nCan you share the code of the calculating articles2vec part?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784423,
          "author_name": "iwatatakuya",
          "author_url": "",
          "post_date": "05/11/2022 07:08:17",
          "content": "<p>Thank you for your comment!<br>\nI used public notebook's articles2vec<br>\n<a href=\"https://www.kaggle.com/code/aerdem4/h-m-rapids-article2vec\" target=\"_blank\">h-m-rapids-article2vec</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1784444,
      "author_name": "quydoanm",
      "author_url": "",
      "post_date": "05/11/2022 07:27:16",
      "content": "<p>Great works! Thank for your sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1785426,
      "author_name": "leewook",
      "author_url": "",
      "post_date": "05/12/2022 05:01:26",
      "content": "<p>Your python code is very beautiful. Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1792745,
      "author_name": "jiahongxie",
      "author_url": "",
      "post_date": "05/17/2022 08:56:40",
      "content": "<p>Congrats! Simple but effective!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1966315,
      "author_name": "arcangelopisa",
      "author_url": "",
      "post_date": "10/01/2022 21:05:14",
      "content": "<p>Congrats and thank you for share!<br>\nCould you please share how did you generate those pkl files? Thank you </p>\n<p>list_model = pd.read_pickle(f\"../input/hmmodel/{iter_train}<em>models</em>{idx_file}_{day_start_val.date()}.pkl\")</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2058604,
      "author_name": "hangcc",
      "author_url": "",
      "post_date": "12/08/2022 04:23:00",
      "content": "<p>Thanks and it helps a lot! From your description: Such large number of candidates makes inference slower but it effects score significantly -- I am wondering how does the number of candidate influence the score quantitatively? What's your thought using a some retrieval stage with various methods? Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1783426": "I'd like to thank this great competition's hosts and congrats winners and all medalists!\n\nSuper thanks @myaun @NARI @Moro @minguin for giving great ideas and solution as a team.\n\n\nIn this post, I would show you my part of solution.\n\nI think my solution is quite simple but takes time to train and infer.\n\nMy solution is published as Kaggle notebook.\nIf you like it ,please post comments and upvote. \n[22nd-place-lgbm-model-single-train](https://www.kaggle.com/code/iwatatakuya/22nd-place-lgbm-model-single-train)\n[22nd-place-lgbm-model-single-infer](https://www.kaggle.com/code/iwatatakuya/22nd-place-lgbm-model-single-infer)\n\n0. environment and validation strategy\nI used mainly Google Colab Pro+ because I need GPU (cudf, cuml) for fast training and inference with heavy data processing.\nIt takes about 15 hours to train and 7 hours to infer (for validation data) and 100 hours to infer (for test data) with 1 GPU.\nGoogle Colab Pro+'s multi-session made it 2 times faster but still it takes too much time.\nI used 2020/9/16~9/22 data for validation and used ~2020/9/15 for training.\n\n1. Dataset for training\nMaybe this part is most characteristic in my solution.\nFirst, I made dataset of customers who bought some items in 2020/9/9~9/15, 9/2~9/8, 8/25~9/1 and 8/18~8/24.\nThe primary key are customer_id and article_id and each pair has target column with 1 (bought) or 0 (did not buy).\nWe cannot use all articles which customer did not buy because they are too much quantity, so I randomly selected 30 articles as negative sample for each customer.\nI think features I made are not so special such as n_buy for customer and article.\nPlease see my notebook for detail.\n\n2. Training\nI used LGBM Classifier (not Ranker).\nIn my case, Classifier was better than Ranker.\n\n3. Inference\nI scored each pair of customer and item and sorted to predict items which the customer seems to buy.\nCandidate items are top sales 15,000 items in 2020/8/18~2020/9/15.\nCandidate items are same among customers.\nSuch large number of candidates makes inference slower but it effects score significantly.\n\n\nThank you for reading this far.\nQuestions are welcome!",
    "1783799": "Congratulations on silver medal!\n> Training\nI used LGBM Classifier (not Ranker).\nIn my case, Classifier was better than Ranker.\n\nCould you please share with us the CV/LB for Classifier VS Ranker? \nI'm curious about that difference",
    "1783810": "Nice, well done! And thanks for sharing code!",
    "1784016": "Congrats! Thanks for sharing :)",
    "1784213": "Congrats! Thanks for sharing great work.\nCan you share the code of the calculating articles2vec part?",
    "1784423": "Thank you for your comment!\nI used public notebook's articles2vec\n[h-m-rapids-article2vec](https://www.kaggle.com/code/aerdem4/h-m-rapids-article2vec)",
    "1784427": "Thank you for your comments!\nIn my case, as below.\nClassifier-> CV:0.03559  Public:0.0303\nRanker->CV:0.03497 Public: not submitted\nI did not submit ranker because it takes tooo  long time to make a submission file...",
    "1784444": "Great works! Thank for your sharing :)",
    "1785426": "Your python code is very beautiful. Thank you for sharing!",
    "1792745": "Congrats! Simple but effective!",
    "1966315": "Congrats and thank you for share!\nCould you please share how did you generate those pkl files? Thank you \n\nlist_model = pd.read_pickle(f\"../input/hmmodel/{iter_train}_models_{idx_file}_{day_start_val.date()}.pkl\")",
    "2058604": "Thanks and it helps a lot! From your description: Such large number of candidates makes inference slower but it effects score significantly -- I am wondering how does the number of candidate influence the score quantitatively? What's your thought using a some retrieval stage with various methods? Thanks!"
  },
  "source": "meta"
}