{
  "id": 324293,
  "title": "25th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/keisuke-25th-place-solution",
  "author_name": "",
  "post_date": "2022-05-11T02:17:36.343Z",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks Kaggle and H&amp;M for hosting the really exciting competition. This is my first time where I devoted significant portion of my spare time on Kaggle platform, and I learned a lot from this competition. I appreciate all the great notebooks and dicussions, which have helped me a lot construct my strategy.</p>\n<p>My strategy was relatively simple, did nothing special, just followed the <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>‘s <a target=\"_blank\">great suggestions</a>.</p>\n<p>Candidate generation</p>\n<ul>\n<li>repurchase (still wonder why this is so important in this competition though…)</li>\n<li>Several collaborative filtering<ul>\n<li>ItemkNN</li>\n<li>Randomwalk with restart</li>\n<li>Matrix factorization with negative sampling</li></ul></li>\n<li>For ItemkNN, I used all the transactions with decaying weight as the timesamp goes past. This weighting reduces cold users while maintaining the recall. ItemkNN gave me the largest gain.</li>\n<li>I use Julia languate to implement them</li>\n<li>Candidates from (global/segmented) popular items also help</li>\n</ul>\n<p>Reranking</p>\n<ul>\n<li>The model is LightGBM LambdaMART and the features are based on popularity, user-item interaction and trending. This part is probably similar to many competitors.</li>\n<li>Rank features (rank of the item in generated candidates) are important.</li>\n<li>As I increased candidates, it became more and more important to adjust positive sample ratio and sampling negatives to deal with the label imbalance (in fact there are several dicussions to suggest this)</li>\n</ul>\n<p>My environment</p>\n<ul>\n<li>At first, I used cudf with the notebook instance attached with a GPU</li>\n<li>As I generated more and more candidates and features, even cudf got difficulty to handle such a large table, so I swithced to the Google BigQuery. I moved all my preprocess codes to SQLs. BigQuery provides me much more scalable way to handle tabular data.</li>\n<li>Since I started to focus on this competition at the bigining of the April, I had only a month to tackle on. After switching to BigQuery, I also moved to GCP vertex training, run lots of experiments in parallel with on-demand high-memory machine (finally with 512GB RAM). It was really useful to perform many trails-and-errors.</li>\n</ul>\n<p>My best model is</p>\n<ul>\n<li>CV: 0.03861</li>\n<li>Public: 0.03225</li>\n<li>Private: 0.03252</li>\n</ul>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1784092",
      "postDate": "05/10/2022 23:39:33",
      "content": "<p>Thanks Kaggle and H&amp;M for hosting the really exciting competition. This is my first time where I devoted significant portion of my spare time on Kaggle platform, and I learned a lot from this competition. I appreciate all the great notebooks and dicussions, which have helped me a lot construct my strategy.</p>\n<p>My strategy was relatively simple, did nothing special, just followed the <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a>‘s <a target=\"_blank\">great suggestions</a>.</p>\n<p>Candidate generation</p>\n<ul>\n<li>repurchase (still wonder why this is so important in this competition though…)</li>\n<li>Several collaborative filtering<ul>\n<li>ItemkNN</li>\n<li>Randomwalk with restart</li>\n<li>Matrix factorization with negative sampling</li></ul></li>\n<li>For ItemkNN, I used all the transactions with decaying weight as the timesamp goes past. This weighting reduces cold users while maintaining the recall. ItemkNN gave me the largest gain.</li>\n<li>I use Julia languate to implement them</li>\n<li>Candidates from (global/segmented) popular items also help</li>\n</ul>\n<p>Reranking</p>\n<ul>\n<li>The model is LightGBM LambdaMART and the features are based on popularity, user-item interaction and trending. This part is probably similar to many competitors.</li>\n<li>Rank features (rank of the item in generated candidates) are important.</li>\n<li>As I increased candidates, it became more and more important to adjust positive sample ratio and sampling negatives to deal with the label imbalance (in fact there are several dicussions to suggest this)</li>\n</ul>\n<p>My environment</p>\n<ul>\n<li>At first, I used cudf with the notebook instance attached with a GPU</li>\n<li>As I generated more and more candidates and features, even cudf got difficulty to handle such a large table, so I swithced to the Google BigQuery. I moved all my preprocess codes to SQLs. BigQuery provides me much more scalable way to handle tabular data.</li>\n<li>Since I started to focus on this competition at the bigining of the April, I had only a month to tackle on. After switching to BigQuery, I also moved to GCP vertex training, run lots of experiments in parallel with on-demand high-memory machine (finally with 512GB RAM). It was really useful to perform many trails-and-errors.</li>\n</ul>\n<p>My best model is</p>\n<ul>\n<li>CV: 0.03861</li>\n<li>Public: 0.03225</li>\n<li>Private: 0.03252</li>\n</ul>\n<p>Thanks.</p>",
      "rawMarkdown": "Thanks Kaggle and H&M for hosting the really exciting competition. This is my first time where I devoted significant portion of my spare time on Kaggle platform, and I learned a lot from this competition. I appreciate all the great notebooks and dicussions, which have helped me a lot construct my strategy.\n\nMy strategy was relatively simple, did nothing special, just followed the [@paweljankiewicz](https://www.kaggle.com/paweljankiewicz)‘s [great suggestions]([https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288#1688473](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288#1688473)).\n\nCandidate generation\n\n- repurchase (still wonder why this is so important in this competition though...)\n- Several collaborative filtering\n    - ItemkNN\n    - Randomwalk with restart\n    - Matrix factorization with negative sampling\n- For ItemkNN, I used all the transactions with decaying weight as the timesamp goes past. This weighting reduces cold users while maintaining the recall. ItemkNN gave me the largest gain.\n- I use Julia languate to implement them\n- Candidates from (global/segmented) popular items also help\n\nReranking\n\n- The model is LightGBM LambdaMART and the features are based on popularity, user-item interaction and trending. This part is probably similar to many competitors.\n- Rank features (rank of the item in generated candidates) are important.\n- As I increased candidates, it became more and more important to adjust positive sample ratio and sampling negatives to deal with the label imbalance (in fact there are several dicussions to suggest this)\n\nMy environment\n\n- At first, I used cudf with the notebook instance attached with a GPU\n- As I generated more and more candidates and features, even cudf got difficulty to handle such a large table, so I swithced to the Google BigQuery. I moved all my preprocess codes to SQLs. BigQuery provides me much more scalable way to handle tabular data.\n- Since I started to focus on this competition at the bigining of the April, I had only a month to tackle on. After switching to BigQuery, I also moved to GCP vertex training, run lots of experiments in parallel with on-demand high-memory machine (finally with 512GB RAM). It was really useful to perform many trails-and-errors.\n\nMy best model is\n\n- CV: 0.03861\n- Public: 0.03225\n- Private: 0.03252\n\nThanks.",
      "votes": null
    },
    {
      "id": "1784123",
      "postDate": "05/11/2022 00:20:07",
      "content": "<p><a href=\"https://www.kaggle.com/keisuke07\" target=\"_blank\">@keisuke07</a> congratulations and thank you for the detailed write up!</p>",
      "rawMarkdown": "keisuke07 congratulations and thank you for the detailed write up!",
      "votes": null
    },
    {
      "id": "1784161",
      "postDate": "05/11/2022 01:31:03",
      "content": "<p><a href=\"https://www.kaggle.com/keisuke07\" target=\"_blank\">@keisuke07</a> Thank you for sharing!</p>",
      "rawMarkdown": "keisuke07 Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1784270",
      "postDate": "05/11/2022 04:33:24",
      "content": "<p>wow this is exhaustive </p>",
      "rawMarkdown": "wow this is exhaustive",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1784123,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/11/2022 00:20:07",
      "content": "<p><a href=\"https://www.kaggle.com/keisuke07\" target=\"_blank\">@keisuke07</a> congratulations and thank you for the detailed write up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784161,
      "author_name": "rayuron",
      "author_url": "",
      "post_date": "05/11/2022 01:31:03",
      "content": "<p><a href=\"https://www.kaggle.com/keisuke07\" target=\"_blank\">@keisuke07</a> Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784270,
      "author_name": "bhavesh0124",
      "author_url": "",
      "post_date": "05/11/2022 04:33:24",
      "content": "<p>wow this is exhaustive </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1784092": "Thanks Kaggle and H&M for hosting the really exciting competition. This is my first time where I devoted significant portion of my spare time on Kaggle platform, and I learned a lot from this competition. I appreciate all the great notebooks and dicussions, which have helped me a lot construct my strategy.\n\nMy strategy was relatively simple, did nothing special, just followed the [@paweljankiewicz](https://www.kaggle.com/paweljankiewicz)‘s [great suggestions]([https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288#1688473](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288#1688473)).\n\nCandidate generation\n\n- repurchase (still wonder why this is so important in this competition though...)\n- Several collaborative filtering\n    - ItemkNN\n    - Randomwalk with restart\n    - Matrix factorization with negative sampling\n- For ItemkNN, I used all the transactions with decaying weight as the timesamp goes past. This weighting reduces cold users while maintaining the recall. ItemkNN gave me the largest gain.\n- I use Julia languate to implement them\n- Candidates from (global/segmented) popular items also help\n\nReranking\n\n- The model is LightGBM LambdaMART and the features are based on popularity, user-item interaction and trending. This part is probably similar to many competitors.\n- Rank features (rank of the item in generated candidates) are important.\n- As I increased candidates, it became more and more important to adjust positive sample ratio and sampling negatives to deal with the label imbalance (in fact there are several dicussions to suggest this)\n\nMy environment\n\n- At first, I used cudf with the notebook instance attached with a GPU\n- As I generated more and more candidates and features, even cudf got difficulty to handle such a large table, so I swithced to the Google BigQuery. I moved all my preprocess codes to SQLs. BigQuery provides me much more scalable way to handle tabular data.\n- Since I started to focus on this competition at the bigining of the April, I had only a month to tackle on. After switching to BigQuery, I also moved to GCP vertex training, run lots of experiments in parallel with on-demand high-memory machine (finally with 512GB RAM). It was really useful to perform many trails-and-errors.\n\nMy best model is\n\n- CV: 0.03861\n- Public: 0.03225\n- Private: 0.03252\n\nThanks.",
    "1784123": "keisuke07 congratulations and thank you for the detailed write up!",
    "1784161": "keisuke07 Thank you for sharing!",
    "1784270": "wow this is exhaustive"
  },
  "source": "meta"
}