{
  "id": 316164,
  "title": "Ranker is important?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/316164",
  "author_name": "",
  "post_date": "2022-03-31T14:41:57.733125800Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>As mentioned in <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288</a>, I tried to make ranking model. However, most of baseline notebooks such as <a href=\"https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021</a> do not use it.<br>\nSo, I want to know whether you use it or not. Do you use ranking model?</p>",
  "messages": [
    {
      "id": "1741139",
      "postDate": "03/31/2022 14:41:57",
      "content": "<p>As mentioned in <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288</a>, I tried to make ranking model. However, most of baseline notebooks such as <a href=\"https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021</a> do not use it.<br>\nSo, I want to know whether you use it or not. Do you use ranking model?</p>",
      "rawMarkdown": "As mentioned in https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288, I tried to make ranking model. However, most of baseline notebooks such as https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021 do not use it.\nSo, I want to know whether you use it or not. Do you use ranking model?",
      "votes": null
    },
    {
      "id": "1741521",
      "postDate": "03/31/2022 21:27:23",
      "content": "<p>I'm experimenting with a LGBMRanker for the first time. The training is up and running, but I had to make some decisions on how to prepare the training data and I'm not sure if I got it right.</p>\n<p>It doesn't seem to be viable to train queries containing all 100K articles for each of the 1.3M customers, and certainly not for more than a handful of weeks. Hence, I'm basically using the actual purchases as an anchor for customers of a few weeks and adding negative instances for each customer to control the size of the data set. However, this way I have only few records (8-20 articles) for each query (the customer). Would be great to hear some opinions, whether this tactic makes sense or is pretty much bullshit.</p>\n<p>I noticed, that once you have the training data prepared for the Ranker you can easily switch to a Regressor or Classifier. So I will try this out as well.</p>\n<p>I'm still coding the actual prediction/validation part. Currently I've got no clue how my training loss converts into MAP@12 :-) For the prediction I'm back to the problem of 1.3M queries with 100K records each. Even if I can reduce the articles significantly with some heuristic this is still a lot of data.</p>\n<p>Maybe an option is a two level stack: a regressor that filters candidate articles and then a ranker that operates on the top 50 articles for each customer.</p>",
      "rawMarkdown": "I'm experimenting with a LGBMRanker for the first time. The training is up and running, but I had to make some decisions on how to prepare the training data and I'm not sure if I got it right.\n\nIt doesn't seem to be viable to train queries containing all 100K articles for each of the 1.3M customers, and certainly not for more than a handful of weeks. Hence, I'm basically using the actual purchases as an anchor for customers of a few weeks and adding negative instances for each customer to control the size of the data set. However, this way I have only few records (8-20 articles) for each query (the customer). Would be great to hear some opinions, whether this tactic makes sense or is pretty much bullshit.\n\nI noticed, that once you have the training data prepared for the Ranker you can easily switch to a Regressor or Classifier. So I will try this out as well.\n\nI'm still coding the actual prediction/validation part. Currently I've got no clue how my training loss converts into MAP@12 :-) For the prediction I'm back to the problem of 1.3M queries with 100K records each. Even if I can reduce the articles significantly with some heuristic this is still a lot of data.\n\nMaybe an option is a two level stack: a regressor that filters candidate articles and then a ranker that operates on the top 50 articles for each customer.",
      "votes": null
    },
    {
      "id": "1742299",
      "postDate": "04/01/2022 17:02:30",
      "content": "<p>It helps me - even without hyperparameter tuning, and just using the same features I was using for sorting the candidates manually.</p>",
      "rawMarkdown": "It helps me - even without hyperparameter tuning, and just using the same features I was using for sorting the candidates manually.",
      "votes": null
    },
    {
      "id": "1743941",
      "postDate": "04/03/2022 12:56:12",
      "content": "<p>Thanks for your reply.<br>\nI have tried the Ranker model but am struggling with the CV method. Could you please tell me how you are doing the CV?</p>",
      "rawMarkdown": "Thanks for your reply.\nI have tried the Ranker model but am struggling with the CV method. Could you please tell me how you are doing the CV?",
      "votes": null
    },
    {
      "id": "1743991",
      "postDate": "04/03/2022 13:34:44",
      "content": "<p>I also struggled with it!<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386\" target=\"_blank\">Here's a post</a> detailing how I'm doing it.</p>",
      "rawMarkdown": "I also struggled with it!\n[Here's a post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386) detailing how I'm doing it.",
      "votes": null
    },
    {
      "id": "1744016",
      "postDate": "04/03/2022 14:04:49",
      "content": "<p>Thank you for sharing great notebook!<br>\nThis has a lot of things I can learn!</p>",
      "rawMarkdown": "Thank you for sharing great notebook!\nThis has a lot of things I can learn!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1741521,
      "author_name": "tbierhance",
      "author_url": "",
      "post_date": "03/31/2022 21:27:23",
      "content": "<p>I'm experimenting with a LGBMRanker for the first time. The training is up and running, but I had to make some decisions on how to prepare the training data and I'm not sure if I got it right.</p>\n<p>It doesn't seem to be viable to train queries containing all 100K articles for each of the 1.3M customers, and certainly not for more than a handful of weeks. Hence, I'm basically using the actual purchases as an anchor for customers of a few weeks and adding negative instances for each customer to control the size of the data set. However, this way I have only few records (8-20 articles) for each query (the customer). Would be great to hear some opinions, whether this tactic makes sense or is pretty much bullshit.</p>\n<p>I noticed, that once you have the training data prepared for the Ranker you can easily switch to a Regressor or Classifier. So I will try this out as well.</p>\n<p>I'm still coding the actual prediction/validation part. Currently I've got no clue how my training loss converts into MAP@12 :-) For the prediction I'm back to the problem of 1.3M queries with 100K records each. Even if I can reduce the articles significantly with some heuristic this is still a lot of data.</p>\n<p>Maybe an option is a two level stack: a regressor that filters candidate articles and then a ranker that operates on the top 50 articles for each customer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1742299,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "04/01/2022 17:02:30",
      "content": "<p>It helps me - even without hyperparameter tuning, and just using the same features I was using for sorting the candidates manually.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1743941,
          "author_name": "tyabanoamami",
          "author_url": "",
          "post_date": "04/03/2022 12:56:12",
          "content": "<p>Thanks for your reply.<br>\nI have tried the Ranker model but am struggling with the CV method. Could you please tell me how you are doing the CV?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1743991,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/03/2022 13:34:44",
          "content": "<p>I also struggled with it!<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386\" target=\"_blank\">Here's a post</a> detailing how I'm doing it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1744016,
          "author_name": "tyabanoamami",
          "author_url": "",
          "post_date": "04/03/2022 14:04:49",
          "content": "<p>Thank you for sharing great notebook!<br>\nThis has a lot of things I can learn!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1741139": "As mentioned in https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288, I tried to make ranking model. However, most of baseline notebooks such as https://www.kaggle.com/code/cdeotte/recommend-items-purchased-together-0-021 do not use it.\nSo, I want to know whether you use it or not. Do you use ranking model?",
    "1741521": "I'm experimenting with a LGBMRanker for the first time. The training is up and running, but I had to make some decisions on how to prepare the training data and I'm not sure if I got it right.\n\nIt doesn't seem to be viable to train queries containing all 100K articles for each of the 1.3M customers, and certainly not for more than a handful of weeks. Hence, I'm basically using the actual purchases as an anchor for customers of a few weeks and adding negative instances for each customer to control the size of the data set. However, this way I have only few records (8-20 articles) for each query (the customer). Would be great to hear some opinions, whether this tactic makes sense or is pretty much bullshit.\n\nI noticed, that once you have the training data prepared for the Ranker you can easily switch to a Regressor or Classifier. So I will try this out as well.\n\nI'm still coding the actual prediction/validation part. Currently I've got no clue how my training loss converts into MAP@12 :-) For the prediction I'm back to the problem of 1.3M queries with 100K records each. Even if I can reduce the articles significantly with some heuristic this is still a lot of data.\n\nMaybe an option is a two level stack: a regressor that filters candidate articles and then a ranker that operates on the top 50 articles for each customer.",
    "1742299": "It helps me - even without hyperparameter tuning, and just using the same features I was using for sorting the candidates manually.",
    "1743941": "Thanks for your reply.\nI have tried the Ranker model but am struggling with the CV method. Could you please tell me how you are doing the CV?",
    "1743991": "I also struggled with it!\n[Here's a post](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/316386) detailing how I'm doing it.",
    "1744016": "Thank you for sharing great notebook!\nThis has a lot of things I can learn!"
  },
  "source": "meta"
}