{
  "id": 321831,
  "title": "How to train the ranking model properly?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/321831",
  "author_name": "",
  "post_date": "2022-04-28T21:38:34.557903700Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My current method contains three steps:</p>\n<ol>\n<li>retrieve/recall items by some basic rules (buying history, popularity, etc)</li>\n<li>label items from step 1 according to the customer's behavior in the valid set</li>\n<li>train an lgb ranker to rank the candidates.</li>\n</ol>\n<p>The recall rate of the result in the first step is about 9% (which is really low). However, when I tried to generate more candidates to improve the overall recall rate, the performance of my lgb ranker dramatically decreased, so as the CV score on the valid set😨. Downsampling negative samples couldn't solve the problem.</p>\n<p>I'm new to this field, did I miss some essential steps in the process? Or do you guys have any suggestions? Thanks a lot!!!</p>\n<p>BTW, I'm looking for teammates who have roughly the same or higher score as me (currently 0.0255) and are willing to spend time on this competition (I can spend around 6 hours each day)😀</p>\n<hr>\n<p>Update: I missed the discussion <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672#1770555\" target=\"_blank\">here</a>, will try the mentioned method and update the result here.</p>\n<p>Update2: Seems like the problem still exists</p>",
  "messages": [
    {
      "id": "1771078",
      "postDate": "04/28/2022 21:38:34",
      "content": "<p>My current method contains three steps:</p>\n<ol>\n<li>retrieve/recall items by some basic rules (buying history, popularity, etc)</li>\n<li>label items from step 1 according to the customer's behavior in the valid set</li>\n<li>train an lgb ranker to rank the candidates.</li>\n</ol>\n<p>The recall rate of the result in the first step is about 9% (which is really low). However, when I tried to generate more candidates to improve the overall recall rate, the performance of my lgb ranker dramatically decreased, so as the CV score on the valid set😨. Downsampling negative samples couldn't solve the problem.</p>\n<p>I'm new to this field, did I miss some essential steps in the process? Or do you guys have any suggestions? Thanks a lot!!!</p>\n<p>BTW, I'm looking for teammates who have roughly the same or higher score as me (currently 0.0255) and are willing to spend time on this competition (I can spend around 6 hours each day)😀</p>\n<hr>\n<p>Update: I missed the discussion <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672#1770555\" target=\"_blank\">here</a>, will try the mentioned method and update the result here.</p>\n<p>Update2: Seems like the problem still exists</p>",
      "rawMarkdown": "My current method contains three steps:\n\n1. retrieve/recall items by some basic rules (buying history, popularity, etc)\n2. label items from step 1 according to the customer's behavior in the valid set\n3. train an lgb ranker to rank the candidates.\n\nThe recall rate of the result in the first step is about 9% (which is really low). However, when I tried to generate more candidates to improve the overall recall rate, the performance of my lgb ranker dramatically decreased, so as the CV score on the valid set😨. Downsampling negative samples couldn't solve the problem.\n\nI'm new to this field, did I miss some essential steps in the process? Or do you guys have any suggestions? Thanks a lot!!!\n\nBTW, I'm looking for teammates who have roughly the same or higher score as me (currently 0.0255) and are willing to spend time on this competition (I can spend around 6 hours each day)😀\n\n--------------------\nUpdate: I missed the discussion [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672#1770555), will try the mentioned method and update the result here.\n\nUpdate2: Seems like the problem still exists",
      "votes": null
    },
    {
      "id": "1771242",
      "postDate": "04/29/2022 03:15:10",
      "content": "<p>I find a dicussion which may be helpful for you.<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672</a></p>",
      "rawMarkdown": "I find a dicussion which may be helpful for you.\nhttps://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672",
      "votes": null
    },
    {
      "id": "1771243",
      "postDate": "04/29/2022 03:18:45",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "1771824",
      "postDate": "04/29/2022 15:54:24",
      "content": "<p>From my understanding the main element that changes the ranker's performance is the negative observations strategy. Each strategy has to be customized for the target group you are looking at - but not to the point of overfitting. The challenge is very complex yet super interesting. You got a nice placement anyway, so you must have gotten something right!</p>",
      "rawMarkdown": "From my understanding the main element that changes the ranker's performance is the negative observations strategy. Each strategy has to be customized for the target group you are looking at - but not to the point of overfitting. The challenge is very complex yet super interesting. You got a nice placement anyway, so you must have gotten something right!",
      "votes": null
    },
    {
      "id": "1772074",
      "postDate": "04/29/2022 19:43:17",
      "content": "<p>thanks for your reply!</p>",
      "rawMarkdown": "thanks for your reply!",
      "votes": null
    },
    {
      "id": "1774645",
      "postDate": "05/02/2022 11:02:22",
      "content": "<p>Our team has currently 0.242 using deep learning models. If you wouldn't mind, we can join your teams and have further communications 😀</p>",
      "rawMarkdown": "Our team has currently 0.242 using deep learning models. If you wouldn't mind, we can join your teams and have further communications 😀",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1771242,
      "author_name": "dengniewei",
      "author_url": "",
      "post_date": "04/29/2022 03:15:10",
      "content": "<p>I find a dicussion which may be helpful for you.<br>\n<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672\" target=\"_blank\">https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1771243,
          "author_name": "weipengzhang",
          "author_url": "",
          "post_date": "04/29/2022 03:18:45",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1771824,
      "author_name": "lorenzopagliaro01",
      "author_url": "",
      "post_date": "04/29/2022 15:54:24",
      "content": "<p>From my understanding the main element that changes the ranker's performance is the negative observations strategy. Each strategy has to be customized for the target group you are looking at - but not to the point of overfitting. The challenge is very complex yet super interesting. You got a nice placement anyway, so you must have gotten something right!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1772074,
          "author_name": "weipengzhang",
          "author_url": "",
          "post_date": "04/29/2022 19:43:17",
          "content": "<p>thanks for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1774645,
      "author_name": "yijunyang123",
      "author_url": "",
      "post_date": "05/02/2022 11:02:22",
      "content": "<p>Our team has currently 0.242 using deep learning models. If you wouldn't mind, we can join your teams and have further communications 😀</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1771078": "My current method contains three steps:\n\n1. retrieve/recall items by some basic rules (buying history, popularity, etc)\n2. label items from step 1 according to the customer's behavior in the valid set\n3. train an lgb ranker to rank the candidates.\n\nThe recall rate of the result in the first step is about 9% (which is really low). However, when I tried to generate more candidates to improve the overall recall rate, the performance of my lgb ranker dramatically decreased, so as the CV score on the valid set😨. Downsampling negative samples couldn't solve the problem.\n\nI'm new to this field, did I miss some essential steps in the process? Or do you guys have any suggestions? Thanks a lot!!!\n\nBTW, I'm looking for teammates who have roughly the same or higher score as me (currently 0.0255) and are willing to spend time on this competition (I can spend around 6 hours each day)😀\n\n--------------------\nUpdate: I missed the discussion [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672#1770555), will try the mentioned method and update the result here.\n\nUpdate2: Seems like the problem still exists",
    "1771242": "I find a dicussion which may be helpful for you.\nhttps://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/321672",
    "1771243": "Thanks a lot!",
    "1771824": "From my understanding the main element that changes the ranker's performance is the negative observations strategy. Each strategy has to be customized for the target group you are looking at - but not to the point of overfitting. The challenge is very complex yet super interesting. You got a nice placement anyway, so you must have gotten something right!",
    "1772074": "thanks for your reply!",
    "1774645": "Our team has currently 0.242 using deep learning models. If you wouldn't mind, we can join your teams and have further communications 😀"
  },
  "source": "meta"
}