{
  "id": 321564,
  "title": "Ranker - Positive and Negative observations",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/321564",
  "author_name": "",
  "post_date": "2022-04-27T12:53:31.504259600Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all, as I work on my <a href=\"https://www.kaggle.com/code/lorenzopagliaro01/h-m-ranker-pyspark-lgbmranker\" target=\"_blank\">LightGBM Ranker</a> I wonder if, besides generating Negative observations, <strong>Positive ones should be added to the Ranker too</strong>. Would it be incorrect to add \"fake bought\" items to the ranker? After all, only suggesting negative items to the model could teach it wrong patterns regarding positive candidates. Therefore I would have \"true bought\", \"suggested bought\", \"negative observations\".<br>\nI believe this could be a step I am currently missing in my work.<br>\nFor reference, I red <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">Pawel's introduction to Ranker</a> and followed his suggestions.<br>\nThank you for your comments,<br>\nLorenzo</p>",
  "messages": [
    {
      "id": "1769621",
      "postDate": "04/27/2022 12:53:31",
      "content": "<p>Hi all, as I work on my <a href=\"https://www.kaggle.com/code/lorenzopagliaro01/h-m-ranker-pyspark-lgbmranker\" target=\"_blank\">LightGBM Ranker</a> I wonder if, besides generating Negative observations, <strong>Positive ones should be added to the Ranker too</strong>. Would it be incorrect to add \"fake bought\" items to the ranker? After all, only suggesting negative items to the model could teach it wrong patterns regarding positive candidates. Therefore I would have \"true bought\", \"suggested bought\", \"negative observations\".<br>\nI believe this could be a step I am currently missing in my work.<br>\nFor reference, I red <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288\" target=\"_blank\">Pawel's introduction to Ranker</a> and followed his suggestions.<br>\nThank you for your comments,<br>\nLorenzo</p>",
      "rawMarkdown": "Hi all, as I work on my [LightGBM Ranker](https://www.kaggle.com/code/lorenzopagliaro01/h-m-ranker-pyspark-lgbmranker) I wonder if, besides generating Negative observations, **Positive ones should be added to the Ranker too**. Would it be incorrect to add \"fake bought\" items to the ranker? After all, only suggesting negative items to the model could teach it wrong patterns regarding positive candidates. Therefore I would have \"true bought\", \"suggested bought\", \"negative observations\".\nI believe this could be a step I am currently missing in my work.\nFor reference, I red [Pawel's introduction to Ranker](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288) and followed his suggestions.\nThank you for your comments,\nLorenzo",
      "votes": null
    },
    {
      "id": "1769670",
      "postDate": "04/27/2022 13:34:40",
      "content": "<p>I don't think it is necessary. Even wrong. The way a ranker works is that queries with all negative observations are useless for training - you must have at least 1 positive observation for the query. It is because when ranker calculates the gradient of some observation assuming it switches position with an observation with a different label. Queries with all negatives have sum of gradient 0.0 and do not contribute anything during training.</p>\n<p>See here</p>\n<p><a href=\"https://github.com/microsoft/LightGBM/blob/master/src/objective/rank_objective.hpp#L173\" target=\"_blank\">https://github.com/microsoft/LightGBM/blob/master/src/objective/rank_objective.hpp#L173</a></p>\n<pre><code>if (label[sorted_idx[i]] == label[sorted_idx[j]]) { continue; }\n</code></pre>\n<p>This line means that pair of candidates with the same label are skipped - if all labels = 0 then whole query is skipped.</p>\n<p>So if your candidates are all negatives you can remove this customer altogether. Only when you are predicting for submission you need to consider the candidates using the strategies you used for training.</p>\n<p>Hope it helps.</p>",
      "rawMarkdown": "I don't think it is necessary. Even wrong. The way a ranker works is that queries with all negative observations are useless for training - you must have at least 1 positive observation for the query. It is because when ranker calculates the gradient of some observation assuming it switches position with an observation with a different label. Queries with all negatives have sum of gradient 0.0 and do not contribute anything during training.\n\nSee here\n\nhttps://github.com/microsoft/LightGBM/blob/master/src/objective/rank_objective.hpp#L173\n\n```\nif (label[sorted_idx[i]] == label[sorted_idx[j]]) { continue; }\n```\n\nThis line means that pair of candidates with the same label are skipped - if all labels = 0 then whole query is skipped.\n\n\nSo if your candidates are all negatives you can remove this customer altogether. Only when you are predicting for submission you need to consider the candidates using the strategies you used for training.\n\nHope it helps.",
      "votes": null
    },
    {
      "id": "1769830",
      "postDate": "04/27/2022 16:42:24",
      "content": "<p>Dear Pawel, thank you for your clarification. I have a question:<br>\n<em>Queries with all negatives have sum of gradient 0.0 and do not contribute anything during training.</em><br>\nDoes this mean that increasing the number of negative tests for each query would not increase the ranker performance?<br>\nI understand that having all negative observations on a customer would not provide any benefit.<br>\nWarm regards,<br>\nLorenzo</p>",
      "rawMarkdown": "Dear Pawel, thank you for your clarification. I have a question:\n*Queries with all negatives have sum of gradient 0.0 and do not contribute anything during training.*\nDoes this mean that increasing the number of negative tests for each query would not increase the ranker performance?\nI understand that having all negative observations on a customer would not provide any benefit.\nWarm regards,\nLorenzo",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1769670,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "04/27/2022 13:34:40",
      "content": "<p>I don't think it is necessary. Even wrong. The way a ranker works is that queries with all negative observations are useless for training - you must have at least 1 positive observation for the query. It is because when ranker calculates the gradient of some observation assuming it switches position with an observation with a different label. Queries with all negatives have sum of gradient 0.0 and do not contribute anything during training.</p>\n<p>See here</p>\n<p><a href=\"https://github.com/microsoft/LightGBM/blob/master/src/objective/rank_objective.hpp#L173\" target=\"_blank\">https://github.com/microsoft/LightGBM/blob/master/src/objective/rank_objective.hpp#L173</a></p>\n<pre><code>if (label[sorted_idx[i]] == label[sorted_idx[j]]) { continue; }\n</code></pre>\n<p>This line means that pair of candidates with the same label are skipped - if all labels = 0 then whole query is skipped.</p>\n<p>So if your candidates are all negatives you can remove this customer altogether. Only when you are predicting for submission you need to consider the candidates using the strategies you used for training.</p>\n<p>Hope it helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1769830,
          "author_name": "lorenzopagliaro01",
          "author_url": "",
          "post_date": "04/27/2022 16:42:24",
          "content": "<p>Dear Pawel, thank you for your clarification. I have a question:<br>\n<em>Queries with all negatives have sum of gradient 0.0 and do not contribute anything during training.</em><br>\nDoes this mean that increasing the number of negative tests for each query would not increase the ranker performance?<br>\nI understand that having all negative observations on a customer would not provide any benefit.<br>\nWarm regards,<br>\nLorenzo</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1769621": "Hi all, as I work on my [LightGBM Ranker](https://www.kaggle.com/code/lorenzopagliaro01/h-m-ranker-pyspark-lgbmranker) I wonder if, besides generating Negative observations, **Positive ones should be added to the Ranker too**. Would it be incorrect to add \"fake bought\" items to the ranker? After all, only suggesting negative items to the model could teach it wrong patterns regarding positive candidates. Therefore I would have \"true bought\", \"suggested bought\", \"negative observations\".\nI believe this could be a step I am currently missing in my work.\nFor reference, I red [Pawel's introduction to Ranker](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288) and followed his suggestions.\nThank you for your comments,\nLorenzo",
    "1769670": "I don't think it is necessary. Even wrong. The way a ranker works is that queries with all negative observations are useless for training - you must have at least 1 positive observation for the query. It is because when ranker calculates the gradient of some observation assuming it switches position with an observation with a different label. Queries with all negatives have sum of gradient 0.0 and do not contribute anything during training.\n\nSee here\n\nhttps://github.com/microsoft/LightGBM/blob/master/src/objective/rank_objective.hpp#L173\n\n```\nif (label[sorted_idx[i]] == label[sorted_idx[j]]) { continue; }\n```\n\nThis line means that pair of candidates with the same label are skipped - if all labels = 0 then whole query is skipped.\n\n\nSo if your candidates are all negatives you can remove this customer altogether. Only when you are predicting for submission you need to consider the candidates using the strategies you used for training.\n\nHope it helps.",
    "1769830": "Dear Pawel, thank you for your clarification. I have a question:\n*Queries with all negatives have sum of gradient 0.0 and do not contribute anything during training.*\nDoes this mean that increasing the number of negative tests for each query would not increase the ranker performance?\nI understand that having all negative observations on a customer would not provide any benefit.\nWarm regards,\nLorenzo"
  },
  "source": "meta"
}