{
  "id": 378604,
  "title": "Dumb question : which LightGBMRanker.predict() doesn't give a prediction between 0 and 1 ?",
  "url": "/competitions/otto-recommender-system/discussion/378604",
  "author_name": "",
  "post_date": "2023-01-16T13:37:45.908893200Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have a question about the actual nature of the prediction given by the ranker. For the LGBMRanker I noticed that the prediction belongs to the R-domain, I was expecting a probability between 0 and 1.</p>\n<p>Any theoretical explanation about it ? </p>",
  "messages": [
    {
      "id": "2102238",
      "postDate": "01/16/2023 13:37:45",
      "content": "<p>I have a question about the actual nature of the prediction given by the ranker. For the LGBMRanker I noticed that the prediction belongs to the R-domain, I was expecting a probability between 0 and 1.</p>\n<p>Any theoretical explanation about it ? </p>",
      "rawMarkdown": "I have a question about the actual nature of the prediction given by the ranker. For the LGBMRanker I noticed that the prediction belongs to the R-domain, I was expecting a probability between 0 and 1.\n\n\nAny theoretical explanation about it ?",
      "votes": null
    },
    {
      "id": "2102428",
      "postDate": "01/16/2023 15:49:45",
      "content": "<p>There is a parameter called <em>group</em> that you use when you fit (model.fit()) your Ranker model. What the model tries to achieve is to sort your train examples inside each group assigning it some specific score. In this particular example, you are not classifying whether the candidate item is good or not, you are trying within each session to find items that likely belong to it. </p>",
      "rawMarkdown": "There is a parameter called *group* that you use when you fit (model.fit()) your Ranker model. What the model tries to achieve is to sort your train examples inside each group assigning it some specific score. In this particular example, you are not classifying whether the candidate item is good or not, you are trying within each session to find items that likely belong to it.",
      "votes": null
    },
    {
      "id": "2102597",
      "postDate": "01/16/2023 18:16:02",
      "content": "<p>I agree that the intuition behind the ranker is quiet easy to understand, but I'm looking into theoretical explanations.</p>\n<p>Basically, I noticed that those numbers can be interpreted as probabilities of a item being relevant (or being at the top) or not.<br>\nAlso, I figured out that the calibration of the model didn't get impacted when applying a sigmoid/softmax function on the prediction.</p>\n<p>Finally, the answer I found for my question is that the predictions can be interpreted as relevance scores, where high values represent an expected high relevance ect..</p>",
      "rawMarkdown": "I agree that the intuition behind the ranker is quiet easy to understand, but I'm looking into theoretical explanations.\n\nBasically, I noticed that those numbers can be interpreted as probabilities of a item being relevant (or being at the top) or not.\nAlso, I figured out that the calibration of the model didn't get impacted when applying a sigmoid/softmax function on the prediction.\n\nFinally, the answer I found for my question is that the predictions can be interpreted as relevance scores, where high values represent an expected high relevance ect..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2102428,
      "author_name": "blindshooter",
      "author_url": "",
      "post_date": "01/16/2023 15:49:45",
      "content": "<p>There is a parameter called <em>group</em> that you use when you fit (model.fit()) your Ranker model. What the model tries to achieve is to sort your train examples inside each group assigning it some specific score. In this particular example, you are not classifying whether the candidate item is good or not, you are trying within each session to find items that likely belong to it. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2102597,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/16/2023 18:16:02",
          "content": "<p>I agree that the intuition behind the ranker is quiet easy to understand, but I'm looking into theoretical explanations.</p>\n<p>Basically, I noticed that those numbers can be interpreted as probabilities of a item being relevant (or being at the top) or not.<br>\nAlso, I figured out that the calibration of the model didn't get impacted when applying a sigmoid/softmax function on the prediction.</p>\n<p>Finally, the answer I found for my question is that the predictions can be interpreted as relevance scores, where high values represent an expected high relevance ect..</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2102238": "I have a question about the actual nature of the prediction given by the ranker. For the LGBMRanker I noticed that the prediction belongs to the R-domain, I was expecting a probability between 0 and 1.\n\n\nAny theoretical explanation about it ?",
    "2102428": "There is a parameter called *group* that you use when you fit (model.fit()) your Ranker model. What the model tries to achieve is to sort your train examples inside each group assigning it some specific score. In this particular example, you are not classifying whether the candidate item is good or not, you are trying within each session to find items that likely belong to it.",
    "2102597": "I agree that the intuition behind the ranker is quiet easy to understand, but I'm looking into theoretical explanations.\n\nBasically, I noticed that those numbers can be interpreted as probabilities of a item being relevant (or being at the top) or not.\nAlso, I figured out that the calibration of the model didn't get impacted when applying a sigmoid/softmax function on the prediction.\n\nFinally, the answer I found for my question is that the predictions can be interpreted as relevance scores, where high values represent an expected high relevance ect.."
  },
  "source": "meta"
}