{
  "id": 329103,
  "title": "Ensembling probabilities with log-odds",
  "url": "/competitions/amex-default-prediction/discussion/329103",
  "author_name": "",
  "post_date": "2022-06-04T20:33:07.914785600Z",
  "votes": 50,
  "comment_count": 14,
  "views": 0,
  "content": "<p>It doesn't seems entirely natural to ensemble propabilities with mean as they are not usually additive (if a model say 0.1, the other 0.01 - is the expectation 0.055 ?). The notion of log-odds that usually appears in logistic regression seems more 'additive'. Without much proof, the general idea would be to average:</p>\n<p>$$ log odds =  log(p/(1-p)) $$</p>\n<ul>\n<li>It seems to provide slightly better ensemble for me (0.005 locally, 0.01 sometimes on LB)</li>\n<li>lgbm usually offers to predict the log_odds directly with the option raw_score = True</li>\n<li>Doing this (and more genrally ensembling) with NN would require checking the model are calibrated in probabilities</li>\n<li>Probability can be obtained with:<br>\n$$ proba= exp(log odds)/(1+exp(log odds)) $$</li>\n<li>The transformation back to probability being strictly monotonic isn't really needed for submission </li>\n</ul>\n<p>What do you think ? Is it usefull for your model ? Anyone with a source for this ?</p>",
  "messages": [
    {
      "id": "1811566",
      "postDate": "06/04/2022 20:33:07",
      "content": "<p>It doesn't seems entirely natural to ensemble propabilities with mean as they are not usually additive (if a model say 0.1, the other 0.01 - is the expectation 0.055 ?). The notion of log-odds that usually appears in logistic regression seems more 'additive'. Without much proof, the general idea would be to average:</p>\n<p>$$ log odds =  log(p/(1-p)) $$</p>\n<ul>\n<li>It seems to provide slightly better ensemble for me (0.005 locally, 0.01 sometimes on LB)</li>\n<li>lgbm usually offers to predict the log_odds directly with the option raw_score = True</li>\n<li>Doing this (and more genrally ensembling) with NN would require checking the model are calibrated in probabilities</li>\n<li>Probability can be obtained with:<br>\n$$ proba= exp(log odds)/(1+exp(log odds)) $$</li>\n<li>The transformation back to probability being strictly monotonic isn't really needed for submission </li>\n</ul>\n<p>What do you think ? Is it usefull for your model ? Anyone with a source for this ?</p>",
      "rawMarkdown": "It doesn't seems entirely natural to ensemble propabilities with mean as they are not usually additive (if a model say 0.1, the other 0.01 - is the expectation 0.055 ?). The notion of log-odds that usually appears in logistic regression seems more 'additive'. Without much proof, the general idea would be to average:\n\n$$ log odds =  log(p/(1-p)) $$\n\n- It seems to provide slightly better ensemble for me (0.005 locally, 0.01 sometimes on LB)\n- lgbm usually offers to predict the log_odds directly with the option raw_score = True\n- Doing this (and more genrally ensembling) with NN would require checking the model are calibrated in probabilities\n- Probability can be obtained with:\n$$ proba= exp(log odds)/(1+exp(log odds)) $$\n- The transformation back to probability being strictly monotonic isn't really needed for submission \n\nWhat do you think ? Is it usefull for your model ? Anyone with a source for this ?",
      "votes": null
    },
    {
      "id": "1811595",
      "postDate": "06/04/2022 21:31:20",
      "content": "<p>I like the idea. I'm confused by your formulas and notation. In your equation, is \"log - odds\" subtraction? or is that how you are writing log odds? Also I believe that the formula for log odds is <code>log_odds = ln( p/(1-p) )</code></p>",
      "rawMarkdown": "I like the idea. I'm confused by your formulas and notation. In your equation, is \"log - odds\" subtraction? or is that how you are writing log odds? Also I believe that the formula for log odds is `log_odds = ln( p/(1-p) )`",
      "votes": null
    },
    {
      "id": "1811602",
      "postDate": "06/04/2022 21:54:40",
      "content": "<p>I was writing log-odds, but Latex render the '-' as a minus; and yes you are right about the formula; my mistake.</p>",
      "rawMarkdown": "I was writing log-odds, but Latex render the '-' as a minus; and yes you are right about the formula; my mistake.",
      "votes": null
    },
    {
      "id": "1811919",
      "postDate": "06/05/2022 09:55:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> Thank you for publishing the idea! I tried the method with the five folds of my LightGBM model; the lb score seems to have improved, but by less than 0.001.</p>",
      "rawMarkdown": "Hi @lucasmorin Thank you for publishing the idea! I tried the method with the five folds of my LightGBM model; the lb score seems to have improved, but by less than 0.001.",
      "votes": null
    },
    {
      "id": "1812219",
      "postDate": "06/05/2022 15:33:48",
      "content": "<p>Great idea, thanks</p>",
      "rawMarkdown": "Great idea, thanks",
      "votes": null
    },
    {
      "id": "1812286",
      "postDate": "06/05/2022 16:28:37",
      "content": "<p>I think it is always a good idea. </p>\n<p>Consider 2 models give you probabilities 99% (very confident about positive label) and 50% (have no confidence positive or negative). Averaging those probas will give you 75% - which is positive but not confident. Sounds not fair -&gt; If you prefer to go to the bar over stay study at home and your friend does not care - you guys indeed will go to the bar  😃</p>\n<p>Logodds solution will give you something like 99% -&gt; 4.59, 50% -&gt; 0, so average will be 2.29 -&gt; 91%. Sounds more realistic 😃</p>",
      "rawMarkdown": "I think it is always a good idea. \n\nConsider 2 models give you probabilities 99% (very confident about positive label) and 50% (have no confidence positive or negative). Averaging those probas will give you 75% - which is positive but not confident. Sounds not fair -> If you prefer to go to the bar over stay study at home and your friend does not care - you guys indeed will go to the bar  😃\n\nLogodds solution will give you something like 99% -> 4.59, 50% -> 0, so average will be 2.29 -> 91%. Sounds more realistic 😃",
      "votes": null
    },
    {
      "id": "1812288",
      "postDate": "06/05/2022 16:30:16",
      "content": "<p>The log-odds transformation has asymptotic behavior near low or high probabilities. Averaging in this space would then behave in such a way that if one of the classifiers is very sure about its prediction (probability 0 or 1), it would dominate the prediction, i.e., other classifiers are ignored.</p>",
      "rawMarkdown": "The log-odds transformation has asymptotic behavior near low or high probabilities. Averaging in this space would then behave in such a way that if one of the classifiers is very sure about its prediction (probability 0 or 1), it would dominate the prediction, i.e., other classifiers are ignored.",
      "votes": null
    },
    {
      "id": "1812821",
      "postDate": "06/06/2022 08:53:09",
      "content": "<p>Thanks for the clarification… so it would 'overweight' extreme values ? I am not sure it is what we want.</p>",
      "rawMarkdown": "Thanks for the clarification... so it would 'overweight' extreme values ? I am not sure it is what we want.",
      "votes": null
    },
    {
      "id": "1813244",
      "postDate": "06/06/2022 16:38:03",
      "content": "<p><a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> I don't know what the bottom line is or what's good for this competition 😄 It is intuitively reasonable to choose the most \"confident\" prediction, or \\(\\mbox{argmax}_p\\max(p,1-p)\\). This is a hard switching. What you are proposing is sort of a soft argmax version.</p>",
      "rawMarkdown": "lucasmorin I don't know what the bottom line is or what's good for this competition 😄 It is intuitively reasonable to choose the most \"confident\" prediction, or \\\\(\\mbox{argmax}_p\\max(p,1-p)\\\\). This is a hard switching. What you are proposing is sort of a soft argmax version.",
      "votes": null
    },
    {
      "id": "1813292",
      "postDate": "06/06/2022 17:14:47",
      "content": "<p>Am I correct in assuming that for your XGB starter, simply switching \"binary:logistic\" to \"binary:logitraw\" should produce output already transformed in exactly this way?</p>\n<p>(More efficiently, since the name implies that's the format it started in. )</p>",
      "rawMarkdown": "Am I correct in assuming that for your XGB starter, simply switching \"binary:logistic\" to \"binary:logitraw\" should produce output already transformed in exactly this way?\n\n(More efficiently, since the name implies that's the format it started in. )",
      "votes": null
    },
    {
      "id": "1813301",
      "postDate": "06/06/2022 17:22:19",
      "content": "<p>Yes, <code>binary:logitraw = log_odds = ln( \"binary:logistic\"/(1-\"binary:logistic\") )</code></p>\n<p>However if we set XGB to this learning objective i'm not sure how it affects training. I have never used it before. i.e. Does it train the same as <code>binary:logistic</code> but then just give a different predict output? Or is the loss and training different? I'm not sure.</p>",
      "rawMarkdown": "Yes, `binary:logitraw = log_odds = ln( \"binary:logistic\"/(1-\"binary:logistic\") )`\n\nHowever if we set XGB to this learning objective i'm not sure how it affects training. I have never used it before. i.e. Does it train the same as `binary:logistic` but then just give a different predict output? Or is the loss and training different? I'm not sure.",
      "votes": null
    },
    {
      "id": "1813412",
      "postDate": "06/06/2022 19:42:20",
      "content": "<p>Inspecting the source code suggests that <code>reg:logistic</code>, <code>binary:logistic</code>, <code>binary:logitraw</code> all compute the same loss function. </p>\n<p><a href=\"https://github.com/dmlc/xgboost/blob/1a33b50a0de15ebc3147710a92877007f0d27e79/src/objective/regression_loss.h#L116\" target=\"_blank\">https://github.com/dmlc/xgboost/blob/1a33b50a0de15ebc3147710a92877007f0d27e79/src/objective/regression_loss.h#L116</a></p>",
      "rawMarkdown": "Inspecting the source code suggests that `reg:logistic`, `binary:logistic`, `binary:logitraw` all compute the same loss function. \n\n[https://github.com/dmlc/xgboost/blob/1a33b50a0de15ebc3147710a92877007f0d27e79/src/objective/regression_loss.h#L116](https://github.com/dmlc/xgboost/blob/1a33b50a0de15ebc3147710a92877007f0d27e79/src/objective/regression_loss.h#L116)",
      "votes": null
    },
    {
      "id": "1813542",
      "postDate": "06/07/2022 01:30:29",
      "content": "<p>Nice explanation. Straightforward and easy to understand</p>",
      "rawMarkdown": "Nice explanation. Straightforward and easy to understand",
      "votes": null
    },
    {
      "id": "1817097",
      "postDate": "06/10/2022 22:35:33",
      "content": "<p>I like the idea. However, by applying this log_odds average did not help with my score in LB. Did I miss something else?  </p>",
      "rawMarkdown": "I like the idea. However, by applying this log_odds average did not help with my score in LB. Did I miss something else?",
      "votes": null
    },
    {
      "id": "1863785",
      "postDate": "07/20/2022 14:02:13",
      "content": "<p>Think this is but logit can even use <code>scipy.special.logit( m_preds.prediction.values)</code> if I am not wrong .</p>",
      "rawMarkdown": "Think this is but logit can even use `scipy.special.logit( m_preds.prediction.values)` if I am not wrong .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1811595,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/04/2022 21:31:20",
      "content": "<p>I like the idea. I'm confused by your formulas and notation. In your equation, is \"log - odds\" subtraction? or is that how you are writing log odds? Also I believe that the formula for log odds is <code>log_odds = ln( p/(1-p) )</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1811602,
          "author_name": "lucasmorin",
          "author_url": "",
          "post_date": "06/04/2022 21:54:40",
          "content": "<p>I was writing log-odds, but Latex render the '-' as a minus; and yes you are right about the formula; my mistake.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1813292,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "06/06/2022 17:14:47",
          "content": "<p>Am I correct in assuming that for your XGB starter, simply switching \"binary:logistic\" to \"binary:logitraw\" should produce output already transformed in exactly this way?</p>\n<p>(More efficiently, since the name implies that's the format it started in. )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1813301,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/06/2022 17:22:19",
          "content": "<p>Yes, <code>binary:logitraw = log_odds = ln( \"binary:logistic\"/(1-\"binary:logistic\") )</code></p>\n<p>However if we set XGB to this learning objective i'm not sure how it affects training. I have never used it before. i.e. Does it train the same as <code>binary:logistic</code> but then just give a different predict output? Or is the loss and training different? I'm not sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1813412,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "06/06/2022 19:42:20",
          "content": "<p>Inspecting the source code suggests that <code>reg:logistic</code>, <code>binary:logistic</code>, <code>binary:logitraw</code> all compute the same loss function. </p>\n<p><a href=\"https://github.com/dmlc/xgboost/blob/1a33b50a0de15ebc3147710a92877007f0d27e79/src/objective/regression_loss.h#L116\" target=\"_blank\">https://github.com/dmlc/xgboost/blob/1a33b50a0de15ebc3147710a92877007f0d27e79/src/objective/regression_loss.h#L116</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1811919,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "06/05/2022 09:55:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> Thank you for publishing the idea! I tried the method with the five folds of my LightGBM model; the lb score seems to have improved, but by less than 0.001.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1812219,
      "author_name": "samanemami",
      "author_url": "",
      "post_date": "06/05/2022 15:33:48",
      "content": "<p>Great idea, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1812286,
      "author_name": "pavelvod",
      "author_url": "",
      "post_date": "06/05/2022 16:28:37",
      "content": "<p>I think it is always a good idea. </p>\n<p>Consider 2 models give you probabilities 99% (very confident about positive label) and 50% (have no confidence positive or negative). Averaging those probas will give you 75% - which is positive but not confident. Sounds not fair -&gt; If you prefer to go to the bar over stay study at home and your friend does not care - you guys indeed will go to the bar  😃</p>\n<p>Logodds solution will give you something like 99% -&gt; 4.59, 50% -&gt; 0, so average will be 2.29 -&gt; 91%. Sounds more realistic 😃</p>",
      "votes": null,
      "replies": [
        {
          "id": 1813542,
          "author_name": "johnnywangzy",
          "author_url": "",
          "post_date": "06/07/2022 01:30:29",
          "content": "<p>Nice explanation. Straightforward and easy to understand</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1812288,
      "author_name": "siukeitin",
      "author_url": "",
      "post_date": "06/05/2022 16:30:16",
      "content": "<p>The log-odds transformation has asymptotic behavior near low or high probabilities. Averaging in this space would then behave in such a way that if one of the classifiers is very sure about its prediction (probability 0 or 1), it would dominate the prediction, i.e., other classifiers are ignored.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1812821,
          "author_name": "lucasmorin",
          "author_url": "",
          "post_date": "06/06/2022 08:53:09",
          "content": "<p>Thanks for the clarification… so it would 'overweight' extreme values ? I am not sure it is what we want.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1813244,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "06/06/2022 16:38:03",
          "content": "<p><a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> I don't know what the bottom line is or what's good for this competition 😄 It is intuitively reasonable to choose the most \"confident\" prediction, or \\(\\mbox{argmax}_p\\max(p,1-p)\\). This is a hard switching. What you are proposing is sort of a soft argmax version.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1817097,
      "author_name": "liaison",
      "author_url": "",
      "post_date": "06/10/2022 22:35:33",
      "content": "<p>I like the idea. However, by applying this log_odds average did not help with my score in LB. Did I miss something else?  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1863785,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "07/20/2022 14:02:13",
      "content": "<p>Think this is but logit can even use <code>scipy.special.logit( m_preds.prediction.values)</code> if I am not wrong .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1811566": "It doesn't seems entirely natural to ensemble propabilities with mean as they are not usually additive (if a model say 0.1, the other 0.01 - is the expectation 0.055 ?). The notion of log-odds that usually appears in logistic regression seems more 'additive'. Without much proof, the general idea would be to average:\n\n$$ log odds =  log(p/(1-p)) $$\n\n- It seems to provide slightly better ensemble for me (0.005 locally, 0.01 sometimes on LB)\n- lgbm usually offers to predict the log_odds directly with the option raw_score = True\n- Doing this (and more genrally ensembling) with NN would require checking the model are calibrated in probabilities\n- Probability can be obtained with:\n$$ proba= exp(log odds)/(1+exp(log odds)) $$\n- The transformation back to probability being strictly monotonic isn't really needed for submission \n\nWhat do you think ? Is it usefull for your model ? Anyone with a source for this ?",
    "1811595": "I like the idea. I'm confused by your formulas and notation. In your equation, is \"log - odds\" subtraction? or is that how you are writing log odds? Also I believe that the formula for log odds is `log_odds = ln( p/(1-p) )`",
    "1811602": "I was writing log-odds, but Latex render the '-' as a minus; and yes you are right about the formula; my mistake.",
    "1811919": "Hi @lucasmorin Thank you for publishing the idea! I tried the method with the five folds of my LightGBM model; the lb score seems to have improved, but by less than 0.001.",
    "1812219": "Great idea, thanks",
    "1812286": "I think it is always a good idea. \n\nConsider 2 models give you probabilities 99% (very confident about positive label) and 50% (have no confidence positive or negative). Averaging those probas will give you 75% - which is positive but not confident. Sounds not fair -> If you prefer to go to the bar over stay study at home and your friend does not care - you guys indeed will go to the bar  😃\n\nLogodds solution will give you something like 99% -> 4.59, 50% -> 0, so average will be 2.29 -> 91%. Sounds more realistic 😃",
    "1812288": "The log-odds transformation has asymptotic behavior near low or high probabilities. Averaging in this space would then behave in such a way that if one of the classifiers is very sure about its prediction (probability 0 or 1), it would dominate the prediction, i.e., other classifiers are ignored.",
    "1812821": "Thanks for the clarification... so it would 'overweight' extreme values ? I am not sure it is what we want.",
    "1813244": "lucasmorin I don't know what the bottom line is or what's good for this competition 😄 It is intuitively reasonable to choose the most \"confident\" prediction, or \\\\(\\mbox{argmax}_p\\max(p,1-p)\\\\). This is a hard switching. What you are proposing is sort of a soft argmax version.",
    "1813292": "Am I correct in assuming that for your XGB starter, simply switching \"binary:logistic\" to \"binary:logitraw\" should produce output already transformed in exactly this way?\n\n(More efficiently, since the name implies that's the format it started in. )",
    "1813301": "Yes, `binary:logitraw = log_odds = ln( \"binary:logistic\"/(1-\"binary:logistic\") )`\n\nHowever if we set XGB to this learning objective i'm not sure how it affects training. I have never used it before. i.e. Does it train the same as `binary:logistic` but then just give a different predict output? Or is the loss and training different? I'm not sure.",
    "1813412": "Inspecting the source code suggests that `reg:logistic`, `binary:logistic`, `binary:logitraw` all compute the same loss function. \n\n[https://github.com/dmlc/xgboost/blob/1a33b50a0de15ebc3147710a92877007f0d27e79/src/objective/regression_loss.h#L116](https://github.com/dmlc/xgboost/blob/1a33b50a0de15ebc3147710a92877007f0d27e79/src/objective/regression_loss.h#L116)",
    "1813542": "Nice explanation. Straightforward and easy to understand",
    "1817097": "I like the idea. However, by applying this log_odds average did not help with my score in LB. Did I miss something else?",
    "1863785": "Think this is but logit can even use `scipy.special.logit( m_preds.prediction.values)` if I am not wrong ."
  },
  "source": "meta"
}