{
  "id": 54092,
  "title": "Probability calibration",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54092",
  "author_name": "",
  "post_date": "2018-04-09T14:26:49.870819300Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I noticed that in this competition neural networks return predictions that are poorly calibrated. Normally they are very reliable in this regard, but here they clearly lag behind other classifiers.</p>\n\n<p>Since AUC scoring cares only about the order of predictions, this will not affect individual models. However, it will likely make blending/stacking unreliable as the scale of NN predictions will be slightly off. I suggest either to re-calibrate predictions, or ensemble on ranked predictions instead.</p>",
  "messages": [
    {
      "id": "311167",
      "postDate": "04/09/2018 14:26:49",
      "content": "<p>I noticed that in this competition neural networks return predictions that are poorly calibrated. Normally they are very reliable in this regard, but here they clearly lag behind other classifiers.</p>\n\n<p>Since AUC scoring cares only about the order of predictions, this will not affect individual models. However, it will likely make blending/stacking unreliable as the scale of NN predictions will be slightly off. I suggest either to re-calibrate predictions, or ensemble on ranked predictions instead.</p>",
      "rawMarkdown": "I noticed that in this competition neural networks return predictions that are poorly calibrated. Normally they are very reliable in this regard, but here they clearly lag behind other classifiers.\n\nSince AUC scoring cares only about the order of predictions, this will not affect individual models. However, it will likely make blending/stacking unreliable as the scale of NN predictions will be slightly off. I suggest either to re-calibrate predictions, or ensemble on ranked predictions instead.",
      "votes": null
    },
    {
      "id": "311169",
      "postDate": "04/09/2018 14:29:25",
      "content": "<p>FWIW I've seen decent improvement with ranked predictions. </p>",
      "rawMarkdown": "FWIW I've seen decent improvement with ranked predictions.",
      "votes": null
    },
    {
      "id": "311178",
      "postDate": "04/09/2018 14:49:12",
      "content": "<p>Is there a standard simple methodology for combining ranks? My first inclination is to do an OLS of actual on predicted ranks, but that's imposing linearity on a system that could be quite nonlinear.</p>",
      "rawMarkdown": "Is there a standard simple methodology for combining ranks? My first inclination is to do an OLS of actual on predicted ranks, but that's imposing linearity on a system that could be quite nonlinear.",
      "votes": null
    },
    {
      "id": "311184",
      "postDate": "04/09/2018 14:58:15",
      "content": "<p>I can post a small R kernel if you'd like - not exactly scientific, just an approach that has served me rather well over the years :-)</p>",
      "rawMarkdown": "I can post a small R kernel if you'd like - not exactly scientific, just an approach that has served me rather well over the years :-)",
      "votes": null
    },
    {
      "id": "311188",
      "postDate": "04/09/2018 15:08:20",
      "content": "<p>@Andy: ok, here you go</p>\n\n<p><a href=\"https://www.kaggle.com/konradb/blended-submission-with-rank-normalization\">https://www.kaggle.com/konradb/blended-submission-with-rank-normalization</a></p>",
      "rawMarkdown": "Andy: ok, here you go\n\nhttps://www.kaggle.com/konradb/blended-submission-with-rank-normalization",
      "votes": null
    },
    {
      "id": "311196",
      "postDate": "04/09/2018 15:32:58",
      "content": "<blockquote>\n  <p>Is there a standard simple methodology for combining ranks? My first inclination is to do an OLS of actual on predicted ranks, but that's imposing linearity on a system that could be quite nonlinear.</p>\n</blockquote>\n\n<p>This is as close to a standard methodology as I know of: <a href=\"http://mlwave.com/kaggle-ensembling-guide/\"><strong>blog</strong></a> and <a href=\"https://github.com/MLWave/Kaggle-Ensemble-Guide\"><strong>code</strong></a>.</p>\n\n<p>Not sure that I understand the nonlinearity concern in this context. AUC scoring cares only about relative order of predictions. The way I see it the goal is to bring all models on to a comparable scale, which can be done by ranking predictions and dividing them by N(max).</p>\n\n<p>If you wish to do scaling without ranking, then OLS would be a bad choice. In that case I suggest <a href=\"https://en.wikipedia.org/wiki/Platt_scaling\"><strong>Platt scaling</strong></a>, but that requires out-of-fold predictions.</p>\n\n<p>For neural networks in this competition, Platt scaling gives quite an improvement:</p>\n\n<pre><code>Calculating log-loss before calibration...\n0.0111558038049\nCalculating AUC before calibration...\n0.974368305078\n\nCalculating log-loss after calibration...\n0.00699372243308\nCalculating AUC after calibration...\n0.974368305078\n</code></pre>\n\n<p>AUC score doesn't change as one would expect.</p>",
      "rawMarkdown": "&gt; Is there a standard simple methodology for combining ranks? My first inclination is to do an OLS of actual on predicted ranks, but that's imposing linearity on a system that could be quite nonlinear.\n\nThis is as close to a standard methodology as I know of: [__blog__](http://mlwave.com/kaggle-ensembling-guide/) and [__code__](https://github.com/MLWave/Kaggle-Ensemble-Guide).\n\nNot sure that I understand the nonlinearity concern in this context. AUC scoring cares only about relative order of predictions. The way I see it the goal is to bring all models on to a comparable scale, which can be done by ranking predictions and dividing them by N(max).\n\nIf you wish to do scaling without ranking, then OLS would be a bad choice. In that case I suggest [__Platt scaling__](https://en.wikipedia.org/wiki/Platt_scaling), but that requires out-of-fold predictions.\n\nFor neural networks in this competition, Platt scaling gives quite an improvement:\n\n    Calculating log-loss before calibration...\n    0.0111558038049\n    Calculating AUC before calibration...\n    0.974368305078\n    \n    Calculating log-loss after calibration...\n    0.00699372243308\n    Calculating AUC after calibration...\n    0.974368305078\n\nAUC score doesn't change as one would expect.",
      "votes": null
    },
    {
      "id": "311920",
      "postDate": "04/10/2018 22:52:41",
      "content": "<p>I guess @andy concerns the distribution of predicted items in the space of predicted scores?</p>",
      "rawMarkdown": "I guess @andy concerns the distribution of predicted items in the space of predicted scores?",
      "votes": null
    },
    {
      "id": "2520377",
      "postDate": "11/10/2023 19:33:20",
      "content": "<p>In 2023 can easily calibrate any model using conformal prediction <a href=\"https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718\" target=\"_blank\">https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718</a></p>\n<p><a href=\"https://github.com/valeman/awesome-conformal-prediction\" target=\"_blank\">https://github.com/valeman/awesome-conformal-prediction</a></p>",
      "rawMarkdown": "In 2023 can easily calibrate any model using conformal prediction https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718\n\nhttps://github.com/valeman/awesome-conformal-prediction",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2520377,
      "author_name": "predaddict",
      "author_url": "",
      "post_date": "11/10/2023 19:33:20",
      "content": "<p>In 2023 can easily calibrate any model using conformal prediction <a href=\"https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718\" target=\"_blank\">https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718</a></p>\n<p><a href=\"https://github.com/valeman/awesome-conformal-prediction\" target=\"_blank\">https://github.com/valeman/awesome-conformal-prediction</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 311169,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "04/09/2018 14:29:25",
      "content": "<p>FWIW I've seen decent improvement with ranked predictions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 311178,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "04/09/2018 14:49:12",
      "content": "<p>Is there a standard simple methodology for combining ranks? My first inclination is to do an OLS of actual on predicted ranks, but that's imposing linearity on a system that could be quite nonlinear.</p>",
      "votes": null,
      "replies": [
        {
          "id": 311184,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "04/09/2018 14:58:15",
          "content": "<p>I can post a small R kernel if you'd like - not exactly scientific, just an approach that has served me rather well over the years :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311188,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "04/09/2018 15:08:20",
          "content": "<p>@Andy: ok, here you go</p>\n\n<p><a href=\"https://www.kaggle.com/konradb/blended-submission-with-rank-normalization\">https://www.kaggle.com/konradb/blended-submission-with-rank-normalization</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311196,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/09/2018 15:32:58",
          "content": "<blockquote>\n  <p>Is there a standard simple methodology for combining ranks? My first inclination is to do an OLS of actual on predicted ranks, but that's imposing linearity on a system that could be quite nonlinear.</p>\n</blockquote>\n\n<p>This is as close to a standard methodology as I know of: <a href=\"http://mlwave.com/kaggle-ensembling-guide/\"><strong>blog</strong></a> and <a href=\"https://github.com/MLWave/Kaggle-Ensemble-Guide\"><strong>code</strong></a>.</p>\n\n<p>Not sure that I understand the nonlinearity concern in this context. AUC scoring cares only about relative order of predictions. The way I see it the goal is to bring all models on to a comparable scale, which can be done by ranking predictions and dividing them by N(max).</p>\n\n<p>If you wish to do scaling without ranking, then OLS would be a bad choice. In that case I suggest <a href=\"https://en.wikipedia.org/wiki/Platt_scaling\"><strong>Platt scaling</strong></a>, but that requires out-of-fold predictions.</p>\n\n<p>For neural networks in this competition, Platt scaling gives quite an improvement:</p>\n\n<pre><code>Calculating log-loss before calibration...\n0.0111558038049\nCalculating AUC before calibration...\n0.974368305078\n\nCalculating log-loss after calibration...\n0.00699372243308\nCalculating AUC after calibration...\n0.974368305078\n</code></pre>\n\n<p>AUC score doesn't change as one would expect.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311920,
          "author_name": "enfeizhan",
          "author_url": "",
          "post_date": "04/10/2018 22:52:41",
          "content": "<p>I guess @andy concerns the distribution of predicted items in the space of predicted scores?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "311167": "I noticed that in this competition neural networks return predictions that are poorly calibrated. Normally they are very reliable in this regard, but here they clearly lag behind other classifiers.\n\nSince AUC scoring cares only about the order of predictions, this will not affect individual models. However, it will likely make blending/stacking unreliable as the scale of NN predictions will be slightly off. I suggest either to re-calibrate predictions, or ensemble on ranked predictions instead.",
    "311169": "FWIW I've seen decent improvement with ranked predictions.",
    "311178": "Is there a standard simple methodology for combining ranks? My first inclination is to do an OLS of actual on predicted ranks, but that's imposing linearity on a system that could be quite nonlinear.",
    "311184": "I can post a small R kernel if you'd like - not exactly scientific, just an approach that has served me rather well over the years :-)",
    "311188": "Andy: ok, here you go\n\nhttps://www.kaggle.com/konradb/blended-submission-with-rank-normalization",
    "311196": "&gt; Is there a standard simple methodology for combining ranks? My first inclination is to do an OLS of actual on predicted ranks, but that's imposing linearity on a system that could be quite nonlinear.\n\nThis is as close to a standard methodology as I know of: [__blog__](http://mlwave.com/kaggle-ensembling-guide/) and [__code__](https://github.com/MLWave/Kaggle-Ensemble-Guide).\n\nNot sure that I understand the nonlinearity concern in this context. AUC scoring cares only about relative order of predictions. The way I see it the goal is to bring all models on to a comparable scale, which can be done by ranking predictions and dividing them by N(max).\n\nIf you wish to do scaling without ranking, then OLS would be a bad choice. In that case I suggest [__Platt scaling__](https://en.wikipedia.org/wiki/Platt_scaling), but that requires out-of-fold predictions.\n\nFor neural networks in this competition, Platt scaling gives quite an improvement:\n\n    Calculating log-loss before calibration...\n    0.0111558038049\n    Calculating AUC before calibration...\n    0.974368305078\n    \n    Calculating log-loss after calibration...\n    0.00699372243308\n    Calculating AUC after calibration...\n    0.974368305078\n\nAUC score doesn't change as one would expect.",
    "311920": "I guess @andy concerns the distribution of predicted items in the space of predicted scores?",
    "2520377": "In 2023 can easily calibrate any model using conformal prediction https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718\n\nhttps://github.com/valeman/awesome-conformal-prediction"
  },
  "source": "meta"
}