{
  "id": 362125,
  "title": "Calibration is all you need!",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/362125",
  "author_name": "",
  "post_date": "2022-10-25T17:40:38.487152900Z",
  "votes": 14,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi colleagues,</p>\n<p>In my previous post <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361813\" target=\"_blank\">\"LogLoss explanation\"</a> we discussed how the data sampling can kill all your efforts to get good LogLoss score. In this topic I would like to show how to fix this problem - how to do the <strong>calibration</strong> for your predictions to survive. I have created the example kernel <a href=\"https://www.kaggle.com/code/alexryzhkov/calibration-is-all-you-need\" target=\"_blank\">\"Calibration is all you need!\"</a> to show how you can do that on practice.</p>\n<p>To do the simple calibration you can use two simple methods, working with almost the same quality:</p>\n<ul>\n<li>Bruteforce the weights for the linear model like <code>A*p+B</code> (where <code>p</code> is the prediction) on prediction for the full validation set to get the best score</li>\n<li>Train <code>LogisticRegression</code> model on the full validation set prediction from sampled train and its target</li>\n</ul>\n<p>After the LogisticRegression model creation you will receive almost the same LogLoss score as on the usual train. But it's important to note that the <strong>LGBM model on the sampled data is trained more than 5x times faster in comparison with the full data model</strong> so you can save a plenty of time to do something valuable - create new features or use augmentations. </p>\n<p>If we take a look on the predictions histograms after validation, we can see that they are now good looking:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fedc27800b26786d2bf8c4e00232afbb3%2F2022-10-25%20%2020.37.47.png?generation=1666719489520585&amp;alt=media\" alt=\"\"></p>\n<p>If you still want to get the good LogLoss score and it's necessary to use sampling for some reason - <strong>don't forget to use calibration!</strong> It's all you need to fix the prediction distribution.</p>\n<p>Hope this helps you!</p>\n<p>Alex</p>",
  "messages": [
    {
      "id": "2003693",
      "postDate": "10/25/2022 17:40:38",
      "content": "<p>Hi colleagues,</p>\n<p>In my previous post <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361813\" target=\"_blank\">\"LogLoss explanation\"</a> we discussed how the data sampling can kill all your efforts to get good LogLoss score. In this topic I would like to show how to fix this problem - how to do the <strong>calibration</strong> for your predictions to survive. I have created the example kernel <a href=\"https://www.kaggle.com/code/alexryzhkov/calibration-is-all-you-need\" target=\"_blank\">\"Calibration is all you need!\"</a> to show how you can do that on practice.</p>\n<p>To do the simple calibration you can use two simple methods, working with almost the same quality:</p>\n<ul>\n<li>Bruteforce the weights for the linear model like <code>A*p+B</code> (where <code>p</code> is the prediction) on prediction for the full validation set to get the best score</li>\n<li>Train <code>LogisticRegression</code> model on the full validation set prediction from sampled train and its target</li>\n</ul>\n<p>After the LogisticRegression model creation you will receive almost the same LogLoss score as on the usual train. But it's important to note that the <strong>LGBM model on the sampled data is trained more than 5x times faster in comparison with the full data model</strong> so you can save a plenty of time to do something valuable - create new features or use augmentations. </p>\n<p>If we take a look on the predictions histograms after validation, we can see that they are now good looking:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fedc27800b26786d2bf8c4e00232afbb3%2F2022-10-25%20%2020.37.47.png?generation=1666719489520585&amp;alt=media\" alt=\"\"></p>\n<p>If you still want to get the good LogLoss score and it's necessary to use sampling for some reason - <strong>don't forget to use calibration!</strong> It's all you need to fix the prediction distribution.</p>\n<p>Hope this helps you!</p>\n<p>Alex</p>",
      "rawMarkdown": "Hi colleagues,\n\nIn my previous post [\"LogLoss explanation\"](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361813) we discussed how the data sampling can kill all your efforts to get good LogLoss score. In this topic I would like to show how to fix this problem - how to do the **calibration** for your predictions to survive. I have created the example kernel [\"Calibration is all you need!\"](https://www.kaggle.com/code/alexryzhkov/calibration-is-all-you-need) to show how you can do that on practice.\n\nTo do the simple calibration you can use two simple methods, working with almost the same quality:\n- Bruteforce the weights for the linear model like `A*p+B` (where `p` is the prediction) on prediction for the full validation set to get the best score\n- Train `LogisticRegression` model on the full validation set prediction from sampled train and its target\n\nAfter the LogisticRegression model creation you will receive almost the same LogLoss score as on the usual train. But it's important to note that the **LGBM model on the sampled data is trained more than 5x times faster in comparison with the full data model** so you can save a plenty of time to do something valuable - create new features or use augmentations. \n\nIf we take a look on the predictions histograms after validation, we can see that they are now good looking:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fedc27800b26786d2bf8c4e00232afbb3%2F2022-10-25%20%2020.37.47.png?generation=1666719489520585&alt=media)\n\nIf you still want to get the good LogLoss score and it's necessary to use sampling for some reason - **don't forget to use calibration!** It's all you need to fix the prediction distribution.\n\nHope this helps you!\n\nAlex",
      "votes": null
    },
    {
      "id": "2003832",
      "postDate": "10/25/2022 20:02:27",
      "content": "<p>Thanks for sharing, looks as a useful way to deal with an imbalanced datasets. One question though - shouldn't the parameters A and B be evaluated on the test set, rather than the validation set? After all, they are in principle part of the model.</p>",
      "rawMarkdown": "Thanks for sharing, looks as a useful way to deal with an imbalanced datasets. One question though - shouldn't the parameters A and B be evaluated on the test set, rather than the validation set? After all, they are in principle part of the model.",
      "votes": null
    },
    {
      "id": "2003880",
      "postDate": "10/25/2022 21:05:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lukamedis\" target=\"_blank\">@lukamedis</a>,<br>\nIn real life you do not have target for the test dataset to do that</p>",
      "rawMarkdown": "Hi @lukamedis,\nIn real life you do not have target for the test dataset to do that",
      "votes": null
    },
    {
      "id": "2013431",
      "postDate": "11/01/2022 22:38:23",
      "content": "<p>Thank you for sharing. Useful article.</p>",
      "rawMarkdown": "Thank you for sharing. Useful article.",
      "votes": null
    },
    {
      "id": "2459961",
      "postDate": "09/28/2023 13:49:08",
      "content": "<p>This method is as hoc and doesn’t have theoretical guarantees.</p>\n<p>For much better framework check conformal prediction in particular Venn-Abers.</p>\n<p><a href=\"https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718\" target=\"_blank\">https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718</a></p>\n<p><a href=\"https://github.com/valeman/awesome-conformal-prediction\" target=\"_blank\">https://github.com/valeman/awesome-conformal-prediction</a></p>",
      "rawMarkdown": "This method is as hoc and doesn’t have theoretical guarantees.\n\nFor much better framework check conformal prediction in particular Venn-Abers.\n\nhttps://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718\n\nhttps://github.com/valeman/awesome-conformal-prediction",
      "votes": null
    },
    {
      "id": "2508373",
      "postDate": "11/01/2023 16:31:15",
      "content": "<p>Your example uses old method Platt's scaling which does not guarantee good calibration and it fact can make it worse. Not only that it has been shown to be equivalent to very restrictive assumption of normality and homoscedasticity (see paper ' Beta calibration').</p>\n<p>Platt's scaler was developed specifically for SVM only and relied on sigmoid shape class scores. To use it for other models is a gamble.</p>",
      "rawMarkdown": "Your example uses old method Platt's scaling which does not guarantee good calibration and it fact can make it worse. Not only that it has been shown to be equivalent to very restrictive assumption of normality and homoscedasticity (see paper ' Beta calibration').\n\nPlatt's scaler was developed specifically for SVM only and relied on sigmoid shape class scores. To use it for other models is a gamble.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2003832,
      "author_name": "lukamedis",
      "author_url": "",
      "post_date": "10/25/2022 20:02:27",
      "content": "<p>Thanks for sharing, looks as a useful way to deal with an imbalanced datasets. One question though - shouldn't the parameters A and B be evaluated on the test set, rather than the validation set? After all, they are in principle part of the model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2003880,
          "author_name": "alexryzhkov",
          "author_url": "",
          "post_date": "10/25/2022 21:05:42",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lukamedis\" target=\"_blank\">@lukamedis</a>,<br>\nIn real life you do not have target for the test dataset to do that</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2013431,
      "author_name": "rafkhat",
      "author_url": "",
      "post_date": "11/01/2022 22:38:23",
      "content": "<p>Thank you for sharing. Useful article.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2459961,
      "author_name": "predaddict",
      "author_url": "",
      "post_date": "09/28/2023 13:49:08",
      "content": "<p>This method is as hoc and doesn’t have theoretical guarantees.</p>\n<p>For much better framework check conformal prediction in particular Venn-Abers.</p>\n<p><a href=\"https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718\" target=\"_blank\">https://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718</a></p>\n<p><a href=\"https://github.com/valeman/awesome-conformal-prediction\" target=\"_blank\">https://github.com/valeman/awesome-conformal-prediction</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2508373,
      "author_name": "predaddict",
      "author_url": "",
      "post_date": "11/01/2023 16:31:15",
      "content": "<p>Your example uses old method Platt's scaling which does not guarantee good calibration and it fact can make it worse. Not only that it has been shown to be equivalent to very restrictive assumption of normality and homoscedasticity (see paper ' Beta calibration').</p>\n<p>Platt's scaler was developed specifically for SVM only and relied on sigmoid shape class scores. To use it for other models is a gamble.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2003693": "Hi colleagues,\n\nIn my previous post [\"LogLoss explanation\"](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361813) we discussed how the data sampling can kill all your efforts to get good LogLoss score. In this topic I would like to show how to fix this problem - how to do the **calibration** for your predictions to survive. I have created the example kernel [\"Calibration is all you need!\"](https://www.kaggle.com/code/alexryzhkov/calibration-is-all-you-need) to show how you can do that on practice.\n\nTo do the simple calibration you can use two simple methods, working with almost the same quality:\n- Bruteforce the weights for the linear model like `A*p+B` (where `p` is the prediction) on prediction for the full validation set to get the best score\n- Train `LogisticRegression` model on the full validation set prediction from sampled train and its target\n\nAfter the LogisticRegression model creation you will receive almost the same LogLoss score as on the usual train. But it's important to note that the **LGBM model on the sampled data is trained more than 5x times faster in comparison with the full data model** so you can save a plenty of time to do something valuable - create new features or use augmentations. \n\nIf we take a look on the predictions histograms after validation, we can see that they are now good looking:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fedc27800b26786d2bf8c4e00232afbb3%2F2022-10-25%20%2020.37.47.png?generation=1666719489520585&alt=media)\n\nIf you still want to get the good LogLoss score and it's necessary to use sampling for some reason - **don't forget to use calibration!** It's all you need to fix the prediction distribution.\n\nHope this helps you!\n\nAlex",
    "2003832": "Thanks for sharing, looks as a useful way to deal with an imbalanced datasets. One question though - shouldn't the parameters A and B be evaluated on the test set, rather than the validation set? After all, they are in principle part of the model.",
    "2003880": "Hi @lukamedis,\nIn real life you do not have target for the test dataset to do that",
    "2013431": "Thank you for sharing. Useful article.",
    "2459961": "This method is as hoc and doesn’t have theoretical guarantees.\n\nFor much better framework check conformal prediction in particular Venn-Abers.\n\nhttps://medium.com/@valeman/how-to-calibrate-your-classifier-in-an-intelligent-way-a996a2faf718\n\nhttps://github.com/valeman/awesome-conformal-prediction",
    "2508373": "Your example uses old method Platt's scaling which does not guarantee good calibration and it fact can make it worse. Not only that it has been shown to be equivalent to very restrictive assumption of normality and homoscedasticity (see paper ' Beta calibration').\n\nPlatt's scaler was developed specifically for SVM only and relied on sigmoid shape class scores. To use it for other models is a gamble."
  },
  "source": "meta"
}