{
  "id": 205737,
  "title": "Label smoothing and post-processing",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205737",
  "author_name": "",
  "post_date": "2020-12-21T16:21:04.223515200Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I tested label smoothing and some simple scaling post-processing here: <a href=\"https://www.kaggle.com/scaomath/riiid-is-there-a-magic-for-sakt\" target=\"_blank\">https://www.kaggle.com/scaomath/riiid-is-there-a-magic-for-sakt</a></p>\n<p>TL;DR: to save you some time for trying, label smoothing does not work. A simple post-processing does not work either.</p>\n<p>I tested locally using LGB's post-processed predictions as smoothed label as well, it does not help either.</p>",
  "messages": [
    {
      "id": "1121415",
      "postDate": "12/21/2020 16:21:04",
      "content": "<p>I tested label smoothing and some simple scaling post-processing here: <a href=\"https://www.kaggle.com/scaomath/riiid-is-there-a-magic-for-sakt\" target=\"_blank\">https://www.kaggle.com/scaomath/riiid-is-there-a-magic-for-sakt</a></p>\n<p>TL;DR: to save you some time for trying, label smoothing does not work. A simple post-processing does not work either.</p>\n<p>I tested locally using LGB's post-processed predictions as smoothed label as well, it does not help either.</p>",
      "rawMarkdown": "I tested label smoothing and some simple scaling post-processing here: https://www.kaggle.com/scaomath/riiid-is-there-a-magic-for-sakt\n\nTL;DR: to save you some time for trying, label smoothing does not work. A simple post-processing does not work either.\n\nI tested locally using LGB's post-processed predictions as smoothed label as well, it does not help either.",
      "votes": null
    },
    {
      "id": "1121485",
      "postDate": "12/21/2020 17:18:55",
      "content": "<p>I can also independently corroborate that label smoothing didn't work in my later models either. Out of the curious, I checked out that notebook you linked about post processing.</p>\n<blockquote>\n  <p>An unconfident prediction is closer to 0.5 than the two ends, the histogram looks like this:</p>\n  <p>Because AUC is the area, the threshold will move from 0 to 1 to check the ration of false positive and false negative, <strong>if your model has a lot of prediction having probability near 0.5, then the AUC cannot be very high</strong>.</p>\n  <p>One simple way, given that your model is somewhat accurate, is to rescale the output of the NN before giving to sigmoid:</p>\n  <p>final probability <code>estimate=1 / (1+exp(𝛽𝑧)</code></p>\n  <p>where 𝑧 is your NN output, and 𝛽 is the scaling.</p>\n</blockquote>\n<p>That doesn't really make any sense to me. AUC doesn't care if you have ALL predictions, e.g. between 0.4999 and 0.5001, as long as they're sorted appropriately. In general, I find that scaling outputs or weighting hurts in situations where AUC is the metric.</p>\n<p>Still interested in other post-processing ideas though, but figure they probably stem from deep EDA rather than straightforward adjustments like this.</p>",
      "rawMarkdown": "I can also independently corroborate that label smoothing didn't work in my later models either. Out of the curious, I checked out that notebook you linked about post processing.\n\n> An unconfident prediction is closer to 0.5 than the two ends, the histogram looks like this:\n>\n> Because AUC is the area, the threshold will move from 0 to 1 to check the ration of false positive and false negative, **if your model has a lot of prediction having probability near 0.5, then the AUC cannot be very high**.\n>\n>One simple way, given that your model is somewhat accurate, is to rescale the output of the NN before giving to sigmoid:\n>\n> final probability `estimate=1 / (1+exp(𝛽𝑧)`\n>\n> where 𝑧 is your NN output, and 𝛽 is the scaling.\n\nThat doesn't really make any sense to me. AUC doesn't care if you have ALL predictions, e.g. between 0.4999 and 0.5001, as long as they're sorted appropriately. In general, I find that scaling outputs or weighting hurts in situations where AUC is the metric.\n\nStill interested in other post-processing ideas though, but figure they probably stem from deep EDA rather than straightforward adjustments like this.",
      "votes": null
    },
    {
      "id": "1122330",
      "postDate": "12/22/2020 11:03:41",
      "content": "<p>Does your auc become better during training?</p>",
      "rawMarkdown": "Does your auc become better during training?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1121485,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/21/2020 17:18:55",
      "content": "<p>I can also independently corroborate that label smoothing didn't work in my later models either. Out of the curious, I checked out that notebook you linked about post processing.</p>\n<blockquote>\n  <p>An unconfident prediction is closer to 0.5 than the two ends, the histogram looks like this:</p>\n  <p>Because AUC is the area, the threshold will move from 0 to 1 to check the ration of false positive and false negative, <strong>if your model has a lot of prediction having probability near 0.5, then the AUC cannot be very high</strong>.</p>\n  <p>One simple way, given that your model is somewhat accurate, is to rescale the output of the NN before giving to sigmoid:</p>\n  <p>final probability <code>estimate=1 / (1+exp(𝛽𝑧)</code></p>\n  <p>where 𝑧 is your NN output, and 𝛽 is the scaling.</p>\n</blockquote>\n<p>That doesn't really make any sense to me. AUC doesn't care if you have ALL predictions, e.g. between 0.4999 and 0.5001, as long as they're sorted appropriately. In general, I find that scaling outputs or weighting hurts in situations where AUC is the metric.</p>\n<p>Still interested in other post-processing ideas though, but figure they probably stem from deep EDA rather than straightforward adjustments like this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122330,
      "author_name": "dwchen",
      "author_url": "",
      "post_date": "12/22/2020 11:03:41",
      "content": "<p>Does your auc become better during training?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1121415": "I tested label smoothing and some simple scaling post-processing here: https://www.kaggle.com/scaomath/riiid-is-there-a-magic-for-sakt\n\nTL;DR: to save you some time for trying, label smoothing does not work. A simple post-processing does not work either.\n\nI tested locally using LGB's post-processed predictions as smoothed label as well, it does not help either.",
    "1121485": "I can also independently corroborate that label smoothing didn't work in my later models either. Out of the curious, I checked out that notebook you linked about post processing.\n\n> An unconfident prediction is closer to 0.5 than the two ends, the histogram looks like this:\n>\n> Because AUC is the area, the threshold will move from 0 to 1 to check the ration of false positive and false negative, **if your model has a lot of prediction having probability near 0.5, then the AUC cannot be very high**.\n>\n>One simple way, given that your model is somewhat accurate, is to rescale the output of the NN before giving to sigmoid:\n>\n> final probability `estimate=1 / (1+exp(𝛽𝑧)`\n>\n> where 𝑧 is your NN output, and 𝛽 is the scaling.\n\nThat doesn't really make any sense to me. AUC doesn't care if you have ALL predictions, e.g. between 0.4999 and 0.5001, as long as they're sorted appropriately. In general, I find that scaling outputs or weighting hurts in situations where AUC is the metric.\n\nStill interested in other post-processing ideas though, but figure they probably stem from deep EDA rather than straightforward adjustments like this.",
    "1122330": "Does your auc become better during training?"
  },
  "source": "meta"
}