{
  "id": 188060,
  "title": "Confidence value of prediction",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/188060",
  "author_name": "",
  "post_date": "2020-10-01T13:34:45.977347Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Suppose I have a DNN which predict the FVC value for given input. How can I calculate the confidence value associated with this prediction. I want to know more about calculating confidence in prediction. How it is calculated for regression, classification problem?</p>",
  "messages": [
    {
      "id": "1034086",
      "postDate": "10/01/2020 13:34:45",
      "content": "<p>Suppose I have a DNN which predict the FVC value for given input. How can I calculate the confidence value associated with this prediction. I want to know more about calculating confidence in prediction. How it is calculated for regression, classification problem?</p>",
      "rawMarkdown": "Suppose I have a DNN which predict the FVC value for given input. How can I calculate the confidence value associated with this prediction. I want to know more about calculating confidence in prediction. How it is calculated for regression, classification problem?",
      "votes": null
    },
    {
      "id": "1034100",
      "postDate": "10/01/2020 13:58:08",
      "content": "<p>With a DNN one popular option is to have further outputs from the model such as 2 or 3 quantiles (e.g. 20th and 80th, or 25th, 50th or 75th) of the distribution chosen in such a manner that the difference between them is a good value for confidence (which can be further improved by adding the competition loss to the quantile/pinball loss function see e.g. <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">this notebook</a> - it's much <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/181505\" target=\"_blank\">harder to use the competition loss on its own</a> without the quantile loss - some implementation details like how to enforce that the quantiles are always correctly ordered is <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/182737#1009826\" target=\"_blank\">discussed here</a>).</p>\n<p>Additionally, with a DNN you could try out leaving drop-out (assuming you are using it) on during inference and to get repeated (randomly varying) predictions and see how helpful that is (see various publication by Yarin Gal such as <a href=\"http://mlg.eng.cam.ac.uk/yarin/thesis/thesis.pdf\" target=\"_blank\">his thesis</a>) - however, with that approach I'm still not 100% clear on whether one ends up getting a distribution for the expected average value (more of a credible or confidence interval for a mean parameter) or a distribution that also additionally includes variation in outcomes (i.e. more of a prediction interval). I'm curious whether someone else has experience with that approach.</p>\n<p>Other approaches that are not all applicable with DNN are e.g. <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/181632#1004612\" target=\"_blank\">discussed here</a>. </p>",
      "rawMarkdown": "With a DNN one popular option is to have further outputs from the model such as 2 or 3 quantiles (e.g. 20th and 80th, or 25th, 50th or 75th) of the distribution chosen in such a manner that the difference between them is a good value for confidence (which can be further improved by adding the competition loss to the quantile/pinball loss function see e.g. [this notebook](https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter) - it's much [harder to use the competition loss on its own](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/181505) without the quantile loss - some implementation details like how to enforce that the quantiles are always correctly ordered is [discussed here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/182737#1009826)).\n\nAdditionally, with a DNN you could try out leaving drop-out (assuming you are using it) on during inference and to get repeated (randomly varying) predictions and see how helpful that is (see various publication by Yarin Gal such as [his thesis](http://mlg.eng.cam.ac.uk/yarin/thesis/thesis.pdf)) - however, with that approach I'm still not 100% clear on whether one ends up getting a distribution for the expected average value (more of a credible or confidence interval for a mean parameter) or a distribution that also additionally includes variation in outcomes (i.e. more of a prediction interval). I'm curious whether someone else has experience with that approach.\n\nOther approaches that are not all applicable with DNN are e.g. [discussed here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/181632#1004612).",
      "votes": null
    },
    {
      "id": "1035556",
      "postDate": "10/02/2020 19:41:20",
      "content": "<p>With the dropout as uncertainty approximation you obviously need to do some tuning of the dropout %.</p>\n<p>I've got good (comparable to quantile regression) results using both dropout and bahesian NN</p>",
      "rawMarkdown": "With the dropout as uncertainty approximation you obviously need to do some tuning of the dropout %.\n\nI've got good (comparable to quantile regression) results using both dropout and bahesian NN",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1034100,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "10/01/2020 13:58:08",
      "content": "<p>With a DNN one popular option is to have further outputs from the model such as 2 or 3 quantiles (e.g. 20th and 80th, or 25th, 50th or 75th) of the distribution chosen in such a manner that the difference between them is a good value for confidence (which can be further improved by adding the competition loss to the quantile/pinball loss function see e.g. <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">this notebook</a> - it's much <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/181505\" target=\"_blank\">harder to use the competition loss on its own</a> without the quantile loss - some implementation details like how to enforce that the quantiles are always correctly ordered is <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/182737#1009826\" target=\"_blank\">discussed here</a>).</p>\n<p>Additionally, with a DNN you could try out leaving drop-out (assuming you are using it) on during inference and to get repeated (randomly varying) predictions and see how helpful that is (see various publication by Yarin Gal such as <a href=\"http://mlg.eng.cam.ac.uk/yarin/thesis/thesis.pdf\" target=\"_blank\">his thesis</a>) - however, with that approach I'm still not 100% clear on whether one ends up getting a distribution for the expected average value (more of a credible or confidence interval for a mean parameter) or a distribution that also additionally includes variation in outcomes (i.e. more of a prediction interval). I'm curious whether someone else has experience with that approach.</p>\n<p>Other approaches that are not all applicable with DNN are e.g. <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/181632#1004612\" target=\"_blank\">discussed here</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1035556,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "10/02/2020 19:41:20",
          "content": "<p>With the dropout as uncertainty approximation you obviously need to do some tuning of the dropout %.</p>\n<p>I've got good (comparable to quantile regression) results using both dropout and bahesian NN</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1034086": "Suppose I have a DNN which predict the FVC value for given input. How can I calculate the confidence value associated with this prediction. I want to know more about calculating confidence in prediction. How it is calculated for regression, classification problem?",
    "1034100": "With a DNN one popular option is to have further outputs from the model such as 2 or 3 quantiles (e.g. 20th and 80th, or 25th, 50th or 75th) of the distribution chosen in such a manner that the difference between them is a good value for confidence (which can be further improved by adding the competition loss to the quantile/pinball loss function see e.g. [this notebook](https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter) - it's much [harder to use the competition loss on its own](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/181505) without the quantile loss - some implementation details like how to enforce that the quantiles are always correctly ordered is [discussed here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/182737#1009826)).\n\nAdditionally, with a DNN you could try out leaving drop-out (assuming you are using it) on during inference and to get repeated (randomly varying) predictions and see how helpful that is (see various publication by Yarin Gal such as [his thesis](http://mlg.eng.cam.ac.uk/yarin/thesis/thesis.pdf)) - however, with that approach I'm still not 100% clear on whether one ends up getting a distribution for the expected average value (more of a credible or confidence interval for a mean parameter) or a distribution that also additionally includes variation in outcomes (i.e. more of a prediction interval). I'm curious whether someone else has experience with that approach.\n\nOther approaches that are not all applicable with DNN are e.g. [discussed here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/181632#1004612).",
    "1035556": "With the dropout as uncertainty approximation you obviously need to do some tuning of the dropout %.\n\nI've got good (comparable to quantile regression) results using both dropout and bahesian NN"
  },
  "source": "meta"
}