{
  "id": 173636,
  "title": "How to calculate the Confidence",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/173636",
  "author_name": "AjayKumar",
  "post_date": "2020-08-10T05:15:48.669000",
  "votes": 16,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I am not sure how the confidence value should be calculated.</p>",
  "messages": [
    {
      "id": 964690,
      "postDate": "2020-08-10T05:15:48.670Z",
      "content": "<p>I am not sure how the confidence value should be calculated.</p>",
      "rawMarkdown": "I am not sure how the confidence value should be calculated.",
      "votes": 16
    },
    {
      "id": 965663,
      "postDate": "2020-08-10T18:49:29.817Z",
      "content": "<p>You don't need the pinball loss/quantile regression to estimate the confidence. The confidence is just that - the ammount of confidence your model has in the given prediction.</p>\n\n<p>To illustrate why the quantile regression is useful (estimating that X% of the data is below a given value). Consider a normal distribution. 68% of the data lies between +/- 1 standard deviation of the mean. So if you want to estimate standard deviation you'd like to know the quantile at 50-34=16% and 50+34=84%. Then your 1 standard deviation estimate is just the difference between these divided by 2. </p>\n\n<p>It's not the only way but it's got a good grounding in theory and has performed well in kaggle competitions in the past. You might for example get the predictions from lots of models and take the standard deviation of their predictions or you might try to predict the standard deviation for each point from the data.</p>\n\n<p>When you see notebooks calculate the optimal value for confidence on the train data they're just solving a set of equations to maximise the objective function. It might be a good exercise for you to do that for a given prediction where you know the true value but it doesn't help in the test data. When you're predicting you don't have this option because you don't KNOW the true value so you can't solve the equation.</p>",
      "rawMarkdown": "You don't need the pinball loss/quantile regression to estimate the confidence. The confidence is just that - the ammount of confidence your model has in the given prediction.\n\nTo illustrate why the quantile regression is useful (estimating that X% of the data is below a given value). Consider a normal distribution. 68% of the data lies between +/- 1 standard deviation of the mean. So if you want to estimate standard deviation you'd like to know the quantile at 50-34=16% and 50+34=84%. Then your 1 standard deviation estimate is just the difference between these divided by 2. \n\nIt's not the only way but it's got a good grounding in theory and has performed well in kaggle competitions in the past. You might for example get the predictions from lots of models and take the standard deviation of their predictions or you might try to predict the standard deviation for each point from the data.\n\nWhen you see notebooks calculate the optimal value for confidence on the train data they're just solving a set of equations to maximise the objective function. It might be a good exercise for you to do that for a given prediction where you know the true value but it doesn't help in the test data. When you're predicting you don't have this option because you don't KNOW the true value so you can't solve the equation.",
      "votes": 13,
      "replies": [
        {
          "id": 965674,
          "postDate": "2020-08-10T18:56:23.600Z",
          "content": "<p>Thanks, James! That's helpful.</p>",
          "rawMarkdown": "Thanks, James! That's helpful."
        },
        {
          "id": 978950,
          "postDate": "2020-08-20T14:20:59.503Z",
          "content": "<p><a href=\"https://www.kaggle.com/jameschapman19\" target=\"_blank\">@jameschapman19</a>  could u help understand why did we use 50 ,34/34 are left n right of 68% area of std deviation.</p>",
          "rawMarkdown": "@jameschapman19  could u help understand why did we use 50 ,34/34 are left n right of 68% area of std deviation.",
          "votes": 1
        },
        {
          "id": 979073,
          "postDate": "2020-08-20T15:46:14.220Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> if you are able to see the picture I attached on the previous comment it's got a diagram but it's just a statistical identity that 68% of the probability mass of a standard normal distribution lies within +-1 standard deviation of the mean.</p>\n<p>The maths to justify it is here <a href=\"https://en.wikipedia.org/wiki/68%E2%80%9395%E2%80%9399.7_rule\" target=\"_blank\">https://en.wikipedia.org/wiki/68%E2%80%9395%E2%80%9399.7_rule</a></p>",
          "rawMarkdown": "Hi @jaideepvalani if you are able to see the picture I attached on the previous comment it's got a diagram but it's just a statistical identity that 68% of the probability mass of a standard normal distribution lies within +\\-1 standard deviation of the mean.\n\nThe maths to justify it is here https://en.wikipedia.org/wiki/68%E2%80%9395%E2%80%9399.7_rule",
          "votes": 1
        },
        {
          "id": 979516,
          "postDate": "2020-08-20T22:56:37.990Z",
          "content": "<p>Thanks, that's very helpful. i forgot dividing by 2.</p>",
          "rawMarkdown": "Thanks, that's very helpful. i forgot dividing by 2.\n"
        }
      ]
    },
    {
      "id": 969142,
      "postDate": "2020-08-13T13:41:05.093Z",
      "content": "<p>In such a way to maximize the objective function 😁<br>\nIf we ignore the min-max limits, we can find that the optimal sigma is about <code>np.sqrt(2) * np.abs(y - p)</code></p>",
      "rawMarkdown": "In such a way to maximize the objective function 😁\nIf we ignore the min-max limits, we can find that the optimal sigma is about `np.sqrt(2) * np.abs(y - p)`",
      "votes": 2
    },
    {
      "id": 988318,
      "postDate": "2020-08-28T01:14:43.037Z",
      "content": "<p>Does the confidence mean Standard Deviation of Laplace distribution?<br>\nOr should we choose the confidence in order to just minimize the metric for each predicted FVC?</p>",
      "rawMarkdown": "Does the confidence mean Standard Deviation of Laplace distribution?\nOr should we choose the confidence in order to just minimize the metric for each predicted FVC?"
    },
    {
      "id": 979088,
      "postDate": "2020-08-20T15:55:49.887Z",
      "content": "<p>Inspired by <a href=\"https://www.kaggle.com/jaideep\" target=\"_blank\">@jaideep</a>'s question, I've taken another look at this.</p>\n<p>If your predictions are normally distributed (let's suppose about the true mean), then the mean absolute deviation of your predictions will be $$E[\\Delta]=\\sqrt{\\frac{2}{\\pi}}\\sigma$$.</p>\n<p>Now let's look at our expected value for the competition metric (without the clipping):</p>\n<p>$$E[-\\frac{\\sqrt{2}\\Delta}{\\hat{\\sigma}} - ln(\\sqrt{2}\\hat{\\sigma})]=$$</p>\n<p>$$E[\\Delta(-\\frac{\\sqrt{2}}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})]=$$</p>\n<p>$$(\\sqrt{\\frac{2}{\\pi}}\\sigma)(-\\frac{\\sqrt{2}}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})=$$</p>\n<p>$$-2\\sqrt{\\frac{1}{\\pi}}\\frac{\\sigma}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})$$</p>\n<p>Where $\\hat{\\sigma}$ is the sigma we can choose and $\\sigma$ is the standard deviation of our predictions which we can improve but can't choose. </p>\n<p>In order to maximise the metric with respect to $\\hat{\\sigma}$, we differentiate and set to zero:<br>\n$$2\\sqrt{\\frac{1}{\\pi}}\\frac{\\sigma}{\\hat{\\sigma}^2}-\\frac{\\sqrt{2}}{\\hat{\\sigma}}=0$$<br>\n$$\\hat{\\sigma}=\\sqrt{\\frac{2}{\\pi}}\\sigma$$<br>\n$$\\hat{\\sigma}=0.8\\sigma$$</p>\n<p>+/-$0.8 \\pi \\sigma$ would capture about 78% of the probability mass of the distribution i.e. between the 21st and 79th quantile.</p>",
      "rawMarkdown": "Inspired by @jaideep's question, I've taken another look at this.\n\nIf your predictions are normally distributed (let's suppose about the true mean), then the mean absolute deviation of your predictions will be $$E[\\Delta]=\\sqrt{\\frac{2}{\\pi}}\\sigma$$.\n\nNow let's look at our expected value for the competition metric (without the clipping):\n\n$$E[-\\frac{\\sqrt{2}\\Delta}{\\hat{\\sigma}} - ln(\\sqrt{2}\\hat{\\sigma})]=$$\n\n$$E[\\Delta(-\\frac{\\sqrt{2}}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})]=$$\n\n$$(\\sqrt{\\frac{2}{\\pi}}\\sigma)(-\\frac{\\sqrt{2}}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})=$$\n\n$$-2\\sqrt{\\frac{1}{\\pi}}\\frac{\\sigma}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})$$\n\nWhere $\\hat{\\sigma}$ is the sigma we can choose and $\\sigma$ is the standard deviation of our predictions which we can improve but can't choose. \n\nIn order to maximise the metric with respect to $\\hat{\\sigma}$, we differentiate and set to zero:\n$$2\\sqrt{\\frac{1}{\\pi}}\\frac{\\sigma}{\\hat{\\sigma}^2}-\\frac{\\sqrt{2}}{\\hat{\\sigma}}=0$$\n$$\\hat{\\sigma}=\\sqrt{\\frac{2}{\\pi}}\\sigma$$\n$$\\hat{\\sigma}=0.8\\sigma$$\n\n+/-$0.8 \\pi \\sigma$ would capture about 78% of the probability mass of the distribution i.e. between the 21st and 79th quantile.",
      "replies": [
        {
          "id": 979237,
          "postDate": "2020-08-20T17:53:26.257Z",
          "content": "<p>Also - from the penultimate line we can see we have actually just recovered:</p>\n<p>$$\\hat{\\sigma}=E[\\Delta]$$</p>\n<p>In other words, we maximise our score when we choose confidence to be as big as our expected absolute error.</p>",
          "rawMarkdown": "Also - from the penultimate line we can see we have actually just recovered:\n\n$$\\hat{\\sigma}=E[\\Delta]$$\n\nIn other words, we maximise our score when we choose confidence to be as big as our expected absolute error.",
          "votes": 1
        }
      ]
    },
    {
      "id": 965572,
      "postDate": "2020-08-10T18:01:59.917Z",
      "content": "<p>This really is quite confusing. They've mentioned that it is the standard deviation of last three week's FVC, but I'm still kinda confused as well. There have been discussions on the past regarding the confidence, you might want to search for those.</p>",
      "rawMarkdown": "This really is quite confusing. They've mentioned that it is the standard deviation of last three week's FVC, but I'm still kinda confused as well. There have been discussions on the past regarding the confidence, you might want to search for those.",
      "replies": [
        {
          "id": 965664,
          "postDate": "2020-08-10T18:51:41.973Z",
          "content": "<p>Hi Jony, I am not sure if this is correct. The uncertainty that is used in the metric calculation is basically the uncertainty of the model predictions in terms of its standard deviation. For ex: if a model says the FVC at this point is 2000, how confident is the model in this prediction? If it is quite confident, its sigma is low, if the model is not confident then sigma is high. Probabilistic models like (Bayesian linear regression for example) allow you to get such uncertainty as it outputs a distribution. Conversely, deep learning models output point estimates, so you can't get uncertainty here. There are many methods that try to get such uncertainty from deep learning models (for example by approximation, ensembles, etc). You will also find several approaches here in the forum. I hope that makes sense and good luck!!</p>",
          "rawMarkdown": "Hi Jony, I am not sure if this is correct. The uncertainty that is used in the metric calculation is basically the uncertainty of the model predictions in terms of its standard deviation. For ex: if a model says the FVC at this point is 2000, how confident is the model in this prediction? If it is quite confident, its sigma is low, if the model is not confident then sigma is high. Probabilistic models like (Bayesian linear regression for example) allow you to get such uncertainty as it outputs a distribution. Conversely, deep learning models output point estimates, so you can't get uncertainty here. There are many methods that try to get such uncertainty from deep learning models (for example by approximation, ensembles, etc). You will also find several approaches here in the forum. I hope that makes sense and good luck!!",
          "votes": 5
        },
        {
          "id": 965781,
          "postDate": "2020-08-10T21:09:38.953Z",
          "content": "<p>Thanks for the reply <a href=\"/ahmedhshahin\">@ahmedhshahin</a> . This makes it clearer.</p>",
          "rawMarkdown": "Thanks for the reply @ahmedhshahin . This makes it clearer."
        },
        {
          "id": 978939,
          "postDate": "2020-08-20T14:13:02.687Z",
          "content": "<p><a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a>  so when do we get penalized more<br>\nEg. Actual 2150 , Prediction 2000 , Confidence 300<br>\n      Actual 2150 , Prediction 2000 , Confidence 100<br>\nSimilarly <br>\n   Actual 3000 ,Prediction 4000 ,confidence 100<br>\n   Actual 3000  Prediction 3500 ,confidence 300</p>",
          "rawMarkdown": "@ahmedhshahin  so when do we get penalized more\nEg. Actual 2150 , Prediction 2000 , Confidence 300\n      Actual 2150 , Prediction 2000 , Confidence 100\nSimilarly \n   Actual 3000 ,Prediction 4000 ,confidence 100\n   Actual 3000  Prediction 3500 ,confidence 300"
        },
        {
          "id": 979250,
          "postDate": "2020-08-20T18:00:44.723Z",
          "content": "<p>From my other comment: we get punished less and less the \"larger\" our confidence value (but the lower our actual confidence) <strong>up until</strong> confidence equals our absolute error. So without calculating the metric: </p>\n<p>In the first two cases the absolute error is 150. This means the optimal confidence value is 150 i.e. the first of your example. Lower confidence value (higher confidence in the prediction) will be penalized more as well as Higher confidence value (lower confidence in the prediction).</p>",
          "rawMarkdown": "From my other comment: we get punished less and less the \"larger\" our confidence value (but the lower our actual confidence) **up until** confidence equals our absolute error. So without calculating the metric: \n\nIn the first two cases the absolute error is 150. This means the optimal confidence value is 150 i.e. the first of your example. Lower confidence value (higher confidence in the prediction) will be penalized more as well as Higher confidence value (lower confidence in the prediction).\n \n\n",
          "votes": 1
        },
        {
          "id": 979518,
          "postDate": "2020-08-20T22:58:44.677Z",
          "content": "<p><a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a> i found using uncertainty makes more as larger the sigma, the more uncertain/less confidence the model about the prediction. </p>",
          "rawMarkdown": "@ahmedhshahin i found using uncertainty makes more as larger the sigma, the more uncertain/less confidence the model about the prediction. \n\n\n "
        },
        {
          "id": 980151,
          "postDate": "2020-08-21T11:09:57.240Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yimacs\" target=\"_blank\">@yimacs</a> , <br>\nThe point is that if the error is large, then typically one would expect that the standard deviation should be large as well.  So, if someone predicts an FVC value that has a large error, then it's good if their confidence in the prediction is low (ie the predicted standard deviation is large).</p>\n<p>If you make a large error but have low confidence in that prediction, then that will not get heavily penalized. However, if you make the same large error and have high confidence in the predictions, then that should get more heavily penalized.</p>",
          "rawMarkdown": "Hi @yimacs , \nThe point is that if the error is large, then typically one would expect that the standard deviation should be large as well.  So, if someone predicts an FVC value that has a large error, then it's good if their confidence in the prediction is low (ie the predicted standard deviation is large).\n\nIf you make a large error but have low confidence in that prediction, then that will not get heavily penalized. However, if you make the same large error and have high confidence in the predictions, then that should get more heavily penalized.",
          "votes": 1
        }
      ]
    },
    {
      "id": 964704,
      "postDate": "2020-08-10T05:37:59.803Z",
      "content": "<p>It took me a few hours of going through this notebook <a href=\"https://www.kaggle.com/mekhdigakhramanian/forked-osic-multiple-quantile-regression\">https://www.kaggle.com/mekhdigakhramanian/forked-osic-multiple-quantile-regression</a> and reading about pinball loss to get a reasonable understanding of how to do it. Thing that it took me a minute to figure out was that all ytrue values need to be set to the truth FVC value to calculate the qloss correctly. I recommend going through this.</p>",
      "rawMarkdown": "It took me a few hours of going through this notebook https://www.kaggle.com/mekhdigakhramanian/forked-osic-multiple-quantile-regression and reading about pinball loss to get a reasonable understanding of how to do it. Thing that it took me a minute to figure out was that all ytrue values need to be set to the truth FVC value to calculate the qloss correctly. I recommend going through this.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 965663,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "2020-08-10T18:49:29.817000",
      "content": "<p>You don't need the pinball loss/quantile regression to estimate the confidence. The confidence is just that - the ammount of confidence your model has in the given prediction.</p>\n\n<p>To illustrate why the quantile regression is useful (estimating that X% of the data is below a given value). Consider a normal distribution. 68% of the data lies between +/- 1 standard deviation of the mean. So if you want to estimate standard deviation you'd like to know the quantile at 50-34=16% and 50+34=84%. Then your 1 standard deviation estimate is just the difference between these divided by 2. </p>\n\n<p>It's not the only way but it's got a good grounding in theory and has performed well in kaggle competitions in the past. You might for example get the predictions from lots of models and take the standard deviation of their predictions or you might try to predict the standard deviation for each point from the data.</p>\n\n<p>When you see notebooks calculate the optimal value for confidence on the train data they're just solving a set of equations to maximise the objective function. It might be a good exercise for you to do that for a given prediction where you know the true value but it doesn't help in the test data. When you're predicting you don't have this option because you don't KNOW the true value so you can't solve the equation.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 965674,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-08-10T18:56:23.600000",
          "content": "<p>Thanks, James! That's helpful.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 978950,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-08-20T14:20:59.503000",
          "content": "<p><a href=\"https://www.kaggle.com/jameschapman19\" target=\"_blank\">@jameschapman19</a>  could u help understand why did we use 50 ,34/34 are left n right of 68% area of std deviation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979073,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "2020-08-20T15:46:14.220000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> if you are able to see the picture I attached on the previous comment it's got a diagram but it's just a statistical identity that 68% of the probability mass of a standard normal distribution lies within +-1 standard deviation of the mean.</p>\n<p>The maths to justify it is here <a href=\"https://en.wikipedia.org/wiki/68%E2%80%9395%E2%80%9399.7_rule\" target=\"_blank\">https://en.wikipedia.org/wiki/68%E2%80%9395%E2%80%9399.7_rule</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979516,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2020-08-20T22:56:37.990000",
          "content": "<p>Thanks, that's very helpful. i forgot dividing by 2.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 969142,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2020-08-13T13:41:05.093000",
      "content": "<p>In such a way to maximize the objective function 😁<br>\nIf we ignore the min-max limits, we can find that the optimal sigma is about <code>np.sqrt(2) * np.abs(y - p)</code></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 988318,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2020-08-28T01:14:43.037000",
      "content": "<p>Does the confidence mean Standard Deviation of Laplace distribution?<br>\nOr should we choose the confidence in order to just minimize the metric for each predicted FVC?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 979088,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "2020-08-20T15:55:49.887000",
      "content": "<p>Inspired by <a href=\"https://www.kaggle.com/jaideep\" target=\"_blank\">@jaideep</a>'s question, I've taken another look at this.</p>\n<p>If your predictions are normally distributed (let's suppose about the true mean), then the mean absolute deviation of your predictions will be $$E[\\Delta]=\\sqrt{\\frac{2}{\\pi}}\\sigma$$.</p>\n<p>Now let's look at our expected value for the competition metric (without the clipping):</p>\n<p>$$E[-\\frac{\\sqrt{2}\\Delta}{\\hat{\\sigma}} - ln(\\sqrt{2}\\hat{\\sigma})]=$$</p>\n<p>$$E[\\Delta(-\\frac{\\sqrt{2}}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})]=$$</p>\n<p>$$(\\sqrt{\\frac{2}{\\pi}}\\sigma)(-\\frac{\\sqrt{2}}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})=$$</p>\n<p>$$-2\\sqrt{\\frac{1}{\\pi}}\\frac{\\sigma}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})$$</p>\n<p>Where $\\hat{\\sigma}$ is the sigma we can choose and $\\sigma$ is the standard deviation of our predictions which we can improve but can't choose. </p>\n<p>In order to maximise the metric with respect to $\\hat{\\sigma}$, we differentiate and set to zero:<br>\n$$2\\sqrt{\\frac{1}{\\pi}}\\frac{\\sigma}{\\hat{\\sigma}^2}-\\frac{\\sqrt{2}}{\\hat{\\sigma}}=0$$<br>\n$$\\hat{\\sigma}=\\sqrt{\\frac{2}{\\pi}}\\sigma$$<br>\n$$\\hat{\\sigma}=0.8\\sigma$$</p>\n<p>+/-$0.8 \\pi \\sigma$ would capture about 78% of the probability mass of the distribution i.e. between the 21st and 79th quantile.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 979237,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "2020-08-20T17:53:26.257000",
          "content": "<p>Also - from the penultimate line we can see we have actually just recovered:</p>\n<p>$$\\hat{\\sigma}=E[\\Delta]$$</p>\n<p>In other words, we maximise our score when we choose confidence to be as big as our expected absolute error.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 965572,
      "author_name": "johnny",
      "author_url": "",
      "post_date": "2020-08-10T18:01:59.917000",
      "content": "<p>This really is quite confusing. They've mentioned that it is the standard deviation of last three week's FVC, but I'm still kinda confused as well. There have been discussions on the past regarding the confidence, you might want to search for those.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 965664,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-08-10T18:51:41.973000",
          "content": "<p>Hi Jony, I am not sure if this is correct. The uncertainty that is used in the metric calculation is basically the uncertainty of the model predictions in terms of its standard deviation. For ex: if a model says the FVC at this point is 2000, how confident is the model in this prediction? If it is quite confident, its sigma is low, if the model is not confident then sigma is high. Probabilistic models like (Bayesian linear regression for example) allow you to get such uncertainty as it outputs a distribution. Conversely, deep learning models output point estimates, so you can't get uncertainty here. There are many methods that try to get such uncertainty from deep learning models (for example by approximation, ensembles, etc). You will also find several approaches here in the forum. I hope that makes sense and good luck!!</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 965781,
          "author_name": "johnny",
          "author_url": "",
          "post_date": "2020-08-10T21:09:38.953000",
          "content": "<p>Thanks for the reply <a href=\"/ahmedhshahin\">@ahmedhshahin</a> . This makes it clearer.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 978939,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-08-20T14:13:02.687000",
          "content": "<p><a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a>  so when do we get penalized more<br>\nEg. Actual 2150 , Prediction 2000 , Confidence 300<br>\n      Actual 2150 , Prediction 2000 , Confidence 100<br>\nSimilarly <br>\n   Actual 3000 ,Prediction 4000 ,confidence 100<br>\n   Actual 3000  Prediction 3500 ,confidence 300</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979250,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "2020-08-20T18:00:44.723000",
          "content": "<p>From my other comment: we get punished less and less the \"larger\" our confidence value (but the lower our actual confidence) <strong>up until</strong> confidence equals our absolute error. So without calculating the metric: </p>\n<p>In the first two cases the absolute error is 150. This means the optimal confidence value is 150 i.e. the first of your example. Lower confidence value (higher confidence in the prediction) will be penalized more as well as Higher confidence value (lower confidence in the prediction).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 979518,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2020-08-20T22:58:44.677000",
          "content": "<p><a href=\"https://www.kaggle.com/ahmedhshahin\" target=\"_blank\">@ahmedhshahin</a> i found using uncertainty makes more as larger the sigma, the more uncertain/less confidence the model about the prediction. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 980151,
          "author_name": "Ahmed Shahin",
          "author_url": "",
          "post_date": "2020-08-21T11:09:57.240000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yimacs\" target=\"_blank\">@yimacs</a> , <br>\nThe point is that if the error is large, then typically one would expect that the standard deviation should be large as well.  So, if someone predicts an FVC value that has a large error, then it's good if their confidence in the prediction is low (ie the predicted standard deviation is large).</p>\n<p>If you make a large error but have low confidence in that prediction, then that will not get heavily penalized. However, if you make the same large error and have high confidence in the predictions, then that should get more heavily penalized.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 964704,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-10T05:37:59.803000",
      "content": "<p>It took me a few hours of going through this notebook <a href=\"https://www.kaggle.com/mekhdigakhramanian/forked-osic-multiple-quantile-regression\">https://www.kaggle.com/mekhdigakhramanian/forked-osic-multiple-quantile-regression</a> and reading about pinball loss to get a reasonable understanding of how to do it. Thing that it took me a minute to figure out was that all ytrue values need to be set to the truth FVC value to calculate the qloss correctly. I recommend going through this.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "964690": "I am not sure how the confidence value should be calculated.",
    "965663": "You don't need the pinball loss/quantile regression to estimate the confidence. The confidence is just that - the ammount of confidence your model has in the given prediction.\n\nTo illustrate why the quantile regression is useful (estimating that X% of the data is below a given value). Consider a normal distribution. 68% of the data lies between +/- 1 standard deviation of the mean. So if you want to estimate standard deviation you'd like to know the quantile at 50-34=16% and 50+34=84%. Then your 1 standard deviation estimate is just the difference between these divided by 2. \n\nIt's not the only way but it's got a good grounding in theory and has performed well in kaggle competitions in the past. You might for example get the predictions from lots of models and take the standard deviation of their predictions or you might try to predict the standard deviation for each point from the data.\n\nWhen you see notebooks calculate the optimal value for confidence on the train data they're just solving a set of equations to maximise the objective function. It might be a good exercise for you to do that for a given prediction where you know the true value but it doesn't help in the test data. When you're predicting you don't have this option because you don't KNOW the true value so you can't solve the equation.",
    "969142": "In such a way to maximize the objective function 😁\nIf we ignore the min-max limits, we can find that the optimal sigma is about `np.sqrt(2) * np.abs(y - p)`",
    "988318": "Does the confidence mean Standard Deviation of Laplace distribution?\nOr should we choose the confidence in order to just minimize the metric for each predicted FVC?",
    "979088": "Inspired by @jaideep's question, I've taken another look at this.\n\nIf your predictions are normally distributed (let's suppose about the true mean), then the mean absolute deviation of your predictions will be $$E[\\Delta]=\\sqrt{\\frac{2}{\\pi}}\\sigma$$.\n\nNow let's look at our expected value for the competition metric (without the clipping):\n\n$$E[-\\frac{\\sqrt{2}\\Delta}{\\hat{\\sigma}} - ln(\\sqrt{2}\\hat{\\sigma})]=$$\n\n$$E[\\Delta(-\\frac{\\sqrt{2}}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})]=$$\n\n$$(\\sqrt{\\frac{2}{\\pi}}\\sigma)(-\\frac{\\sqrt{2}}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})=$$\n\n$$-2\\sqrt{\\frac{1}{\\pi}}\\frac{\\sigma}{\\hat{\\sigma}}) - ln(\\sqrt{2}\\hat{\\sigma})$$\n\nWhere $\\hat{\\sigma}$ is the sigma we can choose and $\\sigma$ is the standard deviation of our predictions which we can improve but can't choose. \n\nIn order to maximise the metric with respect to $\\hat{\\sigma}$, we differentiate and set to zero:\n$$2\\sqrt{\\frac{1}{\\pi}}\\frac{\\sigma}{\\hat{\\sigma}^2}-\\frac{\\sqrt{2}}{\\hat{\\sigma}}=0$$\n$$\\hat{\\sigma}=\\sqrt{\\frac{2}{\\pi}}\\sigma$$\n$$\\hat{\\sigma}=0.8\\sigma$$\n\n+/-$0.8 \\pi \\sigma$ would capture about 78% of the probability mass of the distribution i.e. between the 21st and 79th quantile.",
    "965572": "This really is quite confusing. They've mentioned that it is the standard deviation of last three week's FVC, but I'm still kinda confused as well. There have been discussions on the past regarding the confidence, you might want to search for those.",
    "964704": "It took me a few hours of going through this notebook https://www.kaggle.com/mekhdigakhramanian/forked-osic-multiple-quantile-regression and reading about pinball loss to get a reasonable understanding of how to do it. Thing that it took me a minute to figure out was that all ytrue values need to be set to the truth FVC value to calculate the qloss correctly. I recommend going through this."
  }
}