{
  "id": 180637,
  "title": "Laplace Log Likelihood Score : Understanding the Theory for Beginners",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/180637",
  "author_name": "",
  "post_date": "2020-09-05T20:41:49.994147400Z",
  "votes": 31,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I've seen a bunch of good notebooks visualizing the effect of different FVC and confidence predictions on the Laplace Log Likelihood, especially this one by Volpani:<br>\n<a href=\"https://www.kaggle.com/rohanrao/osic-understanding-laplace-log-likelihood\" target=\"_blank\">https://www.kaggle.com/rohanrao/osic-understanding-laplace-log-likelihood</a></p>\n<p>But since I've seen a few people asking about what it actually means, I wanted to tackle the theory behind how the metric is defined, and how to understand what it means. So what is the Laplace Log Likelihood?</p>\n<h1>Explaining Likelihood</h1>\n<p>First, we need to ask: what is a <strong>Likelihood Function</strong>? A Likelihood Function basically describes the probability of observing a bunch of data points given that they come from a certain distribution. If we assume that each observation is an <em>independent event</em>, that means that the probability of multiple observations is the product of the probability of each of the observations individually. </p>\n<p>$$p(x_1, x_2, x_3) = p(x_1)p(x_2)p(x_3)$$</p>\n<p>That means that the Likelihood Function looks something like this.</p>\n<p>$$\\mathcal{L} = \\prod{p(x|\\theta)}$$</p>\n<p>Theta is just a shorthand for any possible combo of distribution parameters. The | denotes a conditional probability, meaning that this is the probability of x given that the distribution parameters are indeed theta. So we could plug in any distribution's parameters to figure out the Likelihood of the observed data points x coming from that distribution. The closer to 1 L is, the more likely the data came from that distribution. The closer to 0 L is, the less likely.</p>\n<h1>Explaining Log Likelihood</h1>\n<p>But what is <strong>Log Likelihood</strong>, you might ask? Well, that's simple -- we just take the logarithm of the Likelihood function we defined up above. Because of logarithm rules, the product of probabilities turns into a sum of log probabilities. This makes it easier to read the actual number, since the product of a bunch of small numbers will be really close to 0, while a sum does not suffer from that problem.</p>\n<p>$$ln\\mathcal{L} = \\sum{ln(p(x|\\theta))}$$</p>\n<h1>Deriving the Metric</h1>\n<p>Since this competition uses the <strong>Laplace</strong> Log Likelihood, that means that we're assuming the data is coming from a Laplace Distribution, so we'll have to use the Laplace Distribution probability and parameters. From Wikipedia, we can see that the Probability Distribution Function (PDF) of the Laplace distribution is:<br>\n$$p(x| \\mu, b) = \\frac{1}{2b} exp(-\\frac{|x - \\mu|}{b})$$</p>\n<p>Mu is the mean of the distribution, and b is kinda like the standard deviation of the Normal distribution (well, not exactly, but we'll see how it's related). We can plug this PDF into our Likelihood function, giving us:</p>\n<p>$$ln\\mathcal{L} = \\sum{ln(\\frac{1}{2b} exp(-\\frac{|x - \\mu|}{b}))}$$</p>\n<p>Using log rules, we can separate some terms. (I'm going to drop the sum from now on, since it doesn't really matter.)<br>\n$$ln(\\frac{1}{2b}) + ln(exp(-\\frac{|x - \\mu|}{b}))$$<br>\nSince ln is the inverse of exp, we cancel the two out. We also can rewrite the first term thanks to log rules.<br>\n$$-ln(2b) - \\frac{|x - \\mu|}{b}$$</p>\n<p>This is starting to look a lot like our metric now! We just need to substitute in a few terms. We can say that |x - mu| is kinda like our error.</p>\n<p>$$\\Delta = |x - \\mu|$$</p>\n<p>And if we say that the confidence value sigma is our Standard Deviation, we can derive it using b. From Wikipedia, we know that the Variance of the Laplace distribution is related to b:<br>\n$$Var = 2b^2$$</p>\n<p>Standard Deviation is just the square root of Variance, so we can rewrite this as:<br>\n$$\\sigma= \\sqrt{2}b$$<br>\nwhich means:<br>\n$$b= \\frac{\\sigma}{\\sqrt{2}}$$</p>\n<h1>Plugging it All In</h1>\n<p>$$-ln(2b) - \\frac{|x - \\mu|}{b}$$</p>\n<p>Now we can plug all of our substitutions in. This gives us:<br>\n$$metric = -ln(\\sqrt{2} \\sigma) - \\frac{\\sqrt{2} \\Delta}{\\sigma}$$</p>\n<p>Which is exactly what the competition metric is defined as! I'm guessing that the actual competition computes this value over all predictions you make, and then averages them out to give you your Leaderboard Score.</p>\n<p>I hope this helped you understand a bit of the theory behind the competition metric. If anything looks wrong or wasn't clear, feel free to point it out!</p>",
  "messages": [
    {
      "id": "999649",
      "postDate": "09/05/2020 20:41:49",
      "content": "<p>I've seen a bunch of good notebooks visualizing the effect of different FVC and confidence predictions on the Laplace Log Likelihood, especially this one by Volpani:<br>\n<a href=\"https://www.kaggle.com/rohanrao/osic-understanding-laplace-log-likelihood\" target=\"_blank\">https://www.kaggle.com/rohanrao/osic-understanding-laplace-log-likelihood</a></p>\n<p>But since I've seen a few people asking about what it actually means, I wanted to tackle the theory behind how the metric is defined, and how to understand what it means. So what is the Laplace Log Likelihood?</p>\n<h1>Explaining Likelihood</h1>\n<p>First, we need to ask: what is a <strong>Likelihood Function</strong>? A Likelihood Function basically describes the probability of observing a bunch of data points given that they come from a certain distribution. If we assume that each observation is an <em>independent event</em>, that means that the probability of multiple observations is the product of the probability of each of the observations individually. </p>\n<p>$$p(x_1, x_2, x_3) = p(x_1)p(x_2)p(x_3)$$</p>\n<p>That means that the Likelihood Function looks something like this.</p>\n<p>$$\\mathcal{L} = \\prod{p(x|\\theta)}$$</p>\n<p>Theta is just a shorthand for any possible combo of distribution parameters. The | denotes a conditional probability, meaning that this is the probability of x given that the distribution parameters are indeed theta. So we could plug in any distribution's parameters to figure out the Likelihood of the observed data points x coming from that distribution. The closer to 1 L is, the more likely the data came from that distribution. The closer to 0 L is, the less likely.</p>\n<h1>Explaining Log Likelihood</h1>\n<p>But what is <strong>Log Likelihood</strong>, you might ask? Well, that's simple -- we just take the logarithm of the Likelihood function we defined up above. Because of logarithm rules, the product of probabilities turns into a sum of log probabilities. This makes it easier to read the actual number, since the product of a bunch of small numbers will be really close to 0, while a sum does not suffer from that problem.</p>\n<p>$$ln\\mathcal{L} = \\sum{ln(p(x|\\theta))}$$</p>\n<h1>Deriving the Metric</h1>\n<p>Since this competition uses the <strong>Laplace</strong> Log Likelihood, that means that we're assuming the data is coming from a Laplace Distribution, so we'll have to use the Laplace Distribution probability and parameters. From Wikipedia, we can see that the Probability Distribution Function (PDF) of the Laplace distribution is:<br>\n$$p(x| \\mu, b) = \\frac{1}{2b} exp(-\\frac{|x - \\mu|}{b})$$</p>\n<p>Mu is the mean of the distribution, and b is kinda like the standard deviation of the Normal distribution (well, not exactly, but we'll see how it's related). We can plug this PDF into our Likelihood function, giving us:</p>\n<p>$$ln\\mathcal{L} = \\sum{ln(\\frac{1}{2b} exp(-\\frac{|x - \\mu|}{b}))}$$</p>\n<p>Using log rules, we can separate some terms. (I'm going to drop the sum from now on, since it doesn't really matter.)<br>\n$$ln(\\frac{1}{2b}) + ln(exp(-\\frac{|x - \\mu|}{b}))$$<br>\nSince ln is the inverse of exp, we cancel the two out. We also can rewrite the first term thanks to log rules.<br>\n$$-ln(2b) - \\frac{|x - \\mu|}{b}$$</p>\n<p>This is starting to look a lot like our metric now! We just need to substitute in a few terms. We can say that |x - mu| is kinda like our error.</p>\n<p>$$\\Delta = |x - \\mu|$$</p>\n<p>And if we say that the confidence value sigma is our Standard Deviation, we can derive it using b. From Wikipedia, we know that the Variance of the Laplace distribution is related to b:<br>\n$$Var = 2b^2$$</p>\n<p>Standard Deviation is just the square root of Variance, so we can rewrite this as:<br>\n$$\\sigma= \\sqrt{2}b$$<br>\nwhich means:<br>\n$$b= \\frac{\\sigma}{\\sqrt{2}}$$</p>\n<h1>Plugging it All In</h1>\n<p>$$-ln(2b) - \\frac{|x - \\mu|}{b}$$</p>\n<p>Now we can plug all of our substitutions in. This gives us:<br>\n$$metric = -ln(\\sqrt{2} \\sigma) - \\frac{\\sqrt{2} \\Delta}{\\sigma}$$</p>\n<p>Which is exactly what the competition metric is defined as! I'm guessing that the actual competition computes this value over all predictions you make, and then averages them out to give you your Leaderboard Score.</p>\n<p>I hope this helped you understand a bit of the theory behind the competition metric. If anything looks wrong or wasn't clear, feel free to point it out!</p>",
      "rawMarkdown": "I've seen a bunch of good notebooks visualizing the effect of different FVC and confidence predictions on the Laplace Log Likelihood, especially this one by Volpani:\nhttps://www.kaggle.com/rohanrao/osic-understanding-laplace-log-likelihood\n\nBut since I've seen a few people asking about what it actually means, I wanted to tackle the theory behind how the metric is defined, and how to understand what it means. So what is the Laplace Log Likelihood?\n\n# Explaining Likelihood\n\nFirst, we need to ask: what is a **Likelihood Function**? A Likelihood Function basically describes the probability of observing a bunch of data points given that they come from a certain distribution. If we assume that each observation is an *independent event*, that means that the probability of multiple observations is the product of the probability of each of the observations individually. \n\n$$p(x_1, x_2, x_3) = p(x_1)p(x_2)p(x_3)$$\n\nThat means that the Likelihood Function looks something like this.\n\n$$\\mathcal{L} = \\prod{p(x|\\theta)}$$\n\nTheta is just a shorthand for any possible combo of distribution parameters. The | denotes a conditional probability, meaning that this is the probability of x given that the distribution parameters are indeed theta. So we could plug in any distribution's parameters to figure out the Likelihood of the observed data points x coming from that distribution. The closer to 1 L is, the more likely the data came from that distribution. The closer to 0 L is, the less likely.\n\n# Explaining Log Likelihood\n\nBut what is **Log Likelihood**, you might ask? Well, that's simple -- we just take the logarithm of the Likelihood function we defined up above. Because of logarithm rules, the product of probabilities turns into a sum of log probabilities. This makes it easier to read the actual number, since the product of a bunch of small numbers will be really close to 0, while a sum does not suffer from that problem.\n\n$$ln\\mathcal{L} = \\sum{ln(p(x|\\theta))}$$\n\n\n# Deriving the Metric\n\nSince this competition uses the **Laplace** Log Likelihood, that means that we're assuming the data is coming from a Laplace Distribution, so we'll have to use the Laplace Distribution probability and parameters. From Wikipedia, we can see that the Probability Distribution Function (PDF) of the Laplace distribution is:\n$$p(x| \\mu, b) = \\frac{1}{2b} exp(-\\frac{|x - \\mu|}{b})$$\n\nMu is the mean of the distribution, and b is kinda like the standard deviation of the Normal distribution (well, not exactly, but we'll see how it's related). We can plug this PDF into our Likelihood function, giving us:\n\n$$ln\\mathcal{L} = \\sum{ln(\\frac{1}{2b} exp(-\\frac{|x - \\mu|}{b}))}$$\n\nUsing log rules, we can separate some terms. (I'm going to drop the sum from now on, since it doesn't really matter.)\n$$ln(\\frac{1}{2b}) + ln(exp(-\\frac{|x - \\mu|}{b}))$$\nSince ln is the inverse of exp, we cancel the two out. We also can rewrite the first term thanks to log rules.\n$$-ln(2b) - \\frac{|x - \\mu|}{b}$$\n\nThis is starting to look a lot like our metric now! We just need to substitute in a few terms. We can say that |x - mu| is kinda like our error.\n\n$$\\Delta = |x - \\mu|$$\n\nAnd if we say that the confidence value sigma is our Standard Deviation, we can derive it using b. From Wikipedia, we know that the Variance of the Laplace distribution is related to b:\n$$Var = 2b^2$$\n\nStandard Deviation is just the square root of Variance, so we can rewrite this as:\n$$\\sigma= \\sqrt{2}b$$\nwhich means:\n$$b= \\frac{\\sigma}{\\sqrt{2}}$$\n\n\n# Plugging it All In\n$$-ln(2b) - \\frac{|x - \\mu|}{b}$$\n\n\nNow we can plug all of our substitutions in. This gives us:\n$$metric = -ln(\\sqrt{2} \\sigma) - \\frac{\\sqrt{2} \\Delta}{\\sigma}$$\n\nWhich is exactly what the competition metric is defined as! I'm guessing that the actual competition computes this value over all predictions you make, and then averages them out to give you your Leaderboard Score.\n\nI hope this helped you understand a bit of the theory behind the competition metric. If anything looks wrong or wasn't clear, feel free to point it out!",
      "votes": null
    },
    {
      "id": "999940",
      "postDate": "09/06/2020 06:34:29",
      "content": "<p>Thanks for putting this here. Really helpful to know how the metric is defined.</p>",
      "rawMarkdown": "Thanks for putting this here. Really helpful to know how the metric is defined.",
      "votes": null
    },
    {
      "id": "1000075",
      "postDate": "09/06/2020 09:03:32",
      "content": "<p>Thanks! I thought your Watershed Lung Segmentation notebook was really novel and interesting. I'll probably keep that technique in mind when I move on from tabular data to processing the images as well!</p>",
      "rawMarkdown": "Thanks! I thought your Watershed Lung Segmentation notebook was really novel and interesting. I'll probably keep that technique in mind when I move on from tabular data to processing the images as well!",
      "votes": null
    },
    {
      "id": "1000995",
      "postDate": "09/07/2020 01:52:29",
      "content": "<p>Good explanation demystifying the metric</p>",
      "rawMarkdown": "Good explanation demystifying the metric",
      "votes": null
    },
    {
      "id": "1001472",
      "postDate": "09/07/2020 11:00:35",
      "content": "<p>Really helpfull, I haven't been familiar with LLL, thanks! </p>",
      "rawMarkdown": "Really helpfull, I haven't been familiar with LLL, thanks!",
      "votes": null
    },
    {
      "id": "1003401",
      "postDate": "09/08/2020 23:42:02",
      "content": "<p>Great explanation! Thanks!</p>",
      "rawMarkdown": "Great explanation! Thanks!",
      "votes": null
    },
    {
      "id": "1004073",
      "postDate": "09/09/2020 12:50:52",
      "content": "<p>I really needed this as its helpful for a competition. thanks for sharing this</p>",
      "rawMarkdown": "I really needed this as its helpful for a competition. thanks for sharing this",
      "votes": null
    },
    {
      "id": "1006062",
      "postDate": "09/11/2020 03:04:03",
      "content": "<p>Great explanation of the function and how it connects to a Laplace distribution! I think one thing that's worth noting is the minimum of this function:</p>\n<p>Setting \\(\\Delta = 0\\) by getting \\(FVC_{predicted} = FVC_{true},\\) we simplify the metric to $$-\\ln(\\sqrt{2} \\sigma_{clipped}).$$ Since the larger metric the better (vector wise), we can maximize this function with $$\\sigma_{clipped} = 70,$$ yielding a maximum value of \\(-4.595\\). This just gives you a basic understanding of what the goal is for this metric.</p>\n<p>Overall, great post!</p>",
      "rawMarkdown": "Great explanation of the function and how it connects to a Laplace distribution! I think one thing that's worth noting is the minimum of this function:\n\nSetting \\\\(\\Delta = 0\\\\) by getting \\\\(FVC_{predicted} = FVC_{true},\\\\) we simplify the metric to $$-\\ln(\\sqrt{2} \\sigma_{clipped}).$$ Since the larger metric the better (vector wise), we can maximize this function with $$\\sigma_{clipped} = 70,$$ yielding a maximum value of \\\\(-4.595\\\\). This just gives you a basic understanding of what the goal is for this metric.\n\nOverall, great post!",
      "votes": null
    },
    {
      "id": "1006469",
      "postDate": "09/11/2020 09:21:19",
      "content": "<p><a href=\"https://www.kaggle.com/neithermannormachine\" target=\"_blank\">@neithermannormachine</a>  thanks a lot,<br>\n how can we interpret  the likelihood in terms of Sigma /FVC like in broad sense u say likelihood to be as likelihood of a probability for given observed values of X.  so how we could fit sigma/fvc here..</p>",
      "rawMarkdown": "neithermannormachine  thanks a lot,\n how can we interpret  the likelihood in terms of Sigma /FVC like in broad sense u say likelihood to be as likelihood of a probability for given observed values of X.  so how we could fit sigma/fvc here..",
      "votes": null
    },
    {
      "id": "1006739",
      "postDate": "09/11/2020 14:18:32",
      "content": "<p>Sure! The way I think of it, for every \"true\" data point (ie an observation for a specific patient at a certain week), there is a Laplace distribution centered at the true FVC. Going back to the definition of a Laplace distribution, this basically means:</p>\n<p>$$\\mu = FVC_{true}$$<br>\n$$x = FVC_{pred}$$</p>\n<p>For this specific distribution around FVC_true, we want to output FVC_pred and sigma so that we have the highest probability of FVC_pred coming from the Laplace distribution around FVC_true.</p>\n<p>For example, let's say we have two different distributions with the same mean FVC_true = 2000. The one on the left (A) has a <em>lower</em> sigma, while the one on the right (B) has a <em>higher</em> sigma.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5358539%2F699a6dbfdd346f93762f5a15e188a7cf%2F2020-09-11%2008.14.04%20www.kaggle.com%20ec382940eaaa.png?generation=1599837265198666&amp;alt=media\" alt=\"\"></p>\n<p>To maximize this probability, we want to predict FVC_pred so that it's as high up on the curve as possible. However, the value of sigma we output will determine how high up we can get. Notice how when we predict FVC_pred = 2050 on A, we'll get a bigger probability than when we predict FVC_pred = 2050 on B, and so the prediction of 2050 is higher up on curve A. This makes sense, since curve A is the case when we're more confident about the closeness of our prediction to the true value.</p>\n<p>On the other hand, we can get a bigger probability for FVC_pred = 2200 on curve B than on curve A, so the prediction of 2200 is higher up on curve B. This also makes sense, since we're a little further off, but we're willing to recognize that we're less confident about our prediction.</p>\n<p>In general, when you increase sigma, the center of the distribution gets shorter, but the tails get taller. So it's good to be really confident (low sigma) for close predictions, and less confident (high sigma) for far predictions.</p>\n<p>By optimizing for the metric defined in the competition, we are essentially doing this process for each prediction we make. Hope that answered your question!</p>",
      "rawMarkdown": "Sure! The way I think of it, for every \"true\" data point (ie an observation for a specific patient at a certain week), there is a Laplace distribution centered at the true FVC. Going back to the definition of a Laplace distribution, this basically means:\n\n$$\\mu = FVC_{true}$$\n$$x = FVC_{pred}$$\n\n\nFor this specific distribution around FVC_true, we want to output FVC_pred and sigma so that we have the highest probability of FVC_pred coming from the Laplace distribution around FVC_true.\n\nFor example, let's say we have two different distributions with the same mean FVC_true = 2000. The one on the left (A) has a *lower* sigma, while the one on the right (B) has a *higher* sigma.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5358539%2F699a6dbfdd346f93762f5a15e188a7cf%2F2020-09-11%2008.14.04%20www.kaggle.com%20ec382940eaaa.png?generation=1599837265198666&alt=media)\n\nTo maximize this probability, we want to predict FVC_pred so that it's as high up on the curve as possible. However, the value of sigma we output will determine how high up we can get. Notice how when we predict FVC_pred = 2050 on A, we'll get a bigger probability than when we predict FVC_pred = 2050 on B, and so the prediction of 2050 is higher up on curve A. This makes sense, since curve A is the case when we're more confident about the closeness of our prediction to the true value.\n\nOn the other hand, we can get a bigger probability for FVC_pred = 2200 on curve B than on curve A, so the prediction of 2200 is higher up on curve B. This also makes sense, since we're a little further off, but we're willing to recognize that we're less confident about our prediction.\n\nIn general, when you increase sigma, the center of the distribution gets shorter, but the tails get taller. So it's good to be really confident (low sigma) for close predictions, and less confident (high sigma) for far predictions.\n\nBy optimizing for the metric defined in the competition, we are essentially doing this process for each prediction we make. Hope that answered your question!",
      "votes": null
    },
    {
      "id": "1006748",
      "postDate": "09/11/2020 14:28:12",
      "content": "<p>That's a great point. It's definitely good to know what the best possible score is so that we can have a goal in mind. Thanks!</p>",
      "rawMarkdown": "That's a great point. It's definitely good to know what the best possible score is so that we can have a goal in mind. Thanks!",
      "votes": null
    },
    {
      "id": "1006829",
      "postDate": "09/11/2020 15:23:45",
      "content": "<p>No problem! Glad to help everyone here out.</p>",
      "rawMarkdown": "No problem! Glad to help everyone here out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 999940,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "09/06/2020 06:34:29",
      "content": "<p>Thanks for putting this here. Really helpful to know how the metric is defined.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1000075,
          "author_name": "neithermannormachine",
          "author_url": "",
          "post_date": "09/06/2020 09:03:32",
          "content": "<p>Thanks! I thought your Watershed Lung Segmentation notebook was really novel and interesting. I'll probably keep that technique in mind when I move on from tabular data to processing the images as well!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1000995,
      "author_name": "gcspkmdr",
      "author_url": "",
      "post_date": "09/07/2020 01:52:29",
      "content": "<p>Good explanation demystifying the metric</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1001472,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "09/07/2020 11:00:35",
      "content": "<p>Really helpfull, I haven't been familiar with LLL, thanks! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1003401,
      "author_name": "mariapushkareva",
      "author_url": "",
      "post_date": "09/08/2020 23:42:02",
      "content": "<p>Great explanation! Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1006062,
      "author_name": "qwerty71",
      "author_url": "",
      "post_date": "09/11/2020 03:04:03",
      "content": "<p>Great explanation of the function and how it connects to a Laplace distribution! I think one thing that's worth noting is the minimum of this function:</p>\n<p>Setting \\(\\Delta = 0\\) by getting \\(FVC_{predicted} = FVC_{true},\\) we simplify the metric to $$-\\ln(\\sqrt{2} \\sigma_{clipped}).$$ Since the larger metric the better (vector wise), we can maximize this function with $$\\sigma_{clipped} = 70,$$ yielding a maximum value of \\(-4.595\\). This just gives you a basic understanding of what the goal is for this metric.</p>\n<p>Overall, great post!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1006748,
          "author_name": "neithermannormachine",
          "author_url": "",
          "post_date": "09/11/2020 14:28:12",
          "content": "<p>That's a great point. It's definitely good to know what the best possible score is so that we can have a goal in mind. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1006469,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "09/11/2020 09:21:19",
      "content": "<p><a href=\"https://www.kaggle.com/neithermannormachine\" target=\"_blank\">@neithermannormachine</a>  thanks a lot,<br>\n how can we interpret  the likelihood in terms of Sigma /FVC like in broad sense u say likelihood to be as likelihood of a probability for given observed values of X.  so how we could fit sigma/fvc here..</p>",
      "votes": null,
      "replies": [
        {
          "id": 1006739,
          "author_name": "neithermannormachine",
          "author_url": "",
          "post_date": "09/11/2020 14:18:32",
          "content": "<p>Sure! The way I think of it, for every \"true\" data point (ie an observation for a specific patient at a certain week), there is a Laplace distribution centered at the true FVC. Going back to the definition of a Laplace distribution, this basically means:</p>\n<p>$$\\mu = FVC_{true}$$<br>\n$$x = FVC_{pred}$$</p>\n<p>For this specific distribution around FVC_true, we want to output FVC_pred and sigma so that we have the highest probability of FVC_pred coming from the Laplace distribution around FVC_true.</p>\n<p>For example, let's say we have two different distributions with the same mean FVC_true = 2000. The one on the left (A) has a <em>lower</em> sigma, while the one on the right (B) has a <em>higher</em> sigma.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5358539%2F699a6dbfdd346f93762f5a15e188a7cf%2F2020-09-11%2008.14.04%20www.kaggle.com%20ec382940eaaa.png?generation=1599837265198666&amp;alt=media\" alt=\"\"></p>\n<p>To maximize this probability, we want to predict FVC_pred so that it's as high up on the curve as possible. However, the value of sigma we output will determine how high up we can get. Notice how when we predict FVC_pred = 2050 on A, we'll get a bigger probability than when we predict FVC_pred = 2050 on B, and so the prediction of 2050 is higher up on curve A. This makes sense, since curve A is the case when we're more confident about the closeness of our prediction to the true value.</p>\n<p>On the other hand, we can get a bigger probability for FVC_pred = 2200 on curve B than on curve A, so the prediction of 2200 is higher up on curve B. This also makes sense, since we're a little further off, but we're willing to recognize that we're less confident about our prediction.</p>\n<p>In general, when you increase sigma, the center of the distribution gets shorter, but the tails get taller. So it's good to be really confident (low sigma) for close predictions, and less confident (high sigma) for far predictions.</p>\n<p>By optimizing for the metric defined in the competition, we are essentially doing this process for each prediction we make. Hope that answered your question!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1004073,
      "author_name": "blessondensil294",
      "author_url": "",
      "post_date": "09/09/2020 12:50:52",
      "content": "<p>I really needed this as its helpful for a competition. thanks for sharing this</p>",
      "votes": null,
      "replies": [
        {
          "id": 1006829,
          "author_name": "neithermannormachine",
          "author_url": "",
          "post_date": "09/11/2020 15:23:45",
          "content": "<p>No problem! Glad to help everyone here out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "999649": "I've seen a bunch of good notebooks visualizing the effect of different FVC and confidence predictions on the Laplace Log Likelihood, especially this one by Volpani:\nhttps://www.kaggle.com/rohanrao/osic-understanding-laplace-log-likelihood\n\nBut since I've seen a few people asking about what it actually means, I wanted to tackle the theory behind how the metric is defined, and how to understand what it means. So what is the Laplace Log Likelihood?\n\n# Explaining Likelihood\n\nFirst, we need to ask: what is a **Likelihood Function**? A Likelihood Function basically describes the probability of observing a bunch of data points given that they come from a certain distribution. If we assume that each observation is an *independent event*, that means that the probability of multiple observations is the product of the probability of each of the observations individually. \n\n$$p(x_1, x_2, x_3) = p(x_1)p(x_2)p(x_3)$$\n\nThat means that the Likelihood Function looks something like this.\n\n$$\\mathcal{L} = \\prod{p(x|\\theta)}$$\n\nTheta is just a shorthand for any possible combo of distribution parameters. The | denotes a conditional probability, meaning that this is the probability of x given that the distribution parameters are indeed theta. So we could plug in any distribution's parameters to figure out the Likelihood of the observed data points x coming from that distribution. The closer to 1 L is, the more likely the data came from that distribution. The closer to 0 L is, the less likely.\n\n# Explaining Log Likelihood\n\nBut what is **Log Likelihood**, you might ask? Well, that's simple -- we just take the logarithm of the Likelihood function we defined up above. Because of logarithm rules, the product of probabilities turns into a sum of log probabilities. This makes it easier to read the actual number, since the product of a bunch of small numbers will be really close to 0, while a sum does not suffer from that problem.\n\n$$ln\\mathcal{L} = \\sum{ln(p(x|\\theta))}$$\n\n\n# Deriving the Metric\n\nSince this competition uses the **Laplace** Log Likelihood, that means that we're assuming the data is coming from a Laplace Distribution, so we'll have to use the Laplace Distribution probability and parameters. From Wikipedia, we can see that the Probability Distribution Function (PDF) of the Laplace distribution is:\n$$p(x| \\mu, b) = \\frac{1}{2b} exp(-\\frac{|x - \\mu|}{b})$$\n\nMu is the mean of the distribution, and b is kinda like the standard deviation of the Normal distribution (well, not exactly, but we'll see how it's related). We can plug this PDF into our Likelihood function, giving us:\n\n$$ln\\mathcal{L} = \\sum{ln(\\frac{1}{2b} exp(-\\frac{|x - \\mu|}{b}))}$$\n\nUsing log rules, we can separate some terms. (I'm going to drop the sum from now on, since it doesn't really matter.)\n$$ln(\\frac{1}{2b}) + ln(exp(-\\frac{|x - \\mu|}{b}))$$\nSince ln is the inverse of exp, we cancel the two out. We also can rewrite the first term thanks to log rules.\n$$-ln(2b) - \\frac{|x - \\mu|}{b}$$\n\nThis is starting to look a lot like our metric now! We just need to substitute in a few terms. We can say that |x - mu| is kinda like our error.\n\n$$\\Delta = |x - \\mu|$$\n\nAnd if we say that the confidence value sigma is our Standard Deviation, we can derive it using b. From Wikipedia, we know that the Variance of the Laplace distribution is related to b:\n$$Var = 2b^2$$\n\nStandard Deviation is just the square root of Variance, so we can rewrite this as:\n$$\\sigma= \\sqrt{2}b$$\nwhich means:\n$$b= \\frac{\\sigma}{\\sqrt{2}}$$\n\n\n# Plugging it All In\n$$-ln(2b) - \\frac{|x - \\mu|}{b}$$\n\n\nNow we can plug all of our substitutions in. This gives us:\n$$metric = -ln(\\sqrt{2} \\sigma) - \\frac{\\sqrt{2} \\Delta}{\\sigma}$$\n\nWhich is exactly what the competition metric is defined as! I'm guessing that the actual competition computes this value over all predictions you make, and then averages them out to give you your Leaderboard Score.\n\nI hope this helped you understand a bit of the theory behind the competition metric. If anything looks wrong or wasn't clear, feel free to point it out!",
    "999940": "Thanks for putting this here. Really helpful to know how the metric is defined.",
    "1000075": "Thanks! I thought your Watershed Lung Segmentation notebook was really novel and interesting. I'll probably keep that technique in mind when I move on from tabular data to processing the images as well!",
    "1000995": "Good explanation demystifying the metric",
    "1001472": "Really helpfull, I haven't been familiar with LLL, thanks!",
    "1003401": "Great explanation! Thanks!",
    "1004073": "I really needed this as its helpful for a competition. thanks for sharing this",
    "1006062": "Great explanation of the function and how it connects to a Laplace distribution! I think one thing that's worth noting is the minimum of this function:\n\nSetting \\\\(\\Delta = 0\\\\) by getting \\\\(FVC_{predicted} = FVC_{true},\\\\) we simplify the metric to $$-\\ln(\\sqrt{2} \\sigma_{clipped}).$$ Since the larger metric the better (vector wise), we can maximize this function with $$\\sigma_{clipped} = 70,$$ yielding a maximum value of \\\\(-4.595\\\\). This just gives you a basic understanding of what the goal is for this metric.\n\nOverall, great post!",
    "1006469": "neithermannormachine  thanks a lot,\n how can we interpret  the likelihood in terms of Sigma /FVC like in broad sense u say likelihood to be as likelihood of a probability for given observed values of X.  so how we could fit sigma/fvc here..",
    "1006739": "Sure! The way I think of it, for every \"true\" data point (ie an observation for a specific patient at a certain week), there is a Laplace distribution centered at the true FVC. Going back to the definition of a Laplace distribution, this basically means:\n\n$$\\mu = FVC_{true}$$\n$$x = FVC_{pred}$$\n\n\nFor this specific distribution around FVC_true, we want to output FVC_pred and sigma so that we have the highest probability of FVC_pred coming from the Laplace distribution around FVC_true.\n\nFor example, let's say we have two different distributions with the same mean FVC_true = 2000. The one on the left (A) has a *lower* sigma, while the one on the right (B) has a *higher* sigma.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5358539%2F699a6dbfdd346f93762f5a15e188a7cf%2F2020-09-11%2008.14.04%20www.kaggle.com%20ec382940eaaa.png?generation=1599837265198666&alt=media)\n\nTo maximize this probability, we want to predict FVC_pred so that it's as high up on the curve as possible. However, the value of sigma we output will determine how high up we can get. Notice how when we predict FVC_pred = 2050 on A, we'll get a bigger probability than when we predict FVC_pred = 2050 on B, and so the prediction of 2050 is higher up on curve A. This makes sense, since curve A is the case when we're more confident about the closeness of our prediction to the true value.\n\nOn the other hand, we can get a bigger probability for FVC_pred = 2200 on curve B than on curve A, so the prediction of 2200 is higher up on curve B. This also makes sense, since we're a little further off, but we're willing to recognize that we're less confident about our prediction.\n\nIn general, when you increase sigma, the center of the distribution gets shorter, but the tails get taller. So it's good to be really confident (low sigma) for close predictions, and less confident (high sigma) for far predictions.\n\nBy optimizing for the metric defined in the competition, we are essentially doing this process for each prediction we make. Hope that answered your question!",
    "1006748": "That's a great point. It's definitely good to know what the best possible score is so that we can have a goal in mind. Thanks!",
    "1006829": "No problem! Glad to help everyone here out."
  },
  "source": "meta"
}