{
  "id": 65970,
  "title": "We need a better loss function",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/65970",
  "author_name": "",
  "post_date": "2018-09-16T19:20:09.751773100Z",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>If you are pursuing a segmentation approach along the lines of <a href=\"https://www.kaggle.com/jonnedtc/cnn-segmentation-connected-components\">Jonne's kernel</a>, you probably use cross-entropy or IoU (or some combination, as in the kernel) as your loss function. But both of these approaches have problems.  IoU depends on a particular probability threshold, so it won't give the model credit for improving, for example, if a positive pixel goes from p=.3 to p=.4 (if the threshold is .5). [EDIT: It seems the Jaccard loss function, as implemented in the kernel, does in fact give credit for these kinds of improvements, but in my experience it just isn't producing very good results.]  And cross-entropy gives full veto power to the bounding box, even though, by its nature, a rectangle is only an approximation.  For example, even if the radiologist has perfectly located the bounding box, that box may have a region in the corner that doesn't look at all like pneumonia, and your model shouldn't be penalized for a negative prediction in that region.</p>\n\n<p>I'm trying to come up with a better loss function.  I'll start the bidding here, and see if anyone has a better idea.  Define a value <code>p_true</code>, which is equal to zero (or possibly some other value <code>p_min</code>) far outside a bounding box, and equal to <code>p_max</code> at the center of the bounding box (where <code>p_max</code> is close to 1 but not exactly 1), and the values elsewhere will vary depending on how close you are to the center of a box and whether or not you're inside the box.  (For now I'll leave open the questions of how to exactly to vary <code>p_true</code> and how to combine information from different bounding boxes.)  Then define a loss function (where <code>p_pred</code> is the predicted probability for a given pixel):<br><br>\n <code>-log( p_pred*p_true + (1-p_pred)*(1-p_true) )</code><br><br>\nand aggregate that over pixels as you would with cross-entropy.  I think this loss function avoids the problems discussed above.  Maybe it has other problems.  Has it been used elsewhere?  Does it have a name?  Anyhow, assuming you've chosen <code>p_true</code> well, it should penalize <em>inconsistency</em> between the true and predicted probabilities.</p>",
  "messages": [
    {
      "id": "388322",
      "postDate": "09/16/2018 19:20:09",
      "content": "<p>If you are pursuing a segmentation approach along the lines of <a href=\"https://www.kaggle.com/jonnedtc/cnn-segmentation-connected-components\">Jonne's kernel</a>, you probably use cross-entropy or IoU (or some combination, as in the kernel) as your loss function. But both of these approaches have problems.  IoU depends on a particular probability threshold, so it won't give the model credit for improving, for example, if a positive pixel goes from p=.3 to p=.4 (if the threshold is .5). [EDIT: It seems the Jaccard loss function, as implemented in the kernel, does in fact give credit for these kinds of improvements, but in my experience it just isn't producing very good results.]  And cross-entropy gives full veto power to the bounding box, even though, by its nature, a rectangle is only an approximation.  For example, even if the radiologist has perfectly located the bounding box, that box may have a region in the corner that doesn't look at all like pneumonia, and your model shouldn't be penalized for a negative prediction in that region.</p>\n\n<p>I'm trying to come up with a better loss function.  I'll start the bidding here, and see if anyone has a better idea.  Define a value <code>p_true</code>, which is equal to zero (or possibly some other value <code>p_min</code>) far outside a bounding box, and equal to <code>p_max</code> at the center of the bounding box (where <code>p_max</code> is close to 1 but not exactly 1), and the values elsewhere will vary depending on how close you are to the center of a box and whether or not you're inside the box.  (For now I'll leave open the questions of how to exactly to vary <code>p_true</code> and how to combine information from different bounding boxes.)  Then define a loss function (where <code>p_pred</code> is the predicted probability for a given pixel):<br><br>\n <code>-log( p_pred*p_true + (1-p_pred)*(1-p_true) )</code><br><br>\nand aggregate that over pixels as you would with cross-entropy.  I think this loss function avoids the problems discussed above.  Maybe it has other problems.  Has it been used elsewhere?  Does it have a name?  Anyhow, assuming you've chosen <code>p_true</code> well, it should penalize <em>inconsistency</em> between the true and predicted probabilities.</p>",
      "rawMarkdown": "If you are pursuing a segmentation approach along the lines of [Jonne's kernel][1], you probably use cross-entropy or IoU (or some combination, as in the kernel) as your loss function. But both of these approaches have problems.  IoU depends on a particular probability threshold, so it won't give the model credit for improving, for example, if a positive pixel goes from p=.3 to p=.4 (if the threshold is .5). [EDIT: It seems the Jaccard loss function, as implemented in the kernel, does in fact give credit for these kinds of improvements, but in my experience it just isn't producing very good results.]  And cross-entropy gives full veto power to the bounding box, even though, by its nature, a rectangle is only an approximation.  For example, even if the radiologist has perfectly located the bounding box, that box may have a region in the corner that doesn't look at all like pneumonia, and your model shouldn't be penalized for a negative prediction in that region.\n\nI'm trying to come up with a better loss function.  I'll start the bidding here, and see if anyone has a better idea.  Define a value `p_true`, which is equal to zero (or possibly some other value `p_min`) far outside a bounding box, and equal to `p_max` at the center of the bounding box (where `p_max` is close to 1 but not exactly 1), and the values elsewhere will vary depending on how close you are to the center of a box and whether or not you're inside the box.  (For now I'll leave open the questions of how to exactly to vary `p_true` and how to combine information from different bounding boxes.)  Then define a loss function (where `p_pred` is the predicted probability for a given pixel):<br><br>\n `-log( p_pred*p_true + (1-p_pred)*(1-p_true) )`<br><br>\nand aggregate that over pixels as you would with cross-entropy.  I think this loss function avoids the problems discussed above.  Maybe it has other problems.  Has it been used elsewhere?  Does it have a name?  Anyhow, assuming you've chosen `p_true` well, it should penalize *inconsistency* between the true and predicted probabilities.\n\n [1]: https://www.kaggle.com/jonnedtc/cnn-segmentation-connected-components",
      "votes": null
    },
    {
      "id": "388362",
      "postDate": "09/16/2018 21:03:01",
      "content": "<p>what happened to your raw_iou function? it seems to fit the bill?</p>",
      "rawMarkdown": "what happened to your raw_iou function? it seems to fit the bill?",
      "votes": null
    },
    {
      "id": "388374",
      "postDate": "09/16/2018 21:53:47",
      "content": "<p><code>raw_iou</code> is just a metric to keep track for validation purposes. I don't optimize it.</p>",
      "rawMarkdown": "`raw_iou` is just a metric to keep track for validation purposes. I don't optimize it.",
      "votes": null
    },
    {
      "id": "388524",
      "postDate": "09/17/2018 06:10:48",
      "content": "<p>Why don't to use the ordinary cross-entropy? You can however to use it with \"soft\" (not binary) p_true. For example, 0 far away from a box, 1 near the center and something like a linear function near the borders.</p>",
      "rawMarkdown": "Why don't to use the ordinary cross-entropy? You can however to use it with \"soft\" (not binary) p_true. For example, 0 far away from a box, 1 near the center and something like a linear function near the borders.",
      "votes": null
    },
    {
      "id": "388708",
      "postDate": "09/17/2018 13:26:14",
      "content": "<p>The problem with cross-entropy, even with soft p_true, is that it will penalize the model for being confident even when its confidence may be justified. For example, with <code>p_true=0.9</code> and <code>p_pred=0.999</code>, you get <code>CE=-0.9*log(0.999)-0.1*log(0.001)</code>, which gives a huge loss.</p>",
      "rawMarkdown": "The problem with cross-entropy, even with soft p_true, is that it will penalize the model for being confident even when its confidence may be justified. For example, with `p_true=0.9` and `p_pred=0.999`, you get `CE=-0.9*log(0.999)-0.1*log(0.001)`, which gives a huge loss.",
      "votes": null
    },
    {
      "id": "388742",
      "postDate": "09/17/2018 13:54:21",
      "content": "<p>what about -log(1 - (y-p)^2) where (y-p)^2 is the absolute squared prediction error? </p>\n\n<p>-log(1 -abs(y-p)) should also work...</p>",
      "rawMarkdown": "what about -log(1 - (y-p)^2) where (y-p)^2 is the absolute squared prediction error? \n\n-log(1 -abs(y-p)) should also work...",
      "votes": null
    },
    {
      "id": "388788",
      "postDate": "09/17/2018 15:29:59",
      "content": "<p>I don't like anything with <code>y-p</code> in it, because it wants the model to match the ground truth's level of uncertainty even when the model may have good reasons for being more certain.  For example, it will score <code>y=0.49, p=0.99</code> the same as <code>y=0.2, p=0.7</code> even though, in the latter case, the model is probably wrong, whereas in the former case, the model has almost a 50% chance of being right.</p>",
      "rawMarkdown": "I don't like anything with `y-p` in it, because it wants the model to match the ground truth's level of uncertainty even when the model may have good reasons for being more certain.  For example, it will score `y=0.49, p=0.99` the same as `y=0.2, p=0.7` even though, in the latter case, the model is probably wrong, whereas in the former case, the model has almost a 50% chance of being right.",
      "votes": null
    },
    {
      "id": "388803",
      "postDate": "09/17/2018 15:56:30",
      "content": "<p>My original suggestion answers the question, \"Assuming the probabilities given by the model and by ground truth represent independent Bernoulli random variables, what is the probability (expressed as a negative log likelihood) that the two variables will have the same realization?\"  Not sure if there's a name for that.  Maybe \"naive Bayesian negative log likelihood\"?</p>",
      "rawMarkdown": "My original suggestion answers the question, \"Assuming the probabilities given by the model and by ground truth represent independent Bernoulli random variables, what is the probability (expressed as a negative log likelihood) that the two variables will have the same realization?\"  Not sure if there's a name for that.  Maybe \"naive Bayesian negative log likelihood\"?",
      "votes": null
    },
    {
      "id": "388948",
      "postDate": "09/17/2018 21:40:10",
      "content": "<p>I am not sure whether you have tried this one \n<a href=\"https://github.com/bermanmaxim/LovaszSoftmax\">https://github.com/bermanmaxim/LovaszSoftmax</a></p>",
      "rawMarkdown": "I am not sure whether you have tried this one \nhttps://github.com/bermanmaxim/LovaszSoftmax",
      "votes": null
    },
    {
      "id": "389409",
      "postDate": "09/18/2018 17:16:59",
      "content": "<p>Wow! This loss function looks promising on paper, will give it a try. </p>",
      "rawMarkdown": "Wow! This loss function looks promising on paper, will give it a try.",
      "votes": null
    },
    {
      "id": "389794",
      "postDate": "09/19/2018 08:33:57",
      "content": "<p>To my understanding focal loss may be well suited for this kind of problem. Focal loss is not a function in y-p and gives a lower error for almost correctly predicted samples, compared to binary crossentropy. Jeremy Howard explains this loss function in quite some detail in the fast.ai course, part 9.</p>",
      "rawMarkdown": "To my understanding focal loss may be well suited for this kind of problem. Focal loss is not a function in y-p and gives a lower error for almost correctly predicted samples, compared to binary crossentropy. Jeremy Howard explains this loss function in quite some detail in the fast.ai course, part 9.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 388362,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "09/16/2018 21:03:01",
      "content": "<p>what happened to your raw_iou function? it seems to fit the bill?</p>",
      "votes": null,
      "replies": [
        {
          "id": 388374,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/16/2018 21:53:47",
          "content": "<p><code>raw_iou</code> is just a metric to keep track for validation purposes. I don't optimize it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 388524,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "09/17/2018 06:10:48",
      "content": "<p>Why don't to use the ordinary cross-entropy? You can however to use it with \"soft\" (not binary) p_true. For example, 0 far away from a box, 1 near the center and something like a linear function near the borders.</p>",
      "votes": null,
      "replies": [
        {
          "id": 388708,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/17/2018 13:26:14",
          "content": "<p>The problem with cross-entropy, even with soft p_true, is that it will penalize the model for being confident even when its confidence may be justified. For example, with <code>p_true=0.9</code> and <code>p_pred=0.999</code>, you get <code>CE=-0.9*log(0.999)-0.1*log(0.001)</code>, which gives a huge loss.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 388742,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "09/17/2018 13:54:21",
          "content": "<p>what about -log(1 - (y-p)^2) where (y-p)^2 is the absolute squared prediction error? </p>\n\n<p>-log(1 -abs(y-p)) should also work...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 388788,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/17/2018 15:29:59",
          "content": "<p>I don't like anything with <code>y-p</code> in it, because it wants the model to match the ground truth's level of uncertainty even when the model may have good reasons for being more certain.  For example, it will score <code>y=0.49, p=0.99</code> the same as <code>y=0.2, p=0.7</code> even though, in the latter case, the model is probably wrong, whereas in the former case, the model has almost a 50% chance of being right.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 389794,
          "author_name": "maxjeblick",
          "author_url": "",
          "post_date": "09/19/2018 08:33:57",
          "content": "<p>To my understanding focal loss may be well suited for this kind of problem. Focal loss is not a function in y-p and gives a lower error for almost correctly predicted samples, compared to binary crossentropy. Jeremy Howard explains this loss function in quite some detail in the fast.ai course, part 9.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 388803,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "09/17/2018 15:56:30",
      "content": "<p>My original suggestion answers the question, \"Assuming the probabilities given by the model and by ground truth represent independent Bernoulli random variables, what is the probability (expressed as a negative log likelihood) that the two variables will have the same realization?\"  Not sure if there's a name for that.  Maybe \"naive Bayesian negative log likelihood\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 388948,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "09/17/2018 21:40:10",
          "content": "<p>I am not sure whether you have tried this one \n<a href=\"https://github.com/bermanmaxim/LovaszSoftmax\">https://github.com/bermanmaxim/LovaszSoftmax</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 389409,
          "author_name": "nazimgirach",
          "author_url": "",
          "post_date": "09/18/2018 17:16:59",
          "content": "<p>Wow! This loss function looks promising on paper, will give it a try. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "388322": "If you are pursuing a segmentation approach along the lines of [Jonne's kernel][1], you probably use cross-entropy or IoU (or some combination, as in the kernel) as your loss function. But both of these approaches have problems.  IoU depends on a particular probability threshold, so it won't give the model credit for improving, for example, if a positive pixel goes from p=.3 to p=.4 (if the threshold is .5). [EDIT: It seems the Jaccard loss function, as implemented in the kernel, does in fact give credit for these kinds of improvements, but in my experience it just isn't producing very good results.]  And cross-entropy gives full veto power to the bounding box, even though, by its nature, a rectangle is only an approximation.  For example, even if the radiologist has perfectly located the bounding box, that box may have a region in the corner that doesn't look at all like pneumonia, and your model shouldn't be penalized for a negative prediction in that region.\n\nI'm trying to come up with a better loss function.  I'll start the bidding here, and see if anyone has a better idea.  Define a value `p_true`, which is equal to zero (or possibly some other value `p_min`) far outside a bounding box, and equal to `p_max` at the center of the bounding box (where `p_max` is close to 1 but not exactly 1), and the values elsewhere will vary depending on how close you are to the center of a box and whether or not you're inside the box.  (For now I'll leave open the questions of how to exactly to vary `p_true` and how to combine information from different bounding boxes.)  Then define a loss function (where `p_pred` is the predicted probability for a given pixel):<br><br>\n `-log( p_pred*p_true + (1-p_pred)*(1-p_true) )`<br><br>\nand aggregate that over pixels as you would with cross-entropy.  I think this loss function avoids the problems discussed above.  Maybe it has other problems.  Has it been used elsewhere?  Does it have a name?  Anyhow, assuming you've chosen `p_true` well, it should penalize *inconsistency* between the true and predicted probabilities.\n\n [1]: https://www.kaggle.com/jonnedtc/cnn-segmentation-connected-components",
    "388362": "what happened to your raw_iou function? it seems to fit the bill?",
    "388374": "`raw_iou` is just a metric to keep track for validation purposes. I don't optimize it.",
    "388524": "Why don't to use the ordinary cross-entropy? You can however to use it with \"soft\" (not binary) p_true. For example, 0 far away from a box, 1 near the center and something like a linear function near the borders.",
    "388708": "The problem with cross-entropy, even with soft p_true, is that it will penalize the model for being confident even when its confidence may be justified. For example, with `p_true=0.9` and `p_pred=0.999`, you get `CE=-0.9*log(0.999)-0.1*log(0.001)`, which gives a huge loss.",
    "388742": "what about -log(1 - (y-p)^2) where (y-p)^2 is the absolute squared prediction error? \n\n-log(1 -abs(y-p)) should also work...",
    "388788": "I don't like anything with `y-p` in it, because it wants the model to match the ground truth's level of uncertainty even when the model may have good reasons for being more certain.  For example, it will score `y=0.49, p=0.99` the same as `y=0.2, p=0.7` even though, in the latter case, the model is probably wrong, whereas in the former case, the model has almost a 50% chance of being right.",
    "388803": "My original suggestion answers the question, \"Assuming the probabilities given by the model and by ground truth represent independent Bernoulli random variables, what is the probability (expressed as a negative log likelihood) that the two variables will have the same realization?\"  Not sure if there's a name for that.  Maybe \"naive Bayesian negative log likelihood\"?",
    "388948": "I am not sure whether you have tried this one \nhttps://github.com/bermanmaxim/LovaszSoftmax",
    "389409": "Wow! This loss function looks promising on paper, will give it a try.",
    "389794": "To my understanding focal loss may be well suited for this kind of problem. Focal loss is not a function in y-p and gives a lower error for almost correctly predicted samples, compared to binary crossentropy. Jeremy Howard explains this loss function in quite some detail in the fast.ai course, part 9."
  },
  "source": "meta"
}