{
  "id": 40291,
  "title": "Trying to understand the \"smoothness\" in dice loss",
  "url": "/competitions/carvana-image-masking-challenge/discussion/40291",
  "author_name": "Tuatini GODARD",
  "post_date": "2017-09-30T14:57:40.384000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello,\nDuring this competition I used @Heng CherKeng <code>SoftDiceLoss</code> class as my loss function combined with cross-entropy but now I'm trying map the code to the maths. I was wondering why he used a <code>smooth</code> factor here:</p>\n\n<pre><code>class SoftDiceLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(SoftDiceLoss, self).__init__()\n\n    def forward(self, logits, targets):\n        smooth = 1\n        num = targets.size(0)\n        probs = F.sigmoid(logits)\n        m1 = probs.view(num, -1)\n        m2 = targets.view(num, -1)\n        intersection = (m1 * m2)\n\n        score = 2. * (intersection.sum(1) + smooth) / (m1.sum(1) + m2.sum(1) + smooth)\n        score = 1 - score.sum() / num\n        return score\n</code></pre>\n\n<p>Any idea? Thanks</p>",
  "messages": [
    {
      "id": 225978,
      "postDate": "2017-09-30T15:28:16.150Z",
      "content": "<p>For back propagation. If the prediction is hard threshold to 0 and 1, it is difficult to back \npropagate the dice loss gradient. Try to compute the math formula for the gradient for the soft and hard cass</p>",
      "rawMarkdown": "For back propagation. If the prediction is hard threshold to 0 and 1, it is difficult to back \npropagate the dice loss gradient. Try to compute the math formula for the gradient for the soft and hard cass",
      "votes": 4,
      "replies": [
        {
          "id": 225979,
          "postDate": "2017-09-30T15:30:00.557Z",
          "content": "<p>Thanks a lot Heng! I really learnt a lot from you.</p>",
          "rawMarkdown": "Thanks a lot Heng! I really learnt a lot from you.",
          "votes": 1
        },
        {
          "id": 225980,
          "postDate": "2017-09-30T15:34:25.310Z",
          "content": "<p>Tbh I'm not 100% sure I really understand what you are saying. Do you mind giving me a little example? Thank you</p>",
          "rawMarkdown": "Tbh I'm not 100% sure I really understand what you are saying. Do you mind giving me a little example? Thank you",
          "votes": 1
        }
      ]
    },
    {
      "id": 225971,
      "postDate": "2017-09-30T14:57:40.383Z",
      "content": "<p>Hello,\nDuring this competition I used @Heng CherKeng <code>SoftDiceLoss</code> class as my loss function combined with cross-entropy but now I'm trying map the code to the maths. I was wondering why he used a <code>smooth</code> factor here:</p>\n\n<pre><code>class SoftDiceLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(SoftDiceLoss, self).__init__()\n\n    def forward(self, logits, targets):\n        smooth = 1\n        num = targets.size(0)\n        probs = F.sigmoid(logits)\n        m1 = probs.view(num, -1)\n        m2 = targets.view(num, -1)\n        intersection = (m1 * m2)\n\n        score = 2. * (intersection.sum(1) + smooth) / (m1.sum(1) + m2.sum(1) + smooth)\n        score = 1 - score.sum() / num\n        return score\n</code></pre>\n\n<p>Any idea? Thanks</p>",
      "rawMarkdown": "Hello,\nDuring this competition I used @Heng CherKeng `SoftDiceLoss` class as my loss function combined with cross-entropy but now I'm trying map the code to the maths. I was wondering why he used a `smooth` factor here:\n\n    class SoftDiceLoss(nn.Module):\n        def __init__(self, weight=None, size_average=True):\n            super(SoftDiceLoss, self).__init__()\n\n        def forward(self, logits, targets):\n            smooth = 1\n            num = targets.size(0)\n            probs = F.sigmoid(logits)\n            m1 = probs.view(num, -1)\n            m2 = targets.view(num, -1)\n            intersection = (m1 * m2)\n\n            score = 2. * (intersection.sum(1) + smooth) / (m1.sum(1) + m2.sum(1) + smooth)\n            score = 1 - score.sum() / num\n            return score\n\nAny idea? Thanks",
      "votes": 1
    },
    {
      "id": 872865,
      "postDate": "2020-06-03T15:07:54.880Z",
      "content": "<p>I don't know if this is still relevant, but that simple addition actually smooths out the loss function, making it differentiable.\nImagine a situation where, where targets are 0 in binary classification. Then score is always 0, and loss is always constant for network to predict for zeros. Hence, it cant optimize.\nThat simple addition prevents this from happening.</p>",
      "rawMarkdown": "I don't know if this is still relevant, but that simple addition actually smooths out the loss function, making it differentiable.\nImagine a situation where, where targets are 0 in binary classification. Then score is always 0, and loss is always constant for network to predict for zeros. Hence, it cant optimize.\nThat simple addition prevents this from happening."
    },
    {
      "id": 848197,
      "postDate": "2020-05-14T19:49:51.627Z",
      "content": "<p>From the function, the denominator of Dice score is sum of 'probs' and 'targets' which can sometimes be zero. A 'smooth' value is added to avoid this division by zero. </p>",
      "rawMarkdown": "From the function, the denominator of Dice score is sum of 'probs' and 'targets' which can sometimes be zero. A 'smooth' value is added to avoid this division by zero. "
    },
    {
      "id": 424514,
      "postDate": "2018-11-20T08:50:17.063Z",
      "content": "<p>I think smoothing can reduce Overfitting, because the whole coefficient value is made larger, the loss is smaller, and convergence can be achieved faster, avoiding too many training iterations.\nCheck out a this <a href=\"https://github.com/pytorch/pytorch/issues/1249#issuecomment-337999895\">PR</a>.</p>",
      "rawMarkdown": "I think smoothing can reduce Overfitting, because the whole coefficient value is made larger, the loss is smaller, and convergence can be achieved faster, avoiding too many training iterations.\nCheck out a this [PR](https://github.com/pytorch/pytorch/issues/1249#issuecomment-337999895)."
    },
    {
      "id": 364760,
      "postDate": "2018-08-01T07:26:30.693Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 225978,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2017-09-30T15:28:16.150000",
      "content": "<p>For back propagation. If the prediction is hard threshold to 0 and 1, it is difficult to back \npropagate the dice loss gradient. Try to compute the math formula for the gradient for the soft and hard cass</p>",
      "votes": 4,
      "replies": [
        {
          "id": 225979,
          "author_name": "Tuatini GODARD",
          "author_url": "",
          "post_date": "2017-09-30T15:30:00.557000",
          "content": "<p>Thanks a lot Heng! I really learnt a lot from you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 225980,
          "author_name": "Tuatini GODARD",
          "author_url": "",
          "post_date": "2017-09-30T15:34:25.310000",
          "content": "<p>Tbh I'm not 100% sure I really understand what you are saying. Do you mind giving me a little example? Thank you</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 872865,
      "author_name": "Saahil Islam",
      "author_url": "",
      "post_date": "2020-06-03T15:07:54.880000",
      "content": "<p>I don't know if this is still relevant, but that simple addition actually smooths out the loss function, making it differentiable.\nImagine a situation where, where targets are 0 in binary classification. Then score is always 0, and loss is always constant for network to predict for zeros. Hence, it cant optimize.\nThat simple addition prevents this from happening.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 848197,
      "author_name": "Ravi Gunti",
      "author_url": "",
      "post_date": "2020-05-14T19:49:51.627000",
      "content": "<p>From the function, the denominator of Dice score is sum of 'probs' and 'targets' which can sometimes be zero. A 'smooth' value is added to avoid this division by zero. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 424514,
      "author_name": "AnkurShukla",
      "author_url": "",
      "post_date": "2018-11-20T08:50:17.063000",
      "content": "<p>I think smoothing can reduce Overfitting, because the whole coefficient value is made larger, the loss is smaller, and convergence can be achieved faster, avoiding too many training iterations.\nCheck out a this <a href=\"https://github.com/pytorch/pytorch/issues/1249#issuecomment-337999895\">PR</a>.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 364760,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-01T07:26:30.693000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "225978": "For back propagation. If the prediction is hard threshold to 0 and 1, it is difficult to back \npropagate the dice loss gradient. Try to compute the math formula for the gradient for the soft and hard cass",
    "225971": "Hello,\nDuring this competition I used @Heng CherKeng `SoftDiceLoss` class as my loss function combined with cross-entropy but now I'm trying map the code to the maths. I was wondering why he used a `smooth` factor here:\n\n    class SoftDiceLoss(nn.Module):\n        def __init__(self, weight=None, size_average=True):\n            super(SoftDiceLoss, self).__init__()\n\n        def forward(self, logits, targets):\n            smooth = 1\n            num = targets.size(0)\n            probs = F.sigmoid(logits)\n            m1 = probs.view(num, -1)\n            m2 = targets.view(num, -1)\n            intersection = (m1 * m2)\n\n            score = 2. * (intersection.sum(1) + smooth) / (m1.sum(1) + m2.sum(1) + smooth)\n            score = 1 - score.sum() / num\n            return score\n\nAny idea? Thanks",
    "872865": "I don't know if this is still relevant, but that simple addition actually smooths out the loss function, making it differentiable.\nImagine a situation where, where targets are 0 in binary classification. Then score is always 0, and loss is always constant for network to predict for zeros. Hence, it cant optimize.\nThat simple addition prevents this from happening.",
    "848197": "From the function, the denominator of Dice score is sum of 'probs' and 'targets' which can sometimes be zero. A 'smooth' value is added to avoid this division by zero. ",
    "424514": "I think smoothing can reduce Overfitting, because the whole coefficient value is made larger, the loss is smaller, and convergence can be achieved faster, avoiding too many training iterations.\nCheck out a this [PR](https://github.com/pytorch/pytorch/issues/1249#issuecomment-337999895).",
    "364760": ""
  }
}