{
  "id": 233777,
  "title": "What loss function are you using ?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/233777",
  "author_name": "",
  "post_date": "2021-04-21T05:06:43.563335700Z",
  "votes": 6,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I've tried different loss functions like:</p>\n<ol>\n<li>DiceLoss</li>\n<li>BCE with DiceLoss</li>\n<li>Symmetric Lovasz</li>\n<li>Log CosH of DiceLoss<br>\nthese were used when I was using a single class for classification (putting classes = 1, in the Unet parameters), for me the cv was best when I used BCE with DiceLoss, cv was 0.93 but lb was 0.879, but for symmetric lovasz, the cv was 0.91 and lb was 0.888, all were done on same efficientnet b4 with 512x512 tile size, with medium augmentations.<br>\nAfter this I switched to classes  = 2, with CrossEntropyLoss, CrossEntropy with DiceLoss, and CrossEntropy with Log CosH of DiceLoss, and to my surprise the best worked was CrossEntropy with Log CosH of DiceLoss where my cv was 0.9166 and lb was 0.915 (my current score).<br>\nHence, just wanted to know what type of loss functions others are using, because there are number of loss functions that can be used but my GPU quota doesn't allow it :-(</li>\n</ol>",
  "messages": [
    {
      "id": "1279610",
      "postDate": "04/21/2021 05:06:43",
      "content": "<p>I've tried different loss functions like:</p>\n<ol>\n<li>DiceLoss</li>\n<li>BCE with DiceLoss</li>\n<li>Symmetric Lovasz</li>\n<li>Log CosH of DiceLoss<br>\nthese were used when I was using a single class for classification (putting classes = 1, in the Unet parameters), for me the cv was best when I used BCE with DiceLoss, cv was 0.93 but lb was 0.879, but for symmetric lovasz, the cv was 0.91 and lb was 0.888, all were done on same efficientnet b4 with 512x512 tile size, with medium augmentations.<br>\nAfter this I switched to classes  = 2, with CrossEntropyLoss, CrossEntropy with DiceLoss, and CrossEntropy with Log CosH of DiceLoss, and to my surprise the best worked was CrossEntropy with Log CosH of DiceLoss where my cv was 0.9166 and lb was 0.915 (my current score).<br>\nHence, just wanted to know what type of loss functions others are using, because there are number of loss functions that can be used but my GPU quota doesn't allow it :-(</li>\n</ol>",
      "rawMarkdown": "I've tried different loss functions like:\n1. DiceLoss\n2. BCE with DiceLoss\n3. Symmetric Lovasz\n4. Log CosH of DiceLoss\nthese were used when I was using a single class for classification (putting classes = 1, in the Unet parameters), for me the cv was best when I used BCE with DiceLoss, cv was 0.93 but lb was 0.879, but for symmetric lovasz, the cv was 0.91 and lb was 0.888, all were done on same efficientnet b4 with 512x512 tile size, with medium augmentations.\nAfter this I switched to classes  = 2, with CrossEntropyLoss, CrossEntropy with DiceLoss, and CrossEntropy with Log CosH of DiceLoss, and to my surprise the best worked was CrossEntropy with Log CosH of DiceLoss where my cv was 0.9166 and lb was 0.915 (my current score).\nHence, just wanted to know what type of loss functions others are using, because there are number of loss functions that can be used but my GPU quota doesn't allow it :-(",
      "votes": null
    },
    {
      "id": "1279771",
      "postDate": "04/21/2021 08:25:30",
      "content": "<p>I just use BCE.<br>\nI found symmetric Lovasz to converge more slowly, and combined DiceBCE to perform almost identically to plain BCE.<br>\nI was a little surprised given the fairly significant class imbalance between glomeruli and background classes, I thought a Dice/Jaccard base loss would do better.</p>",
      "rawMarkdown": "I just use BCE.\nI found symmetric Lovasz to converge more slowly, and combined DiceBCE to perform almost identically to plain BCE.\nI was a little surprised given the fairly significant class imbalance between glomeruli and background classes, I thought a Dice/Jaccard base loss would do better.",
      "votes": null
    },
    {
      "id": "1279918",
      "postDate": "04/21/2021 11:22:18",
      "content": "<p>Yeah same thought, that is why I also tried Dice Loss in my first attempt, but was surprised to see that other loss functions were working much better.</p>",
      "rawMarkdown": "Yeah same thought, that is why I also tried Dice Loss in my first attempt, but was surprised to see that other loss functions were working much better.",
      "votes": null
    },
    {
      "id": "1280804",
      "postDate": "04/22/2021 11:24:08",
      "content": "<p>Did you change nothing else other than the loss function when experimenting with the variety of loss fn's? </p>",
      "rawMarkdown": "Did you change nothing else other than the loss function when experimenting with the variety of loss fn's?",
      "votes": null
    },
    {
      "id": "1280823",
      "postDate": "04/22/2021 11:44:49",
      "content": "<p>Nope, nothing</p>",
      "rawMarkdown": "Nope, nothing",
      "votes": null
    },
    {
      "id": "1280892",
      "postDate": "04/22/2021 13:02:56",
      "content": "<p>I'm just wondering how you reduced the gap between CV and LB. I changed my Loss Fn to the BCE + Dice, but there is still a 2 - 3% drop. Did you find that changing the classes to 2 reduced the LB drop?</p>",
      "rawMarkdown": "I'm just wondering how you reduced the gap between CV and LB. I changed my Loss Fn to the BCE + Dice, but there is still a 2 - 3% drop. Did you find that changing the classes to 2 reduced the LB drop?",
      "votes": null
    },
    {
      "id": "1280930",
      "postDate": "04/22/2021 13:40:24",
      "content": "<p>Yes, for me atleast, when classes = 1, my cv was around 0.93 with BCE+Dice but lb was 0.879, then used cross entropy with classes = 2, the cv came to 0.911 with lb 0.905 then with cross entropy + log cosh of Dice, cv was 0.9166 with lb 0.915. So I think, changing class to 2 is surely decreasing the gap between lb and cv for me.</p>",
      "rawMarkdown": "Yes, for me atleast, when classes = 1, my cv was around 0.93 with BCE+Dice but lb was 0.879, then used cross entropy with classes = 2, the cv came to 0.911 with lb 0.905 then with cross entropy + log cosh of Dice, cv was 0.9166 with lb 0.915. So I think, changing class to 2 is surely decreasing the gap between lb and cv for me.",
      "votes": null
    },
    {
      "id": "1280940",
      "postDate": "04/22/2021 13:48:17",
      "content": "<p>I see! I'll check it out. This CV-LB gap is killing me.</p>",
      "rawMarkdown": "I see! I'll check it out. This CV-LB gap is killing me.",
      "votes": null
    },
    {
      "id": "1280941",
      "postDate": "04/22/2021 13:48:33",
      "content": "<p>Thanks for all the help by the way!</p>",
      "rawMarkdown": "Thanks for all the help by the way!",
      "votes": null
    },
    {
      "id": "1280961",
      "postDate": "04/22/2021 14:23:11",
      "content": "<p>No problem !</p>",
      "rawMarkdown": "No problem !",
      "votes": null
    },
    {
      "id": "1281095",
      "postDate": "04/22/2021 16:19:52",
      "content": "<p>BCE is the only one that worked for me</p>",
      "rawMarkdown": "BCE is the only one that worked for me",
      "votes": null
    },
    {
      "id": "1281199",
      "postDate": "04/22/2021 17:27:00",
      "content": "<p>How did you implement multi class Dice Loss and Dice Score. I'm a little confused?</p>",
      "rawMarkdown": "How did you implement multi class Dice Loss and Dice Score. I'm a little confused?",
      "votes": null
    },
    {
      "id": "1281434",
      "postDate": "04/22/2021 23:55:30",
      "content": "<p>thanks for your share,but i have a question,if you change the parameter'classes=1' to 'classes = 2',the model output are(batch,2,imgsize,imgsize).How to change the gt_mask size from (batch,1,imgsize,imgsize) to (batch,2,imgsize,imgsize)</p>",
      "rawMarkdown": "thanks for your share,but i have a question,if you change the parameter'classes=1' to 'classes = 2',the model output are(batch,2,imgsize,imgsize).How to change the gt_mask size from (batch,1,imgsize,imgsize) to (batch,2,imgsize,imgsize)",
      "votes": null
    },
    {
      "id": "1281530",
      "postDate": "04/23/2021 04:11:10",
      "content": "<p>` class DiceLoss(nn.Module):<br>\n    def <strong>init</strong>(self, weight=None, size_average=True):<br>\n        super(DiceLoss, self).<strong>init</strong>()</p>\n<pre><code>def forward(self, inputs, targets, smooth=1):\n\n    bs = inputs.size(0)\n\n    inputs = inputs.log_softmax(dim=1).exp()       \n\n    inputs = inputs.view(bs, 2, -1)\n    targets = targets.view(bs, -1)\n    targets = F.one_hot(targets, 2).permute(0,2,1)\n\n    intersection = (inputs * targets).sum()                            \n    dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n\n    return (1 - dice) `\n</code></pre>\n<p>`class DiceCoeff(Metric):<br>\n    def <strong>init</strong>(self, dist_sync_on_step=False, eps = 1e-7):<br>\n        super().<strong>init</strong>(dist_sync_on_step=dist_sync_on_step)</p>\n<pre><code>    self.eps = eps\n\n    self.add_state(\"inter\", default=torch.tensor(0, dtype = torch.float32), dist_reduce_fx=\"sum\")\n    self.add_state(\"union\", default=torch.tensor(0, dtype = torch.float32), dist_reduce_fx=\"sum\")\n\ndef update(self, preds: torch.Tensor, target: torch.Tensor):\n    bs = preds.size(0)\n    preds, target = preds.argmax(dim=1).view(-1), target.view(-1)\n    assert preds.shape == target.shape, f\"Expected output and target to have the same number of elements but got {len(preds)} and {len(target)}.\"\n\n    self.inter += (preds*target).float().sum().item()\n    self.union += (preds+target).float().sum().item()\n\ndef compute(self):\n    return (2.0 * self.inter)/(self.union + 1e-7)`\n</code></pre>\n<p>The Dice Coefficient metric is implemented according to pytorch lightning metric class, so change accordingly 😃</p>",
      "rawMarkdown": "` class DiceLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(DiceLoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1):\n        \n        bs = inputs.size(0)\n        \n        inputs = inputs.log_softmax(dim=1).exp()       \n        \n        inputs = inputs.view(bs, 2, -1)\n        targets = targets.view(bs, -1)\n        targets = F.one_hot(targets, 2).permute(0,2,1)\n        \n        intersection = (inputs * targets).sum()                            \n        dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n        \n        return (1 - dice) `\n\n`class DiceCoeff(Metric):\n    def __init__(self, dist_sync_on_step=False, eps = 1e-7):\n        super().__init__(dist_sync_on_step=dist_sync_on_step)\n        \n        self.eps = eps\n\n        self.add_state(\"inter\", default=torch.tensor(0, dtype = torch.float32), dist_reduce_fx=\"sum\")\n        self.add_state(\"union\", default=torch.tensor(0, dtype = torch.float32), dist_reduce_fx=\"sum\")\n\n    def update(self, preds: torch.Tensor, target: torch.Tensor):\n        bs = preds.size(0)\n        preds, target = preds.argmax(dim=1).view(-1), target.view(-1)\n        assert preds.shape == target.shape, f\"Expected output and target to have the same number of elements but got {len(preds)} and {len(target)}.\"\n\n        self.inter += (preds*target).float().sum().item()\n        self.union += (preds+target).float().sum().item()\n\n    def compute(self):\n        return (2.0 * self.inter)/(self.union + 1e-7)`\n\nThe Dice Coefficient metric is implemented according to pytorch lightning metric class, so change accordingly 😃",
      "votes": null
    },
    {
      "id": "1281536",
      "postDate": "04/23/2021 04:26:34",
      "content": "<p>To calculate the loss, what I've done is first flatten then model outputs by view(batch_size, 2,-1), then for target mask, first flatten it with view(batch_size,-1) then use one_hot as F.one_hot(target,2).permute(0,2,1), then calculate the dice loss as before</p>\n<p>`        bs = inputs.size(0)</p>\n<pre><code>    inputs = inputs.log_softmax(dim=1).exp()       \n\n    inputs = inputs.view(bs,2, -1)\n    targets = targets.view(bs, -1)\n    targets = F.one_hot(targets, 2).permute(0,2,1)\n\n    intersection = (inputs * targets).sum()                            \n    dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n\n    return (1 - dice)`\n</code></pre>",
      "rawMarkdown": "To calculate the loss, what I've done is first flatten then model outputs by view(batch_size, 2,-1), then for target mask, first flatten it with view(batch_size,-1) then use one_hot as F.one_hot(target,2).permute(0,2,1), then calculate the dice loss as before\n\n`        bs = inputs.size(0)\n        \n        inputs = inputs.log_softmax(dim=1).exp()       \n        \n        inputs = inputs.view(bs,2, -1)\n        targets = targets.view(bs, -1)\n        targets = F.one_hot(targets, 2).permute(0,2,1)\n        \n        intersection = (inputs * targets).sum()                            \n        dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n        \n        return (1 - dice)`",
      "votes": null
    },
    {
      "id": "1283055",
      "postDate": "04/24/2021 14:38:07",
      "content": "<p>Don't you mind share your implementation of CrossEntropy with Log CosH of Dice loss? I tried nn.CrossEntropy but I always get type errors. Thanks</p>",
      "rawMarkdown": "Don't you mind share your implementation of CrossEntropy with Log CosH of Dice loss? I tried nn.CrossEntropy but I always get type errors. Thanks",
      "votes": null
    },
    {
      "id": "1283137",
      "postDate": "04/24/2021 15:58:34",
      "content": "<p>use CrossEntropyFlat from fastai</p>",
      "rawMarkdown": "use CrossEntropyFlat from fastai",
      "votes": null
    },
    {
      "id": "1283141",
      "postDate": "04/24/2021 16:04:03",
      "content": "<p>Could I ask(if you don't mind sharing), how you got the LB to not diverge from your CV? My model completely fails on the d48 kidney and I'm not sure how to fix it.</p>",
      "rawMarkdown": "Could I ask(if you don't mind sharing), how you got the LB to not diverge from your CV? My model completely fails on the d48 kidney and I'm not sure how to fix it.",
      "votes": null
    },
    {
      "id": "1283145",
      "postDate": "04/24/2021 16:08:39",
      "content": "<p>I tried different thresholds for predictions, and as mentioned in the d48 discussion, there were some darker areas that had to segmented that's why I kept threshold specifically for this kidney to be low from my default threshold.<br>\nFor me the default threshold was 0.35 and for d48 I kept it 0.15, seems working for me😅</p>",
      "rawMarkdown": "I tried different thresholds for predictions, and as mentioned in the d48 discussion, there were some darker areas that had to segmented that's why I kept threshold specifically for this kidney to be low from my default threshold.\nFor me the default threshold was 0.35 and for d48 I kept it 0.15, seems working for me😅",
      "votes": null
    },
    {
      "id": "1283151",
      "postDate": "04/24/2021 16:13:14",
      "content": "<p>Ohh I see. But, wouldn't this lead to public LB overfit?</p>",
      "rawMarkdown": "Ohh I see. But, wouldn't this lead to public LB overfit?",
      "votes": null
    },
    {
      "id": "1283174",
      "postDate": "04/24/2021 16:42:37",
      "content": "<p>I don't think it would be leading to overfitting because of all the images provided that particular one had to be dealt with some special treatment, that is why we are having pseudo labels for that image :) </p>",
      "rawMarkdown": "I don't think it would be leading to overfitting because of all the images provided that particular one had to be dealt with some special treatment, that is why we are having pseudo labels for that image :)",
      "votes": null
    },
    {
      "id": "1283395",
      "postDate": "04/24/2021 22:12:29",
      "content": "<p>Even with a lower threshold, the performance seems similar. It's really strange, since all other kidneys work fine. If you don't mind sharing, do you have any tips to combat this?</p>",
      "rawMarkdown": "Even with a lower threshold, the performance seems similar. It's really strange, since all other kidneys work fine. If you don't mind sharing, do you have any tips to combat this?",
      "votes": null
    },
    {
      "id": "1283576",
      "postDate": "04/25/2021 04:17:42",
      "content": "<p>Are you using the deepflash dataset ? If not, try using that one</p>",
      "rawMarkdown": "Are you using the deepflash dataset ? If not, try using that one",
      "votes": null
    },
    {
      "id": "1287835",
      "postDate": "04/29/2021 13:01:14",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/deveshdarshan\" target=\"_blank\">@deveshdarshan</a>! Could ypu please explain what Log CosH of DiceLoss is? Do you just apply Log and Cosh functions to Dice score and sum it with CE loss?</p>",
      "rawMarkdown": "Hi, @deveshdarshan! Could ypu please explain what Log CosH of DiceLoss is? Do you just apply Log and Cosh functions to Dice score and sum it with CE loss?",
      "votes": null
    },
    {
      "id": "1289564",
      "postDate": "05/01/2021 06:22:02",
      "content": "<p>Yes yes<br>\n<code>torch.log(torch.cosh(DiceLoss)) + CELoss</code></p>",
      "rawMarkdown": "Yes yes\n`torch.log(torch.cosh(DiceLoss)) + CELoss`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1279771,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "04/21/2021 08:25:30",
      "content": "<p>I just use BCE.<br>\nI found symmetric Lovasz to converge more slowly, and combined DiceBCE to perform almost identically to plain BCE.<br>\nI was a little surprised given the fairly significant class imbalance between glomeruli and background classes, I thought a Dice/Jaccard base loss would do better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1279918,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/21/2021 11:22:18",
          "content": "<p>Yeah same thought, that is why I also tried Dice Loss in my first attempt, but was surprised to see that other loss functions were working much better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1280804,
      "author_name": "andrewshao05",
      "author_url": "",
      "post_date": "04/22/2021 11:24:08",
      "content": "<p>Did you change nothing else other than the loss function when experimenting with the variety of loss fn's? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1280823,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/22/2021 11:44:49",
          "content": "<p>Nope, nothing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280892,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "04/22/2021 13:02:56",
          "content": "<p>I'm just wondering how you reduced the gap between CV and LB. I changed my Loss Fn to the BCE + Dice, but there is still a 2 - 3% drop. Did you find that changing the classes to 2 reduced the LB drop?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280930,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/22/2021 13:40:24",
          "content": "<p>Yes, for me atleast, when classes = 1, my cv was around 0.93 with BCE+Dice but lb was 0.879, then used cross entropy with classes = 2, the cv came to 0.911 with lb 0.905 then with cross entropy + log cosh of Dice, cv was 0.9166 with lb 0.915. So I think, changing class to 2 is surely decreasing the gap between lb and cv for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280940,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "04/22/2021 13:48:17",
          "content": "<p>I see! I'll check it out. This CV-LB gap is killing me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280941,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "04/22/2021 13:48:33",
          "content": "<p>Thanks for all the help by the way!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1280961,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/22/2021 14:23:11",
          "content": "<p>No problem !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1281199,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "04/22/2021 17:27:00",
          "content": "<p>How did you implement multi class Dice Loss and Dice Score. I'm a little confused?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1281530,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/23/2021 04:11:10",
          "content": "<p>` class DiceLoss(nn.Module):<br>\n    def <strong>init</strong>(self, weight=None, size_average=True):<br>\n        super(DiceLoss, self).<strong>init</strong>()</p>\n<pre><code>def forward(self, inputs, targets, smooth=1):\n\n    bs = inputs.size(0)\n\n    inputs = inputs.log_softmax(dim=1).exp()       \n\n    inputs = inputs.view(bs, 2, -1)\n    targets = targets.view(bs, -1)\n    targets = F.one_hot(targets, 2).permute(0,2,1)\n\n    intersection = (inputs * targets).sum()                            \n    dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n\n    return (1 - dice) `\n</code></pre>\n<p>`class DiceCoeff(Metric):<br>\n    def <strong>init</strong>(self, dist_sync_on_step=False, eps = 1e-7):<br>\n        super().<strong>init</strong>(dist_sync_on_step=dist_sync_on_step)</p>\n<pre><code>    self.eps = eps\n\n    self.add_state(\"inter\", default=torch.tensor(0, dtype = torch.float32), dist_reduce_fx=\"sum\")\n    self.add_state(\"union\", default=torch.tensor(0, dtype = torch.float32), dist_reduce_fx=\"sum\")\n\ndef update(self, preds: torch.Tensor, target: torch.Tensor):\n    bs = preds.size(0)\n    preds, target = preds.argmax(dim=1).view(-1), target.view(-1)\n    assert preds.shape == target.shape, f\"Expected output and target to have the same number of elements but got {len(preds)} and {len(target)}.\"\n\n    self.inter += (preds*target).float().sum().item()\n    self.union += (preds+target).float().sum().item()\n\ndef compute(self):\n    return (2.0 * self.inter)/(self.union + 1e-7)`\n</code></pre>\n<p>The Dice Coefficient metric is implemented according to pytorch lightning metric class, so change accordingly 😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1281095,
      "author_name": "andrasferenczi",
      "author_url": "",
      "post_date": "04/22/2021 16:19:52",
      "content": "<p>BCE is the only one that worked for me</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1281434,
      "author_name": "fatliuyun",
      "author_url": "",
      "post_date": "04/22/2021 23:55:30",
      "content": "<p>thanks for your share,but i have a question,if you change the parameter'classes=1' to 'classes = 2',the model output are(batch,2,imgsize,imgsize).How to change the gt_mask size from (batch,1,imgsize,imgsize) to (batch,2,imgsize,imgsize)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1281536,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/23/2021 04:26:34",
          "content": "<p>To calculate the loss, what I've done is first flatten then model outputs by view(batch_size, 2,-1), then for target mask, first flatten it with view(batch_size,-1) then use one_hot as F.one_hot(target,2).permute(0,2,1), then calculate the dice loss as before</p>\n<p>`        bs = inputs.size(0)</p>\n<pre><code>    inputs = inputs.log_softmax(dim=1).exp()       \n\n    inputs = inputs.view(bs,2, -1)\n    targets = targets.view(bs, -1)\n    targets = F.one_hot(targets, 2).permute(0,2,1)\n\n    intersection = (inputs * targets).sum()                            \n    dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n\n    return (1 - dice)`\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1283055,
      "author_name": "nosanmark",
      "author_url": "",
      "post_date": "04/24/2021 14:38:07",
      "content": "<p>Don't you mind share your implementation of CrossEntropy with Log CosH of Dice loss? I tried nn.CrossEntropy but I always get type errors. Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1283137,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/24/2021 15:58:34",
          "content": "<p>use CrossEntropyFlat from fastai</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283141,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "04/24/2021 16:04:03",
          "content": "<p>Could I ask(if you don't mind sharing), how you got the LB to not diverge from your CV? My model completely fails on the d48 kidney and I'm not sure how to fix it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283145,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/24/2021 16:08:39",
          "content": "<p>I tried different thresholds for predictions, and as mentioned in the d48 discussion, there were some darker areas that had to segmented that's why I kept threshold specifically for this kidney to be low from my default threshold.<br>\nFor me the default threshold was 0.35 and for d48 I kept it 0.15, seems working for me😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283151,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "04/24/2021 16:13:14",
          "content": "<p>Ohh I see. But, wouldn't this lead to public LB overfit?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283174,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/24/2021 16:42:37",
          "content": "<p>I don't think it would be leading to overfitting because of all the images provided that particular one had to be dealt with some special treatment, that is why we are having pseudo labels for that image :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283395,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "04/24/2021 22:12:29",
          "content": "<p>Even with a lower threshold, the performance seems similar. It's really strange, since all other kidneys work fine. If you don't mind sharing, do you have any tips to combat this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1283576,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "04/25/2021 04:17:42",
          "content": "<p>Are you using the deepflash dataset ? If not, try using that one</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1287835,
      "author_name": "etagiev",
      "author_url": "",
      "post_date": "04/29/2021 13:01:14",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/deveshdarshan\" target=\"_blank\">@deveshdarshan</a>! Could ypu please explain what Log CosH of DiceLoss is? Do you just apply Log and Cosh functions to Dice score and sum it with CE loss?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1289564,
          "author_name": "deveshdarshan",
          "author_url": "",
          "post_date": "05/01/2021 06:22:02",
          "content": "<p>Yes yes<br>\n<code>torch.log(torch.cosh(DiceLoss)) + CELoss</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1279610": "I've tried different loss functions like:\n1. DiceLoss\n2. BCE with DiceLoss\n3. Symmetric Lovasz\n4. Log CosH of DiceLoss\nthese were used when I was using a single class for classification (putting classes = 1, in the Unet parameters), for me the cv was best when I used BCE with DiceLoss, cv was 0.93 but lb was 0.879, but for symmetric lovasz, the cv was 0.91 and lb was 0.888, all were done on same efficientnet b4 with 512x512 tile size, with medium augmentations.\nAfter this I switched to classes  = 2, with CrossEntropyLoss, CrossEntropy with DiceLoss, and CrossEntropy with Log CosH of DiceLoss, and to my surprise the best worked was CrossEntropy with Log CosH of DiceLoss where my cv was 0.9166 and lb was 0.915 (my current score).\nHence, just wanted to know what type of loss functions others are using, because there are number of loss functions that can be used but my GPU quota doesn't allow it :-(",
    "1279771": "I just use BCE.\nI found symmetric Lovasz to converge more slowly, and combined DiceBCE to perform almost identically to plain BCE.\nI was a little surprised given the fairly significant class imbalance between glomeruli and background classes, I thought a Dice/Jaccard base loss would do better.",
    "1279918": "Yeah same thought, that is why I also tried Dice Loss in my first attempt, but was surprised to see that other loss functions were working much better.",
    "1280804": "Did you change nothing else other than the loss function when experimenting with the variety of loss fn's?",
    "1280823": "Nope, nothing",
    "1280892": "I'm just wondering how you reduced the gap between CV and LB. I changed my Loss Fn to the BCE + Dice, but there is still a 2 - 3% drop. Did you find that changing the classes to 2 reduced the LB drop?",
    "1280930": "Yes, for me atleast, when classes = 1, my cv was around 0.93 with BCE+Dice but lb was 0.879, then used cross entropy with classes = 2, the cv came to 0.911 with lb 0.905 then with cross entropy + log cosh of Dice, cv was 0.9166 with lb 0.915. So I think, changing class to 2 is surely decreasing the gap between lb and cv for me.",
    "1280940": "I see! I'll check it out. This CV-LB gap is killing me.",
    "1280941": "Thanks for all the help by the way!",
    "1280961": "No problem !",
    "1281095": "BCE is the only one that worked for me",
    "1281199": "How did you implement multi class Dice Loss and Dice Score. I'm a little confused?",
    "1281434": "thanks for your share,but i have a question,if you change the parameter'classes=1' to 'classes = 2',the model output are(batch,2,imgsize,imgsize).How to change the gt_mask size from (batch,1,imgsize,imgsize) to (batch,2,imgsize,imgsize)",
    "1281530": "` class DiceLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(DiceLoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1):\n        \n        bs = inputs.size(0)\n        \n        inputs = inputs.log_softmax(dim=1).exp()       \n        \n        inputs = inputs.view(bs, 2, -1)\n        targets = targets.view(bs, -1)\n        targets = F.one_hot(targets, 2).permute(0,2,1)\n        \n        intersection = (inputs * targets).sum()                            \n        dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n        \n        return (1 - dice) `\n\n`class DiceCoeff(Metric):\n    def __init__(self, dist_sync_on_step=False, eps = 1e-7):\n        super().__init__(dist_sync_on_step=dist_sync_on_step)\n        \n        self.eps = eps\n\n        self.add_state(\"inter\", default=torch.tensor(0, dtype = torch.float32), dist_reduce_fx=\"sum\")\n        self.add_state(\"union\", default=torch.tensor(0, dtype = torch.float32), dist_reduce_fx=\"sum\")\n\n    def update(self, preds: torch.Tensor, target: torch.Tensor):\n        bs = preds.size(0)\n        preds, target = preds.argmax(dim=1).view(-1), target.view(-1)\n        assert preds.shape == target.shape, f\"Expected output and target to have the same number of elements but got {len(preds)} and {len(target)}.\"\n\n        self.inter += (preds*target).float().sum().item()\n        self.union += (preds+target).float().sum().item()\n\n    def compute(self):\n        return (2.0 * self.inter)/(self.union + 1e-7)`\n\nThe Dice Coefficient metric is implemented according to pytorch lightning metric class, so change accordingly 😃",
    "1281536": "To calculate the loss, what I've done is first flatten then model outputs by view(batch_size, 2,-1), then for target mask, first flatten it with view(batch_size,-1) then use one_hot as F.one_hot(target,2).permute(0,2,1), then calculate the dice loss as before\n\n`        bs = inputs.size(0)\n        \n        inputs = inputs.log_softmax(dim=1).exp()       \n        \n        inputs = inputs.view(bs,2, -1)\n        targets = targets.view(bs, -1)\n        targets = F.one_hot(targets, 2).permute(0,2,1)\n        \n        intersection = (inputs * targets).sum()                            \n        dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n        \n        return (1 - dice)`",
    "1283055": "Don't you mind share your implementation of CrossEntropy with Log CosH of Dice loss? I tried nn.CrossEntropy but I always get type errors. Thanks",
    "1283137": "use CrossEntropyFlat from fastai",
    "1283141": "Could I ask(if you don't mind sharing), how you got the LB to not diverge from your CV? My model completely fails on the d48 kidney and I'm not sure how to fix it.",
    "1283145": "I tried different thresholds for predictions, and as mentioned in the d48 discussion, there were some darker areas that had to segmented that's why I kept threshold specifically for this kidney to be low from my default threshold.\nFor me the default threshold was 0.35 and for d48 I kept it 0.15, seems working for me😅",
    "1283151": "Ohh I see. But, wouldn't this lead to public LB overfit?",
    "1283174": "I don't think it would be leading to overfitting because of all the images provided that particular one had to be dealt with some special treatment, that is why we are having pseudo labels for that image :)",
    "1283395": "Even with a lower threshold, the performance seems similar. It's really strange, since all other kidneys work fine. If you don't mind sharing, do you have any tips to combat this?",
    "1283576": "Are you using the deepflash dataset ? If not, try using that one",
    "1287835": "Hi, @deveshdarshan! Could ypu please explain what Log CosH of DiceLoss is? Do you just apply Log and Cosh functions to Dice score and sum it with CE loss?",
    "1289564": "Yes yes\n`torch.log(torch.cosh(DiceLoss)) + CELoss`"
  },
  "source": "meta"
}