{
  "id": 212978,
  "title": "FocalLoss or BCELoss? (share my FocalLoss pytorch version)",
  "url": "/competitions/rfcx-species-audio-detection/discussion/212978",
  "author_name": "Hao",
  "post_date": "2021-01-21T01:51:19.643000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I see most people choose FocalLoss or BCELoss for this competition, ideally FocalLoss has advantage on imbalance targets(via APLPHA) and hard samples(via GAMMA), but in my pratice so far I did not see obvious improvement when using FocalLoss. So I`m wondering which loss is best for this competition, or anything wrong in my FocalLoss implementation?</p>\n<p>I`m using below version 2 with default hyperparamters value: ALPHA = 0.75, GAMMA = 2, which is told to be the best combination according to the Focal Loss paper. </p>\n<p>And I also have doubt in Version 1 below which is  an alpha-modified version which set same ALPHA value on both positive and negative samples, I am wondering what`s the reason behind this.</p>\n<h1>Version 1 - From: <a href=\"https://www.kaggle.com/bigironsphere/loss-function-library-keras-pytorch/notebook\" target=\"_blank\">https://www.kaggle.com/bigironsphere/loss-function-library-keras-pytorch/notebook</a></h1>\n<pre><code>ALPHA = 0.8\nGAMMA = 2\nclass FocalLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(FocalLoss, self).__init__()\n\n    def forward(self, inputs, targets, alpha=ALPHA, gamma=GAMMA, smooth=1):\n        # comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = torch.sigmoid(inputs)\n\n        # flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n\n        # first compute binary cross-entropy\n        _BCE = F.binary_cross_entropy(inputs, targets, reduction='mean')\n        _BCE_EXP = torch.exp(-_BCE)\n        focal_loss = alpha * (1 - _BCE_EXP) ** gamma * _BCE\n\n        return focal_loss\n</code></pre>\n<h1>Version 2 - From: <a href=\"https://github.com/open-mmlab/mmdetection\" target=\"_blank\">https://github.com/open-mmlab/mmdetection</a></h1>\n<pre><code>class FocalLoss(nn.Module):  # gamma=0, alpha=0.75 or gamma=2, alpha=0.25\n    def __init__(self, alpha=0.25, gamma=2.0, reduction='mean', label_smoothing=1):\n        super(FocalLoss, self).__init__()\n        self.gamma = gamma\n        self.alpha = alpha\n        self.reduction = reduction\n        self.label_smoothing = label_smoothing\n\n    def forward(self, pred, target): \n        pred_sigmoid = pred.sigmoid()\n        target = target.type_as(pred)  # to torch.float32\n        pt = (1 - pred_sigmoid) * target + pred_sigmoid * (1 - target)\n        focal_weight = (self.alpha * target + (1 - self.alpha) * (1 - target)) * pt.pow(self.gamma)\n        loss = F.binary_cross_entropy_with_logits(pred, target, reduction='none') * focal_weight\n        # loss = weight_reduce_loss(loss, reduction)\n        loss = loss.mean()\n\n        return loss\n</code></pre>",
  "messages": [
    {
      "id": 1162097,
      "postDate": "2021-01-21T01:51:19.643Z",
      "content": "<p>I see most people choose FocalLoss or BCELoss for this competition, ideally FocalLoss has advantage on imbalance targets(via APLPHA) and hard samples(via GAMMA), but in my pratice so far I did not see obvious improvement when using FocalLoss. So I`m wondering which loss is best for this competition, or anything wrong in my FocalLoss implementation?</p>\n<p>I`m using below version 2 with default hyperparamters value: ALPHA = 0.75, GAMMA = 2, which is told to be the best combination according to the Focal Loss paper. </p>\n<p>And I also have doubt in Version 1 below which is  an alpha-modified version which set same ALPHA value on both positive and negative samples, I am wondering what`s the reason behind this.</p>\n<h1>Version 1 - From: <a href=\"https://www.kaggle.com/bigironsphere/loss-function-library-keras-pytorch/notebook\" target=\"_blank\">https://www.kaggle.com/bigironsphere/loss-function-library-keras-pytorch/notebook</a></h1>\n<pre><code>ALPHA = 0.8\nGAMMA = 2\nclass FocalLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(FocalLoss, self).__init__()\n\n    def forward(self, inputs, targets, alpha=ALPHA, gamma=GAMMA, smooth=1):\n        # comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = torch.sigmoid(inputs)\n\n        # flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n\n        # first compute binary cross-entropy\n        _BCE = F.binary_cross_entropy(inputs, targets, reduction='mean')\n        _BCE_EXP = torch.exp(-_BCE)\n        focal_loss = alpha * (1 - _BCE_EXP) ** gamma * _BCE\n\n        return focal_loss\n</code></pre>\n<h1>Version 2 - From: <a href=\"https://github.com/open-mmlab/mmdetection\" target=\"_blank\">https://github.com/open-mmlab/mmdetection</a></h1>\n<pre><code>class FocalLoss(nn.Module):  # gamma=0, alpha=0.75 or gamma=2, alpha=0.25\n    def __init__(self, alpha=0.25, gamma=2.0, reduction='mean', label_smoothing=1):\n        super(FocalLoss, self).__init__()\n        self.gamma = gamma\n        self.alpha = alpha\n        self.reduction = reduction\n        self.label_smoothing = label_smoothing\n\n    def forward(self, pred, target): \n        pred_sigmoid = pred.sigmoid()\n        target = target.type_as(pred)  # to torch.float32\n        pt = (1 - pred_sigmoid) * target + pred_sigmoid * (1 - target)\n        focal_weight = (self.alpha * target + (1 - self.alpha) * (1 - target)) * pt.pow(self.gamma)\n        loss = F.binary_cross_entropy_with_logits(pred, target, reduction='none') * focal_weight\n        # loss = weight_reduce_loss(loss, reduction)\n        loss = loss.mean()\n\n        return loss\n</code></pre>",
      "rawMarkdown": "I see most people choose FocalLoss or BCELoss for this competition, ideally FocalLoss has advantage on imbalance targets(via APLPHA) and hard samples(via GAMMA), but in my pratice so far I did not see obvious improvement when using FocalLoss. So I`m wondering which loss is best for this competition, or anything wrong in my FocalLoss implementation?\n\nI`m using below version 2 with default hyperparamters value: ALPHA = 0.75, GAMMA = 2, which is told to be the best combination according to the Focal Loss paper. \n\nAnd I also have doubt in Version 1 below which is  an alpha-modified version which set same ALPHA value on both positive and negative samples, I am wondering what`s the reason behind this.\n\n# Version 1 - From: [https://www.kaggle.com/bigironsphere/loss-function-library-keras-pytorch/notebook](https://www.kaggle.com/bigironsphere/loss-function-library-keras-pytorch/notebook) \n```\nALPHA = 0.8\nGAMMA = 2\nclass FocalLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(FocalLoss, self).__init__()\n\n    def forward(self, inputs, targets, alpha=ALPHA, gamma=GAMMA, smooth=1):\n        # comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = torch.sigmoid(inputs)\n\n        # flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n\n        # first compute binary cross-entropy\n        _BCE = F.binary_cross_entropy(inputs, targets, reduction='mean')\n        _BCE_EXP = torch.exp(-_BCE)\n        focal_loss = alpha * (1 - _BCE_EXP) ** gamma * _BCE\n\n        return focal_loss\n```\n# Version 2 - From: [https://github.com/open-mmlab/mmdetection](https://github.com/open-mmlab/mmdetection)\n```\nclass FocalLoss(nn.Module):  # gamma=0, alpha=0.75 or gamma=2, alpha=0.25\n    def __init__(self, alpha=0.25, gamma=2.0, reduction='mean', label_smoothing=1):\n        super(FocalLoss, self).__init__()\n        self.gamma = gamma\n        self.alpha = alpha\n        self.reduction = reduction\n        self.label_smoothing = label_smoothing\n\n    def forward(self, pred, target): \n        pred_sigmoid = pred.sigmoid()\n        target = target.type_as(pred)  # to torch.float32\n        pt = (1 - pred_sigmoid) * target + pred_sigmoid * (1 - target)\n        focal_weight = (self.alpha * target + (1 - self.alpha) * (1 - target)) * pt.pow(self.gamma)\n        loss = F.binary_cross_entropy_with_logits(pred, target, reduction='none') * focal_weight\n        # loss = weight_reduce_loss(loss, reduction)\n        loss = loss.mean()\n\n        return loss\n```",
      "votes": 4
    },
    {
      "id": 1163464,
      "postDate": "2021-01-21T17:26:11.847Z",
      "content": "<p>Your second version looks very similar to the one I posted. Thanks for sharing it.  I think the first one is wrong.</p>",
      "rawMarkdown": "Your second version looks very similar to the one I posted. Thanks for sharing it.  I think the first one is wrong.",
      "replies": [
        {
          "id": 1163865,
          "postDate": "2021-01-22T02:07:20.910Z",
          "content": "<p>Version 2 somehow did not use label smoothing?👀</p>",
          "rawMarkdown": "Version 2 somehow did not use label smoothing?👀"
        },
        {
          "id": 1163869,
          "postDate": "2021-01-22T02:19:04.263Z",
          "content": "<p>Not yet, will try add label smoothing later.  My thought is that Focal Loss already has penalty on easy sample s(high confident predit), is label smoothing duplicated with such penalty?</p>",
          "rawMarkdown": "Not yet, will try add label smoothing later.  My thought is that Focal Loss already has penalty on easy sample s(high confident predit), is label smoothing duplicated with such penalty?"
        }
      ]
    },
    {
      "id": 1173954,
      "postDate": "2021-01-28T07:44:14.980Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1163464,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-01-21T17:26:11.847000",
      "content": "<p>Your second version looks very similar to the one I posted. Thanks for sharing it.  I think the first one is wrong.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1163865,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2021-01-22T02:07:20.910000",
          "content": "<p>Version 2 somehow did not use label smoothing?👀</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1163869,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-01-22T02:19:04.263000",
          "content": "<p>Not yet, will try add label smoothing later.  My thought is that Focal Loss already has penalty on easy sample s(high confident predit), is label smoothing duplicated with such penalty?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1173954,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-28T07:44:14.980000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1162097": "I see most people choose FocalLoss or BCELoss for this competition, ideally FocalLoss has advantage on imbalance targets(via APLPHA) and hard samples(via GAMMA), but in my pratice so far I did not see obvious improvement when using FocalLoss. So I`m wondering which loss is best for this competition, or anything wrong in my FocalLoss implementation?\n\nI`m using below version 2 with default hyperparamters value: ALPHA = 0.75, GAMMA = 2, which is told to be the best combination according to the Focal Loss paper. \n\nAnd I also have doubt in Version 1 below which is  an alpha-modified version which set same ALPHA value on both positive and negative samples, I am wondering what`s the reason behind this.\n\n# Version 1 - From: [https://www.kaggle.com/bigironsphere/loss-function-library-keras-pytorch/notebook](https://www.kaggle.com/bigironsphere/loss-function-library-keras-pytorch/notebook) \n```\nALPHA = 0.8\nGAMMA = 2\nclass FocalLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(FocalLoss, self).__init__()\n\n    def forward(self, inputs, targets, alpha=ALPHA, gamma=GAMMA, smooth=1):\n        # comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = torch.sigmoid(inputs)\n\n        # flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n\n        # first compute binary cross-entropy\n        _BCE = F.binary_cross_entropy(inputs, targets, reduction='mean')\n        _BCE_EXP = torch.exp(-_BCE)\n        focal_loss = alpha * (1 - _BCE_EXP) ** gamma * _BCE\n\n        return focal_loss\n```\n# Version 2 - From: [https://github.com/open-mmlab/mmdetection](https://github.com/open-mmlab/mmdetection)\n```\nclass FocalLoss(nn.Module):  # gamma=0, alpha=0.75 or gamma=2, alpha=0.25\n    def __init__(self, alpha=0.25, gamma=2.0, reduction='mean', label_smoothing=1):\n        super(FocalLoss, self).__init__()\n        self.gamma = gamma\n        self.alpha = alpha\n        self.reduction = reduction\n        self.label_smoothing = label_smoothing\n\n    def forward(self, pred, target): \n        pred_sigmoid = pred.sigmoid()\n        target = target.type_as(pred)  # to torch.float32\n        pt = (1 - pred_sigmoid) * target + pred_sigmoid * (1 - target)\n        focal_weight = (self.alpha * target + (1 - self.alpha) * (1 - target)) * pt.pow(self.gamma)\n        loss = F.binary_cross_entropy_with_logits(pred, target, reduction='none') * focal_weight\n        # loss = weight_reduce_loss(loss, reduction)\n        loss = loss.mean()\n\n        return loss\n```",
    "1163464": "Your second version looks very similar to the one I posted. Thanks for sharing it.  I think the first one is wrong.",
    "1173954": ""
  }
}