{
  "id": 128665,
  "title": "[pytorch] focal loss + ohem implementation",
  "url": "/competitions/bengaliai-cv19/discussion/128665",
  "author_name": "",
  "post_date": "2020-02-02T09:26:56.894026500Z",
  "votes": 23,
  "comment_count": 7,
  "views": 0,
  "content": "<p>```python</p>\n\n<p>import torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.autograd import Variable\ndevice = torch.device('cuda:0' if torch.cuda.is_available() else 'cpu')</p>\n\n<p>class FocalLoss(nn.Module):\n    def <strong>init</strong>(self, class_num, alpha=None, gamma=2, size_average=True):\n        super(FocalLoss, self).<strong>init</strong>()\n        if alpha is None:\n            self.alpha = Variable(torch.ones(class_num, 1))\n        else:\n            if isinstance(alpha, Variable):\n                self.alpha = alpha\n            else:\n                self.alpha = Variable(alpha)\n        self.gamma = gamma\n        self.class_num = class_num\n        self.size_average = size_average</p>\n\n<pre><code>def forward(self, inputs, targets):\n    N = inputs.size(0)\n    C = inputs.size(1)\n    P = F.softmax(inputs)\n\n    class_mask = inputs.data.new(N, C).fill_(0)\n    class_mask = Variable(class_mask)\n    ids = targets.view(-1, 1)\n    class_mask.scatter_(1, ids.data, 1.)\n    #print(class_mask)\n\n\n    if inputs.is_cuda and not self.alpha.is_cuda:\n        self.alpha = self.alpha.to(device)\n    alpha = self.alpha[ids.data.view(-1)]\n\n    probs = (P*class_mask).sum(1).view(-1,1)\n\n    log_p = probs.log()\n    #print('probs size= {}'.format(probs.size()))\n    #print(probs)\n\n    batch_loss = -alpha*(torch.pow((1-probs), self.gamma))*log_p \n    #print('-----bacth_loss------')\n    #print(batch_loss)\n\n\n    if self.size_average:\n        loss = batch_loss.mean()\n    else:\n        loss = batch_loss.sum()\n    return loss\n</code></pre>\n\n<p>F1 = FocalLoss(168)\nF2 = FocalLoss(11)\nF3 = FocalLoss(7)</p>\n\n<p>```</p>\n\n<p>```python\ndef ohem_loss( rate, cls_pred, cls_target ):\n    batch_size = cls_pred.size(0) \n    # ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n    ohem_cls_loss = F1(cls_pred, cls_target)</p>\n\n<pre><code>sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\nkeep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\nif keep_num &amp;lt; sorted_ohem_loss.size()[0]:\n    keep_idx_cuda = idx[:keep_num]\n    ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\ncls_loss = ohem_cls_loss.sum() / keep_num\nreturn cls_loss\n</code></pre>\n\n<p>```</p>\n\n<p>reference：\n（1）<a href=\"https://github.com/clcarwin/focal_loss_pytorch/blob/master/focalloss.py\">https://github.com/clcarwin/focal_loss_pytorch/blob/master/focalloss.py</a>\n（2）<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a> \n（3）<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128592\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128592</a></p>",
  "messages": [
    {
      "id": "734951",
      "postDate": "02/02/2020 09:26:56",
      "content": "<p>```python</p>\n\n<p>import torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.autograd import Variable\ndevice = torch.device('cuda:0' if torch.cuda.is_available() else 'cpu')</p>\n\n<p>class FocalLoss(nn.Module):\n    def <strong>init</strong>(self, class_num, alpha=None, gamma=2, size_average=True):\n        super(FocalLoss, self).<strong>init</strong>()\n        if alpha is None:\n            self.alpha = Variable(torch.ones(class_num, 1))\n        else:\n            if isinstance(alpha, Variable):\n                self.alpha = alpha\n            else:\n                self.alpha = Variable(alpha)\n        self.gamma = gamma\n        self.class_num = class_num\n        self.size_average = size_average</p>\n\n<pre><code>def forward(self, inputs, targets):\n    N = inputs.size(0)\n    C = inputs.size(1)\n    P = F.softmax(inputs)\n\n    class_mask = inputs.data.new(N, C).fill_(0)\n    class_mask = Variable(class_mask)\n    ids = targets.view(-1, 1)\n    class_mask.scatter_(1, ids.data, 1.)\n    #print(class_mask)\n\n\n    if inputs.is_cuda and not self.alpha.is_cuda:\n        self.alpha = self.alpha.to(device)\n    alpha = self.alpha[ids.data.view(-1)]\n\n    probs = (P*class_mask).sum(1).view(-1,1)\n\n    log_p = probs.log()\n    #print('probs size= {}'.format(probs.size()))\n    #print(probs)\n\n    batch_loss = -alpha*(torch.pow((1-probs), self.gamma))*log_p \n    #print('-----bacth_loss------')\n    #print(batch_loss)\n\n\n    if self.size_average:\n        loss = batch_loss.mean()\n    else:\n        loss = batch_loss.sum()\n    return loss\n</code></pre>\n\n<p>F1 = FocalLoss(168)\nF2 = FocalLoss(11)\nF3 = FocalLoss(7)</p>\n\n<p>```</p>\n\n<p>```python\ndef ohem_loss( rate, cls_pred, cls_target ):\n    batch_size = cls_pred.size(0) \n    # ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n    ohem_cls_loss = F1(cls_pred, cls_target)</p>\n\n<pre><code>sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\nkeep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\nif keep_num &amp;lt; sorted_ohem_loss.size()[0]:\n    keep_idx_cuda = idx[:keep_num]\n    ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\ncls_loss = ohem_cls_loss.sum() / keep_num\nreturn cls_loss\n</code></pre>\n\n<p>```</p>\n\n<p>reference：\n（1）<a href=\"https://github.com/clcarwin/focal_loss_pytorch/blob/master/focalloss.py\">https://github.com/clcarwin/focal_loss_pytorch/blob/master/focalloss.py</a>\n（2）<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a> \n（3）<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128592\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128592</a></p>",
      "rawMarkdown": "```python\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.autograd import Variable\ndevice = torch.device('cuda:0' if torch.cuda.is_available() else 'cpu')\n\nclass FocalLoss(nn.Module):\n    def __init__(self, class_num, alpha=None, gamma=2, size_average=True):\n        super(FocalLoss, self).__init__()\n        if alpha is None:\n            self.alpha = Variable(torch.ones(class_num, 1))\n        else:\n            if isinstance(alpha, Variable):\n                self.alpha = alpha\n            else:\n                self.alpha = Variable(alpha)\n        self.gamma = gamma\n        self.class_num = class_num\n        self.size_average = size_average\n\n    def forward(self, inputs, targets):\n        N = inputs.size(0)\n        C = inputs.size(1)\n        P = F.softmax(inputs)\n\n        class_mask = inputs.data.new(N, C).fill_(0)\n        class_mask = Variable(class_mask)\n        ids = targets.view(-1, 1)\n        class_mask.scatter_(1, ids.data, 1.)\n        #print(class_mask)\n\n\n        if inputs.is_cuda and not self.alpha.is_cuda:\n            self.alpha = self.alpha.to(device)\n        alpha = self.alpha[ids.data.view(-1)]\n\n        probs = (P*class_mask).sum(1).view(-1,1)\n\n        log_p = probs.log()\n        #print('probs size= {}'.format(probs.size()))\n        #print(probs)\n\n        batch_loss = -alpha*(torch.pow((1-probs), self.gamma))*log_p \n        #print('-----bacth_loss------')\n        #print(batch_loss)\n\n\n        if self.size_average:\n            loss = batch_loss.mean()\n        else:\n            loss = batch_loss.sum()\n        return loss\n\nF1 = FocalLoss(168)\nF2 = FocalLoss(11)\nF3 = FocalLoss(7)\n\n```\n\n```python\ndef ohem_loss( rate, cls_pred, cls_target ):\n    batch_size = cls_pred.size(0) \n    # ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n    ohem_cls_loss = F1(cls_pred, cls_target)\n\n    sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\n    keep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\n    if keep_num &lt; sorted_ohem_loss.size()[0]:\n        keep_idx_cuda = idx[:keep_num]\n        ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\n    cls_loss = ohem_cls_loss.sum() / keep_num\n    return cls_loss\n```\n\nreference：\n（1）https://github.com/clcarwin/focal_loss_pytorch/blob/master/focalloss.py\n（2）https://www.kaggle.com/c/bengaliai-cv19/discussion/128637 \n（3）https://www.kaggle.com/c/bengaliai-cv19/discussion/128592",
      "votes": null
    },
    {
      "id": "735228",
      "postDate": "02/02/2020 18:46:06",
      "content": "<p>Have you tried focal loss? It seems like it didnt give me better results :\\ what about you?</p>",
      "rawMarkdown": "Have you tried focal loss? It seems like it didnt give me better results :\\ what about you?",
      "votes": null
    },
    {
      "id": "737339",
      "postDate": "02/05/2020 07:52:19",
      "content": "<p>ohem loss didn't give me better results either. I'm afraid they are more useful in a object detection task than in a classification task.</p>",
      "rawMarkdown": "ohem loss didn't give me better results either. I'm afraid they are more useful in a object detection task than in a classification task.",
      "votes": null
    },
    {
      "id": "737465",
      "postDate": "02/05/2020 11:29:22",
      "content": "<p>Thanks for sharing <a href=\"/machinelp\">@machinelp</a> </p>",
      "rawMarkdown": "Thanks for sharing @machinelp",
      "votes": null
    },
    {
      "id": "740822",
      "postDate": "02/09/2020 20:58:29",
      "content": "<p>Thanks for sharing! I have a few questions regarding the implementation and the paper. </p>\n\n<p>1) I'm not very familiar with pytorch Variable, seems like it can do autograd (although now seems like its deprecated, and we can pass in requires_grad=True), so based on the code, seems like you are setting <code>alpha</code> to require gradient i.e. to be backproped and updated, Is that correct? In the paper they only showed binary example and they did bunch experiments using a range of fixed <code>alpha</code> (if i understood correctly), e.g. gamma=2 and alpha=0.25 they found that works the best, they didnt mention to train <code>alpha</code> (although i imagine one could, and initialized as what u showed here to be ones)</p>\n\n<p>2) <code>class_mask = Variable(class_mask)</code> are you also setting class_mask to be backproped/updated? (sorry again if im not understanding something silly here). And to just confirm, the class_mask here is to just make the computation of ce more efficient right? as it does not need to sum up the terms that multiplied by zeros in the mask.</p>\n\n<p>EDIT: Just realized in order for alpha to be trainable, it also needs to be passed into the optimizer, and the code is not doing it, but also i wonder if simply making it trainable, the model would just learn a naive alpha e.g. +inf, since in the formula the alpha is just a constant multiplier on the loss.</p>",
      "rawMarkdown": "Thanks for sharing! I have a few questions regarding the implementation and the paper. \n\n1) I'm not very familiar with pytorch Variable, seems like it can do autograd (although now seems like its deprecated, and we can pass in requires_grad=True), so based on the code, seems like you are setting `alpha` to require gradient i.e. to be backproped and updated, Is that correct? In the paper they only showed binary example and they did bunch experiments using a range of fixed `alpha` (if i understood correctly), e.g. gamma=2 and alpha=0.25 they found that works the best, they didnt mention to train `alpha` (although i imagine one could, and initialized as what u showed here to be ones)\n\n2) `class_mask = Variable(class_mask)` are you also setting class_mask to be backproped/updated? (sorry again if im not understanding something silly here). And to just confirm, the class_mask here is to just make the computation of ce more efficient right? as it does not need to sum up the terms that multiplied by zeros in the mask.\n\nEDIT: Just realized in order for alpha to be trainable, it also needs to be passed into the optimizer, and the code is not doing it, but also i wonder if simply making it trainable, the model would just learn a naive alpha e.g. +inf, since in the formula the alpha is just a constant multiplier on the loss.",
      "votes": null
    },
    {
      "id": "744748",
      "postDate": "02/13/2020 05:17:07",
      "content": "<p>Hi <a href=\"/machinelp\">@machinelp</a> . Your Focal loss is returning either average or sum of the batch_loss, i.e, a single value for each batch. But for ohem we want loss for each image. I think focal loss here shld return only batch_loss and not batch_loss.mean() or batch_loss.sum(). Am I right ? </p>",
      "rawMarkdown": "Hi @machinelp . Your Focal loss is returning either average or sum of the batch_loss, i.e, a single value for each batch. But for ohem we want loss for each image. I think focal loss here shld return only batch_loss and not batch_loss.mean() or batch_loss.sum(). Am I right ?",
      "votes": null
    },
    {
      "id": "749836",
      "postDate": "02/18/2020 23:41:21",
      "content": "<p>thx</p>",
      "rawMarkdown": "thx",
      "votes": null
    },
    {
      "id": "761205",
      "postDate": "03/02/2020 08:53:01",
      "content": "<p>FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133322\">https://www.kaggle.com/c/bengaliai-cv19/discussion/133322</a></p>",
      "rawMarkdown": "FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：https://www.kaggle.com/c/bengaliai-cv19/discussion/133322",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 735228,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "02/02/2020 18:46:06",
      "content": "<p>Have you tried focal loss? It seems like it didnt give me better results :\\ what about you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 737339,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "02/05/2020 07:52:19",
          "content": "<p>ohem loss didn't give me better results either. I'm afraid they are more useful in a object detection task than in a classification task.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 737465,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "02/05/2020 11:29:22",
      "content": "<p>Thanks for sharing <a href=\"/machinelp\">@machinelp</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 749836,
          "author_name": "machinelp",
          "author_url": "",
          "post_date": "02/18/2020 23:41:21",
          "content": "<p>thx</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 740822,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "02/09/2020 20:58:29",
      "content": "<p>Thanks for sharing! I have a few questions regarding the implementation and the paper. </p>\n\n<p>1) I'm not very familiar with pytorch Variable, seems like it can do autograd (although now seems like its deprecated, and we can pass in requires_grad=True), so based on the code, seems like you are setting <code>alpha</code> to require gradient i.e. to be backproped and updated, Is that correct? In the paper they only showed binary example and they did bunch experiments using a range of fixed <code>alpha</code> (if i understood correctly), e.g. gamma=2 and alpha=0.25 they found that works the best, they didnt mention to train <code>alpha</code> (although i imagine one could, and initialized as what u showed here to be ones)</p>\n\n<p>2) <code>class_mask = Variable(class_mask)</code> are you also setting class_mask to be backproped/updated? (sorry again if im not understanding something silly here). And to just confirm, the class_mask here is to just make the computation of ce more efficient right? as it does not need to sum up the terms that multiplied by zeros in the mask.</p>\n\n<p>EDIT: Just realized in order for alpha to be trainable, it also needs to be passed into the optimizer, and the code is not doing it, but also i wonder if simply making it trainable, the model would just learn a naive alpha e.g. +inf, since in the formula the alpha is just a constant multiplier on the loss.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 744748,
      "author_name": "virajbagal",
      "author_url": "",
      "post_date": "02/13/2020 05:17:07",
      "content": "<p>Hi <a href=\"/machinelp\">@machinelp</a> . Your Focal loss is returning either average or sum of the batch_loss, i.e, a single value for each batch. But for ohem we want loss for each image. I think focal loss here shld return only batch_loss and not batch_loss.mean() or batch_loss.sum(). Am I right ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 761205,
      "author_name": "machinelp",
      "author_url": "",
      "post_date": "03/02/2020 08:53:01",
      "content": "<p>FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133322\">https://www.kaggle.com/c/bengaliai-cv19/discussion/133322</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "734951": "```python\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.autograd import Variable\ndevice = torch.device('cuda:0' if torch.cuda.is_available() else 'cpu')\n\nclass FocalLoss(nn.Module):\n    def __init__(self, class_num, alpha=None, gamma=2, size_average=True):\n        super(FocalLoss, self).__init__()\n        if alpha is None:\n            self.alpha = Variable(torch.ones(class_num, 1))\n        else:\n            if isinstance(alpha, Variable):\n                self.alpha = alpha\n            else:\n                self.alpha = Variable(alpha)\n        self.gamma = gamma\n        self.class_num = class_num\n        self.size_average = size_average\n\n    def forward(self, inputs, targets):\n        N = inputs.size(0)\n        C = inputs.size(1)\n        P = F.softmax(inputs)\n\n        class_mask = inputs.data.new(N, C).fill_(0)\n        class_mask = Variable(class_mask)\n        ids = targets.view(-1, 1)\n        class_mask.scatter_(1, ids.data, 1.)\n        #print(class_mask)\n\n\n        if inputs.is_cuda and not self.alpha.is_cuda:\n            self.alpha = self.alpha.to(device)\n        alpha = self.alpha[ids.data.view(-1)]\n\n        probs = (P*class_mask).sum(1).view(-1,1)\n\n        log_p = probs.log()\n        #print('probs size= {}'.format(probs.size()))\n        #print(probs)\n\n        batch_loss = -alpha*(torch.pow((1-probs), self.gamma))*log_p \n        #print('-----bacth_loss------')\n        #print(batch_loss)\n\n\n        if self.size_average:\n            loss = batch_loss.mean()\n        else:\n            loss = batch_loss.sum()\n        return loss\n\nF1 = FocalLoss(168)\nF2 = FocalLoss(11)\nF3 = FocalLoss(7)\n\n```\n\n```python\ndef ohem_loss( rate, cls_pred, cls_target ):\n    batch_size = cls_pred.size(0) \n    # ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n    ohem_cls_loss = F1(cls_pred, cls_target)\n\n    sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\n    keep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\n    if keep_num &lt; sorted_ohem_loss.size()[0]:\n        keep_idx_cuda = idx[:keep_num]\n        ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\n    cls_loss = ohem_cls_loss.sum() / keep_num\n    return cls_loss\n```\n\nreference：\n（1）https://github.com/clcarwin/focal_loss_pytorch/blob/master/focalloss.py\n（2）https://www.kaggle.com/c/bengaliai-cv19/discussion/128637 \n（3）https://www.kaggle.com/c/bengaliai-cv19/discussion/128592",
    "735228": "Have you tried focal loss? It seems like it didnt give me better results :\\ what about you?",
    "737339": "ohem loss didn't give me better results either. I'm afraid they are more useful in a object detection task than in a classification task.",
    "737465": "Thanks for sharing @machinelp",
    "740822": "Thanks for sharing! I have a few questions regarding the implementation and the paper. \n\n1) I'm not very familiar with pytorch Variable, seems like it can do autograd (although now seems like its deprecated, and we can pass in requires_grad=True), so based on the code, seems like you are setting `alpha` to require gradient i.e. to be backproped and updated, Is that correct? In the paper they only showed binary example and they did bunch experiments using a range of fixed `alpha` (if i understood correctly), e.g. gamma=2 and alpha=0.25 they found that works the best, they didnt mention to train `alpha` (although i imagine one could, and initialized as what u showed here to be ones)\n\n2) `class_mask = Variable(class_mask)` are you also setting class_mask to be backproped/updated? (sorry again if im not understanding something silly here). And to just confirm, the class_mask here is to just make the computation of ce more efficient right? as it does not need to sum up the terms that multiplied by zeros in the mask.\n\nEDIT: Just realized in order for alpha to be trainable, it also needs to be passed into the optimizer, and the code is not doing it, but also i wonder if simply making it trainable, the model would just learn a naive alpha e.g. +inf, since in the formula the alpha is just a constant multiplier on the loss.",
    "744748": "Hi @machinelp . Your Focal loss is returning either average or sum of the batch_loss, i.e, a single value for each batch. But for ohem we want loss for each image. I think focal loss here shld return only batch_loss and not batch_loss.mean() or batch_loss.sum(). Am I right ?",
    "749836": "thx",
    "761205": "FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：https://www.kaggle.com/c/bengaliai-cv19/discussion/133322"
  },
  "source": "meta"
}