{
  "id": 88370,
  "title": "A Way to Solve the Unbalanced Data Problem -Focal Loss",
  "url": "/competitions/imet-2019-fgvc6/discussion/88370",
  "author_name": "",
  "post_date": "2019-04-08T08:31:33.876101200Z",
  "votes": 9,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The data provided  by FGVC6 is highly unbalanced data, some labels have many image to train but some labels only have few image to train. The traditional loss such as \"binary-crossentropy\" (\"BCEWithLogitsLoss\" in pytorch), \"categorical_crossentropy\"(\"CrossEntropyLoss\" in pytorch), can't solve this problem well. They will significantly cause the overfit problem, and the score is very low(about 0.17). This kernel show the problem: \n<a href=\"https://www.kaggle.com/xiuchengwang/cnn-keras-starter-senet-50\">https://www.kaggle.com/xiuchengwang/cnn-keras-starter-senet-50</a>\nSo how to solve the problem caused by  unbalanced data?\nI find a simple way : using Focal Loss. By using thin loss function my score improved form 0.17 to 0.575.</p>\n\n<p>And this is the way to use Focal Loss in <strong>keras</strong></p>\n\n<p>`\n    def focal_loss(y_true, y_pred)</p>\n\n<pre><code>pt = y_pred * y_true + (1-y_pred) * (1-y_true)\n\npt = K.clip(pt, epsilon, 1-epsilon)\n\nCE = -K.log(pt)\n\nFL = K.pow(1-pt, gamma) * CE\n\nloss = K.sum(FL, axis=1)\n\nreturn loss\n</code></pre>\n\n<p>`</p>\n\n<p>In <strong>fastsi</strong></p>\n\n<p>`\n        <strong>def forward(self, logit, target)</strong>：</p>\n\n<pre><code>    target = target.float()\n\n    max_val = (-logit).clamp(min=0)\n\n    loss = logit - logit * target + max_val + \\\n\n           ((-max_val).exp() + (-logit - max_val).exp()).log()\n\n    invprobs = F.logsigmoid(-logit * (target * 2.0 - 1.0))\n\n    loss = (invprobs * self.gamma).exp() * loss\n\n    if len(loss.size())==2:\n\n        loss = loss.sum(dim=1)\n\n    return loss.mean()`\n</code></pre>\n\n<p>There  are some bugs when I edit the code in Discussion, you can see  the detail in the following kernel. </p>\n\n<p><a href=\"https://www.kaggle.com/xiuchengwang/keras-xception-fine-turning-facol-loss\">https://www.kaggle.com/xiuchengwang/keras-xception-fine-turning-facol-loss</a></p>",
  "messages": [
    {
      "id": "509751",
      "postDate": "04/08/2019 08:31:33",
      "content": "<p>The data provided  by FGVC6 is highly unbalanced data, some labels have many image to train but some labels only have few image to train. The traditional loss such as \"binary-crossentropy\" (\"BCEWithLogitsLoss\" in pytorch), \"categorical_crossentropy\"(\"CrossEntropyLoss\" in pytorch), can't solve this problem well. They will significantly cause the overfit problem, and the score is very low(about 0.17). This kernel show the problem: \n<a href=\"https://www.kaggle.com/xiuchengwang/cnn-keras-starter-senet-50\">https://www.kaggle.com/xiuchengwang/cnn-keras-starter-senet-50</a>\nSo how to solve the problem caused by  unbalanced data?\nI find a simple way : using Focal Loss. By using thin loss function my score improved form 0.17 to 0.575.</p>\n\n<p>And this is the way to use Focal Loss in <strong>keras</strong></p>\n\n<p>`\n    def focal_loss(y_true, y_pred)</p>\n\n<pre><code>pt = y_pred * y_true + (1-y_pred) * (1-y_true)\n\npt = K.clip(pt, epsilon, 1-epsilon)\n\nCE = -K.log(pt)\n\nFL = K.pow(1-pt, gamma) * CE\n\nloss = K.sum(FL, axis=1)\n\nreturn loss\n</code></pre>\n\n<p>`</p>\n\n<p>In <strong>fastsi</strong></p>\n\n<p>`\n        <strong>def forward(self, logit, target)</strong>：</p>\n\n<pre><code>    target = target.float()\n\n    max_val = (-logit).clamp(min=0)\n\n    loss = logit - logit * target + max_val + \\\n\n           ((-max_val).exp() + (-logit - max_val).exp()).log()\n\n    invprobs = F.logsigmoid(-logit * (target * 2.0 - 1.0))\n\n    loss = (invprobs * self.gamma).exp() * loss\n\n    if len(loss.size())==2:\n\n        loss = loss.sum(dim=1)\n\n    return loss.mean()`\n</code></pre>\n\n<p>There  are some bugs when I edit the code in Discussion, you can see  the detail in the following kernel. </p>\n\n<p><a href=\"https://www.kaggle.com/xiuchengwang/keras-xception-fine-turning-facol-loss\">https://www.kaggle.com/xiuchengwang/keras-xception-fine-turning-facol-loss</a></p>",
      "rawMarkdown": "The data provided  by FGVC6 is highly unbalanced data, some labels have many image to train but some labels only have few image to train. The traditional loss such as \"binary-crossentropy\" (\"BCEWithLogitsLoss\" in pytorch), \"categorical_crossentropy\"(\"CrossEntropyLoss\" in pytorch), can't solve this problem well. They will significantly cause the overfit problem, and the score is very low(about 0.17). This kernel show the problem: \n[https://www.kaggle.com/xiuchengwang/cnn-keras-starter-senet-50](https://www.kaggle.com/xiuchengwang/cnn-keras-starter-senet-50)\nSo how to solve the problem caused by  unbalanced data?\nI find a simple way : using Focal Loss. By using thin loss function my score improved form 0.17 to 0.575.\n\nAnd this is the way to use Focal Loss in **keras**\n\n`\n    def focal_loss(y_true, y_pred)\n\n    pt = y_pred * y_true + (1-y_pred) * (1-y_true)\n\n    pt = K.clip(pt, epsilon, 1-epsilon)\n\n    CE = -K.log(pt)\n\n    FL = K.pow(1-pt, gamma) * CE\n\n    loss = K.sum(FL, axis=1)\n\n    return loss\n`\n\n\n\nIn **fastsi**\n\n`\n        **def forward(self, logit, target)**：\n\n        target = target.float()\n\n        max_val = (-logit).clamp(min=0)\n\n        loss = logit - logit * target + max_val + \\\n\n               ((-max_val).exp() + (-logit - max_val).exp()).log()\n\n        invprobs = F.logsigmoid(-logit * (target * 2.0 - 1.0))\n\n        loss = (invprobs * self.gamma).exp() * loss\n\n        if len(loss.size())==2:\n\n            loss = loss.sum(dim=1)\n\n        return loss.mean()`\n\n\nThere  are some bugs when I edit the code in Discussion, you can see  the detail in the following kernel. \n\n\n[https://www.kaggle.com/xiuchengwang/keras-xception-fine-turning-facol-loss](https://www.kaggle.com/xiuchengwang/keras-xception-fine-turning-facol-loss)",
      "votes": null
    },
    {
      "id": "510620",
      "postDate": "04/09/2019 09:47:56",
      "content": "<p>I tried to use binary-crossentropy and got 0.602 on LB</p>",
      "rawMarkdown": "I tried to use binary-crossentropy and got 0.602 on LB",
      "votes": null
    },
    {
      "id": "510640",
      "postDate": "04/09/2019 10:16:07",
      "content": "<p>I tried binary-crossentropy in Xception too， but the score is very low. What kind of base-model do you use?</p>",
      "rawMarkdown": "I tried binary-crossentropy in Xception too， but the score is very low. What kind of base-model do you use?",
      "votes": null
    },
    {
      "id": "510924",
      "postDate": "04/09/2019 15:52:20",
      "content": "<p>In my experiments, I found that BCE and focal loss give similar performance. Also, it seems that the top public kernel (Konstantin’s 0.597LB) still use BCE, not focal. (Please correct me if I am wrong)</p>",
      "rawMarkdown": "In my experiments, I found that BCE and focal loss give similar performance. Also, it seems that the top public kernel (Konstantin’s 0.597LB) still use BCE, not focal. (Please correct me if I am wrong)",
      "votes": null
    },
    {
      "id": "511159",
      "postDate": "04/09/2019 20:04:36",
      "content": "<p>My observation is that any pure-classification approach can rather easily reach the current top-3 score (~0.63), but further improvements require metric learning/meta learning or other tricks....</p>",
      "rawMarkdown": "My observation is that any pure-classification approach can rather easily reach the current top-3 score (~0.63), but further improvements require metric learning/meta learning or other tricks....",
      "votes": null
    },
    {
      "id": "511479",
      "postDate": "04/10/2019 04:17:34",
      "content": "<p>ResNet</p>",
      "rawMarkdown": "ResNet",
      "votes": null
    },
    {
      "id": "516134",
      "postDate": "04/13/2019 16:59:14",
      "content": "<p>torch frame?</p>",
      "rawMarkdown": "torch frame?",
      "votes": null
    },
    {
      "id": "516139",
      "postDate": "04/13/2019 17:01:38",
      "content": "<p>If you only run once, you can’t prove good or bad.</p>",
      "rawMarkdown": "If you only run once, you can’t prove good or bad.",
      "votes": null
    },
    {
      "id": "519759",
      "postDate": "04/19/2019 15:45:49",
      "content": "<p>do you use keras or pytorch framework?</p>",
      "rawMarkdown": "do you use keras or pytorch framework?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 510620,
      "author_name": "",
      "author_url": "",
      "post_date": "04/09/2019 09:47:56",
      "content": "<p>I tried to use binary-crossentropy and got 0.602 on LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 510640,
          "author_name": "xiuchengwang",
          "author_url": "",
          "post_date": "04/09/2019 10:16:07",
          "content": "<p>I tried binary-crossentropy in Xception too， but the score is very low. What kind of base-model do you use?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 511479,
          "author_name": "",
          "author_url": "",
          "post_date": "04/10/2019 04:17:34",
          "content": "<p>ResNet</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 516134,
          "author_name": "",
          "author_url": "",
          "post_date": "04/13/2019 16:59:14",
          "content": "<p>torch frame?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 516139,
          "author_name": "",
          "author_url": "",
          "post_date": "04/13/2019 17:01:38",
          "content": "<p>If you only run once, you can’t prove good or bad.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 510924,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "04/09/2019 15:52:20",
      "content": "<p>In my experiments, I found that BCE and focal loss give similar performance. Also, it seems that the top public kernel (Konstantin’s 0.597LB) still use BCE, not focal. (Please correct me if I am wrong)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 511159,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "04/09/2019 20:04:36",
      "content": "<p>My observation is that any pure-classification approach can rather easily reach the current top-3 score (~0.63), but further improvements require metric learning/meta learning or other tricks....</p>",
      "votes": null,
      "replies": [
        {
          "id": 519759,
          "author_name": "harleys",
          "author_url": "",
          "post_date": "04/19/2019 15:45:49",
          "content": "<p>do you use keras or pytorch framework?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "509751": "The data provided  by FGVC6 is highly unbalanced data, some labels have many image to train but some labels only have few image to train. The traditional loss such as \"binary-crossentropy\" (\"BCEWithLogitsLoss\" in pytorch), \"categorical_crossentropy\"(\"CrossEntropyLoss\" in pytorch), can't solve this problem well. They will significantly cause the overfit problem, and the score is very low(about 0.17). This kernel show the problem: \n[https://www.kaggle.com/xiuchengwang/cnn-keras-starter-senet-50](https://www.kaggle.com/xiuchengwang/cnn-keras-starter-senet-50)\nSo how to solve the problem caused by  unbalanced data?\nI find a simple way : using Focal Loss. By using thin loss function my score improved form 0.17 to 0.575.\n\nAnd this is the way to use Focal Loss in **keras**\n\n`\n    def focal_loss(y_true, y_pred)\n\n    pt = y_pred * y_true + (1-y_pred) * (1-y_true)\n\n    pt = K.clip(pt, epsilon, 1-epsilon)\n\n    CE = -K.log(pt)\n\n    FL = K.pow(1-pt, gamma) * CE\n\n    loss = K.sum(FL, axis=1)\n\n    return loss\n`\n\n\n\nIn **fastsi**\n\n`\n        **def forward(self, logit, target)**：\n\n        target = target.float()\n\n        max_val = (-logit).clamp(min=0)\n\n        loss = logit - logit * target + max_val + \\\n\n               ((-max_val).exp() + (-logit - max_val).exp()).log()\n\n        invprobs = F.logsigmoid(-logit * (target * 2.0 - 1.0))\n\n        loss = (invprobs * self.gamma).exp() * loss\n\n        if len(loss.size())==2:\n\n            loss = loss.sum(dim=1)\n\n        return loss.mean()`\n\n\nThere  are some bugs when I edit the code in Discussion, you can see  the detail in the following kernel. \n\n\n[https://www.kaggle.com/xiuchengwang/keras-xception-fine-turning-facol-loss](https://www.kaggle.com/xiuchengwang/keras-xception-fine-turning-facol-loss)",
    "510620": "I tried to use binary-crossentropy and got 0.602 on LB",
    "510640": "I tried binary-crossentropy in Xception too， but the score is very low. What kind of base-model do you use?",
    "510924": "In my experiments, I found that BCE and focal loss give similar performance. Also, it seems that the top public kernel (Konstantin’s 0.597LB) still use BCE, not focal. (Please correct me if I am wrong)",
    "511159": "My observation is that any pure-classification approach can rather easily reach the current top-3 score (~0.63), but further improvements require metric learning/meta learning or other tricks....",
    "511479": "ResNet",
    "516134": "torch frame?",
    "516139": "If you only run once, you can’t prove good or bad.",
    "519759": "do you use keras or pytorch framework?"
  },
  "source": "meta"
}