{
  "id": 77289,
  "title": "11th place solution(custom f1 loss)",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/77289",
  "author_name": "",
  "post_date": "2019-01-11T06:21:41.322456100Z",
  "votes": 32,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Dream comes true! Here I briefly describe the approach I used.  <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77282\">Here is the details of our team's solution.</a></p>\n\n<p>I use custom f1 loss to train model, and ensemble with my teammates' models which use bce loss.</p>\n\n<p>Single se-resnext50(5 folds ensemble) with f1 loss gets 0.562 on private LB(which is higher than our final ensemble score, but strangely, it's public LB score is low, only 0.606, so we didn't use it for final submission). Resnet18(5 folds ensemble) gets 0.612 on public LB, and 0.536 on private LB. All these scores come from 512*512 RGBY input, incluing HPA v18 external data.</p>\n\n<p>When using f1 loss, I oversample the rare classes. I apply data augmentation based on class frequency(images that have rare class get more chances to be augmented).</p>\n\n<p>I noticed that when a batch of images lacks some classes(which is common for rare classes), there is no gradient backpropagates to those classes. So I add bce to those classes in that case. In addition, I clamp the network output at 0.01 to force network focus on the hard samples.</p>\n\n<p>I use thresholds 0.205 for all classes.</p>\n\n<p>Inspired by <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74068\">Tilli's idea</a>, we select about 300 pairs of similar images in test set, and replace poor quality images with good quality ones. This approach yields about 0.002 score improvement.</p>\n\n<pre><code>def f1_loss(predict, target):\n    loss = 0\n    lack_cls = target.sum(dim=0) == 0\n    if lack_cls.any():\n        loss += F.binary_cross_entropy_with_logits(\n            predict[:, lack_cls], target[:, lack_cls])\n    predict = torch.sigmoid(predict)\n    predict = torch.clamp(predict * (1-target), min=0.01) + predict * target\n    tp = predict * target\n    tp = tp.sum(dim=0)\n    precision = tp / (predict.sum(dim=0) + 1e-8)\n    recall = tp / (target.sum(dim=0) + 1e-8)\n    f1 = 2 * (precision * recall / (precision + recall + 1e-8))\n    return 1 - f1.mean() + loss\n</code></pre>",
  "messages": [
    {
      "id": "454106",
      "postDate": "01/11/2019 06:21:41",
      "content": "<p>Dream comes true! Here I briefly describe the approach I used.  <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77282\">Here is the details of our team's solution.</a></p>\n\n<p>I use custom f1 loss to train model, and ensemble with my teammates' models which use bce loss.</p>\n\n<p>Single se-resnext50(5 folds ensemble) with f1 loss gets 0.562 on private LB(which is higher than our final ensemble score, but strangely, it's public LB score is low, only 0.606, so we didn't use it for final submission). Resnet18(5 folds ensemble) gets 0.612 on public LB, and 0.536 on private LB. All these scores come from 512*512 RGBY input, incluing HPA v18 external data.</p>\n\n<p>When using f1 loss, I oversample the rare classes. I apply data augmentation based on class frequency(images that have rare class get more chances to be augmented).</p>\n\n<p>I noticed that when a batch of images lacks some classes(which is common for rare classes), there is no gradient backpropagates to those classes. So I add bce to those classes in that case. In addition, I clamp the network output at 0.01 to force network focus on the hard samples.</p>\n\n<p>I use thresholds 0.205 for all classes.</p>\n\n<p>Inspired by <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74068\">Tilli's idea</a>, we select about 300 pairs of similar images in test set, and replace poor quality images with good quality ones. This approach yields about 0.002 score improvement.</p>\n\n<pre><code>def f1_loss(predict, target):\n    loss = 0\n    lack_cls = target.sum(dim=0) == 0\n    if lack_cls.any():\n        loss += F.binary_cross_entropy_with_logits(\n            predict[:, lack_cls], target[:, lack_cls])\n    predict = torch.sigmoid(predict)\n    predict = torch.clamp(predict * (1-target), min=0.01) + predict * target\n    tp = predict * target\n    tp = tp.sum(dim=0)\n    precision = tp / (predict.sum(dim=0) + 1e-8)\n    recall = tp / (target.sum(dim=0) + 1e-8)\n    f1 = 2 * (precision * recall / (precision + recall + 1e-8))\n    return 1 - f1.mean() + loss\n</code></pre>",
      "rawMarkdown": "Dream comes true! Here I briefly describe the approach I used.  [Here is the details of our team's solution.][1]\n\nI use custom f1 loss to train model, and ensemble with my teammates' models which use bce loss.\n\nSingle se-resnext50(5 folds ensemble) with f1 loss gets 0.562 on private LB(which is higher than our final ensemble score, but strangely, it's public LB score is low, only 0.606, so we didn't use it for final submission). Resnet18(5 folds ensemble) gets 0.612 on public LB, and 0.536 on private LB. All these scores come from 512*512 RGBY input, incluing HPA v18 external data.\n\nWhen using f1 loss, I oversample the rare classes. I apply data augmentation based on class frequency(images that have rare class get more chances to be augmented).\n\nI noticed that when a batch of images lacks some classes(which is common for rare classes), there is no gradient backpropagates to those classes. So I add bce to those classes in that case. In addition, I clamp the network output at 0.01 to force network focus on the hard samples.\n\nI use thresholds 0.205 for all classes.\n\nInspired by [Tilli's idea][2], we select about 300 pairs of similar images in test set, and replace poor quality images with good quality ones. This approach yields about 0.002 score improvement.\n\n    def f1_loss(predict, target):\n        loss = 0\n        lack_cls = target.sum(dim=0) == 0\n        if lack_cls.any():\n            loss += F.binary_cross_entropy_with_logits(\n                predict[:, lack_cls], target[:, lack_cls])\n        predict = torch.sigmoid(predict)\n        predict = torch.clamp(predict * (1-target), min=0.01) + predict * target\n        tp = predict * target\n        tp = tp.sum(dim=0)\n        precision = tp / (predict.sum(dim=0) + 1e-8)\n        recall = tp / (target.sum(dim=0) + 1e-8)\n        f1 = 2 * (precision * recall / (precision + recall + 1e-8))\n        return 1 - f1.mean() + loss\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77282\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74068",
      "votes": null
    },
    {
      "id": "454116",
      "postDate": "01/11/2019 06:28:48",
      "content": "<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77282\">The details of our team's solution.</a></p>",
      "rawMarkdown": "[The details of our team's solution.][1]\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77282",
      "votes": null
    },
    {
      "id": "454146",
      "postDate": "01/11/2019 07:08:22",
      "content": "<p>Congratulations and thanks!</p>",
      "rawMarkdown": "Congratulations and thanks!",
      "votes": null
    },
    {
      "id": "454196",
      "postDate": "01/11/2019 08:23:02",
      "content": "<p>I used a very similar loss function, but only started testing it this week! A mix of focal loss/CE and a continuos L1_loss function. Good work!</p>",
      "rawMarkdown": "I used a very similar loss function, but only started testing it this week! A mix of focal loss/CE and a continuos L1_loss function. Good work!",
      "votes": null
    },
    {
      "id": "454348",
      "postDate": "01/11/2019 13:34:38",
      "content": "<p>Good work and your custom Loss is very new for me and I have learn from it.. Thanks once again for sharing..</p>",
      "rawMarkdown": "Good work and your custom Loss is very new for me and I have learn from it.. Thanks once again for sharing..",
      "votes": null
    },
    {
      "id": "465501",
      "postDate": "02/03/2019 09:32:19",
      "content": "<p>你好，我想问一下在反向传播的时候，f1.mean()是不是也会回传梯度呢？还是只是当作一个常量？谢谢</p>",
      "rawMarkdown": "你好，我想问一下在反向传播的时候，f1.mean()是不是也会回传梯度呢？还是只是当作一个常量？谢谢",
      "votes": null
    },
    {
      "id": "488775",
      "postDate": "03/13/2019 01:41:12",
      "content": "<p>会回传梯度的</p>",
      "rawMarkdown": "会回传梯度的",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454116,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "01/11/2019 06:28:48",
      "content": "<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77282\">The details of our team's solution.</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454146,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "01/11/2019 07:08:22",
      "content": "<p>Congratulations and thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454196,
      "author_name": "tcapelle",
      "author_url": "",
      "post_date": "01/11/2019 08:23:02",
      "content": "<p>I used a very similar loss function, but only started testing it this week! A mix of focal loss/CE and a continuos L1_loss function. Good work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454348,
      "author_name": "viswanathravindran",
      "author_url": "",
      "post_date": "01/11/2019 13:34:38",
      "content": "<p>Good work and your custom Loss is very new for me and I have learn from it.. Thanks once again for sharing..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 465501,
      "author_name": "beitadoge",
      "author_url": "",
      "post_date": "02/03/2019 09:32:19",
      "content": "<p>你好，我想问一下在反向传播的时候，f1.mean()是不是也会回传梯度呢？还是只是当作一个常量？谢谢</p>",
      "votes": null,
      "replies": [
        {
          "id": 488775,
          "author_name": "shisususu",
          "author_url": "",
          "post_date": "03/13/2019 01:41:12",
          "content": "<p>会回传梯度的</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454106": "Dream comes true! Here I briefly describe the approach I used.  [Here is the details of our team's solution.][1]\n\nI use custom f1 loss to train model, and ensemble with my teammates' models which use bce loss.\n\nSingle se-resnext50(5 folds ensemble) with f1 loss gets 0.562 on private LB(which is higher than our final ensemble score, but strangely, it's public LB score is low, only 0.606, so we didn't use it for final submission). Resnet18(5 folds ensemble) gets 0.612 on public LB, and 0.536 on private LB. All these scores come from 512*512 RGBY input, incluing HPA v18 external data.\n\nWhen using f1 loss, I oversample the rare classes. I apply data augmentation based on class frequency(images that have rare class get more chances to be augmented).\n\nI noticed that when a batch of images lacks some classes(which is common for rare classes), there is no gradient backpropagates to those classes. So I add bce to those classes in that case. In addition, I clamp the network output at 0.01 to force network focus on the hard samples.\n\nI use thresholds 0.205 for all classes.\n\nInspired by [Tilli's idea][2], we select about 300 pairs of similar images in test set, and replace poor quality images with good quality ones. This approach yields about 0.002 score improvement.\n\n    def f1_loss(predict, target):\n        loss = 0\n        lack_cls = target.sum(dim=0) == 0\n        if lack_cls.any():\n            loss += F.binary_cross_entropy_with_logits(\n                predict[:, lack_cls], target[:, lack_cls])\n        predict = torch.sigmoid(predict)\n        predict = torch.clamp(predict * (1-target), min=0.01) + predict * target\n        tp = predict * target\n        tp = tp.sum(dim=0)\n        precision = tp / (predict.sum(dim=0) + 1e-8)\n        recall = tp / (target.sum(dim=0) + 1e-8)\n        f1 = 2 * (precision * recall / (precision + recall + 1e-8))\n        return 1 - f1.mean() + loss\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77282\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74068",
    "454116": "[The details of our team's solution.][1]\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/77282",
    "454146": "Congratulations and thanks!",
    "454196": "I used a very similar loss function, but only started testing it this week! A mix of focal loss/CE and a continuos L1_loss function. Good work!",
    "454348": "Good work and your custom Loss is very new for me and I have learn from it.. Thanks once again for sharing..",
    "465501": "你好，我想问一下在反向传播的时候，f1.mean()是不是也会回传梯度呢？还是只是当作一个常量？谢谢",
    "488775": "会回传梯度的"
  },
  "source": "meta"
}