{
  "id": 81821,
  "title": "SphereFace, CosFace, ArcFace or other AngularSoftmax loss",
  "url": "/competitions/humpback-whale-identification/discussion/81821",
  "author_name": "Eduardo Rocha de Andrade",
  "post_date": "2019-02-25T10:12:30.280000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Is any of you guys having success in using one of these Angular/Cosine Softmax loss functions?\nThey seem pretty nice in the papers and although I haven't explored them intensively, for me they yielded better results than triplet/contrastive loss and in less time. However, they were overperformed by the ProtoNets. Anyone getting above 0.9LB with them?</p>\n\n<p><a href=\"https://arxiv.org/abs/1704.08063\">SphereFace</a></p>\n\n<p><a href=\"https://arxiv.org/abs/1801.09414\">CosFace</a></p>\n\n<p><a href=\"https://arxiv.org/abs/1801.07698\">ArcFace</a></p>",
  "messages": [
    {
      "id": 480287,
      "postDate": "2019-02-28T02:30:40.660Z",
      "content": "<p>Arcface works in spite of its simplicity.\nPyTorch, arcface + over-sampling, se_resnext101_32x4d, input size 768x256 -&gt; LB: 0.91</p>",
      "rawMarkdown": "Arcface works in spite of its simplicity.\nPyTorch, arcface + over-sampling, se_resnext101_32x4d, input size 768x256 -&gt; LB: 0.91",
      "votes": 4
    },
    {
      "id": 478337,
      "postDate": "2019-02-26T03:14:35.443Z",
      "content": "<p>in face recognition, intra-class variation is small if the face are of same pose.</p>\n\n<p>if you have good feature extraction,  such metric-learning based softmax like cosface, sphereface, etc should work too in our case.</p>\n\n<p>if e.g. sphereface+softmax can achieve &gt;lb.0.90, then softmax alone must achieve &gt;lb0.88 first i think.</p>\n\n<p>i can get large margin center loss to work with lb&gt;0.90\n<a href=\"https://arxiv.org/abs/1803.02988\">https://arxiv.org/abs/1803.02988</a></p>",
      "rawMarkdown": "in face recognition, intra-class variation is small if the face are of same pose.\n\nif you have good feature extraction,  such metric-learning based softmax like cosface, sphereface, etc should work too in our case.\n\nif e.g. sphereface+softmax can achieve &gt;lb.0.90, then softmax alone must achieve &gt;lb0.88 first i think.\n\ni can get large margin center loss to work with lb&gt;0.90\nhttps://arxiv.org/abs/1803.02988\n\n",
      "votes": 4
    },
    {
      "id": 477822,
      "postDate": "2019-02-25T10:12:30.280Z",
      "content": "<p>Is any of you guys having success in using one of these Angular/Cosine Softmax loss functions?\nThey seem pretty nice in the papers and although I haven't explored them intensively, for me they yielded better results than triplet/contrastive loss and in less time. However, they were overperformed by the ProtoNets. Anyone getting above 0.9LB with them?</p>\n\n<p><a href=\"https://arxiv.org/abs/1704.08063\">SphereFace</a></p>\n\n<p><a href=\"https://arxiv.org/abs/1801.09414\">CosFace</a></p>\n\n<p><a href=\"https://arxiv.org/abs/1801.07698\">ArcFace</a></p>",
      "rawMarkdown": "Is any of you guys having success in using one of these Angular/Cosine Softmax loss functions?\nThey seem pretty nice in the papers and although I haven't explored them intensively, for me they yielded better results than triplet/contrastive loss and in less time. However, they were overperformed by the ProtoNets. Anyone getting above 0.9LB with them?\n\n[SphereFace][1]\n\n[CosFace][2]\n\n[ArcFace][3]\n\n  [1]: https://arxiv.org/abs/1704.08063\n  [2]: https://arxiv.org/abs/1801.09414\n  [3]: https://arxiv.org/abs/1801.07698",
      "votes": 4
    },
    {
      "id": 478452,
      "postDate": "2019-02-26T07:03:25.953Z",
      "content": "<p>You got me:) So I can confirm that you can get model &gt; 0.9 with this technique, even ~ 0.93. More after competition.</p>",
      "rawMarkdown": "You got me:) So I can confirm that you can get model &gt; 0.9 with this technique, even ~ 0.93. More after competition.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 478609,
          "postDate": "2019-02-26T11:48:04.800Z",
          "content": "<p>Haha that's funny. \nFor me the ProtoNet is better in public LB and equal at local validation.\nI must be doing something wrong because I can't go over 0.9 with these methods. It could be the way I'm doing the query or the threshold I'm using to input new_whales, idk..\nI'd appreciate if you could tell me about your method when the competition ends! Good luck :)</p>",
          "rawMarkdown": "Haha that's funny. \nFor me the ProtoNet is better in public LB and equal at local validation.\nI must be doing something wrong because I can't go over 0.9 with these methods. It could be the way I'm doing the query or the threshold I'm using to input new_whales, idk..\nI'd appreciate if you could tell me about your method when the competition ends! Good luck :)"
        },
        {
          "id": 478633,
          "postDate": "2019-02-26T12:23:15.047Z",
          "content": "<p>To tell you more now, the single model without TTA and without newwhale (so pure classification on 5004  whales) can get 0.682 LB. So you can check how good your classification is (without checking if you are predicting new_whale correctly)</p>",
          "rawMarkdown": "To tell you more now, the single model without TTA and without newwhale (so pure classification on 5004  whales) can get 0.682 LB. So you can check how good your classification is (without checking if you are predicting new_whale correctly)",
          "votes": 3
        }
      ]
    },
    {
      "id": 478125,
      "postDate": "2019-02-25T19:11:35.703Z",
      "content": "<p>ArcFace for us didn't work.</p>",
      "rawMarkdown": "ArcFace for us didn't work."
    },
    {
      "id": 480331,
      "postDate": "2019-02-28T03:57:24.187Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 480287,
      "author_name": "yu4u",
      "author_url": "",
      "post_date": "2019-02-28T02:30:40.660000",
      "content": "<p>Arcface works in spite of its simplicity.\nPyTorch, arcface + over-sampling, se_resnext101_32x4d, input size 768x256 -&gt; LB: 0.91</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 478337,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-02-26T03:14:35.443000",
      "content": "<p>in face recognition, intra-class variation is small if the face are of same pose.</p>\n\n<p>if you have good feature extraction,  such metric-learning based softmax like cosface, sphereface, etc should work too in our case.</p>\n\n<p>if e.g. sphereface+softmax can achieve &gt;lb.0.90, then softmax alone must achieve &gt;lb0.88 first i think.</p>\n\n<p>i can get large margin center loss to work with lb&gt;0.90\n<a href=\"https://arxiv.org/abs/1803.02988\">https://arxiv.org/abs/1803.02988</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 478452,
      "author_name": "Bartek",
      "author_url": "",
      "post_date": "2019-02-26T07:03:25.953000",
      "content": "<p>You got me:) So I can confirm that you can get model &gt; 0.9 with this technique, even ~ 0.93. More after competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 478609,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-02-26T11:48:04.800000",
          "content": "<p>Haha that's funny. \nFor me the ProtoNet is better in public LB and equal at local validation.\nI must be doing something wrong because I can't go over 0.9 with these methods. It could be the way I'm doing the query or the threshold I'm using to input new_whales, idk..\nI'd appreciate if you could tell me about your method when the competition ends! Good luck :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478633,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-02-26T12:23:15.047000",
          "content": "<p>To tell you more now, the single model without TTA and without newwhale (so pure classification on 5004  whales) can get 0.682 LB. So you can check how good your classification is (without checking if you are predicting new_whale correctly)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 478125,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2019-02-25T19:11:35.703000",
      "content": "<p>ArcFace for us didn't work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 480331,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-28T03:57:24.187000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "480287": "Arcface works in spite of its simplicity.\nPyTorch, arcface + over-sampling, se_resnext101_32x4d, input size 768x256 -&gt; LB: 0.91",
    "478337": "in face recognition, intra-class variation is small if the face are of same pose.\n\nif you have good feature extraction,  such metric-learning based softmax like cosface, sphereface, etc should work too in our case.\n\nif e.g. sphereface+softmax can achieve &gt;lb.0.90, then softmax alone must achieve &gt;lb0.88 first i think.\n\ni can get large margin center loss to work with lb&gt;0.90\nhttps://arxiv.org/abs/1803.02988\n\n",
    "477822": "Is any of you guys having success in using one of these Angular/Cosine Softmax loss functions?\nThey seem pretty nice in the papers and although I haven't explored them intensively, for me they yielded better results than triplet/contrastive loss and in less time. However, they were overperformed by the ProtoNets. Anyone getting above 0.9LB with them?\n\n[SphereFace][1]\n\n[CosFace][2]\n\n[ArcFace][3]\n\n  [1]: https://arxiv.org/abs/1704.08063\n  [2]: https://arxiv.org/abs/1801.09414\n  [3]: https://arxiv.org/abs/1801.07698",
    "478452": "You got me:) So I can confirm that you can get model &gt; 0.9 with this technique, even ~ 0.93. More after competition.\n\n",
    "478125": "ArcFace for us didn't work.",
    "480331": ""
  }
}