{
  "id": 109987,
  "title": "Sharing problems I encountered training Arcface models",
  "url": "/competitions/recursion-cellular-image-classification/discussion/109987",
  "author_name": "",
  "post_date": "2019-09-24T02:36:37.392645900Z",
  "votes": 48,
  "comment_count": 11,
  "views": 0,
  "content": "<p>From the start of this competition I have been experimenting with metric learning and Arcface in particular, because of excellent results in previous competitions reported by bestfitting, pudae and others. I learned a great deal from their works, but it still took me weeks to actually get my Arcface model going. Below are some of the problems I've encountered along the way. </p>\n\n<ol>\n<li><p>Numerical instability. I use tensorflow Keras, and when I first started to train my Arcface model, the training loss was stuck at 16.xxx and never changes. It took me a few days hopelessly debugging through various places in my code to finally realize that the problem was numerical instability. When defining the Arcface layer, there is a hyperparameter scaling factor s that rescales the features before classification, which is seldom mentioned and presumably helps with numerical stability. The s is defaulted to 30.0 in most implementations, however with 1108 classes, this s value is actually too big, and resulted in numerical instability when calculating softmax downstream. Once I lowered s to 10.0, training loss started to decrease.</p></li>\n<li><p>Local minimum. Once I got the model training, I got my Arcface model training accuracy to 100% in almost no time, even though the training loss was still quite high. However the prediction was complete garbage. After quite a bit of angry debugging, it turned out that I fell into a local minimum of the Arcface model itself. If you follow the math of the Arcface layer, it turns out that when you have an embedding model that makes all cosine logits -1.0, you get 100% accuracy after the Arcface layer because of the interaction with the true label passed into the layer, so that's actually a stable local minimum. I don't know exactly what I did that got me over this problem. Maybe some changes I made in my model definition made it less likely to predict a whole bunch of -1.0. </p></li>\n<li><p>Training from scratch. Arcface model seems to need a good starting embedding model. When I trained it from scratch, it was very difficult for the model to converge to anything useful. So for my subsequent training, I always train a softmax model first, then train Arcface model based on the embedding from the softmax model.</p></li>\n</ol>\n\n<p>My results with Arcface models are markedly better compared to softmax models, although my softmax models and Arcface models are not exactly identical so some of the difference might come from that. So I consider my efforts getting it working well worth it, although these tricky problems were borderline infuriating while I was trying to work through them. Except for 3, I haven't seen these mentioned anywhere, so I am sharing these here to hopefully be helpful for someone else trying to get Arcface to work.</p>",
  "messages": [
    {
      "id": "632748",
      "postDate": "09/24/2019 02:36:37",
      "content": "<p>From the start of this competition I have been experimenting with metric learning and Arcface in particular, because of excellent results in previous competitions reported by bestfitting, pudae and others. I learned a great deal from their works, but it still took me weeks to actually get my Arcface model going. Below are some of the problems I've encountered along the way. </p>\n\n<ol>\n<li><p>Numerical instability. I use tensorflow Keras, and when I first started to train my Arcface model, the training loss was stuck at 16.xxx and never changes. It took me a few days hopelessly debugging through various places in my code to finally realize that the problem was numerical instability. When defining the Arcface layer, there is a hyperparameter scaling factor s that rescales the features before classification, which is seldom mentioned and presumably helps with numerical stability. The s is defaulted to 30.0 in most implementations, however with 1108 classes, this s value is actually too big, and resulted in numerical instability when calculating softmax downstream. Once I lowered s to 10.0, training loss started to decrease.</p></li>\n<li><p>Local minimum. Once I got the model training, I got my Arcface model training accuracy to 100% in almost no time, even though the training loss was still quite high. However the prediction was complete garbage. After quite a bit of angry debugging, it turned out that I fell into a local minimum of the Arcface model itself. If you follow the math of the Arcface layer, it turns out that when you have an embedding model that makes all cosine logits -1.0, you get 100% accuracy after the Arcface layer because of the interaction with the true label passed into the layer, so that's actually a stable local minimum. I don't know exactly what I did that got me over this problem. Maybe some changes I made in my model definition made it less likely to predict a whole bunch of -1.0. </p></li>\n<li><p>Training from scratch. Arcface model seems to need a good starting embedding model. When I trained it from scratch, it was very difficult for the model to converge to anything useful. So for my subsequent training, I always train a softmax model first, then train Arcface model based on the embedding from the softmax model.</p></li>\n</ol>\n\n<p>My results with Arcface models are markedly better compared to softmax models, although my softmax models and Arcface models are not exactly identical so some of the difference might come from that. So I consider my efforts getting it working well worth it, although these tricky problems were borderline infuriating while I was trying to work through them. Except for 3, I haven't seen these mentioned anywhere, so I am sharing these here to hopefully be helpful for someone else trying to get Arcface to work.</p>",
      "rawMarkdown": "From the start of this competition I have been experimenting with metric learning and Arcface in particular, because of excellent results in previous competitions reported by bestfitting, pudae and others. I learned a great deal from their works, but it still took me weeks to actually get my Arcface model going. Below are some of the problems I've encountered along the way. \n\n1. Numerical instability. I use tensorflow Keras, and when I first started to train my Arcface model, the training loss was stuck at 16.xxx and never changes. It took me a few days hopelessly debugging through various places in my code to finally realize that the problem was numerical instability. When defining the Arcface layer, there is a hyperparameter scaling factor s that rescales the features before classification, which is seldom mentioned and presumably helps with numerical stability. The s is defaulted to 30.0 in most implementations, however with 1108 classes, this s value is actually too big, and resulted in numerical instability when calculating softmax downstream. Once I lowered s to 10.0, training loss started to decrease.\n\n2. Local minimum. Once I got the model training, I got my Arcface model training accuracy to 100% in almost no time, even though the training loss was still quite high. However the prediction was complete garbage. After quite a bit of angry debugging, it turned out that I fell into a local minimum of the Arcface model itself. If you follow the math of the Arcface layer, it turns out that when you have an embedding model that makes all cosine logits -1.0, you get 100% accuracy after the Arcface layer because of the interaction with the true label passed into the layer, so that's actually a stable local minimum. I don't know exactly what I did that got me over this problem. Maybe some changes I made in my model definition made it less likely to predict a whole bunch of -1.0. \n\n3. Training from scratch. Arcface model seems to need a good starting embedding model. When I trained it from scratch, it was very difficult for the model to converge to anything useful. So for my subsequent training, I always train a softmax model first, then train Arcface model based on the embedding from the softmax model.\n\nMy results with Arcface models are markedly better compared to softmax models, although my softmax models and Arcface models are not exactly identical so some of the difference might come from that. So I consider my efforts getting it working well worth it, although these tricky problems were borderline infuriating while I was trying to work through them. Except for 3, I haven't seen these mentioned anywhere, so I am sharing these here to hopefully be helpful for someone else trying to get Arcface to work.",
      "votes": null
    },
    {
      "id": "632762",
      "postDate": "09/24/2019 03:02:34",
      "content": "<p>Great insight! I'll try this out for the time remaining. Expecting improvements:)</p>",
      "rawMarkdown": "Great insight! I'll try this out for the time remaining. Expecting improvements:)",
      "votes": null
    },
    {
      "id": "632799",
      "postDate": "09/24/2019 04:27:17",
      "content": "<p>This is awesome! Thanks for posting!</p>",
      "rawMarkdown": "This is awesome! Thanks for posting!",
      "votes": null
    },
    {
      "id": "632820",
      "postDate": "09/24/2019 05:25:04",
      "content": "<p>Thank you for sharing!  Very interesting insights. \nI tried to tune hyperparameters in Arcface, including s, but none of them gave me a significant improvement.. All of them performed somehow fine, slightly better than softmax models :/</p>",
      "rawMarkdown": "Thank you for sharing!  Very interesting insights. \nI tried to tune hyperparameters in Arcface, including s, but none of them gave me a significant improvement.. All of them performed somehow fine, slightly better than softmax models :/",
      "votes": null
    },
    {
      "id": "632827",
      "postDate": "09/24/2019 05:44:30",
      "content": "<p>Many thanks for your insights. Looking forward to your code-based solutions at the end of this competition. :-)</p>",
      "rawMarkdown": "Many thanks for your insights. Looking forward to your code-based solutions at the end of this competition. :-)",
      "votes": null
    },
    {
      "id": "632952",
      "postDate": "09/24/2019 08:48:17",
      "content": "<p>Please let me add to the list, from my experience</p>\n\n<ol>\n<li>Batch normalization layer is vital just before the L2-normalized features multiplication by the weights. It is indeed mentioned in the original article, but I didn't pay attention at first, and ArcFace started working for me only after I added it.</li>\n</ol>\n\n<p>On the other hand I experienced no problems with <code>s=30</code>. In fact, I have not experienced neither of 1,2,3 from the OP. What <code>m</code> value do you use?</p>",
      "rawMarkdown": "Please let me add to the list, from my experience\n\n4. Batch normalization layer is vital just before the L2-normalized features multiplication by the weights. It is indeed mentioned in the original article, but I didn't pay attention at first, and ArcFace started working for me only after I added it.\n\nOn the other hand I experienced no problems with `s=30`. In fact, I have not experienced neither of 1,2,3 from the OP. What `m` value do you use?",
      "votes": null
    },
    {
      "id": "632988",
      "postDate": "09/24/2019 09:58:52",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> interesting finding about BatchNorm! I wonder what happened when you didn't use it before the L2-normalized features multiplication by the weights. Did network diverge? Or you just get worse results? </p>",
      "rawMarkdown": "zaharch interesting finding about BatchNorm! I wonder what happened when you didn't use it before the L2-normalized features multiplication by the weights. Did network diverge? Or you just get worse results?",
      "votes": null
    },
    {
      "id": "633001",
      "postDate": "09/24/2019 10:24:25",
      "content": "<p>It was worse results, both training and validation. Specifically, the training accuracy was not improving above some middle range. Usually, the first thing I want from a network is to be able to overfit on the training data, otherwise it is lacking descriptive power. And here it was not able to get high training scores, staying in mid range. </p>",
      "rawMarkdown": "It was worse results, both training and validation. Specifically, the training accuracy was not improving above some middle range. Usually, the first thing I want from a network is to be able to overfit on the training data, otherwise it is lacking descriptive power. And here it was not able to get high training scores, staying in mid range.",
      "votes": null
    },
    {
      "id": "633166",
      "postDate": "09/24/2019 14:01:37",
      "content": "<p>I have read someone talking about following the exact sequence of layers as the original paper was important, so possibly the position of the batchnorm layer is important. </p>\n\n<p>The s=30 problem might be specific to Tensorflow/Keras. It’s also possibly related to the random initial values of the network. Otherwise I would assume more people would have seen it. I was about to give up on Keras and switch to PyTorch before I figured out what it was.</p>\n\n<p>The local minimum issue was probably related to model structure. I might happened to have a model that was very prone to move that way.</p>\n\n<p>I found a github issue where one of the suggestions for training arcface was to try train a softmax model first, so at least I wasn’t the only one having problem getting it converge.</p>\n\n<p>I don’t  think these problems are common, otherwise I would have found something about them and didn’t have to debug through it myself. Hopefully this might help someone else having similar issues.</p>",
      "rawMarkdown": "I have read someone talking about following the exact sequence of layers as the original paper was important, so possibly the position of the batchnorm layer is important. \n\nThe s=30 problem might be specific to Tensorflow/Keras. It’s also possibly related to the random initial values of the network. Otherwise I would assume more people would have seen it. I was about to give up on Keras and switch to PyTorch before I figured out what it was.\n\nThe local minimum issue was probably related to model structure. I might happened to have a model that was very prone to move that way.\n\nI found a github issue where one of the suggestions for training arcface was to try train a softmax model first, so at least I wasn’t the only one having problem getting it converge.\n\nI don’t  think these problems are common, otherwise I would have found something about them and didn’t have to debug through it myself. Hopefully this might help someone else having similar issues.",
      "votes": null
    },
    {
      "id": "633312",
      "postDate": "09/24/2019 17:59:46",
      "content": "<blockquote>\n  <p>The s=30 problem might be specific to Tensorflow/Keras</p>\n</blockquote>\n\n<p>I agree with this statement, because using PyTorch and <code>s=30</code> successfully converged with different architectures (DenseNet, ResNet, SE ResNets) and with <code>m=0.35, m=0.5</code> at least for me:)</p>\n\n<p>And also, I didn't use any pre-training with Vanilla Softmax, so most probably it's the problem of particular Tensorflow/Keras implementation.</p>",
      "rawMarkdown": "&gt; The s=30 problem might be specific to Tensorflow/Keras\n\nI agree with this statement, because using PyTorch and `s=30` successfully converged with different architectures (DenseNet, ResNet, SE ResNets) and with `m=0.35, m=0.5` at least for me:)\n\nAnd also, I didn't use any pre-training with Vanilla Softmax, so most probably it's the problem of particular Tensorflow/Keras implementation.",
      "votes": null
    },
    {
      "id": "633349",
      "postDate": "09/24/2019 19:42:12",
      "content": "<p>Actually, it seems like \"30.0 in most implementations, however with 1108 classes, this s value is actually too big\"\nI remember training Arcface on Face Recognition datasets with large amounts of classes, and it worked quite well regardless of s being 30 or 60.\nThanks for your insights! It's been really confusing for me since everyone was talking about Arcface in this competition, but when I tried it it made no sense at all. Maybe, I should try harder!</p>\n\n<p>Btw, here is the code that I adopted: <a href=\"https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py\">https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py</a></p>\n\n<p>It's somewhat different from Arcface that Bestfitting used, but I think still ok</p>",
      "rawMarkdown": "Actually, it seems like \"30.0 in most implementations, however with 1108 classes, this s value is actually too big\"\nI remember training Arcface on Face Recognition datasets with large amounts of classes, and it worked quite well regardless of s being 30 or 60.\nThanks for your insights! It's been really confusing for me since everyone was talking about Arcface in this competition, but when I tried it it made no sense at all. Maybe, I should try harder!\n\nBtw, here is the code that I adopted: https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py\n\nIt's somewhat different from Arcface that Bestfitting used, but I think still ok",
      "votes": null
    },
    {
      "id": "1736148",
      "postDate": "03/27/2022 03:35:34",
      "content": "<p>\"Training from scratch. Arcface model seems to need a good starting embedding model. When I trained it from scratch, it was very difficult for the model to converge to anything useful\"</p>\n<p>this is because of the code:</p>\n<pre><code>if self.easy_margin:\nphi = torch.where(cosine &gt; 0, phi, cosine)\nelse:\nphi = torch.where(cosine &gt; self.th, phi, cosine - self.mm)\nin which, self.mm = math.sin(math.pi - m) *m\n</code></pre>\n<p><a href=\"https://github.com/deepinsight/insightface/issues/247\" target=\"_blank\">https://github.com/deepinsight/insightface/issues/247</a><br>\nArcFace loss code: the meaning of mm #247</p>\n<p><a href=\"https://github.com/ronghuaiyang/arcface-pytorch/issues/48\" target=\"_blank\">https://github.com/ronghuaiyang/arcface-pytorch/issues/48</a><br>\n<a href=\"https://zhuanlan.zhihu.com/p/103766001\" target=\"_blank\">https://zhuanlan.zhihu.com/p/103766001</a></p>",
      "rawMarkdown": "\"Training from scratch. Arcface model seems to need a good starting embedding model. When I trained it from scratch, it was very difficult for the model to converge to anything useful\"\n\nthis is because of the code:\n```\nif self.easy_margin:\nphi = torch.where(cosine > 0, phi, cosine)\nelse:\nphi = torch.where(cosine > self.th, phi, cosine - self.mm)\nin which, self.mm = math.sin(math.pi - m) *m\n\n```\n\nhttps://github.com/deepinsight/insightface/issues/247\nArcFace loss code: the meaning of mm #247\n\nhttps://github.com/ronghuaiyang/arcface-pytorch/issues/48\nhttps://zhuanlan.zhihu.com/p/103766001",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1736148,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/27/2022 03:35:34",
      "content": "<p>\"Training from scratch. Arcface model seems to need a good starting embedding model. When I trained it from scratch, it was very difficult for the model to converge to anything useful\"</p>\n<p>this is because of the code:</p>\n<pre><code>if self.easy_margin:\nphi = torch.where(cosine &gt; 0, phi, cosine)\nelse:\nphi = torch.where(cosine &gt; self.th, phi, cosine - self.mm)\nin which, self.mm = math.sin(math.pi - m) *m\n</code></pre>\n<p><a href=\"https://github.com/deepinsight/insightface/issues/247\" target=\"_blank\">https://github.com/deepinsight/insightface/issues/247</a><br>\nArcFace loss code: the meaning of mm #247</p>\n<p><a href=\"https://github.com/ronghuaiyang/arcface-pytorch/issues/48\" target=\"_blank\">https://github.com/ronghuaiyang/arcface-pytorch/issues/48</a><br>\n<a href=\"https://zhuanlan.zhihu.com/p/103766001\" target=\"_blank\">https://zhuanlan.zhihu.com/p/103766001</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 632762,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "09/24/2019 03:02:34",
      "content": "<p>Great insight! I'll try this out for the time remaining. Expecting improvements:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 632799,
      "author_name": "kitkatbar0429",
      "author_url": "",
      "post_date": "09/24/2019 04:27:17",
      "content": "<p>This is awesome! Thanks for posting!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 632820,
      "author_name": "analokamus",
      "author_url": "",
      "post_date": "09/24/2019 05:25:04",
      "content": "<p>Thank you for sharing!  Very interesting insights. \nI tried to tune hyperparameters in Arcface, including s, but none of them gave me a significant improvement.. All of them performed somehow fine, slightly better than softmax models :/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 632827,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "09/24/2019 05:44:30",
      "content": "<p>Many thanks for your insights. Looking forward to your code-based solutions at the end of this competition. :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 632952,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "09/24/2019 08:48:17",
      "content": "<p>Please let me add to the list, from my experience</p>\n\n<ol>\n<li>Batch normalization layer is vital just before the L2-normalized features multiplication by the weights. It is indeed mentioned in the original article, but I didn't pay attention at first, and ArcFace started working for me only after I added it.</li>\n</ol>\n\n<p>On the other hand I experienced no problems with <code>s=30</code>. In fact, I have not experienced neither of 1,2,3 from the OP. What <code>m</code> value do you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 632988,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "09/24/2019 09:58:52",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> interesting finding about BatchNorm! I wonder what happened when you didn't use it before the L2-normalized features multiplication by the weights. Did network diverge? Or you just get worse results? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 633001,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "09/24/2019 10:24:25",
          "content": "<p>It was worse results, both training and validation. Specifically, the training accuracy was not improving above some middle range. Usually, the first thing I want from a network is to be able to overfit on the training data, otherwise it is lacking descriptive power. And here it was not able to get high training scores, staying in mid range. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 633166,
          "author_name": "mingzhao03",
          "author_url": "",
          "post_date": "09/24/2019 14:01:37",
          "content": "<p>I have read someone talking about following the exact sequence of layers as the original paper was important, so possibly the position of the batchnorm layer is important. </p>\n\n<p>The s=30 problem might be specific to Tensorflow/Keras. It’s also possibly related to the random initial values of the network. Otherwise I would assume more people would have seen it. I was about to give up on Keras and switch to PyTorch before I figured out what it was.</p>\n\n<p>The local minimum issue was probably related to model structure. I might happened to have a model that was very prone to move that way.</p>\n\n<p>I found a github issue where one of the suggestions for training arcface was to try train a softmax model first, so at least I wasn’t the only one having problem getting it converge.</p>\n\n<p>I don’t  think these problems are common, otherwise I would have found something about them and didn’t have to debug through it myself. Hopefully this might help someone else having similar issues.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 633312,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "09/24/2019 17:59:46",
          "content": "<blockquote>\n  <p>The s=30 problem might be specific to Tensorflow/Keras</p>\n</blockquote>\n\n<p>I agree with this statement, because using PyTorch and <code>s=30</code> successfully converged with different architectures (DenseNet, ResNet, SE ResNets) and with <code>m=0.35, m=0.5</code> at least for me:)</p>\n\n<p>And also, I didn't use any pre-training with Vanilla Softmax, so most probably it's the problem of particular Tensorflow/Keras implementation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 633349,
      "author_name": "devvindan",
      "author_url": "",
      "post_date": "09/24/2019 19:42:12",
      "content": "<p>Actually, it seems like \"30.0 in most implementations, however with 1108 classes, this s value is actually too big\"\nI remember training Arcface on Face Recognition datasets with large amounts of classes, and it worked quite well regardless of s being 30 or 60.\nThanks for your insights! It's been really confusing for me since everyone was talking about Arcface in this competition, but when I tried it it made no sense at all. Maybe, I should try harder!</p>\n\n<p>Btw, here is the code that I adopted: <a href=\"https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py\">https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py</a></p>\n\n<p>It's somewhat different from Arcface that Bestfitting used, but I think still ok</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "632748": "From the start of this competition I have been experimenting with metric learning and Arcface in particular, because of excellent results in previous competitions reported by bestfitting, pudae and others. I learned a great deal from their works, but it still took me weeks to actually get my Arcface model going. Below are some of the problems I've encountered along the way. \n\n1. Numerical instability. I use tensorflow Keras, and when I first started to train my Arcface model, the training loss was stuck at 16.xxx and never changes. It took me a few days hopelessly debugging through various places in my code to finally realize that the problem was numerical instability. When defining the Arcface layer, there is a hyperparameter scaling factor s that rescales the features before classification, which is seldom mentioned and presumably helps with numerical stability. The s is defaulted to 30.0 in most implementations, however with 1108 classes, this s value is actually too big, and resulted in numerical instability when calculating softmax downstream. Once I lowered s to 10.0, training loss started to decrease.\n\n2. Local minimum. Once I got the model training, I got my Arcface model training accuracy to 100% in almost no time, even though the training loss was still quite high. However the prediction was complete garbage. After quite a bit of angry debugging, it turned out that I fell into a local minimum of the Arcface model itself. If you follow the math of the Arcface layer, it turns out that when you have an embedding model that makes all cosine logits -1.0, you get 100% accuracy after the Arcface layer because of the interaction with the true label passed into the layer, so that's actually a stable local minimum. I don't know exactly what I did that got me over this problem. Maybe some changes I made in my model definition made it less likely to predict a whole bunch of -1.0. \n\n3. Training from scratch. Arcface model seems to need a good starting embedding model. When I trained it from scratch, it was very difficult for the model to converge to anything useful. So for my subsequent training, I always train a softmax model first, then train Arcface model based on the embedding from the softmax model.\n\nMy results with Arcface models are markedly better compared to softmax models, although my softmax models and Arcface models are not exactly identical so some of the difference might come from that. So I consider my efforts getting it working well worth it, although these tricky problems were borderline infuriating while I was trying to work through them. Except for 3, I haven't seen these mentioned anywhere, so I am sharing these here to hopefully be helpful for someone else trying to get Arcface to work.",
    "632762": "Great insight! I'll try this out for the time remaining. Expecting improvements:)",
    "632799": "This is awesome! Thanks for posting!",
    "632820": "Thank you for sharing!  Very interesting insights. \nI tried to tune hyperparameters in Arcface, including s, but none of them gave me a significant improvement.. All of them performed somehow fine, slightly better than softmax models :/",
    "632827": "Many thanks for your insights. Looking forward to your code-based solutions at the end of this competition. :-)",
    "632952": "Please let me add to the list, from my experience\n\n4. Batch normalization layer is vital just before the L2-normalized features multiplication by the weights. It is indeed mentioned in the original article, but I didn't pay attention at first, and ArcFace started working for me only after I added it.\n\nOn the other hand I experienced no problems with `s=30`. In fact, I have not experienced neither of 1,2,3 from the OP. What `m` value do you use?",
    "632988": "zaharch interesting finding about BatchNorm! I wonder what happened when you didn't use it before the L2-normalized features multiplication by the weights. Did network diverge? Or you just get worse results?",
    "633001": "It was worse results, both training and validation. Specifically, the training accuracy was not improving above some middle range. Usually, the first thing I want from a network is to be able to overfit on the training data, otherwise it is lacking descriptive power. And here it was not able to get high training scores, staying in mid range.",
    "633166": "I have read someone talking about following the exact sequence of layers as the original paper was important, so possibly the position of the batchnorm layer is important. \n\nThe s=30 problem might be specific to Tensorflow/Keras. It’s also possibly related to the random initial values of the network. Otherwise I would assume more people would have seen it. I was about to give up on Keras and switch to PyTorch before I figured out what it was.\n\nThe local minimum issue was probably related to model structure. I might happened to have a model that was very prone to move that way.\n\nI found a github issue where one of the suggestions for training arcface was to try train a softmax model first, so at least I wasn’t the only one having problem getting it converge.\n\nI don’t  think these problems are common, otherwise I would have found something about them and didn’t have to debug through it myself. Hopefully this might help someone else having similar issues.",
    "633312": "&gt; The s=30 problem might be specific to Tensorflow/Keras\n\nI agree with this statement, because using PyTorch and `s=30` successfully converged with different architectures (DenseNet, ResNet, SE ResNets) and with `m=0.35, m=0.5` at least for me:)\n\nAnd also, I didn't use any pre-training with Vanilla Softmax, so most probably it's the problem of particular Tensorflow/Keras implementation.",
    "633349": "Actually, it seems like \"30.0 in most implementations, however with 1108 classes, this s value is actually too big\"\nI remember training Arcface on Face Recognition datasets with large amounts of classes, and it worked quite well regardless of s being 30 or 60.\nThanks for your insights! It's been really confusing for me since everyone was talking about Arcface in this competition, but when I tried it it made no sense at all. Maybe, I should try harder!\n\nBtw, here is the code that I adopted: https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py\n\nIt's somewhat different from Arcface that Bestfitting used, but I think still ok",
    "1736148": "\"Training from scratch. Arcface model seems to need a good starting embedding model. When I trained it from scratch, it was very difficult for the model to converge to anything useful\"\n\nthis is because of the code:\n```\nif self.easy_margin:\nphi = torch.where(cosine > 0, phi, cosine)\nelse:\nphi = torch.where(cosine > self.th, phi, cosine - self.mm)\nin which, self.mm = math.sin(math.pi - m) *m\n\n```\n\nhttps://github.com/deepinsight/insightface/issues/247\nArcFace loss code: the meaning of mm #247\n\nhttps://github.com/ronghuaiyang/arcface-pytorch/issues/48\nhttps://zhuanlan.zhihu.com/p/103766001"
  },
  "source": "meta"
}