{
  "id": 307486,
  "title": "Problem about ArcMarginProduct",
  "url": "/competitions/happy-whale-and-dolphin/discussion/307486",
  "author_name": "",
  "post_date": "2022-02-14T11:49:05.107855300Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>``Hi, I'm new here, and I noticed many public notebooks use ArcMarginProduct as the last layer of their model, so I also tried that, but when I want to do the inference step, I got a problem: the ArcMarginProduct layer need to use the label as one of the input, but I don't know the label at this stage. How to deal with this problem? Following is code of ArcMarginProduct layer</p>\n<p>'<br>\nclass ArcMarginProduct(nn.Module):</p>\n<pre><code>def __init__(self, in_features, out_features, s=30.0, \n             m=0.50, easy_margin=False, ls_eps=0.0):\n    super(ArcMarginProduct, self).__init__()\n    self.in_features = in_features\n    self.out_features = out_features\n    self.s = s\n    self.m = m\n    self.ls_eps = ls_eps  # label smoothing\n    self.weight = nn.Parameter(torch.FloatTensor(out_features, in_features))\n    nn.init.xavier_uniform_(self.weight)\n\n    self.easy_margin = easy_margin\n    self.cos_m = math.cos(m)\n    self.sin_m = math.sin(m)\n    self.th = math.cos(math.pi - m)\n    self.mm = math.sin(math.pi - m) * m\n\ndef forward(self, input, label):\n    # --------------------------- cos(theta) &amp; phi(theta) ---------------------\n    cosine = F.linear(F.normalize(input), F.normalize(self.weight))\n    sine = torch.sqrt(1.0 - torch.pow(cosine, 2))\n    phi = cosine * self.cos_m - sine * self.sin_m\n    if self.easy_margin:\n        phi = torch.where(cosine &gt; 0, phi, cosine)\n    else:\n        phi = torch.where(cosine &gt; self.th, phi, cosine - self.mm)\n    # --------------------------- convert label to one-hot ---------------------\n    # one_hot = torch.zeros(cosine.size(), requires_grad=True, device='cuda')\n    one_hot = torch.zeros(cosine.size(), device=CONFIG['device'])\n    one_hot.scatter_(1, label.view(-1, 1).long(), 1)\n    if self.ls_eps &gt; 0:\n        one_hot = (1 - self.ls_eps) * one_hot + self.ls_eps / self.out_features\n    # -------------torch.where(out_i = {x_i if condition_i else y_i) ------------\n    output = (one_hot * phi) + ((1.0 - one_hot) * cosine)\n    output *= self.s\n\n    return output'\n</code></pre>",
  "messages": [
    {
      "id": "1689662",
      "postDate": "02/14/2022 11:49:05",
      "content": "<p>``Hi, I'm new here, and I noticed many public notebooks use ArcMarginProduct as the last layer of their model, so I also tried that, but when I want to do the inference step, I got a problem: the ArcMarginProduct layer need to use the label as one of the input, but I don't know the label at this stage. How to deal with this problem? Following is code of ArcMarginProduct layer</p>\n<p>'<br>\nclass ArcMarginProduct(nn.Module):</p>\n<pre><code>def __init__(self, in_features, out_features, s=30.0, \n             m=0.50, easy_margin=False, ls_eps=0.0):\n    super(ArcMarginProduct, self).__init__()\n    self.in_features = in_features\n    self.out_features = out_features\n    self.s = s\n    self.m = m\n    self.ls_eps = ls_eps  # label smoothing\n    self.weight = nn.Parameter(torch.FloatTensor(out_features, in_features))\n    nn.init.xavier_uniform_(self.weight)\n\n    self.easy_margin = easy_margin\n    self.cos_m = math.cos(m)\n    self.sin_m = math.sin(m)\n    self.th = math.cos(math.pi - m)\n    self.mm = math.sin(math.pi - m) * m\n\ndef forward(self, input, label):\n    # --------------------------- cos(theta) &amp; phi(theta) ---------------------\n    cosine = F.linear(F.normalize(input), F.normalize(self.weight))\n    sine = torch.sqrt(1.0 - torch.pow(cosine, 2))\n    phi = cosine * self.cos_m - sine * self.sin_m\n    if self.easy_margin:\n        phi = torch.where(cosine &gt; 0, phi, cosine)\n    else:\n        phi = torch.where(cosine &gt; self.th, phi, cosine - self.mm)\n    # --------------------------- convert label to one-hot ---------------------\n    # one_hot = torch.zeros(cosine.size(), requires_grad=True, device='cuda')\n    one_hot = torch.zeros(cosine.size(), device=CONFIG['device'])\n    one_hot.scatter_(1, label.view(-1, 1).long(), 1)\n    if self.ls_eps &gt; 0:\n        one_hot = (1 - self.ls_eps) * one_hot + self.ls_eps / self.out_features\n    # -------------torch.where(out_i = {x_i if condition_i else y_i) ------------\n    output = (one_hot * phi) + ((1.0 - one_hot) * cosine)\n    output *= self.s\n\n    return output'\n</code></pre>",
      "rawMarkdown": "``Hi, I'm new here, and I noticed many public notebooks use ArcMarginProduct as the last layer of their model, so I also tried that, but when I want to do the inference step, I got a problem: the ArcMarginProduct layer need to use the label as one of the input, but I don't know the label at this stage. How to deal with this problem? Following is code of ArcMarginProduct layer\n\n\n'\nclass ArcMarginProduct(nn.Module):\n\n    def __init__(self, in_features, out_features, s=30.0, \n                 m=0.50, easy_margin=False, ls_eps=0.0):\n        super(ArcMarginProduct, self).__init__()\n        self.in_features = in_features\n        self.out_features = out_features\n        self.s = s\n        self.m = m\n        self.ls_eps = ls_eps  # label smoothing\n        self.weight = nn.Parameter(torch.FloatTensor(out_features, in_features))\n        nn.init.xavier_uniform_(self.weight)\n\n        self.easy_margin = easy_margin\n        self.cos_m = math.cos(m)\n        self.sin_m = math.sin(m)\n        self.th = math.cos(math.pi - m)\n        self.mm = math.sin(math.pi - m) * m\n\n    def forward(self, input, label):\n        # --------------------------- cos(theta) & phi(theta) ---------------------\n        cosine = F.linear(F.normalize(input), F.normalize(self.weight))\n        sine = torch.sqrt(1.0 - torch.pow(cosine, 2))\n        phi = cosine * self.cos_m - sine * self.sin_m\n        if self.easy_margin:\n            phi = torch.where(cosine > 0, phi, cosine)\n        else:\n            phi = torch.where(cosine > self.th, phi, cosine - self.mm)\n        # --------------------------- convert label to one-hot ---------------------\n        # one_hot = torch.zeros(cosine.size(), requires_grad=True, device='cuda')\n        one_hot = torch.zeros(cosine.size(), device=CONFIG['device'])\n        one_hot.scatter_(1, label.view(-1, 1).long(), 1)\n        if self.ls_eps > 0:\n            one_hot = (1 - self.ls_eps) * one_hot + self.ls_eps / self.out_features\n        # -------------torch.where(out_i = {x_i if condition_i else y_i) ------------\n        output = (one_hot * phi) + ((1.0 - one_hot) * cosine)\n        output *= self.s\n\n        return output'",
      "votes": null
    },
    {
      "id": "1689732",
      "postDate": "02/14/2022 13:09:58",
      "content": "<p>from various notebooks I have seen and I've also created one notebook with arcmargin product here we are providing the labels as the species column in train dataframe </p>\n<p>while test data frame doesn't have any species column there you just have to take outputs from model before taking arc margin product that will give out embeddings for you. </p>\n<p><code>here with using arcmargin product our goal is to find the vector space between multiple images after that we can use KNN for finding the nearest 5 images</code> </p>",
      "rawMarkdown": "from various notebooks I have seen and I've also created one notebook with arcmargin product here we are providing the labels as the species column in train dataframe \n\nwhile test data frame doesn't have any species column there you just have to take outputs from model before taking arc margin product that will give out embeddings for you. \n\n```  here with using arcmargin product our goal is to find the vector space between multiple images after that we can use KNN for finding the nearest 5 images ```",
      "votes": null
    },
    {
      "id": "1689739",
      "postDate": "02/14/2022 13:14:02",
      "content": "<p>I've had similar doubt before <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a>  gave this answer I think it'll be helpful to you too </p>\n<p>I will suggest you reading this paper to get better idea of ArcFace: <a href=\"https://arxiv.org/pdf/1801.07698.pdf\" target=\"_blank\">https://arxiv.org/pdf/1801.07698.pdf</a>. If you want to get in more depth, start with SphereFace: <a href=\"https://arxiv.org/abs/1704.08063\" target=\"_blank\">https://arxiv.org/abs/1704.08063</a>. I should have mentioned it at the top of notebook. The implementation here is similar for easy_margin case. In short, we pass image &amp; target both as an input to arcface model. We l2 normalize the embedding space and ArcFace weights to project it to a sphere of radius 1. The ArcFace weights [w1,w2,…,wk] is of dimension N_CLASSES*EMBED_DIM. Each of wi represent the class centre in the vector space over 512-dim sphere of radius 1. We compute cosine similarity, scale the values and then apply Softmax.<br>\nNow, for ArcFace, we add constant angular margin to the target class for computing cosine similarities. The idea to add extra penalty for target class is to simultaneously enhance the intra-class compactness and inter-class discrepancy. Those labels in ArcMarginProduct implementation are the same target class as the output.<br>\nSo, here, we are not interested in directly classifying targets but to learn distances between images. For embed_model, the output is 512-dim representation that is passed to ArcMarginProduct layer which helps it to learn cosine distances. So, \"model\" in my code is used for training. \"embed_model\" is used for inference. There is an alternative way for doing this, that is to add entire logic for ArcFace as a custom loss function. Then, you can have same model for training and inference</p>",
      "rawMarkdown": "I've had similar doubt before @ks2019  gave this answer I think it'll be helpful to you too \n\nI will suggest you reading this paper to get better idea of ArcFace: [https://arxiv.org/pdf/1801.07698.pdf](https://arxiv.org/pdf/1801.07698.pdf). If you want to get in more depth, start with SphereFace: [https://arxiv.org/abs/1704.08063](https://arxiv.org/abs/1704.08063). I should have mentioned it at the top of notebook. The implementation here is similar for easy_margin case. In short, we pass image & target both as an input to arcface model. We l2 normalize the embedding space and ArcFace weights to project it to a sphere of radius 1. The ArcFace weights [w1,w2,…,wk] is of dimension N_CLASSES*EMBED_DIM. Each of wi represent the class centre in the vector space over 512-dim sphere of radius 1. We compute cosine similarity, scale the values and then apply Softmax.\nNow, for ArcFace, we add constant angular margin to the target class for computing cosine similarities. The idea to add extra penalty for target class is to simultaneously enhance the intra-class compactness and inter-class discrepancy. Those labels in ArcMarginProduct implementation are the same target class as the output.\nSo, here, we are not interested in directly classifying targets but to learn distances between images. For embed_model, the output is 512-dim representation that is passed to ArcMarginProduct layer which helps it to learn cosine distances. So, \"model\" in my code is used for training. \"embed_model\" is used for inference. There is an alternative way for doing this, that is to add entire logic for ArcFace as a custom loss function. Then, you can have same model for training and inference",
      "votes": null
    },
    {
      "id": "1689740",
      "postDate": "02/14/2022 13:14:08",
      "content": "<p>One of the possible solutions is to perform retrieval using the nearest neighbors search.</p>",
      "rawMarkdown": "One of the possible solutions is to perform retrieval using the nearest neighbors search.",
      "votes": null
    },
    {
      "id": "1689754",
      "postDate": "02/14/2022 13:31:11",
      "content": "<p>Is this method kind of like an ensemble method? It sounds like stacking a neural network upon a kNN classifier.</p>",
      "rawMarkdown": "Is this method kind of like an ensemble method? It sounds like stacking a neural network upon a kNN classifier.",
      "votes": null
    },
    {
      "id": "1689758",
      "postDate": "02/14/2022 13:33:15",
      "content": "<p>yes it's similar to </p>",
      "rawMarkdown": "yes it's similar to",
      "votes": null
    },
    {
      "id": "1689759",
      "postDate": "02/14/2022 13:34:30",
      "content": "<p>Thank you for your patient reply. I think it takes a while for me to read these two papers. Do you mind me asking you some questions about the papers later?</p>",
      "rawMarkdown": "Thank you for your patient reply. I think it takes a while for me to read these two papers. Do you mind me asking you some questions about the papers later?",
      "votes": null
    },
    {
      "id": "1689766",
      "postDate": "02/14/2022 13:41:11",
      "content": "<p>yes sure you can mail me at <a href=\"someshfengde@gmail.com\" target=\"_blank\">someshfengde@gmail.com</a> or connect on discord if you have some doubts I am not expert but I'll try to resolve :)  </p>",
      "rawMarkdown": "yes sure you can mail me at [someshfengde@gmail.com](someshfengde@gmail.com) or connect on discord if you have some doubts I am not expert but I'll try to resolve :)",
      "votes": null
    },
    {
      "id": "1689777",
      "postDate": "02/14/2022 13:52:03",
      "content": "<p><a href=\"https://www.kaggle.com/cheinting\" target=\"_blank\">@cheinting</a> <a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> Please be careful not to discuss anything privately related to the competition, its against the rules :) </p>",
      "rawMarkdown": "cheinting @somesh88 Please be careful not to discuss anything privately related to the competition, its against the rules :)",
      "votes": null
    },
    {
      "id": "1689794",
      "postDate": "02/14/2022 13:56:26",
      "content": "<p>For sure, only discuss something related to those papers. </p>",
      "rawMarkdown": "For sure, only discuss something related to those papers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1689732,
      "author_name": "somesh88",
      "author_url": "",
      "post_date": "02/14/2022 13:09:58",
      "content": "<p>from various notebooks I have seen and I've also created one notebook with arcmargin product here we are providing the labels as the species column in train dataframe </p>\n<p>while test data frame doesn't have any species column there you just have to take outputs from model before taking arc margin product that will give out embeddings for you. </p>\n<p><code>here with using arcmargin product our goal is to find the vector space between multiple images after that we can use KNN for finding the nearest 5 images</code> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1689739,
      "author_name": "somesh88",
      "author_url": "",
      "post_date": "02/14/2022 13:14:02",
      "content": "<p>I've had similar doubt before <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a>  gave this answer I think it'll be helpful to you too </p>\n<p>I will suggest you reading this paper to get better idea of ArcFace: <a href=\"https://arxiv.org/pdf/1801.07698.pdf\" target=\"_blank\">https://arxiv.org/pdf/1801.07698.pdf</a>. If you want to get in more depth, start with SphereFace: <a href=\"https://arxiv.org/abs/1704.08063\" target=\"_blank\">https://arxiv.org/abs/1704.08063</a>. I should have mentioned it at the top of notebook. The implementation here is similar for easy_margin case. In short, we pass image &amp; target both as an input to arcface model. We l2 normalize the embedding space and ArcFace weights to project it to a sphere of radius 1. The ArcFace weights [w1,w2,…,wk] is of dimension N_CLASSES*EMBED_DIM. Each of wi represent the class centre in the vector space over 512-dim sphere of radius 1. We compute cosine similarity, scale the values and then apply Softmax.<br>\nNow, for ArcFace, we add constant angular margin to the target class for computing cosine similarities. The idea to add extra penalty for target class is to simultaneously enhance the intra-class compactness and inter-class discrepancy. Those labels in ArcMarginProduct implementation are the same target class as the output.<br>\nSo, here, we are not interested in directly classifying targets but to learn distances between images. For embed_model, the output is 512-dim representation that is passed to ArcMarginProduct layer which helps it to learn cosine distances. So, \"model\" in my code is used for training. \"embed_model\" is used for inference. There is an alternative way for doing this, that is to add entire logic for ArcFace as a custom loss function. Then, you can have same model for training and inference</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689759,
          "author_name": "cheinting",
          "author_url": "",
          "post_date": "02/14/2022 13:34:30",
          "content": "<p>Thank you for your patient reply. I think it takes a while for me to read these two papers. Do you mind me asking you some questions about the papers later?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689766,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/14/2022 13:41:11",
          "content": "<p>yes sure you can mail me at <a href=\"someshfengde@gmail.com\" target=\"_blank\">someshfengde@gmail.com</a> or connect on discord if you have some doubts I am not expert but I'll try to resolve :)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689777,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/14/2022 13:52:03",
          "content": "<p><a href=\"https://www.kaggle.com/cheinting\" target=\"_blank\">@cheinting</a> <a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> Please be careful not to discuss anything privately related to the competition, its against the rules :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689794,
          "author_name": "cheinting",
          "author_url": "",
          "post_date": "02/14/2022 13:56:26",
          "content": "<p>For sure, only discuss something related to those papers. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689740,
      "author_name": "igorkrashenyi",
      "author_url": "",
      "post_date": "02/14/2022 13:14:08",
      "content": "<p>One of the possible solutions is to perform retrieval using the nearest neighbors search.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689754,
          "author_name": "cheinting",
          "author_url": "",
          "post_date": "02/14/2022 13:31:11",
          "content": "<p>Is this method kind of like an ensemble method? It sounds like stacking a neural network upon a kNN classifier.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689758,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/14/2022 13:33:15",
          "content": "<p>yes it's similar to </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1689662": "``Hi, I'm new here, and I noticed many public notebooks use ArcMarginProduct as the last layer of their model, so I also tried that, but when I want to do the inference step, I got a problem: the ArcMarginProduct layer need to use the label as one of the input, but I don't know the label at this stage. How to deal with this problem? Following is code of ArcMarginProduct layer\n\n\n'\nclass ArcMarginProduct(nn.Module):\n\n    def __init__(self, in_features, out_features, s=30.0, \n                 m=0.50, easy_margin=False, ls_eps=0.0):\n        super(ArcMarginProduct, self).__init__()\n        self.in_features = in_features\n        self.out_features = out_features\n        self.s = s\n        self.m = m\n        self.ls_eps = ls_eps  # label smoothing\n        self.weight = nn.Parameter(torch.FloatTensor(out_features, in_features))\n        nn.init.xavier_uniform_(self.weight)\n\n        self.easy_margin = easy_margin\n        self.cos_m = math.cos(m)\n        self.sin_m = math.sin(m)\n        self.th = math.cos(math.pi - m)\n        self.mm = math.sin(math.pi - m) * m\n\n    def forward(self, input, label):\n        # --------------------------- cos(theta) & phi(theta) ---------------------\n        cosine = F.linear(F.normalize(input), F.normalize(self.weight))\n        sine = torch.sqrt(1.0 - torch.pow(cosine, 2))\n        phi = cosine * self.cos_m - sine * self.sin_m\n        if self.easy_margin:\n            phi = torch.where(cosine > 0, phi, cosine)\n        else:\n            phi = torch.where(cosine > self.th, phi, cosine - self.mm)\n        # --------------------------- convert label to one-hot ---------------------\n        # one_hot = torch.zeros(cosine.size(), requires_grad=True, device='cuda')\n        one_hot = torch.zeros(cosine.size(), device=CONFIG['device'])\n        one_hot.scatter_(1, label.view(-1, 1).long(), 1)\n        if self.ls_eps > 0:\n            one_hot = (1 - self.ls_eps) * one_hot + self.ls_eps / self.out_features\n        # -------------torch.where(out_i = {x_i if condition_i else y_i) ------------\n        output = (one_hot * phi) + ((1.0 - one_hot) * cosine)\n        output *= self.s\n\n        return output'",
    "1689732": "from various notebooks I have seen and I've also created one notebook with arcmargin product here we are providing the labels as the species column in train dataframe \n\nwhile test data frame doesn't have any species column there you just have to take outputs from model before taking arc margin product that will give out embeddings for you. \n\n```  here with using arcmargin product our goal is to find the vector space between multiple images after that we can use KNN for finding the nearest 5 images ```",
    "1689739": "I've had similar doubt before @ks2019  gave this answer I think it'll be helpful to you too \n\nI will suggest you reading this paper to get better idea of ArcFace: [https://arxiv.org/pdf/1801.07698.pdf](https://arxiv.org/pdf/1801.07698.pdf). If you want to get in more depth, start with SphereFace: [https://arxiv.org/abs/1704.08063](https://arxiv.org/abs/1704.08063). I should have mentioned it at the top of notebook. The implementation here is similar for easy_margin case. In short, we pass image & target both as an input to arcface model. We l2 normalize the embedding space and ArcFace weights to project it to a sphere of radius 1. The ArcFace weights [w1,w2,…,wk] is of dimension N_CLASSES*EMBED_DIM. Each of wi represent the class centre in the vector space over 512-dim sphere of radius 1. We compute cosine similarity, scale the values and then apply Softmax.\nNow, for ArcFace, we add constant angular margin to the target class for computing cosine similarities. The idea to add extra penalty for target class is to simultaneously enhance the intra-class compactness and inter-class discrepancy. Those labels in ArcMarginProduct implementation are the same target class as the output.\nSo, here, we are not interested in directly classifying targets but to learn distances between images. For embed_model, the output is 512-dim representation that is passed to ArcMarginProduct layer which helps it to learn cosine distances. So, \"model\" in my code is used for training. \"embed_model\" is used for inference. There is an alternative way for doing this, that is to add entire logic for ArcFace as a custom loss function. Then, you can have same model for training and inference",
    "1689740": "One of the possible solutions is to perform retrieval using the nearest neighbors search.",
    "1689754": "Is this method kind of like an ensemble method? It sounds like stacking a neural network upon a kNN classifier.",
    "1689758": "yes it's similar to",
    "1689759": "Thank you for your patient reply. I think it takes a while for me to read these two papers. Do you mind me asking you some questions about the papers later?",
    "1689766": "yes sure you can mail me at [someshfengde@gmail.com](someshfengde@gmail.com) or connect on discord if you have some doubts I am not expert but I'll try to resolve :)",
    "1689777": "cheinting @somesh88 Please be careful not to discuss anything privately related to the competition, its against the rules :)",
    "1689794": "For sure, only discuss something related to those papers."
  },
  "source": "meta"
}