{
  "id": 165685,
  "title": "Learning objective",
  "url": "/competitions/landmark-retrieval-2020/discussion/165685",
  "author_name": "",
  "post_date": "2020-07-10T15:24:00.410058100Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>First of all, I'm new to this type of machine learning problem. So what I'm doing right now is digging through some literature (+ GitHub) to find good algorithms to train a Landmark retrieval model. What I've tried so far is the <a href=\"https://arxiv.org/pdf/1503.03832.pdf\">Triplet loss</a> and <a href=\"https://arxiv.org/pdf/1801.07698.pdf\">ArcFace</a>.</p>\n\n<p>With <code>Triplet loss</code> it seems to always converge at the margin --- so if margin is set to 1.0, my loss gets stuck at 1.0. With <code>ArcFace</code> the loss gets kind of stuck at ~10.0 (btw, what is the expected loss when using ArcFace?). Right now I'm trying hard to tune the hyperparameters of <code>Arcface</code>, and I'm also trying various implementations of this algorithm (without any success).</p>\n\n<p>ArcFace hyperparameters:\n* Scale (default seems to be ~30.0). \n* Margin (default seems to be ~0.5)</p>\n\n<p>Triplet loss hyperparameters:\n* Margin (default seems to be ~0.5)\n* K (either 2 or 4; actually a hyperparameter of my <code>DataGenerator</code>), which is the number examples of each ID in the minibatch. So if K=4, a minibatch of 16 would have labels like this <code>[1,1,1,1,7,7,7,7,15,15,15,15,28,28,28,28]</code>. Then the loss function would figure out the best triplets (Anchor (e.g. 1), Negative (e.g. 7), Positive (e.g. 1)) based on the embeddings.</p>\n\n<p><strong>So I'm looking for you professionals</strong> for some advice on how to make my model learn! I'm sure someone has struggled through this like I'm doing now :-) </p>\n\n<p>Additional info:\nMy <code>backbone/ConvNet</code> is just a regular <code>ResNet50</code> or <code>EfficientNetB0</code> followed by a global average pooling and a dense layer. The output of this <code>ConvNet+avg+dense</code> network is then fed to <code>ArcFace -&amp;gt; CategoricalCrossentopy</code>, or, in the case of the triplet loss, the output of the <code>ConvNet+avg+dense</code> is passed to the <code>triplet_hard loss function</code>. I've also tried <code>triplet_semihard</code>.</p>",
  "messages": [
    {
      "id": "923166",
      "postDate": "07/10/2020 15:24:00",
      "content": "<p>Hi,</p>\n\n<p>First of all, I'm new to this type of machine learning problem. So what I'm doing right now is digging through some literature (+ GitHub) to find good algorithms to train a Landmark retrieval model. What I've tried so far is the <a href=\"https://arxiv.org/pdf/1503.03832.pdf\">Triplet loss</a> and <a href=\"https://arxiv.org/pdf/1801.07698.pdf\">ArcFace</a>.</p>\n\n<p>With <code>Triplet loss</code> it seems to always converge at the margin --- so if margin is set to 1.0, my loss gets stuck at 1.0. With <code>ArcFace</code> the loss gets kind of stuck at ~10.0 (btw, what is the expected loss when using ArcFace?). Right now I'm trying hard to tune the hyperparameters of <code>Arcface</code>, and I'm also trying various implementations of this algorithm (without any success).</p>\n\n<p>ArcFace hyperparameters:\n* Scale (default seems to be ~30.0). \n* Margin (default seems to be ~0.5)</p>\n\n<p>Triplet loss hyperparameters:\n* Margin (default seems to be ~0.5)\n* K (either 2 or 4; actually a hyperparameter of my <code>DataGenerator</code>), which is the number examples of each ID in the minibatch. So if K=4, a minibatch of 16 would have labels like this <code>[1,1,1,1,7,7,7,7,15,15,15,15,28,28,28,28]</code>. Then the loss function would figure out the best triplets (Anchor (e.g. 1), Negative (e.g. 7), Positive (e.g. 1)) based on the embeddings.</p>\n\n<p><strong>So I'm looking for you professionals</strong> for some advice on how to make my model learn! I'm sure someone has struggled through this like I'm doing now :-) </p>\n\n<p>Additional info:\nMy <code>backbone/ConvNet</code> is just a regular <code>ResNet50</code> or <code>EfficientNetB0</code> followed by a global average pooling and a dense layer. The output of this <code>ConvNet+avg+dense</code> network is then fed to <code>ArcFace -&amp;gt; CategoricalCrossentopy</code>, or, in the case of the triplet loss, the output of the <code>ConvNet+avg+dense</code> is passed to the <code>triplet_hard loss function</code>. I've also tried <code>triplet_semihard</code>.</p>",
      "rawMarkdown": "Hi,\n\nFirst of all, I'm new to this type of machine learning problem. So what I'm doing right now is digging through some literature (+ GitHub) to find good algorithms to train a Landmark retrieval model. What I've tried so far is the [Triplet loss](https://arxiv.org/pdf/1503.03832.pdf) and [ArcFace](https://arxiv.org/pdf/1801.07698.pdf).\n\nWith `Triplet loss` it seems to always converge at the margin --- so if margin is set to 1.0, my loss gets stuck at 1.0. With `ArcFace` the loss gets kind of stuck at ~10.0 (btw, what is the expected loss when using ArcFace?). Right now I'm trying hard to tune the hyperparameters of `Arcface`, and I'm also trying various implementations of this algorithm (without any success).\n\nArcFace hyperparameters:\n* Scale (default seems to be ~30.0). \n* Margin (default seems to be ~0.5)\n\nTriplet loss hyperparameters:\n* Margin (default seems to be ~0.5)\n* K (either 2 or 4; actually a hyperparameter of my `DataGenerator`), which is the number examples of each ID in the minibatch. So if K=4, a minibatch of 16 would have labels like this `[1,1,1,1,7,7,7,7,15,15,15,15,28,28,28,28]`. Then the loss function would figure out the best triplets (Anchor (e.g. 1), Negative (e.g. 7), Positive (e.g. 1)) based on the embeddings.\n\n**So I'm looking for you professionals** for some advice on how to make my model learn! I'm sure someone has struggled through this like I'm doing now :-) \n\nAdditional info:\nMy `backbone/ConvNet` is just a regular `ResNet50` or `EfficientNetB0` followed by a global average pooling and a dense layer. The output of this `ConvNet+avg+dense` network is then fed to `ArcFace -&gt; CategoricalCrossentopy`, or, in the case of the triplet loss, the output of the `ConvNet+avg+dense` is passed to the `triplet_hard loss function`. I've also tried `triplet_semihard`.",
      "votes": null
    },
    {
      "id": "923261",
      "postDate": "07/10/2020 16:50:18",
      "content": "<p>[EDITED/UPDATED] It seems that the loss is now decreasing steadily (going below 10) with ArcFace. The loss starts really high (~21) and goes down to ~5 after 50k iterations (batch size of 32). For you interested, here's the implementation I use (feel free to give feedback!):</p>\n\n<p>```\nclass ArcMarginProduct(tf.keras.layers.Layer):\n    '''\n    References:\n        <a href=\"https://arxiv.org/pdf/1801.07698.pdf\">https://arxiv.org/pdf/1801.07698.pdf</a>\n        <a href=\"https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution/\">https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution/</a>\n            blob/master/src/modeling/metric_learning.py\n    '''\n    def <strong>init</strong>(self, n_classes, s=30, m=0.30, easy_margin=False,\n                 ls_eps=0.0, **kwargs):</p>\n\n<pre><code>    super(ArcMarginProduct, self).__init__(**kwargs)\n\n    self.n_classes = n_classes\n    self.s = s\n    self.m = m\n    self.ls_eps = ls_eps\n    self.easy_margin = easy_margin\n    self.cos_m = tf.math.cos(m)\n    self.sin_m = tf.math.sin(m)\n    self.th = tf.math.cos(math.pi - m)\n    self.mm = tf.math.sin(math.pi - m) * m\n\ndef build(self, input_shape):\n    super(ArcMarginProduct, self).build(input_shape[0])\n    self.W = self.add_weight(\n        name='W',\n        shape=(int(input_shape[0][-1]), self.n_classes),\n        initializer='glorot_uniform',\n        dtype='float32',\n        trainable=True)\n\ndef call(self, inputs):\n    X, y = inputs\n    y = tf.cast(y, dtype=tf.int32)\n    cosine = tf.matmul(\n        tf.math.l2_normalize(X, axis=1),\n        tf.math.l2_normalize(self.W, axis=0)\n    )\n    sine = tf.math.sqrt(1.0 - tf.math.pow(cosine, 2))\n    phi = cosine * self.cos_m - sine * self.sin_m\n    if self.easy_margin:\n        phi = tf.where(cosine &amp;gt; 0, phi, cosine)\n    else:\n        phi = tf.where(cosine &amp;gt; self.th, phi, cosine - self.mm)\n    one_hot = tf.cast(\n        tf.one_hot(y, depth=self.n_classes),\n        dtype=cosine.dtype\n    )\n    if self.ls_eps &amp;gt; 0:\n        one_hot = (1 - self.ls_eps) * one_hot + self.ls_eps / self.n_classes\n\n    output = (one_hot * phi) + ((1.0 - one_hot) * cosine)\n    output *= self.s\n    return output\n</code></pre>\n\n<p>class NeuralNet(tf.keras.Model):</p>\n\n<pre><code>def __init__(self, input_shape, num_classes):\n\n    super(NeuralNet, self).__init__()\n\n    self.engine = ResNet50(\n        include_top=False,\n        input_shape=input_shape,\n        weights=\"imagenet\")\n\n    self.pool = tf.keras.layers.GlobalAveragePooling2D()\n    self.batch_norm = tf.keras.layers.BatchNormalization()\n    self.dropout = tf.keras.layers.Dropout(0.0)\n    self.dense = tf.keras.layers.Dense(512)\n    self.arcface = ArcMarginProduct(scale=30, margin=0.3,\n          n_classes=n_classes, dtype='float32', name='arcface')\n\ndef call(self, inputs, **kwargs):\n    x = self.engine(inputs[0])\n    x = self.pool(x)\n    x = self.dropout(x)\n    x = self.dense(x)\n    x = self.batch_norm(x)\n    x = self.arcface([x, inputs[1]])\n    return x\n</code></pre>\n\n<h1>loss_function = SparseCategoricalCrossentropy(..., from_logits=True)</h1>\n\n<h1>optimizer = tf.keras.optimizers.SGD(1e-3, momentum=0.9) # with decay schedule</h1>\n\n<p>```</p>",
      "rawMarkdown": "[EDITED/UPDATED] It seems that the loss is now decreasing steadily (going below 10) with ArcFace. The loss starts really high (~21) and goes down to ~5 after 50k iterations (batch size of 32). For you interested, here's the implementation I use (feel free to give feedback!):\n\n```\nclass ArcMarginProduct(tf.keras.layers.Layer):\n    '''\n    References:\n        https://arxiv.org/pdf/1801.07698.pdf\n        https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution/\n            blob/master/src/modeling/metric_learning.py\n    '''\n    def __init__(self, n_classes, s=30, m=0.30, easy_margin=False,\n                 ls_eps=0.0, **kwargs):\n\n        super(ArcMarginProduct, self).__init__(**kwargs)\n\n        self.n_classes = n_classes\n        self.s = s\n        self.m = m\n        self.ls_eps = ls_eps\n        self.easy_margin = easy_margin\n        self.cos_m = tf.math.cos(m)\n        self.sin_m = tf.math.sin(m)\n        self.th = tf.math.cos(math.pi - m)\n        self.mm = tf.math.sin(math.pi - m) * m\n\n    def build(self, input_shape):\n        super(ArcMarginProduct, self).build(input_shape[0])\n        self.W = self.add_weight(\n            name='W',\n            shape=(int(input_shape[0][-1]), self.n_classes),\n            initializer='glorot_uniform',\n            dtype='float32',\n            trainable=True)\n\n    def call(self, inputs):\n        X, y = inputs\n        y = tf.cast(y, dtype=tf.int32)\n        cosine = tf.matmul(\n            tf.math.l2_normalize(X, axis=1),\n            tf.math.l2_normalize(self.W, axis=0)\n        )\n        sine = tf.math.sqrt(1.0 - tf.math.pow(cosine, 2))\n        phi = cosine * self.cos_m - sine * self.sin_m\n        if self.easy_margin:\n            phi = tf.where(cosine &gt; 0, phi, cosine)\n        else:\n            phi = tf.where(cosine &gt; self.th, phi, cosine - self.mm)\n        one_hot = tf.cast(\n            tf.one_hot(y, depth=self.n_classes),\n            dtype=cosine.dtype\n        )\n        if self.ls_eps &gt; 0:\n            one_hot = (1 - self.ls_eps) * one_hot + self.ls_eps / self.n_classes\n\n        output = (one_hot * phi) + ((1.0 - one_hot) * cosine)\n        output *= self.s\n        return output\n\n\nclass NeuralNet(tf.keras.Model):\n\n    def __init__(self, input_shape, num_classes):\n\n        super(NeuralNet, self).__init__()\n\n        self.engine = ResNet50(\n            include_top=False,\n            input_shape=input_shape,\n            weights=\"imagenet\")\n\n        self.pool = tf.keras.layers.GlobalAveragePooling2D()\n        self.batch_norm = tf.keras.layers.BatchNormalization()\n        self.dropout = tf.keras.layers.Dropout(0.0)\n        self.dense = tf.keras.layers.Dense(512)\n        self.arcface = ArcMarginProduct(scale=30, margin=0.3,\n              n_classes=n_classes, dtype='float32', name='arcface')\n\n    def call(self, inputs, **kwargs):\n        x = self.engine(inputs[0])\n        x = self.pool(x)\n        x = self.dropout(x)\n        x = self.dense(x)\n        x = self.batch_norm(x)\n        x = self.arcface([x, inputs[1]])\n        return x\n\n# loss_function = SparseCategoricalCrossentropy(..., from_logits=True)\n# optimizer = tf.keras.optimizers.SGD(1e-3, momentum=0.9) # with decay schedule\n```",
      "votes": null
    },
    {
      "id": "923358",
      "postDate": "07/10/2020 18:47:48",
      "content": "<p>Hi <a href=\"/akensert\">@akensert</a> ,\nHere we met again. xD\nI wanted to know about dataset. \nLike right now, I am not blindly training. Just wanted to know landmark_ids of training set and their corresponding indexes? Like labels for index files?</p>",
      "rawMarkdown": "Hi @akensert ,\nHere we met again. xD\nI wanted to know about dataset. \nLike right now, I am not blindly training. Just wanted to know landmark_ids of training set and their corresponding indexes? Like labels for index files?",
      "votes": null
    },
    {
      "id": "932148",
      "postDate": "07/16/2020 18:38:16",
      "content": "<p>what's your accuracy when the loss is below 10, and what epoch is it on?</p>\n\n<p>btw DELG paper used margin of 0.1. I think scale is 1, but I may be wrong. </p>\n\n<p>For me, when the margin is 0.1, the accuracy keeps going down in the first epoch (from 0.2% to 0.02%). </p>",
      "rawMarkdown": "what's your accuracy when the loss is below 10, and what epoch is it on?\n\nbtw DELG paper used margin of 0.1. I think scale is 1, but I may be wrong. \n\nFor me, when the margin is 0.1, the accuracy keeps going down in the first epoch (from 0.2% to 0.02%).",
      "votes": null
    },
    {
      "id": "932717",
      "postDate": "07/17/2020 08:32:14",
      "content": "<p>Thanks <a href=\"/suruili\">@suruili</a>, I'm trying these hyperparameters now (according to the DELG paper) and the loss is pretty much stuck already from the beginning. Right now I only look at the loss, perhaps I should implement some kind of accuracy function so that I can see the accuracy throughout training. </p>",
      "rawMarkdown": "Thanks @suruili, I'm trying these hyperparameters now (according to the DELG paper) and the loss is pretty much stuck already from the beginning. Right now I only look at the loss, perhaps I should implement some kind of accuracy function so that I can see the accuracy throughout training.",
      "votes": null
    },
    {
      "id": "933203",
      "postDate": "07/17/2020 15:02:28",
      "content": "<p>Yeah, I find that setting the scale larger help. I am trying with 30 and 60</p>",
      "rawMarkdown": "Yeah, I find that setting the scale larger help. I am trying with 30 and 60",
      "votes": null
    },
    {
      "id": "933433",
      "postDate": "07/17/2020 17:42:01",
      "content": "<p>Indeed, the DELG paper did not specify the initial value for the scale factor: we initialize it to sqrt(embedding_dimension), which in our case would lead to sqrt(2048) ~= 45</p>\n<p>Will include this in the next updated version of the paper, thanks for bringing it up.</p>",
      "rawMarkdown": "Indeed, the DELG paper did not specify the initial value for the scale factor: we initialize it to sqrt(embedding_dimension), which in our case would lead to sqrt(2048) ~= 45\n\nWill include this in the next updated version of the paper, thanks for bringing it up.",
      "votes": null
    },
    {
      "id": "933551",
      "postDate": "07/17/2020 19:13:51",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> Thanks for the info! I have some questions about the scale factor. Why do you think it affect training? Is it for numerical stability, or is the relevant scale of the scale factor vs margin matter? And why would you set it to sqrt of embedding dimension?</p>",
      "rawMarkdown": "andrefaraujo Thanks for the info! I have some questions about the scale factor. Why do you think it affect training? Is it for numerical stability, or is the relevant scale of the scale factor vs margin matter? And why would you set it to sqrt of embedding dimension?",
      "votes": null
    },
    {
      "id": "937255",
      "postDate": "07/20/2020 21:27:27",
      "content": "<p>The scale factor is necessary due to the cross-entropy loss becoming bounded for normalized class weights / features. See proposition 2 in the <a href=\"https://arxiv.org/pdf/1704.06369.pdf\" target=\"_blank\">NormFace paper</a>.</p>",
      "rawMarkdown": "The scale factor is necessary due to the cross-entropy loss becoming bounded for normalized class weights / features. See proposition 2 in the [NormFace paper](https://arxiv.org/pdf/1704.06369.pdf).",
      "votes": null
    },
    {
      "id": "937490",
      "postDate": "07/21/2020 03:50:33",
      "content": "<p>Thanks for sharing ArcMarginProduct code.</p>\n\n<p>How diverse are your predictions?\nIn my case, predictions are too concentrated in one class (which has more than 6000 images) and output embeddings \n are similar after 2 epoch\nmargin penalty(0.5) might be too big</p>\n\n<p>(2020 7 23 add)\nLast years chmpion might set margin=0.3\n<a href=\"https://arxiv.org/pdf/1906.04087.pdf\">https://arxiv.org/pdf/1906.04087.pdf</a></p>",
      "rawMarkdown": "Thanks for sharing ArcMarginProduct code.\n\n\nHow diverse are your predictions?\nIn my case, predictions are too concentrated in one class (which has more than 6000 images) and output embeddings \n are similar after 2 epoch\nmargin penalty(0.5) might be too big\n\n(2020 7 23 add)\nLast years chmpion might set margin=0.3\nhttps://arxiv.org/pdf/1906.04087.pdf",
      "votes": null
    },
    {
      "id": "981532",
      "postDate": "08/22/2020 14:16:49",
      "content": "<p><a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> can you please share how you used this model on prediction?<br>\nLike how you removed last layer from forward function. <br>\nI am pytorch user new to tensorflow. Kindly help. xD</p>",
      "rawMarkdown": "akensert can you please share how you used this model on prediction?\nLike how you removed last layer from forward function. \nI am pytorch user new to tensorflow. Kindly help. xD",
      "votes": null
    },
    {
      "id": "981830",
      "postDate": "08/22/2020 18:18:00",
      "content": "<p><a href=\"https://www.kaggle.com/micheomaano\" target=\"_blank\">@micheomaano</a> check out the public kernel. Or somethnig like this:</p>\n<pre><code>...\nmodel.load_weights('your_finetuned_weights.h5')\n\nnewmodel = tf.keras.Model(\n    inputs=model.model.get_layer('input/image').input,\n    outputs=model.model.get_layer('head/dense').output)\n\noutput = newmodel(image[tf.newaxis])[0]\nprint(output.shape)\n\n# &gt;&gt; out:\n# (512,)\n</code></pre>\n<p>Assuming that you've named the layers 'input/image' and 'head/dense': <code>tf.keras.layers.Input(..., name='input/image')</code> and <code>tf.keras.layers.Dense(..., name='head/dense')</code>  :)</p>\n<p>Hope it helps!</p>",
      "rawMarkdown": "micheomaano check out the public kernel. Or somethnig like this:\n```\n\n...\nmodel.load_weights('your_finetuned_weights.h5')\n\nnewmodel = tf.keras.Model(\n    inputs=model.model.get_layer('input/image').input,\n    outputs=model.model.get_layer('head/dense').output)\n\noutput = newmodel(image[tf.newaxis])[0]\nprint(output.shape)\n\n# >> out:\n# (512,)\n\n```\nAssuming that you've named the layers 'input/image' and 'head/dense': `tf.keras.layers.Input(..., name='input/image')` and `tf.keras.layers.Dense(..., name='head/dense')`  :)\n\n\nHope it helps!",
      "votes": null
    },
    {
      "id": "981888",
      "postDate": "08/22/2020 19:47:42",
      "content": "<p>Got it. Thanks a lot.</p>",
      "rawMarkdown": "Got it. Thanks a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937490,
      "author_name": "yasagure",
      "author_url": "",
      "post_date": "07/21/2020 03:50:33",
      "content": "<p>Thanks for sharing ArcMarginProduct code.</p>\n\n<p>How diverse are your predictions?\nIn my case, predictions are too concentrated in one class (which has more than 6000 images) and output embeddings \n are similar after 2 epoch\nmargin penalty(0.5) might be too big</p>\n\n<p>(2020 7 23 add)\nLast years chmpion might set margin=0.3\n<a href=\"https://arxiv.org/pdf/1906.04087.pdf\">https://arxiv.org/pdf/1906.04087.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 923261,
      "author_name": "akensert",
      "author_url": "",
      "post_date": "07/10/2020 16:50:18",
      "content": "<p>[EDITED/UPDATED] It seems that the loss is now decreasing steadily (going below 10) with ArcFace. The loss starts really high (~21) and goes down to ~5 after 50k iterations (batch size of 32). For you interested, here's the implementation I use (feel free to give feedback!):</p>\n\n<p>```\nclass ArcMarginProduct(tf.keras.layers.Layer):\n    '''\n    References:\n        <a href=\"https://arxiv.org/pdf/1801.07698.pdf\">https://arxiv.org/pdf/1801.07698.pdf</a>\n        <a href=\"https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution/\">https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution/</a>\n            blob/master/src/modeling/metric_learning.py\n    '''\n    def <strong>init</strong>(self, n_classes, s=30, m=0.30, easy_margin=False,\n                 ls_eps=0.0, **kwargs):</p>\n\n<pre><code>    super(ArcMarginProduct, self).__init__(**kwargs)\n\n    self.n_classes = n_classes\n    self.s = s\n    self.m = m\n    self.ls_eps = ls_eps\n    self.easy_margin = easy_margin\n    self.cos_m = tf.math.cos(m)\n    self.sin_m = tf.math.sin(m)\n    self.th = tf.math.cos(math.pi - m)\n    self.mm = tf.math.sin(math.pi - m) * m\n\ndef build(self, input_shape):\n    super(ArcMarginProduct, self).build(input_shape[0])\n    self.W = self.add_weight(\n        name='W',\n        shape=(int(input_shape[0][-1]), self.n_classes),\n        initializer='glorot_uniform',\n        dtype='float32',\n        trainable=True)\n\ndef call(self, inputs):\n    X, y = inputs\n    y = tf.cast(y, dtype=tf.int32)\n    cosine = tf.matmul(\n        tf.math.l2_normalize(X, axis=1),\n        tf.math.l2_normalize(self.W, axis=0)\n    )\n    sine = tf.math.sqrt(1.0 - tf.math.pow(cosine, 2))\n    phi = cosine * self.cos_m - sine * self.sin_m\n    if self.easy_margin:\n        phi = tf.where(cosine &amp;gt; 0, phi, cosine)\n    else:\n        phi = tf.where(cosine &amp;gt; self.th, phi, cosine - self.mm)\n    one_hot = tf.cast(\n        tf.one_hot(y, depth=self.n_classes),\n        dtype=cosine.dtype\n    )\n    if self.ls_eps &amp;gt; 0:\n        one_hot = (1 - self.ls_eps) * one_hot + self.ls_eps / self.n_classes\n\n    output = (one_hot * phi) + ((1.0 - one_hot) * cosine)\n    output *= self.s\n    return output\n</code></pre>\n\n<p>class NeuralNet(tf.keras.Model):</p>\n\n<pre><code>def __init__(self, input_shape, num_classes):\n\n    super(NeuralNet, self).__init__()\n\n    self.engine = ResNet50(\n        include_top=False,\n        input_shape=input_shape,\n        weights=\"imagenet\")\n\n    self.pool = tf.keras.layers.GlobalAveragePooling2D()\n    self.batch_norm = tf.keras.layers.BatchNormalization()\n    self.dropout = tf.keras.layers.Dropout(0.0)\n    self.dense = tf.keras.layers.Dense(512)\n    self.arcface = ArcMarginProduct(scale=30, margin=0.3,\n          n_classes=n_classes, dtype='float32', name='arcface')\n\ndef call(self, inputs, **kwargs):\n    x = self.engine(inputs[0])\n    x = self.pool(x)\n    x = self.dropout(x)\n    x = self.dense(x)\n    x = self.batch_norm(x)\n    x = self.arcface([x, inputs[1]])\n    return x\n</code></pre>\n\n<h1>loss_function = SparseCategoricalCrossentropy(..., from_logits=True)</h1>\n\n<h1>optimizer = tf.keras.optimizers.SGD(1e-3, momentum=0.9) # with decay schedule</h1>\n\n<p>```</p>",
      "votes": null,
      "replies": [
        {
          "id": 932148,
          "author_name": "suruili",
          "author_url": "",
          "post_date": "07/16/2020 18:38:16",
          "content": "<p>what's your accuracy when the loss is below 10, and what epoch is it on?</p>\n\n<p>btw DELG paper used margin of 0.1. I think scale is 1, but I may be wrong. </p>\n\n<p>For me, when the margin is 0.1, the accuracy keeps going down in the first epoch (from 0.2% to 0.02%). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932717,
          "author_name": "akensert",
          "author_url": "",
          "post_date": "07/17/2020 08:32:14",
          "content": "<p>Thanks <a href=\"/suruili\">@suruili</a>, I'm trying these hyperparameters now (according to the DELG paper) and the loss is pretty much stuck already from the beginning. Right now I only look at the loss, perhaps I should implement some kind of accuracy function so that I can see the accuracy throughout training. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933203,
          "author_name": "suruili",
          "author_url": "",
          "post_date": "07/17/2020 15:02:28",
          "content": "<p>Yeah, I find that setting the scale larger help. I am trying with 30 and 60</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933433,
          "author_name": "andrefaraujo",
          "author_url": "",
          "post_date": "07/17/2020 17:42:01",
          "content": "<p>Indeed, the DELG paper did not specify the initial value for the scale factor: we initialize it to sqrt(embedding_dimension), which in our case would lead to sqrt(2048) ~= 45</p>\n<p>Will include this in the next updated version of the paper, thanks for bringing it up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933551,
          "author_name": "suruili",
          "author_url": "",
          "post_date": "07/17/2020 19:13:51",
          "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> Thanks for the info! I have some questions about the scale factor. Why do you think it affect training? Is it for numerical stability, or is the relevant scale of the scale factor vs margin matter? And why would you set it to sqrt of embedding dimension?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937255,
          "author_name": "andrefaraujo",
          "author_url": "",
          "post_date": "07/20/2020 21:27:27",
          "content": "<p>The scale factor is necessary due to the cross-entropy loss becoming bounded for normalized class weights / features. See proposition 2 in the <a href=\"https://arxiv.org/pdf/1704.06369.pdf\" target=\"_blank\">NormFace paper</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 981532,
          "author_name": "micheomaano",
          "author_url": "",
          "post_date": "08/22/2020 14:16:49",
          "content": "<p><a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> can you please share how you used this model on prediction?<br>\nLike how you removed last layer from forward function. <br>\nI am pytorch user new to tensorflow. Kindly help. xD</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 981830,
          "author_name": "akensert",
          "author_url": "",
          "post_date": "08/22/2020 18:18:00",
          "content": "<p><a href=\"https://www.kaggle.com/micheomaano\" target=\"_blank\">@micheomaano</a> check out the public kernel. Or somethnig like this:</p>\n<pre><code>...\nmodel.load_weights('your_finetuned_weights.h5')\n\nnewmodel = tf.keras.Model(\n    inputs=model.model.get_layer('input/image').input,\n    outputs=model.model.get_layer('head/dense').output)\n\noutput = newmodel(image[tf.newaxis])[0]\nprint(output.shape)\n\n# &gt;&gt; out:\n# (512,)\n</code></pre>\n<p>Assuming that you've named the layers 'input/image' and 'head/dense': <code>tf.keras.layers.Input(..., name='input/image')</code> and <code>tf.keras.layers.Dense(..., name='head/dense')</code>  :)</p>\n<p>Hope it helps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 981888,
          "author_name": "micheomaano",
          "author_url": "",
          "post_date": "08/22/2020 19:47:42",
          "content": "<p>Got it. Thanks a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 923358,
      "author_name": "micheomaano",
      "author_url": "",
      "post_date": "07/10/2020 18:47:48",
      "content": "<p>Hi <a href=\"/akensert\">@akensert</a> ,\nHere we met again. xD\nI wanted to know about dataset. \nLike right now, I am not blindly training. Just wanted to know landmark_ids of training set and their corresponding indexes? Like labels for index files?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "923166": "Hi,\n\nFirst of all, I'm new to this type of machine learning problem. So what I'm doing right now is digging through some literature (+ GitHub) to find good algorithms to train a Landmark retrieval model. What I've tried so far is the [Triplet loss](https://arxiv.org/pdf/1503.03832.pdf) and [ArcFace](https://arxiv.org/pdf/1801.07698.pdf).\n\nWith `Triplet loss` it seems to always converge at the margin --- so if margin is set to 1.0, my loss gets stuck at 1.0. With `ArcFace` the loss gets kind of stuck at ~10.0 (btw, what is the expected loss when using ArcFace?). Right now I'm trying hard to tune the hyperparameters of `Arcface`, and I'm also trying various implementations of this algorithm (without any success).\n\nArcFace hyperparameters:\n* Scale (default seems to be ~30.0). \n* Margin (default seems to be ~0.5)\n\nTriplet loss hyperparameters:\n* Margin (default seems to be ~0.5)\n* K (either 2 or 4; actually a hyperparameter of my `DataGenerator`), which is the number examples of each ID in the minibatch. So if K=4, a minibatch of 16 would have labels like this `[1,1,1,1,7,7,7,7,15,15,15,15,28,28,28,28]`. Then the loss function would figure out the best triplets (Anchor (e.g. 1), Negative (e.g. 7), Positive (e.g. 1)) based on the embeddings.\n\n**So I'm looking for you professionals** for some advice on how to make my model learn! I'm sure someone has struggled through this like I'm doing now :-) \n\nAdditional info:\nMy `backbone/ConvNet` is just a regular `ResNet50` or `EfficientNetB0` followed by a global average pooling and a dense layer. The output of this `ConvNet+avg+dense` network is then fed to `ArcFace -&gt; CategoricalCrossentopy`, or, in the case of the triplet loss, the output of the `ConvNet+avg+dense` is passed to the `triplet_hard loss function`. I've also tried `triplet_semihard`.",
    "923261": "[EDITED/UPDATED] It seems that the loss is now decreasing steadily (going below 10) with ArcFace. The loss starts really high (~21) and goes down to ~5 after 50k iterations (batch size of 32). For you interested, here's the implementation I use (feel free to give feedback!):\n\n```\nclass ArcMarginProduct(tf.keras.layers.Layer):\n    '''\n    References:\n        https://arxiv.org/pdf/1801.07698.pdf\n        https://github.com/lyakaap/Landmark2019-1st-and-3rd-Place-Solution/\n            blob/master/src/modeling/metric_learning.py\n    '''\n    def __init__(self, n_classes, s=30, m=0.30, easy_margin=False,\n                 ls_eps=0.0, **kwargs):\n\n        super(ArcMarginProduct, self).__init__(**kwargs)\n\n        self.n_classes = n_classes\n        self.s = s\n        self.m = m\n        self.ls_eps = ls_eps\n        self.easy_margin = easy_margin\n        self.cos_m = tf.math.cos(m)\n        self.sin_m = tf.math.sin(m)\n        self.th = tf.math.cos(math.pi - m)\n        self.mm = tf.math.sin(math.pi - m) * m\n\n    def build(self, input_shape):\n        super(ArcMarginProduct, self).build(input_shape[0])\n        self.W = self.add_weight(\n            name='W',\n            shape=(int(input_shape[0][-1]), self.n_classes),\n            initializer='glorot_uniform',\n            dtype='float32',\n            trainable=True)\n\n    def call(self, inputs):\n        X, y = inputs\n        y = tf.cast(y, dtype=tf.int32)\n        cosine = tf.matmul(\n            tf.math.l2_normalize(X, axis=1),\n            tf.math.l2_normalize(self.W, axis=0)\n        )\n        sine = tf.math.sqrt(1.0 - tf.math.pow(cosine, 2))\n        phi = cosine * self.cos_m - sine * self.sin_m\n        if self.easy_margin:\n            phi = tf.where(cosine &gt; 0, phi, cosine)\n        else:\n            phi = tf.where(cosine &gt; self.th, phi, cosine - self.mm)\n        one_hot = tf.cast(\n            tf.one_hot(y, depth=self.n_classes),\n            dtype=cosine.dtype\n        )\n        if self.ls_eps &gt; 0:\n            one_hot = (1 - self.ls_eps) * one_hot + self.ls_eps / self.n_classes\n\n        output = (one_hot * phi) + ((1.0 - one_hot) * cosine)\n        output *= self.s\n        return output\n\n\nclass NeuralNet(tf.keras.Model):\n\n    def __init__(self, input_shape, num_classes):\n\n        super(NeuralNet, self).__init__()\n\n        self.engine = ResNet50(\n            include_top=False,\n            input_shape=input_shape,\n            weights=\"imagenet\")\n\n        self.pool = tf.keras.layers.GlobalAveragePooling2D()\n        self.batch_norm = tf.keras.layers.BatchNormalization()\n        self.dropout = tf.keras.layers.Dropout(0.0)\n        self.dense = tf.keras.layers.Dense(512)\n        self.arcface = ArcMarginProduct(scale=30, margin=0.3,\n              n_classes=n_classes, dtype='float32', name='arcface')\n\n    def call(self, inputs, **kwargs):\n        x = self.engine(inputs[0])\n        x = self.pool(x)\n        x = self.dropout(x)\n        x = self.dense(x)\n        x = self.batch_norm(x)\n        x = self.arcface([x, inputs[1]])\n        return x\n\n# loss_function = SparseCategoricalCrossentropy(..., from_logits=True)\n# optimizer = tf.keras.optimizers.SGD(1e-3, momentum=0.9) # with decay schedule\n```",
    "923358": "Hi @akensert ,\nHere we met again. xD\nI wanted to know about dataset. \nLike right now, I am not blindly training. Just wanted to know landmark_ids of training set and their corresponding indexes? Like labels for index files?",
    "932148": "what's your accuracy when the loss is below 10, and what epoch is it on?\n\nbtw DELG paper used margin of 0.1. I think scale is 1, but I may be wrong. \n\nFor me, when the margin is 0.1, the accuracy keeps going down in the first epoch (from 0.2% to 0.02%).",
    "932717": "Thanks @suruili, I'm trying these hyperparameters now (according to the DELG paper) and the loss is pretty much stuck already from the beginning. Right now I only look at the loss, perhaps I should implement some kind of accuracy function so that I can see the accuracy throughout training.",
    "933203": "Yeah, I find that setting the scale larger help. I am trying with 30 and 60",
    "933433": "Indeed, the DELG paper did not specify the initial value for the scale factor: we initialize it to sqrt(embedding_dimension), which in our case would lead to sqrt(2048) ~= 45\n\nWill include this in the next updated version of the paper, thanks for bringing it up.",
    "933551": "andrefaraujo Thanks for the info! I have some questions about the scale factor. Why do you think it affect training? Is it for numerical stability, or is the relevant scale of the scale factor vs margin matter? And why would you set it to sqrt of embedding dimension?",
    "937255": "The scale factor is necessary due to the cross-entropy loss becoming bounded for normalized class weights / features. See proposition 2 in the [NormFace paper](https://arxiv.org/pdf/1704.06369.pdf).",
    "937490": "Thanks for sharing ArcMarginProduct code.\n\n\nHow diverse are your predictions?\nIn my case, predictions are too concentrated in one class (which has more than 6000 images) and output embeddings \n are similar after 2 epoch\nmargin penalty(0.5) might be too big\n\n(2020 7 23 add)\nLast years chmpion might set margin=0.3\nhttps://arxiv.org/pdf/1906.04087.pdf",
    "981532": "akensert can you please share how you used this model on prediction?\nLike how you removed last layer from forward function. \nI am pytorch user new to tensorflow. Kindly help. xD",
    "981830": "micheomaano check out the public kernel. Or somethnig like this:\n```\n\n...\nmodel.load_weights('your_finetuned_weights.h5')\n\nnewmodel = tf.keras.Model(\n    inputs=model.model.get_layer('input/image').input,\n    outputs=model.model.get_layer('head/dense').output)\n\noutput = newmodel(image[tf.newaxis])[0]\nprint(output.shape)\n\n# >> out:\n# (512,)\n\n```\nAssuming that you've named the layers 'input/image' and 'head/dense': `tf.keras.layers.Input(..., name='input/image')` and `tf.keras.layers.Dense(..., name='head/dense')`  :)\n\n\nHope it helps!",
    "981888": "Got it. Thanks a lot."
  },
  "source": "meta"
}