{
  "id": 176099,
  "title": "Small question about stopping gradients in DELF implementation",
  "url": "/competitions/landmark-recognition-2020/discussion/176099",
  "author_name": "",
  "post_date": "2020-08-20T14:04:11.632631700Z",
  "votes": 5,
  "comment_count": 9,
  "views": 0,
  "content": "<p>So I've studied the source code of the organizers DELF implementation, e.g. <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/train.py\" target=\"_blank\">here (train.py)</a> and <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/delf_model.py\" target=\"_blank\">here (delf_model.py)</a>. (Many thanks for sharing the source code, has been of great help!)</p>\n<p>In the <a href=\"https://arxiv.org/pdf/2001.05027.pdf\" target=\"_blank\">paper</a> you/they mention about the benefit of stopping the gradients of the attention model to back-propagate through the backbone. Here's the code snippet of the attention model's train step (from train.py; some code excluded):</p>\n<pre><code> with tf.GradientTape() as attn_tape:\n    block3 = blocks['block3']  # pytype: disable=key-error\n    # Stopping gradients according to DELG paper:\n    # (https://arxiv.org/abs/2001.05027).\n    block3 = tf.stop_gradient(block3)\n    prelogits, scores, _ = model.attention(block3, training=True)\n    logits = model.attn_classification(prelogits)\n    attn_loss = compute_loss(labels, logits)\n# Backprop only through attention weights.\n_backprop_loss(attn_tape, attn_loss, model.attn_trainable_weights)\n</code></pre>\n<p>However, does the <code>tf.stop_gradient()</code> do anything at all here? To my understanding, the attention model inputs <code>block3</code> and only [forward] propagates through the attention layers (and only <code>attn_trainable_weights</code> are passed to <code>_backprop_loss</code>). Thus, whether or not the <code>tf.stop_gradient</code> is removed, the weights that will be updated are the attention weights only, no?</p>\n<p>Lastly, it doesn't seem that the source code have the autoencoder (and the reconstruction loss) implemented, correct?</p>",
  "messages": [
    {
      "id": "978922",
      "postDate": "08/20/2020 14:04:11",
      "content": "<p>So I've studied the source code of the organizers DELF implementation, e.g. <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/train.py\" target=\"_blank\">here (train.py)</a> and <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/delf_model.py\" target=\"_blank\">here (delf_model.py)</a>. (Many thanks for sharing the source code, has been of great help!)</p>\n<p>In the <a href=\"https://arxiv.org/pdf/2001.05027.pdf\" target=\"_blank\">paper</a> you/they mention about the benefit of stopping the gradients of the attention model to back-propagate through the backbone. Here's the code snippet of the attention model's train step (from train.py; some code excluded):</p>\n<pre><code> with tf.GradientTape() as attn_tape:\n    block3 = blocks['block3']  # pytype: disable=key-error\n    # Stopping gradients according to DELG paper:\n    # (https://arxiv.org/abs/2001.05027).\n    block3 = tf.stop_gradient(block3)\n    prelogits, scores, _ = model.attention(block3, training=True)\n    logits = model.attn_classification(prelogits)\n    attn_loss = compute_loss(labels, logits)\n# Backprop only through attention weights.\n_backprop_loss(attn_tape, attn_loss, model.attn_trainable_weights)\n</code></pre>\n<p>However, does the <code>tf.stop_gradient()</code> do anything at all here? To my understanding, the attention model inputs <code>block3</code> and only [forward] propagates through the attention layers (and only <code>attn_trainable_weights</code> are passed to <code>_backprop_loss</code>). Thus, whether or not the <code>tf.stop_gradient</code> is removed, the weights that will be updated are the attention weights only, no?</p>\n<p>Lastly, it doesn't seem that the source code have the autoencoder (and the reconstruction loss) implemented, correct?</p>",
      "rawMarkdown": "So I've studied the source code of the organizers DELF implementation, e.g. [here (train.py)](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/train.py) and [here (delf_model.py)](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/delf_model.py). (Many thanks for sharing the source code, has been of great help!)\n\nIn the [paper](https://arxiv.org/pdf/2001.05027.pdf) you/they mention about the benefit of stopping the gradients of the attention model to back-propagate through the backbone. Here's the code snippet of the attention model's train step (from train.py; some code excluded):\n```\n with tf.GradientTape() as attn_tape:\n    block3 = blocks['block3']  # pytype: disable=key-error\n    # Stopping gradients according to DELG paper:\n    # (https://arxiv.org/abs/2001.05027).\n    block3 = tf.stop_gradient(block3)\n    prelogits, scores, _ = model.attention(block3, training=True)\n    logits = model.attn_classification(prelogits)\n    attn_loss = compute_loss(labels, logits)\n# Backprop only through attention weights.\n_backprop_loss(attn_tape, attn_loss, model.attn_trainable_weights)\n```\nHowever, does the `tf.stop_gradient()` do anything at all here? To my understanding, the attention model inputs `block3` and only [forward] propagates through the attention layers (and only `attn_trainable_weights` are passed to `_backprop_loss`). Thus, whether or not the `tf.stop_gradient` is removed, the weights that will be updated are the attention weights only, no?\n\nLastly, it doesn't seem that the source code have the autoencoder (and the reconstruction loss) implemented, correct?",
      "votes": null
    },
    {
      "id": "979017",
      "postDate": "08/20/2020 15:11:40",
      "content": "<blockquote>\n  <p>However, does the tf.stop_gradient() do anything at all here?</p>\n</blockquote>\n<p><code>tf.stop_gradient</code> \"hides\" block 3 inputs from the gradient generator so they aren't watched  and so not taken into account during gradient computation.</p>\n<p>Reference: <a href=\"https://www.tensorflow.org/guide/advanced_autodiff\" target=\"_blank\">advanced_autodiff</a></p>",
      "rawMarkdown": "> However, does the tf.stop_gradient() do anything at all here?\n\n`tf.stop_gradient` \"hides\" block 3 inputs from the gradient generator so they aren't watched  and so not taken into account during gradient computation.\n\nReference: [advanced_autodiff](https://www.tensorflow.org/guide/advanced_autodiff)",
      "votes": null
    },
    {
      "id": "979100",
      "postDate": "08/20/2020 16:06:37",
      "content": "<p>Hmm okay, perhaps some confusion from my part then. So if <code>tf.stop_gradient</code> was removed, block3, <em>but not the preceding layers (block2, block1, etc.)</em>, would be updated based on the attention loss? </p>\n<p>To clarify, at first I thought the purpose of stopping the gradient for block3 was to stop the gradient for block3 + all preceeding layers (based on the attention loss), which didn't make sense [to me] by looking at the code. Although isn't that what it says in the paper, something like (very rough paraphrasing here): \"to stop the gradient to backpropagate through the backbone\".</p>",
      "rawMarkdown": "Hmm okay, perhaps some confusion from my part then. So if `tf.stop_gradient` was removed, block3, *but not the preceding layers (block2, block1, etc.)*, would be updated based on the attention loss? \n\nTo clarify, at first I thought the purpose of stopping the gradient for block3 was to stop the gradient for block3 + all preceeding layers (based on the attention loss), which didn't make sense [to me] by looking at the code. Although isn't that what it says in the paper, something like (very rough paraphrasing here): \"to stop the gradient to backpropagate through the backbone\".",
      "votes": null
    },
    {
      "id": "979129",
      "postDate": "08/20/2020 16:28:06",
      "content": "<p><code>tf.stop_gradient</code> will prevent back-propagation <strong>from continuing past a given node in the graph</strong>. In otherwords, gradients will be computed up-to the node where they are stopped, then no further updates will occur beyond that node. <br>\nThis should add clarity to my earlier response.</p>\n<blockquote>\n  <p>at first I thought the purpose of stopping the gradient for block3 was to stop the gradient for block3 + all preceeding layers</p>\n</blockquote>\n<p>Your right</p>",
      "rawMarkdown": "`tf.stop_gradient` will prevent back-propagation **from continuing past a given node in the graph**. In otherwords, gradients will be computed up-to the node where they are stopped, then no further updates will occur beyond that node. \nThis should add clarity to my earlier response.\n> at first I thought the purpose of stopping the gradient for block3 was to stop the gradient for block3 + all preceeding layers\n\nYour right",
      "votes": null
    },
    {
      "id": "980209",
      "postDate": "08/21/2020 12:08:05",
      "content": "<blockquote>\n  <p>\"to stop the gradient to backpropagate through the backbone\".</p>\n</blockquote>\n<p>In a network usually the shallow layers learn denser features with little semantic content, i.e. localized features agnostic to classes. Whereas denser layers do the opposite, i.e. sparse features very sensitive to classes. </p>\n<p>Since the local head of Delg is trained on shallow layer (although not that shallow), if we allow the gradients to propagate to the backbone, the shallow layers now will learn characteristics of deep layers (because they are much closer to the loss now), which disrupts the structure which is usually expected for a convnet. This is especially bad for the global descriptor, which is located at the deepest layers of the net.</p>",
      "rawMarkdown": "> \"to stop the gradient to backpropagate through the backbone\".\n\nIn a network usually the shallow layers learn denser features with little semantic content, i.e. localized features agnostic to classes. Whereas denser layers do the opposite, i.e. sparse features very sensitive to classes. \n\nSince the local head of Delg is trained on shallow layer (although not that shallow), if we allow the gradients to propagate to the backbone, the shallow layers now will learn characteristics of deep layers (because they are much closer to the loss now), which disrupts the structure which is usually expected for a convnet. This is especially bad for the global descriptor, which is located at the deepest layers of the net.",
      "votes": null
    },
    {
      "id": "980756",
      "postDate": "08/21/2020 20:25:39",
      "content": "<p>Thanks for the replies so far, they clarified some things for me. Although (sorry for persisting a bit), one of my main questions/points was that it doesn't matter if <code>tf.stop_gradient(block3)</code> is used or not in <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/train.py\" target=\"_blank\">their implementation</a>. Some output below, based on my implementation, which is similar to theirs (I obtain block4 (input to the attention model) and block5 from keras resnet-50):</p>\n<pre><code>with tf.GradientTape() as attn_tape:  \n     feat_block4 = tf.stop_gradient(feat_block4)\n     probs = self.model.forward_prop_attn(feat_block4, training=True)\n     attn_loss = self._compute_loss(labels, probs)\n     self._backprop_loss(attn_tape, attn_loss, self.model.get_attention_weights) # printing out gradients (list of tensors) inside here\n\n&gt;&gt; out:\n# gradients when removing tf.stop_gradient\n[&lt;tf.Tensor 'mul_217:0' shape=(1, 1, 1024, 512) dtype=float32&gt;, &lt;tf.Tensor 'mul_218:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_219:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_220:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_221:0' shape=(1, 1, 512, 1) dtype=float32&gt;, &lt;tf.Tensor 'mul_222:0' shape=(1,) dtype=float32&gt;, &lt;tf.Tensor 'mul_223:0' shape=(1024, 81313) dtype=float32&gt;, &lt;tf.Tensor 'mul_224:0' shape=(81313,) dtype=float32&gt;]\n\n# gradients with tf.stop_gradient\n[&lt;tf.Tensor 'mul_217:0' shape=(1, 1, 1024, 512) dtype=float32&gt;, &lt;tf.Tensor 'mul_218:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_219:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_220:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_221:0' shape=(1, 1, 512, 1) dtype=float32&gt;, &lt;tf.Tensor 'mul_222:0' shape=(1,) dtype=float32&gt;, &lt;tf.Tensor 'mul_223:0' shape=(1024, 81313) dtype=float32&gt;, &lt;tf.Tensor 'mul_224:0' shape=(81313,) dtype=float32&gt;]\n</code></pre>",
      "rawMarkdown": "Thanks for the replies so far, they clarified some things for me. Although (sorry for persisting a bit), one of my main questions/points was that it doesn't matter if `tf.stop_gradient(block3)` is used or not in [their implementation](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/train.py). Some output below, based on my implementation, which is similar to theirs (I obtain block4 (input to the attention model) and block5 from keras resnet-50):\n```\nwith tf.GradientTape() as attn_tape:  \n     feat_block4 = tf.stop_gradient(feat_block4)\n     probs = self.model.forward_prop_attn(feat_block4, training=True)\n     attn_loss = self._compute_loss(labels, probs)\n     self._backprop_loss(attn_tape, attn_loss, self.model.get_attention_weights) # printing out gradients (list of tensors) inside here\n\n>> out:\n# gradients when removing tf.stop_gradient\n[<tf.Tensor 'mul_217:0' shape=(1, 1, 1024, 512) dtype=float32>, <tf.Tensor 'mul_218:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_219:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_220:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_221:0' shape=(1, 1, 512, 1) dtype=float32>, <tf.Tensor 'mul_222:0' shape=(1,) dtype=float32>, <tf.Tensor 'mul_223:0' shape=(1024, 81313) dtype=float32>, <tf.Tensor 'mul_224:0' shape=(81313,) dtype=float32>]\n\n# gradients with tf.stop_gradient\n[<tf.Tensor 'mul_217:0' shape=(1, 1, 1024, 512) dtype=float32>, <tf.Tensor 'mul_218:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_219:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_220:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_221:0' shape=(1, 1, 512, 1) dtype=float32>, <tf.Tensor 'mul_222:0' shape=(1,) dtype=float32>, <tf.Tensor 'mul_223:0' shape=(1024, 81313) dtype=float32>, <tf.Tensor 'mul_224:0' shape=(81313,) dtype=float32>]\n```",
      "votes": null
    },
    {
      "id": "980890",
      "postDate": "08/22/2020 01:16:52",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> , good question.</p>\n<p>For the DELG paper, all models and experiments were based on an internal TF1-based implementation. We are currently rewriting DELG training in TF2, and pushing updates to the github codebase little by little (we're also learning TF2 in the process :).</p>\n<p>I am looking at your comments, and I think you are right actually: since in the first operation within the attention GradientTape is the attention layer, I think it will only backpropagate up until that point -- so, in this specific implementation, the stop_gradient call is not required.</p>\n<p>As you may have seen, there is a related TODO in the same file, to unify the backprop into a single pass (which is actually what we do in the internal codebase):</p>\n<pre><code># TODO(andrearaujo): we should try to unify the backprop into a single\n# function, instead of applying once to descriptor then to attention.\n</code></pre>\n<p>This will simplify this TF2 training code, and in this case the stop_gradient call will be needed, since everything would be covered by the same GradientTape.</p>\n<p>Thanks for pointing this out! I guess our current code can be made one line cleaner :)</p>",
      "rawMarkdown": "Hi @akensert , good question.\n\nFor the DELG paper, all models and experiments were based on an internal TF1-based implementation. We are currently rewriting DELG training in TF2, and pushing updates to the github codebase little by little (we're also learning TF2 in the process :).\n\nI am looking at your comments, and I think you are right actually: since in the first operation within the attention GradientTape is the attention layer, I think it will only backpropagate up until that point -- so, in this specific implementation, the stop_gradient call is not required.\n\nAs you may have seen, there is a related TODO in the same file, to unify the backprop into a single pass (which is actually what we do in the internal codebase):\n\n```\n# TODO(andrearaujo): we should try to unify the backprop into a single\n# function, instead of applying once to descriptor then to attention.\n```\n\nThis will simplify this TF2 training code, and in this case the stop_gradient call will be needed, since everything would be covered by the same GradientTape.\n\nThanks for pointing this out! I guess our current code can be made one line cleaner :)",
      "votes": null
    },
    {
      "id": "981826",
      "postDate": "08/22/2020 18:11:10",
      "content": "<p>Thanks for clarifying/confirming <a href=\"https://www.kaggle.com/andrefaraujo\" target=\"_blank\">@andrefaraujo</a>! And again, thanks for sharing your work and source code :) I think it's great to both have the paper and the code when learning about your model/implementation. (I'm only missing the code for the additional autoencoder/reconstruction loss ;) Although you describe it nicely in the paper so it should be fine!)</p>",
      "rawMarkdown": "Thanks for clarifying/confirming @andrefaraujo! And again, thanks for sharing your work and source code :) I think it's great to both have the paper and the code when learning about your model/implementation. (I'm only missing the code for the additional autoencoder/reconstruction loss ;) Although you describe it nicely in the paper so it should be fine!)",
      "votes": null
    },
    {
      "id": "985995",
      "postDate": "08/26/2020 06:26:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a>, how you train DELF? local or colab? </p>",
      "rawMarkdown": "Hi @akensert, how you train DELF? local or colab?",
      "votes": null
    },
    {
      "id": "986070",
      "postDate": "08/26/2020 07:25:44",
      "content": "<p><a href=\"https://www.kaggle.com/cswwp347724\" target=\"_blank\">@cswwp347724</a> Local</p>",
      "rawMarkdown": "cswwp347724 Local",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 979017,
      "author_name": "stephenmugisha",
      "author_url": "",
      "post_date": "08/20/2020 15:11:40",
      "content": "<blockquote>\n  <p>However, does the tf.stop_gradient() do anything at all here?</p>\n</blockquote>\n<p><code>tf.stop_gradient</code> \"hides\" block 3 inputs from the gradient generator so they aren't watched  and so not taken into account during gradient computation.</p>\n<p>Reference: <a href=\"https://www.tensorflow.org/guide/advanced_autodiff\" target=\"_blank\">advanced_autodiff</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 979100,
          "author_name": "akensert",
          "author_url": "",
          "post_date": "08/20/2020 16:06:37",
          "content": "<p>Hmm okay, perhaps some confusion from my part then. So if <code>tf.stop_gradient</code> was removed, block3, <em>but not the preceding layers (block2, block1, etc.)</em>, would be updated based on the attention loss? </p>\n<p>To clarify, at first I thought the purpose of stopping the gradient for block3 was to stop the gradient for block3 + all preceeding layers (based on the attention loss), which didn't make sense [to me] by looking at the code. Although isn't that what it says in the paper, something like (very rough paraphrasing here): \"to stop the gradient to backpropagate through the backbone\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979129,
          "author_name": "stephenmugisha",
          "author_url": "",
          "post_date": "08/20/2020 16:28:06",
          "content": "<p><code>tf.stop_gradient</code> will prevent back-propagation <strong>from continuing past a given node in the graph</strong>. In otherwords, gradients will be computed up-to the node where they are stopped, then no further updates will occur beyond that node. <br>\nThis should add clarity to my earlier response.</p>\n<blockquote>\n  <p>at first I thought the purpose of stopping the gradient for block3 was to stop the gradient for block3 + all preceeding layers</p>\n</blockquote>\n<p>Your right</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980209,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "08/21/2020 12:08:05",
          "content": "<blockquote>\n  <p>\"to stop the gradient to backpropagate through the backbone\".</p>\n</blockquote>\n<p>In a network usually the shallow layers learn denser features with little semantic content, i.e. localized features agnostic to classes. Whereas denser layers do the opposite, i.e. sparse features very sensitive to classes. </p>\n<p>Since the local head of Delg is trained on shallow layer (although not that shallow), if we allow the gradients to propagate to the backbone, the shallow layers now will learn characteristics of deep layers (because they are much closer to the loss now), which disrupts the structure which is usually expected for a convnet. This is especially bad for the global descriptor, which is located at the deepest layers of the net.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 980756,
      "author_name": "akensert",
      "author_url": "",
      "post_date": "08/21/2020 20:25:39",
      "content": "<p>Thanks for the replies so far, they clarified some things for me. Although (sorry for persisting a bit), one of my main questions/points was that it doesn't matter if <code>tf.stop_gradient(block3)</code> is used or not in <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/train.py\" target=\"_blank\">their implementation</a>. Some output below, based on my implementation, which is similar to theirs (I obtain block4 (input to the attention model) and block5 from keras resnet-50):</p>\n<pre><code>with tf.GradientTape() as attn_tape:  \n     feat_block4 = tf.stop_gradient(feat_block4)\n     probs = self.model.forward_prop_attn(feat_block4, training=True)\n     attn_loss = self._compute_loss(labels, probs)\n     self._backprop_loss(attn_tape, attn_loss, self.model.get_attention_weights) # printing out gradients (list of tensors) inside here\n\n&gt;&gt; out:\n# gradients when removing tf.stop_gradient\n[&lt;tf.Tensor 'mul_217:0' shape=(1, 1, 1024, 512) dtype=float32&gt;, &lt;tf.Tensor 'mul_218:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_219:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_220:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_221:0' shape=(1, 1, 512, 1) dtype=float32&gt;, &lt;tf.Tensor 'mul_222:0' shape=(1,) dtype=float32&gt;, &lt;tf.Tensor 'mul_223:0' shape=(1024, 81313) dtype=float32&gt;, &lt;tf.Tensor 'mul_224:0' shape=(81313,) dtype=float32&gt;]\n\n# gradients with tf.stop_gradient\n[&lt;tf.Tensor 'mul_217:0' shape=(1, 1, 1024, 512) dtype=float32&gt;, &lt;tf.Tensor 'mul_218:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_219:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_220:0' shape=(512,) dtype=float32&gt;, &lt;tf.Tensor 'mul_221:0' shape=(1, 1, 512, 1) dtype=float32&gt;, &lt;tf.Tensor 'mul_222:0' shape=(1,) dtype=float32&gt;, &lt;tf.Tensor 'mul_223:0' shape=(1024, 81313) dtype=float32&gt;, &lt;tf.Tensor 'mul_224:0' shape=(81313,) dtype=float32&gt;]\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 980890,
          "author_name": "andrefaraujo",
          "author_url": "",
          "post_date": "08/22/2020 01:16:52",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> , good question.</p>\n<p>For the DELG paper, all models and experiments were based on an internal TF1-based implementation. We are currently rewriting DELG training in TF2, and pushing updates to the github codebase little by little (we're also learning TF2 in the process :).</p>\n<p>I am looking at your comments, and I think you are right actually: since in the first operation within the attention GradientTape is the attention layer, I think it will only backpropagate up until that point -- so, in this specific implementation, the stop_gradient call is not required.</p>\n<p>As you may have seen, there is a related TODO in the same file, to unify the backprop into a single pass (which is actually what we do in the internal codebase):</p>\n<pre><code># TODO(andrearaujo): we should try to unify the backprop into a single\n# function, instead of applying once to descriptor then to attention.\n</code></pre>\n<p>This will simplify this TF2 training code, and in this case the stop_gradient call will be needed, since everything would be covered by the same GradientTape.</p>\n<p>Thanks for pointing this out! I guess our current code can be made one line cleaner :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 981826,
          "author_name": "akensert",
          "author_url": "",
          "post_date": "08/22/2020 18:11:10",
          "content": "<p>Thanks for clarifying/confirming <a href=\"https://www.kaggle.com/andrefaraujo\" target=\"_blank\">@andrefaraujo</a>! And again, thanks for sharing your work and source code :) I think it's great to both have the paper and the code when learning about your model/implementation. (I'm only missing the code for the additional autoencoder/reconstruction loss ;) Although you describe it nicely in the paper so it should be fine!)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 985995,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "08/26/2020 06:26:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a>, how you train DELF? local or colab? </p>",
      "votes": null,
      "replies": [
        {
          "id": 986070,
          "author_name": "akensert",
          "author_url": "",
          "post_date": "08/26/2020 07:25:44",
          "content": "<p><a href=\"https://www.kaggle.com/cswwp347724\" target=\"_blank\">@cswwp347724</a> Local</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "978922": "So I've studied the source code of the organizers DELF implementation, e.g. [here (train.py)](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/train.py) and [here (delf_model.py)](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/delf_model.py). (Many thanks for sharing the source code, has been of great help!)\n\nIn the [paper](https://arxiv.org/pdf/2001.05027.pdf) you/they mention about the benefit of stopping the gradients of the attention model to back-propagate through the backbone. Here's the code snippet of the attention model's train step (from train.py; some code excluded):\n```\n with tf.GradientTape() as attn_tape:\n    block3 = blocks['block3']  # pytype: disable=key-error\n    # Stopping gradients according to DELG paper:\n    # (https://arxiv.org/abs/2001.05027).\n    block3 = tf.stop_gradient(block3)\n    prelogits, scores, _ = model.attention(block3, training=True)\n    logits = model.attn_classification(prelogits)\n    attn_loss = compute_loss(labels, logits)\n# Backprop only through attention weights.\n_backprop_loss(attn_tape, attn_loss, model.attn_trainable_weights)\n```\nHowever, does the `tf.stop_gradient()` do anything at all here? To my understanding, the attention model inputs `block3` and only [forward] propagates through the attention layers (and only `attn_trainable_weights` are passed to `_backprop_loss`). Thus, whether or not the `tf.stop_gradient` is removed, the weights that will be updated are the attention weights only, no?\n\nLastly, it doesn't seem that the source code have the autoencoder (and the reconstruction loss) implemented, correct?",
    "979017": "> However, does the tf.stop_gradient() do anything at all here?\n\n`tf.stop_gradient` \"hides\" block 3 inputs from the gradient generator so they aren't watched  and so not taken into account during gradient computation.\n\nReference: [advanced_autodiff](https://www.tensorflow.org/guide/advanced_autodiff)",
    "979100": "Hmm okay, perhaps some confusion from my part then. So if `tf.stop_gradient` was removed, block3, *but not the preceding layers (block2, block1, etc.)*, would be updated based on the attention loss? \n\nTo clarify, at first I thought the purpose of stopping the gradient for block3 was to stop the gradient for block3 + all preceeding layers (based on the attention loss), which didn't make sense [to me] by looking at the code. Although isn't that what it says in the paper, something like (very rough paraphrasing here): \"to stop the gradient to backpropagate through the backbone\".",
    "979129": "`tf.stop_gradient` will prevent back-propagation **from continuing past a given node in the graph**. In otherwords, gradients will be computed up-to the node where they are stopped, then no further updates will occur beyond that node. \nThis should add clarity to my earlier response.\n> at first I thought the purpose of stopping the gradient for block3 was to stop the gradient for block3 + all preceeding layers\n\nYour right",
    "980209": "> \"to stop the gradient to backpropagate through the backbone\".\n\nIn a network usually the shallow layers learn denser features with little semantic content, i.e. localized features agnostic to classes. Whereas denser layers do the opposite, i.e. sparse features very sensitive to classes. \n\nSince the local head of Delg is trained on shallow layer (although not that shallow), if we allow the gradients to propagate to the backbone, the shallow layers now will learn characteristics of deep layers (because they are much closer to the loss now), which disrupts the structure which is usually expected for a convnet. This is especially bad for the global descriptor, which is located at the deepest layers of the net.",
    "980756": "Thanks for the replies so far, they clarified some things for me. Although (sorry for persisting a bit), one of my main questions/points was that it doesn't matter if `tf.stop_gradient(block3)` is used or not in [their implementation](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/train.py). Some output below, based on my implementation, which is similar to theirs (I obtain block4 (input to the attention model) and block5 from keras resnet-50):\n```\nwith tf.GradientTape() as attn_tape:  \n     feat_block4 = tf.stop_gradient(feat_block4)\n     probs = self.model.forward_prop_attn(feat_block4, training=True)\n     attn_loss = self._compute_loss(labels, probs)\n     self._backprop_loss(attn_tape, attn_loss, self.model.get_attention_weights) # printing out gradients (list of tensors) inside here\n\n>> out:\n# gradients when removing tf.stop_gradient\n[<tf.Tensor 'mul_217:0' shape=(1, 1, 1024, 512) dtype=float32>, <tf.Tensor 'mul_218:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_219:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_220:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_221:0' shape=(1, 1, 512, 1) dtype=float32>, <tf.Tensor 'mul_222:0' shape=(1,) dtype=float32>, <tf.Tensor 'mul_223:0' shape=(1024, 81313) dtype=float32>, <tf.Tensor 'mul_224:0' shape=(81313,) dtype=float32>]\n\n# gradients with tf.stop_gradient\n[<tf.Tensor 'mul_217:0' shape=(1, 1, 1024, 512) dtype=float32>, <tf.Tensor 'mul_218:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_219:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_220:0' shape=(512,) dtype=float32>, <tf.Tensor 'mul_221:0' shape=(1, 1, 512, 1) dtype=float32>, <tf.Tensor 'mul_222:0' shape=(1,) dtype=float32>, <tf.Tensor 'mul_223:0' shape=(1024, 81313) dtype=float32>, <tf.Tensor 'mul_224:0' shape=(81313,) dtype=float32>]\n```",
    "980890": "Hi @akensert , good question.\n\nFor the DELG paper, all models and experiments were based on an internal TF1-based implementation. We are currently rewriting DELG training in TF2, and pushing updates to the github codebase little by little (we're also learning TF2 in the process :).\n\nI am looking at your comments, and I think you are right actually: since in the first operation within the attention GradientTape is the attention layer, I think it will only backpropagate up until that point -- so, in this specific implementation, the stop_gradient call is not required.\n\nAs you may have seen, there is a related TODO in the same file, to unify the backprop into a single pass (which is actually what we do in the internal codebase):\n\n```\n# TODO(andrearaujo): we should try to unify the backprop into a single\n# function, instead of applying once to descriptor then to attention.\n```\n\nThis will simplify this TF2 training code, and in this case the stop_gradient call will be needed, since everything would be covered by the same GradientTape.\n\nThanks for pointing this out! I guess our current code can be made one line cleaner :)",
    "981826": "Thanks for clarifying/confirming @andrefaraujo! And again, thanks for sharing your work and source code :) I think it's great to both have the paper and the code when learning about your model/implementation. (I'm only missing the code for the additional autoencoder/reconstruction loss ;) Although you describe it nicely in the paper so it should be fine!)",
    "985995": "Hi @akensert, how you train DELF? local or colab?",
    "986070": "cswwp347724 Local"
  },
  "source": "meta"
}