{
  "id": 187344,
  "title": "Last-minute question: local descriptor keypoints in the DELG paper?",
  "url": "/competitions/landmark-recognition-2020/discussion/187344",
  "author_name": "",
  "post_date": "2020-09-28T14:55:56.628494800Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm doing some last-minute training (just started to train the local head few hours ago, and my global models haven't converged yet, panic mode activated 😃).</p>\n<hr>\n<p>I'm following the DELG paper and basically my local feature extraction head outputs an <strong>attention map</strong> of shape (h,w) and <strong>local descriptors</strong> of shape (h,w,c). I got stuck on how to compute the location of corresponding <strong>keypoints</strong> or <strong>bounding-boxes</strong>.</p>\n<p>I've found the following code:<br>\n<a href=\"https://github.com/tensorflow/models/blob/63620f4ce55a608e76c0ff4a76d64d90507bf997/research/delf/delf/python/training/model/export_model_utils.py#L277\" target=\"_blank\">https://github.com/tensorflow/models/blob/63620f4ce55a608e76c0ff4a76d64d90507bf997/research/delf/delf/python/training/model/export_model_utils.py#L277</a></p>\n<p>that seems to have what I need. However, I'm struggling with setting parameters:</p>\n<ul>\n<li><p>The DELF repo hard-codes the parameters:<br>\n<code>rf, stride, padding = [291.0, 16.0 * stride_factor, 145.0]</code><br>\nwhich are basically receptive field, stride, and padding for <code>block_3</code> outputs of ResNet50. So, my question is:</p>\n<ul>\n<li>What are the corresponding parameters for <code>conv5_block1_preact_relu</code> (basically activations after the outputs of <code>block_4</code> of ResNet152 V2?</li>\n<li>Same question, for <code>block7a_expand_activation</code> (basically activation after <code>block_6</code> expansion) of EfficientNetB6?</li></ul></li>\n<li><p>The default parameters of <code>IoU</code> for local branch is <code>1.0</code>. I'm wondering what was used in the DELG paper?</p></li>\n<li><p>I see that the attention score threshold is set to <code>175.0</code> in the baseline code. How it was chosen in the DELG paper? </p></li>\n</ul>\n<hr>\n<p>A huuuuuuge thanks in advance to anyone who will respond to this thread!</p>",
  "messages": [
    {
      "id": "1030312",
      "postDate": "09/28/2020 14:55:56",
      "content": "<p>I'm doing some last-minute training (just started to train the local head few hours ago, and my global models haven't converged yet, panic mode activated 😃).</p>\n<hr>\n<p>I'm following the DELG paper and basically my local feature extraction head outputs an <strong>attention map</strong> of shape (h,w) and <strong>local descriptors</strong> of shape (h,w,c). I got stuck on how to compute the location of corresponding <strong>keypoints</strong> or <strong>bounding-boxes</strong>.</p>\n<p>I've found the following code:<br>\n<a href=\"https://github.com/tensorflow/models/blob/63620f4ce55a608e76c0ff4a76d64d90507bf997/research/delf/delf/python/training/model/export_model_utils.py#L277\" target=\"_blank\">https://github.com/tensorflow/models/blob/63620f4ce55a608e76c0ff4a76d64d90507bf997/research/delf/delf/python/training/model/export_model_utils.py#L277</a></p>\n<p>that seems to have what I need. However, I'm struggling with setting parameters:</p>\n<ul>\n<li><p>The DELF repo hard-codes the parameters:<br>\n<code>rf, stride, padding = [291.0, 16.0 * stride_factor, 145.0]</code><br>\nwhich are basically receptive field, stride, and padding for <code>block_3</code> outputs of ResNet50. So, my question is:</p>\n<ul>\n<li>What are the corresponding parameters for <code>conv5_block1_preact_relu</code> (basically activations after the outputs of <code>block_4</code> of ResNet152 V2?</li>\n<li>Same question, for <code>block7a_expand_activation</code> (basically activation after <code>block_6</code> expansion) of EfficientNetB6?</li></ul></li>\n<li><p>The default parameters of <code>IoU</code> for local branch is <code>1.0</code>. I'm wondering what was used in the DELG paper?</p></li>\n<li><p>I see that the attention score threshold is set to <code>175.0</code> in the baseline code. How it was chosen in the DELG paper? </p></li>\n</ul>\n<hr>\n<p>A huuuuuuge thanks in advance to anyone who will respond to this thread!</p>",
      "rawMarkdown": "I'm doing some last-minute training (just started to train the local head few hours ago, and my global models haven't converged yet, panic mode activated 😃).\n\n--------------------------------------------------------\n\nI'm following the DELG paper and basically my local feature extraction head outputs an **attention map** of shape (h,w) and **local descriptors** of shape (h,w,c). I got stuck on how to compute the location of corresponding **keypoints** or **bounding-boxes**.\n\nI've found the following code:\nhttps://github.com/tensorflow/models/blob/63620f4ce55a608e76c0ff4a76d64d90507bf997/research/delf/delf/python/training/model/export_model_utils.py#L277\n\nthat seems to have what I need. However, I'm struggling with setting parameters:\n- The DELF repo hard-codes the parameters:\n  `rf, stride, padding = [291.0, 16.0 * stride_factor, 145.0]`\n  which are basically receptive field, stride, and padding for `block_3` outputs of ResNet50. So, my question is:\n  - What are the corresponding parameters for `conv5_block1_preact_relu` (basically activations after the outputs of `block_4` of ResNet152 V2?\n  - Same question, for `block7a_expand_activation` (basically activation after `block_6` expansion) of EfficientNetB6?\n\n- The default parameters of `IoU` for local branch is `1.0`. I'm wondering what was used in the DELG paper?\n- I see that the attention score threshold is set to `175.0` in the baseline code. How it was chosen in the DELG paper? \n\n-------------------------------------------------------------------------\n\nA huuuuuuge thanks in advance to anyone who will respond to this thread!",
      "votes": null
    },
    {
      "id": "1031018",
      "postDate": "09/29/2020 07:13:39",
      "content": "<p>What I recall seeing in one of the papers(either DELG or DELF), is that in the activation map's 75th percentile value was chosen as threshold . Anything above it to be selected as important activation and thus to be selcted as keypoint.</p>",
      "rawMarkdown": "What I recall seeing in one of the papers(either DELG or DELF), is that in the activation map's 75th percentile value was chosen as threshold . Anything above it to be selected as important activation and thus to be selcted as keypoint.",
      "votes": null
    },
    {
      "id": "1031579",
      "postDate": "09/29/2020 14:42:04",
      "content": "<p>Thanks for the answer! At the end, I did managed to train a local descriptor in the last minute (literally last minute) ^^ I used your threshold value, so my score gets increased in the end - I would need to thank you :D</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2270821%2Ff9c4e85e1dc6377ea2ff4af05ebba89a%2Fdownload%20(4).png?generation=1601390734073323&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for the answer! At the end, I did managed to train a local descriptor in the last minute (literally last minute) ^^ I used your threshold value, so my score gets increased in the end - I would need to thank you :D\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2270821%2Ff9c4e85e1dc6377ea2ff4af05ebba89a%2Fdownload%20(4).png?generation=1601390734073323&alt=media)",
      "votes": null
    },
    {
      "id": "1031680",
      "postDate": "09/29/2020 15:51:37",
      "content": "<p>Glad to hear it helped! You also might want to make scaled activation maps of the scaled inputs (kind of TTA for activations). Collect the keypoints and descriptors for scaled inputs separately and send it to the pydegensac , altogether. It'll increase the score for sure. However, very less time is left. You might want to give it a try. Time overhead for the additional attention map related process is quite less.</p>",
      "rawMarkdown": "Glad to hear it helped! You also might want to make scaled activation maps of the scaled inputs (kind of TTA for activations). Collect the keypoints and descriptors for scaled inputs separately and send it to the pydegensac , altogether. It'll increase the score for sure. However, very less time is left. You might want to give it a try. Time overhead for the additional attention map related process is quite less.",
      "votes": null
    },
    {
      "id": "1031684",
      "postDate": "09/29/2020 15:56:03",
      "content": "<p>Haha, thanks, but I'm done :) I do export local descriptors on 3 different scales though. I expect my solution to take a full 12 hours to run (my non-landmark removal pipeline is quite heavy, too), and there's only 8 hours left, so there's nothing I could do more right now 😄</p>\n<p>I also failed to include pre-computed embeddings in-time, because my global models haven't converged until this morning :)</p>",
      "rawMarkdown": "Haha, thanks, but I'm done :) I do export local descriptors on 3 different scales though. I expect my solution to take a full 12 hours to run (my non-landmark removal pipeline is quite heavy, too), and there's only 8 hours left, so there's nothing I could do more right now 😄\n\nI also failed to include pre-computed embeddings in-time, because my global models haven't converged until this morning :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1031018,
      "author_name": "dsnil87",
      "author_url": "",
      "post_date": "09/29/2020 07:13:39",
      "content": "<p>What I recall seeing in one of the papers(either DELG or DELF), is that in the activation map's 75th percentile value was chosen as threshold . Anything above it to be selected as important activation and thus to be selcted as keypoint.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1031579,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "09/29/2020 14:42:04",
          "content": "<p>Thanks for the answer! At the end, I did managed to train a local descriptor in the last minute (literally last minute) ^^ I used your threshold value, so my score gets increased in the end - I would need to thank you :D</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2270821%2Ff9c4e85e1dc6377ea2ff4af05ebba89a%2Fdownload%20(4).png?generation=1601390734073323&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1031680,
          "author_name": "dsnil87",
          "author_url": "",
          "post_date": "09/29/2020 15:51:37",
          "content": "<p>Glad to hear it helped! You also might want to make scaled activation maps of the scaled inputs (kind of TTA for activations). Collect the keypoints and descriptors for scaled inputs separately and send it to the pydegensac , altogether. It'll increase the score for sure. However, very less time is left. You might want to give it a try. Time overhead for the additional attention map related process is quite less.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1031684,
          "author_name": "chankhavu",
          "author_url": "",
          "post_date": "09/29/2020 15:56:03",
          "content": "<p>Haha, thanks, but I'm done :) I do export local descriptors on 3 different scales though. I expect my solution to take a full 12 hours to run (my non-landmark removal pipeline is quite heavy, too), and there's only 8 hours left, so there's nothing I could do more right now 😄</p>\n<p>I also failed to include pre-computed embeddings in-time, because my global models haven't converged until this morning :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1030312": "I'm doing some last-minute training (just started to train the local head few hours ago, and my global models haven't converged yet, panic mode activated 😃).\n\n--------------------------------------------------------\n\nI'm following the DELG paper and basically my local feature extraction head outputs an **attention map** of shape (h,w) and **local descriptors** of shape (h,w,c). I got stuck on how to compute the location of corresponding **keypoints** or **bounding-boxes**.\n\nI've found the following code:\nhttps://github.com/tensorflow/models/blob/63620f4ce55a608e76c0ff4a76d64d90507bf997/research/delf/delf/python/training/model/export_model_utils.py#L277\n\nthat seems to have what I need. However, I'm struggling with setting parameters:\n- The DELF repo hard-codes the parameters:\n  `rf, stride, padding = [291.0, 16.0 * stride_factor, 145.0]`\n  which are basically receptive field, stride, and padding for `block_3` outputs of ResNet50. So, my question is:\n  - What are the corresponding parameters for `conv5_block1_preact_relu` (basically activations after the outputs of `block_4` of ResNet152 V2?\n  - Same question, for `block7a_expand_activation` (basically activation after `block_6` expansion) of EfficientNetB6?\n\n- The default parameters of `IoU` for local branch is `1.0`. I'm wondering what was used in the DELG paper?\n- I see that the attention score threshold is set to `175.0` in the baseline code. How it was chosen in the DELG paper? \n\n-------------------------------------------------------------------------\n\nA huuuuuuge thanks in advance to anyone who will respond to this thread!",
    "1031018": "What I recall seeing in one of the papers(either DELG or DELF), is that in the activation map's 75th percentile value was chosen as threshold . Anything above it to be selected as important activation and thus to be selcted as keypoint.",
    "1031579": "Thanks for the answer! At the end, I did managed to train a local descriptor in the last minute (literally last minute) ^^ I used your threshold value, so my score gets increased in the end - I would need to thank you :D\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2270821%2Ff9c4e85e1dc6377ea2ff4af05ebba89a%2Fdownload%20(4).png?generation=1601390734073323&alt=media)",
    "1031680": "Glad to hear it helped! You also might want to make scaled activation maps of the scaled inputs (kind of TTA for activations). Collect the keypoints and descriptors for scaled inputs separately and send it to the pydegensac , altogether. It'll increase the score for sure. However, very less time is left. You might want to give it a try. Time overhead for the additional attention map related process is quite less.",
    "1031684": "Haha, thanks, but I'm done :) I do export local descriptors on 3 different scales though. I expect my solution to take a full 12 hours to run (my non-landmark removal pipeline is quite heavy, too), and there's only 8 hours left, so there's nothing I could do more right now 😄\n\nI also failed to include pre-computed embeddings in-time, because my global models haven't converged until this morning :)"
  },
  "source": "meta"
}