{
  "id": 420471,
  "title": "12nd Place Solution: SP+SG",
  "url": "/competitions/image-matching-challenge-2023/writeups/niubiren-12nd-place-solution-sp-sg",
  "author_name": "",
  "post_date": "2023-07-01T02:11:25.887Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to the competition organizers and Kaggle staff for hosting this amazing competition and solid support. Thanks to each my team member for their creative perspective and effort which gain us this place in competition.</p>\n<ol>\n<li><p>Overview<br>\nOur final model is rather simple. The solution is based on a modular toolbox named Hierarchical-Localization (implementation of typical SP+SG structure), and the key modification is to set the input resolution to 2000.<br>\nAs the resolution increased from 1024 to 2000, the public LB scores increased from 0.24 to 0.46. As the resolution further increased to 3000, there is still improvement in the trainset (especially \"wall\").</p></li>\n<li><p>Configuration in detail </p>\n<pre><code>confs = {\n    'superpoint_aachen': {\n        'output': 'feats-superpoint-n-r',\n        'model': {\n            'name': 'superpoint',\n            'nms_radius': ,\n            'max_keypoints': ,\n        },\n        'preprocessing': {\n            'grayscale': True,\n            'resize_max': , \n        },\n    }\n}\n</code></pre>\n<pre><code>'superglue': {\n    'name': 'superglue',\n    'weights': 'outdoor',\n    'sinkhorn_iterations': ,\n}\n</code></pre></li>\n<li><p>Tricks didn't work<br>\nOur code suffers from randomness; the difference in public LB score for the same code can be up to 0.04! Unfortunately, during the whole competition, we failed to find a way to eliminate randomness. We decide a modification useless if no obvious improvement observed during repetitive submissions, but still, there are possibly mistakes.</p>\n<ul>\n<li>Image retrieval<br>\nAfter incremental mapping , try to add the images mapped failed to the successfully reconstructed model.</li>\n<li>TTA<br>\nReverse and concatenate and apply NMS to features extracted from [original left-right-flip 10-deg-rotation], however, computation time hugely increased while no obvious improvement observed.</li>\n<li>Ensemble SIFT to SP<br>\nWe learned that SP fails when the image is extremely in-plane rotated. So we tried to perform SIFT extraction and matching using pycolmap and extract the match from the database and concatenate with SPSG. This trick can improve the score for \"cyprus\" but is useless in this year's test set. Worth noting is that it doesn't harm the LB score either.</li>\n<li>Multi-models and multi-resolution<br>\nWe tried to ensemble SPSG with DKM, Loftr, Quadtree, silk, but these models perform poorly despite being in the same resolution with SPSG. Some of them are one-stage models, which made them inefficient for the task. We tried to concatenate [800, 1500, 2000] resolutions calculated by SPSG; the result is nearly the same with 2000 alone.</li>\n<li>Suppress randomness by more iterations<br>\nHloc uses geometry verification (<code>pycolmap.verify_matches</code>), set <code>max_num_trials=40000</code> didn't work.</li></ul></li>\n<li><p>Some experience from the competition<br>\nReading images using cv2 can lose EXIF information, which is not conducive to reconstruction. But there's a counter-example, \"theater\" can score even higher after the information removed.</p></li>\n<li><p>Referrence</p></li>\n</ol>\n<ul>\n<li>Hloc<br>\n<a href=\"https://github.com/cvg/Hierarchical-Localization/\" target=\"_blank\">https://github.com/cvg/Hierarchical-Localization/</a></li>\n<li>Superpoint introduction<br>\n<a href=\"https://github.com/magicleap/SuperPointPretrainedNetwork/blob/master/assets/DL4VSLAM_talk.pdf\" target=\"_blank\">https://github.com/magicleap/SuperPointPretrainedNetwork/blob/master/assets/DL4VSLAM_talk.pdf</a></li>\n</ul>",
  "messages": [
    {
      "id": "2324902",
      "postDate": "07/01/2023 01:53:12",
      "content": "<p>Thanks to the competition organizers and Kaggle staff for hosting this amazing competition and solid support. Thanks to each my team member for their creative perspective and effort which gain us this place in competition.</p>\n<ol>\n<li><p>Overview<br>\nOur final model is rather simple. The solution is based on a modular toolbox named Hierarchical-Localization (implementation of typical SP+SG structure), and the key modification is to set the input resolution to 2000.<br>\nAs the resolution increased from 1024 to 2000, the public LB scores increased from 0.24 to 0.46. As the resolution further increased to 3000, there is still improvement in the trainset (especially \"wall\").</p></li>\n<li><p>Configuration in detail </p>\n<pre><code>confs = {\n    'superpoint_aachen': {\n        'output': 'feats-superpoint-n-r',\n        'model': {\n            'name': 'superpoint',\n            'nms_radius': ,\n            'max_keypoints': ,\n        },\n        'preprocessing': {\n            'grayscale': True,\n            'resize_max': , \n        },\n    }\n}\n</code></pre>\n<pre><code>'superglue': {\n    'name': 'superglue',\n    'weights': 'outdoor',\n    'sinkhorn_iterations': ,\n}\n</code></pre></li>\n<li><p>Tricks didn't work<br>\nOur code suffers from randomness; the difference in public LB score for the same code can be up to 0.04! Unfortunately, during the whole competition, we failed to find a way to eliminate randomness. We decide a modification useless if no obvious improvement observed during repetitive submissions, but still, there are possibly mistakes.</p>\n<ul>\n<li>Image retrieval<br>\nAfter incremental mapping , try to add the images mapped failed to the successfully reconstructed model.</li>\n<li>TTA<br>\nReverse and concatenate and apply NMS to features extracted from [original left-right-flip 10-deg-rotation], however, computation time hugely increased while no obvious improvement observed.</li>\n<li>Ensemble SIFT to SP<br>\nWe learned that SP fails when the image is extremely in-plane rotated. So we tried to perform SIFT extraction and matching using pycolmap and extract the match from the database and concatenate with SPSG. This trick can improve the score for \"cyprus\" but is useless in this year's test set. Worth noting is that it doesn't harm the LB score either.</li>\n<li>Multi-models and multi-resolution<br>\nWe tried to ensemble SPSG with DKM, Loftr, Quadtree, silk, but these models perform poorly despite being in the same resolution with SPSG. Some of them are one-stage models, which made them inefficient for the task. We tried to concatenate [800, 1500, 2000] resolutions calculated by SPSG; the result is nearly the same with 2000 alone.</li>\n<li>Suppress randomness by more iterations<br>\nHloc uses geometry verification (<code>pycolmap.verify_matches</code>), set <code>max_num_trials=40000</code> didn't work.</li></ul></li>\n<li><p>Some experience from the competition<br>\nReading images using cv2 can lose EXIF information, which is not conducive to reconstruction. But there's a counter-example, \"theater\" can score even higher after the information removed.</p></li>\n<li><p>Referrence</p></li>\n</ol>\n<ul>\n<li>Hloc<br>\n<a href=\"https://github.com/cvg/Hierarchical-Localization/\" target=\"_blank\">https://github.com/cvg/Hierarchical-Localization/</a></li>\n<li>Superpoint introduction<br>\n<a href=\"https://github.com/magicleap/SuperPointPretrainedNetwork/blob/master/assets/DL4VSLAM_talk.pdf\" target=\"_blank\">https://github.com/magicleap/SuperPointPretrainedNetwork/blob/master/assets/DL4VSLAM_talk.pdf</a></li>\n</ul>",
      "rawMarkdown": "Thanks to the competition organizers and Kaggle staff for hosting this amazing competition and solid support. Thanks to each my team member for their creative perspective and effort which gain us this place in competition.\n\n1. Overview\n Our final model is rather simple. The solution is based on a modular toolbox named Hierarchical-Localization (implementation of typical SP+SG structure), and the key modification is to set the input resolution to 2000.\n As the resolution increased from 1024 to 2000, the public LB scores increased from 0.24 to 0.46. As the resolution further increased to 3000, there is still improvement in the trainset (especially \"wall\").\n2. Configuration in detail \n\n    ```\n    confs = {\n        'superpoint_aachen': {\n            'output': 'feats-superpoint-n4096-r1024',\n            'model': {\n                'name': 'superpoint',\n                'nms_radius': 3,\n                'max_keypoints': 4096,\n            },\n            'preprocessing': {\n                'grayscale': True,\n                'resize_max': 2000, # Scaling longside to 2000\n            },\n        }\n    }\n    ```\n\n    ```\n    'superglue': {\n        'name': 'superglue',\n        'weights': 'outdoor',\n        'sinkhorn_iterations': 50,\n    }\n    ```\n\n\n3. Tricks didn't work\n   Our code suffers from randomness; the difference in public LB score for the same code can be up to 0.04! Unfortunately, during the whole competition, we failed to find a way to eliminate randomness. We decide a modification useless if no obvious improvement observed during repetitive submissions, but still, there are possibly mistakes.\n   - Image retrieval\n     After incremental mapping , try to add the images mapped failed to the successfully reconstructed model.\n   - TTA\n     Reverse and concatenate and apply NMS to features extracted from [original left-right-flip 10-deg-rotation], however, computation time hugely increased while no obvious improvement observed.\n   - Ensemble SIFT to SP\n     We learned that SP fails when the image is extremely in-plane rotated. So we tried to perform SIFT extraction and matching using pycolmap and extract the match from the database and concatenate with SPSG. This trick can improve the score for \"cyprus\" but is useless in this year's test set. Worth noting is that it doesn't harm the LB score either.\n   - Multi-models and multi-resolution\n     We tried to ensemble SPSG with DKM, Loftr, Quadtree, silk, but these models perform poorly despite being in the same resolution with SPSG. Some of them are one-stage models, which made them inefficient for the task. We tried to concatenate [800, 1500, 2000] resolutions calculated by SPSG; the result is nearly the same with 2000 alone.\n   - Suppress randomness by more iterations\n     Hloc uses geometry verification (`pycolmap.verify_matches`), set `max_num_trials=40000` didn't work.\n4. Some experience from the competition\n   Reading images using cv2 can lose EXIF information, which is not conducive to reconstruction. But there's a counter-example, \"theater\" can score even higher after the information removed.\n5. Referrence\n - Hloc\n  https://github.com/cvg/Hierarchical-Localization/\n - Superpoint introduction\n  https://github.com/magicleap/SuperPointPretrainedNetwork/blob/master/assets/DL4VSLAM_talk.pdf",
      "votes": null
    },
    {
      "id": "2325208",
      "postDate": "07/01/2023 07:01:44",
      "content": "<p>Congrats on coming 12th <a href=\"https://www.kaggle.com/niubiren\" target=\"_blank\">@niubiren</a> 🎉🎉🎉</p>",
      "rawMarkdown": "Congrats on coming 12th @niubiren 🎉🎉🎉",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2325208,
      "author_name": "swapnilchowdhury",
      "author_url": "",
      "post_date": "07/01/2023 07:01:44",
      "content": "<p>Congrats on coming 12th <a href=\"https://www.kaggle.com/niubiren\" target=\"_blank\">@niubiren</a> 🎉🎉🎉</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2324902": "Thanks to the competition organizers and Kaggle staff for hosting this amazing competition and solid support. Thanks to each my team member for their creative perspective and effort which gain us this place in competition.\n\n1. Overview\n Our final model is rather simple. The solution is based on a modular toolbox named Hierarchical-Localization (implementation of typical SP+SG structure), and the key modification is to set the input resolution to 2000.\n As the resolution increased from 1024 to 2000, the public LB scores increased from 0.24 to 0.46. As the resolution further increased to 3000, there is still improvement in the trainset (especially \"wall\").\n2. Configuration in detail \n\n    ```\n    confs = {\n        'superpoint_aachen': {\n            'output': 'feats-superpoint-n4096-r1024',\n            'model': {\n                'name': 'superpoint',\n                'nms_radius': 3,\n                'max_keypoints': 4096,\n            },\n            'preprocessing': {\n                'grayscale': True,\n                'resize_max': 2000, # Scaling longside to 2000\n            },\n        }\n    }\n    ```\n\n    ```\n    'superglue': {\n        'name': 'superglue',\n        'weights': 'outdoor',\n        'sinkhorn_iterations': 50,\n    }\n    ```\n\n\n3. Tricks didn't work\n   Our code suffers from randomness; the difference in public LB score for the same code can be up to 0.04! Unfortunately, during the whole competition, we failed to find a way to eliminate randomness. We decide a modification useless if no obvious improvement observed during repetitive submissions, but still, there are possibly mistakes.\n   - Image retrieval\n     After incremental mapping , try to add the images mapped failed to the successfully reconstructed model.\n   - TTA\n     Reverse and concatenate and apply NMS to features extracted from [original left-right-flip 10-deg-rotation], however, computation time hugely increased while no obvious improvement observed.\n   - Ensemble SIFT to SP\n     We learned that SP fails when the image is extremely in-plane rotated. So we tried to perform SIFT extraction and matching using pycolmap and extract the match from the database and concatenate with SPSG. This trick can improve the score for \"cyprus\" but is useless in this year's test set. Worth noting is that it doesn't harm the LB score either.\n   - Multi-models and multi-resolution\n     We tried to ensemble SPSG with DKM, Loftr, Quadtree, silk, but these models perform poorly despite being in the same resolution with SPSG. Some of them are one-stage models, which made them inefficient for the task. We tried to concatenate [800, 1500, 2000] resolutions calculated by SPSG; the result is nearly the same with 2000 alone.\n   - Suppress randomness by more iterations\n     Hloc uses geometry verification (`pycolmap.verify_matches`), set `max_num_trials=40000` didn't work.\n4. Some experience from the competition\n   Reading images using cv2 can lose EXIF information, which is not conducive to reconstruction. But there's a counter-example, \"theater\" can score even higher after the information removed.\n5. Referrence\n - Hloc\n  https://github.com/cvg/Hierarchical-Localization/\n - Superpoint introduction\n  https://github.com/magicleap/SuperPointPretrainedNetwork/blob/master/assets/DL4VSLAM_talk.pdf",
    "2325208": "Congrats on coming 12th @niubiren 🎉🎉🎉"
  },
  "source": "meta"
}