{
  "id": 416847,
  "title": "39th place solution - SuperGlue + SIFT",
  "url": "/competitions/image-matching-challenge-2023/discussion/416847",
  "author_name": "Gunes Evitan",
  "post_date": "2023-06-13T07:26:48.912000",
  "votes": 29,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>Overview</h2>\n<p>We used SuperGlue or SIFT on different scenes based on a heuristic and you can see our scores below.</p>\n<p>Notebook: <a href=\"https://www.kaggle.com/code/gunesevitan/image-matching-challenge-2023-inference\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/image-matching-challenge-2023-inference</a><br>\nCode: <a href=\"https://github.com/gunesevitan/image-matching-challenge-2023\" target=\"_blank\">https://github.com/gunesevitan/image-matching-challenge-2023</a></p>\n<h3>Scene Scores</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>mAA</th>\n<th>mAA Rotation</th>\n<th>mAA Translation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>bike</td>\n<td>0.9228</td>\n<td>0.9904</td>\n<td>0.9228</td>\n</tr>\n<tr>\n<td>chairs</td>\n<td>0.9775</td>\n<td>0.9916</td>\n<td>0.9775</td>\n</tr>\n<tr>\n<td>fountain</td>\n<td>1.0</td>\n<td>1.0</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>dioscuri</td>\n<td>0.5062</td>\n<td>0.5220</td>\n<td>0.5236</td>\n</tr>\n<tr>\n<td>cyprus</td>\n<td>0.6523</td>\n<td>0.7887</td>\n<td>0.6586</td>\n</tr>\n<tr>\n<td>wall</td>\n<td>0.8150</td>\n<td>0.9359</td>\n<td>0.8317</td>\n</tr>\n<tr>\n<td>kyiv-puppet-theater</td>\n<td>0.7704</td>\n<td>0.8756</td>\n<td>0.7895</td>\n</tr>\n</tbody>\n</table>\n<h3>Dataset Scores</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>mAA</th>\n<th>mAA Rotation</th>\n<th>mAA Translation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>haiper</td>\n<td>0.9667</td>\n<td>0.9994</td>\n<td>0.9667</td>\n</tr>\n<tr>\n<td>heritage</td>\n<td>0.6578</td>\n<td>0.7489</td>\n<td>0.6713</td>\n</tr>\n<tr>\n<td>urban</td>\n<td>0.7704</td>\n<td>0.8756</td>\n<td>0.7895</td>\n</tr>\n</tbody>\n</table>\n<h3>Global Scores</h3>\n<table>\n<thead>\n<tr>\n<th>mAA</th>\n<th>mAA Rotation</th>\n<th>mAA Translation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.7983</td>\n<td>0.8746</td>\n<td>0.8092</td>\n</tr>\n</tbody>\n</table>\n<h3>LB Scores</h3>\n<p>Public LB Score: <strong>0.415</strong><br>\nPrivate LB Score: <strong>0.465</strong></p>\n<h2>SuperPoint &amp; SuperGlue</h2>\n<p>SuperPoint and SuperGlue models are used with almost default parameters except <code>keypoint_threshold</code> is set to 0.01. We found that SuperGlue works better with raw sizes but some of the scenes had very large images that didn't fit into GPU memory. We resized images to 2560 (maximum longest edge that can be used safely on Kaggle) longest edge if any of the edges exceed that number. Otherwise, raw sizes are used.</p>\n<h2>SIFT</h2>\n<p>We initially started with COLMAP's SIFT implementation and it was working pretty good as a baseline. It was performing better on some scenes with very large images and strong rotations compared to deep models. There was a score trade-off between cyprus and wall while enabling <code>estimate_affine_shape</code> and <code>upright</code> and we ended up disabling both of them.</p>\n<pre><code>sift_extraction_options.max_image_size = \nsift_extraction_options.max_num_features = \nsift_extraction_options.estimate_affine_shape = \nsift_extraction_options.upright = \nsift_extraction_options.normalization = \n</code></pre>\n<h2>Model Selection</h2>\n<p>We noticed that large images with EXIF metadata has very high memory consumption and those are the images that have 90 degree rotations because of DSLR camera orientation. We add a simple if block that was checking the mean memory consumption of each scene. If that was greater than 16 megabytes, we used SIFT. Otherwise, we used SuperGlue on that scene.</p>\n<h2>Incremental Mapper</h2>\n<p>We used COLMAP's incremental mapper for reconstruction with almost default parameters except <code>min_model_size</code> is set to 3. Best reconstruction is selected based on registered image count and unregistered images are filled with scene mean rotation matrix and translation vector.</p>\n<h2>Thing that didn't work</h2>\n<ul>\n<li>OpenCV SIFT (COLMAP's SIFT implementation was working way better for some reason)</li>\n<li>DSP-SIFT (domain size pooling was boosting my validation score on local but it was throwing an error on Kaggle)</li>\n<li>LoFTR (too slow)</li>\n<li>KeyNet, AffNet, HardNet (bad score)</li>\n<li>DISK (bad score)</li>\n<li>SiLK (too slow)</li>\n<li>ASpanFormer (too slow)</li>\n<li>Rotation Correction (we probably didn't make it correctly within the limited timeframe)</li>\n<li>Two stage matching (it was boosting my validation score but didn't have enough time to fit it into pipeline)</li>\n</ul>",
  "messages": [
    {
      "id": 2300404,
      "postDate": "2023-06-13T07:26:48.913Z",
      "content": "<h2>Overview</h2>\n<p>We used SuperGlue or SIFT on different scenes based on a heuristic and you can see our scores below.</p>\n<p>Notebook: <a href=\"https://www.kaggle.com/code/gunesevitan/image-matching-challenge-2023-inference\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/image-matching-challenge-2023-inference</a><br>\nCode: <a href=\"https://github.com/gunesevitan/image-matching-challenge-2023\" target=\"_blank\">https://github.com/gunesevitan/image-matching-challenge-2023</a></p>\n<h3>Scene Scores</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>mAA</th>\n<th>mAA Rotation</th>\n<th>mAA Translation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>bike</td>\n<td>0.9228</td>\n<td>0.9904</td>\n<td>0.9228</td>\n</tr>\n<tr>\n<td>chairs</td>\n<td>0.9775</td>\n<td>0.9916</td>\n<td>0.9775</td>\n</tr>\n<tr>\n<td>fountain</td>\n<td>1.0</td>\n<td>1.0</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>dioscuri</td>\n<td>0.5062</td>\n<td>0.5220</td>\n<td>0.5236</td>\n</tr>\n<tr>\n<td>cyprus</td>\n<td>0.6523</td>\n<td>0.7887</td>\n<td>0.6586</td>\n</tr>\n<tr>\n<td>wall</td>\n<td>0.8150</td>\n<td>0.9359</td>\n<td>0.8317</td>\n</tr>\n<tr>\n<td>kyiv-puppet-theater</td>\n<td>0.7704</td>\n<td>0.8756</td>\n<td>0.7895</td>\n</tr>\n</tbody>\n</table>\n<h3>Dataset Scores</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>mAA</th>\n<th>mAA Rotation</th>\n<th>mAA Translation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>haiper</td>\n<td>0.9667</td>\n<td>0.9994</td>\n<td>0.9667</td>\n</tr>\n<tr>\n<td>heritage</td>\n<td>0.6578</td>\n<td>0.7489</td>\n<td>0.6713</td>\n</tr>\n<tr>\n<td>urban</td>\n<td>0.7704</td>\n<td>0.8756</td>\n<td>0.7895</td>\n</tr>\n</tbody>\n</table>\n<h3>Global Scores</h3>\n<table>\n<thead>\n<tr>\n<th>mAA</th>\n<th>mAA Rotation</th>\n<th>mAA Translation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.7983</td>\n<td>0.8746</td>\n<td>0.8092</td>\n</tr>\n</tbody>\n</table>\n<h3>LB Scores</h3>\n<p>Public LB Score: <strong>0.415</strong><br>\nPrivate LB Score: <strong>0.465</strong></p>\n<h2>SuperPoint &amp; SuperGlue</h2>\n<p>SuperPoint and SuperGlue models are used with almost default parameters except <code>keypoint_threshold</code> is set to 0.01. We found that SuperGlue works better with raw sizes but some of the scenes had very large images that didn't fit into GPU memory. We resized images to 2560 (maximum longest edge that can be used safely on Kaggle) longest edge if any of the edges exceed that number. Otherwise, raw sizes are used.</p>\n<h2>SIFT</h2>\n<p>We initially started with COLMAP's SIFT implementation and it was working pretty good as a baseline. It was performing better on some scenes with very large images and strong rotations compared to deep models. There was a score trade-off between cyprus and wall while enabling <code>estimate_affine_shape</code> and <code>upright</code> and we ended up disabling both of them.</p>\n<pre><code>sift_extraction_options.max_image_size = \nsift_extraction_options.max_num_features = \nsift_extraction_options.estimate_affine_shape = \nsift_extraction_options.upright = \nsift_extraction_options.normalization = \n</code></pre>\n<h2>Model Selection</h2>\n<p>We noticed that large images with EXIF metadata has very high memory consumption and those are the images that have 90 degree rotations because of DSLR camera orientation. We add a simple if block that was checking the mean memory consumption of each scene. If that was greater than 16 megabytes, we used SIFT. Otherwise, we used SuperGlue on that scene.</p>\n<h2>Incremental Mapper</h2>\n<p>We used COLMAP's incremental mapper for reconstruction with almost default parameters except <code>min_model_size</code> is set to 3. Best reconstruction is selected based on registered image count and unregistered images are filled with scene mean rotation matrix and translation vector.</p>\n<h2>Thing that didn't work</h2>\n<ul>\n<li>OpenCV SIFT (COLMAP's SIFT implementation was working way better for some reason)</li>\n<li>DSP-SIFT (domain size pooling was boosting my validation score on local but it was throwing an error on Kaggle)</li>\n<li>LoFTR (too slow)</li>\n<li>KeyNet, AffNet, HardNet (bad score)</li>\n<li>DISK (bad score)</li>\n<li>SiLK (too slow)</li>\n<li>ASpanFormer (too slow)</li>\n<li>Rotation Correction (we probably didn't make it correctly within the limited timeframe)</li>\n<li>Two stage matching (it was boosting my validation score but didn't have enough time to fit it into pipeline)</li>\n</ul>",
      "rawMarkdown": "## Overview\n\nWe used SuperGlue or SIFT on different scenes based on a heuristic and you can see our scores below.\n\nNotebook: https://www.kaggle.com/code/gunesevitan/image-matching-challenge-2023-inference\nCode: https://github.com/gunesevitan/image-matching-challenge-2023\n\n### Scene Scores\n\n|                     | mAA    | mAA Rotation | mAA Translation |\n|---------------------|--------|--------------|-----------------|\n| bike                | 0.9228 | 0.9904       | 0.9228          |\n| chairs              | 0.9775 | 0.9916       | 0.9775          |\n| fountain            | 1.0    | 1.0          | 1.0             |\n| dioscuri            | 0.5062 | 0.5220       | 0.5236          |\n| cyprus              | 0.6523 | 0.7887       | 0.6586          |\n| wall                | 0.8150 | 0.9359       | 0.8317          |\n| kyiv-puppet-theater | 0.7704 | 0.8756       | 0.7895          |\n\n### Dataset Scores\n\n|          | mAA    | mAA Rotation | mAA Translation |\n|----------|--------|--------------|-----------------|\n| haiper   | 0.9667 | 0.9994       | 0.9667          |\n| heritage | 0.6578 | 0.7489       | 0.6713          |\n| urban    | 0.7704 | 0.8756       | 0.7895          |\n\n### Global Scores\n\n| mAA    | mAA Rotation | mAA Translation |\n|--------|--------------|-----------------|\n| 0.7983 | 0.8746       | 0.8092          |\n\n### LB Scores\n\nPublic LB Score: **0.415**\nPrivate LB Score: **0.465**\n\n## SuperPoint & SuperGlue\n\nSuperPoint and SuperGlue models are used with almost default parameters except `keypoint_threshold` is set to 0.01. We found that SuperGlue works better with raw sizes but some of the scenes had very large images that didn't fit into GPU memory. We resized images to 2560 (maximum longest edge that can be used safely on Kaggle) longest edge if any of the edges exceed that number. Otherwise, raw sizes are used.\n\n## SIFT\n\nWe initially started with COLMAP's SIFT implementation and it was working pretty good as a baseline. It was performing better on some scenes with very large images and strong rotations compared to deep models. There was a score trade-off between cyprus and wall while enabling `estimate_affine_shape` and `upright` and we ended up disabling both of them.\n\n```python\nsift_extraction_options.max_image_size = 1400\nsift_extraction_options.max_num_features = 8192\nsift_extraction_options.estimate_affine_shape = False\nsift_extraction_options.upright = False\nsift_extraction_options.normalization = 'L2'\n```\n\n## Model Selection\n\nWe noticed that large images with EXIF metadata has very high memory consumption and those are the images that have 90 degree rotations because of DSLR camera orientation. We add a simple if block that was checking the mean memory consumption of each scene. If that was greater than 16 megabytes, we used SIFT. Otherwise, we used SuperGlue on that scene.\n\n## Incremental Mapper\n\nWe used COLMAP's incremental mapper for reconstruction with almost default parameters except `min_model_size` is set to 3. Best reconstruction is selected based on registered image count and unregistered images are filled with scene mean rotation matrix and translation vector.\n\n## Thing that didn't work\n* OpenCV SIFT (COLMAP's SIFT implementation was working way better for some reason)\n* DSP-SIFT (domain size pooling was boosting my validation score on local but it was throwing an error on Kaggle)\n* LoFTR (too slow)\n* KeyNet, AffNet, HardNet (bad score)\n* DISK (bad score)\n* SiLK (too slow)\n* ASpanFormer (too slow)\n* Rotation Correction (we probably didn't make it correctly within the limited timeframe)\n* Two stage matching (it was boosting my validation score but didn't have enough time to fit it into pipeline)\n\n",
      "votes": 29
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2300404": "## Overview\n\nWe used SuperGlue or SIFT on different scenes based on a heuristic and you can see our scores below.\n\nNotebook: https://www.kaggle.com/code/gunesevitan/image-matching-challenge-2023-inference\nCode: https://github.com/gunesevitan/image-matching-challenge-2023\n\n### Scene Scores\n\n|                     | mAA    | mAA Rotation | mAA Translation |\n|---------------------|--------|--------------|-----------------|\n| bike                | 0.9228 | 0.9904       | 0.9228          |\n| chairs              | 0.9775 | 0.9916       | 0.9775          |\n| fountain            | 1.0    | 1.0          | 1.0             |\n| dioscuri            | 0.5062 | 0.5220       | 0.5236          |\n| cyprus              | 0.6523 | 0.7887       | 0.6586          |\n| wall                | 0.8150 | 0.9359       | 0.8317          |\n| kyiv-puppet-theater | 0.7704 | 0.8756       | 0.7895          |\n\n### Dataset Scores\n\n|          | mAA    | mAA Rotation | mAA Translation |\n|----------|--------|--------------|-----------------|\n| haiper   | 0.9667 | 0.9994       | 0.9667          |\n| heritage | 0.6578 | 0.7489       | 0.6713          |\n| urban    | 0.7704 | 0.8756       | 0.7895          |\n\n### Global Scores\n\n| mAA    | mAA Rotation | mAA Translation |\n|--------|--------------|-----------------|\n| 0.7983 | 0.8746       | 0.8092          |\n\n### LB Scores\n\nPublic LB Score: **0.415**\nPrivate LB Score: **0.465**\n\n## SuperPoint & SuperGlue\n\nSuperPoint and SuperGlue models are used with almost default parameters except `keypoint_threshold` is set to 0.01. We found that SuperGlue works better with raw sizes but some of the scenes had very large images that didn't fit into GPU memory. We resized images to 2560 (maximum longest edge that can be used safely on Kaggle) longest edge if any of the edges exceed that number. Otherwise, raw sizes are used.\n\n## SIFT\n\nWe initially started with COLMAP's SIFT implementation and it was working pretty good as a baseline. It was performing better on some scenes with very large images and strong rotations compared to deep models. There was a score trade-off between cyprus and wall while enabling `estimate_affine_shape` and `upright` and we ended up disabling both of them.\n\n```python\nsift_extraction_options.max_image_size = 1400\nsift_extraction_options.max_num_features = 8192\nsift_extraction_options.estimate_affine_shape = False\nsift_extraction_options.upright = False\nsift_extraction_options.normalization = 'L2'\n```\n\n## Model Selection\n\nWe noticed that large images with EXIF metadata has very high memory consumption and those are the images that have 90 degree rotations because of DSLR camera orientation. We add a simple if block that was checking the mean memory consumption of each scene. If that was greater than 16 megabytes, we used SIFT. Otherwise, we used SuperGlue on that scene.\n\n## Incremental Mapper\n\nWe used COLMAP's incremental mapper for reconstruction with almost default parameters except `min_model_size` is set to 3. Best reconstruction is selected based on registered image count and unregistered images are filled with scene mean rotation matrix and translation vector.\n\n## Thing that didn't work\n* OpenCV SIFT (COLMAP's SIFT implementation was working way better for some reason)\n* DSP-SIFT (domain size pooling was boosting my validation score on local but it was throwing an error on Kaggle)\n* LoFTR (too slow)\n* KeyNet, AffNet, HardNet (bad score)\n* DISK (bad score)\n* SiLK (too slow)\n* ASpanFormer (too slow)\n* Rotation Correction (we probably didn't make it correctly within the limited timeframe)\n* Two stage matching (it was boosting my validation score but didn't have enough time to fit it into pipeline)\n\n"
  }
}