{
  "id": 416842,
  "title": "9th Solution",
  "url": "/competitions/image-matching-challenge-2023/writeups/mts-ai-itmo-university-9th-solution",
  "author_name": "",
  "post_date": "2023-06-21T00:46:18.287Z",
  "votes": 29,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Great thanks to the organizers and Kaggle staff for this amazing competition.<br>\nOur solution shares architectural similarities with the baseline provided by the organizers. <br>\nFirstly, we addressed the issue of rotation in-variance by implementing a rotation model to standardize the orientation of input images. Next, we employed a neural network for the retrieval task, enabling us to extract matching pairs for the feature extraction and matching process. This process generated a database which served as the input for the incremental mapping in Colmap.</p>\n<h2>Orientation Model</h2>\n<p>We would like to express our gratitude to <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> for developing this great <a href=\"https://github.com/ternaus/check_orientation\">model</a> , which was trained on a large dataset and exhibited good performance in both (LB) and (CV) evaluations. We made a slight modification to this model, which we refer to as self-ensembling.<br>\nIn our proposed adjustment, we utilized the model iteratively on the query image and its different rotated versions. Subsequently, we tuned the threshold by conducting validation tests on more than 1500 images with different rotation angles. Through experimentation, we discovered that using a threshold between 0.8 and 0.87 might result in incorrect orientation predictions. To address this issue, we applied self-ensembling by further rotating the image and checking if the predicted class probability fell within this range of thresholds. We repeated this last step twice to ensure accuracy.</p>\n<h2>Retrieval Method</h2>\n<h3>Retrieval Model</h3>\n<p>In order to address the challenges posed by the cost inefficiency of exhaustive matching and the limitations of sequential matching for 3D reconstruction, we sought alternative methods for image retrieval. <br>\nAfter careful evaluation, we chose to utilize NetVlad as our chosen method due to its superior performance compared to openibl and cosplace. To further enhance the results, we employed various techniques, including:<br>\n1) We passed the original image and its horizontally flipped version through the model. By summing the descriptors obtained from both passes, we achieved improved performance, particularly for highly similar images. This technique is especially effective in cases where the scene exhibits symmetry. By redefining a new point in the n-dimensional descriptive space, which effectively increase the similarity distance between two distinct parts of the scene.</p>\n<p>2) Re-ranking: After calculating the similarity scores, we performed re-ranking by selecting the top 1 match. We then re-queried the retrieval process using these two images instead of just one. The similarity scores of the resulting matches were summed together after raising them to a specific power \"m\". This manipulation of probabilities ensures that if the best match for one of the query images is found, it will be favored over an image that is similar to both where the sum of the similarities on both scenarios is equal.</p>\n<p>We repeated this procedure twice using NetVlad with Test Time Augmentation (TTA) on the image size. The results were nearly perfect, and the approach even enabled the correct ordering of The Wall scene pairs, which can be best matched by doing sequential matching.</p>\n<h3>Number of Matches</h3>\n<p>Determining the number of image pairs to select is a critical factor that directly affects both validation and leaderboard performance. This becomes particularly important for large scenes where there may be no or very few common images between subsets of the scene.</p>\n<p>To address this challenge, we devised a strategy based on a complete graph representation. In this graph, each edge represents the similarity between two images. The goal is to choose a subset of edges where each image has an equal number of connected nodes.</p>\n<p>We employed a Binary Search approach to determine the number of matches for each image, with a check function to verify if the resulting graph is connected or not. The lower bound of the binary search was set to half the number of images, ensuring that we consider common matches and prevent incomplete 3D model reconstruction. Additionally, we made sure that the approach remains exhaustive for small scenes containing less than 40 images.</p>\n<p>By employing this method, we aimed to strike a balance between capturing sufficient matching pairs for accurate 3D reconstruction while avoiding redundant or disconnected image subsets.</p>\n<h2>Feature Extraction and Matching</h2>\n<p>In our selected submissions, we have utilized the SuperPoint and SuperGlue algorithms followed by MagSac filtering. Unfortunately SuperGlue is not licensed for commercial use. However, we have also achieved highly promising results using GlueStick. We have made modifications to the GlueStick architecture to integrate it with SuperPoint, and we achieved a score of approximately 0.450 on the public leaderboard and 0.513 on private leaderboard without employing our best tuning parameters. It is worth noting that this modified architecture is permitted for commercial use and offers improved processing speed. We anticipate that further tuning can yield even better results with GlueStick, but didn't choose it as our last submissions.</p>\n<h2>Refinement</h2>\n<p>Although not included in our selected submissions, we would like to mention an approach that significantly improved validation results across various scenes. We employed Pixel-Perfect-SFM in conjunction with sd2net for dense matching, but it didn't improved the results on public leaderboard.</p>\n<h2>Registering Unregistered Images</h2>\n<p>While not part of our selected submissions, we made attempts to register unregistered images using various techniques. However, these attempts did not yield significant improvements on the leaderboard. We explored the following strategies:</p>\n<ul>\n<li>Utilizing different orientations and attempting registration.</li>\n<li>Applying different extractor-matchers, such as LoFTR and R2D2, for the unregistered images.</li>\n<li>Adjusting parameters for SuperPoint and SuperGlue to optimize the registration process.</li>\n</ul>\n<h2>Tried but not worked</h2>\n<ul>\n<li>Pixel-Perfect-SFM</li>\n<li>Semantic Segmentation masks</li>\n<li>Illumenation enhancement</li>\n<li>PnP localization for unregistered Images</li>\n<li>LoFTR/ SILK / DAC / Wireframe</li>\n<li>CosPlace / OpenIBL</li>\n<li>Large Image splitting into Parts</li>\n<li>Grid-based Point Sampling (equally spatial replacement of points in an image).</li>\n<li>Rotation self-ensemble averaging</li>\n<li>Filtering the extracted pairs using a threshold from a binarysearch function or a fixed threshold.</li>\n</ul>\n<h2>Important Notes:</h2>\n<p>1) Most participants were using colmap reconstruction, which is nondeterministic. We were able to modify the code so that it became deterministic but working only one CPU thread, which helped us observe our improvement and avoid randomness.<br>\n2) We found out that OpenCV and PIL libs use EXIF Information when reading images, unless providing flags to prevent it. By that, we mean that if the orientation in EXIF will rotate the image automatically before processing. This was confusing for us, as there are missing information about how the GT were collected for rotation part (with or without considering them), that's why one of  our chosen solutions included this correction to overcome such issue, plus we lost a lot of submission to check the effect of this issue on leaderboard. It would have been more helpful if there was better explanation about how the GT calculated.<br>\n3) Our validation scored 0.485 on public and 0.514 on private, with local score of 0.87.<br>\n4) The fact that validation and leaderboard were not correlated made things more difficult and random, as it was clear there will be a shake up due to the fact that for some specific scene some solutions might fail which will drastically impact the mAA score.</p>\n<h2>Acknowldgment</h2>\n<ul>\n<li>I wanted to take a moment to express my sincere appreciation for the exceptional contribution and dedication to my friend and teammate <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> . His hard work and commitment have played a crucial role in our collective success, and I am truly grateful for the opportunity to work alongside such a remarkable individual. we had fun and we learned a lot.</li>\n<li>Special thanks to the <a href=\"https://github.com/cvg\" target=\"_blank\">Computer Vision and Geometry Lab</a> Group for (hloc, gluestick, pixel perfect sfm ..)</li>\n</ul>\n<h2> </h2>",
  "messages": [
    {
      "id": "2300383",
      "postDate": "06/13/2023 07:15:50",
      "content": "<p>Great thanks to the organizers and Kaggle staff for this amazing competition.<br>\nOur solution shares architectural similarities with the baseline provided by the organizers. <br>\nFirstly, we addressed the issue of rotation in-variance by implementing a rotation model to standardize the orientation of input images. Next, we employed a neural network for the retrieval task, enabling us to extract matching pairs for the feature extraction and matching process. This process generated a database which served as the input for the incremental mapping in Colmap.</p>\n<h2>Orientation Model</h2>\n<p>We would like to express our gratitude to <a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> for developing this great <a href=\"https://github.com/ternaus/check_orientation\">model</a> , which was trained on a large dataset and exhibited good performance in both (LB) and (CV) evaluations. We made a slight modification to this model, which we refer to as self-ensembling.<br>\nIn our proposed adjustment, we utilized the model iteratively on the query image and its different rotated versions. Subsequently, we tuned the threshold by conducting validation tests on more than 1500 images with different rotation angles. Through experimentation, we discovered that using a threshold between 0.8 and 0.87 might result in incorrect orientation predictions. To address this issue, we applied self-ensembling by further rotating the image and checking if the predicted class probability fell within this range of thresholds. We repeated this last step twice to ensure accuracy.</p>\n<h2>Retrieval Method</h2>\n<h3>Retrieval Model</h3>\n<p>In order to address the challenges posed by the cost inefficiency of exhaustive matching and the limitations of sequential matching for 3D reconstruction, we sought alternative methods for image retrieval. <br>\nAfter careful evaluation, we chose to utilize NetVlad as our chosen method due to its superior performance compared to openibl and cosplace. To further enhance the results, we employed various techniques, including:<br>\n1) We passed the original image and its horizontally flipped version through the model. By summing the descriptors obtained from both passes, we achieved improved performance, particularly for highly similar images. This technique is especially effective in cases where the scene exhibits symmetry. By redefining a new point in the n-dimensional descriptive space, which effectively increase the similarity distance between two distinct parts of the scene.</p>\n<p>2) Re-ranking: After calculating the similarity scores, we performed re-ranking by selecting the top 1 match. We then re-queried the retrieval process using these two images instead of just one. The similarity scores of the resulting matches were summed together after raising them to a specific power \"m\". This manipulation of probabilities ensures that if the best match for one of the query images is found, it will be favored over an image that is similar to both where the sum of the similarities on both scenarios is equal.</p>\n<p>We repeated this procedure twice using NetVlad with Test Time Augmentation (TTA) on the image size. The results were nearly perfect, and the approach even enabled the correct ordering of The Wall scene pairs, which can be best matched by doing sequential matching.</p>\n<h3>Number of Matches</h3>\n<p>Determining the number of image pairs to select is a critical factor that directly affects both validation and leaderboard performance. This becomes particularly important for large scenes where there may be no or very few common images between subsets of the scene.</p>\n<p>To address this challenge, we devised a strategy based on a complete graph representation. In this graph, each edge represents the similarity between two images. The goal is to choose a subset of edges where each image has an equal number of connected nodes.</p>\n<p>We employed a Binary Search approach to determine the number of matches for each image, with a check function to verify if the resulting graph is connected or not. The lower bound of the binary search was set to half the number of images, ensuring that we consider common matches and prevent incomplete 3D model reconstruction. Additionally, we made sure that the approach remains exhaustive for small scenes containing less than 40 images.</p>\n<p>By employing this method, we aimed to strike a balance between capturing sufficient matching pairs for accurate 3D reconstruction while avoiding redundant or disconnected image subsets.</p>\n<h2>Feature Extraction and Matching</h2>\n<p>In our selected submissions, we have utilized the SuperPoint and SuperGlue algorithms followed by MagSac filtering. Unfortunately SuperGlue is not licensed for commercial use. However, we have also achieved highly promising results using GlueStick. We have made modifications to the GlueStick architecture to integrate it with SuperPoint, and we achieved a score of approximately 0.450 on the public leaderboard and 0.513 on private leaderboard without employing our best tuning parameters. It is worth noting that this modified architecture is permitted for commercial use and offers improved processing speed. We anticipate that further tuning can yield even better results with GlueStick, but didn't choose it as our last submissions.</p>\n<h2>Refinement</h2>\n<p>Although not included in our selected submissions, we would like to mention an approach that significantly improved validation results across various scenes. We employed Pixel-Perfect-SFM in conjunction with sd2net for dense matching, but it didn't improved the results on public leaderboard.</p>\n<h2>Registering Unregistered Images</h2>\n<p>While not part of our selected submissions, we made attempts to register unregistered images using various techniques. However, these attempts did not yield significant improvements on the leaderboard. We explored the following strategies:</p>\n<ul>\n<li>Utilizing different orientations and attempting registration.</li>\n<li>Applying different extractor-matchers, such as LoFTR and R2D2, for the unregistered images.</li>\n<li>Adjusting parameters for SuperPoint and SuperGlue to optimize the registration process.</li>\n</ul>\n<h2>Tried but not worked</h2>\n<ul>\n<li>Pixel-Perfect-SFM</li>\n<li>Semantic Segmentation masks</li>\n<li>Illumenation enhancement</li>\n<li>PnP localization for unregistered Images</li>\n<li>LoFTR/ SILK / DAC / Wireframe</li>\n<li>CosPlace / OpenIBL</li>\n<li>Large Image splitting into Parts</li>\n<li>Grid-based Point Sampling (equally spatial replacement of points in an image).</li>\n<li>Rotation self-ensemble averaging</li>\n<li>Filtering the extracted pairs using a threshold from a binarysearch function or a fixed threshold.</li>\n</ul>\n<h2>Important Notes:</h2>\n<p>1) Most participants were using colmap reconstruction, which is nondeterministic. We were able to modify the code so that it became deterministic but working only one CPU thread, which helped us observe our improvement and avoid randomness.<br>\n2) We found out that OpenCV and PIL libs use EXIF Information when reading images, unless providing flags to prevent it. By that, we mean that if the orientation in EXIF will rotate the image automatically before processing. This was confusing for us, as there are missing information about how the GT were collected for rotation part (with or without considering them), that's why one of  our chosen solutions included this correction to overcome such issue, plus we lost a lot of submission to check the effect of this issue on leaderboard. It would have been more helpful if there was better explanation about how the GT calculated.<br>\n3) Our validation scored 0.485 on public and 0.514 on private, with local score of 0.87.<br>\n4) The fact that validation and leaderboard were not correlated made things more difficult and random, as it was clear there will be a shake up due to the fact that for some specific scene some solutions might fail which will drastically impact the mAA score.</p>\n<h2>Acknowldgment</h2>\n<ul>\n<li>I wanted to take a moment to express my sincere appreciation for the exceptional contribution and dedication to my friend and teammate <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> . His hard work and commitment have played a crucial role in our collective success, and I am truly grateful for the opportunity to work alongside such a remarkable individual. we had fun and we learned a lot.</li>\n<li>Special thanks to the <a href=\"https://github.com/cvg\" target=\"_blank\">Computer Vision and Geometry Lab</a> Group for (hloc, gluestick, pixel perfect sfm ..)</li>\n</ul>\n<h2> </h2>",
      "rawMarkdown": "Great thanks to the organizers and Kaggle staff for this amazing competition.\nOur solution shares architectural similarities with the baseline provided by the organizers. \nFirstly, we addressed the issue of rotation in-variance by implementing a rotation model to standardize the orientation of input images. Next, we employed a neural network for the retrieval task, enabling us to extract matching pairs for the feature extraction and matching process. This process generated a database which served as the input for the incremental mapping in Colmap.\n## Orientation Model\nWe would like to express our gratitude to @iglovikov for developing this great <a href=\"https://github.com/ternaus/check_orientation\">model</a> , which was trained on a large dataset and exhibited good performance in both (LB) and (CV) evaluations. We made a slight modification to this model, which we refer to as self-ensembling.\nIn our proposed adjustment, we utilized the model iteratively on the query image and its different rotated versions. Subsequently, we tuned the threshold by conducting validation tests on more than 1500 images with different rotation angles. Through experimentation, we discovered that using a threshold between 0.8 and 0.87 might result in incorrect orientation predictions. To address this issue, we applied self-ensembling by further rotating the image and checking if the predicted class probability fell within this range of thresholds. We repeated this last step twice to ensure accuracy.\n## Retrieval Method\n### Retrieval Model\nIn order to address the challenges posed by the cost inefficiency of exhaustive matching and the limitations of sequential matching for 3D reconstruction, we sought alternative methods for image retrieval. \nAfter careful evaluation, we chose to utilize NetVlad as our chosen method due to its superior performance compared to openibl and cosplace. To further enhance the results, we employed various techniques, including:\n1) We passed the original image and its horizontally flipped version through the model. By summing the descriptors obtained from both passes, we achieved improved performance, particularly for highly similar images. This technique is especially effective in cases where the scene exhibits symmetry. By redefining a new point in the n-dimensional descriptive space, which effectively increase the similarity distance between two distinct parts of the scene.\n\n2) Re-ranking: After calculating the similarity scores, we performed re-ranking by selecting the top 1 match. We then re-queried the retrieval process using these two images instead of just one. The similarity scores of the resulting matches were summed together after raising them to a specific power \"m\". This manipulation of probabilities ensures that if the best match for one of the query images is found, it will be favored over an image that is similar to both where the sum of the similarities on both scenarios is equal.\n\nWe repeated this procedure twice using NetVlad with Test Time Augmentation (TTA) on the image size. The results were nearly perfect, and the approach even enabled the correct ordering of The Wall scene pairs, which can be best matched by doing sequential matching.\n### Number of Matches\nDetermining the number of image pairs to select is a critical factor that directly affects both validation and leaderboard performance. This becomes particularly important for large scenes where there may be no or very few common images between subsets of the scene.\n\nTo address this challenge, we devised a strategy based on a complete graph representation. In this graph, each edge represents the similarity between two images. The goal is to choose a subset of edges where each image has an equal number of connected nodes.\n\nWe employed a Binary Search approach to determine the number of matches for each image, with a check function to verify if the resulting graph is connected or not. The lower bound of the binary search was set to half the number of images, ensuring that we consider common matches and prevent incomplete 3D model reconstruction. Additionally, we made sure that the approach remains exhaustive for small scenes containing less than 40 images.\n\nBy employing this method, we aimed to strike a balance between capturing sufficient matching pairs for accurate 3D reconstruction while avoiding redundant or disconnected image subsets.\n## Feature Extraction and Matching\nIn our selected submissions, we have utilized the SuperPoint and SuperGlue algorithms followed by MagSac filtering. Unfortunately SuperGlue is not licensed for commercial use. However, we have also achieved highly promising results using GlueStick. We have made modifications to the GlueStick architecture to integrate it with SuperPoint, and we achieved a score of approximately 0.450 on the public leaderboard and 0.513 on private leaderboard without employing our best tuning parameters. It is worth noting that this modified architecture is permitted for commercial use and offers improved processing speed. We anticipate that further tuning can yield even better results with GlueStick, but didn't choose it as our last submissions.\n## Refinement\nAlthough not included in our selected submissions, we would like to mention an approach that significantly improved validation results across various scenes. We employed Pixel-Perfect-SFM in conjunction with sd2net for dense matching, but it didn't improved the results on public leaderboard.\n\n## Registering Unregistered Images\nWhile not part of our selected submissions, we made attempts to register unregistered images using various techniques. However, these attempts did not yield significant improvements on the leaderboard. We explored the following strategies:\n- Utilizing different orientations and attempting registration.\n- Applying different extractor-matchers, such as LoFTR and R2D2, for the unregistered images.\n- Adjusting parameters for SuperPoint and SuperGlue to optimize the registration process.\n## Tried but not worked\n- Pixel-Perfect-SFM\n- Semantic Segmentation masks\n- Illumenation enhancement\n- PnP localization for unregistered Images\n- LoFTR/ SILK / DAC / Wireframe\n- CosPlace / OpenIBL\n- Large Image splitting into Parts\n- Grid-based Point Sampling (equally spatial replacement of points in an image).\n- Rotation self-ensemble averaging\n- Filtering the extracted pairs using a threshold from a binarysearch function or a fixed threshold.\n## Important Notes:\n1) Most participants were using colmap reconstruction, which is nondeterministic. We were able to modify the code so that it became deterministic but working only one CPU thread, which helped us observe our improvement and avoid randomness.\n2) We found out that OpenCV and PIL libs use EXIF Information when reading images, unless providing flags to prevent it. By that, we mean that if the orientation in EXIF will rotate the image automatically before processing. This was confusing for us, as there are missing information about how the GT were collected for rotation part (with or without considering them), that's why one of  our chosen solutions included this correction to overcome such issue, plus we lost a lot of submission to check the effect of this issue on leaderboard. It would have been more helpful if there was better explanation about how the GT calculated.\n3) Our validation scored 0.485 on public and 0.514 on private, with local score of 0.87.\n4) The fact that validation and leaderboard were not correlated made things more difficult and random, as it was clear there will be a shake up due to the fact that for some specific scene some solutions might fail which will drastically impact the mAA score.\n## Acknowldgment\n- I wanted to take a moment to express my sincere appreciation for the exceptional contribution and dedication to my friend and teammate @ammarali32 . His hard work and commitment have played a crucial role in our collective success, and I am truly grateful for the opportunity to work alongside such a remarkable individual. we had fun and we learned a lot.\n- Special thanks to the [Computer Vision and Geometry Lab](https://github.com/cvg) Group for (hloc, gluestick, pixel perfect sfm ..)\n##",
      "votes": null
    },
    {
      "id": "2300496",
      "postDate": "06/13/2023 08:18:03",
      "content": "<p>Guys, great job and congratulations with gold and Kaggle Master level! We followed each other last weeks :) I'm impressed of your interesting approach for determining N of matches. Thanks for detailed write-up.</p>",
      "rawMarkdown": "Guys, great job and congratulations with gold and Kaggle Master level! We followed each other last weeks :) I'm impressed of your interesting approach for determining N of matches. Thanks for detailed write-up.",
      "votes": null
    },
    {
      "id": "2300502",
      "postDate": "06/13/2023 08:19:51",
      "content": "<p>Thanks! Great work as well! Congrats on the 2nd place🎉! </p>",
      "rawMarkdown": "Thanks! Great work as well! Congrats on the 2nd place🎉!",
      "votes": null
    },
    {
      "id": "2300519",
      "postDate": "06/13/2023 08:32:06",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/jaafarmahmoud1\" target=\"_blank\">@jaafarmahmoud1</a> and <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a>. Will you share the COLMAP modification that made it deterministic?</p>",
      "rawMarkdown": "Congrats @jaafarmahmoud1 and @ammarali32. Will you share the COLMAP modification that made it deterministic?",
      "votes": null
    },
    {
      "id": "2300544",
      "postDate": "06/13/2023 08:46:51",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>! Yes, you can check this wheel for colmap 0.4.0. Basically, changing everything inside to one thread, and skipping <a href=\"https://github.com/colmap/pycolmap/blob/743a4ac305183f96d2a4cfce7c7f6418b31b8598/pipeline/match_features.cc#L76\" target=\"_blank\">geometric verification </a> made things deterministic. because this process is somehow random, then we replaced it with MagSac and provided a specific seed as USAC parameters. <br>\n<a href=\"https://www.kaggle.com/datasets/jaafarmahmoud1/pycolmap-040\" target=\"_blank\">https://www.kaggle.com/datasets/jaafarmahmoud1/pycolmap-040</a></p>",
      "rawMarkdown": "Thanks @gunesevitan! Yes, you can check this wheel for colmap 0.4.0. Basically, changing everything inside to one thread, and skipping [geometric verification ](https://github.com/colmap/pycolmap/blob/743a4ac305183f96d2a4cfce7c7f6418b31b8598/pipeline/match_features.cc#L76) made things deterministic. because this process is somehow random, then we replaced it with MagSac and provided a specific seed as USAC parameters. \nhttps://www.kaggle.com/datasets/jaafarmahmoud1/pycolmap-040",
      "votes": null
    },
    {
      "id": "2301061",
      "postDate": "06/13/2023 15:42:09",
      "content": "<p>Thanks! you have done an amazing work congrats on the second place !! </p>",
      "rawMarkdown": "Thanks! you have done an amazing work congrats on the second place !!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2300496,
      "author_name": "igorlashkov",
      "author_url": "",
      "post_date": "06/13/2023 08:18:03",
      "content": "<p>Guys, great job and congratulations with gold and Kaggle Master level! We followed each other last weeks :) I'm impressed of your interesting approach for determining N of matches. Thanks for detailed write-up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2300502,
          "author_name": "jaafarmahmoud1",
          "author_url": "",
          "post_date": "06/13/2023 08:19:51",
          "content": "<p>Thanks! Great work as well! Congrats on the 2nd place🎉! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2301061,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "06/13/2023 15:42:09",
          "content": "<p>Thanks! you have done an amazing work congrats on the second place !! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2300519,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "06/13/2023 08:32:06",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/jaafarmahmoud1\" target=\"_blank\">@jaafarmahmoud1</a> and <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a>. Will you share the COLMAP modification that made it deterministic?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2300544,
          "author_name": "jaafarmahmoud1",
          "author_url": "",
          "post_date": "06/13/2023 08:46:51",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>! Yes, you can check this wheel for colmap 0.4.0. Basically, changing everything inside to one thread, and skipping <a href=\"https://github.com/colmap/pycolmap/blob/743a4ac305183f96d2a4cfce7c7f6418b31b8598/pipeline/match_features.cc#L76\" target=\"_blank\">geometric verification </a> made things deterministic. because this process is somehow random, then we replaced it with MagSac and provided a specific seed as USAC parameters. <br>\n<a href=\"https://www.kaggle.com/datasets/jaafarmahmoud1/pycolmap-040\" target=\"_blank\">https://www.kaggle.com/datasets/jaafarmahmoud1/pycolmap-040</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2300383": "Great thanks to the organizers and Kaggle staff for this amazing competition.\nOur solution shares architectural similarities with the baseline provided by the organizers. \nFirstly, we addressed the issue of rotation in-variance by implementing a rotation model to standardize the orientation of input images. Next, we employed a neural network for the retrieval task, enabling us to extract matching pairs for the feature extraction and matching process. This process generated a database which served as the input for the incremental mapping in Colmap.\n## Orientation Model\nWe would like to express our gratitude to @iglovikov for developing this great <a href=\"https://github.com/ternaus/check_orientation\">model</a> , which was trained on a large dataset and exhibited good performance in both (LB) and (CV) evaluations. We made a slight modification to this model, which we refer to as self-ensembling.\nIn our proposed adjustment, we utilized the model iteratively on the query image and its different rotated versions. Subsequently, we tuned the threshold by conducting validation tests on more than 1500 images with different rotation angles. Through experimentation, we discovered that using a threshold between 0.8 and 0.87 might result in incorrect orientation predictions. To address this issue, we applied self-ensembling by further rotating the image and checking if the predicted class probability fell within this range of thresholds. We repeated this last step twice to ensure accuracy.\n## Retrieval Method\n### Retrieval Model\nIn order to address the challenges posed by the cost inefficiency of exhaustive matching and the limitations of sequential matching for 3D reconstruction, we sought alternative methods for image retrieval. \nAfter careful evaluation, we chose to utilize NetVlad as our chosen method due to its superior performance compared to openibl and cosplace. To further enhance the results, we employed various techniques, including:\n1) We passed the original image and its horizontally flipped version through the model. By summing the descriptors obtained from both passes, we achieved improved performance, particularly for highly similar images. This technique is especially effective in cases where the scene exhibits symmetry. By redefining a new point in the n-dimensional descriptive space, which effectively increase the similarity distance between two distinct parts of the scene.\n\n2) Re-ranking: After calculating the similarity scores, we performed re-ranking by selecting the top 1 match. We then re-queried the retrieval process using these two images instead of just one. The similarity scores of the resulting matches were summed together after raising them to a specific power \"m\". This manipulation of probabilities ensures that if the best match for one of the query images is found, it will be favored over an image that is similar to both where the sum of the similarities on both scenarios is equal.\n\nWe repeated this procedure twice using NetVlad with Test Time Augmentation (TTA) on the image size. The results were nearly perfect, and the approach even enabled the correct ordering of The Wall scene pairs, which can be best matched by doing sequential matching.\n### Number of Matches\nDetermining the number of image pairs to select is a critical factor that directly affects both validation and leaderboard performance. This becomes particularly important for large scenes where there may be no or very few common images between subsets of the scene.\n\nTo address this challenge, we devised a strategy based on a complete graph representation. In this graph, each edge represents the similarity between two images. The goal is to choose a subset of edges where each image has an equal number of connected nodes.\n\nWe employed a Binary Search approach to determine the number of matches for each image, with a check function to verify if the resulting graph is connected or not. The lower bound of the binary search was set to half the number of images, ensuring that we consider common matches and prevent incomplete 3D model reconstruction. Additionally, we made sure that the approach remains exhaustive for small scenes containing less than 40 images.\n\nBy employing this method, we aimed to strike a balance between capturing sufficient matching pairs for accurate 3D reconstruction while avoiding redundant or disconnected image subsets.\n## Feature Extraction and Matching\nIn our selected submissions, we have utilized the SuperPoint and SuperGlue algorithms followed by MagSac filtering. Unfortunately SuperGlue is not licensed for commercial use. However, we have also achieved highly promising results using GlueStick. We have made modifications to the GlueStick architecture to integrate it with SuperPoint, and we achieved a score of approximately 0.450 on the public leaderboard and 0.513 on private leaderboard without employing our best tuning parameters. It is worth noting that this modified architecture is permitted for commercial use and offers improved processing speed. We anticipate that further tuning can yield even better results with GlueStick, but didn't choose it as our last submissions.\n## Refinement\nAlthough not included in our selected submissions, we would like to mention an approach that significantly improved validation results across various scenes. We employed Pixel-Perfect-SFM in conjunction with sd2net for dense matching, but it didn't improved the results on public leaderboard.\n\n## Registering Unregistered Images\nWhile not part of our selected submissions, we made attempts to register unregistered images using various techniques. However, these attempts did not yield significant improvements on the leaderboard. We explored the following strategies:\n- Utilizing different orientations and attempting registration.\n- Applying different extractor-matchers, such as LoFTR and R2D2, for the unregistered images.\n- Adjusting parameters for SuperPoint and SuperGlue to optimize the registration process.\n## Tried but not worked\n- Pixel-Perfect-SFM\n- Semantic Segmentation masks\n- Illumenation enhancement\n- PnP localization for unregistered Images\n- LoFTR/ SILK / DAC / Wireframe\n- CosPlace / OpenIBL\n- Large Image splitting into Parts\n- Grid-based Point Sampling (equally spatial replacement of points in an image).\n- Rotation self-ensemble averaging\n- Filtering the extracted pairs using a threshold from a binarysearch function or a fixed threshold.\n## Important Notes:\n1) Most participants were using colmap reconstruction, which is nondeterministic. We were able to modify the code so that it became deterministic but working only one CPU thread, which helped us observe our improvement and avoid randomness.\n2) We found out that OpenCV and PIL libs use EXIF Information when reading images, unless providing flags to prevent it. By that, we mean that if the orientation in EXIF will rotate the image automatically before processing. This was confusing for us, as there are missing information about how the GT were collected for rotation part (with or without considering them), that's why one of  our chosen solutions included this correction to overcome such issue, plus we lost a lot of submission to check the effect of this issue on leaderboard. It would have been more helpful if there was better explanation about how the GT calculated.\n3) Our validation scored 0.485 on public and 0.514 on private, with local score of 0.87.\n4) The fact that validation and leaderboard were not correlated made things more difficult and random, as it was clear there will be a shake up due to the fact that for some specific scene some solutions might fail which will drastically impact the mAA score.\n## Acknowldgment\n- I wanted to take a moment to express my sincere appreciation for the exceptional contribution and dedication to my friend and teammate @ammarali32 . His hard work and commitment have played a crucial role in our collective success, and I am truly grateful for the opportunity to work alongside such a remarkable individual. we had fun and we learned a lot.\n- Special thanks to the [Computer Vision and Geometry Lab](https://github.com/cvg) Group for (hloc, gluestick, pixel perfect sfm ..)\n##",
    "2300496": "Guys, great job and congratulations with gold and Kaggle Master level! We followed each other last weeks :) I'm impressed of your interesting approach for determining N of matches. Thanks for detailed write-up.",
    "2300502": "Thanks! Great work as well! Congrats on the 2nd place🎉!",
    "2300519": "Congrats @jaafarmahmoud1 and @ammarali32. Will you share the COLMAP modification that made it deterministic?",
    "2300544": "Thanks @gunesevitan! Yes, you can check this wheel for colmap 0.4.0. Basically, changing everything inside to one thread, and skipping [geometric verification ](https://github.com/colmap/pycolmap/blob/743a4ac305183f96d2a4cfce7c7f6418b31b8598/pipeline/match_features.cc#L76) made things deterministic. because this process is somehow random, then we replaced it with MagSac and provided a specific seed as USAC parameters. \nhttps://www.kaggle.com/datasets/jaafarmahmoud1/pycolmap-040",
    "2301061": "Thanks! you have done an amazing work congrats on the second place !!"
  },
  "source": "meta"
}