{
  "id": 510603,
  "title": "5th Place Solution: Customized Scene Matching",
  "url": "/competitions/image-matching-challenge-2024/writeups/khoa-ngo-5th-place-solution-customized-scene-match",
  "author_name": "",
  "post_date": "2024-06-06T20:10:16.743Z",
  "votes": 31,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all, thanks to the hosts for organizing the Image Matching Challenge for the third time on Kaggle. Also, congratulate the winners and wish the best to everyone who has gone through this tough journey safe and sound.</p>\n<p>I have to admit that I attended the competition quite late. Therefore, I just did whatever I “felt” right and improvised a lot. So, you might find my write-up a little bit tricky at some points :)). Let's go!</p>\n<h1>1. Overview</h1>\n<p>My solution architecture can be summarized in the following figure:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2F198ac7d486076bea7b708f4f0fc7682f%2Fimc24%20arch.png?generation=1717683469569615&amp;alt=media\" alt=\"arch\"></p>\n<p>I will split my summary into 4 main sections:</p>\n<ul>\n<li>Build local evaluation datasets</li>\n<li>Build a general SfM pipeline</li>\n<li>Customize the pipeline for each specific category</li>\n<li>Results</li>\n</ul>\n<h1>2. Local evaluation datasets</h1>\n<p>The original dataset is obviously too large and unrealistic to evaluate. To tackle this, I built 3 versions of subsets (by applying some random sampling strategies). The results on those 3 versions were highly correlated. Hence, I chose one as my main local validation dataset:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset/Scene</th>\n<th>Number of samples</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Pond</td>\n<td>100</td>\n</tr>\n<tr>\n<td>Lizard</td>\n<td>90</td>\n</tr>\n<tr>\n<td>Church</td>\n<td>80</td>\n</tr>\n<tr>\n<td>Dioscuri</td>\n<td>70</td>\n</tr>\n<tr>\n<td>Multi-temporal-temple-baalshamin</td>\n<td>68</td>\n</tr>\n<tr>\n<td>Transp_obj_glass_cup</td>\n<td>36</td>\n</tr>\n<tr>\n<td>Transp_obj_glass_cylinder</td>\n<td>36</td>\n</tr>\n</tbody>\n</table>\n<p>I will report my local CV results on this dataset.</p>\n<h1>3. SfM pipeline</h1>\n<p>For a better explanation, I will divide the pipeline into 3 modules:</p>\n<ul>\n<li>Proposing pair candidates by global descriptors</li>\n<li>Matching pairs in the candidate list</li>\n<li>Reconstruction with Colmap</li>\n</ul>\n<h2>3.1. Finding pair candidates</h2>\n<p>I used 3 pretrained models from <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm</a> to extract global features:</p>\n<pre><code>EVA-CLIP    \\\nConvNeXt     | --&gt;[]--&gt; [fc ]\nDinov2 ViT  /\n</code></pre>\n<p>I also customized similarity thresholds for different types of scenes (note that I used cosine similarity instead of distance). For example, with highly diverse scenes like Lizard, I used a small threshold of 0.6. In contrast, for object scenes like Cylinder and Cup, a high threshold of 0.95 would make much more sense.</p>\n<h2>3.2. Image matching</h2>\n<p>Following the spirit of my last year’s solution, I only focused on <strong>detector-based</strong> methods, because they are much more lightweight than semi-dense or dense models. Due to the lack of experiments on both CV and LB, I cannot give detailed numbers for each method, but only tell the sense of their performance in general. Here are some combinations that I tried during the competitions:</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>CV performance</th>\n<th>LB performance</th>\n<th>Remark</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SuperPoint + SuperGlue</td>\n<td>good</td>\n<td>not good</td>\n<td>slow compared to LightGlue</td>\n</tr>\n<tr>\n<td>GlueStick</td>\n<td>normal</td>\n<td>-</td>\n<td>slow</td>\n</tr>\n<tr>\n<td><strong>SuperPoint + LightGlue</strong></td>\n<td><strong>very good</strong></td>\n<td>not good</td>\n<td>no ideas why it is not good on LB</td>\n</tr>\n<tr>\n<td><strong>ALIKED + LightGlue</strong></td>\n<td>good</td>\n<td><strong>very good</strong></td>\n<td>best on LB</td>\n</tr>\n<tr>\n<td>DISK + LightGlue</td>\n<td>good</td>\n<td>normal</td>\n<td></td>\n</tr>\n<tr>\n<td>OmniGlue</td>\n<td>not good</td>\n<td>-</td>\n<td>slow</td>\n</tr>\n</tbody>\n</table>\n<p>In addition, I also tried SIFT + NN as a light ensemble model added to LightGlue. It is claimed in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/427143\" target=\"_blank\">IMC23 7th solution</a> that ensembling SIFT with LightGlue could boost the performance. I think it is the case in some scenes in CV, but it harmed the LB results a lot (-0.03).</p>\n<p>In summary, I could say that <strong>LightGlue</strong> is one of the best models in Image Matching problem at present (hats off to the authors)!</p>\n<h2>3.3. Reconstruction</h2>\n<p>This year I spent some time trying to make use of colmap better than the host baseline. Here are some of my trials:</p>\n<ul>\n<li><strong>Single camera for specific scenes</strong>: good on transparent object scenes in local CV, but worsens the result on LB.</li>\n<li><strong>Manual initial pair</strong>: I chose the pair with most matches and most keypoints and set it as the initial pair for colmap. But it made the results worse in my CV (maybe my code is wrong?!!).</li>\n<li><strong>Run Incremental Mapping multiple times</strong>: it made the results on my CV more consistent and a little better. However, I didn’t see it had any effect on LB.</li>\n</ul>\n<p>Sadly to say that I could not make any noticeable improvement with colmap.</p>\n<h1>4. Deal with specific categories</h1>\n<h2>4.1. Transparent objects</h2>\n<p>I think this category decides the <strong>gold medal</strong>. Because when I cracked the problem, I moved from top 300 to top 10 immediately!!!<br>\nFirst, I will present the process of how I created this solution, hope that it will be more vivid :).<br>\nIf you open one image and move to the next image-by-image, you can easily see that it is like a sequence of frames cut from a video captured by only one camera.<br>\nWhat a human might do to match those images? If it were me, I would focus on the <strong>object</strong> only and ignore the whole background. In addition, I only need to match a <strong>sequence</strong> (or a <strong>ring</strong>) of frames: (0, 1), (1, 2),…,(35, 0), which means there is no need to match (1, 20) or (3, 29)!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2Fbcdc3ca4607ce1c2b024a7e9b7f03aef%2Fimc24%20ring%20match.png?generation=1717692702373754&amp;alt=media\" alt=\"ring\"></p>\n<p>By adding these two lines of code before matching the cylinder, I went from ~0.0 to <strong>0.86</strong> on CV with LightGlue:</p>\n<pre><code> = [:, :, :]\nindex_pairs = [(i, i+)  i  (len()-)] + [(len()-, )]\n</code></pre>\n<p>So the problem now can be broken down into 2 smaller problems:</p>\n<ul>\n<li>Detect the object</li>\n<li>Find consecutive pairs</li>\n</ul>\n<h3>4.1.1. Object detection</h3>\n<p><strong>Method 1</strong> (not worked): Use object detection models: I tried <strong>Yolo</strong> family. Nonetheless, it is hard to make the prompt for the model to find out the exact object location.</p>\n<ul>\n<li><strong>Approach 1</strong>: Use <strong>Yolov8</strong> to produce all bounding boxes =&gt; it found no boxes!</li>\n<li><strong>Approach 2</strong>: Use <strong>YoloWorld</strong> with keyword “transparent object” =&gt; not detect anything.</li>\n<li><strong>Approach 3</strong>: Use <strong>YoloWorld</strong> with keyword “cup” =&gt; it could detect the cup. Yet, it is not generalized enough to apply to LB dataset.</li>\n</ul>\n<p><strong>Method 2</strong> (final solution): Use segmentation models:</p>\n<ul>\n<li><strong>Step 1</strong>: Use <strong><a href=\"https://github.com/ChaoningZhang/MobileSAM\" target=\"_blank\">MobileSAM</a></strong> to detect all the masks in the images.</li>\n<li><strong>Step 2</strong>: Use keypoint extractors above (SuperPoint or ALIKED) to extract object keypoints. It will produce some noise in the background.</li>\n<li><strong>Step 3</strong>: Find the best mask: the smallest mask which contains most of the keypoints.</li>\n<li><strong>Step 4</strong>: Find the smallest bounding box which includes the best mask.</li>\n</ul>\n<p>I tried DBSCAN too. However, due to the limitation of development time, I did not go to the end to see if it could replace SAM.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2Fe3891c005ff8a58801e207f7f7e3309e%2Fimc24%20mask.png?generation=1717688739099600&amp;alt=media\" alt=\"sam_mask\"></p>\n<h3>4.1.2. Finding the best pairs</h3>\n<p><strong>Problem definition</strong>:</p>\n<ul>\n<li>Version 1: Find the perfect sequence of images.</li>\n<li>Version 2 (simplified): For each image, find out two other best images for matching.</li>\n</ul>\n<p><strong>Method 1</strong> (not good): Stem from the fact that two consecutive frames would have the smallest movement among pixels, I calculated Optical Flow with <a href=\"https://pytorch.org/vision/0.12/auto_examples/plot_optical_flow.html\" target=\"_blank\">RAFT</a> for every possible pairs, then got the average magnitudes. I further combined with cosine similarity of Global embeddings above to build an NxN matrix:</p>\n<pre><code>pair_score = -of_mag + cos_sim\n</code></pre>\n<p>TSP solver could not produce the perfect ring (~90% accuracy). Thus, I think it is not good enough.</p>\n<p><strong>Method 2</strong> (simpler but better): I performed exhaustive matching on all pairs. Then, build an NxN matrix with elements that are the number of matches between each pair:</p>\n<pre><code>pair_score = n_matches\n</code></pre>\n<p>Subsequently, I just picked out the top-2 largest values on each row. This simple approach astonishingly produced nearly 100% accuracy (I tried different matchers, and sometimes there were 1-2 wrong pairs).</p>\n<h2>4.2. Day-night</h2>\n<p><strong>Method 1</strong>: Find dark images and enhance them.<br>\nDark detection is very easy, just find those images that have the mean value &lt; threshold (can be 65, 70).<br>\nI used <strong>CLAHE</strong> to enhance the dark images. Unfortunately, it downgraded my LB score by 0.01.<br>\n<strong>Method 2</strong>: Use <a href=\"https://github.com/THU-LYJ-Lab/DarkFeat\" target=\"_blank\">DarkFeat</a> matcher.<br>\nMy idea is to ensemble LightGlue with a more robust matcher for dark images. I found out that DarkFeat might be a promising candidate. Unluckily, I didn’t see any benefits of using it.</p>\n<h2>4.3. Symmetries-and-repeats</h2>\n<p>I assume that these scenes may have plenty of false positive matches because of their natural properties. <br>\n<strong>Method 1</strong>: First, I tried tuning parameters with much more strict values (e.g., increasing the matching threshold to 0.5, etc.). Nonetheless, it didn’t show any improvement.</p>\n<p><strong>Method 2</strong>: Next, I move to a more promising method: <a href=\"https://github.com/RuojinCai/Doppelgangers\" target=\"_blank\">Doppelgangers: Learning to Disambiguate Images of Similar Structures</a>. The main idea of this method is:</p>\n<ul>\n<li>First, run SfM one time to get all the matching pairs.</li>\n<li>Use the Doppelganger model to filter out pairs that have high probabilities as false positives (doppelgangers).<br>\nIt didn’t show any improvement on the church scene on my local CV.</li>\n</ul>\n<p>So basically, I could not make any improvement on this category.</p>\n<h2>4.4. Others</h2>\n<p>The remaining categories are nature, air-to-ground, historical_preservation, temporal.<br>\nActually, it is hard to customize these scenes, since they seem too general.<br>\nTherefore, I just applied the <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">rotation checking method</a> to correct the image orientations. It made a huge boost on the Dioscuri scene in local CV, and added about +0.005 to the LB.</p>\n<h1>5. Results</h1>\n<p><strong>Cross validation</strong></p>\n<table>\n<thead>\n<tr>\n<th>Scene</th>\n<th>Method</th>\n<th>Result</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Pond</td>\n<td>ALIKED+LightGlue</td>\n<td>0.46</td>\n</tr>\n<tr>\n<td>Lizard</td>\n<td>SP+LightGlue</td>\n<td>0.73</td>\n</tr>\n<tr>\n<td>Church</td>\n<td>ALIKED+LightGlue</td>\n<td>0.18</td>\n</tr>\n<tr>\n<td>Dioscuri</td>\n<td>ALIKED+LightGlue</td>\n<td>0.49</td>\n</tr>\n<tr>\n<td>Multi-temporal-temple-baalshamin</td>\n<td>ALIKED+LightGlue</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>Transp_obj_glass_cup</td>\n<td>DISK+LightGlue</td>\n<td>0.32</td>\n</tr>\n<tr>\n<td>Transp_obj_glass_cylinder</td>\n<td>ALIKED+LightGlue</td>\n<td>0.81</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Leaderboard</strong></p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>My IMC23: SP+SG</td>\n<td>0.104</td>\n<td>0.111</td>\n</tr>\n<tr>\n<td>SP+LG</td>\n<td>0.099</td>\n<td>0.115</td>\n</tr>\n<tr>\n<td>SP+LG+rot</td>\n<td>0.103</td>\n<td>0.119</td>\n</tr>\n<tr>\n<td>ALIKED+LG+rot</td>\n<td>0.135</td>\n<td>0.147</td>\n</tr>\n<tr>\n<td>ALIKED+LG+rot+transparent_custom</td>\n<td>0.18</td>\n<td>0.19</td>\n</tr>\n<tr>\n<td><strong>ALIKED+LG+rot+transparent_custom+tuning</strong></td>\n<td><strong>0.186</strong></td>\n<td><strong>0.195</strong></td>\n</tr>\n<tr>\n<td>ALIKED+LG+rot+transparent_custom+tuning+dark_enhance</td>\n<td>0.179</td>\n<td>0.183</td>\n</tr>\n</tbody>\n</table>\n<h1>6. Conclusion</h1>\n<p>A fun fact: I participated in all three Image Matching Challenges on Kaggle and earned medals of all three different colors 😂:</p>\n<ul>\n<li>2022: bronze🥉</li>\n<li>2023: silver 🥈</li>\n<li>2024: gold 🥇</li>\n</ul>\n<p>I feel exhausted after going solo in a tough competition like this. But finally, the sweet fruits have come after tireless efforts 😁.</p>\n<p>Thank you for your reading. I hope you can find something interesting in my write-up.</p>\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": "2859107",
      "postDate": "06/06/2024 19:37:59",
      "content": "<p>First of all, thanks to the hosts for organizing the Image Matching Challenge for the third time on Kaggle. Also, congratulate the winners and wish the best to everyone who has gone through this tough journey safe and sound.</p>\n<p>I have to admit that I attended the competition quite late. Therefore, I just did whatever I “felt” right and improvised a lot. So, you might find my write-up a little bit tricky at some points :)). Let's go!</p>\n<h1>1. Overview</h1>\n<p>My solution architecture can be summarized in the following figure:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2F198ac7d486076bea7b708f4f0fc7682f%2Fimc24%20arch.png?generation=1717683469569615&amp;alt=media\" alt=\"arch\"></p>\n<p>I will split my summary into 4 main sections:</p>\n<ul>\n<li>Build local evaluation datasets</li>\n<li>Build a general SfM pipeline</li>\n<li>Customize the pipeline for each specific category</li>\n<li>Results</li>\n</ul>\n<h1>2. Local evaluation datasets</h1>\n<p>The original dataset is obviously too large and unrealistic to evaluate. To tackle this, I built 3 versions of subsets (by applying some random sampling strategies). The results on those 3 versions were highly correlated. Hence, I chose one as my main local validation dataset:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset/Scene</th>\n<th>Number of samples</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Pond</td>\n<td>100</td>\n</tr>\n<tr>\n<td>Lizard</td>\n<td>90</td>\n</tr>\n<tr>\n<td>Church</td>\n<td>80</td>\n</tr>\n<tr>\n<td>Dioscuri</td>\n<td>70</td>\n</tr>\n<tr>\n<td>Multi-temporal-temple-baalshamin</td>\n<td>68</td>\n</tr>\n<tr>\n<td>Transp_obj_glass_cup</td>\n<td>36</td>\n</tr>\n<tr>\n<td>Transp_obj_glass_cylinder</td>\n<td>36</td>\n</tr>\n</tbody>\n</table>\n<p>I will report my local CV results on this dataset.</p>\n<h1>3. SfM pipeline</h1>\n<p>For a better explanation, I will divide the pipeline into 3 modules:</p>\n<ul>\n<li>Proposing pair candidates by global descriptors</li>\n<li>Matching pairs in the candidate list</li>\n<li>Reconstruction with Colmap</li>\n</ul>\n<h2>3.1. Finding pair candidates</h2>\n<p>I used 3 pretrained models from <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm</a> to extract global features:</p>\n<pre><code>EVA-CLIP    \\\nConvNeXt     | --&gt;[]--&gt; [fc ]\nDinov2 ViT  /\n</code></pre>\n<p>I also customized similarity thresholds for different types of scenes (note that I used cosine similarity instead of distance). For example, with highly diverse scenes like Lizard, I used a small threshold of 0.6. In contrast, for object scenes like Cylinder and Cup, a high threshold of 0.95 would make much more sense.</p>\n<h2>3.2. Image matching</h2>\n<p>Following the spirit of my last year’s solution, I only focused on <strong>detector-based</strong> methods, because they are much more lightweight than semi-dense or dense models. Due to the lack of experiments on both CV and LB, I cannot give detailed numbers for each method, but only tell the sense of their performance in general. Here are some combinations that I tried during the competitions:</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>CV performance</th>\n<th>LB performance</th>\n<th>Remark</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>SuperPoint + SuperGlue</td>\n<td>good</td>\n<td>not good</td>\n<td>slow compared to LightGlue</td>\n</tr>\n<tr>\n<td>GlueStick</td>\n<td>normal</td>\n<td>-</td>\n<td>slow</td>\n</tr>\n<tr>\n<td><strong>SuperPoint + LightGlue</strong></td>\n<td><strong>very good</strong></td>\n<td>not good</td>\n<td>no ideas why it is not good on LB</td>\n</tr>\n<tr>\n<td><strong>ALIKED + LightGlue</strong></td>\n<td>good</td>\n<td><strong>very good</strong></td>\n<td>best on LB</td>\n</tr>\n<tr>\n<td>DISK + LightGlue</td>\n<td>good</td>\n<td>normal</td>\n<td></td>\n</tr>\n<tr>\n<td>OmniGlue</td>\n<td>not good</td>\n<td>-</td>\n<td>slow</td>\n</tr>\n</tbody>\n</table>\n<p>In addition, I also tried SIFT + NN as a light ensemble model added to LightGlue. It is claimed in <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/427143\" target=\"_blank\">IMC23 7th solution</a> that ensembling SIFT with LightGlue could boost the performance. I think it is the case in some scenes in CV, but it harmed the LB results a lot (-0.03).</p>\n<p>In summary, I could say that <strong>LightGlue</strong> is one of the best models in Image Matching problem at present (hats off to the authors)!</p>\n<h2>3.3. Reconstruction</h2>\n<p>This year I spent some time trying to make use of colmap better than the host baseline. Here are some of my trials:</p>\n<ul>\n<li><strong>Single camera for specific scenes</strong>: good on transparent object scenes in local CV, but worsens the result on LB.</li>\n<li><strong>Manual initial pair</strong>: I chose the pair with most matches and most keypoints and set it as the initial pair for colmap. But it made the results worse in my CV (maybe my code is wrong?!!).</li>\n<li><strong>Run Incremental Mapping multiple times</strong>: it made the results on my CV more consistent and a little better. However, I didn’t see it had any effect on LB.</li>\n</ul>\n<p>Sadly to say that I could not make any noticeable improvement with colmap.</p>\n<h1>4. Deal with specific categories</h1>\n<h2>4.1. Transparent objects</h2>\n<p>I think this category decides the <strong>gold medal</strong>. Because when I cracked the problem, I moved from top 300 to top 10 immediately!!!<br>\nFirst, I will present the process of how I created this solution, hope that it will be more vivid :).<br>\nIf you open one image and move to the next image-by-image, you can easily see that it is like a sequence of frames cut from a video captured by only one camera.<br>\nWhat a human might do to match those images? If it were me, I would focus on the <strong>object</strong> only and ignore the whole background. In addition, I only need to match a <strong>sequence</strong> (or a <strong>ring</strong>) of frames: (0, 1), (1, 2),…,(35, 0), which means there is no need to match (1, 20) or (3, 29)!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2Fbcdc3ca4607ce1c2b024a7e9b7f03aef%2Fimc24%20ring%20match.png?generation=1717692702373754&amp;alt=media\" alt=\"ring\"></p>\n<p>By adding these two lines of code before matching the cylinder, I went from ~0.0 to <strong>0.86</strong> on CV with LightGlue:</p>\n<pre><code> = [:, :, :]\nindex_pairs = [(i, i+)  i  (len()-)] + [(len()-, )]\n</code></pre>\n<p>So the problem now can be broken down into 2 smaller problems:</p>\n<ul>\n<li>Detect the object</li>\n<li>Find consecutive pairs</li>\n</ul>\n<h3>4.1.1. Object detection</h3>\n<p><strong>Method 1</strong> (not worked): Use object detection models: I tried <strong>Yolo</strong> family. Nonetheless, it is hard to make the prompt for the model to find out the exact object location.</p>\n<ul>\n<li><strong>Approach 1</strong>: Use <strong>Yolov8</strong> to produce all bounding boxes =&gt; it found no boxes!</li>\n<li><strong>Approach 2</strong>: Use <strong>YoloWorld</strong> with keyword “transparent object” =&gt; not detect anything.</li>\n<li><strong>Approach 3</strong>: Use <strong>YoloWorld</strong> with keyword “cup” =&gt; it could detect the cup. Yet, it is not generalized enough to apply to LB dataset.</li>\n</ul>\n<p><strong>Method 2</strong> (final solution): Use segmentation models:</p>\n<ul>\n<li><strong>Step 1</strong>: Use <strong><a href=\"https://github.com/ChaoningZhang/MobileSAM\" target=\"_blank\">MobileSAM</a></strong> to detect all the masks in the images.</li>\n<li><strong>Step 2</strong>: Use keypoint extractors above (SuperPoint or ALIKED) to extract object keypoints. It will produce some noise in the background.</li>\n<li><strong>Step 3</strong>: Find the best mask: the smallest mask which contains most of the keypoints.</li>\n<li><strong>Step 4</strong>: Find the smallest bounding box which includes the best mask.</li>\n</ul>\n<p>I tried DBSCAN too. However, due to the limitation of development time, I did not go to the end to see if it could replace SAM.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2Fe3891c005ff8a58801e207f7f7e3309e%2Fimc24%20mask.png?generation=1717688739099600&amp;alt=media\" alt=\"sam_mask\"></p>\n<h3>4.1.2. Finding the best pairs</h3>\n<p><strong>Problem definition</strong>:</p>\n<ul>\n<li>Version 1: Find the perfect sequence of images.</li>\n<li>Version 2 (simplified): For each image, find out two other best images for matching.</li>\n</ul>\n<p><strong>Method 1</strong> (not good): Stem from the fact that two consecutive frames would have the smallest movement among pixels, I calculated Optical Flow with <a href=\"https://pytorch.org/vision/0.12/auto_examples/plot_optical_flow.html\" target=\"_blank\">RAFT</a> for every possible pairs, then got the average magnitudes. I further combined with cosine similarity of Global embeddings above to build an NxN matrix:</p>\n<pre><code>pair_score = -of_mag + cos_sim\n</code></pre>\n<p>TSP solver could not produce the perfect ring (~90% accuracy). Thus, I think it is not good enough.</p>\n<p><strong>Method 2</strong> (simpler but better): I performed exhaustive matching on all pairs. Then, build an NxN matrix with elements that are the number of matches between each pair:</p>\n<pre><code>pair_score = n_matches\n</code></pre>\n<p>Subsequently, I just picked out the top-2 largest values on each row. This simple approach astonishingly produced nearly 100% accuracy (I tried different matchers, and sometimes there were 1-2 wrong pairs).</p>\n<h2>4.2. Day-night</h2>\n<p><strong>Method 1</strong>: Find dark images and enhance them.<br>\nDark detection is very easy, just find those images that have the mean value &lt; threshold (can be 65, 70).<br>\nI used <strong>CLAHE</strong> to enhance the dark images. Unfortunately, it downgraded my LB score by 0.01.<br>\n<strong>Method 2</strong>: Use <a href=\"https://github.com/THU-LYJ-Lab/DarkFeat\" target=\"_blank\">DarkFeat</a> matcher.<br>\nMy idea is to ensemble LightGlue with a more robust matcher for dark images. I found out that DarkFeat might be a promising candidate. Unluckily, I didn’t see any benefits of using it.</p>\n<h2>4.3. Symmetries-and-repeats</h2>\n<p>I assume that these scenes may have plenty of false positive matches because of their natural properties. <br>\n<strong>Method 1</strong>: First, I tried tuning parameters with much more strict values (e.g., increasing the matching threshold to 0.5, etc.). Nonetheless, it didn’t show any improvement.</p>\n<p><strong>Method 2</strong>: Next, I move to a more promising method: <a href=\"https://github.com/RuojinCai/Doppelgangers\" target=\"_blank\">Doppelgangers: Learning to Disambiguate Images of Similar Structures</a>. The main idea of this method is:</p>\n<ul>\n<li>First, run SfM one time to get all the matching pairs.</li>\n<li>Use the Doppelganger model to filter out pairs that have high probabilities as false positives (doppelgangers).<br>\nIt didn’t show any improvement on the church scene on my local CV.</li>\n</ul>\n<p>So basically, I could not make any improvement on this category.</p>\n<h2>4.4. Others</h2>\n<p>The remaining categories are nature, air-to-ground, historical_preservation, temporal.<br>\nActually, it is hard to customize these scenes, since they seem too general.<br>\nTherefore, I just applied the <a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">rotation checking method</a> to correct the image orientations. It made a huge boost on the Dioscuri scene in local CV, and added about +0.005 to the LB.</p>\n<h1>5. Results</h1>\n<p><strong>Cross validation</strong></p>\n<table>\n<thead>\n<tr>\n<th>Scene</th>\n<th>Method</th>\n<th>Result</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Pond</td>\n<td>ALIKED+LightGlue</td>\n<td>0.46</td>\n</tr>\n<tr>\n<td>Lizard</td>\n<td>SP+LightGlue</td>\n<td>0.73</td>\n</tr>\n<tr>\n<td>Church</td>\n<td>ALIKED+LightGlue</td>\n<td>0.18</td>\n</tr>\n<tr>\n<td>Dioscuri</td>\n<td>ALIKED+LightGlue</td>\n<td>0.49</td>\n</tr>\n<tr>\n<td>Multi-temporal-temple-baalshamin</td>\n<td>ALIKED+LightGlue</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>Transp_obj_glass_cup</td>\n<td>DISK+LightGlue</td>\n<td>0.32</td>\n</tr>\n<tr>\n<td>Transp_obj_glass_cylinder</td>\n<td>ALIKED+LightGlue</td>\n<td>0.81</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Leaderboard</strong></p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>My IMC23: SP+SG</td>\n<td>0.104</td>\n<td>0.111</td>\n</tr>\n<tr>\n<td>SP+LG</td>\n<td>0.099</td>\n<td>0.115</td>\n</tr>\n<tr>\n<td>SP+LG+rot</td>\n<td>0.103</td>\n<td>0.119</td>\n</tr>\n<tr>\n<td>ALIKED+LG+rot</td>\n<td>0.135</td>\n<td>0.147</td>\n</tr>\n<tr>\n<td>ALIKED+LG+rot+transparent_custom</td>\n<td>0.18</td>\n<td>0.19</td>\n</tr>\n<tr>\n<td><strong>ALIKED+LG+rot+transparent_custom+tuning</strong></td>\n<td><strong>0.186</strong></td>\n<td><strong>0.195</strong></td>\n</tr>\n<tr>\n<td>ALIKED+LG+rot+transparent_custom+tuning+dark_enhance</td>\n<td>0.179</td>\n<td>0.183</td>\n</tr>\n</tbody>\n</table>\n<h1>6. Conclusion</h1>\n<p>A fun fact: I participated in all three Image Matching Challenges on Kaggle and earned medals of all three different colors 😂:</p>\n<ul>\n<li>2022: bronze🥉</li>\n<li>2023: silver 🥈</li>\n<li>2024: gold 🥇</li>\n</ul>\n<p>I feel exhausted after going solo in a tough competition like this. But finally, the sweet fruits have come after tireless efforts 😁.</p>\n<p>Thank you for your reading. I hope you can find something interesting in my write-up.</p>\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "First of all, thanks to the hosts for organizing the Image Matching Challenge for the third time on Kaggle. Also, congratulate the winners and wish the best to everyone who has gone through this tough journey safe and sound.\n\nI have to admit that I attended the competition quite late. Therefore, I just did whatever I “felt” right and improvised a lot. So, you might find my write-up a little bit tricky at some points :)). Let's go!\n\n# 1. Overview \nMy solution architecture can be summarized in the following figure:\n![arch](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2F198ac7d486076bea7b708f4f0fc7682f%2Fimc24%20arch.png?generation=1717683469569615&alt=media)\n\nI will split my summary into 4 main sections:\n- Build local evaluation datasets\n- Build a general SfM pipeline\n- Customize the pipeline for each specific category\n- Results\n\n# 2. Local evaluation datasets\nThe original dataset is obviously too large and unrealistic to evaluate. To tackle this, I built 3 versions of subsets (by applying some random sampling strategies). The results on those 3 versions were highly correlated. Hence, I chose one as my main local validation dataset:\n\n| Dataset/Scene |Number of samples  |\n| :--- | :---: |\n| Pond | 100 |\n|Lizard|90|\n|Church|80|\n|Dioscuri|70|\n|Multi-temporal-temple-baalshamin|68|\n|Transp_obj_glass_cup|36|\n|Transp_obj_glass_cylinder|36|\n\n\nI will report my local CV results on this dataset.\n\n# 3. SfM pipeline\nFor a better explanation, I will divide the pipeline into 3 modules:\n- Proposing pair candidates by global descriptors\n- Matching pairs in the candidate list\n- Reconstruction with Colmap\n\n## 3.1. Finding pair candidates\nI used 3 pretrained models from [timm](https://github.com/huggingface/pytorch-image-models) to extract global features:\n```\nEVA-CLIP Base   \\\nConvNeXt Base    | -->[concat]--> [fc 2560]\nDinov2 ViT Base /\n```\nI also customized similarity thresholds for different types of scenes (note that I used cosine similarity instead of distance). For example, with highly diverse scenes like Lizard, I used a small threshold of 0.6. In contrast, for object scenes like Cylinder and Cup, a high threshold of 0.95 would make much more sense.\n\n## 3.2. Image matching\nFollowing the spirit of my last year’s solution, I only focused on **detector-based** methods, because they are much more lightweight than semi-dense or dense models. Due to the lack of experiments on both CV and LB, I cannot give detailed numbers for each method, but only tell the sense of their performance in general. Here are some combinations that I tried during the competitions:\n\n| Method                           |CV performance|LB performance|Remark                                              |\n| :------------------------|:----------------:|:---------------:|:-----------------------------------|\n|SuperPoint + SuperGlue|       good             |     not good      |     slow compared to LightGlue      |\n|GlueStick                         |         normal        |            -             |                       slow                         |\n|**SuperPoint + LightGlue** |    **very good**       |     not good      | no ideas why it is not good on LB  |\n|**ALIKED + LightGlue**|good|**very good**|best on LB|\n|DISK + LightGlue| good|normal||\n|OmniGlue|not good|-|slow|\n\n\nIn addition, I also tried SIFT + NN as a light ensemble model added to LightGlue. It is claimed in [IMC23 7th solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/427143) that ensembling SIFT with LightGlue could boost the performance. I think it is the case in some scenes in CV, but it harmed the LB results a lot (-0.03).\n\nIn summary, I could say that **LightGlue** is one of the best models in Image Matching problem at present (hats off to the authors)!\n\n## 3.3. Reconstruction\nThis year I spent some time trying to make use of colmap better than the host baseline. Here are some of my trials:\n- **Single camera for specific scenes**: good on transparent object scenes in local CV, but worsens the result on LB.\n- **Manual initial pair**: I chose the pair with most matches and most keypoints and set it as the initial pair for colmap. But it made the results worse in my CV (maybe my code is wrong?!!).\n- **Run Incremental Mapping multiple times**: it made the results on my CV more consistent and a little better. However, I didn’t see it had any effect on LB.\n\nSadly to say that I could not make any noticeable improvement with colmap.\n\n# 4. Deal with specific categories\n## 4.1. Transparent objects\nI think this category decides the **gold medal**. Because when I cracked the problem, I moved from top 300 to top 10 immediately!!!\nFirst, I will present the process of how I created this solution, hope that it will be more vivid :).\nIf you open one image and move to the next image-by-image, you can easily see that it is like a sequence of frames cut from a video captured by only one camera.\nWhat a human might do to match those images? If it were me, I would focus on the **object** only and ignore the whole background. In addition, I only need to match a **sequence** (or a **ring**) of frames: (0, 1), (1, 2),...,(35, 0), which means there is no need to match (1, 20) or (3, 29)!\n\n![ring](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2Fbcdc3ca4607ce1c2b024a7e9b7f03aef%2Fimc24%20ring%20match.png?generation=1717692702373754&alt=media)\n\nBy adding these two lines of code before matching the cylinder, I went from ~0.0 to **0.86** on CV with LightGlue:\n```\nimage = image[950:2850, :, :]\nindex_pairs = [(i, i+1) for i in range(len(scene)-1)] + [(len(scene)-1, 0)]\n```\nSo the problem now can be broken down into 2 smaller problems:\n- Detect the object\n- Find consecutive pairs\n### 4.1.1. Object detection\n**Method 1** (not worked): Use object detection models: I tried **Yolo** family. Nonetheless, it is hard to make the prompt for the model to find out the exact object location.\n- **Approach 1**: Use **Yolov8** to produce all bounding boxes => it found no boxes!\n- **Approach 2**: Use **YoloWorld** with keyword “transparent object” => not detect anything.\n- **Approach 3**: Use **YoloWorld** with keyword “cup” => it could detect the cup. Yet, it is not generalized enough to apply to LB dataset.\n\n**Method 2** (final solution): Use segmentation models:\n- **Step 1**: Use **[MobileSAM](https://github.com/ChaoningZhang/MobileSAM)** to detect all the masks in the images.\n- **Step 2**: Use keypoint extractors above (SuperPoint or ALIKED) to extract object keypoints. It will produce some noise in the background.\n- **Step 3**: Find the best mask: the smallest mask which contains most of the keypoints.\n- **Step 4**: Find the smallest bounding box which includes the best mask.\n\nI tried DBSCAN too. However, due to the limitation of development time, I did not go to the end to see if it could replace SAM.\n\n![sam_mask](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2Fe3891c005ff8a58801e207f7f7e3309e%2Fimc24%20mask.png?generation=1717688739099600&alt=media)\n\n### 4.1.2. Finding the best pairs\n**Problem definition**:\n- Version 1: Find the perfect sequence of images.\n- Version 2 (simplified): For each image, find out two other best images for matching.\n\n**Method 1** (not good): Stem from the fact that two consecutive frames would have the smallest movement among pixels, I calculated Optical Flow with [RAFT](https://pytorch.org/vision/0.12/auto_examples/plot_optical_flow.html) for every possible pairs, then got the average magnitudes. I further combined with cosine similarity of Global embeddings above to build an NxN matrix:\n```\npair_score[i, j] = -of_mag[i, j] + cos_sim[i, j]\n```\n\nTSP solver could not produce the perfect ring (~90% accuracy). Thus, I think it is not good enough.\n\n**Method 2** (simpler but better): I performed exhaustive matching on all pairs. Then, build an NxN matrix with elements that are the number of matches between each pair:\n```\npair_score[i, j] = n_matches[i, j]\n```\n\nSubsequently, I just picked out the top-2 largest values on each row. This simple approach astonishingly produced nearly 100% accuracy (I tried different matchers, and sometimes there were 1-2 wrong pairs).\n\n## 4.2. Day-night\n**Method 1**: Find dark images and enhance them.\nDark detection is very easy, just find those images that have the mean value < threshold (can be 65, 70).\nI used **CLAHE** to enhance the dark images. Unfortunately, it downgraded my LB score by 0.01.\n**Method 2**: Use [DarkFeat](https://github.com/THU-LYJ-Lab/DarkFeat) matcher.\nMy idea is to ensemble LightGlue with a more robust matcher for dark images. I found out that DarkFeat might be a promising candidate. Unluckily, I didn’t see any benefits of using it.\n\n## 4.3. Symmetries-and-repeats\nI assume that these scenes may have plenty of false positive matches because of their natural properties. \n**Method 1**: First, I tried tuning parameters with much more strict values (e.g., increasing the matching threshold to 0.5, etc.). Nonetheless, it didn’t show any improvement.\n\n**Method 2**: Next, I move to a more promising method: [Doppelgangers: Learning to Disambiguate Images of Similar Structures](https://github.com/RuojinCai/Doppelgangers). The main idea of this method is:\n- First, run SfM one time to get all the matching pairs.\n- Use the Doppelganger model to filter out pairs that have high probabilities as false positives (doppelgangers).\nIt didn’t show any improvement on the church scene on my local CV.\n\nSo basically, I could not make any improvement on this category.\n\n## 4.4. Others\nThe remaining categories are nature, air-to-ground, historical_preservation, temporal.\nActually, it is hard to customize these scenes, since they seem too general.\nTherefore, I just applied the [rotation checking method](https://github.com/ternaus/check_orientation) to correct the image orientations. It made a huge boost on the Dioscuri scene in local CV, and added about +0.005 to the LB.\n\n# 5. Results\n**Cross validation**\n\n|Scene|Method|Result|\n|---|---|---|\n|Pond|ALIKED+LightGlue|0.46|\n|Lizard|SP+LightGlue|0.73|\n|Church|ALIKED+LightGlue|0.18|\n|Dioscuri|ALIKED+LightGlue|0.49|\n|Multi-temporal-temple-baalshamin|ALIKED+LightGlue|0.48|\n|Transp_obj_glass_cup|DISK+LightGlue|0.32|\n|Transp_obj_glass_cylinder|ALIKED+LightGlue|0.81|\n\n**Leaderboard**\n\n|Method|Public LB|Private LB|\n|---|---|---|\n|My IMC23: SP+SG|0.104|0.111|\n|SP+LG|0.099|0.115|\n|SP+LG+rot|0.103|0.119|\n|ALIKED+LG+rot|0.135|0.147|\n|ALIKED+LG+rot+transparent_custom|0.18|0.19|\n|**ALIKED+LG+rot+transparent_custom+tuning**|**0.186**|**0.195**|\n|ALIKED+LG+rot+transparent_custom+tuning+dark_enhance|0.179|0.183|\n\n# 6. Conclusion\nA fun fact: I participated in all three Image Matching Challenges on Kaggle and earned medals of all three different colors 😂:\n- 2022: bronze🥉\n- 2023: silver 🥈\n- 2024: gold 🥇\n\nI feel exhausted after going solo in a tough competition like this. But finally, the sweet fruits have come after tireless efforts 😁.\n\nThank you for your reading. I hope you can find something interesting in my write-up.\n\nHappy Kaggling!",
      "votes": null
    },
    {
      "id": "2859255",
      "postDate": "06/06/2024 22:16:17",
      "content": "<p>Impressive work, especially for solo, congrats!</p>",
      "rawMarkdown": "Impressive work, especially for solo, congrats!",
      "votes": null
    },
    {
      "id": "2859362",
      "postDate": "06/07/2024 01:53:21",
      "content": "<p>Thank you. Congrats your 1st place, too. Your team is truly amazing!</p>",
      "rawMarkdown": "Thank you. Congrats your 1st place, too. Your team is truly amazing!",
      "votes": null
    },
    {
      "id": "2864247",
      "postDate": "06/10/2024 04:23:08",
      "content": "<p>Congrat e <a href=\"https://www.kaggle.com/nejicool96\" target=\"_blank\">@nejicool96</a> for solo gold medal ! Chờ mãi thành quả rồi cũng đến nhỉ 🎉</p>",
      "rawMarkdown": "Congrat e @nejicool96 for solo gold medal ! Chờ mãi thành quả rồi cũng đến nhỉ 🎉",
      "votes": null
    },
    {
      "id": "2864460",
      "postDate": "06/10/2024 06:43:42",
      "content": "<p>Thanks anh. Fighting so long for this.<br>\nAnh P. đúng ko ạ :)).</p>",
      "rawMarkdown": "Thanks anh. Fighting so long for this.\nAnh P. đúng ko ạ :)).",
      "votes": null
    },
    {
      "id": "2885286",
      "postDate": "06/23/2024 02:41:36",
      "content": "<p>Congrats e. Kiên trì quá </p>",
      "rawMarkdown": "Congrats e. Kiên trì quá",
      "votes": null
    },
    {
      "id": "2887013",
      "postDate": "06/24/2024 01:46:23",
      "content": "<p>Thanks anh 😄.</p>",
      "rawMarkdown": "Thanks anh 😄.",
      "votes": null
    },
    {
      "id": "2900458",
      "postDate": "07/02/2024 09:59:32",
      "content": "<p>detailed solution guys, so happy to see you score all 3 color medals :D</p>",
      "rawMarkdown": "detailed solution guys, so happy to see you score all 3 color medals :D",
      "votes": null
    },
    {
      "id": "2900553",
      "postDate": "07/02/2024 10:59:38",
      "content": "<p>Thank you guys 😁</p>",
      "rawMarkdown": "Thank you guys 😁",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2859255,
      "author_name": "vostankovich",
      "author_url": "",
      "post_date": "06/06/2024 22:16:17",
      "content": "<p>Impressive work, especially for solo, congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2859362,
          "author_name": "nejicool96",
          "author_url": "",
          "post_date": "06/07/2024 01:53:21",
          "content": "<p>Thank you. Congrats your 1st place, too. Your team is truly amazing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2864247,
      "author_name": "nocturnebflat123",
      "author_url": "",
      "post_date": "06/10/2024 04:23:08",
      "content": "<p>Congrat e <a href=\"https://www.kaggle.com/nejicool96\" target=\"_blank\">@nejicool96</a> for solo gold medal ! Chờ mãi thành quả rồi cũng đến nhỉ 🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 2864460,
          "author_name": "nejicool96",
          "author_url": "",
          "post_date": "06/10/2024 06:43:42",
          "content": "<p>Thanks anh. Fighting so long for this.<br>\nAnh P. đúng ko ạ :)).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2885286,
      "author_name": "truonghoang",
      "author_url": "",
      "post_date": "06/23/2024 02:41:36",
      "content": "<p>Congrats e. Kiên trì quá </p>",
      "votes": null,
      "replies": [
        {
          "id": 2887013,
          "author_name": "nejicool96",
          "author_url": "",
          "post_date": "06/24/2024 01:46:23",
          "content": "<p>Thanks anh 😄.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2900458,
      "author_name": "shyamgupta196",
      "author_url": "",
      "post_date": "07/02/2024 09:59:32",
      "content": "<p>detailed solution guys, so happy to see you score all 3 color medals :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 2900553,
          "author_name": "nejicool96",
          "author_url": "",
          "post_date": "07/02/2024 10:59:38",
          "content": "<p>Thank you guys 😁</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2859107": "First of all, thanks to the hosts for organizing the Image Matching Challenge for the third time on Kaggle. Also, congratulate the winners and wish the best to everyone who has gone through this tough journey safe and sound.\n\nI have to admit that I attended the competition quite late. Therefore, I just did whatever I “felt” right and improvised a lot. So, you might find my write-up a little bit tricky at some points :)). Let's go!\n\n# 1. Overview \nMy solution architecture can be summarized in the following figure:\n![arch](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2F198ac7d486076bea7b708f4f0fc7682f%2Fimc24%20arch.png?generation=1717683469569615&alt=media)\n\nI will split my summary into 4 main sections:\n- Build local evaluation datasets\n- Build a general SfM pipeline\n- Customize the pipeline for each specific category\n- Results\n\n# 2. Local evaluation datasets\nThe original dataset is obviously too large and unrealistic to evaluate. To tackle this, I built 3 versions of subsets (by applying some random sampling strategies). The results on those 3 versions were highly correlated. Hence, I chose one as my main local validation dataset:\n\n| Dataset/Scene |Number of samples  |\n| :--- | :---: |\n| Pond | 100 |\n|Lizard|90|\n|Church|80|\n|Dioscuri|70|\n|Multi-temporal-temple-baalshamin|68|\n|Transp_obj_glass_cup|36|\n|Transp_obj_glass_cylinder|36|\n\n\nI will report my local CV results on this dataset.\n\n# 3. SfM pipeline\nFor a better explanation, I will divide the pipeline into 3 modules:\n- Proposing pair candidates by global descriptors\n- Matching pairs in the candidate list\n- Reconstruction with Colmap\n\n## 3.1. Finding pair candidates\nI used 3 pretrained models from [timm](https://github.com/huggingface/pytorch-image-models) to extract global features:\n```\nEVA-CLIP Base   \\\nConvNeXt Base    | -->[concat]--> [fc 2560]\nDinov2 ViT Base /\n```\nI also customized similarity thresholds for different types of scenes (note that I used cosine similarity instead of distance). For example, with highly diverse scenes like Lizard, I used a small threshold of 0.6. In contrast, for object scenes like Cylinder and Cup, a high threshold of 0.95 would make much more sense.\n\n## 3.2. Image matching\nFollowing the spirit of my last year’s solution, I only focused on **detector-based** methods, because they are much more lightweight than semi-dense or dense models. Due to the lack of experiments on both CV and LB, I cannot give detailed numbers for each method, but only tell the sense of their performance in general. Here are some combinations that I tried during the competitions:\n\n| Method                           |CV performance|LB performance|Remark                                              |\n| :------------------------|:----------------:|:---------------:|:-----------------------------------|\n|SuperPoint + SuperGlue|       good             |     not good      |     slow compared to LightGlue      |\n|GlueStick                         |         normal        |            -             |                       slow                         |\n|**SuperPoint + LightGlue** |    **very good**       |     not good      | no ideas why it is not good on LB  |\n|**ALIKED + LightGlue**|good|**very good**|best on LB|\n|DISK + LightGlue| good|normal||\n|OmniGlue|not good|-|slow|\n\n\nIn addition, I also tried SIFT + NN as a light ensemble model added to LightGlue. It is claimed in [IMC23 7th solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/427143) that ensembling SIFT with LightGlue could boost the performance. I think it is the case in some scenes in CV, but it harmed the LB results a lot (-0.03).\n\nIn summary, I could say that **LightGlue** is one of the best models in Image Matching problem at present (hats off to the authors)!\n\n## 3.3. Reconstruction\nThis year I spent some time trying to make use of colmap better than the host baseline. Here are some of my trials:\n- **Single camera for specific scenes**: good on transparent object scenes in local CV, but worsens the result on LB.\n- **Manual initial pair**: I chose the pair with most matches and most keypoints and set it as the initial pair for colmap. But it made the results worse in my CV (maybe my code is wrong?!!).\n- **Run Incremental Mapping multiple times**: it made the results on my CV more consistent and a little better. However, I didn’t see it had any effect on LB.\n\nSadly to say that I could not make any noticeable improvement with colmap.\n\n# 4. Deal with specific categories\n## 4.1. Transparent objects\nI think this category decides the **gold medal**. Because when I cracked the problem, I moved from top 300 to top 10 immediately!!!\nFirst, I will present the process of how I created this solution, hope that it will be more vivid :).\nIf you open one image and move to the next image-by-image, you can easily see that it is like a sequence of frames cut from a video captured by only one camera.\nWhat a human might do to match those images? If it were me, I would focus on the **object** only and ignore the whole background. In addition, I only need to match a **sequence** (or a **ring**) of frames: (0, 1), (1, 2),...,(35, 0), which means there is no need to match (1, 20) or (3, 29)!\n\n![ring](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2Fbcdc3ca4607ce1c2b024a7e9b7f03aef%2Fimc24%20ring%20match.png?generation=1717692702373754&alt=media)\n\nBy adding these two lines of code before matching the cylinder, I went from ~0.0 to **0.86** on CV with LightGlue:\n```\nimage = image[950:2850, :, :]\nindex_pairs = [(i, i+1) for i in range(len(scene)-1)] + [(len(scene)-1, 0)]\n```\nSo the problem now can be broken down into 2 smaller problems:\n- Detect the object\n- Find consecutive pairs\n### 4.1.1. Object detection\n**Method 1** (not worked): Use object detection models: I tried **Yolo** family. Nonetheless, it is hard to make the prompt for the model to find out the exact object location.\n- **Approach 1**: Use **Yolov8** to produce all bounding boxes => it found no boxes!\n- **Approach 2**: Use **YoloWorld** with keyword “transparent object” => not detect anything.\n- **Approach 3**: Use **YoloWorld** with keyword “cup” => it could detect the cup. Yet, it is not generalized enough to apply to LB dataset.\n\n**Method 2** (final solution): Use segmentation models:\n- **Step 1**: Use **[MobileSAM](https://github.com/ChaoningZhang/MobileSAM)** to detect all the masks in the images.\n- **Step 2**: Use keypoint extractors above (SuperPoint or ALIKED) to extract object keypoints. It will produce some noise in the background.\n- **Step 3**: Find the best mask: the smallest mask which contains most of the keypoints.\n- **Step 4**: Find the smallest bounding box which includes the best mask.\n\nI tried DBSCAN too. However, due to the limitation of development time, I did not go to the end to see if it could replace SAM.\n\n![sam_mask](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2198466%2Fe3891c005ff8a58801e207f7f7e3309e%2Fimc24%20mask.png?generation=1717688739099600&alt=media)\n\n### 4.1.2. Finding the best pairs\n**Problem definition**:\n- Version 1: Find the perfect sequence of images.\n- Version 2 (simplified): For each image, find out two other best images for matching.\n\n**Method 1** (not good): Stem from the fact that two consecutive frames would have the smallest movement among pixels, I calculated Optical Flow with [RAFT](https://pytorch.org/vision/0.12/auto_examples/plot_optical_flow.html) for every possible pairs, then got the average magnitudes. I further combined with cosine similarity of Global embeddings above to build an NxN matrix:\n```\npair_score[i, j] = -of_mag[i, j] + cos_sim[i, j]\n```\n\nTSP solver could not produce the perfect ring (~90% accuracy). Thus, I think it is not good enough.\n\n**Method 2** (simpler but better): I performed exhaustive matching on all pairs. Then, build an NxN matrix with elements that are the number of matches between each pair:\n```\npair_score[i, j] = n_matches[i, j]\n```\n\nSubsequently, I just picked out the top-2 largest values on each row. This simple approach astonishingly produced nearly 100% accuracy (I tried different matchers, and sometimes there were 1-2 wrong pairs).\n\n## 4.2. Day-night\n**Method 1**: Find dark images and enhance them.\nDark detection is very easy, just find those images that have the mean value < threshold (can be 65, 70).\nI used **CLAHE** to enhance the dark images. Unfortunately, it downgraded my LB score by 0.01.\n**Method 2**: Use [DarkFeat](https://github.com/THU-LYJ-Lab/DarkFeat) matcher.\nMy idea is to ensemble LightGlue with a more robust matcher for dark images. I found out that DarkFeat might be a promising candidate. Unluckily, I didn’t see any benefits of using it.\n\n## 4.3. Symmetries-and-repeats\nI assume that these scenes may have plenty of false positive matches because of their natural properties. \n**Method 1**: First, I tried tuning parameters with much more strict values (e.g., increasing the matching threshold to 0.5, etc.). Nonetheless, it didn’t show any improvement.\n\n**Method 2**: Next, I move to a more promising method: [Doppelgangers: Learning to Disambiguate Images of Similar Structures](https://github.com/RuojinCai/Doppelgangers). The main idea of this method is:\n- First, run SfM one time to get all the matching pairs.\n- Use the Doppelganger model to filter out pairs that have high probabilities as false positives (doppelgangers).\nIt didn’t show any improvement on the church scene on my local CV.\n\nSo basically, I could not make any improvement on this category.\n\n## 4.4. Others\nThe remaining categories are nature, air-to-ground, historical_preservation, temporal.\nActually, it is hard to customize these scenes, since they seem too general.\nTherefore, I just applied the [rotation checking method](https://github.com/ternaus/check_orientation) to correct the image orientations. It made a huge boost on the Dioscuri scene in local CV, and added about +0.005 to the LB.\n\n# 5. Results\n**Cross validation**\n\n|Scene|Method|Result|\n|---|---|---|\n|Pond|ALIKED+LightGlue|0.46|\n|Lizard|SP+LightGlue|0.73|\n|Church|ALIKED+LightGlue|0.18|\n|Dioscuri|ALIKED+LightGlue|0.49|\n|Multi-temporal-temple-baalshamin|ALIKED+LightGlue|0.48|\n|Transp_obj_glass_cup|DISK+LightGlue|0.32|\n|Transp_obj_glass_cylinder|ALIKED+LightGlue|0.81|\n\n**Leaderboard**\n\n|Method|Public LB|Private LB|\n|---|---|---|\n|My IMC23: SP+SG|0.104|0.111|\n|SP+LG|0.099|0.115|\n|SP+LG+rot|0.103|0.119|\n|ALIKED+LG+rot|0.135|0.147|\n|ALIKED+LG+rot+transparent_custom|0.18|0.19|\n|**ALIKED+LG+rot+transparent_custom+tuning**|**0.186**|**0.195**|\n|ALIKED+LG+rot+transparent_custom+tuning+dark_enhance|0.179|0.183|\n\n# 6. Conclusion\nA fun fact: I participated in all three Image Matching Challenges on Kaggle and earned medals of all three different colors 😂:\n- 2022: bronze🥉\n- 2023: silver 🥈\n- 2024: gold 🥇\n\nI feel exhausted after going solo in a tough competition like this. But finally, the sweet fruits have come after tireless efforts 😁.\n\nThank you for your reading. I hope you can find something interesting in my write-up.\n\nHappy Kaggling!",
    "2859255": "Impressive work, especially for solo, congrats!",
    "2859362": "Thank you. Congrats your 1st place, too. Your team is truly amazing!",
    "2864247": "Congrat e @nejicool96 for solo gold medal ! Chờ mãi thành quả rồi cũng đến nhỉ 🎉",
    "2864460": "Thanks anh. Fighting so long for this.\nAnh P. đúng ko ạ :)).",
    "2885286": "Congrats e. Kiên trì quá",
    "2887013": "Thanks anh 😄.",
    "2900458": "detailed solution guys, so happy to see you score all 3 color medals :D",
    "2900553": "Thank you guys 😁"
  },
  "source": "meta"
}