{
  "id": 583076,
  "title": "6th Place Solution",
  "url": "/competitions/image-matching-challenge-2025/writeups/never-happen-6th-place-solution",
  "author_name": "",
  "post_date": "2025-06-04T14:56:11.174660Z",
  "votes": 30,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I'd like to express my gratitude to the hosts and Kaggle staffs for organizing such an exciting and practical competition!<br>\nI would also like to thank the competitors who pointed out the metric bugs, and the staff for promptly addressing them.</p>\n<h1>Overview</h1>\n<ul>\n<li><p>To increase the number of matching pairs with consistent orientation, the orientation of all images was detected and then aligned by rotating the images accordingly.</p></li>\n<li><p>As with many top competitors from last year, I primarily used ALIKED for keypoint detection and LightGlue for matching.</p></li>\n<li><p>Instead of using global features, I matched all image pairs and filtered the top k * log(n)/(n-1)% image pairs to extract matches within each scene (cluster) accurately. Here, k is the parameter to adjust the ratio of selection, and n is the number of images in the dataset.</p></li>\n<li><p>To improve mAA scores using a small number of accurate pairs, I adopted two complementary strategies: <strong>locally</strong>, I applied high-density crop matching by cropping regions with densely matched keypoints; <strong>globally</strong>, I applied image 4splits by splitting each image into four parts, which significantly increased the number of matched keypoints across the entire image.</p></li>\n<li><p>To mitigate mAA score fluctuations due to image scale, pair filtering and high-density crop matching were executed at multiple scales(resize_to = 1024, 1280, 1536, 2048) and ensembled into the colmap.</p></li>\n</ul>\n<h1>Pipeline</h1>\n<p>I submitted two notebooks.</p>\n<table>\n<thead>\n<tr>\n<th>Notebook</th>\n<th>CV*</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>① Best CV</td>\n<td>58.64</td>\n<td>46.67</td>\n<td>46.61</td>\n</tr>\n<tr>\n<td>② Best LB</td>\n<td>54.86</td>\n<td>46.98</td>\n<td>45.33</td>\n</tr>\n</tbody>\n</table>\n<p>*CV used all datasets in the train folder.</p>\n<h3>① Best CV</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12213964%2F48c439d9fb58c4ef6dd446b10c1614d4%2F1.png?generation=1749048960243147&amp;alt=media\" alt=\"\"></p>\n<h4>Steps</h4>\n<ol>\n<li><p>Detect the orientation of each image and rotate them to align their directions consistently.(<a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">https://github.com/ternaus/check_orientation</a>)</p></li>\n<li><p>Match for all image pairs using ALIKED (resize_to=1024) + LightGlue.</p></li>\n<li><p>Select the top k * log(n)/(n-1)% pairs based on the number of matches. k is the parameter to adjust the ratio of selection, and n is the number of images in the dataset.<br>\nk was experimentally determined to achieve high scores on both CV and LB. The same rationale applies to the subsequent steps.</p></li>\n<li><p>With the filtered image pairs from step 3 (k=7), match using ALIKED (resize_to=1280, 1536, 2048) + LightGlue, and again select the top k * log(n)/(n-1)% pairs.</p></li>\n<li><p>From the merged results of all scales (resize_to=1024,1280,1536,2048), select the top k=1.5 pairs.</p></li>\n</ol>\n<p>6-1. Split each image into four parts, and match using ALIKED (resize_to=1024) + LightGlue.</p>\n<p>6-2. Crop regions with high match density from each image and match using ALIKED (resize_to=1024,1280,1536,2048) + LightGlue.</p>\n<h3>② Best LB (description only)</h3>\n<h4>Steps</h4>\n<ol>\n<li><p>Detect the orientation of each image and rotate them to align their directions consistently.</p></li>\n<li><p>Match for all image pairs using ALIKED (resize_to=1024, 1280, 1536, 2048) + LightGlue.</p></li>\n<li><p>Build an undirected graph where nodes are images and edge weights are the average number of matches.</p></li>\n<li><p>Cluster the graph using the Louvain method. For each cluster, select the top k * log(n)/(n-1)% image pairs (k is the parameter to adjust the ratio of selection, n is the number of images in the cluster).<br>\nThe Louvain resolution parameter was set to a small value (resolution=0.1) to avoid over-segmentation.</p></li>\n<li><p>From the merged results of all scales (resize_to=1024,1280,1536,2048), select the top k=1.5 pairs.</p></li>\n</ol>\n<p>6-1. Split each image into four parts, and match using ALIKED (resize_to=1024) + LightGlue.</p>\n<p>6-2. Crop regions with high match density from each image and match using ALIKED (resize_to=1024,1280,1536,2048) + LightGlue.</p>\n<h1>What didn't work or were not tried</h1>\n<ul>\n<li>Clustering completely based on the number of matches; scenes such as <code>fbk_vineyard</code>and <code>stairs</code> were difficult to classify even with DBSCAN or graph partitioning.</li>\n<li>Tuning COLMAP parameters didn’t improve the score.</li>\n<li>Detector-free matchers such as OmniGlue or DKM were slower than ALIKED + LightGlue and thus not adopted.</li>\n</ul>\n<h4>In the end</h4>\n<p>Last year, due to a large number of medal sellers and buyers getting banned and the resulting shifts in team population, my medal was downgraded from gold to silver.<br>\nBut this year, I’m very happy to have won a gold medal!</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "3217123",
      "postDate": "06/04/2025 14:56:11",
      "content": "<p>I'd like to express my gratitude to the hosts and Kaggle staffs for organizing such an exciting and practical competition!<br>\nI would also like to thank the competitors who pointed out the metric bugs, and the staff for promptly addressing them.</p>\n<h1>Overview</h1>\n<ul>\n<li><p>To increase the number of matching pairs with consistent orientation, the orientation of all images was detected and then aligned by rotating the images accordingly.</p></li>\n<li><p>As with many top competitors from last year, I primarily used ALIKED for keypoint detection and LightGlue for matching.</p></li>\n<li><p>Instead of using global features, I matched all image pairs and filtered the top k * log(n)/(n-1)% image pairs to extract matches within each scene (cluster) accurately. Here, k is the parameter to adjust the ratio of selection, and n is the number of images in the dataset.</p></li>\n<li><p>To improve mAA scores using a small number of accurate pairs, I adopted two complementary strategies: <strong>locally</strong>, I applied high-density crop matching by cropping regions with densely matched keypoints; <strong>globally</strong>, I applied image 4splits by splitting each image into four parts, which significantly increased the number of matched keypoints across the entire image.</p></li>\n<li><p>To mitigate mAA score fluctuations due to image scale, pair filtering and high-density crop matching were executed at multiple scales(resize_to = 1024, 1280, 1536, 2048) and ensembled into the colmap.</p></li>\n</ul>\n<h1>Pipeline</h1>\n<p>I submitted two notebooks.</p>\n<table>\n<thead>\n<tr>\n<th>Notebook</th>\n<th>CV*</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>① Best CV</td>\n<td>58.64</td>\n<td>46.67</td>\n<td>46.61</td>\n</tr>\n<tr>\n<td>② Best LB</td>\n<td>54.86</td>\n<td>46.98</td>\n<td>45.33</td>\n</tr>\n</tbody>\n</table>\n<p>*CV used all datasets in the train folder.</p>\n<h3>① Best CV</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12213964%2F48c439d9fb58c4ef6dd446b10c1614d4%2F1.png?generation=1749048960243147&amp;alt=media\" alt=\"\"></p>\n<h4>Steps</h4>\n<ol>\n<li><p>Detect the orientation of each image and rotate them to align their directions consistently.(<a href=\"https://github.com/ternaus/check_orientation\" target=\"_blank\">https://github.com/ternaus/check_orientation</a>)</p></li>\n<li><p>Match for all image pairs using ALIKED (resize_to=1024) + LightGlue.</p></li>\n<li><p>Select the top k * log(n)/(n-1)% pairs based on the number of matches. k is the parameter to adjust the ratio of selection, and n is the number of images in the dataset.<br>\nk was experimentally determined to achieve high scores on both CV and LB. The same rationale applies to the subsequent steps.</p></li>\n<li><p>With the filtered image pairs from step 3 (k=7), match using ALIKED (resize_to=1280, 1536, 2048) + LightGlue, and again select the top k * log(n)/(n-1)% pairs.</p></li>\n<li><p>From the merged results of all scales (resize_to=1024,1280,1536,2048), select the top k=1.5 pairs.</p></li>\n</ol>\n<p>6-1. Split each image into four parts, and match using ALIKED (resize_to=1024) + LightGlue.</p>\n<p>6-2. Crop regions with high match density from each image and match using ALIKED (resize_to=1024,1280,1536,2048) + LightGlue.</p>\n<h3>② Best LB (description only)</h3>\n<h4>Steps</h4>\n<ol>\n<li><p>Detect the orientation of each image and rotate them to align their directions consistently.</p></li>\n<li><p>Match for all image pairs using ALIKED (resize_to=1024, 1280, 1536, 2048) + LightGlue.</p></li>\n<li><p>Build an undirected graph where nodes are images and edge weights are the average number of matches.</p></li>\n<li><p>Cluster the graph using the Louvain method. For each cluster, select the top k * log(n)/(n-1)% image pairs (k is the parameter to adjust the ratio of selection, n is the number of images in the cluster).<br>\nThe Louvain resolution parameter was set to a small value (resolution=0.1) to avoid over-segmentation.</p></li>\n<li><p>From the merged results of all scales (resize_to=1024,1280,1536,2048), select the top k=1.5 pairs.</p></li>\n</ol>\n<p>6-1. Split each image into four parts, and match using ALIKED (resize_to=1024) + LightGlue.</p>\n<p>6-2. Crop regions with high match density from each image and match using ALIKED (resize_to=1024,1280,1536,2048) + LightGlue.</p>\n<h1>What didn't work or were not tried</h1>\n<ul>\n<li>Clustering completely based on the number of matches; scenes such as <code>fbk_vineyard</code>and <code>stairs</code> were difficult to classify even with DBSCAN or graph partitioning.</li>\n<li>Tuning COLMAP parameters didn’t improve the score.</li>\n<li>Detector-free matchers such as OmniGlue or DKM were slower than ALIKED + LightGlue and thus not adopted.</li>\n</ul>\n<h4>In the end</h4>\n<p>Last year, due to a large number of medal sellers and buyers getting banned and the resulting shifts in team population, my medal was downgraded from gold to silver.<br>\nBut this year, I’m very happy to have won a gold medal!</p>\n<p>Thank you</p>",
      "rawMarkdown": "I'd like to express my gratitude to the hosts and Kaggle staffs for organizing such an exciting and practical competition!\nI would also like to thank the competitors who pointed out the metric bugs, and the staff for promptly addressing them.\n\n# Overview\n* To increase the number of matching pairs with consistent orientation, the orientation of all images was detected and then aligned by rotating the images accordingly.\n\n* As with many top competitors from last year, I primarily used ALIKED for keypoint detection and LightGlue for matching.\n\n* Instead of using global features, I matched all image pairs and filtered the top k * log(n)/(n-1)% image pairs to extract matches within each scene (cluster) accurately. Here, k is the parameter to adjust the ratio of selection, and n is the number of images in the dataset.\n\n* To improve mAA scores using a small number of accurate pairs, I adopted two complementary strategies: **locally**, I applied high-density crop matching by cropping regions with densely matched keypoints; **globally**, I applied image 4splits by splitting each image into four parts, which significantly increased the number of matched keypoints across the entire image.\n\n* To mitigate mAA score fluctuations due to image scale, pair filtering and high-density crop matching were executed at multiple scales(resize_to = 1024, 1280, 1536, 2048) and ensembled into the colmap.\n\n# Pipeline\nI submitted two notebooks.\n\n| Notebook | CV* | Public LB | Private LB|\n| --- | --- | --- | --- |\n| ① Best CV | 58.64 | 46.67 | 46.61 |\n| ② Best LB | 54.86 | 46.98 | 45.33 |\n\n*CV used all datasets in the train folder.\n\n### ① Best CV\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12213964%2F48c439d9fb58c4ef6dd446b10c1614d4%2F1.png?generation=1749048960243147&alt=media)\n\n#### Steps\n1. Detect the orientation of each image and rotate them to align their directions consistently.(https://github.com/ternaus/check_orientation)\n\n2. Match for all image pairs using ALIKED (resize_to=1024) + LightGlue.\n\n3. Select the top k * log(n)/(n-1)% pairs based on the number of matches. k is the parameter to adjust the ratio of selection, and n is the number of images in the dataset.\nk was experimentally determined to achieve high scores on both CV and LB. The same rationale applies to the subsequent steps.\n\n4. With the filtered image pairs from step 3 (k=7), match using ALIKED (resize_to=1280, 1536, 2048) + LightGlue, and again select the top k * log(n)/(n-1)% pairs.\n\n5. From the merged results of all scales (resize_to=1024,1280,1536,2048), select the top k=1.5 pairs.\n\n6-1. Split each image into four parts, and match using ALIKED (resize_to=1024) + LightGlue.\n\n6-2. Crop regions with high match density from each image and match using ALIKED (resize_to=1024,1280,1536,2048) + LightGlue.\n\n### ② Best LB (description only)\n\n#### Steps\n1. Detect the orientation of each image and rotate them to align their directions consistently.\n\n2. Match for all image pairs using ALIKED (resize_to=1024, 1280, 1536, 2048) + LightGlue.\n\n3. Build an undirected graph where nodes are images and edge weights are the average number of matches.\n\n4. Cluster the graph using the Louvain method. For each cluster, select the top k * log(n)/(n-1)% image pairs (k is the parameter to adjust the ratio of selection, n is the number of images in the cluster).\nThe Louvain resolution parameter was set to a small value (resolution=0.1) to avoid over-segmentation.\n\n5. From the merged results of all scales (resize_to=1024,1280,1536,2048), select the top k=1.5 pairs.\n\n6-1. Split each image into four parts, and match using ALIKED (resize_to=1024) + LightGlue.\n\n6-2. Crop regions with high match density from each image and match using ALIKED (resize_to=1024,1280,1536,2048) + LightGlue.\n\n# What didn't work or were not tried\n* Clustering completely based on the number of matches; scenes such as `fbk_vineyard `and `stairs` were difficult to classify even with DBSCAN or graph partitioning.\n* Tuning COLMAP parameters didn’t improve the score.\n* Detector-free matchers such as OmniGlue or DKM were slower than ALIKED + LightGlue and thus not adopted.\n\n#### In the end\nLast year, due to a large number of medal sellers and buyers getting banned and the resulting shifts in team population, my medal was downgraded from gold to silver.\nBut this year, I’m very happy to have won a gold medal!\n\nThank you",
      "votes": null
    },
    {
      "id": "3217167",
      "postDate": "06/04/2025 15:40:48",
      "content": "<p>Congratulations on your well-deserved gold medal! May I have a question, how did you define this formula: k * log(n)/(n-1)%?<br>\nIs there a theory behind it, or did you achieve it through experiments?</p>",
      "rawMarkdown": "Congratulations on your well-deserved gold medal! May I have a question, how did you define this formula: k * log(n)/(n-1)%?\nIs there a theory behind it, or did you achieve it through experiments?",
      "votes": null
    },
    {
      "id": "3217191",
      "postDate": "06/04/2025 16:26:39",
      "content": "<p>Thank you for your comment.</p>\n<p>There isn’t a particularly detailed theoretical background, but the numerator represents the number of images that a single image can potentially be matched with, which is (n - 1). The denominator is an estimate of how many images a single image can accurately match with. The ratio of these two gives the desired value.<br>\nI took the logarithm of n because the number of image pairs grows as nC2 = O(n²), so we applied a correction to make the relationship linear.</p>",
      "rawMarkdown": "Thank you for your comment.\n\nThere isn’t a particularly detailed theoretical background, but the numerator represents the number of images that a single image can potentially be matched with, which is (n - 1). The denominator is an estimate of how many images a single image can accurately match with. The ratio of these two gives the desired value.\nI took the logarithm of n because the number of image pairs grows as nC2 = O(n²), so we applied a correction to make the relationship linear.",
      "votes": null
    },
    {
      "id": "3217215",
      "postDate": "06/04/2025 17:07:47",
      "content": "<p>Thank you for the explanation, very creative!<br>\nYou did a great comeback against those medal cheaters 💪.</p>",
      "rawMarkdown": "Thank you for the explanation, very creative!\nYou did a great comeback against those medal cheaters 💪.",
      "votes": null
    },
    {
      "id": "3217768",
      "postDate": "06/05/2025 11:45:30",
      "content": "<p>Good creativity. But it doesn't relate to similarity score, may I ask how much will it improve comparing without this method?</p>",
      "rawMarkdown": "Good creativity. But it doesn't relate to similarity score, may I ask how much will it improve comparing without this method?",
      "votes": null
    },
    {
      "id": "3218630",
      "postDate": "06/06/2025 13:15:43",
      "content": "<p>Thank you for your comment.</p>\n<p>First, if I do not perform filtering based on the number of images as represented in this formula, the subsequent multi-scale dense crop matching and four-part image matching will easily exceed the time limit.</p>\n<p>Without using the above filtering method, the following pipeline barely stayed within the time limit, achieving the best score (Private LB = 40.31, Public LB = 38.95):</p>\n<ol>\n<li>Match on all image pairs (resize_to=1024), and keep only those pairs with 50 or more matches.</li>\n<li>Perform dense crop matching (resize_to=1280) + four-part matching (resize_to=1024).</li>\n</ol>\n<p>Thus, making it possible to ensemble multi-scale filtering match and high-density crop match likely contributed significantly to the remaining improvement of the score.</p>",
      "rawMarkdown": "Thank you for your comment.\n\nFirst, if I do not perform filtering based on the number of images as represented in this formula, the subsequent multi-scale dense crop matching and four-part image matching will easily exceed the time limit.\n\nWithout using the above filtering method, the following pipeline barely stayed within the time limit, achieving the best score (Private LB = 40.31, Public LB = 38.95):\n\n1. Match on all image pairs (resize_to=1024), and keep only those pairs with 50 or more matches.\n2. Perform dense crop matching (resize_to=1280) + four-part matching (resize_to=1024).\n\nThus, making it possible to ensemble multi-scale filtering match and high-density crop match likely contributed significantly to the remaining improvement of the score.",
      "votes": null
    },
    {
      "id": "3218656",
      "postDate": "06/06/2025 14:03:40",
      "content": "<p>Congratulations on the gold medal!<br>\nI have a question, how much did the CV/LB scores improve after adding cropping and 4-split matching?</p>\n<p>I actually tried something similar locally, but since there was almost no change in CV,  I ended up dropping it.<br>\nMaybe the difference comes down to resolution settings or how pair filtering was done, but I’m curious how much of a score gain could be expected if these techniques were used effectively.</p>",
      "rawMarkdown": "Congratulations on the gold medal!\nI have a question, how much did the CV/LB scores improve after adding cropping and 4-split matching?\n\nI actually tried something similar locally, but since there was almost no change in CV,  I ended up dropping it.\nMaybe the difference comes down to resolution settings or how pair filtering was done, but I’m curious how much of a score gain could be expected if these techniques were used effectively.",
      "votes": null
    },
    {
      "id": "3218699",
      "postDate": "06/06/2025 15:35:40",
      "content": "<p>Thank you for your comment.</p>\n<p>At first, I couldn’t see any correlation between the CV and LB scores. Given that Public LB data accounted for 50% and that there's usually a strong correlation between Public and Private scores in past competitions, I concentrated on improving my score by relying on the LB without CV.</p>\n<p>Looking at the submissions I can currently check, there are some fluctuations depending on parameters such as \"resize_to\" and \"min_matches\", but the results are roughly as follows:</p>\n<table>\n<thead>\n<tr>\n<th>methold</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>all pair filtering(≧min_matches)* + 4splits image match</td>\n<td>32–35</td>\n</tr>\n<tr>\n<td>all pair filtering(≧min_matches) + High-density crop (resize_to = 1024, 1280, 1536, 2048)</td>\n<td>37(only example)</td>\n</tr>\n<tr>\n<td>Original pair filtering + High-density crop (resize_to = 1024, 1280, 1536, 2048)</td>\n<td>41-44</td>\n</tr>\n<tr>\n<td>Original pair filtering + High-density crop (resize_to = 1024, 1280, 1536, 2048) + 4splits image match</td>\n<td>43-46</td>\n</tr>\n</tbody>\n</table>\n<p>*all pair filtering(≧min_matches) is constant to all datasets.</p>\n<p>Maybe, I think the key factor that pair filtering  focused on a small number of highly reliable pairs and concentrating 4-split image matching and high-density cropping on these pairs,thus I was able to improve the mAA while maintaining a high clustering score.</p>",
      "rawMarkdown": "Thank you for your comment.\n\nAt first, I couldn’t see any correlation between the CV and LB scores. Given that Public LB data accounted for 50% and that there's usually a strong correlation between Public and Private scores in past competitions, I concentrated on improving my score by relying on the LB without CV.\n\nLooking at the submissions I can currently check, there are some fluctuations depending on parameters such as \"resize_to\" and \"min_matches\", but the results are roughly as follows:\n\n| methold | Public LB |\n| --- | --- |\n| all pair filtering(≧min_matches)* + 4splits image match | 32–35 |\n| all pair filtering(≧min_matches) + High-density crop (resize_to = 1024, 1280, 1536, 2048) | 37(only example) |\n| Original pair filtering + High-density crop (resize_to = 1024, 1280, 1536, 2048) | 41-44 |\n| Original pair filtering + High-density crop (resize_to = 1024, 1280, 1536, 2048) + 4splits image match | 43-46 |\n\n*all pair filtering(≧min_matches) is constant to all datasets.\n\nMaybe, I think the key factor that pair filtering  focused on a small number of highly reliable pairs and concentrating 4-split image matching and high-density cropping on these pairs,thus I was able to improve the mAA while maintaining a high clustering score.",
      "votes": null
    },
    {
      "id": "3218940",
      "postDate": "06/07/2025 00:26:33",
      "content": "<p>Thank you for the detailed explanation.<br>\nWhat I tried was also quite close to all-pair filtering, so just as you thought, the key might have been applying cropping and 4-split only to high-confidence pairs.</p>\n<p>It’s a great solution, with nice ideas throughout despite its simplicity.<br>\nOnce again, congratulations.</p>",
      "rawMarkdown": "Thank you for the detailed explanation.\nWhat I tried was also quite close to all-pair filtering, so just as you thought, the key might have been applying cropping and 4-split only to high-confidence pairs.\n\nIt’s a great solution, with nice ideas throughout despite its simplicity.\nOnce again, congratulations.",
      "votes": null
    },
    {
      "id": "3218956",
      "postDate": "06/07/2025 01:14:10",
      "content": "<p>Thanks for the details! </p>",
      "rawMarkdown": "Thanks for the details!",
      "votes": null
    },
    {
      "id": "3218965",
      "postDate": "06/07/2025 01:30:47",
      "content": "<p>Thank you for sharing your excellent solution!<br>\nI'm curious about the parameter k in your filtering approach; how did you determine its optimal value?</p>",
      "rawMarkdown": "Thank you for sharing your excellent solution!\nI'm curious about the parameter k in your filtering approach; how did you determine its optimal value?",
      "votes": null
    },
    {
      "id": "3219339",
      "postDate": "06/07/2025 14:35:01",
      "content": "<p>Thank you for your comment.</p>\n<p>I determined by integrating the following two considerations:</p>\n<ol>\n<li>During performing CV on the entire training dataset, I kept the Clustering Score from downing while improving mAA as high as possible.</li>\n<li>Improve the score on the Public LB.</li>\n</ol>",
      "rawMarkdown": "Thank you for your comment.\n\nI determined by integrating the following two considerations:\n\n1.  During performing CV on the entire training dataset, I kept the Clustering Score from downing while improving mAA as high as possible.\n2.  Improve the score on the Public LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217167,
      "author_name": "nejicool96",
      "author_url": "",
      "post_date": "06/04/2025 15:40:48",
      "content": "<p>Congratulations on your well-deserved gold medal! May I have a question, how did you define this formula: k * log(n)/(n-1)%?<br>\nIs there a theory behind it, or did you achieve it through experiments?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3217191,
          "author_name": "kzkknmt",
          "author_url": "",
          "post_date": "06/04/2025 16:26:39",
          "content": "<p>Thank you for your comment.</p>\n<p>There isn’t a particularly detailed theoretical background, but the numerator represents the number of images that a single image can potentially be matched with, which is (n - 1). The denominator is an estimate of how many images a single image can accurately match with. The ratio of these two gives the desired value.<br>\nI took the logarithm of n because the number of image pairs grows as nC2 = O(n²), so we applied a correction to make the relationship linear.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3217215,
              "author_name": "nejicool96",
              "author_url": "",
              "post_date": "06/04/2025 17:07:47",
              "content": "<p>Thank you for the explanation, very creative!<br>\nYou did a great comeback against those medal cheaters 💪.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3217768,
              "author_name": "kurisew",
              "author_url": "",
              "post_date": "06/05/2025 11:45:30",
              "content": "<p>Good creativity. But it doesn't relate to similarity score, may I ask how much will it improve comparing without this method?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3218630,
                  "author_name": "kzkknmt",
                  "author_url": "",
                  "post_date": "06/06/2025 13:15:43",
                  "content": "<p>Thank you for your comment.</p>\n<p>First, if I do not perform filtering based on the number of images as represented in this formula, the subsequent multi-scale dense crop matching and four-part image matching will easily exceed the time limit.</p>\n<p>Without using the above filtering method, the following pipeline barely stayed within the time limit, achieving the best score (Private LB = 40.31, Public LB = 38.95):</p>\n<ol>\n<li>Match on all image pairs (resize_to=1024), and keep only those pairs with 50 or more matches.</li>\n<li>Perform dense crop matching (resize_to=1280) + four-part matching (resize_to=1024).</li>\n</ol>\n<p>Thus, making it possible to ensemble multi-scale filtering match and high-density crop match likely contributed significantly to the remaining improvement of the score.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3218956,
                      "author_name": "kurisew",
                      "author_url": "",
                      "post_date": "06/07/2025 01:14:10",
                      "content": "<p>Thanks for the details! </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 3218965,
              "author_name": "tyyuki",
              "author_url": "",
              "post_date": "06/07/2025 01:30:47",
              "content": "<p>Thank you for sharing your excellent solution!<br>\nI'm curious about the parameter k in your filtering approach; how did you determine its optimal value?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3219339,
                  "author_name": "kzkknmt",
                  "author_url": "",
                  "post_date": "06/07/2025 14:35:01",
                  "content": "<p>Thank you for your comment.</p>\n<p>I determined by integrating the following two considerations:</p>\n<ol>\n<li>During performing CV on the entire training dataset, I kept the Clustering Score from downing while improving mAA as high as possible.</li>\n<li>Improve the score on the Public LB.</li>\n</ol>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3218656,
      "author_name": "kashiwaba",
      "author_url": "",
      "post_date": "06/06/2025 14:03:40",
      "content": "<p>Congratulations on the gold medal!<br>\nI have a question, how much did the CV/LB scores improve after adding cropping and 4-split matching?</p>\n<p>I actually tried something similar locally, but since there was almost no change in CV,  I ended up dropping it.<br>\nMaybe the difference comes down to resolution settings or how pair filtering was done, but I’m curious how much of a score gain could be expected if these techniques were used effectively.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218699,
          "author_name": "kzkknmt",
          "author_url": "",
          "post_date": "06/06/2025 15:35:40",
          "content": "<p>Thank you for your comment.</p>\n<p>At first, I couldn’t see any correlation between the CV and LB scores. Given that Public LB data accounted for 50% and that there's usually a strong correlation between Public and Private scores in past competitions, I concentrated on improving my score by relying on the LB without CV.</p>\n<p>Looking at the submissions I can currently check, there are some fluctuations depending on parameters such as \"resize_to\" and \"min_matches\", but the results are roughly as follows:</p>\n<table>\n<thead>\n<tr>\n<th>methold</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>all pair filtering(≧min_matches)* + 4splits image match</td>\n<td>32–35</td>\n</tr>\n<tr>\n<td>all pair filtering(≧min_matches) + High-density crop (resize_to = 1024, 1280, 1536, 2048)</td>\n<td>37(only example)</td>\n</tr>\n<tr>\n<td>Original pair filtering + High-density crop (resize_to = 1024, 1280, 1536, 2048)</td>\n<td>41-44</td>\n</tr>\n<tr>\n<td>Original pair filtering + High-density crop (resize_to = 1024, 1280, 1536, 2048) + 4splits image match</td>\n<td>43-46</td>\n</tr>\n</tbody>\n</table>\n<p>*all pair filtering(≧min_matches) is constant to all datasets.</p>\n<p>Maybe, I think the key factor that pair filtering  focused on a small number of highly reliable pairs and concentrating 4-split image matching and high-density cropping on these pairs,thus I was able to improve the mAA while maintaining a high clustering score.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3218940,
              "author_name": "kashiwaba",
              "author_url": "",
              "post_date": "06/07/2025 00:26:33",
              "content": "<p>Thank you for the detailed explanation.<br>\nWhat I tried was also quite close to all-pair filtering, so just as you thought, the key might have been applying cropping and 4-split only to high-confidence pairs.</p>\n<p>It’s a great solution, with nice ideas throughout despite its simplicity.<br>\nOnce again, congratulations.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3217123": "I'd like to express my gratitude to the hosts and Kaggle staffs for organizing such an exciting and practical competition!\nI would also like to thank the competitors who pointed out the metric bugs, and the staff for promptly addressing them.\n\n# Overview\n* To increase the number of matching pairs with consistent orientation, the orientation of all images was detected and then aligned by rotating the images accordingly.\n\n* As with many top competitors from last year, I primarily used ALIKED for keypoint detection and LightGlue for matching.\n\n* Instead of using global features, I matched all image pairs and filtered the top k * log(n)/(n-1)% image pairs to extract matches within each scene (cluster) accurately. Here, k is the parameter to adjust the ratio of selection, and n is the number of images in the dataset.\n\n* To improve mAA scores using a small number of accurate pairs, I adopted two complementary strategies: **locally**, I applied high-density crop matching by cropping regions with densely matched keypoints; **globally**, I applied image 4splits by splitting each image into four parts, which significantly increased the number of matched keypoints across the entire image.\n\n* To mitigate mAA score fluctuations due to image scale, pair filtering and high-density crop matching were executed at multiple scales(resize_to = 1024, 1280, 1536, 2048) and ensembled into the colmap.\n\n# Pipeline\nI submitted two notebooks.\n\n| Notebook | CV* | Public LB | Private LB|\n| --- | --- | --- | --- |\n| ① Best CV | 58.64 | 46.67 | 46.61 |\n| ② Best LB | 54.86 | 46.98 | 45.33 |\n\n*CV used all datasets in the train folder.\n\n### ① Best CV\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12213964%2F48c439d9fb58c4ef6dd446b10c1614d4%2F1.png?generation=1749048960243147&alt=media)\n\n#### Steps\n1. Detect the orientation of each image and rotate them to align their directions consistently.(https://github.com/ternaus/check_orientation)\n\n2. Match for all image pairs using ALIKED (resize_to=1024) + LightGlue.\n\n3. Select the top k * log(n)/(n-1)% pairs based on the number of matches. k is the parameter to adjust the ratio of selection, and n is the number of images in the dataset.\nk was experimentally determined to achieve high scores on both CV and LB. The same rationale applies to the subsequent steps.\n\n4. With the filtered image pairs from step 3 (k=7), match using ALIKED (resize_to=1280, 1536, 2048) + LightGlue, and again select the top k * log(n)/(n-1)% pairs.\n\n5. From the merged results of all scales (resize_to=1024,1280,1536,2048), select the top k=1.5 pairs.\n\n6-1. Split each image into four parts, and match using ALIKED (resize_to=1024) + LightGlue.\n\n6-2. Crop regions with high match density from each image and match using ALIKED (resize_to=1024,1280,1536,2048) + LightGlue.\n\n### ② Best LB (description only)\n\n#### Steps\n1. Detect the orientation of each image and rotate them to align their directions consistently.\n\n2. Match for all image pairs using ALIKED (resize_to=1024, 1280, 1536, 2048) + LightGlue.\n\n3. Build an undirected graph where nodes are images and edge weights are the average number of matches.\n\n4. Cluster the graph using the Louvain method. For each cluster, select the top k * log(n)/(n-1)% image pairs (k is the parameter to adjust the ratio of selection, n is the number of images in the cluster).\nThe Louvain resolution parameter was set to a small value (resolution=0.1) to avoid over-segmentation.\n\n5. From the merged results of all scales (resize_to=1024,1280,1536,2048), select the top k=1.5 pairs.\n\n6-1. Split each image into four parts, and match using ALIKED (resize_to=1024) + LightGlue.\n\n6-2. Crop regions with high match density from each image and match using ALIKED (resize_to=1024,1280,1536,2048) + LightGlue.\n\n# What didn't work or were not tried\n* Clustering completely based on the number of matches; scenes such as `fbk_vineyard `and `stairs` were difficult to classify even with DBSCAN or graph partitioning.\n* Tuning COLMAP parameters didn’t improve the score.\n* Detector-free matchers such as OmniGlue or DKM were slower than ALIKED + LightGlue and thus not adopted.\n\n#### In the end\nLast year, due to a large number of medal sellers and buyers getting banned and the resulting shifts in team population, my medal was downgraded from gold to silver.\nBut this year, I’m very happy to have won a gold medal!\n\nThank you",
    "3217167": "Congratulations on your well-deserved gold medal! May I have a question, how did you define this formula: k * log(n)/(n-1)%?\nIs there a theory behind it, or did you achieve it through experiments?",
    "3217191": "Thank you for your comment.\n\nThere isn’t a particularly detailed theoretical background, but the numerator represents the number of images that a single image can potentially be matched with, which is (n - 1). The denominator is an estimate of how many images a single image can accurately match with. The ratio of these two gives the desired value.\nI took the logarithm of n because the number of image pairs grows as nC2 = O(n²), so we applied a correction to make the relationship linear.",
    "3217215": "Thank you for the explanation, very creative!\nYou did a great comeback against those medal cheaters 💪.",
    "3217768": "Good creativity. But it doesn't relate to similarity score, may I ask how much will it improve comparing without this method?",
    "3218630": "Thank you for your comment.\n\nFirst, if I do not perform filtering based on the number of images as represented in this formula, the subsequent multi-scale dense crop matching and four-part image matching will easily exceed the time limit.\n\nWithout using the above filtering method, the following pipeline barely stayed within the time limit, achieving the best score (Private LB = 40.31, Public LB = 38.95):\n\n1. Match on all image pairs (resize_to=1024), and keep only those pairs with 50 or more matches.\n2. Perform dense crop matching (resize_to=1280) + four-part matching (resize_to=1024).\n\nThus, making it possible to ensemble multi-scale filtering match and high-density crop match likely contributed significantly to the remaining improvement of the score.",
    "3218656": "Congratulations on the gold medal!\nI have a question, how much did the CV/LB scores improve after adding cropping and 4-split matching?\n\nI actually tried something similar locally, but since there was almost no change in CV,  I ended up dropping it.\nMaybe the difference comes down to resolution settings or how pair filtering was done, but I’m curious how much of a score gain could be expected if these techniques were used effectively.",
    "3218699": "Thank you for your comment.\n\nAt first, I couldn’t see any correlation between the CV and LB scores. Given that Public LB data accounted for 50% and that there's usually a strong correlation between Public and Private scores in past competitions, I concentrated on improving my score by relying on the LB without CV.\n\nLooking at the submissions I can currently check, there are some fluctuations depending on parameters such as \"resize_to\" and \"min_matches\", but the results are roughly as follows:\n\n| methold | Public LB |\n| --- | --- |\n| all pair filtering(≧min_matches)* + 4splits image match | 32–35 |\n| all pair filtering(≧min_matches) + High-density crop (resize_to = 1024, 1280, 1536, 2048) | 37(only example) |\n| Original pair filtering + High-density crop (resize_to = 1024, 1280, 1536, 2048) | 41-44 |\n| Original pair filtering + High-density crop (resize_to = 1024, 1280, 1536, 2048) + 4splits image match | 43-46 |\n\n*all pair filtering(≧min_matches) is constant to all datasets.\n\nMaybe, I think the key factor that pair filtering  focused on a small number of highly reliable pairs and concentrating 4-split image matching and high-density cropping on these pairs,thus I was able to improve the mAA while maintaining a high clustering score.",
    "3218940": "Thank you for the detailed explanation.\nWhat I tried was also quite close to all-pair filtering, so just as you thought, the key might have been applying cropping and 4-split only to high-confidence pairs.\n\nIt’s a great solution, with nice ideas throughout despite its simplicity.\nOnce again, congratulations.",
    "3218956": "Thanks for the details!",
    "3218965": "Thank you for sharing your excellent solution!\nI'm curious about the parameter k in your filtering approach; how did you determine its optimal value?",
    "3219339": "Thank you for your comment.\n\nI determined by integrating the following two considerations:\n\n1.  During performing CV on the entire training dataset, I kept the Clustering Score from downing while improving mAA as high as possible.\n2.  Improve the score on the Public LB."
  },
  "source": "meta"
}