{
  "id": 583877,
  "title": "17th Place Solution: Connected Component Clustering ",
  "url": "/competitions/image-matching-challenge-2025/discussion/583877",
  "author_name": "Samrat Thapa",
  "post_date": "2025-06-10T06:33:42.382000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, congratulations to all the winners and participants! It was a fantastic opportunity to dive deep into large-scale image matching across diverse datasets. My approach is quite simple, and differnt from the common DNN feature based clustering approaches. Please leave an upvote if you find it useful.</p>\n<p>I started by building upon the <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">1st place solution (thank you very much to the authors)</a> from the <strong>Image Matching Challenge 2024</strong>. The robust pipeline from the previous year provided a solid baseline, and my goal was to improve scene clustering on top of it.</p>\n<h2>Connected Component Clustering</h2>\n<p>For all datasets (with all scenes, no clustering yet), I ran the first place solution pipeline to obtain precise image pairs. From these, I constructed a <strong>graph</strong>, where nodes represent images and edges represent matching pairs. I then computed <strong>connected components</strong> of this graph. Each connected component was treated as a separate scene, and components with fewer than <strong>5 images</strong> were considered outliers and discarded. Then, I ran reconstruction on each clusters using pycolmap.<br>\nThis basic strategy already gave a <strong>public leaderboard score of 42</strong>, with a <strong>CV score</strong> breakdown as follows:</p>\n<p>CV Mean: 60.79  <br>\nPer-dataset Scores:</p>\n<pre><code>{\n  'imc_haiper' \n  'imc_heritage' \n  'imc_theather_imc_church' \n  'imc_dioscuri_baalshamin' \n  'imc_lizard_pond' \n  'pt_brandenburg_british_buckingham' \n  'pt_piazzasanmarco_grandplace' \n  'pt_sacrecoeur_trevi_tajmahal' \n  'pt_stpeters_stpauls' \n  'amy_gardens' \n  'fbk_vineyard' \n  'ETs' \n  'stairs' \n}\n</code></pre>\n<p>The major drawback of this approach is it is slower than using DNN features for clustering especially as the number of images in the dataset increases. In that case, this approach should be combined with DNN feature-based approach balancing speed-accuracy tradeoff.</p>\n<h2>Refinement – Visual Layout-Based Clustering</h2>\n<p>While analyzing datasets such as <code>pt_brandenburg_british_buckingham</code> and <code>pt_piazzasanmarco_grandplace</code>, I noticed that the <strong>connected component graphs had high average node degree</strong>, and false positive image pairs were causing <strong>inter-scene connections</strong> that standard connected-component clustering couldn't resolve.</p>\n<p>By visualizing the graph using <strong><code>networkx.spring_layout</code></strong>, it became clear that there were <strong>distinct visual clusters</strong> that weren't being captured by simple graph traversal.</p>\n<h3>➕ What I Did:</h3>\n<ul>\n<li>Computed spring layout coordinates of each node (image).</li>\n<li>Clustered these 2D layout coordinates using an unsupervised clustering method (hdbscan).</li>\n<li>Applied this refinement only on datasets with <strong>high average degree</strong>, where naive connected components failed.</li>\n</ul>\n<h3>📈 Result:</h3>\n<p>CV Mean: 66.23  <br>\nPer-dataset Scores:</p>\n<pre><code>{\n  'imc_haiper' \n  'imc_heritage' \n  'imc_theather_imc_church' \n  'imc_dioscuri_baalshamin' \n  'imc_lizard_pond' \n  'pt_brandenburg_british_buckingham' \n  'pt_piazzasanmarco_grandplace' \n  'pt_sacrecoeur_trevi_tajmahal' \n  'pt_stpeters_stpauls' \n  'amy_gardens' \n  'fbk_vineyard' \n  'ETs' \n  'stairs' \n}\n</code></pre>\n<p>This refinement <strong>significantly improved my CV</strong>, especially for public landmark datasets. However, the <strong>public LB score remained unchanged</strong>, likely due to the test set not containing such datasets.</p>\n<h2>What did not work</h2>\n<ul>\n<li>Refining the image pairs using dense feature matchers like <a href=\"https://arxiv.org/abs/2305.15404\" target=\"_blank\">RoMA</a> </li>\n<li><a href=\"https://snap.stanford.edu/node2vec/\" target=\"_blank\">Node2Vec</a> for clustering the graphs.</li>\n<li>Using <a href=\"https://github.com/IDEA-Research/Grounded-Segment-Anything\" target=\"_blank\">GroudingSAM</a> to filter out high-frequency and low-information regions from the amy_gardens dataset. I tried segmenting out classes like <code>sky, ground, grass</code>, but the approach failed to improve the reconstruction results.</li>\n</ul>\n<p>The full notebook is available <a href=\"https://www.kaggle.com/code/samratthapa/imc2025-submission-tuning?scriptVersionId=242256949\" target=\"_blank\">here</a> (link your Kaggle notebook if published)</p>\n<p>```</p>",
  "messages": [
    {
      "id": 3220916,
      "postDate": "2025-06-10T06:33:42.383Z",
      "content": "<p>First of all, congratulations to all the winners and participants! It was a fantastic opportunity to dive deep into large-scale image matching across diverse datasets. My approach is quite simple, and differnt from the common DNN feature based clustering approaches. Please leave an upvote if you find it useful.</p>\n<p>I started by building upon the <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084\" target=\"_blank\">1st place solution (thank you very much to the authors)</a> from the <strong>Image Matching Challenge 2024</strong>. The robust pipeline from the previous year provided a solid baseline, and my goal was to improve scene clustering on top of it.</p>\n<h2>Connected Component Clustering</h2>\n<p>For all datasets (with all scenes, no clustering yet), I ran the first place solution pipeline to obtain precise image pairs. From these, I constructed a <strong>graph</strong>, where nodes represent images and edges represent matching pairs. I then computed <strong>connected components</strong> of this graph. Each connected component was treated as a separate scene, and components with fewer than <strong>5 images</strong> were considered outliers and discarded. Then, I ran reconstruction on each clusters using pycolmap.<br>\nThis basic strategy already gave a <strong>public leaderboard score of 42</strong>, with a <strong>CV score</strong> breakdown as follows:</p>\n<p>CV Mean: 60.79  <br>\nPer-dataset Scores:</p>\n<pre><code>{\n  'imc_haiper' \n  'imc_heritage' \n  'imc_theather_imc_church' \n  'imc_dioscuri_baalshamin' \n  'imc_lizard_pond' \n  'pt_brandenburg_british_buckingham' \n  'pt_piazzasanmarco_grandplace' \n  'pt_sacrecoeur_trevi_tajmahal' \n  'pt_stpeters_stpauls' \n  'amy_gardens' \n  'fbk_vineyard' \n  'ETs' \n  'stairs' \n}\n</code></pre>\n<p>The major drawback of this approach is it is slower than using DNN features for clustering especially as the number of images in the dataset increases. In that case, this approach should be combined with DNN feature-based approach balancing speed-accuracy tradeoff.</p>\n<h2>Refinement – Visual Layout-Based Clustering</h2>\n<p>While analyzing datasets such as <code>pt_brandenburg_british_buckingham</code> and <code>pt_piazzasanmarco_grandplace</code>, I noticed that the <strong>connected component graphs had high average node degree</strong>, and false positive image pairs were causing <strong>inter-scene connections</strong> that standard connected-component clustering couldn't resolve.</p>\n<p>By visualizing the graph using <strong><code>networkx.spring_layout</code></strong>, it became clear that there were <strong>distinct visual clusters</strong> that weren't being captured by simple graph traversal.</p>\n<h3>➕ What I Did:</h3>\n<ul>\n<li>Computed spring layout coordinates of each node (image).</li>\n<li>Clustered these 2D layout coordinates using an unsupervised clustering method (hdbscan).</li>\n<li>Applied this refinement only on datasets with <strong>high average degree</strong>, where naive connected components failed.</li>\n</ul>\n<h3>📈 Result:</h3>\n<p>CV Mean: 66.23  <br>\nPer-dataset Scores:</p>\n<pre><code>{\n  'imc_haiper' \n  'imc_heritage' \n  'imc_theather_imc_church' \n  'imc_dioscuri_baalshamin' \n  'imc_lizard_pond' \n  'pt_brandenburg_british_buckingham' \n  'pt_piazzasanmarco_grandplace' \n  'pt_sacrecoeur_trevi_tajmahal' \n  'pt_stpeters_stpauls' \n  'amy_gardens' \n  'fbk_vineyard' \n  'ETs' \n  'stairs' \n}\n</code></pre>\n<p>This refinement <strong>significantly improved my CV</strong>, especially for public landmark datasets. However, the <strong>public LB score remained unchanged</strong>, likely due to the test set not containing such datasets.</p>\n<h2>What did not work</h2>\n<ul>\n<li>Refining the image pairs using dense feature matchers like <a href=\"https://arxiv.org/abs/2305.15404\" target=\"_blank\">RoMA</a> </li>\n<li><a href=\"https://snap.stanford.edu/node2vec/\" target=\"_blank\">Node2Vec</a> for clustering the graphs.</li>\n<li>Using <a href=\"https://github.com/IDEA-Research/Grounded-Segment-Anything\" target=\"_blank\">GroudingSAM</a> to filter out high-frequency and low-information regions from the amy_gardens dataset. I tried segmenting out classes like <code>sky, ground, grass</code>, but the approach failed to improve the reconstruction results.</li>\n</ul>\n<p>The full notebook is available <a href=\"https://www.kaggle.com/code/samratthapa/imc2025-submission-tuning?scriptVersionId=242256949\" target=\"_blank\">here</a> (link your Kaggle notebook if published)</p>\n<p>```</p>",
      "rawMarkdown": "First of all, congratulations to all the winners and participants! It was a fantastic opportunity to dive deep into large-scale image matching across diverse datasets. My approach is quite simple, and differnt from the common DNN feature based clustering approaches. Please leave an upvote if you find it useful.\n\nI started by building upon the [1st place solution (thank you very much to the authors)](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084) from the **Image Matching Challenge 2024**. The robust pipeline from the previous year provided a solid baseline, and my goal was to improve scene clustering on top of it.\n\n## Connected Component Clustering\n\nFor all datasets (with all scenes, no clustering yet), I ran the first place solution pipeline to obtain precise image pairs. From these, I constructed a **graph**, where nodes represent images and edges represent matching pairs. I then computed **connected components** of this graph. Each connected component was treated as a separate scene, and components with fewer than **5 images** were considered outliers and discarded. Then, I ran reconstruction on each clusters using pycolmap.\nThis basic strategy already gave a **public leaderboard score of 42**, with a **CV score** breakdown as follows:\n\nCV Mean: 60.79  \nPer-dataset Scores:\n```\n{\n  'imc2023_haiper': 67.1,\n  'imc2023_heritage': 93.0,\n  'imc2023_theather_imc2024_church': 67.9,\n  'imc2024_dioscuri_baalshamin': 91.9,\n  'imc2024_lizard_pond': 74.3,\n  'pt_brandenburg_british_buckingham': 46.9,\n  'pt_piazzasanmarco_grandplace': 58.4,\n  'pt_sacrecoeur_trevi_tajmahal': 92.7,\n  'pt_stpeters_stpauls': 61.2,\n  'amy_gardens': 28.9,\n  'fbk_vineyard': 46.8,\n  'ETs': 61.3,\n  'stairs': 0.0\n}\n```\nThe major drawback of this approach is it is slower than using DNN features for clustering especially as the number of images in the dataset increases. In that case, this approach should be combined with DNN feature-based approach balancing speed-accuracy tradeoff.\n\n##  Refinement – Visual Layout-Based Clustering\n\nWhile analyzing datasets such as `pt_brandenburg_british_buckingham` and `pt_piazzasanmarco_grandplace`, I noticed that the **connected component graphs had high average node degree**, and false positive image pairs were causing **inter-scene connections** that standard connected-component clustering couldn't resolve.\n\nBy visualizing the graph using **`networkx.spring_layout`**, it became clear that there were **distinct visual clusters** that weren't being captured by simple graph traversal.\n\n### ➕ What I Did:\n- Computed spring layout coordinates of each node (image).\n- Clustered these 2D layout coordinates using an unsupervised clustering method (hdbscan).\n- Applied this refinement only on datasets with **high average degree**, where naive connected components failed.\n\n### 📈 Result:\n\nCV Mean: 66.23  \nPer-dataset Scores:\n```\n{\n  'imc2023_haiper': 67.1,\n  'imc2023_heritage': 93.0,\n  'imc2023_theather_imc2024_church': 67.6,\n  'imc2024_dioscuri_baalshamin': 91.7,\n  'imc2024_lizard_pond': 74.4,\n  'pt_brandenburg_british_buckingham': 85.6,\n  'pt_piazzasanmarco_grandplace': 60.9,\n  'pt_sacrecoeur_trevi_tajmahal': 92.9,\n  'pt_stpeters_stpauls': 88.3,\n  'amy_gardens': 28.9,\n  'fbk_vineyard': 44.9,\n  'ETs': 61.3,\n  'stairs': 4.3\n}\n```\nThis refinement **significantly improved my CV**, especially for public landmark datasets. However, the **public LB score remained unchanged**, likely due to the test set not containing such datasets.\n\n## What did not work\n\n- Refining the image pairs using dense feature matchers like [RoMA](https://arxiv.org/abs/2305.15404) \n- [Node2Vec](https://snap.stanford.edu/node2vec/) for clustering the graphs.\n- Using [GroudingSAM](https://github.com/IDEA-Research/Grounded-Segment-Anything) to filter out high-frequency and low-information regions from the amy_gardens dataset. I tried segmenting out classes like `sky, ground, grass`, but the approach failed to improve the reconstruction results.\n\nThe full notebook is available [here](https://www.kaggle.com/code/samratthapa/imc2025-submission-tuning?scriptVersionId=242256949) (link your Kaggle notebook if published)\n\n```\n",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3220916": "First of all, congratulations to all the winners and participants! It was a fantastic opportunity to dive deep into large-scale image matching across diverse datasets. My approach is quite simple, and differnt from the common DNN feature based clustering approaches. Please leave an upvote if you find it useful.\n\nI started by building upon the [1st place solution (thank you very much to the authors)](https://www.kaggle.com/competitions/image-matching-challenge-2024/discussion/510084) from the **Image Matching Challenge 2024**. The robust pipeline from the previous year provided a solid baseline, and my goal was to improve scene clustering on top of it.\n\n## Connected Component Clustering\n\nFor all datasets (with all scenes, no clustering yet), I ran the first place solution pipeline to obtain precise image pairs. From these, I constructed a **graph**, where nodes represent images and edges represent matching pairs. I then computed **connected components** of this graph. Each connected component was treated as a separate scene, and components with fewer than **5 images** were considered outliers and discarded. Then, I ran reconstruction on each clusters using pycolmap.\nThis basic strategy already gave a **public leaderboard score of 42**, with a **CV score** breakdown as follows:\n\nCV Mean: 60.79  \nPer-dataset Scores:\n```\n{\n  'imc2023_haiper': 67.1,\n  'imc2023_heritage': 93.0,\n  'imc2023_theather_imc2024_church': 67.9,\n  'imc2024_dioscuri_baalshamin': 91.9,\n  'imc2024_lizard_pond': 74.3,\n  'pt_brandenburg_british_buckingham': 46.9,\n  'pt_piazzasanmarco_grandplace': 58.4,\n  'pt_sacrecoeur_trevi_tajmahal': 92.7,\n  'pt_stpeters_stpauls': 61.2,\n  'amy_gardens': 28.9,\n  'fbk_vineyard': 46.8,\n  'ETs': 61.3,\n  'stairs': 0.0\n}\n```\nThe major drawback of this approach is it is slower than using DNN features for clustering especially as the number of images in the dataset increases. In that case, this approach should be combined with DNN feature-based approach balancing speed-accuracy tradeoff.\n\n##  Refinement – Visual Layout-Based Clustering\n\nWhile analyzing datasets such as `pt_brandenburg_british_buckingham` and `pt_piazzasanmarco_grandplace`, I noticed that the **connected component graphs had high average node degree**, and false positive image pairs were causing **inter-scene connections** that standard connected-component clustering couldn't resolve.\n\nBy visualizing the graph using **`networkx.spring_layout`**, it became clear that there were **distinct visual clusters** that weren't being captured by simple graph traversal.\n\n### ➕ What I Did:\n- Computed spring layout coordinates of each node (image).\n- Clustered these 2D layout coordinates using an unsupervised clustering method (hdbscan).\n- Applied this refinement only on datasets with **high average degree**, where naive connected components failed.\n\n### 📈 Result:\n\nCV Mean: 66.23  \nPer-dataset Scores:\n```\n{\n  'imc2023_haiper': 67.1,\n  'imc2023_heritage': 93.0,\n  'imc2023_theather_imc2024_church': 67.6,\n  'imc2024_dioscuri_baalshamin': 91.7,\n  'imc2024_lizard_pond': 74.4,\n  'pt_brandenburg_british_buckingham': 85.6,\n  'pt_piazzasanmarco_grandplace': 60.9,\n  'pt_sacrecoeur_trevi_tajmahal': 92.9,\n  'pt_stpeters_stpauls': 88.3,\n  'amy_gardens': 28.9,\n  'fbk_vineyard': 44.9,\n  'ETs': 61.3,\n  'stairs': 4.3\n}\n```\nThis refinement **significantly improved my CV**, especially for public landmark datasets. However, the **public LB score remained unchanged**, likely due to the test set not containing such datasets.\n\n## What did not work\n\n- Refining the image pairs using dense feature matchers like [RoMA](https://arxiv.org/abs/2305.15404) \n- [Node2Vec](https://snap.stanford.edu/node2vec/) for clustering the graphs.\n- Using [GroudingSAM](https://github.com/IDEA-Research/Grounded-Segment-Anything) to filter out high-frequency and low-information regions from the amy_gardens dataset. I tried segmenting out classes like `sky, ground, grass`, but the approach failed to improve the reconstruction results.\n\nThe full notebook is available [here](https://www.kaggle.com/code/samratthapa/imc2025-submission-tuning?scriptVersionId=242256949) (link your Kaggle notebook if published)\n\n```\n"
  }
}