{
  "id": 582844,
  "title": "8th Place Solution",
  "url": "/competitions/image-matching-challenge-2025/discussion/582844",
  "author_name": "yangyefd",
  "post_date": "2025-06-03T07:05:14.014000",
  "votes": 32,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thank you very much for organizing this wonderful competition. I sincerely thank the organizers and contestants. At the same time, I want to say that copilot helped me a lot, and our code relies heavily on copilot.</p>\n<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/overview\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2025/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/data\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2025/data</a></p>\n<h1>Our main optimization points are as follows:</h1>\n<ol>\n<li>Use gimlightglue: PB score from 32.17=&gt;33.75, we found that gimlightglue produces more accurate matching pairs, but fewer matching pairs.</li>\n<li>Use CLIP to replace DINO for screening: 33.75 =&gt;36.98, CLIP performs exceptionally well in the training set, we use 0.76 cosine similarity, CLIP produces perfect segmentation in most scenes, but there is confusion between scenes in difficult scenes such as stairs.</li>\n<li>Use secondary matching, gimlightglue matches and then uses alike_lightglue (baseline) to match the matching area again: PB score increases to 41.7. This involves matching pair filtering. We first perform dbscan clustering on the matching points separately, then merge the clustering results of the two images, and finally select the points near the cluster centers of the two images for secondary matching. If clustering fails during this process, the matching pair is directly abandoned.</li>\n<li>Use loop checking to filter matching pairs that can form a loop, filter out matching pairs with a loop error greater than 30, and remove the matching pairs with the least number of matches for matching pairs with an average loop error greater than 30.</li>\n<li>Limit the number of matching pairs (take the top 1500). We found that this not only improves performance but also increases matching efficiency.</li>\n<li>Initial matching uses gimlightglue and alike_lightglue for integration, and secondary matching uses alike_lightglue.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22204153%2F852c7c90f39f42cab9ef5ec0692f0b30%2Fyuque_diagram.jpg?generation=1749010125569432&amp;alt=media\" alt=\"\"></li>\n</ol>\n<h1>Useless attempts</h1>\n<ol>\n<li>TTT: Fine-tune during testing. In order to further enhance the matching ability of Lightglue, we tried to fine-tune the model with the test data set during testing. We used the image and the transformation of the image itself for self-supervised learning (rotation of the same image, projection transformation, etc.). We tried this solution for half a month. We found that it can significantly increase the number of matching pairs in the stair scene, but too strong matching leads to more matching between scenes (so we also spent a lot of time studying the design of the classifier), and the PB score dropped after use.</li>\n<li>Classifier training: We believe that classification is the key to this image matching competition. We can use CLIP to produce perfect segmentation in most scenes, but there is some confusion in the stair scene. There are some errors that are easy to distinguish with the naked eye, but Lightglue will match the two. Therefore, we tried CNN image classification model and logistic regression model for classification respectively. We found that the classifier can work well in the training set, which can improve about two points, but PB did not improve. This solution also took a lot of time.</li>\n<li>Yolo mask: We found that some building scenes in the training set contained a large number of people. We thought of using masks to remove them to reduce the interference of irrelevant scenes, but experiments found that masks would produce incorrect masks for ETs and human sculptures of some buildings. Finally, we added color checks to avoid such false masks as much as possible, but PB did not improve at all.</li>\n<li>Multi-resolution, image rotation angle correction: We tried the multi-resolution and rotation correction commonly used in previous solutions, but the PB score did not improve. It may be that there is a problem with the method we used.</li>\n</ol>\n<h1>Some additional information</h1>\n<h2>1. CLIP vs DINOv2</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22204153%2Fbcd9b27a7ed4964a613cc63b722fde98%2FCLIP%20vs%20DINO.png?generation=1749087797727038&amp;alt=media\" alt=\"\"><br>\nIt is easy to see that CLIP produces a more accurate segmentation than DINO</p>\n<h1>The location of our code</h1>\n<p><a href=\"https://github.com/yangyefd/IMC2025\" target=\"_blank\">https://github.com/yangyefd/IMC2025</a></p>\n<h1>Links</h1>\n<p><a href=\"https://github.com/xuelunshen/gim\" target=\"_blank\">https://github.com/xuelunshen/gim</a><br>\n<a href=\"https://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue\" target=\"_blank\">https://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue</a></p>",
  "messages": [
    {
      "id": 3216145,
      "postDate": "2025-06-03T07:05:14.013Z",
      "content": "<p>Thank you very much for organizing this wonderful competition. I sincerely thank the organizers and contestants. At the same time, I want to say that copilot helped me a lot, and our code relies heavily on copilot.</p>\n<h1>Context</h1>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/overview\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2025/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2025/data\" target=\"_blank\">https://www.kaggle.com/competitions/image-matching-challenge-2025/data</a></p>\n<h1>Our main optimization points are as follows:</h1>\n<ol>\n<li>Use gimlightglue: PB score from 32.17=&gt;33.75, we found that gimlightglue produces more accurate matching pairs, but fewer matching pairs.</li>\n<li>Use CLIP to replace DINO for screening: 33.75 =&gt;36.98, CLIP performs exceptionally well in the training set, we use 0.76 cosine similarity, CLIP produces perfect segmentation in most scenes, but there is confusion between scenes in difficult scenes such as stairs.</li>\n<li>Use secondary matching, gimlightglue matches and then uses alike_lightglue (baseline) to match the matching area again: PB score increases to 41.7. This involves matching pair filtering. We first perform dbscan clustering on the matching points separately, then merge the clustering results of the two images, and finally select the points near the cluster centers of the two images for secondary matching. If clustering fails during this process, the matching pair is directly abandoned.</li>\n<li>Use loop checking to filter matching pairs that can form a loop, filter out matching pairs with a loop error greater than 30, and remove the matching pairs with the least number of matches for matching pairs with an average loop error greater than 30.</li>\n<li>Limit the number of matching pairs (take the top 1500). We found that this not only improves performance but also increases matching efficiency.</li>\n<li>Initial matching uses gimlightglue and alike_lightglue for integration, and secondary matching uses alike_lightglue.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22204153%2F852c7c90f39f42cab9ef5ec0692f0b30%2Fyuque_diagram.jpg?generation=1749010125569432&amp;alt=media\" alt=\"\"></li>\n</ol>\n<h1>Useless attempts</h1>\n<ol>\n<li>TTT: Fine-tune during testing. In order to further enhance the matching ability of Lightglue, we tried to fine-tune the model with the test data set during testing. We used the image and the transformation of the image itself for self-supervised learning (rotation of the same image, projection transformation, etc.). We tried this solution for half a month. We found that it can significantly increase the number of matching pairs in the stair scene, but too strong matching leads to more matching between scenes (so we also spent a lot of time studying the design of the classifier), and the PB score dropped after use.</li>\n<li>Classifier training: We believe that classification is the key to this image matching competition. We can use CLIP to produce perfect segmentation in most scenes, but there is some confusion in the stair scene. There are some errors that are easy to distinguish with the naked eye, but Lightglue will match the two. Therefore, we tried CNN image classification model and logistic regression model for classification respectively. We found that the classifier can work well in the training set, which can improve about two points, but PB did not improve. This solution also took a lot of time.</li>\n<li>Yolo mask: We found that some building scenes in the training set contained a large number of people. We thought of using masks to remove them to reduce the interference of irrelevant scenes, but experiments found that masks would produce incorrect masks for ETs and human sculptures of some buildings. Finally, we added color checks to avoid such false masks as much as possible, but PB did not improve at all.</li>\n<li>Multi-resolution, image rotation angle correction: We tried the multi-resolution and rotation correction commonly used in previous solutions, but the PB score did not improve. It may be that there is a problem with the method we used.</li>\n</ol>\n<h1>Some additional information</h1>\n<h2>1. CLIP vs DINOv2</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22204153%2Fbcd9b27a7ed4964a613cc63b722fde98%2FCLIP%20vs%20DINO.png?generation=1749087797727038&amp;alt=media\" alt=\"\"><br>\nIt is easy to see that CLIP produces a more accurate segmentation than DINO</p>\n<h1>The location of our code</h1>\n<p><a href=\"https://github.com/yangyefd/IMC2025\" target=\"_blank\">https://github.com/yangyefd/IMC2025</a></p>\n<h1>Links</h1>\n<p><a href=\"https://github.com/xuelunshen/gim\" target=\"_blank\">https://github.com/xuelunshen/gim</a><br>\n<a href=\"https://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue\" target=\"_blank\">https://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue</a></p>",
      "rawMarkdown": "Thank you very much for organizing this wonderful competition. I sincerely thank the organizers and contestants. At the same time, I want to say that copilot helped me a lot, and our code relies heavily on copilot.\n\n# Context\nBusiness context: https://www.kaggle.com/competitions/image-matching-challenge-2025/overview\nData context: https://www.kaggle.com/competitions/image-matching-challenge-2025/data\n\n# Our main optimization points are as follows:\n1.  Use gimlightglue: PB score from 32.17=>33.75, we found that gimlightglue produces more accurate matching pairs, but fewer matching pairs.\n2.  Use CLIP to replace DINO for screening: 33.75 =>36.98, CLIP performs exceptionally well in the training set, we use 0.76 cosine similarity, CLIP produces perfect segmentation in most scenes, but there is confusion between scenes in difficult scenes such as stairs.\n3.  Use secondary matching, gimlightglue matches and then uses alike_lightglue (baseline) to match the matching area again: PB score increases to 41.7. This involves matching pair filtering. We first perform dbscan clustering on the matching points separately, then merge the clustering results of the two images, and finally select the points near the cluster centers of the two images for secondary matching. If clustering fails during this process, the matching pair is directly abandoned.\n4.  Use loop checking to filter matching pairs that can form a loop, filter out matching pairs with a loop error greater than 30, and remove the matching pairs with the least number of matches for matching pairs with an average loop error greater than 30.\n5.  Limit the number of matching pairs (take the top 1500). We found that this not only improves performance but also increases matching efficiency.\n6.  Initial matching uses gimlightglue and alike_lightglue for integration, and secondary matching uses alike_lightglue.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22204153%2F852c7c90f39f42cab9ef5ec0692f0b30%2Fyuque_diagram.jpg?generation=1749010125569432&alt=media)\n\n# Useless attempts\n1.  TTT: Fine-tune during testing. In order to further enhance the matching ability of Lightglue, we tried to fine-tune the model with the test data set during testing. We used the image and the transformation of the image itself for self-supervised learning (rotation of the same image, projection transformation, etc.). We tried this solution for half a month. We found that it can significantly increase the number of matching pairs in the stair scene, but too strong matching leads to more matching between scenes (so we also spent a lot of time studying the design of the classifier), and the PB score dropped after use.\n2.  Classifier training: We believe that classification is the key to this image matching competition. We can use CLIP to produce perfect segmentation in most scenes, but there is some confusion in the stair scene. There are some errors that are easy to distinguish with the naked eye, but Lightglue will match the two. Therefore, we tried CNN image classification model and logistic regression model for classification respectively. We found that the classifier can work well in the training set, which can improve about two points, but PB did not improve. This solution also took a lot of time.\n3. Yolo mask: We found that some building scenes in the training set contained a large number of people. We thought of using masks to remove them to reduce the interference of irrelevant scenes, but experiments found that masks would produce incorrect masks for ETs and human sculptures of some buildings. Finally, we added color checks to avoid such false masks as much as possible, but PB did not improve at all.\n4.  Multi-resolution, image rotation angle correction: We tried the multi-resolution and rotation correction commonly used in previous solutions, but the PB score did not improve. It may be that there is a problem with the method we used.\n\n# Some additional information\n## 1. CLIP vs DINOv2\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22204153%2Fbcd9b27a7ed4964a613cc63b722fde98%2FCLIP%20vs%20DINO.png?generation=1749087797727038&alt=media)\nIt is easy to see that CLIP produces a more accurate segmentation than DINO\n\n# The location of our code\nhttps://github.com/yangyefd/IMC2025\n\n# Links\nhttps://github.com/xuelunshen/gim\nhttps://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue",
      "votes": 32
    },
    {
      "id": 3216173,
      "postDate": "2025-06-03T07:39:10.833Z",
      "content": "<p>In the ETs dataset, most top-performing teams are able to achieve a 100% MAA score. I would like to ask whether your offline validation process is stable. Have you observed any specific techniques or strategies that consistently lead to achieving both 100% MAA and 100% clusterness scores?</p>\n<p>For reference, my current model only reaches approximately 50% MAA on local validation using the ETs dataset, and I have not yet found an effective way to improve it.</p>\n<p>If possible, could you kindly share your approach to improving offline cross-validation (CV) performance, for example on the ETs dataset? I would greatly appreciate any insights into how model validation strategies can be enhanced in this context. Thank you in advance.</p>",
      "rawMarkdown": "In the ETs dataset, most top-performing teams are able to achieve a 100% MAA score. I would like to ask whether your offline validation process is stable. Have you observed any specific techniques or strategies that consistently lead to achieving both 100% MAA and 100% clusterness scores?\n\nFor reference, my current model only reaches approximately 50% MAA on local validation using the ETs dataset, and I have not yet found an effective way to improve it.\n\nIf possible, could you kindly share your approach to improving offline cross-validation (CV) performance, for example on the ETs dataset? I would greatly appreciate any insights into how model validation strategies can be enhanced in this context. Thank you in advance.",
      "votes": 2,
      "replies": [
        {
          "id": 3216177,
          "postDate": "2025-06-03T07:50:17Z",
          "content": "<p>In fact, our score on ETs is also around 50. I thought it was a problem with the accuracy of the points, so I tried lofter and dkm, but neither worked.</p>",
          "rawMarkdown": "In fact, our score on ETs is also around 50. I thought it was a problem with the accuracy of the points, so I tried lofter and dkm, but neither worked."
        },
        {
          "id": 3217555,
          "postDate": "2025-06-05T06:13:24.740Z",
          "content": "<p>During this competition, there was a metric bug for a while, which might have affected the results. In fact, when exploiting the bug, the mAA of ETs could exceed 100%. However, the actual mAA is around 50%.</p>",
          "rawMarkdown": "During this competition, there was a metric bug for a while, which might have affected the results. In fact, when exploiting the bug, the mAA of ETs could exceed 100%. However, the actual mAA is around 50%."
        }
      ]
    },
    {
      "id": 3217265,
      "postDate": "2025-06-04T18:38:12.470Z",
      "content": "<p>Congrats! And thanks for sharing your solution so early — I had a look at your GitHub code and it's super clean and well-organized. I’ve been meaning to learn from it, so I ported it over to a Kaggle notebook and got it running — funny enough, the score went up a tiny bit, which just shows how solid your pipeline is!</p>\n<p>Quick question — how did you come up with the 0.76 threshold? Did you tune it on the whole training set? Were you going for best accuracy or something else?</p>\n<p>Also, would it be okay if I made the notebook public? I’ve linked your post and credited you clearly. Totally fine if not — just wanted to ask first. Thanks again!</p>",
      "rawMarkdown": "Congrats! And thanks for sharing your solution so early — I had a look at your GitHub code and it's super clean and well-organized. I’ve been meaning to learn from it, so I ported it over to a Kaggle notebook and got it running — funny enough, the score went up a tiny bit, which just shows how solid your pipeline is!\n\nQuick question — how did you come up with the 0.76 threshold? Did you tune it on the whole training set? Were you going for best accuracy or something else?\n\nAlso, would it be okay if I made the notebook public? I’ve linked your post and credited you clearly. Totally fine if not — just wanted to ask first. Thanks again!",
      "replies": [
        {
          "id": 3217398,
          "postDate": "2025-06-05T01:40:28.720Z",
          "content": "<p>Thank you for your appreciation. I used CLIP to calculate the N*N similarity matrix for the images in each folder of the training set. It is easy to find that a threshold of about 0.75 can well separate the scenes. Then I selected some thresholds around 0.75 and conducted experiments and found that 0.76 would be better.<br>\nThe code can be shared freely as long as the source is added.</p>",
          "rawMarkdown": "Thank you for your appreciation. I used CLIP to calculate the N*N similarity matrix for the images in each folder of the training set. It is easy to find that a threshold of about 0.75 can well separate the scenes. Then I selected some thresholds around 0.75 and conducted experiments and found that 0.76 would be better.\nThe code can be shared freely as long as the source is added.",
          "votes": 1,
          "replies": [
            {
              "id": 3218002,
              "postDate": "2025-06-05T16:45:49.763Z",
              "content": "<p>Okay, thanks for the clarification! Smart move!</p>",
              "rawMarkdown": "Okay, thanks for the clarification! Smart move!"
            }
          ]
        },
        {
          "id": 3217402,
          "postDate": "2025-06-05T01:47:30.697Z",
          "content": "<p>I put a comparison chart of CLIP and DINO above for your reference. By the way, I gave the above picture to LLM and asked him about the best threshold. LLM told me that 0.75 is more suitable, haha</p>",
          "rawMarkdown": "I put a comparison chart of CLIP and DINO above for your reference. By the way, I gave the above picture to LLM and asked him about the best threshold. LLM told me that 0.75 is more suitable, haha",
          "replies": [
            {
              "id": 3218005,
              "postDate": "2025-06-05T16:46:56.407Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3218006,
              "postDate": "2025-06-05T16:47:36.373Z",
              "content": "<p>That’s really clear now, thanks a lot! LLM could do that? Amazing!</p>",
              "rawMarkdown": "That’s really clear now, thanks a lot! LLM could do that? Amazing!"
            },
            {
              "id": 3218008,
              "postDate": "2025-06-05T16:50:02.460Z",
              "content": "<p>I test it, gemini could do that for real! Impressive!</p>",
              "rawMarkdown": "I test it, gemini could do that for real! Impressive!"
            }
          ]
        }
      ]
    },
    {
      "id": 3217260,
      "postDate": "2025-06-04T18:17:04.357Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3217259,
      "postDate": "2025-06-04T18:16:12.317Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3216173,
      "author_name": "shanzhong8",
      "author_url": "",
      "post_date": "2025-06-03T07:39:10.833000",
      "content": "<p>In the ETs dataset, most top-performing teams are able to achieve a 100% MAA score. I would like to ask whether your offline validation process is stable. Have you observed any specific techniques or strategies that consistently lead to achieving both 100% MAA and 100% clusterness scores?</p>\n<p>For reference, my current model only reaches approximately 50% MAA on local validation using the ETs dataset, and I have not yet found an effective way to improve it.</p>\n<p>If possible, could you kindly share your approach to improving offline cross-validation (CV) performance, for example on the ETs dataset? I would greatly appreciate any insights into how model validation strategies can be enhanced in this context. Thank you in advance.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3216177,
          "author_name": "yangyefd",
          "author_url": "",
          "post_date": "2025-06-03T07:50:17",
          "content": "<p>In fact, our score on ETs is also around 50. I thought it was a problem with the accuracy of the points, so I tried lofter and dkm, but neither worked.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3217555,
          "author_name": "HayatoFujihara",
          "author_url": "",
          "post_date": "2025-06-05T06:13:24.740000",
          "content": "<p>During this competition, there was a metric bug for a while, which might have affected the results. In fact, when exploiting the bug, the mAA of ETs could exceed 100%. However, the actual mAA is around 50%.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3217265,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2025-06-04T18:38:12.470000",
      "content": "<p>Congrats! And thanks for sharing your solution so early — I had a look at your GitHub code and it's super clean and well-organized. I’ve been meaning to learn from it, so I ported it over to a Kaggle notebook and got it running — funny enough, the score went up a tiny bit, which just shows how solid your pipeline is!</p>\n<p>Quick question — how did you come up with the 0.76 threshold? Did you tune it on the whole training set? Were you going for best accuracy or something else?</p>\n<p>Also, would it be okay if I made the notebook public? I’ve linked your post and credited you clearly. Totally fine if not — just wanted to ask first. Thanks again!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3217398,
          "author_name": "yangyefd",
          "author_url": "",
          "post_date": "2025-06-05T01:40:28.720000",
          "content": "<p>Thank you for your appreciation. I used CLIP to calculate the N*N similarity matrix for the images in each folder of the training set. It is easy to find that a threshold of about 0.75 can well separate the scenes. Then I selected some thresholds around 0.75 and conducted experiments and found that 0.76 would be better.<br>\nThe code can be shared freely as long as the source is added.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3218002,
              "author_name": "Helen",
              "author_url": "",
              "post_date": "2025-06-05T16:45:49.763000",
              "content": "<p>Okay, thanks for the clarification! Smart move!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3217402,
          "author_name": "yangyefd",
          "author_url": "",
          "post_date": "2025-06-05T01:47:30.697000",
          "content": "<p>I put a comparison chart of CLIP and DINO above for your reference. By the way, I gave the above picture to LLM and asked him about the best threshold. LLM told me that 0.75 is more suitable, haha</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3218005,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-06-05T16:46:56.407000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3218006,
              "author_name": "Helen",
              "author_url": "",
              "post_date": "2025-06-05T16:47:36.373000",
              "content": "<p>That’s really clear now, thanks a lot! LLM could do that? Amazing!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3218008,
              "author_name": "Helen",
              "author_url": "",
              "post_date": "2025-06-05T16:50:02.460000",
              "content": "<p>I test it, gemini could do that for real! Impressive!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3217260,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-04T18:17:04.357000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3217259,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-04T18:16:12.317000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3216145": "Thank you very much for organizing this wonderful competition. I sincerely thank the organizers and contestants. At the same time, I want to say that copilot helped me a lot, and our code relies heavily on copilot.\n\n# Context\nBusiness context: https://www.kaggle.com/competitions/image-matching-challenge-2025/overview\nData context: https://www.kaggle.com/competitions/image-matching-challenge-2025/data\n\n# Our main optimization points are as follows:\n1.  Use gimlightglue: PB score from 32.17=>33.75, we found that gimlightglue produces more accurate matching pairs, but fewer matching pairs.\n2.  Use CLIP to replace DINO for screening: 33.75 =>36.98, CLIP performs exceptionally well in the training set, we use 0.76 cosine similarity, CLIP produces perfect segmentation in most scenes, but there is confusion between scenes in difficult scenes such as stairs.\n3.  Use secondary matching, gimlightglue matches and then uses alike_lightglue (baseline) to match the matching area again: PB score increases to 41.7. This involves matching pair filtering. We first perform dbscan clustering on the matching points separately, then merge the clustering results of the two images, and finally select the points near the cluster centers of the two images for secondary matching. If clustering fails during this process, the matching pair is directly abandoned.\n4.  Use loop checking to filter matching pairs that can form a loop, filter out matching pairs with a loop error greater than 30, and remove the matching pairs with the least number of matches for matching pairs with an average loop error greater than 30.\n5.  Limit the number of matching pairs (take the top 1500). We found that this not only improves performance but also increases matching efficiency.\n6.  Initial matching uses gimlightglue and alike_lightglue for integration, and secondary matching uses alike_lightglue.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22204153%2F852c7c90f39f42cab9ef5ec0692f0b30%2Fyuque_diagram.jpg?generation=1749010125569432&alt=media)\n\n# Useless attempts\n1.  TTT: Fine-tune during testing. In order to further enhance the matching ability of Lightglue, we tried to fine-tune the model with the test data set during testing. We used the image and the transformation of the image itself for self-supervised learning (rotation of the same image, projection transformation, etc.). We tried this solution for half a month. We found that it can significantly increase the number of matching pairs in the stair scene, but too strong matching leads to more matching between scenes (so we also spent a lot of time studying the design of the classifier), and the PB score dropped after use.\n2.  Classifier training: We believe that classification is the key to this image matching competition. We can use CLIP to produce perfect segmentation in most scenes, but there is some confusion in the stair scene. There are some errors that are easy to distinguish with the naked eye, but Lightglue will match the two. Therefore, we tried CNN image classification model and logistic regression model for classification respectively. We found that the classifier can work well in the training set, which can improve about two points, but PB did not improve. This solution also took a lot of time.\n3. Yolo mask: We found that some building scenes in the training set contained a large number of people. We thought of using masks to remove them to reduce the interference of irrelevant scenes, but experiments found that masks would produce incorrect masks for ETs and human sculptures of some buildings. Finally, we added color checks to avoid such false masks as much as possible, but PB did not improve at all.\n4.  Multi-resolution, image rotation angle correction: We tried the multi-resolution and rotation correction commonly used in previous solutions, but the PB score did not improve. It may be that there is a problem with the method we used.\n\n# Some additional information\n## 1. CLIP vs DINOv2\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22204153%2Fbcd9b27a7ed4964a613cc63b722fde98%2FCLIP%20vs%20DINO.png?generation=1749087797727038&alt=media)\nIt is easy to see that CLIP produces a more accurate segmentation than DINO\n\n# The location of our code\nhttps://github.com/yangyefd/IMC2025\n\n# Links\nhttps://github.com/xuelunshen/gim\nhttps://www.kaggle.com/code/octaviograu/baseline-dinov2-aliked-lightglue",
    "3216173": "In the ETs dataset, most top-performing teams are able to achieve a 100% MAA score. I would like to ask whether your offline validation process is stable. Have you observed any specific techniques or strategies that consistently lead to achieving both 100% MAA and 100% clusterness scores?\n\nFor reference, my current model only reaches approximately 50% MAA on local validation using the ETs dataset, and I have not yet found an effective way to improve it.\n\nIf possible, could you kindly share your approach to improving offline cross-validation (CV) performance, for example on the ETs dataset? I would greatly appreciate any insights into how model validation strategies can be enhanced in this context. Thank you in advance.",
    "3217265": "Congrats! And thanks for sharing your solution so early — I had a look at your GitHub code and it's super clean and well-organized. I’ve been meaning to learn from it, so I ported it over to a Kaggle notebook and got it running — funny enough, the score went up a tiny bit, which just shows how solid your pipeline is!\n\nQuick question — how did you come up with the 0.76 threshold? Did you tune it on the whole training set? Were you going for best accuracy or something else?\n\nAlso, would it be okay if I made the notebook public? I’ve linked your post and credited you clearly. Totally fine if not — just wanted to ask first. Thanks again!",
    "3217260": "",
    "3217259": ""
  }
}