{
  "id": 510295,
  "title": "13th Place Solution",
  "url": "/competitions/image-matching-challenge-2024/discussion/510295",
  "author_name": "Jow",
  "post_date": "2024-06-05T15:12:22.896000",
  "votes": 21,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, I would like to express my gratitude to the competition organizers and Kaggle staff. I have learned a lot from this competition.</p>\n<h1>Overview of 13th place solution</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fbbb555507832683f3bb53e995bf5380b%2F2.png?generation=1717599897657757&amp;alt=media\"></p>\n<p>Our basic strategy involves using a rotation-resistant model based on <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\" target=\"_blank\">IMC2023 6th place method</a>, along with a multi-model approach using Alike-LightGlue. Moreover, we made several adjustments depending on the scene.</p>\n<h1>Measures for Transparent Objects</h1>\n<h2>1. Cropping the Center Portion</h2>\n<p>Our teammate <a href=\"https://www.kaggle.com/sugupoko\" target=\"_blank\">@sugupoko</a>, discovered that using only the central part of the image improves accuracy for transparent objects. This approach was based on the assumption that transparent objects do not move from the center. As a result, our local score for the cylinder exceeded 0.2. Additionally, by setting the camera model to single-camera, our cylinder score reached <strong>0.463</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F54a639f705f938781ecde5faa806e1cf%2F3.png?generation=1717600002927540&amp;alt=media\"></p>\n<h2>2. Removing Keypoints with Close Pixel Coordinates</h2>\n<p>We observed that the cylinder overreacted to reflected light and tended to match incorrect keypoints due to minimal image variation. Therefore, after matching, we removed keypoints where the x and y coordinate differences were both within 5 pixels, as these were likely incorrect keypoints.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F4e690c419a844e454009816936fb7c7f%2F4.png?generation=1717600034499906&amp;alt=media\"></p>\n<h1>Measures for Church (symmetries-and-repeats)</h1>\n<p>The challenge with church was distinguishing between the front and back, causing the camera to focus on the front. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F30cc61a03702e1107976131439cd5859%2F5.png?generation=1717600111598243&amp;alt=media\"><br>\nTo capture finer details beyond just the clock, we implemented an approach to detect keypoints by dividing the image into four sections. Ultimately, with a model ensemble and setting the camera model to simple-pinhole, we achieved a score of <strong>0.3561</strong> on the train data.<br>\n(Note: Our final submission used simple-radial for all scenes.)<br>\nHowever, this alone did not resolve the issue of distinguishing between the front and back as mentioned above.</p>\n<h1>Approaches We Tried but Did Not Work</h1>\n<ul>\n<li>I had high expectations for end-to-end matching models. However, due to inference time constraints, I could not include RoMA and OmniGlue in the final submission, and Xfeat did not achieve high scores locally.</li>\n<li>Similar to IMC2023, I implemented an approach to determine and crop the Region of Interest (RoI) based on matched points, but this worsened our scores.</li>\n<li>I tried image correction using CLAHE to deal with dark images, but this also did not work.</li>\n</ul>",
  "messages": [
    {
      "id": 2856878,
      "postDate": "2024-06-05T15:12:22.897Z",
      "content": "<p>First, I would like to express my gratitude to the competition organizers and Kaggle staff. I have learned a lot from this competition.</p>\n<h1>Overview of 13th place solution</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fbbb555507832683f3bb53e995bf5380b%2F2.png?generation=1717599897657757&amp;alt=media\"></p>\n<p>Our basic strategy involves using a rotation-resistant model based on <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\" target=\"_blank\">IMC2023 6th place method</a>, along with a multi-model approach using Alike-LightGlue. Moreover, we made several adjustments depending on the scene.</p>\n<h1>Measures for Transparent Objects</h1>\n<h2>1. Cropping the Center Portion</h2>\n<p>Our teammate <a href=\"https://www.kaggle.com/sugupoko\" target=\"_blank\">@sugupoko</a>, discovered that using only the central part of the image improves accuracy for transparent objects. This approach was based on the assumption that transparent objects do not move from the center. As a result, our local score for the cylinder exceeded 0.2. Additionally, by setting the camera model to single-camera, our cylinder score reached <strong>0.463</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F54a639f705f938781ecde5faa806e1cf%2F3.png?generation=1717600002927540&amp;alt=media\"></p>\n<h2>2. Removing Keypoints with Close Pixel Coordinates</h2>\n<p>We observed that the cylinder overreacted to reflected light and tended to match incorrect keypoints due to minimal image variation. Therefore, after matching, we removed keypoints where the x and y coordinate differences were both within 5 pixels, as these were likely incorrect keypoints.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F4e690c419a844e454009816936fb7c7f%2F4.png?generation=1717600034499906&amp;alt=media\"></p>\n<h1>Measures for Church (symmetries-and-repeats)</h1>\n<p>The challenge with church was distinguishing between the front and back, causing the camera to focus on the front. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F30cc61a03702e1107976131439cd5859%2F5.png?generation=1717600111598243&amp;alt=media\"><br>\nTo capture finer details beyond just the clock, we implemented an approach to detect keypoints by dividing the image into four sections. Ultimately, with a model ensemble and setting the camera model to simple-pinhole, we achieved a score of <strong>0.3561</strong> on the train data.<br>\n(Note: Our final submission used simple-radial for all scenes.)<br>\nHowever, this alone did not resolve the issue of distinguishing between the front and back as mentioned above.</p>\n<h1>Approaches We Tried but Did Not Work</h1>\n<ul>\n<li>I had high expectations for end-to-end matching models. However, due to inference time constraints, I could not include RoMA and OmniGlue in the final submission, and Xfeat did not achieve high scores locally.</li>\n<li>Similar to IMC2023, I implemented an approach to determine and crop the Region of Interest (RoI) based on matched points, but this worsened our scores.</li>\n<li>I tried image correction using CLAHE to deal with dark images, but this also did not work.</li>\n</ul>",
      "rawMarkdown": "First, I would like to express my gratitude to the competition organizers and Kaggle staff. I have learned a lot from this competition.\n\n# Overview of 13th place solution\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fbbb555507832683f3bb53e995bf5380b%2F2.png?generation=1717599897657757&alt=media)\n\nOur basic strategy involves using a rotation-resistant model based on [IMC2023 6th place method](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045), along with a multi-model approach using Alike-LightGlue. Moreover, we made several adjustments depending on the scene.\n\n# Measures for Transparent Objects\n## 1. Cropping the Center Portion\nOur teammate @sugupoko, discovered that using only the central part of the image improves accuracy for transparent objects. This approach was based on the assumption that transparent objects do not move from the center. As a result, our local score for the cylinder exceeded 0.2. Additionally, by setting the camera model to single-camera, our cylinder score reached **0.463**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F54a639f705f938781ecde5faa806e1cf%2F3.png?generation=1717600002927540&alt=media)\n\n## 2. Removing Keypoints with Close Pixel Coordinates\nWe observed that the cylinder overreacted to reflected light and tended to match incorrect keypoints due to minimal image variation. Therefore, after matching, we removed keypoints where the x and y coordinate differences were both within 5 pixels, as these were likely incorrect keypoints.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F4e690c419a844e454009816936fb7c7f%2F4.png?generation=1717600034499906&alt=media)\n\n# Measures for Church (symmetries-and-repeats)\nThe challenge with church was distinguishing between the front and back, causing the camera to focus on the front. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F30cc61a03702e1107976131439cd5859%2F5.png?generation=1717600111598243&alt=media)\nTo capture finer details beyond just the clock, we implemented an approach to detect keypoints by dividing the image into four sections. Ultimately, with a model ensemble and setting the camera model to simple-pinhole, we achieved a score of **0.3561** on the train data.\n(Note: Our final submission used simple-radial for all scenes.)\nHowever, this alone did not resolve the issue of distinguishing between the front and back as mentioned above.\n\n# Approaches We Tried but Did Not Work\n- I had high expectations for end-to-end matching models. However, due to inference time constraints, I could not include RoMA and OmniGlue in the final submission, and Xfeat did not achieve high scores locally.\n- Similar to IMC2023, I implemented an approach to determine and crop the Region of Interest (RoI) based on matched points, but this worsened our scores.\n- I tried image correction using CLAHE to deal with dark images, but this also did not work.\n",
      "votes": 21
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2856878": "First, I would like to express my gratitude to the competition organizers and Kaggle staff. I have learned a lot from this competition.\n\n# Overview of 13th place solution\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2Fbbb555507832683f3bb53e995bf5380b%2F2.png?generation=1717599897657757&alt=media)\n\nOur basic strategy involves using a rotation-resistant model based on [IMC2023 6th place method](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045), along with a multi-model approach using Alike-LightGlue. Moreover, we made several adjustments depending on the scene.\n\n# Measures for Transparent Objects\n## 1. Cropping the Center Portion\nOur teammate @sugupoko, discovered that using only the central part of the image improves accuracy for transparent objects. This approach was based on the assumption that transparent objects do not move from the center. As a result, our local score for the cylinder exceeded 0.2. Additionally, by setting the camera model to single-camera, our cylinder score reached **0.463**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F54a639f705f938781ecde5faa806e1cf%2F3.png?generation=1717600002927540&alt=media)\n\n## 2. Removing Keypoints with Close Pixel Coordinates\nWe observed that the cylinder overreacted to reflected light and tended to match incorrect keypoints due to minimal image variation. Therefore, after matching, we removed keypoints where the x and y coordinate differences were both within 5 pixels, as these were likely incorrect keypoints.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F4e690c419a844e454009816936fb7c7f%2F4.png?generation=1717600034499906&alt=media)\n\n# Measures for Church (symmetries-and-repeats)\nThe challenge with church was distinguishing between the front and back, causing the camera to focus on the front. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640938%2F30cc61a03702e1107976131439cd5859%2F5.png?generation=1717600111598243&alt=media)\nTo capture finer details beyond just the clock, we implemented an approach to detect keypoints by dividing the image into four sections. Ultimately, with a model ensemble and setting the camera model to simple-pinhole, we achieved a score of **0.3561** on the train data.\n(Note: Our final submission used simple-radial for all scenes.)\nHowever, this alone did not resolve the issue of distinguishing between the front and back as mentioned above.\n\n# Approaches We Tried but Did Not Work\n- I had high expectations for end-to-end matching models. However, due to inference time constraints, I could not include RoMA and OmniGlue in the final submission, and Xfeat did not achieve high scores locally.\n- Similar to IMC2023, I implemented an approach to determine and crop the Region of Interest (RoI) based on matched points, but this worsened our scores.\n- I tried image correction using CLAHE to deal with dark images, but this also did not work.\n"
  }
}