{
  "id": 510084,
  "title": "1st Place Solution – High Image Resolution ALIKED/LightGlue + Transparent Trick [Prize Eligible]",
  "url": "/competitions/image-matching-challenge-2024/discussion/510084",
  "author_name": "Igor Lashkov",
  "post_date": "2024-06-04T21:48:16.710000",
  "votes": 96,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Our team would like to deeply appreciate the Kaggle team, Czech Technical University in Prague, and other people helping to carry out the series of exciting image matching challenges, as well as everyone here who competes. My congratulations to all the participants!</p>\n<p>The work we describe here is truly a joint effort of <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>, <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a>, <a href=\"https://www.kaggle.com/jaafarmahmoud1\" target=\"_blank\">@jaafarmahmoud1</a>, <a href=\"https://www.kaggle.com/kashiwaba\" target=\"_blank\">@kashiwaba</a>, and <a href=\"https://www.kaggle.com/igorlashkov\" target=\"_blank\">@igorlashkov</a>. I’m grateful for being a part of this hardworking, cohesive, and skilled team.</p>\n<p>As the final submissions we selected:<br>\n•    <strong>best public LB notebook</strong> which scored <strong>0.28</strong><br>\n•    submission with the <strong>best local CV</strong> with <strong>0.24</strong> in LB</p>\n<h1>Overview</h1>\n<p>Our final solution consists of the 3D image reconstruction (I3DR) module powered by COLMAP for non-transparent scenes and a simplified direct image pose estimation (DIP) module for transparent scenes (category “transparent” in “categories.csv”). We stay only with sparse detectors and matchers by employing an ensemble of ALIKED extractors and LightGlue matchers on the high-resolution image pairs. As for the DIP module, we estimate a 3D pose of every image in the scene by placing images in the correct order of object rotation and computing a rotation matrix and a translation vector.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F98595408fef16e793d7a77b247dee746%2FReference%20model%20pipeline.png?generation=1717536474163029&amp;alt=media\" alt=\"architecture\"></p>\n<h1>CV</h1>\n<p>Just like in many other Kaggle competitions, it is important to have a good cross-validation pipeline to improve the LB score consistently. We used transparent and non-transparent scenes for the validation separately. Only images having a GT were included in the test set. As for the validation of the I3DR module, it was hard to improve a pipeline step by step as the correlation with LB was uncertain. It is important to limit COLMAP by 1 thread. On the contrary, the algorithm operating on transparent scenes was easier to debug. In this case, CV correlated better with LB, but not in every case.</p>\n<h1>1. I3DR Module [non-transparent scenes]</h1>\n<p>For the image pair selection, we decided not to use any image retrieval method. Instead, we rely only on the number of matches generated by ALIKED and LightGlue (LG). Specifically, we apply a threshold of 30 matches for a single detector and a value of 100 for matches produced by all detectors on an entire image in the ensemble. It is easy to notice that some scenes (e.g., dioscuri) have images not in the natural orientation and need to be re-rotated to find more matches. Similar to previous IMCs, we rely on the use of the matches extracted not only from the entire image but also from the cropped overlap regions. DBSCAN aids in finding dense point clusters with most of the matches. We compute the two-view geometry from image point correspondences employing RANSAC instead of COLMAP internal implementation.</p>\n<h2>Key Takeaways:</h2>\n<p>•    <strong>Gradually filter unique image pairs</strong> within the scene based on the number of matches<br>\n•    <strong>ALIKED+LightGlue finetuned settings</strong> with unlimited N of keypoints produced by ALIKED n16 and the LG parameters providing accurate results.<br>\n•    <strong>Cache keypoints &amp; descriptors</strong>, generated by ALIKED for each image. It helped to reduce the running time a lot.<br>\n•    <strong>Multi-GPU acceleration. Mixed precision with GPU T4x2</strong> hardware helps to reduce the time of image matching stage significantly. We employ both GPUs to perform SfM in parallel.<br>\n•    <strong>Matches TTA.</strong> An ensemble of matches extracted from high-resolution images of different scales, as well as matches extracted from both original and cropped images. In LB experiments, the best results were obtained with a combination of 1280 and 2048 for both original and cropped images.<br>\n•    <strong>New crop method.</strong> Instead of the method proposed at IMC2022, where crop areas are calculated for each image pair, we have adopted a new cropping technique that calculates crop areas for each image individually. In this method, key points that are matched to other images with a frequency above a certain threshold are clustered using DBSCAN to crop representative regions of the image. The traditional cropping per pair carried the risk of discarding important areas if they did not match in the original images, but this method reduces the risk of overlooking important areas by using information from more image pairs. Local experiments consistently showed better performance with this method, so we adopted it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F1015c0b267363e4efe2fde995554275f%2FNew%20crop%20method.png?generation=1717536578615603&amp;alt=media\" alt=\"new_crop_method\"></p>\n<p>•    <strong>Rotate one of each image pair</strong> by 0 degrees, 90 degrees, 180 degrees, or 270 degrees, and adopt the rotation that results in the most matches.<br>\n•    <strong>Repeat scene reconstruction.</strong> We relax the number of image matches threshold to address scenes with a low number of registered images. We select a reconstruction having a greater number of registered images.<br>\n•    <strong>Merge Multiple reconstructions.</strong> Hence COLMAP returns multiple reconstructions, we used Horn alignment to estimate the transformation matrix between them, and then we projected other reconstructions to the best one to register as many images as possible.<br>\n•    <strong>OmniGlue.</strong> The best private submission is a merge between ALIKED+LG and OmniGlue, unfortunately, it was not selected.</p>\n<h2>Ideas that did not work out or not fully tested:</h2>\n<p>•    <strong>TTA multi-crop</strong> for image pair filtering. The idea was to split an image into multiple crops and extract matches in order to find similar images in the scene and, then, determine the best image pairs.<br>\n•    <strong>Different detectors, matchers.</strong> We tested SP/SG, LoFTR, DKM, RoMa, OmniGlue, XFeat, KeyNetAffNetHardNet, DISK, and SIFT. Eventually, ALIKED and LightGlue were the best LB choice for us. Apparently, the reason why many dense-based methods did not work out for us is because of the low performance of the “repeatability” across different image pairs and sometimes high noise in the matches. OmniGlue showed promising local results, but no noticeable improvement in LB. RoMa surprisingly showed high performance on the <strong>lizard</strong> scene (84.78 %).<br>\n•    <strong>Different CNNs to find the most similar images</strong> in the scene and generate corresponding image pairs (NetVLAD, Dino, Dino Salad, etc.), we evaluated the the score for these methods using ground truth and intrinsics by calculating the average  IOU of volumetric frustums of each camera with its candidates [image below], but this wasn’t reflected on LB. so Eventually, we gave up using this strategy at all.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2Fe28e557bcafb34c54b52aadfc0e43894%2FVolumetric%20frustums.png?generation=1717536717556995&amp;alt=media\" alt=\"volumetric_frustums\"></p>\n<p>•    <strong>Different keypoint/matching refinement methods</strong>: Based on our experience from IMC23, we knew that PixSFM might help, but due to updates on the Kaggle environment since last year, we only could have tested with the old COLMAP version. So, when the IMC23 winners shared their code for DFSFM, we tested the refinement part above our final pipeline. Unfortunately, even with parallel computing, due to the time limit the submission didn’t pass, and in general the DFSFM refinement improved the validation only a little, so we didn’t use it. [It was really hard to make it run on Kaggle].</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F99ab25e5727fdeda1c26c5c82bb6e9f4%2Fdfsfm_table.png?generation=1717536833178076&amp;alt=media\" alt=\"dfsfm_table\"></p>\n<p>•    <strong>NMS.</strong> We implemented the strategy described by the team “ZJU3DV” in IMC2023. We got some improvement on church locally, but no boost on LB.<br>\n•    <strong>Disambiguation.</strong> There are multiple papers that tackle symmetrical scenes, we have tried <a href=\"https://arxiv.org/pdf/2309.02420\" target=\"_blank\">Doppelgangers</a> and <a href=\"https://yanqingan.github.io/docs/cvpr17_distinguishing.pdf\" target=\"_blank\">Yan</a> methods, the latter had a tiny local improvement on CV but a drop on LB, so we decided to skip them.<br>\n•    <strong>3D RANSAC cleaning.</strong> We utilized 3D RANSAC cleaning to eliminate incorrect matches from symmetric scenes, taking advantage of depth information. For every image, we created depth masks. Then, for each pair of matches, we considered the x,z projection, which is similar to a bird's-eye view of the matched points. Subsequently, we employed RANSAC to identify and remove the outliers.<br>\n•    <strong>Image motion de-blurring.</strong> We recognized that the quality of some images in train was reduced possibly due to the camera shake. To mitigate this effect, we experimented with the single image restoration de-blur method but did not get noticeable improvement on LB.</p>\n<hr>\n<h1>2. DIP Module [transparent scenes]</h1>\n<p>We quickly realized that the SfM pipeline with image matching and COLMAP \"as is\" does not work with transparent scenes. Therefore, we decided to experiment with different strategies. We assumed that the direct pose estimation of the object in each image may help to compute the rotation matrix. We started to make some assumptions about the transparent category: “Given very low metric thresholds for the transparent category, the camera positions have to be somewhere very close to the object, and most probably the object is shot from all sides”. <br>\nSo, we simply throw the cameras on a circle around the object, that looks towards the object as in this image.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F520a9375c05ebdb0c8039d38c0134006%2Fdip_camera_pose_circle.png?generation=1717536931811867&amp;alt=media\" alt=\"camera pose circle\"></p>\n<p>Lately, we’ve also found a very interesting scientific paper which describes approaches to solve transparent categories, that included a tiny dataset of 4 transparent objects, yet we didn’t find it publicly available.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F0b7bb8a9fb2a11fb9c2a2346d567a660%2Ftransparent_objects_poses_paper.png?generation=1717537022273030&amp;alt=media\" alt=\"transparent_objects_paper\"></p>\n<p>Image is taken from the <a href=\"https://isprs-archives.copernicus.org/articles/XLVIII-2-W2-2022/77/2022/isprs-archives-XLVIII-2-W2-2022-77-2022.pdf\" target=\"_blank\">paper</a> [Morelli, L., et al. \"Orientation of Images with Low Contrast Textures and Transparent Objects.\" The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences 48 (2022): 77-84.] which shares interesting experiment results about transparent scenes.</p>\n<p>We assumed that if we manage to solve these transparent scenes to some extent, we would be good with LB. For validation purposes, we even <strong>generated our own transparent scene dataset</strong> with some plastic &amp; glass bottles:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F241f6199e349cc56c206f2b075484c53%2Fvlad_two_bottles_dataset.png?generation=1717537121183118&amp;alt=media\" alt=\"two_bottles_dataset\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F42bf710a8be051e95f38b03ee132c408%2Fvlad%20imc24%20bottles%20slideshow.png?generation=1717537147849661&amp;alt=media\" alt=\"bottles_slideshow\"></p>\n<p>Basically, we came up with the following two approaches:</p>\n<h3>Approach #1</h3>\n<p>To make our assumption work, we needed to sort the images in the right order, then assign the camera pose for each image as to be located on the ideal circle around the object with even spaces between them (e.g. for 36 images its 10 degrees). We tried several approaches to sort the images, all of them include solving the Travelling Salesman Problem as the final step:</p>\n<ol>\n<li><strong>Optical Flow.</strong> Calculate the magnitude of OF for each image pair and assign each pair a weight equal to the standard deviation of magnitude.</li>\n<li><strong>Pixel-level difference.</strong> A simple difference of grayscaled images, the weight for each pair is equal to the difference value.</li>\n<li><strong>SSIM score.</strong> Calculating the <a href=\"https://en.wikipedia.org/wiki/Structural_similarity_index_measure\" target=\"_blank\">SSIM index</a> for each pair and assigning pair weight equal to 1 - ssim.</li>\n<li><strong>ALIKED+LG matching.</strong> Again calculating the number of matches for each pair, and assigning a pair weight equal to (1 / num_matches).</li>\n</ol>\n<p>In each of the four above-mentioned approaches the pair weight indicates how close images are to each other (the less the weight, the more similar the images are). We built a distance matrix based on these pair weights and solved the final ordering problem through TSP.</p>\n<h3>Approach #2</h3>\n<p>A reasonable score was obtained by estimating the order of the images, pairing them with the previous and following images, and performing a matching process at a very high resolution (4096~).<br>\nImage order was estimated based on the number of matches. Experimental results showed that the number of matches tended to be higher for the before and after images. This tendency is used to estimate image order using a kNN-like method.</p>\n<p><strong>CV results</strong>, cylinder: 77.78%, cup: 41.92%.<br>\n<strong>LB boost</strong> using a transparent trick: +0.03</p>\n<h2>Event timeline with progression:</h2>\n<p>•    <strong>Optical Flow</strong> for image ordering using the mean average of the flow which is basically presented by pixel displacement u and v. Standard deviation scored better.<br>\n•    <strong>Grayscale pixel difference.</strong><br>\n•    <strong>SIMM score</strong> in a range [-1, 1], where 1 indicates perfect similarity, 0 indicates no similarity, and -1 indicates perfect anti-correlation.<br>\n•    <strong>SIMM + \"Matching Flow\"</strong> is the final selected ensemble. It gave us <strong>+0.09</strong> on public LB.<br>\n<strong>CV results</strong>, cylinder: <strong>~92%</strong>, cup: <strong>~62%</strong></p>\n<h3>Ideas that did not work out:</h3>\n<p>•    Edge extraction methods (e.g., CLAHE, Canny, Gaussian smoothing, Laplacian).<br>\n•    Depth masks.<br>\n•    Segmentation of transparent objects before flow calculation.</p>\n<hr>\n<h2>Facts and Numbers:</h2>\n<p>•    <strong>0.28/0.25</strong> is the final score in public/private LB<br>\n•    <strong>0.28</strong>4871 is the best public score<br>\n•    <strong>0.26</strong>7720 is the best private score<br>\n•    <strong>+0.09</strong> with a transparent trick in public/private LB<br>\n•    <strong>16914 sec (4h 42min)</strong> is the execution time of the selected best submission<br>\n•    <strong>2 scenes with &lt;50% of registered images</strong> yet in LB<br>\n•    <strong>335</strong> submissions by our team<br>\n•    <strong>∞ cups of coffee and dedication</strong></p>\n<h2>Local Validation:</h2>\n<p>Our validation image dataset used the following data for each scene:<br>\n•    <strong>\"church\", \"lizard\"</strong>: Image sets that were provided as test data in the competition.<br>\n•    <strong>\"dioscuri\", \"multi-temporal-temple\"</strong>: Image sets that were provided as train data.<br>\n•    <strong>\"pond\"</strong>: A set of 65 images randomly extracted from the training dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2Fb0388aa318697244ad212f25bffbccc9%2Fyumeneko_final_metrics_jafaar_v2.PNG?generation=1717546332981607&amp;alt=media\" alt=\"final_metrics\"><br>\n<strong>Final metric -&gt; mAA=41.23 (Best LB) / 43.50 (Best CV)</strong></p>\n<p><strong>UPD</strong> We shared our <a href=\"https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution\" target=\"_blank\">winning notebook</a></p>",
  "messages": [
    {
      "id": 2855591,
      "postDate": "2024-06-04T21:48:16.710Z",
      "content": "<p>Our team would like to deeply appreciate the Kaggle team, Czech Technical University in Prague, and other people helping to carry out the series of exciting image matching challenges, as well as everyone here who competes. My congratulations to all the participants!</p>\n<p>The work we describe here is truly a joint effort of <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a>, <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a>, <a href=\"https://www.kaggle.com/jaafarmahmoud1\" target=\"_blank\">@jaafarmahmoud1</a>, <a href=\"https://www.kaggle.com/kashiwaba\" target=\"_blank\">@kashiwaba</a>, and <a href=\"https://www.kaggle.com/igorlashkov\" target=\"_blank\">@igorlashkov</a>. I’m grateful for being a part of this hardworking, cohesive, and skilled team.</p>\n<p>As the final submissions we selected:<br>\n•    <strong>best public LB notebook</strong> which scored <strong>0.28</strong><br>\n•    submission with the <strong>best local CV</strong> with <strong>0.24</strong> in LB</p>\n<h1>Overview</h1>\n<p>Our final solution consists of the 3D image reconstruction (I3DR) module powered by COLMAP for non-transparent scenes and a simplified direct image pose estimation (DIP) module for transparent scenes (category “transparent” in “categories.csv”). We stay only with sparse detectors and matchers by employing an ensemble of ALIKED extractors and LightGlue matchers on the high-resolution image pairs. As for the DIP module, we estimate a 3D pose of every image in the scene by placing images in the correct order of object rotation and computing a rotation matrix and a translation vector.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F98595408fef16e793d7a77b247dee746%2FReference%20model%20pipeline.png?generation=1717536474163029&amp;alt=media\" alt=\"architecture\"></p>\n<h1>CV</h1>\n<p>Just like in many other Kaggle competitions, it is important to have a good cross-validation pipeline to improve the LB score consistently. We used transparent and non-transparent scenes for the validation separately. Only images having a GT were included in the test set. As for the validation of the I3DR module, it was hard to improve a pipeline step by step as the correlation with LB was uncertain. It is important to limit COLMAP by 1 thread. On the contrary, the algorithm operating on transparent scenes was easier to debug. In this case, CV correlated better with LB, but not in every case.</p>\n<h1>1. I3DR Module [non-transparent scenes]</h1>\n<p>For the image pair selection, we decided not to use any image retrieval method. Instead, we rely only on the number of matches generated by ALIKED and LightGlue (LG). Specifically, we apply a threshold of 30 matches for a single detector and a value of 100 for matches produced by all detectors on an entire image in the ensemble. It is easy to notice that some scenes (e.g., dioscuri) have images not in the natural orientation and need to be re-rotated to find more matches. Similar to previous IMCs, we rely on the use of the matches extracted not only from the entire image but also from the cropped overlap regions. DBSCAN aids in finding dense point clusters with most of the matches. We compute the two-view geometry from image point correspondences employing RANSAC instead of COLMAP internal implementation.</p>\n<h2>Key Takeaways:</h2>\n<p>•    <strong>Gradually filter unique image pairs</strong> within the scene based on the number of matches<br>\n•    <strong>ALIKED+LightGlue finetuned settings</strong> with unlimited N of keypoints produced by ALIKED n16 and the LG parameters providing accurate results.<br>\n•    <strong>Cache keypoints &amp; descriptors</strong>, generated by ALIKED for each image. It helped to reduce the running time a lot.<br>\n•    <strong>Multi-GPU acceleration. Mixed precision with GPU T4x2</strong> hardware helps to reduce the time of image matching stage significantly. We employ both GPUs to perform SfM in parallel.<br>\n•    <strong>Matches TTA.</strong> An ensemble of matches extracted from high-resolution images of different scales, as well as matches extracted from both original and cropped images. In LB experiments, the best results were obtained with a combination of 1280 and 2048 for both original and cropped images.<br>\n•    <strong>New crop method.</strong> Instead of the method proposed at IMC2022, where crop areas are calculated for each image pair, we have adopted a new cropping technique that calculates crop areas for each image individually. In this method, key points that are matched to other images with a frequency above a certain threshold are clustered using DBSCAN to crop representative regions of the image. The traditional cropping per pair carried the risk of discarding important areas if they did not match in the original images, but this method reduces the risk of overlooking important areas by using information from more image pairs. Local experiments consistently showed better performance with this method, so we adopted it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F1015c0b267363e4efe2fde995554275f%2FNew%20crop%20method.png?generation=1717536578615603&amp;alt=media\" alt=\"new_crop_method\"></p>\n<p>•    <strong>Rotate one of each image pair</strong> by 0 degrees, 90 degrees, 180 degrees, or 270 degrees, and adopt the rotation that results in the most matches.<br>\n•    <strong>Repeat scene reconstruction.</strong> We relax the number of image matches threshold to address scenes with a low number of registered images. We select a reconstruction having a greater number of registered images.<br>\n•    <strong>Merge Multiple reconstructions.</strong> Hence COLMAP returns multiple reconstructions, we used Horn alignment to estimate the transformation matrix between them, and then we projected other reconstructions to the best one to register as many images as possible.<br>\n•    <strong>OmniGlue.</strong> The best private submission is a merge between ALIKED+LG and OmniGlue, unfortunately, it was not selected.</p>\n<h2>Ideas that did not work out or not fully tested:</h2>\n<p>•    <strong>TTA multi-crop</strong> for image pair filtering. The idea was to split an image into multiple crops and extract matches in order to find similar images in the scene and, then, determine the best image pairs.<br>\n•    <strong>Different detectors, matchers.</strong> We tested SP/SG, LoFTR, DKM, RoMa, OmniGlue, XFeat, KeyNetAffNetHardNet, DISK, and SIFT. Eventually, ALIKED and LightGlue were the best LB choice for us. Apparently, the reason why many dense-based methods did not work out for us is because of the low performance of the “repeatability” across different image pairs and sometimes high noise in the matches. OmniGlue showed promising local results, but no noticeable improvement in LB. RoMa surprisingly showed high performance on the <strong>lizard</strong> scene (84.78 %).<br>\n•    <strong>Different CNNs to find the most similar images</strong> in the scene and generate corresponding image pairs (NetVLAD, Dino, Dino Salad, etc.), we evaluated the the score for these methods using ground truth and intrinsics by calculating the average  IOU of volumetric frustums of each camera with its candidates [image below], but this wasn’t reflected on LB. so Eventually, we gave up using this strategy at all.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2Fe28e557bcafb34c54b52aadfc0e43894%2FVolumetric%20frustums.png?generation=1717536717556995&amp;alt=media\" alt=\"volumetric_frustums\"></p>\n<p>•    <strong>Different keypoint/matching refinement methods</strong>: Based on our experience from IMC23, we knew that PixSFM might help, but due to updates on the Kaggle environment since last year, we only could have tested with the old COLMAP version. So, when the IMC23 winners shared their code for DFSFM, we tested the refinement part above our final pipeline. Unfortunately, even with parallel computing, due to the time limit the submission didn’t pass, and in general the DFSFM refinement improved the validation only a little, so we didn’t use it. [It was really hard to make it run on Kaggle].</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F99ab25e5727fdeda1c26c5c82bb6e9f4%2Fdfsfm_table.png?generation=1717536833178076&amp;alt=media\" alt=\"dfsfm_table\"></p>\n<p>•    <strong>NMS.</strong> We implemented the strategy described by the team “ZJU3DV” in IMC2023. We got some improvement on church locally, but no boost on LB.<br>\n•    <strong>Disambiguation.</strong> There are multiple papers that tackle symmetrical scenes, we have tried <a href=\"https://arxiv.org/pdf/2309.02420\" target=\"_blank\">Doppelgangers</a> and <a href=\"https://yanqingan.github.io/docs/cvpr17_distinguishing.pdf\" target=\"_blank\">Yan</a> methods, the latter had a tiny local improvement on CV but a drop on LB, so we decided to skip them.<br>\n•    <strong>3D RANSAC cleaning.</strong> We utilized 3D RANSAC cleaning to eliminate incorrect matches from symmetric scenes, taking advantage of depth information. For every image, we created depth masks. Then, for each pair of matches, we considered the x,z projection, which is similar to a bird's-eye view of the matched points. Subsequently, we employed RANSAC to identify and remove the outliers.<br>\n•    <strong>Image motion de-blurring.</strong> We recognized that the quality of some images in train was reduced possibly due to the camera shake. To mitigate this effect, we experimented with the single image restoration de-blur method but did not get noticeable improvement on LB.</p>\n<hr>\n<h1>2. DIP Module [transparent scenes]</h1>\n<p>We quickly realized that the SfM pipeline with image matching and COLMAP \"as is\" does not work with transparent scenes. Therefore, we decided to experiment with different strategies. We assumed that the direct pose estimation of the object in each image may help to compute the rotation matrix. We started to make some assumptions about the transparent category: “Given very low metric thresholds for the transparent category, the camera positions have to be somewhere very close to the object, and most probably the object is shot from all sides”. <br>\nSo, we simply throw the cameras on a circle around the object, that looks towards the object as in this image.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F520a9375c05ebdb0c8039d38c0134006%2Fdip_camera_pose_circle.png?generation=1717536931811867&amp;alt=media\" alt=\"camera pose circle\"></p>\n<p>Lately, we’ve also found a very interesting scientific paper which describes approaches to solve transparent categories, that included a tiny dataset of 4 transparent objects, yet we didn’t find it publicly available.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F0b7bb8a9fb2a11fb9c2a2346d567a660%2Ftransparent_objects_poses_paper.png?generation=1717537022273030&amp;alt=media\" alt=\"transparent_objects_paper\"></p>\n<p>Image is taken from the <a href=\"https://isprs-archives.copernicus.org/articles/XLVIII-2-W2-2022/77/2022/isprs-archives-XLVIII-2-W2-2022-77-2022.pdf\" target=\"_blank\">paper</a> [Morelli, L., et al. \"Orientation of Images with Low Contrast Textures and Transparent Objects.\" The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences 48 (2022): 77-84.] which shares interesting experiment results about transparent scenes.</p>\n<p>We assumed that if we manage to solve these transparent scenes to some extent, we would be good with LB. For validation purposes, we even <strong>generated our own transparent scene dataset</strong> with some plastic &amp; glass bottles:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F241f6199e349cc56c206f2b075484c53%2Fvlad_two_bottles_dataset.png?generation=1717537121183118&amp;alt=media\" alt=\"two_bottles_dataset\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F42bf710a8be051e95f38b03ee132c408%2Fvlad%20imc24%20bottles%20slideshow.png?generation=1717537147849661&amp;alt=media\" alt=\"bottles_slideshow\"></p>\n<p>Basically, we came up with the following two approaches:</p>\n<h3>Approach #1</h3>\n<p>To make our assumption work, we needed to sort the images in the right order, then assign the camera pose for each image as to be located on the ideal circle around the object with even spaces between them (e.g. for 36 images its 10 degrees). We tried several approaches to sort the images, all of them include solving the Travelling Salesman Problem as the final step:</p>\n<ol>\n<li><strong>Optical Flow.</strong> Calculate the magnitude of OF for each image pair and assign each pair a weight equal to the standard deviation of magnitude.</li>\n<li><strong>Pixel-level difference.</strong> A simple difference of grayscaled images, the weight for each pair is equal to the difference value.</li>\n<li><strong>SSIM score.</strong> Calculating the <a href=\"https://en.wikipedia.org/wiki/Structural_similarity_index_measure\" target=\"_blank\">SSIM index</a> for each pair and assigning pair weight equal to 1 - ssim.</li>\n<li><strong>ALIKED+LG matching.</strong> Again calculating the number of matches for each pair, and assigning a pair weight equal to (1 / num_matches).</li>\n</ol>\n<p>In each of the four above-mentioned approaches the pair weight indicates how close images are to each other (the less the weight, the more similar the images are). We built a distance matrix based on these pair weights and solved the final ordering problem through TSP.</p>\n<h3>Approach #2</h3>\n<p>A reasonable score was obtained by estimating the order of the images, pairing them with the previous and following images, and performing a matching process at a very high resolution (4096~).<br>\nImage order was estimated based on the number of matches. Experimental results showed that the number of matches tended to be higher for the before and after images. This tendency is used to estimate image order using a kNN-like method.</p>\n<p><strong>CV results</strong>, cylinder: 77.78%, cup: 41.92%.<br>\n<strong>LB boost</strong> using a transparent trick: +0.03</p>\n<h2>Event timeline with progression:</h2>\n<p>•    <strong>Optical Flow</strong> for image ordering using the mean average of the flow which is basically presented by pixel displacement u and v. Standard deviation scored better.<br>\n•    <strong>Grayscale pixel difference.</strong><br>\n•    <strong>SIMM score</strong> in a range [-1, 1], where 1 indicates perfect similarity, 0 indicates no similarity, and -1 indicates perfect anti-correlation.<br>\n•    <strong>SIMM + \"Matching Flow\"</strong> is the final selected ensemble. It gave us <strong>+0.09</strong> on public LB.<br>\n<strong>CV results</strong>, cylinder: <strong>~92%</strong>, cup: <strong>~62%</strong></p>\n<h3>Ideas that did not work out:</h3>\n<p>•    Edge extraction methods (e.g., CLAHE, Canny, Gaussian smoothing, Laplacian).<br>\n•    Depth masks.<br>\n•    Segmentation of transparent objects before flow calculation.</p>\n<hr>\n<h2>Facts and Numbers:</h2>\n<p>•    <strong>0.28/0.25</strong> is the final score in public/private LB<br>\n•    <strong>0.28</strong>4871 is the best public score<br>\n•    <strong>0.26</strong>7720 is the best private score<br>\n•    <strong>+0.09</strong> with a transparent trick in public/private LB<br>\n•    <strong>16914 sec (4h 42min)</strong> is the execution time of the selected best submission<br>\n•    <strong>2 scenes with &lt;50% of registered images</strong> yet in LB<br>\n•    <strong>335</strong> submissions by our team<br>\n•    <strong>∞ cups of coffee and dedication</strong></p>\n<h2>Local Validation:</h2>\n<p>Our validation image dataset used the following data for each scene:<br>\n•    <strong>\"church\", \"lizard\"</strong>: Image sets that were provided as test data in the competition.<br>\n•    <strong>\"dioscuri\", \"multi-temporal-temple\"</strong>: Image sets that were provided as train data.<br>\n•    <strong>\"pond\"</strong>: A set of 65 images randomly extracted from the training dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2Fb0388aa318697244ad212f25bffbccc9%2Fyumeneko_final_metrics_jafaar_v2.PNG?generation=1717546332981607&amp;alt=media\" alt=\"final_metrics\"><br>\n<strong>Final metric -&gt; mAA=41.23 (Best LB) / 43.50 (Best CV)</strong></p>\n<p><strong>UPD</strong> We shared our <a href=\"https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution\" target=\"_blank\">winning notebook</a></p>",
      "rawMarkdown": "Our team would like to deeply appreciate the Kaggle team, Czech Technical University in Prague, and other people helping to carry out the series of exciting image matching challenges, as well as everyone here who competes. My congratulations to all the participants!\n\nThe work we describe here is truly a joint effort of @vostankovich, @ammarali32, @jaafarmahmoud1, @kashiwaba, and @igorlashkov. I’m grateful for being a part of this hardworking, cohesive, and skilled team.\n\nAs the final submissions we selected:\n•\t**best public LB notebook** which scored **0.28**\n•\tsubmission with the **best local CV** with **0.24** in LB\n\n# Overview\n\nOur final solution consists of the 3D image reconstruction (I3DR) module powered by COLMAP for non-transparent scenes and a simplified direct image pose estimation (DIP) module for transparent scenes (category “transparent” in “categories.csv”). We stay only with sparse detectors and matchers by employing an ensemble of ALIKED extractors and LightGlue matchers on the high-resolution image pairs. As for the DIP module, we estimate a 3D pose of every image in the scene by placing images in the correct order of object rotation and computing a rotation matrix and a translation vector.\n\n ![architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F98595408fef16e793d7a77b247dee746%2FReference%20model%20pipeline.png?generation=1717536474163029&alt=media)\n\n# CV\n\nJust like in many other Kaggle competitions, it is important to have a good cross-validation pipeline to improve the LB score consistently. We used transparent and non-transparent scenes for the validation separately. Only images having a GT were included in the test set. As for the validation of the I3DR module, it was hard to improve a pipeline step by step as the correlation with LB was uncertain. It is important to limit COLMAP by 1 thread. On the contrary, the algorithm operating on transparent scenes was easier to debug. In this case, CV correlated better with LB, but not in every case.\n\n# 1. I3DR Module [non-transparent scenes]\n\nFor the image pair selection, we decided not to use any image retrieval method. Instead, we rely only on the number of matches generated by ALIKED and LightGlue (LG). Specifically, we apply a threshold of 30 matches for a single detector and a value of 100 for matches produced by all detectors on an entire image in the ensemble. It is easy to notice that some scenes (e.g., dioscuri) have images not in the natural orientation and need to be re-rotated to find more matches. Similar to previous IMCs, we rely on the use of the matches extracted not only from the entire image but also from the cropped overlap regions. DBSCAN aids in finding dense point clusters with most of the matches. We compute the two-view geometry from image point correspondences employing RANSAC instead of COLMAP internal implementation.\n\n## Key Takeaways:\n\n•\t**Gradually filter unique image pairs** within the scene based on the number of matches\n•\t**ALIKED+LightGlue finetuned settings** with unlimited N of keypoints produced by ALIKED n16 and the LG parameters providing accurate results.\n•\t**Cache keypoints & descriptors**, generated by ALIKED for each image. It helped to reduce the running time a lot.\n•\t**Multi-GPU acceleration. Mixed precision with GPU T4x2** hardware helps to reduce the time of image matching stage significantly. We employ both GPUs to perform SfM in parallel.\n•\t**Matches TTA.** An ensemble of matches extracted from high-resolution images of different scales, as well as matches extracted from both original and cropped images. In LB experiments, the best results were obtained with a combination of 1280 and 2048 for both original and cropped images.\n•\t**New crop method.** Instead of the method proposed at IMC2022, where crop areas are calculated for each image pair, we have adopted a new cropping technique that calculates crop areas for each image individually. In this method, key points that are matched to other images with a frequency above a certain threshold are clustered using DBSCAN to crop representative regions of the image. The traditional cropping per pair carried the risk of discarding important areas if they did not match in the original images, but this method reduces the risk of overlooking important areas by using information from more image pairs. Local experiments consistently showed better performance with this method, so we adopted it.\n \n![new_crop_method](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F1015c0b267363e4efe2fde995554275f%2FNew%20crop%20method.png?generation=1717536578615603&alt=media)\n\n•\t**Rotate one of each image pair** by 0 degrees, 90 degrees, 180 degrees, or 270 degrees, and adopt the rotation that results in the most matches.\n•\t**Repeat scene reconstruction.** We relax the number of image matches threshold to address scenes with a low number of registered images. We select a reconstruction having a greater number of registered images.\n•\t**Merge Multiple reconstructions.** Hence COLMAP returns multiple reconstructions, we used Horn alignment to estimate the transformation matrix between them, and then we projected other reconstructions to the best one to register as many images as possible.\n•\t**OmniGlue.** The best private submission is a merge between ALIKED+LG and OmniGlue, unfortunately, it was not selected.\n\n## Ideas that did not work out or not fully tested:\n\n•\t**TTA multi-crop** for image pair filtering. The idea was to split an image into multiple crops and extract matches in order to find similar images in the scene and, then, determine the best image pairs.\n•\t**Different detectors, matchers.** We tested SP/SG, LoFTR, DKM, RoMa, OmniGlue, XFeat, KeyNetAffNetHardNet, DISK, and SIFT. Eventually, ALIKED and LightGlue were the best LB choice for us. Apparently, the reason why many dense-based methods did not work out for us is because of the low performance of the “repeatability” across different image pairs and sometimes high noise in the matches. OmniGlue showed promising local results, but no noticeable improvement in LB. RoMa surprisingly showed high performance on the **lizard** scene (84.78 %).\n•\t**Different CNNs to find the most similar images** in the scene and generate corresponding image pairs (NetVLAD, Dino, Dino Salad, etc.), we evaluated the the score for these methods using ground truth and intrinsics by calculating the average  IOU of volumetric frustums of each camera with its candidates [image below], but this wasn’t reflected on LB. so Eventually, we gave up using this strategy at all.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2Fe28e557bcafb34c54b52aadfc0e43894%2FVolumetric%20frustums.png?generation=1717536717556995&alt=media\" height=\"150px\" alt=\"volumetric_frustums\">\n \n•\t**Different keypoint/matching refinement methods**: Based on our experience from IMC23, we knew that PixSFM might help, but due to updates on the Kaggle environment since last year, we only could have tested with the old COLMAP version. So, when the IMC23 winners shared their code for DFSFM, we tested the refinement part above our final pipeline. Unfortunately, even with parallel computing, due to the time limit the submission didn’t pass, and in general the DFSFM refinement improved the validation only a little, so we didn’t use it. [It was really hard to make it run on Kaggle].\n \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F99ab25e5727fdeda1c26c5c82bb6e9f4%2Fdfsfm_table.png?generation=1717536833178076&alt=media\" height=\"150px\" alt=\"dfsfm_table\">\n\n•\t**NMS.** We implemented the strategy described by the team “ZJU3DV” in IMC2023. We got some improvement on church locally, but no boost on LB.\n•\t**Disambiguation.** There are multiple papers that tackle symmetrical scenes, we have tried [Doppelgangers](https://arxiv.org/pdf/2309.02420) and [Yan](https://yanqingan.github.io/docs/cvpr17_distinguishing.pdf) methods, the latter had a tiny local improvement on CV but a drop on LB, so we decided to skip them.\n•\t**3D RANSAC cleaning.** We utilized 3D RANSAC cleaning to eliminate incorrect matches from symmetric scenes, taking advantage of depth information. For every image, we created depth masks. Then, for each pair of matches, we considered the x,z projection, which is similar to a bird's-eye view of the matched points. Subsequently, we employed RANSAC to identify and remove the outliers.\n•\t**Image motion de-blurring.** We recognized that the quality of some images in train was reduced possibly due to the camera shake. To mitigate this effect, we experimented with the single image restoration de-blur method but did not get noticeable improvement on LB.\n\n---\n# 2. DIP Module [transparent scenes]\n\nWe quickly realized that the SfM pipeline with image matching and COLMAP \"as is\" does not work with transparent scenes. Therefore, we decided to experiment with different strategies. We assumed that the direct pose estimation of the object in each image may help to compute the rotation matrix. We started to make some assumptions about the transparent category: “Given very low metric thresholds for the transparent category, the camera positions have to be somewhere very close to the object, and most probably the object is shot from all sides”. \nSo, we simply throw the cameras on a circle around the object, that looks towards the object as in this image.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F520a9375c05ebdb0c8039d38c0134006%2Fdip_camera_pose_circle.png?generation=1717536931811867&alt=media\" height=\"300px\" alt=\"camera pose circle\">\n \nLately, we’ve also found a very interesting scientific paper which describes approaches to solve transparent categories, that included a tiny dataset of 4 transparent objects, yet we didn’t find it publicly available.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F0b7bb8a9fb2a11fb9c2a2346d567a660%2Ftransparent_objects_poses_paper.png?generation=1717537022273030&alt=media\" height=\"240px\" alt=\"transparent_objects_paper\">\n \nImage is taken from the [paper](https://isprs-archives.copernicus.org/articles/XLVIII-2-W2-2022/77/2022/isprs-archives-XLVIII-2-W2-2022-77-2022.pdf) [Morelli, L., et al. \"Orientation of Images with Low Contrast Textures and Transparent Objects.\" The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences 48 (2022): 77-84.] which shares interesting experiment results about transparent scenes.\n\nWe assumed that if we manage to solve these transparent scenes to some extent, we would be good with LB. For validation purposes, we even **generated our own transparent scene dataset** with some plastic & glass bottles:\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F241f6199e349cc56c206f2b075484c53%2Fvlad_two_bottles_dataset.png?generation=1717537121183118&alt=media\" height=\"300px\" alt=\"two_bottles_dataset\">\n  \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F42bf710a8be051e95f38b03ee132c408%2Fvlad%20imc24%20bottles%20slideshow.png?generation=1717537147849661&alt=media\" height=\"300px\" alt=\"bottles_slideshow\">\n\nBasically, we came up with the following two approaches:\n\n### Approach #1 \nTo make our assumption work, we needed to sort the images in the right order, then assign the camera pose for each image as to be located on the ideal circle around the object with even spaces between them (e.g. for 36 images its 10 degrees). We tried several approaches to sort the images, all of them include solving the Travelling Salesman Problem as the final step:\n1.\t**Optical Flow.** Calculate the magnitude of OF for each image pair and assign each pair a weight equal to the standard deviation of magnitude.\n2.\t**Pixel-level difference.** A simple difference of grayscaled images, the weight for each pair is equal to the difference value.\n3.\t**SSIM score.** Calculating the [SSIM index](https://en.wikipedia.org/wiki/Structural_similarity_index_measure) for each pair and assigning pair weight equal to 1 - ssim.\n4.\t**ALIKED+LG matching.** Again calculating the number of matches for each pair, and assigning a pair weight equal to (1 / num_matches).\n\nIn each of the four above-mentioned approaches the pair weight indicates how close images are to each other (the less the weight, the more similar the images are). We built a distance matrix based on these pair weights and solved the final ordering problem through TSP.\n\n### Approach #2 \nA reasonable score was obtained by estimating the order of the images, pairing them with the previous and following images, and performing a matching process at a very high resolution (4096~).\nImage order was estimated based on the number of matches. Experimental results showed that the number of matches tended to be higher for the before and after images. This tendency is used to estimate image order using a kNN-like method.\n\n**CV results**, cylinder: 77.78%, cup: 41.92%.\n**LB boost** using a transparent trick: +0.03\n\n## Event timeline with progression:\n•\t**Optical Flow** for image ordering using the mean average of the flow which is basically presented by pixel displacement u and v. Standard deviation scored better.\n•\t**Grayscale pixel difference.**\n•\t**SIMM score** in a range [-1, 1], where 1 indicates perfect similarity, 0 indicates no similarity, and -1 indicates perfect anti-correlation.\n•\t**SIMM + \"Matching Flow\"** is the final selected ensemble. It gave us **+0.09** on public LB.\n**CV results**, cylinder: **~92%**, cup: **~62%**\n\n### Ideas that did not work out:\n•\tEdge extraction methods (e.g., CLAHE, Canny, Gaussian smoothing, Laplacian).\n•\tDepth masks.\n•\tSegmentation of transparent objects before flow calculation.\n\n---\n## Facts and Numbers:\n\n•\t**0.28/0.25** is the final score in public/private LB\n•\t**0.28**4871 is the best public score\n•\t**0.26**7720 is the best private score\n•\t**+0.09** with a transparent trick in public/private LB\n•\t**16914 sec (4h 42min)** is the execution time of the selected best submission\n•\t**2 scenes with <50% of registered images** yet in LB\n•\t**335** submissions by our team\n•\t**∞ cups of coffee and dedication**\n\n## Local Validation:\n\nOur validation image dataset used the following data for each scene:\n•\t**\"church\", \"lizard\"**: Image sets that were provided as test data in the competition.\n•\t**\"dioscuri\", \"multi-temporal-temple\"**: Image sets that were provided as train data.\n•\t**\"pond\"**: A set of 65 images randomly extracted from the training dataset.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2Fb0388aa318697244ad212f25bffbccc9%2Fyumeneko_final_metrics_jafaar_v2.PNG?generation=1717546332981607&alt=media\" alt=\"final_metrics\">\n**Final metric -> mAA=41.23 (Best LB) / 43.50 (Best CV)**\n\n\n**UPD** We shared our [winning notebook](https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution)\n",
      "votes": 96
    },
    {
      "id": 2871706,
      "postDate": "2024-06-14T10:56:02.340Z",
      "content": "<p>We have released the <a href=\"https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution\" target=\"_blank\">code</a></p>",
      "rawMarkdown": "We have released the [code](https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution)",
      "votes": 7
    },
    {
      "id": 2856950,
      "postDate": "2024-06-05T16:19:53.147Z",
      "content": "<p>Wow great work guys! I must admit it is both exciting and painful to see how successful using similarity to differentiate glass and using separate ideas on those scenes was. Based on some research I did very early on I was able to find that this could provide a value and it did match up with the 0.03 gain that some top teams got early on. I put a ton of time into getting a solution with it to work but I had an error in submission that I simply could not solve for. Glad to see the idea worked as well as I thought it could. Also great to see just how elegant the solution you crafted was a very well deserved win!</p>",
      "rawMarkdown": "Wow great work guys! I must admit it is both exciting and painful to see how successful using similarity to differentiate glass and using separate ideas on those scenes was. Based on some research I did very early on I was able to find that this could provide a value and it did match up with the 0.03 gain that some top teams got early on. I put a ton of time into getting a solution with it to work but I had an error in submission that I simply could not solve for. Glad to see the idea worked as well as I thought it could. Also great to see just how elegant the solution you crafted was a very well deserved win!",
      "votes": 7
    },
    {
      "id": 2858274,
      "postDate": "2024-06-06T11:44:25.547Z",
      "content": "<p>Impressive solution! The way you combined ALIKED extractors and LightGlue matchers with high-resolution images is brilliant. Your dedication to optimizing the pipeline and handling transparent scenes is evident. Great job!</p>",
      "rawMarkdown": "Impressive solution! The way you combined ALIKED extractors and LightGlue matchers with high-resolution images is brilliant. Your dedication to optimizing the pipeline and handling transparent scenes is evident. Great job!",
      "votes": 5
    },
    {
      "id": 2857817,
      "postDate": "2024-06-06T06:19:32.840Z",
      "content": "<p>Congratulations on winning the competition! Your approach to handling transparent scenes is very impressive. We also processed transparent scenes in a similar way, trying to restore the order of the images, but unfortunately, we didn't make more in-depth efforts in this aspect. Congratulations again to you.</p>",
      "rawMarkdown": "Congratulations on winning the competition! Your approach to handling transparent scenes is very impressive. We also processed transparent scenes in a similar way, trying to restore the order of the images, but unfortunately, we didn't make more in-depth efforts in this aspect. Congratulations again to you.",
      "votes": 5
    },
    {
      "id": 2856425,
      "postDate": "2024-06-05T09:08:53.473Z",
      "content": "<p>Congratulations!!! Thanks for sharing. Reading your approach makes me realize that I still have a lot to learn about 3D CV! But it's inspiring on the contrary, it means there's a direction to keep pumping up my skills.</p>",
      "rawMarkdown": "Congratulations!!! Thanks for sharing. Reading your approach makes me realize that I still have a lot to learn about 3D CV! But it's inspiring on the contrary, it means there's a direction to keep pumping up my skills.",
      "votes": 5
    },
    {
      "id": 2855655,
      "postDate": "2024-06-05T00:41:08.633Z",
      "content": "<p>Congratulations with the strong 1st place! And congratulations to <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> and <a href=\"https://www.kaggle.com/kashiwaba\" target=\"_blank\">@kashiwaba</a> for becoming new GMs!</p>\n<p>Loved the part where you created a validation set with just a phone camera, some bottles, and household objects 😂 You might be the first team on Kaggle in a while that has done some manual data collection like that. Gives me the vibes of old-school Computer Vision research before the deep-learning era 😁</p>",
      "rawMarkdown": "Congratulations with the strong 1st place! And congratulations to @vostankovich and @kashiwaba for becoming new GMs!\n\nLoved the part where you created a validation set with just a phone camera, some bottles, and household objects 😂 You might be the first team on Kaggle in a while that has done some manual data collection like that. Gives me the vibes of old-school Computer Vision research before the deep-learning era 😁",
      "votes": 5,
      "replies": [
        {
          "id": 2855843,
          "postDate": "2024-06-05T04:26:55.990Z",
          "content": "<p>Thanks a lot! It was not easy to find bottles of ideal symmetry in the shop, yet another problem was to remove all the tags and glue xD</p>",
          "rawMarkdown": "Thanks a lot! It was not easy to find bottles of ideal symmetry in the shop, yet another problem was to remove all the tags and glue xD",
          "votes": 7
        }
      ]
    },
    {
      "id": 2866243,
      "postDate": "2024-06-11T08:10:04.013Z",
      "content": "<p>Congratulations!! It is so impressive</p>",
      "rawMarkdown": "Congratulations!! It is so impressive",
      "votes": 3
    },
    {
      "id": 2861363,
      "postDate": "2024-06-08T06:43:14.977Z",
      "content": "<p>Congrats, very impressive</p>",
      "rawMarkdown": "Congrats, very impressive",
      "votes": 3
    },
    {
      "id": 2860253,
      "postDate": "2024-06-07T13:56:24.440Z",
      "content": "<p>congratulations for winning the compitation and very good work</p>",
      "rawMarkdown": "congratulations for winning the compitation and very good work",
      "votes": 3
    },
    {
      "id": 2860048,
      "postDate": "2024-06-07T11:34:24.590Z",
      "content": "<p>Impressive and very insightful!</p>",
      "rawMarkdown": "Impressive and very insightful!",
      "votes": 3
    },
    {
      "id": 2859808,
      "postDate": "2024-06-07T09:10:32.413Z",
      "content": "<p>Congratulations on winning the competition! 🎉</p>",
      "rawMarkdown": "Congratulations on winning the competition! 🎉",
      "votes": 3
    },
    {
      "id": 2858549,
      "postDate": "2024-06-06T14:53:33.607Z",
      "content": "<p>Congratulations 🎉</p>",
      "rawMarkdown": "Congratulations 🎉",
      "votes": 4
    },
    {
      "id": 2857745,
      "postDate": "2024-06-06T05:38:18.920Z",
      "content": "<p>Congratulations!! It is very helpful for me to understand the method of 3D image reconstruction. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations!! It is very helpful for me to understand the method of 3D image reconstruction. Thanks for sharing.",
      "votes": 4
    },
    {
      "id": 2857492,
      "postDate": "2024-06-05T23:08:40.430Z",
      "content": "<p>This is amazing. Great work!</p>",
      "rawMarkdown": "This is amazing. Great work!",
      "votes": 4
    },
    {
      "id": 2856649,
      "postDate": "2024-06-05T12:26:27.247Z",
      "content": "<p>Congr, This is amazing, thanks for sharing. </p>",
      "rawMarkdown": "Congr, This is amazing, thanks for sharing. ",
      "votes": 4
    },
    {
      "id": 2856385,
      "postDate": "2024-06-05T08:43:44.903Z",
      "content": "<p>Innovative approach, excellent results speak for themselves 👍. Can't wait to learn more details on your approach in slides and the paper at CVPR. Impressive trick on transparent scene and great write-up!</p>",
      "rawMarkdown": "Innovative approach, excellent results speak for themselves 👍. Can't wait to learn more details on your approach in slides and the paper at CVPR. Impressive trick on transparent scene and great write-up!",
      "votes": 4
    },
    {
      "id": 2855789,
      "postDate": "2024-06-05T03:26:47.770Z",
      "content": "<p>Nice work! Congratulations 🎉</p>",
      "rawMarkdown": "Nice work! Congratulations 🎉",
      "votes": 3
    },
    {
      "id": 2897171,
      "postDate": "2024-06-30T09:47:34.823Z",
      "content": "<p>the transparency approach was very well executed, i will try to read that paper you mentioned.</p>",
      "rawMarkdown": "the transparency approach was very well executed, i will try to read that paper you mentioned.",
      "votes": 1
    },
    {
      "id": 2867727,
      "postDate": "2024-06-12T04:04:19.527Z",
      "content": "<p>Congratulations!! Your explanation was detailed and amazing.</p>",
      "rawMarkdown": "Congratulations!! Your explanation was detailed and amazing.",
      "votes": 1
    },
    {
      "id": 2868654,
      "postDate": "2024-06-12T14:36:08.283Z",
      "content": "<p>congrats man!</p>",
      "rawMarkdown": "congrats man!",
      "votes": 2
    },
    {
      "id": 2862071,
      "postDate": "2024-06-08T14:41:32.513Z",
      "content": "<p>very impressive work, well done!</p>",
      "rawMarkdown": "very impressive work, well done!",
      "votes": 2
    },
    {
      "id": 2862046,
      "postDate": "2024-06-08T14:13:20.120Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 2
    },
    {
      "id": 2856046,
      "postDate": "2024-06-05T06:08:17.413Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2855626,
      "postDate": "2024-06-04T23:19:33.100Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true
    },
    {
      "id": 2855703,
      "postDate": "2024-06-05T01:14:13.153Z",
      "content": "<p>thanks for sharing.  </p>",
      "rawMarkdown": "thanks for sharing.  ",
      "votes": 4
    },
    {
      "id": 2860235,
      "postDate": "2024-06-07T13:39:38.960Z",
      "content": "<p>Thank you for your sharing</p>",
      "rawMarkdown": "Thank you for your sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2871706,
      "author_name": "Vladislav Ostankovich",
      "author_url": "",
      "post_date": "2024-06-14T10:56:02.340000",
      "content": "<p>We have released the <a href=\"https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution\" target=\"_blank\">code</a></p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2856950,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2024-06-05T16:19:53.147000",
      "content": "<p>Wow great work guys! I must admit it is both exciting and painful to see how successful using similarity to differentiate glass and using separate ideas on those scenes was. Based on some research I did very early on I was able to find that this could provide a value and it did match up with the 0.03 gain that some top teams got early on. I put a ton of time into getting a solution with it to work but I had an error in submission that I simply could not solve for. Glad to see the idea worked as well as I thought it could. Also great to see just how elegant the solution you crafted was a very well deserved win!</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2858274,
      "author_name": "Krens",
      "author_url": "",
      "post_date": "2024-06-06T11:44:25.547000",
      "content": "<p>Impressive solution! The way you combined ALIKED extractors and LightGlue matchers with high-resolution images is brilliant. Your dedication to optimizing the pipeline and handling transparent scenes is evident. Great job!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2857817,
      "author_name": "Hao Yu",
      "author_url": "",
      "post_date": "2024-06-06T06:19:32.840000",
      "content": "<p>Congratulations on winning the competition! Your approach to handling transparent scenes is very impressive. We also processed transparent scenes in a similar way, trying to restore the order of the images, but unfortunately, we didn't make more in-depth efforts in this aspect. Congratulations again to you.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2856425,
      "author_name": "Evgeny Bratkovsky",
      "author_url": "",
      "post_date": "2024-06-05T09:08:53.473000",
      "content": "<p>Congratulations!!! Thanks for sharing. Reading your approach makes me realize that I still have a lot to learn about 3D CV! But it's inspiring on the contrary, it means there's a direction to keep pumping up my skills.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2855655,
      "author_name": "Chan Kha Vu",
      "author_url": "",
      "post_date": "2024-06-05T00:41:08.633000",
      "content": "<p>Congratulations with the strong 1st place! And congratulations to <a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> and <a href=\"https://www.kaggle.com/kashiwaba\" target=\"_blank\">@kashiwaba</a> for becoming new GMs!</p>\n<p>Loved the part where you created a validation set with just a phone camera, some bottles, and household objects 😂 You might be the first team on Kaggle in a while that has done some manual data collection like that. Gives me the vibes of old-school Computer Vision research before the deep-learning era 😁</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2855843,
          "author_name": "Vladislav Ostankovich",
          "author_url": "",
          "post_date": "2024-06-05T04:26:55.990000",
          "content": "<p>Thanks a lot! It was not easy to find bottles of ideal symmetry in the shop, yet another problem was to remove all the tags and glue xD</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 2866243,
      "author_name": "KoSanberg",
      "author_url": "",
      "post_date": "2024-06-11T08:10:04.013000",
      "content": "<p>Congratulations!! It is so impressive</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2861363,
      "author_name": "Abdallah_Gaber333",
      "author_url": "",
      "post_date": "2024-06-08T06:43:14.977000",
      "content": "<p>Congrats, very impressive</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2860253,
      "author_name": "Muhammad Irfan Sajid",
      "author_url": "",
      "post_date": "2024-06-07T13:56:24.440000",
      "content": "<p>congratulations for winning the compitation and very good work</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2860048,
      "author_name": "Fatima Azfar Ziya",
      "author_url": "",
      "post_date": "2024-06-07T11:34:24.590000",
      "content": "<p>Impressive and very insightful!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2859808,
      "author_name": "zhigan hou",
      "author_url": "",
      "post_date": "2024-06-07T09:10:32.413000",
      "content": "<p>Congratulations on winning the competition! 🎉</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2858549,
      "author_name": "cccs ff",
      "author_url": "",
      "post_date": "2024-06-06T14:53:33.607000",
      "content": "<p>Congratulations 🎉</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2857745,
      "author_name": "Masaya Nakanishi",
      "author_url": "",
      "post_date": "2024-06-06T05:38:18.920000",
      "content": "<p>Congratulations!! It is very helpful for me to understand the method of 3D image reconstruction. Thanks for sharing.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2857492,
      "author_name": "Joshua W Robertson",
      "author_url": "",
      "post_date": "2024-06-05T23:08:40.430000",
      "content": "<p>This is amazing. Great work!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2856649,
      "author_name": "Muhammed Tausif",
      "author_url": "",
      "post_date": "2024-06-05T12:26:27.247000",
      "content": "<p>Congr, This is amazing, thanks for sharing. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2856385,
      "author_name": "Nomchek",
      "author_url": "",
      "post_date": "2024-06-05T08:43:44.903000",
      "content": "<p>Innovative approach, excellent results speak for themselves 👍. Can't wait to learn more details on your approach in slides and the paper at CVPR. Impressive trick on transparent scene and great write-up!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2855789,
      "author_name": "Jiatengyu",
      "author_url": "",
      "post_date": "2024-06-05T03:26:47.770000",
      "content": "<p>Nice work! Congratulations 🎉</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2897171,
      "author_name": "SHYAM GUPTA",
      "author_url": "",
      "post_date": "2024-06-30T09:47:34.823000",
      "content": "<p>the transparency approach was very well executed, i will try to read that paper you mentioned.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2867727,
      "author_name": "Muhammad Azeem",
      "author_url": "",
      "post_date": "2024-06-12T04:04:19.527000",
      "content": "<p>Congratulations!! Your explanation was detailed and amazing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2868654,
      "author_name": "Huy G Le",
      "author_url": "",
      "post_date": "2024-06-12T14:36:08.283000",
      "content": "<p>congrats man!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2862071,
      "author_name": "Zakariae Youssefi",
      "author_url": "",
      "post_date": "2024-06-08T14:41:32.513000",
      "content": "<p>very impressive work, well done!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2862046,
      "author_name": "Finneogan",
      "author_url": "",
      "post_date": "2024-06-08T14:13:20.120000",
      "content": "<p>Congratulations!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2856046,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-05T06:08:17.413000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2855626,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-04T23:19:33.100000",
      "content": "",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2855703,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2024-06-05T01:14:13.153000",
      "content": "<p>thanks for sharing.  </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2860235,
      "author_name": "angelwill",
      "author_url": "",
      "post_date": "2024-06-07T13:39:38.960000",
      "content": "<p>Thank you for your sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2855591": "Our team would like to deeply appreciate the Kaggle team, Czech Technical University in Prague, and other people helping to carry out the series of exciting image matching challenges, as well as everyone here who competes. My congratulations to all the participants!\n\nThe work we describe here is truly a joint effort of @vostankovich, @ammarali32, @jaafarmahmoud1, @kashiwaba, and @igorlashkov. I’m grateful for being a part of this hardworking, cohesive, and skilled team.\n\nAs the final submissions we selected:\n•\t**best public LB notebook** which scored **0.28**\n•\tsubmission with the **best local CV** with **0.24** in LB\n\n# Overview\n\nOur final solution consists of the 3D image reconstruction (I3DR) module powered by COLMAP for non-transparent scenes and a simplified direct image pose estimation (DIP) module for transparent scenes (category “transparent” in “categories.csv”). We stay only with sparse detectors and matchers by employing an ensemble of ALIKED extractors and LightGlue matchers on the high-resolution image pairs. As for the DIP module, we estimate a 3D pose of every image in the scene by placing images in the correct order of object rotation and computing a rotation matrix and a translation vector.\n\n ![architecture](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F98595408fef16e793d7a77b247dee746%2FReference%20model%20pipeline.png?generation=1717536474163029&alt=media)\n\n# CV\n\nJust like in many other Kaggle competitions, it is important to have a good cross-validation pipeline to improve the LB score consistently. We used transparent and non-transparent scenes for the validation separately. Only images having a GT were included in the test set. As for the validation of the I3DR module, it was hard to improve a pipeline step by step as the correlation with LB was uncertain. It is important to limit COLMAP by 1 thread. On the contrary, the algorithm operating on transparent scenes was easier to debug. In this case, CV correlated better with LB, but not in every case.\n\n# 1. I3DR Module [non-transparent scenes]\n\nFor the image pair selection, we decided not to use any image retrieval method. Instead, we rely only on the number of matches generated by ALIKED and LightGlue (LG). Specifically, we apply a threshold of 30 matches for a single detector and a value of 100 for matches produced by all detectors on an entire image in the ensemble. It is easy to notice that some scenes (e.g., dioscuri) have images not in the natural orientation and need to be re-rotated to find more matches. Similar to previous IMCs, we rely on the use of the matches extracted not only from the entire image but also from the cropped overlap regions. DBSCAN aids in finding dense point clusters with most of the matches. We compute the two-view geometry from image point correspondences employing RANSAC instead of COLMAP internal implementation.\n\n## Key Takeaways:\n\n•\t**Gradually filter unique image pairs** within the scene based on the number of matches\n•\t**ALIKED+LightGlue finetuned settings** with unlimited N of keypoints produced by ALIKED n16 and the LG parameters providing accurate results.\n•\t**Cache keypoints & descriptors**, generated by ALIKED for each image. It helped to reduce the running time a lot.\n•\t**Multi-GPU acceleration. Mixed precision with GPU T4x2** hardware helps to reduce the time of image matching stage significantly. We employ both GPUs to perform SfM in parallel.\n•\t**Matches TTA.** An ensemble of matches extracted from high-resolution images of different scales, as well as matches extracted from both original and cropped images. In LB experiments, the best results were obtained with a combination of 1280 and 2048 for both original and cropped images.\n•\t**New crop method.** Instead of the method proposed at IMC2022, where crop areas are calculated for each image pair, we have adopted a new cropping technique that calculates crop areas for each image individually. In this method, key points that are matched to other images with a frequency above a certain threshold are clustered using DBSCAN to crop representative regions of the image. The traditional cropping per pair carried the risk of discarding important areas if they did not match in the original images, but this method reduces the risk of overlooking important areas by using information from more image pairs. Local experiments consistently showed better performance with this method, so we adopted it.\n \n![new_crop_method](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F1015c0b267363e4efe2fde995554275f%2FNew%20crop%20method.png?generation=1717536578615603&alt=media)\n\n•\t**Rotate one of each image pair** by 0 degrees, 90 degrees, 180 degrees, or 270 degrees, and adopt the rotation that results in the most matches.\n•\t**Repeat scene reconstruction.** We relax the number of image matches threshold to address scenes with a low number of registered images. We select a reconstruction having a greater number of registered images.\n•\t**Merge Multiple reconstructions.** Hence COLMAP returns multiple reconstructions, we used Horn alignment to estimate the transformation matrix between them, and then we projected other reconstructions to the best one to register as many images as possible.\n•\t**OmniGlue.** The best private submission is a merge between ALIKED+LG and OmniGlue, unfortunately, it was not selected.\n\n## Ideas that did not work out or not fully tested:\n\n•\t**TTA multi-crop** for image pair filtering. The idea was to split an image into multiple crops and extract matches in order to find similar images in the scene and, then, determine the best image pairs.\n•\t**Different detectors, matchers.** We tested SP/SG, LoFTR, DKM, RoMa, OmniGlue, XFeat, KeyNetAffNetHardNet, DISK, and SIFT. Eventually, ALIKED and LightGlue were the best LB choice for us. Apparently, the reason why many dense-based methods did not work out for us is because of the low performance of the “repeatability” across different image pairs and sometimes high noise in the matches. OmniGlue showed promising local results, but no noticeable improvement in LB. RoMa surprisingly showed high performance on the **lizard** scene (84.78 %).\n•\t**Different CNNs to find the most similar images** in the scene and generate corresponding image pairs (NetVLAD, Dino, Dino Salad, etc.), we evaluated the the score for these methods using ground truth and intrinsics by calculating the average  IOU of volumetric frustums of each camera with its candidates [image below], but this wasn’t reflected on LB. so Eventually, we gave up using this strategy at all.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2Fe28e557bcafb34c54b52aadfc0e43894%2FVolumetric%20frustums.png?generation=1717536717556995&alt=media\" height=\"150px\" alt=\"volumetric_frustums\">\n \n•\t**Different keypoint/matching refinement methods**: Based on our experience from IMC23, we knew that PixSFM might help, but due to updates on the Kaggle environment since last year, we only could have tested with the old COLMAP version. So, when the IMC23 winners shared their code for DFSFM, we tested the refinement part above our final pipeline. Unfortunately, even with parallel computing, due to the time limit the submission didn’t pass, and in general the DFSFM refinement improved the validation only a little, so we didn’t use it. [It was really hard to make it run on Kaggle].\n \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F99ab25e5727fdeda1c26c5c82bb6e9f4%2Fdfsfm_table.png?generation=1717536833178076&alt=media\" height=\"150px\" alt=\"dfsfm_table\">\n\n•\t**NMS.** We implemented the strategy described by the team “ZJU3DV” in IMC2023. We got some improvement on church locally, but no boost on LB.\n•\t**Disambiguation.** There are multiple papers that tackle symmetrical scenes, we have tried [Doppelgangers](https://arxiv.org/pdf/2309.02420) and [Yan](https://yanqingan.github.io/docs/cvpr17_distinguishing.pdf) methods, the latter had a tiny local improvement on CV but a drop on LB, so we decided to skip them.\n•\t**3D RANSAC cleaning.** We utilized 3D RANSAC cleaning to eliminate incorrect matches from symmetric scenes, taking advantage of depth information. For every image, we created depth masks. Then, for each pair of matches, we considered the x,z projection, which is similar to a bird's-eye view of the matched points. Subsequently, we employed RANSAC to identify and remove the outliers.\n•\t**Image motion de-blurring.** We recognized that the quality of some images in train was reduced possibly due to the camera shake. To mitigate this effect, we experimented with the single image restoration de-blur method but did not get noticeable improvement on LB.\n\n---\n# 2. DIP Module [transparent scenes]\n\nWe quickly realized that the SfM pipeline with image matching and COLMAP \"as is\" does not work with transparent scenes. Therefore, we decided to experiment with different strategies. We assumed that the direct pose estimation of the object in each image may help to compute the rotation matrix. We started to make some assumptions about the transparent category: “Given very low metric thresholds for the transparent category, the camera positions have to be somewhere very close to the object, and most probably the object is shot from all sides”. \nSo, we simply throw the cameras on a circle around the object, that looks towards the object as in this image.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F520a9375c05ebdb0c8039d38c0134006%2Fdip_camera_pose_circle.png?generation=1717536931811867&alt=media\" height=\"300px\" alt=\"camera pose circle\">\n \nLately, we’ve also found a very interesting scientific paper which describes approaches to solve transparent categories, that included a tiny dataset of 4 transparent objects, yet we didn’t find it publicly available.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F0b7bb8a9fb2a11fb9c2a2346d567a660%2Ftransparent_objects_poses_paper.png?generation=1717537022273030&alt=media\" height=\"240px\" alt=\"transparent_objects_paper\">\n \nImage is taken from the [paper](https://isprs-archives.copernicus.org/articles/XLVIII-2-W2-2022/77/2022/isprs-archives-XLVIII-2-W2-2022-77-2022.pdf) [Morelli, L., et al. \"Orientation of Images with Low Contrast Textures and Transparent Objects.\" The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences 48 (2022): 77-84.] which shares interesting experiment results about transparent scenes.\n\nWe assumed that if we manage to solve these transparent scenes to some extent, we would be good with LB. For validation purposes, we even **generated our own transparent scene dataset** with some plastic & glass bottles:\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F241f6199e349cc56c206f2b075484c53%2Fvlad_two_bottles_dataset.png?generation=1717537121183118&alt=media\" height=\"300px\" alt=\"two_bottles_dataset\">\n  \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2F42bf710a8be051e95f38b03ee132c408%2Fvlad%20imc24%20bottles%20slideshow.png?generation=1717537147849661&alt=media\" height=\"300px\" alt=\"bottles_slideshow\">\n\nBasically, we came up with the following two approaches:\n\n### Approach #1 \nTo make our assumption work, we needed to sort the images in the right order, then assign the camera pose for each image as to be located on the ideal circle around the object with even spaces between them (e.g. for 36 images its 10 degrees). We tried several approaches to sort the images, all of them include solving the Travelling Salesman Problem as the final step:\n1.\t**Optical Flow.** Calculate the magnitude of OF for each image pair and assign each pair a weight equal to the standard deviation of magnitude.\n2.\t**Pixel-level difference.** A simple difference of grayscaled images, the weight for each pair is equal to the difference value.\n3.\t**SSIM score.** Calculating the [SSIM index](https://en.wikipedia.org/wiki/Structural_similarity_index_measure) for each pair and assigning pair weight equal to 1 - ssim.\n4.\t**ALIKED+LG matching.** Again calculating the number of matches for each pair, and assigning a pair weight equal to (1 / num_matches).\n\nIn each of the four above-mentioned approaches the pair weight indicates how close images are to each other (the less the weight, the more similar the images are). We built a distance matrix based on these pair weights and solved the final ordering problem through TSP.\n\n### Approach #2 \nA reasonable score was obtained by estimating the order of the images, pairing them with the previous and following images, and performing a matching process at a very high resolution (4096~).\nImage order was estimated based on the number of matches. Experimental results showed that the number of matches tended to be higher for the before and after images. This tendency is used to estimate image order using a kNN-like method.\n\n**CV results**, cylinder: 77.78%, cup: 41.92%.\n**LB boost** using a transparent trick: +0.03\n\n## Event timeline with progression:\n•\t**Optical Flow** for image ordering using the mean average of the flow which is basically presented by pixel displacement u and v. Standard deviation scored better.\n•\t**Grayscale pixel difference.**\n•\t**SIMM score** in a range [-1, 1], where 1 indicates perfect similarity, 0 indicates no similarity, and -1 indicates perfect anti-correlation.\n•\t**SIMM + \"Matching Flow\"** is the final selected ensemble. It gave us **+0.09** on public LB.\n**CV results**, cylinder: **~92%**, cup: **~62%**\n\n### Ideas that did not work out:\n•\tEdge extraction methods (e.g., CLAHE, Canny, Gaussian smoothing, Laplacian).\n•\tDepth masks.\n•\tSegmentation of transparent objects before flow calculation.\n\n---\n## Facts and Numbers:\n\n•\t**0.28/0.25** is the final score in public/private LB\n•\t**0.28**4871 is the best public score\n•\t**0.26**7720 is the best private score\n•\t**+0.09** with a transparent trick in public/private LB\n•\t**16914 sec (4h 42min)** is the execution time of the selected best submission\n•\t**2 scenes with <50% of registered images** yet in LB\n•\t**335** submissions by our team\n•\t**∞ cups of coffee and dedication**\n\n## Local Validation:\n\nOur validation image dataset used the following data for each scene:\n•\t**\"church\", \"lizard\"**: Image sets that were provided as test data in the competition.\n•\t**\"dioscuri\", \"multi-temporal-temple\"**: Image sets that were provided as train data.\n•\t**\"pond\"**: A set of 65 images randomly extracted from the training dataset.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5065877%2Fb0388aa318697244ad212f25bffbccc9%2Fyumeneko_final_metrics_jafaar_v2.PNG?generation=1717546332981607&alt=media\" alt=\"final_metrics\">\n**Final metric -> mAA=41.23 (Best LB) / 43.50 (Best CV)**\n\n\n**UPD** We shared our [winning notebook](https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution)\n",
    "2871706": "We have released the [code](https://www.kaggle.com/code/vostankovich/imc2024-1st-place-solution)",
    "2856950": "Wow great work guys! I must admit it is both exciting and painful to see how successful using similarity to differentiate glass and using separate ideas on those scenes was. Based on some research I did very early on I was able to find that this could provide a value and it did match up with the 0.03 gain that some top teams got early on. I put a ton of time into getting a solution with it to work but I had an error in submission that I simply could not solve for. Glad to see the idea worked as well as I thought it could. Also great to see just how elegant the solution you crafted was a very well deserved win!",
    "2858274": "Impressive solution! The way you combined ALIKED extractors and LightGlue matchers with high-resolution images is brilliant. Your dedication to optimizing the pipeline and handling transparent scenes is evident. Great job!",
    "2857817": "Congratulations on winning the competition! Your approach to handling transparent scenes is very impressive. We also processed transparent scenes in a similar way, trying to restore the order of the images, but unfortunately, we didn't make more in-depth efforts in this aspect. Congratulations again to you.",
    "2856425": "Congratulations!!! Thanks for sharing. Reading your approach makes me realize that I still have a lot to learn about 3D CV! But it's inspiring on the contrary, it means there's a direction to keep pumping up my skills.",
    "2855655": "Congratulations with the strong 1st place! And congratulations to @vostankovich and @kashiwaba for becoming new GMs!\n\nLoved the part where you created a validation set with just a phone camera, some bottles, and household objects 😂 You might be the first team on Kaggle in a while that has done some manual data collection like that. Gives me the vibes of old-school Computer Vision research before the deep-learning era 😁",
    "2866243": "Congratulations!! It is so impressive",
    "2861363": "Congrats, very impressive",
    "2860253": "congratulations for winning the compitation and very good work",
    "2860048": "Impressive and very insightful!",
    "2859808": "Congratulations on winning the competition! 🎉",
    "2858549": "Congratulations 🎉",
    "2857745": "Congratulations!! It is very helpful for me to understand the method of 3D image reconstruction. Thanks for sharing.",
    "2857492": "This is amazing. Great work!",
    "2856649": "Congr, This is amazing, thanks for sharing. ",
    "2856385": "Innovative approach, excellent results speak for themselves 👍. Can't wait to learn more details on your approach in slides and the paper at CVPR. Impressive trick on transparent scene and great write-up!",
    "2855789": "Nice work! Congratulations 🎉",
    "2897171": "the transparency approach was very well executed, i will try to read that paper you mentioned.",
    "2867727": "Congratulations!! Your explanation was detailed and amazing.",
    "2868654": "congrats man!",
    "2862071": "very impressive work, well done!",
    "2862046": "Congratulations!",
    "2856046": "",
    "2855626": "",
    "2855703": "thanks for sharing.  ",
    "2860235": "Thank you for your sharing"
  }
}