{
  "id": 510373,
  "title": "[39th place] Superpoint+Lightglue & Keynet_Affnet_Hardnet_Adalam",
  "url": "/competitions/image-matching-challenge-2024/discussion/510373",
  "author_name": "Roy Wei",
  "post_date": "2024-06-06T00:38:16.945000",
  "votes": 19,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I would like to first thank our organizers and Kaggle staff for organizing this competition, authors of Hloc, Lightglue and pycolmap, my teammate <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> for his great work. Even though the preliminary standing doesn't meet our expectation, we have learnt a lot as a team. On behalf of the team, I also want to thank <a href=\"https://www.kaggle.com/maxchen303\" target=\"_blank\">@maxchen303</a> for the <a href=\"https://www.kaggle.com/code/maxchen303/imc2023-final-pub\" target=\"_blank\">notebook</a> he has published last year. It is the baseline with which we experiment, and I suppose a lot of teams owe him a big thank you!</p>\n<p>P.S. Our team has experimented with lots of models and has gained some insights. Plus, I will be in Seattle area in person during the CVPR, and I really would like to be invited and present our work in workshop☺️ <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a></p>\n<hr>\n<p>For an overview of Max's <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\" target=\"_blank\">baseline</a></p>\n<h3>Overview of our solution</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fae8aa8fa69386a5e24495a6eaf67f8bf%2FIMC2024_solution-4.png?generation=1717591891360933&amp;alt=media\" alt=\"solution\"></p>\n<h3>Thoughts</h3>\n<p>The greatest impression our team has of IMC2024 is <strong>randomness</strong>. It is very hard to set up a convincing local pipeline, because colmap's randomness accounts for differences in scores, obscuring the effectiveness of the trivial trick. Moreover, trick that has zero risk of harming the result, such as unregistered image localization, multiple reconstructions don't boost the score significantly. </p>\n<p>Despite having great potential in camera position estimation, Dense matching is inherently not compatible with the reconstruction of the entire scene as mentioned in last year winning <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">solution</a>. We have tried quantization and some miscellaneous tricks so that view tracks may be established but they all failed. </p>\n<h3>Local Result</h3>\n<p>Average mAA of an ensemble of Keynet+Affnets+Hardnet+Adalam<em>: 23.5% ~ 24%\nAverage mAA of SP+LG</em>: 25.7% ~ 26%<br>\nAverage mAA of LoFTR: &lt; 15%<br>\nAverage mAA of dense matchers + quantization: &lt;10%</p>\n<p>*Our experiment suggests that DoG + Affnet works the best among all, on a par with the ensemble in terms of accuracy. </p>\n<p>*higher score on average using SP+LG is attributed to much better result on <code>lizard</code> (more 10 additional images registered)  and slight boost on <code>pond</code>. Yet, it doesn't perform as well as Keynet+Affnets+Hardnet+Adalam in reconstructing architectures, like the rest of dataset. This also inspires us to reconstruct models once with and another time without SP+LG to squeeze some extra accuracy. </p>\n<h3>What works for us:</h3>\n<ol>\n<li><p>Choose a different and right resolution before sending to the model -   For Superpoint and Lightglue, the best resolution is 2000 when we use Hloc (We can confirm the boost only because score slight boost in every dataset). For Keynes-Affnet-Hardnet, we stack key points extracted from model based on 1024 and 1600. Essentially we have used same resolution for images smaller than 1024 since enlargement is set to false.</p></li>\n<li><p>Image localization using Hloc  - For datasets such as lizard and pond, many images are not registered. We manage to restore 2 or 3 images in extra under the easiest threshold. </p></li>\n<li><p>Multi-time reconstruction:  We set up different thresholds to filter out image pairs that have too few matched points, and then reconstruct them. We estimate the score of each model using this formula: <code>score = num_images * num_3dpoints / projection_error</code>. This helps us to select the best model but it does enhance the score much on LB compared to local (~1%)</p></li>\n</ol>\n<p>Incidentally, out best sub doesn't even use the later two tricks that locally work, suggesting again that being lucky is somehow the key 😬</p>\n<h3>What doesn't work</h3>\n<ol>\n<li><p>Dense matching always ends up with 0 image registered. This is not surprising as keypoints in an image vary when that image is paired up with different images, so there are always only two 2D coordinates that can be projected to a 3D point and the model thus couldn't be robustly constructed. Regretfully, RoMa has quite great performance as I check the image pairs it produces manually. We have tried <a href=\"https://github.com/zju3dv/DetectorFreeSfM\" target=\"_blank\">detector-free sfm</a> to handle this issue, but we aren't able to compile it on Kaggle, unfortunately. </p></li>\n<li><p><a href=\"https://github.com/google-research/omniglue\" target=\"_blank\">Omniglue</a> - I thought this will give a boost to the current SP+LG pipeline, but it doesn't. After trying different configurations, it still couldn't surpass LG on local datasets (approximately 30~10 more images are unregistered compared to LG). </p></li>\n<li><p>DISK + LG - concating with keynet+affnet+hardnet key points lowers the score. Moreover, we also find that <code>kornia.feature</code> has a bug that prevents model to be loaded from local machine. <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a></p></li>\n</ol>\n<pre><code> ():\n     ():\n        (KF.LightGlueMatcher, self).__init__(feature_name)\n        self.feature_name = feature_name\n        self.params = params\n        self.matcher = KF.LightGlue(, **params)\n\n    KF.LightGlueMatcher.__init__ = new_init\n</code></pre>\n<h3>What we didn't do</h3>\n<ol>\n<li>Tricks on glass - it seems that higher resolution gives a better performance on glass in local testing, but we have never implemented it online. We have also made the assumption that camera position for glassware images in test set is similar to camera position of the one in train set, except that it's not in order. Yet, we didn't know how to restore the order… (hence thumb up to 1st place for their innovative approaches!)</li>\n</ol>\n<p>If you have any question about our solution, you are welcome to comment below. ❤️</p>",
  "messages": [
    {
      "id": 2857529,
      "postDate": "2024-06-06T00:38:16.947Z",
      "content": "<p>I would like to first thank our organizers and Kaggle staff for organizing this competition, authors of Hloc, Lightglue and pycolmap, my teammate <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> for his great work. Even though the preliminary standing doesn't meet our expectation, we have learnt a lot as a team. On behalf of the team, I also want to thank <a href=\"https://www.kaggle.com/maxchen303\" target=\"_blank\">@maxchen303</a> for the <a href=\"https://www.kaggle.com/code/maxchen303/imc2023-final-pub\" target=\"_blank\">notebook</a> he has published last year. It is the baseline with which we experiment, and I suppose a lot of teams owe him a big thank you!</p>\n<p>P.S. Our team has experimented with lots of models and has gained some insights. Plus, I will be in Seattle area in person during the CVPR, and I really would like to be invited and present our work in workshop☺️ <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a></p>\n<hr>\n<p>For an overview of Max's <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045\" target=\"_blank\">baseline</a></p>\n<h3>Overview of our solution</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fae8aa8fa69386a5e24495a6eaf67f8bf%2FIMC2024_solution-4.png?generation=1717591891360933&amp;alt=media\" alt=\"solution\"></p>\n<h3>Thoughts</h3>\n<p>The greatest impression our team has of IMC2024 is <strong>randomness</strong>. It is very hard to set up a convincing local pipeline, because colmap's randomness accounts for differences in scores, obscuring the effectiveness of the trivial trick. Moreover, trick that has zero risk of harming the result, such as unregistered image localization, multiple reconstructions don't boost the score significantly. </p>\n<p>Despite having great potential in camera position estimation, Dense matching is inherently not compatible with the reconstruction of the entire scene as mentioned in last year winning <a href=\"https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407\" target=\"_blank\">solution</a>. We have tried quantization and some miscellaneous tricks so that view tracks may be established but they all failed. </p>\n<h3>Local Result</h3>\n<p>Average mAA of an ensemble of Keynet+Affnets+Hardnet+Adalam<em>: 23.5% ~ 24%\nAverage mAA of SP+LG</em>: 25.7% ~ 26%<br>\nAverage mAA of LoFTR: &lt; 15%<br>\nAverage mAA of dense matchers + quantization: &lt;10%</p>\n<p>*Our experiment suggests that DoG + Affnet works the best among all, on a par with the ensemble in terms of accuracy. </p>\n<p>*higher score on average using SP+LG is attributed to much better result on <code>lizard</code> (more 10 additional images registered)  and slight boost on <code>pond</code>. Yet, it doesn't perform as well as Keynet+Affnets+Hardnet+Adalam in reconstructing architectures, like the rest of dataset. This also inspires us to reconstruct models once with and another time without SP+LG to squeeze some extra accuracy. </p>\n<h3>What works for us:</h3>\n<ol>\n<li><p>Choose a different and right resolution before sending to the model -   For Superpoint and Lightglue, the best resolution is 2000 when we use Hloc (We can confirm the boost only because score slight boost in every dataset). For Keynes-Affnet-Hardnet, we stack key points extracted from model based on 1024 and 1600. Essentially we have used same resolution for images smaller than 1024 since enlargement is set to false.</p></li>\n<li><p>Image localization using Hloc  - For datasets such as lizard and pond, many images are not registered. We manage to restore 2 or 3 images in extra under the easiest threshold. </p></li>\n<li><p>Multi-time reconstruction:  We set up different thresholds to filter out image pairs that have too few matched points, and then reconstruct them. We estimate the score of each model using this formula: <code>score = num_images * num_3dpoints / projection_error</code>. This helps us to select the best model but it does enhance the score much on LB compared to local (~1%)</p></li>\n</ol>\n<p>Incidentally, out best sub doesn't even use the later two tricks that locally work, suggesting again that being lucky is somehow the key 😬</p>\n<h3>What doesn't work</h3>\n<ol>\n<li><p>Dense matching always ends up with 0 image registered. This is not surprising as keypoints in an image vary when that image is paired up with different images, so there are always only two 2D coordinates that can be projected to a 3D point and the model thus couldn't be robustly constructed. Regretfully, RoMa has quite great performance as I check the image pairs it produces manually. We have tried <a href=\"https://github.com/zju3dv/DetectorFreeSfM\" target=\"_blank\">detector-free sfm</a> to handle this issue, but we aren't able to compile it on Kaggle, unfortunately. </p></li>\n<li><p><a href=\"https://github.com/google-research/omniglue\" target=\"_blank\">Omniglue</a> - I thought this will give a boost to the current SP+LG pipeline, but it doesn't. After trying different configurations, it still couldn't surpass LG on local datasets (approximately 30~10 more images are unregistered compared to LG). </p></li>\n<li><p>DISK + LG - concating with keynet+affnet+hardnet key points lowers the score. Moreover, we also find that <code>kornia.feature</code> has a bug that prevents model to be loaded from local machine. <a href=\"https://www.kaggle.com/oldufo\" target=\"_blank\">@oldufo</a></p></li>\n</ol>\n<pre><code> ():\n     ():\n        (KF.LightGlueMatcher, self).__init__(feature_name)\n        self.feature_name = feature_name\n        self.params = params\n        self.matcher = KF.LightGlue(, **params)\n\n    KF.LightGlueMatcher.__init__ = new_init\n</code></pre>\n<h3>What we didn't do</h3>\n<ol>\n<li>Tricks on glass - it seems that higher resolution gives a better performance on glass in local testing, but we have never implemented it online. We have also made the assumption that camera position for glassware images in test set is similar to camera position of the one in train set, except that it's not in order. Yet, we didn't know how to restore the order… (hence thumb up to 1st place for their innovative approaches!)</li>\n</ol>\n<p>If you have any question about our solution, you are welcome to comment below. ❤️</p>",
      "rawMarkdown": "I would like to first thank our organizers and Kaggle staff for organizing this competition, authors of Hloc, Lightglue and pycolmap, my teammate @cody11null for his great work. Even though the preliminary standing doesn't meet our expectation, we have learnt a lot as a team. On behalf of the team, I also want to thank @maxchen303 for the [notebook](https://www.kaggle.com/code/maxchen303/imc2023-final-pub) he has published last year. It is the baseline with which we experiment, and I suppose a lot of teams owe him a big thank you!\n\nP.S. Our team has experimented with lots of models and has gained some insights. Plus, I will be in Seattle area in person during the CVPR, and I really would like to be invited and present our work in workshop☺️ @oldufo\n\n---\n\nFor an overview of Max's [baseline](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045)\n\n### Overview of our solution\n![solution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fae8aa8fa69386a5e24495a6eaf67f8bf%2FIMC2024_solution-4.png?generation=1717591891360933&alt=media)\n\n### Thoughts\n\nThe greatest impression our team has of IMC2024 is **randomness**. It is very hard to set up a convincing local pipeline, because colmap's randomness accounts for differences in scores, obscuring the effectiveness of the trivial trick. Moreover, trick that has zero risk of harming the result, such as unregistered image localization, multiple reconstructions don't boost the score significantly. \n\nDespite having great potential in camera position estimation, Dense matching is inherently not compatible with the reconstruction of the entire scene as mentioned in last year winning [solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407). We have tried quantization and some miscellaneous tricks so that view tracks may be established but they all failed. \n\n### Local Result\n\nAverage mAA of an ensemble of Keynet+Affnets+Hardnet+Adalam*: 23.5% ~ 24%\nAverage mAA of SP+LG*: 25.7% ~ 26%\nAverage mAA of LoFTR: < 15%\nAverage mAA of dense matchers + quantization: <10%\n\n*Our experiment suggests that DoG + Affnet works the best among all, on a par with the ensemble in terms of accuracy. \n\n*higher score on average using SP+LG is attributed to much better result on `lizard` (more 10 additional images registered)  and slight boost on `pond`. Yet, it doesn't perform as well as Keynet+Affnets+Hardnet+Adalam in reconstructing architectures, like the rest of dataset. This also inspires us to reconstruct models once with and another time without SP+LG to squeeze some extra accuracy. \n\n### What works for us:\n1. Choose a different and right resolution before sending to the model -   For Superpoint and Lightglue, the best resolution is 2000 when we use Hloc (We can confirm the boost only because score slight boost in every dataset). For Keynes-Affnet-Hardnet, we stack key points extracted from model based on 1024 and 1600. Essentially we have used same resolution for images smaller than 1024 since enlargement is set to false.\n\n2. Image localization using Hloc  - For datasets such as lizard and pond, many images are not registered. We manage to restore 2 or 3 images in extra under the easiest threshold. \n\n3. Multi-time reconstruction:  We set up different thresholds to filter out image pairs that have too few matched points, and then reconstruct them. We estimate the score of each model using this formula: `score = num_images * num_3dpoints / projection_error`. This helps us to select the best model but it does enhance the score much on LB compared to local (~1%)\n\nIncidentally, out best sub doesn't even use the later two tricks that locally work, suggesting again that being lucky is somehow the key 😬\n\n### What doesn't work\n1. Dense matching always ends up with 0 image registered. This is not surprising as keypoints in an image vary when that image is paired up with different images, so there are always only two 2D coordinates that can be projected to a 3D point and the model thus couldn't be robustly constructed. Regretfully, RoMa has quite great performance as I check the image pairs it produces manually. We have tried [detector-free sfm](https://github.com/zju3dv/DetectorFreeSfM) to handle this issue, but we aren't able to compile it on Kaggle, unfortunately. \n\n2. [Omniglue](https://github.com/google-research/omniglue) - I thought this will give a boost to the current SP+LG pipeline, but it doesn't. After trying different configurations, it still couldn't surpass LG on local datasets (approximately 30~10 more images are unregistered compared to LG). \n\n3.  DISK + LG - concating with keynet+affnet+hardnet key points lowers the score. Moreover, we also find that `kornia.feature` has a bug that prevents model to be loaded from local machine. @oldufo\n\n```python\ndef modify_lightgluematcher():\n    def new_init(self, feature_name: str = \"disk\", params = {}):\n        super(KF.LightGlueMatcher, self).__init__(feature_name)\n        self.feature_name = feature_name\n        self.params = params\n        self.matcher = KF.LightGlue(None, **params)\n        \n    KF.LightGlueMatcher.__init__ = new_init\n```\n\n### What we didn't do\n\n1. Tricks on glass - it seems that higher resolution gives a better performance on glass in local testing, but we have never implemented it online. We have also made the assumption that camera position for glassware images in test set is similar to camera position of the one in train set, except that it's not in order. Yet, we didn't know how to restore the order... (hence thumb up to 1st place for their innovative approaches!)\n\n\nIf you have any question about our solution, you are welcome to comment below. ❤️\n",
      "votes": 19
    },
    {
      "id": 2857537,
      "postDate": "2024-06-06T00:48:40.780Z",
      "content": "<p>Hopefully not 59th by finalized results haha yes it was an interesting competition. One thing that really sucks is that glass being difficult. We found a way to statistically separate these images reliably before re noticed the large gap in resolution between the 2. The original plan was to use higher resolution and set configs specifically based off these glass images. This worked fine locally giving me a significant bump and I estimated it to give 0.02-0.04 early on in competition but struggled with implementation in submission for about 3 weeks before opting to move on to other ideas. I simply could not figure out the error.</p>",
      "rawMarkdown": "Hopefully not 59th by finalized results haha yes it was an interesting competition. One thing that really sucks is that glass being difficult. We found a way to statistically separate these images reliably before re noticed the large gap in resolution between the 2. The original plan was to use higher resolution and set configs specifically based off these glass images. This worked fine locally giving me a significant bump and I estimated it to give 0.02-0.04 early on in competition but struggled with implementation in submission for about 3 weeks before opting to move on to other ideas. I simply could not figure out the error.",
      "votes": 1
    },
    {
      "id": 2863206,
      "postDate": "2024-06-09T09:17:38.087Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2857537,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2024-06-06T00:48:40.780000",
      "content": "<p>Hopefully not 59th by finalized results haha yes it was an interesting competition. One thing that really sucks is that glass being difficult. We found a way to statistically separate these images reliably before re noticed the large gap in resolution between the 2. The original plan was to use higher resolution and set configs specifically based off these glass images. This worked fine locally giving me a significant bump and I estimated it to give 0.02-0.04 early on in competition but struggled with implementation in submission for about 3 weeks before opting to move on to other ideas. I simply could not figure out the error.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2863206,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-09T09:17:38.087000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2857529": "I would like to first thank our organizers and Kaggle staff for organizing this competition, authors of Hloc, Lightglue and pycolmap, my teammate @cody11null for his great work. Even though the preliminary standing doesn't meet our expectation, we have learnt a lot as a team. On behalf of the team, I also want to thank @maxchen303 for the [notebook](https://www.kaggle.com/code/maxchen303/imc2023-final-pub) he has published last year. It is the baseline with which we experiment, and I suppose a lot of teams owe him a big thank you!\n\nP.S. Our team has experimented with lots of models and has gained some insights. Plus, I will be in Seattle area in person during the CVPR, and I really would like to be invited and present our work in workshop☺️ @oldufo\n\n---\n\nFor an overview of Max's [baseline](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417045)\n\n### Overview of our solution\n![solution](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fae8aa8fa69386a5e24495a6eaf67f8bf%2FIMC2024_solution-4.png?generation=1717591891360933&alt=media)\n\n### Thoughts\n\nThe greatest impression our team has of IMC2024 is **randomness**. It is very hard to set up a convincing local pipeline, because colmap's randomness accounts for differences in scores, obscuring the effectiveness of the trivial trick. Moreover, trick that has zero risk of harming the result, such as unregistered image localization, multiple reconstructions don't boost the score significantly. \n\nDespite having great potential in camera position estimation, Dense matching is inherently not compatible with the reconstruction of the entire scene as mentioned in last year winning [solution](https://www.kaggle.com/competitions/image-matching-challenge-2023/discussion/417407). We have tried quantization and some miscellaneous tricks so that view tracks may be established but they all failed. \n\n### Local Result\n\nAverage mAA of an ensemble of Keynet+Affnets+Hardnet+Adalam*: 23.5% ~ 24%\nAverage mAA of SP+LG*: 25.7% ~ 26%\nAverage mAA of LoFTR: < 15%\nAverage mAA of dense matchers + quantization: <10%\n\n*Our experiment suggests that DoG + Affnet works the best among all, on a par with the ensemble in terms of accuracy. \n\n*higher score on average using SP+LG is attributed to much better result on `lizard` (more 10 additional images registered)  and slight boost on `pond`. Yet, it doesn't perform as well as Keynet+Affnets+Hardnet+Adalam in reconstructing architectures, like the rest of dataset. This also inspires us to reconstruct models once with and another time without SP+LG to squeeze some extra accuracy. \n\n### What works for us:\n1. Choose a different and right resolution before sending to the model -   For Superpoint and Lightglue, the best resolution is 2000 when we use Hloc (We can confirm the boost only because score slight boost in every dataset). For Keynes-Affnet-Hardnet, we stack key points extracted from model based on 1024 and 1600. Essentially we have used same resolution for images smaller than 1024 since enlargement is set to false.\n\n2. Image localization using Hloc  - For datasets such as lizard and pond, many images are not registered. We manage to restore 2 or 3 images in extra under the easiest threshold. \n\n3. Multi-time reconstruction:  We set up different thresholds to filter out image pairs that have too few matched points, and then reconstruct them. We estimate the score of each model using this formula: `score = num_images * num_3dpoints / projection_error`. This helps us to select the best model but it does enhance the score much on LB compared to local (~1%)\n\nIncidentally, out best sub doesn't even use the later two tricks that locally work, suggesting again that being lucky is somehow the key 😬\n\n### What doesn't work\n1. Dense matching always ends up with 0 image registered. This is not surprising as keypoints in an image vary when that image is paired up with different images, so there are always only two 2D coordinates that can be projected to a 3D point and the model thus couldn't be robustly constructed. Regretfully, RoMa has quite great performance as I check the image pairs it produces manually. We have tried [detector-free sfm](https://github.com/zju3dv/DetectorFreeSfM) to handle this issue, but we aren't able to compile it on Kaggle, unfortunately. \n\n2. [Omniglue](https://github.com/google-research/omniglue) - I thought this will give a boost to the current SP+LG pipeline, but it doesn't. After trying different configurations, it still couldn't surpass LG on local datasets (approximately 30~10 more images are unregistered compared to LG). \n\n3.  DISK + LG - concating with keynet+affnet+hardnet key points lowers the score. Moreover, we also find that `kornia.feature` has a bug that prevents model to be loaded from local machine. @oldufo\n\n```python\ndef modify_lightgluematcher():\n    def new_init(self, feature_name: str = \"disk\", params = {}):\n        super(KF.LightGlueMatcher, self).__init__(feature_name)\n        self.feature_name = feature_name\n        self.params = params\n        self.matcher = KF.LightGlue(None, **params)\n        \n    KF.LightGlueMatcher.__init__ = new_init\n```\n\n### What we didn't do\n\n1. Tricks on glass - it seems that higher resolution gives a better performance on glass in local testing, but we have never implemented it online. We have also made the assumption that camera position for glassware images in test set is similar to camera position of the one in train set, except that it's not in order. Yet, we didn't know how to restore the order... (hence thumb up to 1st place for their innovative approaches!)\n\n\nIf you have any question about our solution, you are welcome to comment below. ❤️\n",
    "2857537": "Hopefully not 59th by finalized results haha yes it was an interesting competition. One thing that really sucks is that glass being difficult. We found a way to statistically separate these images reliably before re noticed the large gap in resolution between the 2. The original plan was to use higher resolution and set configs specifically based off these glass images. This worked fine locally giving me a significant bump and I estimated it to give 0.02-0.04 early on in competition but struggled with implementation in submission for about 3 weeks before opting to move on to other ideas. I simply could not figure out the error.",
    "2863206": ""
  }
}