{
  "id": 571224,
  "title": "\"outliers\" present in train scenes — how should we treat them?",
  "url": "/competitions/image-matching-challenge-2025/discussion/571224",
  "author_name": "",
  "post_date": "2025-04-02T01:56:53.974432600Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<h2>\"outliers\" scene in train_labels.csv</h2>\n<p>While inspecting <code>train_labels.csv</code>, I noticed that some training images, in <code>imc2023_heritage</code> dataset, are assigned to a scene called <strong>\"outliers\"</strong>.  <br>\nThis caught my attention because I assumed all training images had well-defined scene labels for supervised learning.</p>\n<p>Here’s a quick way to check:</p>\n<pre><code> pandas  pd\ndf = pd.read_csv()\n(df[df[] == ].head())\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2Feb6e90c32e90e7b89cf5d397c13ee68a%2F2025-04-02%20105212.png?generation=1743558899738703&amp;alt=media\" alt=\"\"></p>\n<h2>My understanding of \"outliers\"</h2>\n<p>Based on the dataset description, I thought \"outliers\" referred only to <strong>test images</strong>.<br>\nHowever, seeing \"outliers\" also in the train set makes me wonder.</p>\n<h2>Questions</h2>\n<ul>\n<li><p>Are you using images labeled \"outliers\" in your training pipeline?</p></li>\n<li><p>Do you treat them as a valid scene class, or ignore them entirely?</p></li>\n<li><p>Any ideas on how to handle similar outliers in the test phase?</p></li>\n</ul>\n<p>I'd love to hear how others are dealing with this!</p>",
  "messages": [
    {
      "id": "3167979",
      "postDate": "04/02/2025 01:56:53",
      "content": "<h2>\"outliers\" scene in train_labels.csv</h2>\n<p>While inspecting <code>train_labels.csv</code>, I noticed that some training images, in <code>imc2023_heritage</code> dataset, are assigned to a scene called <strong>\"outliers\"</strong>.  <br>\nThis caught my attention because I assumed all training images had well-defined scene labels for supervised learning.</p>\n<p>Here’s a quick way to check:</p>\n<pre><code> pandas  pd\ndf = pd.read_csv()\n(df[df[] == ].head())\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2Feb6e90c32e90e7b89cf5d397c13ee68a%2F2025-04-02%20105212.png?generation=1743558899738703&amp;alt=media\" alt=\"\"></p>\n<h2>My understanding of \"outliers\"</h2>\n<p>Based on the dataset description, I thought \"outliers\" referred only to <strong>test images</strong>.<br>\nHowever, seeing \"outliers\" also in the train set makes me wonder.</p>\n<h2>Questions</h2>\n<ul>\n<li><p>Are you using images labeled \"outliers\" in your training pipeline?</p></li>\n<li><p>Do you treat them as a valid scene class, or ignore them entirely?</p></li>\n<li><p>Any ideas on how to handle similar outliers in the test phase?</p></li>\n</ul>\n<p>I'd love to hear how others are dealing with this!</p>",
      "rawMarkdown": "## \"outliers\" scene in train_labels.csv\nWhile inspecting `train_labels.csv`, I noticed that some training images, in `imc2023_heritage` dataset, are assigned to a scene called **\"outliers\"**.  \nThis caught my attention because I assumed all training images had well-defined scene labels for supervised learning.\n\nHere’s a quick way to check:\n\n```python\nimport pandas as pd\ndf = pd.read_csv(\"/kaggle/input/image-matching-challenge-2025/train_labels.csv\")\nprint(df[df['scene'] == 'outliers'].head())\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2Feb6e90c32e90e7b89cf5d397c13ee68a%2F2025-04-02%20105212.png?generation=1743558899738703&alt=media)\n\n## My understanding of \"outliers\"\nBased on the dataset description, I thought \"outliers\" referred only to **test images**.\nHowever, seeing \"outliers\" also in the train set makes me wonder.\n\n## Questions\n- Are you using images labeled \"outliers\" in your training pipeline?\n\n- Do you treat them as a valid scene class, or ignore them entirely?\n\n- Any ideas on how to handle similar outliers in the test phase?\n\nI'd love to hear how others are dealing with this!",
      "votes": null
    },
    {
      "id": "3168041",
      "postDate": "04/02/2025 04:07:19",
      "content": "<p>The training data mirrors the test data: every dataset contains multiple scenes, grouped together, and \"outliers\". \"Outliers\" (a better word would be \"distractors\", because \"outliers\" have different meanings in the image matching literature) are samples that do not correspond to any others, so they should not be part of your scene reconstructions.</p>\n<p>In previous years winning solutions have focused on combining off-the-shelf ML models and 3D reconstruction frameworks in smart ways, rather than fine-tuning them. That's not to say that you CAN'T do that: you may actually want to, particularly since there's a new dimension this year in that you have to cluster scenes together and remove distractors / outliers.</p>\n<p>But in principle, I would look at the training set for this competition less as a collection of data to do supervised learning with, and more as a way to get a better understanding of the problem before submitting directly on the hidden test set. This has been one of our friction points in previous editions: I believe this should help.</p>",
      "rawMarkdown": "The training data mirrors the test data: every dataset contains multiple scenes, grouped together, and \"outliers\". \"Outliers\" (a better word would be \"distractors\", because \"outliers\" have different meanings in the image matching literature) are samples that do not correspond to any others, so they should not be part of your scene reconstructions.\n\nIn previous years winning solutions have focused on combining off-the-shelf ML models and 3D reconstruction frameworks in smart ways, rather than fine-tuning them. That's not to say that you CAN'T do that: you may actually want to, particularly since there's a new dimension this year in that you have to cluster scenes together and remove distractors / outliers.\n\nBut in principle, I would look at the training set for this competition less as a collection of data to do supervised learning with, and more as a way to get a better understanding of the problem before submitting directly on the hidden test set. This has been one of our friction points in previous editions: I believe this should help.",
      "votes": null
    },
    {
      "id": "3171222",
      "postDate": "04/05/2025 12:19:50",
      "content": "<p>Thank you for your reply.<br>\nI’ve resolved my confusion by carefully reading the dataset description — my apologies for the oversight.<br>\nI now understand that when you say \"the training data mirrors the test data\", it means that both are structured in similar ways with datasets, scenes, and outliers.</p>\n<p>As far as I understand, the only training data that can be truly used for learning are the images corresponding to scenes listed in <code>train_labels.csv</code>. That makes sense now.</p>",
      "rawMarkdown": "Thank you for your reply.\nI’ve resolved my confusion by carefully reading the dataset description — my apologies for the oversight.\nI now understand that when you say \"the training data mirrors the test data\", it means that both are structured in similar ways with datasets, scenes, and outliers.\n\nAs far as I understand, the only training data that can be truly used for learning are the images corresponding to scenes listed in `train_labels.csv`. That makes sense now.",
      "votes": null
    },
    {
      "id": "3171397",
      "postDate": "04/05/2025 16:39:54",
      "content": "<p>It does seem that some of the outliers are of the same scene though.  For instance, <code>imc2024_lizard_pond/outliers_img_20230617_181536.png</code> and <code>imc2024_lizard_pond/outliers_img_20230617_181545.png</code> certainly seem like the same scene to me.  (and it seems like there are 24 images of that scene all with a scene of <code>outliers</code>)</p>\n<p>Is this a labeling mistake or is this some distinction that makes a group of images of the same location a \"scene\" vs all outliers?</p>",
      "rawMarkdown": "It does seem that some of the outliers are of the same scene though.  For instance, `imc2024_lizard_pond/outliers_img_20230617_181536.png` and `imc2024_lizard_pond/outliers_img_20230617_181545.png` certainly seem like the same scene to me.  (and it seems like there are 24 images of that scene all with a scene of `outliers`)\n\nIs this a labeling mistake or is this some distinction that makes a group of images of the same location a \"scene\" vs all outliers?",
      "votes": null
    },
    {
      "id": "3171962",
      "postDate": "04/06/2025 09:22:21",
      "content": "<p>Yes, they are. They are outliers for the photos from the other scene, and they can form their own cluster. <br>\nIf you cluster them, it is not a mistake, it is just that the poses may not matter. </p>",
      "rawMarkdown": "Yes, they are. They are outliers for the photos from the other scene, and they can form their own cluster. \nIf you cluster them, it is not a mistake, it is just that the poses may not matter.",
      "votes": null
    },
    {
      "id": "3171970",
      "postDate": "04/06/2025 09:29:10",
      "content": "<p><a href=\"https://www.kaggle.com/pabyrnes\" target=\"_blank\">@pabyrnes</a> Please notice also that this is smallest cluster in the dataset. In any case the score will be maximized since image tagged outliers are not considered within the score formulation (see the relative function in the utils package).</p>",
      "rawMarkdown": "pabyrnes Please notice also that this is smallest cluster in the dataset. In any case the score will be maximized since image tagged outliers are not considered within the score formulation (see the relative function in the utils package).",
      "votes": null
    },
    {
      "id": "3186355",
      "postDate": "04/24/2025 14:33:23",
      "content": "<p>Image tagged outliers by us won't be considered in evaluation?</p>",
      "rawMarkdown": "Image tagged outliers by us won't be considered in evaluation?",
      "votes": null
    },
    {
      "id": "3186376",
      "postDate": "04/24/2025 15:19:08",
      "content": "<p>No, they will be excluded in the evaluation. Please see the metric script <a href=\"https://www.kaggle.com/datasets/eduardtrulls/imc25-utils\" target=\"_blank\">here</a> for details about how the metric works.</p>",
      "rawMarkdown": "No, they will be excluded in the evaluation. Please see the metric script [here](https://www.kaggle.com/datasets/eduardtrulls/imc25-utils) for details about how the metric works.",
      "votes": null
    },
    {
      "id": "3186394",
      "postDate": "04/24/2025 15:38:10",
      "content": "<p>Thanks for the clarification 😅</p>",
      "rawMarkdown": "Thanks for the clarification 😅",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3168041,
      "author_name": "eduardtrulls",
      "author_url": "",
      "post_date": "04/02/2025 04:07:19",
      "content": "<p>The training data mirrors the test data: every dataset contains multiple scenes, grouped together, and \"outliers\". \"Outliers\" (a better word would be \"distractors\", because \"outliers\" have different meanings in the image matching literature) are samples that do not correspond to any others, so they should not be part of your scene reconstructions.</p>\n<p>In previous years winning solutions have focused on combining off-the-shelf ML models and 3D reconstruction frameworks in smart ways, rather than fine-tuning them. That's not to say that you CAN'T do that: you may actually want to, particularly since there's a new dimension this year in that you have to cluster scenes together and remove distractors / outliers.</p>\n<p>But in principle, I would look at the training set for this competition less as a collection of data to do supervised learning with, and more as a way to get a better understanding of the problem before submitting directly on the hidden test set. This has been one of our friction points in previous editions: I believe this should help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3171222,
          "author_name": "thoth000",
          "author_url": "",
          "post_date": "04/05/2025 12:19:50",
          "content": "<p>Thank you for your reply.<br>\nI’ve resolved my confusion by carefully reading the dataset description — my apologies for the oversight.<br>\nI now understand that when you say \"the training data mirrors the test data\", it means that both are structured in similar ways with datasets, scenes, and outliers.</p>\n<p>As far as I understand, the only training data that can be truly used for learning are the images corresponding to scenes listed in <code>train_labels.csv</code>. That makes sense now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3171397,
          "author_name": "pabyrnes",
          "author_url": "",
          "post_date": "04/05/2025 16:39:54",
          "content": "<p>It does seem that some of the outliers are of the same scene though.  For instance, <code>imc2024_lizard_pond/outliers_img_20230617_181536.png</code> and <code>imc2024_lizard_pond/outliers_img_20230617_181545.png</code> certainly seem like the same scene to me.  (and it seems like there are 24 images of that scene all with a scene of <code>outliers</code>)</p>\n<p>Is this a labeling mistake or is this some distinction that makes a group of images of the same location a \"scene\" vs all outliers?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3171962,
              "author_name": "oldufo",
              "author_url": "",
              "post_date": "04/06/2025 09:22:21",
              "content": "<p>Yes, they are. They are outliers for the photos from the other scene, and they can form their own cluster. <br>\nIf you cluster them, it is not a mistake, it is just that the poses may not matter. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3171970,
                  "author_name": "fabiobellavia",
                  "author_url": "",
                  "post_date": "04/06/2025 09:29:10",
                  "content": "<p><a href=\"https://www.kaggle.com/pabyrnes\" target=\"_blank\">@pabyrnes</a> Please notice also that this is smallest cluster in the dataset. In any case the score will be maximized since image tagged outliers are not considered within the score formulation (see the relative function in the utils package).</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3186355,
                      "author_name": "bhavesjain",
                      "author_url": "",
                      "post_date": "04/24/2025 14:33:23",
                      "content": "<p>Image tagged outliers by us won't be considered in evaluation?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3186376,
                          "author_name": "fabiobellavia",
                          "author_url": "",
                          "post_date": "04/24/2025 15:19:08",
                          "content": "<p>No, they will be excluded in the evaluation. Please see the metric script <a href=\"https://www.kaggle.com/datasets/eduardtrulls/imc25-utils\" target=\"_blank\">here</a> for details about how the metric works.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3186394,
                              "author_name": "bhavesjain",
                              "author_url": "",
                              "post_date": "04/24/2025 15:38:10",
                              "content": "<p>Thanks for the clarification 😅</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3167979": "## \"outliers\" scene in train_labels.csv\nWhile inspecting `train_labels.csv`, I noticed that some training images, in `imc2023_heritage` dataset, are assigned to a scene called **\"outliers\"**.  \nThis caught my attention because I assumed all training images had well-defined scene labels for supervised learning.\n\nHere’s a quick way to check:\n\n```python\nimport pandas as pd\ndf = pd.read_csv(\"/kaggle/input/image-matching-challenge-2025/train_labels.csv\")\nprint(df[df['scene'] == 'outliers'].head())\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6855922%2Feb6e90c32e90e7b89cf5d397c13ee68a%2F2025-04-02%20105212.png?generation=1743558899738703&alt=media)\n\n## My understanding of \"outliers\"\nBased on the dataset description, I thought \"outliers\" referred only to **test images**.\nHowever, seeing \"outliers\" also in the train set makes me wonder.\n\n## Questions\n- Are you using images labeled \"outliers\" in your training pipeline?\n\n- Do you treat them as a valid scene class, or ignore them entirely?\n\n- Any ideas on how to handle similar outliers in the test phase?\n\nI'd love to hear how others are dealing with this!",
    "3168041": "The training data mirrors the test data: every dataset contains multiple scenes, grouped together, and \"outliers\". \"Outliers\" (a better word would be \"distractors\", because \"outliers\" have different meanings in the image matching literature) are samples that do not correspond to any others, so they should not be part of your scene reconstructions.\n\nIn previous years winning solutions have focused on combining off-the-shelf ML models and 3D reconstruction frameworks in smart ways, rather than fine-tuning them. That's not to say that you CAN'T do that: you may actually want to, particularly since there's a new dimension this year in that you have to cluster scenes together and remove distractors / outliers.\n\nBut in principle, I would look at the training set for this competition less as a collection of data to do supervised learning with, and more as a way to get a better understanding of the problem before submitting directly on the hidden test set. This has been one of our friction points in previous editions: I believe this should help.",
    "3171222": "Thank you for your reply.\nI’ve resolved my confusion by carefully reading the dataset description — my apologies for the oversight.\nI now understand that when you say \"the training data mirrors the test data\", it means that both are structured in similar ways with datasets, scenes, and outliers.\n\nAs far as I understand, the only training data that can be truly used for learning are the images corresponding to scenes listed in `train_labels.csv`. That makes sense now.",
    "3171397": "It does seem that some of the outliers are of the same scene though.  For instance, `imc2024_lizard_pond/outliers_img_20230617_181536.png` and `imc2024_lizard_pond/outliers_img_20230617_181545.png` certainly seem like the same scene to me.  (and it seems like there are 24 images of that scene all with a scene of `outliers`)\n\nIs this a labeling mistake or is this some distinction that makes a group of images of the same location a \"scene\" vs all outliers?",
    "3171962": "Yes, they are. They are outliers for the photos from the other scene, and they can form their own cluster. \nIf you cluster them, it is not a mistake, it is just that the poses may not matter.",
    "3171970": "pabyrnes Please notice also that this is smallest cluster in the dataset. In any case the score will be maximized since image tagged outliers are not considered within the score formulation (see the relative function in the utils package).",
    "3186355": "Image tagged outliers by us won't be considered in evaluation?",
    "3186376": "No, they will be excluded in the evaluation. Please see the metric script [here](https://www.kaggle.com/datasets/eduardtrulls/imc25-utils) for details about how the metric works.",
    "3186394": "Thanks for the clarification 😅"
  },
  "source": "meta"
}