{
  "id": 572714,
  "title": "Error in image labeling in training data",
  "url": "/competitions/image-matching-challenge-2025/discussion/572714",
  "author_name": "",
  "post_date": "2025-04-10T23:23:04.440114800Z",
  "votes": -1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>In some folders, the images that are separated by labeling have not been correctly separated.<br>\nFor example, in the training data and in the 'imc2023_heritage' folder, the number of image clusters is higher.</p>",
  "messages": [
    {
      "id": "3176063",
      "postDate": "04/10/2025 23:23:04",
      "content": "<p>In some folders, the images that are separated by labeling have not been correctly separated.<br>\nFor example, in the training data and in the 'imc2023_heritage' folder, the number of image clusters is higher.</p>",
      "rawMarkdown": "In some folders, the images that are separated by labeling have not been correctly separated.\n\nFor example, in the training data and in the 'imc2023_heritage' folder, the number of image clusters is higher.",
      "votes": null
    },
    {
      "id": "3176244",
      "postDate": "04/11/2025 06:02:07",
      "content": "<p>What do you mean? <code>imc2023_heritage</code> contains three scenes. You can look at <code>thresholds.csv</code>:</p>\n<pre><code>,cyprus,.;.;.;.;.;.\n,dioscuri,.;.;.;.;.;.\n,wall,.;.;.;.;.;.\n</code></pre>\n<p>On top of this there are outlier images that are not part of any scene. Look at <code>train_labels.csv</code>:</p>\n<pre><code>,outliers,outliers_img_9085.png,nan;nan;nan;nan;nan;nan;nan;nan;nan,nan;nan;nan\n,dioscuri,dioscuri_archive_0069.png,-.;.;.;.;.;.;-.;.;-.,-.;-.;.\n</code></pre>\n<p>Does this make sense?</p>",
      "rawMarkdown": "What do you mean? `imc2023_heritage` contains three scenes. You can look at `thresholds.csv`:\n\n```\nimc2023_heritage,cyprus,0.025;0.05;0.1;0.2;0.5;1.0\nimc2023_heritage,dioscuri,0.025;0.05;0.1;0.2;0.5;1.0\nimc2023_heritage,wall,0.025;0.05;0.1;0.2;0.5;1.0\n```\n\nOn top of this there are outlier images that are not part of any scene. Look at `train_labels.csv`:\n\n```\nimc2023_heritage,outliers,outliers_img_9085.png,nan;nan;nan;nan;nan;nan;nan;nan;nan,nan;nan;nan\nimc2023_heritage,dioscuri,dioscuri_archive_0069.png,-0.979195550;0.007108398;0.202794343;0.075310782;0.940738779;0.330664235;-0.188426009;0.339057548;-0.921702565,-0.291553977;-0.946437595;13.408536680\n```\n\nDoes this make sense?",
      "votes": null
    },
    {
      "id": "3176479",
      "postDate": "04/11/2025 11:50:01",
      "content": "<p>In the original data used to evaluate the code for the competition, the image names do not indicate the clusters, and the algorithm clusters solely based on the image features. Noisy data should not be related to each other, because the algorithm will group them into one cluster—unless we remove small clusters. In the submitted image, the clustering was performed on this same folder, and even among the data labeled as noisy and unrelated, there were actually related samples.</p>",
      "rawMarkdown": "In the original data used to evaluate the code for the competition, the image names do not indicate the clusters, and the algorithm clusters solely based on the image features. Noisy data should not be related to each other, because the algorithm will group them into one cluster—unless we remove small clusters. In the submitted image, the clustering was performed on this same folder, and even among the data labeled as noisy and unrelated, there were actually related samples.",
      "votes": null
    },
    {
      "id": "3176495",
      "postDate": "04/11/2025 12:10:36",
      "content": "<p>What is the algorithm are you referring to? The only source of truth is spatial location of the camera</p>",
      "rawMarkdown": "What is the algorithm are you referring to? The only source of truth is spatial location of the camera",
      "votes": null
    },
    {
      "id": "3176520",
      "postDate": "04/11/2025 12:29:43",
      "content": "<p>\"Does this mean that only images taken by the camera from different angles are considered the main clusters, and the rest are assumed to be noise data?\"</p>",
      "rawMarkdown": "\"Does this mean that only images taken by the camera from different angles are considered the main clusters, and the rest are assumed to be noise data?\"",
      "votes": null
    },
    {
      "id": "3176521",
      "postDate": "04/11/2025 12:33:33",
      "content": "<p>I'm sorry, but I still don't understand the question. We added the scene name to each image to make it easier for you to work with the training data (this information is obviously not available in the hidden test set). So for instance, all images like <code>imc2023_heritage/cyprus_*</code> correspond to the same scene, and your algorithm should ideally group them into a single cluster.</p>\n<p>If I understand your images correctly, it looks like you're producing very small clusters. You could tweak your settings, or aggregate them into larger clusters - I'm not sure what's the best way to go about it.</p>",
      "rawMarkdown": "I'm sorry, but I still don't understand the question. We added the scene name to each image to make it easier for you to work with the training data (this information is obviously not available in the hidden test set). So for instance, all images like `imc2023_heritage/cyprus_*` correspond to the same scene, and your algorithm should ideally group them into a single cluster.\n\nIf I understand your images correctly, it looks like you're producing very small clusters. You could tweak your settings, or aggregate them into larger clusters - I'm not sure what's the best way to go about it.",
      "votes": null
    },
    {
      "id": "3176522",
      "postDate": "04/11/2025 12:35:12",
      "content": "<p>A dataset contains N scenes, plus (optionally) outliers. All the images in one scene create a \"cluster\". Outliers are what you call \"noise\".</p>",
      "rawMarkdown": "A dataset contains N scenes, plus (optionally) outliers. All the images in one scene create a \"cluster\". Outliers are what you call \"noise\".",
      "votes": null
    },
    {
      "id": "3176526",
      "postDate": "04/11/2025 12:41:46",
      "content": "<p>A cluster is a set of images that can be registered together in terms of rotation an translation. If you have the same \"scene\" but there are holes (lack of overlap within images), you will be unable to register all together and you will obtain more clusters. The challenge is to maximize the overall correct registration.</p>",
      "rawMarkdown": "A cluster is a set of images that can be registered together in terms of rotation an translation. If you have the same \"scene\" but there are holes (lack of overlap within images), you will be unable to register all together and you will obtain more clusters. The challenge is to maximize the overall correct registration.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3176244,
      "author_name": "eduardtrulls",
      "author_url": "",
      "post_date": "04/11/2025 06:02:07",
      "content": "<p>What do you mean? <code>imc2023_heritage</code> contains three scenes. You can look at <code>thresholds.csv</code>:</p>\n<pre><code>,cyprus,.;.;.;.;.;.\n,dioscuri,.;.;.;.;.;.\n,wall,.;.;.;.;.;.\n</code></pre>\n<p>On top of this there are outlier images that are not part of any scene. Look at <code>train_labels.csv</code>:</p>\n<pre><code>,outliers,outliers_img_9085.png,nan;nan;nan;nan;nan;nan;nan;nan;nan,nan;nan;nan\n,dioscuri,dioscuri_archive_0069.png,-.;.;.;.;.;.;-.;.;-.,-.;-.;.\n</code></pre>\n<p>Does this make sense?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3176479,
          "author_name": "ali78kabirzadeh",
          "author_url": "",
          "post_date": "04/11/2025 11:50:01",
          "content": "<p>In the original data used to evaluate the code for the competition, the image names do not indicate the clusters, and the algorithm clusters solely based on the image features. Noisy data should not be related to each other, because the algorithm will group them into one cluster—unless we remove small clusters. In the submitted image, the clustering was performed on this same folder, and even among the data labeled as noisy and unrelated, there were actually related samples.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3176495,
              "author_name": "oldufo",
              "author_url": "",
              "post_date": "04/11/2025 12:10:36",
              "content": "<p>What is the algorithm are you referring to? The only source of truth is spatial location of the camera</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3176520,
                  "author_name": "ali78kabirzadeh",
                  "author_url": "",
                  "post_date": "04/11/2025 12:29:43",
                  "content": "<p>\"Does this mean that only images taken by the camera from different angles are considered the main clusters, and the rest are assumed to be noise data?\"</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3176522,
                      "author_name": "eduardtrulls",
                      "author_url": "",
                      "post_date": "04/11/2025 12:35:12",
                      "content": "<p>A dataset contains N scenes, plus (optionally) outliers. All the images in one scene create a \"cluster\". Outliers are what you call \"noise\".</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 3176521,
              "author_name": "eduardtrulls",
              "author_url": "",
              "post_date": "04/11/2025 12:33:33",
              "content": "<p>I'm sorry, but I still don't understand the question. We added the scene name to each image to make it easier for you to work with the training data (this information is obviously not available in the hidden test set). So for instance, all images like <code>imc2023_heritage/cyprus_*</code> correspond to the same scene, and your algorithm should ideally group them into a single cluster.</p>\n<p>If I understand your images correctly, it looks like you're producing very small clusters. You could tweak your settings, or aggregate them into larger clusters - I'm not sure what's the best way to go about it.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3176526,
                  "author_name": "fabiobellavia",
                  "author_url": "",
                  "post_date": "04/11/2025 12:41:46",
                  "content": "<p>A cluster is a set of images that can be registered together in terms of rotation an translation. If you have the same \"scene\" but there are holes (lack of overlap within images), you will be unable to register all together and you will obtain more clusters. The challenge is to maximize the overall correct registration.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3176063": "In some folders, the images that are separated by labeling have not been correctly separated.\n\nFor example, in the training data and in the 'imc2023_heritage' folder, the number of image clusters is higher.",
    "3176244": "What do you mean? `imc2023_heritage` contains three scenes. You can look at `thresholds.csv`:\n\n```\nimc2023_heritage,cyprus,0.025;0.05;0.1;0.2;0.5;1.0\nimc2023_heritage,dioscuri,0.025;0.05;0.1;0.2;0.5;1.0\nimc2023_heritage,wall,0.025;0.05;0.1;0.2;0.5;1.0\n```\n\nOn top of this there are outlier images that are not part of any scene. Look at `train_labels.csv`:\n\n```\nimc2023_heritage,outliers,outliers_img_9085.png,nan;nan;nan;nan;nan;nan;nan;nan;nan,nan;nan;nan\nimc2023_heritage,dioscuri,dioscuri_archive_0069.png,-0.979195550;0.007108398;0.202794343;0.075310782;0.940738779;0.330664235;-0.188426009;0.339057548;-0.921702565,-0.291553977;-0.946437595;13.408536680\n```\n\nDoes this make sense?",
    "3176479": "In the original data used to evaluate the code for the competition, the image names do not indicate the clusters, and the algorithm clusters solely based on the image features. Noisy data should not be related to each other, because the algorithm will group them into one cluster—unless we remove small clusters. In the submitted image, the clustering was performed on this same folder, and even among the data labeled as noisy and unrelated, there were actually related samples.",
    "3176495": "What is the algorithm are you referring to? The only source of truth is spatial location of the camera",
    "3176520": "\"Does this mean that only images taken by the camera from different angles are considered the main clusters, and the rest are assumed to be noise data?\"",
    "3176521": "I'm sorry, but I still don't understand the question. We added the scene name to each image to make it easier for you to work with the training data (this information is obviously not available in the hidden test set). So for instance, all images like `imc2023_heritage/cyprus_*` correspond to the same scene, and your algorithm should ideally group them into a single cluster.\n\nIf I understand your images correctly, it looks like you're producing very small clusters. You could tweak your settings, or aggregate them into larger clusters - I'm not sure what's the best way to go about it.",
    "3176522": "A dataset contains N scenes, plus (optionally) outliers. All the images in one scene create a \"cluster\". Outliers are what you call \"noise\".",
    "3176526": "A cluster is a set of images that can be registered together in terms of rotation an translation. If you have the same \"scene\" but there are holes (lack of overlap within images), you will be unable to register all together and you will obtain more clusters. The challenge is to maximize the overall correct registration."
  },
  "source": "meta"
}