{
  "id": 573028,
  "title": "Help Needed: How Do Threshold Points Work in IMC 2025?",
  "url": "/competitions/image-matching-challenge-2025/discussion/573028",
  "author_name": "Bishwa3901",
  "post_date": "2025-04-13T05:20:58.857000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone! I'm new to the Image Matching Challenge 2025 and I'm having trouble understanding how the threshold points work. I'd really appreciate it if someone could explain it to me.</p>",
  "messages": [
    {
      "id": 3177658,
      "postDate": "2025-04-13T05:20:58.857Z",
      "content": "<p>Hi everyone! I'm new to the Image Matching Challenge 2025 and I'm having trouble understanding how the threshold points work. I'd really appreciate it if someone could explain it to me.</p>",
      "rawMarkdown": "Hi everyone! I'm new to the Image Matching Challenge 2025 and I'm having trouble understanding how the threshold points work. I'd really appreciate it if someone could explain it to me.",
      "votes": 4
    },
    {
      "id": 3184488,
      "postDate": "2025-04-22T06:41:16.840Z",
      "content": "<p>We follow the standard procedure of evaluating accuracy at different thresholds. We register the cameras against the ground truth and then compute how many are closer than a threshold, in spatial terms (you can think of this as classifying a camera as \"correct\" or \"incorrect\"). We do this for each camera, and average to compute the accuracy. Then average again over multiple thresholds: this is why the metric is called \"mean average accuracy\". Finally we average over the datasets. The thresholds are not the same for every dataset simply because the scale of the scene can vary a lot: think of a small object (where photos are captured within a few meters, or even less) vs a large square (which can span hundreds of meters).</p>\n<p>It may sound complicated but it's pretty simple. The metric code is public, and you can read it <a href=\"https://www.kaggle.com/datasets/eduardtrulls/imc25-utils\" target=\"_blank\">here</a>. This one is pretty long because it also solves the cluster assignment, so if you primarily want to understand how the thresholds work you can look at the <a href=\"https://www.kaggle.com/code/fabiobellavia/imc2024-3d-metric-evaluation-example\" target=\"_blank\">metric code</a> for last year's competition, which is essentially the same.</p>",
      "rawMarkdown": "We follow the standard procedure of evaluating accuracy at different thresholds. We register the cameras against the ground truth and then compute how many are closer than a threshold, in spatial terms (you can think of this as classifying a camera as \"correct\" or \"incorrect\"). We do this for each camera, and average to compute the accuracy. Then average again over multiple thresholds: this is why the metric is called \"mean average accuracy\". Finally we average over the datasets. The thresholds are not the same for every dataset simply because the scale of the scene can vary a lot: think of a small object (where photos are captured within a few meters, or even less) vs a large square (which can span hundreds of meters).\n\nIt may sound complicated but it's pretty simple. The metric code is public, and you can read it [here](https://www.kaggle.com/datasets/eduardtrulls/imc25-utils). This one is pretty long because it also solves the cluster assignment, so if you primarily want to understand how the thresholds work you can look at the [metric code](https://www.kaggle.com/code/fabiobellavia/imc2024-3d-metric-evaluation-example) for last year's competition, which is essentially the same.",
      "votes": 1
    },
    {
      "id": 3177930,
      "postDate": "2025-04-13T13:28:13.227Z",
      "content": "<p><a href=\"https://www.kaggle.com/pragyatripathiii23\" target=\"_blank\">@pragyatripathiii23</a> just answered my question. I'm sharing it here for anyone else who might be wondering the same thing.  So all the credits of the below explanation goes to <a href=\"https://www.kaggle.com/pragyatripathiii23\" target=\"_blank\">@pragyatripathiii23</a>  </p>\n<p>Explanation: <br>\n\"The threshold CSV file contains error limits (thresholds) defined for each scene in every dataset. These thresholds are used to evaluate how accurate the model’s predictions are.</p>\n<p>A simple way to think about it: imagine trying to guess someone’s location on a map. The threshold tells us how close our guess must be to the actual location for it to count as “correct.”</p>\n<p>For example:</p>\n<p>If the threshold is 0.01 meters, the guess must be very precise.</p>\n<p>If the threshold is 1.0 meter, the guess can be more relaxed, and still count as correct.</p>\n<p>A fun analogy: Think of a number guessing game where one person picks a number, and the other tries to guess it. If the threshold is 20, and the correct number is 50, then a guess of 57 or 40 would be accepted — but 71 would be outside the threshold, and count as incorrect.</p>\n<p>Each scene in the dataset is evaluated using six different thresholds. This multi-threshold evaluation is important because:</p>\n<p>It lets us assess performance at various levels of precision.</p>\n<p>Some scenes are easy (e.g. clear, simple environments), while others are much harder (e.g. cluttered, repetitive, or ambiguous scenes).</p>\n<p>By checking the model’s accuracy across these thresholds, we understand how well it performs in both strict and forgiving scenarios.</p>\n<p>For example:</p>\n<p>Lower thresholds (e.g. 0.002, 0.01) require extremely accurate matches.</p>\n<p>Higher thresholds (e.g. 0.5, 1.0, 2.0) allow more error, and test whether the model can still be useful under real-world constraints.</p>\n<p>Ultimately, these thresholds enable a comprehensive and fair evaluation of the model across diverse datasets and difficulty levels. Hope this clears it up!😊\"</p>",
      "rawMarkdown": "@pragyatripathiii23 just answered my question. I'm sharing it here for anyone else who might be wondering the same thing.  So all the credits of the below explanation goes to @pragyatripathiii23  \n\nExplanation: \n\"The threshold CSV file contains error limits (thresholds) defined for each scene in every dataset. These thresholds are used to evaluate how accurate the model’s predictions are.\n\nA simple way to think about it: imagine trying to guess someone’s location on a map. The threshold tells us how close our guess must be to the actual location for it to count as “correct.”\n\nFor example:\n\nIf the threshold is 0.01 meters, the guess must be very precise.\n\nIf the threshold is 1.0 meter, the guess can be more relaxed, and still count as correct.\n\nA fun analogy: Think of a number guessing game where one person picks a number, and the other tries to guess it. If the threshold is 20, and the correct number is 50, then a guess of 57 or 40 would be accepted — but 71 would be outside the threshold, and count as incorrect.\n\nEach scene in the dataset is evaluated using six different thresholds. This multi-threshold evaluation is important because:\n\nIt lets us assess performance at various levels of precision.\n\nSome scenes are easy (e.g. clear, simple environments), while others are much harder (e.g. cluttered, repetitive, or ambiguous scenes).\n\nBy checking the model’s accuracy across these thresholds, we understand how well it performs in both strict and forgiving scenarios.\n\nFor example:\n\nLower thresholds (e.g. 0.002, 0.01) require extremely accurate matches.\n\nHigher thresholds (e.g. 0.5, 1.0, 2.0) allow more error, and test whether the model can still be useful under real-world constraints.\n\nUltimately, these thresholds enable a comprehensive and fair evaluation of the model across diverse datasets and difficulty levels. Hope this clears it up!😊\"",
      "votes": 1
    },
    {
      "id": 3177849,
      "postDate": "2025-04-13T10:52:27.790Z",
      "content": "<p>Hi everyone! I'm new to the Image Matching Challenge 2025 and I'm having trouble understanding how the threshold points work. I'd really appreciate it if someone could explain it to me. </p>\n<p>Hey guys if anyone have any idea, please share that with me. </p>",
      "rawMarkdown": "Hi everyone! I'm new to the Image Matching Challenge 2025 and I'm having trouble understanding how the threshold points work. I'd really appreciate it if someone could explain it to me. \n\n\nHey guys if anyone have any idea, please share that with me. "
    },
    {
      "id": 3178397,
      "postDate": "2025-04-14T05:55:03.250Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3184488,
      "author_name": "Eduard Trulls",
      "author_url": "",
      "post_date": "2025-04-22T06:41:16.840000",
      "content": "<p>We follow the standard procedure of evaluating accuracy at different thresholds. We register the cameras against the ground truth and then compute how many are closer than a threshold, in spatial terms (you can think of this as classifying a camera as \"correct\" or \"incorrect\"). We do this for each camera, and average to compute the accuracy. Then average again over multiple thresholds: this is why the metric is called \"mean average accuracy\". Finally we average over the datasets. The thresholds are not the same for every dataset simply because the scale of the scene can vary a lot: think of a small object (where photos are captured within a few meters, or even less) vs a large square (which can span hundreds of meters).</p>\n<p>It may sound complicated but it's pretty simple. The metric code is public, and you can read it <a href=\"https://www.kaggle.com/datasets/eduardtrulls/imc25-utils\" target=\"_blank\">here</a>. This one is pretty long because it also solves the cluster assignment, so if you primarily want to understand how the thresholds work you can look at the <a href=\"https://www.kaggle.com/code/fabiobellavia/imc2024-3d-metric-evaluation-example\" target=\"_blank\">metric code</a> for last year's competition, which is essentially the same.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3177930,
      "author_name": "Bishwa3901",
      "author_url": "",
      "post_date": "2025-04-13T13:28:13.227000",
      "content": "<p><a href=\"https://www.kaggle.com/pragyatripathiii23\" target=\"_blank\">@pragyatripathiii23</a> just answered my question. I'm sharing it here for anyone else who might be wondering the same thing.  So all the credits of the below explanation goes to <a href=\"https://www.kaggle.com/pragyatripathiii23\" target=\"_blank\">@pragyatripathiii23</a>  </p>\n<p>Explanation: <br>\n\"The threshold CSV file contains error limits (thresholds) defined for each scene in every dataset. These thresholds are used to evaluate how accurate the model’s predictions are.</p>\n<p>A simple way to think about it: imagine trying to guess someone’s location on a map. The threshold tells us how close our guess must be to the actual location for it to count as “correct.”</p>\n<p>For example:</p>\n<p>If the threshold is 0.01 meters, the guess must be very precise.</p>\n<p>If the threshold is 1.0 meter, the guess can be more relaxed, and still count as correct.</p>\n<p>A fun analogy: Think of a number guessing game where one person picks a number, and the other tries to guess it. If the threshold is 20, and the correct number is 50, then a guess of 57 or 40 would be accepted — but 71 would be outside the threshold, and count as incorrect.</p>\n<p>Each scene in the dataset is evaluated using six different thresholds. This multi-threshold evaluation is important because:</p>\n<p>It lets us assess performance at various levels of precision.</p>\n<p>Some scenes are easy (e.g. clear, simple environments), while others are much harder (e.g. cluttered, repetitive, or ambiguous scenes).</p>\n<p>By checking the model’s accuracy across these thresholds, we understand how well it performs in both strict and forgiving scenarios.</p>\n<p>For example:</p>\n<p>Lower thresholds (e.g. 0.002, 0.01) require extremely accurate matches.</p>\n<p>Higher thresholds (e.g. 0.5, 1.0, 2.0) allow more error, and test whether the model can still be useful under real-world constraints.</p>\n<p>Ultimately, these thresholds enable a comprehensive and fair evaluation of the model across diverse datasets and difficulty levels. Hope this clears it up!😊\"</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3177849,
      "author_name": "Bishwa3901",
      "author_url": "",
      "post_date": "2025-04-13T10:52:27.790000",
      "content": "<p>Hi everyone! I'm new to the Image Matching Challenge 2025 and I'm having trouble understanding how the threshold points work. I'd really appreciate it if someone could explain it to me. </p>\n<p>Hey guys if anyone have any idea, please share that with me. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3178397,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-14T05:55:03.250000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3177658": "Hi everyone! I'm new to the Image Matching Challenge 2025 and I'm having trouble understanding how the threshold points work. I'd really appreciate it if someone could explain it to me.",
    "3184488": "We follow the standard procedure of evaluating accuracy at different thresholds. We register the cameras against the ground truth and then compute how many are closer than a threshold, in spatial terms (you can think of this as classifying a camera as \"correct\" or \"incorrect\"). We do this for each camera, and average to compute the accuracy. Then average again over multiple thresholds: this is why the metric is called \"mean average accuracy\". Finally we average over the datasets. The thresholds are not the same for every dataset simply because the scale of the scene can vary a lot: think of a small object (where photos are captured within a few meters, or even less) vs a large square (which can span hundreds of meters).\n\nIt may sound complicated but it's pretty simple. The metric code is public, and you can read it [here](https://www.kaggle.com/datasets/eduardtrulls/imc25-utils). This one is pretty long because it also solves the cluster assignment, so if you primarily want to understand how the thresholds work you can look at the [metric code](https://www.kaggle.com/code/fabiobellavia/imc2024-3d-metric-evaluation-example) for last year's competition, which is essentially the same.",
    "3177930": "@pragyatripathiii23 just answered my question. I'm sharing it here for anyone else who might be wondering the same thing.  So all the credits of the below explanation goes to @pragyatripathiii23  \n\nExplanation: \n\"The threshold CSV file contains error limits (thresholds) defined for each scene in every dataset. These thresholds are used to evaluate how accurate the model’s predictions are.\n\nA simple way to think about it: imagine trying to guess someone’s location on a map. The threshold tells us how close our guess must be to the actual location for it to count as “correct.”\n\nFor example:\n\nIf the threshold is 0.01 meters, the guess must be very precise.\n\nIf the threshold is 1.0 meter, the guess can be more relaxed, and still count as correct.\n\nA fun analogy: Think of a number guessing game where one person picks a number, and the other tries to guess it. If the threshold is 20, and the correct number is 50, then a guess of 57 or 40 would be accepted — but 71 would be outside the threshold, and count as incorrect.\n\nEach scene in the dataset is evaluated using six different thresholds. This multi-threshold evaluation is important because:\n\nIt lets us assess performance at various levels of precision.\n\nSome scenes are easy (e.g. clear, simple environments), while others are much harder (e.g. cluttered, repetitive, or ambiguous scenes).\n\nBy checking the model’s accuracy across these thresholds, we understand how well it performs in both strict and forgiving scenarios.\n\nFor example:\n\nLower thresholds (e.g. 0.002, 0.01) require extremely accurate matches.\n\nHigher thresholds (e.g. 0.5, 1.0, 2.0) allow more error, and test whether the model can still be useful under real-world constraints.\n\nUltimately, these thresholds enable a comprehensive and fair evaluation of the model across diverse datasets and difficulty levels. Hope this clears it up!😊\"",
    "3177849": "Hi everyone! I'm new to the Image Matching Challenge 2025 and I'm having trouble understanding how the threshold points work. I'd really appreciate it if someone could explain it to me. \n\n\nHey guys if anyone have any idea, please share that with me. ",
    "3178397": ""
  }
}