{
  "id": 225316,
  "title": "Benchmarks and Evaluation",
  "url": "/competitions/iwildcam2021-fgvc8/discussion/225316",
  "author_name": "",
  "post_date": "2021-03-11T17:14:41.649539600Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>One of the most challenging decisions when building this competition was the choice of error metric. We needed a metric that would (1) penalize species categorization mistakes, (2) penalize counting mistakes appropriately, so predicting 4 when there were 5 of a species is penalized less than predicting 1, and (3) penalize false positives in empty sequences. The only metric in the set of metrics supported by Kaggle that would capture all of these mistakes was Mean Columnwise Root Mean Squared Error (MCRMSE, you can see the definition in the Overview-&gt;Evaluation tab). However, this metric was not really designed for our situation, where any given sequence contains only a few species out of the set, and many sequences are empty. This means that the normalization of the metric (the “mean”s you see in the metric name, over both the number of sequences and the number of classes) causes us to divide our error by two very large numbers. For example, look at the “all zeros” benchmark on the leaderboard! By guessing there are no animals of any species in any of the sequences, MCRMSE is ~0.038! </p>\n<p>So to sum up: our metric captures all the types of mistakes we care about, but the error values look quite low because of the nature of the data. To map the error to something more ecologically relevant and easier to interpret, you can un-normalize the leaderboard error by scaling with a factor of number of classes * sqrt(number of test sequences). After that scaling step, you’re left with something that will be a reasonable estimate of the total error in your species counts, the Summed Columnwise Root Summed Squared Error (SCRSSE, also defined in Overview-&gt;Evaluation).</p>\n<p>As our second naive benchmark (besides just all zeros 😛) we took our <a href=\"https://www.kaggle.com/c/iwildcam-2020-fgvc7/discussion/136755\" target=\"_blank\">benchmark class predictions from last year's iWildCam</a> and the MegaDetector detections, used majority vote over the images in a sequence that did not have detected boxes to determine the species (note this only allows for one species per sequence, and there are sequences with multiple species in the dataset!), and used the maximum number of boxes across any image in the sequence as our guess for the total count across the sequence. This performs even worse than the all-zero benchmark, due to a combination of missed classifications and over- and under-predictions.</p>\n<p>Let me know if you have any questions! </p>",
  "messages": [
    {
      "id": "1234904",
      "postDate": "03/11/2021 17:14:41",
      "content": "<p>One of the most challenging decisions when building this competition was the choice of error metric. We needed a metric that would (1) penalize species categorization mistakes, (2) penalize counting mistakes appropriately, so predicting 4 when there were 5 of a species is penalized less than predicting 1, and (3) penalize false positives in empty sequences. The only metric in the set of metrics supported by Kaggle that would capture all of these mistakes was Mean Columnwise Root Mean Squared Error (MCRMSE, you can see the definition in the Overview-&gt;Evaluation tab). However, this metric was not really designed for our situation, where any given sequence contains only a few species out of the set, and many sequences are empty. This means that the normalization of the metric (the “mean”s you see in the metric name, over both the number of sequences and the number of classes) causes us to divide our error by two very large numbers. For example, look at the “all zeros” benchmark on the leaderboard! By guessing there are no animals of any species in any of the sequences, MCRMSE is ~0.038! </p>\n<p>So to sum up: our metric captures all the types of mistakes we care about, but the error values look quite low because of the nature of the data. To map the error to something more ecologically relevant and easier to interpret, you can un-normalize the leaderboard error by scaling with a factor of number of classes * sqrt(number of test sequences). After that scaling step, you’re left with something that will be a reasonable estimate of the total error in your species counts, the Summed Columnwise Root Summed Squared Error (SCRSSE, also defined in Overview-&gt;Evaluation).</p>\n<p>As our second naive benchmark (besides just all zeros 😛) we took our <a href=\"https://www.kaggle.com/c/iwildcam-2020-fgvc7/discussion/136755\" target=\"_blank\">benchmark class predictions from last year's iWildCam</a> and the MegaDetector detections, used majority vote over the images in a sequence that did not have detected boxes to determine the species (note this only allows for one species per sequence, and there are sequences with multiple species in the dataset!), and used the maximum number of boxes across any image in the sequence as our guess for the total count across the sequence. This performs even worse than the all-zero benchmark, due to a combination of missed classifications and over- and under-predictions.</p>\n<p>Let me know if you have any questions! </p>",
      "rawMarkdown": "One of the most challenging decisions when building this competition was the choice of error metric. We needed a metric that would (1) penalize species categorization mistakes, (2) penalize counting mistakes appropriately, so predicting 4 when there were 5 of a species is penalized less than predicting 1, and (3) penalize false positives in empty sequences. The only metric in the set of metrics supported by Kaggle that would capture all of these mistakes was Mean Columnwise Root Mean Squared Error (MCRMSE, you can see the definition in the Overview->Evaluation tab). However, this metric was not really designed for our situation, where any given sequence contains only a few species out of the set, and many sequences are empty. This means that the normalization of the metric (the “mean”s you see in the metric name, over both the number of sequences and the number of classes) causes us to divide our error by two very large numbers. For example, look at the “all zeros” benchmark on the leaderboard! By guessing there are no animals of any species in any of the sequences, MCRMSE is ~0.038! \n\nSo to sum up: our metric captures all the types of mistakes we care about, but the error values look quite low because of the nature of the data. To map the error to something more ecologically relevant and easier to interpret, you can un-normalize the leaderboard error by scaling with a factor of number of classes * sqrt(number of test sequences). After that scaling step, you’re left with something that will be a reasonable estimate of the total error in your species counts, the Summed Columnwise Root Summed Squared Error (SCRSSE, also defined in Overview->Evaluation).\n\nAs our second naive benchmark (besides just all zeros 😛) we took our [benchmark class predictions from last year's iWildCam](https://www.kaggle.com/c/iwildcam-2020-fgvc7/discussion/136755) and the MegaDetector detections, used majority vote over the images in a sequence that did not have detected boxes to determine the species (note this only allows for one species per sequence, and there are sequences with multiple species in the dataset!), and used the maximum number of boxes across any image in the sequence as our guess for the total count across the sequence. This performs even worse than the all-zero benchmark, due to a combination of missed classifications and over- and under-predictions.\n\nLet me know if you have any questions!",
      "votes": null
    },
    {
      "id": "1258853",
      "postDate": "04/01/2021 00:21:07",
      "content": "<p>I've just added an additional benchmark, for this one I used the same heuristics as the above (MegaDetector detections, used majority vote over the images in a sequence that did not have detected boxes to determine the species (note this only allows for one species per sequence, and there are sequences with multiple species in the dataset!), and used the maximum number of boxes across any image in the sequence as our guess for the total count across the sequence), but this time instead of species predictions from our vanilla benchmark from the 2020 competition, I used the species predictions from the winning 2020 submission :)</p>",
      "rawMarkdown": "I've just added an additional benchmark, for this one I used the same heuristics as the above (MegaDetector detections, used majority vote over the images in a sequence that did not have detected boxes to determine the species (note this only allows for one species per sequence, and there are sequences with multiple species in the dataset!), and used the maximum number of boxes across any image in the sequence as our guess for the total count across the sequence), but this time instead of species predictions from our vanilla benchmark from the 2020 competition, I used the species predictions from the winning 2020 submission :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1258853,
      "author_name": "sbeery",
      "author_url": "",
      "post_date": "04/01/2021 00:21:07",
      "content": "<p>I've just added an additional benchmark, for this one I used the same heuristics as the above (MegaDetector detections, used majority vote over the images in a sequence that did not have detected boxes to determine the species (note this only allows for one species per sequence, and there are sequences with multiple species in the dataset!), and used the maximum number of boxes across any image in the sequence as our guess for the total count across the sequence), but this time instead of species predictions from our vanilla benchmark from the 2020 competition, I used the species predictions from the winning 2020 submission :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1234904": "One of the most challenging decisions when building this competition was the choice of error metric. We needed a metric that would (1) penalize species categorization mistakes, (2) penalize counting mistakes appropriately, so predicting 4 when there were 5 of a species is penalized less than predicting 1, and (3) penalize false positives in empty sequences. The only metric in the set of metrics supported by Kaggle that would capture all of these mistakes was Mean Columnwise Root Mean Squared Error (MCRMSE, you can see the definition in the Overview->Evaluation tab). However, this metric was not really designed for our situation, where any given sequence contains only a few species out of the set, and many sequences are empty. This means that the normalization of the metric (the “mean”s you see in the metric name, over both the number of sequences and the number of classes) causes us to divide our error by two very large numbers. For example, look at the “all zeros” benchmark on the leaderboard! By guessing there are no animals of any species in any of the sequences, MCRMSE is ~0.038! \n\nSo to sum up: our metric captures all the types of mistakes we care about, but the error values look quite low because of the nature of the data. To map the error to something more ecologically relevant and easier to interpret, you can un-normalize the leaderboard error by scaling with a factor of number of classes * sqrt(number of test sequences). After that scaling step, you’re left with something that will be a reasonable estimate of the total error in your species counts, the Summed Columnwise Root Summed Squared Error (SCRSSE, also defined in Overview->Evaluation).\n\nAs our second naive benchmark (besides just all zeros 😛) we took our [benchmark class predictions from last year's iWildCam](https://www.kaggle.com/c/iwildcam-2020-fgvc7/discussion/136755) and the MegaDetector detections, used majority vote over the images in a sequence that did not have detected boxes to determine the species (note this only allows for one species per sequence, and there are sequences with multiple species in the dataset!), and used the maximum number of boxes across any image in the sequence as our guess for the total count across the sequence. This performs even worse than the all-zero benchmark, due to a combination of missed classifications and over- and under-predictions.\n\nLet me know if you have any questions!",
    "1258853": "I've just added an additional benchmark, for this one I used the same heuristics as the above (MegaDetector detections, used majority vote over the images in a sequence that did not have detected boxes to determine the species (note this only allows for one species per sequence, and there are sequences with multiple species in the dataset!), and used the maximum number of boxes across any image in the sequence as our guess for the total count across the sequence), but this time instead of species predictions from our vanilla benchmark from the 2020 competition, I used the species predictions from the winning 2020 submission :)"
  },
  "source": "meta"
}