{
  "id": 403314,
  "title": "Evaluation: Averaging datasets",
  "url": "/competitions/image-matching-challenge-2023/discussion/403314",
  "author_name": "",
  "post_date": "2023-04-22T11:43:35.629269300Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My question is mostly for competition hosts, but not limited to them and open for a discussion.</p>\n<p>Current metric averages scores over scenes first, then averages it over datasets. I wonder why is it so important to average among datasets as a final step? Isn't it enough to average only among scenes? When I look at data that we have (labeled train) I notice that some datasets are much harder to reconstruct because of poor scenes (simple objects) and lack of images (like 15 images) (Example haiper-bike or haiper-chairs). The score for that kind of scenes is small. If some dataset contains such kind of scenes it will negatively affect the score when averaging. The issue with averaging over datasets is also that datasets with different number of scenes has same weight.</p>",
  "messages": [
    {
      "id": "2230469",
      "postDate": "04/22/2023 11:43:35",
      "content": "<p>My question is mostly for competition hosts, but not limited to them and open for a discussion.</p>\n<p>Current metric averages scores over scenes first, then averages it over datasets. I wonder why is it so important to average among datasets as a final step? Isn't it enough to average only among scenes? When I look at data that we have (labeled train) I notice that some datasets are much harder to reconstruct because of poor scenes (simple objects) and lack of images (like 15 images) (Example haiper-bike or haiper-chairs). The score for that kind of scenes is small. If some dataset contains such kind of scenes it will negatively affect the score when averaging. The issue with averaging over datasets is also that datasets with different number of scenes has same weight.</p>",
      "rawMarkdown": "My question is mostly for competition hosts, but not limited to them and open for a discussion.\n\nCurrent metric averages scores over scenes first, then averages it over datasets. I wonder why is it so important to average among datasets as a final step? Isn't it enough to average only among scenes? When I look at data that we have (labeled train) I notice that some datasets are much harder to reconstruct because of poor scenes (simple objects) and lack of images (like 15 images) (Example haiper-bike or haiper-chairs). The score for that kind of scenes is small. If some dataset contains such kind of scenes it will negatively affect the score when averaging. The issue with averaging over datasets is also that datasets with different number of scenes has same weight.",
      "votes": null
    },
    {
      "id": "2230492",
      "postDate": "04/22/2023 12:18:33",
      "content": "<p>It's not a bug, it's a feature :)</p>",
      "rawMarkdown": "It's not a bug, it's a feature :)",
      "votes": null
    },
    {
      "id": "2230501",
      "postDate": "04/22/2023 12:35:52",
      "content": "<blockquote>\n  <p>The score for that kind of scenes is small.<br>\n  Yes, that's the idea :)</p>\n</blockquote>",
      "rawMarkdown": ">The score for that kind of scenes is small.\nYes, that's the idea :)",
      "votes": null
    },
    {
      "id": "2230510",
      "postDate": "04/22/2023 12:44:49",
      "content": "<p>Got it! Thanks for the fast reply)</p>",
      "rawMarkdown": "Got it! Thanks for the fast reply)",
      "votes": null
    },
    {
      "id": "2230511",
      "postDate": "04/22/2023 12:45:04",
      "content": "<p>Thank you for the fast answer!</p>",
      "rawMarkdown": "Thank you for the fast answer!",
      "votes": null
    },
    {
      "id": "2251666",
      "postDate": "05/09/2023 14:51:59",
      "content": "<p>I guess the intention was weighting each dataset equally like you said. It seemed little bit weird to me as well.</p>",
      "rawMarkdown": "I guess the intention was weighting each dataset equally like you said. It seemed little bit weird to me as well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2230492,
      "author_name": "fabiobellavia",
      "author_url": "",
      "post_date": "04/22/2023 12:18:33",
      "content": "<p>It's not a bug, it's a feature :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2230511,
          "author_name": "vostankovich",
          "author_url": "",
          "post_date": "04/22/2023 12:45:04",
          "content": "<p>Thank you for the fast answer!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2230501,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "04/22/2023 12:35:52",
      "content": "<blockquote>\n  <p>The score for that kind of scenes is small.<br>\n  Yes, that's the idea :)</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 2230510,
          "author_name": "vostankovich",
          "author_url": "",
          "post_date": "04/22/2023 12:44:49",
          "content": "<p>Got it! Thanks for the fast reply)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2251666,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "05/09/2023 14:51:59",
      "content": "<p>I guess the intention was weighting each dataset equally like you said. It seemed little bit weird to me as well.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2230469": "My question is mostly for competition hosts, but not limited to them and open for a discussion.\n\nCurrent metric averages scores over scenes first, then averages it over datasets. I wonder why is it so important to average among datasets as a final step? Isn't it enough to average only among scenes? When I look at data that we have (labeled train) I notice that some datasets are much harder to reconstruct because of poor scenes (simple objects) and lack of images (like 15 images) (Example haiper-bike or haiper-chairs). The score for that kind of scenes is small. If some dataset contains such kind of scenes it will negatively affect the score when averaging. The issue with averaging over datasets is also that datasets with different number of scenes has same weight.",
    "2230492": "It's not a bug, it's a feature :)",
    "2230501": ">The score for that kind of scenes is small.\nYes, that's the idea :)",
    "2230510": "Got it! Thanks for the fast reply)",
    "2230511": "Thank you for the fast answer!",
    "2251666": "I guess the intention was weighting each dataset equally like you said. It seemed little bit weird to me as well."
  },
  "source": "meta"
}