{
  "id": 294869,
  "title": "Better model but lower score",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/294869",
  "author_name": "",
  "post_date": "2021-12-13T09:44:44.992559700Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nI have an issue that I didn't understand. I have a model that has 0.519 score on the public leaderboard and I have another model that has 0.337 score. When I tested and compared them with different images on the internet the model which has lower score, seems better than the other. Am I doing something wrong or is there anyone who is run into same problem? Can anyone explain this?(I don't have much experience on object detection competitions.)</p>",
  "messages": [
    {
      "id": "1616269",
      "postDate": "12/13/2021 09:44:44",
      "content": "<p>Hello everyone,<br>\nI have an issue that I didn't understand. I have a model that has 0.519 score on the public leaderboard and I have another model that has 0.337 score. When I tested and compared them with different images on the internet the model which has lower score, seems better than the other. Am I doing something wrong or is there anyone who is run into same problem? Can anyone explain this?(I don't have much experience on object detection competitions.)</p>",
      "rawMarkdown": "Hello everyone,\nI have an issue that I didn't understand. I have a model that has 0.519 score on the public leaderboard and I have another model that has 0.337 score. When I tested and compared them with different images on the internet the model which has lower score, seems better than the other. Am I doing something wrong or is there anyone who is run into same problem? Can anyone explain this?(I don't have much experience on object detection competitions.)",
      "votes": null
    },
    {
      "id": "1616291",
      "postDate": "12/13/2021 10:07:28",
      "content": "<p>Imo probably your model is more accurate (the 0.337) but competition metric is F2 Score which prefers Recall over precission. Just find them as many as possible because its plague  :)</p>",
      "rawMarkdown": "Imo probably your model is more accurate (the 0.337) but competition metric is F2 Score which prefers Recall over precission. Just find them as many as possible because its plague  :)",
      "votes": null
    },
    {
      "id": "1620593",
      "postDate": "12/16/2021 23:26:27",
      "content": "<p>Don't forget that the public leaderboard is only measuring your score on the public test set (which represents 25% of the whole test set).<br>\nWhen the competition comes to an end, there will be a new Leaderboard (Private LB) and it will show your score on the rest of the test set (75%). That's why many Kagglers don't trust the public LB; because sometimes, the public test set doesn't represent the test distribution.</p>",
      "rawMarkdown": "Don't forget that the public leaderboard is only measuring your score on the public test set (which represents 25% of the whole test set).\nWhen the competition comes to an end, there will be a new Leaderboard (Private LB) and it will show your score on the rest of the test set (75%). That's why many Kagglers don't trust the public LB; because sometimes, the public test set doesn't represent the test distribution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1616291,
      "author_name": "lukaszborecki",
      "author_url": "",
      "post_date": "12/13/2021 10:07:28",
      "content": "<p>Imo probably your model is more accurate (the 0.337) but competition metric is F2 Score which prefers Recall over precission. Just find them as many as possible because its plague  :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1620593,
      "author_name": "rinnqd",
      "author_url": "",
      "post_date": "12/16/2021 23:26:27",
      "content": "<p>Don't forget that the public leaderboard is only measuring your score on the public test set (which represents 25% of the whole test set).<br>\nWhen the competition comes to an end, there will be a new Leaderboard (Private LB) and it will show your score on the rest of the test set (75%). That's why many Kagglers don't trust the public LB; because sometimes, the public test set doesn't represent the test distribution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1616269": "Hello everyone,\nI have an issue that I didn't understand. I have a model that has 0.519 score on the public leaderboard and I have another model that has 0.337 score. When I tested and compared them with different images on the internet the model which has lower score, seems better than the other. Am I doing something wrong or is there anyone who is run into same problem? Can anyone explain this?(I don't have much experience on object detection competitions.)",
    "1616291": "Imo probably your model is more accurate (the 0.337) but competition metric is F2 Score which prefers Recall over precission. Just find them as many as possible because its plague  :)",
    "1620593": "Don't forget that the public leaderboard is only measuring your score on the public test set (which represents 25% of the whole test set).\nWhen the competition comes to an end, there will be a new Leaderboard (Private LB) and it will show your score on the rest of the test set (75%). That's why many Kagglers don't trust the public LB; because sometimes, the public test set doesn't represent the test distribution."
  },
  "source": "meta"
}