{
  "id": 171337,
  "title": "[Offtopic] On Comps Evaluation",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171337",
  "author_name": "",
  "post_date": "2020-07-31T12:09:55.692082600Z",
  "votes": -3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey guys! \nA little off topic about competition evaluation. Taking into consideration recently ended competition (Panda) with some shakeup and some people went up with a weak LB solution occasionally overfittng private. Why do you think Kaggle and organizers don't consider to evaluate results on the best paired solution (public and private)? Do you think this could be reasonable or not? Why?</p>\n\n<p>PS. Really sorry for off topic!</p>",
  "messages": [
    {
      "id": "952994",
      "postDate": "07/31/2020 12:09:55",
      "content": "<p>Hey guys! \nA little off topic about competition evaluation. Taking into consideration recently ended competition (Panda) with some shakeup and some people went up with a weak LB solution occasionally overfittng private. Why do you think Kaggle and organizers don't consider to evaluate results on the best paired solution (public and private)? Do you think this could be reasonable or not? Why?</p>\n\n<p>PS. Really sorry for off topic!</p>",
      "rawMarkdown": "Hey guys! \nA little off topic about competition evaluation. Taking into consideration recently ended competition (Panda) with some shakeup and some people went up with a weak LB solution occasionally overfittng private. Why do you think Kaggle and organizers don't consider to evaluate results on the best paired solution (public and private)? Do you think this could be reasonable or not? Why?\n\nPS. Really sorry for off topic!",
      "votes": null
    },
    {
      "id": "953043",
      "postDate": "07/31/2020 13:23:51",
      "content": "<p>I can see no reason to do this, it would bypass the whole reason of having a hold-out test set and encourage work which follows bad ML practice.</p>",
      "rawMarkdown": "I can see no reason to do this, it would bypass the whole reason of having a hold-out test set and encourage work which follows bad ML practice.",
      "votes": null
    },
    {
      "id": "953211",
      "postDate": "07/31/2020 15:50:02",
      "content": "<p>It is too easy to purposefully overfit the public LB</p>",
      "rawMarkdown": "It is too easy to purposefully overfit the public LB",
      "votes": null
    },
    {
      "id": "953244",
      "postDate": "07/31/2020 16:12:08",
      "content": "<p>How do you overfit private? You can overfit public, but private, how??</p>",
      "rawMarkdown": "How do you overfit private? You can overfit public, but private, how??",
      "votes": null
    },
    {
      "id": "953253",
      "postDate": "07/31/2020 16:16:58",
      "content": "<p>I meant, still has public and private. Just after competition ends, take into consideration results of both sets. The higher results on both and the lower the difference between them (say lower generalization error) the better.</p>",
      "rawMarkdown": "I meant, still has public and private. Just after competition ends, take into consideration results of both sets. The higher results on both and the lower the difference between them (say lower generalization error) the better.",
      "votes": null
    },
    {
      "id": "953255",
      "postDate": "07/31/2020 16:20:46",
      "content": "<p>What I meant is more or less by chance. For example, consider the case. Person A made a 2 submissions at the beginning of the comp and decided to stop. His LB = 0.87 and private = 0.93. \nPerson B work through the comp and get LB = 0.91 and private = 0.92.\nWhat result would be considered better?</p>",
      "rawMarkdown": "What I meant is more or less by chance. For example, consider the case. Person A made a 2 submissions at the beginning of the comp and decided to stop. His LB = 0.87 and private = 0.93. \nPerson B work through the comp and get LB = 0.91 and private = 0.92.\nWhat result would be considered better?",
      "votes": null
    },
    {
      "id": "953352",
      "postDate": "07/31/2020 17:59:34",
      "content": "<p>A more representative example would be:</p>\n\n<p>Person C:\nLB = 0.93, Private = 0.91</p>\n\n<p>and Person D:\nLB = 1.00, Private = 0.90</p>\n\n<p>Clearly person C has the highest performing and 'better' model but how would you take into account Person D in your scheme when they clearly overfit the Public leader board?</p>\n\n<p>What you suggest makes no sense in the competition, by the very fact that the public LB is public and can be trivially overfit.</p>",
      "rawMarkdown": "A more representative example would be:\n\nPerson C:\nLB = 0.93, Private = 0.91\n\nand Person D:\nLB = 1.00, Private = 0.90\n\nClearly person C has the highest performing and 'better' model but how would you take into account Person D in your scheme when they clearly overfit the Public leader board?\n\nWhat you suggest makes no sense in the competition, by the very fact that the public LB is public and can be trivially overfit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 953043,
      "author_name": "fchmiel",
      "author_url": "",
      "post_date": "07/31/2020 13:23:51",
      "content": "<p>I can see no reason to do this, it would bypass the whole reason of having a hold-out test set and encourage work which follows bad ML practice.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953211,
      "author_name": "abiolatti",
      "author_url": "",
      "post_date": "07/31/2020 15:50:02",
      "content": "<p>It is too easy to purposefully overfit the public LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953244,
      "author_name": "ks2019",
      "author_url": "",
      "post_date": "07/31/2020 16:12:08",
      "content": "<p>How do you overfit private? You can overfit public, but private, how??</p>",
      "votes": null,
      "replies": [
        {
          "id": 953255,
          "author_name": "ademyanchuk",
          "author_url": "",
          "post_date": "07/31/2020 16:20:46",
          "content": "<p>What I meant is more or less by chance. For example, consider the case. Person A made a 2 submissions at the beginning of the comp and decided to stop. His LB = 0.87 and private = 0.93. \nPerson B work through the comp and get LB = 0.91 and private = 0.92.\nWhat result would be considered better?</p>",
          "votes": null,
          "replies": [
            {
              "id": 953352,
              "author_name": "fchmiel",
              "author_url": "",
              "post_date": "07/31/2020 17:59:34",
              "content": "<p>A more representative example would be:</p>\n\n<p>Person C:\nLB = 0.93, Private = 0.91</p>\n\n<p>and Person D:\nLB = 1.00, Private = 0.90</p>\n\n<p>Clearly person C has the highest performing and 'better' model but how would you take into account Person D in your scheme when they clearly overfit the Public leader board?</p>\n\n<p>What you suggest makes no sense in the competition, by the very fact that the public LB is public and can be trivially overfit.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 953253,
      "author_name": "ademyanchuk",
      "author_url": "",
      "post_date": "07/31/2020 16:16:58",
      "content": "<p>I meant, still has public and private. Just after competition ends, take into consideration results of both sets. The higher results on both and the lower the difference between them (say lower generalization error) the better.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "952994": "Hey guys! \nA little off topic about competition evaluation. Taking into consideration recently ended competition (Panda) with some shakeup and some people went up with a weak LB solution occasionally overfittng private. Why do you think Kaggle and organizers don't consider to evaluate results on the best paired solution (public and private)? Do you think this could be reasonable or not? Why?\n\nPS. Really sorry for off topic!",
    "953043": "I can see no reason to do this, it would bypass the whole reason of having a hold-out test set and encourage work which follows bad ML practice.",
    "953211": "It is too easy to purposefully overfit the public LB",
    "953244": "How do you overfit private? You can overfit public, but private, how??",
    "953253": "I meant, still has public and private. Just after competition ends, take into consideration results of both sets. The higher results on both and the lower the difference between them (say lower generalization error) the better.",
    "953255": "What I meant is more or less by chance. For example, consider the case. Person A made a 2 submissions at the beginning of the comp and decided to stop. His LB = 0.87 and private = 0.93. \nPerson B work through the comp and get LB = 0.91 and private = 0.92.\nWhat result would be considered better?",
    "953352": "A more representative example would be:\n\nPerson C:\nLB = 0.93, Private = 0.91\n\nand Person D:\nLB = 1.00, Private = 0.90\n\nClearly person C has the highest performing and 'better' model but how would you take into account Person D in your scheme when they clearly overfit the Public leader board?\n\nWhat you suggest makes no sense in the competition, by the very fact that the public LB is public and can be trivially overfit."
  },
  "source": "meta"
}