{
  "id": 127326,
  "title": "About private stage",
  "url": "/competitions/bengaliai-cv19/discussion/127326",
  "author_name": "",
  "post_date": "2020-01-23T10:12:22.175658100Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Good day,</p>\n\n<p>If my kernel succeeded during public stage does it mean it will succeed on private? Thanks</p>",
  "messages": [
    {
      "id": "726911",
      "postDate": "01/23/2020 10:12:22",
      "content": "<p>Good day,</p>\n\n<p>If my kernel succeeded during public stage does it mean it will succeed on private? Thanks</p>",
      "rawMarkdown": "Good day,\n\nIf my kernel succeeded during public stage does it mean it will succeed on private? Thanks",
      "votes": null
    },
    {
      "id": "727094",
      "postDate": "01/23/2020 13:04:13",
      "content": "<p>Dear John,</p>\n\n<p>Your score has already been calculated for the private dataset, it is just not shown to you yet ! What you see as being your public score represents subset of the whole test dataset. Hence, it is not representative of your overall performance on the entire dataset. </p>\n\n<p>So in simple words, if your kernel has succeeded on the public submission, it has already succeeded on the private submission.</p>",
      "rawMarkdown": "Dear John,\n\nYour score has already been calculated for the private dataset, it is just not shown to you yet ! What you see as being your public score represents subset of the whole test dataset. Hence, it is not representative of your overall performance on the entire dataset. \n\nSo in simple words, if your kernel has succeeded on the public submission, it has already succeeded on the private submission.",
      "votes": null
    },
    {
      "id": "727200",
      "postDate": "01/23/2020 14:27:39",
      "content": "<p>I see, thanks, Thomas! </p>",
      "rawMarkdown": "I see, thanks, Thomas!",
      "votes": null
    },
    {
      "id": "727379",
      "postDate": "01/23/2020 17:13:02",
      "content": "<p>To add to what <a href=\"/dimartinot\">@dimartinot</a> said, I think it’s also interesting to know why there’s a public &amp; private leader board.</p>\n\n<p>One of them is to simulate a «&nbsp;deployment&nbsp;» of the model in production: if we deploy your model toady, and feed it new data you’ve never been able to test again, how would it perform?</p>\n\n<p>Then, these are competitions, done by smart people. People are going to try everything they can to try to gain every little advantage they can (and that’s great, that pushes the field ahead!) but not when it comes to «&nbsp;grey-area&nbsp;» practices. One such practice is leader-board dataset probing: Let’s say we’re trying to make a binary classifier, only 2 classes. But you don’t know the distribution of both classes, it could be 50/50, but maybe 20/80 or even 99/1. Knowing this would help you greatly! So if you could know your score on the leaderboard, even without seeing the data you could get some info. For example, make a submission with everything class 1, and then another one with every class 2. if you get an accuracy of 60% the 1st time and then 40%, all of a sudden you know the distribution without having ever seen the data! Hence the probing.</p>\n\n<p>Overall, it’s to prevent from fitting a model specifically on this dataset, but pushing people to make a model that generalises well.</p>\n\n<p>In this competition for example, there are going to be unseen characters in the private LB. So combinations of vowel, consonants and roots that we haven’t seen in the training data (but with the same building blocks, no new vowels, consonants or roots). This will probably lead to some shake-up: when there’s a difference between public leaderboard, and private. This is pretty common in competitions, and can sometimes be drops (or rises) of &gt; 1,000 places!</p>\n\n<p>Hope that helps!</p>",
      "rawMarkdown": "To add to what @dimartinot said, I think it’s also interesting to know why there’s a public &amp; private leader board.\n\nOne of them is to simulate a «&nbsp;deployment&nbsp;» of the model in production: if we deploy your model toady, and feed it new data you’ve never been able to test again, how would it perform?\n\nThen, these are competitions, done by smart people. People are going to try everything they can to try to gain every little advantage they can (and that’s great, that pushes the field ahead!) but not when it comes to «&nbsp;grey-area&nbsp;» practices. One such practice is leader-board dataset probing: Let’s say we’re trying to make a binary classifier, only 2 classes. But you don’t know the distribution of both classes, it could be 50/50, but maybe 20/80 or even 99/1. Knowing this would help you greatly! So if you could know your score on the leaderboard, even without seeing the data you could get some info. For example, make a submission with everything class 1, and then another one with every class 2. if you get an accuracy of 60% the 1st time and then 40%, all of a sudden you know the distribution without having ever seen the data! Hence the probing.\n\nOverall, it’s to prevent from fitting a model specifically on this dataset, but pushing people to make a model that generalises well.\n\nIn this competition for example, there are going to be unseen characters in the private LB. So combinations of vowel, consonants and roots that we haven’t seen in the training data (but with the same building blocks, no new vowels, consonants or roots). This will probably lead to some shake-up: when there’s a difference between public leaderboard, and private. This is pretty common in competitions, and can sometimes be drops (or rises) of &gt; 1,000 places!\n\nHope that helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 727094,
      "author_name": "dimartinot",
      "author_url": "",
      "post_date": "01/23/2020 13:04:13",
      "content": "<p>Dear John,</p>\n\n<p>Your score has already been calculated for the private dataset, it is just not shown to you yet ! What you see as being your public score represents subset of the whole test dataset. Hence, it is not representative of your overall performance on the entire dataset. </p>\n\n<p>So in simple words, if your kernel has succeeded on the public submission, it has already succeeded on the private submission.</p>",
      "votes": null,
      "replies": [
        {
          "id": 727200,
          "author_name": "johnndoea",
          "author_url": "",
          "post_date": "01/23/2020 14:27:39",
          "content": "<p>I see, thanks, Thomas! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 727379,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "01/23/2020 17:13:02",
          "content": "<p>To add to what <a href=\"/dimartinot\">@dimartinot</a> said, I think it’s also interesting to know why there’s a public &amp; private leader board.</p>\n\n<p>One of them is to simulate a «&nbsp;deployment&nbsp;» of the model in production: if we deploy your model toady, and feed it new data you’ve never been able to test again, how would it perform?</p>\n\n<p>Then, these are competitions, done by smart people. People are going to try everything they can to try to gain every little advantage they can (and that’s great, that pushes the field ahead!) but not when it comes to «&nbsp;grey-area&nbsp;» practices. One such practice is leader-board dataset probing: Let’s say we’re trying to make a binary classifier, only 2 classes. But you don’t know the distribution of both classes, it could be 50/50, but maybe 20/80 or even 99/1. Knowing this would help you greatly! So if you could know your score on the leaderboard, even without seeing the data you could get some info. For example, make a submission with everything class 1, and then another one with every class 2. if you get an accuracy of 60% the 1st time and then 40%, all of a sudden you know the distribution without having ever seen the data! Hence the probing.</p>\n\n<p>Overall, it’s to prevent from fitting a model specifically on this dataset, but pushing people to make a model that generalises well.</p>\n\n<p>In this competition for example, there are going to be unseen characters in the private LB. So combinations of vowel, consonants and roots that we haven’t seen in the training data (but with the same building blocks, no new vowels, consonants or roots). This will probably lead to some shake-up: when there’s a difference between public leaderboard, and private. This is pretty common in competitions, and can sometimes be drops (or rises) of &gt; 1,000 places!</p>\n\n<p>Hope that helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "726911": "Good day,\n\nIf my kernel succeeded during public stage does it mean it will succeed on private? Thanks",
    "727094": "Dear John,\n\nYour score has already been calculated for the private dataset, it is just not shown to you yet ! What you see as being your public score represents subset of the whole test dataset. Hence, it is not representative of your overall performance on the entire dataset. \n\nSo in simple words, if your kernel has succeeded on the public submission, it has already succeeded on the private submission.",
    "727200": "I see, thanks, Thomas!",
    "727379": "To add to what @dimartinot said, I think it’s also interesting to know why there’s a public &amp; private leader board.\n\nOne of them is to simulate a «&nbsp;deployment&nbsp;» of the model in production: if we deploy your model toady, and feed it new data you’ve never been able to test again, how would it perform?\n\nThen, these are competitions, done by smart people. People are going to try everything they can to try to gain every little advantage they can (and that’s great, that pushes the field ahead!) but not when it comes to «&nbsp;grey-area&nbsp;» practices. One such practice is leader-board dataset probing: Let’s say we’re trying to make a binary classifier, only 2 classes. But you don’t know the distribution of both classes, it could be 50/50, but maybe 20/80 or even 99/1. Knowing this would help you greatly! So if you could know your score on the leaderboard, even without seeing the data you could get some info. For example, make a submission with everything class 1, and then another one with every class 2. if you get an accuracy of 60% the 1st time and then 40%, all of a sudden you know the distribution without having ever seen the data! Hence the probing.\n\nOverall, it’s to prevent from fitting a model specifically on this dataset, but pushing people to make a model that generalises well.\n\nIn this competition for example, there are going to be unseen characters in the private LB. So combinations of vowel, consonants and roots that we haven’t seen in the training data (but with the same building blocks, no new vowels, consonants or roots). This will probably lead to some shake-up: when there’s a difference between public leaderboard, and private. This is pretty common in competitions, and can sometimes be drops (or rises) of &gt; 1,000 places!\n\nHope that helps!"
  },
  "source": "meta"
}