{
  "id": 8107,
  "title": "Clarification please: Public and Private split on the leaderboard",
  "url": "/competitions/decoding-the-human-brain/discussion/8107",
  "author_name": "",
  "post_date": "2014-05-10T07:07:16.440Z",
  "votes": null,
  "comment_count": 4,
  "views": 1460,
  "content": "<p>I apology if the following question can be easily inferred or has been asked before...</p>\n<p>On the evaluation page, it is said that:</p>\n<p>Notice that the (public) score on the leaderboard and the final (private) score of the competition are computed on trials from <em>different</em> sets of subjects.</p>\n<p>I got two possible interpretations:</p>\n<p>1) public and private splits are uniform randomly sampled from all trials of all test subjects. So, some of the trials of a specific subject can go into public split, and the rest trials of this subject go into private split.</p>\n<p>2) public and private splits contain different subjects. For example, all trials of test_subject17-19 go into public split, and those of test_subject20-23 go into private split.</p>\n<p>Though I personally think 2) is the most possible option, a clarification would be very appreciated.</p>\n<p>If 2) is the case, I think one can use a few submissions to infer which of the test subjects are in public split (and also&nbsp;the private split). If one can also manage to get (most of) the correct labels for those test subjects, and then use these samples together with the provided training set to train his model and get scored, is this approach allowed? Just being curious...</p>",
  "messages": [
    {
      "id": "44332",
      "postDate": "05/10/2014 07:07:16",
      "content": "<p>I apology if the following question can be easily inferred or has been asked before...</p>\n<p>On the evaluation page, it is said that:</p>\n<p>Notice that the (public) score on the leaderboard and the final (private) score of the competition are computed on trials from <em>different</em> sets of subjects.</p>\n<p>I got two possible interpretations:</p>\n<p>1) public and private splits are uniform randomly sampled from all trials of all test subjects. So, some of the trials of a specific subject can go into public split, and the rest trials of this subject go into private split.</p>\n<p>2) public and private splits contain different subjects. For example, all trials of test_subject17-19 go into public split, and those of test_subject20-23 go into private split.</p>\n<p>Though I personally think 2) is the most possible option, a clarification would be very appreciated.</p>\n<p>If 2) is the case, I think one can use a few submissions to infer which of the test subjects are in public split (and also&nbsp;the private split). If one can also manage to get (most of) the correct labels for those test subjects, and then use these samples together with the provided training set to train his model and get scored, is this approach allowed? Just being curious...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44400",
      "postDate": "05/12/2014 10:36:55",
      "content": "<p>Hi,</p>\n<p>In the evaluation page it is written that the public test set is made of&nbsp;the trials of some of the test subjects. The private test set is made of the trials of the remaining test subjects. So all the trials of a given test subject, either belong to the public test set or (XOR) to the private test set.</p>\n<p>If you want to use your past predictions to improve your future submissions, you are free to do that. In the specific case of this competition, I guess that this practice might have little scientific value. Nevertheless I expect&nbsp;that it would be extremely unlikely&nbsp;to get a significant&nbsp;improvement in the private score in that way.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44401",
      "postDate": "05/12/2014 11:23:40",
      "content": "<p>Thanks for the clarification!</p>\n<p>I just wanted to make sure my cross-validation methodology is appropriate with respect to the public and private split. Regarding your description, leave-one-subject out cross-validation seems the right way to go, which was also adopted in your paper:&nbsp;http://arxiv.org/abs/1404.4175</p>\n<p>In case anyone interested in the CV score and the public score, I got a difference around 0.025~0.04 (public LB is always better than local CV), though it is an observation of a few submissions so far. Would like to hear results from others.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44490",
      "postDate": "05/13/2014 21:17:49",
      "content": "<p>[quote=yr;44401]</p>\n<p>In case anyone interested in the CV score and the public score, I got a difference around 0.025~0.04 (public LB is always better than local CV), though it is an observation of a few submissions so far. Would like to hear results from others.</p>\n<p>[/quote]</p>\n<p>My current best method does a mean of 0.702 in leave-one-subject-out cross-validation (LOSOCV). This gives me a leaderboard score of&nbsp;0.676 (diff:&nbsp;0.026). In general I have seen&nbsp;0.01-0.03 variations between LOSOCV and the public leaderboard. My leaderboard scores are always lower than LOSOCV results.</p>\n<p>I would tend to trust LOSOCV more than the leaderboard as it represents a test on a larger pool of subjects. The risk, though, is to develop a method that is overfit to the train subjects.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44492",
      "postDate": "05/14/2014 00:55:01",
      "content": "<p>@fchollet, Thanks for sharing! In other comp, I heard people using the public LB score as another fold in the local CV and try to combine that score with the local CV score. I think that is doable here too. By doing so, it might give a better estimate of the model performance.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 44400,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "05/12/2014 10:36:55",
      "content": "<p>Hi,</p>\n<p>In the evaluation page it is written that the public test set is made of&nbsp;the trials of some of the test subjects. The private test set is made of the trials of the remaining test subjects. So all the trials of a given test subject, either belong to the public test set or (XOR) to the private test set.</p>\n<p>If you want to use your past predictions to improve your future submissions, you are free to do that. In the specific case of this competition, I guess that this practice might have little scientific value. Nevertheless I expect&nbsp;that it would be extremely unlikely&nbsp;to get a significant&nbsp;improvement in the private score in that way.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44401,
      "author_name": "chenglongchen",
      "author_url": "",
      "post_date": "05/12/2014 11:23:40",
      "content": "<p>Thanks for the clarification!</p>\n<p>I just wanted to make sure my cross-validation methodology is appropriate with respect to the public and private split. Regarding your description, leave-one-subject out cross-validation seems the right way to go, which was also adopted in your paper:&nbsp;http://arxiv.org/abs/1404.4175</p>\n<p>In case anyone interested in the CV score and the public score, I got a difference around 0.025~0.04 (public LB is always better than local CV), though it is an observation of a few submissions so far. Would like to hear results from others.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44490,
      "author_name": "fchollet",
      "author_url": "",
      "post_date": "05/13/2014 21:17:49",
      "content": "<p>[quote=yr;44401]</p>\n<p>In case anyone interested in the CV score and the public score, I got a difference around 0.025~0.04 (public LB is always better than local CV), though it is an observation of a few submissions so far. Would like to hear results from others.</p>\n<p>[/quote]</p>\n<p>My current best method does a mean of 0.702 in leave-one-subject-out cross-validation (LOSOCV). This gives me a leaderboard score of&nbsp;0.676 (diff:&nbsp;0.026). In general I have seen&nbsp;0.01-0.03 variations between LOSOCV and the public leaderboard. My leaderboard scores are always lower than LOSOCV results.</p>\n<p>I would tend to trust LOSOCV more than the leaderboard as it represents a test on a larger pool of subjects. The risk, though, is to develop a method that is overfit to the train subjects.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44492,
      "author_name": "chenglongchen",
      "author_url": "",
      "post_date": "05/14/2014 00:55:01",
      "content": "<p>@fchollet, Thanks for sharing! In other comp, I heard people using the public LB score as another fold in the local CV and try to combine that score with the local CV score. I think that is doable here too. By doing so, it might give a better estimate of the model performance.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "44332": "",
    "44400": "",
    "44401": "",
    "44490": "",
    "44492": ""
  },
  "source": "meta"
}