{
  "id": 8058,
  "title": "on the calculated scores.",
  "url": "/competitions/decoding-the-human-brain/discussion/8058",
  "author_name": "",
  "post_date": "2014-05-06T16:42:01.590Z",
  "votes": null,
  "comment_count": 2,
  "views": 984,
  "content": "<p>Does this&nbsp;43% of the test data on which the score is calculated includes samples of all test subjects?</p>\n<p>It seems to me it uses only 3 out of 7 subjects, and always the same ones. If this is the case the final score with all subjects may vary substantially. &nbsp;</p>",
  "messages": [
    {
      "id": "44063",
      "postDate": "05/06/2014 16:42:01",
      "content": "<p>Does this&nbsp;43% of the test data on which the score is calculated includes samples of all test subjects?</p>\n<p>It seems to me it uses only 3 out of 7 subjects, and always the same ones. If this is the case the final score with all subjects may vary substantially. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44157",
      "postDate": "05/07/2014 09:26:42",
      "content": "<p>Hi Luis,</p>\n<p>As you can read in the&nbsp;<a href=\"http://www.kaggle.com/c/decoding-the-human-brain/details/evaluation\">evaluation page</a>, &quot;the (public) score on the leaderboard and the final (private) score of the competition are computed on trials from different sets of subjects&quot;. We specifically designed the competition to be like that because we are interested in assessing how good&nbsp;can be the prediction on future trials from&nbsp;<em>unseen</em> subjects.</p>\n<p>If you keep submitting and taking decision from the public score, you will end up overfitting the part of the test set on which that score is computed. In the long term this may harm your private score, which is computed on other subjects.</p>\n<p>Needless to say, among Kagglers it is common experience to &quot;overfit&quot; the leaderboard and to suffer from it when the private score is revealed and the final ranking differs from the public leaderboard.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44204",
      "postDate": "05/07/2014 20:05:55",
      "content": "<p>Thank you Emanuele. I understand and agree this is the correct approach.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 44157,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "05/07/2014 09:26:42",
      "content": "<p>Hi Luis,</p>\n<p>As you can read in the&nbsp;<a href=\"http://www.kaggle.com/c/decoding-the-human-brain/details/evaluation\">evaluation page</a>, &quot;the (public) score on the leaderboard and the final (private) score of the competition are computed on trials from different sets of subjects&quot;. We specifically designed the competition to be like that because we are interested in assessing how good&nbsp;can be the prediction on future trials from&nbsp;<em>unseen</em> subjects.</p>\n<p>If you keep submitting and taking decision from the public score, you will end up overfitting the part of the test set on which that score is computed. In the long term this may harm your private score, which is computed on other subjects.</p>\n<p>Needless to say, among Kagglers it is common experience to &quot;overfit&quot; the leaderboard and to suffer from it when the private score is revealed and the final ranking differs from the public leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44204,
      "author_name": "luisgarciadominguez",
      "author_url": "",
      "post_date": "05/07/2014 20:05:55",
      "content": "<p>Thank you Emanuele. I understand and agree this is the correct approach.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "44063": "",
    "44157": "",
    "44204": ""
  },
  "source": "meta"
}