{
  "id": 4269,
  "title": "Final submission for consideration on private test data",
  "url": "/competitions/icdar2013-gender-prediction-from-handwriting/discussion/4269",
  "author_name": "",
  "post_date": "2013-04-11T12:13:36.557Z",
  "votes": 1,
  "comment_count": 2,
  "views": 2008,
  "content": "<p>Given that the number of observations in the training and test data sets is so small, there is a high chance of over-fitting the public leaderboard. In this regard, selecting only one entry for final submission seems to be a risky game.</p>\r\n<p>Any reason why we have the option of choosing only one entry as opposed to the standard 5 that is there in most Kaggle competitions?</p>",
  "messages": [
    {
      "id": "22617",
      "postDate": "04/11/2013 12:13:36",
      "content": "<p>Given that the number of observations in the training and test data sets is so small, there is a high chance of over-fitting the public leaderboard. In this regard, selecting only one entry for final submission seems to be a risky game.</p>\r\n<p>Any reason why we have the option of choosing only one entry as opposed to the standard 5 that is there in most Kaggle competitions?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22673",
      "postDate": "04/12/2013 02:16:35",
      "content": "<p>i agree. this is my first formal competition, and i found that my score on the public leaderboard is much different from the cross-validation score on the training set. Some parameter set works good on the training set, while some others works good on the\r\n leaderboard, it is hard to decide which one is better.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22696",
      "postDate": "04/12/2013 16:09:32",
      "content": "<p>Here is Ben Hamner's answer:</p>\r\n<p>[quote]</p>\r\n<p><span>No, doesn't make sense. If it's a lottery or overfitting with 1 then it's still a lottery or overfitting with 5. You don't have the luxury in production systems to select your best model after the fact, so we've changed our defaults to 1 to account\r\n for this.</span></p>\r\n<p>[/quote]</p>\r\n<p>I personally have a different opinion, not to prevent participants from overfitting, but because in such competitions, it is allowed to submit more than one system.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 22673,
      "author_name": "zouxiaochuan",
      "author_url": "",
      "post_date": "04/12/2013 02:16:35",
      "content": "<p>i agree. this is my first formal competition, and i found that my score on the public leaderboard is much different from the cross-validation score on the training set. Some parameter set works good on the training set, while some others works good on the\r\n leaderboard, it is hard to decide which one is better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22696,
      "author_name": "ahassaine",
      "author_url": "",
      "post_date": "04/12/2013 16:09:32",
      "content": "<p>Here is Ben Hamner's answer:</p>\r\n<p>[quote]</p>\r\n<p><span>No, doesn't make sense. If it's a lottery or overfitting with 1 then it's still a lottery or overfitting with 5. You don't have the luxury in production systems to select your best model after the fact, so we've changed our defaults to 1 to account\r\n for this.</span></p>\r\n<p>[/quote]</p>\r\n<p>I personally have a different opinion, not to prevent participants from overfitting, but because in such competitions, it is allowed to submit more than one system.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "22617": "",
    "22673": "",
    "22696": ""
  },
  "source": "meta"
}