{
  "id": 4590,
  "title": "Cross validation results and test results",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4590",
  "author_name": "",
  "post_date": "2013-05-14T04:36:08.230Z",
  "votes": null,
  "comment_count": 4,
  "views": 1975,
  "content": "<p>Maybe because of the small labeled training examples, cross validation results are somewhat different from test results.</p>\r\n<p>For examples, the model 2% better than the previous model on cross validation is not better on test results.</p>\r\n<p>Does everyone suffer from similar problems? Any idea to predict test results accurately before submission?</p>\r\n<p></p>\r\n<p></p>",
  "messages": [
    {
      "id": "24289",
      "postDate": "05/14/2013 04:36:08",
      "content": "<p>Maybe because of the small labeled training examples, cross validation results are somewhat different from test results.</p>\r\n<p>For examples, the model 2% better than the previous model on cross validation is not better on test results.</p>\r\n<p>Does everyone suffer from similar problems? Any idea to predict test results accurately before submission?</p>\r\n<p></p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24290",
      "postDate": "05/14/2013 05:50:53",
      "content": "<p>Me too.</p>\r\n<p>I think it is about the inbalance of the data.</p>\r\n<p>The training data has more 1,2,3 than other label, but random submission will get score of 0.118, which is near 1/9, means that the final test data is balanced.</p>\r\n<p></p>\r\n<p>&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24291",
      "postDate": "05/14/2013 05:57:01",
      "content": "<p>https://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/forums/t/4302/methods-and-algorithms/23561#post23561</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24309",
      "postDate": "05/14/2013 16:58:19",
      "content": "<p>Random submission will have score 1/9 regardless of distribution</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24327",
      "postDate": "05/14/2013 23:09:17",
      "content": "<p>I have 4 different test harnesses, stupid, I know, but motived by the various methods I've tried. The N of local/leaderboard paired scores is small for each group (below). You can see most of by submissions are based on models from my python (pylearn2) harness\r\n (5 fold cross validation).</p>\r\n<ul>\r\n<li>caret_10fold_cv 5 </li><li>py_5fold_cv 15 </li><li>py_90_10_split 4 </li><li>r_70_30_split 6 </li></ul>\r\n<p>The pearson correlation scores for each group are as follows:</p>\r\n<ul>\r\n<li>caret_10fold_cv 0.2362019 </li><li>py_5fold_cv 0.6465968 </li><li>py_90_10_split 0.6512473 </li><li>r_70_30_split 0.4717204 </li></ul>\r\n<p>The spearman correlations for the groups are as follows:</p>\r\n<ul>\r\n<li>caret_10fold_cv 0.1000000 </li><li>py_5fold_cv 0.6961456 </li><li>py_90_10_split 0.8000000 </li><li>r_70_30_split 0.4857143 </li></ul>\r\n<p>Attached is a scatter plot of my local test harness scores against leaderboard scores, grouped by test harness. You can see in the graph that the the largest group is tightly bunched, basically because I've been submitting variations on the same model over\r\n and over for the last week - not enough variation to be interesting. The other groups have samples that are too small to be interesting I think.</p>\r\n<p>I don't think there is much to learn from this data, sadly so I can't say anything intelligent. I'd love to see a mass-adopted &quot;standardized&quot; test harness for a comp (one day...) and get correlation scores in aggregate. That would be cool.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 24290,
      "author_name": "binghsu",
      "author_url": "",
      "post_date": "05/14/2013 05:50:53",
      "content": "<p>Me too.</p>\r\n<p>I think it is about the inbalance of the data.</p>\r\n<p>The training data has more 1,2,3 than other label, but random submission will get score of 0.118, which is near 1/9, means that the final test data is balanced.</p>\r\n<p></p>\r\n<p>&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24291,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "05/14/2013 05:57:01",
      "content": "<p>https://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/forums/t/4302/methods-and-algorithms/23561#post23561</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24309,
      "author_name": "ccccat",
      "author_url": "",
      "post_date": "05/14/2013 16:58:19",
      "content": "<p>Random submission will have score 1/9 regardless of distribution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24327,
      "author_name": "jasonbrownlee",
      "author_url": "",
      "post_date": "05/14/2013 23:09:17",
      "content": "<p>I have 4 different test harnesses, stupid, I know, but motived by the various methods I've tried. The N of local/leaderboard paired scores is small for each group (below). You can see most of by submissions are based on models from my python (pylearn2) harness\r\n (5 fold cross validation).</p>\r\n<ul>\r\n<li>caret_10fold_cv 5 </li><li>py_5fold_cv 15 </li><li>py_90_10_split 4 </li><li>r_70_30_split 6 </li></ul>\r\n<p>The pearson correlation scores for each group are as follows:</p>\r\n<ul>\r\n<li>caret_10fold_cv 0.2362019 </li><li>py_5fold_cv 0.6465968 </li><li>py_90_10_split 0.6512473 </li><li>r_70_30_split 0.4717204 </li></ul>\r\n<p>The spearman correlations for the groups are as follows:</p>\r\n<ul>\r\n<li>caret_10fold_cv 0.1000000 </li><li>py_5fold_cv 0.6961456 </li><li>py_90_10_split 0.8000000 </li><li>r_70_30_split 0.4857143 </li></ul>\r\n<p>Attached is a scatter plot of my local test harness scores against leaderboard scores, grouped by test harness. You can see in the graph that the the largest group is tightly bunched, basically because I've been submitting variations on the same model over\r\n and over for the last week - not enough variation to be interesting. The other groups have samples that are too small to be interesting I think.</p>\r\n<p>I don't think there is much to learn from this data, sadly so I can't say anything intelligent. I'd love to see a mass-adopted &quot;standardized&quot; test harness for a comp (one day...) and get correlation scores in aggregate. That would be cool.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "24289": "",
    "24290": "",
    "24291": "",
    "24309": "",
    "24327": ""
  },
  "source": "meta"
}