{
  "id": 34061,
  "title": "How different is the distribution of training data from the test data?",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/34061",
  "author_name": "",
  "post_date": "2017-06-02T20:25:27.387830500Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I split my training into train and validation set, however, a model that performs well on my validation performs worse on the leaderboard while a model that does a poor job on my validation performs well on the leaderboard. What do you think might be the reason? my train and validation split is 70-30.</p>",
  "messages": [
    {
      "id": "188417",
      "postDate": "06/02/2017 20:25:27",
      "content": "<p>I split my training into train and validation set, however, a model that performs well on my validation performs worse on the leaderboard while a model that does a poor job on my validation performs well on the leaderboard. What do you think might be the reason? my train and validation split is 70-30.</p>",
      "rawMarkdown": "I split my training into train and validation set, however, a model that performs well on my validation performs worse on the leaderboard while a model that does a poor job on my validation performs well on the leaderboard. What do you think might be the reason? my train and validation split is 70-30.",
      "votes": null
    },
    {
      "id": "188423",
      "postDate": "06/02/2017 20:36:54",
      "content": "<p>Hmmm.... did you use stratified split? </p>",
      "rawMarkdown": "Hmmm.... did you use stratified split?",
      "votes": null
    },
    {
      "id": "188433",
      "postDate": "06/02/2017 20:57:11",
      "content": "<p>Probably not.. I am using the train validatioin split of keras here: <a href=\"https://www.kaggle.com/the1owl/artificial-intelligence-for-cc-screening\">https://www.kaggle.com/the1owl/artificial-intelligence-for-cc-screening</a>\nwhich I am guessing shuffles the data before splitting.</p>\n\n<p>Do you think that might be the reason? I will check the stratified split.</p>",
      "rawMarkdown": "Probably not.. I am using the train validatioin split of keras here: https://www.kaggle.com/the1owl/artificial-intelligence-for-cc-screening\nwhich I am guessing shuffles the data before splitting.\n\nDo you think that might be the reason? I will check the stratified split.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 188423,
      "author_name": "oysteijo",
      "author_url": "",
      "post_date": "06/02/2017 20:36:54",
      "content": "<p>Hmmm.... did you use stratified split? </p>",
      "votes": null,
      "replies": [
        {
          "id": 188433,
          "author_name": "charmchi",
          "author_url": "",
          "post_date": "06/02/2017 20:57:11",
          "content": "<p>Probably not.. I am using the train validatioin split of keras here: <a href=\"https://www.kaggle.com/the1owl/artificial-intelligence-for-cc-screening\">https://www.kaggle.com/the1owl/artificial-intelligence-for-cc-screening</a>\nwhich I am guessing shuffles the data before splitting.</p>\n\n<p>Do you think that might be the reason? I will check the stratified split.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "188417": "I split my training into train and validation set, however, a model that performs well on my validation performs worse on the leaderboard while a model that does a poor job on my validation performs well on the leaderboard. What do you think might be the reason? my train and validation split is 70-30.",
    "188423": "Hmmm.... did you use stratified split?",
    "188433": "Probably not.. I am using the train validatioin split of keras here: https://www.kaggle.com/the1owl/artificial-intelligence-for-cc-screening\nwhich I am guessing shuffles the data before splitting.\n\nDo you think that might be the reason? I will check the stratified split."
  },
  "source": "meta"
}