{
  "id": 94199,
  "title": "High Accuracy on train and validation data set, but low accuracy on test data, Why???",
  "url": "/competitions/histopathologic-cancer-detection/discussion/94199",
  "author_name": "",
  "post_date": "2019-06-03T05:03:30.909103100Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I get 95%+ accuracy on train/validation data (I split the train data into 0.8-0.2), but only 78% accuracy on the test data, why???</p>",
  "messages": [
    {
      "id": "541837",
      "postDate": "06/03/2019 05:03:30",
      "content": "<p>I get 95%+ accuracy on train/validation data (I split the train data into 0.8-0.2), but only 78% accuracy on the test data, why???</p>",
      "rawMarkdown": "I get 95%+ accuracy on train/validation data (I split the train data into 0.8-0.2), but only 78% accuracy on the test data, why???",
      "votes": null
    },
    {
      "id": "542096",
      "postDate": "06/03/2019 13:03:44",
      "content": "<p>I think this is the case of Overfitting. You need to check again.</p>",
      "rawMarkdown": "I think this is the case of Overfitting. You need to check again.",
      "votes": null
    },
    {
      "id": "542160",
      "postDate": "06/03/2019 14:18:16",
      "content": "<p>Whenever we observe high accuracy on training set but low accuracy on test set, it is likely the case of overfitting. </p>\n\n<p>After train-test split, the model learn the details from the training set and predict on the test set. In case of overfitting, the model learn the details along with noise and random fluctuations in the training data. So, the noise and random fluctuations is learned as concepts by the model. </p>\n\n<p>The problem of overfitting occurs as these noise and random fluctuations do not apply to test set. It negatively impacts the model ability to generalize and we get low test accuracy. </p>",
      "rawMarkdown": "Whenever we observe high accuracy on training set but low accuracy on test set, it is likely the case of overfitting. \n\nAfter train-test split, the model learn the details from the training set and predict on the test set. In case of overfitting, the model learn the details along with noise and random fluctuations in the training data. So, the noise and random fluctuations is learned as concepts by the model. \n\nThe problem of overfitting occurs as these noise and random fluctuations do not apply to test set. It negatively impacts the model ability to generalize and we get low test accuracy.",
      "votes": null
    },
    {
      "id": "542167",
      "postDate": "06/03/2019 14:27:38",
      "content": "<p>Perhaps your model is overfitted</p>",
      "rawMarkdown": "Perhaps your model is overfitted",
      "votes": null
    },
    {
      "id": "552979",
      "postDate": "06/14/2019 22:43:42",
      "content": "<p>It is not over-fitting but a leak. Check discussion on <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\">WSIs</a>. Essentially, the thing u need is doing train/val spit based on WSI indexes. Without this CV is not reliable: it is quite simple to get 0.995-0.997 without WSI based split (that corresponds to 0.96-0.97 LB), though if you continue training you val would increase but LB score drop.</p>",
      "rawMarkdown": "It is not over-fitting but a leak. Check discussion on [WSIs](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760). Essentially, the thing u need is doing train/val spit based on WSI indexes. Without this CV is not reliable: it is quite simple to get 0.995-0.997 without WSI based split (that corresponds to 0.96-0.97 LB), though if you continue training you val would increase but LB score drop.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 542096,
      "author_name": "shwetagoyal4",
      "author_url": "",
      "post_date": "06/03/2019 13:03:44",
      "content": "<p>I think this is the case of Overfitting. You need to check again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542160,
      "author_name": "prashant111",
      "author_url": "",
      "post_date": "06/03/2019 14:18:16",
      "content": "<p>Whenever we observe high accuracy on training set but low accuracy on test set, it is likely the case of overfitting. </p>\n\n<p>After train-test split, the model learn the details from the training set and predict on the test set. In case of overfitting, the model learn the details along with noise and random fluctuations in the training data. So, the noise and random fluctuations is learned as concepts by the model. </p>\n\n<p>The problem of overfitting occurs as these noise and random fluctuations do not apply to test set. It negatively impacts the model ability to generalize and we get low test accuracy. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 542167,
      "author_name": "yunwanx1",
      "author_url": "",
      "post_date": "06/03/2019 14:27:38",
      "content": "<p>Perhaps your model is overfitted</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 552979,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "06/14/2019 22:43:42",
      "content": "<p>It is not over-fitting but a leak. Check discussion on <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\">WSIs</a>. Essentially, the thing u need is doing train/val spit based on WSI indexes. Without this CV is not reliable: it is quite simple to get 0.995-0.997 without WSI based split (that corresponds to 0.96-0.97 LB), though if you continue training you val would increase but LB score drop.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "541837": "I get 95%+ accuracy on train/validation data (I split the train data into 0.8-0.2), but only 78% accuracy on the test data, why???",
    "542096": "I think this is the case of Overfitting. You need to check again.",
    "542160": "Whenever we observe high accuracy on training set but low accuracy on test set, it is likely the case of overfitting. \n\nAfter train-test split, the model learn the details from the training set and predict on the test set. In case of overfitting, the model learn the details along with noise and random fluctuations in the training data. So, the noise and random fluctuations is learned as concepts by the model. \n\nThe problem of overfitting occurs as these noise and random fluctuations do not apply to test set. It negatively impacts the model ability to generalize and we get low test accuracy.",
    "542167": "Perhaps your model is overfitted",
    "552979": "It is not over-fitting but a leak. Check discussion on [WSIs](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760). Essentially, the thing u need is doing train/val spit based on WSI indexes. Without this CV is not reliable: it is quite simple to get 0.995-0.997 without WSI based split (that corresponds to 0.96-0.97 LB), though if you continue training you val would increase but LB score drop."
  },
  "source": "meta"
}