{
  "id": 73383,
  "title": "big performance gap between validation and test",
  "url": "/competitions/histopathologic-cancer-detection/discussion/73383",
  "author_name": "Xiuchao",
  "post_date": "2018-12-02T14:49:52.921000",
  "votes": 0,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I used densenet121 as the base CNN and got 99% AUC on 20% held-out data, but on the test data it's only around 96%. Anyone has any idea to explain this gap? If the training and test data were from the same distribution, as the 20% validation data are big enough, the performance should be very close to that on the test data. Hmmm...</p>",
  "messages": [
    {
      "id": 482700,
      "postDate": "2019-03-03T14:37:28.563Z",
      "content": "<p>Experiencing the same issue. AUC on train 99.7, AUC on valid 99.7, AUC on test 96.7. <br>\nScore changes on train/valid does not correlate with score changes on test. </p>\n\n<p>The reason, I guess, is that test set consists of new, unseen WSIs parsed into patches. Random train/validation splits would share patches of the same WSI (<a href=\"https://camelyon16.grand-challenge.org/Data/\">https://camelyon16.grand-challenge.org/Data/</a>). Thats the leak. To make proper train/val split we need to have WSI ids for every patch id. </p>",
      "rawMarkdown": "Experiencing the same issue. AUC on train 99.7, AUC on valid 99.7, AUC on test 96.7.  \nScore changes on train/valid does not correlate with score changes on test. \n\nThe reason, I guess, is that test set consists of new, unseen WSIs parsed into patches. Random train/validation splits would share patches of the same WSI (https://camelyon16.grand-challenge.org/Data/). Thats the leak. To make proper train/val split we need to have WSI ids for every patch id. ",
      "votes": 4,
      "replies": [
        {
          "id": 487712,
          "postDate": "2019-03-11T11:25:24.920Z",
          "content": "<p>It inspired me. Thank you a lot.</p>",
          "rawMarkdown": "It inspired me. Thank you a lot.",
          "votes": 3
        },
        {
          "id": 488462,
          "postDate": "2019-03-12T13:46:00.533Z",
          "content": "<p>How do you split train/val ? Could you share your idea with us?</p>",
          "rawMarkdown": "How do you split train/val ? Could you share your idea with us?",
          "votes": 3
        },
        {
          "id": 488494,
          "postDate": "2019-03-12T14:45:50.340Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 488552,
          "postDate": "2019-03-12T16:54:09.783Z",
          "content": "<p><a href=\"https://drive.google.com/open?id=1NgE2Uuwhr3yDPVwVNpmSwJwQG1RIIJsC\">enjoy</a>\nUPDT: link updated</p>",
          "rawMarkdown": "[enjoy](https://drive.google.com/open?id=1NgE2Uuwhr3yDPVwVNpmSwJwQG1RIIJsC)\nUPDT: link updated",
          "votes": 1
        },
        {
          "id": 488563,
          "postDate": "2019-03-12T17:21:58.740Z",
          "content": "<p>@SM Thank you, but the link does not work. \nUPDT: Thank you, it works! Does this replace the original trains_label? </p>",
          "rawMarkdown": "@SM Thank you, but the link does not work. \nUPDT: Thank you, it works! Does this replace the original trains_label? "
        },
        {
          "id": 488587,
          "postDate": "2019-03-12T18:04:50.243Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 450109,
      "postDate": "2019-01-04T10:19:48.083Z",
      "content": "<p>Well looks like you might overfitted your validation set. Say you have 100% on train and 99% on val, that's a little overfitting to the train data. Then if you pick your best model on val then this overfitted model doesn't generalize well to the test set. So maybe you can pick a model that is 98% on train and 98% on val, then the test might be close to 98%.</p>",
      "rawMarkdown": "Well looks like you might overfitted your validation set. Say you have 100% on train and 99% on val, that's a little overfitting to the train data. Then if you pick your best model on val then this overfitted model doesn't generalize well to the test set. So maybe you can pick a model that is 98% on train and 98% on val, then the test might be close to 98%.",
      "votes": 1
    },
    {
      "id": 431911,
      "postDate": "2018-12-03T04:47:36.927Z",
      "content": "<p>It seems that this is caused by the randomness inherit in any data set. Testing over different sets of held out data can result in different accuracies or scoring method. A way to get a more robust and cohesive estimation of scoring would be to use cross-validation and average the results. I hope that helps.</p>",
      "rawMarkdown": "It seems that this is caused by the randomness inherit in any data set. Testing over different sets of held out data can result in different accuracies or scoring method. A way to get a more robust and cohesive estimation of scoring would be to use cross-validation and average the results. I hope that helps.",
      "votes": 2,
      "replies": [
        {
          "id": 435662,
          "postDate": "2018-12-08T13:53:45.157Z",
          "content": "<p>Thanks for your suggestion. I've run the code several times and each time the training/val are split randomly. The results are always around 99% on the val data, but on LB the results are always around 96%. So seems the gap is there consistently, which is really weird...</p>",
          "rawMarkdown": "Thanks for your suggestion. I've run the code several times and each time the training/val are split randomly. The results are always around 99% on the val data, but on LB the results are always around 96%. So seems the gap is there consistently, which is really weird..."
        }
      ]
    },
    {
      "id": 481792,
      "postDate": "2019-03-01T21:37:55.763Z",
      "content": "<p>How come no one has mentioned it:\nAs you said: \"if the training and test data were from the same distribution, as the 20% validation data are big enough, the performance should be very close to that on the test data.\" Then the fact that the scores are different means your assumption is not true, right? I am experiencing the same thing. So I decided to make the two have a more similar distribution. I will do some experiment before coming back. Have you try anything else?</p>",
      "rawMarkdown": "How come no one has mentioned it:\nAs you said: \"if the training and test data were from the same distribution, as the 20% validation data are big enough, the performance should be very close to that on the test data.\" Then the fact that the scores are different means your assumption is not true, right? I am experiencing the same thing. So I decided to make the two have a more similar distribution. I will do some experiment before coming back. Have you try anything else?"
    },
    {
      "id": 456292,
      "postDate": "2019-01-15T14:03:47.903Z",
      "content": "<p>Have tried training set augmentation ? Flip some images and other common transforms to increase the variety in your training set ?</p>",
      "rawMarkdown": "Have tried training set augmentation ? Flip some images and other common transforms to increase the variety in your training set ?"
    },
    {
      "id": 433427,
      "postDate": "2018-12-05T04:06:44.903Z",
      "content": "<p>I thought this is a typical problem of overfitting. Did you try to use some dropouts?</p>",
      "rawMarkdown": "I thought this is a typical problem of overfitting. Did you try to use some dropouts?",
      "replies": [
        {
          "id": 435665,
          "postDate": "2018-12-08T13:57:58.530Z",
          "content": "<p>Thanks. Yeah I did use dropout and its variants... to little avail. As I replied Marco's comment, each training I split training/val randomly and the performance gap between val/test is always there. </p>",
          "rawMarkdown": "Thanks. Yeah I did use dropout and its variants... to little avail. As I replied Marco's comment, each training I split training/val randomly and the performance gap between val/test is always there. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 432926,
      "postDate": "2018-12-04T13:31:18.827Z",
      "content": "<p>maybe you can try models stacking</p>",
      "rawMarkdown": "maybe you can try models stacking",
      "replies": [
        {
          "id": 435664,
          "postDate": "2018-12-08T13:55:34.423Z",
          "content": "<p>Thanks. I'm considering to train two models: one in the first pass, and the second on those instances that the first model doesn't perform well. Then combine their predictions. Not sure if that will improve.</p>",
          "rawMarkdown": "Thanks. I'm considering to train two models: one in the first pass, and the second on those instances that the first model doesn't perform well. Then combine their predictions. Not sure if that will improve."
        }
      ]
    },
    {
      "id": 431582,
      "postDate": "2018-12-02T14:49:52.920Z",
      "content": "<p>I used densenet121 as the base CNN and got 99% AUC on 20% held-out data, but on the test data it's only around 96%. Anyone has any idea to explain this gap? If the training and test data were from the same distribution, as the 20% validation data are big enough, the performance should be very close to that on the test data. Hmmm...</p>",
      "rawMarkdown": "I used densenet121 as the base CNN and got 99% AUC on 20% held-out data, but on the test data it's only around 96%. Anyone has any idea to explain this gap? If the training and test data were from the same distribution, as the 20% validation data are big enough, the performance should be very close to that on the test data. Hmmm...\n"
    }
  ],
  "comments": [
    {
      "id": 482700,
      "author_name": "SM",
      "author_url": "",
      "post_date": "2019-03-03T14:37:28.563000",
      "content": "<p>Experiencing the same issue. AUC on train 99.7, AUC on valid 99.7, AUC on test 96.7. <br>\nScore changes on train/valid does not correlate with score changes on test. </p>\n\n<p>The reason, I guess, is that test set consists of new, unseen WSIs parsed into patches. Random train/validation splits would share patches of the same WSI (<a href=\"https://camelyon16.grand-challenge.org/Data/\">https://camelyon16.grand-challenge.org/Data/</a>). Thats the leak. To make proper train/val split we need to have WSI ids for every patch id. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 487712,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "2019-03-11T11:25:24.920000",
          "content": "<p>It inspired me. Thank you a lot.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 488462,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "2019-03-12T13:46:00.533000",
          "content": "<p>How do you split train/val ? Could you share your idea with us?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 488494,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-03-12T14:45:50.340000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488552,
          "author_name": "SM",
          "author_url": "",
          "post_date": "2019-03-12T16:54:09.783000",
          "content": "<p><a href=\"https://drive.google.com/open?id=1NgE2Uuwhr3yDPVwVNpmSwJwQG1RIIJsC\">enjoy</a>\nUPDT: link updated</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 488563,
          "author_name": "William Green",
          "author_url": "",
          "post_date": "2019-03-12T17:21:58.740000",
          "content": "<p>@SM Thank you, but the link does not work. \nUPDT: Thank you, it works! Does this replace the original trains_label? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488587,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-03-12T18:04:50.243000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 450109,
      "author_name": "JayChen",
      "author_url": "",
      "post_date": "2019-01-04T10:19:48.083000",
      "content": "<p>Well looks like you might overfitted your validation set. Say you have 100% on train and 99% on val, that's a little overfitting to the train data. Then if you pick your best model on val then this overfitted model doesn't generalize well to the test set. So maybe you can pick a model that is 98% on train and 98% on val, then the test might be close to 98%.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 431911,
      "author_name": "Marco Gancitano",
      "author_url": "",
      "post_date": "2018-12-03T04:47:36.927000",
      "content": "<p>It seems that this is caused by the randomness inherit in any data set. Testing over different sets of held out data can result in different accuracies or scoring method. A way to get a more robust and cohesive estimation of scoring would be to use cross-validation and average the results. I hope that helps.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 435662,
          "author_name": "Xiuchao",
          "author_url": "",
          "post_date": "2018-12-08T13:53:45.157000",
          "content": "<p>Thanks for your suggestion. I've run the code several times and each time the training/val are split randomly. The results are always around 99% on the val data, but on LB the results are always around 96%. So seems the gap is there consistently, which is really weird...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 481792,
      "author_name": "Hanke Chen",
      "author_url": "",
      "post_date": "2019-03-01T21:37:55.763000",
      "content": "<p>How come no one has mentioned it:\nAs you said: \"if the training and test data were from the same distribution, as the 20% validation data are big enough, the performance should be very close to that on the test data.\" Then the fact that the scores are different means your assumption is not true, right? I am experiencing the same thing. So I decided to make the two have a more similar distribution. I will do some experiment before coming back. Have you try anything else?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 456292,
      "author_name": "Manjunath Sripadarao",
      "author_url": "",
      "post_date": "2019-01-15T14:03:47.903000",
      "content": "<p>Have tried training set augmentation ? Flip some images and other common transforms to increase the variety in your training set ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433427,
      "author_name": "Dr. Devasia Kurian",
      "author_url": "",
      "post_date": "2018-12-05T04:06:44.903000",
      "content": "<p>I thought this is a typical problem of overfitting. Did you try to use some dropouts?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 435665,
          "author_name": "Xiuchao",
          "author_url": "",
          "post_date": "2018-12-08T13:57:58.530000",
          "content": "<p>Thanks. Yeah I did use dropout and its variants... to little avail. As I replied Marco's comment, each training I split training/val randomly and the performance gap between val/test is always there. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 432926,
      "author_name": "Luque do",
      "author_url": "",
      "post_date": "2018-12-04T13:31:18.827000",
      "content": "<p>maybe you can try models stacking</p>",
      "votes": 0,
      "replies": [
        {
          "id": 435664,
          "author_name": "Xiuchao",
          "author_url": "",
          "post_date": "2018-12-08T13:55:34.423000",
          "content": "<p>Thanks. I'm considering to train two models: one in the first pass, and the second on those instances that the first model doesn't perform well. Then combine their predictions. Not sure if that will improve.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "482700": "Experiencing the same issue. AUC on train 99.7, AUC on valid 99.7, AUC on test 96.7.  \nScore changes on train/valid does not correlate with score changes on test. \n\nThe reason, I guess, is that test set consists of new, unseen WSIs parsed into patches. Random train/validation splits would share patches of the same WSI (https://camelyon16.grand-challenge.org/Data/). Thats the leak. To make proper train/val split we need to have WSI ids for every patch id. ",
    "450109": "Well looks like you might overfitted your validation set. Say you have 100% on train and 99% on val, that's a little overfitting to the train data. Then if you pick your best model on val then this overfitted model doesn't generalize well to the test set. So maybe you can pick a model that is 98% on train and 98% on val, then the test might be close to 98%.",
    "431911": "It seems that this is caused by the randomness inherit in any data set. Testing over different sets of held out data can result in different accuracies or scoring method. A way to get a more robust and cohesive estimation of scoring would be to use cross-validation and average the results. I hope that helps.",
    "481792": "How come no one has mentioned it:\nAs you said: \"if the training and test data were from the same distribution, as the 20% validation data are big enough, the performance should be very close to that on the test data.\" Then the fact that the scores are different means your assumption is not true, right? I am experiencing the same thing. So I decided to make the two have a more similar distribution. I will do some experiment before coming back. Have you try anything else?",
    "456292": "Have tried training set augmentation ? Flip some images and other common transforms to increase the variety in your training set ?",
    "433427": "I thought this is a typical problem of overfitting. Did you try to use some dropouts?",
    "432926": "maybe you can try models stacking",
    "431582": "I used densenet121 as the base CNN and got 99% AUC on 20% held-out data, but on the test data it's only around 96%. Anyone has any idea to explain this gap? If the training and test data were from the same distribution, as the 20% validation data are big enough, the performance should be very close to that on the test data. Hmmm...\n"
  }
}