{
  "id": 165598,
  "title": "Good validation accuracy, but low test score",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165598",
  "author_name": "",
  "post_date": "2020-07-10T10:11:56.265484Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Here is my first post on Kaggle :)\nEven when I get training and validation accuracies higher than 95%, test score always stucks below 85%. What do you think about the possible causes of this issue? Here are some details about my implementation:\n-- First, I resize each image to 128x128 pixels (plan to increase it later).\n-- Since the data is quite imbalanced, I apply augmentation on the positive samples such that I generate 10 new samples out of each positive labeled training image.\n-- After the augmentation, there still exists a 5:1 ratio between the numbers of negative and positive samples. So, in the loss function, I use different weights for each class, i.e. class_weights = {0:0.20, 1:1.0}\n-- I use Keras' pretrained VGG-16 network (without the dense layers) for the images. Then, the output of the convolutional layers is concatenated with the contextual information provided in .csv file. Finally, we have 2 more dense layers and the output layer.\n-- I split the training dataset by 80-20% for training and validation (stratification is used).</p>\n\n<p>In conclusion, my model performs well (&gt;95%) on the validation set which is never seen by the model, but poorly (&lt;85%) on the test set. What do you think about it? Can it be about the unrealistic augmentations such that the model learns funny irrelevant features that does not exist in test images?\nSorry for the long post. Thanks in advance..</p>",
  "messages": [
    {
      "id": "922751",
      "postDate": "07/10/2020 10:11:56",
      "content": "<p>Here is my first post on Kaggle :)\nEven when I get training and validation accuracies higher than 95%, test score always stucks below 85%. What do you think about the possible causes of this issue? Here are some details about my implementation:\n-- First, I resize each image to 128x128 pixels (plan to increase it later).\n-- Since the data is quite imbalanced, I apply augmentation on the positive samples such that I generate 10 new samples out of each positive labeled training image.\n-- After the augmentation, there still exists a 5:1 ratio between the numbers of negative and positive samples. So, in the loss function, I use different weights for each class, i.e. class_weights = {0:0.20, 1:1.0}\n-- I use Keras' pretrained VGG-16 network (without the dense layers) for the images. Then, the output of the convolutional layers is concatenated with the contextual information provided in .csv file. Finally, we have 2 more dense layers and the output layer.\n-- I split the training dataset by 80-20% for training and validation (stratification is used).</p>\n\n<p>In conclusion, my model performs well (&gt;95%) on the validation set which is never seen by the model, but poorly (&lt;85%) on the test set. What do you think about it? Can it be about the unrealistic augmentations such that the model learns funny irrelevant features that does not exist in test images?\nSorry for the long post. Thanks in advance..</p>",
      "rawMarkdown": "Here is my first post on Kaggle :)\nEven when I get training and validation accuracies higher than 95%, test score always stucks below 85%. What do you think about the possible causes of this issue? Here are some details about my implementation:\n-- First, I resize each image to 128x128 pixels (plan to increase it later).\n-- Since the data is quite imbalanced, I apply augmentation on the positive samples such that I generate 10 new samples out of each positive labeled training image.\n-- After the augmentation, there still exists a 5:1 ratio between the numbers of negative and positive samples. So, in the loss function, I use different weights for each class, i.e. class_weights = {0:0.20, 1:1.0}\n-- I use Keras' pretrained VGG-16 network (without the dense layers) for the images. Then, the output of the convolutional layers is concatenated with the contextual information provided in .csv file. Finally, we have 2 more dense layers and the output layer.\n-- I split the training dataset by 80-20% for training and validation (stratification is used).\n\nIn conclusion, my model performs well (&gt;95%) on the validation set which is never seen by the model, but poorly (&lt;85%) on the test set. What do you think about it? Can it be about the unrealistic augmentations such that the model learns funny irrelevant features that does not exist in test images?\nSorry for the long post. Thanks in advance..",
      "votes": null
    },
    {
      "id": "922862",
      "postDate": "07/10/2020 11:20:48",
      "content": "<p>you should check the performance using the ROC curve instead of accuracy to get indication on the test set score, see the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/evaluation\">evaluation</a> page. High accuracy in this case is pretty easy because of the class imbalance problem. Hope this clears your doubts.</p>",
      "rawMarkdown": "you should check the performance using the ROC curve instead of accuracy to get indication on the test set score, see the [evaluation](https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/evaluation) page. High accuracy in this case is pretty easy because of the class imbalance problem. Hope this clears your doubts.",
      "votes": null
    },
    {
      "id": "922904",
      "postDate": "07/10/2020 12:05:52",
      "content": "<p>Thanks a lot! I understand what you mean. I was (mis)thinking that the test score was simply the accuracy after using threshold of 0.5 on the prediction probabilities.</p>",
      "rawMarkdown": "Thanks a lot! I understand what you mean. I was (mis)thinking that the test score was simply the accuracy after using threshold of 0.5 on the prediction probabilities.",
      "votes": null
    },
    {
      "id": "956671",
      "postDate": "08/03/2020 17:19:45",
      "content": "<p>You must make sure that all samples coming from an original image are kept in same fold (trian or validation).  If you augment data before fold split then you end up having the same image in train and valid data which inflates your score.</p>",
      "rawMarkdown": "You must make sure that all samples coming from an original image are kept in same fold (trian or validation).  If you augment data before fold split then you end up having the same image in train and valid data which inflates your score.",
      "votes": null
    },
    {
      "id": "957451",
      "postDate": "08/04/2020 10:17:13",
      "content": "<p>That perfectly makes sense. Thanks for the warning!</p>",
      "rawMarkdown": "That perfectly makes sense. Thanks for the warning!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 922862,
      "author_name": "pranavkasela",
      "author_url": "",
      "post_date": "07/10/2020 11:20:48",
      "content": "<p>you should check the performance using the ROC curve instead of accuracy to get indication on the test set score, see the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/evaluation\">evaluation</a> page. High accuracy in this case is pretty easy because of the class imbalance problem. Hope this clears your doubts.</p>",
      "votes": null,
      "replies": [
        {
          "id": 922904,
          "author_name": "oguzhansevim",
          "author_url": "",
          "post_date": "07/10/2020 12:05:52",
          "content": "<p>Thanks a lot! I understand what you mean. I was (mis)thinking that the test score was simply the accuracy after using threshold of 0.5 on the prediction probabilities.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 956671,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/03/2020 17:19:45",
      "content": "<p>You must make sure that all samples coming from an original image are kept in same fold (trian or validation).  If you augment data before fold split then you end up having the same image in train and valid data which inflates your score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 957451,
          "author_name": "oguzhansevim",
          "author_url": "",
          "post_date": "08/04/2020 10:17:13",
          "content": "<p>That perfectly makes sense. Thanks for the warning!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "922751": "Here is my first post on Kaggle :)\nEven when I get training and validation accuracies higher than 95%, test score always stucks below 85%. What do you think about the possible causes of this issue? Here are some details about my implementation:\n-- First, I resize each image to 128x128 pixels (plan to increase it later).\n-- Since the data is quite imbalanced, I apply augmentation on the positive samples such that I generate 10 new samples out of each positive labeled training image.\n-- After the augmentation, there still exists a 5:1 ratio between the numbers of negative and positive samples. So, in the loss function, I use different weights for each class, i.e. class_weights = {0:0.20, 1:1.0}\n-- I use Keras' pretrained VGG-16 network (without the dense layers) for the images. Then, the output of the convolutional layers is concatenated with the contextual information provided in .csv file. Finally, we have 2 more dense layers and the output layer.\n-- I split the training dataset by 80-20% for training and validation (stratification is used).\n\nIn conclusion, my model performs well (&gt;95%) on the validation set which is never seen by the model, but poorly (&lt;85%) on the test set. What do you think about it? Can it be about the unrealistic augmentations such that the model learns funny irrelevant features that does not exist in test images?\nSorry for the long post. Thanks in advance..",
    "922862": "you should check the performance using the ROC curve instead of accuracy to get indication on the test set score, see the [evaluation](https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/evaluation) page. High accuracy in this case is pretty easy because of the class imbalance problem. Hope this clears your doubts.",
    "922904": "Thanks a lot! I understand what you mean. I was (mis)thinking that the test score was simply the accuracy after using threshold of 0.5 on the prediction probabilities.",
    "956671": "You must make sure that all samples coming from an original image are kept in same fold (trian or validation).  If you augment data before fold split then you end up having the same image in train and valid data which inflates your score.",
    "957451": "That perfectly makes sense. Thanks for the warning!"
  },
  "source": "meta"
}