{
  "id": 273639,
  "title": "High validation AUC vs. Low score (85% vs. 57%!)",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273639",
  "author_name": "",
  "post_date": "2021-09-21T23:51:39.270478700Z",
  "votes": 14,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello every one. After challenging for a long time with the dataset and trying different models, I have finally found a model that  is being trained well (based on Training &amp; Validation accuracy and AUC curves). But when I submit the result to the competition, the score is not comparable. I am confused that why is it happening? </p>\n<p><img src=\"https://i.ibb.co/xSFkB6W/image-2021-09-20-T19-11-22-574-Z.png\" alt=\"Here\"> is a figure that shows the Loss, accuracy and AUC curves of my model. You can see that the maximum validation and training AUC is arrived to ~85% but when submitting, my score is 57%! </p>\n<p>Any opinion would be appreciated!</p>",
  "messages": [
    {
      "id": "1519807",
      "postDate": "09/21/2021 23:51:39",
      "content": "<p>Hello every one. After challenging for a long time with the dataset and trying different models, I have finally found a model that  is being trained well (based on Training &amp; Validation accuracy and AUC curves). But when I submit the result to the competition, the score is not comparable. I am confused that why is it happening? </p>\n<p><img src=\"https://i.ibb.co/xSFkB6W/image-2021-09-20-T19-11-22-574-Z.png\" alt=\"Here\"> is a figure that shows the Loss, accuracy and AUC curves of my model. You can see that the maximum validation and training AUC is arrived to ~85% but when submitting, my score is 57%! </p>\n<p>Any opinion would be appreciated!</p>",
      "rawMarkdown": "Hello every one. After challenging for a long time with the dataset and trying different models, I have finally found a model that  is being trained well (based on Training & Validation accuracy and AUC curves). But when I submit the result to the competition, the score is not comparable. I am confused that why is it happening? \n\n![Here](https://i.ibb.co/xSFkB6W/image-2021-09-20-T19-11-22-574-Z.png) is a figure that shows the Loss, accuracy and AUC curves of my model. You can see that the maximum validation and training AUC is arrived to ~85% but when submitting, my score is 57%! \n\nAny opinion would be appreciated!",
      "votes": null
    },
    {
      "id": "1519889",
      "postDate": "09/22/2021 02:48:32",
      "content": "<p><a href=\"https://www.kaggle.com/niloofarrahmani\" target=\"_blank\">@niloofarrahmani</a> , Some notes below. </p>\n<ol>\n<li>I had something similar when training with images of size 64x64x64. </li>\n<li>My validation AUC is not as good as yours (~85% is amazing ); </li>\n<li>mine was ~72% and then the public test submission was less than 0.6. </li>\n<li>I am not sure why this is happening. I have only done 80/20 split and still need to do 3/5-fold CV. </li>\n<li>are you getting 85% with 3/5-fold CV ? If yes that is quite impressive. </li>\n<li>My thinking is that you should get good valid AUC and good submission score.</li>\n<li>So if you have a model which is 67% in both cases maybe that is a more reliable model ? Not sure, curious to hear what others say. </li>\n<li>Also when you pick a model do you use valid_AUC or valid_loss ? maybe valid_loss is a better criteria ? </li>\n</ol>",
      "rawMarkdown": "niloofarrahmani , Some notes below. \n1. I had something similar when training with images of size 64x64x64. \n2. My validation AUC is not as good as yours (~85% is amazing ); \n3. mine was ~72% and then the public test submission was less than 0.6. \n4. I am not sure why this is happening. I have only done 80/20 split and still need to do 3/5-fold CV. \n5. are you getting 85% with 3/5-fold CV ? If yes that is quite impressive. \n6. My thinking is that you should get good valid AUC and good submission score.\n7.  So if you have a model which is 67% in both cases maybe that is a more reliable model ? Not sure, curious to hear what others say. \n8. Also when you pick a model do you use valid_AUC or valid_loss ? maybe valid_loss is a better criteria ?",
      "votes": null
    },
    {
      "id": "1520016",
      "postDate": "09/22/2021 05:09:04",
      "content": "<p>hi ,<br>\nThe error is being caused because the public lb is only being measured on 87 samples not so good to be honest , so when you run inference on 87 samples which is very low , the metrics are also not that equivalent to what you measured over more samples , it can be more it can be less . But dont worry trust your cv or the score you are seeing , the private lb has 400 samples , in the private lb we will see a huge shake . I have a question to ask why is your validation loss same always ?</p>",
      "rawMarkdown": "hi ,\nThe error is being caused because the public lb is only being measured on 87 samples not so good to be honest , so when you run inference on 87 samples which is very low , the metrics are also not that equivalent to what you measured over more samples , it can be more it can be less . But dont worry trust your cv or the score you are seeing , the private lb has 400 samples , in the private lb we will see a huge shake . I have a question to ask why is your validation loss same always ?",
      "votes": null
    },
    {
      "id": "1520025",
      "postDate": "09/22/2021 05:16:45",
      "content": "<p>Hmm. Good point. But still if you have a good model that generalizes well; why would it only be bad on these 87 cases ? <br>\nPerhaps it is possible and I am missing something ? </p>",
      "rawMarkdown": "Hmm. Good point. But still if you have a good model that generalizes well; why would it only be bad on these 87 cases ? \nPerhaps it is possible and I am missing something ?",
      "votes": null
    },
    {
      "id": "1520033",
      "postDate": "09/22/2021 05:28:29",
      "content": "<p>The testing set may have some differences from valid . I would recommend you to split the data into a custom test set , sometimes people overfit to valid set . Thus three sets train valid and test</p>",
      "rawMarkdown": "The testing set may have some differences from valid . I would recommend you to split the data into a custom test set , sometimes people overfit to valid set . Thus three sets train valid and test",
      "votes": null
    },
    {
      "id": "1520037",
      "postDate": "09/22/2021 05:31:22",
      "content": "<p>also if you are using something like 2d cnn then you should measure CV not only parallely for each slice also how you are predicting a single prediction for all slices</p>",
      "rawMarkdown": "also if you are using something like 2d cnn then you should measure CV not only parallely for each slice also how you are predicting a single prediction for all slices",
      "votes": null
    },
    {
      "id": "1520504",
      "postDate": "09/22/2021 11:08:58",
      "content": "<p>Your validation loss is quite high, and it rises during the training. This can be a cause of low score on LB. In fact, I'm not sure if validation AUC is a good metric here.</p>",
      "rawMarkdown": "Your validation loss is quite high, and it rises during the training. This can be a cause of low score on LB. In fact, I'm not sure if validation AUC is a good metric here.",
      "votes": null
    },
    {
      "id": "1520841",
      "postDate": "09/22/2021 16:04:17",
      "content": "<p>Thanks for your reply. <br>\nActually I haven't use CV fold yet and I should consider it. Also I am using valid_AUC now, maybe valid_loss will be more helpful.</p>",
      "rawMarkdown": "Thanks for your reply. \nActually I haven't use CV fold yet and I should consider it. Also I am using valid_AUC now, maybe valid_loss will be more helpful.",
      "votes": null
    },
    {
      "id": "1521029",
      "postDate": "09/22/2021 20:02:31",
      "content": "<p>Yes Araik you are right. I also can't get why is that happening. Some time before I had a same problem because of different augmentation on validation and train dataset, but now I am sure this not occurring here. why do you think is it happening? </p>",
      "rawMarkdown": "Yes Araik you are right. I also can't get why is that happening. Some time before I had a same problem because of different augmentation on validation and train dataset, but now I am sure this not occurring here. why do you think is it happening?",
      "votes": null
    },
    {
      "id": "1521094",
      "postDate": "09/22/2021 21:55:27",
      "content": "<p>This is one of the challenges and mysteries of this competition. We have very few rows of data to train on - fewer than 600 records. And when we have few rows of data, weird stuff happens. </p>\n<p>For example, you uniquely need to make sure you are not overfitting to TWO datasets: your training dataset AND your validation set. Every tweak you make to your model and retrain, you learn a little bit more information on how to fit better to the validation set…do it too much, and you can risk overfitting. If you try changing the seed, do you still get the amazing AUC?</p>\n<p>Another thing that most have noticed in this competition is that the AUC may improve, but the Log Loss gets worse. Log Loss is a smooth metric - it's measuring how good are the predictions you make , like how far away they are from the targets 0 and 1. AUC is not smooth - it is just measuring the ranking of your predictions. So your predictions could all suck, but if they're ranked correctly, you'll get a great AUC.</p>\n<p>Add to these problems that the public LB is fewer than 100 examples and we have a really tough competition on our hands. Since the public LB is small, we don't know if a bad public LB is because of the small dataset size, or if it's because our model doesn't generalize well. This is a problem that everyone is battling with right now.</p>\n<p>Most people recommend to trust your cross-validation over the public LB… but when you have such few amount of training rows, is even the CV trustworthy? We will find out on the private LB.</p>",
      "rawMarkdown": "This is one of the challenges and mysteries of this competition. We have very few rows of data to train on - fewer than 600 records. And when we have few rows of data, weird stuff happens. \n\nFor example, you uniquely need to make sure you are not overfitting to TWO datasets: your training dataset AND your validation set. Every tweak you make to your model and retrain, you learn a little bit more information on how to fit better to the validation set...do it too much, and you can risk overfitting. If you try changing the seed, do you still get the amazing AUC?\n\nAnother thing that most have noticed in this competition is that the AUC may improve, but the Log Loss gets worse. Log Loss is a smooth metric - it's measuring how good are the predictions you make , like how far away they are from the targets 0 and 1. AUC is not smooth - it is just measuring the ranking of your predictions. So your predictions could all suck, but if they're ranked correctly, you'll get a great AUC.\n\nAdd to these problems that the public LB is fewer than 100 examples and we have a really tough competition on our hands. Since the public LB is small, we don't know if a bad public LB is because of the small dataset size, or if it's because our model doesn't generalize well. This is a problem that everyone is battling with right now.\n\nMost people recommend to trust your cross-validation over the public LB... but when you have such few amount of training rows, is even the CV trustworthy? We will find out on the private LB.",
      "votes": null
    },
    {
      "id": "1521362",
      "postDate": "09/23/2021 06:52:43",
      "content": "<p>Your comment clearly shows the problem we are facing in this competition.<br>\nIt is possible that my model is already over-fitting to a particular Fold and MRI type. The first Fold's FLAIR CV improves, but the others tend to stay the same or get worse. And when we change the seed, the scenery changes completely.<br>\nAnd I'm interested in what we can learn from this competition in the end. I'm new to working with voxel data and that alone is valuable to me, but if there is a way to do robust and consistent experiments with less data, I would like to learn from the solutions after the competition is over.</p>",
      "rawMarkdown": "Your comment clearly shows the problem we are facing in this competition.\nIt is possible that my model is already over-fitting to a particular Fold and MRI type. The first Fold's FLAIR CV improves, but the others tend to stay the same or get worse. And when we change the seed, the scenery changes completely.\nAnd I'm interested in what we can learn from this competition in the end. I'm new to working with voxel data and that alone is valuable to me, but if there is a way to do robust and consistent experiments with less data, I would like to learn from the solutions after the competition is over.",
      "votes": null
    },
    {
      "id": "1523022",
      "postDate": "09/24/2021 21:07:12",
      "content": "<p>I think the model simply doesn't learn anything. It's the same case for me.</p>",
      "rawMarkdown": "I think the model simply doesn't learn anything. It's the same case for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1519889,
      "author_name": "mpsampat",
      "author_url": "",
      "post_date": "09/22/2021 02:48:32",
      "content": "<p><a href=\"https://www.kaggle.com/niloofarrahmani\" target=\"_blank\">@niloofarrahmani</a> , Some notes below. </p>\n<ol>\n<li>I had something similar when training with images of size 64x64x64. </li>\n<li>My validation AUC is not as good as yours (~85% is amazing ); </li>\n<li>mine was ~72% and then the public test submission was less than 0.6. </li>\n<li>I am not sure why this is happening. I have only done 80/20 split and still need to do 3/5-fold CV. </li>\n<li>are you getting 85% with 3/5-fold CV ? If yes that is quite impressive. </li>\n<li>My thinking is that you should get good valid AUC and good submission score.</li>\n<li>So if you have a model which is 67% in both cases maybe that is a more reliable model ? Not sure, curious to hear what others say. </li>\n<li>Also when you pick a model do you use valid_AUC or valid_loss ? maybe valid_loss is a better criteria ? </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1520841,
          "author_name": "niloofarrahmani",
          "author_url": "",
          "post_date": "09/22/2021 16:04:17",
          "content": "<p>Thanks for your reply. <br>\nActually I haven't use CV fold yet and I should consider it. Also I am using valid_AUC now, maybe valid_loss will be more helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1520016,
      "author_name": "swaralipibose",
      "author_url": "",
      "post_date": "09/22/2021 05:09:04",
      "content": "<p>hi ,<br>\nThe error is being caused because the public lb is only being measured on 87 samples not so good to be honest , so when you run inference on 87 samples which is very low , the metrics are also not that equivalent to what you measured over more samples , it can be more it can be less . But dont worry trust your cv or the score you are seeing , the private lb has 400 samples , in the private lb we will see a huge shake . I have a question to ask why is your validation loss same always ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1520025,
          "author_name": "mpsampat",
          "author_url": "",
          "post_date": "09/22/2021 05:16:45",
          "content": "<p>Hmm. Good point. But still if you have a good model that generalizes well; why would it only be bad on these 87 cases ? <br>\nPerhaps it is possible and I am missing something ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1520033,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "09/22/2021 05:28:29",
          "content": "<p>The testing set may have some differences from valid . I would recommend you to split the data into a custom test set , sometimes people overfit to valid set . Thus three sets train valid and test</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1520037,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "09/22/2021 05:31:22",
          "content": "<p>also if you are using something like 2d cnn then you should measure CV not only parallely for each slice also how you are predicting a single prediction for all slices</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1520504,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "09/22/2021 11:08:58",
      "content": "<p>Your validation loss is quite high, and it rises during the training. This can be a cause of low score on LB. In fact, I'm not sure if validation AUC is a good metric here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1521029,
          "author_name": "niloofarrahmani",
          "author_url": "",
          "post_date": "09/22/2021 20:02:31",
          "content": "<p>Yes Araik you are right. I also can't get why is that happening. Some time before I had a same problem because of different augmentation on validation and train dataset, but now I am sure this not occurring here. why do you think is it happening? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1523022,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "09/24/2021 21:07:12",
          "content": "<p>I think the model simply doesn't learn anything. It's the same case for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1521094,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "09/22/2021 21:55:27",
      "content": "<p>This is one of the challenges and mysteries of this competition. We have very few rows of data to train on - fewer than 600 records. And when we have few rows of data, weird stuff happens. </p>\n<p>For example, you uniquely need to make sure you are not overfitting to TWO datasets: your training dataset AND your validation set. Every tweak you make to your model and retrain, you learn a little bit more information on how to fit better to the validation set…do it too much, and you can risk overfitting. If you try changing the seed, do you still get the amazing AUC?</p>\n<p>Another thing that most have noticed in this competition is that the AUC may improve, but the Log Loss gets worse. Log Loss is a smooth metric - it's measuring how good are the predictions you make , like how far away they are from the targets 0 and 1. AUC is not smooth - it is just measuring the ranking of your predictions. So your predictions could all suck, but if they're ranked correctly, you'll get a great AUC.</p>\n<p>Add to these problems that the public LB is fewer than 100 examples and we have a really tough competition on our hands. Since the public LB is small, we don't know if a bad public LB is because of the small dataset size, or if it's because our model doesn't generalize well. This is a problem that everyone is battling with right now.</p>\n<p>Most people recommend to trust your cross-validation over the public LB… but when you have such few amount of training rows, is even the CV trustworthy? We will find out on the private LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1521362,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "09/23/2021 06:52:43",
          "content": "<p>Your comment clearly shows the problem we are facing in this competition.<br>\nIt is possible that my model is already over-fitting to a particular Fold and MRI type. The first Fold's FLAIR CV improves, but the others tend to stay the same or get worse. And when we change the seed, the scenery changes completely.<br>\nAnd I'm interested in what we can learn from this competition in the end. I'm new to working with voxel data and that alone is valuable to me, but if there is a way to do robust and consistent experiments with less data, I would like to learn from the solutions after the competition is over.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1519807": "Hello every one. After challenging for a long time with the dataset and trying different models, I have finally found a model that  is being trained well (based on Training & Validation accuracy and AUC curves). But when I submit the result to the competition, the score is not comparable. I am confused that why is it happening? \n\n![Here](https://i.ibb.co/xSFkB6W/image-2021-09-20-T19-11-22-574-Z.png) is a figure that shows the Loss, accuracy and AUC curves of my model. You can see that the maximum validation and training AUC is arrived to ~85% but when submitting, my score is 57%! \n\nAny opinion would be appreciated!",
    "1519889": "niloofarrahmani , Some notes below. \n1. I had something similar when training with images of size 64x64x64. \n2. My validation AUC is not as good as yours (~85% is amazing ); \n3. mine was ~72% and then the public test submission was less than 0.6. \n4. I am not sure why this is happening. I have only done 80/20 split and still need to do 3/5-fold CV. \n5. are you getting 85% with 3/5-fold CV ? If yes that is quite impressive. \n6. My thinking is that you should get good valid AUC and good submission score.\n7.  So if you have a model which is 67% in both cases maybe that is a more reliable model ? Not sure, curious to hear what others say. \n8. Also when you pick a model do you use valid_AUC or valid_loss ? maybe valid_loss is a better criteria ?",
    "1520016": "hi ,\nThe error is being caused because the public lb is only being measured on 87 samples not so good to be honest , so when you run inference on 87 samples which is very low , the metrics are also not that equivalent to what you measured over more samples , it can be more it can be less . But dont worry trust your cv or the score you are seeing , the private lb has 400 samples , in the private lb we will see a huge shake . I have a question to ask why is your validation loss same always ?",
    "1520025": "Hmm. Good point. But still if you have a good model that generalizes well; why would it only be bad on these 87 cases ? \nPerhaps it is possible and I am missing something ?",
    "1520033": "The testing set may have some differences from valid . I would recommend you to split the data into a custom test set , sometimes people overfit to valid set . Thus three sets train valid and test",
    "1520037": "also if you are using something like 2d cnn then you should measure CV not only parallely for each slice also how you are predicting a single prediction for all slices",
    "1520504": "Your validation loss is quite high, and it rises during the training. This can be a cause of low score on LB. In fact, I'm not sure if validation AUC is a good metric here.",
    "1520841": "Thanks for your reply. \nActually I haven't use CV fold yet and I should consider it. Also I am using valid_AUC now, maybe valid_loss will be more helpful.",
    "1521029": "Yes Araik you are right. I also can't get why is that happening. Some time before I had a same problem because of different augmentation on validation and train dataset, but now I am sure this not occurring here. why do you think is it happening?",
    "1521094": "This is one of the challenges and mysteries of this competition. We have very few rows of data to train on - fewer than 600 records. And when we have few rows of data, weird stuff happens. \n\nFor example, you uniquely need to make sure you are not overfitting to TWO datasets: your training dataset AND your validation set. Every tweak you make to your model and retrain, you learn a little bit more information on how to fit better to the validation set...do it too much, and you can risk overfitting. If you try changing the seed, do you still get the amazing AUC?\n\nAnother thing that most have noticed in this competition is that the AUC may improve, but the Log Loss gets worse. Log Loss is a smooth metric - it's measuring how good are the predictions you make , like how far away they are from the targets 0 and 1. AUC is not smooth - it is just measuring the ranking of your predictions. So your predictions could all suck, but if they're ranked correctly, you'll get a great AUC.\n\nAdd to these problems that the public LB is fewer than 100 examples and we have a really tough competition on our hands. Since the public LB is small, we don't know if a bad public LB is because of the small dataset size, or if it's because our model doesn't generalize well. This is a problem that everyone is battling with right now.\n\nMost people recommend to trust your cross-validation over the public LB... but when you have such few amount of training rows, is even the CV trustworthy? We will find out on the private LB.",
    "1521362": "Your comment clearly shows the problem we are facing in this competition.\nIt is possible that my model is already over-fitting to a particular Fold and MRI type. The first Fold's FLAIR CV improves, but the others tend to stay the same or get worse. And when we change the seed, the scenery changes completely.\nAnd I'm interested in what we can learn from this competition in the end. I'm new to working with voxel data and that alone is valuable to me, but if there is a way to do robust and consistent experiments with less data, I would like to learn from the solutions after the competition is over.",
    "1523022": "I think the model simply doesn't learn anything. It's the same case for me."
  },
  "source": "meta"
}