{
  "id": 206403,
  "title": "Difference between val accuracy and submission score",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/206403",
  "author_name": "",
  "post_date": "2020-12-24T12:36:03.633506200Z",
  "votes": 1,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Why this big difference happening between Val accuracy and in the submission score ?</p>",
  "messages": [
    {
      "id": "1125157",
      "postDate": "12/24/2020 12:36:03",
      "content": "<p>Why this big difference happening between Val accuracy and in the submission score ?</p>",
      "rawMarkdown": "Why this big difference happening between Val accuracy and in the submission score ?",
      "votes": null
    },
    {
      "id": "1125280",
      "postDate": "12/24/2020 15:22:38",
      "content": "<p>This sounds like you are overfitting your validation set or you might have a data leakage in the cross-validation folds. Check if you had done the cross-validation in a right manner ..</p>",
      "rawMarkdown": "This sounds like you are overfitting your validation set or you might have a data leakage in the cross-validation folds. Check if you had done the cross-validation in a right manner ..",
      "votes": null
    },
    {
      "id": "1125957",
      "postDate": "12/25/2020 07:46:18",
      "content": "<p>I think they have test set from different distribution.</p>",
      "rawMarkdown": "I think they have test set from different distribution.",
      "votes": null
    },
    {
      "id": "1127077",
      "postDate": "12/26/2020 08:18:14",
      "content": "<p>Possibility of overfitting <a href=\"https://www.kaggle.com/jeelgondaliya\" target=\"_blank\">@jeelgondaliya</a> !!!</p>",
      "rawMarkdown": "Possibility of overfitting @jeelgondaliya !!!",
      "votes": null
    },
    {
      "id": "1127949",
      "postDate": "12/27/2020 04:13:47",
      "content": "<p>Since you did not share the values for your val accuracy and the LB score you have posed a pretty generic question :)  Your \"big\" might be my \"small\".   </p>\n<p>To get a better answer - make your question a bit more specific.</p>\n<p>My val_accuracy have been in the 0.87 to 0.89 range and my LB score pretty close to that same range.  Always when I get a val_accuracy higher than 0.90 than I have made a mistake.  As already mentioned data leakage is the root of most of my \"too good to be true\" val_accuracy.  Next biggest sources have been mistakes in my TTA code.</p>\n<p>I do not think that the test set distribution being different is the root cause of the small differences I have seen.</p>",
      "rawMarkdown": "Since you did not share the values for your val accuracy and the LB score you have posed a pretty generic question :)  Your \"big\" might be my \"small\".   \n\nTo get a better answer - make your question a bit more specific.\n\nMy val_accuracy have been in the 0.87 to 0.89 range and my LB score pretty close to that same range.  Always when I get a val_accuracy higher than 0.90 than I have made a mistake.  As already mentioned data leakage is the root of most of my \"too good to be true\" val_accuracy.  Next biggest sources have been mistakes in my TTA code.\n\nI do not think that the test set distribution being different is the root cause of the small differences I have seen.",
      "votes": null
    },
    {
      "id": "1128106",
      "postDate": "12/27/2020 07:14:01",
      "content": "<p>My CV score is 0.922 while LB is 0.839. What can be the possible issues?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2648341%2F677981eb78f10a9cf7af5d078a68cbee%2FScreenshot_2020-12-27%20try1%20-%20Jupyter%20Notebook.png?generation=1609053806420347&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "My CV score is 0.922 while LB is 0.839. What can be the possible issues?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2648341%2F677981eb78f10a9cf7af5d078a68cbee%2FScreenshot_2020-12-27%20try1%20-%20Jupyter%20Notebook.png?generation=1609053806420347&alt=media)",
      "votes": null
    },
    {
      "id": "1128108",
      "postDate": "12/27/2020 07:16:12",
      "content": "<p>If the model was over fitting on train the results on validation should also be low. Why are the results on Val high while on test they are low.</p>",
      "rawMarkdown": "If the model was over fitting on train the results on validation should also be low. Why are the results on Val high while on test they are low.",
      "votes": null
    },
    {
      "id": "1128111",
      "postDate": "12/27/2020 07:17:27",
      "content": "<p>How can one over fit his validation set? <a href=\"https://www.kaggle.com/atharvaingle\" target=\"_blank\">@atharvaingle</a> </p>",
      "rawMarkdown": "How can one over fit his validation set? @atharvaingle",
      "votes": null
    },
    {
      "id": "1129209",
      "postDate": "12/28/2020 06:36:57",
      "content": "<p>Hmmm.  Doing this on iPad and the result confusing.  Will look again soon when Ihave computer running.  </p>",
      "rawMarkdown": "Hmmm.  Doing this on iPad and the result confusing.  Will look again soon when Ihave computer running.",
      "votes": null
    },
    {
      "id": "1130250",
      "postDate": "12/28/2020 21:50:02",
      "content": "<p>Your accuracy numbers are a bit confusing.  Couple of things I see that might be the issue.</p>\n<ol>\n<li>Your validation set looks like it's 5%.  </li>\n<li>You only ran for 2 epochs.</li>\n</ol>\n<p>With only 2 epochs the LB score makes sense.  Running more epochs with a higher percentage (more images) for validation should bring the numbers back to a \"makes sense\" level.</p>",
      "rawMarkdown": "Your accuracy numbers are a bit confusing.  Couple of things I see that might be the issue.\n1.  Your validation set looks like it's 5%.  \n2.  You only ran for 2 epochs.\n\nWith only 2 epochs the LB score makes sense.  Running more epochs with a higher percentage (more images) for validation should bring the numbers back to a \"makes sense\" level.",
      "votes": null
    },
    {
      "id": "1132273",
      "postDate": "12/30/2020 09:27:14",
      "content": "<p>1) I am making 5 folds. Val is not small, train batch size is 4 while val batch size is 16.</p>\n<p>I used this same pipeline and here were the results:<br>\nResnet50           CV:0.89  LB:0.872 (Train batch size=8, Val batch size=32)<br>\nEfficientNetb3:  CV:0.922  LB:0.839 (Train batch size=4, Val batch size=16)</p>\n<p>Do you think there is some code error? What are your CV results and the batch size you are using? Should my batch size be consistent for all experiments to compare results?</p>",
      "rawMarkdown": "1) I am making 5 folds. Val is not small, train batch size is 4 while val batch size is 16.\n\nI used this same pipeline and here were the results:\nResnet50           CV:0.89  LB:0.872 (Train batch size=8, Val batch size=32)\nEfficientNetb3:  CV:0.922  LB:0.839 (Train batch size=4, Val batch size=16)\n\nDo you think there is some code error? What are your CV results and the batch size you are using? Should my batch size be consistent for all experiments to compare results?",
      "votes": null
    },
    {
      "id": "1132285",
      "postDate": "12/30/2020 09:33:14",
      "content": "<p>Discussions like <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017\" target=\"_blank\">these</a>, suggest that noise also has a role in the CV LB Gap. Since the test dataset is also noisy, it is suggested here not to choose the lowest validation loss.</p>",
      "rawMarkdown": "Discussions like [these](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017), suggest that noise also has a role in the CV LB Gap. Since the test dataset is also noisy, it is suggested here not to choose the lowest validation loss.",
      "votes": null
    },
    {
      "id": "1132850",
      "postDate": "12/30/2020 18:08:18",
      "content": "<p>Batch size can make a difference.  For some competitions my LB score has been best with a small batch size and for other's it has been best with a large batch; and there I been some were it made no change.  I have only made a limited number of submissions in this competition (trusting my cv on 4 local machines right now) so not sure which way batch size will take things but seem to recall a discussion post or two were smaller batch was better.  IMO - believe that the more noise in the data that smaller batch is better.   I normally explore the batch size response only after I got the model I want - until than it's as large as memory will let me take it for speed.</p>\n<p>I did do industrial experimentation for the last thirty years of my working career and do know your at huge risk of bad decisions when you leave important features not being controlled.  </p>\n<p>The Resnet50 numbers you have are not unreasonable - so I would repeat the Efficient with larger batch and both train and validation at same level.</p>",
      "rawMarkdown": "Batch size can make a difference.  For some competitions my LB score has been best with a small batch size and for other's it has been best with a large batch; and there I been some were it made no change.  I have only made a limited number of submissions in this competition (trusting my cv on 4 local machines right now) so not sure which way batch size will take things but seem to recall a discussion post or two were smaller batch was better.  IMO - believe that the more noise in the data that smaller batch is better.   I normally explore the batch size response only after I got the model I want - until than it's as large as memory will let me take it for speed.\n\nI did do industrial experimentation for the last thirty years of my working career and do know your at huge risk of bad decisions when you leave important features not being controlled.  \n\nThe Resnet50 numbers you have are not unreasonable - so I would repeat the Efficient with larger batch and both train and validation at same level.",
      "votes": null
    },
    {
      "id": "1175732",
      "postDate": "01/29/2021 10:01:23",
      "content": "<blockquote>\n  <p>IMO - believe that the more noise in the data that smaller batch is better.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> </p>\n<p>you mean that small batch_size is better for noisy data ?</p>",
      "rawMarkdown": "> IMO - believe that the more noise in the data that smaller batch is better.\n\n@pcjimmmy \n\nyou mean that small batch_size is better for noisy data ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1125280,
      "author_name": "atharvaingle",
      "author_url": "",
      "post_date": "12/24/2020 15:22:38",
      "content": "<p>This sounds like you are overfitting your validation set or you might have a data leakage in the cross-validation folds. Check if you had done the cross-validation in a right manner ..</p>",
      "votes": null,
      "replies": [
        {
          "id": 1125957,
          "author_name": "jeelgondaliya",
          "author_url": "",
          "post_date": "12/25/2020 07:46:18",
          "content": "<p>I think they have test set from different distribution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128111,
          "author_name": "deepdreamx",
          "author_url": "",
          "post_date": "12/27/2020 07:17:27",
          "content": "<p>How can one over fit his validation set? <a href=\"https://www.kaggle.com/atharvaingle\" target=\"_blank\">@atharvaingle</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127077,
      "author_name": "saurabhshahane",
      "author_url": "",
      "post_date": "12/26/2020 08:18:14",
      "content": "<p>Possibility of overfitting <a href=\"https://www.kaggle.com/jeelgondaliya\" target=\"_blank\">@jeelgondaliya</a> !!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1128108,
          "author_name": "deepdreamx",
          "author_url": "",
          "post_date": "12/27/2020 07:16:12",
          "content": "<p>If the model was over fitting on train the results on validation should also be low. Why are the results on Val high while on test they are low.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127949,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "12/27/2020 04:13:47",
      "content": "<p>Since you did not share the values for your val accuracy and the LB score you have posed a pretty generic question :)  Your \"big\" might be my \"small\".   </p>\n<p>To get a better answer - make your question a bit more specific.</p>\n<p>My val_accuracy have been in the 0.87 to 0.89 range and my LB score pretty close to that same range.  Always when I get a val_accuracy higher than 0.90 than I have made a mistake.  As already mentioned data leakage is the root of most of my \"too good to be true\" val_accuracy.  Next biggest sources have been mistakes in my TTA code.</p>\n<p>I do not think that the test set distribution being different is the root cause of the small differences I have seen.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1128106,
          "author_name": "deepdreamx",
          "author_url": "",
          "post_date": "12/27/2020 07:14:01",
          "content": "<p>My CV score is 0.922 while LB is 0.839. What can be the possible issues?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2648341%2F677981eb78f10a9cf7af5d078a68cbee%2FScreenshot_2020-12-27%20try1%20-%20Jupyter%20Notebook.png?generation=1609053806420347&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1129209,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/28/2020 06:36:57",
          "content": "<p>Hmmm.  Doing this on iPad and the result confusing.  Will look again soon when Ihave computer running.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130250,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/28/2020 21:50:02",
          "content": "<p>Your accuracy numbers are a bit confusing.  Couple of things I see that might be the issue.</p>\n<ol>\n<li>Your validation set looks like it's 5%.  </li>\n<li>You only ran for 2 epochs.</li>\n</ol>\n<p>With only 2 epochs the LB score makes sense.  Running more epochs with a higher percentage (more images) for validation should bring the numbers back to a \"makes sense\" level.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132273,
          "author_name": "deepdreamx",
          "author_url": "",
          "post_date": "12/30/2020 09:27:14",
          "content": "<p>1) I am making 5 folds. Val is not small, train batch size is 4 while val batch size is 16.</p>\n<p>I used this same pipeline and here were the results:<br>\nResnet50           CV:0.89  LB:0.872 (Train batch size=8, Val batch size=32)<br>\nEfficientNetb3:  CV:0.922  LB:0.839 (Train batch size=4, Val batch size=16)</p>\n<p>Do you think there is some code error? What are your CV results and the batch size you are using? Should my batch size be consistent for all experiments to compare results?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132850,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/30/2020 18:08:18",
          "content": "<p>Batch size can make a difference.  For some competitions my LB score has been best with a small batch size and for other's it has been best with a large batch; and there I been some were it made no change.  I have only made a limited number of submissions in this competition (trusting my cv on 4 local machines right now) so not sure which way batch size will take things but seem to recall a discussion post or two were smaller batch was better.  IMO - believe that the more noise in the data that smaller batch is better.   I normally explore the batch size response only after I got the model I want - until than it's as large as memory will let me take it for speed.</p>\n<p>I did do industrial experimentation for the last thirty years of my working career and do know your at huge risk of bad decisions when you leave important features not being controlled.  </p>\n<p>The Resnet50 numbers you have are not unreasonable - so I would repeat the Efficient with larger batch and both train and validation at same level.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1175732,
          "author_name": "joshi98kishan",
          "author_url": "",
          "post_date": "01/29/2021 10:01:23",
          "content": "<blockquote>\n  <p>IMO - believe that the more noise in the data that smaller batch is better.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> </p>\n<p>you mean that small batch_size is better for noisy data ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1132285,
      "author_name": "yerramvarun",
      "author_url": "",
      "post_date": "12/30/2020 09:33:14",
      "content": "<p>Discussions like <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017\" target=\"_blank\">these</a>, suggest that noise also has a role in the CV LB Gap. Since the test dataset is also noisy, it is suggested here not to choose the lowest validation loss.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1125157": "Why this big difference happening between Val accuracy and in the submission score ?",
    "1125280": "This sounds like you are overfitting your validation set or you might have a data leakage in the cross-validation folds. Check if you had done the cross-validation in a right manner ..",
    "1125957": "I think they have test set from different distribution.",
    "1127077": "Possibility of overfitting @jeelgondaliya !!!",
    "1127949": "Since you did not share the values for your val accuracy and the LB score you have posed a pretty generic question :)  Your \"big\" might be my \"small\".   \n\nTo get a better answer - make your question a bit more specific.\n\nMy val_accuracy have been in the 0.87 to 0.89 range and my LB score pretty close to that same range.  Always when I get a val_accuracy higher than 0.90 than I have made a mistake.  As already mentioned data leakage is the root of most of my \"too good to be true\" val_accuracy.  Next biggest sources have been mistakes in my TTA code.\n\nI do not think that the test set distribution being different is the root cause of the small differences I have seen.",
    "1128106": "My CV score is 0.922 while LB is 0.839. What can be the possible issues?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2648341%2F677981eb78f10a9cf7af5d078a68cbee%2FScreenshot_2020-12-27%20try1%20-%20Jupyter%20Notebook.png?generation=1609053806420347&alt=media)",
    "1128108": "If the model was over fitting on train the results on validation should also be low. Why are the results on Val high while on test they are low.",
    "1128111": "How can one over fit his validation set? @atharvaingle",
    "1129209": "Hmmm.  Doing this on iPad and the result confusing.  Will look again soon when Ihave computer running.",
    "1130250": "Your accuracy numbers are a bit confusing.  Couple of things I see that might be the issue.\n1.  Your validation set looks like it's 5%.  \n2.  You only ran for 2 epochs.\n\nWith only 2 epochs the LB score makes sense.  Running more epochs with a higher percentage (more images) for validation should bring the numbers back to a \"makes sense\" level.",
    "1132273": "1) I am making 5 folds. Val is not small, train batch size is 4 while val batch size is 16.\n\nI used this same pipeline and here were the results:\nResnet50           CV:0.89  LB:0.872 (Train batch size=8, Val batch size=32)\nEfficientNetb3:  CV:0.922  LB:0.839 (Train batch size=4, Val batch size=16)\n\nDo you think there is some code error? What are your CV results and the batch size you are using? Should my batch size be consistent for all experiments to compare results?",
    "1132285": "Discussions like [these](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202017), suggest that noise also has a role in the CV LB Gap. Since the test dataset is also noisy, it is suggested here not to choose the lowest validation loss.",
    "1132850": "Batch size can make a difference.  For some competitions my LB score has been best with a small batch size and for other's it has been best with a large batch; and there I been some were it made no change.  I have only made a limited number of submissions in this competition (trusting my cv on 4 local machines right now) so not sure which way batch size will take things but seem to recall a discussion post or two were smaller batch was better.  IMO - believe that the more noise in the data that smaller batch is better.   I normally explore the batch size response only after I got the model I want - until than it's as large as memory will let me take it for speed.\n\nI did do industrial experimentation for the last thirty years of my working career and do know your at huge risk of bad decisions when you leave important features not being controlled.  \n\nThe Resnet50 numbers you have are not unreasonable - so I would repeat the Efficient with larger batch and both train and validation at same level.",
    "1175732": "> IMO - believe that the more noise in the data that smaller batch is better.\n\n@pcjimmmy \n\nyou mean that small batch_size is better for noisy data ?"
  },
  "source": "meta"
}