{
  "id": 215566,
  "title": "There are two questions for experts to answer ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215566",
  "author_name": "",
  "post_date": "2021-01-30T12:23:31.745142700Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p><strong>1 The accuracy of the test set is always higher than that of the training set？</strong><br>\nI use resnet50 pretraining network. I observe that there is no dropout in the network structure. I use k-fold, k = 5. During training, the accuracy of the test set is always higher than that of the verifier. What is the reason? The difficulty is that the k-fold partition data set is not uniform? Can the StratifiedKfold function improve it?</p>\n<p><strong>2 How accurate is it to submit results？</strong><br>\nWith k-fold, after I trained the model for 9 hours, the accuracy of the training set and the test set was 90.5% and 91.5%, respectively. May I ask the experts, how much accuracy can I submit and the final result is more than 90%?</p>",
  "messages": [
    {
      "id": "1177642",
      "postDate": "01/30/2021 12:23:31",
      "content": "<p><strong>1 The accuracy of the test set is always higher than that of the training set？</strong><br>\nI use resnet50 pretraining network. I observe that there is no dropout in the network structure. I use k-fold, k = 5. During training, the accuracy of the test set is always higher than that of the verifier. What is the reason? The difficulty is that the k-fold partition data set is not uniform? Can the StratifiedKfold function improve it?</p>\n<p><strong>2 How accurate is it to submit results？</strong><br>\nWith k-fold, after I trained the model for 9 hours, the accuracy of the training set and the test set was 90.5% and 91.5%, respectively. May I ask the experts, how much accuracy can I submit and the final result is more than 90%?</p>",
      "rawMarkdown": "**1 The accuracy of the test set is always higher than that of the training set？**\nI use resnet50 pretraining network. I observe that there is no dropout in the network structure. I use k-fold, k = 5. During training, the accuracy of the test set is always higher than that of the verifier. What is the reason? The difficulty is that the k-fold partition data set is not uniform? Can the StratifiedKfold function improve it?\n\n**2 How accurate is it to submit results？**\nWith k-fold, after I trained the model for 9 hours, the accuracy of the training set and the test set was 90.5% and 91.5%, respectively. May I ask the experts, how much accuracy can I submit and the final result is more than 90%?",
      "votes": null
    },
    {
      "id": "1177867",
      "postDate": "01/30/2021 14:59:04",
      "content": "<p>help me thank</p>",
      "rawMarkdown": "help me thank",
      "votes": null
    },
    {
      "id": "1178106",
      "postDate": "01/30/2021 16:45:58",
      "content": "<p>If I understand your 1st question you are seeing the validation accuracy higher than the training accuracy ???</p>\n<p>For this competition EVERY model I have run the training accuracy is worse than validation after a sufficient number of epochs until early stopping kicks in (with only a few epochs the validation is sometimes higher than train accuracy).  So first option - run more epochs.   I do use StratifiedKfold all the time but not sure that's the root cause for validation accuracy to be lower than training after the full number of epochs.  Since your using the full 9 hours you probably need to code and use TPU's or go offline from Kaggle to a compute with no 9 hour limit.</p>\n<p>You could also break the training into two or more kernels if you don't want TPU or don't have an off-line compute - for TensorFlow it's very easy to resume more training after loading the previous model and resuming training.</p>\n<p>For your second question - 5 free submissions per day so submit often and figure this question out yourself.</p>",
      "rawMarkdown": "If I understand your 1st question you are seeing the validation accuracy higher than the training accuracy ???\n\nFor this competition EVERY model I have run the training accuracy is worse than validation after a sufficient number of epochs until early stopping kicks in (with only a few epochs the validation is sometimes higher than train accuracy).  So first option - run more epochs.   I do use StratifiedKfold all the time but not sure that's the root cause for validation accuracy to be lower than training after the full number of epochs.  Since your using the full 9 hours you probably need to code and use TPU's or go offline from Kaggle to a compute with no 9 hour limit.\n\nYou could also break the training into two or more kernels if you don't want TPU or don't have an off-line compute - for TensorFlow it's very easy to resume more training after loading the previous model and resuming training.\n\nFor your second question - 5 free submissions per day so submit often and figure this question out yourself.",
      "votes": null
    },
    {
      "id": "1178577",
      "postDate": "01/31/2021 01:45:18",
      "content": "<p>Thank you very much for your reply. I will try to solve the first problem. Thanks again</p>",
      "rawMarkdown": "Thank you very much for your reply. I will try to solve the first problem. Thanks again",
      "votes": null
    },
    {
      "id": "1178897",
      "postDate": "01/31/2021 06:51:44",
      "content": "<p>WOW! IT IS REALLY INTRESTING<br>\nUPVOTED!<br>\n<a href=\"https://www.kaggle.com/ashokkumarbibbab/covid-19-vaccination-process\" target=\"_blank\">https://www.kaggle.com/ashokkumarbibbab/covid-19-vaccination-process</a><br>\nSEE THIS NOTEBOOK🙄</p>",
      "rawMarkdown": "WOW! IT IS REALLY INTRESTING\nUPVOTED!\nhttps://www.kaggle.com/ashokkumarbibbab/covid-19-vaccination-process\nSEE THIS NOTEBOOK🙄",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1177867,
      "author_name": "henini",
      "author_url": "",
      "post_date": "01/30/2021 14:59:04",
      "content": "<p>help me thank</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1178106,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/30/2021 16:45:58",
      "content": "<p>If I understand your 1st question you are seeing the validation accuracy higher than the training accuracy ???</p>\n<p>For this competition EVERY model I have run the training accuracy is worse than validation after a sufficient number of epochs until early stopping kicks in (with only a few epochs the validation is sometimes higher than train accuracy).  So first option - run more epochs.   I do use StratifiedKfold all the time but not sure that's the root cause for validation accuracy to be lower than training after the full number of epochs.  Since your using the full 9 hours you probably need to code and use TPU's or go offline from Kaggle to a compute with no 9 hour limit.</p>\n<p>You could also break the training into two or more kernels if you don't want TPU or don't have an off-line compute - for TensorFlow it's very easy to resume more training after loading the previous model and resuming training.</p>\n<p>For your second question - 5 free submissions per day so submit often and figure this question out yourself.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1178577,
          "author_name": "henini",
          "author_url": "",
          "post_date": "01/31/2021 01:45:18",
          "content": "<p>Thank you very much for your reply. I will try to solve the first problem. Thanks again</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1178897,
      "author_name": "ashokkumarbibbab",
      "author_url": "",
      "post_date": "01/31/2021 06:51:44",
      "content": "<p>WOW! IT IS REALLY INTRESTING<br>\nUPVOTED!<br>\n<a href=\"https://www.kaggle.com/ashokkumarbibbab/covid-19-vaccination-process\" target=\"_blank\">https://www.kaggle.com/ashokkumarbibbab/covid-19-vaccination-process</a><br>\nSEE THIS NOTEBOOK🙄</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1177642": "**1 The accuracy of the test set is always higher than that of the training set？**\nI use resnet50 pretraining network. I observe that there is no dropout in the network structure. I use k-fold, k = 5. During training, the accuracy of the test set is always higher than that of the verifier. What is the reason? The difficulty is that the k-fold partition data set is not uniform? Can the StratifiedKfold function improve it?\n\n**2 How accurate is it to submit results？**\nWith k-fold, after I trained the model for 9 hours, the accuracy of the training set and the test set was 90.5% and 91.5%, respectively. May I ask the experts, how much accuracy can I submit and the final result is more than 90%?",
    "1177867": "help me thank",
    "1178106": "If I understand your 1st question you are seeing the validation accuracy higher than the training accuracy ???\n\nFor this competition EVERY model I have run the training accuracy is worse than validation after a sufficient number of epochs until early stopping kicks in (with only a few epochs the validation is sometimes higher than train accuracy).  So first option - run more epochs.   I do use StratifiedKfold all the time but not sure that's the root cause for validation accuracy to be lower than training after the full number of epochs.  Since your using the full 9 hours you probably need to code and use TPU's or go offline from Kaggle to a compute with no 9 hour limit.\n\nYou could also break the training into two or more kernels if you don't want TPU or don't have an off-line compute - for TensorFlow it's very easy to resume more training after loading the previous model and resuming training.\n\nFor your second question - 5 free submissions per day so submit often and figure this question out yourself.",
    "1178577": "Thank you very much for your reply. I will try to solve the first problem. Thanks again",
    "1178897": "WOW! IT IS REALLY INTRESTING\nUPVOTED!\nhttps://www.kaggle.com/ashokkumarbibbab/covid-19-vaccination-process\nSEE THIS NOTEBOOK🙄"
  },
  "source": "meta"
}