{
  "id": 218375,
  "title": "Eval Data Percentage",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/218375",
  "author_name": "",
  "post_date": "2021-02-10T10:57:17.069210Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>How much of the training data do you people usually use for validation? In most of my experiments (no cross validation, just single model) I used 15% which is about 4500 images. </p>\n<p>The advantage was it gave me an extremely accuracte estimate of the leaderboard score.<br>\nThe disadvantage was that I probably lost performance because getting another 1000-2000 images for training would probably improve my score.</p>\n<p>In a competition like this, how much do you tend to take for validation?</p>",
  "messages": [
    {
      "id": "1194759",
      "postDate": "02/10/2021 10:57:17",
      "content": "<p>How much of the training data do you people usually use for validation? In most of my experiments (no cross validation, just single model) I used 15% which is about 4500 images. </p>\n<p>The advantage was it gave me an extremely accuracte estimate of the leaderboard score.<br>\nThe disadvantage was that I probably lost performance because getting another 1000-2000 images for training would probably improve my score.</p>\n<p>In a competition like this, how much do you tend to take for validation?</p>",
      "rawMarkdown": "How much of the training data do you people usually use for validation? In most of my experiments (no cross validation, just single model) I used 15% which is about 4500 images. \n\nThe advantage was it gave me an extremely accuracte estimate of the leaderboard score.\nThe disadvantage was that I probably lost performance because getting another 1000-2000 images for training would probably improve my score.\n\nIn a competition like this, how much do you tend to take for validation?",
      "votes": null
    },
    {
      "id": "1194866",
      "postDate": "02/10/2021 12:18:31",
      "content": "<p>Hello!</p>\n<p>A rule of thumb I heard from my tutors is to hold out 33% of the data if there are lots of. </p>\n<p>As for me, I mostly use cross-validation without a hold-out subset, as CV score is supposed to be more trustworthy. It is also a rough indicator of the ensemble lower PB boundary (?). </p>\n<p>I suggest trying out a trick I've come across in the book by <a href=\"https://www.kaggle.com/fchollet\" target=\"_blank\">@fchollet</a>: train your model with a hold-out set and after you finish tuning it, retrain it on all available data for the same number of epochs.</p>",
      "rawMarkdown": "Hello!\n\nA rule of thumb I heard from my tutors is to hold out 33% of the data if there are lots of. \n\nAs for me, I mostly use cross-validation without a hold-out subset, as CV score is supposed to be more trustworthy. It is also a rough indicator of the ensemble lower PB boundary (?). \n\nI suggest trying out a trick I've come across in the book by @fchollet: train your model with a hold-out set and after you finish tuning it, retrain it on all available data for the same number of epochs.",
      "votes": null
    },
    {
      "id": "1194894",
      "postDate": "02/10/2021 12:52:29",
      "content": "<p>Ahh yes thats a good trick thanks. I should try that out again. Have been scared of it after failing at it once a long time ago :) </p>",
      "rawMarkdown": "Ahh yes thats a good trick thanks. I should try that out again. Have been scared of it after failing at it once a long time ago :)",
      "votes": null
    },
    {
      "id": "1196853",
      "postDate": "02/11/2021 17:50:29",
      "content": "<p>Hi,</p>\n<p>I use 5 fold CV instead of just one validation set to make sure to see all the examples in the dataset. IMHO, this approach also seems more robust in order to evaluate our results. Also if you find a good validation scheme which reflects the LB well, you can be much more confident about your improvements and decrease the potential noisy information. </p>",
      "rawMarkdown": "Hi,\n\nI use 5 fold CV instead of just one validation set to make sure to see all the examples in the dataset. IMHO, this approach also seems more robust in order to evaluate our results. Also if you find a good validation scheme which reflects the LB well, you can be much more confident about your improvements and decrease the potential noisy information.",
      "votes": null
    },
    {
      "id": "1200503",
      "postDate": "02/14/2021 18:14:02",
      "content": "<p>Yeah. You're definitely right. I have just been hesitant to switch from single split to 5 fold CV because it would increases my training/experimentation time by 5x. This increase in training time sounds like it would really hinder my progression. How do you and other people normally deal with this? Or is that just the sacrifice you make?</p>",
      "rawMarkdown": "Yeah. You're definitely right. I have just been hesitant to switch from single split to 5 fold CV because it would increases my training/experimentation time by 5x. This increase in training time sounds like it would really hinder my progression. How do you and other people normally deal with this? Or is that just the sacrifice you make?",
      "votes": null
    },
    {
      "id": "1200596",
      "postDate": "02/14/2021 19:10:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raivokoot\" target=\"_blank\">@raivokoot</a> it is a good idea to sperate the training and the inference. <br>\nAlso, I use colab for inference as the GPU times runs out in Kaggle</p>\n<p>colab is slower but you have unlimited GPU</p>",
      "rawMarkdown": "Hi @raivokoot it is a good idea to sperate the training and the inference. \nAlso, I use colab for inference as the GPU times runs out in Kaggle\n\ncolab is slower but you have unlimited GPU",
      "votes": null
    },
    {
      "id": "1200729",
      "postDate": "02/14/2021 23:29:48",
      "content": "<p>Hi again <a href=\"https://www.kaggle.com/raivokoot\" target=\"_blank\">@raivokoot</a> </p>\n<p>As I said, 5 fold will give you much more robust estimation of your performance. You might end up choosing a \"lucky\" subset for your validation set and your performance might not generalize well to unseen data (Let's remind ourselves that we will also be tested with the private test set). You basically eliminate the chance of choosing that one particular good validation set. In K-Fold CV instead, you use all your data for validation once. </p>\n<p>And yes it increases the training time, but we become much more confident about our model's performance. I'm not an experienced kaggler, but this is the general scheme I've seen people follow.</p>",
      "rawMarkdown": "Hi again @raivokoot \n\nAs I said, 5 fold will give you much more robust estimation of your performance. You might end up choosing a \"lucky\" subset for your validation set and your performance might not generalize well to unseen data (Let's remind ourselves that we will also be tested with the private test set). You basically eliminate the chance of choosing that one particular good validation set. In K-Fold CV instead, you use all your data for validation once. \n\nAnd yes it increases the training time, but we become much more confident about our model's performance. I'm not an experienced kaggler, but this is the general scheme I've seen people follow.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1194866,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "02/10/2021 12:18:31",
      "content": "<p>Hello!</p>\n<p>A rule of thumb I heard from my tutors is to hold out 33% of the data if there are lots of. </p>\n<p>As for me, I mostly use cross-validation without a hold-out subset, as CV score is supposed to be more trustworthy. It is also a rough indicator of the ensemble lower PB boundary (?). </p>\n<p>I suggest trying out a trick I've come across in the book by <a href=\"https://www.kaggle.com/fchollet\" target=\"_blank\">@fchollet</a>: train your model with a hold-out set and after you finish tuning it, retrain it on all available data for the same number of epochs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1194894,
          "author_name": "raivokoot",
          "author_url": "",
          "post_date": "02/10/2021 12:52:29",
          "content": "<p>Ahh yes thats a good trick thanks. I should try that out again. Have been scared of it after failing at it once a long time ago :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1196853,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/11/2021 17:50:29",
      "content": "<p>Hi,</p>\n<p>I use 5 fold CV instead of just one validation set to make sure to see all the examples in the dataset. IMHO, this approach also seems more robust in order to evaluate our results. Also if you find a good validation scheme which reflects the LB well, you can be much more confident about your improvements and decrease the potential noisy information. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1200503,
          "author_name": "raivokoot",
          "author_url": "",
          "post_date": "02/14/2021 18:14:02",
          "content": "<p>Yeah. You're definitely right. I have just been hesitant to switch from single split to 5 fold CV because it would increases my training/experimentation time by 5x. This increase in training time sounds like it would really hinder my progression. How do you and other people normally deal with this? Or is that just the sacrifice you make?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1200596,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/14/2021 19:10:57",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/raivokoot\" target=\"_blank\">@raivokoot</a> it is a good idea to sperate the training and the inference. <br>\nAlso, I use colab for inference as the GPU times runs out in Kaggle</p>\n<p>colab is slower but you have unlimited GPU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1200729,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "02/14/2021 23:29:48",
          "content": "<p>Hi again <a href=\"https://www.kaggle.com/raivokoot\" target=\"_blank\">@raivokoot</a> </p>\n<p>As I said, 5 fold will give you much more robust estimation of your performance. You might end up choosing a \"lucky\" subset for your validation set and your performance might not generalize well to unseen data (Let's remind ourselves that we will also be tested with the private test set). You basically eliminate the chance of choosing that one particular good validation set. In K-Fold CV instead, you use all your data for validation once. </p>\n<p>And yes it increases the training time, but we become much more confident about our model's performance. I'm not an experienced kaggler, but this is the general scheme I've seen people follow.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1194759": "How much of the training data do you people usually use for validation? In most of my experiments (no cross validation, just single model) I used 15% which is about 4500 images. \n\nThe advantage was it gave me an extremely accuracte estimate of the leaderboard score.\nThe disadvantage was that I probably lost performance because getting another 1000-2000 images for training would probably improve my score.\n\nIn a competition like this, how much do you tend to take for validation?",
    "1194866": "Hello!\n\nA rule of thumb I heard from my tutors is to hold out 33% of the data if there are lots of. \n\nAs for me, I mostly use cross-validation without a hold-out subset, as CV score is supposed to be more trustworthy. It is also a rough indicator of the ensemble lower PB boundary (?). \n\nI suggest trying out a trick I've come across in the book by @fchollet: train your model with a hold-out set and after you finish tuning it, retrain it on all available data for the same number of epochs.",
    "1194894": "Ahh yes thats a good trick thanks. I should try that out again. Have been scared of it after failing at it once a long time ago :)",
    "1196853": "Hi,\n\nI use 5 fold CV instead of just one validation set to make sure to see all the examples in the dataset. IMHO, this approach also seems more robust in order to evaluate our results. Also if you find a good validation scheme which reflects the LB well, you can be much more confident about your improvements and decrease the potential noisy information.",
    "1200503": "Yeah. You're definitely right. I have just been hesitant to switch from single split to 5 fold CV because it would increases my training/experimentation time by 5x. This increase in training time sounds like it would really hinder my progression. How do you and other people normally deal with this? Or is that just the sacrifice you make?",
    "1200596": "Hi @raivokoot it is a good idea to sperate the training and the inference. \nAlso, I use colab for inference as the GPU times runs out in Kaggle\n\ncolab is slower but you have unlimited GPU",
    "1200729": "Hi again @raivokoot \n\nAs I said, 5 fold will give you much more robust estimation of your performance. You might end up choosing a \"lucky\" subset for your validation set and your performance might not generalize well to unseen data (Let's remind ourselves that we will also be tested with the private test set). You basically eliminate the chance of choosing that one particular good validation set. In K-Fold CV instead, you use all your data for validation once. \n\nAnd yes it increases the training time, but we become much more confident about our model's performance. I'm not an experienced kaggler, but this is the general scheme I've seen people follow."
  },
  "source": "meta"
}