{
  "id": 397432,
  "title": "what percentage of data do you use for training?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/397432",
  "author_name": "",
  "post_date": "2023-03-25T15:51:40.863732300Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>After 10 percent of the entire training set, the accuracy of my model stops growing. What about you? </p>",
  "messages": [
    {
      "id": "2196698",
      "postDate": "03/25/2023 15:51:40",
      "content": "<p>After 10 percent of the entire training set, the accuracy of my model stops growing. What about you? </p>",
      "rawMarkdown": "After 10 percent of the entire training set, the accuracy of my model stops growing. What about you?",
      "votes": null
    },
    {
      "id": "2196793",
      "postDate": "03/25/2023 16:59:00",
      "content": "<p>Do you use the 10 percent one time, or use it repeatedly, like an epoch?<br>\nI am using 8-15% of the data for training, as a compromise between preprocessing storage requirements (and time to redo it when a different preprocessing is called for) and not getting enough data.</p>",
      "rawMarkdown": "Do you use the 10 percent one time, or use it repeatedly, like an epoch?\nI am using 8-15% of the data for training, as a compromise between preprocessing storage requirements (and time to redo it when a different preprocessing is called for) and not getting enough data.",
      "votes": null
    },
    {
      "id": "2203553",
      "postDate": "03/30/2023 23:44:32",
      "content": "<p>Due to the size of the data, I'm doing validation with a single batch reserved at the beginning when I shuffle the batch indexes prior to training. Every time I run a batch through training, that batch's metrics are evaluated against the validation batch. Each batch has about 200000 events each, so this is PLENTY. I tend to throw out the standard 80/20-90/10 rule if there's a ton of data to work with.</p>",
      "rawMarkdown": "Due to the size of the data, I'm doing validation with a single batch reserved at the beginning when I shuffle the batch indexes prior to training. Every time I run a batch through training, that batch's metrics are evaluated against the validation batch. Each batch has about 200000 events each, so this is PLENTY. I tend to throw out the standard 80/20-90/10 rule if there's a ton of data to work with.",
      "votes": null
    },
    {
      "id": "2205147",
      "postDate": "04/01/2023 09:50:10",
      "content": "<p>I use the first 3 batches as validation data. I've put this fixed for all my experiments.</p>\n<p>For training … trying out new hyperparameters, model etc with 5 to 10% of the data. If turns out to be an improvement… then a full run with up to 50% of the data.</p>\n<p>Doing a training run with all data I will perform in the last week…. not sure if there is much benefit though.</p>",
      "rawMarkdown": "I use the first 3 batches as validation data. I've put this fixed for all my experiments.\n\nFor training ... trying out new hyperparameters, model etc with 5 to 10% of the data. If turns out to be an improvement... then a full run with up to 50% of the data.\n\nDoing a training run with all data I will perform in the last week.... not sure if there is much benefit though.",
      "votes": null
    },
    {
      "id": "2220042",
      "postDate": "04/13/2023 04:51:03",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Just have a question. Feel free to not answer it if you don't want.<br>\nUsing more data is really time and resources consuming. Just want to know if it worth spending time on it or just it is better to make better preprocessing?</p>",
      "rawMarkdown": "rsmits Just have a question. Feel free to not answer it if you don't want.\nUsing more data is really time and resources consuming. Just want to know if it worth spending time on it or just it is better to make better preprocessing?",
      "votes": null
    },
    {
      "id": "2220633",
      "postDate": "04/13/2023 14:58:20",
      "content": "<p><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a> No problem. Thanks for the question. Based on current experiments I guess using about 100 training files is for most models likely more than enough.</p>\n<p>Using more files - I agree that is really time and resources consuming - but it does offer some small performance increase.</p>\n<p>If that increase is worth it depends on your model setup, preprocessing etc.</p>\n<p>You could increase the usage of training data. Also monitor the validation carefully and decide for yourself if it is still beneficial.</p>",
      "rawMarkdown": "mohammad2012191 No problem. Thanks for the question. Based on current experiments I guess using about 100 training files is for most models likely more than enough.\n\nUsing more files - I agree that is really time and resources consuming - but it does offer some small performance increase.\n\nIf that increase is worth it depends on your model setup, preprocessing etc.\n\nYou could increase the usage of training data. Also monitor the validation carefully and decide for yourself if it is still beneficial.",
      "votes": null
    },
    {
      "id": "2220849",
      "postDate": "04/13/2023 17:59:30",
      "content": "<p>Using more data is better, once you start training for a long time it'll eventually start to overfit. If you use partial data loading and fast on-line preprocessing, it uses the exact same resources while being able to use more batches.</p>",
      "rawMarkdown": "Using more data is better, once you start training for a long time it'll eventually start to overfit. If you use partial data loading and fast on-line preprocessing, it uses the exact same resources while being able to use more batches.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2196793,
      "author_name": "solverworld",
      "author_url": "",
      "post_date": "03/25/2023 16:59:00",
      "content": "<p>Do you use the 10 percent one time, or use it repeatedly, like an epoch?<br>\nI am using 8-15% of the data for training, as a compromise between preprocessing storage requirements (and time to redo it when a different preprocessing is called for) and not getting enough data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2203553,
      "author_name": "kwierman",
      "author_url": "",
      "post_date": "03/30/2023 23:44:32",
      "content": "<p>Due to the size of the data, I'm doing validation with a single batch reserved at the beginning when I shuffle the batch indexes prior to training. Every time I run a batch through training, that batch's metrics are evaluated against the validation batch. Each batch has about 200000 events each, so this is PLENTY. I tend to throw out the standard 80/20-90/10 rule if there's a ton of data to work with.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2205147,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "04/01/2023 09:50:10",
      "content": "<p>I use the first 3 batches as validation data. I've put this fixed for all my experiments.</p>\n<p>For training … trying out new hyperparameters, model etc with 5 to 10% of the data. If turns out to be an improvement… then a full run with up to 50% of the data.</p>\n<p>Doing a training run with all data I will perform in the last week…. not sure if there is much benefit though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2220042,
          "author_name": "mohammad2012191",
          "author_url": "",
          "post_date": "04/13/2023 04:51:03",
          "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Just have a question. Feel free to not answer it if you don't want.<br>\nUsing more data is really time and resources consuming. Just want to know if it worth spending time on it or just it is better to make better preprocessing?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2220633,
              "author_name": "rsmits",
              "author_url": "",
              "post_date": "04/13/2023 14:58:20",
              "content": "<p><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a> No problem. Thanks for the question. Based on current experiments I guess using about 100 training files is for most models likely more than enough.</p>\n<p>Using more files - I agree that is really time and resources consuming - but it does offer some small performance increase.</p>\n<p>If that increase is worth it depends on your model setup, preprocessing etc.</p>\n<p>You could increase the usage of training data. Also monitor the validation carefully and decide for yourself if it is still beneficial.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2220849,
              "author_name": "dipamc77",
              "author_url": "",
              "post_date": "04/13/2023 17:59:30",
              "content": "<p>Using more data is better, once you start training for a long time it'll eventually start to overfit. If you use partial data loading and fast on-line preprocessing, it uses the exact same resources while being able to use more batches.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2196698": "After 10 percent of the entire training set, the accuracy of my model stops growing. What about you?",
    "2196793": "Do you use the 10 percent one time, or use it repeatedly, like an epoch?\nI am using 8-15% of the data for training, as a compromise between preprocessing storage requirements (and time to redo it when a different preprocessing is called for) and not getting enough data.",
    "2203553": "Due to the size of the data, I'm doing validation with a single batch reserved at the beginning when I shuffle the batch indexes prior to training. Every time I run a batch through training, that batch's metrics are evaluated against the validation batch. Each batch has about 200000 events each, so this is PLENTY. I tend to throw out the standard 80/20-90/10 rule if there's a ton of data to work with.",
    "2205147": "I use the first 3 batches as validation data. I've put this fixed for all my experiments.\n\nFor training ... trying out new hyperparameters, model etc with 5 to 10% of the data. If turns out to be an improvement... then a full run with up to 50% of the data.\n\nDoing a training run with all data I will perform in the last week.... not sure if there is much benefit though.",
    "2220042": "rsmits Just have a question. Feel free to not answer it if you don't want.\nUsing more data is really time and resources consuming. Just want to know if it worth spending time on it or just it is better to make better preprocessing?",
    "2220633": "mohammad2012191 No problem. Thanks for the question. Based on current experiments I guess using about 100 training files is for most models likely more than enough.\n\nUsing more files - I agree that is really time and resources consuming - but it does offer some small performance increase.\n\nIf that increase is worth it depends on your model setup, preprocessing etc.\n\nYou could increase the usage of training data. Also monitor the validation carefully and decide for yourself if it is still beneficial.",
    "2220849": "Using more data is better, once you start training for a long time it'll eventually start to overfit. If you use partial data loading and fast on-line preprocessing, it uses the exact same resources while being able to use more batches."
  },
  "source": "meta"
}