{
  "id": 584560,
  "title": "How should I split data for cross-validation with 940 files from 10 distinct families?",
  "url": "/competitions/waveform-inversion/discussion/584560",
  "author_name": "",
  "post_date": "2025-06-14T08:55:46.672229600Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>With dataset consisting of 940 files, with each file containing 500 samples. This data comes from 10 different families, and I need advice on the best strategy for splitting the data for cross-validation. What is the most appropriate split size and methodology to ensure a robust evaluation of my model, given the familial structure of the data?</p>",
  "messages": [
    {
      "id": "3224029",
      "postDate": "06/14/2025 08:55:46",
      "content": "<p>With dataset consisting of 940 files, with each file containing 500 samples. This data comes from 10 different families, and I need advice on the best strategy for splitting the data for cross-validation. What is the most appropriate split size and methodology to ensure a robust evaluation of my model, given the familial structure of the data?</p>",
      "rawMarkdown": "With dataset consisting of 940 files, with each file containing 500 samples. This data comes from 10 different families, and I need advice on the best strategy for splitting the data for cross-validation. What is the most appropriate split size and methodology to ensure a robust evaluation of my model, given the familial structure of the data?",
      "votes": null
    },
    {
      "id": "3224051",
      "postDate": "06/14/2025 10:03:30",
      "content": "<p>Your goal is to have local validation score changes reflect accurately on private LB. For this competition, we can be pretty sure (as long as host did not lie or have a huge bug) that if local is reflected in public LB, it will also be reflected in private. Shuffling and taking 1K samples for validation from each family (10K total) was good for me.  <br>\np.s. No need for cross-validation. Regular validation (1 split) is enough and is the best practice for large data.</p>",
      "rawMarkdown": "Your goal is to have local validation score changes reflect accurately on private LB. For this competition, we can be pretty sure (as long as host did not lie or have a huge bug) that if local is reflected in public LB, it will also be reflected in private. Shuffling and taking 1K samples for validation from each family (10K total) was good for me.  \np.s. No need for cross-validation. Regular validation (1 split) is enough and is the best practice for large data.",
      "votes": null
    },
    {
      "id": "3224094",
      "postDate": "06/14/2025 11:30:34",
      "content": "<p>Thanks! I'm using 5-fold cross-validation, stratified by family. I'm training two models, which could take several days to complete. So far, I haven’t been able to outperform the public best score.<br>\nBy the way, you mentioned it worked for you, what kind of results did you get? What’s your CV or public score?</p>",
      "rawMarkdown": "Thanks! I'm using 5-fold cross-validation, stratified by family. I'm training two models, which could take several days to complete. So far, I haven’t been able to outperform the public best score.\nBy the way, you mentioned it worked for you, what kind of results did you get? What’s your CV or public score?",
      "votes": null
    },
    {
      "id": "3224131",
      "postDate": "06/14/2025 12:55:51",
      "content": "<p>There is no CV. As I said, I only do validation. And it's very constant with public LB with changes.</p>",
      "rawMarkdown": "There is no CV. As I said, I only do validation. And it's very constant with public LB with changes.",
      "votes": null
    },
    {
      "id": "3224195",
      "postDate": "06/14/2025 14:28:29",
      "content": "<p>if you use <code>n_splits = 5</code> would the model be unstable due to small fold sizes? maybe LOGO might help in that case? not sure tho just something to consider depending on how much variance there is between families. if families differ a lot, LOGO might give a better picture of generalization.</p>",
      "rawMarkdown": "if you use `n_splits = 5` would the model be unstable due to small fold sizes? maybe LOGO might help in that case? not sure tho just something to consider depending on how much variance there is between families. if families differ a lot, LOGO might give a better picture of generalization.",
      "votes": null
    },
    {
      "id": "3224243",
      "postDate": "06/14/2025 15:03:53",
      "content": "<p>I'm finding it really difficult to get good results on the evaluation set with different types of data splits. Right now, I'm using a random split, one fold with 10% validation, stratified by family. Testing two backbones with different seeds for my model takes around 4 days, but I'm going ahead with it for the remaining time. Meanwhile, I'm also trying out different model architectures for other experiments after training done. maybe 4 different models to go. difficult to come with new idea with limited compute power. </p>",
      "rawMarkdown": "I'm finding it really difficult to get good results on the evaluation set with different types of data splits. Right now, I'm using a random split, one fold with 10% validation, stratified by family. Testing two backbones with different seeds for my model takes around 4 days, but I'm going ahead with it for the remaining time. Meanwhile, I'm also trying out different model architectures for other experiments after training done. maybe 4 different models to go. difficult to come with new idea with limited compute power.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3224051,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "06/14/2025 10:03:30",
      "content": "<p>Your goal is to have local validation score changes reflect accurately on private LB. For this competition, we can be pretty sure (as long as host did not lie or have a huge bug) that if local is reflected in public LB, it will also be reflected in private. Shuffling and taking 1K samples for validation from each family (10K total) was good for me.  <br>\np.s. No need for cross-validation. Regular validation (1 split) is enough and is the best practice for large data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3224094,
          "author_name": "jamalsaeedi",
          "author_url": "",
          "post_date": "06/14/2025 11:30:34",
          "content": "<p>Thanks! I'm using 5-fold cross-validation, stratified by family. I'm training two models, which could take several days to complete. So far, I haven’t been able to outperform the public best score.<br>\nBy the way, you mentioned it worked for you, what kind of results did you get? What’s your CV or public score?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3224131,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "06/14/2025 12:55:51",
              "content": "<p>There is no CV. As I said, I only do validation. And it's very constant with public LB with changes.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3224195,
      "author_name": "dakshbhatnagar08",
      "author_url": "",
      "post_date": "06/14/2025 14:28:29",
      "content": "<p>if you use <code>n_splits = 5</code> would the model be unstable due to small fold sizes? maybe LOGO might help in that case? not sure tho just something to consider depending on how much variance there is between families. if families differ a lot, LOGO might give a better picture of generalization.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3224243,
          "author_name": "jamalsaeedi",
          "author_url": "",
          "post_date": "06/14/2025 15:03:53",
          "content": "<p>I'm finding it really difficult to get good results on the evaluation set with different types of data splits. Right now, I'm using a random split, one fold with 10% validation, stratified by family. Testing two backbones with different seeds for my model takes around 4 days, but I'm going ahead with it for the remaining time. Meanwhile, I'm also trying out different model architectures for other experiments after training done. maybe 4 different models to go. difficult to come with new idea with limited compute power. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3224029": "With dataset consisting of 940 files, with each file containing 500 samples. This data comes from 10 different families, and I need advice on the best strategy for splitting the data for cross-validation. What is the most appropriate split size and methodology to ensure a robust evaluation of my model, given the familial structure of the data?",
    "3224051": "Your goal is to have local validation score changes reflect accurately on private LB. For this competition, we can be pretty sure (as long as host did not lie or have a huge bug) that if local is reflected in public LB, it will also be reflected in private. Shuffling and taking 1K samples for validation from each family (10K total) was good for me.  \np.s. No need for cross-validation. Regular validation (1 split) is enough and is the best practice for large data.",
    "3224094": "Thanks! I'm using 5-fold cross-validation, stratified by family. I'm training two models, which could take several days to complete. So far, I haven’t been able to outperform the public best score.\nBy the way, you mentioned it worked for you, what kind of results did you get? What’s your CV or public score?",
    "3224131": "There is no CV. As I said, I only do validation. And it's very constant with public LB with changes.",
    "3224195": "if you use `n_splits = 5` would the model be unstable due to small fold sizes? maybe LOGO might help in that case? not sure tho just something to consider depending on how much variance there is between families. if families differ a lot, LOGO might give a better picture of generalization.",
    "3224243": "I'm finding it really difficult to get good results on the evaluation set with different types of data splits. Right now, I'm using a random split, one fold with 10% validation, stratified by family. Testing two backbones with different seeds for my model takes around 4 days, but I'm going ahead with it for the remaining time. Meanwhile, I'm also trying out different model architectures for other experiments after training done. maybe 4 different models to go. difficult to come with new idea with limited compute power."
  },
  "source": "meta"
}