{
  "id": 574495,
  "title": "Scaling up results- Kaggle data vs full data",
  "url": "/competitions/waveform-inversion/discussion/574495",
  "author_name": "greySnow",
  "post_date": "2025-04-22T07:58:30.799000",
  "votes": 27,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Some people probably wondering how much more data help.  <br>\nFor the same model, same validation dataset, BS 64, overfitting on data the size of Kaggle data (=10K samples, 1000 from each sub-dataset. But not the same data exactly): 50 train/227 val.  <br>\nOn all train data, BS 256: 60 train/70 val/80 LB.  <br>\nIt's not a comprehensive study. Just to get a feel.  </p>",
  "messages": [
    {
      "id": 3184523,
      "postDate": "2025-04-22T07:58:30.800Z",
      "content": "<p>Some people probably wondering how much more data help.  <br>\nFor the same model, same validation dataset, BS 64, overfitting on data the size of Kaggle data (=10K samples, 1000 from each sub-dataset. But not the same data exactly): 50 train/227 val.  <br>\nOn all train data, BS 256: 60 train/70 val/80 LB.  <br>\nIt's not a comprehensive study. Just to get a feel.  </p>",
      "rawMarkdown": "Some people probably wondering how much more data help.  \nFor the same model, same validation dataset, BS 64, overfitting on data the size of Kaggle data (=10K samples, 1000 from each sub-dataset. But not the same data exactly): 50 train/227 val.  \nOn all train data, BS 256: 60 train/70 val/80 LB.  \nIt's not a comprehensive study. Just to get a feel.  ",
      "votes": 27
    },
    {
      "id": 3184895,
      "postDate": "2025-04-22T15:54:15.833Z",
      "content": "<p>all train data: 55 train，60 val，64+ LB.</p>",
      "rawMarkdown": "all train data: 55 train，60 val，64+ LB.",
      "votes": 3
    },
    {
      "id": 3185496,
      "postDate": "2025-04-23T12:02:18.750Z",
      "content": "<p>Thanks for sharing your results. Could you please share how many epochs you used for training your model on the full dataset?</p>",
      "rawMarkdown": "Thanks for sharing your results. Could you please share how many epochs you used for training your model on the full dataset?",
      "replies": [
        {
          "id": 3185686,
          "postDate": "2025-04-23T17:31:17.570Z",
          "rawMarkdown": "",
          "votes": -5,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3184833,
      "postDate": "2025-04-22T14:05:06.650Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true,
      "replies": [
        {
          "id": 3185656,
          "postDate": "2025-04-23T16:51:03.447Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3185660,
          "postDate": "2025-04-23T16:54:55.580Z",
          "content": "<p>Hi, I just wanted to clarify my earlier comment. I'm new to this platform and a first-year BTech student, eager to learn more and engage with the Kaggle community. I initially thought that comments like 'Thank you' or 'Well done' were a good way to engage, but after reading more, I realized that it's more valuable to share thoughts related to data science.<br>\nTo help me get started, I used AI to draft my comment, but in hindsight, I see that it may have come across as too polished or like I was trying to appear overly knowledgeable. Thank you for the feedbacks, and I'll be more mindful moving forward in contributing thoughts that are more aligned with the community's standards.<br>\nThank you all for your understanding. I'm genuinely looking forward to learning and growing here. I’ll be deleting my earlier comment to avoid clutter, but I truly value being a part of this community.</p>",
          "rawMarkdown": "Hi, I just wanted to clarify my earlier comment. I'm new to this platform and a first-year BTech student, eager to learn more and engage with the Kaggle community. I initially thought that comments like 'Thank you' or 'Well done' were a good way to engage, but after reading more, I realized that it's more valuable to share thoughts related to data science.\nTo help me get started, I used AI to draft my comment, but in hindsight, I see that it may have come across as too polished or like I was trying to appear overly knowledgeable. Thank you for the feedbacks, and I'll be more mindful moving forward in contributing thoughts that are more aligned with the community's standards.\nThank you all for your understanding. I'm genuinely looking forward to learning and growing here. I’ll be deleting my earlier comment to avoid clutter, but I truly value being a part of this community.",
          "votes": 1,
          "replies": [
            {
              "id": 3185708,
              "postDate": "2025-04-23T17:54:11.383Z",
              "content": "<p>Just a small advice. Don't try to engage. Instead, compete. Then, from the hardships, engages would come naturally.  <br>\nPeople may disagree, but in my experience, on Kaggle: engaging without competing = mostly useless spam.  </p>",
              "rawMarkdown": "Just a small advice. Don't try to engage. Instead, compete. Then, from the hardships, engages would come naturally.  \nPeople may disagree, but in my experience, on Kaggle: engaging without competing = mostly useless spam.  ",
              "votes": 9
            },
            {
              "id": 3185714,
              "postDate": "2025-04-23T18:05:02.573Z",
              "content": "<p>👍 thank you</p>",
              "rawMarkdown": "👍 thank you",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3184895,
      "author_name": "lin-kukuoreoa",
      "author_url": "",
      "post_date": "2025-04-22T15:54:15.833000",
      "content": "<p>all train data: 55 train，60 val，64+ LB.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3185496,
      "author_name": "Arash karoudi",
      "author_url": "",
      "post_date": "2025-04-23T12:02:18.750000",
      "content": "<p>Thanks for sharing your results. Could you please share how many epochs you used for training your model on the full dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3185686,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-04-23T17:31:17.570000",
          "content": "",
          "votes": -5,
          "replies": []
        }
      ]
    },
    {
      "id": 3184833,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-22T14:05:06.650000",
      "content": "",
      "votes": -4,
      "replies": [
        {
          "id": 3185656,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-04-23T16:51:03.447000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3185660,
          "author_name": "Gopi Krishnan R",
          "author_url": "",
          "post_date": "2025-04-23T16:54:55.580000",
          "content": "<p>Hi, I just wanted to clarify my earlier comment. I'm new to this platform and a first-year BTech student, eager to learn more and engage with the Kaggle community. I initially thought that comments like 'Thank you' or 'Well done' were a good way to engage, but after reading more, I realized that it's more valuable to share thoughts related to data science.<br>\nTo help me get started, I used AI to draft my comment, but in hindsight, I see that it may have come across as too polished or like I was trying to appear overly knowledgeable. Thank you for the feedbacks, and I'll be more mindful moving forward in contributing thoughts that are more aligned with the community's standards.<br>\nThank you all for your understanding. I'm genuinely looking forward to learning and growing here. I’ll be deleting my earlier comment to avoid clutter, but I truly value being a part of this community.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3185708,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-04-23T17:54:11.383000",
              "content": "<p>Just a small advice. Don't try to engage. Instead, compete. Then, from the hardships, engages would come naturally.  <br>\nPeople may disagree, but in my experience, on Kaggle: engaging without competing = mostly useless spam.  </p>",
              "votes": 9,
              "replies": []
            },
            {
              "id": 3185714,
              "author_name": "Gopi Krishnan R",
              "author_url": "",
              "post_date": "2025-04-23T18:05:02.573000",
              "content": "<p>👍 thank you</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3184523": "Some people probably wondering how much more data help.  \nFor the same model, same validation dataset, BS 64, overfitting on data the size of Kaggle data (=10K samples, 1000 from each sub-dataset. But not the same data exactly): 50 train/227 val.  \nOn all train data, BS 256: 60 train/70 val/80 LB.  \nIt's not a comprehensive study. Just to get a feel.  ",
    "3184895": "all train data: 55 train，60 val，64+ LB.",
    "3185496": "Thanks for sharing your results. Could you please share how many epochs you used for training your model on the full dataset?",
    "3184833": ""
  }
}