{
  "id": 614693,
  "title": "Large Data Size",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/614693",
  "author_name": "",
  "post_date": "2025-11-05T17:33:06.628323200Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This is an interesting competition, but 20 GB (!) of data seems prohibitively large.  How are people dealing with such a large dataset?</p>",
  "messages": [
    {
      "id": "3311755",
      "postDate": "11/05/2025 17:33:06",
      "content": "<p>This is an interesting competition, but 20 GB (!) of data seems prohibitively large.  How are people dealing with such a large dataset?</p>",
      "rawMarkdown": "This is an interesting competition, but 20 GB (!) of data seems prohibitively large.  How are people dealing with such a large dataset?",
      "votes": null
    },
    {
      "id": "3312056",
      "postDate": "11/06/2025 10:35:18",
      "content": "<p>Good point! The dataset size reflects a careful balance between practicality and realism. In fact, datasets used in real-world applications within this field are often much larger than what we included in the competition. To encourage meaningful innovation, we aimed to evaluate methods across a wide range of heterogeneous characteristics, which naturally requires more data. The size of the dataset is therefore intentional and necessary to support a fair and comprehensive assessment of different approaches.</p>",
      "rawMarkdown": "Good point! The dataset size reflects a careful balance between practicality and realism. In fact, datasets used in real-world applications within this field are often much larger than what we included in the competition. To encourage meaningful innovation, we aimed to evaluate methods across a wide range of heterogeneous characteristics, which naturally requires more data. The size of the dataset is therefore intentional and necessary to support a fair and comprehensive assessment of different approaches.",
      "votes": null
    },
    {
      "id": "3317989",
      "postDate": "11/11/2025 00:22:27",
      "content": "<p>Is it allowed to use other, external datasets for training? Because if yes, that'd probably mean that this competition becomes quite pay-to-win depending on how much compute you have available. Additionally I was wondering, since the submission only contains the output and not the model, if there's a risk of people manually solving it or figuring out how the data was generated and then just entering this?</p>",
      "rawMarkdown": "Is it allowed to use other, external datasets for training? Because if yes, that'd probably mean that this competition becomes quite pay-to-win depending on how much compute you have available. Additionally I was wondering, since the submission only contains the output and not the model, if there's a risk of people manually solving it or figuring out how the data was generated and then just entering this?",
      "votes": null
    },
    {
      "id": "3318800",
      "postDate": "11/11/2025 12:05:10",
      "content": "<p>These are valid concerns. We discussed these extensively during the study design and peer review of the now \"registered\" protocol. </p>\n<blockquote>\n  <p>Additionally I was wondering, since the submission only contains the output and not the model, if there's a risk of people manually solving it or figuring out how the data was generated and then just entering this?</p>\n</blockquote>\n<p>Under the \"Prizes\" details on the overview page, we talk about monetary rewards and a scientific manuscript authorship. As stated, to win the prize money, a prerequisite is that the participants make their code open-source. Additionally, as stated under \"Scientific manuscript authorship\", the top 10 submissions will be invited to co-author a scientific manuscript describing the competition's coutcome with their model descriptions, code, and related discussions. Under the \"Code requirements\", we strongly encourage everyone to adhere to a code template that we provided that enables a uniform way of running models. This is because, the top 10 models will be further stress-tested in a subsequent phase on many other datasets outside Kaggle platform. Thus, the top submissions cannot be just the predictions without open-source code.</p>\n<blockquote>\n  <p>Is it allowed to use other, external datasets for training? Because if yes, that'd probably mean that this competition becomes quite pay-to-win depending on how much compute you have available.</p>\n</blockquote>\n<p>We initially considered restricting external data augmentation to ensure fairness. However, all the three external peer-reviewers (domian experts) of the now frozen protocol had good arguments against this, which we agreed with. \nTo quote some of their arguments: \"Verifying compliance with these restrictions is nearly impossible, but even if it were possible, is it desirable? In real applications, developers will utilize all available pretraining, augmentation, and fine-tuning techniques to optimize performance. Restricting these methods might be artificial and based on a traditional notion of fairness that does not align with practical ML development. I suggest that the authors reconsider their approach and focus on transparency instead, i.e. allowing these techniques as long as they are properly reported and the code is reproducible.\" The open-source requirement of code to win the prize money, and to be part of the scientific manuscript strives for such transparency. If training on external data is what it really takes to top the leaderboard, that is also useful knowledge for the field. </p>",
      "rawMarkdown": "These are valid concerns. We discussed these extensively during the study design and peer review of the now \"registered\" protocol. \n\n>Additionally I was wondering, since the submission only contains the output and not the model, if there's a risk of people manually solving it or figuring out how the data was generated and then just entering this?\n\nUnder the \"Prizes\" details on the overview page, we talk about monetary rewards and a scientific manuscript authorship. As stated, to win the prize money, a prerequisite is that the participants make their code open-source. Additionally, as stated under \"Scientific manuscript authorship\", the top 10 submissions will be invited to co-author a scientific manuscript describing the competition's coutcome with their model descriptions, code, and related discussions. Under the \"Code requirements\", we strongly encourage everyone to adhere to a code template that we provided that enables a uniform way of running models. This is because, the top 10 models will be further stress-tested in a subsequent phase on many other datasets outside Kaggle platform. Thus, the top submissions cannot be just the predictions without open-source code.\n\n>Is it allowed to use other, external datasets for training? Because if yes, that'd probably mean that this competition becomes quite pay-to-win depending on how much compute you have available.\n\nWe initially considered restricting external data augmentation to ensure fairness. However, all the three external peer-reviewers (domian experts) of the now frozen protocol had good arguments against this, which we agreed with. \nTo quote some of their arguments: \"Verifying compliance with these restrictions is nearly impossible, but even if it were possible, is it desirable? In real applications, developers will utilize all available pretraining, augmentation, and fine-tuning techniques to optimize performance. Restricting these methods might be artificial and based on a traditional notion of fairness that does not align with practical ML development. I suggest that the authors reconsider their approach and focus on transparency instead, i.e. allowing these techniques as long as they are properly reported and the code is reproducible.\" The open-source requirement of code to win the prize money, and to be part of the scientific manuscript strives for such transparency. If training on external data is what it really takes to top the leaderboard, that is also useful knowledge for the field.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3312056,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "11/06/2025 10:35:18",
      "content": "<p>Good point! The dataset size reflects a careful balance between practicality and realism. In fact, datasets used in real-world applications within this field are often much larger than what we included in the competition. To encourage meaningful innovation, we aimed to evaluate methods across a wide range of heterogeneous characteristics, which naturally requires more data. The size of the dataset is therefore intentional and necessary to support a fair and comprehensive assessment of different approaches.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3317989,
          "author_name": "jonaskreutz",
          "author_url": "",
          "post_date": "11/11/2025 00:22:27",
          "content": "<p>Is it allowed to use other, external datasets for training? Because if yes, that'd probably mean that this competition becomes quite pay-to-win depending on how much compute you have available. Additionally I was wondering, since the submission only contains the output and not the model, if there's a risk of people manually solving it or figuring out how the data was generated and then just entering this?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3318800,
              "author_name": "ckanduri",
              "author_url": "",
              "post_date": "11/11/2025 12:05:10",
              "content": "<p>These are valid concerns. We discussed these extensively during the study design and peer review of the now \"registered\" protocol. </p>\n<blockquote>\n  <p>Additionally I was wondering, since the submission only contains the output and not the model, if there's a risk of people manually solving it or figuring out how the data was generated and then just entering this?</p>\n</blockquote>\n<p>Under the \"Prizes\" details on the overview page, we talk about monetary rewards and a scientific manuscript authorship. As stated, to win the prize money, a prerequisite is that the participants make their code open-source. Additionally, as stated under \"Scientific manuscript authorship\", the top 10 submissions will be invited to co-author a scientific manuscript describing the competition's coutcome with their model descriptions, code, and related discussions. Under the \"Code requirements\", we strongly encourage everyone to adhere to a code template that we provided that enables a uniform way of running models. This is because, the top 10 models will be further stress-tested in a subsequent phase on many other datasets outside Kaggle platform. Thus, the top submissions cannot be just the predictions without open-source code.</p>\n<blockquote>\n  <p>Is it allowed to use other, external datasets for training? Because if yes, that'd probably mean that this competition becomes quite pay-to-win depending on how much compute you have available.</p>\n</blockquote>\n<p>We initially considered restricting external data augmentation to ensure fairness. However, all the three external peer-reviewers (domian experts) of the now frozen protocol had good arguments against this, which we agreed with. \nTo quote some of their arguments: \"Verifying compliance with these restrictions is nearly impossible, but even if it were possible, is it desirable? In real applications, developers will utilize all available pretraining, augmentation, and fine-tuning techniques to optimize performance. Restricting these methods might be artificial and based on a traditional notion of fairness that does not align with practical ML development. I suggest that the authors reconsider their approach and focus on transparency instead, i.e. allowing these techniques as long as they are properly reported and the code is reproducible.\" The open-source requirement of code to win the prize money, and to be part of the scientific manuscript strives for such transparency. If training on external data is what it really takes to top the leaderboard, that is also useful knowledge for the field. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3311755": "This is an interesting competition, but 20 GB (!) of data seems prohibitively large.  How are people dealing with such a large dataset?",
    "3312056": "Good point! The dataset size reflects a careful balance between practicality and realism. In fact, datasets used in real-world applications within this field are often much larger than what we included in the competition. To encourage meaningful innovation, we aimed to evaluate methods across a wide range of heterogeneous characteristics, which naturally requires more data. The size of the dataset is therefore intentional and necessary to support a fair and comprehensive assessment of different approaches.",
    "3317989": "Is it allowed to use other, external datasets for training? Because if yes, that'd probably mean that this competition becomes quite pay-to-win depending on how much compute you have available. Additionally I was wondering, since the submission only contains the output and not the model, if there's a risk of people manually solving it or figuring out how the data was generated and then just entering this?",
    "3318800": "These are valid concerns. We discussed these extensively during the study design and peer review of the now \"registered\" protocol. \n\n>Additionally I was wondering, since the submission only contains the output and not the model, if there's a risk of people manually solving it or figuring out how the data was generated and then just entering this?\n\nUnder the \"Prizes\" details on the overview page, we talk about monetary rewards and a scientific manuscript authorship. As stated, to win the prize money, a prerequisite is that the participants make their code open-source. Additionally, as stated under \"Scientific manuscript authorship\", the top 10 submissions will be invited to co-author a scientific manuscript describing the competition's coutcome with their model descriptions, code, and related discussions. Under the \"Code requirements\", we strongly encourage everyone to adhere to a code template that we provided that enables a uniform way of running models. This is because, the top 10 models will be further stress-tested in a subsequent phase on many other datasets outside Kaggle platform. Thus, the top submissions cannot be just the predictions without open-source code.\n\n>Is it allowed to use other, external datasets for training? Because if yes, that'd probably mean that this competition becomes quite pay-to-win depending on how much compute you have available.\n\nWe initially considered restricting external data augmentation to ensure fairness. However, all the three external peer-reviewers (domian experts) of the now frozen protocol had good arguments against this, which we agreed with. \nTo quote some of their arguments: \"Verifying compliance with these restrictions is nearly impossible, but even if it were possible, is it desirable? In real applications, developers will utilize all available pretraining, augmentation, and fine-tuning techniques to optimize performance. Restricting these methods might be artificial and based on a traditional notion of fairness that does not align with practical ML development. I suggest that the authors reconsider their approach and focus on transparency instead, i.e. allowing these techniques as long as they are properly reported and the code is reproducible.\" The open-source requirement of code to win the prize money, and to be part of the scientific manuscript strives for such transparency. If training on external data is what it really takes to top the leaderboard, that is also useful knowledge for the field."
  },
  "source": "meta"
}