{
  "id": 617864,
  "title": "What is “train_dataset_1”?",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/617864",
  "author_name": "",
  "post_date": "2025-11-12T12:17:34.708851900Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The training dataset seems to be divided into eight subgroups. Why is that?\nAre the subgroups separated by different diseases, or are they from different facilities but with the same disease?</p>\n<p>I’m currently struggling with how to construct my training dataset for model development.\nDo you have any information or insights about this?</p>",
  "messages": [
    {
      "id": "3320293",
      "postDate": "11/12/2025 12:17:34",
      "content": "<p>The training dataset seems to be divided into eight subgroups. Why is that?\nAre the subgroups separated by different diseases, or are they from different facilities but with the same disease?</p>\n<p>I’m currently struggling with how to construct my training dataset for model development.\nDo you have any information or insights about this?</p>",
      "rawMarkdown": "The training dataset seems to be divided into eight subgroups. Why is that?\nAre the subgroups separated by different diseases, or are they from different facilities but with the same disease?\n\nI’m currently struggling with how to construct my training dataset for model development.\nDo you have any information or insights about this?",
      "votes": null
    },
    {
      "id": "3320733",
      "postDate": "11/12/2025 18:46:38",
      "content": "<p>That is a reasonable question 😀 Yes, each training dataset can be considered as separated by a disease, infection or similar. Note that we provided one or more corresponding test dataset(s) for each training dataset and they both follow the same naming convention pattern (train_dataset_X, test_dataset_X). When multiple test datasets are provided per training dataset, they follow a naming convention like “test_dataset_7_1”, “test_dataset_7_2” and so on. </p>",
      "rawMarkdown": "That is a reasonable question 😀 Yes, each training dataset can be considered as separated by a disease, infection or similar. Note that we provided one or more corresponding test dataset(s) for each training dataset and they both follow the same naming convention pattern (train_dataset_X, test_dataset_X). When multiple test datasets are provided per training dataset, they follow a naming convention like “test_dataset_7_1”, “test_dataset_7_2” and so on.",
      "votes": null
    },
    {
      "id": "3322244",
      "postDate": "11/13/2025 12:13:35",
      "content": "<p>I also have a related question: if the training data is divided by disease or infection, why are the label_positive=False samples, which are expected to be healthy, also divided into subgroups?</p>",
      "rawMarkdown": "I also have a related question: if the training data is divided by disease or infection, why are the label_positive=False samples, which are expected to be healthy, also divided into subgroups?",
      "votes": null
    },
    {
      "id": "3322450",
      "postDate": "11/13/2025 14:25:07",
      "content": "<p>That is an excellent question.</p>\n<p>The structure you have noticed, where each training dataset has its own set of <code>label_positive=False</code> controls, reflects the typical data collection processes in observational studies with a case-control design. Each dataset (e.g., <code>train_dataset_1</code>, <code>train_dataset_2</code>) can be thought of as a distinct study or cohort. In one study, we do not know all labels (e.g, disease-1 is known only for <code>train_dataset_1</code>, but not <code>train_dataset_2</code>). In such studies, study organisers often aim to minimise systematic differences between comparison groups except for the outcome of interest. This is because systematic differences between comparison groups, such as variations in lab protocols, equipment, or patient demographics (e.g. age, sex, etc), can lead to misleading conclusions. However, not all studies succeed in eliminating such unwanted differences, and in such cases, the knowledge of systematic differences provides an opportunity for the modellers to account for such biases. Where available, we provided such metadata. </p>\n<p>The <code>'label_positive=False'</code> controls in one dataset are the specific comparators for the <code>'label_positive=True'</code> cases from that same study. As is common with real-world data, the degree of systematic difference within a study can vary, and this is one of the key challenges also in this competition. It is often observed that systematic differences <em>between</em> distinct studies (e.g., <code>train_dataset_1</code> vs <code>train_dataset_2</code>) are larger than the differences <em>within</em> a single study. Therefore, a common concern in modelling is to avoid having a model inadvertently learn to distinguish between the studies themselves, rather than the underlying biology of the disease.</p>\n<p>How the participants choose to use this information in their modelling strategy is entirely up to them. We are excited to see the innovative ways participants tackle these complexities.</p>",
      "rawMarkdown": "That is an excellent question.\n\nThe structure you have noticed, where each training dataset has its own set of `label_positive=False` controls, reflects the typical data collection processes in observational studies with a case-control design. Each dataset (e.g., `train_dataset_1`, `train_dataset_2`) can be thought of as a distinct study or cohort. In one study, we do not know all labels (e.g, disease-1 is known only for `train_dataset_1`, but not `train_dataset_2`). In such studies, study organisers often aim to minimise systematic differences between comparison groups except for the outcome of interest. This is because systematic differences between comparison groups, such as variations in lab protocols, equipment, or patient demographics (e.g. age, sex, etc), can lead to misleading conclusions. However, not all studies succeed in eliminating such unwanted differences, and in such cases, the knowledge of systematic differences provides an opportunity for the modellers to account for such biases. Where available, we provided such metadata. \n\nThe `'label_positive=False'` controls in one dataset are the specific comparators for the `'label_positive=True'` cases from that same study. As is common with real-world data, the degree of systematic difference within a study can vary, and this is one of the key challenges also in this competition. It is often observed that systematic differences *between* distinct studies (e.g., `train_dataset_1` vs `train_dataset_2`) are larger than the differences *within* a single study. Therefore, a common concern in modelling is to avoid having a model inadvertently learn to distinguish between the studies themselves, rather than the underlying biology of the disease.\n\nHow the participants choose to use this information in their modelling strategy is entirely up to them. We are excited to see the innovative ways participants tackle these complexities.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3320733,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "11/12/2025 18:46:38",
      "content": "<p>That is a reasonable question 😀 Yes, each training dataset can be considered as separated by a disease, infection or similar. Note that we provided one or more corresponding test dataset(s) for each training dataset and they both follow the same naming convention pattern (train_dataset_X, test_dataset_X). When multiple test datasets are provided per training dataset, they follow a naming convention like “test_dataset_7_1”, “test_dataset_7_2” and so on. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3322244,
          "author_name": "haru22",
          "author_url": "",
          "post_date": "11/13/2025 12:13:35",
          "content": "<p>I also have a related question: if the training data is divided by disease or infection, why are the label_positive=False samples, which are expected to be healthy, also divided into subgroups?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3322450,
              "author_name": "ckanduri",
              "author_url": "",
              "post_date": "11/13/2025 14:25:07",
              "content": "<p>That is an excellent question.</p>\n<p>The structure you have noticed, where each training dataset has its own set of <code>label_positive=False</code> controls, reflects the typical data collection processes in observational studies with a case-control design. Each dataset (e.g., <code>train_dataset_1</code>, <code>train_dataset_2</code>) can be thought of as a distinct study or cohort. In one study, we do not know all labels (e.g, disease-1 is known only for <code>train_dataset_1</code>, but not <code>train_dataset_2</code>). In such studies, study organisers often aim to minimise systematic differences between comparison groups except for the outcome of interest. This is because systematic differences between comparison groups, such as variations in lab protocols, equipment, or patient demographics (e.g. age, sex, etc), can lead to misleading conclusions. However, not all studies succeed in eliminating such unwanted differences, and in such cases, the knowledge of systematic differences provides an opportunity for the modellers to account for such biases. Where available, we provided such metadata. </p>\n<p>The <code>'label_positive=False'</code> controls in one dataset are the specific comparators for the <code>'label_positive=True'</code> cases from that same study. As is common with real-world data, the degree of systematic difference within a study can vary, and this is one of the key challenges also in this competition. It is often observed that systematic differences <em>between</em> distinct studies (e.g., <code>train_dataset_1</code> vs <code>train_dataset_2</code>) are larger than the differences <em>within</em> a single study. Therefore, a common concern in modelling is to avoid having a model inadvertently learn to distinguish between the studies themselves, rather than the underlying biology of the disease.</p>\n<p>How the participants choose to use this information in their modelling strategy is entirely up to them. We are excited to see the innovative ways participants tackle these complexities.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3320293": "The training dataset seems to be divided into eight subgroups. Why is that?\nAre the subgroups separated by different diseases, or are they from different facilities but with the same disease?\n\nI’m currently struggling with how to construct my training dataset for model development.\nDo you have any information or insights about this?",
    "3320733": "That is a reasonable question 😀 Yes, each training dataset can be considered as separated by a disease, infection or similar. Note that we provided one or more corresponding test dataset(s) for each training dataset and they both follow the same naming convention pattern (train_dataset_X, test_dataset_X). When multiple test datasets are provided per training dataset, they follow a naming convention like “test_dataset_7_1”, “test_dataset_7_2” and so on.",
    "3322244": "I also have a related question: if the training data is divided by disease or infection, why are the label_positive=False samples, which are expected to be healthy, also divided into subgroups?",
    "3322450": "That is an excellent question.\n\nThe structure you have noticed, where each training dataset has its own set of `label_positive=False` controls, reflects the typical data collection processes in observational studies with a case-control design. Each dataset (e.g., `train_dataset_1`, `train_dataset_2`) can be thought of as a distinct study or cohort. In one study, we do not know all labels (e.g, disease-1 is known only for `train_dataset_1`, but not `train_dataset_2`). In such studies, study organisers often aim to minimise systematic differences between comparison groups except for the outcome of interest. This is because systematic differences between comparison groups, such as variations in lab protocols, equipment, or patient demographics (e.g. age, sex, etc), can lead to misleading conclusions. However, not all studies succeed in eliminating such unwanted differences, and in such cases, the knowledge of systematic differences provides an opportunity for the modellers to account for such biases. Where available, we provided such metadata. \n\nThe `'label_positive=False'` controls in one dataset are the specific comparators for the `'label_positive=True'` cases from that same study. As is common with real-world data, the degree of systematic difference within a study can vary, and this is one of the key challenges also in this competition. It is often observed that systematic differences *between* distinct studies (e.g., `train_dataset_1` vs `train_dataset_2`) are larger than the differences *within* a single study. Therefore, a common concern in modelling is to avoid having a model inadvertently learn to distinguish between the studies themselves, rather than the underlying biology of the disease.\n\nHow the participants choose to use this information in their modelling strategy is entirely up to them. We are excited to see the innovative ways participants tackle these complexities."
  },
  "source": "meta"
}