{
  "id": 658560,
  "title": "Clarifying two points about the data and LB",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/658560",
  "author_name": "",
  "post_date": "2025-12-11T12:17:06.652858700Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi! I'd like to clarify two things:</p>\n<ol>\n<li><p>In the synthetic datasets where immune signals come from real antigen-specific sequences (from Parse Biosciences), am I correct that these immune-associated sequences should not appear in negative repertoires? or is it possible for them to show up in negatives as well, if a same sequence happens to occur in the background distribution? The same question about very similar clonotypes to  the to immune-associated ones.</p></li>\n<li><p>For the test datasets that have non-zero weights for public leaderboard: are these datasets used in full, with all their test repertoires included in the public leaderboard score, without any subsampling or partial usage?</p></li>\n</ol>",
  "messages": [
    {
      "id": "3371336",
      "postDate": "12/11/2025 12:17:06",
      "content": "<p>Hi! I'd like to clarify two things:</p>\n<ol>\n<li><p>In the synthetic datasets where immune signals come from real antigen-specific sequences (from Parse Biosciences), am I correct that these immune-associated sequences should not appear in negative repertoires? or is it possible for them to show up in negatives as well, if a same sequence happens to occur in the background distribution? The same question about very similar clonotypes to  the to immune-associated ones.</p></li>\n<li><p>For the test datasets that have non-zero weights for public leaderboard: are these datasets used in full, with all their test repertoires included in the public leaderboard score, without any subsampling or partial usage?</p></li>\n</ol>",
      "rawMarkdown": "Hi! I'd like to clarify two things:\n\n1. In the synthetic datasets where immune signals come from real antigen-specific sequences (from Parse Biosciences), am I correct that these immune-associated sequences should not appear in negative repertoires? or is it possible for them to show up in negatives as well, if a same sequence happens to occur in the background distribution? The same question about very similar clonotypes to  the to immune-associated ones.\n\n2. For the test datasets that have non-zero weights for public leaderboard: are these datasets used in full, with all their test repertoires included in the public leaderboard score, without any subsampling or partial usage?",
      "votes": null
    },
    {
      "id": "3371508",
      "postDate": "12/11/2025 14:03:04",
      "content": "<blockquote>\n  <p>In the synthetic datasets where immune signals come from real antigen-specific sequences (from Parse Biosciences), am I correct that these immune-associated sequences should not appear in negative repertoires? or is it possible for them to show up in negatives as well, if a same sequence happens to occur in the background distribution? The same question about very similar clonotypes to the to immune-associated ones.</p>\n</blockquote>\n<p>In the real-world, immune state-associated sequences can appear in negative-labeled individuals for various reasons. All the details that we intended to provide on the synthetic datasets are described in the provided <a href=\"https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf\" target=\"_blank\">pre-registered protocol</a>. </p>\n<blockquote>\n  <p>For the test datasets that have non-zero weights for public leaderboard: are these datasets used in full, with all their test repertoires included in the public leaderboard score, without any subsampling or partial usage?</p>\n</blockquote>\n<p>The test datasets that have non-zero weights for the public leaderboard are used in full for the Public leaderboard 😀. </p>",
      "rawMarkdown": ">In the synthetic datasets where immune signals come from real antigen-specific sequences (from Parse Biosciences), am I correct that these immune-associated sequences should not appear in negative repertoires? or is it possible for them to show up in negatives as well, if a same sequence happens to occur in the background distribution? The same question about very similar clonotypes to the to immune-associated ones.\n\nIn the real-world, immune state-associated sequences can appear in negative-labeled individuals for various reasons. All the details that we intended to provide on the synthetic datasets are described in the provided [pre-registered protocol](https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf). \n\n\n>For the test datasets that have non-zero weights for public leaderboard: are these datasets used in full, with all their test repertoires included in the public leaderboard score, without any subsampling or partial usage?\n\nThe test datasets that have non-zero weights for the public leaderboard are used in full for the Public leaderboard 😀.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3371508,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "12/11/2025 14:03:04",
      "content": "<blockquote>\n  <p>In the synthetic datasets where immune signals come from real antigen-specific sequences (from Parse Biosciences), am I correct that these immune-associated sequences should not appear in negative repertoires? or is it possible for them to show up in negatives as well, if a same sequence happens to occur in the background distribution? The same question about very similar clonotypes to the to immune-associated ones.</p>\n</blockquote>\n<p>In the real-world, immune state-associated sequences can appear in negative-labeled individuals for various reasons. All the details that we intended to provide on the synthetic datasets are described in the provided <a href=\"https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf\" target=\"_blank\">pre-registered protocol</a>. </p>\n<blockquote>\n  <p>For the test datasets that have non-zero weights for public leaderboard: are these datasets used in full, with all their test repertoires included in the public leaderboard score, without any subsampling or partial usage?</p>\n</blockquote>\n<p>The test datasets that have non-zero weights for the public leaderboard are used in full for the Public leaderboard 😀. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3371336": "Hi! I'd like to clarify two things:\n\n1. In the synthetic datasets where immune signals come from real antigen-specific sequences (from Parse Biosciences), am I correct that these immune-associated sequences should not appear in negative repertoires? or is it possible for them to show up in negatives as well, if a same sequence happens to occur in the background distribution? The same question about very similar clonotypes to  the to immune-associated ones.\n\n2. For the test datasets that have non-zero weights for public leaderboard: are these datasets used in full, with all their test repertoires included in the public leaderboard score, without any subsampling or partial usage?",
    "3371508": ">In the synthetic datasets where immune signals come from real antigen-specific sequences (from Parse Biosciences), am I correct that these immune-associated sequences should not appear in negative repertoires? or is it possible for them to show up in negatives as well, if a same sequence happens to occur in the background distribution? The same question about very similar clonotypes to the to immune-associated ones.\n\nIn the real-world, immune state-associated sequences can appear in negative-labeled individuals for various reasons. All the details that we intended to provide on the synthetic datasets are described in the provided [pre-registered protocol](https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf). \n\n\n>For the test datasets that have non-zero weights for public leaderboard: are these datasets used in full, with all their test repertoires included in the public leaderboard score, without any subsampling or partial usage?\n\nThe test datasets that have non-zero weights for the public leaderboard are used in full for the Public leaderboard 😀."
  },
  "source": "meta"
}