{
  "id": 639975,
  "title": "Important Update: Frozen research plan of this challenge now made public!",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/639975",
  "author_name": "Chakravarthi Kanduri",
  "post_date": "2025-11-25T11:00:36.245000",
  "votes": 5,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>We have an exciting update! To foster greater collaboration and transparency, we are now sharing the frozen research plan, according to which this Kaggle challenge is being organised. Note that this research plan has been peer-reviewed through a scientific journal by independent domain experts and approved <em>before</em> the launch of this challenge. You can find it here: <a href=\"https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf\" target=\"_blank\">https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf</a></p>\n<p>By providing you with all the available information, we hope to empower you to develop truly generalizable and state-of-the-art models.</p>\n<p>Among other things, the registered report contains in-depth details on:</p>\n<ul>\n<li><strong>Scoring Metrics:</strong> A comprehensive explanation of how submissions are evaluated.</li>\n<li><strong>Datasets:</strong> Detailed descriptions of the training and testing datasets.</li>\n<li><strong>Pilot Data &amp; Baselines:</strong> Insights from our pilot data, including the performance of baseline methods.</li>\n</ul>\n<p>We hope that this information will help you better understand the challenge and develop novel solutions. We are excited to see the innovative approaches you will bring to the competition.</p>\n<p>Happy modeling!</p>\n<p>Best Wishes,\nThe AIRR-ML-25 Organizers</p>",
  "messages": [
    {
      "id": 3347710,
      "postDate": "2025-11-25T11:00:36.247Z",
      "content": "<p>Hi everyone,</p>\n<p>We have an exciting update! To foster greater collaboration and transparency, we are now sharing the frozen research plan, according to which this Kaggle challenge is being organised. Note that this research plan has been peer-reviewed through a scientific journal by independent domain experts and approved <em>before</em> the launch of this challenge. You can find it here: <a href=\"https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf\" target=\"_blank\">https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf</a></p>\n<p>By providing you with all the available information, we hope to empower you to develop truly generalizable and state-of-the-art models.</p>\n<p>Among other things, the registered report contains in-depth details on:</p>\n<ul>\n<li><strong>Scoring Metrics:</strong> A comprehensive explanation of how submissions are evaluated.</li>\n<li><strong>Datasets:</strong> Detailed descriptions of the training and testing datasets.</li>\n<li><strong>Pilot Data &amp; Baselines:</strong> Insights from our pilot data, including the performance of baseline methods.</li>\n</ul>\n<p>We hope that this information will help you better understand the challenge and develop novel solutions. We are excited to see the innovative approaches you will bring to the competition.</p>\n<p>Happy modeling!</p>\n<p>Best Wishes,\nThe AIRR-ML-25 Organizers</p>",
      "rawMarkdown": "Hi everyone,\n\nWe have an exciting update! To foster greater collaboration and transparency, we are now sharing the frozen research plan, according to which this Kaggle challenge is being organised. Note that this research plan has been peer-reviewed through a scientific journal by independent domain experts and approved *before* the launch of this challenge. You can find it here: https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf\n\nBy providing you with all the available information, we hope to empower you to develop truly generalizable and state-of-the-art models.\n\nAmong other things, the registered report contains in-depth details on:\n*   **Scoring Metrics:** A comprehensive explanation of how submissions are evaluated.\n*   **Datasets:** Detailed descriptions of the training and testing datasets.\n*   **Pilot Data & Baselines:** Insights from our pilot data, including the performance of baseline methods.\n\nWe hope that this information will help you better understand the challenge and develop novel solutions. We are excited to see the innovative approaches you will bring to the competition.\n\nHappy modeling!\n\nBest Wishes,\nThe AIRR-ML-25 Organizers",
      "votes": 5
    },
    {
      "id": 3366425,
      "postDate": "2025-12-07T17:58:43.907Z",
      "content": "<p>Hello, i have a question regarding Table 2. Are the datasets 1-6 from the table 2 in the same order as the ones that were provided to us? I mean ds1 is the Dataset 1 referred to Table 2, ds2 the Dataset 2 referred to Table 2 etc.. ??</p>",
      "rawMarkdown": "Hello, i have a question regarding Table 2. Are the datasets 1-6 from the table 2 in the same order as the ones that were provided to us? I mean ds1 is the Dataset 1 referred to Table 2, ds2 the Dataset 2 referred to Table 2 etc.. ??",
      "votes": 1,
      "replies": [
        {
          "id": 3367207,
          "postDate": "2025-12-08T09:26:35.817Z",
          "content": "<p>Good question 😀. We have intentionally not revealed the mapping between the datasets provided on Kaggle and their descriptions in the registered report. </p>",
          "rawMarkdown": "Good question 😀. We have intentionally not revealed the mapping between the datasets provided on Kaggle and their descriptions in the registered report. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3356337,
      "postDate": "2025-11-30T22:20:46.987Z",
      "content": "<p>Why train data is devided in 8 parts? Why not just 2(synthetic and experimental)? Do you recommend training all 8 separately like in a baseline that you provided?</p>",
      "rawMarkdown": "Why train data is devided in 8 parts? Why not just 2(synthetic and experimental)? Do you recommend training all 8 separately like in a baseline that you provided?",
      "replies": [
        {
          "id": 3356957,
          "postDate": "2025-12-01T05:30:36.723Z",
          "content": "<p>Does this thread answer your question 😀 : <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/617864\" target=\"_blank\">https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/617864</a></p>",
          "rawMarkdown": "Does this thread answer your question 😀 : https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/617864"
        }
      ]
    },
    {
      "id": 3348369,
      "postDate": "2025-11-25T19:18:13.463Z",
      "content": "<p>\"Modern Hopfield Networks and Attention for Immune Repertoire Classification\"</p>\n<p>in this paper scores around 0.8 are reported with deeprc and svm model. i have been trying to implement deeprc and svm, but maybe it was not correct because i got score worse than logistic regression. </p>\n<p>is score like that possible to achieve with the datasets that we have here in this competition?</p>",
      "rawMarkdown": "\"Modern Hopfield Networks and Attention for Immune Repertoire Classification\"\n\nin this paper scores around 0.8 are reported with deeprc and svm model. i have been trying to implement deeprc and svm, but maybe it was not correct because i got score worse than logistic regression. \n\nis score like that possible to achieve with the datasets that we have here in this competition?",
      "replies": [
        {
          "id": 3348441,
          "postDate": "2025-11-25T21:27:19.190Z",
          "content": "<p>It is a very good observation! Not necessarily all immune states (e.g. different infections, autoimmune diseases etc) exhibit similar level and types of changes in immune repertoires, and thus the performances of methods can vary. We have not evaluated DeepRC or SVM on these datasets and thus cannot comment on their performance. But we are equally curious to know what approaches will be more generalizable and robust, and in that context, very glad that the participants are trying different approaches and being curious also about existing methods/approaches; kudos 🎉</p>",
          "rawMarkdown": "It is a very good observation! Not necessarily all immune states (e.g. different infections, autoimmune diseases etc) exhibit similar level and types of changes in immune repertoires, and thus the performances of methods can vary. We have not evaluated DeepRC or SVM on these datasets and thus cannot comment on their performance. But we are equally curious to know what approaches will be more generalizable and robust, and in that context, very glad that the participants are trying different approaches and being curious also about existing methods/approaches; kudos 🎉"
        }
      ]
    },
    {
      "id": 3347867,
      "postDate": "2025-11-25T13:07:09.790Z",
      "content": "<p>Thanks for the detailed information. </p>\n<p>While it is interesting to get more information about the datasets and possible approaches, I think this unfortunately also allows hacking the final leaderboard by using information on the test datasets. </p>\n<p>E.g. this \"…we generated three datasets with increasing signal abundance (5, 18 20, and 100 immune state-associated sequences per positive labeled repertoires), where shared complex 19 sequence patterns constitute the immune signals. Specifically, the immune state-associated sequences share20 short sub-sequence patterns spanning 4 to 6 amino acid residues with up to a maximum of 3 gapped positions obtained through rejection sampling of sequences from known V(D)J recombination models\" seems like something that could quite easily be abused to win the competition while it will not generalize to real-world data. </p>\n<p>Will solutions doing this be disqualified or is this allowed/expected? - To what degree is using knowledge of the dataset-generation process (e.g., fixed signal abundance, motif length constraints, gapped-motif structure) to optimize prediction allowed, since it is now provided?</p>",
      "rawMarkdown": "Thanks for the detailed information. \n\nWhile it is interesting to get more information about the datasets and possible approaches, I think this unfortunately also allows hacking the final leaderboard by using information on the test datasets. \n\nE.g. this \"...we generated three datasets with increasing signal abundance (5, 18 20, and 100 immune state-associated sequences per positive labeled repertoires), where shared complex 19 sequence patterns constitute the immune signals. Specifically, the immune state-associated sequences share20 short sub-sequence patterns spanning 4 to 6 amino acid residues with up to a maximum of 3 gapped positions obtained through rejection sampling of sequences from known V(D)J recombination models\" seems like something that could quite easily be abused to win the competition while it will not generalize to real-world data. \n\nWill solutions doing this be disqualified or is this allowed/expected? - To what degree is using knowledge of the dataset-generation process (e.g., fixed signal abundance, motif length constraints, gapped-motif structure) to optimize prediction allowed, since it is now provided?\n",
      "replies": [
        {
          "id": 3348464,
          "postDate": "2025-11-25T22:21:14.673Z",
          "content": "<p>This is an understandable concern. We, from the outset, wanted to make this information public (for the reasons stated below), and \"leaderboard hacking\" is something we discussed in that context and addressed carefully in the competition's design.</p>\n<p>While it is true that providing details on the synthetic data generation process (DGP) gives participants more information, we believe the competition is designed to reward consistently generalizable models. Here is our perspective:</p>\n<ul>\n<li><p><strong>Topping the leaderboard requires real-world performance:</strong> A key feature of the scoring is that <strong>50% of the leaderboard score comes from performance on real-world experimental data</strong>. For this data, the underlying \"rules\" are unknown. A model that only learns to exploit the specific structure of the synthetic datasets will perform poorly on this half of the challenge and will not be able to win. To secure a top position, a method must demonstrate that it can learn generalizable patterns from complex, real-world biological data.</p></li>\n<li><p><strong>Signal Recovery is a \"Blind\" Task:</strong> The task of recovering true immune signals (measured by the Jaccard score) is evaluated \"blind.\" This component accounts for <strong>25% of the final private leaderboard score, but we provide no continuous feedback on it in the public leaderboard.</strong> This design makes it impossible to iteratively \"hack\" one's way to a high score on this task. One has to build a model that they know is fundamentally good at signal discovery, as they only get one shot in the final evaluation.</p></li>\n<li><p><strong>Leveling the Playing Field to Foster Innovation:</strong> Real-world machine learning development always involves using domain knowledge. Experts in this field would likely incorporate this knowledge into their models anyway. By making the registered report public, our goal is to provide this domain knowledge to <em>all</em> participants, creating a more level playing field and encouraging everyone to innovate. The challenge, nevertheless, is not necessarily the lack of information often, but using it cleverly to build a model that generalizes well. Therefore, using this knowledge is absolutely encouraged and will not be a reason for disqualification. In our experience, building a high-performing model is a non-trivial task, even with this information. The true test of a winning model in this challenge will be its ability to perform well across a diverse set of problems, both synthetic and real.</p></li>\n</ul>\n<p>Thank you again for raising your concern and allowing us to clarify. We are very excited to see the creative and powerful solutions you and other participants develop!</p>",
          "rawMarkdown": "This is an understandable concern. We, from the outset, wanted to make this information public (for the reasons stated below), and \"leaderboard hacking\" is something we discussed in that context and addressed carefully in the competition's design.\n\nWhile it is true that providing details on the synthetic data generation process (DGP) gives participants more information, we believe the competition is designed to reward consistently generalizable models. Here is our perspective:\n\n- **Topping the leaderboard requires real-world performance:** A key feature of the scoring is that **50% of the leaderboard score comes from performance on real-world experimental data**. For this data, the underlying \"rules\" are unknown. A model that only learns to exploit the specific structure of the synthetic datasets will perform poorly on this half of the challenge and will not be able to win. To secure a top position, a method must demonstrate that it can learn generalizable patterns from complex, real-world biological data.\n\n- **Signal Recovery is a \"Blind\" Task:** The task of recovering true immune signals (measured by the Jaccard score) is evaluated \"blind.\" This component accounts for **25% of the final private leaderboard score, but we provide no continuous feedback on it in the public leaderboard.** This design makes it impossible to iteratively \"hack\" one's way to a high score on this task. One has to build a model that they know is fundamentally good at signal discovery, as they only get one shot in the final evaluation.\n\n- **Leveling the Playing Field to Foster Innovation:** Real-world machine learning development always involves using domain knowledge. Experts in this field would likely incorporate this knowledge into their models anyway. By making the registered report public, our goal is to provide this domain knowledge to *all* participants, creating a more level playing field and encouraging everyone to innovate. The challenge, nevertheless, is not necessarily the lack of information often, but using it cleverly to build a model that generalizes well. Therefore, using this knowledge is absolutely encouraged and will not be a reason for disqualification. In our experience, building a high-performing model is a non-trivial task, even with this information. The true test of a winning model in this challenge will be its ability to perform well across a diverse set of problems, both synthetic and real.\n\nThank you again for raising your concern and allowing us to clarify. We are very excited to see the creative and powerful solutions you and other participants develop!",
          "replies": [
            {
              "id": 3351509,
              "postDate": "2025-11-28T13:20:55.347Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3366425,
      "author_name": "aggelos papadopoulos",
      "author_url": "",
      "post_date": "2025-12-07T17:58:43.907000",
      "content": "<p>Hello, i have a question regarding Table 2. Are the datasets 1-6 from the table 2 in the same order as the ones that were provided to us? I mean ds1 is the Dataset 1 referred to Table 2, ds2 the Dataset 2 referred to Table 2 etc.. ??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3367207,
          "author_name": "Chakravarthi Kanduri",
          "author_url": "",
          "post_date": "2025-12-08T09:26:35.817000",
          "content": "<p>Good question 😀. We have intentionally not revealed the mapping between the datasets provided on Kaggle and their descriptions in the registered report. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3356337,
      "author_name": "Meet Babariya",
      "author_url": "",
      "post_date": "2025-11-30T22:20:46.987000",
      "content": "<p>Why train data is devided in 8 parts? Why not just 2(synthetic and experimental)? Do you recommend training all 8 separately like in a baseline that you provided?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3356957,
          "author_name": "Chakravarthi Kanduri",
          "author_url": "",
          "post_date": "2025-12-01T05:30:36.723000",
          "content": "<p>Does this thread answer your question 😀 : <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/617864\" target=\"_blank\">https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/617864</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3348369,
      "author_name": "Meet Babariya",
      "author_url": "",
      "post_date": "2025-11-25T19:18:13.463000",
      "content": "<p>\"Modern Hopfield Networks and Attention for Immune Repertoire Classification\"</p>\n<p>in this paper scores around 0.8 are reported with deeprc and svm model. i have been trying to implement deeprc and svm, but maybe it was not correct because i got score worse than logistic regression. </p>\n<p>is score like that possible to achieve with the datasets that we have here in this competition?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3348441,
          "author_name": "Chakravarthi Kanduri",
          "author_url": "",
          "post_date": "2025-11-25T21:27:19.190000",
          "content": "<p>It is a very good observation! Not necessarily all immune states (e.g. different infections, autoimmune diseases etc) exhibit similar level and types of changes in immune repertoires, and thus the performances of methods can vary. We have not evaluated DeepRC or SVM on these datasets and thus cannot comment on their performance. But we are equally curious to know what approaches will be more generalizable and robust, and in that context, very glad that the participants are trying different approaches and being curious also about existing methods/approaches; kudos 🎉</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3347867,
      "author_name": "Jokrasa",
      "author_url": "",
      "post_date": "2025-11-25T13:07:09.790000",
      "content": "<p>Thanks for the detailed information. </p>\n<p>While it is interesting to get more information about the datasets and possible approaches, I think this unfortunately also allows hacking the final leaderboard by using information on the test datasets. </p>\n<p>E.g. this \"…we generated three datasets with increasing signal abundance (5, 18 20, and 100 immune state-associated sequences per positive labeled repertoires), where shared complex 19 sequence patterns constitute the immune signals. Specifically, the immune state-associated sequences share20 short sub-sequence patterns spanning 4 to 6 amino acid residues with up to a maximum of 3 gapped positions obtained through rejection sampling of sequences from known V(D)J recombination models\" seems like something that could quite easily be abused to win the competition while it will not generalize to real-world data. </p>\n<p>Will solutions doing this be disqualified or is this allowed/expected? - To what degree is using knowledge of the dataset-generation process (e.g., fixed signal abundance, motif length constraints, gapped-motif structure) to optimize prediction allowed, since it is now provided?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3348464,
          "author_name": "Chakravarthi Kanduri",
          "author_url": "",
          "post_date": "2025-11-25T22:21:14.673000",
          "content": "<p>This is an understandable concern. We, from the outset, wanted to make this information public (for the reasons stated below), and \"leaderboard hacking\" is something we discussed in that context and addressed carefully in the competition's design.</p>\n<p>While it is true that providing details on the synthetic data generation process (DGP) gives participants more information, we believe the competition is designed to reward consistently generalizable models. Here is our perspective:</p>\n<ul>\n<li><p><strong>Topping the leaderboard requires real-world performance:</strong> A key feature of the scoring is that <strong>50% of the leaderboard score comes from performance on real-world experimental data</strong>. For this data, the underlying \"rules\" are unknown. A model that only learns to exploit the specific structure of the synthetic datasets will perform poorly on this half of the challenge and will not be able to win. To secure a top position, a method must demonstrate that it can learn generalizable patterns from complex, real-world biological data.</p></li>\n<li><p><strong>Signal Recovery is a \"Blind\" Task:</strong> The task of recovering true immune signals (measured by the Jaccard score) is evaluated \"blind.\" This component accounts for <strong>25% of the final private leaderboard score, but we provide no continuous feedback on it in the public leaderboard.</strong> This design makes it impossible to iteratively \"hack\" one's way to a high score on this task. One has to build a model that they know is fundamentally good at signal discovery, as they only get one shot in the final evaluation.</p></li>\n<li><p><strong>Leveling the Playing Field to Foster Innovation:</strong> Real-world machine learning development always involves using domain knowledge. Experts in this field would likely incorporate this knowledge into their models anyway. By making the registered report public, our goal is to provide this domain knowledge to <em>all</em> participants, creating a more level playing field and encouraging everyone to innovate. The challenge, nevertheless, is not necessarily the lack of information often, but using it cleverly to build a model that generalizes well. Therefore, using this knowledge is absolutely encouraged and will not be a reason for disqualification. In our experience, building a high-performing model is a non-trivial task, even with this information. The true test of a winning model in this challenge will be its ability to perform well across a diverse set of problems, both synthetic and real.</p></li>\n</ul>\n<p>Thank you again for raising your concern and allowing us to clarify. We are very excited to see the creative and powerful solutions you and other participants develop!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3351509,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-11-28T13:20:55.347000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3347710": "Hi everyone,\n\nWe have an exciting update! To foster greater collaboration and transparency, we are now sharing the frozen research plan, according to which this Kaggle challenge is being organised. Note that this research plan has been peer-reviewed through a scientific journal by independent domain experts and approved *before* the launch of this challenge. You can find it here: https://github.com/uio-bmi/adaptive_immune_profiling_challenge_2025/blob/main/registered_report.pdf\n\nBy providing you with all the available information, we hope to empower you to develop truly generalizable and state-of-the-art models.\n\nAmong other things, the registered report contains in-depth details on:\n*   **Scoring Metrics:** A comprehensive explanation of how submissions are evaluated.\n*   **Datasets:** Detailed descriptions of the training and testing datasets.\n*   **Pilot Data & Baselines:** Insights from our pilot data, including the performance of baseline methods.\n\nWe hope that this information will help you better understand the challenge and develop novel solutions. We are excited to see the innovative approaches you will bring to the competition.\n\nHappy modeling!\n\nBest Wishes,\nThe AIRR-ML-25 Organizers",
    "3366425": "Hello, i have a question regarding Table 2. Are the datasets 1-6 from the table 2 in the same order as the ones that were provided to us? I mean ds1 is the Dataset 1 referred to Table 2, ds2 the Dataset 2 referred to Table 2 etc.. ??",
    "3356337": "Why train data is devided in 8 parts? Why not just 2(synthetic and experimental)? Do you recommend training all 8 separately like in a baseline that you provided?",
    "3348369": "\"Modern Hopfield Networks and Attention for Immune Repertoire Classification\"\n\nin this paper scores around 0.8 are reported with deeprc and svm model. i have been trying to implement deeprc and svm, but maybe it was not correct because i got score worse than logistic regression. \n\nis score like that possible to achieve with the datasets that we have here in this competition?",
    "3347867": "Thanks for the detailed information. \n\nWhile it is interesting to get more information about the datasets and possible approaches, I think this unfortunately also allows hacking the final leaderboard by using information on the test datasets. \n\nE.g. this \"...we generated three datasets with increasing signal abundance (5, 18 20, and 100 immune state-associated sequences per positive labeled repertoires), where shared complex 19 sequence patterns constitute the immune signals. Specifically, the immune state-associated sequences share20 short sub-sequence patterns spanning 4 to 6 amino acid residues with up to a maximum of 3 gapped positions obtained through rejection sampling of sequences from known V(D)J recombination models\" seems like something that could quite easily be abused to win the competition while it will not generalize to real-world data. \n\nWill solutions doing this be disqualified or is this allowed/expected? - To what degree is using knowledge of the dataset-generation process (e.g., fixed signal abundance, motif length constraints, gapped-motif structure) to optimize prediction allowed, since it is now provided?\n"
  }
}