{
  "id": 663316,
  "title": "Comment on scoring scheme",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/663316",
  "author_name": "",
  "post_date": "2025-12-17T11:28:19.844430700Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>One comment is that the difference between public and private scores is too big, which seems to indicate that the scoring setting for this competition is not appropriate.</p>\n<p>For example, many teams may see that the (public) scores of other teams are too high, and so they just give up while seeing that their (public) score is too far away. Suddenly, with the private score, many teams that were ahead of them became of lesser scores. </p>\n<p>Or in the opposite, a team saw that their public score is very high and believes that their private score will be similar, and becomes disappointed. </p>\n<p>There are extreme cases where the public and private rankings differ by 90. In our case, we went up more than 40. Which seems (and is) insane.      </p>\n<p>Even the code by the organisers, which boasted public score of 0.72, turned out to be only 0.52 in private scores. Which is very misleading! Of course, the organisers knew about this fact long ago, which make the fact more misleading to attendants. </p>\n<p>This points out that the way the validation data is used for public scores does not seem to have similar distribution as the validation data used for private scores. Also, the way the submissions are scored (for public and private ones) are too different, which make it difficult for teams. </p>",
  "messages": [
    {
      "id": "3378015",
      "postDate": "12/17/2025 11:28:19",
      "content": "<p>One comment is that the difference between public and private scores is too big, which seems to indicate that the scoring setting for this competition is not appropriate.</p>\n<p>For example, many teams may see that the (public) scores of other teams are too high, and so they just give up while seeing that their (public) score is too far away. Suddenly, with the private score, many teams that were ahead of them became of lesser scores. </p>\n<p>Or in the opposite, a team saw that their public score is very high and believes that their private score will be similar, and becomes disappointed. </p>\n<p>There are extreme cases where the public and private rankings differ by 90. In our case, we went up more than 40. Which seems (and is) insane.      </p>\n<p>Even the code by the organisers, which boasted public score of 0.72, turned out to be only 0.52 in private scores. Which is very misleading! Of course, the organisers knew about this fact long ago, which make the fact more misleading to attendants. </p>\n<p>This points out that the way the validation data is used for public scores does not seem to have similar distribution as the validation data used for private scores. Also, the way the submissions are scored (for public and private ones) are too different, which make it difficult for teams. </p>",
      "rawMarkdown": "One comment is that the difference between public and private scores is too big, which seems to indicate that the scoring setting for this competition is not appropriate.\n\nFor example, many teams may see that the (public) scores of other teams are too high, and so they just give up while seeing that their (public) score is too far away. Suddenly, with the private score, many teams that were ahead of them became of lesser scores. \n\nOr in the opposite, a team saw that their public score is very high and believes that their private score will be similar, and becomes disappointed. \n\nThere are extreme cases where the public and private rankings differ by 90. In our case, we went up more than 40. Which seems (and is) insane.      \n\nEven the code by the organisers, which boasted public score of 0.72, turned out to be only 0.52 in private scores. Which is very misleading! Of course, the organisers knew about this fact long ago, which make the fact more misleading to attendants. \n\nThis points out that the way the validation data is used for public scores does not seem to have similar distribution as the validation data used for private scores. Also, the way the submissions are scored (for public and private ones) are too different, which make it difficult for teams.",
      "votes": null
    },
    {
      "id": "3378329",
      "postDate": "12/17/2025 22:42:30",
      "content": "<p>I agree with kangarooooragnak - it seems to me that using only 1% of the test data is too little to reflect the performance of the models. Of course, it makes sense not to use the entire test set for public scoring during the competition, but the discrepancy between the public and private rankings is very large. It is a bit disappointing to discover at the end that the leaderboard we relied on throughout the competition was pretty uninformative. I hope this can be reconsidered for future competitions.</p>\n<p>In any case, thank you very much to the organizers for organizing this excellent and thoughtfully designed competition! The field will definitely benefit from this great initiative and the incredible efforts you made!</p>",
      "rawMarkdown": "I agree with kangarooooragnak - it seems to me that using only 1% of the test data is too little to reflect the performance of the models. Of course, it makes sense not to use the entire test set for public scoring during the competition, but the discrepancy between the public and private rankings is very large. It is a bit disappointing to discover at the end that the leaderboard we relied on throughout the competition was pretty uninformative. I hope this can be reconsidered for future competitions.\n\nIn any case, thank you very much to the organizers for organizing this excellent and thoughtfully designed competition! The field will definitely benefit from this great initiative and the incredible efforts you made!",
      "votes": null
    },
    {
      "id": "3378453",
      "postDate": "12/18/2025 08:28:35",
      "content": "<p>1% just means that Jaccard distance wasn’t evaluated on the public leaderboard. It doesn’t mean that only 1% of the data was used for validation. For rocauc, the validation was actually quite extensive, maybe even too much, in my opinion.</p>",
      "rawMarkdown": "1% just means that Jaccard distance wasn’t evaluated on the public leaderboard. It doesn’t mean that only 1% of the data was used for validation. For rocauc, the validation was actually quite extensive, maybe even too much, in my opinion.",
      "votes": null
    },
    {
      "id": "3378456",
      "postDate": "12/18/2025 08:37:38",
      "content": "<p>What you said could be true, but let's hear from the organisers for to be certain. Anyway, probably you do not deny that this is very misleading. For example, the team who went up 90 in ranking, maybe they already gave up from very early, and should they know that their ranking was much more promising than it was shown publicly, they could try and get into top 10. I can say that my team already kind of giving up 2 weeks ago (it was always publicly shown that the code by the organisers has better score than ours), and did not try to implement many ideas which we had. </p>",
      "rawMarkdown": "What you said could be true, but let's hear from the organisers for to be certain. Anyway, probably you do not deny that this is very misleading. For example, the team who went up 90 in ranking, maybe they already gave up from very early, and should they know that their ranking was much more promising than it was shown publicly, they could try and get into top 10. I can say that my team already kind of giving up 2 weeks ago (it was always publicly shown that the code by the organisers has better score than ours), and did not try to implement many ideas which we had.",
      "votes": null
    },
    {
      "id": "3378493",
      "postDate": "12/18/2025 11:01:38",
      "content": "<p>There's a public report that was published that clearly establishes that the public score was just based on classification task, while the private score includes both. This explains most of the score difference.</p>\n<p>As for the 1% part, the discussion at <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/651754#3362698\" target=\"_blank\">https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/651754#3362698</a> has your confirmation from the organizers that the 1% value was just because number of rows in task 2 outnumbering task 1.</p>",
      "rawMarkdown": "There's a public report that was published that clearly establishes that the public score was just based on classification task, while the private score includes both. This explains most of the score difference.\n\nAs for the 1% part, the discussion at https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/651754#3362698 has your confirmation from the organizers that the 1% value was just because number of rows in task 2 outnumbering task 1.",
      "votes": null
    },
    {
      "id": "3378501",
      "postDate": "12/18/2025 11:34:34",
      "content": "<p>That I know. However, it does not contradict the fact that the scoring scheme is misleading, and make it difficult for teams. </p>",
      "rawMarkdown": "That I know. However, it does not contradict the fact that the scoring scheme is misleading, and make it difficult for teams.",
      "votes": null
    },
    {
      "id": "3398853",
      "postDate": "01/29/2026 19:53:49",
      "content": "<p>Thanks for your patience on this thread; we are working through the backlog after Christmas and New Year break. We understand that rank changes can be surprising, and we thank you for sharing your perspective. With this, we also appreciate the opportunity to clarify our scoring design.</p>\n<p>Our primary goal was to create a rigorous challenge that rewards generalizable methods capable of advancing the field. The scoring system was carefully designed and described in our peer-reviewed registered report.</p>\n<p>Here are the key aspects addressing the points made in this thread:</p>\n<p><strong>Substantial feedback on public leaderboard:</strong> The public leaderboard provided continuous feedback on 50% of the test data for Task-1 (Repertoire-level classification). This task carried a 75% weight in the final private leaderboard score, offering substantial feedback. As clarified in the discussion forums previously, the statement of \"1% test data\" is auto-generated by Kaggle (which we are not allowed to correct) and is stemming from the fact that the solution file also included rows from Task-2 in a disproportionate way.</p>\n<p><strong>Indirect feedback on a blind Task:</strong> Task-2 (ranking label-associated sequences) was deliberately a \"blind\" task scored on the private leaderboard to prevent shortcuts (even unintentionally). However, methods that effectively learnt true label-associated sequences for Task-2 were also likely to perform well on Task-1 (where we provided 50% feedback), providing an indirect measure of success.</p>\n<p><strong>Rewarding generalizability:</strong> The private leaderboard was designed to test for true generalizability. It included held-out experimental datasets with sufficient variation to ensure that models relying on confounding signals would not perform well. </p>\n<p><strong>Relative rank over absolute score:</strong> The final scoring used different weights for the composite score and the results of the blind Task-2, as opposed to Public leaderboard scoring. Consequently, the public and private absolute scores are not directly comparable by design. The most meaningful comparison is the relative shift in ranking, which reflects a model's generalizability on unseen data. The difference in absolute scores could not have misled/confused the participants during the competition, as private leaderboard rankings were only revealed at the conclusion.</p>\n<p><strong>Leaderboard stability for top performers:</strong> We observed that the public leaderboard was a strong indicator for the top teams, where ~ 90% of the teams in the top-10 experienced minimal changes in their ranking. Although we did not yet analyze the performance of top-teams across individual datasets and tasks, this observation hints that their methods were potentially well-generalized.</p>\n<p><strong>Factors influencing rank changes:</strong> Several factors can contribute to shifts in the leaderboard, including the ability to maintain performance on unseen data, the density of scores at different leaderboard percentiles (gains from a dense neighborhood are not the same as gains in less-dense neighborhood), how many on a particular sample of leaderboard in a typical competition over-tune to the public leaderboard etc. </p>\n<p>In conclusion, our main interest is to advance the field of adaptive immune profiling, and we believe this competition structure has successfully spurred the interest we envisioned. We plan to share more detailed lessons learnt after a thorough analysis of the top-performing solutions.</p>\n<p>We appreciate your engagement, your valuable submissions, and your contribution to this important scientific endeavor. </p>",
      "rawMarkdown": "Thanks for your patience on this thread; we are working through the backlog after Christmas and New Year break. We understand that rank changes can be surprising, and we thank you for sharing your perspective. With this, we also appreciate the opportunity to clarify our scoring design.\n\nOur primary goal was to create a rigorous challenge that rewards generalizable methods capable of advancing the field. The scoring system was carefully designed and described in our peer-reviewed registered report.\n\nHere are the key aspects addressing the points made in this thread:\n\n**Substantial feedback on public leaderboard:** The public leaderboard provided continuous feedback on 50% of the test data for Task-1 (Repertoire-level classification). This task carried a 75% weight in the final private leaderboard score, offering substantial feedback. As clarified in the discussion forums previously, the statement of \"1% test data\" is auto-generated by Kaggle (which we are not allowed to correct) and is stemming from the fact that the solution file also included rows from Task-2 in a disproportionate way.\n\n**Indirect feedback on a blind Task:** Task-2 (ranking label-associated sequences) was deliberately a \"blind\" task scored on the private leaderboard to prevent shortcuts (even unintentionally). However, methods that effectively learnt true label-associated sequences for Task-2 were also likely to perform well on Task-1 (where we provided 50% feedback), providing an indirect measure of success.\n\n**Rewarding generalizability:** The private leaderboard was designed to test for true generalizability. It included held-out experimental datasets with sufficient variation to ensure that models relying on confounding signals would not perform well. \n\n**Relative rank over absolute score:** The final scoring used different weights for the composite score and the results of the blind Task-2, as opposed to Public leaderboard scoring. Consequently, the public and private absolute scores are not directly comparable by design. The most meaningful comparison is the relative shift in ranking, which reflects a model's generalizability on unseen data. The difference in absolute scores could not have misled/confused the participants during the competition, as private leaderboard rankings were only revealed at the conclusion.\n\n**Leaderboard stability for top performers:** We observed that the public leaderboard was a strong indicator for the top teams, where ~ 90% of the teams in the top-10 experienced minimal changes in their ranking. Although we did not yet analyze the performance of top-teams across individual datasets and tasks, this observation hints that their methods were potentially well-generalized.\n\n**Factors influencing rank changes:** Several factors can contribute to shifts in the leaderboard, including the ability to maintain performance on unseen data, the density of scores at different leaderboard percentiles (gains from a dense neighborhood are not the same as gains in less-dense neighborhood), how many on a particular sample of leaderboard in a typical competition over-tune to the public leaderboard etc. \n\nIn conclusion, our main interest is to advance the field of adaptive immune profiling, and we believe this competition structure has successfully spurred the interest we envisioned. We plan to share more detailed lessons learnt after a thorough analysis of the top-performing solutions.\n\nWe appreciate your engagement, your valuable submissions, and your contribution to this important scientific endeavor.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3378329,
      "author_name": "lielcohenlavi",
      "author_url": "",
      "post_date": "12/17/2025 22:42:30",
      "content": "<p>I agree with kangarooooragnak - it seems to me that using only 1% of the test data is too little to reflect the performance of the models. Of course, it makes sense not to use the entire test set for public scoring during the competition, but the discrepancy between the public and private rankings is very large. It is a bit disappointing to discover at the end that the leaderboard we relied on throughout the competition was pretty uninformative. I hope this can be reconsidered for future competitions.</p>\n<p>In any case, thank you very much to the organizers for organizing this excellent and thoughtfully designed competition! The field will definitely benefit from this great initiative and the incredible efforts you made!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3378453,
          "author_name": "gottalottarock",
          "author_url": "",
          "post_date": "12/18/2025 08:28:35",
          "content": "<p>1% just means that Jaccard distance wasn’t evaluated on the public leaderboard. It doesn’t mean that only 1% of the data was used for validation. For rocauc, the validation was actually quite extensive, maybe even too much, in my opinion.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3378456,
              "author_name": "kangarooooragnak",
              "author_url": "",
              "post_date": "12/18/2025 08:37:38",
              "content": "<p>What you said could be true, but let's hear from the organisers for to be certain. Anyway, probably you do not deny that this is very misleading. For example, the team who went up 90 in ranking, maybe they already gave up from very early, and should they know that their ranking was much more promising than it was shown publicly, they could try and get into top 10. I can say that my team already kind of giving up 2 weeks ago (it was always publicly shown that the code by the organisers has better score than ours), and did not try to implement many ideas which we had. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3378493,
                  "author_name": "sajayr",
                  "author_url": "",
                  "post_date": "12/18/2025 11:01:38",
                  "content": "<p>There's a public report that was published that clearly establishes that the public score was just based on classification task, while the private score includes both. This explains most of the score difference.</p>\n<p>As for the 1% part, the discussion at <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/651754#3362698\" target=\"_blank\">https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/651754#3362698</a> has your confirmation from the organizers that the 1% value was just because number of rows in task 2 outnumbering task 1.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3378501,
                      "author_name": "kangarooooragnak",
                      "author_url": "",
                      "post_date": "12/18/2025 11:34:34",
                      "content": "<p>That I know. However, it does not contradict the fact that the scoring scheme is misleading, and make it difficult for teams. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3398853,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "01/29/2026 19:53:49",
      "content": "<p>Thanks for your patience on this thread; we are working through the backlog after Christmas and New Year break. We understand that rank changes can be surprising, and we thank you for sharing your perspective. With this, we also appreciate the opportunity to clarify our scoring design.</p>\n<p>Our primary goal was to create a rigorous challenge that rewards generalizable methods capable of advancing the field. The scoring system was carefully designed and described in our peer-reviewed registered report.</p>\n<p>Here are the key aspects addressing the points made in this thread:</p>\n<p><strong>Substantial feedback on public leaderboard:</strong> The public leaderboard provided continuous feedback on 50% of the test data for Task-1 (Repertoire-level classification). This task carried a 75% weight in the final private leaderboard score, offering substantial feedback. As clarified in the discussion forums previously, the statement of \"1% test data\" is auto-generated by Kaggle (which we are not allowed to correct) and is stemming from the fact that the solution file also included rows from Task-2 in a disproportionate way.</p>\n<p><strong>Indirect feedback on a blind Task:</strong> Task-2 (ranking label-associated sequences) was deliberately a \"blind\" task scored on the private leaderboard to prevent shortcuts (even unintentionally). However, methods that effectively learnt true label-associated sequences for Task-2 were also likely to perform well on Task-1 (where we provided 50% feedback), providing an indirect measure of success.</p>\n<p><strong>Rewarding generalizability:</strong> The private leaderboard was designed to test for true generalizability. It included held-out experimental datasets with sufficient variation to ensure that models relying on confounding signals would not perform well. </p>\n<p><strong>Relative rank over absolute score:</strong> The final scoring used different weights for the composite score and the results of the blind Task-2, as opposed to Public leaderboard scoring. Consequently, the public and private absolute scores are not directly comparable by design. The most meaningful comparison is the relative shift in ranking, which reflects a model's generalizability on unseen data. The difference in absolute scores could not have misled/confused the participants during the competition, as private leaderboard rankings were only revealed at the conclusion.</p>\n<p><strong>Leaderboard stability for top performers:</strong> We observed that the public leaderboard was a strong indicator for the top teams, where ~ 90% of the teams in the top-10 experienced minimal changes in their ranking. Although we did not yet analyze the performance of top-teams across individual datasets and tasks, this observation hints that their methods were potentially well-generalized.</p>\n<p><strong>Factors influencing rank changes:</strong> Several factors can contribute to shifts in the leaderboard, including the ability to maintain performance on unseen data, the density of scores at different leaderboard percentiles (gains from a dense neighborhood are not the same as gains in less-dense neighborhood), how many on a particular sample of leaderboard in a typical competition over-tune to the public leaderboard etc. </p>\n<p>In conclusion, our main interest is to advance the field of adaptive immune profiling, and we believe this competition structure has successfully spurred the interest we envisioned. We plan to share more detailed lessons learnt after a thorough analysis of the top-performing solutions.</p>\n<p>We appreciate your engagement, your valuable submissions, and your contribution to this important scientific endeavor. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3378015": "One comment is that the difference between public and private scores is too big, which seems to indicate that the scoring setting for this competition is not appropriate.\n\nFor example, many teams may see that the (public) scores of other teams are too high, and so they just give up while seeing that their (public) score is too far away. Suddenly, with the private score, many teams that were ahead of them became of lesser scores. \n\nOr in the opposite, a team saw that their public score is very high and believes that their private score will be similar, and becomes disappointed. \n\nThere are extreme cases where the public and private rankings differ by 90. In our case, we went up more than 40. Which seems (and is) insane.      \n\nEven the code by the organisers, which boasted public score of 0.72, turned out to be only 0.52 in private scores. Which is very misleading! Of course, the organisers knew about this fact long ago, which make the fact more misleading to attendants. \n\nThis points out that the way the validation data is used for public scores does not seem to have similar distribution as the validation data used for private scores. Also, the way the submissions are scored (for public and private ones) are too different, which make it difficult for teams.",
    "3378329": "I agree with kangarooooragnak - it seems to me that using only 1% of the test data is too little to reflect the performance of the models. Of course, it makes sense not to use the entire test set for public scoring during the competition, but the discrepancy between the public and private rankings is very large. It is a bit disappointing to discover at the end that the leaderboard we relied on throughout the competition was pretty uninformative. I hope this can be reconsidered for future competitions.\n\nIn any case, thank you very much to the organizers for organizing this excellent and thoughtfully designed competition! The field will definitely benefit from this great initiative and the incredible efforts you made!",
    "3378453": "1% just means that Jaccard distance wasn’t evaluated on the public leaderboard. It doesn’t mean that only 1% of the data was used for validation. For rocauc, the validation was actually quite extensive, maybe even too much, in my opinion.",
    "3378456": "What you said could be true, but let's hear from the organisers for to be certain. Anyway, probably you do not deny that this is very misleading. For example, the team who went up 90 in ranking, maybe they already gave up from very early, and should they know that their ranking was much more promising than it was shown publicly, they could try and get into top 10. I can say that my team already kind of giving up 2 weeks ago (it was always publicly shown that the code by the organisers has better score than ours), and did not try to implement many ideas which we had.",
    "3378493": "There's a public report that was published that clearly establishes that the public score was just based on classification task, while the private score includes both. This explains most of the score difference.\n\nAs for the 1% part, the discussion at https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/651754#3362698 has your confirmation from the organizers that the 1% value was just because number of rows in task 2 outnumbering task 1.",
    "3378501": "That I know. However, it does not contradict the fact that the scoring scheme is misleading, and make it difficult for teams.",
    "3398853": "Thanks for your patience on this thread; we are working through the backlog after Christmas and New Year break. We understand that rank changes can be surprising, and we thank you for sharing your perspective. With this, we also appreciate the opportunity to clarify our scoring design.\n\nOur primary goal was to create a rigorous challenge that rewards generalizable methods capable of advancing the field. The scoring system was carefully designed and described in our peer-reviewed registered report.\n\nHere are the key aspects addressing the points made in this thread:\n\n**Substantial feedback on public leaderboard:** The public leaderboard provided continuous feedback on 50% of the test data for Task-1 (Repertoire-level classification). This task carried a 75% weight in the final private leaderboard score, offering substantial feedback. As clarified in the discussion forums previously, the statement of \"1% test data\" is auto-generated by Kaggle (which we are not allowed to correct) and is stemming from the fact that the solution file also included rows from Task-2 in a disproportionate way.\n\n**Indirect feedback on a blind Task:** Task-2 (ranking label-associated sequences) was deliberately a \"blind\" task scored on the private leaderboard to prevent shortcuts (even unintentionally). However, methods that effectively learnt true label-associated sequences for Task-2 were also likely to perform well on Task-1 (where we provided 50% feedback), providing an indirect measure of success.\n\n**Rewarding generalizability:** The private leaderboard was designed to test for true generalizability. It included held-out experimental datasets with sufficient variation to ensure that models relying on confounding signals would not perform well. \n\n**Relative rank over absolute score:** The final scoring used different weights for the composite score and the results of the blind Task-2, as opposed to Public leaderboard scoring. Consequently, the public and private absolute scores are not directly comparable by design. The most meaningful comparison is the relative shift in ranking, which reflects a model's generalizability on unseen data. The difference in absolute scores could not have misled/confused the participants during the competition, as private leaderboard rankings were only revealed at the conclusion.\n\n**Leaderboard stability for top performers:** We observed that the public leaderboard was a strong indicator for the top teams, where ~ 90% of the teams in the top-10 experienced minimal changes in their ranking. Although we did not yet analyze the performance of top-teams across individual datasets and tasks, this observation hints that their methods were potentially well-generalized.\n\n**Factors influencing rank changes:** Several factors can contribute to shifts in the leaderboard, including the ability to maintain performance on unseen data, the density of scores at different leaderboard percentiles (gains from a dense neighborhood are not the same as gains in less-dense neighborhood), how many on a particular sample of leaderboard in a typical competition over-tune to the public leaderboard etc. \n\nIn conclusion, our main interest is to advance the field of adaptive immune profiling, and we believe this competition structure has successfully spurred the interest we envisioned. We plan to share more detailed lessons learnt after a thorough analysis of the top-performing solutions.\n\nWe appreciate your engagement, your valuable submissions, and your contribution to this important scientific endeavor."
  },
  "source": "meta"
}