{
  "id": 663586,
  "title": "Thought on score distribution",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/663586",
  "author_name": "",
  "post_date": "2025-12-18T22:52:42.015344400Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Looking at the final leaderboard and the distribution of per-dataset AUC scores, one thing that stood out to me is that many AUCs fall systematically below 0.5 rather than clustering tightly around chance. In fact, across datasets, AUCs are below 0.5 substantially more often than above (&gt;70%), which strongly departs from symmetry around 0.5. From a signal-detection perspective, this suggests that several methods may be learning real, reproducible structure but with inconsistent or inverted score orientation across datasets. Since an AUC of 0.35 is mathematically equivalent to 0.65 under score inversion, it may be worth discussing how such orientation ambiguity should be interpreted when benchmarking repertoire-level models.</p>",
  "messages": [
    {
      "id": "3378963",
      "postDate": "12/18/2025 22:52:42",
      "content": "<p>Looking at the final leaderboard and the distribution of per-dataset AUC scores, one thing that stood out to me is that many AUCs fall systematically below 0.5 rather than clustering tightly around chance. In fact, across datasets, AUCs are below 0.5 substantially more often than above (&gt;70%), which strongly departs from symmetry around 0.5. From a signal-detection perspective, this suggests that several methods may be learning real, reproducible structure but with inconsistent or inverted score orientation across datasets. Since an AUC of 0.35 is mathematically equivalent to 0.65 under score inversion, it may be worth discussing how such orientation ambiguity should be interpreted when benchmarking repertoire-level models.</p>",
      "rawMarkdown": "Looking at the final leaderboard and the distribution of per-dataset AUC scores, one thing that stood out to me is that many AUCs fall systematically below 0.5 rather than clustering tightly around chance. In fact, across datasets, AUCs are below 0.5 substantially more often than above (>70%), which strongly departs from symmetry around 0.5. From a signal-detection perspective, this suggests that several methods may be learning real, reproducible structure but with inconsistent or inverted score orientation across datasets. Since an AUC of 0.35 is mathematically equivalent to 0.65 under score inversion, it may be worth discussing how such orientation ambiguity should be interpreted when benchmarking repertoire-level models.",
      "votes": null
    },
    {
      "id": "3379001",
      "postDate": "12/19/2025 01:44:07",
      "content": "<p>I think this is explained by the final weighting used: 75% AUC (across multiple datasets) + 25% Jaccard index for the top sequences (see <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408)\" target=\"_blank\">https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408)</a>. A random classifier would have 0.75 * 0.50 = 0.375 expected performance if the Jaccard index is 0, which is around where the lowest scores fall in the final leaderboard (with, say, +/- 5% AUC explained by chance).</p>",
      "rawMarkdown": "I think this is explained by the final weighting used: 75% AUC (across multiple datasets) + 25% Jaccard index for the top sequences (see https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408). A random classifier would have 0.75 * 0.50 = 0.375 expected performance if the Jaccard index is 0, which is around where the lowest scores fall in the final leaderboard (with, say, +/- 5% AUC explained by chance).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3379001,
      "author_name": "abuendia",
      "author_url": "",
      "post_date": "12/19/2025 01:44:07",
      "content": "<p>I think this is explained by the final weighting used: 75% AUC (across multiple datasets) + 25% Jaccard index for the top sequences (see <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408)\" target=\"_blank\">https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408)</a>. A random classifier would have 0.75 * 0.50 = 0.375 expected performance if the Jaccard index is 0, which is around where the lowest scores fall in the final leaderboard (with, say, +/- 5% AUC explained by chance).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3378963": "Looking at the final leaderboard and the distribution of per-dataset AUC scores, one thing that stood out to me is that many AUCs fall systematically below 0.5 rather than clustering tightly around chance. In fact, across datasets, AUCs are below 0.5 substantially more often than above (>70%), which strongly departs from symmetry around 0.5. From a signal-detection perspective, this suggests that several methods may be learning real, reproducible structure but with inconsistent or inverted score orientation across datasets. Since an AUC of 0.35 is mathematically equivalent to 0.65 under score inversion, it may be worth discussing how such orientation ambiguity should be interpreted when benchmarking repertoire-level models.",
    "3379001": "I think this is explained by the final weighting used: 75% AUC (across multiple datasets) + 25% Jaccard index for the top sequences (see https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408). A random classifier would have 0.75 * 0.50 = 0.375 expected performance if the Jaccard index is 0, which is around where the lowest scores fall in the final leaderboard (with, say, +/- 5% AUC explained by chance)."
  },
  "source": "meta"
}