{
  "id": 630408,
  "title": "Evaluation Metrics and Weight",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408",
  "author_name": "",
  "post_date": "2025-11-18T09:59:38.716082800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<blockquote>\n  <p>These will be used  to compute the performance metrics area under the ROC curve and Jaccard similarity, respectively, for each of the datasets. A <strong>weighted average</strong> of both measures across all the included datasets will be used as the basis for ranking on the leaderboard for the competition&gt;</p>\n</blockquote>\n<p>Can you explain more about weighted average? Is the leaderboard score calculated using 0.5<em>ROC + 0.5</em>Jaccard (or  0.4<em>ROC + 0.6</em>Jaccard ?)</p>",
  "messages": [
    {
      "id": "3336049",
      "postDate": "11/18/2025 09:59:38",
      "content": "<blockquote>\n  <p>These will be used  to compute the performance metrics area under the ROC curve and Jaccard similarity, respectively, for each of the datasets. A <strong>weighted average</strong> of both measures across all the included datasets will be used as the basis for ranking on the leaderboard for the competition&gt;</p>\n</blockquote>\n<p>Can you explain more about weighted average? Is the leaderboard score calculated using 0.5<em>ROC + 0.5</em>Jaccard (or  0.4<em>ROC + 0.6</em>Jaccard ?)</p>",
      "rawMarkdown": ">These will be used  to compute the performance metrics area under the ROC curve and Jaccard similarity, respectively, for each of the datasets. A **weighted average** of both measures across all the included datasets will be used as the basis for ranking on the leaderboard for the competition>\n\nCan you explain more about weighted average? Is the leaderboard score calculated using 0.5*ROC + 0.5*Jaccard (or  0.4*ROC + 0.6*Jaccard ?)",
      "votes": null
    },
    {
      "id": "3336553",
      "postDate": "11/18/2025 14:20:44",
      "content": "<p>That is a very good question. As you rightly refer to the quoted text from the overview page, the leaderboard score is based on a weighted average of both the Area Under the ROC Curve (AUC) for repertoire label prediction and the Jaccard Similarity for the recovery of immune signals. </p>\n<p>The weighting is applied more granularly across <em>each individual dataset</em> and <em>each metric</em>. Since there are 11 test datasets (corresponding to AUC) and 8 training datasets (corresponding to Jaccard), we assign different weights to each of those metrics for the Public and Private leaderboards. For the final Private leaderboard, the repertoire label prediction (ROC AUCs of individual datasets) contributes 75% towards the final score, whereas recovery of immune signals (Jaccard) contributes 25%. Hope this information helps you strategise 😀.</p>",
      "rawMarkdown": "That is a very good question. As you rightly refer to the quoted text from the overview page, the leaderboard score is based on a weighted average of both the Area Under the ROC Curve (AUC) for repertoire label prediction and the Jaccard Similarity for the recovery of immune signals. \n\nThe weighting is applied more granularly across *each individual dataset* and *each metric*. Since there are 11 test datasets (corresponding to AUC) and 8 training datasets (corresponding to Jaccard), we assign different weights to each of those metrics for the Public and Private leaderboards. For the final Private leaderboard, the repertoire label prediction (ROC AUCs of individual datasets) contributes 75% towards the final score, whereas recovery of immune signals (Jaccard) contributes 25%. Hope this information helps you strategise 😀.",
      "votes": null
    },
    {
      "id": "3343604",
      "postDate": "11/21/2025 19:22:46",
      "content": "<p>Can you give us weights for AUC ROC across individual datasets? This will help with local CV and testing.</p>",
      "rawMarkdown": "Can you give us weights for AUC ROC across individual datasets? This will help with local CV and testing.",
      "votes": null
    },
    {
      "id": "3347783",
      "postDate": "11/25/2025 12:04:45",
      "content": "<p>As you may have noticed on the discussion forum, we have now made the detailed research plan of this challenge public, which contains information on the weights we assign for individual datasets for Public and Private leaderboards. One would notice that we consciously assign zero as weights to some test sets, and the interpretability part (Jaccard) of the challenge for the Public leaderboard. A good strategy, thus, could be to use the weights we assign for the Private leaderboard. Happy modelling 🎉</p>",
      "rawMarkdown": "As you may have noticed on the discussion forum, we have now made the detailed research plan of this challenge public, which contains information on the weights we assign for individual datasets for Public and Private leaderboards. One would notice that we consciously assign zero as weights to some test sets, and the interpretability part (Jaccard) of the challenge for the Public leaderboard. A good strategy, thus, could be to use the weights we assign for the Private leaderboard. Happy modelling 🎉",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3336553,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "11/18/2025 14:20:44",
      "content": "<p>That is a very good question. As you rightly refer to the quoted text from the overview page, the leaderboard score is based on a weighted average of both the Area Under the ROC Curve (AUC) for repertoire label prediction and the Jaccard Similarity for the recovery of immune signals. </p>\n<p>The weighting is applied more granularly across <em>each individual dataset</em> and <em>each metric</em>. Since there are 11 test datasets (corresponding to AUC) and 8 training datasets (corresponding to Jaccard), we assign different weights to each of those metrics for the Public and Private leaderboards. For the final Private leaderboard, the repertoire label prediction (ROC AUCs of individual datasets) contributes 75% towards the final score, whereas recovery of immune signals (Jaccard) contributes 25%. Hope this information helps you strategise 😀.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3343604,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "11/21/2025 19:22:46",
          "content": "<p>Can you give us weights for AUC ROC across individual datasets? This will help with local CV and testing.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3347783,
              "author_name": "ckanduri",
              "author_url": "",
              "post_date": "11/25/2025 12:04:45",
              "content": "<p>As you may have noticed on the discussion forum, we have now made the detailed research plan of this challenge public, which contains information on the weights we assign for individual datasets for Public and Private leaderboards. One would notice that we consciously assign zero as weights to some test sets, and the interpretability part (Jaccard) of the challenge for the Public leaderboard. A good strategy, thus, could be to use the weights we assign for the Private leaderboard. Happy modelling 🎉</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3336049": ">These will be used  to compute the performance metrics area under the ROC curve and Jaccard similarity, respectively, for each of the datasets. A **weighted average** of both measures across all the included datasets will be used as the basis for ranking on the leaderboard for the competition>\n\nCan you explain more about weighted average? Is the leaderboard score calculated using 0.5*ROC + 0.5*Jaccard (or  0.4*ROC + 0.6*Jaccard ?)",
    "3336553": "That is a very good question. As you rightly refer to the quoted text from the overview page, the leaderboard score is based on a weighted average of both the Area Under the ROC Curve (AUC) for repertoire label prediction and the Jaccard Similarity for the recovery of immune signals. \n\nThe weighting is applied more granularly across *each individual dataset* and *each metric*. Since there are 11 test datasets (corresponding to AUC) and 8 training datasets (corresponding to Jaccard), we assign different weights to each of those metrics for the Public and Private leaderboards. For the final Private leaderboard, the repertoire label prediction (ROC AUCs of individual datasets) contributes 75% towards the final score, whereas recovery of immune signals (Jaccard) contributes 25%. Hope this information helps you strategise 😀.",
    "3343604": "Can you give us weights for AUC ROC across individual datasets? This will help with local CV and testing.",
    "3347783": "As you may have noticed on the discussion forum, we have now made the detailed research plan of this challenge public, which contains information on the weights we assign for individual datasets for Public and Private leaderboards. One would notice that we consciously assign zero as weights to some test sets, and the interpretability part (Jaccard) of the challenge for the Public leaderboard. A good strategy, thus, could be to use the weights we assign for the Private leaderboard. Happy modelling 🎉"
  },
  "source": "meta"
}