{
  "id": 617673,
  "title": "Ranking top sequences",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/617673",
  "author_name": "",
  "post_date": "2025-11-11T14:35:25.000660600Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>As we are being asked to rank top sequences and there are 50000 such rankings in each train dataset, it seems obvious that there will be no \"perfect\" matches. My question is, as there are 400,000 such rows in the submission file, what is the weight of the ranking compared to the weight of the classification?</p>",
  "messages": [
    {
      "id": "3319057",
      "postDate": "11/11/2025 14:35:25",
      "content": "<p>As we are being asked to rank top sequences and there are 50000 such rankings in each train dataset, it seems obvious that there will be no \"perfect\" matches. My question is, as there are 400,000 such rows in the submission file, what is the weight of the ranking compared to the weight of the classification?</p>",
      "rawMarkdown": "As we are being asked to rank top sequences and there are 50000 such rankings in each train dataset, it seems obvious that there will be no \"perfect\" matches. My question is, as there are 400,000 such rows in the submission file, what is the weight of the ranking compared to the weight of the classification?",
      "votes": null
    },
    {
      "id": "3320108",
      "postDate": "11/12/2025 08:43:41",
      "content": "<p>Hi! I am not sure if I fully follow your question 😀. Could you please elaborate more?</p>\n<p>Since there will be many more rows from task-2 (almost 100 times more) than from task-1 in the final submission file, how will these in the end contribute towards the leaderboard? (Here referring to task-1 and task-2 from the image shown under \"Evaluation\" in the Overview page). Is that your question?</p>\n<p>Or were you wondering about how important it is to get the ordering (ranking) correct and whether that affects the leaderboard score?</p>",
      "rawMarkdown": "Hi! I am not sure if I fully follow your question 😀. Could you please elaborate more?\n\nSince there will be many more rows from task-2 (almost 100 times more) than from task-1 in the final submission file, how will these in the end contribute towards the leaderboard? (Here referring to task-1 and task-2 from the image shown under \"Evaluation\" in the Overview page). Is that your question?\n\nOr were you wondering about how important it is to get the ordering (ranking) correct and whether that affects the leaderboard score?",
      "votes": null
    },
    {
      "id": "3320149",
      "postDate": "11/12/2025 09:19:22",
      "content": "<p>Apologies, I should have framed that more clearly. As you noted, having “100 times more” rows from task 2 could heavily bias the results toward that task. In addition, for the T1D data, there is no established “truth” in your ranking, especially if one of the competitors outperforms your benchmark. Continuing with that example, if someone does outperform your benchmark on the T1D classification task, their classification performance would justifiably improve their overall score. However, their ranking, which must differ from yours, would then contribute negatively to their ranking score, which seems counterproductive.</p>",
      "rawMarkdown": "Apologies, I should have framed that more clearly. As you noted, having “100 times more” rows from task 2 could heavily bias the results toward that task. In addition, for the T1D data, there is no established “truth” in your ranking, especially if one of the competitors outperforms your benchmark. Continuing with that example, if someone does outperform your benchmark on the T1D classification task, their classification performance would justifiably improve their overall score. However, their ranking, which must differ from yours, would then contribute negatively to their ranking score, which seems counterproductive.",
      "votes": null
    },
    {
      "id": "3320238",
      "postDate": "11/12/2025 10:52:41",
      "content": "<p>Good questions, and a very understandable dilemma. As we tried to show in the illustration on the overview page, we compute ROC AUC for each test dataset and Jaccard similarity for each training dataset. It is at this stage, we take a weighted average of these performance metrics to compute a leaderboard score. In other words, the leaderboard score is not heavily biased towards task-2. We asked for a ranked list only because each dataset may have different sizes of true immune signals, e.g., one dataset may have only 100 true immune signals, whereas another dataset may have 50,000. We take care of extracting the right number of predicted immune signals per each training dataset from the submission file. hope that clarifies 😀  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14003908%2F13d3d8ca174bb1cb8efd2bb9df671813%2Fevaluation_illustration_for_kaggle_page.png?generation=1756380205170091&amp;alt=media\" alt=\"\">, </p>",
      "rawMarkdown": "Good questions, and a very understandable dilemma. As we tried to show in the illustration on the overview page, we compute ROC AUC for each test dataset and Jaccard similarity for each training dataset. It is at this stage, we take a weighted average of these performance metrics to compute a leaderboard score. In other words, the leaderboard score is not heavily biased towards task-2. We asked for a ranked list only because each dataset may have different sizes of true immune signals, e.g., one dataset may have only 100 true immune signals, whereas another dataset may have 50,000. We take care of extracting the right number of predicted immune signals per each training dataset from the submission file. hope that clarifies 😀  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14003908%2F13d3d8ca174bb1cb8efd2bb9df671813%2Fevaluation_illustration_for_kaggle_page.png?generation=1756380205170091&alt=media),",
      "votes": null
    },
    {
      "id": "3320853",
      "postDate": "11/12/2025 21:06:42",
      "content": "<p>Much appreciated. I would be grateful if you could please clarify what is meant by the following: “We take care of extracting the right number of predicted immune signals per each training dataset from the submission file.”</p>\n<p>Also, I believe my earlier question about the potential contradiction between success in sample prediction and success in sequence ranking may not have been fully addressed. For instance, if your own benchmark for T1D achieves 0.9 on the test set, and another model achieves 0.91, the corresponding sequence rankings would necessarily differ. Wouldn’t this mean that an improved model could be penalized for outperforming the benchmark?</p>",
      "rawMarkdown": "Much appreciated. I would be grateful if you could please clarify what is meant by the following: “We take care of extracting the right number of predicted immune signals per each training dataset from the submission file.”\n\nAlso, I believe my earlier question about the potential contradiction between success in sample prediction and success in sequence ranking may not have been fully addressed. For instance, if your own benchmark for T1D achieves 0.9 on the test set, and another model achieves 0.91, the corresponding sequence rankings would necessarily differ. Wouldn’t this mean that an improved model could be penalized for outperforming the benchmark?",
      "votes": null
    },
    {
      "id": "3322134",
      "postDate": "11/13/2025 10:47:49",
      "content": "<p>Great follow-up questions, as this will also help clarification to other participants 😀. </p>\n<p>Short answer is: <strong>Submissions will not be penalized for creating a better model.</strong> A submission's ranking is <strong>not</strong> compared to any benchmark's ranking.</p>\n<p>Below is a more elaborate answer.</p>\n<h3>The \"ground truth\" is NOT based on any benchmark</h3>\n<ul>\n<li>The 'ground truth' for the ranking task (Task 2) is <strong>not</strong> based on any benchmark or model's results.</li>\n<li>Instead, the organizers have a <strong>fixed, finite 'ground truth' set</strong> of rows that are <em>known</em> to harbor 'true immune signals.'</li>\n<li>Think of this as a \"bag\" of true signals, not a ranked list. This knowledge is not derived from any model, and there is no \"ranking\" within this true set.</li>\n</ul>\n<p>Therefore, an improved model (like your 0.91 ROC AUC example) that finds more of these 'true' rows and ranks them highly will simply get a <strong>better</strong> score. It is never compared to any benchmark.</p>\n<h3>Why we ask for the top 50,000 rows (the \"top-N\")</h3>\n<p>This answers your first question about \"extracting the right number.\"</p>\n<ul>\n<li>The <em>size</em> of that 'ground truth set' (let's call its size 'N') is <strong>different for each training dataset.</strong></li>\n<li>For example, the true size 'N' for each dataset might be:<ul>\n<li><strong>Dataset 1:</strong> N = 40,943</li>\n<li><strong>Dataset 2:</strong> N = 32,664</li>\n<li><strong>Dataset 3:</strong> N = 9,923</li>\n<li>…and so on.</li></ul></li>\n<li>Instead of asking participants to guess these different 'N' values, we uniformly ask everyone to submit a <strong>top 50,000 ranked list.</strong> This is just a simple way to ensure one has submitted <em>enough</em> sequences to cover the largest possible 'N'.</li>\n</ul>\n<p><strong>How we score a submission:</strong></p>\n<p>When we calculate the Jaccard score for Dataset 1, our scoring script knows the true size is <strong>N = 40,943</strong>.</p>\n<ol>\n<li>It takes <strong>only the top 40,943</strong> ranked sequences from the participant's submission file.</li>\n<li>It treats those 40,943 sequences as a <strong>set</strong> (the order <em>within</em> this set no longer matters).</li>\n<li>It compares participants' predicted set to our 'ground truth set' (which also has 40,943 items) and computes the Jaccard similarity.</li>\n</ol>\n<p><strong>In summary:</strong> The <em>only</em> purpose of ranking is to determine which sequences land in participants' 'Top-N' set for each dataset. A better model will naturally place more 'true' sequences at the top of its list, creating a 'Top-N' set that overlaps more with the ground truth and resulting in a higher Jaccard score.</p>\n<p>Hope that makes it clear!</p>",
      "rawMarkdown": "Great follow-up questions, as this will also help clarification to other participants 😀. \n\nShort answer is: **Submissions will not be penalized for creating a better model.** A submission's ranking is **not** compared to any benchmark's ranking.\n\nBelow is a more elaborate answer.\n\n### The \"ground truth\" is NOT based on any benchmark\n\n* The 'ground truth' for the ranking task (Task 2) is **not** based on any benchmark or model's results.\n* Instead, the organizers have a **fixed, finite 'ground truth' set** of rows that are *known* to harbor 'true immune signals.'\n* Think of this as a \"bag\" of true signals, not a ranked list. This knowledge is not derived from any model, and there is no \"ranking\" within this true set.\n\nTherefore, an improved model (like your 0.91 ROC AUC example) that finds more of these 'true' rows and ranks them highly will simply get a **better** score. It is never compared to any benchmark.\n\n### Why we ask for the top 50,000 rows (the \"top-N\")\n\nThis answers your first question about \"extracting the right number.\"\n\n* The *size* of that 'ground truth set' (let's call its size 'N') is **different for each training dataset.**\n* For example, the true size 'N' for each dataset might be:\n    * **Dataset 1:** N = 40,943\n    * **Dataset 2:** N = 32,664\n    * **Dataset 3:** N = 9,923\n    * ...and so on.\n* Instead of asking participants to guess these different 'N' values, we uniformly ask everyone to submit a **top 50,000 ranked list.** This is just a simple way to ensure one has submitted *enough* sequences to cover the largest possible 'N'.\n\n**How we score a submission:**\n\nWhen we calculate the Jaccard score for Dataset 1, our scoring script knows the true size is **N = 40,943**.\n1.  It takes **only the top 40,943** ranked sequences from the participant's submission file.\n2.  It treats those 40,943 sequences as a **set** (the order *within* this set no longer matters).\n3.  It compares participants' predicted set to our 'ground truth set' (which also has 40,943 items) and computes the Jaccard similarity.\n\n**In summary:** The *only* purpose of ranking is to determine which sequences land in participants' 'Top-N' set for each dataset. A better model will naturally place more 'true' sequences at the top of its list, creating a 'Top-N' set that overlaps more with the ground truth and resulting in a higher Jaccard score.\n\nHope that makes it clear!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3320108,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "11/12/2025 08:43:41",
      "content": "<p>Hi! I am not sure if I fully follow your question 😀. Could you please elaborate more?</p>\n<p>Since there will be many more rows from task-2 (almost 100 times more) than from task-1 in the final submission file, how will these in the end contribute towards the leaderboard? (Here referring to task-1 and task-2 from the image shown under \"Evaluation\" in the Overview page). Is that your question?</p>\n<p>Or were you wondering about how important it is to get the ordering (ranking) correct and whether that affects the leaderboard score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3320149,
          "author_name": "parejepareje",
          "author_url": "",
          "post_date": "11/12/2025 09:19:22",
          "content": "<p>Apologies, I should have framed that more clearly. As you noted, having “100 times more” rows from task 2 could heavily bias the results toward that task. In addition, for the T1D data, there is no established “truth” in your ranking, especially if one of the competitors outperforms your benchmark. Continuing with that example, if someone does outperform your benchmark on the T1D classification task, their classification performance would justifiably improve their overall score. However, their ranking, which must differ from yours, would then contribute negatively to their ranking score, which seems counterproductive.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3320238,
              "author_name": "ckanduri",
              "author_url": "",
              "post_date": "11/12/2025 10:52:41",
              "content": "<p>Good questions, and a very understandable dilemma. As we tried to show in the illustration on the overview page, we compute ROC AUC for each test dataset and Jaccard similarity for each training dataset. It is at this stage, we take a weighted average of these performance metrics to compute a leaderboard score. In other words, the leaderboard score is not heavily biased towards task-2. We asked for a ranked list only because each dataset may have different sizes of true immune signals, e.g., one dataset may have only 100 true immune signals, whereas another dataset may have 50,000. We take care of extracting the right number of predicted immune signals per each training dataset from the submission file. hope that clarifies 😀  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14003908%2F13d3d8ca174bb1cb8efd2bb9df671813%2Fevaluation_illustration_for_kaggle_page.png?generation=1756380205170091&amp;alt=media\" alt=\"\">, </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3320853,
                  "author_name": "parejepareje",
                  "author_url": "",
                  "post_date": "11/12/2025 21:06:42",
                  "content": "<p>Much appreciated. I would be grateful if you could please clarify what is meant by the following: “We take care of extracting the right number of predicted immune signals per each training dataset from the submission file.”</p>\n<p>Also, I believe my earlier question about the potential contradiction between success in sample prediction and success in sequence ranking may not have been fully addressed. For instance, if your own benchmark for T1D achieves 0.9 on the test set, and another model achieves 0.91, the corresponding sequence rankings would necessarily differ. Wouldn’t this mean that an improved model could be penalized for outperforming the benchmark?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3322134,
                      "author_name": "ckanduri",
                      "author_url": "",
                      "post_date": "11/13/2025 10:47:49",
                      "content": "<p>Great follow-up questions, as this will also help clarification to other participants 😀. </p>\n<p>Short answer is: <strong>Submissions will not be penalized for creating a better model.</strong> A submission's ranking is <strong>not</strong> compared to any benchmark's ranking.</p>\n<p>Below is a more elaborate answer.</p>\n<h3>The \"ground truth\" is NOT based on any benchmark</h3>\n<ul>\n<li>The 'ground truth' for the ranking task (Task 2) is <strong>not</strong> based on any benchmark or model's results.</li>\n<li>Instead, the organizers have a <strong>fixed, finite 'ground truth' set</strong> of rows that are <em>known</em> to harbor 'true immune signals.'</li>\n<li>Think of this as a \"bag\" of true signals, not a ranked list. This knowledge is not derived from any model, and there is no \"ranking\" within this true set.</li>\n</ul>\n<p>Therefore, an improved model (like your 0.91 ROC AUC example) that finds more of these 'true' rows and ranks them highly will simply get a <strong>better</strong> score. It is never compared to any benchmark.</p>\n<h3>Why we ask for the top 50,000 rows (the \"top-N\")</h3>\n<p>This answers your first question about \"extracting the right number.\"</p>\n<ul>\n<li>The <em>size</em> of that 'ground truth set' (let's call its size 'N') is <strong>different for each training dataset.</strong></li>\n<li>For example, the true size 'N' for each dataset might be:<ul>\n<li><strong>Dataset 1:</strong> N = 40,943</li>\n<li><strong>Dataset 2:</strong> N = 32,664</li>\n<li><strong>Dataset 3:</strong> N = 9,923</li>\n<li>…and so on.</li></ul></li>\n<li>Instead of asking participants to guess these different 'N' values, we uniformly ask everyone to submit a <strong>top 50,000 ranked list.</strong> This is just a simple way to ensure one has submitted <em>enough</em> sequences to cover the largest possible 'N'.</li>\n</ul>\n<p><strong>How we score a submission:</strong></p>\n<p>When we calculate the Jaccard score for Dataset 1, our scoring script knows the true size is <strong>N = 40,943</strong>.</p>\n<ol>\n<li>It takes <strong>only the top 40,943</strong> ranked sequences from the participant's submission file.</li>\n<li>It treats those 40,943 sequences as a <strong>set</strong> (the order <em>within</em> this set no longer matters).</li>\n<li>It compares participants' predicted set to our 'ground truth set' (which also has 40,943 items) and computes the Jaccard similarity.</li>\n</ol>\n<p><strong>In summary:</strong> The <em>only</em> purpose of ranking is to determine which sequences land in participants' 'Top-N' set for each dataset. A better model will naturally place more 'true' sequences at the top of its list, creating a 'Top-N' set that overlaps more with the ground truth and resulting in a higher Jaccard score.</p>\n<p>Hope that makes it clear!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3319057": "As we are being asked to rank top sequences and there are 50000 such rankings in each train dataset, it seems obvious that there will be no \"perfect\" matches. My question is, as there are 400,000 such rows in the submission file, what is the weight of the ranking compared to the weight of the classification?",
    "3320108": "Hi! I am not sure if I fully follow your question 😀. Could you please elaborate more?\n\nSince there will be many more rows from task-2 (almost 100 times more) than from task-1 in the final submission file, how will these in the end contribute towards the leaderboard? (Here referring to task-1 and task-2 from the image shown under \"Evaluation\" in the Overview page). Is that your question?\n\nOr were you wondering about how important it is to get the ordering (ranking) correct and whether that affects the leaderboard score?",
    "3320149": "Apologies, I should have framed that more clearly. As you noted, having “100 times more” rows from task 2 could heavily bias the results toward that task. In addition, for the T1D data, there is no established “truth” in your ranking, especially if one of the competitors outperforms your benchmark. Continuing with that example, if someone does outperform your benchmark on the T1D classification task, their classification performance would justifiably improve their overall score. However, their ranking, which must differ from yours, would then contribute negatively to their ranking score, which seems counterproductive.",
    "3320238": "Good questions, and a very understandable dilemma. As we tried to show in the illustration on the overview page, we compute ROC AUC for each test dataset and Jaccard similarity for each training dataset. It is at this stage, we take a weighted average of these performance metrics to compute a leaderboard score. In other words, the leaderboard score is not heavily biased towards task-2. We asked for a ranked list only because each dataset may have different sizes of true immune signals, e.g., one dataset may have only 100 true immune signals, whereas another dataset may have 50,000. We take care of extracting the right number of predicted immune signals per each training dataset from the submission file. hope that clarifies 😀  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14003908%2F13d3d8ca174bb1cb8efd2bb9df671813%2Fevaluation_illustration_for_kaggle_page.png?generation=1756380205170091&alt=media),",
    "3320853": "Much appreciated. I would be grateful if you could please clarify what is meant by the following: “We take care of extracting the right number of predicted immune signals per each training dataset from the submission file.”\n\nAlso, I believe my earlier question about the potential contradiction between success in sample prediction and success in sequence ranking may not have been fully addressed. For instance, if your own benchmark for T1D achieves 0.9 on the test set, and another model achieves 0.91, the corresponding sequence rankings would necessarily differ. Wouldn’t this mean that an improved model could be penalized for outperforming the benchmark?",
    "3322134": "Great follow-up questions, as this will also help clarification to other participants 😀. \n\nShort answer is: **Submissions will not be penalized for creating a better model.** A submission's ranking is **not** compared to any benchmark's ranking.\n\nBelow is a more elaborate answer.\n\n### The \"ground truth\" is NOT based on any benchmark\n\n* The 'ground truth' for the ranking task (Task 2) is **not** based on any benchmark or model's results.\n* Instead, the organizers have a **fixed, finite 'ground truth' set** of rows that are *known* to harbor 'true immune signals.'\n* Think of this as a \"bag\" of true signals, not a ranked list. This knowledge is not derived from any model, and there is no \"ranking\" within this true set.\n\nTherefore, an improved model (like your 0.91 ROC AUC example) that finds more of these 'true' rows and ranks them highly will simply get a **better** score. It is never compared to any benchmark.\n\n### Why we ask for the top 50,000 rows (the \"top-N\")\n\nThis answers your first question about \"extracting the right number.\"\n\n* The *size* of that 'ground truth set' (let's call its size 'N') is **different for each training dataset.**\n* For example, the true size 'N' for each dataset might be:\n    * **Dataset 1:** N = 40,943\n    * **Dataset 2:** N = 32,664\n    * **Dataset 3:** N = 9,923\n    * ...and so on.\n* Instead of asking participants to guess these different 'N' values, we uniformly ask everyone to submit a **top 50,000 ranked list.** This is just a simple way to ensure one has submitted *enough* sequences to cover the largest possible 'N'.\n\n**How we score a submission:**\n\nWhen we calculate the Jaccard score for Dataset 1, our scoring script knows the true size is **N = 40,943**.\n1.  It takes **only the top 40,943** ranked sequences from the participant's submission file.\n2.  It treats those 40,943 sequences as a **set** (the order *within* this set no longer matters).\n3.  It compares participants' predicted set to our 'ground truth set' (which also has 40,943 items) and computes the Jaccard similarity.\n\n**In summary:** The *only* purpose of ranking is to determine which sequences land in participants' 'Top-N' set for each dataset. A better model will naturally place more 'true' sequences at the top of its list, creating a 'Top-N' set that overlaps more with the ground truth and resulting in a higher Jaccard score.\n\nHope that makes it clear!"
  },
  "source": "meta"
}