{
  "id": 651754,
  "title": "How will test data run with our code",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/651754",
  "author_name": "",
  "post_date": "2025-12-05T04:45:52.713932900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi Organizer,\n     Thanks a lot for your effort on hosting this challenge. I have a question that how will our code run on the test data on final leaderboard? Since you mention that only 1% test data is utilized in continuous leaderboard now. Does that mean our final submission should be our code and feature importance scores for the 8 training datasets? \n      But the strange thing is that in your given report file about this challenge, you show how final leaderboard calculated weights of testing data. From that table, our current testing data is much more than 1% among all, which is quite confusing. \n      If you will test our submitted code on your own extra data, do we need to consider the memory, time that our model needs to run the extra 99% test data? Thanks!</p>\n<p>Best,</p>",
  "messages": [
    {
      "id": "3362649",
      "postDate": "12/05/2025 04:45:52",
      "content": "<p>Hi Organizer,\n     Thanks a lot for your effort on hosting this challenge. I have a question that how will our code run on the test data on final leaderboard? Since you mention that only 1% test data is utilized in continuous leaderboard now. Does that mean our final submission should be our code and feature importance scores for the 8 training datasets? \n      But the strange thing is that in your given report file about this challenge, you show how final leaderboard calculated weights of testing data. From that table, our current testing data is much more than 1% among all, which is quite confusing. \n      If you will test our submitted code on your own extra data, do we need to consider the memory, time that our model needs to run the extra 99% test data? Thanks!</p>\n<p>Best,</p>",
      "rawMarkdown": "Hi Organizer,\n     Thanks a lot for your effort on hosting this challenge. I have a question that how will our code run on the test data on final leaderboard? Since you mention that only 1% test data is utilized in continuous leaderboard now. Does that mean our final submission should be our code and feature importance scores for the 8 training datasets? \n      But the strange thing is that in your given report file about this challenge, you show how final leaderboard calculated weights of testing data. From that table, our current testing data is much more than 1% among all, which is quite confusing. \n      If you will test our submitted code on your own extra data, do we need to consider the memory, time that our model needs to run the extra 99% test data? Thanks!\n\n\nBest,",
      "votes": null
    },
    {
      "id": "3362698",
      "postDate": "12/05/2025 06:26:00",
      "content": "<p>Hi! </p>\n<blockquote>\n  <p>I have a question that how will our code run on the test data on final leaderboard? Since you mention that only 1% test data is utilized in continuous leaderboard now. Does that mean our final submission should be our code and feature importance scores for the 8 training datasets?</p>\n</blockquote>\n<p>Every participant's submissions are continuously evaluated for both Public and Private leaderboards using the <code>submission.csv</code> they submit. The Private leaderboard is hidden for the participants until the challenge concludes. The final submission is still the <code>submission.csv</code>, although open-source code is a pre-requisite to win the prize money and/or be invited to co-author a scientific manuscript describing the challenge outcome. We will contact the top-10 participants after the Private leaderboard is revealed regarding making their code open-source. </p>\n<blockquote>\n  <p>But the strange thing is that in your given report file about this challenge, you show how final leaderboard calculated weights of testing data. From that table, our current testing data is much more than 1% among all, which is quite confusing. If you will test our submitted code on your own extra data, do we need to consider the memory, time that our model needs to run the extra 99% test data? </p>\n</blockquote>\n<p>You are correct in your observation that public leaderboard is much more than 1%. This <code>1%</code> statement on the leaderboard is automatically generated by Kaggle based on number of rows, where rows from Task-2 outnumber rows from Task-1. This was also discussed in this other <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/640452\" target=\"_blank\">thread</a>, and this <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408\" target=\"_blank\">thread</a> is also relevant. Having said that, we will indeed test the top models on a large benchmarking suite for the scientific manuscript after the conclusion of this challenge, in collaboration with the top 10 teams, where it is great if participants consider memory and compute requirements of their code -- but that will not affect the prize money and winning of this Kaggle challenge itself. Hope that clarified your questions 😀?</p>",
      "rawMarkdown": "Hi! \n\n>I have a question that how will our code run on the test data on final leaderboard? Since you mention that only 1% test data is utilized in continuous leaderboard now. Does that mean our final submission should be our code and feature importance scores for the 8 training datasets?\n\nEvery participant's submissions are continuously evaluated for both Public and Private leaderboards using the `submission.csv` they submit. The Private leaderboard is hidden for the participants until the challenge concludes. The final submission is still the `submission.csv`, although open-source code is a pre-requisite to win the prize money and/or be invited to co-author a scientific manuscript describing the challenge outcome. We will contact the top-10 participants after the Private leaderboard is revealed regarding making their code open-source. \n\n>But the strange thing is that in your given report file about this challenge, you show how final leaderboard calculated weights of testing data. From that table, our current testing data is much more than 1% among all, which is quite confusing. If you will test our submitted code on your own extra data, do we need to consider the memory, time that our model needs to run the extra 99% test data? \n\nYou are correct in your observation that public leaderboard is much more than 1%. This `1%` statement on the leaderboard is automatically generated by Kaggle based on number of rows, where rows from Task-2 outnumber rows from Task-1. This was also discussed in this other [thread](https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/640452), and this [thread](https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408) is also relevant. Having said that, we will indeed test the top models on a large benchmarking suite for the scientific manuscript after the conclusion of this challenge, in collaboration with the top 10 teams, where it is great if participants consider memory and compute requirements of their code -- but that will not affect the prize money and winning of this Kaggle challenge itself. Hope that clarified your questions 😀?",
      "votes": null
    },
    {
      "id": "3362898",
      "postDate": "12/05/2025 14:35:44",
      "content": "<p>Yes, it clarifies. Thanks a lot for your response!</p>",
      "rawMarkdown": "Yes, it clarifies. Thanks a lot for your response!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3362698,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "12/05/2025 06:26:00",
      "content": "<p>Hi! </p>\n<blockquote>\n  <p>I have a question that how will our code run on the test data on final leaderboard? Since you mention that only 1% test data is utilized in continuous leaderboard now. Does that mean our final submission should be our code and feature importance scores for the 8 training datasets?</p>\n</blockquote>\n<p>Every participant's submissions are continuously evaluated for both Public and Private leaderboards using the <code>submission.csv</code> they submit. The Private leaderboard is hidden for the participants until the challenge concludes. The final submission is still the <code>submission.csv</code>, although open-source code is a pre-requisite to win the prize money and/or be invited to co-author a scientific manuscript describing the challenge outcome. We will contact the top-10 participants after the Private leaderboard is revealed regarding making their code open-source. </p>\n<blockquote>\n  <p>But the strange thing is that in your given report file about this challenge, you show how final leaderboard calculated weights of testing data. From that table, our current testing data is much more than 1% among all, which is quite confusing. If you will test our submitted code on your own extra data, do we need to consider the memory, time that our model needs to run the extra 99% test data? </p>\n</blockquote>\n<p>You are correct in your observation that public leaderboard is much more than 1%. This <code>1%</code> statement on the leaderboard is automatically generated by Kaggle based on number of rows, where rows from Task-2 outnumber rows from Task-1. This was also discussed in this other <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/640452\" target=\"_blank\">thread</a>, and this <a href=\"https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408\" target=\"_blank\">thread</a> is also relevant. Having said that, we will indeed test the top models on a large benchmarking suite for the scientific manuscript after the conclusion of this challenge, in collaboration with the top 10 teams, where it is great if participants consider memory and compute requirements of their code -- but that will not affect the prize money and winning of this Kaggle challenge itself. Hope that clarified your questions 😀?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3362898,
          "author_name": "xinluoumich",
          "author_url": "",
          "post_date": "12/05/2025 14:35:44",
          "content": "<p>Yes, it clarifies. Thanks a lot for your response!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3362649": "Hi Organizer,\n     Thanks a lot for your effort on hosting this challenge. I have a question that how will our code run on the test data on final leaderboard? Since you mention that only 1% test data is utilized in continuous leaderboard now. Does that mean our final submission should be our code and feature importance scores for the 8 training datasets? \n      But the strange thing is that in your given report file about this challenge, you show how final leaderboard calculated weights of testing data. From that table, our current testing data is much more than 1% among all, which is quite confusing. \n      If you will test our submitted code on your own extra data, do we need to consider the memory, time that our model needs to run the extra 99% test data? Thanks!\n\n\nBest,",
    "3362698": "Hi! \n\n>I have a question that how will our code run on the test data on final leaderboard? Since you mention that only 1% test data is utilized in continuous leaderboard now. Does that mean our final submission should be our code and feature importance scores for the 8 training datasets?\n\nEvery participant's submissions are continuously evaluated for both Public and Private leaderboards using the `submission.csv` they submit. The Private leaderboard is hidden for the participants until the challenge concludes. The final submission is still the `submission.csv`, although open-source code is a pre-requisite to win the prize money and/or be invited to co-author a scientific manuscript describing the challenge outcome. We will contact the top-10 participants after the Private leaderboard is revealed regarding making their code open-source. \n\n>But the strange thing is that in your given report file about this challenge, you show how final leaderboard calculated weights of testing data. From that table, our current testing data is much more than 1% among all, which is quite confusing. If you will test our submitted code on your own extra data, do we need to consider the memory, time that our model needs to run the extra 99% test data? \n\nYou are correct in your observation that public leaderboard is much more than 1%. This `1%` statement on the leaderboard is automatically generated by Kaggle based on number of rows, where rows from Task-2 outnumber rows from Task-1. This was also discussed in this other [thread](https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/640452), and this [thread](https://www.kaggle.com/competitions/adaptive-immune-profiling-challenge-2025/discussion/630408) is also relevant. Having said that, we will indeed test the top models on a large benchmarking suite for the scientific manuscript after the conclusion of this challenge, in collaboration with the top 10 teams, where it is great if participants consider memory and compute requirements of their code -- but that will not affect the prize money and winning of this Kaggle challenge itself. Hope that clarified your questions 😀?",
    "3362898": "Yes, it clarifies. Thanks a lot for your response!"
  },
  "source": "meta"
}