{
  "id": 613060,
  "title": "Competition Evaluation - Clarification",
  "url": "/competitions/acm-icaif-25-ai-agentic-retrieval-grand-challenge/discussion/613060",
  "author_name": "DanT",
  "post_date": "2025-10-24T00:39:44.716000",
  "votes": 0,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Dear organizers, thank you for your significant effort in running this challenge. Now that submissions have closed for 3 days, I have a few questions about the evaluation process.</p>\n<hr>\n<p><strong>1. Evaluation Timeline</strong></p>\n<p>What is the timeline for: (a) private leaderboard release, and (b) notebook submission deadline?</p>\n<p>A specified deadline (e.g., \"notebooks due 3 days after submission close\") would ensure fairness and compliance, consistent with standard Kaggle and ML competition practices.</p>\n<p><strong>2. Evaluation Components</strong></p>\n<p>How do different components map to final standing: private leaderboard rank, reproducibility, methodology explanation?</p>\n<p>Official statements emphasize that final outcomes depend on methodology justification and reproducible code, not solely on leaderboard scores.\n\"Every submission must include a working notebook\": there is no other notebook disclosed by now. </p>\n<p>Without reproduction check, public leaderboard does not reflect reality: not only it's based only on 30% of eval data (small sample), it's not the correct metric per the rules (only MAP@5 vs. average of 3 scores).</p>\n<p>We're curious to see the winning solutions and learn from them.</p>\n<hr>\n<p>Thank you again for organizing the competition. I think all participants have enjoyed the challenge would appreciate some clarification. Of course I might misunderstand some rules, I'll be happy to hear your thoughts.</p>\n<hr>\n<h2>Side-note: why reproducible notebook is the standard of evaluation</h2>\n<p>Just in case some participants are curious about this.</p>\n<p>Fact: most LLMs competitions, including the other ones currently held at ICAIF'25 (e.g. <a href=\"https://finddr2025.github.io/\" target=\"_blank\">finddr2025</a>) have same standards for validating submissions (reproducible code + explanation)</p>\n<p>Reason: easy ways to overfit both public and private leaderboard, as the eval-sets (the queries) are known at training time. </p>\n<p>Implication: accurate evaluation framework = re-run code on eval set (or a hidden one). This is why most LLMs competitions are 'code' competition.</p>\n<p>So the rules of this competition are in fact well designed and conform to best practices.</p>\n<p>Happy to hear your thoughts on this if anyone has any question or idea.</p>\n<p>All the bests.</p>",
  "messages": [
    {
      "id": 3306634,
      "postDate": "2025-10-25T00:44:24.517Z",
      "content": "<p><strong>What is the timeline for: (a) private leaderboard release, and (b) notebook submission deadline?</strong></p>\n<p>We plan to release the private leaderboard this weekend.<br>\nParticipants were expected to submit their notebooks by the official competition deadline.<br>\nHowever, since only few teams have uploaded their notebooks so far and the conference will be held on November 15, we have limited time for the review and presentation preparation process.</p>\n<p>In this context, it was necessary to first contact the top-performing teams and those who submitted their notebooks early.<br>\nWe will strongly encourage all invited teams to share their code and methodological details afterward to ensure the reproducibility and integrity of their work.</p>\n<p><strong>How do different components map to final standing: private leaderboard rank, reproducibility, methodology explanation?</strong></p>\n<p>Due to the time constraints, we kindly ask for your understanding that the evaluation of reproducibility and methodology will be primarily based on the materials shared and the explanations provided during the invited presentations and Q&amp;A sessions.</p>\n<p>In summary, the final standings will reflect:</p>\n<ul>\n<li>Private Leaderboard Rank -- main performance indicator</li>\n<li>Reproducibility &amp; Methodology -- assessed through shared materials and presentation discussions</li>\n</ul>\n<p>We appreciate the community’s understanding and cooperation in this process.</p>",
      "rawMarkdown": "**What is the timeline for: (a) private leaderboard release, and (b) notebook submission deadline?**\n\nWe plan to release the private leaderboard this weekend.\nParticipants were expected to submit their notebooks by the official competition deadline.\nHowever, since only few teams have uploaded their notebooks so far and the conference will be held on November 15, we have limited time for the review and presentation preparation process.\n\nIn this context, it was necessary to first contact the top-performing teams and those who submitted their notebooks early.\nWe will strongly encourage all invited teams to share their code and methodological details afterward to ensure the reproducibility and integrity of their work.\n\n**How do different components map to final standing: private leaderboard rank, reproducibility, methodology explanation?**\n\nDue to the time constraints, we kindly ask for your understanding that the evaluation of reproducibility and methodology will be primarily based on the materials shared and the explanations provided during the invited presentations and Q&A sessions.\n\nIn summary, the final standings will reflect:\n- Private Leaderboard Rank -- main performance indicator\n- Reproducibility & Methodology -- assessed through shared materials and presentation discussions\n\nWe appreciate the community’s understanding and cooperation in this process.",
      "replies": [
        {
          "id": 3307945,
          "postDate": "2025-10-28T07:16:11.313Z",
          "content": "<p>Thank you for clarifying. I appreciate your transparency. </p>\n<p>Although there is a significant shift in evaluation method, I believe that you got to make a difficult decision given the circumstance.</p>\n<p>Best wishes for future challenges.</p>",
          "rawMarkdown": "Thank you for clarifying. I appreciate your transparency. \n\nAlthough there is a significant shift in evaluation method, I believe that you got to make a difficult decision given the circumstance.\n\nBest wishes for future challenges.\n\n",
          "replies": [
            {
              "id": 3307947,
              "postDate": "2025-10-28T07:20:27.170Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3307948,
          "postDate": "2025-10-28T07:21:48.630Z",
          "content": "<p>P/S: I'll send you an email for administrative reason</p>",
          "rawMarkdown": "P/S: I'll send you an email for administrative reason"
        }
      ]
    },
    {
      "id": 3306074,
      "postDate": "2025-10-24T00:39:44.717Z",
      "content": "<p>Dear organizers, thank you for your significant effort in running this challenge. Now that submissions have closed for 3 days, I have a few questions about the evaluation process.</p>\n<hr>\n<p><strong>1. Evaluation Timeline</strong></p>\n<p>What is the timeline for: (a) private leaderboard release, and (b) notebook submission deadline?</p>\n<p>A specified deadline (e.g., \"notebooks due 3 days after submission close\") would ensure fairness and compliance, consistent with standard Kaggle and ML competition practices.</p>\n<p><strong>2. Evaluation Components</strong></p>\n<p>How do different components map to final standing: private leaderboard rank, reproducibility, methodology explanation?</p>\n<p>Official statements emphasize that final outcomes depend on methodology justification and reproducible code, not solely on leaderboard scores.\n\"Every submission must include a working notebook\": there is no other notebook disclosed by now. </p>\n<p>Without reproduction check, public leaderboard does not reflect reality: not only it's based only on 30% of eval data (small sample), it's not the correct metric per the rules (only MAP@5 vs. average of 3 scores).</p>\n<p>We're curious to see the winning solutions and learn from them.</p>\n<hr>\n<p>Thank you again for organizing the competition. I think all participants have enjoyed the challenge would appreciate some clarification. Of course I might misunderstand some rules, I'll be happy to hear your thoughts.</p>\n<hr>\n<h2>Side-note: why reproducible notebook is the standard of evaluation</h2>\n<p>Just in case some participants are curious about this.</p>\n<p>Fact: most LLMs competitions, including the other ones currently held at ICAIF'25 (e.g. <a href=\"https://finddr2025.github.io/\" target=\"_blank\">finddr2025</a>) have same standards for validating submissions (reproducible code + explanation)</p>\n<p>Reason: easy ways to overfit both public and private leaderboard, as the eval-sets (the queries) are known at training time. </p>\n<p>Implication: accurate evaluation framework = re-run code on eval set (or a hidden one). This is why most LLMs competitions are 'code' competition.</p>\n<p>So the rules of this competition are in fact well designed and conform to best practices.</p>\n<p>Happy to hear your thoughts on this if anyone has any question or idea.</p>\n<p>All the bests.</p>",
      "rawMarkdown": "Dear organizers, thank you for your significant effort in running this challenge. Now that submissions have closed for 3 days, I have a few questions about the evaluation process.\n\n---\n\n**1. Evaluation Timeline**\n\nWhat is the timeline for: (a) private leaderboard release, and (b) notebook submission deadline?\n\nA specified deadline (e.g., \"notebooks due 3 days after submission close\") would ensure fairness and compliance, consistent with standard Kaggle and ML competition practices.\n\n**2. Evaluation Components**\n\nHow do different components map to final standing: private leaderboard rank, reproducibility, methodology explanation?\n\nOfficial statements emphasize that final outcomes depend on methodology justification and reproducible code, not solely on leaderboard scores.\n\"Every submission must include a working notebook\": there is no other notebook disclosed by now. \n\nWithout reproduction check, public leaderboard does not reflect reality: not only it's based only on 30% of eval data (small sample), it's not the correct metric per the rules (only MAP@5 vs. average of 3 scores).\n\nWe're curious to see the winning solutions and learn from them.\n\n---\n\nThank you again for organizing the competition. I think all participants have enjoyed the challenge would appreciate some clarification. Of course I might misunderstand some rules, I'll be happy to hear your thoughts.\n\n---\n\n## Side-note: why reproducible notebook is the standard of evaluation\n\nJust in case some participants are curious about this.\n\nFact: most LLMs competitions, including the other ones currently held at ICAIF'25 (e.g. [finddr2025](https://finddr2025.github.io/)) have same standards for validating submissions (reproducible code + explanation)\n\nReason: easy ways to overfit both public and private leaderboard, as the eval-sets (the queries) are known at training time. \n\nImplication: accurate evaluation framework = re-run code on eval set (or a hidden one). This is why most LLMs competitions are 'code' competition.\n\nSo the rules of this competition are in fact well designed and conform to best practices.\n\nHappy to hear your thoughts on this if anyone has any question or idea.\n\nAll the bests.\n"
    },
    {
      "id": 3307946,
      "postDate": "2025-10-28T07:19:26.147Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3306634,
      "author_name": "JIHOONKWON",
      "author_url": "",
      "post_date": "2025-10-25T00:44:24.517000",
      "content": "<p><strong>What is the timeline for: (a) private leaderboard release, and (b) notebook submission deadline?</strong></p>\n<p>We plan to release the private leaderboard this weekend.<br>\nParticipants were expected to submit their notebooks by the official competition deadline.<br>\nHowever, since only few teams have uploaded their notebooks so far and the conference will be held on November 15, we have limited time for the review and presentation preparation process.</p>\n<p>In this context, it was necessary to first contact the top-performing teams and those who submitted their notebooks early.<br>\nWe will strongly encourage all invited teams to share their code and methodological details afterward to ensure the reproducibility and integrity of their work.</p>\n<p><strong>How do different components map to final standing: private leaderboard rank, reproducibility, methodology explanation?</strong></p>\n<p>Due to the time constraints, we kindly ask for your understanding that the evaluation of reproducibility and methodology will be primarily based on the materials shared and the explanations provided during the invited presentations and Q&amp;A sessions.</p>\n<p>In summary, the final standings will reflect:</p>\n<ul>\n<li>Private Leaderboard Rank -- main performance indicator</li>\n<li>Reproducibility &amp; Methodology -- assessed through shared materials and presentation discussions</li>\n</ul>\n<p>We appreciate the community’s understanding and cooperation in this process.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3307945,
          "author_name": "DanT",
          "author_url": "",
          "post_date": "2025-10-28T07:16:11.313000",
          "content": "<p>Thank you for clarifying. I appreciate your transparency. </p>\n<p>Although there is a significant shift in evaluation method, I believe that you got to make a difficult decision given the circumstance.</p>\n<p>Best wishes for future challenges.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3307947,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-10-28T07:20:27.170000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3307948,
          "author_name": "DanT",
          "author_url": "",
          "post_date": "2025-10-28T07:21:48.630000",
          "content": "<p>P/S: I'll send you an email for administrative reason</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3307946,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-10-28T07:19:26.147000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3306634": "**What is the timeline for: (a) private leaderboard release, and (b) notebook submission deadline?**\n\nWe plan to release the private leaderboard this weekend.\nParticipants were expected to submit their notebooks by the official competition deadline.\nHowever, since only few teams have uploaded their notebooks so far and the conference will be held on November 15, we have limited time for the review and presentation preparation process.\n\nIn this context, it was necessary to first contact the top-performing teams and those who submitted their notebooks early.\nWe will strongly encourage all invited teams to share their code and methodological details afterward to ensure the reproducibility and integrity of their work.\n\n**How do different components map to final standing: private leaderboard rank, reproducibility, methodology explanation?**\n\nDue to the time constraints, we kindly ask for your understanding that the evaluation of reproducibility and methodology will be primarily based on the materials shared and the explanations provided during the invited presentations and Q&A sessions.\n\nIn summary, the final standings will reflect:\n- Private Leaderboard Rank -- main performance indicator\n- Reproducibility & Methodology -- assessed through shared materials and presentation discussions\n\nWe appreciate the community’s understanding and cooperation in this process.",
    "3306074": "Dear organizers, thank you for your significant effort in running this challenge. Now that submissions have closed for 3 days, I have a few questions about the evaluation process.\n\n---\n\n**1. Evaluation Timeline**\n\nWhat is the timeline for: (a) private leaderboard release, and (b) notebook submission deadline?\n\nA specified deadline (e.g., \"notebooks due 3 days after submission close\") would ensure fairness and compliance, consistent with standard Kaggle and ML competition practices.\n\n**2. Evaluation Components**\n\nHow do different components map to final standing: private leaderboard rank, reproducibility, methodology explanation?\n\nOfficial statements emphasize that final outcomes depend on methodology justification and reproducible code, not solely on leaderboard scores.\n\"Every submission must include a working notebook\": there is no other notebook disclosed by now. \n\nWithout reproduction check, public leaderboard does not reflect reality: not only it's based only on 30% of eval data (small sample), it's not the correct metric per the rules (only MAP@5 vs. average of 3 scores).\n\nWe're curious to see the winning solutions and learn from them.\n\n---\n\nThank you again for organizing the competition. I think all participants have enjoyed the challenge would appreciate some clarification. Of course I might misunderstand some rules, I'll be happy to hear your thoughts.\n\n---\n\n## Side-note: why reproducible notebook is the standard of evaluation\n\nJust in case some participants are curious about this.\n\nFact: most LLMs competitions, including the other ones currently held at ICAIF'25 (e.g. [finddr2025](https://finddr2025.github.io/)) have same standards for validating submissions (reproducible code + explanation)\n\nReason: easy ways to overfit both public and private leaderboard, as the eval-sets (the queries) are known at training time. \n\nImplication: accurate evaluation framework = re-run code on eval set (or a hidden one). This is why most LLMs competitions are 'code' competition.\n\nSo the rules of this competition are in fact well designed and conform to best practices.\n\nHappy to hear your thoughts on this if anyone has any question or idea.\n\nAll the bests.\n",
    "3307946": ""
  }
}