{
  "id": 607246,
  "title": "Organizer Clarification: External Training Data, Model Weights in DBFE, and Off-Platform Training",
  "url": "/competitions/acm-icaif-25-ai-agentic-retrieval-grand-challenge/discussion/607246",
  "author_name": "",
  "post_date": "2025-09-12T20:23:56.525937400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi organizers — thanks for hosting this challenge! I’d like to confirm three rule details so we can keep our submissions compliant and reproducible:</p>\n<p>1) <strong>External public data for training</strong>  <br>\n   May we use additional public-domain SEC/EDGAR filings <strong>for training only</strong> open-source LLMs, with the understanding that <strong>evaluation/ranking is strictly over the competition’s provided corpus</strong>? If allowed, we’ll disclose sources, snapshot dates, and licenses.</p>\n<p>2) <strong>Fine-tuned model weights inside Databricks Free Edition (DBFE)</strong>  <br>\n   Are we permitted to <strong>upload fine-tuned open-source model weights</strong> (e.g., a ZIP/artifact) into DBFE storage (e.g., a volume/DBFS) and load them from the <strong>reproducible DBFE notebook</strong>? If so, are there any <strong>file-size limits</strong> or preferred storage locations we should follow?</p>\n<p>3) <strong>Training off-platform</strong>  <br>\n   Is it acceptable to <strong>train models on our own cloud/GPU</strong> (outside DBFE) and perform <strong>inference + reproduction in the DBFE notebook</strong> only?</p>\n<p>(We will not expand the evaluation candidate set or call external sources at evaluation time; any external data would be used solely for training if permitted.)</p>\n<p>Thanks in advance for confirming!</p>",
  "messages": [
    {
      "id": "3288012",
      "postDate": "09/12/2025 20:23:56",
      "content": "<p>Hi organizers — thanks for hosting this challenge! I’d like to confirm three rule details so we can keep our submissions compliant and reproducible:</p>\n<p>1) <strong>External public data for training</strong>  <br>\n   May we use additional public-domain SEC/EDGAR filings <strong>for training only</strong> open-source LLMs, with the understanding that <strong>evaluation/ranking is strictly over the competition’s provided corpus</strong>? If allowed, we’ll disclose sources, snapshot dates, and licenses.</p>\n<p>2) <strong>Fine-tuned model weights inside Databricks Free Edition (DBFE)</strong>  <br>\n   Are we permitted to <strong>upload fine-tuned open-source model weights</strong> (e.g., a ZIP/artifact) into DBFE storage (e.g., a volume/DBFS) and load them from the <strong>reproducible DBFE notebook</strong>? If so, are there any <strong>file-size limits</strong> or preferred storage locations we should follow?</p>\n<p>3) <strong>Training off-platform</strong>  <br>\n   Is it acceptable to <strong>train models on our own cloud/GPU</strong> (outside DBFE) and perform <strong>inference + reproduction in the DBFE notebook</strong> only?</p>\n<p>(We will not expand the evaluation candidate set or call external sources at evaluation time; any external data would be used solely for training if permitted.)</p>\n<p>Thanks in advance for confirming!</p>",
      "rawMarkdown": "Hi organizers — thanks for hosting this challenge! I’d like to confirm three rule details so we can keep our submissions compliant and reproducible:\n\n1) **External public data for training**  \n   May we use additional public-domain SEC/EDGAR filings **for training only** open-source LLMs, with the understanding that **evaluation/ranking is strictly over the competition’s provided corpus**? If allowed, we’ll disclose sources, snapshot dates, and licenses.\n\n2) **Fine-tuned model weights inside Databricks Free Edition (DBFE)**  \n   Are we permitted to **upload fine-tuned open-source model weights** (e.g., a ZIP/artifact) into DBFE storage (e.g., a volume/DBFS) and load them from the **reproducible DBFE notebook**? If so, are there any **file-size limits** or preferred storage locations we should follow?\n\n3) **Training off-platform**  \n   Is it acceptable to **train models on our own cloud/GPU** (outside DBFE) and perform **inference + reproduction in the DBFE notebook** only?\n\n(We will not expand the evaluation candidate set or call external sources at evaluation time; any external data would be used solely for training if permitted.)\n\nThanks in advance for confirming!",
      "votes": null
    },
    {
      "id": "3288126",
      "postDate": "09/13/2025 07:06:49",
      "content": "<p>Thank you for your questions and interest in the challenge! Please see our clarifications below:</p>\n<h3>1.  External Public Data for Training</h3>\n<p>Yes, you may use additional public-domain SEC/EDGAR filings (or similar open public data) for training open-source LLMs, provided that evaluation and ranking remain strictly over the competition’s provided corpus. Please disclose all external sources, snapshot dates, and licenses in your submission on <strong><code>Discussion</code></strong> for transparency and reproducibility.</p>\n<h3>2.  Fine-Tuned Model Weights in Databricks Free Edition (DBFE)</h3>\n<p>Yes, you may upload fine-tuned open-source model weights into DBFE storage (e.g., DBFS volumes). Alternatively, hosting the model on Hugging Face and sharing the link for loading inside the DBFE notebook is also perfectly acceptable and often more convenient. Please follow standard DBFE file-size and storage guidelines.</p>\n<h3>3.  Training Off-Platform</h3>\n<p>Yes, it is acceptable to train models on your own cloud/GPU infrastructure and use the DBFE notebook only for inference and reproducibility, as long as evaluation and submission strictly follow the competition rules and no external sources are accessed at evaluation time.</p>\n<p>We hope this helps, and we look forward to your submissions!</p>",
      "rawMarkdown": "Thank you for your questions and interest in the challenge! Please see our clarifications below:\n###\t1.\tExternal Public Data for Training\nYes, you may use additional public-domain SEC/EDGAR filings (or similar open public data) for training open-source LLMs, provided that evaluation and ranking remain strictly over the competition’s provided corpus. Please disclose all external sources, snapshot dates, and licenses in your submission on **`Discussion`** for transparency and reproducibility.\n###\t2.\tFine-Tuned Model Weights in Databricks Free Edition (DBFE)\nYes, you may upload fine-tuned open-source model weights into DBFE storage (e.g., DBFS volumes). Alternatively, hosting the model on Hugging Face and sharing the link for loading inside the DBFE notebook is also perfectly acceptable and often more convenient. Please follow standard DBFE file-size and storage guidelines.\n###\t3.\tTraining Off-Platform\nYes, it is acceptable to train models on your own cloud/GPU infrastructure and use the DBFE notebook only for inference and reproducibility, as long as evaluation and submission strictly follow the competition rules and no external sources are accessed at evaluation time.\n\nWe hope this helps, and we look forward to your submissions!",
      "votes": null
    },
    {
      "id": "3289441",
      "postDate": "09/16/2025 04:08:15",
      "content": "<p>What LLM models are allowed in this challenge?</p>",
      "rawMarkdown": "What LLM models are allowed in this challenge?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3288126,
      "author_name": "jihoonkwon",
      "author_url": "",
      "post_date": "09/13/2025 07:06:49",
      "content": "<p>Thank you for your questions and interest in the challenge! Please see our clarifications below:</p>\n<h3>1.  External Public Data for Training</h3>\n<p>Yes, you may use additional public-domain SEC/EDGAR filings (or similar open public data) for training open-source LLMs, provided that evaluation and ranking remain strictly over the competition’s provided corpus. Please disclose all external sources, snapshot dates, and licenses in your submission on <strong><code>Discussion</code></strong> for transparency and reproducibility.</p>\n<h3>2.  Fine-Tuned Model Weights in Databricks Free Edition (DBFE)</h3>\n<p>Yes, you may upload fine-tuned open-source model weights into DBFE storage (e.g., DBFS volumes). Alternatively, hosting the model on Hugging Face and sharing the link for loading inside the DBFE notebook is also perfectly acceptable and often more convenient. Please follow standard DBFE file-size and storage guidelines.</p>\n<h3>3.  Training Off-Platform</h3>\n<p>Yes, it is acceptable to train models on your own cloud/GPU infrastructure and use the DBFE notebook only for inference and reproducibility, as long as evaluation and submission strictly follow the competition rules and no external sources are accessed at evaluation time.</p>\n<p>We hope this helps, and we look forward to your submissions!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3289441,
      "author_name": "yimingqian1",
      "author_url": "",
      "post_date": "09/16/2025 04:08:15",
      "content": "<p>What LLM models are allowed in this challenge?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3288012": "Hi organizers — thanks for hosting this challenge! I’d like to confirm three rule details so we can keep our submissions compliant and reproducible:\n\n1) **External public data for training**  \n   May we use additional public-domain SEC/EDGAR filings **for training only** open-source LLMs, with the understanding that **evaluation/ranking is strictly over the competition’s provided corpus**? If allowed, we’ll disclose sources, snapshot dates, and licenses.\n\n2) **Fine-tuned model weights inside Databricks Free Edition (DBFE)**  \n   Are we permitted to **upload fine-tuned open-source model weights** (e.g., a ZIP/artifact) into DBFE storage (e.g., a volume/DBFS) and load them from the **reproducible DBFE notebook**? If so, are there any **file-size limits** or preferred storage locations we should follow?\n\n3) **Training off-platform**  \n   Is it acceptable to **train models on our own cloud/GPU** (outside DBFE) and perform **inference + reproduction in the DBFE notebook** only?\n\n(We will not expand the evaluation candidate set or call external sources at evaluation time; any external data would be used solely for training if permitted.)\n\nThanks in advance for confirming!",
    "3288126": "Thank you for your questions and interest in the challenge! Please see our clarifications below:\n###\t1.\tExternal Public Data for Training\nYes, you may use additional public-domain SEC/EDGAR filings (or similar open public data) for training open-source LLMs, provided that evaluation and ranking remain strictly over the competition’s provided corpus. Please disclose all external sources, snapshot dates, and licenses in your submission on **`Discussion`** for transparency and reproducibility.\n###\t2.\tFine-Tuned Model Weights in Databricks Free Edition (DBFE)\nYes, you may upload fine-tuned open-source model weights into DBFE storage (e.g., DBFS volumes). Alternatively, hosting the model on Hugging Face and sharing the link for loading inside the DBFE notebook is also perfectly acceptable and often more convenient. Please follow standard DBFE file-size and storage guidelines.\n###\t3.\tTraining Off-Platform\nYes, it is acceptable to train models on your own cloud/GPU infrastructure and use the DBFE notebook only for inference and reproducibility, as long as evaluation and submission strictly follow the competition rules and no external sources are accessed at evaluation time.\n\nWe hope this helps, and we look forward to your submissions!",
    "3289441": "What LLM models are allowed in this challenge?"
  },
  "source": "meta"
}