{
  "id": 573798,
  "title": "Can We Use LLMs/Open Source Models to Generate Training/Validation Metadata (Not Test Data), If Fully Documented?",
  "url": "/competitions/fungi-clef-2025/discussion/573798",
  "author_name": "",
  "post_date": "2025-04-18T00:06:36.492251800Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/picekl\" target=\"_blank\">@picekl</a> ,</p>\n<p>I am participating in FungiCLEF25 and am seeking clarification regarding the use of large language models (LLM) and open source model APIs for metadata augmentation, specifically for the provided training and validation sets (NOT the test set).</p>\n<p>My scenario:</p>\n<p>I would like to use either an open source language model (like Llama, Mistral, etc.) or a commercial LLM API (like ChatGPT, Gemini, etc.) to generate additional metadata fields for the training and validation data.<br>\nThis could include rephrased descriptions, summarized content, or possibly additional factual fields generated by the LLM based on the existing info in training/validation.<br>\nI do NOT intend to use these tools for the test set in any way.<br>\nAll steps involved in augmentation, including all prompts, parameters, and code used for the metadata generation, will be fully documented and included in my working notes for complete transparency and reproducibility.</p>\n<p>Specifically:</p>\n<p>If all data processing, LLM prompting, and field generation are shared and documented, is it permitted to use LLM outputs or open source model outputs to augment the training/validation metadata?<br>\nIf not, is there any way to use such techniques for training/validation set feature engineering that falls within the rules (for example, are autogenerated paraphrases or classifications admissible if disclosed)?<br>\nAgain, no external data or model output will be used for the test set, and all steps will be made public for review.<br>\nYour clarification would be much appreciated, as the competition rules state that no external data can be used, but the definition of “external” in relation to LLM-generated fields is unclear especially when full transparency is offered and the process is only applied to the training and validation sets.</p>\n<p>Thank you for your time and for helping keep the competition fair and clear for everyone!</p>\n<p>Best regards,</p>",
  "messages": [
    {
      "id": "3181481",
      "postDate": "04/18/2025 00:06:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/picekl\" target=\"_blank\">@picekl</a> ,</p>\n<p>I am participating in FungiCLEF25 and am seeking clarification regarding the use of large language models (LLM) and open source model APIs for metadata augmentation, specifically for the provided training and validation sets (NOT the test set).</p>\n<p>My scenario:</p>\n<p>I would like to use either an open source language model (like Llama, Mistral, etc.) or a commercial LLM API (like ChatGPT, Gemini, etc.) to generate additional metadata fields for the training and validation data.<br>\nThis could include rephrased descriptions, summarized content, or possibly additional factual fields generated by the LLM based on the existing info in training/validation.<br>\nI do NOT intend to use these tools for the test set in any way.<br>\nAll steps involved in augmentation, including all prompts, parameters, and code used for the metadata generation, will be fully documented and included in my working notes for complete transparency and reproducibility.</p>\n<p>Specifically:</p>\n<p>If all data processing, LLM prompting, and field generation are shared and documented, is it permitted to use LLM outputs or open source model outputs to augment the training/validation metadata?<br>\nIf not, is there any way to use such techniques for training/validation set feature engineering that falls within the rules (for example, are autogenerated paraphrases or classifications admissible if disclosed)?<br>\nAgain, no external data or model output will be used for the test set, and all steps will be made public for review.<br>\nYour clarification would be much appreciated, as the competition rules state that no external data can be used, but the definition of “external” in relation to LLM-generated fields is unclear especially when full transparency is offered and the process is only applied to the training and validation sets.</p>\n<p>Thank you for your time and for helping keep the competition fair and clear for everyone!</p>\n<p>Best regards,</p>",
      "rawMarkdown": "Hi @picekl ,\n\nI am participating in FungiCLEF25 and am seeking clarification regarding the use of large language models (LLM) and open source model APIs for metadata augmentation, specifically for the provided training and validation sets (NOT the test set).\n\nMy scenario:\n\nI would like to use either an open source language model (like Llama, Mistral, etc.) or a commercial LLM API (like ChatGPT, Gemini, etc.) to generate additional metadata fields for the training and validation data.\nThis could include rephrased descriptions, summarized content, or possibly additional factual fields generated by the LLM based on the existing info in training/validation.\nI do NOT intend to use these tools for the test set in any way.\nAll steps involved in augmentation, including all prompts, parameters, and code used for the metadata generation, will be fully documented and included in my working notes for complete transparency and reproducibility.\n\nSpecifically:\n\nIf all data processing, LLM prompting, and field generation are shared and documented, is it permitted to use LLM outputs or open source model outputs to augment the training/validation metadata?\nIf not, is there any way to use such techniques for training/validation set feature engineering that falls within the rules (for example, are autogenerated paraphrases or classifications admissible if disclosed)?\nAgain, no external data or model output will be used for the test set, and all steps will be made public for review.\nYour clarification would be much appreciated, as the competition rules state that no external data can be used, but the definition of “external” in relation to LLM-generated fields is unclear especially when full transparency is offered and the process is only applied to the training and validation sets.\n\nThank you for your time and for helping keep the competition fair and clear for everyone!\n\nBest regards,",
      "votes": null
    },
    {
      "id": "3183875",
      "postDate": "04/21/2025 11:56:21",
      "content": "<p>I look for an answer too</p>",
      "rawMarkdown": "I look for an answer too",
      "votes": null
    },
    {
      "id": "3184522",
      "postDate": "04/22/2025 07:56:31",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/abdrah\" target=\"_blank\">@abdrah</a> and <a href=\"https://www.kaggle.com/gonginmichail\" target=\"_blank\">@gonginmichail</a>,</p>\n<p>Thanks for asking, this sounds great! These are exactly the kinds of approaches we're excited to see.</p>\n<p>Given the interesting direction you're exploring, we’d also encourage you to consider writing a Working Note. It’s a great way to share insights and contribute to the broader community.</p>\n<p>You can find more details about Working Notes in this discussion thread: [LINK](<a href=\"https://www.kaggle.com/competitions/fungi-clef-2025/discussion/569\" target=\"_blank\">https://www.kaggle.com/competitions/fungi-clef-2025/discussion/569</a></p>",
      "rawMarkdown": "Dear @abdrah and @gonginmichail,\n\nThanks for asking, this sounds great! These are exactly the kinds of approaches we're excited to see.\n\nGiven the interesting direction you're exploring, we’d also encourage you to consider writing a Working Note. It’s a great way to share insights and contribute to the broader community.\n\nYou can find more details about Working Notes in this discussion thread: [LINK](https://www.kaggle.com/competitions/fungi-clef-2025/discussion/569",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3183875,
      "author_name": "",
      "author_url": "",
      "post_date": "04/21/2025 11:56:21",
      "content": "<p>I look for an answer too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3184522,
      "author_name": "picekl",
      "author_url": "",
      "post_date": "04/22/2025 07:56:31",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/abdrah\" target=\"_blank\">@abdrah</a> and <a href=\"https://www.kaggle.com/gonginmichail\" target=\"_blank\">@gonginmichail</a>,</p>\n<p>Thanks for asking, this sounds great! These are exactly the kinds of approaches we're excited to see.</p>\n<p>Given the interesting direction you're exploring, we’d also encourage you to consider writing a Working Note. It’s a great way to share insights and contribute to the broader community.</p>\n<p>You can find more details about Working Notes in this discussion thread: [LINK](<a href=\"https://www.kaggle.com/competitions/fungi-clef-2025/discussion/569\" target=\"_blank\">https://www.kaggle.com/competitions/fungi-clef-2025/discussion/569</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3181481": "Hi @picekl ,\n\nI am participating in FungiCLEF25 and am seeking clarification regarding the use of large language models (LLM) and open source model APIs for metadata augmentation, specifically for the provided training and validation sets (NOT the test set).\n\nMy scenario:\n\nI would like to use either an open source language model (like Llama, Mistral, etc.) or a commercial LLM API (like ChatGPT, Gemini, etc.) to generate additional metadata fields for the training and validation data.\nThis could include rephrased descriptions, summarized content, or possibly additional factual fields generated by the LLM based on the existing info in training/validation.\nI do NOT intend to use these tools for the test set in any way.\nAll steps involved in augmentation, including all prompts, parameters, and code used for the metadata generation, will be fully documented and included in my working notes for complete transparency and reproducibility.\n\nSpecifically:\n\nIf all data processing, LLM prompting, and field generation are shared and documented, is it permitted to use LLM outputs or open source model outputs to augment the training/validation metadata?\nIf not, is there any way to use such techniques for training/validation set feature engineering that falls within the rules (for example, are autogenerated paraphrases or classifications admissible if disclosed)?\nAgain, no external data or model output will be used for the test set, and all steps will be made public for review.\nYour clarification would be much appreciated, as the competition rules state that no external data can be used, but the definition of “external” in relation to LLM-generated fields is unclear especially when full transparency is offered and the process is only applied to the training and validation sets.\n\nThank you for your time and for helping keep the competition fair and clear for everyone!\n\nBest regards,",
    "3183875": "I look for an answer too",
    "3184522": "Dear @abdrah and @gonginmichail,\n\nThanks for asking, this sounds great! These are exactly the kinds of approaches we're excited to see.\n\nGiven the interesting direction you're exploring, we’d also encourage you to consider writing a Working Note. It’s a great way to share insights and contribute to the broader community.\n\nYou can find more details about Working Notes in this discussion thread: [LINK](https://www.kaggle.com/competitions/fungi-clef-2025/discussion/569"
  },
  "source": "meta"
}