{
  "id": 566854,
  "title": "Clarification on Dataset Terminology",
  "url": "/competitions/stanford-rna-3d-folding/discussion/566854",
  "author_name": "",
  "post_date": "2025-03-07T04:44:04.230210400Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, everyone!</p>\n<p>I’m a bit confused about the different dataset terms used in the description (see quoted text below) and their relationship to the public and private leaderboards. Not sure if others feel the same way.</p>\n<p>Could someone kindly clarify how the \"public test set\", \"hidden test set\", \"private test set\", \"public set\", \"test set\",  relate to each leaderboard across different phases?</p>\n<p>Thanks!</p>\n<blockquote>\n  <p>Initial model training phase. <br>\n  At launch, expect approximately 25 sequences in the <strong>hidden test set</strong> . Some of those sequences are used for a private leaderboard to allow the host to track progress on wholly unseen data. During this phase the <strong>public test set</strong> sequences includes–but is not limited to–targets from the 2024 CASP16 competition whose structures have not yet been publicly released in the PDB database.<br>\n  Model training phase 2.<br>\n  On April 23rd we will update the <strong>hidden test set</strong> and reset the leaderboard. Sequences in the current <strong>public test set</strong> will be added to the train data, all sequences currently in the <strong>private set</strong> will be rolled into the new <strong>public set</strong> , and new sequences will be added to the <strong>public test set</strong> .<br>\n  Future data phase. <br>\n  Your selected submissions will be run against a completely new <strong>private test set</strong> generated after the end of the model training phases. There will be up to 40 sequences in the <strong>test set</strong> , all of them used for the private leaderboard.</p>\n  <hr>\n</blockquote>",
  "messages": [
    {
      "id": "3143292",
      "postDate": "03/07/2025 04:44:04",
      "content": "<p>Hi, everyone!</p>\n<p>I’m a bit confused about the different dataset terms used in the description (see quoted text below) and their relationship to the public and private leaderboards. Not sure if others feel the same way.</p>\n<p>Could someone kindly clarify how the \"public test set\", \"hidden test set\", \"private test set\", \"public set\", \"test set\",  relate to each leaderboard across different phases?</p>\n<p>Thanks!</p>\n<blockquote>\n  <p>Initial model training phase. <br>\n  At launch, expect approximately 25 sequences in the <strong>hidden test set</strong> . Some of those sequences are used for a private leaderboard to allow the host to track progress on wholly unseen data. During this phase the <strong>public test set</strong> sequences includes–but is not limited to–targets from the 2024 CASP16 competition whose structures have not yet been publicly released in the PDB database.<br>\n  Model training phase 2.<br>\n  On April 23rd we will update the <strong>hidden test set</strong> and reset the leaderboard. Sequences in the current <strong>public test set</strong> will be added to the train data, all sequences currently in the <strong>private set</strong> will be rolled into the new <strong>public set</strong> , and new sequences will be added to the <strong>public test set</strong> .<br>\n  Future data phase. <br>\n  Your selected submissions will be run against a completely new <strong>private test set</strong> generated after the end of the model training phases. There will be up to 40 sequences in the <strong>test set</strong> , all of them used for the private leaderboard.</p>\n  <hr>\n</blockquote>",
      "rawMarkdown": "Hi, everyone!\n\nI’m a bit confused about the different dataset terms used in the description (see quoted text below) and their relationship to the public and private leaderboards. Not sure if others feel the same way.\n\nCould someone kindly clarify how the \"public test set\", \"hidden test set\", \"private test set\", \"public set\", \"test set\",  relate to each leaderboard across different phases?\n\nThanks!\n\n>Initial model training phase. \nAt launch, expect approximately 25 sequences in the **hidden test set** . Some of those sequences are used for a private leaderboard to allow the host to track progress on wholly unseen data. During this phase the **public test set** sequences includes–but is not limited to–targets from the 2024 CASP16 competition whose structures have not yet been publicly released in the PDB database.\nModel training phase 2.\nOn April 23rd we will update the **hidden test set** and reset the leaderboard. Sequences in the current **public test set** will be added to the train data, all sequences currently in the **private set** will be rolled into the new **public set** , and new sequences will be added to the **public test set** .\nFuture data phase. \nYour selected submissions will be run against a completely new **private test set** generated after the end of the model training phases. There will be up to 40 sequences in the **test set** , all of them used for the private leaderboard.\n****",
      "votes": null
    },
    {
      "id": "3148528",
      "postDate": "03/13/2025 08:42:32",
      "content": "<p>The terms describe time‐phased sets of sequences that transition between “public” (visible to participants) and “hidden/private” (used only for final scoring). Here’s a brief timeline:</p>\n<ol>\n<li><p><strong>Initial Model Training Phase:</strong></p>\n<ul>\n<li><strong>Hidden test set (~25 sequences):</strong> Used internally by the host for a private leaderboard on unseen data.</li>\n<li><strong>Public test set:</strong> Visible to participants, includes some CASP16 targets. This set supports the public leaderboard initially.</li></ul></li>\n<li><p><strong>Model Training Phase 2 (April 23rd update):</strong></p>\n<ul>\n<li>The <strong>current public test set</strong> is moved into training data.</li>\n<li>The <strong>current hidden test set</strong> becomes “rolled into” the new public set so participants can see it.</li>\n<li><strong>New sequences</strong> form a fresh public test set for a new public leaderboard.</li></ul></li>\n<li><p><strong>Future Data Phase:</strong></p>\n<ul>\n<li>A <strong>completely new private test set</strong> (~40 sequences) is introduced and remains hidden.</li>\n<li>This final hidden set is used for the ultimate private leaderboard at competition’s end.</li></ul></li>\n</ol>\n<p>In short:</p>\n<ul>\n<li>“Public test sets” are released to participants and used for the <strong>public</strong> leaderboard.</li>\n<li>“Hidden/private test sets” stay unseen until they become public or are used for the final <strong>private</strong> leaderboard scoring.</li>\n<li>Old public sets can transition into train data, and old hidden sets may become public in the next phase. Ultimately, a new hidden set is used for the final unbiased scoring.</li>\n</ul>",
      "rawMarkdown": "The terms describe time‐phased sets of sequences that transition between “public” (visible to participants) and “hidden/private” (used only for final scoring). Here’s a brief timeline:\n\n1. **Initial Model Training Phase:**\n   - **Hidden test set (~25 sequences):** Used internally by the host for a private leaderboard on unseen data.\n   - **Public test set:** Visible to participants, includes some CASP16 targets. This set supports the public leaderboard initially.\n\n2. **Model Training Phase 2 (April 23rd update):**\n   - The **current public test set** is moved into training data.\n   - The **current hidden test set** becomes “rolled into” the new public set so participants can see it.\n   - **New sequences** form a fresh public test set for a new public leaderboard.\n\n3. **Future Data Phase:**\n   - A **completely new private test set** (~40 sequences) is introduced and remains hidden.\n   - This final hidden set is used for the ultimate private leaderboard at competition’s end.\n\nIn short:\n- “Public test sets” are released to participants and used for the **public** leaderboard.\n- “Hidden/private test sets” stay unseen until they become public or are used for the final **private** leaderboard scoring.\n- Old public sets can transition into train data, and old hidden sets may become public in the next phase. Ultimately, a new hidden set is used for the final unbiased scoring.",
      "votes": null
    },
    {
      "id": "3149266",
      "postDate": "03/14/2025 02:26:41",
      "content": "<p>Thank you so much for the clarification! </p>",
      "rawMarkdown": "Thank you so much for the clarification!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3148528,
      "author_name": "younusmohamed",
      "author_url": "",
      "post_date": "03/13/2025 08:42:32",
      "content": "<p>The terms describe time‐phased sets of sequences that transition between “public” (visible to participants) and “hidden/private” (used only for final scoring). Here’s a brief timeline:</p>\n<ol>\n<li><p><strong>Initial Model Training Phase:</strong></p>\n<ul>\n<li><strong>Hidden test set (~25 sequences):</strong> Used internally by the host for a private leaderboard on unseen data.</li>\n<li><strong>Public test set:</strong> Visible to participants, includes some CASP16 targets. This set supports the public leaderboard initially.</li></ul></li>\n<li><p><strong>Model Training Phase 2 (April 23rd update):</strong></p>\n<ul>\n<li>The <strong>current public test set</strong> is moved into training data.</li>\n<li>The <strong>current hidden test set</strong> becomes “rolled into” the new public set so participants can see it.</li>\n<li><strong>New sequences</strong> form a fresh public test set for a new public leaderboard.</li></ul></li>\n<li><p><strong>Future Data Phase:</strong></p>\n<ul>\n<li>A <strong>completely new private test set</strong> (~40 sequences) is introduced and remains hidden.</li>\n<li>This final hidden set is used for the ultimate private leaderboard at competition’s end.</li></ul></li>\n</ol>\n<p>In short:</p>\n<ul>\n<li>“Public test sets” are released to participants and used for the <strong>public</strong> leaderboard.</li>\n<li>“Hidden/private test sets” stay unseen until they become public or are used for the final <strong>private</strong> leaderboard scoring.</li>\n<li>Old public sets can transition into train data, and old hidden sets may become public in the next phase. Ultimately, a new hidden set is used for the final unbiased scoring.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3149266,
          "author_name": "achenge07",
          "author_url": "",
          "post_date": "03/14/2025 02:26:41",
          "content": "<p>Thank you so much for the clarification! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3143292": "Hi, everyone!\n\nI’m a bit confused about the different dataset terms used in the description (see quoted text below) and their relationship to the public and private leaderboards. Not sure if others feel the same way.\n\nCould someone kindly clarify how the \"public test set\", \"hidden test set\", \"private test set\", \"public set\", \"test set\",  relate to each leaderboard across different phases?\n\nThanks!\n\n>Initial model training phase. \nAt launch, expect approximately 25 sequences in the **hidden test set** . Some of those sequences are used for a private leaderboard to allow the host to track progress on wholly unseen data. During this phase the **public test set** sequences includes–but is not limited to–targets from the 2024 CASP16 competition whose structures have not yet been publicly released in the PDB database.\nModel training phase 2.\nOn April 23rd we will update the **hidden test set** and reset the leaderboard. Sequences in the current **public test set** will be added to the train data, all sequences currently in the **private set** will be rolled into the new **public set** , and new sequences will be added to the **public test set** .\nFuture data phase. \nYour selected submissions will be run against a completely new **private test set** generated after the end of the model training phases. There will be up to 40 sequences in the **test set** , all of them used for the private leaderboard.\n****",
    "3148528": "The terms describe time‐phased sets of sequences that transition between “public” (visible to participants) and “hidden/private” (used only for final scoring). Here’s a brief timeline:\n\n1. **Initial Model Training Phase:**\n   - **Hidden test set (~25 sequences):** Used internally by the host for a private leaderboard on unseen data.\n   - **Public test set:** Visible to participants, includes some CASP16 targets. This set supports the public leaderboard initially.\n\n2. **Model Training Phase 2 (April 23rd update):**\n   - The **current public test set** is moved into training data.\n   - The **current hidden test set** becomes “rolled into” the new public set so participants can see it.\n   - **New sequences** form a fresh public test set for a new public leaderboard.\n\n3. **Future Data Phase:**\n   - A **completely new private test set** (~40 sequences) is introduced and remains hidden.\n   - This final hidden set is used for the ultimate private leaderboard at competition’s end.\n\nIn short:\n- “Public test sets” are released to participants and used for the **public** leaderboard.\n- “Hidden/private test sets” stay unseen until they become public or are used for the final **private** leaderboard scoring.\n- Old public sets can transition into train data, and old hidden sets may become public in the next phase. Ultimately, a new hidden set is used for the final unbiased scoring.",
    "3149266": "Thank you so much for the clarification!"
  },
  "source": "meta"
}