{
  "id": 575745,
  "title": "Questions Regarding the Hidden Test Set for Final Evaluation",
  "url": "/competitions/stanford-rna-3d-folding/discussion/575745",
  "author_name": "",
  "post_date": "2025-04-30T16:34:57.297481500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> ，I have some questions about the hidden test sets，can you provide some information?</p>\n<p>First, I would like to know the total number of RNA sequences in this hidden test set.</p>\n<p>I am particularly interested in three types of RNA and would like to know more information about them:</p>\n<p>1.Particularly long RNA sequences: For example, those with &gt;480 residues. How many such long sequences are present? This question arises because many participants have experienced issues with their models running excessively long or encountering memory overflow after submission, even though they ran correctly on the provided test sequence data. If participants could be informed about the length range of the RNA sequences in the hidden set beforehand, they could implement certain constraints in their models before submission. This would help avoid repeated trial-and-error submissions and potentially increase the number of valid submissions.</p>\n<p>2.Particularly short RNA sequences: With sequence lengths &lt;30 residues. This type of RNA is often mentioned in literature as being less suitable for evaluation using USalign alignment and similarity scoring. These sequences frequently yield very low scores (&lt;0.2), which can significantly lower the overall average score.</p>\n<p>3.Synthetic RNA: Sequences that do not exist in nature. These often have more complex structures, and some longer ones can cause models to completely crash and fail prediction. Furthermore, scores obtained using existing prediction models on these sequences are often not very high. I would like to know how many of these synthetic RNA sequences are included in the hidden dataset. This information is crucial for me to decide whether it is necessary to perform de novo prediction training using a large dataset of 400,000 synthetic RNA sequences. This is a difficult decision because de novo training requires significant time for data processing and substantial computational resources.</p>",
  "messages": [
    {
      "id": "3190414",
      "postDate": "04/30/2025 16:34:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> ，I have some questions about the hidden test sets，can you provide some information?</p>\n<p>First, I would like to know the total number of RNA sequences in this hidden test set.</p>\n<p>I am particularly interested in three types of RNA and would like to know more information about them:</p>\n<p>1.Particularly long RNA sequences: For example, those with &gt;480 residues. How many such long sequences are present? This question arises because many participants have experienced issues with their models running excessively long or encountering memory overflow after submission, even though they ran correctly on the provided test sequence data. If participants could be informed about the length range of the RNA sequences in the hidden set beforehand, they could implement certain constraints in their models before submission. This would help avoid repeated trial-and-error submissions and potentially increase the number of valid submissions.</p>\n<p>2.Particularly short RNA sequences: With sequence lengths &lt;30 residues. This type of RNA is often mentioned in literature as being less suitable for evaluation using USalign alignment and similarity scoring. These sequences frequently yield very low scores (&lt;0.2), which can significantly lower the overall average score.</p>\n<p>3.Synthetic RNA: Sequences that do not exist in nature. These often have more complex structures, and some longer ones can cause models to completely crash and fail prediction. Furthermore, scores obtained using existing prediction models on these sequences are often not very high. I would like to know how many of these synthetic RNA sequences are included in the hidden dataset. This information is crucial for me to decide whether it is necessary to perform de novo prediction training using a large dataset of 400,000 synthetic RNA sequences. This is a difficult decision because de novo training requires significant time for data processing and substantial computational resources.</p>",
      "rawMarkdown": "Hi @rhijudas ，I have some questions about the hidden test sets，can you provide some information?\n\nFirst, I would like to know the total number of RNA sequences in this hidden test set.\n\nI am particularly interested in three types of RNA and would like to know more information about them:\n\n1.Particularly long RNA sequences: For example, those with >480 residues. How many such long sequences are present? This question arises because many participants have experienced issues with their models running excessively long or encountering memory overflow after submission, even though they ran correctly on the provided test sequence data. If participants could be informed about the length range of the RNA sequences in the hidden set beforehand, they could implement certain constraints in their models before submission. This would help avoid repeated trial-and-error submissions and potentially increase the number of valid submissions.\n\n2.Particularly short RNA sequences: With sequence lengths <30 residues. This type of RNA is often mentioned in literature as being less suitable for evaluation using USalign alignment and similarity scoring. These sequences frequently yield very low scores (<0.2), which can significantly lower the overall average score.\n\n3.Synthetic RNA: Sequences that do not exist in nature. These often have more complex structures, and some longer ones can cause models to completely crash and fail prediction. Furthermore, scores obtained using existing prediction models on these sequences are often not very high. I would like to know how many of these synthetic RNA sequences are included in the hidden dataset. This information is crucial for me to decide whether it is necessary to perform de novo prediction training using a large dataset of 400,000 synthetic RNA sequences. This is a difficult decision because de novo training requires significant time for data processing and substantial computational resources.",
      "votes": null
    },
    {
      "id": "3190449",
      "postDate": "04/30/2025 17:12:25",
      "content": "<p>I think point number 1 would be the most relevant as if they add too many long sequences, codes that execute well in the allowed time might exceed the allowed limit of 9 hours depending on the amount of long sequences. So perhaps you have a gold medal solution that ends up being disqualified because your code did not finish in the allowed time.</p>",
      "rawMarkdown": "I think point number 1 would be the most relevant as if they add too many long sequences, codes that execute well in the allowed time might exceed the allowed limit of 9 hours depending on the amount of long sequences. So perhaps you have a gold medal solution that ends up being disqualified because your code did not finish in the allowed time.",
      "votes": null
    },
    {
      "id": "3190487",
      "postDate": "04/30/2025 18:04:25",
      "content": "<blockquote>\n  <p>I would like to know the total number of RNA sequences in this hidden test set.</p>\n</blockquote>\n<p>It's up to 40 according to the competition page.</p>\n<blockquote>\n  <p>if they add too many long sequences, codes that execute well in the allowed time might exceed the allowed limit</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> I 100% agree.  But let's put ourselves in the host's shoes.</p>\n<p>(1) I pick only up to 40 RNA's for hidden test.</p>\n<p>(2) Short and medium-length RNA predictions are already somewhat good using current SOTA models (compute resources aside).  It's the longer RNA's that are the suckers.</p>\n<p>So, if I were the host, given (1) and (2) as facts, it makes sense as an RNA scientist to focus more on longer RNA's.</p>\n<p>Below are the frequency charts of sequence lengths of the old and new training sets, approximately adjusted to compensate for training set size differences.</p>\n<p>The intention for longer RNA sequences is pretty obvious.</p>\n<p>And my crystal ball says the up-to-40 final test set will be even more disproportionately for longer RNA sequences.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1114184%2F926d1c9d5ad94470e9b73680aed03c9a%2FNew_vs_Old_Length.png?generation=1746037568390731&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": ">I would like to know the total number of RNA sequences in this hidden test set.\n\nIt's up to 40 according to the competition page.\n\n>if they add too many long sequences, codes that execute well in the allowed time might exceed the allowed limit\n\n@alejopaullier I 100% agree.  But let's put ourselves in the host's shoes.\n\n(1) I pick only up to 40 RNA's for hidden test.\n\n(2) Short and medium-length RNA predictions are already somewhat good using current SOTA models (compute resources aside).  It's the longer RNA's that are the suckers.\n\nSo, if I were the host, given (1) and (2) as facts, it makes sense as an RNA scientist to focus more on longer RNA's.\n\nBelow are the frequency charts of sequence lengths of the old and new training sets, approximately adjusted to compensate for training set size differences.\n\nThe intention for longer RNA sequences is pretty obvious.\n\nAnd my crystal ball says the up-to-40 final test set will be even more disproportionately for longer RNA sequences.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1114184%2F926d1c9d5ad94470e9b73680aed03c9a%2FNew_vs_Old_Length.png?generation=1746037568390731&alt=media)",
      "votes": null
    },
    {
      "id": "3190503",
      "postDate": "04/30/2025 18:34:48",
      "content": "<p>Thanks for your questions and the discussion!  </p>\n<p>For the final leaderboard prizes, we will use RNA targets that are not in the current leaderboard whose structures will be released between May 30, 2025 and Sep. 2025.  </p>\n<ul>\n<li>We hosts do not actually know what those targets will be -- we are intentionally scoring the competition on 'future' data to avoid leakage or bias.</li>\n<li>Hosts will, however, curate targets from this future set where the sequences are not exact copies of prior RNA structures, to really test whether models can predict novel RNA structure. So likely there will be no ribosome targets which are often the largest RNA's solved experimentally these days, but do not present particularly novel RNA structures.</li>\n<li>We will filter for targets where the structure is mostly or all RNA (as opposed to small RNA bound to DNA and proteins). The current <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=validation_sequences.csv\" target=\"_blank\">validation test set</a> is drawn from <a href=\"https://predictioncenter.org/casp15/targetlist.cgi?view=rna\" target=\"_blank\">CASP15</a> and provides examples of expected distribution.  </li>\n<li>Note that lengths may well be up to 700-800 nucleotides, as we saw for some targets in <a href=\"https://predictioncenter.org/casp15/target.cgi?id=91&amp;view=rna\" target=\"_blank\">CASP15</a>.</li>\n<li>There will likely be some synthetic RNA's -- they are being solved by several labs in recent years and so there will likely be more in May-Sep 2025.</li>\n<li><a href=\"https://www.rnapuzzles.org/results/\" target=\"_blank\">RNA-Puzzles</a>  provides another example of the kind of test set we can expect -- here's a recent [paper] on 23 RNA targets (<a href=\"https://www.nature.com/articles/s41592-024-02543-9)\" target=\"_blank\">https://www.nature.com/articles/s41592-024-02543-9)</a>.</li>\n<li><a href=\"https://predictioncenter.org/casp16/targetlist.cgi?view=rna&amp;phase=1\" target=\"_blank\">CASP16</a> provides a last example of a test set that may have similarities to the future test set. You can check out this <a href=\"https://www.biorxiv.org/content/10.1101/2025.04.15.649049v1\" target=\"_blank\">CASP16 assessment by structure providers</a> for descriptions. Also stay tuned for a more comprehensive nucleic acid assessment paper in upcoming days.</li>\n</ul>",
      "rawMarkdown": "Thanks for your questions and the discussion!  \n\nFor the final leaderboard prizes, we will use RNA targets that are not in the current leaderboard whose structures will be released between May 30, 2025 and Sep. 2025.  \n - We hosts do not actually know what those targets will be -- we are intentionally scoring the competition on 'future' data to avoid leakage or bias.\n - Hosts will, however, curate targets from this future set where the sequences are not exact copies of prior RNA structures, to really test whether models can predict novel RNA structure. So likely there will be no ribosome targets which are often the largest RNA's solved experimentally these days, but do not present particularly novel RNA structures.\n - We will filter for targets where the structure is mostly or all RNA (as opposed to small RNA bound to DNA and proteins). The current [validation test set](https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=validation_sequences.csv) is drawn from [CASP15](https://predictioncenter.org/casp15/targetlist.cgi?view=rna) and provides examples of expected distribution.  \n - Note that lengths may well be up to 700-800 nucleotides, as we saw for some targets in [CASP15](https://predictioncenter.org/casp15/target.cgi?id=91&view=rna).\n - There will likely be some synthetic RNA's -- they are being solved by several labs in recent years and so there will likely be more in May-Sep 2025.\n - [RNA-Puzzles](https://www.rnapuzzles.org/results/)  provides another example of the kind of test set we can expect -- here's a recent [paper] on 23 RNA targets (https://www.nature.com/articles/s41592-024-02543-9).\n - [CASP16](https://predictioncenter.org/casp16/targetlist.cgi?view=rna&phase=1) provides a last example of a test set that may have similarities to the future test set. You can check out this [CASP16 assessment by structure providers](https://www.biorxiv.org/content/10.1101/2025.04.15.649049v1) for descriptions. Also stay tuned for a more comprehensive nucleic acid assessment paper in upcoming days.",
      "votes": null
    },
    {
      "id": "3190516",
      "postDate": "04/30/2025 18:55:45",
      "content": "<blockquote>\n  <p>We hosts do not actually know what those targets will be -- we are intentionally scoring the competition on 'future' data to avoid leakage or bias.</p>\n</blockquote>\n<p>Will this step be \"blinded\" so that the choice of the final test set be committed and set in stone prior to running any submitted notebooks on it?</p>\n<p>As an example, the host picks the 40 test set, run them on the submitted notebooks, and find that 95% of the notebooks fail (for whatever reason, perhaps time limit).  Then goes back to reselect the 40 test set because 95% of failed notebooks is obviously a bad thing.</p>\n<p>Wouldn't this introduce some unintentional bias, esp against those 5% of the notebooks that ran fine in the first shot?</p>\n<p>Please replace \"95% of the notebooks fail\" above with any unexpected unsatisfactory happenstance, but the essence of my question remains.</p>",
      "rawMarkdown": ">We hosts do not actually know what those targets will be -- we are intentionally scoring the competition on 'future' data to avoid leakage or bias.\n\nWill this step be \"blinded\" so that the choice of the final test set be committed and set in stone prior to running any submitted notebooks on it?\n\nAs an example, the host picks the 40 test set, run them on the submitted notebooks, and find that 95% of the notebooks fail (for whatever reason, perhaps time limit).  Then goes back to reselect the 40 test set because 95% of failed notebooks is obviously a bad thing.\n\nWouldn't this introduce some unintentional bias, esp against those 5% of the notebooks that ran fine in the first shot?\n\nPlease replace \"95% of the notebooks fail\" above with any unexpected unsatisfactory happenstance, but the essence of my question remains.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3190449,
      "author_name": "alejopaullier",
      "author_url": "",
      "post_date": "04/30/2025 17:12:25",
      "content": "<p>I think point number 1 would be the most relevant as if they add too many long sequences, codes that execute well in the allowed time might exceed the allowed limit of 9 hours depending on the amount of long sequences. So perhaps you have a gold medal solution that ends up being disqualified because your code did not finish in the allowed time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3190487,
      "author_name": "revealer",
      "author_url": "",
      "post_date": "04/30/2025 18:04:25",
      "content": "<blockquote>\n  <p>I would like to know the total number of RNA sequences in this hidden test set.</p>\n</blockquote>\n<p>It's up to 40 according to the competition page.</p>\n<blockquote>\n  <p>if they add too many long sequences, codes that execute well in the allowed time might exceed the allowed limit</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> I 100% agree.  But let's put ourselves in the host's shoes.</p>\n<p>(1) I pick only up to 40 RNA's for hidden test.</p>\n<p>(2) Short and medium-length RNA predictions are already somewhat good using current SOTA models (compute resources aside).  It's the longer RNA's that are the suckers.</p>\n<p>So, if I were the host, given (1) and (2) as facts, it makes sense as an RNA scientist to focus more on longer RNA's.</p>\n<p>Below are the frequency charts of sequence lengths of the old and new training sets, approximately adjusted to compensate for training set size differences.</p>\n<p>The intention for longer RNA sequences is pretty obvious.</p>\n<p>And my crystal ball says the up-to-40 final test set will be even more disproportionately for longer RNA sequences.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1114184%2F926d1c9d5ad94470e9b73680aed03c9a%2FNew_vs_Old_Length.png?generation=1746037568390731&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3190503,
      "author_name": "rhijudas",
      "author_url": "",
      "post_date": "04/30/2025 18:34:48",
      "content": "<p>Thanks for your questions and the discussion!  </p>\n<p>For the final leaderboard prizes, we will use RNA targets that are not in the current leaderboard whose structures will be released between May 30, 2025 and Sep. 2025.  </p>\n<ul>\n<li>We hosts do not actually know what those targets will be -- we are intentionally scoring the competition on 'future' data to avoid leakage or bias.</li>\n<li>Hosts will, however, curate targets from this future set where the sequences are not exact copies of prior RNA structures, to really test whether models can predict novel RNA structure. So likely there will be no ribosome targets which are often the largest RNA's solved experimentally these days, but do not present particularly novel RNA structures.</li>\n<li>We will filter for targets where the structure is mostly or all RNA (as opposed to small RNA bound to DNA and proteins). The current <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=validation_sequences.csv\" target=\"_blank\">validation test set</a> is drawn from <a href=\"https://predictioncenter.org/casp15/targetlist.cgi?view=rna\" target=\"_blank\">CASP15</a> and provides examples of expected distribution.  </li>\n<li>Note that lengths may well be up to 700-800 nucleotides, as we saw for some targets in <a href=\"https://predictioncenter.org/casp15/target.cgi?id=91&amp;view=rna\" target=\"_blank\">CASP15</a>.</li>\n<li>There will likely be some synthetic RNA's -- they are being solved by several labs in recent years and so there will likely be more in May-Sep 2025.</li>\n<li><a href=\"https://www.rnapuzzles.org/results/\" target=\"_blank\">RNA-Puzzles</a>  provides another example of the kind of test set we can expect -- here's a recent [paper] on 23 RNA targets (<a href=\"https://www.nature.com/articles/s41592-024-02543-9)\" target=\"_blank\">https://www.nature.com/articles/s41592-024-02543-9)</a>.</li>\n<li><a href=\"https://predictioncenter.org/casp16/targetlist.cgi?view=rna&amp;phase=1\" target=\"_blank\">CASP16</a> provides a last example of a test set that may have similarities to the future test set. You can check out this <a href=\"https://www.biorxiv.org/content/10.1101/2025.04.15.649049v1\" target=\"_blank\">CASP16 assessment by structure providers</a> for descriptions. Also stay tuned for a more comprehensive nucleic acid assessment paper in upcoming days.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3190516,
          "author_name": "revealer",
          "author_url": "",
          "post_date": "04/30/2025 18:55:45",
          "content": "<blockquote>\n  <p>We hosts do not actually know what those targets will be -- we are intentionally scoring the competition on 'future' data to avoid leakage or bias.</p>\n</blockquote>\n<p>Will this step be \"blinded\" so that the choice of the final test set be committed and set in stone prior to running any submitted notebooks on it?</p>\n<p>As an example, the host picks the 40 test set, run them on the submitted notebooks, and find that 95% of the notebooks fail (for whatever reason, perhaps time limit).  Then goes back to reselect the 40 test set because 95% of failed notebooks is obviously a bad thing.</p>\n<p>Wouldn't this introduce some unintentional bias, esp against those 5% of the notebooks that ran fine in the first shot?</p>\n<p>Please replace \"95% of the notebooks fail\" above with any unexpected unsatisfactory happenstance, but the essence of my question remains.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3190414": "Hi @rhijudas ，I have some questions about the hidden test sets，can you provide some information?\n\nFirst, I would like to know the total number of RNA sequences in this hidden test set.\n\nI am particularly interested in three types of RNA and would like to know more information about them:\n\n1.Particularly long RNA sequences: For example, those with >480 residues. How many such long sequences are present? This question arises because many participants have experienced issues with their models running excessively long or encountering memory overflow after submission, even though they ran correctly on the provided test sequence data. If participants could be informed about the length range of the RNA sequences in the hidden set beforehand, they could implement certain constraints in their models before submission. This would help avoid repeated trial-and-error submissions and potentially increase the number of valid submissions.\n\n2.Particularly short RNA sequences: With sequence lengths <30 residues. This type of RNA is often mentioned in literature as being less suitable for evaluation using USalign alignment and similarity scoring. These sequences frequently yield very low scores (<0.2), which can significantly lower the overall average score.\n\n3.Synthetic RNA: Sequences that do not exist in nature. These often have more complex structures, and some longer ones can cause models to completely crash and fail prediction. Furthermore, scores obtained using existing prediction models on these sequences are often not very high. I would like to know how many of these synthetic RNA sequences are included in the hidden dataset. This information is crucial for me to decide whether it is necessary to perform de novo prediction training using a large dataset of 400,000 synthetic RNA sequences. This is a difficult decision because de novo training requires significant time for data processing and substantial computational resources.",
    "3190449": "I think point number 1 would be the most relevant as if they add too many long sequences, codes that execute well in the allowed time might exceed the allowed limit of 9 hours depending on the amount of long sequences. So perhaps you have a gold medal solution that ends up being disqualified because your code did not finish in the allowed time.",
    "3190487": ">I would like to know the total number of RNA sequences in this hidden test set.\n\nIt's up to 40 according to the competition page.\n\n>if they add too many long sequences, codes that execute well in the allowed time might exceed the allowed limit\n\n@alejopaullier I 100% agree.  But let's put ourselves in the host's shoes.\n\n(1) I pick only up to 40 RNA's for hidden test.\n\n(2) Short and medium-length RNA predictions are already somewhat good using current SOTA models (compute resources aside).  It's the longer RNA's that are the suckers.\n\nSo, if I were the host, given (1) and (2) as facts, it makes sense as an RNA scientist to focus more on longer RNA's.\n\nBelow are the frequency charts of sequence lengths of the old and new training sets, approximately adjusted to compensate for training set size differences.\n\nThe intention for longer RNA sequences is pretty obvious.\n\nAnd my crystal ball says the up-to-40 final test set will be even more disproportionately for longer RNA sequences.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1114184%2F926d1c9d5ad94470e9b73680aed03c9a%2FNew_vs_Old_Length.png?generation=1746037568390731&alt=media)",
    "3190503": "Thanks for your questions and the discussion!  \n\nFor the final leaderboard prizes, we will use RNA targets that are not in the current leaderboard whose structures will be released between May 30, 2025 and Sep. 2025.  \n - We hosts do not actually know what those targets will be -- we are intentionally scoring the competition on 'future' data to avoid leakage or bias.\n - Hosts will, however, curate targets from this future set where the sequences are not exact copies of prior RNA structures, to really test whether models can predict novel RNA structure. So likely there will be no ribosome targets which are often the largest RNA's solved experimentally these days, but do not present particularly novel RNA structures.\n - We will filter for targets where the structure is mostly or all RNA (as opposed to small RNA bound to DNA and proteins). The current [validation test set](https://www.kaggle.com/competitions/stanford-rna-3d-folding/data?select=validation_sequences.csv) is drawn from [CASP15](https://predictioncenter.org/casp15/targetlist.cgi?view=rna) and provides examples of expected distribution.  \n - Note that lengths may well be up to 700-800 nucleotides, as we saw for some targets in [CASP15](https://predictioncenter.org/casp15/target.cgi?id=91&view=rna).\n - There will likely be some synthetic RNA's -- they are being solved by several labs in recent years and so there will likely be more in May-Sep 2025.\n - [RNA-Puzzles](https://www.rnapuzzles.org/results/)  provides another example of the kind of test set we can expect -- here's a recent [paper] on 23 RNA targets (https://www.nature.com/articles/s41592-024-02543-9).\n - [CASP16](https://predictioncenter.org/casp16/targetlist.cgi?view=rna&phase=1) provides a last example of a test set that may have similarities to the future test set. You can check out this [CASP16 assessment by structure providers](https://www.biorxiv.org/content/10.1101/2025.04.15.649049v1) for descriptions. Also stay tuned for a more comprehensive nucleic acid assessment paper in upcoming days.",
    "3190516": ">We hosts do not actually know what those targets will be -- we are intentionally scoring the competition on 'future' data to avoid leakage or bias.\n\nWill this step be \"blinded\" so that the choice of the final test set be committed and set in stone prior to running any submitted notebooks on it?\n\nAs an example, the host picks the 40 test set, run them on the submitted notebooks, and find that 95% of the notebooks fail (for whatever reason, perhaps time limit).  Then goes back to reselect the 40 test set because 95% of failed notebooks is obviously a bad thing.\n\nWouldn't this introduce some unintentional bias, esp against those 5% of the notebooks that ran fine in the first shot?\n\nPlease replace \"95% of the notebooks fail\" above with any unexpected unsatisfactory happenstance, but the essence of my question remains."
  },
  "source": "meta"
}