{
  "id": 568794,
  "title": "Temporal Cutoff and Data Leak",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568794",
  "author_name": "",
  "post_date": "2025-03-18T01:51:13.596274400Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>From my understanding, an RNA is typically \"solved\" (meaning its true nucleotide 3D coordinates obtained) by some lab out there with expensive chemistry equipment.</p>\n<p>Say RNA X was solved, then <strong>sometime later</strong>, RNA Y was solved.</p>\n<p>(A) Say we use Y as training data trying to predict X as validation or test data, would that incur a leak?  Assume training from scratch afresh without any pre-trained models.</p>\n<p>(B) As far as nature goes, the answer should be no.  Just because we poor human mortals solved Y only after solving X, it doesn't mean that the structure of Y depends on X in any inherent way (other than any possible co-evolution or whatever happened naturally between RNA X and RNA Y, but that dependence wouldn't be a \"leak\" and certainly has nothing to do with Y being solved after X).</p>\n<p>But, if the lab/team that solved Y somehow used the solution of X in any way when solving Y, then there could potentially be a leak if we do (A).</p>\n<p>Could someone please clarify if doing (A) is safe and leak-free?</p>\n<p>If (A) is not leak-free, then this becomes a bit like time-series prediction, where we should never use a \"future\" data to make a \"past\" prediction.  As far as ML pipeline goes, not a problem.  We do it all the time.  But for RNA structure, this \"time-series\" likeness really bothers me due to (B).  </p>",
  "messages": [
    {
      "id": "3152599",
      "postDate": "03/18/2025 01:51:13",
      "content": "<p>From my understanding, an RNA is typically \"solved\" (meaning its true nucleotide 3D coordinates obtained) by some lab out there with expensive chemistry equipment.</p>\n<p>Say RNA X was solved, then <strong>sometime later</strong>, RNA Y was solved.</p>\n<p>(A) Say we use Y as training data trying to predict X as validation or test data, would that incur a leak?  Assume training from scratch afresh without any pre-trained models.</p>\n<p>(B) As far as nature goes, the answer should be no.  Just because we poor human mortals solved Y only after solving X, it doesn't mean that the structure of Y depends on X in any inherent way (other than any possible co-evolution or whatever happened naturally between RNA X and RNA Y, but that dependence wouldn't be a \"leak\" and certainly has nothing to do with Y being solved after X).</p>\n<p>But, if the lab/team that solved Y somehow used the solution of X in any way when solving Y, then there could potentially be a leak if we do (A).</p>\n<p>Could someone please clarify if doing (A) is safe and leak-free?</p>\n<p>If (A) is not leak-free, then this becomes a bit like time-series prediction, where we should never use a \"future\" data to make a \"past\" prediction.  As far as ML pipeline goes, not a problem.  We do it all the time.  But for RNA structure, this \"time-series\" likeness really bothers me due to (B).  </p>",
      "rawMarkdown": "From my understanding, an RNA is typically \"solved\" (meaning its true nucleotide 3D coordinates obtained) by some lab out there with expensive chemistry equipment.\n\nSay RNA X was solved, then **sometime later**, RNA Y was solved.\n\n(A) Say we use Y as training data trying to predict X as validation or test data, would that incur a leak?  Assume training from scratch afresh without any pre-trained models.\n\n(B) As far as nature goes, the answer should be no.  Just because we poor human mortals solved Y only after solving X, it doesn't mean that the structure of Y depends on X in any inherent way (other than any possible co-evolution or whatever happened naturally between RNA X and RNA Y, but that dependence wouldn't be a \"leak\" and certainly has nothing to do with Y being solved after X).\n\nBut, if the lab/team that solved Y somehow used the solution of X in any way when solving Y, then there could potentially be a leak if we do (A).\n\nCould someone please clarify if doing (A) is safe and leak-free?\n\nIf (A) is not leak-free, then this becomes a bit like time-series prediction, where we should never use a \"future\" data to make a \"past\" prediction.  As far as ML pipeline goes, not a problem.  We do it all the time.  But for RNA structure, this \"time-series\" likeness really bothers me due to (B).",
      "votes": null
    },
    {
      "id": "3152627",
      "postDate": "03/18/2025 02:31:58",
      "content": "<p>Hi. But as labs out there with expensive chemistry equipment doesn't like to waste resources they only elucidate brand new structures. So time cutoff grants that none model trained before cutoff had the oportunity to train with those labels.</p>",
      "rawMarkdown": "Hi. But as labs out there with expensive chemistry equipment doesn't like to waste resources they only elucidate brand new structures. So time cutoff grants that none model trained before cutoff had the oportunity to train with those labels.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3152627,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "03/18/2025 02:31:58",
      "content": "<p>Hi. But as labs out there with expensive chemistry equipment doesn't like to waste resources they only elucidate brand new structures. So time cutoff grants that none model trained before cutoff had the oportunity to train with those labels.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3152599": "From my understanding, an RNA is typically \"solved\" (meaning its true nucleotide 3D coordinates obtained) by some lab out there with expensive chemistry equipment.\n\nSay RNA X was solved, then **sometime later**, RNA Y was solved.\n\n(A) Say we use Y as training data trying to predict X as validation or test data, would that incur a leak?  Assume training from scratch afresh without any pre-trained models.\n\n(B) As far as nature goes, the answer should be no.  Just because we poor human mortals solved Y only after solving X, it doesn't mean that the structure of Y depends on X in any inherent way (other than any possible co-evolution or whatever happened naturally between RNA X and RNA Y, but that dependence wouldn't be a \"leak\" and certainly has nothing to do with Y being solved after X).\n\nBut, if the lab/team that solved Y somehow used the solution of X in any way when solving Y, then there could potentially be a leak if we do (A).\n\nCould someone please clarify if doing (A) is safe and leak-free?\n\nIf (A) is not leak-free, then this becomes a bit like time-series prediction, where we should never use a \"future\" data to make a \"past\" prediction.  As far as ML pipeline goes, not a problem.  We do it all the time.  But for RNA structure, this \"time-series\" likeness really bothers me due to (B).",
    "3152627": "Hi. But as labs out there with expensive chemistry equipment doesn't like to waste resources they only elucidate brand new structures. So time cutoff grants that none model trained before cutoff had the oportunity to train with those labels."
  },
  "source": "meta"
}