{
  "id": 575109,
  "title": "Training data update",
  "url": "/competitions/stanford-rna-3d-folding/discussion/575109",
  "author_name": "Shujun",
  "post_date": "2025-04-26T03:25:45.359000",
  "votes": 13,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Hi Kagglers, we have released an updated version of training data with more relaxed filters. The main changes are </p>\n<ol>\n<li>we downloaded pdbs from the protein data bank with full text search for keyword RNA</li>\n<li>we relaxed filter for unstructured RNAs based on pairwise C1' distances, where 20% of residues have to be close to some other residue that is over 4 bases apart</li>\n</ol>\n<p>As a result, we now have &gt;5000 sequences in total and around 3300 before the CASP15 temporal cutoff (2022-05-01). The new files are in competition training data named <code>train_sequences.v2.csv</code> and <code>train_labels.v2.csv</code></p>\n<p>I have done a rerun of the previously releasesd Rnet2 3D finetuning notebook and score improved to around 0.27 on the public leaderboard. I have updated the inference notebook: <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference</a></p>\n<p>Have fun modeling! </p>",
  "messages": [
    {
      "id": 3187439,
      "postDate": "2025-04-26T03:25:45.360Z",
      "content": "<p>Hi Kagglers, we have released an updated version of training data with more relaxed filters. The main changes are </p>\n<ol>\n<li>we downloaded pdbs from the protein data bank with full text search for keyword RNA</li>\n<li>we relaxed filter for unstructured RNAs based on pairwise C1' distances, where 20% of residues have to be close to some other residue that is over 4 bases apart</li>\n</ol>\n<p>As a result, we now have &gt;5000 sequences in total and around 3300 before the CASP15 temporal cutoff (2022-05-01). The new files are in competition training data named <code>train_sequences.v2.csv</code> and <code>train_labels.v2.csv</code></p>\n<p>I have done a rerun of the previously releasesd Rnet2 3D finetuning notebook and score improved to around 0.27 on the public leaderboard. I have updated the inference notebook: <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference</a></p>\n<p>Have fun modeling! </p>",
      "rawMarkdown": "Hi Kagglers, we have released an updated version of training data with more relaxed filters. The main changes are \n1.  we downloaded pdbs from the protein data bank with full text search for keyword RNA\n2. we relaxed filter for unstructured RNAs based on pairwise C1' distances, where 20% of residues have to be close to some other residue that is over 4 bases apart\n\nAs a result, we now have >5000 sequences in total and around 3300 before the CASP15 temporal cutoff (2022-05-01). The new files are in competition training data named ```train_sequences.v2.csv``` and ```train_labels.v2.csv```\n\nI have done a rerun of the previously releasesd Rnet2 3D finetuning notebook and score improved to around 0.27 on the public leaderboard. I have updated the inference notebook: [https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference](https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference)\n\nHave fun modeling! \n",
      "votes": 13
    },
    {
      "id": 3187508,
      "postDate": "2025-04-26T06:05:38.520Z",
      "content": "<p>Thanks a lot for updating the training data!</p>\n<p>Could you share the training data preprocessing pipeline as well like <a href=\"https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing\" target=\"_blank\">https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing</a> </p>",
      "rawMarkdown": "Thanks a lot for updating the training data!\n\nCould you share the training data preprocessing pipeline as well like https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing ",
      "votes": 1,
      "replies": [
        {
          "id": 3187512,
          "postDate": "2025-04-26T06:14:55.253Z",
          "content": "<p>yes it's in the v2 branch</p>",
          "rawMarkdown": "yes it's in the v2 branch",
          "votes": 1
        }
      ]
    },
    {
      "id": 3194413,
      "postDate": "2025-05-05T20:59:46.523Z",
      "content": "<p>Hey everyone, in addition to <a href=\"https://www.kaggle.com/ckjoshi9\" target=\"_blank\">@ckjoshi9</a> 's repo (<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556)\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556)</a>, hosts are looking into providing more info for training as well as during testing. </p>\n<p>As you know the final leaderboard will involve running your notebooks for future targets from June to Sep 2025. </p>\n<p>As hosts,  we need to decide what future information we will retrieve for your notebooks in that last phase -- perhaps you want something more than MSA's? </p>\n<p>Ideally we'd compile this information for current train data sets so you can make sure your notebooks can use the information. </p>\n<p>We have some questions for the Kaggle community.</p>\n<ul>\n<li><p>We are preparing MSA's for train data v2 We are currently using rMSA, but perhaps you have found a better one and would like it run for future targets. Does anyone have pipelines set up that produce better MSA's? </p></li>\n<li><p>Is anyone exploring the direct identification of template structures? If so, can you share a pipeline? We can look into providing templates for train data and also supplying templates for future targets.</p></li>\n<li><p>Would any other information be useful for prediction that doesn't require actual knowledge of the RNA 3D structure?</p></li>\n</ul>",
      "rawMarkdown": "Hey everyone, in addition to @ckjoshi9 's repo (https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556), hosts are looking into providing more info for training as well as during testing. \n\nAs you know the final leaderboard will involve running your notebooks for future targets from June to Sep 2025. \n\nAs hosts,  we need to decide what future information we will retrieve for your notebooks in that last phase -- perhaps you want something more than MSA's? \n\nIdeally we'd compile this information for current train data sets so you can make sure your notebooks can use the information. \n\nWe have some questions for the Kaggle community.\n\n- We are preparing MSA's for train data v2 We are currently using rMSA, but perhaps you have found a better one and would like it run for future targets. Does anyone have pipelines set up that produce better MSA's? \n\n- Is anyone exploring the direct identification of template structures? If so, can you share a pipeline? We can look into providing templates for train data and also supplying templates for future targets.\n\n- Would any other information be useful for prediction that doesn't require actual knowledge of the RNA 3D structure?\n",
      "votes": 2
    },
    {
      "id": 3200552,
      "postDate": "2025-05-12T18:27:13.003Z",
      "content": "<p>Have you made any progress best of luck looking forward with your team</p>",
      "rawMarkdown": "Have you made any progress best of luck looking forward with your team\n"
    },
    {
      "id": 3195209,
      "postDate": "2025-05-06T18:59:48.753Z",
      "content": "<p>Great! Looking forward for the MSAs too1</p>",
      "rawMarkdown": "Great! Looking forward for the MSAs too1"
    },
    {
      "id": 3194344,
      "postDate": "2025-05-05T18:10:14.947Z",
      "content": "<p>Here's some more instructions on how you can use all known RNA 3D structures clustered by sequence and structural similarity to create your own datasets!</p>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556</a></p>",
      "rawMarkdown": "Here's some more instructions on how you can use all known RNA 3D structures clustered by sequence and structural similarity to create your own datasets!\n\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556"
    },
    {
      "id": 3192895,
      "postDate": "2025-05-03T13:58:49.190Z",
      "content": "<p>Hi do you happen to have the secondary strutures  of the data ?</p>",
      "rawMarkdown": "Hi do you happen to have the secondary strutures  of the data ?"
    },
    {
      "id": 3188836,
      "postDate": "2025-04-28T09:28:09.827Z",
      "content": "<p>Hi thanks for sharing,  the inference notebook: <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference</a> does not have /kaggle/input/ribonanzanet2-ddpm-v2/diffusion_config.yaml in it. </p>",
      "rawMarkdown": "Hi thanks for sharing,  the inference notebook: https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference does not have /kaggle/input/ribonanzanet2-ddpm-v2/diffusion_config.yaml in it. ",
      "replies": [
        {
          "id": 3189260,
          "postDate": "2025-04-29T03:15:45.887Z",
          "content": "<p>fixed it. forgot to make the dataset public</p>",
          "rawMarkdown": "fixed it. forgot to make the dataset public",
          "votes": 1,
          "replies": [
            {
              "id": 3189656,
              "postDate": "2025-04-29T14:43:27.560Z",
              "content": "<p>Ok , thank you</p>",
              "rawMarkdown": "Ok , thank you"
            }
          ]
        }
      ]
    },
    {
      "id": 3188501,
      "postDate": "2025-04-27T17:47:56.143Z",
      "content": "<p>Me and my teammates are new to the competition, so are we now exclusively using the v2 data instead of v1?</p>",
      "rawMarkdown": "Me and my teammates are new to the competition, so are we now exclusively using the v2 data instead of v1?"
    },
    {
      "id": 3187698,
      "postDate": "2025-04-26T12:03:29.813Z",
      "content": "<p>v1 data has MSA result by rMSA, will  v2 add and update?</p>",
      "rawMarkdown": "v1 data has MSA result by rMSA, will  v2 add and update?",
      "replies": [
        {
          "id": 3187733,
          "postDate": "2025-04-26T12:51:15.820Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> Have you resubmitted your old notebook and gotten a score? as w I tried to resubmit my old one I g I get susubmission scoring error ? i tried many thing like sorting the id etc but has issues</p>",
          "rawMarkdown": "@lihaoweicvch Have you resubmitted your old notebook and gotten a score? as w I tried to resubmit my old one I g I get susubmission scoring error ? i tried many thing like sorting the id etc but has issues",
          "replies": [
            {
              "id": 3188054,
              "postDate": "2025-04-27T01:37:59.280Z",
              "content": "<p>MSA's for train v2 are not ready yet but will be soon</p>",
              "rawMarkdown": "MSA's for train v2 are not ready yet but will be soon"
            },
            {
              "id": 3188069,
              "postDate": "2025-04-27T02:18:12.913Z",
              "content": "<p><a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a>  yes, I tried some old  notebooks,  I got timeout sometimes,  no other submission scoring error,  maybe check if your format strictly  match sample submission.</p>",
              "rawMarkdown": "@arunodhayan  yes, I tried some old  notebooks,  I got timeout sometimes,  no other submission scoring error,  maybe check if your format strictly  match sample submission."
            },
            {
              "id": 3199493,
              "postDate": "2025-05-11T04:26:08.903Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3199494,
              "postDate": "2025-05-11T04:26:21.517Z",
              "content": "<p>Will there be an update for v2 of the MSA?</p>",
              "rawMarkdown": "Will there be an update for v2 of the MSA?"
            },
            {
              "id": 3200404,
              "postDate": "2025-05-12T14:25:37.673Z",
              "content": "<p>Yes, just posted: <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578548\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578548</a></p>",
              "rawMarkdown": "Yes, just posted: https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578548"
            }
          ]
        }
      ]
    },
    {
      "id": 3187598,
      "postDate": "2025-04-26T09:02:09.647Z",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> after new testset , my old script i try to rerun and submit with the test_sequences.csv present i get submission scoring error where as same script with previous testset has scored results,  what should i do ?  </p>\n<p>Notebook Rerun for Submission [] on 04/25/2025  - it took my worst scored notebook and provided the results , it didnt select the top 2 notebooks</p>",
      "rawMarkdown": "@shujun717 after new testset , my old script i try to rerun and submit with the test_sequences.csv present i get submission scoring error where as same script with previous testset has scored results,  what should i do ?  \n\nNotebook Rerun for Submission [] on 04/25/2025  - it took my worst scored notebook and provided the results , it didnt select the top 2 notebooks",
      "replies": [
        {
          "id": 3188083,
          "postDate": "2025-04-27T02:53:05.003Z",
          "content": "<p>new test dataset takes longer but not much, nothing wrong with my old scirpt when inferencing this time.</p>\n<p>Maybe you can use some train dataset to test locally and see what's wrong?</p>",
          "rawMarkdown": "new test dataset takes longer but not much, nothing wrong with my old scirpt when inferencing this time.\n\nMaybe you can use some train dataset to test locally and see what's wrong?"
        }
      ]
    },
    {
      "id": 3187485,
      "postDate": "2025-04-26T05:10:16.267Z",
      "content": "<p>Let's forget for a moment about this competition or winning it.</p>\n<p>Forget about how little or how much data we have.</p>\n<p>Forget about the fact that gazillion-size large data is a typical requirement for deep learning to work well.</p>\n<p>Heck, let's even forget for a moment about machine learning altogether.</p>\n<p>From a purely \"scientific discovery\" and \"novel RNA applications for human benefits and such\" point of view, what is the difference between \"structured RNA\" and the complement of it?</p>\n<p>Surely, the original filtered set of 856 RNAs were filtered for a reason, since otherwise we'd have been given a lot more on day 1.</p>\n<p>As an RNA scientist, without regard to machine learning or improving a score, what is that <em>reason</em>?</p>",
      "rawMarkdown": "Let's forget for a moment about this competition or winning it.\n\nForget about how little or how much data we have.\n\nForget about the fact that gazillion-size large data is a typical requirement for deep learning to work well.\n\nHeck, let's even forget for a moment about machine learning altogether.\n\nFrom a purely \"scientific discovery\" and \"novel RNA applications for human benefits and such\" point of view, what is the difference between \"structured RNA\" and the complement of it?\n\nSurely, the original filtered set of 856 RNAs were filtered for a reason, since otherwise we'd have been given a lot more on day 1.\n\nAs an RNA scientist, without regard to machine learning or improving a score, what is that *reason*?",
      "replies": [
        {
          "id": 3187791,
          "postDate": "2025-04-26T14:45:44.367Z",
          "content": "<p>The original data set was curated a bit conservatively and looked for the word <code>RNA</code> in the title of the PDB entry. </p>\n<p>The newer set is expanded to include PDB entries with the word <code>RNA</code> anywhere in the file along with some checks that the extracted chains are RNA. We just wanted to be careful before making the newer data available.</p>",
          "rawMarkdown": "The original data set was curated a bit conservatively and looked for the word `RNA` in the title of the PDB entry. \n\nThe newer set is expanded to include PDB entries with the word `RNA` anywhere in the file along with some checks that the extracted chains are RNA. We just wanted to be careful before making the newer data available.",
          "replies": [
            {
              "id": 3188504,
              "postDate": "2025-04-27T17:57:56.093Z",
              "content": "<p>The collection logic of the final test set of up to 40 RNAs will be closer (even if not identical) to the newer training set curation than it is to the original training set curation.</p>\n<p>Is it reasonable to assume the statement above is true?</p>\n<p>As test set size is naturally small for RNA, I'd like to minimize any potential for sloppiness on my part in designing my training methodology.</p>",
              "rawMarkdown": "The collection logic of the final test set of up to 40 RNAs will be closer (even if not identical) to the newer training set curation than it is to the original training set curation.\n\nIs it reasonable to assume the statement above is true?\n\nAs test set size is naturally small for RNA, I'd like to minimize any potential for sloppiness on my part in designing my training methodology."
            },
            {
              "id": 3189244,
              "postDate": "2025-04-29T02:23:57.710Z",
              "content": "<p>Yes correct.</p>",
              "rawMarkdown": "Yes correct."
            },
            {
              "id": 3192126,
              "postDate": "2025-05-02T13:32:09.193Z",
              "content": "<p>I am seeking clarification on the following statement:</p>\n<pre><code>Only notebooks that beat  CASP16 expert baseline  also paying  attention  temporal_cutoff will be eligible   Early Sharing prizes!\n</code></pre>\n<p>Does this rule apply to the early sharing prize only?<br>\nIf I prompt LLMs that were released recently, and submit for the final private leaderboard, it is ok right?</p>",
              "rawMarkdown": "I am seeking clarification on the following statement:\n\n```\nOnly notebooks that beat the CASP16 expert baseline while also paying close attention to temporal_cutoff will be eligible for the Early Sharing prizes!\n```\nDoes this rule apply to the early sharing prize only?\nIf I prompt LLMs that were released recently, and submit for the final private leaderboard, it is ok right?"
            },
            {
              "id": 3192449,
              "postDate": "2025-05-02T20:53:41.583Z",
              "content": "<p>Yes the <code>temporal_cutoff</code> is just for Early Sharing prize. If you compete for final private leaderboard, your notebook will be run on targets released in public databases between June and September 2025 -- after final notebook deadline of May 29, 2025.</p>\n<p>Note: if you'd like to use recent LLM's even for Early Sharing prize, this would be allowed. Your notebook could prompt the LLM to ask for answers as if they are responding the day before the <code>temporal_cutoff</code> for each target. (We as hosts would double-check the LLM answers if you make your notebook public in a bid for Early Sharing prize.)</p>",
              "rawMarkdown": "Yes the `temporal_cutoff` is just for Early Sharing prize. If you compete for final private leaderboard, your notebook will be run on targets released in public databases between June and September 2025 -- after final notebook deadline of May 29, 2025.\n\nNote: if you'd like to use recent LLM's even for Early Sharing prize, this would be allowed. Your notebook could prompt the LLM to ask for answers as if they are responding the day before the `temporal_cutoff` for each target. (We as hosts would double-check the LLM answers if you make your notebook public in a bid for Early Sharing prize.)"
            },
            {
              "id": 3192452,
              "postDate": "2025-05-02T21:02:22.360Z",
              "content": "<p>Thanks for the clarification </p>",
              "rawMarkdown": "Thanks for the clarification "
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3187508,
      "author_name": "Shuxian Zou",
      "author_url": "",
      "post_date": "2025-04-26T06:05:38.520000",
      "content": "<p>Thanks a lot for updating the training data!</p>\n<p>Could you share the training data preprocessing pipeline as well like <a href=\"https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing\" target=\"_blank\">https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3187512,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2025-04-26T06:14:55.253000",
          "content": "<p>yes it's in the v2 branch</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3194413,
      "author_name": "Rhiju Das",
      "author_url": "",
      "post_date": "2025-05-05T20:59:46.523000",
      "content": "<p>Hey everyone, in addition to <a href=\"https://www.kaggle.com/ckjoshi9\" target=\"_blank\">@ckjoshi9</a> 's repo (<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556)\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556)</a>, hosts are looking into providing more info for training as well as during testing. </p>\n<p>As you know the final leaderboard will involve running your notebooks for future targets from June to Sep 2025. </p>\n<p>As hosts,  we need to decide what future information we will retrieve for your notebooks in that last phase -- perhaps you want something more than MSA's? </p>\n<p>Ideally we'd compile this information for current train data sets so you can make sure your notebooks can use the information. </p>\n<p>We have some questions for the Kaggle community.</p>\n<ul>\n<li><p>We are preparing MSA's for train data v2 We are currently using rMSA, but perhaps you have found a better one and would like it run for future targets. Does anyone have pipelines set up that produce better MSA's? </p></li>\n<li><p>Is anyone exploring the direct identification of template structures? If so, can you share a pipeline? We can look into providing templates for train data and also supplying templates for future targets.</p></li>\n<li><p>Would any other information be useful for prediction that doesn't require actual knowledge of the RNA 3D structure?</p></li>\n</ul>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3200552,
      "author_name": "nikhil Goswami0174",
      "author_url": "",
      "post_date": "2025-05-12T18:27:13.003000",
      "content": "<p>Have you made any progress best of luck looking forward with your team</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3195209,
      "author_name": "Elena Angeli",
      "author_url": "",
      "post_date": "2025-05-06T18:59:48.753000",
      "content": "<p>Great! Looking forward for the MSAs too1</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3194344,
      "author_name": "Chaitanya Joshi",
      "author_url": "",
      "post_date": "2025-05-05T18:10:14.947000",
      "content": "<p>Here's some more instructions on how you can use all known RNA 3D structures clustered by sequence and structural similarity to create your own datasets!</p>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3192895,
      "author_name": "soltani abdellatif",
      "author_url": "",
      "post_date": "2025-05-03T13:58:49.190000",
      "content": "<p>Hi do you happen to have the secondary strutures  of the data ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3188836,
      "author_name": "soltani abdellatif",
      "author_url": "",
      "post_date": "2025-04-28T09:28:09.827000",
      "content": "<p>Hi thanks for sharing,  the inference notebook: <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference</a> does not have /kaggle/input/ribonanzanet2-ddpm-v2/diffusion_config.yaml in it. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3189260,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2025-04-29T03:15:45.887000",
          "content": "<p>fixed it. forgot to make the dataset public</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3189656,
              "author_name": "soltani abdellatif",
              "author_url": "",
              "post_date": "2025-04-29T14:43:27.560000",
              "content": "<p>Ok , thank you</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3188501,
      "author_name": "Max Chang_kK",
      "author_url": "",
      "post_date": "2025-04-27T17:47:56.143000",
      "content": "<p>Me and my teammates are new to the competition, so are we now exclusively using the v2 data instead of v1?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3187698,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2025-04-26T12:03:29.813000",
      "content": "<p>v1 data has MSA result by rMSA, will  v2 add and update?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3187733,
          "author_name": "Arunodhayan",
          "author_url": "",
          "post_date": "2025-04-26T12:51:15.820000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> Have you resubmitted your old notebook and gotten a score? as w I tried to resubmit my old one I g I get susubmission scoring error ? i tried many thing like sorting the id etc but has issues</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3188054,
              "author_name": "Shujun",
              "author_url": "",
              "post_date": "2025-04-27T01:37:59.280000",
              "content": "<p>MSA's for train v2 are not ready yet but will be soon</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3188069,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2025-04-27T02:18:12.913000",
              "content": "<p><a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a>  yes, I tried some old  notebooks,  I got timeout sometimes,  no other submission scoring error,  maybe check if your format strictly  match sample submission.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3199493,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-11T04:26:08.903000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3199494,
              "author_name": "DaoHe Liu",
              "author_url": "",
              "post_date": "2025-05-11T04:26:21.517000",
              "content": "<p>Will there be an update for v2 of the MSA?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3200404,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2025-05-12T14:25:37.673000",
              "content": "<p>Yes, just posted: <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578548\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/578548</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3187598,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2025-04-26T09:02:09.647000",
      "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> after new testset , my old script i try to rerun and submit with the test_sequences.csv present i get submission scoring error where as same script with previous testset has scored results,  what should i do ?  </p>\n<p>Notebook Rerun for Submission [] on 04/25/2025  - it took my worst scored notebook and provided the results , it didnt select the top 2 notebooks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3188083,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2025-04-27T02:53:05.003000",
          "content": "<p>new test dataset takes longer but not much, nothing wrong with my old scirpt when inferencing this time.</p>\n<p>Maybe you can use some train dataset to test locally and see what's wrong?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3187485,
      "author_name": "Truth Seeker",
      "author_url": "",
      "post_date": "2025-04-26T05:10:16.267000",
      "content": "<p>Let's forget for a moment about this competition or winning it.</p>\n<p>Forget about how little or how much data we have.</p>\n<p>Forget about the fact that gazillion-size large data is a typical requirement for deep learning to work well.</p>\n<p>Heck, let's even forget for a moment about machine learning altogether.</p>\n<p>From a purely \"scientific discovery\" and \"novel RNA applications for human benefits and such\" point of view, what is the difference between \"structured RNA\" and the complement of it?</p>\n<p>Surely, the original filtered set of 856 RNAs were filtered for a reason, since otherwise we'd have been given a lot more on day 1.</p>\n<p>As an RNA scientist, without regard to machine learning or improving a score, what is that <em>reason</em>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3187791,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2025-04-26T14:45:44.367000",
          "content": "<p>The original data set was curated a bit conservatively and looked for the word <code>RNA</code> in the title of the PDB entry. </p>\n<p>The newer set is expanded to include PDB entries with the word <code>RNA</code> anywhere in the file along with some checks that the extracted chains are RNA. We just wanted to be careful before making the newer data available.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3188504,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2025-04-27T17:57:56.093000",
              "content": "<p>The collection logic of the final test set of up to 40 RNAs will be closer (even if not identical) to the newer training set curation than it is to the original training set curation.</p>\n<p>Is it reasonable to assume the statement above is true?</p>\n<p>As test set size is naturally small for RNA, I'd like to minimize any potential for sloppiness on my part in designing my training methodology.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3189244,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2025-04-29T02:23:57.710000",
              "content": "<p>Yes correct.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3192126,
              "author_name": "JK-Piece",
              "author_url": "",
              "post_date": "2025-05-02T13:32:09.193000",
              "content": "<p>I am seeking clarification on the following statement:</p>\n<pre><code>Only notebooks that beat  CASP16 expert baseline  also paying  attention  temporal_cutoff will be eligible   Early Sharing prizes!\n</code></pre>\n<p>Does this rule apply to the early sharing prize only?<br>\nIf I prompt LLMs that were released recently, and submit for the final private leaderboard, it is ok right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3192449,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2025-05-02T20:53:41.583000",
              "content": "<p>Yes the <code>temporal_cutoff</code> is just for Early Sharing prize. If you compete for final private leaderboard, your notebook will be run on targets released in public databases between June and September 2025 -- after final notebook deadline of May 29, 2025.</p>\n<p>Note: if you'd like to use recent LLM's even for Early Sharing prize, this would be allowed. Your notebook could prompt the LLM to ask for answers as if they are responding the day before the <code>temporal_cutoff</code> for each target. (We as hosts would double-check the LLM answers if you make your notebook public in a bid for Early Sharing prize.)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3192452,
              "author_name": "JK-Piece",
              "author_url": "",
              "post_date": "2025-05-02T21:02:22.360000",
              "content": "<p>Thanks for the clarification </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3187439": "Hi Kagglers, we have released an updated version of training data with more relaxed filters. The main changes are \n1.  we downloaded pdbs from the protein data bank with full text search for keyword RNA\n2. we relaxed filter for unstructured RNAs based on pairwise C1' distances, where 20% of residues have to be close to some other residue that is over 4 bases apart\n\nAs a result, we now have >5000 sequences in total and around 3300 before the CASP15 temporal cutoff (2022-05-01). The new files are in competition training data named ```train_sequences.v2.csv``` and ```train_labels.v2.csv```\n\nI have done a rerun of the previously releasesd Rnet2 3D finetuning notebook and score improved to around 0.27 on the public leaderboard. I have updated the inference notebook: [https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference](https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference)\n\nHave fun modeling! \n",
    "3187508": "Thanks a lot for updating the training data!\n\nCould you share the training data preprocessing pipeline as well like https://github.com/Shujun-He/Stanford3Dfolding_dataprocessing ",
    "3194413": "Hey everyone, in addition to @ckjoshi9 's repo (https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556), hosts are looking into providing more info for training as well as during testing. \n\nAs you know the final leaderboard will involve running your notebooks for future targets from June to Sep 2025. \n\nAs hosts,  we need to decide what future information we will retrieve for your notebooks in that last phase -- perhaps you want something more than MSA's? \n\nIdeally we'd compile this information for current train data sets so you can make sure your notebooks can use the information. \n\nWe have some questions for the Kaggle community.\n\n- We are preparing MSA's for train data v2 We are currently using rMSA, but perhaps you have found a better one and would like it run for future targets. Does anyone have pipelines set up that produce better MSA's? \n\n- Is anyone exploring the direct identification of template structures? If so, can you share a pipeline? We can look into providing templates for train data and also supplying templates for future targets.\n\n- Would any other information be useful for prediction that doesn't require actual knowledge of the RNA 3D structure?\n",
    "3200552": "Have you made any progress best of luck looking forward with your team\n",
    "3195209": "Great! Looking forward for the MSAs too1",
    "3194344": "Here's some more instructions on how you can use all known RNA 3D structures clustered by sequence and structural similarity to create your own datasets!\n\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/576556",
    "3192895": "Hi do you happen to have the secondary strutures  of the data ?",
    "3188836": "Hi thanks for sharing,  the inference notebook: https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference does not have /kaggle/input/ribonanzanet2-ddpm-v2/diffusion_config.yaml in it. ",
    "3188501": "Me and my teammates are new to the competition, so are we now exclusively using the v2 data instead of v1?",
    "3187698": "v1 data has MSA result by rMSA, will  v2 add and update?",
    "3187598": "@shujun717 after new testset , my old script i try to rerun and submit with the test_sequences.csv present i get submission scoring error where as same script with previous testset has scored results,  what should i do ?  \n\nNotebook Rerun for Submission [] on 04/25/2025  - it took my worst scored notebook and provided the results , it didnt select the top 2 notebooks",
    "3187485": "Let's forget for a moment about this competition or winning it.\n\nForget about how little or how much data we have.\n\nForget about the fact that gazillion-size large data is a typical requirement for deep learning to work well.\n\nHeck, let's even forget for a moment about machine learning altogether.\n\nFrom a purely \"scientific discovery\" and \"novel RNA applications for human benefits and such\" point of view, what is the difference between \"structured RNA\" and the complement of it?\n\nSurely, the original filtered set of 856 RNAs were filtered for a reason, since otherwise we'd have been given a lot more on day 1.\n\nAs an RNA scientist, without regard to machine learning or improving a score, what is that *reason*?"
  }
}