{
  "id": 565064,
  "title": "Welcome to the Stanford RNA 3D Folding Challenge!",
  "url": "/competitions/stanford-rna-3d-folding/discussion/565064",
  "author_name": "Rhiju Das",
  "post_date": "2025-02-26T14:51:20.021000",
  "votes": 50,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Slightly more than a year ago, Kaggle set out to solve one of the biggest challenges remaining in biology.</p>\n<p>RNA chains are the basis for new medicines and the oldest forms of life — and the most important RNAs fold up into beautiful three-dimensional structures, which underlie their functions.  </p>\n<p>Unfortunately, the world's efforts to advance biology and biotechnology are being slowed down by our inability to computationally predict these RNA 3D structures.</p>\n<p>In 2023-24, Kaggle’s <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Ribonanza challenge</a> elicited models from ‘chemical mapping’ profiles that provide indirect readouts of RNA structure. And a foundation model <a href=\"https://www.biorxiv.org/content/10.1101/2024.02.24.581671v2\" target=\"_blank\">RibonanzaNet</a> – itself distilled from Kagglers’ top models – appears state-of-the-art in most RNA modeling tasks for which it’s been fine-tuned.</p>\n<p>However, we and other expert groups have still not figured out how to take RibonanzaNet–or any other models–and accurately predict RNA structures in three dimensions. </p>\n<p>In fact, in the most recent community-wide <a href=\"https://predictioncenter.org/casp16\" target=\"_blank\">Critical Assessment of Structure Prediction</a> and <a href=\"https://www.nature.com/articles/s41592-024-02543-9\" target=\"_blank\">RNA-Puzzles</a> structure prediction challenges, human predictors using semiautomated methods remain better than any fully automated model!</p>\n<p>In this competition, we are seeking your help to achieve a critical breakthrough  in this problem: <strong>to achieve an automated RNA 3D structure prediction method that can outperform human experts</strong>.</p>\n<p>There are many strategies to explore. Some ideas:</p>\n<ul>\n<li>Are there better ways to fine-tune RibonanzaNet to predict 3D structures?</li>\n<li>Humans detect information on ‘multiple sequence alignments’ of RNA that current automated methods can’t — how do we get artificial neural networks to see these patterns?</li>\n<li>Human experts have a rich knowledge of the scientific literature on prior RNA structures – can neural nets achieve and leverage this written knowledge too?</li>\n<li>For RNA's with multiple structures, are there better ways to generate 5 predictions than just different random seeds?</li>\n<li>Human experts invest substantial efforts at inference time, sometimes creating 1000s of 'decoy' structures per RNA target and then using reasoning to iterate to additional structures -- can't that be automated?</li>\n<li>Can synthetic RNA structures be used to augment model training?</li>\n</ul>\n<p>And of course there’s more that we haven't even thought of!</p>\n<p>We’re looking forward to seeing what Kagglers come up with. We are especially excited to meet first-time Kagglers, whose fresh insights and high-placing models have had a huge impact in prior RNA competitions.</p>\n<p>In addition, we'll be releasing some Easter eggs throughout the competition–and incentivizing sharing with Early Sharing prizes based on who beats the baseline first!–so make sure to stay tuned to these forums.</p>\n<p>Let’s get to superhuman performance for RNA structure prediction! See you in the challenge!</p>\n<p>Your hosts<br>\nRhiju Das&nbsp;@rhijudas and Shujun He&nbsp;@shujun717</p>\n<p>Extra thanks to: </p>\n<ul>\n<li>Members of the Das lab providing new structures and analysis,  <a href=\"https://www.kaggle.com/rkretsch\" target=\"_blank\">@rkretsch</a> and <a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a></li>\n<li>CASP16 organizers, <a href=\"https://www.kaggle.com/andriyca\" target=\"_blank\">@andriyca</a> and John Moult</li>\n<li>RNA-Puzzles organizers Chichau Miao and Eric Westhof</li>\n<li>The global RNA structural biology community providing blind prediction targets </li>\n<li>The RNA RFdiffusion/MPNN team at the Institute of Protein Design, Andrew Favor <a href=\"https://www.kaggle.com/andrewfavor\" target=\"_blank\">@andrewfavor</a>, Andrew Kubaney, and David Baker</li>\n<li>Shi-Jie Chen and the VFOLD team for providing <code>VFOLD_human_expert</code> baseline predictions.</li>\n</ul>\n<p>And a huge thanks to our long-time collaborators at Kaggle:<br>\nMaggie Demkin&nbsp;@maggiemd<br>\nInversion&nbsp;@inversion</p>",
  "messages": [
    {
      "id": 3134622,
      "postDate": "2025-02-26T14:51:20.020Z",
      "content": "<p>Slightly more than a year ago, Kaggle set out to solve one of the biggest challenges remaining in biology.</p>\n<p>RNA chains are the basis for new medicines and the oldest forms of life — and the most important RNAs fold up into beautiful three-dimensional structures, which underlie their functions.  </p>\n<p>Unfortunately, the world's efforts to advance biology and biotechnology are being slowed down by our inability to computationally predict these RNA 3D structures.</p>\n<p>In 2023-24, Kaggle’s <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Ribonanza challenge</a> elicited models from ‘chemical mapping’ profiles that provide indirect readouts of RNA structure. And a foundation model <a href=\"https://www.biorxiv.org/content/10.1101/2024.02.24.581671v2\" target=\"_blank\">RibonanzaNet</a> – itself distilled from Kagglers’ top models – appears state-of-the-art in most RNA modeling tasks for which it’s been fine-tuned.</p>\n<p>However, we and other expert groups have still not figured out how to take RibonanzaNet–or any other models–and accurately predict RNA structures in three dimensions. </p>\n<p>In fact, in the most recent community-wide <a href=\"https://predictioncenter.org/casp16\" target=\"_blank\">Critical Assessment of Structure Prediction</a> and <a href=\"https://www.nature.com/articles/s41592-024-02543-9\" target=\"_blank\">RNA-Puzzles</a> structure prediction challenges, human predictors using semiautomated methods remain better than any fully automated model!</p>\n<p>In this competition, we are seeking your help to achieve a critical breakthrough  in this problem: <strong>to achieve an automated RNA 3D structure prediction method that can outperform human experts</strong>.</p>\n<p>There are many strategies to explore. Some ideas:</p>\n<ul>\n<li>Are there better ways to fine-tune RibonanzaNet to predict 3D structures?</li>\n<li>Humans detect information on ‘multiple sequence alignments’ of RNA that current automated methods can’t — how do we get artificial neural networks to see these patterns?</li>\n<li>Human experts have a rich knowledge of the scientific literature on prior RNA structures – can neural nets achieve and leverage this written knowledge too?</li>\n<li>For RNA's with multiple structures, are there better ways to generate 5 predictions than just different random seeds?</li>\n<li>Human experts invest substantial efforts at inference time, sometimes creating 1000s of 'decoy' structures per RNA target and then using reasoning to iterate to additional structures -- can't that be automated?</li>\n<li>Can synthetic RNA structures be used to augment model training?</li>\n</ul>\n<p>And of course there’s more that we haven't even thought of!</p>\n<p>We’re looking forward to seeing what Kagglers come up with. We are especially excited to meet first-time Kagglers, whose fresh insights and high-placing models have had a huge impact in prior RNA competitions.</p>\n<p>In addition, we'll be releasing some Easter eggs throughout the competition–and incentivizing sharing with Early Sharing prizes based on who beats the baseline first!–so make sure to stay tuned to these forums.</p>\n<p>Let’s get to superhuman performance for RNA structure prediction! See you in the challenge!</p>\n<p>Your hosts<br>\nRhiju Das&nbsp;@rhijudas and Shujun He&nbsp;@shujun717</p>\n<p>Extra thanks to: </p>\n<ul>\n<li>Members of the Das lab providing new structures and analysis,  <a href=\"https://www.kaggle.com/rkretsch\" target=\"_blank\">@rkretsch</a> and <a href=\"https://www.kaggle.com/alissahummer\" target=\"_blank\">@alissahummer</a></li>\n<li>CASP16 organizers, <a href=\"https://www.kaggle.com/andriyca\" target=\"_blank\">@andriyca</a> and John Moult</li>\n<li>RNA-Puzzles organizers Chichau Miao and Eric Westhof</li>\n<li>The global RNA structural biology community providing blind prediction targets </li>\n<li>The RNA RFdiffusion/MPNN team at the Institute of Protein Design, Andrew Favor <a href=\"https://www.kaggle.com/andrewfavor\" target=\"_blank\">@andrewfavor</a>, Andrew Kubaney, and David Baker</li>\n<li>Shi-Jie Chen and the VFOLD team for providing <code>VFOLD_human_expert</code> baseline predictions.</li>\n</ul>\n<p>And a huge thanks to our long-time collaborators at Kaggle:<br>\nMaggie Demkin&nbsp;@maggiemd<br>\nInversion&nbsp;@inversion</p>",
      "rawMarkdown": "Slightly more than a year ago, Kaggle set out to solve one of the biggest challenges remaining in biology.\n\nRNA chains are the basis for new medicines and the oldest forms of life — and the most important RNAs fold up into beautiful three-dimensional structures, which underlie their functions.  \n\nUnfortunately, the world's efforts to advance biology and biotechnology are being slowed down by our inability to computationally predict these RNA 3D structures.\n\nIn 2023-24, Kaggle’s [Ribonanza challenge](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding) elicited models from ‘chemical mapping’ profiles that provide indirect readouts of RNA structure. And a foundation model [RibonanzaNet](https://www.biorxiv.org/content/10.1101/2024.02.24.581671v2) – itself distilled from Kagglers’ top models – appears state-of-the-art in most RNA modeling tasks for which it’s been fine-tuned.\n\nHowever, we and other expert groups have still not figured out how to take RibonanzaNet–or any other models–and accurately predict RNA structures in three dimensions. \n\nIn fact, in the most recent community-wide [Critical Assessment of Structure Prediction](https://predictioncenter.org/casp16) and [RNA-Puzzles](https://www.nature.com/articles/s41592-024-02543-9) structure prediction challenges, human predictors using semiautomated methods remain better than any fully automated model!\n\nIn this competition, we are seeking your help to achieve a critical breakthrough  in this problem: **to achieve an automated RNA 3D structure prediction method that can outperform human experts**.\n\nThere are many strategies to explore. Some ideas:\n\n- Are there better ways to fine-tune RibonanzaNet to predict 3D structures?\n- Humans detect information on ‘multiple sequence alignments’ of RNA that current automated methods can’t — how do we get artificial neural networks to see these patterns?\n- Human experts have a rich knowledge of the scientific literature on prior RNA structures – can neural nets achieve and leverage this written knowledge too?\n- For RNA's with multiple structures, are there better ways to generate 5 predictions than just different random seeds?\n- Human experts invest substantial efforts at inference time, sometimes creating 1000s of 'decoy' structures per RNA target and then using reasoning to iterate to additional structures -- can't that be automated?\n- Can synthetic RNA structures be used to augment model training?\n\nAnd of course there’s more that we haven't even thought of!\n\nWe’re looking forward to seeing what Kagglers come up with. We are especially excited to meet first-time Kagglers, whose fresh insights and high-placing models have had a huge impact in prior RNA competitions.\n\nIn addition, we'll be releasing some Easter eggs throughout the competition–and incentivizing sharing with Early Sharing prizes based on who beats the baseline first!–so make sure to stay tuned to these forums.\n\nLet’s get to superhuman performance for RNA structure prediction! See you in the challenge!\n\nYour hosts\nRhiju Das @rhijudas and Shujun He @shujun717\n\nExtra thanks to: \n- Members of the Das lab providing new structures and analysis,  @rkretsch and @alissahummer\n- CASP16 organizers, @andriyca and John Moult\n- RNA-Puzzles organizers Chichau Miao and Eric Westhof\n- The global RNA structural biology community providing blind prediction targets \n- The RNA RFdiffusion/MPNN team at the Institute of Protein Design, Andrew Favor @andrewfavor, Andrew Kubaney, and David Baker\n- Shi-Jie Chen and the VFOLD team for providing `VFOLD_human_expert` baseline predictions.\n\nAnd a huge thanks to our long-time collaborators at Kaggle:\nMaggie Demkin @maggiemd\nInversion @inversion",
      "votes": 50
    },
    {
      "id": 3138203,
      "postDate": "2025-03-02T10:46:20.290Z",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>\n<p>I have a question regarding the training data for this competition. This time, the dataset contains 844 sequences, whereas the previous competition had around 800,000 sequences. This seems like a significant reduction in data size.</p>\n<p>Could you share the reasoning behind this difference? Additionally, I would appreciate more details on how the training dataset was created.<br>\nThank you for your time and clarification.</p>\n<p>Best regards,</p>",
      "rawMarkdown": "@rhijudas @shujun717 \n\nI have a question regarding the training data for this competition. This time, the dataset contains 844 sequences, whereas the previous competition had around 800,000 sequences. This seems like a significant reduction in data size.\n\nCould you share the reasoning behind this difference? Additionally, I would appreciate more details on how the training dataset was created.\nThank you for your time and clarification.\n\nBest regards,",
      "votes": 10
    },
    {
      "id": 3213708,
      "postDate": "2025-05-30T10:54:40.163Z",
      "content": "<p>After the end of submission:</p>\n<p>I genuinely admire your group's contributions to this competition. Most kagglers, myself included, simply applied existing sota methods to compete. I consider myself merely a clumsy imitator—truth be told, our academic breakthroughs have been quite limited.</p>\n<p>I've followed all your papers, tweets, and even media coverage since the competition began. In my eyes, your group is the true victor of this challenge.</p>",
      "rawMarkdown": "After the end of submission:\n\nI genuinely admire your group's contributions to this competition. Most kagglers, myself included, simply applied existing sota methods to compete. I consider myself merely a clumsy imitator—truth be told, our academic breakthroughs have been quite limited.\n\nI've followed all your papers, tweets, and even media coverage since the competition began. In my eyes, your group is the true victor of this challenge.",
      "votes": 1
    },
    {
      "id": 3207660,
      "postDate": "2025-05-23T05:29:55.477Z",
      "content": "<p>Two month working hard, learnt a lot from this competition! </p>",
      "rawMarkdown": "Two month working hard, learnt a lot from this competition! ",
      "votes": 1,
      "replies": [
        {
          "id": 3209507,
          "postDate": "2025-05-26T00:13:11.953Z",
          "content": "<p>Hi, can you guide us on how to submit my predictions?</p>",
          "rawMarkdown": "Hi, can you guide us on how to submit my predictions?",
          "votes": 1,
          "replies": [
            {
              "id": 3209876,
              "postDate": "2025-05-26T12:46:44.120Z",
              "content": "<p>You can edit your notebook on kaggle, save version for your notebook first, then you can select the saved version csv to submit. </p>",
              "rawMarkdown": "You can edit your notebook on kaggle, save version for your notebook first, then you can select the saved version csv to submit. ",
              "votes": 2
            },
            {
              "id": 3210854,
              "postDate": "2025-05-27T18:47:24.230Z",
              "content": "<p>Hi Alex, I have tried this and keep encountering an error message, 'Submission scoring error'. Any tips on how to navigate this?</p>",
              "rawMarkdown": "Hi Alex, I have tried this and keep encountering an error message, 'Submission scoring error'. Any tips on how to navigate this?"
            }
          ]
        }
      ]
    },
    {
      "id": 3147146,
      "postDate": "2025-03-11T17:53:14.530Z",
      "content": "<p>Hello,<br>\nI was wondering if we can use predictions from existing methods like AlphaFold 3 or RhoFold during the inference process? <br>\nIf so, are we supposed to install those tools in the notebook? </p>",
      "rawMarkdown": "Hello,\nI was wondering if we can use predictions from existing methods like AlphaFold 3 or RhoFold during the inference process? \nIf so, are we supposed to install those tools in the notebook? ",
      "votes": 3,
      "replies": [
        {
          "id": 3147149,
          "postDate": "2025-03-11T17:57:42.070Z",
          "content": "<p>Yes you should but I don’t think you can</p>",
          "rawMarkdown": "Yes you should but I don’t think you can",
          "replies": [
            {
              "id": 3147213,
              "postDate": "2025-03-11T19:01:56.517Z",
              "content": "<p>Distill from alphafold etc</p>\n<p>qoute above: \"And a foundation model RibonanzaNet – itself distilled from Kagglers’ top models – appears state-of-the-art in most RNA modeling tasks for which it’s been fine-tune …\"</p>\n<p>Deepseek r1 qwen models is another example to show distillation work</p>",
              "rawMarkdown": "Distill from alphafold etc\n\nqoute above: \"And a foundation model RibonanzaNet – itself distilled from Kagglers’ top models – appears state-of-the-art in most RNA modeling tasks for which it’s been fine-tune ...\"\n\nDeepseek r1 qwen models is another example to show distillation work"
            },
            {
              "id": 3147219,
              "postDate": "2025-03-11T19:11:21.237Z",
              "content": "<p>But how are we supposed to use AF3 in the notebook? Or even models that need MSAs as inputs? We can't actually in the notebook. </p>",
              "rawMarkdown": "But how are we supposed to use AF3 in the notebook? Or even models that need MSAs as inputs? We can't actually in the notebook. "
            },
            {
              "id": 3147229,
              "postDate": "2025-03-11T19:22:52.177Z",
              "content": "<p>MSA is my current experiment. you can</p>\n<ul>\n<li>used processed/compressed/lite database</li>\n<li>LLM to predict MSA or evolution vector</li>\n<li>maybe secondary structure help</li>\n</ul>\n<p>ESM shows that if you have many many many and many train samples, MSA is not required (at least of inlier data). ESM has comparable performance as (old) alphafold </p>",
              "rawMarkdown": "MSA is my current experiment. you can\n- used processed/compressed/lite database\n- LLM to predict MSA or evolution vector\n- maybe secondary structure help\n\nESM shows that if you have many many many and many train samples, MSA is not required (at least of inlier data). ESM has comparable performance as (old) alphafold "
            },
            {
              "id": 3147237,
              "postDate": "2025-03-11T19:34:53.267Z",
              "content": "<p>Yes but in the requirements of the notebook, it is said that:\"Internet access disabled\"<br>\nHow are you supposed to use let's say AF3? Upload from local installation? </p>",
              "rawMarkdown": "Yes but in the requirements of the notebook, it is said that:\"Internet access disabled\"\nHow are you supposed to use let's say AF3? Upload from local installation? "
            },
            {
              "id": 3147283,
              "postDate": "2025-03-11T21:32:06.940Z",
              "content": "<p>One more thing to add to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's insightful tips. </p>\n<p>Outputs of AF3 (and it's server) are great to use for benchmarking…. but the licensing for AlphaFold3 prevents use of its output for training new models.  🙃 </p>\n<p>Luckily, AlphaFold3 reproductions like <a href=\"https://github.com/chaidiscovery/chai-lab\" target=\"_blank\">Chai-1</a> and <a href=\"https://github.com/jwohlwend/boltz\" target=\"_blank\">Boltz-1</a> are available under open source licenses. It is not yet clear whether these reproductions have achieved parity with AF3 -- if you find out, please do post!  </p>",
              "rawMarkdown": "One more thing to add to @hengck23 's insightful tips. \n\nOutputs of AF3 (and it's server) are great to use for benchmarking.... but the licensing for AlphaFold3 prevents use of its output for training new models.  🙃 \n\nLuckily, AlphaFold3 reproductions like [Chai-1](https://github.com/chaidiscovery/chai-lab) and [Boltz-1](https://github.com/jwohlwend/boltz) are available under open source licenses. It is not yet clear whether these reproductions have achieved parity with AF3 -- if you find out, please do post!  \n",
              "votes": 1
            },
            {
              "id": 3147322,
              "postDate": "2025-03-11T23:10:19.363Z",
              "content": "<p>i am using</p>\n<p><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">https://github.com/bytedance/Protenix</a><br>\nA trainable PyTorch reproduction of AlphaFold 3.<br>\nApache 2.0 License</p>\n<p>its casp15 RNA results (also for DNA, etc) matches deepmind AF3</p>\n<p>paper: <a href=\"https://www.biorxiv.org/content/10.1101/2025.01.08.631967v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.01.08.631967v1.full.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F93a5364c7386f7fcad0fb034179ba5a1%2FSelection_039.png?generation=1741734944290200&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "i am using\n\nhttps://github.com/bytedance/Protenix\nA trainable PyTorch reproduction of AlphaFold 3.\nApache 2.0 License\n\nits casp15 RNA results (also for DNA, etc) matches deepmind AF3\n\npaper: https://www.biorxiv.org/content/10.1101/2025.01.08.631967v1.full.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F93a5364c7386f7fcad0fb034179ba5a1%2FSelection_039.png?generation=1741734944290200&alt=media)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3175689,
      "postDate": "2025-04-10T13:39:25.157Z",
      "content": "<p>Good things :)</p>",
      "rawMarkdown": "Good things :)",
      "votes": 2
    },
    {
      "id": 3153793,
      "postDate": "2025-03-19T07:19:41.940Z",
      "content": "<p>Any useful links to understanding how humans predict RNA folding or books, especially for non-biochemists.</p>",
      "rawMarkdown": "Any useful links to understanding how humans predict RNA folding or books, especially for non-biochemists.",
      "votes": 2,
      "replies": [
        {
          "id": 3154124,
          "postDate": "2025-03-19T14:53:34.813Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ravimohanms\" target=\"_blank\">@ravimohanms</a> ,</p>\n<p>sorry, I dont say anything about books but you can have a look here:</p>\n<p><a href=\"https://www.youtube.com/watch?v=P_fHJIYENdI\" target=\"_blank\">https://www.youtube.com/watch?v=P_fHJIYENdI</a></p>\n<p>About RNA and data, I tried also to understand it <br>\ncheckout my nb -&gt; <a href=\"https://www.kaggle.com/code/dantheshark/rna-3d-folding-understand-data\" target=\"_blank\">https://www.kaggle.com/code/dantheshark/rna-3d-folding-understand-data</a></p>\n<p>Also check this big post please -&gt; <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565292\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565292</a></p>\n<p>Good luck and have fun :)</p>",
          "rawMarkdown": "Hi @ravimohanms ,\n\nsorry, I dont say anything about books but you can have a look here:\n\nhttps://www.youtube.com/watch?v=P_fHJIYENdI\n\nAbout RNA and data, I tried also to understand it \ncheckout my nb -> https://www.kaggle.com/code/dantheshark/rna-3d-folding-understand-data\n\n\nAlso check this big post please -> https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565292\n\nGood luck and have fun :)",
          "votes": 5,
          "replies": [
            {
              "id": 3175265,
              "postDate": "2025-04-10T01:00:21.540Z",
              "content": "<p>Thanks for this!</p>",
              "rawMarkdown": "Thanks for this!\n"
            }
          ]
        }
      ]
    },
    {
      "id": 3206884,
      "postDate": "2025-05-21T22:11:32.387Z",
      "content": "<p>………………</p>",
      "rawMarkdown": "..................",
      "votes": -1
    },
    {
      "id": 3175264,
      "postDate": "2025-04-10T01:00:00.317Z",
      "content": "<p>The idea of fine-tuning RibonanzaNet to predict 3D structures seems promising. Exciting challenge. </p>",
      "rawMarkdown": "The idea of fine-tuning RibonanzaNet to predict 3D structures seems promising. Exciting challenge. "
    },
    {
      "id": 3170476,
      "postDate": "2025-04-04T18:58:33.193Z",
      "content": "<p>is the VFOLD_human_expert baseline what won the CASP16? Or how was it obtained?</p>",
      "rawMarkdown": "is the VFOLD_human_expert baseline what won the CASP16? Or how was it obtained?",
      "replies": [
        {
          "id": 3175618,
          "postDate": "2025-04-10T11:47:55.847Z",
          "content": "<p>Yes the VFOLD team (Chen lab, U. Missouri) is collaborating with us. They have kindly agreed to make blind, confidential predictions for the leaderboard targets to set the human expert baseline for the Early Sharing Prize. We are all very much hoping to see progress against the baseline after the data refresh later in April 2025!</p>",
          "rawMarkdown": "Yes the VFOLD team (Chen lab, U. Missouri) is collaborating with us. They have kindly agreed to make blind, confidential predictions for the leaderboard targets to set the human expert baseline for the Early Sharing Prize. We are all very much hoping to see progress against the baseline after the data refresh later in April 2025!"
        }
      ]
    },
    {
      "id": 3170377,
      "postDate": "2025-04-04T16:45:27.800Z",
      "content": "<p>Good things :)</p>",
      "rawMarkdown": "Good things :)"
    },
    {
      "id": 3167337,
      "postDate": "2025-04-01T13:02:19.137Z",
      "content": "<p>Great competition!</p>",
      "rawMarkdown": "Great competition!"
    },
    {
      "id": 3153917,
      "postDate": "2025-03-19T10:10:23.383Z",
      "content": "<p>wow good job kkk</p>",
      "rawMarkdown": "wow good job kkk"
    },
    {
      "id": 3146650,
      "postDate": "2025-03-11T05:08:35.667Z",
      "content": "<p>will there be distribution shift intended? Or dataset is generated under the same distribution?</p>",
      "rawMarkdown": "will there be distribution shift intended? Or dataset is generated under the same distribution?",
      "replies": [
        {
          "id": 3147276,
          "postDate": "2025-03-11T21:15:17.597Z",
          "content": "<p>It's a great question.</p>\n<p>The future data set will be based on what new RNA structures the world decides to release in the public database between May and September. </p>\n<p>We can expect the lengths to be similar to what is in the <code>validation_sequences.csv</code> as well as the currently hidden test sequences used for the public leaderboard.</p>\n<p>Other than that, we don't know -- but feel free to speculate about what new RNAs will have their 3D structures experimentally solved in upcoming months! </p>\n<p>You can take a look at very recent PDB structures with RNA chains for a guess:</p>\n<p><a href=\"https://www.rcsb.org/structure/8T5O\" target=\"_blank\">https://www.rcsb.org/structure/8T5O</a><br>\n<a href=\"https://www.rcsb.org/structure/9CXF\" target=\"_blank\">https://www.rcsb.org/structure/9CXF</a><br>\n<a href=\"https://www.rcsb.org/structure/9IS7\" target=\"_blank\">https://www.rcsb.org/structure/9IS7</a></p>",
          "rawMarkdown": "It's a great question.\n\nThe future data set will be based on what new RNA structures the world decides to release in the public database between May and September. \n\nWe can expect the lengths to be similar to what is in the `validation_sequences.csv` as well as the currently hidden test sequences used for the public leaderboard.\n\nOther than that, we don't know -- but feel free to speculate about what new RNAs will have their 3D structures experimentally solved in upcoming months! \n\nYou can take a look at very recent PDB structures with RNA chains for a guess:\n\nhttps://www.rcsb.org/structure/8T5O\nhttps://www.rcsb.org/structure/9CXF\nhttps://www.rcsb.org/structure/9IS7",
          "votes": 4,
          "replies": [
            {
              "id": 3163057,
              "postDate": "2025-03-30T08:39:00.407Z",
              "content": "<p>i think, this competision have purpose to develop AI from Suqquence to human disease</p>",
              "rawMarkdown": "i think, this competision have purpose to develop AI from Suqquence to human disease"
            }
          ]
        }
      ]
    },
    {
      "id": 3145401,
      "postDate": "2025-03-09T18:37:57.707Z",
      "content": "<p>Here we go</p>",
      "rawMarkdown": "Here we go"
    },
    {
      "id": 3143761,
      "postDate": "2025-03-07T15:00:55.653Z",
      "content": "<p>Hello,everyone<br>\nCould you please explain the difference between [train/validation/test]_sequences.csv and [train/validation]_labels.csv? I want to understand how they are used in the dataset ?</p>",
      "rawMarkdown": "Hello,everyone\nCould you please explain the difference between [train/validation/test]_sequences.csv and [train/validation]_labels.csv? I want to understand how they are used in the dataset ?",
      "replies": [
        {
          "id": 3146370,
          "postDate": "2025-03-10T19:27:01.227Z",
          "content": "<p>Can someone help me with this question, please?</p>",
          "rawMarkdown": "Can someone help me with this question, please?",
          "replies": [
            {
              "id": 3146984,
              "postDate": "2025-03-11T14:27:48.117Z",
              "content": "<p>*_sequences.csv contain the unfolded nucleotide sequence (A, C, G, U) of the RNA structure, whereas *_labels.csv contain the actual structure of the folded RNA by providing the x-,y- and z-coordinate of all the C1' atoms of all nucleotides in the sequence. So from what I understand, the sequence is the input for the model, and the labels are the output, i.e the structure of the folded RNA.</p>",
              "rawMarkdown": "*_sequences.csv contain the unfolded nucleotide sequence (A, C, G, U) of the RNA structure, whereas *_labels.csv contain the actual structure of the folded RNA by providing the x-,y- and z-coordinate of all the C1' atoms of all nucleotides in the sequence. So from what I understand, the sequence is the input for the model, and the labels are the output, i.e the structure of the folded RNA."
            },
            {
              "id": 3146996,
              "postDate": "2025-03-11T14:42:36.277Z",
              "content": "<p>thank you so much for pointing out that it's C1' atom. I have been extracting data from model like trRosetteRNA, Alphafold, etc. when I use those data to train my model the score drop maybe because I average out the coordinate of all the atom in a nucleotide.</p>",
              "rawMarkdown": "thank you so much for pointing out that it's C1' atom. I have been extracting data from model like trRosetteRNA, Alphafold, etc. when I use those data to train my model the score drop maybe because I average out the coordinate of all the atom in a nucleotide."
            },
            {
              "id": 3148732,
              "postDate": "2025-03-13T13:23:17.797Z",
              "content": "<p>Thank you so much</p>",
              "rawMarkdown": "Thank you so much"
            }
          ]
        }
      ]
    },
    {
      "id": 3143733,
      "postDate": "2025-03-07T14:36:03.410Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> .  I haven't run into this before, but when I attempt to attach the uw_synthetic_rna_structures dataset that is linked in the competition materials, I get the following error:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11924794%2F003f8b732d460b6ed557ef652d50f6d1%2Fcannot_submit_error.png?generation=1741358101013391&amp;alt=media\" alt=\"\"></p>\n<p>Just to clarify, we are able to use this dataset?  </p>",
      "rawMarkdown": "Hi @rhijudas .  I haven't run into this before, but when I attempt to attach the uw_synthetic_rna_structures dataset that is linked in the competition materials, I get the following error:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11924794%2F003f8b732d460b6ed557ef652d50f6d1%2Fcannot_submit_error.png?generation=1741358101013391&alt=media)\n\nJust to clarify, we are able to use this dataset?  ",
      "replies": [
        {
          "id": 3146312,
          "postDate": "2025-03-10T18:17:22.677Z",
          "content": "<p>Hi, you are totally able to use that dataset. The error that you've reported seems to be related to that dataset only having the latest version accessible. We have created (and now redirect to) an identical dataset, <a href=\"https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final/data\" target=\"_blank\">here</a>, which has only a single-version (but with \"all versions visible\"). <br>\nThis should fix the reported problem, but please let me know if any issues with the synthetic RNA dataset persist.</p>",
          "rawMarkdown": "Hi, you are totally able to use that dataset. The error that you've reported seems to be related to that dataset only having the latest version accessible. We have created (and now redirect to) an identical dataset, [here](https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final/data), which has only a single-version (but with \"all versions visible\"). \nThis should fix the reported problem, but please let me know if any issues with the synthetic RNA dataset persist.",
          "votes": 1,
          "replies": [
            {
              "id": 3146943,
              "postDate": "2025-03-11T13:01:12.413Z",
              "content": "<p><a href=\"https://www.kaggle.com/andrewfavor\" target=\"_blank\">@andrewfavor</a> , that seems to work now, thanks!</p>",
              "rawMarkdown": "@andrewfavor , that seems to work now, thanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 3141180,
      "postDate": "2025-03-05T09:47:57.430Z",
      "content": "<p>Thank you, this is my first competition, not an easy one to start but a very interesting challenge. </p>",
      "rawMarkdown": "Thank you, this is my first competition, not an easy one to start but a very interesting challenge. ",
      "replies": [
        {
          "id": 3141373,
          "postDate": "2025-03-05T14:04:07.113Z",
          "content": "<p>Yes, this challenge is very tough but you can always learn from others. This is a great place to start.</p>",
          "rawMarkdown": "Yes, this challenge is very tough but you can always learn from others. This is a great place to start."
        }
      ]
    },
    {
      "id": 3138637,
      "postDate": "2025-03-02T18:55:58.103Z",
      "content": "<p>Hi I'm new </p>\n<p>I saw that in train_label only has x1 y1 z1 but you want us to predict 5 set ? so how does that work I don't really get it please clarify. thanks in advance.</p>",
      "rawMarkdown": "Hi I'm new \n\nI saw that in train_label only has x1 y1 z1 but you want us to predict 5 set ? so how does that work I don't really get it please clarify. thanks in advance.",
      "replies": [
        {
          "id": 3138755,
          "postDate": "2025-03-02T21:50:46.470Z",
          "content": "<p>For example you can train 5 different models.<br>\nYou can also predict the same coordinates 5 times, but its less efficient.<br>\nThe 'predict 5' allows us more freedom/diversity.</p>",
          "rawMarkdown": "For example you can train 5 different models.\nYou can also predict the same coordinates 5 times, but its less efficient.\nThe 'predict 5' allows us more freedom/diversity.",
          "votes": 2,
          "replies": [
            {
              "id": 3139255,
              "postDate": "2025-03-03T10:06:55.987Z",
              "content": "<p>So basically allow us for more flexibility. I thought that I have to group data in set of 5 then predict 5 set of output</p>",
              "rawMarkdown": "So basically allow us for more flexibility. I thought that I have to group data in set of 5 then predict 5 set of output"
            },
            {
              "id": 3139831,
              "postDate": "2025-03-04T00:01:25.920Z",
              "content": "<p>Yeah since you have the 'predict 5' like greySnow mentioned, it will give you more flexibility in your model.</p>",
              "rawMarkdown": "Yeah since you have the 'predict 5' like greySnow mentioned, it will give you more flexibility in your model."
            }
          ]
        }
      ]
    },
    {
      "id": 3305592,
      "postDate": "2025-10-23T03:45:52.427Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3138203,
      "author_name": "213tubo",
      "author_url": "",
      "post_date": "2025-03-02T10:46:20.290000",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> </p>\n<p>I have a question regarding the training data for this competition. This time, the dataset contains 844 sequences, whereas the previous competition had around 800,000 sequences. This seems like a significant reduction in data size.</p>\n<p>Could you share the reasoning behind this difference? Additionally, I would appreciate more details on how the training dataset was created.<br>\nThank you for your time and clarification.</p>\n<p>Best regards,</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 3213708,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2025-05-30T10:54:40.163000",
      "content": "<p>After the end of submission:</p>\n<p>I genuinely admire your group's contributions to this competition. Most kagglers, myself included, simply applied existing sota methods to compete. I consider myself merely a clumsy imitator—truth be told, our academic breakthroughs have been quite limited.</p>\n<p>I've followed all your papers, tweets, and even media coverage since the competition began. In my eyes, your group is the true victor of this challenge.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3207660,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-23T05:29:55.477000",
      "content": "<p>Two month working hard, learnt a lot from this competition! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3209507,
          "author_name": "Obiageli Nwachukwu",
          "author_url": "",
          "post_date": "2025-05-26T00:13:11.953000",
          "content": "<p>Hi, can you guide us on how to submit my predictions?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3209876,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-26T12:46:44.120000",
              "content": "<p>You can edit your notebook on kaggle, save version for your notebook first, then you can select the saved version csv to submit. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3210854,
              "author_name": "Obiageli Nwachukwu",
              "author_url": "",
              "post_date": "2025-05-27T18:47:24.230000",
              "content": "<p>Hi Alex, I have tried this and keep encountering an error message, 'Submission scoring error'. Any tips on how to navigate this?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3147146,
      "author_name": "clemesdu77 clemesdu77",
      "author_url": "",
      "post_date": "2025-03-11T17:53:14.530000",
      "content": "<p>Hello,<br>\nI was wondering if we can use predictions from existing methods like AlphaFold 3 or RhoFold during the inference process? <br>\nIf so, are we supposed to install those tools in the notebook? </p>",
      "votes": 3,
      "replies": [
        {
          "id": 3147149,
          "author_name": "TheOneTheOnly",
          "author_url": "",
          "post_date": "2025-03-11T17:57:42.070000",
          "content": "<p>Yes you should but I don’t think you can</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3147213,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-11T19:01:56.517000",
              "content": "<p>Distill from alphafold etc</p>\n<p>qoute above: \"And a foundation model RibonanzaNet – itself distilled from Kagglers’ top models – appears state-of-the-art in most RNA modeling tasks for which it’s been fine-tune …\"</p>\n<p>Deepseek r1 qwen models is another example to show distillation work</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3147219,
              "author_name": "clemesdu77 clemesdu77",
              "author_url": "",
              "post_date": "2025-03-11T19:11:21.237000",
              "content": "<p>But how are we supposed to use AF3 in the notebook? Or even models that need MSAs as inputs? We can't actually in the notebook. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3147229,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-11T19:22:52.177000",
              "content": "<p>MSA is my current experiment. you can</p>\n<ul>\n<li>used processed/compressed/lite database</li>\n<li>LLM to predict MSA or evolution vector</li>\n<li>maybe secondary structure help</li>\n</ul>\n<p>ESM shows that if you have many many many and many train samples, MSA is not required (at least of inlier data). ESM has comparable performance as (old) alphafold </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3147237,
              "author_name": "clemesdu77 clemesdu77",
              "author_url": "",
              "post_date": "2025-03-11T19:34:53.267000",
              "content": "<p>Yes but in the requirements of the notebook, it is said that:\"Internet access disabled\"<br>\nHow are you supposed to use let's say AF3? Upload from local installation? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3147283,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2025-03-11T21:32:06.940000",
              "content": "<p>One more thing to add to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's insightful tips. </p>\n<p>Outputs of AF3 (and it's server) are great to use for benchmarking…. but the licensing for AlphaFold3 prevents use of its output for training new models.  🙃 </p>\n<p>Luckily, AlphaFold3 reproductions like <a href=\"https://github.com/chaidiscovery/chai-lab\" target=\"_blank\">Chai-1</a> and <a href=\"https://github.com/jwohlwend/boltz\" target=\"_blank\">Boltz-1</a> are available under open source licenses. It is not yet clear whether these reproductions have achieved parity with AF3 -- if you find out, please do post!  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3147322,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-11T23:10:19.363000",
              "content": "<p>i am using</p>\n<p><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">https://github.com/bytedance/Protenix</a><br>\nA trainable PyTorch reproduction of AlphaFold 3.<br>\nApache 2.0 License</p>\n<p>its casp15 RNA results (also for DNA, etc) matches deepmind AF3</p>\n<p>paper: <a href=\"https://www.biorxiv.org/content/10.1101/2025.01.08.631967v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.01.08.631967v1.full.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F93a5364c7386f7fcad0fb034179ba5a1%2FSelection_039.png?generation=1741734944290200&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3175689,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-10T13:39:25.157000",
      "content": "<p>Good things :)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3153793,
      "author_name": "Ravi Mohan",
      "author_url": "",
      "post_date": "2025-03-19T07:19:41.940000",
      "content": "<p>Any useful links to understanding how humans predict RNA folding or books, especially for non-biochemists.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3154124,
          "author_name": "dan",
          "author_url": "",
          "post_date": "2025-03-19T14:53:34.813000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ravimohanms\" target=\"_blank\">@ravimohanms</a> ,</p>\n<p>sorry, I dont say anything about books but you can have a look here:</p>\n<p><a href=\"https://www.youtube.com/watch?v=P_fHJIYENdI\" target=\"_blank\">https://www.youtube.com/watch?v=P_fHJIYENdI</a></p>\n<p>About RNA and data, I tried also to understand it <br>\ncheckout my nb -&gt; <a href=\"https://www.kaggle.com/code/dantheshark/rna-3d-folding-understand-data\" target=\"_blank\">https://www.kaggle.com/code/dantheshark/rna-3d-folding-understand-data</a></p>\n<p>Also check this big post please -&gt; <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565292\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565292</a></p>\n<p>Good luck and have fun :)</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3175265,
              "author_name": "Justin",
              "author_url": "",
              "post_date": "2025-04-10T01:00:21.540000",
              "content": "<p>Thanks for this!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3206884,
      "author_name": "Ines FILALii",
      "author_url": "",
      "post_date": "2025-05-21T22:11:32.387000",
      "content": "<p>………………</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3175264,
      "author_name": "Justin",
      "author_url": "",
      "post_date": "2025-04-10T01:00:00.317000",
      "content": "<p>The idea of fine-tuning RibonanzaNet to predict 3D structures seems promising. Exciting challenge. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3170476,
      "author_name": "weke",
      "author_url": "",
      "post_date": "2025-04-04T18:58:33.193000",
      "content": "<p>is the VFOLD_human_expert baseline what won the CASP16? Or how was it obtained?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3175618,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2025-04-10T11:47:55.847000",
          "content": "<p>Yes the VFOLD team (Chen lab, U. Missouri) is collaborating with us. They have kindly agreed to make blind, confidential predictions for the leaderboard targets to set the human expert baseline for the Early Sharing Prize. We are all very much hoping to see progress against the baseline after the data refresh later in April 2025!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3170377,
      "author_name": "루디브리엄주민",
      "author_url": "",
      "post_date": "2025-04-04T16:45:27.800000",
      "content": "<p>Good things :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3167337,
      "author_name": "Chinmay Gupta",
      "author_url": "",
      "post_date": "2025-04-01T13:02:19.137000",
      "content": "<p>Great competition!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3153917,
      "author_name": "KhoaNgo",
      "author_url": "",
      "post_date": "2025-03-19T10:10:23.383000",
      "content": "<p>wow good job kkk</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3146650,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2025-03-11T05:08:35.667000",
      "content": "<p>will there be distribution shift intended? Or dataset is generated under the same distribution?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3147276,
          "author_name": "Rhiju Das",
          "author_url": "",
          "post_date": "2025-03-11T21:15:17.597000",
          "content": "<p>It's a great question.</p>\n<p>The future data set will be based on what new RNA structures the world decides to release in the public database between May and September. </p>\n<p>We can expect the lengths to be similar to what is in the <code>validation_sequences.csv</code> as well as the currently hidden test sequences used for the public leaderboard.</p>\n<p>Other than that, we don't know -- but feel free to speculate about what new RNAs will have their 3D structures experimentally solved in upcoming months! </p>\n<p>You can take a look at very recent PDB structures with RNA chains for a guess:</p>\n<p><a href=\"https://www.rcsb.org/structure/8T5O\" target=\"_blank\">https://www.rcsb.org/structure/8T5O</a><br>\n<a href=\"https://www.rcsb.org/structure/9CXF\" target=\"_blank\">https://www.rcsb.org/structure/9CXF</a><br>\n<a href=\"https://www.rcsb.org/structure/9IS7\" target=\"_blank\">https://www.rcsb.org/structure/9IS7</a></p>",
          "votes": 4,
          "replies": [
            {
              "id": 3163057,
              "author_name": "mayanurrachmawati",
              "author_url": "",
              "post_date": "2025-03-30T08:39:00.407000",
              "content": "<p>i think, this competision have purpose to develop AI from Suqquence to human disease</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3145401,
      "author_name": "Engineer Wanga",
      "author_url": "",
      "post_date": "2025-03-09T18:37:57.707000",
      "content": "<p>Here we go</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3143761,
      "author_name": "Ali Ahmed",
      "author_url": "",
      "post_date": "2025-03-07T15:00:55.653000",
      "content": "<p>Hello,everyone<br>\nCould you please explain the difference between [train/validation/test]_sequences.csv and [train/validation]_labels.csv? I want to understand how they are used in the dataset ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3146370,
          "author_name": "Ali Ahmed",
          "author_url": "",
          "post_date": "2025-03-10T19:27:01.227000",
          "content": "<p>Can someone help me with this question, please?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3146984,
              "author_name": "Felix Basilisk",
              "author_url": "",
              "post_date": "2025-03-11T14:27:48.117000",
              "content": "<p>*_sequences.csv contain the unfolded nucleotide sequence (A, C, G, U) of the RNA structure, whereas *_labels.csv contain the actual structure of the folded RNA by providing the x-,y- and z-coordinate of all the C1' atoms of all nucleotides in the sequence. So from what I understand, the sequence is the input for the model, and the labels are the output, i.e the structure of the folded RNA.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3146996,
              "author_name": "TheOneTheOnly",
              "author_url": "",
              "post_date": "2025-03-11T14:42:36.277000",
              "content": "<p>thank you so much for pointing out that it's C1' atom. I have been extracting data from model like trRosetteRNA, Alphafold, etc. when I use those data to train my model the score drop maybe because I average out the coordinate of all the atom in a nucleotide.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3148732,
              "author_name": "Ali Ahmed",
              "author_url": "",
              "post_date": "2025-03-13T13:23:17.797000",
              "content": "<p>Thank you so much</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3143733,
      "author_name": "ChrisK321",
      "author_url": "",
      "post_date": "2025-03-07T14:36:03.410000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> .  I haven't run into this before, but when I attempt to attach the uw_synthetic_rna_structures dataset that is linked in the competition materials, I get the following error:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11924794%2F003f8b732d460b6ed557ef652d50f6d1%2Fcannot_submit_error.png?generation=1741358101013391&amp;alt=media\" alt=\"\"></p>\n<p>Just to clarify, we are able to use this dataset?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3146312,
          "author_name": "Andrew Favor",
          "author_url": "",
          "post_date": "2025-03-10T18:17:22.677000",
          "content": "<p>Hi, you are totally able to use that dataset. The error that you've reported seems to be related to that dataset only having the latest version accessible. We have created (and now redirect to) an identical dataset, <a href=\"https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-final/data\" target=\"_blank\">here</a>, which has only a single-version (but with \"all versions visible\"). <br>\nThis should fix the reported problem, but please let me know if any issues with the synthetic RNA dataset persist.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3146943,
              "author_name": "ChrisK321",
              "author_url": "",
              "post_date": "2025-03-11T13:01:12.413000",
              "content": "<p><a href=\"https://www.kaggle.com/andrewfavor\" target=\"_blank\">@andrewfavor</a> , that seems to work now, thanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3141180,
      "author_name": "Romain Lejeune",
      "author_url": "",
      "post_date": "2025-03-05T09:47:57.430000",
      "content": "<p>Thank you, this is my first competition, not an easy one to start but a very interesting challenge. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3141373,
          "author_name": "Swastik Nanda",
          "author_url": "",
          "post_date": "2025-03-05T14:04:07.113000",
          "content": "<p>Yes, this challenge is very tough but you can always learn from others. This is a great place to start.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3138637,
      "author_name": "TheOneTheOnly",
      "author_url": "",
      "post_date": "2025-03-02T18:55:58.103000",
      "content": "<p>Hi I'm new </p>\n<p>I saw that in train_label only has x1 y1 z1 but you want us to predict 5 set ? so how does that work I don't really get it please clarify. thanks in advance.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3138755,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2025-03-02T21:50:46.470000",
          "content": "<p>For example you can train 5 different models.<br>\nYou can also predict the same coordinates 5 times, but its less efficient.<br>\nThe 'predict 5' allows us more freedom/diversity.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3139255,
              "author_name": "TheOneTheOnly",
              "author_url": "",
              "post_date": "2025-03-03T10:06:55.987000",
              "content": "<p>So basically allow us for more flexibility. I thought that I have to group data in set of 5 then predict 5 set of output</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3139831,
              "author_name": "Swastik Nanda",
              "author_url": "",
              "post_date": "2025-03-04T00:01:25.920000",
              "content": "<p>Yeah since you have the 'predict 5' like greySnow mentioned, it will give you more flexibility in your model.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3305592,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-10-23T03:45:52.427000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3134622": "Slightly more than a year ago, Kaggle set out to solve one of the biggest challenges remaining in biology.\n\nRNA chains are the basis for new medicines and the oldest forms of life — and the most important RNAs fold up into beautiful three-dimensional structures, which underlie their functions.  \n\nUnfortunately, the world's efforts to advance biology and biotechnology are being slowed down by our inability to computationally predict these RNA 3D structures.\n\nIn 2023-24, Kaggle’s [Ribonanza challenge](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding) elicited models from ‘chemical mapping’ profiles that provide indirect readouts of RNA structure. And a foundation model [RibonanzaNet](https://www.biorxiv.org/content/10.1101/2024.02.24.581671v2) – itself distilled from Kagglers’ top models – appears state-of-the-art in most RNA modeling tasks for which it’s been fine-tuned.\n\nHowever, we and other expert groups have still not figured out how to take RibonanzaNet–or any other models–and accurately predict RNA structures in three dimensions. \n\nIn fact, in the most recent community-wide [Critical Assessment of Structure Prediction](https://predictioncenter.org/casp16) and [RNA-Puzzles](https://www.nature.com/articles/s41592-024-02543-9) structure prediction challenges, human predictors using semiautomated methods remain better than any fully automated model!\n\nIn this competition, we are seeking your help to achieve a critical breakthrough  in this problem: **to achieve an automated RNA 3D structure prediction method that can outperform human experts**.\n\nThere are many strategies to explore. Some ideas:\n\n- Are there better ways to fine-tune RibonanzaNet to predict 3D structures?\n- Humans detect information on ‘multiple sequence alignments’ of RNA that current automated methods can’t — how do we get artificial neural networks to see these patterns?\n- Human experts have a rich knowledge of the scientific literature on prior RNA structures – can neural nets achieve and leverage this written knowledge too?\n- For RNA's with multiple structures, are there better ways to generate 5 predictions than just different random seeds?\n- Human experts invest substantial efforts at inference time, sometimes creating 1000s of 'decoy' structures per RNA target and then using reasoning to iterate to additional structures -- can't that be automated?\n- Can synthetic RNA structures be used to augment model training?\n\nAnd of course there’s more that we haven't even thought of!\n\nWe’re looking forward to seeing what Kagglers come up with. We are especially excited to meet first-time Kagglers, whose fresh insights and high-placing models have had a huge impact in prior RNA competitions.\n\nIn addition, we'll be releasing some Easter eggs throughout the competition–and incentivizing sharing with Early Sharing prizes based on who beats the baseline first!–so make sure to stay tuned to these forums.\n\nLet’s get to superhuman performance for RNA structure prediction! See you in the challenge!\n\nYour hosts\nRhiju Das @rhijudas and Shujun He @shujun717\n\nExtra thanks to: \n- Members of the Das lab providing new structures and analysis,  @rkretsch and @alissahummer\n- CASP16 organizers, @andriyca and John Moult\n- RNA-Puzzles organizers Chichau Miao and Eric Westhof\n- The global RNA structural biology community providing blind prediction targets \n- The RNA RFdiffusion/MPNN team at the Institute of Protein Design, Andrew Favor @andrewfavor, Andrew Kubaney, and David Baker\n- Shi-Jie Chen and the VFOLD team for providing `VFOLD_human_expert` baseline predictions.\n\nAnd a huge thanks to our long-time collaborators at Kaggle:\nMaggie Demkin @maggiemd\nInversion @inversion",
    "3138203": "@rhijudas @shujun717 \n\nI have a question regarding the training data for this competition. This time, the dataset contains 844 sequences, whereas the previous competition had around 800,000 sequences. This seems like a significant reduction in data size.\n\nCould you share the reasoning behind this difference? Additionally, I would appreciate more details on how the training dataset was created.\nThank you for your time and clarification.\n\nBest regards,",
    "3213708": "After the end of submission:\n\nI genuinely admire your group's contributions to this competition. Most kagglers, myself included, simply applied existing sota methods to compete. I consider myself merely a clumsy imitator—truth be told, our academic breakthroughs have been quite limited.\n\nI've followed all your papers, tweets, and even media coverage since the competition began. In my eyes, your group is the true victor of this challenge.",
    "3207660": "Two month working hard, learnt a lot from this competition! ",
    "3147146": "Hello,\nI was wondering if we can use predictions from existing methods like AlphaFold 3 or RhoFold during the inference process? \nIf so, are we supposed to install those tools in the notebook? ",
    "3175689": "Good things :)",
    "3153793": "Any useful links to understanding how humans predict RNA folding or books, especially for non-biochemists.",
    "3206884": "..................",
    "3175264": "The idea of fine-tuning RibonanzaNet to predict 3D structures seems promising. Exciting challenge. ",
    "3170476": "is the VFOLD_human_expert baseline what won the CASP16? Or how was it obtained?",
    "3170377": "Good things :)",
    "3167337": "Great competition!",
    "3153917": "wow good job kkk",
    "3146650": "will there be distribution shift intended? Or dataset is generated under the same distribution?",
    "3145401": "Here we go",
    "3143761": "Hello,everyone\nCould you please explain the difference between [train/validation/test]_sequences.csv and [train/validation]_labels.csv? I want to understand how they are used in the dataset ?",
    "3143733": "Hi @rhijudas .  I haven't run into this before, but when I attempt to attach the uw_synthetic_rna_structures dataset that is linked in the competition materials, I get the following error:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11924794%2F003f8b732d460b6ed557ef652d50f6d1%2Fcannot_submit_error.png?generation=1741358101013391&alt=media)\n\nJust to clarify, we are able to use this dataset?  ",
    "3141180": "Thank you, this is my first competition, not an easy one to start but a very interesting challenge. ",
    "3138637": "Hi I'm new \n\nI saw that in train_label only has x1 y1 z1 but you want us to predict 5 set ? so how does that work I don't really get it please clarify. thanks in advance.",
    "3305592": ""
  }
}