{
  "id": 571704,
  "title": "RibonanzaNet2 alpha release",
  "url": "/competitions/stanford-rna-3d-folding/discussion/571704",
  "author_name": "Shujun",
  "post_date": "2025-04-05T01:38:11.642000",
  "votes": 40,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi kagglers, today we are excited to release an alpha version of RibonanzaNet2 (Rnet2)! Link to kaggle models is <a href=\"https://www.kaggle.com/models/shujun717/ribonanzanet2\" target=\"_blank\">https://www.kaggle.com/models/shujun717/ribonanzanet2</a>. Compared to our previous work Ribonanza, we have scaled up data by 100x and the model by 10x. Overall we have collected data on 40 million 100mers and 200 billion rawreads to train Rnet2. We have tested Rnet2 on downstream RNA structure tasks and it has significantly improved performance compared to RibonanzaNet1.</p>\n<p>We hope you will try out and RibonanzaNet2 and let us know if you have any feedback! Enjoy modeling and let us know if you have any questions!</p>\n<p>I have created some starter non-equivariant diffusion finetuning noteooks as follows:</p>\n<p>Starter training notebook w Rnet2: <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training</a><br>\nStarter inference notebook w Rnet2 (after training for 50 epochs): <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference</a></p>\n<h1>Performance improvements</h1>\n<h2>Secondary structure</h2>\n<p>We are getting significant improvements in model downstream performance on RNA secondary structure, particularly with CP_F1, F1 score of crossed pairs or pseudo knots, which are hallmarks of RNA tertiary structure. On pseudoknots, it looks like we're getting accuracy boost of roughly 0.1 for every 10x scaleup (Rnet1-&gt;Rnet2 4M-&gt;Rnet2 30M)!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F871976fa7f1075c037eb9ae6eae81c3e%2FRnet2_SS.png?generation=1743816541079688&amp;alt=media\" alt=\"\"></p>\n<h2>3D structure</h2>\n<p>We compared Rnet2 vs Rnet1 vs Rnet2 no pretrain (randomly initialized) 3D finetuning runs using my DDPM non equivarinat diffusion code. These models are finetuned for 50 epochs with max crop 384 using the competition train data. </p>\n<p>Train val loss curves, Rnet2 does better than Rnet1 in all metrics! Rnet2 without pretraining does not do very well expected. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fee3ea670f693c14667775d02436ee96b%2FRnet2_3d_train.png?generation=1743816679948987&amp;alt=media\" alt=\"\"></p>\n<p>On casp15, Rnet2 has huge signal on casp15 and gets 0.31 TM vs Rnet1 which gets 0.23<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F22a0f68a0e0e16697300344eaba48c57%2FRnet2_CASP15.png?generation=1743816647142829&amp;alt=media\" alt=\"\"></p>\n<p>On a CASP15 target R1116, Rnet2 gets 0.67 on R1116, better than top casp15 submission! On the same target Rnet1 only gets 0.34. </p>\n<p>Rnet2 (red) prediction superimposed with target (green)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F0a5377df5cddbccbed283bcf2c0633d5%2Fimage%20(5).png?generation=1743816710040770&amp;alt=media\" alt=\"\"></p>\n<p>Rnet1 (red) prediction superimposed with target (green)<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F079721176f3bf54f6d710057b6622b37%2Fimage%20(6).png?generation=1743816728175379&amp;alt=media\" alt=\"\"></p>\n<p>Rnet2 no pretrain does not do well<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F00132f300e157f9a8dc8afea8e1d4659%2Fimage%20(7).png?generation=1743816760978140&amp;alt=media\" alt=\"\"></p>\n<p>Another interesting case is R1128 paranemic cross over triangle Rnet2 clearly gets the triangle shape and Rnet1 does not </p>\n<p>Rnet2<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5341fee1f1caac44fcdb40a804bfff5c%2Fimage%20(8).png?generation=1743816884457590&amp;alt=media\" alt=\"\"></p>\n<p>Rnet1<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fb8e9baeff219b3dee3254d39790842a6%2Fimage%20(9).png?generation=1743816914453029&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3170763,
      "postDate": "2025-04-05T01:38:11.643Z",
      "content": "<p>Hi kagglers, today we are excited to release an alpha version of RibonanzaNet2 (Rnet2)! Link to kaggle models is <a href=\"https://www.kaggle.com/models/shujun717/ribonanzanet2\" target=\"_blank\">https://www.kaggle.com/models/shujun717/ribonanzanet2</a>. Compared to our previous work Ribonanza, we have scaled up data by 100x and the model by 10x. Overall we have collected data on 40 million 100mers and 200 billion rawreads to train Rnet2. We have tested Rnet2 on downstream RNA structure tasks and it has significantly improved performance compared to RibonanzaNet1.</p>\n<p>We hope you will try out and RibonanzaNet2 and let us know if you have any feedback! Enjoy modeling and let us know if you have any questions!</p>\n<p>I have created some starter non-equivariant diffusion finetuning noteooks as follows:</p>\n<p>Starter training notebook w Rnet2: <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training</a><br>\nStarter inference notebook w Rnet2 (after training for 50 epochs): <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference\" target=\"_blank\">https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference</a></p>\n<h1>Performance improvements</h1>\n<h2>Secondary structure</h2>\n<p>We are getting significant improvements in model downstream performance on RNA secondary structure, particularly with CP_F1, F1 score of crossed pairs or pseudo knots, which are hallmarks of RNA tertiary structure. On pseudoknots, it looks like we're getting accuracy boost of roughly 0.1 for every 10x scaleup (Rnet1-&gt;Rnet2 4M-&gt;Rnet2 30M)!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F871976fa7f1075c037eb9ae6eae81c3e%2FRnet2_SS.png?generation=1743816541079688&amp;alt=media\" alt=\"\"></p>\n<h2>3D structure</h2>\n<p>We compared Rnet2 vs Rnet1 vs Rnet2 no pretrain (randomly initialized) 3D finetuning runs using my DDPM non equivarinat diffusion code. These models are finetuned for 50 epochs with max crop 384 using the competition train data. </p>\n<p>Train val loss curves, Rnet2 does better than Rnet1 in all metrics! Rnet2 without pretraining does not do very well expected. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fee3ea670f693c14667775d02436ee96b%2FRnet2_3d_train.png?generation=1743816679948987&amp;alt=media\" alt=\"\"></p>\n<p>On casp15, Rnet2 has huge signal on casp15 and gets 0.31 TM vs Rnet1 which gets 0.23<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F22a0f68a0e0e16697300344eaba48c57%2FRnet2_CASP15.png?generation=1743816647142829&amp;alt=media\" alt=\"\"></p>\n<p>On a CASP15 target R1116, Rnet2 gets 0.67 on R1116, better than top casp15 submission! On the same target Rnet1 only gets 0.34. </p>\n<p>Rnet2 (red) prediction superimposed with target (green)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F0a5377df5cddbccbed283bcf2c0633d5%2Fimage%20(5).png?generation=1743816710040770&amp;alt=media\" alt=\"\"></p>\n<p>Rnet1 (red) prediction superimposed with target (green)<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F079721176f3bf54f6d710057b6622b37%2Fimage%20(6).png?generation=1743816728175379&amp;alt=media\" alt=\"\"></p>\n<p>Rnet2 no pretrain does not do well<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F00132f300e157f9a8dc8afea8e1d4659%2Fimage%20(7).png?generation=1743816760978140&amp;alt=media\" alt=\"\"></p>\n<p>Another interesting case is R1128 paranemic cross over triangle Rnet2 clearly gets the triangle shape and Rnet1 does not </p>\n<p>Rnet2<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5341fee1f1caac44fcdb40a804bfff5c%2Fimage%20(8).png?generation=1743816884457590&amp;alt=media\" alt=\"\"></p>\n<p>Rnet1<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fb8e9baeff219b3dee3254d39790842a6%2Fimage%20(9).png?generation=1743816914453029&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi kagglers, today we are excited to release an alpha version of RibonanzaNet2 (Rnet2)! Link to kaggle models is https://www.kaggle.com/models/shujun717/ribonanzanet2. Compared to our previous work Ribonanza, we have scaled up data by 100x and the model by 10x. Overall we have collected data on 40 million 100mers and 200 billion rawreads to train Rnet2. We have tested Rnet2 on downstream RNA structure tasks and it has significantly improved performance compared to RibonanzaNet1.\n\nWe hope you will try out and RibonanzaNet2 and let us know if you have any feedback! Enjoy modeling and let us know if you have any questions!\n\nI have created some starter non-equivariant diffusion finetuning noteooks as follows:\n\nStarter training notebook w Rnet2: https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training\nStarter inference notebook w Rnet2 (after training for 50 epochs): https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference\n# Performance improvements\n\n## Secondary structure\n\nWe are getting significant improvements in model downstream performance on RNA secondary structure, particularly with CP_F1, F1 score of crossed pairs or pseudo knots, which are hallmarks of RNA tertiary structure. On pseudoknots, it looks like we're getting accuracy boost of roughly 0.1 for every 10x scaleup (Rnet1->Rnet2 4M->Rnet2 30M)!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F871976fa7f1075c037eb9ae6eae81c3e%2FRnet2_SS.png?generation=1743816541079688&alt=media)\n\n\n## 3D structure\n\nWe compared Rnet2 vs Rnet1 vs Rnet2 no pretrain (randomly initialized) 3D finetuning runs using my DDPM non equivarinat diffusion code. These models are finetuned for 50 epochs with max crop 384 using the competition train data. \n\nTrain val loss curves, Rnet2 does better than Rnet1 in all metrics! Rnet2 without pretraining does not do very well expected. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fee3ea670f693c14667775d02436ee96b%2FRnet2_3d_train.png?generation=1743816679948987&alt=media)\n\n\nOn casp15, Rnet2 has huge signal on casp15 and gets 0.31 TM vs Rnet1 which gets 0.23![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F22a0f68a0e0e16697300344eaba48c57%2FRnet2_CASP15.png?generation=1743816647142829&alt=media)\n\nOn a CASP15 target R1116, Rnet2 gets 0.67 on R1116, better than top casp15 submission! On the same target Rnet1 only gets 0.34. \n\nRnet2 (red) prediction superimposed with target (green)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F0a5377df5cddbccbed283bcf2c0633d5%2Fimage%20(5).png?generation=1743816710040770&alt=media)\n\nRnet1 (red) prediction superimposed with target (green)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F079721176f3bf54f6d710057b6622b37%2Fimage%20(6).png?generation=1743816728175379&alt=media)\n\nRnet2 no pretrain does not do well![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F00132f300e157f9a8dc8afea8e1d4659%2Fimage%20(7).png?generation=1743816760978140&alt=media)\n\nAnother interesting case is R1128 paranemic cross over triangle Rnet2 clearly gets the triangle shape and Rnet1 does not \n\nRnet2\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5341fee1f1caac44fcdb40a804bfff5c%2Fimage%20(8).png?generation=1743816884457590&alt=media)\n\nRnet1\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fb8e9baeff219b3dee3254d39790842a6%2Fimage%20(9).png?generation=1743816914453029&alt=media)",
      "votes": 40
    },
    {
      "id": 3171756,
      "postDate": "2025-04-06T04:16:32.967Z",
      "content": "<p>Maxlength trained using Rnet2?   </p>",
      "rawMarkdown": "Maxlength trained using Rnet2?   ",
      "votes": 1,
      "replies": [
        {
          "id": 3173329,
          "postDate": "2025-04-07T19:16:11.023Z",
          "content": "<p>Pretraining has max len 206 but you can use longer for finetuning on 3d. </p>",
          "rawMarkdown": "Pretraining has max len 206 but you can use longer for finetuning on 3d. ",
          "replies": [
            {
              "id": 3173332,
              "postDate": "2025-04-07T19:17:34.627Z",
              "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> i tried 800, trained for 100 epoch  obtained lb 0.236</p>",
              "rawMarkdown": "@shujun717 i tried 800, trained for 100 epoch  obtained lb 0.236",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3181880,
      "postDate": "2025-04-18T13:24:47.157Z",
      "content": "<p>this ishould be related<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3181877\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3181877</a></p>\n<p>trRosettaRNA2 (yang-server) from ss-pretrain to 3d strcuture prediction</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F31abb0b7d081270b8a3968b5a92afcf2%2FSelection_999(8114).png?generation=1744982642327329&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "this ishould be related\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3181877\n\ntrRosettaRNA2 (yang-server) from ss-pretrain to 3d strcuture prediction\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F31abb0b7d081270b8a3968b5a92afcf2%2FSelection_999(8114).png?generation=1744982642327329&alt=media)\n"
    },
    {
      "id": 3176045,
      "postDate": "2025-04-10T21:58:46.597Z",
      "content": "<p>amazing work</p>",
      "rawMarkdown": "amazing work\n"
    },
    {
      "id": 3175357,
      "postDate": "2025-04-10T04:05:47.203Z",
      "content": "<p>Thanks for sharing this exciting work! Could you also share the LDDT scores for these targets? </p>",
      "rawMarkdown": "Thanks for sharing this exciting work! Could you also share the LDDT scores for these targets? "
    },
    {
      "id": 3171227,
      "postDate": "2025-04-05T12:25:32.640Z",
      "content": "<p>thanks for your amazing work!</p>\n<blockquote>\n  <p>On casp15, Rnet2 has huge signal on casp15 and gets 0.31 TM vs Rnet1 which gets 0.23</p>\n</blockquote>\n<p>By modifying backbone and configuration, even Rnet1 can reach around 0.255 without leakage in the leaderboard, and I'm still trying. Maybe RNet2 can be boosted to 0.300. I love this lightweight model, and it's powerful enough to compete with some sota in some datasets.</p>",
      "rawMarkdown": "thanks for your amazing work!\n\n>On casp15, Rnet2 has huge signal on casp15 and gets 0.31 TM vs Rnet1 which gets 0.23\n\nBy modifying backbone and configuration, even Rnet1 can reach around 0.255 without leakage in the leaderboard, and I'm still trying. Maybe RNet2 can be boosted to 0.300. I love this lightweight model, and it's powerful enough to compete with some sota in some datasets.\n\n",
      "replies": [
        {
          "id": 3172160,
          "postDate": "2025-04-06T14:24:00.910Z",
          "content": "<p>i think the host will be the winner of this competition !!</p>\n<p>btw, here is the x post on this<br>\n\"Today we’re releasing an alpha version of RibonanzaNet2. Rnet2 is a 100M-parameter foundation model for #RNA structure, trained on <strong>chemical mapping profiles</strong> for 30M RNAs with complex structure. \"</p>\n<p>\"Rnet2 training data are on <strong>synthetic designs from RFdiffusion</strong> <a href=\"https://www.kaggle.com/UWproteindesign\" target=\"_blank\">@UWproteindesign</a>, gRNade@chaitjo,  genome scans, and a new design method Shujun <a href=\"https://www.kaggle.com/SJ\" target=\"_blank\">@SJ</a>_He. Sequences mapped thanks to key experimental innovations by Ann Kladwang and Hamish Blair <a href=\"https://www.kaggle.com/rdaslab\" target=\"_blank\">@rdaslab</a>\"<br>\n.</p>\n<p><a href=\"https://x.com/RDasLab/status/1908637244279415014\" target=\"_blank\">https://x.com/RDasLab/status/1908637244279415014</a></p>",
          "rawMarkdown": "i think the host will be the winner of this competition !!\n\nbtw, here is the x post on this\n\"Today we’re releasing an alpha version of RibonanzaNet2. Rnet2 is a 100M-parameter foundation model for #RNA structure, trained on **chemical mapping profiles** for 30M RNAs with complex structure. \"\n\n\"Rnet2 training data are on **synthetic designs from RFdiffusion** @UWproteindesign, gRNade@chaitjo,  genome scans, and a new design method Shujun @SJ_He. Sequences mapped thanks to key experimental innovations by Ann Kladwang and Hamish Blair @rdaslab\"\n.\n\n\nhttps://x.com/RDasLab/status/1908637244279415014",
          "votes": 2
        }
      ]
    },
    {
      "id": 3172554,
      "postDate": "2025-04-07T00:22:23.157Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3172601,
      "postDate": "2025-04-07T02:18:18.553Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3171756,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2025-04-06T04:16:32.967000",
      "content": "<p>Maxlength trained using Rnet2?   </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3173329,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2025-04-07T19:16:11.023000",
          "content": "<p>Pretraining has max len 206 but you can use longer for finetuning on 3d. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3173332,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-04-07T19:17:34.627000",
              "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> i tried 800, trained for 100 epoch  obtained lb 0.236</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3181880,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-18T13:24:47.157000",
      "content": "<p>this ishould be related<br>\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3181877\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3181877</a></p>\n<p>trRosettaRNA2 (yang-server) from ss-pretrain to 3d strcuture prediction</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F31abb0b7d081270b8a3968b5a92afcf2%2FSelection_999(8114).png?generation=1744982642327329&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3176045,
      "author_name": "Ch ali hassan",
      "author_url": "",
      "post_date": "2025-04-10T21:58:46.597000",
      "content": "<p>amazing work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3175357,
      "author_name": "asiwu0103",
      "author_url": "",
      "post_date": "2025-04-10T04:05:47.203000",
      "content": "<p>Thanks for sharing this exciting work! Could you also share the LDDT scores for these targets? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3171227,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2025-04-05T12:25:32.640000",
      "content": "<p>thanks for your amazing work!</p>\n<blockquote>\n  <p>On casp15, Rnet2 has huge signal on casp15 and gets 0.31 TM vs Rnet1 which gets 0.23</p>\n</blockquote>\n<p>By modifying backbone and configuration, even Rnet1 can reach around 0.255 without leakage in the leaderboard, and I'm still trying. Maybe RNet2 can be boosted to 0.300. I love this lightweight model, and it's powerful enough to compete with some sota in some datasets.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3172160,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-04-06T14:24:00.910000",
          "content": "<p>i think the host will be the winner of this competition !!</p>\n<p>btw, here is the x post on this<br>\n\"Today we’re releasing an alpha version of RibonanzaNet2. Rnet2 is a 100M-parameter foundation model for #RNA structure, trained on <strong>chemical mapping profiles</strong> for 30M RNAs with complex structure. \"</p>\n<p>\"Rnet2 training data are on <strong>synthetic designs from RFdiffusion</strong> <a href=\"https://www.kaggle.com/UWproteindesign\" target=\"_blank\">@UWproteindesign</a>, gRNade@chaitjo,  genome scans, and a new design method Shujun <a href=\"https://www.kaggle.com/SJ\" target=\"_blank\">@SJ</a>_He. Sequences mapped thanks to key experimental innovations by Ann Kladwang and Hamish Blair <a href=\"https://www.kaggle.com/rdaslab\" target=\"_blank\">@rdaslab</a>\"<br>\n.</p>\n<p><a href=\"https://x.com/RDasLab/status/1908637244279415014\" target=\"_blank\">https://x.com/RDasLab/status/1908637244279415014</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3172554,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-07T00:22:23.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3172601,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-07T02:18:18.553000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3170763": "Hi kagglers, today we are excited to release an alpha version of RibonanzaNet2 (Rnet2)! Link to kaggle models is https://www.kaggle.com/models/shujun717/ribonanzanet2. Compared to our previous work Ribonanza, we have scaled up data by 100x and the model by 10x. Overall we have collected data on 40 million 100mers and 200 billion rawreads to train Rnet2. We have tested Rnet2 on downstream RNA structure tasks and it has significantly improved performance compared to RibonanzaNet1.\n\nWe hope you will try out and RibonanzaNet2 and let us know if you have any feedback! Enjoy modeling and let us know if you have any questions!\n\nI have created some starter non-equivariant diffusion finetuning noteooks as follows:\n\nStarter training notebook w Rnet2: https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-training\nStarter inference notebook w Rnet2 (after training for 50 epochs): https://www.kaggle.com/code/shujun717/ribonanzanet2-ddpm-inference\n# Performance improvements\n\n## Secondary structure\n\nWe are getting significant improvements in model downstream performance on RNA secondary structure, particularly with CP_F1, F1 score of crossed pairs or pseudo knots, which are hallmarks of RNA tertiary structure. On pseudoknots, it looks like we're getting accuracy boost of roughly 0.1 for every 10x scaleup (Rnet1->Rnet2 4M->Rnet2 30M)!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F871976fa7f1075c037eb9ae6eae81c3e%2FRnet2_SS.png?generation=1743816541079688&alt=media)\n\n\n## 3D structure\n\nWe compared Rnet2 vs Rnet1 vs Rnet2 no pretrain (randomly initialized) 3D finetuning runs using my DDPM non equivarinat diffusion code. These models are finetuned for 50 epochs with max crop 384 using the competition train data. \n\nTrain val loss curves, Rnet2 does better than Rnet1 in all metrics! Rnet2 without pretraining does not do very well expected. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fee3ea670f693c14667775d02436ee96b%2FRnet2_3d_train.png?generation=1743816679948987&alt=media)\n\n\nOn casp15, Rnet2 has huge signal on casp15 and gets 0.31 TM vs Rnet1 which gets 0.23![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F22a0f68a0e0e16697300344eaba48c57%2FRnet2_CASP15.png?generation=1743816647142829&alt=media)\n\nOn a CASP15 target R1116, Rnet2 gets 0.67 on R1116, better than top casp15 submission! On the same target Rnet1 only gets 0.34. \n\nRnet2 (red) prediction superimposed with target (green)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F0a5377df5cddbccbed283bcf2c0633d5%2Fimage%20(5).png?generation=1743816710040770&alt=media)\n\nRnet1 (red) prediction superimposed with target (green)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F079721176f3bf54f6d710057b6622b37%2Fimage%20(6).png?generation=1743816728175379&alt=media)\n\nRnet2 no pretrain does not do well![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F00132f300e157f9a8dc8afea8e1d4659%2Fimage%20(7).png?generation=1743816760978140&alt=media)\n\nAnother interesting case is R1128 paranemic cross over triangle Rnet2 clearly gets the triangle shape and Rnet1 does not \n\nRnet2\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2F5341fee1f1caac44fcdb40a804bfff5c%2Fimage%20(8).png?generation=1743816884457590&alt=media)\n\nRnet1\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3355848%2Fb8e9baeff219b3dee3254d39790842a6%2Fimage%20(9).png?generation=1743816914453029&alt=media)",
    "3171756": "Maxlength trained using Rnet2?   ",
    "3181880": "this ishould be related\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566906#3181877\n\ntrRosettaRNA2 (yang-server) from ss-pretrain to 3d strcuture prediction\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F31abb0b7d081270b8a3968b5a92afcf2%2FSelection_999(8114).png?generation=1744982642327329&alt=media)\n",
    "3176045": "amazing work\n",
    "3175357": "Thanks for sharing this exciting work! Could you also share the LDDT scores for these targets? ",
    "3171227": "thanks for your amazing work!\n\n>On casp15, Rnet2 has huge signal on casp15 and gets 0.31 TM vs Rnet1 which gets 0.23\n\nBy modifying backbone and configuration, even Rnet1 can reach around 0.255 without leakage in the leaderboard, and I'm still trying. Maybe RNet2 can be boosted to 0.300. I love this lightweight model, and it's powerful enough to compete with some sota in some datasets.\n\n",
    "3172554": "",
    "3172601": "Thanks for sharing"
  }
}