{
  "id": 568125,
  "title": "[draft] example to use MAS provided: case of rhofold+",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568125",
  "author_name": "hengck23",
  "post_date": "2025-03-14T00:37:58.701000",
  "votes": 16,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Each step will be described in detail below. The focus is on how to check whether your usage is correct. Fortunately, there is an online web server for Rhofold+ that allows verification against our local results</p>\n<p>here are the summary:<br>\n1)  <br>\ni think kaggle MAS already in a3m format. the extension fasta should be a3m<br>\n2) predict using Rhofold+ using web server and local<br>\n3) compare the results<br>\n4) other important note, eg effects of num of msa used, conformations using different msa input<br>\n3) generate your own msa</p>",
  "messages": [
    {
      "id": 3149222,
      "postDate": "2025-03-14T00:37:58.700Z",
      "content": "<p>Each step will be described in detail below. The focus is on how to check whether your usage is correct. Fortunately, there is an online web server for Rhofold+ that allows verification against our local results</p>\n<p>here are the summary:<br>\n1)  <br>\ni think kaggle MAS already in a3m format. the extension fasta should be a3m<br>\n2) predict using Rhofold+ using web server and local<br>\n3) compare the results<br>\n4) other important note, eg effects of num of msa used, conformations using different msa input<br>\n3) generate your own msa</p>",
      "rawMarkdown": "Each step will be described in detail below. The focus is on how to check whether your usage is correct. Fortunately, there is an online web server for Rhofold+ that allows verification against our local results\n\nhere are the summary:\n1) ~~understand alignment format: fasta and a3m. Convert to a3m.~~ \ni think kaggle MAS already in a3m format. the extension fasta should be a3m\n2) predict using Rhofold+ using web server and local\n3) compare the results\n4) other important note, eg effects of num of msa used, conformations using different msa input\n3) generate your own msa",
      "votes": 16
    },
    {
      "id": 3150011,
      "postDate": "2025-03-15T00:53:39.307Z",
      "content": "<p>good news.<br>\nwe have another benchmark paper, their a3m files is provided and can be downloaded</p>\n<p>Systematic benchmarking of deep-learning methods for tertiary RNA structure prediction<br>\n<a href=\"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715#sec035\" target=\"_blank\">https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715#sec035</a><br>\n<a href=\"https://github.com/akashbahai/rna_benchmarking/tree/main/msa/r1108\" target=\"_blank\">https://github.com/akashbahai/rna_benchmarking/tree/main/msa/r1108</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5091071a931c5eec3b4849717a11c612%2FSelection_052.png?generation=1741999960906042&amp;alt=media\" alt=\"\"></p>\n<p>but this is for rhofold repo (and not rhofold+ repo)<br>\nbut both rhofold and rhofold+ repot seem to point to the same pretrain model</p>\n<p>rhofold+:<br>\n<a href=\"https://github.com/ml4bio/RhoFold?tab=readme-ov-file\" target=\"_blank\">https://github.com/ml4bio/RhoFold?tab=readme-ov-file</a></p>\n<p>rhofold:<br>\n<a href=\"https://github.com/Dharmogata/RhoFold?tab=readme-ov-file\" target=\"_blank\">https://github.com/Dharmogata/RhoFold?tab=readme-ov-file</a></p>",
      "rawMarkdown": "good news.\nwe have another benchmark paper, their a3m files is provided and can be downloaded\n\nSystematic benchmarking of deep-learning methods for tertiary RNA structure prediction\nhttps://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715#sec035\nhttps://github.com/akashbahai/rna_benchmarking/tree/main/msa/r1108\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5091071a931c5eec3b4849717a11c612%2FSelection_052.png?generation=1741999960906042&alt=media)\n\nbut this is for rhofold repo (and not rhofold+ repo)\nbut both rhofold and rhofold+ repot seem to point to the same pretrain model\n\n\nrhofold+:\nhttps://github.com/ml4bio/RhoFold?tab=readme-ov-file\n\nrhofold:\nhttps://github.com/Dharmogata/RhoFold?tab=readme-ov-file",
      "votes": 1
    },
    {
      "id": 3149991,
      "postDate": "2025-03-14T23:59:54.737Z",
      "content": "<p>let's \"algorithmically debug\" our implementation.<br>\nwe use kaggle validation csv (which is casp15).</p>\n<p>first we need reference results and we will use rhofold+ paper (supplementary material)<br>\n<a href=\"https://www.nature.com/articles/s41592-024-02487-0#additional-information\" target=\"_blank\">https://www.nature.com/articles/s41592-024-02487-0#additional-information</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc85b403de5259c2a67a340fff86577fb%2FSelection_051.png?generation=1741996415341156&amp;alt=media\" alt=\"\"></p>\n<p>if you run my example notebook[1] with max num of MAS you should be getting<br>\n[1] <a href=\"https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa</a>  </p>\n<pre><code> relax\n\n R1107 num_msa=   len(seq)= df.shape=(, )   seq = GGGGGCCACA...\n .\n\n R1108   (, ) GGGGGCCACA...\n .\n\n R1116   (, ) CGCCCGGAUA...\n .\n\n R1117v2   (, ) UUGGGUUCCC...\n .\n\n R1149   (, ) GGACACGAGU...\n .\n\n R1156   (, ) GGAGCAUCGU...\n .\n</code></pre>\n<p>our results are definitely lower!<br>\nlet's investigate why and improve on yt!</p>\n<p>… to be updated …</p>\n<p>some quite observation:</p>\n<ul>\n<li>is TM-score in paper for backbone or all atoms (i.e. paper TM score is the same as ours?)</li>\n<li>I also note that the num of MSA is different from the paper. luckliy the paper provide details of max match seq …)</li>\n<li>I think the comment from a3m header like \"&gt;000|DF967972.1/3772384-3772312/1-73 [subseq from] Longilinea arvoryzae DNA, scaffold: Larv_scaffold_1, strain: KOME-1.\" would help. maybe we can use a LLM model for text analysis for this </li>\n</ul>",
      "rawMarkdown": "let's \"algorithmically debug\" our implementation.\nwe use kaggle validation csv (which is casp15).\n\nfirst we need reference results and we will use rhofold+ paper (supplementary material)\nhttps://www.nature.com/articles/s41592-024-02487-0#additional-information\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc85b403de5259c2a67a340fff86577fb%2FSelection_051.png?generation=1741996415341156&alt=media)\n\n\nif you run my example notebook[1] with max num of MAS you should be getting\n[1] https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa  \n\n```\nno relax\n\n0 R1107 num_msa=21   len(seq)=69 df.shape=(69, 6)   seq = GGGGGCCACA...\ntm_score_unrelax 0.32113\n\n1 R1108 26 69 (69, 6) GGGGGCCACA...\ntm_score_unrelax 0.41223\n\n2 R1116 68 157 (157, 6) CGCCCGGAUA...\ntm_score_unrelax 0.52747\n\n3 R1117v2 1 30 (30, 6) UUGGGUUCCC...\ntm_score_unrelax 0.20945\n\n8 R1149 9 124 (124, 6) GGACACGAGU...\ntm_score_unrelax 0.25613\n\n9 R1156 11 135 (135, 6) GGAGCAUCGU...\ntm_score_unrelax 0.34896\n```\n\nour results are definitely lower!\nlet's investigate why and improve on yt!\n\n... to be updated ...\n\nsome quite observation:\n- is TM-score in paper for backbone or all atoms (i.e. paper TM score is the same as ours?)\n- I also note that the num of MSA is different from the paper. luckliy the paper provide details of max match seq ...)\n- I think the comment from a3m header like \">000|DF967972.1/3772384-3772312/1-73 [subseq from] Longilinea arvoryzae DNA, scaffold: Larv_scaffold_1, strain: KOME-1.\" would help. maybe we can use a LLM model for text analysis for this ",
      "votes": 1,
      "replies": [
        {
          "id": 3150060,
          "postDate": "2025-03-15T03:02:23.320Z",
          "content": "<p>i can call rhofold server to do MSA. <br>\nthen we can compare the quality of MSA. will report results later</p>",
          "rawMarkdown": "i can call rhofold server to do MSA. \nthen we can compare the quality of MSA. will report results later",
          "votes": 1,
          "replies": [
            {
              "id": 3150216,
              "postDate": "2025-03-15T07:19:15.810Z",
              "content": "<p>We got RhoFold LB score for MSAs created from rnacentral.fasta a bit lower than for precomputed MSAs from organizers (0.208 vs 0.214).</p>",
              "rawMarkdown": "We got RhoFold LB score for MSAs created from rnacentral.fasta a bit lower than for precomputed MSAs from organizers (0.208 vs 0.214)."
            },
            {
              "id": 3150242,
              "postDate": "2025-03-15T08:11:54.673Z",
              "content": "<p>base on my experiments and \"guess\", for natural rna of length less than 200, SOTA like rhofold+ etc should be able to get tm score 0.40 to 0.50 for casp15/16 and rna puzzle test samples. We are still not getting the implementation correct. the devils are still hiding</p>",
              "rawMarkdown": "base on my experiments and \"guess\", for natural rna of length less than 200, SOTA like rhofold+ etc should be able to get tm score 0.40 to 0.50 for casp15/16 and rna puzzle test samples. We are still not getting the implementation correct. the devils are still hiding",
              "votes": 1
            },
            {
              "id": 3150255,
              "postDate": "2025-03-15T08:43:28.150Z",
              "content": "<p>Validation scores for RhoFold are not so impressive:</p>\n<blockquote>\n  <h3>0.256 for single_seq_pred</h3>\n  <h3>0.295 for rnacentral.fasta MSAs</h3>\n  <h3>0.286 for kaggle MSAs</h3>\n</blockquote>\n<p>validation_sequences.csv (and test_sequences.csv which is the same file) contains all but one sequences &lt;500 nt, so RhoFold inference runs without any issues despite it was trained on shorter RNAs.</p>",
              "rawMarkdown": "Validation scores for RhoFold are not so impressive:\n>### 0.256 for single_seq_pred\n### 0.295 for rnacentral.fasta MSAs\n### 0.286 for kaggle MSAs\n\nvalidation_sequences.csv (and test_sequences.csv which is the same file) contains all but one sequences <500 nt, so RhoFold inference runs without any issues despite it was trained on shorter RNAs."
            },
            {
              "id": 3150345,
              "postDate": "2025-03-15T11:37:02.120Z",
              "content": "<p>validation_sequences is casp15. if you check rhofold papers and other benchmark papers, it can get higher<br>\n(you need to split score for natural and sythetic rna)</p>",
              "rawMarkdown": "validation_sequences is casp15. if you check rhofold papers and other benchmark papers, it can get higher\n(you need to split score for natural and sythetic rna)"
            },
            {
              "id": 3150573,
              "postDate": "2025-03-15T16:12:38.917Z",
              "content": "<p><a href=\"https://www.kaggle.com/ogurtsov\" target=\"_blank\">@ogurtsov</a></p>\n<p>the rhoflod code only provide blastN for construct MSA.<br>\nBut the paper actually uses rMSA and infernal which gives better results</p>",
              "rawMarkdown": "@ogurtsov\n\nthe rhoflod code only provide blastN for construct MSA.\nBut the paper actually uses rMSA and infernal which gives better results",
              "votes": 1
            },
            {
              "id": 3150611,
              "postDate": "2025-03-15T16:51:45.270Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I hope it's possible to make <code>rMSA</code> &amp; <code>infernal</code> work on kaggle! </p>",
              "rawMarkdown": "@hengck23 I hope it's possible to make `rMSA` & `infernal` work on kaggle! "
            },
            {
              "id": 3150613,
              "postDate": "2025-03-15T16:54:32.267Z",
              "content": "<p>at least one have to prove that that is the cuase of performance.<br>\nif the results are really good, we can ask the host to provide precomputed mas again</p>",
              "rawMarkdown": "at least one have to prove that that is the cuase of performance.\nif the results are really good, we can ask the host to provide precomputed mas again",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3149461,
      "postDate": "2025-03-14T08:41:41.373Z",
      "content": "<h2>2) predict using Rhofold+ using local</h2>\n<p>follow this notebook link to use rhofold+ on kaggle data for 1EIY_C (5 MSA)<br>\n<a href=\"https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fac2110a027e0fea1e42531f8f170febb%2FSelection_049.png?generation=1741941642841589&amp;alt=media\" alt=\"\"></p>\n<p>render on <a href=\"https://molstar.org/viewer/\" target=\"_blank\">https://molstar.org/viewer/</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcd2c3b43c290599db28d456144dee661%2FSelection_048.png?generation=1741941699751604&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "##2) predict using Rhofold+ using local##\n\nfollow this notebook link to use rhofold+ on kaggle data for 1EIY_C (5 MSA)\nhttps://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fac2110a027e0fea1e42531f8f170febb%2FSelection_049.png?generation=1741941642841589&alt=media)\n\nrender on https://molstar.org/viewer/\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcd2c3b43c290599db28d456144dee661%2FSelection_048.png?generation=1741941699751604&alt=media)",
      "votes": 2
    },
    {
      "id": 3157211,
      "postDate": "2025-03-23T06:01:53.480Z",
      "content": "<p>In the method section of the Rhofold+ paper, the authors state: \"By default, the top 256 MSAs are chosen as input features for predicting the standard structure, which we refer to as standard RhoFold+. RhoFold+ (TopK) refers to the optimal model selected from K different models generated using distinct sampled MSAs.\"</p>\n<p>Does this imply that the CASP15 results were selected from multiple attempts? Additionally, Supplementary Table 6 indicates that the MSA depths used in their experiments are deeper than those available in the Kaggle dataset. </p>",
      "rawMarkdown": "In the method section of the Rhofold+ paper, the authors state: \"By default, the top 256 MSAs are chosen as input features for predicting the standard structure, which we refer to as standard RhoFold+. RhoFold+ (TopK) refers to the optimal model selected from K different models generated using distinct sampled MSAs.\"\n\nDoes this imply that the CASP15 results were selected from multiple attempts? Additionally, Supplementary Table 6 indicates that the MSA depths used in their experiments are deeper than those available in the Kaggle dataset. ",
      "replies": [
        {
          "id": 3157217,
          "postDate": "2025-03-23T06:08:43.727Z",
          "content": "<p>Msa is not unique process. There are many ways Eg blastn rmsa mmseq etc. not sure what is exactly use. U will need to experiment </p>",
          "rawMarkdown": "Msa is not unique process. There are many ways Eg blastn rmsa mmseq etc. not sure what is exactly use. U will need to experiment \n",
          "replies": [
            {
              "id": 3157232,
              "postDate": "2025-03-23T06:32:34.887Z",
              "content": "<p>Yes, I plan to use Infernal to generate deeper MSAs and evaluate whether this can lead Rhofold+ to a higher score.</p>",
              "rawMarkdown": "Yes, I plan to use Infernal to generate deeper MSAs and evaluate whether this can lead Rhofold+ to a higher score."
            },
            {
              "id": 3157394,
              "postDate": "2025-03-23T11:55:59.290Z",
              "content": "<p>don't over focus on one method rhofold. try af3, nufold, trRosettaRNA, etc … maybe other authors describe their mas pipline in more details and their results are more reproducible.</p>",
              "rawMarkdown": "don't over focus on one method rhofold. try af3, nufold, trRosettaRNA, etc ... maybe other authors describe their mas pipline in more details and their results are more reproducible."
            },
            {
              "id": 3157425,
              "postDate": "2025-03-23T12:29:44.873Z",
              "content": "<blockquote>\n  <p>evaluate whether this can lead Rhofold+ to a higher score.</p>\n</blockquote>\n<p>It depends on huge datasets:</p>\n<blockquote>\n  <p>To support MSA construction, 3 sequence databases (RNAcentral, Rfam, and nt) totaling about 900GB need to be downloaded.</p>\n</blockquote>",
              "rawMarkdown": ">evaluate whether this can lead Rhofold+ to a higher score.\n\nIt depends on huge datasets:\n\n>To support MSA construction, 3 sequence databases (RNAcentral, Rfam, and nt) totaling about 900GB need to be downloaded."
            },
            {
              "id": 3157448,
              "postDate": "2025-03-23T12:56:43.673Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3157449,
              "postDate": "2025-03-23T12:57:08.160Z",
              "content": "<p>I agree that my current storage capacity is insufficient for handling such large datasets. I should likely focus on other methods.</p>",
              "rawMarkdown": "I agree that my current storage capacity is insufficient for handling such large datasets. I should likely focus on other methods."
            },
            {
              "id": 3157457,
              "postDate": "2025-03-23T13:08:32.023Z",
              "content": "<p>Alternatively, try online msa server. Key is to get pipeline correct. Can get data to local later.</p>\n<p>To submit with online server, submit offline predicted xyz, see my other notebook</p>",
              "rawMarkdown": "Alternatively, try online msa server. Key is to get pipeline correct. Can get data to local later.\n\nTo submit with online server, submit offline predicted xyz, see my other notebook"
            }
          ]
        }
      ]
    },
    {
      "id": 3150555,
      "postDate": "2025-03-15T16:00:55.957Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff1b85c9b31b9245c6da161fee2f18001%2FSelection_058.png?generation=1742054324943370&amp;alt=media\" alt=\"\"></p>\n<ol>\n<li><p>i need to implement or find online server for blastN, rMSA, infernal to verify if the paper better results is due to better  MSA search.</p></li>\n<li><p>It is interesting to note that server results is better then local for same input.<br>\n(did the server use a better pretrain model? or is it due to different GPU, python package  version)</p></li>\n</ol>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff1b85c9b31b9245c6da161fee2f18001%2FSelection_058.png?generation=1742054324943370&alt=media)\n\n1. i need to implement or find online server for blastN, rMSA, infernal to verify if the paper better results is due to better  MSA search.\n\n2. It is interesting to note that server results is better then local for same input.\n(did the server use a better pretrain model? or is it due to different GPU, python package  version)"
    },
    {
      "id": 3149224,
      "postDate": "2025-03-14T00:43:52.903Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3207667,
      "postDate": "2025-05-23T05:37:40.287Z",
      "content": "<p>great work! </p>",
      "rawMarkdown": "great work! ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3150011,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-15T00:53:39.307000",
      "content": "<p>good news.<br>\nwe have another benchmark paper, their a3m files is provided and can be downloaded</p>\n<p>Systematic benchmarking of deep-learning methods for tertiary RNA structure prediction<br>\n<a href=\"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715#sec035\" target=\"_blank\">https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715#sec035</a><br>\n<a href=\"https://github.com/akashbahai/rna_benchmarking/tree/main/msa/r1108\" target=\"_blank\">https://github.com/akashbahai/rna_benchmarking/tree/main/msa/r1108</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5091071a931c5eec3b4849717a11c612%2FSelection_052.png?generation=1741999960906042&amp;alt=media\" alt=\"\"></p>\n<p>but this is for rhofold repo (and not rhofold+ repo)<br>\nbut both rhofold and rhofold+ repot seem to point to the same pretrain model</p>\n<p>rhofold+:<br>\n<a href=\"https://github.com/ml4bio/RhoFold?tab=readme-ov-file\" target=\"_blank\">https://github.com/ml4bio/RhoFold?tab=readme-ov-file</a></p>\n<p>rhofold:<br>\n<a href=\"https://github.com/Dharmogata/RhoFold?tab=readme-ov-file\" target=\"_blank\">https://github.com/Dharmogata/RhoFold?tab=readme-ov-file</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3149991,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-14T23:59:54.737000",
      "content": "<p>let's \"algorithmically debug\" our implementation.<br>\nwe use kaggle validation csv (which is casp15).</p>\n<p>first we need reference results and we will use rhofold+ paper (supplementary material)<br>\n<a href=\"https://www.nature.com/articles/s41592-024-02487-0#additional-information\" target=\"_blank\">https://www.nature.com/articles/s41592-024-02487-0#additional-information</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc85b403de5259c2a67a340fff86577fb%2FSelection_051.png?generation=1741996415341156&amp;alt=media\" alt=\"\"></p>\n<p>if you run my example notebook[1] with max num of MAS you should be getting<br>\n[1] <a href=\"https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa</a>  </p>\n<pre><code> relax\n\n R1107 num_msa=   len(seq)= df.shape=(, )   seq = GGGGGCCACA...\n .\n\n R1108   (, ) GGGGGCCACA...\n .\n\n R1116   (, ) CGCCCGGAUA...\n .\n\n R1117v2   (, ) UUGGGUUCCC...\n .\n\n R1149   (, ) GGACACGAGU...\n .\n\n R1156   (, ) GGAGCAUCGU...\n .\n</code></pre>\n<p>our results are definitely lower!<br>\nlet's investigate why and improve on yt!</p>\n<p>… to be updated …</p>\n<p>some quite observation:</p>\n<ul>\n<li>is TM-score in paper for backbone or all atoms (i.e. paper TM score is the same as ours?)</li>\n<li>I also note that the num of MSA is different from the paper. luckliy the paper provide details of max match seq …)</li>\n<li>I think the comment from a3m header like \"&gt;000|DF967972.1/3772384-3772312/1-73 [subseq from] Longilinea arvoryzae DNA, scaffold: Larv_scaffold_1, strain: KOME-1.\" would help. maybe we can use a LLM model for text analysis for this </li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 3150060,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-03-15T03:02:23.320000",
          "content": "<p>i can call rhofold server to do MSA. <br>\nthen we can compare the quality of MSA. will report results later</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3150216,
              "author_name": "Ogurtsov",
              "author_url": "",
              "post_date": "2025-03-15T07:19:15.810000",
              "content": "<p>We got RhoFold LB score for MSAs created from rnacentral.fasta a bit lower than for precomputed MSAs from organizers (0.208 vs 0.214).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3150242,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-15T08:11:54.673000",
              "content": "<p>base on my experiments and \"guess\", for natural rna of length less than 200, SOTA like rhofold+ etc should be able to get tm score 0.40 to 0.50 for casp15/16 and rna puzzle test samples. We are still not getting the implementation correct. the devils are still hiding</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3150255,
              "author_name": "Ogurtsov",
              "author_url": "",
              "post_date": "2025-03-15T08:43:28.150000",
              "content": "<p>Validation scores for RhoFold are not so impressive:</p>\n<blockquote>\n  <h3>0.256 for single_seq_pred</h3>\n  <h3>0.295 for rnacentral.fasta MSAs</h3>\n  <h3>0.286 for kaggle MSAs</h3>\n</blockquote>\n<p>validation_sequences.csv (and test_sequences.csv which is the same file) contains all but one sequences &lt;500 nt, so RhoFold inference runs without any issues despite it was trained on shorter RNAs.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3150345,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-15T11:37:02.120000",
              "content": "<p>validation_sequences is casp15. if you check rhofold papers and other benchmark papers, it can get higher<br>\n(you need to split score for natural and sythetic rna)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3150573,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-15T16:12:38.917000",
              "content": "<p><a href=\"https://www.kaggle.com/ogurtsov\" target=\"_blank\">@ogurtsov</a></p>\n<p>the rhoflod code only provide blastN for construct MSA.<br>\nBut the paper actually uses rMSA and infernal which gives better results</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3150611,
              "author_name": "Ogurtsov",
              "author_url": "",
              "post_date": "2025-03-15T16:51:45.270000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I hope it's possible to make <code>rMSA</code> &amp; <code>infernal</code> work on kaggle! </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3150613,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-15T16:54:32.267000",
              "content": "<p>at least one have to prove that that is the cuase of performance.<br>\nif the results are really good, we can ask the host to provide precomputed mas again</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3149461,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-14T08:41:41.373000",
      "content": "<h2>2) predict using Rhofold+ using local</h2>\n<p>follow this notebook link to use rhofold+ on kaggle data for 1EIY_C (5 MSA)<br>\n<a href=\"https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa\" target=\"_blank\">https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fac2110a027e0fea1e42531f8f170febb%2FSelection_049.png?generation=1741941642841589&amp;alt=media\" alt=\"\"></p>\n<p>render on <a href=\"https://molstar.org/viewer/\" target=\"_blank\">https://molstar.org/viewer/</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcd2c3b43c290599db28d456144dee661%2FSelection_048.png?generation=1741941699751604&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3157211,
      "author_name": "Bianco Chiu",
      "author_url": "",
      "post_date": "2025-03-23T06:01:53.480000",
      "content": "<p>In the method section of the Rhofold+ paper, the authors state: \"By default, the top 256 MSAs are chosen as input features for predicting the standard structure, which we refer to as standard RhoFold+. RhoFold+ (TopK) refers to the optimal model selected from K different models generated using distinct sampled MSAs.\"</p>\n<p>Does this imply that the CASP15 results were selected from multiple attempts? Additionally, Supplementary Table 6 indicates that the MSA depths used in their experiments are deeper than those available in the Kaggle dataset. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3157217,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-03-23T06:08:43.727000",
          "content": "<p>Msa is not unique process. There are many ways Eg blastn rmsa mmseq etc. not sure what is exactly use. U will need to experiment </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3157232,
              "author_name": "Bianco Chiu",
              "author_url": "",
              "post_date": "2025-03-23T06:32:34.887000",
              "content": "<p>Yes, I plan to use Infernal to generate deeper MSAs and evaluate whether this can lead Rhofold+ to a higher score.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157394,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-23T11:55:59.290000",
              "content": "<p>don't over focus on one method rhofold. try af3, nufold, trRosettaRNA, etc … maybe other authors describe their mas pipline in more details and their results are more reproducible.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157425,
              "author_name": "Timmy Juicehouse",
              "author_url": "",
              "post_date": "2025-03-23T12:29:44.873000",
              "content": "<blockquote>\n  <p>evaluate whether this can lead Rhofold+ to a higher score.</p>\n</blockquote>\n<p>It depends on huge datasets:</p>\n<blockquote>\n  <p>To support MSA construction, 3 sequence databases (RNAcentral, Rfam, and nt) totaling about 900GB need to be downloaded.</p>\n</blockquote>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157448,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-23T12:56:43.673000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157449,
              "author_name": "Bianco Chiu",
              "author_url": "",
              "post_date": "2025-03-23T12:57:08.160000",
              "content": "<p>I agree that my current storage capacity is insufficient for handling such large datasets. I should likely focus on other methods.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157457,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-03-23T13:08:32.023000",
              "content": "<p>Alternatively, try online msa server. Key is to get pipeline correct. Can get data to local later.</p>\n<p>To submit with online server, submit offline predicted xyz, see my other notebook</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3150555,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-15T16:00:55.957000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff1b85c9b31b9245c6da161fee2f18001%2FSelection_058.png?generation=1742054324943370&amp;alt=media\" alt=\"\"></p>\n<ol>\n<li><p>i need to implement or find online server for blastN, rMSA, infernal to verify if the paper better results is due to better  MSA search.</p></li>\n<li><p>It is interesting to note that server results is better then local for same input.<br>\n(did the server use a better pretrain model? or is it due to different GPU, python package  version)</p></li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3149224,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-14T00:43:52.903000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3207667,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-23T05:37:40.287000",
      "content": "<p>great work! </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3149222": "Each step will be described in detail below. The focus is on how to check whether your usage is correct. Fortunately, there is an online web server for Rhofold+ that allows verification against our local results\n\nhere are the summary:\n1) ~~understand alignment format: fasta and a3m. Convert to a3m.~~ \ni think kaggle MAS already in a3m format. the extension fasta should be a3m\n2) predict using Rhofold+ using web server and local\n3) compare the results\n4) other important note, eg effects of num of msa used, conformations using different msa input\n3) generate your own msa",
    "3150011": "good news.\nwe have another benchmark paper, their a3m files is provided and can be downloaded\n\nSystematic benchmarking of deep-learning methods for tertiary RNA structure prediction\nhttps://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012715#sec035\nhttps://github.com/akashbahai/rna_benchmarking/tree/main/msa/r1108\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F5091071a931c5eec3b4849717a11c612%2FSelection_052.png?generation=1741999960906042&alt=media)\n\nbut this is for rhofold repo (and not rhofold+ repo)\nbut both rhofold and rhofold+ repot seem to point to the same pretrain model\n\n\nrhofold+:\nhttps://github.com/ml4bio/RhoFold?tab=readme-ov-file\n\nrhofold:\nhttps://github.com/Dharmogata/RhoFold?tab=readme-ov-file",
    "3149991": "let's \"algorithmically debug\" our implementation.\nwe use kaggle validation csv (which is casp15).\n\nfirst we need reference results and we will use rhofold+ paper (supplementary material)\nhttps://www.nature.com/articles/s41592-024-02487-0#additional-information\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc85b403de5259c2a67a340fff86577fb%2FSelection_051.png?generation=1741996415341156&alt=media)\n\n\nif you run my example notebook[1] with max num of MAS you should be getting\n[1] https://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa  \n\n```\nno relax\n\n0 R1107 num_msa=21   len(seq)=69 df.shape=(69, 6)   seq = GGGGGCCACA...\ntm_score_unrelax 0.32113\n\n1 R1108 26 69 (69, 6) GGGGGCCACA...\ntm_score_unrelax 0.41223\n\n2 R1116 68 157 (157, 6) CGCCCGGAUA...\ntm_score_unrelax 0.52747\n\n3 R1117v2 1 30 (30, 6) UUGGGUUCCC...\ntm_score_unrelax 0.20945\n\n8 R1149 9 124 (124, 6) GGACACGAGU...\ntm_score_unrelax 0.25613\n\n9 R1156 11 135 (135, 6) GGAGCAUCGU...\ntm_score_unrelax 0.34896\n```\n\nour results are definitely lower!\nlet's investigate why and improve on yt!\n\n... to be updated ...\n\nsome quite observation:\n- is TM-score in paper for backbone or all atoms (i.e. paper TM score is the same as ours?)\n- I also note that the num of MSA is different from the paper. luckliy the paper provide details of max match seq ...)\n- I think the comment from a3m header like \">000|DF967972.1/3772384-3772312/1-73 [subseq from] Longilinea arvoryzae DNA, scaffold: Larv_scaffold_1, strain: KOME-1.\" would help. maybe we can use a LLM model for text analysis for this ",
    "3149461": "##2) predict using Rhofold+ using local##\n\nfollow this notebook link to use rhofold+ on kaggle data for 1EIY_C (5 MSA)\nhttps://www.kaggle.com/code/hengck23/demo-for-rhofold-plus-with-kaggle-msa\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fac2110a027e0fea1e42531f8f170febb%2FSelection_049.png?generation=1741941642841589&alt=media)\n\nrender on https://molstar.org/viewer/\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcd2c3b43c290599db28d456144dee661%2FSelection_048.png?generation=1741941699751604&alt=media)",
    "3157211": "In the method section of the Rhofold+ paper, the authors state: \"By default, the top 256 MSAs are chosen as input features for predicting the standard structure, which we refer to as standard RhoFold+. RhoFold+ (TopK) refers to the optimal model selected from K different models generated using distinct sampled MSAs.\"\n\nDoes this imply that the CASP15 results were selected from multiple attempts? Additionally, Supplementary Table 6 indicates that the MSA depths used in their experiments are deeper than those available in the Kaggle dataset. ",
    "3150555": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff1b85c9b31b9245c6da161fee2f18001%2FSelection_058.png?generation=1742054324943370&alt=media)\n\n1. i need to implement or find online server for blastN, rMSA, infernal to verify if the paper better results is due to better  MSA search.\n\n2. It is interesting to note that server results is better then local for same input.\n(did the server use a better pretrain model? or is it due to different GPU, python package  version)",
    "3149224": "",
    "3207667": "great work! "
  }
}