{
  "id": 567821,
  "title": "Is Multiple Sequence Alignment (MSA) Essential?",
  "url": "/competitions/stanford-rna-3d-folding/discussion/567821",
  "author_name": "",
  "post_date": "2025-03-12T09:13:02.884990300Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've been reading about CASP models and it seems like a lot of models use MSA, is it mandatory to use this?</p>",
  "messages": [
    {
      "id": "3147666",
      "postDate": "03/12/2025 09:13:02",
      "content": "<p>I've been reading about CASP models and it seems like a lot of models use MSA, is it mandatory to use this?</p>",
      "rawMarkdown": "I've been reading about CASP models and it seems like a lot of models use MSA, is it mandatory to use this?",
      "votes": null
    },
    {
      "id": "3147677",
      "postDate": "03/12/2025 09:30:13",
      "content": "<p>For the best result, yes. MSAs provide information about co-variation, which is essential for proper pairing of nucleotides. See <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566208#3140627\" target=\"_blank\"><strong>here</strong></a> for more information.</p>",
      "rawMarkdown": "For the best result, yes. MSAs provide information about co-variation, which is essential for proper pairing of nucleotides. See [**here**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566208#3140627) for more information.",
      "votes": null
    },
    {
      "id": "3147694",
      "postDate": "03/12/2025 10:04:11",
      "content": "<p>Thank you for the great answer confirming that my reasoning wasn't incorrect.</p>",
      "rawMarkdown": "Thank you for the great answer confirming that my reasoning wasn't incorrect.",
      "votes": null
    },
    {
      "id": "3148194",
      "postDate": "03/12/2025 20:18:34",
      "content": "<p>thought experiment:</p>\n<ul>\n<li>say if alphafold3 uses dataset A,B,C for rna msa, and train dataset T.  A,B,C has about 100 million rna seq</li>\n<li>now you use alphafold3 to predict 3d structure for all rna in A,B,C </li>\n<li>then you use (rna for A,B,C + distilled alphafold3 3d structure) + (rna T + truth 3d structure) to train your LLM rna predictor</li>\n</ul>\n<p>maybe you do not need MSA ?</p>\n<p>further assume, you check align score (test = casp15,casp16,RNA puzzle …. | reference = A,B,A) &gt;0.6 (?), then may be you are safe.</p>\n<hr>\n<p>note that for protein, ESM trained on 60 million protein can captured evolution information and has same performance as alphafold1. there are many opensource foundation  RNA LLM model around (trained with &gt;25 million), maybe useful?</p>",
      "rawMarkdown": "thought experiment:\n- say if alphafold3 uses dataset A,B,C for rna msa, and train dataset T.  A,B,C has about 100 million rna seq\n- now you use alphafold3 to predict 3d structure for all rna in A,B,C \n- then you use (rna for A,B,C + distilled alphafold3 3d structure) + (rna T + truth 3d structure) to train your LLM rna predictor\n\nmaybe you do not need MSA ?\n\nfurther assume, you check align score (test = casp15,casp16,RNA puzzle .... | reference = A,B,A) >0.6 (?), then may be you are safe.\n\n---\n\nnote that for protein, ESM trained on 60 million protein can captured evolution information and has same performance as alphafold1. there are many opensource foundation  RNA LLM model around (trained with >25 million), maybe useful?",
      "votes": null
    },
    {
      "id": "3148260",
      "postDate": "03/12/2025 23:44:29",
      "content": "<p>All the things you suggested would yield useful models, but this is a competition. RNA LLM models are okay, but not as good as models derived from MSAs. Same is true for protein LLM models versus MSA-based models used by AF2/AF3.</p>\n<p>It is possible, and maybe even likely, that a hybrid model based on LLMs and MSAs would be the best. Either way, MSAs will be needed for most competitive solutions.</p>",
      "rawMarkdown": "All the things you suggested would yield useful models, but this is a competition. RNA LLM models are okay, but not as good as models derived from MSAs. Same is true for protein LLM models versus MSA-based models used by AF2/AF3.\n\nIt is possible, and maybe even likely, that a hybrid model based on LLMs and MSAs would be the best. Either way, MSAs will be needed for most competitive solutions.",
      "votes": null
    },
    {
      "id": "3148263",
      "postDate": "03/12/2025 23:51:03",
      "content": "<p>i agree with your comments.<br>\nbut since internet is probhibited, we cannot use online MSA server.<br>\nsetting local offline MSA server seems impossible due to the large size of MSA database (and complexity of seeting up of the alignment tools)</p>\n<p>maybe the host can allow of use of \"online\" MSA server by providing MSA as dataset to add?</p>",
      "rawMarkdown": "i agree with your comments.\nbut since internet is probhibited, we cannot use online MSA server.\nsetting local offline MSA server seems impossible due to the large size of MSA database (and complexity of seeting up of the alignment tools)\n\nmaybe the host can allow of use of \"online\" MSA server by providing MSA as dataset to add?",
      "votes": null
    },
    {
      "id": "3148273",
      "postDate": "03/13/2025 00:45:34",
      "content": "<blockquote>\n  <p>maybe the host can allow of use of \"online\" MSA server by providing MSA as dataset to add?</p>\n</blockquote>\n<p>Good point. This will likely be needed to get the best performance.</p>\n<p>This is what Kaggle dataset page says:</p>\n<p>Kaggle Datasets allows you to publish and share datasets privately or publicly. We provide resources for storing and processing datasets, but there are certain technical specifications:</p>\n<ul>\n<li>200GB per dataset limit</li>\n<li>200GB max private datasets (if you exceed this, either make your datasets public or delete unused datasets)</li>\n</ul>\n<p>According to this page a good database could be created locally. Of course, it would be wasteful if all of us built these databases individually when a host could provide a single, comprehensive database.</p>\n<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> </p>",
      "rawMarkdown": "> maybe the host can allow of use of \"online\" MSA server by providing MSA as dataset to add?\n\nGood point. This will likely be needed to get the best performance.\n\nThis is what Kaggle dataset page says:\n\nKaggle Datasets allows you to publish and share datasets privately or publicly. We provide resources for storing and processing datasets, but there are certain technical specifications:\n\n- 200GB per dataset limit\n- 200GB max private datasets (if you exceed this, either make your datasets public or delete unused datasets)\n\nAccording to this page a good database could be created locally. Of course, it would be wasteful if all of us built these databases individually when a host could provide a single, comprehensive database.\n\n@shujun717 @rhijudas",
      "votes": null
    },
    {
      "id": "3148496",
      "postDate": "03/13/2025 07:32:58",
      "content": "<p>Interesting discussion. As mentioned, MSAs still generally outperform in terms of accuracy. However, given the practical constraints of internet limitations in this competition, exploring a hybrid approach combining LLMs with MSAs could indeed be a valuable strategy. Additionally, leveraging knowledge distillation from AlphaFold3's predictions may enhance the quality of training data. If the organizers allow the use of provided MSA datasets, it could lead to a more efficient and effective approach for everyone involved.</p>",
      "rawMarkdown": "Interesting discussion. As mentioned, MSAs still generally outperform in terms of accuracy. However, given the practical constraints of internet limitations in this competition, exploring a hybrid approach combining LLMs with MSAs could indeed be a valuable strategy. Additionally, leveraging knowledge distillation from AlphaFold3's predictions may enhance the quality of training data. If the organizers allow the use of provided MSA datasets, it could lead to a more efficient and effective approach for everyone involved.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3147677,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "03/12/2025 09:30:13",
      "content": "<p>For the best result, yes. MSAs provide information about co-variation, which is essential for proper pairing of nucleotides. See <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566208#3140627\" target=\"_blank\"><strong>here</strong></a> for more information.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3147694,
          "author_name": "junhanzangai",
          "author_url": "",
          "post_date": "03/12/2025 10:04:11",
          "content": "<p>Thank you for the great answer confirming that my reasoning wasn't incorrect.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3148194,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/12/2025 20:18:34",
      "content": "<p>thought experiment:</p>\n<ul>\n<li>say if alphafold3 uses dataset A,B,C for rna msa, and train dataset T.  A,B,C has about 100 million rna seq</li>\n<li>now you use alphafold3 to predict 3d structure for all rna in A,B,C </li>\n<li>then you use (rna for A,B,C + distilled alphafold3 3d structure) + (rna T + truth 3d structure) to train your LLM rna predictor</li>\n</ul>\n<p>maybe you do not need MSA ?</p>\n<p>further assume, you check align score (test = casp15,casp16,RNA puzzle …. | reference = A,B,A) &gt;0.6 (?), then may be you are safe.</p>\n<hr>\n<p>note that for protein, ESM trained on 60 million protein can captured evolution information and has same performance as alphafold1. there are many opensource foundation  RNA LLM model around (trained with &gt;25 million), maybe useful?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3148260,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/12/2025 23:44:29",
          "content": "<p>All the things you suggested would yield useful models, but this is a competition. RNA LLM models are okay, but not as good as models derived from MSAs. Same is true for protein LLM models versus MSA-based models used by AF2/AF3.</p>\n<p>It is possible, and maybe even likely, that a hybrid model based on LLMs and MSAs would be the best. Either way, MSAs will be needed for most competitive solutions.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3148263,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/12/2025 23:51:03",
              "content": "<p>i agree with your comments.<br>\nbut since internet is probhibited, we cannot use online MSA server.<br>\nsetting local offline MSA server seems impossible due to the large size of MSA database (and complexity of seeting up of the alignment tools)</p>\n<p>maybe the host can allow of use of \"online\" MSA server by providing MSA as dataset to add?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3148273,
                  "author_name": "tilii7",
                  "author_url": "",
                  "post_date": "03/13/2025 00:45:34",
                  "content": "<blockquote>\n  <p>maybe the host can allow of use of \"online\" MSA server by providing MSA as dataset to add?</p>\n</blockquote>\n<p>Good point. This will likely be needed to get the best performance.</p>\n<p>This is what Kaggle dataset page says:</p>\n<p>Kaggle Datasets allows you to publish and share datasets privately or publicly. We provide resources for storing and processing datasets, but there are certain technical specifications:</p>\n<ul>\n<li>200GB per dataset limit</li>\n<li>200GB max private datasets (if you exceed this, either make your datasets public or delete unused datasets)</li>\n</ul>\n<p>According to this page a good database could be created locally. Of course, it would be wasteful if all of us built these databases individually when a host could provide a single, comprehensive database.</p>\n<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3148496,
                      "author_name": "junhanzangai",
                      "author_url": "",
                      "post_date": "03/13/2025 07:32:58",
                      "content": "<p>Interesting discussion. As mentioned, MSAs still generally outperform in terms of accuracy. However, given the practical constraints of internet limitations in this competition, exploring a hybrid approach combining LLMs with MSAs could indeed be a valuable strategy. Additionally, leveraging knowledge distillation from AlphaFold3's predictions may enhance the quality of training data. If the organizers allow the use of provided MSA datasets, it could lead to a more efficient and effective approach for everyone involved.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3147666": "I've been reading about CASP models and it seems like a lot of models use MSA, is it mandatory to use this?",
    "3147677": "For the best result, yes. MSAs provide information about co-variation, which is essential for proper pairing of nucleotides. See [**here**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/566208#3140627) for more information.",
    "3147694": "Thank you for the great answer confirming that my reasoning wasn't incorrect.",
    "3148194": "thought experiment:\n- say if alphafold3 uses dataset A,B,C for rna msa, and train dataset T.  A,B,C has about 100 million rna seq\n- now you use alphafold3 to predict 3d structure for all rna in A,B,C \n- then you use (rna for A,B,C + distilled alphafold3 3d structure) + (rna T + truth 3d structure) to train your LLM rna predictor\n\nmaybe you do not need MSA ?\n\nfurther assume, you check align score (test = casp15,casp16,RNA puzzle .... | reference = A,B,A) >0.6 (?), then may be you are safe.\n\n---\n\nnote that for protein, ESM trained on 60 million protein can captured evolution information and has same performance as alphafold1. there are many opensource foundation  RNA LLM model around (trained with >25 million), maybe useful?",
    "3148260": "All the things you suggested would yield useful models, but this is a competition. RNA LLM models are okay, but not as good as models derived from MSAs. Same is true for protein LLM models versus MSA-based models used by AF2/AF3.\n\nIt is possible, and maybe even likely, that a hybrid model based on LLMs and MSAs would be the best. Either way, MSAs will be needed for most competitive solutions.",
    "3148263": "i agree with your comments.\nbut since internet is probhibited, we cannot use online MSA server.\nsetting local offline MSA server seems impossible due to the large size of MSA database (and complexity of seeting up of the alignment tools)\n\nmaybe the host can allow of use of \"online\" MSA server by providing MSA as dataset to add?",
    "3148273": "> maybe the host can allow of use of \"online\" MSA server by providing MSA as dataset to add?\n\nGood point. This will likely be needed to get the best performance.\n\nThis is what Kaggle dataset page says:\n\nKaggle Datasets allows you to publish and share datasets privately or publicly. We provide resources for storing and processing datasets, but there are certain technical specifications:\n\n- 200GB per dataset limit\n- 200GB max private datasets (if you exceed this, either make your datasets public or delete unused datasets)\n\nAccording to this page a good database could be created locally. Of course, it would be wasteful if all of us built these databases individually when a host could provide a single, comprehensive database.\n\n@shujun717 @rhijudas",
    "3148496": "Interesting discussion. As mentioned, MSAs still generally outperform in terms of accuracy. However, given the practical constraints of internet limitations in this competition, exploring a hybrid approach combining LLMs with MSAs could indeed be a valuable strategy. Additionally, leveraging knowledge distillation from AlphaFold3's predictions may enhance the quality of training data. If the organizers allow the use of provided MSA datasets, it could lead to a more efficient and effective approach for everyone involved."
  },
  "source": "meta"
}