{
  "id": 566203,
  "title": "Foundation models for RNA sequences - how can we use  them for this challenge",
  "url": "/competitions/stanford-rna-3d-folding/discussion/566203",
  "author_name": "",
  "post_date": "2025-03-04T10:07:47.244435700Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all,<br>\nFoundation models for biology are the equivalent of LLMS, trained on DNA sequences or other biological entities instead of natural language. They are generally based on the transformer or similar architectures, and trained on massive amounts of data. During the training, the model learns relationships between different parts of the sequence, similarly to the way the attention mechanism identifies relations between words in a text. </p>\n<p>Many excellent foundation models have been published, trained on sequences or other biological datasets. I think some of these models may be useful in solving this challenge, but I need some help to investigate it and brainstorm it.</p>\n<p>One of the most recent models to be published is <a href=\"https://arcinstitute.org/news/blog/evo2\" target=\"_blank\">evo-2</a>. This model has been trained on more than 100,000 DNA sequences and almost 8T nucleotides. During the training, the model learned the relationship between DNA nucleotides. As shown in Fig 4D of the paper (<a href=\"https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1)\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1)</a>, the model also learned 3D features such as α-helices, β-sheet associated to RNA structures. Even if the model is trained only on DNA sequences, it learned something about RNA.<br>\nHowever, installing Evo2 on a Kaggle notebook is complex. I did not manage to get it installed. </p>\n<p>Another model is <a href=\"https://github.com/helicalAI/helical/tree/release/helical/models/helix_mrna\" target=\"_blank\">Helix-mRNA</a>, which is a smaller model trained only on RNA sequences. This model is easier to install, and may be used to generate the embeddings for training a further model to predict the structure. I managed to get it installed in <a href=\"https://www.kaggle.com/code/dalloliogm/computing-embeddings-using-helix-mrna\" target=\"_blank\">this notebook</a>, but now I am unsure on how to proceed for the prediction. </p>",
  "messages": [
    {
      "id": "3140134",
      "postDate": "03/04/2025 10:07:47",
      "content": "<p>Hi all,<br>\nFoundation models for biology are the equivalent of LLMS, trained on DNA sequences or other biological entities instead of natural language. They are generally based on the transformer or similar architectures, and trained on massive amounts of data. During the training, the model learns relationships between different parts of the sequence, similarly to the way the attention mechanism identifies relations between words in a text. </p>\n<p>Many excellent foundation models have been published, trained on sequences or other biological datasets. I think some of these models may be useful in solving this challenge, but I need some help to investigate it and brainstorm it.</p>\n<p>One of the most recent models to be published is <a href=\"https://arcinstitute.org/news/blog/evo2\" target=\"_blank\">evo-2</a>. This model has been trained on more than 100,000 DNA sequences and almost 8T nucleotides. During the training, the model learned the relationship between DNA nucleotides. As shown in Fig 4D of the paper (<a href=\"https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1)\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1)</a>, the model also learned 3D features such as α-helices, β-sheet associated to RNA structures. Even if the model is trained only on DNA sequences, it learned something about RNA.<br>\nHowever, installing Evo2 on a Kaggle notebook is complex. I did not manage to get it installed. </p>\n<p>Another model is <a href=\"https://github.com/helicalAI/helical/tree/release/helical/models/helix_mrna\" target=\"_blank\">Helix-mRNA</a>, which is a smaller model trained only on RNA sequences. This model is easier to install, and may be used to generate the embeddings for training a further model to predict the structure. I managed to get it installed in <a href=\"https://www.kaggle.com/code/dalloliogm/computing-embeddings-using-helix-mrna\" target=\"_blank\">this notebook</a>, but now I am unsure on how to proceed for the prediction. </p>",
      "rawMarkdown": "Hi all,\nFoundation models for biology are the equivalent of LLMS, trained on DNA sequences or other biological entities instead of natural language. They are generally based on the transformer or similar architectures, and trained on massive amounts of data. During the training, the model learns relationships between different parts of the sequence, similarly to the way the attention mechanism identifies relations between words in a text. \n\nMany excellent foundation models have been published, trained on sequences or other biological datasets. I think some of these models may be useful in solving this challenge, but I need some help to investigate it and brainstorm it.\n\nOne of the most recent models to be published is [evo-2](https://arcinstitute.org/news/blog/evo2). This model has been trained on more than 100,000 DNA sequences and almost 8T nucleotides. During the training, the model learned the relationship between DNA nucleotides. As shown in Fig 4D of the paper (https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1), the model also learned 3D features such as α-helices, β-sheet associated to RNA structures. Even if the model is trained only on DNA sequences, it learned something about RNA.\nHowever, installing Evo2 on a Kaggle notebook is complex. I did not manage to get it installed. \n\nAnother model is [Helix-mRNA](https://github.com/helicalAI/helical/tree/release/helical/models/helix_mrna), which is a smaller model trained only on RNA sequences. This model is easier to install, and may be used to generate the embeddings for training a further model to predict the structure. I managed to get it installed in [this notebook](https://www.kaggle.com/code/dalloliogm/computing-embeddings-using-helix-mrna), but now I am unsure on how to proceed for the prediction.",
      "votes": null
    },
    {
      "id": "3140192",
      "postDate": "03/04/2025 11:26:45",
      "content": "<p>You can try to attach a new prediction head and finetune it for our task. Same as the host did with RibonanzaNet</p>",
      "rawMarkdown": "You can try to attach a new prediction head and finetune it for our task. Same as the host did with RibonanzaNet",
      "votes": null
    },
    {
      "id": "3140608",
      "postDate": "03/04/2025 19:09:47",
      "content": "<blockquote>\n  <p>Even if the model is trained only on DNA sequences, it learned something about RNA.</p>\n</blockquote>\n<p>We are dealing here with something known as structured RNAs (AKA, non-coding RNAs). Those would be RNA molecules that form stable structures based on the patterns of short- and long-distance base-pairing. They are the minority of RNA molecules coded in any genome, so any general language model is more likely to learn about coding RNAs rather than non-coding (structured) RNAs. The model may have learned to distinguish between coding and non-coding RNAs, but that doesn't really help with how to fold non-coding RNAs.</p>\n<blockquote>\n  <p>However, installing Evo2 on a Kaggle notebook is complex. I did not manage to get it installed.</p>\n</blockquote>\n<p>My understanding is that Evo2 requires &gt;=8.9 GPU computing architecture, which excludes most GPUs that are in wide use. The largest Evo2 model requires more memory than any single commercially available GPU has at this moment, so I'd be surprised if anyone managed to install it on Kaggle. I don't think it is worth even trying given a low chance it would help with this problem.</p>\n<blockquote>\n  <p>Another model is Helix-mRNA, which is a smaller model trained only on RNA sequences.</p>\n</blockquote>\n<p>On a quick reading it seems that this model is tuned for mRNA molecules, which are coding RNAs.</p>",
      "rawMarkdown": "> Even if the model is trained only on DNA sequences, it learned something about RNA.\n\nWe are dealing here with something known as structured RNAs (AKA, non-coding RNAs). Those would be RNA molecules that form stable structures based on the patterns of short- and long-distance base-pairing. They are the minority of RNA molecules coded in any genome, so any general language model is more likely to learn about coding RNAs rather than non-coding (structured) RNAs. The model may have learned to distinguish between coding and non-coding RNAs, but that doesn't really help with how to fold non-coding RNAs.\n\n> However, installing Evo2 on a Kaggle notebook is complex. I did not manage to get it installed.\n\nMy understanding is that Evo2 requires >=8.9 GPU computing architecture, which excludes most GPUs that are in wide use. The largest Evo2 model requires more memory than any single commercially available GPU has at this moment, so I'd be surprised if anyone managed to install it on Kaggle. I don't think it is worth even trying given a low chance it would help with this problem.\n\n> Another model is Helix-mRNA, which is a smaller model trained only on RNA sequences.\n\nOn a quick reading it seems that this model is tuned for mRNA molecules, which are coding RNAs.",
      "votes": null
    },
    {
      "id": "3140690",
      "postDate": "03/04/2025 20:27:42",
      "content": "<p>Thanks! This is really good feedback.<br>\nI didn't understand that we were dealing with non-coding RNAs, so this is very useful.<br>\nIn the case of Evo2, the model learned sequences associated with specific structures such as alpha-helices, beta-folds, etc.. Wouldn't these be the same for coding and non-coding regions?</p>",
      "rawMarkdown": "Thanks! This is really good feedback.\nI didn't understand that we were dealing with non-coding RNAs, so this is very useful.\nIn the case of Evo2, the model learned sequences associated with specific structures such as alpha-helices, beta-folds, etc.. Wouldn't these be the same for coding and non-coding regions?",
      "votes": null
    },
    {
      "id": "3140691",
      "postDate": "03/04/2025 20:28:16",
      "content": "<p>Thanks! That is my plan, in the long term. I just need to understand how to code it.</p>",
      "rawMarkdown": "Thanks! That is my plan, in the long term. I just need to understand how to code it.",
      "votes": null
    },
    {
      "id": "3140708",
      "postDate": "03/04/2025 20:58:11",
      "content": "<blockquote>\n  <p>In the case of Evo2, the model learned sequences associated with specific structures such as alpha-helices, beta-folds, etc.. Wouldn't these be the same for coding and non-coding regions?</p>\n</blockquote>\n<p>Those are protein building blocks, so not relevant here. The only Evo2 thing somewhat relevant to this competition is the ability to tease out transfer-RNAs (tRNAs), which are of the non-coding variety. However, predicting that something is a non-coding RNA is completely different from predicting how it folds, and I don't think Evo2 would be very helpful in that regard.</p>",
      "rawMarkdown": "> In the case of Evo2, the model learned sequences associated with specific structures such as alpha-helices, beta-folds, etc.. Wouldn't these be the same for coding and non-coding regions?\n\nThose are protein building blocks, so not relevant here. The only Evo2 thing somewhat relevant to this competition is the ability to tease out transfer-RNAs (tRNAs), which are of the non-coding variety. However, predicting that something is a non-coding RNA is completely different from predicting how it folds, and I don't think Evo2 would be very helpful in that regard.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3140192,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "03/04/2025 11:26:45",
      "content": "<p>You can try to attach a new prediction head and finetune it for our task. Same as the host did with RibonanzaNet</p>",
      "votes": null,
      "replies": [
        {
          "id": 3140691,
          "author_name": "dalloliogm",
          "author_url": "",
          "post_date": "03/04/2025 20:28:16",
          "content": "<p>Thanks! That is my plan, in the long term. I just need to understand how to code it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3140608,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "03/04/2025 19:09:47",
      "content": "<blockquote>\n  <p>Even if the model is trained only on DNA sequences, it learned something about RNA.</p>\n</blockquote>\n<p>We are dealing here with something known as structured RNAs (AKA, non-coding RNAs). Those would be RNA molecules that form stable structures based on the patterns of short- and long-distance base-pairing. They are the minority of RNA molecules coded in any genome, so any general language model is more likely to learn about coding RNAs rather than non-coding (structured) RNAs. The model may have learned to distinguish between coding and non-coding RNAs, but that doesn't really help with how to fold non-coding RNAs.</p>\n<blockquote>\n  <p>However, installing Evo2 on a Kaggle notebook is complex. I did not manage to get it installed.</p>\n</blockquote>\n<p>My understanding is that Evo2 requires &gt;=8.9 GPU computing architecture, which excludes most GPUs that are in wide use. The largest Evo2 model requires more memory than any single commercially available GPU has at this moment, so I'd be surprised if anyone managed to install it on Kaggle. I don't think it is worth even trying given a low chance it would help with this problem.</p>\n<blockquote>\n  <p>Another model is Helix-mRNA, which is a smaller model trained only on RNA sequences.</p>\n</blockquote>\n<p>On a quick reading it seems that this model is tuned for mRNA molecules, which are coding RNAs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3140690,
          "author_name": "dalloliogm",
          "author_url": "",
          "post_date": "03/04/2025 20:27:42",
          "content": "<p>Thanks! This is really good feedback.<br>\nI didn't understand that we were dealing with non-coding RNAs, so this is very useful.<br>\nIn the case of Evo2, the model learned sequences associated with specific structures such as alpha-helices, beta-folds, etc.. Wouldn't these be the same for coding and non-coding regions?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3140708,
              "author_name": "tilii7",
              "author_url": "",
              "post_date": "03/04/2025 20:58:11",
              "content": "<blockquote>\n  <p>In the case of Evo2, the model learned sequences associated with specific structures such as alpha-helices, beta-folds, etc.. Wouldn't these be the same for coding and non-coding regions?</p>\n</blockquote>\n<p>Those are protein building blocks, so not relevant here. The only Evo2 thing somewhat relevant to this competition is the ability to tease out transfer-RNAs (tRNAs), which are of the non-coding variety. However, predicting that something is a non-coding RNA is completely different from predicting how it folds, and I don't think Evo2 would be very helpful in that regard.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3140134": "Hi all,\nFoundation models for biology are the equivalent of LLMS, trained on DNA sequences or other biological entities instead of natural language. They are generally based on the transformer or similar architectures, and trained on massive amounts of data. During the training, the model learns relationships between different parts of the sequence, similarly to the way the attention mechanism identifies relations between words in a text. \n\nMany excellent foundation models have been published, trained on sequences or other biological datasets. I think some of these models may be useful in solving this challenge, but I need some help to investigate it and brainstorm it.\n\nOne of the most recent models to be published is [evo-2](https://arcinstitute.org/news/blog/evo2). This model has been trained on more than 100,000 DNA sequences and almost 8T nucleotides. During the training, the model learned the relationship between DNA nucleotides. As shown in Fig 4D of the paper (https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1), the model also learned 3D features such as α-helices, β-sheet associated to RNA structures. Even if the model is trained only on DNA sequences, it learned something about RNA.\nHowever, installing Evo2 on a Kaggle notebook is complex. I did not manage to get it installed. \n\nAnother model is [Helix-mRNA](https://github.com/helicalAI/helical/tree/release/helical/models/helix_mrna), which is a smaller model trained only on RNA sequences. This model is easier to install, and may be used to generate the embeddings for training a further model to predict the structure. I managed to get it installed in [this notebook](https://www.kaggle.com/code/dalloliogm/computing-embeddings-using-helix-mrna), but now I am unsure on how to proceed for the prediction.",
    "3140192": "You can try to attach a new prediction head and finetune it for our task. Same as the host did with RibonanzaNet",
    "3140608": "> Even if the model is trained only on DNA sequences, it learned something about RNA.\n\nWe are dealing here with something known as structured RNAs (AKA, non-coding RNAs). Those would be RNA molecules that form stable structures based on the patterns of short- and long-distance base-pairing. They are the minority of RNA molecules coded in any genome, so any general language model is more likely to learn about coding RNAs rather than non-coding (structured) RNAs. The model may have learned to distinguish between coding and non-coding RNAs, but that doesn't really help with how to fold non-coding RNAs.\n\n> However, installing Evo2 on a Kaggle notebook is complex. I did not manage to get it installed.\n\nMy understanding is that Evo2 requires >=8.9 GPU computing architecture, which excludes most GPUs that are in wide use. The largest Evo2 model requires more memory than any single commercially available GPU has at this moment, so I'd be surprised if anyone managed to install it on Kaggle. I don't think it is worth even trying given a low chance it would help with this problem.\n\n> Another model is Helix-mRNA, which is a smaller model trained only on RNA sequences.\n\nOn a quick reading it seems that this model is tuned for mRNA molecules, which are coding RNAs.",
    "3140690": "Thanks! This is really good feedback.\nI didn't understand that we were dealing with non-coding RNAs, so this is very useful.\nIn the case of Evo2, the model learned sequences associated with specific structures such as alpha-helices, beta-folds, etc.. Wouldn't these be the same for coding and non-coding regions?",
    "3140691": "Thanks! That is my plan, in the long term. I just need to understand how to code it.",
    "3140708": "> In the case of Evo2, the model learned sequences associated with specific structures such as alpha-helices, beta-folds, etc.. Wouldn't these be the same for coding and non-coding regions?\n\nThose are protein building blocks, so not relevant here. The only Evo2 thing somewhat relevant to this competition is the ability to tease out transfer-RNAs (tRNAs), which are of the non-coding variety. However, predicting that something is a non-coding RNA is completely different from predicting how it folds, and I don't think Evo2 would be very helpful in that regard."
  },
  "source": "meta"
}