{
  "id": 437794,
  "title": "🥇Papers for this Competition [with Code]🔥",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/437794",
  "author_name": "",
  "post_date": "2023-09-08T06:59:00.525785900Z",
  "votes": 44,
  "comment_count": 5,
  "views": 0,
  "content": "<ol>\n<li><a href=\"https://www.nature.com/articles/s41586-021-03819-2\" target=\"_blank\">Highly accurate protein structure prediction with AlphaFold</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/google-deepmind/alphafold\" target=\"_blank\">https://github.com/google-deepmind/alphafold</a></li>\n<li><strong>Summary:</strong> Proteins are vital to life, and understanding their structure helps decipher their function. While about 100,000 unique protein structures have been experimentally identified, this is only a tiny fraction of known protein sequences. Determining a single protein's structure is a long and arduous process. Hence, computational methods are crucial to bridge this gap. For over 50 years, predicting protein structures based on their amino acid sequences has been a significant challenge. Current methods lack atomic precision, particularly when no similar structure exists. However, an improved version of the AlphaFold neural network model can now predict protein structures with atomic-level accuracy, even without a known analogous structure. This achievement surpasses other methods, with the latest AlphaFold version incorporating deep learning combined with knowledge of protein structure.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2207.13921v3.pdf\" target=\"_blank\">HelixFold-Single: MSA-free Protein Structure Prediction by Using Protein Language Model as an Alternative</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/PaddlePaddle/PaddleHelix\" target=\"_blank\">https://github.com/PaddlePaddle/PaddleHelix</a></li>\n<li><strong>Summary:</strong> Advanced AI tools like AlphaFold2 can predict protein structures with near-experimental precision using Multiple Sequence Alignments (MSAs) to learn from related sequences. However, obtaining MSAs is time-intensive. HelixFold-Single aims for rapid protein structure prediction using only primary protein sequences. It combines a large protein language model (PLM), trained on billions of sequences, with AlphaFold2's geometry learning. This PLM replaces the need for MSAs. When tested, HelixFold-Single matched the accuracy of MSA-based methods on certain datasets while being much faster.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2207.01586.pdf\" target=\"_blank\">E2Efold-3D: End-to-End Deep Learning Method for accurate de novo RNA 3D Structure Prediction</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/RFOLD/RhoFold\" target=\"_blank\">https://github.com/RFOLD/RhoFold</a></li>\n<li><strong>Summary:</strong> RNA structure determination is crucial for drug development and synthetic element design. Traditional methods like X-ray crystallography, NMR, and Cryo-EM face challenges due to RNA's structural flexibility, resulting in few resolved RNA structures. Current computational predictions are not deep learning-based due to limited available structures and often use time-consuming sampling strategies with stagnant performance. This study introduces E2Efold-3D, the first deep learning method for RNA structure prediction. With innovative features to address data scarcity, it outperforms existing methods, achieving remarkable accuracy on a test dataset. It also uniquely predicts RNA complex structures. Integrating E2Efold-3D with experimental techniques can significantly advance the RNA structure prediction field.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2001.04020v1.pdf\" target=\"_blank\">LinearFold: linear-time approximate RNA folding by 5'-to-3' dynamic programming and beam search</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/LinearFold/LinearFold\" target=\"_blank\">https://github.com/LinearFold/LinearFold</a></li>\n<li><strong>Summary:</strong> Existing RNA secondary structure prediction tools have a limitation in speed as their runtimes increase significantly with RNA length. A new method is introduced that uses an alternative dynamic programming approach inspired by computational linguistics, which scans sequences from left-to-right. This allows for linear runtime and space without limiting the output structure, an achievement not seen in prior algorithms. Unexpectedly, this method not only works faster but also shows higher accuracy, especially in predicting long sequences and long-range base pairs, which were challenges for previous models.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2205.13927v2.pdf\" target=\"_blank\">Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/automl/probtransformer\" target=\"_blank\">https://github.com/automl/probtransformer</a></li>\n<li><strong>Summary:</strong> The data we use for training algorithms often reflects the ambiguity of our world, especially in natural processes like RNA folding where one sequence can lead to multiple structures. To better model such ambiguities, a hierarchical latent distribution is introduced to enhance the Transformer, a leading deep learning model. This enhanced approach shows promising results in understanding hidden data distributions, excelling in RNA folding predictions for ambiguous data, and in generating property-based molecule designs by learning underlying distributions, surpassing previous methods.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2212.14041v2.pdf\" target=\"_blank\">RFold: RNA Secondary Structure Prediction with Decoupled Optimization</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/a4bio/rfold\" target=\"_blank\">https://github.com/a4bio/rfold</a></li>\n<li><strong>Summary:</strong> RNA's secondary structure is crucial for functional prediction. While deep learning has been promising in predicting this, existing techniques often struggle with generalization and complexity. We introduce RFold, an end-to-end solution that uses a decoupled optimization process, breaking down the problem for easier solving and ensuring accurate results. Instead of hand-crafted features, RFold uses attention maps. Tests show that RFold is as accurate as leading methods but runs about eight times faster</li></ul></li>\n<li><a href=\"https://www.nature.com/articles/s41467-019-13395-9\" target=\"_blank\">RNA secondary structure prediction using an ensemble of two-dimensional deep neural networks and transfer learning</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/jaswindersingh2/SPOT-RNA\" target=\"_blank\">https://github.com/jaswindersingh2/SPOT-RNA</a></li>\n<li><strong>Summary:</strong> The human genome largely transcribes into noncoding RNAs with unclear structures and roles. Current methods for predicting their structures have seen limited progress over the past decade. This study introduces a deep contextual learning method, SPOT-RNA, to predict base pairs, including the noncanonical and pseudoknot types. By using transfer learning from a dataset of over 10,000 RNAs, the method showcases notable accuracy improvements, especially for noncanonical and non-nested pairs. SPOT-RNA is available as a free server and software, offering potential enhancements in RNA modeling, alignment, and function identification.</li></ul></li>\n</ol>\n<p><strong>Check out other discussions about AI trends, papers etc:</strong><br>\n<a href=\"https://www.kaggle.com/discussions/general/437852\" target=\"_blank\">Recent AI Releases &amp; Announcements</a><br>\n<a href=\"https://www.kaggle.com/discussions/general/437099\" target=\"_blank\">Trending AI Research</a><br>\n<a href=\"https://www.kaggle.com/discussions/general/437038\" target=\"_blank\">Top 5 AI Web Tools Of The Week</a></p>",
  "messages": [
    {
      "id": "2428797",
      "postDate": "09/08/2023 06:59:00",
      "content": "<ol>\n<li><a href=\"https://www.nature.com/articles/s41586-021-03819-2\" target=\"_blank\">Highly accurate protein structure prediction with AlphaFold</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/google-deepmind/alphafold\" target=\"_blank\">https://github.com/google-deepmind/alphafold</a></li>\n<li><strong>Summary:</strong> Proteins are vital to life, and understanding their structure helps decipher their function. While about 100,000 unique protein structures have been experimentally identified, this is only a tiny fraction of known protein sequences. Determining a single protein's structure is a long and arduous process. Hence, computational methods are crucial to bridge this gap. For over 50 years, predicting protein structures based on their amino acid sequences has been a significant challenge. Current methods lack atomic precision, particularly when no similar structure exists. However, an improved version of the AlphaFold neural network model can now predict protein structures with atomic-level accuracy, even without a known analogous structure. This achievement surpasses other methods, with the latest AlphaFold version incorporating deep learning combined with knowledge of protein structure.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2207.13921v3.pdf\" target=\"_blank\">HelixFold-Single: MSA-free Protein Structure Prediction by Using Protein Language Model as an Alternative</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/PaddlePaddle/PaddleHelix\" target=\"_blank\">https://github.com/PaddlePaddle/PaddleHelix</a></li>\n<li><strong>Summary:</strong> Advanced AI tools like AlphaFold2 can predict protein structures with near-experimental precision using Multiple Sequence Alignments (MSAs) to learn from related sequences. However, obtaining MSAs is time-intensive. HelixFold-Single aims for rapid protein structure prediction using only primary protein sequences. It combines a large protein language model (PLM), trained on billions of sequences, with AlphaFold2's geometry learning. This PLM replaces the need for MSAs. When tested, HelixFold-Single matched the accuracy of MSA-based methods on certain datasets while being much faster.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2207.01586.pdf\" target=\"_blank\">E2Efold-3D: End-to-End Deep Learning Method for accurate de novo RNA 3D Structure Prediction</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/RFOLD/RhoFold\" target=\"_blank\">https://github.com/RFOLD/RhoFold</a></li>\n<li><strong>Summary:</strong> RNA structure determination is crucial for drug development and synthetic element design. Traditional methods like X-ray crystallography, NMR, and Cryo-EM face challenges due to RNA's structural flexibility, resulting in few resolved RNA structures. Current computational predictions are not deep learning-based due to limited available structures and often use time-consuming sampling strategies with stagnant performance. This study introduces E2Efold-3D, the first deep learning method for RNA structure prediction. With innovative features to address data scarcity, it outperforms existing methods, achieving remarkable accuracy on a test dataset. It also uniquely predicts RNA complex structures. Integrating E2Efold-3D with experimental techniques can significantly advance the RNA structure prediction field.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2001.04020v1.pdf\" target=\"_blank\">LinearFold: linear-time approximate RNA folding by 5'-to-3' dynamic programming and beam search</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/LinearFold/LinearFold\" target=\"_blank\">https://github.com/LinearFold/LinearFold</a></li>\n<li><strong>Summary:</strong> Existing RNA secondary structure prediction tools have a limitation in speed as their runtimes increase significantly with RNA length. A new method is introduced that uses an alternative dynamic programming approach inspired by computational linguistics, which scans sequences from left-to-right. This allows for linear runtime and space without limiting the output structure, an achievement not seen in prior algorithms. Unexpectedly, this method not only works faster but also shows higher accuracy, especially in predicting long sequences and long-range base pairs, which were challenges for previous models.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2205.13927v2.pdf\" target=\"_blank\">Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/automl/probtransformer\" target=\"_blank\">https://github.com/automl/probtransformer</a></li>\n<li><strong>Summary:</strong> The data we use for training algorithms often reflects the ambiguity of our world, especially in natural processes like RNA folding where one sequence can lead to multiple structures. To better model such ambiguities, a hierarchical latent distribution is introduced to enhance the Transformer, a leading deep learning model. This enhanced approach shows promising results in understanding hidden data distributions, excelling in RNA folding predictions for ambiguous data, and in generating property-based molecule designs by learning underlying distributions, surpassing previous methods.</li></ul></li>\n<li><a href=\"https://arxiv.org/pdf/2212.14041v2.pdf\" target=\"_blank\">RFold: RNA Secondary Structure Prediction with Decoupled Optimization</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/a4bio/rfold\" target=\"_blank\">https://github.com/a4bio/rfold</a></li>\n<li><strong>Summary:</strong> RNA's secondary structure is crucial for functional prediction. While deep learning has been promising in predicting this, existing techniques often struggle with generalization and complexity. We introduce RFold, an end-to-end solution that uses a decoupled optimization process, breaking down the problem for easier solving and ensuring accurate results. Instead of hand-crafted features, RFold uses attention maps. Tests show that RFold is as accurate as leading methods but runs about eight times faster</li></ul></li>\n<li><a href=\"https://www.nature.com/articles/s41467-019-13395-9\" target=\"_blank\">RNA secondary structure prediction using an ensemble of two-dimensional deep neural networks and transfer learning</a><ul>\n<li><strong>Code:</strong> <a href=\"https://github.com/jaswindersingh2/SPOT-RNA\" target=\"_blank\">https://github.com/jaswindersingh2/SPOT-RNA</a></li>\n<li><strong>Summary:</strong> The human genome largely transcribes into noncoding RNAs with unclear structures and roles. Current methods for predicting their structures have seen limited progress over the past decade. This study introduces a deep contextual learning method, SPOT-RNA, to predict base pairs, including the noncanonical and pseudoknot types. By using transfer learning from a dataset of over 10,000 RNAs, the method showcases notable accuracy improvements, especially for noncanonical and non-nested pairs. SPOT-RNA is available as a free server and software, offering potential enhancements in RNA modeling, alignment, and function identification.</li></ul></li>\n</ol>\n<p><strong>Check out other discussions about AI trends, papers etc:</strong><br>\n<a href=\"https://www.kaggle.com/discussions/general/437852\" target=\"_blank\">Recent AI Releases &amp; Announcements</a><br>\n<a href=\"https://www.kaggle.com/discussions/general/437099\" target=\"_blank\">Trending AI Research</a><br>\n<a href=\"https://www.kaggle.com/discussions/general/437038\" target=\"_blank\">Top 5 AI Web Tools Of The Week</a></p>",
      "rawMarkdown": "1. [Highly accurate protein structure prediction with AlphaFold](https://www.nature.com/articles/s41586-021-03819-2)\n   - **Code:** https://github.com/google-deepmind/alphafold\n   - **Summary:** Proteins are vital to life, and understanding their structure helps decipher their function. While about 100,000 unique protein structures have been experimentally identified, this is only a tiny fraction of known protein sequences. Determining a single protein's structure is a long and arduous process. Hence, computational methods are crucial to bridge this gap. For over 50 years, predicting protein structures based on their amino acid sequences has been a significant challenge. Current methods lack atomic precision, particularly when no similar structure exists. However, an improved version of the AlphaFold neural network model can now predict protein structures with atomic-level accuracy, even without a known analogous structure. This achievement surpasses other methods, with the latest AlphaFold version incorporating deep learning combined with knowledge of protein structure.\n2. [HelixFold-Single: MSA-free Protein Structure Prediction by Using Protein Language Model as an Alternative](https://arxiv.org/pdf/2207.13921v3.pdf)\n   - **Code:** https://github.com/PaddlePaddle/PaddleHelix\n   - **Summary:** Advanced AI tools like AlphaFold2 can predict protein structures with near-experimental precision using Multiple Sequence Alignments (MSAs) to learn from related sequences. However, obtaining MSAs is time-intensive. HelixFold-Single aims for rapid protein structure prediction using only primary protein sequences. It combines a large protein language model (PLM), trained on billions of sequences, with AlphaFold2's geometry learning. This PLM replaces the need for MSAs. When tested, HelixFold-Single matched the accuracy of MSA-based methods on certain datasets while being much faster.\n2. [E2Efold-3D: End-to-End Deep Learning Method for accurate de novo RNA 3D Structure Prediction](https://arxiv.org/pdf/2207.01586.pdf)\n   - **Code:** https://github.com/RFOLD/RhoFold\n   - **Summary:** RNA structure determination is crucial for drug development and synthetic element design. Traditional methods like X-ray crystallography, NMR, and Cryo-EM face challenges due to RNA's structural flexibility, resulting in few resolved RNA structures. Current computational predictions are not deep learning-based due to limited available structures and often use time-consuming sampling strategies with stagnant performance. This study introduces E2Efold-3D, the first deep learning method for RNA structure prediction. With innovative features to address data scarcity, it outperforms existing methods, achieving remarkable accuracy on a test dataset. It also uniquely predicts RNA complex structures. Integrating E2Efold-3D with experimental techniques can significantly advance the RNA structure prediction field.\n2. [LinearFold: linear-time approximate RNA folding by 5'-to-3' dynamic programming and beam search](https://arxiv.org/pdf/2001.04020v1.pdf)\n   - **Code:** https://github.com/LinearFold/LinearFold\n   - **Summary:** Existing RNA secondary structure prediction tools have a limitation in speed as their runtimes increase significantly with RNA length. A new method is introduced that uses an alternative dynamic programming approach inspired by computational linguistics, which scans sequences from left-to-right. This allows for linear runtime and space without limiting the output structure, an achievement not seen in prior algorithms. Unexpectedly, this method not only works faster but also shows higher accuracy, especially in predicting long sequences and long-range base pairs, which were challenges for previous models.\n3. [Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design](https://arxiv.org/pdf/2205.13927v2.pdf)\n   - **Code:** https://github.com/automl/probtransformer\n   - **Summary:** The data we use for training algorithms often reflects the ambiguity of our world, especially in natural processes like RNA folding where one sequence can lead to multiple structures. To better model such ambiguities, a hierarchical latent distribution is introduced to enhance the Transformer, a leading deep learning model. This enhanced approach shows promising results in understanding hidden data distributions, excelling in RNA folding predictions for ambiguous data, and in generating property-based molecule designs by learning underlying distributions, surpassing previous methods.\n4. [RFold: RNA Secondary Structure Prediction with Decoupled Optimization](https://arxiv.org/pdf/2212.14041v2.pdf)\n   - **Code:** https://github.com/a4bio/rfold\n   - **Summary:** RNA's secondary structure is crucial for functional prediction. While deep learning has been promising in predicting this, existing techniques often struggle with generalization and complexity. We introduce RFold, an end-to-end solution that uses a decoupled optimization process, breaking down the problem for easier solving and ensuring accurate results. Instead of hand-crafted features, RFold uses attention maps. Tests show that RFold is as accurate as leading methods but runs about eight times faster\n5. [RNA secondary structure prediction using an ensemble of two-dimensional deep neural networks and transfer learning](https://www.nature.com/articles/s41467-019-13395-9)\n   - **Code:** https://github.com/jaswindersingh2/SPOT-RNA\n   - **Summary:** The human genome largely transcribes into noncoding RNAs with unclear structures and roles. Current methods for predicting their structures have seen limited progress over the past decade. This study introduces a deep contextual learning method, SPOT-RNA, to predict base pairs, including the noncanonical and pseudoknot types. By using transfer learning from a dataset of over 10,000 RNAs, the method showcases notable accuracy improvements, especially for noncanonical and non-nested pairs. SPOT-RNA is available as a free server and software, offering potential enhancements in RNA modeling, alignment, and function identification.\n\n**Check out other discussions about AI trends, papers etc:**\n[Recent AI Releases & Announcements](https://www.kaggle.com/discussions/general/437852)\n[Trending AI Research](https://www.kaggle.com/discussions/general/437099)\n[Top 5 AI Web Tools Of The Week](https://www.kaggle.com/discussions/general/437038)",
      "votes": null
    },
    {
      "id": "2428817",
      "postDate": "09/08/2023 07:16:24",
      "content": "<p>great summary, thanks for the code</p>",
      "rawMarkdown": "great summary, thanks for the code",
      "votes": null
    },
    {
      "id": "2428848",
      "postDate": "09/08/2023 07:39:59",
      "content": "<p>RhoFold: Fast and Accurate RNA 3D Structure Prediction with Deep Learning <a href=\"https://github.com/RFOLD/RhoFold\" target=\"_blank\">https://github.com/RFOLD/RhoFold</a></p>",
      "rawMarkdown": "RhoFold: Fast and Accurate RNA 3D Structure Prediction with Deep Learning https://github.com/RFOLD/RhoFold",
      "votes": null
    },
    {
      "id": "2428869",
      "postDate": "09/08/2023 08:00:08",
      "content": "<p>There are really a lot of papers in this area. Meta is also researching and publishing some models. I will add your recommended paper to the list.</p>",
      "rawMarkdown": "There are really a lot of papers in this area. Meta is also researching and publishing some models. I will add your recommended paper to the list.",
      "votes": null
    },
    {
      "id": "2514009",
      "postDate": "11/06/2023 00:15:23",
      "content": "<p>Thanks! Just need to add that the link to the code for E2EFold isn't working :(</p>",
      "rawMarkdown": "Thanks! Just need to add that the link to the code for E2EFold isn't working :(",
      "votes": null
    },
    {
      "id": "2514291",
      "postDate": "11/06/2023 07:03:51",
      "content": "<p><a href=\"https://www.kaggle.com/code/rhijudas/interactive-rhofold\" target=\"_blank\">https://www.kaggle.com/code/rhijudas/interactive-rhofold</a></p>",
      "rawMarkdown": "https://www.kaggle.com/code/rhijudas/interactive-rhofold",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2428817,
      "author_name": "t0xicwast3",
      "author_url": "",
      "post_date": "09/08/2023 07:16:24",
      "content": "<p>great summary, thanks for the code</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2428848,
      "author_name": "shentao",
      "author_url": "",
      "post_date": "09/08/2023 07:39:59",
      "content": "<p>RhoFold: Fast and Accurate RNA 3D Structure Prediction with Deep Learning <a href=\"https://github.com/RFOLD/RhoFold\" target=\"_blank\">https://github.com/RFOLD/RhoFold</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2428869,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/08/2023 08:00:08",
          "content": "<p>There are really a lot of papers in this area. Meta is also researching and publishing some models. I will add your recommended paper to the list.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2514009,
      "author_name": "samuelzxu",
      "author_url": "",
      "post_date": "11/06/2023 00:15:23",
      "content": "<p>Thanks! Just need to add that the link to the code for E2EFold isn't working :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 2514291,
          "author_name": "shentao",
          "author_url": "",
          "post_date": "11/06/2023 07:03:51",
          "content": "<p><a href=\"https://www.kaggle.com/code/rhijudas/interactive-rhofold\" target=\"_blank\">https://www.kaggle.com/code/rhijudas/interactive-rhofold</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2428797": "1. [Highly accurate protein structure prediction with AlphaFold](https://www.nature.com/articles/s41586-021-03819-2)\n   - **Code:** https://github.com/google-deepmind/alphafold\n   - **Summary:** Proteins are vital to life, and understanding their structure helps decipher their function. While about 100,000 unique protein structures have been experimentally identified, this is only a tiny fraction of known protein sequences. Determining a single protein's structure is a long and arduous process. Hence, computational methods are crucial to bridge this gap. For over 50 years, predicting protein structures based on their amino acid sequences has been a significant challenge. Current methods lack atomic precision, particularly when no similar structure exists. However, an improved version of the AlphaFold neural network model can now predict protein structures with atomic-level accuracy, even without a known analogous structure. This achievement surpasses other methods, with the latest AlphaFold version incorporating deep learning combined with knowledge of protein structure.\n2. [HelixFold-Single: MSA-free Protein Structure Prediction by Using Protein Language Model as an Alternative](https://arxiv.org/pdf/2207.13921v3.pdf)\n   - **Code:** https://github.com/PaddlePaddle/PaddleHelix\n   - **Summary:** Advanced AI tools like AlphaFold2 can predict protein structures with near-experimental precision using Multiple Sequence Alignments (MSAs) to learn from related sequences. However, obtaining MSAs is time-intensive. HelixFold-Single aims for rapid protein structure prediction using only primary protein sequences. It combines a large protein language model (PLM), trained on billions of sequences, with AlphaFold2's geometry learning. This PLM replaces the need for MSAs. When tested, HelixFold-Single matched the accuracy of MSA-based methods on certain datasets while being much faster.\n2. [E2Efold-3D: End-to-End Deep Learning Method for accurate de novo RNA 3D Structure Prediction](https://arxiv.org/pdf/2207.01586.pdf)\n   - **Code:** https://github.com/RFOLD/RhoFold\n   - **Summary:** RNA structure determination is crucial for drug development and synthetic element design. Traditional methods like X-ray crystallography, NMR, and Cryo-EM face challenges due to RNA's structural flexibility, resulting in few resolved RNA structures. Current computational predictions are not deep learning-based due to limited available structures and often use time-consuming sampling strategies with stagnant performance. This study introduces E2Efold-3D, the first deep learning method for RNA structure prediction. With innovative features to address data scarcity, it outperforms existing methods, achieving remarkable accuracy on a test dataset. It also uniquely predicts RNA complex structures. Integrating E2Efold-3D with experimental techniques can significantly advance the RNA structure prediction field.\n2. [LinearFold: linear-time approximate RNA folding by 5'-to-3' dynamic programming and beam search](https://arxiv.org/pdf/2001.04020v1.pdf)\n   - **Code:** https://github.com/LinearFold/LinearFold\n   - **Summary:** Existing RNA secondary structure prediction tools have a limitation in speed as their runtimes increase significantly with RNA length. A new method is introduced that uses an alternative dynamic programming approach inspired by computational linguistics, which scans sequences from left-to-right. This allows for linear runtime and space without limiting the output structure, an achievement not seen in prior algorithms. Unexpectedly, this method not only works faster but also shows higher accuracy, especially in predicting long sequences and long-range base pairs, which were challenges for previous models.\n3. [Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design](https://arxiv.org/pdf/2205.13927v2.pdf)\n   - **Code:** https://github.com/automl/probtransformer\n   - **Summary:** The data we use for training algorithms often reflects the ambiguity of our world, especially in natural processes like RNA folding where one sequence can lead to multiple structures. To better model such ambiguities, a hierarchical latent distribution is introduced to enhance the Transformer, a leading deep learning model. This enhanced approach shows promising results in understanding hidden data distributions, excelling in RNA folding predictions for ambiguous data, and in generating property-based molecule designs by learning underlying distributions, surpassing previous methods.\n4. [RFold: RNA Secondary Structure Prediction with Decoupled Optimization](https://arxiv.org/pdf/2212.14041v2.pdf)\n   - **Code:** https://github.com/a4bio/rfold\n   - **Summary:** RNA's secondary structure is crucial for functional prediction. While deep learning has been promising in predicting this, existing techniques often struggle with generalization and complexity. We introduce RFold, an end-to-end solution that uses a decoupled optimization process, breaking down the problem for easier solving and ensuring accurate results. Instead of hand-crafted features, RFold uses attention maps. Tests show that RFold is as accurate as leading methods but runs about eight times faster\n5. [RNA secondary structure prediction using an ensemble of two-dimensional deep neural networks and transfer learning](https://www.nature.com/articles/s41467-019-13395-9)\n   - **Code:** https://github.com/jaswindersingh2/SPOT-RNA\n   - **Summary:** The human genome largely transcribes into noncoding RNAs with unclear structures and roles. Current methods for predicting their structures have seen limited progress over the past decade. This study introduces a deep contextual learning method, SPOT-RNA, to predict base pairs, including the noncanonical and pseudoknot types. By using transfer learning from a dataset of over 10,000 RNAs, the method showcases notable accuracy improvements, especially for noncanonical and non-nested pairs. SPOT-RNA is available as a free server and software, offering potential enhancements in RNA modeling, alignment, and function identification.\n\n**Check out other discussions about AI trends, papers etc:**\n[Recent AI Releases & Announcements](https://www.kaggle.com/discussions/general/437852)\n[Trending AI Research](https://www.kaggle.com/discussions/general/437099)\n[Top 5 AI Web Tools Of The Week](https://www.kaggle.com/discussions/general/437038)",
    "2428817": "great summary, thanks for the code",
    "2428848": "RhoFold: Fast and Accurate RNA 3D Structure Prediction with Deep Learning https://github.com/RFOLD/RhoFold",
    "2428869": "There are really a lot of papers in this area. Meta is also researching and publishing some models. I will add your recommended paper to the list.",
    "2514009": "Thanks! Just need to add that the link to the code for E2EFold isn't working :(",
    "2514291": "https://www.kaggle.com/code/rhijudas/interactive-rhofold"
  },
  "source": "meta"
}